Files
foxhunt/docs
jgrusewski 6da9d262db feat(ml): MAMBA-2 P0 fixes + hyperparameter optimization (13 params)
CRITICAL P0 FIXES (Validated - Loss 0.87 → 0.07):
- Add sigmoid activation to inference and training (ml/src/mamba/mod.rs:798, 1538)
- Fix config.total_decay_steps (was hardcoded 10000) (ml/src/mamba/mod.rs:2271)
- Update d_state: 16→64, 32→64 (Mamba-2 spec) (ml/src/mamba/mod.rs:178, 730)

HYPERPARAMETER OPTIMIZATION:
- Implement 13-parameter Bayesian optimization with argmin
- Add async data loading with 3-batch prefetch (+20-30% speedup)
- Create hyperopt adapter: ml/src/hyperopt/adapters/mamba2.rs
- Add example: ml/examples/hyperopt_mamba2_demo.rs

VALIDATION:
- Local test: Loss 0.07 vs 0.87 (12× improvement)
- Val loss: 0.04-0.14 vs 1.2 (27× improvement)
- Accuracy: 12-30% vs 1-5% (3-6× improvement)
- All binaries rebuilt and uploaded to Runpod S3

DEPLOYMENT:
- RTX 4090 pod active (n0fq2ikt4uk0zy)
- Training: 10 trials × 50 epochs, batch_size=256
- Expected: 1.3 days, $10.41 cost

Fixes #P0-sigmoid #P0-decay-steps #hyperopt-mamba2
2025-10-28 14:11:18 +01:00
..

Foxhunt Documentation Index

Last Updated: 2025-10-22 Status: Organized and Indexed + Wave 5 Operational Documentation Complete Total Documentation: 940 files, 12.4 MB (includes 28 new operational docs)


🎯 Start Here

New to Foxhunt?

  1. CLAUDE.md - System overview, architecture, current status (MUST READ)
  2. README.md - Project introduction
  3. ML Infrastructure Guide - Master documentation index

Quick Start Guides

  1. Quick Start: Training - Train your first model (5-7 weeks)
  2. Quick Start: Tuning - Optimize hyperparameters (3-4 days)

📁 Documentation Categories

Operational Guides (deployment/, runbooks/, troubleshooting/, monitoring/, templates/) 🆕

28 documents (Wave 5) - Production deployment, incident response, troubleshooting

  • Deployment Guides (5 docs): Docker, Kubernetes, Cloud (AWS/GCP/Azure), Zero-Downtime, Rollback Procedures
  • Operational Runbooks (6 docs): Incident Response (P0-P4), Service Restart, Database Migration, Disaster Recovery, Scaling, Security Incidents
  • Troubleshooting Guides (7 docs): High Latency, Memory Leaks, Service Crashes, Database Issues, GPU Errors, Network Errors, Circuit Breakers
  • Monitoring Playbooks (5 docs): Prometheus Setup, Grafana Setup, Alerting Rules, SLO/SLI Tracking, Log Aggregation
  • Templates & Checklists (5 docs): Deployment Checklist (25 items), Incident Report, Change Request, Runbook Template, On-Call Handoff

Key Files:

Training Guides (training/)

371 documents - ML model training, checkpoints, hyperparameters

  • DQN, PPO, MAMBA-2, TFT training
  • Checkpoint management
  • Feature engineering
  • GPU optimization

Key Files:

Deployment Guides (deployment/)

546 documents - Production deployment, infrastructure, operations

  • Production runbooks
  • Docker deployment
  • Infrastructure scaling
  • Security hardening

Key Files:

Analysis & Reports (analysis/)

738 documents - Performance analysis, audits, investigations

  • Wave reports (488 files)
  • Agent reports
  • Performance benchmarks
  • Security audits

Key Files:

API Reference (api/)

716 documents - gRPC endpoints, integrations, service interfaces

  • API Gateway (22 methods)
  • Trading Service
  • Backtesting Service
  • ML Training Service

Key Files:

Quick Start Guides (guides/)

129 documents - Getting started, tutorials, runbooks

  • Training guides
  • Tuning guides
  • Deployment guides
  • Troubleshooting guides

Key Files:

Troubleshooting (troubleshooting/)

667 documents - Debug guides, fixes, known issues

  • Port conflicts
  • GPU/CUDA issues
  • Database connection
  • Service health

Key Files:

Archive (archive/)

50+ candidates - Obsolete and historical documentation

  • Superseded versions
  • Completed wave reports
  • Temporary handoffs
  • Duplicate content

🔍 Find Documentation By...

By Topic

  • Authentication → Security section
  • Backtesting → Training guides + Deployment
  • Checkpoints → Training guides
  • Deployment → Deployment guides
  • GPU/CUDA → Training guides
  • Hyperparameters → Tuning guides
  • Models (DQN/PPO/MAMBA-2/TFT) → Training guides
  • Performance → Analysis section
  • Security → Deployment guides
  • Testing → Analysis section

By Use Case

I want to... Start here
Train a model Quick Start: Training
Optimize hyperparameters Quick Start: Tuning
Deploy to production Production Deployment Runbook V3
Troubleshoot an issue Troubleshooting Guide
Understand the API ML Infrastructure Guide - API Section
Set up paper trading Paper Trading Deployment Plan

📊 Documentation Statistics

By Category

  • Analysis/Reports: 738 files (80.9%)
  • API Reference: 716 files (78.5%)
  • Troubleshooting: 667 files (73.1%)
  • Deployment: 546 files (59.9%)
  • Wave Reports: 488 files (53.5%)
  • Architecture: 463 files (50.8%)
  • Training: 371 files (40.7%)

By Size

  • Total: 11.7 MB (404,079 lines)
  • Largest: DATA_PLAN.md (99.3K)
  • Average: 13.1K per file

By Location

  • Root directory: 421 files (46%)
  • Docs directory: 334 files (37%)
  • Other directories: 157 files (17%)

🔧 Contributing to Documentation

Adding New Documentation

  1. Choose appropriate category directory
  2. Follow naming convention (UPPERCASE_SNAKE_CASE.md)
  3. Add entry to ML_INFRASTRUCTURE_GUIDE.md
  4. Include cross-references to related docs
  5. Update this README if adding new category

Updating Existing Documentation

  1. Update file content
  2. Update "Last Updated" date
  3. Update cross-references if structure changes
  4. Update ML_INFRASTRUCTURE_GUIDE.md if major changes

Archiving Documentation

  1. Move to docs/archive/YYYY-MM-DD-reason/
  2. Create README in archive directory
  3. Update ML_INFRASTRUCTURE_GUIDE.md
  4. Remove from this index

📅 Recent Updates

2025-10-14 (Documentation Consolidation)

  • Created ML Infrastructure Guide (master index)
  • Created 2 quick-start guides (Training, Tuning)
  • Organized directory structure (7 categories)
  • Added 200+ cross-references
  • Identified 50+ archive candidates

2025-10-13 (Wave 160 Phase 4)

  • ML training pipeline complete
  • 19 agents, 4 models trained
  • System 100% production ready

🎯 Next Steps

Phase 2 (Short-term - 1-2 weeks)

  1. Move files to category directories
  2. Create consolidated guides (API, Training, Deployment)
  3. Archive obsolete documentation
  4. Add more cross-references

Phase 3 (Medium-term - 1 month)

  1. Consolidate wave reports (488 → 20 phase summaries)
  2. Enhance troubleshooting guide
  3. Search optimization (keywords, metadata)
  4. Documentation tests (link validation)

📞 Support

Documentation Issues

  • Missing documentation? Create GitHub issue with docs label
  • Broken links? Submit PR with fix
  • Outdated content? File issue with current status

Technical Support

  • Development: See Troubleshooting Guide
  • Deployment: Review production runbooks
  • ML Training: Consult training guides
  • Performance: See performance benchmarks

Document Version: 1.0 Created: 2025-10-14 Last Updated: 2025-10-14 Maintained by: Foxhunt Development Team