# PPO Hyperparameter Tuning - Quick Start **Duration**: 8-12 hours | **Trials**: 50 | **Objective**: 0.7 × Sharpe + 0.3 × ExplainedVar --- ## One-Line Execution ```bash cd /home/jgrusewski/Work/foxhunt && ./run_ppo_comprehensive_tuning.sh ``` That's it! The script handles everything. --- ## What Gets Optimized | Hyperparameter | Choices | Current Best (Epoch 380) | |----------------|---------|--------------------------| | Learning Rate | [0.0001, 0.0003, 0.001] | 1e-4 | | Batch Size | [32, 64, 128, 256] | 64 | | Gamma | [0.95, 0.99] | 0.99 | | GAE Lambda | [0.9, 0.95, 0.98] | 0.95 | | Clip Epsilon | [0.1, 0.2, 0.3] | 0.2 | | Entropy Coef | [0.001, 0.01, 0.1] | 0.05 | **Search Space**: 648 combinations → 50 intelligent trials (TPE sampling) --- ## Timeline | Time | Status | |------|--------| | 0:00 | Setup + prerequisites check | | 0:05 | Trial 1/50 starts | | 1:00 | Trial 5/50 (baseline established) | | 2:00 | Trial 10/50 (pruning active) | | 5:00 | Trial 25/50 (halfway) | | 8:00 | Trial 40/50 (late-stage) | | 10:00 | Trial 50/50 complete | | 10:10 | Results analysis + report generation | **Total**: 8-12 hours --- ## Monitoring Progress ### Real-Time Monitoring ```bash # Progress bar with ETA ./run_ppo_comprehensive_tuning.sh # (automatically monitors progress) ``` ### Manual Status Check ```bash # Get job ID from job_id.txt export JOB_ID=$(cat ml/trained_models/tuning/ppo_comprehensive/job_id.txt) # Check status cargo run -p tli -- tune status --job-id $JOB_ID ``` --- ## Results Location After completion (8-12 hours): ``` ml/trained_models/tuning/ppo_comprehensive/ ├── best_hyperparameters.txt ← USE THIS FOR PRODUCTION ├── TUNING_SUMMARY_REPORT.md ← SHARE WITH TEAM └── tuning_execution.log ← DEBUG IF NEEDED ``` --- ## Success Criteria ✅ **Target**: 5-10% improvement over baseline ✅ **Baseline**: Epoch 380 (explained_var=0.4469, EXCELLENT) ✅ **Expected**: Sharpe > 1.5, ExplainedVar > 0.45 --- ## After Tuning ### Step 1: Production Training (6-8 hours) ```bash # Use best hyperparameters for 500-epoch training cargo run -p ml --example train_ppo_production \ --config best_hyperparameters.txt \ --epochs 500 ``` ### Step 2: Checkpoint Analysis ```bash # Find optimal checkpoint (may not be epoch 500) cargo run -p ml --example analyze_ppo_checkpoints \ --checkpoint-dir ml/trained_models/production/ppo_tuned/ ``` ### Step 3: Backtesting ```bash # Test on all 4 symbols cargo run -p backtesting_service --example comprehensive_backtest \ --model ppo_tuned/ppo_final_epoch500.safetensors \ --symbols 6E.FUT,ZN.FUT,ES.FUT,NQ.FUT ``` --- ## Troubleshooting | Issue | Fix | |-------|-----| | GPU OOM | Auto-handled (batch size <= 230) | | Service down | `cargo run -p ml_training_service --release &` | | Missing data | Check `test_data/*.dbn.zst` files | | Slow progress | Check `nvidia-smi` (GPU utilization) | --- ## Documentation - **Full Guide**: `PPO_COMPREHENSIVE_TUNING_GUIDE.md` (20+ pages) - **Handoff Doc**: `AGENT_79_PPO_TUNING_HANDOFF.md` (technical details) - **Config File**: `tuning_config_ppo_comprehensive.yaml` (YAML) - **Execution Script**: `run_ppo_comprehensive_tuning.sh` (Bash) --- **Ready to Run** ✅ | **Configuration Complete** ✅ | **Expected: 8-12 hours** ⏱️ ```bash ./run_ppo_comprehensive_tuning.sh ```