Wave 1 (Architecture & Design - 5 agents): - Multi-model training orchestration (DQN, PPO, MAMBA-2, TFT-INT8) - Sequential training strategy (95.9% GPU headroom, 6.3min total) - Hybrid multi-asset strategy (2x parallel, 22% GPU usage, 12-18min) - Backward compatible gRPC API design with oneof pattern - TDD test pyramid (67 tests: 24 unit + 28 integration + 15 E2E) - Implementation roadmap (20 agents, 2.5 weeks, 13,280 LOC) Wave 2 (Core TLI Commands - 5 agents): - tli train start: Multi-model, multi-asset job submission (14 tests ✅) - tli train watch: Real-time streaming with weighted progress (10 tests ✅) - tli train status: Color-coded formatted status display (10 tests ✅) - tli train list: Filtering, sorting, pagination support (12 tests ✅) - tli train stop: Graceful cancellation with checkpoints (11 tests ✅) Status: - 57/57 tests passing (100% TDD compliance) - ~4,095 LOC (tests + implementation + docs) - 3.5 hours actual vs 15-20 hours estimated (78% faster) - Zero compilation errors, production-ready code - Full documentation: WAVE_2_TLI_COMMANDS_COMPLETE.md Next: Wave 3 (Multi-Asset Multi-Model Backend Logic - 5 agents) 🤖 Generated with Claude Code Co-Authored-By: Claude <noreply@anthropic.com>
274 lines
8.2 KiB
Markdown
274 lines
8.2 KiB
Markdown
# Agent 152: E2E 225-Feature Training Validation Checklist
|
|
|
|
**Date**: 2025-10-22
|
|
**Status**: ⏳ Pending Execution
|
|
**Test Plan**: See `AGENT_152_PHASE_5_E2E_TEST_PLAN.md`
|
|
**Quick Guide**: See `AGENT_152_PHASE_5_QUICK_SUMMARY.md`
|
|
|
|
---
|
|
|
|
## 🎯 Quick Test (Scenario 1) - **MANDATORY**
|
|
|
|
**Estimated Time**: 15 minutes
|
|
**Model**: PPO
|
|
**Dataset**: ES_FUT_small.parquet (~500 bars)
|
|
**Command**:
|
|
```bash
|
|
cargo run --release -p ml --example train_ppo_parquet -- \
|
|
--parquet-file test_data/ES_FUT_small.parquet \
|
|
--epochs 1 --batch-size 8 --no-early-stopping
|
|
```
|
|
|
|
### Pre-Flight Checks
|
|
- [ ] Working directory: `/home/jgrusewski/Work/foxhunt`
|
|
- [ ] File exists: `test_data/ES_FUT_small.parquet` (~25KB)
|
|
- [ ] Rust toolchain: `cargo --version` works
|
|
- [ ] (Optional) GPU available: `nvidia-smi` works
|
|
|
|
### Execution
|
|
- [ ] **STARTED**: Training command executed at: `____:____ (HH:MM)`
|
|
- [ ] Compilation completed without errors
|
|
- [ ] Training started without errors
|
|
|
|
### Critical Validation Points (IN ORDER)
|
|
|
|
#### 1. Data Loading
|
|
- [ ] ✅ Console shows: "📊 Loading market data from Parquet file..."
|
|
- [ ] ✅ Console shows: "✅ Loaded XXX OHLCV bars" (XXX should be ~100-500)
|
|
- [ ] ❌ No errors: "Failed to open Parquet file"
|
|
- [ ] ❌ No errors: "No bars loaded from Parquet file"
|
|
|
|
**Data loading result**: [ ] PASS [ ] FAIL
|
|
|
|
---
|
|
|
|
#### 2. Feature Extraction (**MOST CRITICAL**)
|
|
- [ ] ✅ Console shows: "🏗️ Extracting 225-dimensional feature vectors..."
|
|
- [ ] ✅ Console shows: "✅ Extracted XXX feature vectors **(dim=225, warmup bars skipped=50)**"
|
|
- **Required text**: Must see **"dim=225"** explicitly
|
|
- Feature count should be ~50 bars less than loaded bars (due to warmup)
|
|
- [ ] ✅ Console shows: "✅ Feature extraction complete: XXX samples"
|
|
- [ ] ❌ No errors: "State dimension mismatch: expected 225, got X"
|
|
- [ ] ❌ No errors: "No features extracted"
|
|
|
|
**Feature extraction result**: [ ] PASS [ ] FAIL
|
|
|
|
**If FAIL**: See Troubleshooting → Issue 2 in test plan
|
|
|
|
---
|
|
|
|
#### 3. Model Initialization
|
|
- [ ] ✅ Console shows: "✅ PPO trainer initialized **(state_dim=225)**"
|
|
- **Required text**: Must see **"state_dim=225"** explicitly
|
|
- [ ] ❌ No errors: "Failed to create PPO trainer"
|
|
- [ ] ❌ No errors: "CUDA out of memory" (if GPU is used)
|
|
|
|
**Model initialization result**: [ ] PASS [ ] FAIL
|
|
|
|
**If GPU OOM**: Retry with `--batch-size 4`
|
|
|
|
---
|
|
|
|
#### 4. Training Execution
|
|
- [ ] ✅ Console shows: "🏋️ Starting training..."
|
|
- [ ] ✅ Console shows epoch progress: "📊 Epoch 1/1: policy_loss=X.XXXX, value_loss=X.XXXX, ..."
|
|
- [ ] ✅ Training completes: "✅ Training completed successfully!"
|
|
- [ ] ❌ No errors: "Training failed"
|
|
- [ ] ❌ No NaN/Inf errors: "loss=nan" or "loss=inf"
|
|
|
|
**Training execution result**: [ ] PASS [ ] FAIL
|
|
|
|
---
|
|
|
|
#### 5. Model Checkpointing
|
|
- [ ] ✅ Console shows: "💾 Final checkpoint saved to: ml/trained_models/ppo_checkpoint_epoch_1.safetensors"
|
|
- [ ] ✅ Console shows: "📁 Model files saved to: ml/trained_models"
|
|
- [ ] ❌ No errors: "Failed to save checkpoint"
|
|
|
|
**Model checkpointing result**: [ ] PASS [ ] FAIL
|
|
|
|
---
|
|
|
|
#### 6. Policy Convergence (Optional, but indicates quality)
|
|
- [ ] Console shows: "🔍 Policy Convergence Analysis:"
|
|
- [ ] Console shows: "• Policy updates (KL > 0): X" where X > 0
|
|
- [ ] Console shows: "✅ PASS: Policy updates detected (KL divergence > 0)"
|
|
- [ ] Console shows: "✅ PASS: Value network learning (explained variance > 0.5)"
|
|
|
|
**Policy convergence result**: [ ] PASS [ ] PARTIAL [ ] FAIL
|
|
|
|
**Note**: Even if FAIL, this is non-critical for quick validation
|
|
|
|
---
|
|
|
|
### Post-Execution Verification
|
|
|
|
#### Filesystem Checks
|
|
```bash
|
|
# Run these commands:
|
|
ls -lh ml/trained_models/ppo_*.safetensors | tail -5
|
|
```
|
|
|
|
- [ ] ✅ File exists: `ml/trained_models/ppo_actor_epoch_1.safetensors`
|
|
- [ ] ✅ File exists: `ml/trained_models/ppo_critic_epoch_1.safetensors`
|
|
- [ ] ✅ File exists: `ml/trained_models/ppo_checkpoint_epoch_1.safetensors`
|
|
- [ ] ✅ All 3 files have size >1MB (non-zero)
|
|
- [ ] ✅ Timestamps match training run time
|
|
|
|
**Filesystem result**: [ ] PASS [ ] FAIL
|
|
|
|
---
|
|
|
|
### Training Metrics
|
|
|
|
**Actual Results** (fill in from console output):
|
|
- **Loaded bars**: ______
|
|
- **Feature vectors**: ______ (should be ~50 less than loaded bars)
|
|
- **Feature dimension**: ______ (MUST be 225)
|
|
- **State dimension**: ______ (MUST be 225)
|
|
- **Epochs completed**: 1 / 1
|
|
- **Training time**: ______ seconds
|
|
- **Policy updates (KL > 0)**: ______ / 1
|
|
- **Final policy loss**: ______
|
|
- **Final value loss**: ______
|
|
- **Explained variance**: ______
|
|
- **GPU used**: [ ] Yes [ ] No (CPU fallback)
|
|
|
|
---
|
|
|
|
## 🎯 Overall Scenario 1 Result
|
|
|
|
**PASSED** if ALL of the following are true:
|
|
1. ✅ Feature extraction shows "dim=225"
|
|
2. ✅ Model initialization shows "state_dim=225"
|
|
3. ✅ Training completes 1 epoch without errors
|
|
4. ✅ All 3 model files saved successfully
|
|
5. ✅ File sizes are non-zero (>1MB)
|
|
|
|
**Final Result**: [ ] ✅ PASS [ ] ⚠️ PARTIAL [ ] ❌ FAIL
|
|
|
|
**Completion Time**: ____:____ (HH:MM)
|
|
**Total Duration**: ______ minutes
|
|
|
|
---
|
|
|
|
## 📝 Notes / Issues Encountered
|
|
|
|
**Console Output** (paste key lines):
|
|
```
|
|
[Paste critical console output here, especially lines showing "225-dimensional" and "state_dim=225"]
|
|
```
|
|
|
|
**Errors Encountered**:
|
|
```
|
|
[Paste any error messages here]
|
|
```
|
|
|
|
**Troubleshooting Steps Taken**:
|
|
- [ ] N/A - No issues
|
|
- [ ] Reduced batch size from 8 to 4 (GPU OOM)
|
|
- [ ] Rebuilt ML crate (`cargo clean -p ml && cargo build --release -p ml`)
|
|
- [ ] Other: ___________________________________
|
|
|
|
---
|
|
|
|
## 🔬 Medium Test (Scenario 2) - **OPTIONAL**
|
|
|
|
**Status**: [ ] Not Started [ ] In Progress [ ] Complete [ ] Skipped
|
|
|
|
**Model**: MAMBA-2
|
|
**Dataset**: ZN_FUT_90d_clean.parquet (~3.5K bars)
|
|
**Epochs**: 10
|
|
**Estimated Time**: 30 minutes
|
|
|
|
### Key Validations
|
|
- [ ] Wave D configuration detected: `feature_count: 225`
|
|
- [ ] Input shape validated: `[1, 60, 225]`
|
|
- [ ] All 10 epochs complete without errors
|
|
- [ ] Checkpoints saved at epoch 10
|
|
- [ ] `training_losses.csv` and `training_metrics.json` created
|
|
|
|
**Result**: [ ] PASS [ ] FAIL [ ] SKIPPED
|
|
|
|
**Notes**:
|
|
```
|
|
[Optional: Paste any notes from Scenario 2]
|
|
```
|
|
|
|
---
|
|
|
|
## 🏆 Full Test (Scenario 3) - **OPTIONAL**
|
|
|
|
**Status**: [ ] Not Started [ ] In Progress [ ] Complete [ ] Skipped
|
|
|
|
**Models**: TFT + PPO
|
|
**Dataset**: ES_FUT_180d.parquet (~12.5K bars)
|
|
**Epochs**: 20 (TFT) + 30 (PPO)
|
|
**Estimated Time**: 1 hour
|
|
|
|
### Key Validations
|
|
- [ ] TFT trained for 20 epochs with 225 features
|
|
- [ ] PPO trained for 30 epochs with 225 features
|
|
- [ ] Policy convergence achieved (KL divergence > 0)
|
|
- [ ] Value network learning (explained variance > 0.5)
|
|
- [ ] Cross-model validation passes (TFT + PPO + MAMBA-2)
|
|
|
|
**Result**: [ ] PASS [ ] FAIL [ ] SKIPPED
|
|
|
|
**Notes**:
|
|
```
|
|
[Optional: Paste any notes from Scenario 3]
|
|
```
|
|
|
|
---
|
|
|
|
## ✅ Final Certification
|
|
|
|
**Agent 152 Phase 5 Validation**: [ ] ✅ CERTIFIED [ ] ⚠️ PARTIAL [ ] ❌ FAILED
|
|
|
|
**Certification Requirements** (all must be checked):
|
|
- [ ] Scenario 1 (Quick Test) completed successfully
|
|
- [ ] 225-dimensional feature extraction confirmed
|
|
- [ ] Model training with 225 input features confirmed
|
|
- [ ] Model checkpoints saved and verified
|
|
- [ ] No critical errors or blockers encountered
|
|
|
|
**Certified By**: ________________ (Your name/ID)
|
|
**Date**: 2025-10-__
|
|
**Time**: ____:____ (HH:MM)
|
|
|
|
---
|
|
|
|
## 📤 Deliverables
|
|
|
|
After successful validation, attach:
|
|
1. ✅ This completed checklist
|
|
2. ✅ Console output log (full training session)
|
|
3. ✅ Screenshot showing "225-dimensional feature vectors" line
|
|
4. ✅ `ls -lh ml/trained_models/` output showing saved models
|
|
5. ⏳ (Optional) `training_metrics.json` from MAMBA-2 if Scenario 2 completed
|
|
|
|
---
|
|
|
|
## 🚀 Next Steps
|
|
|
|
**After Scenario 1 PASS**:
|
|
1. [ ] Proceed to production training (90-180 day datasets)
|
|
2. [ ] Integrate with ML Training Service (gRPC testing)
|
|
3. [ ] Deploy to cloud GPU (if local GPU is insufficient)
|
|
4. [ ] Test TLI commands: `tli tune start --model PPO`
|
|
|
|
**After Scenario 2 PASS** (optional):
|
|
1. [ ] Multi-model retraining (DQN, PPO, MAMBA-2, TFT)
|
|
2. [ ] Hyperparameter tuning with Optuna
|
|
3. [ ] Wave Comparison Backtest (Wave C vs Wave D)
|
|
|
|
**After Scenario 3 PASS** (optional):
|
|
1. [ ] Full production deployment
|
|
2. [ ] Live paper trading with 225-feature models
|
|
3. [ ] Performance monitoring and validation
|
|
|
|
---
|
|
|
|
**END OF CHECKLIST** - Good luck with validation! 🚀
|