# ML Training Service Validation - Quick Start Guide **Date**: 2025-10-22 **Status**: ⏳ PENDING (Waiting for agents 1-4) --- ## What You Can Do RIGHT NOW (No Dependencies) ### ✅ Phase 1: Service Health Check (15 minutes) #### Step 1: Start ML Training Service ```bash # Terminal 1: Start the service cd /home/jgrusewski/Work/foxhunt cargo run -p ml_training_service --release serve # Expected output (wait 10-15 seconds): # ✅ "ML Training Service ready" # ✅ "gRPC server listening on 0.0.0.0:50053" # ✅ "Prometheus metrics endpoint listening on http://0.0.0.0:9094" # ✅ "Health check server on 0.0.0.0:8080" ``` #### Step 2: Check Health Endpoints ```bash # Terminal 2: Test HTTP health endpoint curl -v http://localhost:8080/health # Expected: HTTP 200, {"status": "healthy", ...} # Test Prometheus metrics endpoint curl http://localhost:9094/metrics | grep ml_training # Expected: ml_training_service_uptime_seconds, ml_training_jobs_total, etc. ``` #### Step 3: Verify Database Connection ```bash # Check service logs for database connection # Look for: "Database connection established" # Verify database is running docker-compose ps postgres # Expected: State = Up ``` #### Step 4: Check GPU Detection (Optional) ```bash # If GPU available nvidia-smi # Check service logs for: # "GPU configuration loaded: device=0, memory=4GB" # OR "GPU validation issues detected" (fallback to CPU is OK) ``` --- ## What is BLOCKED (Wait for Other Agents) ### ⏳ Phase 2: TLI Integration (30 min) - **BLOCKED by Agent 1** Cannot test until `tli ml train` commands are implemented. **Required Commands**: - `tli ml train submit` - Submit training job - `tli ml train status ` - Get job status - `tli ml train logs ` - Stream logs - `tli ml train cancel ` - Cancel job - `tli ml train list` - List all jobs **ETA**: Agent 1 completes in 4-6 hours --- ### ⏳ Phase 3: Parquet Training (1 hour) - **BLOCKED by Agent 1** Cannot test via TLI until Agent 1 completes. **Workaround**: Can run standalone training examples NOW: ```bash # Test Parquet training directly (bypasses TLI) cargo run -p ml --example train_tft_parquet --release -- \ --parquet-file test_data/ES_FUT_small.parquet \ --epochs 3 \ --batch-size 2 # Expected: Training completes in ~5 minutes ``` --- ### ⏳ Phase 4: Hyperparameter Tuning (2 hours) - **BLOCKED by Agents 1 + 2** Cannot test until: - Agent 1: `tli ml tune` commands - Agent 2: gRPC tuning endpoints **ETA**: Agents 1 + 2 complete in 10-14 hours --- ### ⏳ Phase 6: Resilience (30 min) - **BLOCKED by Agent 1** Cannot test checkpoint recovery until TLI job submission works. --- ## Success Criteria (Phase 1) ### ✅ PASS (All Green) - [ ] Service starts without errors - [ ] No port conflicts (50053, 8080, 9094) - [ ] HTTP health endpoint returns 200 - [ ] Metrics endpoint returns Prometheus data - [ ] Database connection established - [ ] Logs show "Training orchestrator started successfully" ### ❌ FAIL (Red Flags) - Service crashes on startup - Port 50053 already in use (conflict) - Database connection failed - Health endpoint returns 500 or times out --- ## Troubleshooting ### Problem: Port 50053 already in use ```bash # Check what's using the port lsof -i :50053 # Kill the conflicting process kill -9 # Restart ML Training Service cargo run -p ml_training_service --release serve ``` --- ### Problem: Database connection failed ```bash # Check Docker services docker-compose ps # Expected: postgres, redis, vault all "Up" # If postgres is down: docker-compose up -d postgres # Check DATABASE_URL environment variable echo $DATABASE_URL # Expected: postgresql://foxhunt:foxhunt_dev_password@localhost:5432/foxhunt # Run migrations if needed cargo sqlx migrate run ``` --- ### Problem: GPU validation failed ```bash # Check if GPU is available nvidia-smi # If GPU not needed (CPU-only validation): # This warning is EXPECTED and non-blocking # Service will fall back to CPU training ``` --- ### Problem: Health endpoint times out ```bash # Check if service is actually running ps aux | grep ml_training_service # Check service logs for errors tail -f service.log # Check firewall (if applicable) sudo ufw status ``` --- ## What Happens Next? ### After Phase 1 Passes ✅ 1. Wait for Agent 1 to complete (TLI commands) 2. Execute Phase 2: TLI Integration (30 min) 3. Execute Phase 3: Parquet Training via TLI (1 hour) 4. Wait for Agent 2 to complete (Tuning gRPC) 5. Execute Phase 4: Hyperparameter Tuning (2 hours) 6. Execute Phase 6: Resilience (30 min) **Total Time**: 15-26 hours from now --- ### After All Phases Pass ✅ **GO FOR CLOUD GPU DEPLOYMENT** Next steps: 1. Create cloud GPU instance (e.g., AWS p3.2xlarge, $3.06/hour) 2. Run smoke tests ($5-10) 3. Deploy production workloads --- ### If Any Phase Fails ❌ **NO-GO - Fix Locally First** Do NOT deploy to cloud GPU until all critical tests pass locally. **Why**: Local failures → cloud GPU failures → wasted money ($50-$200 per attempt) --- ## Full Documentation - **Validation Plan** (43KB, 1,460 lines): `/home/jgrusewski/Work/foxhunt/AGENT_E2E_VALIDATION_PLAN.md` - Detailed test cases for all 6 phases - Expected results and success criteria - Failure scenarios and debugging steps - Risk assessment and cost impact - **Executive Summary** (13KB, 446 lines): `/home/jgrusewski/Work/foxhunt/AGENT_E2E_VALIDATION_SUMMARY.md` - High-level overview of validation strategy - Dependency analysis - Go/No-Go decision framework - Time estimates and recommendations --- ## Quick Reference ### Service Endpoints - **gRPC**: `localhost:50053` (ML Training Service API) - **HTTP Health**: `http://localhost:8080/health` - **Prometheus Metrics**: `http://localhost:9094/metrics` ### Environment Variables (Required) ```bash export DATABASE_URL="postgresql://foxhunt:foxhunt_dev_password@localhost:5432/foxhunt" export GRPC_PORT=50053 export HEALTH_PORT=8080 export METRICS_PORT=9094 ``` ### Test Data Files ``` test_data/ ├── ES_FUT_small.parquet # 25 KB, ~1,000 bars (quick tests) ├── NQ_FUT_small.parquet # 27 KB, ~1,000 bars ├── 6E_FUT_small.parquet # 23 KB, ~1,000 bars ├── ZN_FUT_90d_clean.parquet # 65 KB, ~30,000 bars (medium test) ├── ES_FUT_180d.parquet # 2.9 MB, ~90,000 bars (full training) ``` --- ## Bottom Line **Can you deploy to cloud GPU now?** ❌ **NO** - Validation is blocked by Agents 1-4. **Can you start Phase 1 now?** ✅ **YES** - No dependencies, takes 15 minutes. **Time to deployment-ready?** ⏳ **15-26 hours** (Agent 1 + Agent 2 + validation execution) **Expected ROI?** 💰 **$85-$370 saved** (by catching failures locally before cloud GPU) --- **Next Action**: Execute Phase 1 health check (15 min) while waiting for agents 1-4 to complete. **Author**: AGENT-VALIDATION-PLAN **Date**: 2025-10-22