Files
foxhunt/VALIDATION_QUICK_START.md
jgrusewski 4d0efa82df feat(wave1-2): Complete multi-model training architecture + TLI commands
Wave 1 (Architecture & Design - 5 agents):
- Multi-model training orchestration (DQN, PPO, MAMBA-2, TFT-INT8)
- Sequential training strategy (95.9% GPU headroom, 6.3min total)
- Hybrid multi-asset strategy (2x parallel, 22% GPU usage, 12-18min)
- Backward compatible gRPC API design with oneof pattern
- TDD test pyramid (67 tests: 24 unit + 28 integration + 15 E2E)
- Implementation roadmap (20 agents, 2.5 weeks, 13,280 LOC)

Wave 2 (Core TLI Commands - 5 agents):
- tli train start: Multi-model, multi-asset job submission (14 tests )
- tli train watch: Real-time streaming with weighted progress (10 tests )
- tli train status: Color-coded formatted status display (10 tests )
- tli train list: Filtering, sorting, pagination support (12 tests )
- tli train stop: Graceful cancellation with checkpoints (11 tests )

Status:
- 57/57 tests passing (100% TDD compliance)
- ~4,095 LOC (tests + implementation + docs)
- 3.5 hours actual vs 15-20 hours estimated (78% faster)
- Zero compilation errors, production-ready code
- Full documentation: WAVE_2_TLI_COMMANDS_COMPLETE.md

Next: Wave 3 (Multi-Asset Multi-Model Backend Logic - 5 agents)

🤖 Generated with Claude Code
Co-Authored-By: Claude <noreply@anthropic.com>
2025-10-22 20:50:43 +02:00

6.8 KiB

ML Training Service Validation - Quick Start Guide

Date: 2025-10-22 Status: PENDING (Waiting for agents 1-4)


What You Can Do RIGHT NOW (No Dependencies)

Phase 1: Service Health Check (15 minutes)

Step 1: Start ML Training Service

# Terminal 1: Start the service
cd /home/jgrusewski/Work/foxhunt
cargo run -p ml_training_service --release serve

# Expected output (wait 10-15 seconds):
# ✅ "ML Training Service ready"
# ✅ "gRPC server listening on 0.0.0.0:50053"
# ✅ "Prometheus metrics endpoint listening on http://0.0.0.0:9094"
# ✅ "Health check server on 0.0.0.0:8080"

Step 2: Check Health Endpoints

# Terminal 2: Test HTTP health endpoint
curl -v http://localhost:8080/health

# Expected: HTTP 200, {"status": "healthy", ...}

# Test Prometheus metrics endpoint
curl http://localhost:9094/metrics | grep ml_training

# Expected: ml_training_service_uptime_seconds, ml_training_jobs_total, etc.

Step 3: Verify Database Connection

# Check service logs for database connection
# Look for: "Database connection established"

# Verify database is running
docker-compose ps postgres

# Expected: State = Up

Step 4: Check GPU Detection (Optional)

# If GPU available
nvidia-smi

# Check service logs for:
# "GPU configuration loaded: device=0, memory=4GB"
# OR "GPU validation issues detected" (fallback to CPU is OK)

What is BLOCKED (Wait for Other Agents)

Phase 2: TLI Integration (30 min) - BLOCKED by Agent 1

Cannot test until tli ml train commands are implemented.

Required Commands:

  • tli ml train submit - Submit training job
  • tli ml train status <JOB_ID> - Get job status
  • tli ml train logs <JOB_ID> - Stream logs
  • tli ml train cancel <JOB_ID> - Cancel job
  • tli ml train list - List all jobs

ETA: Agent 1 completes in 4-6 hours


Phase 3: Parquet Training (1 hour) - BLOCKED by Agent 1

Cannot test via TLI until Agent 1 completes.

Workaround: Can run standalone training examples NOW:

# Test Parquet training directly (bypasses TLI)
cargo run -p ml --example train_tft_parquet --release -- \
  --parquet-file test_data/ES_FUT_small.parquet \
  --epochs 3 \
  --batch-size 2

# Expected: Training completes in ~5 minutes

Phase 4: Hyperparameter Tuning (2 hours) - BLOCKED by Agents 1 + 2

Cannot test until:

  • Agent 1: tli ml tune commands
  • Agent 2: gRPC tuning endpoints

ETA: Agents 1 + 2 complete in 10-14 hours


Phase 6: Resilience (30 min) - BLOCKED by Agent 1

Cannot test checkpoint recovery until TLI job submission works.


Success Criteria (Phase 1)

PASS (All Green)

  • Service starts without errors
  • No port conflicts (50053, 8080, 9094)
  • HTTP health endpoint returns 200
  • Metrics endpoint returns Prometheus data
  • Database connection established
  • Logs show "Training orchestrator started successfully"

FAIL (Red Flags)

  • Service crashes on startup
  • Port 50053 already in use (conflict)
  • Database connection failed
  • Health endpoint returns 500 or times out

Troubleshooting

Problem: Port 50053 already in use

# Check what's using the port
lsof -i :50053

# Kill the conflicting process
kill -9 <PID>

# Restart ML Training Service
cargo run -p ml_training_service --release serve

Problem: Database connection failed

# Check Docker services
docker-compose ps

# Expected: postgres, redis, vault all "Up"

# If postgres is down:
docker-compose up -d postgres

# Check DATABASE_URL environment variable
echo $DATABASE_URL

# Expected: postgresql://foxhunt:foxhunt_dev_password@localhost:5432/foxhunt

# Run migrations if needed
cargo sqlx migrate run

Problem: GPU validation failed

# Check if GPU is available
nvidia-smi

# If GPU not needed (CPU-only validation):
# This warning is EXPECTED and non-blocking
# Service will fall back to CPU training

Problem: Health endpoint times out

# Check if service is actually running
ps aux | grep ml_training_service

# Check service logs for errors
tail -f service.log

# Check firewall (if applicable)
sudo ufw status

What Happens Next?

After Phase 1 Passes

  1. Wait for Agent 1 to complete (TLI commands)
  2. Execute Phase 2: TLI Integration (30 min)
  3. Execute Phase 3: Parquet Training via TLI (1 hour)
  4. Wait for Agent 2 to complete (Tuning gRPC)
  5. Execute Phase 4: Hyperparameter Tuning (2 hours)
  6. Execute Phase 6: Resilience (30 min)

Total Time: 15-26 hours from now


After All Phases Pass

GO FOR CLOUD GPU DEPLOYMENT

Next steps:

  1. Create cloud GPU instance (e.g., AWS p3.2xlarge, $3.06/hour)
  2. Run smoke tests ($5-10)
  3. Deploy production workloads

If Any Phase Fails

NO-GO - Fix Locally First

Do NOT deploy to cloud GPU until all critical tests pass locally.

Why: Local failures → cloud GPU failures → wasted money ($50-$200 per attempt)


Full Documentation

  • Validation Plan (43KB, 1,460 lines): /home/jgrusewski/Work/foxhunt/AGENT_E2E_VALIDATION_PLAN.md

    • Detailed test cases for all 6 phases
    • Expected results and success criteria
    • Failure scenarios and debugging steps
    • Risk assessment and cost impact
  • Executive Summary (13KB, 446 lines): /home/jgrusewski/Work/foxhunt/AGENT_E2E_VALIDATION_SUMMARY.md

    • High-level overview of validation strategy
    • Dependency analysis
    • Go/No-Go decision framework
    • Time estimates and recommendations

Quick Reference

Service Endpoints

  • gRPC: localhost:50053 (ML Training Service API)
  • HTTP Health: http://localhost:8080/health
  • Prometheus Metrics: http://localhost:9094/metrics

Environment Variables (Required)

export DATABASE_URL="postgresql://foxhunt:foxhunt_dev_password@localhost:5432/foxhunt"
export GRPC_PORT=50053
export HEALTH_PORT=8080
export METRICS_PORT=9094

Test Data Files

test_data/
├── ES_FUT_small.parquet        # 25 KB, ~1,000 bars (quick tests)
├── NQ_FUT_small.parquet        # 27 KB, ~1,000 bars
├── 6E_FUT_small.parquet        # 23 KB, ~1,000 bars
├── ZN_FUT_90d_clean.parquet    # 65 KB, ~30,000 bars (medium test)
├── ES_FUT_180d.parquet         # 2.9 MB, ~90,000 bars (full training)

Bottom Line

Can you deploy to cloud GPU now? NO - Validation is blocked by Agents 1-4.

Can you start Phase 1 now? YES - No dependencies, takes 15 minutes.

Time to deployment-ready? 15-26 hours (Agent 1 + Agent 2 + validation execution)

Expected ROI? 💰 $85-$370 saved (by catching failures locally before cloud GPU)


Next Action: Execute Phase 1 health check (15 min) while waiting for agents 1-4 to complete.

Author: AGENT-VALIDATION-PLAN Date: 2025-10-22