Wave 1 (Architecture & Design - 5 agents): - Multi-model training orchestration (DQN, PPO, MAMBA-2, TFT-INT8) - Sequential training strategy (95.9% GPU headroom, 6.3min total) - Hybrid multi-asset strategy (2x parallel, 22% GPU usage, 12-18min) - Backward compatible gRPC API design with oneof pattern - TDD test pyramid (67 tests: 24 unit + 28 integration + 15 E2E) - Implementation roadmap (20 agents, 2.5 weeks, 13,280 LOC) Wave 2 (Core TLI Commands - 5 agents): - tli train start: Multi-model, multi-asset job submission (14 tests ✅) - tli train watch: Real-time streaming with weighted progress (10 tests ✅) - tli train status: Color-coded formatted status display (10 tests ✅) - tli train list: Filtering, sorting, pagination support (12 tests ✅) - tli train stop: Graceful cancellation with checkpoints (11 tests ✅) Status: - 57/57 tests passing (100% TDD compliance) - ~4,095 LOC (tests + implementation + docs) - 3.5 hours actual vs 15-20 hours estimated (78% faster) - Zero compilation errors, production-ready code - Full documentation: WAVE_2_TLI_COMMANDS_COMPLETE.md Next: Wave 3 (Multi-Asset Multi-Model Backend Logic - 5 agents) 🤖 Generated with Claude Code Co-Authored-By: Claude <noreply@anthropic.com>
44 KiB
AGENT-35: Cloud GPU Recommendation for Production ML Training
Date: 2025-10-21 Status: ✅ COMPLETE Agent: AGENT-35 (Cloud GPU Research & Recommendation) Context: Wave 12 ML Production - Cloud GPU strategy for 90-180 day training datasets
🎯 Executive Summary
Recommendation: RunPod Community Cloud (RTX 4090) for immediate production training.
Key Findings:
- Best Value: RunPod RTX 4090 @ $0.34/hr (spot pricing)
- Estimated 4-Model Training Cost: $3.40 - $6.80 (10-20 hours total)
- ROI vs Local GPU Upgrade: Cloud wins by $1,293 - $1,596 (avoiding $1,600-$2,400 RTX 4090 purchase)
- Memory Headroom: RTX 4090 (24GB) provides 5.5x more VRAM than current RTX 3050 Ti (4GB)
- Speedup: Expected 3-5x faster training vs RTX 3050 Ti (based on CUDA core count: 16,384 vs 2,560)
📊 Cloud GPU Provider Comparison
Tier 1: Budget-Friendly Options (Recommended)
| Provider | GPU Model | VRAM | $/hour (Spot) | $/hour (On-Demand) | Billing | Docker | SSH | CUDA 11.8+ | Rust Support | Notes |
|---|---|---|---|---|---|---|---|---|---|---|
| 🏆 RunPod | RTX 4090 | 24GB | $0.34 | $0.59 | Per-second | ✅ | ✅ | ✅ 12.1 | ✅ | BEST VALUE - Pay-per-second, 31 regions |
| Vast.ai | RTX 4090 | 24GB | $0.29 | $0.61 | Per-hour | ✅ | ✅ | ✅ 12.x | ✅ | Marketplace (variable availability) |
| RunPod | RTX 3090 | 24GB | $0.27 | $0.43 | Per-second | ✅ | ✅ | ✅ 11.8 | ✅ | Lower CUDA cores (10,496 vs 16,384) |
| Vast.ai | A100 PCIe | 40GB | $0.67 | $1.20 | Per-hour | ✅ | ✅ | ✅ 11.8+ | ✅ | More VRAM, better for large models |
| RunPod | A100 80GB | 80GB | $1.64 | $1.99 | Per-second | ✅ | ✅ | ✅ 11.8+ | ✅ | Overkill for our 440MB memory budget |
Tier 2: Mid-Range (Enterprise Support)
| Provider | GPU Model | VRAM | $/hour (On-Demand) | Billing | Support | CUDA | Notes |
|---|---|---|---|---|---|---|---|
| Lambda Labs | A100 80GB (1x GPU) | 80GB | $3.29 | Per-hour | Premium | ✅ 11.8+ | ML-focused, JupyterLab pre-configured |
| Lambda Labs | A100 80GB (8x GPU) | 640GB | $1.79/GPU ($14.32 total) | Per-hour | Premium | ✅ 11.8+ | Multi-GPU training (not needed for us) |
| Lambda Labs | H100 80GB (1x GPU) | 80GB | $3.29 | Per-hour | Premium | ✅ 12.0+ | Latest generation (overkill) |
| DigitalOcean | H100 80GB | 80GB | $1.99 | Per-hour | Excellent | ✅ 12.0+ | New offering, great developer UX |
Tier 3: Major Cloud Providers (AWS/GCP/Azure)
| Provider | Instance Type | GPU Model | VRAM | $/hour (On-Demand) | $/hour (Spot) | Savings Plans | Notes |
|---|---|---|---|---|---|---|---|
| AWS | ml.g5.xlarge | A10G | 24GB | $1.41 | ~$0.42 (70% off) | Up to 64% off | SageMaker integration |
| AWS | ml.p4d.24xlarge | A100 (8x) | 320GB | $32.77 | ~$9.83 (70% off) | Up to 64% off | Multi-GPU (overkill) |
| AWS | p3.2xlarge | V100 | 16GB | $3.06 | ~$0.92 (70% off) | N/A | Older generation |
| GCP | n1-standard-8 + T4 | T4 | 16GB | $0.56 | ~$0.17 (70% off) | N/A | Budget option |
| GCP | a2-highgpu-1g | A100 | 40GB | $3.67 | ~$1.10 (70% off) | N/A | Enterprise support |
| Azure | NC A100 v4 | A100 | 80GB | ~$3.67 | Variable | Reserved | Enterprise compliance |
Tier 4: Specialized ML Platforms
| Provider | GPU Options | $/hour Range | Strengths | Weaknesses |
|---|---|---|---|---|
| Paperspace | RTX 4000, A100 | $0.51 - $2.00 | Free tier, pre-configured ML | Limited GPU selection |
| Jarvis Labs | RTX 3090, A100 | $0.49 - $1.89 | Developer-focused | Smaller community |
| Genesis Cloud | RTX 3090, A100 | $0.40 - $1.80 | Reliable uptime | Less flexible than RunPod |
| Hyperstack | A100, H100 | $1.42 - $2.74 | Production reliability | Higher base price |
💰 Cost Analysis: 4-Model Training (90-180 Day Datasets)
Assumptions
- Dataset: 90-180 days per asset (ES.FUT, NQ.FUT, 6E.FUT, ZN.FUT)
- Training Time Estimates (based on current RTX 3050 Ti benchmarks, scaled by CUDA cores):
- MAMBA-2: 1.9 min (current) → 0.5 min (RTX 4090, 6.4x CUDA cores)
- DQN: 15 sec (current) → 5 sec (RTX 4090)
- PPO: 7 sec (current) → 2 sec (RTX 4090)
- TFT-INT8: 3.5 min (current) → 1.0 min (RTX 4090)
- Total per asset: ~5.5 min (current) → ~1.5 min (RTX 4090)
- 4 Assets: 4 × 1.5 min = 6 min (optimistic)
- 50 Epochs: 6 min × 50 = 300 min = 5 hours
- Safety Buffer (2x): 5 hours × 2 = 10 hours (realistic estimate)
- Hyperparameter Tuning: 10 hours × 2 = 20 hours (with tuning)
Cost Breakdown by Provider
| Provider | GPU | $/hour | 10h Cost | 20h Cost | 100h Cost (Monthly) | Notes |
|---|---|---|---|---|---|---|
| 🥇 RunPod (Spot) | RTX 4090 | $0.34 | $3.40 | $6.80 | $34.00 | BEST VALUE |
| 🥈 Vast.ai | RTX 4090 | $0.29 | $2.90 | $5.80 | $29.00 | Slightly cheaper, less stable |
| 🥉 RunPod (Spot) | RTX 3090 | $0.27 | $2.70 | $5.40 | $27.00 | Slower (40% fewer CUDA cores) |
| Vast.ai | A100 PCIe | $0.67 | $6.70 | $13.40 | $67.00 | More VRAM, overkill for us |
| RunPod (On-Demand) | RTX 4090 | $0.59 | $5.90 | $11.80 | $59.00 | Guaranteed availability |
| Lambda Labs | A100 80GB | $3.29 | $32.90 | $65.80 | $329.00 | Premium support |
| AWS (Spot) | ml.g5.xlarge | $0.42 | $4.20 | $8.40 | $42.00 | SageMaker integration |
| AWS (On-Demand) | ml.g5.xlarge | $1.41 | $14.10 | $28.20 | $141.00 | Enterprise support |
Annual Cost Projection (Monthly Retraining)
| Provider | GPU | $/month (10h) | $/year | vs Local GPU ($1,600) | ROI Year 1 |
|---|---|---|---|---|---|
| RunPod | RTX 4090 | $34.00 | $408 | +$1,192 savings | 293% ROI |
| Vast.ai | RTX 4090 | $29.00 | $348 | +$1,252 savings | 360% ROI |
| Lambda Labs | A100 80GB | $329.00 | $3,948 | -$2,348 loss | -147% ROI |
| AWS (Spot) | ml.g5.xlarge | $42.00 | $504 | +$1,096 savings | 217% ROI |
Verdict: Cloud GPU (RunPod/Vast.ai) is drastically cheaper than buying local hardware for infrequent training.
🏆 Top 3 Recommendations
🥇 #1: RunPod Community Cloud (RTX 4090) - RECOMMENDED
Why This is the Best Choice:
- ✅ Best Price-Performance: $0.34/hr spot pricing (68% cheaper than on-demand)
- ✅ Per-Second Billing: Only pay for actual training time (no waste)
- ✅ Proven Rust Support: Docker with CUDA 12.1 pre-configured
- ✅ 24GB VRAM: 5.5x more than RTX 3050 Ti (handles all 4 models easily)
- ✅ 31 Global Regions: Low latency, high availability
- ✅ SSH + Jupyter: Full development environment access
- ✅ Auto-Shutdown: No accidental overcharges
- ✅ Community Support: Large user base, active Discord
Cost for Our Use Case:
- Initial Training (50 epochs, 4 models): $3.40 - $6.80 (10-20 hours)
- Monthly Retraining (10h/month): $34.00/month
- Annual Cost: $408/year (vs $1,600 local GPU purchase)
Integration Complexity: ⭐⭐⭐⭐☆ (4/5 - Easy)
- Upload Parquet files via SSH/SCP
- Run training commands via SSH
- Download checkpoints via SCP/rsync
- Minimal ML Training Service changes required
Limitations:
- ⚠️ Spot instances can be preempted (rare, but possible)
- ⚠️ Requires manual instance management (start/stop)
- ⚠️ No native integration with ML Training Service (requires custom script)
When to Use:
- ✅ Production training (monthly retraining cycles)
- ✅ Hyperparameter tuning experiments
- ✅ Cost-sensitive projects
- ✅ Rapid iteration workflows
🥈 #2: AWS SageMaker (ml.g5.xlarge Spot) - ENTERPRISE CHOICE
Why Consider This:
- ✅ AWS Ecosystem Integration: S3, IAM, CloudWatch, EventBridge
- ✅ Spot Instances: 70% savings ($1.41 → $0.42/hr)
- ✅ Enterprise Support: 24/7 support, SLAs, compliance (SOC2, HIPAA)
- ✅ Managed ML Pipeline: SageMaker Training Jobs API
- ✅ Auto-Scaling: Scale to multiple GPUs if needed
- ✅ Built-in Monitoring: CloudWatch metrics, logs, alarms
Cost for Our Use Case:
- Initial Training (50 epochs, 4 models): $4.20 - $8.40 (10-20 hours)
- Monthly Retraining (10h/month): $42.00/month
- Annual Cost: $504/year
Integration Complexity: ⭐⭐⭐☆☆ (3/5 - Moderate)
- Requires AWS SDK integration in ML Training Service
- Dockerfile for SageMaker container
- S3 for data storage (upload/download)
- IAM roles and policies
- EventBridge for scheduling
Limitations:
- ⚠️ Higher cost than RunPod ($42/month vs $34/month)
- ⚠️ Spot instances less predictable than on-demand
- ⚠️ AWS learning curve (IAM, S3, SageMaker APIs)
- ⚠️ Minimum 1-minute billing (vs RunPod's per-second)
When to Use:
- ✅ Already using AWS infrastructure
- ✅ Need enterprise compliance (SOC2, HIPAA)
- ✅ Want managed ML pipeline
- ✅ Scaling to multi-GPU training in future
🥉 #3: Lambda Labs (A100 80GB) - PREMIUM CHOICE
Why Consider This:
- ✅ ML-Optimized: Pre-configured PyTorch, TensorFlow, CUDA 11.8+
- ✅ Zero Setup: JupyterLab, SSH, Docker ready on boot
- ✅ Premium Support: Fast response times, ML expertise
- ✅ Latest Hardware: H100, H200 early access
- ✅ 80GB VRAM: Future-proof for larger models
- ✅ Persistent Storage: NVMe SSD included (19.5TB)
Cost for Our Use Case:
- Initial Training (50 epochs, 4 models): $32.90 - $65.80 (10-20 hours)
- Monthly Retraining (10h/month): $329.00/month
- Annual Cost: $3,948/year
Integration Complexity: ⭐⭐⭐⭐⭐ (5/5 - Very Easy)
- Pre-configured CUDA environment
- One-click JupyterLab access
- SSH access for automation
- Minimal setup required
Limitations:
- ⚠️ 10x more expensive than RunPod ($329/month vs $34/month)
- ⚠️ Overkill for our 440MB memory budget
- ⚠️ Per-hour billing (no per-second option)
- ⚠️ Limited global regions (US-only for some GPUs)
When to Use:
- ✅ Need premium support
- ✅ Training very large models (>10GB weights)
- ✅ Budget is not a constraint
- ✅ Want zero-friction ML development
🏗️ Integration Architecture
Current Architecture (Local GPU)
┌─────────────────────────────────────────────────────────────┐
│ Foxhunt System (Local) │
│ │
│ ┌──────────────┐ ┌──────────────┐ ┌──────────┐ │
│ │ ML Training │─────▶│ RTX 3050 Ti │─────▶│ Models │ │
│ │ Service │ │ (4GB VRAM) │ │ (.safetensors)│
│ │ (Port 50054)│ │ │ │ │ │
│ └──────────────┘ └──────────────┘ └──────────┘ │
│ │ │ │
│ │ │ │
│ ▼ ▼ │
│ ┌──────────────┐ ┌──────────────┐ │
│ │ Parquet │ │ PostgreSQL │ │
│ │ Files │ │ (Metadata) │ │
│ │ (test_data/) │ │ │ │
│ └──────────────┘ └──────────────┘ │
└─────────────────────────────────────────────────────────────┘
Hybrid Architecture (Cloud GPU Training)
┌─────────────────────────────────────────────────────────────────────┐
│ Foxhunt System (Local) │
│ │
│ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐ │
│ │ ML Training │─────▶│ Cloud GPU │─────▶│ Models │ │
│ │ Service │ gRPC │ Orchestrator│ │ (.safetensors)│ │
│ │ (Port 50054)│ │ │ │ │ │
│ └──────────────┘ └──────────────┘ └──────────────┘ │
│ │ │ │ │
│ │ │ │ │
│ │ ▼ ▼ │
│ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐ │
│ │ Parquet │─────▶│ S3 Bucket │ │ PostgreSQL │ │
│ │ Files │ rsync│ (Training │ │ (Metadata) │ │
│ │ (test_data/) │ │ Data) │ │ │ │
│ └──────────────┘ └──────────────┘ └──────────────┘ │
└─────────────────────────┼────────────────────────────────────────────┘
│
│ SSH/API
▼
┌─────────────────────────────────────────────────────────────────────┐
│ RunPod Cloud GPU │
│ │
│ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐ │
│ │ Docker │ │ RTX 4090 │ │ Training │ │
│ │ Container │─────▶│ (24GB VRAM) │─────▶│ Checkpoint │ │
│ │ (Rust+CUDA) │ │ │ │ (.safetensors)│ │
│ └──────────────┘ └──────────────┘ └──────────────┘ │
│ │ │ │
│ │ │ │
│ ▼ ▼ │
│ ┌──────────────┐ ┌──────────────┐ │
│ │ Parquet │ │ Auto- │ │
│ │ Files │ │ Shutdown │ │
│ │ (downloaded) │ │ (No Waste) │ │
│ └──────────────┘ └──────────────┘ │
└─────────────────────────────────────────────────────────────────────┘
TLI Integration (Future Enhancement)
┌──────────────────────────────────────────────────────────────────┐
│ TLI (Terminal Client) │
│ │
│ $ tli ml train --cloud runpod --gpu rtx-4090 \ │
│ --model mamba2 --epochs 50 --asset ES.FUT │
│ │
│ ┌─────────────────────────────────────────────────────────┐ │
│ │ 1. Upload Parquet files to RunPod via rsync │ │
│ │ 2. Spin up RTX 4090 instance (spot, $0.34/hr) │ │
│ │ 3. Execute training command via SSH │ │
│ │ 4. Monitor progress (live logs streaming) │ │
│ │ 5. Download checkpoints to ml/trained_models/ │ │
│ │ 6. Auto-shutdown instance (stop billing) │ │
│ └─────────────────────────────────────────────────────────┘ │
│ │
│ Output: │
│ ✅ Training complete in 10.5 hours │
│ 💾 Checkpoints downloaded to ml/trained_models/ │
│ 💰 Total cost: $3.57 (10.5h × $0.34/hr) │
│ 🔌 Instance terminated (no ongoing charges) │
└──────────────────────────────────────────────────────────────────┘
🛣️ Migration Path (Local → Cloud Training)
Phase 1: Manual Cloud Training (Immediate - 1 week)
Goal: Validate cloud GPU workflow without code changes
Steps:
-
Setup RunPod Account
# Sign up at https://www.runpod.io/ # Add payment method # Generate API key -
Create Docker Image with Rust + CUDA
# Dockerfile FROM nvidia/cuda:12.1.0-devel-ubuntu22.04 # Install Rust RUN curl --proto '=https' --tlsv1.2 -sSf https://sh.rustup.rs | sh -s -- -y ENV PATH="/root/.cargo/bin:${PATH}" # Install dependencies RUN apt-get update && apt-get install -y \ build-essential \ libssl-dev \ pkg-config \ git \ rsync # Copy Foxhunt codebase WORKDIR /workspace COPY . . # Pre-compile (cache dependencies) RUN cargo build --release --features cuda -p ml ENTRYPOINT ["/bin/bash"] -
Upload Parquet Files
# Local machine rsync -avz --progress test_data/*.parquet \ root@runpod-instance:/workspace/test_data/ -
Run Training via SSH
# SSH into RunPod instance ssh root@runpod-instance # Inside RunPod container cd /workspace # Train MAMBA-2 cargo run -p ml --example train_mamba2_parquet --release --features cuda -- \ --parquet-file test_data/ES_FUT_180d.parquet \ --epochs 50 # Train DQN cargo run -p ml --example train_dqn --release --features cuda -- \ --parquet-file test_data/NQ_FUT_180d.parquet \ --epochs 100 # Train PPO cargo run -p ml --example train_ppo_parquet --release --features cuda -- \ --parquet-file test_data/ZN_FUT_90d_clean.parquet \ --epochs 50 # Train TFT cargo run -p ml --example train_tft_parquet --release --features cuda -- \ --parquet-file test_data/6E_FUT_180d.parquet \ --epochs 50 -
Download Checkpoints
# Local machine rsync -avz --progress \ root@runpod-instance:/workspace/ml/trained_models/ \ ml/trained_models/ -
Terminate Instance
# RunPod web console or CLI runpod stop <instance-id>
Validation:
- ✅ All 4 models train successfully on cloud GPU
- ✅ Checkpoints download cleanly
- ✅ Cost matches estimate ($3-7 for 10-20 hours)
- ✅ Performance is 3-5x faster than local GPU
Estimated Time: 1 week (includes testing, validation)
Phase 2: Automated Cloud Training (2-3 weeks)
Goal: TLI command triggers cloud training automatically
Implementation:
-
Add RunPod Integration to Config
# config/config.toml [cloud_gpu] provider = "runpod" api_key = "${RUNPOD_API_KEY}" # From Vault gpu_type = "rtx-4090" instance_type = "community" # or "secure" auto_shutdown = true [cloud_gpu.storage] type = "s3" # or "runpod_network" bucket = "foxhunt-training-data" checkpoint_dir = "ml/trained_models" -
Create Cloud GPU Orchestrator
// services/ml_training_service/src/cloud_gpu.rs pub struct CloudGpuOrchestrator { provider: CloudGpuProvider, config: CloudGpuConfig, storage: Arc<dyn CloudStorage>, } impl CloudGpuOrchestrator { /// Spin up cloud GPU instance pub async fn provision_instance(&self) -> Result<CloudInstance> { // 1. Call RunPod API to create instance // 2. Wait for instance to be ready // 3. Return SSH connection details } /// Upload training data to cloud pub async fn upload_data(&self, parquet_files: Vec<PathBuf>) -> Result<()> { // rsync or S3 upload } /// Execute training command via SSH pub async fn run_training(&self, instance: &CloudInstance, cmd: &str) -> Result<()> { // SSH into instance, run command, stream logs } /// Download trained models pub async fn download_models(&self, instance: &CloudInstance) -> Result<()> { // rsync or S3 download } /// Terminate instance (stop billing) pub async fn terminate_instance(&self, instance: CloudInstance) -> Result<()> { // Call RunPod API to stop instance } } -
Add TLI Command
// tli/src/commands/ml.rs #[derive(Parser)] pub enum MlCommand { // ... existing commands ... /// Train ML models on cloud GPU TrainCloud { /// GPU provider (runpod, aws, lambda) #[arg(long, default_value = "runpod")] provider: String, /// GPU type (rtx-4090, a100, h100) #[arg(long, default_value = "rtx-4090")] gpu: String, /// Model to train (mamba2, dqn, ppo, tft, all) #[arg(long)] model: String, /// Parquet file or asset symbol #[arg(long)] asset: String, /// Training epochs #[arg(long, default_value = "50")] epochs: usize, /// Keep instance running after training (for debugging) #[arg(long)] keep_alive: bool, }, } -
End-to-End Workflow
# User executes TLI command tli ml train-cloud --model mamba2 --asset ES.FUT --epochs 50 # TLI → ML Training Service → Cloud GPU Orchestrator: # 1. Provision RunPod RTX 4090 instance ($0.34/hr spot) # 2. Upload ES_FUT_180d.parquet to instance # 3. Run training command via SSH # 4. Stream logs to TLI (live progress) # 5. Download checkpoint when complete # 6. Terminate instance (stop billing) # Output: # ✅ Training complete in 2.3 hours # 💾 Model saved: ml/trained_models/mamba2_ES_FUT_epoch50.safetensors # 💰 Total cost: $0.78 (2.3h × $0.34/hr)
Validation:
- ✅ TLI command successfully provisions cloud GPU
- ✅ Training executes without manual SSH intervention
- ✅ Logs stream to TLI in real-time
- ✅ Checkpoints download automatically
- ✅ Instance terminates after training (no waste)
Estimated Time: 2-3 weeks (includes testing, error handling)
Phase 3: Production Deployment (1 week)
Goal: Scheduled retraining on cloud GPU
Implementation:
-
Add Scheduler to ML Training Service
// services/ml_training_service/src/scheduler.rs pub struct TrainingScheduler { schedule: CronSchedule, // e.g., "0 2 * * 0" (2am every Sunday) orchestrator: Arc<CloudGpuOrchestrator>, } impl TrainingScheduler { pub async fn run(&self) -> Result<()> { // 1. Wait for next scheduled time // 2. Provision cloud GPU // 3. Train all 4 models // 4. Download checkpoints // 5. Update PostgreSQL metadata // 6. Terminate instance // 7. Send Slack notification } } -
Monitoring & Alerting
// Prometheus metrics training_duration_seconds{model="mamba2", gpu="rtx-4090"} training_cost_usd{model="mamba2", gpu="rtx-4090"} gpu_utilization_percent{gpu="rtx-4090"} instance_uptime_seconds{gpu="rtx-4090"} // Grafana alerts - Alert: Training failed Condition: training_duration_seconds > 86400 (24h timeout) - Alert: Instance not terminated Condition: instance_uptime_seconds > 43200 (12h max) - Alert: High cost Condition: training_cost_usd > 50 (safety threshold) -
Cost Tracking
-- PostgreSQL table for cost tracking CREATE TABLE ml_training_costs ( id SERIAL PRIMARY KEY, model_name TEXT NOT NULL, gpu_type TEXT NOT NULL, provider TEXT NOT NULL, duration_seconds INTEGER NOT NULL, cost_usd DECIMAL(10, 2) NOT NULL, checkpoint_path TEXT NOT NULL, trained_at TIMESTAMP NOT NULL DEFAULT NOW() ); -- Query monthly costs SELECT DATE_TRUNC('month', trained_at) AS month, SUM(cost_usd) AS total_cost, COUNT(*) AS training_runs FROM ml_training_costs WHERE trained_at >= NOW() - INTERVAL '1 year' GROUP BY month ORDER BY month DESC;
Validation:
- ✅ Weekly retraining runs automatically
- ✅ All 4 models update every week
- ✅ Costs track below $50/month budget
- ✅ Alerts trigger on failures or cost overruns
- ✅ Grafana dashboard shows training metrics
Estimated Time: 1 week (includes monitoring setup)
💡 Cost Optimization Tips
1. Use Spot Instances Aggressively
Savings: 40-70% vs on-demand
# RunPod: Spot instances are default (no flag needed)
# AWS: Add --spot flag to SageMaker training jobs
# GCP: Use --preemptible flag
Risk Mitigation:
- ✅ Checkpoint every 10 epochs (already implemented)
- ✅ Resume from last checkpoint on preemption
- ✅ Retry up to 3 times before falling back to on-demand
2. Auto-Shutdown After Training
Savings: 100% of idle costs
# RunPod: Enable auto-shutdown in config
auto_shutdown = true
max_idle_minutes = 10 # Shutdown after 10 min idle
# AWS: Use SageMaker Training Jobs (auto-terminate on completion)
# GCP: Use Cloud Functions to monitor instance state
Implementation:
// Cloud GPU Orchestrator
pub async fn monitor_training(&self, instance: &CloudInstance) -> Result<()> {
loop {
let status = self.get_training_status(instance).await?;
match status {
TrainingStatus::Complete => {
self.download_models(instance).await?;
self.terminate_instance(instance).await?;
break;
}
TrainingStatus::Failed => {
self.retry_training(instance).await?;
}
TrainingStatus::Running => {
tokio::time::sleep(Duration::from_secs(60)).await;
}
}
}
Ok(())
}
3. Batch Multiple Training Runs
Savings: Reduce startup overhead (5-10 min per instance)
# Instead of:
# 1. Train MAMBA-2 → terminate → restart → train DQN → ...
# Do this:
# 1. Train all 4 models sequentially in one instance session
# Example:
ssh root@runpod-instance << 'EOF'
cargo run --example train_mamba2_parquet --release --features cuda -- --epochs 50
cargo run --example train_dqn --release --features cuda -- --epochs 100
cargo run --example train_ppo_parquet --release --features cuda -- --epochs 50
cargo run --example train_tft_parquet --release --features cuda -- --epochs 50
EOF
Time Savings:
- Before: 4 models × (2 min training + 5 min startup) = 28 min
- After: 1 startup (5 min) + 4 models × 2 min = 13 min
- Savings: 15 min = 54% reduction
Cost Savings (RunPod @ $0.34/hr):
- Before: 28 min × $0.34/hr ÷ 60 = $0.16
- After: 13 min × $0.34/hr ÷ 60 = $0.07
- Savings: $0.09 per training session (56% reduction)
4. Dataset Caching on Cloud Storage
Savings: Avoid re-uploading 100MB+ Parquet files every run
# Option 1: RunPod Network Storage (persistent volume)
# - $0.10/GB/month
# - 1GB Parquet files = $0.10/month
# - Saves 5-10 min upload time per run
# Option 2: S3 (if using AWS)
# - $0.023/GB/month (standard storage)
# - 1GB Parquet files = $0.023/month
# - Fast download from S3 to EC2 (same region)
# Implementation:
# 1. Upload Parquet files once to persistent storage
# 2. Mount volume on instance startup
# 3. No re-upload needed for subsequent runs
Cost-Benefit Analysis:
- Storage Cost: $0.10/month (RunPod) or $0.023/month (S3)
- Upload Time Saved: 5-10 min per run × $0.34/hr = $0.03-$0.06 per run
- Break-Even: 2-3 runs per month
Verdict: Worth it if retraining >2x per month.
5. Reserved Instances (Long-Term)
Savings: 30-50% vs on-demand (if using cloud GPU continuously)
| Provider | GPU | On-Demand | Reserved (1-year) | Savings |
|---|---|---|---|---|
| AWS | ml.g5.xlarge | $1.41/hr | $0.85/hr | 40% |
| Lambda Labs | A100 80GB | $3.29/hr | $2.00/hr | 39% |
| RunPod | RTX 4090 | $0.59/hr | N/A (spot only) | N/A |
When to Use:
- ✅ Training >100 hours/month consistently
- ✅ Multi-tenant workloads (shared GPU across projects)
- ❌ NOT for Foxhunt (infrequent training, <20h/month)
Verdict: Stick with spot instances for now.
6. GPU Selection Strategy
Heuristic: Match GPU to model memory requirements
| Model | GPU Memory | Recommended GPU | $/hour | Notes |
|---|---|---|---|---|
| DQN | 6MB | RTX 3090 (24GB) | $0.27 | Smallest GPU works |
| PPO | 145MB | RTX 3090 (24GB) | $0.27 | Smallest GPU works |
| TFT-INT8 | 125MB | RTX 4090 (24GB) | $0.34 | Need CUDA 12.x for INT8 |
| MAMBA-2 | 164MB | RTX 4090 (24GB) | $0.34 | Best performance |
Optimization:
- Train DQN + PPO on RTX 3090 ($0.27/hr)
- Train TFT + MAMBA-2 on RTX 4090 ($0.34/hr)
- Total Savings: 20% vs using RTX 4090 for all models
Verdict: Micro-optimization, not worth the complexity. Stick with RTX 4090 for simplicity.
🔒 Security & Compliance
Data Security
Threat Model:
- ✅ Parquet files contain public market data (no PII, no secrets)
- ⚠️ Trained models are proprietary IP (protect checkpoints)
- ⚠️ API keys (RunPod, AWS) must be secured in Vault
Mitigations:
- Encrypt Parquet Files at Rest (S3 SSE or RunPod volume encryption)
- Encrypt Model Checkpoints (AES-256 before download)
- Rotate API Keys (every 90 days, stored in Vault)
- Audit Logs (track all cloud GPU provisioning events)
Compliance
Foxhunt Requirements:
- ✅ SOC2 Type II (if deploying to production)
- ✅ GDPR (not applicable - no EU customer data)
- ✅ PCI-DSS (not applicable - no payment card data)
Provider Compliance:
| Provider | SOC2 | ISO 27001 | GDPR | HIPAA | Notes |
|---|---|---|---|---|---|
| AWS | ✅ | ✅ | ✅ | ✅ | Full compliance suite |
| GCP | ✅ | ✅ | ✅ | ✅ | Full compliance suite |
| Azure | ✅ | ✅ | ✅ | ✅ | Full compliance suite |
| Lambda Labs | ⚠️ | ❌ | ⚠️ | ❌ | Limited compliance |
| RunPod | ⚠️ | ❌ | ⚠️ | ❌ | Limited compliance |
| Vast.ai | ❌ | ❌ | ❌ | ❌ | No compliance certifications |
Verdict:
- Development/Testing: RunPod is fine (public data only)
- Production (if regulated): Use AWS SageMaker for compliance
📈 Performance Validation
Expected Speedup: RTX 4090 vs RTX 3050 Ti
Hardware Comparison:
| Metric | RTX 3050 Ti (Mobile) | RTX 4090 | Improvement |
|---|---|---|---|
| CUDA Cores | 2,560 | 16,384 | 6.4x |
| Tensor Cores | 80 (3rd Gen) | 512 (4th Gen) | 6.4x |
| VRAM | 4GB GDDR6 | 24GB GDDR6X | 6.0x |
| Memory Bandwidth | 112 GB/s | 1,008 GB/s | 9.0x |
| TDP | 60W | 450W | 7.5x |
| FP16 TFLOPS | 9.0 | 82.6 | 9.2x |
| Architecture | Ampere | Ada Lovelace | 1 gen newer |
Estimated Training Time Reduction:
| Model | RTX 3050 Ti (Current) | RTX 4090 (Estimated) | Speedup |
|---|---|---|---|
| MAMBA-2 (30 epochs) | 1.9 min | 0.3-0.5 min | 3.8-6.3x |
| DQN (100 epochs) | 15 sec | 2-4 sec | 3.8-7.5x |
| PPO (30 epochs) | 7 sec | 1-2 sec | 3.5-7.0x |
| TFT-INT8 (50 epochs) | 3.5 min | 0.5-1.0 min | 3.5-7.0x |
| Total (4 models) | 5.5 min | 1.0-1.5 min | 3.7-5.5x |
Assumptions:
- Speedup limited by memory bandwidth (9x theoretical → 4-6x practical)
- Rust
candleframework scales well with CUDA cores - INT8 quantization benefits from 4th Gen Tensor Cores
🎯 ROI Analysis: Cloud GPU vs Local GPU Upgrade
Scenario 1: Buy RTX 4090 Local ($1,600)
Upfront Cost: $1,600 (GPU only, excluding PSU upgrade, cooling, installation)
Ongoing Costs:
- Electricity: ~450W × 10h/month × $0.12/kWh = $5.40/month
- Depreciation: $1,600 ÷ 24 months = $66.67/month
- Total: $72/month
Pros:
- ✅ Zero network latency
- ✅ Immediate access (no provisioning delay)
- ✅ Unlimited training time (no per-hour charges)
Cons:
- ❌ High upfront cost ($1,600)
- ❌ Obsolescence risk (GPU depreciates 50% in 2 years)
- ❌ Power/cooling requirements (450W TDP)
- ❌ Single GPU (no scaling to multi-GPU)
Scenario 2: Use RunPod Cloud GPU ($34/month)
Upfront Cost: $0 (pay-as-you-go)
Ongoing Costs:
- Training: 10h/month × $0.34/hr = $3.40/month
- Storage: 1GB Parquet × $0.10/GB/month = $0.10/month
- Total: $3.50/month
Pros:
- ✅ Zero upfront cost
- ✅ Pay only for actual usage (no idle waste)
- ✅ Scalable (upgrade to A100/H100 if needed)
- ✅ No hardware maintenance
Cons:
- ❌ Network upload/download overhead (5-10 min)
- ❌ Spot instance preemption risk (rare)
- ❌ Requires cloud GPU orchestration code
ROI Comparison (2-Year Horizon)
| Metric | Local RTX 4090 | RunPod Cloud GPU | Winner |
|---|---|---|---|
| Upfront Cost | $1,600 | $0 | ☁️ Cloud |
| Year 1 Cost | $1,600 + $65 = $1,665 | $42 | ☁️ Cloud |
| Year 2 Cost | $65/month × 12 = $780 | $42 | ☁️ Cloud |
| 2-Year Total | $2,445 | $84 | ☁️ Cloud |
| Savings | - | $2,361 | ☁️ Cloud |
| ROI | - | 2,815% | ☁️ Cloud |
| Performance | 6.4x faster than RTX 3050 Ti | 6.4x faster than RTX 3050 Ti | 🟰 Tie |
| Flexibility | Single GPU, fixed | Upgrade to A100/H100 anytime | ☁️ Cloud |
Verdict: Cloud GPU (RunPod) wins decisively for infrequent training workloads (<100h/month).
🚨 Risk Assessment
Technical Risks
| Risk | Likelihood | Impact | Mitigation |
|---|---|---|---|
| Spot instance preemption | Medium | Low | Checkpoint every 10 epochs, auto-retry |
| Network upload/download failures | Low | Medium | Use rsync with resume, retry 3x |
| RunPod API downtime | Low | High | Fallback to local GPU, multi-provider support |
| CUDA version mismatch | Low | High | Pin CUDA 12.1 in Dockerfile, test before production |
| Cost overrun | Medium | Medium | Set hard cost limits ($50/month), alerts on Grafana |
| Model checkpoint corruption | Low | High | Validate checksums, keep last 3 checkpoints |
Business Risks
| Risk | Likelihood | Impact | Mitigation |
|---|---|---|---|
| RunPod pricing increase | Medium | Low | Multi-provider strategy (add AWS as backup) |
| Vendor lock-in | Low | Medium | Abstract cloud provider interface |
| Compliance issues | Low | High | Use AWS SageMaker for regulated workloads |
| Data exfiltration | Low | Critical | Encrypt checkpoints, audit access logs |
Overall Risk Score: Low (acceptable for production deployment)
📋 Implementation Checklist
Phase 1: Manual Cloud Training (Week 1)
- Create RunPod account + add payment method
- Build Docker image with Rust + CUDA 12.1
- Push Docker image to DockerHub or RunPod registry
- Provision RTX 4090 spot instance
- Upload test Parquet files (ES_FUT_small.parquet)
- Run MAMBA-2 training via SSH
- Validate GPU utilization (>70%)
- Download checkpoint, verify integrity
- Terminate instance, confirm cost ($0.34/hr)
- Document workflow in
docs/cloud_gpu_manual_workflow.md
Phase 2: Automated Cloud Training (Weeks 2-4)
- Add
cloud_gpusection toconfig/config.toml - Implement
CloudGpuOrchestratorin ML Training Service - Add RunPod API integration (provision/terminate instances)
- Implement SSH command execution with log streaming
- Add rsync integration for data upload/download
- Create TLI command:
tli ml train-cloud - Add unit tests for cloud GPU orchestration
- Add integration test (end-to-end workflow)
- Document TLI usage in
ML_TRAINING_PARQUET_GUIDE.md
Phase 3: Production Deployment (Week 5)
- Add training scheduler (cron: every Sunday 2am)
- Implement Prometheus metrics (cost, duration, GPU utilization)
- Create Grafana dashboard for cloud GPU training
- Set up Slack/email alerts for failures
- Add PostgreSQL table for cost tracking
- Run dry-run production test (all 4 models, 50 epochs)
- Validate costs (<$50/month)
- Update CLAUDE.md with cloud GPU status
- Create production runbook (
docs/cloud_gpu_runbook.md)
🎉 Success Criteria
Technical Success
- ✅ All 4 models (MAMBA-2, DQN, PPO, TFT) train successfully on cloud GPU
- ✅ Training time 3-5x faster than local RTX 3050 Ti
- ✅ Checkpoints download cleanly and validate via checksum
- ✅ GPU utilization >70% during training (no bottlenecks)
- ✅ Zero CUDA out-of-memory errors (24GB VRAM sufficient)
Cost Success
- ✅ 4-model training session costs <$10 (20 hours max)
- ✅ Monthly retraining costs <$50/month
- ✅ Cloud GPU 2,815% ROI vs local GPU purchase ($84 vs $2,445 over 2 years)
- ✅ Spot instances used >80% of the time (vs on-demand)
Operational Success
- ✅ TLI command (
tli ml train-cloud) provisions GPU in <5 min - ✅ Training logs stream to TLI in real-time
- ✅ Auto-shutdown after training (zero idle costs)
- ✅ Grafana alerts trigger on cost overruns or failures
- ✅ Runbook created for common issues (preemption, failures, cost spikes)
📚 References
Documentation
- ML_TRAINING_PARQUET_GUIDE.md - Parquet training guide
- CLAUDE.md - System architecture
- AGENT_33_TFT_INT8_QUANTIZATION_FIX.md - INT8 optimizations
Cloud GPU Providers
- RunPod: https://www.runpod.io/pricing
- Lambda Labs: https://lambda.ai/pricing
- Vast.ai: https://vast.ai/pricing
- AWS SageMaker: https://aws.amazon.com/sagemaker/pricing/
- GCP Vertex AI: https://cloud.google.com/vertex-ai/pricing
- Azure ML: https://azure.microsoft.com/en-us/pricing/details/machine-learning/
Research Sources
- 7 Cheapest Cloud GPU Providers in 2025
- Cloud GPU Pricing Comparison 2025
- NVIDIA A100 Pricing Showdown
- AWS SageMaker Pricing Guide
🏁 Conclusion
Final Recommendation: RunPod Community Cloud (RTX 4090 Spot Instances)
Why:
- ✅ Best Price: $0.34/hr (68% cheaper than on-demand)
- ✅ Perfect Fit: 24GB VRAM handles all 4 models with 5.5x headroom
- ✅ Performance: 6.4x more CUDA cores than RTX 3050 Ti
- ✅ Flexibility: Per-second billing, 31 regions, auto-shutdown
- ✅ ROI: 2,815% return vs buying local RTX 4090 ($84 vs $2,445 over 2 years)
- ✅ Integration: Docker + SSH + rsync = minimal code changes
Next Steps:
- Immediate (This Week): Sign up for RunPod, run manual training test
- Short-Term (2-3 Weeks): Implement
tli ml train-cloudcommand - Long-Term (4-6 Weeks): Deploy automated weekly retraining
Total Implementation Time: 5 weeks (Phase 1-3) Total Cost: $3.40 - $6.80 (initial training) + $34/month (ongoing) Expected Speedup: 3-5x faster training Expected Savings: $2,361 over 2 years vs local GPU
End of AGENT-35 Report
📎 Appendix A: RunPod Quick Start Guide
# 1. Create RunPod account
https://www.runpod.io/signup
# 2. Add payment method
https://www.runpod.io/console/billing
# 3. Deploy RTX 4090 instance
# - GPU: RTX 4090 (24GB)
# - Type: Community Cloud (spot pricing)
# - Template: CUDA 12.1 + Ubuntu 22.04
# - Disk: 50GB (sufficient for Foxhunt codebase)
# 4. SSH into instance
ssh root@<instance-ip>
# 5. Install Rust
curl --proto '=https' --tlsv1.2 -sSf https://sh.rustup.rs | sh -s -- -y
source $HOME/.cargo/env
# 6. Clone Foxhunt repo
git clone https://github.com/<your-org>/foxhunt.git
cd foxhunt
# 7. Upload Parquet files
# (On local machine)
rsync -avz --progress test_data/*.parquet root@<instance-ip>:~/foxhunt/test_data/
# 8. Build with CUDA
cargo build --release --features cuda -p ml
# 9. Train models
cargo run -p ml --example train_mamba2_parquet --release --features cuda -- \
--parquet-file test_data/ES_FUT_180d.parquet --epochs 50
# 10. Download checkpoints
# (On local machine)
rsync -avz --progress root@<instance-ip>:~/foxhunt/ml/trained_models/ \
ml/trained_models/
# 11. Terminate instance
# (RunPod web console)
# Click "Stop" → Confirm → Billing stops
📎 Appendix B: Cost Calculator
#!/usr/bin/env python3
"""
Cloud GPU Cost Calculator for Foxhunt ML Training
Usage:
python cost_calculator.py --hours 10 --gpu rtx-4090 --provider runpod
"""
PRICING = {
"runpod": {
"rtx-4090": {"spot": 0.34, "on_demand": 0.59},
"rtx-3090": {"spot": 0.27, "on_demand": 0.43},
"a100-80gb": {"spot": 1.64, "on_demand": 1.99},
},
"lambda": {
"a100-80gb": {"on_demand": 3.29},
"h100-80gb": {"on_demand": 3.29},
},
"aws": {
"ml.g5.xlarge": {"spot": 0.42, "on_demand": 1.41},
"ml.p4d.24xlarge": {"spot": 9.83, "on_demand": 32.77},
},
}
def calculate_cost(hours, gpu, provider, instance_type="spot"):
price = PRICING[provider][gpu][instance_type]
total = hours * price
return {
"hours": hours,
"gpu": gpu,
"provider": provider,
"instance_type": instance_type,
"price_per_hour": price,
"total_cost": round(total, 2),
}
# Example: 10 hours on RunPod RTX 4090 (spot)
result = calculate_cost(10, "rtx-4090", "runpod", "spot")
print(f"Total cost: ${result['total_cost']} ({result['hours']}h × ${result['price_per_hour']}/hr)")
# Example: Monthly cost (10h/month for 12 months)
monthly = result['total_cost']
annual = monthly * 12
print(f"Monthly cost: ${monthly}")
print(f"Annual cost: ${annual}")
Output:
Total cost: $3.40 (10h × $0.34/hr)
Monthly cost: $3.40
Annual cost: $40.80