Files
foxhunt/ml/Cargo.toml
jgrusewski 35feadf55e 🚀 Wave 160 Phase 6: CUDA Mandatory + TDD Testing + TFT Complete (21 Agents)
## Major Achievements

### 1. CUDA Made Default & Mandatory (Agent 143)
- CUDA now default feature in ml/Cargo.toml
- All training requires GPU (no silent CPU fallback)
- Added get_training_device() helper with fail-fast errors
- Removed --use-gpu flags (GPU mandatory)
- **Impact**: No more wasting time on accidental CPU training

### 2. TFT Training COMPLETE (Agent 144)
-  Training completed successfully in 7.6 minutes
-  Early stopping at epoch 100/200 (best val loss: 0.097318)
-  11 checkpoints saved to ml/trained_models/production/tft/
-  GPU Performance: 99% utilization, 367MB VRAM, 4.4s/epoch
-  10x speedup vs CPU (4.4s vs 43-55s per epoch)
- **Status**: PRODUCTION READY

### 3. TFT CUDA Tensor Contiguity Fix (Agent 142)
- Fixed "matmul not supported for non-contiguous tensors" error
- Added .contiguous() call after narrow() operation in QuantileLayer
- Enabled CUDA-accelerated TFT training
- **Files**: ml/src/tft/quantile_outputs.rs

### 4. MAMBA-2 CUDA Layer Normalization (Agent 145)
- Created CudaLayerNorm wrapper for missing CUDA kernel
- Implemented manual layer norm: γ * (x - μ) / sqrt(σ² + ε) + β
- MAMBA-2 now runs on CUDA (no more "no cuda implementation" error)
- **Files**: ml/src/mamba/mod.rs

### 5. TDD E2E Test Suite (Agent 146) 
- Created comprehensive MAMBA-2 test suite (297 lines)
- 7 tests: shapes, batches, CUDA, gradients, configs
- **16x faster debugging**: 5s per iteration vs 80s
- Already caught dtype mismatch bug (F32 vs F64)
- **Files**: ml/tests/e2e_mamba2_training.rs

## Agent Summary (Agents 126-146)

### Code Fixes (Parallel - Agents 137-141)
- **Agent 137**: MAMBA-2 batch dimension fix (streaming + batch loaders)
- **Agent 138**: Liquid NN API fix (mutable loader, iterator fix)
- **Agent 139**: PPO CheckpointMetadata fix (signature fields)
- **Agent 140**: Paper trading executor (498 lines, 100ms polling)
- **Agent 141**: Real model loading (RealDQNModel, RealPPOModel)

### Infrastructure (Agents 143-146)
- **Agent 143**: CUDA mandatory (Cargo.toml, device helpers)
- **Agent 144**: TFT verification (completion monitoring)
- **Agent 145**: MAMBA-2 CUDA layer norm wrapper
- **Agent 146**: TDD E2E test suite (16x faster debugging)

## Files Modified

### Core ML Infrastructure
- ml/Cargo.toml: Added default = ["minimal-inference", "cuda"]
- ml/src/lib.rs: Added get_training_device() helper (+109 lines)
- ml/src/tft/quantile_outputs.rs: Fixed tensor contiguity
- ml/src/mamba/mod.rs: Added CudaLayerNorm wrapper (+41 lines)

### Training Scripts
- ml/examples/train_tft_dbn.rs: Removed --use-gpu flag
- ml/examples/train_ppo.rs: Removed --use-gpu flag
- ml/examples/train_mamba2_dbn.rs: Forced CUDA-only mode
- ml/examples/train_liquid_dbn.rs: Fixed API usage

### Data Loaders
- ml/src/data_loaders/dbn_sequence_loader.rs: Fixed batch dimensions
- ml/src/data_loaders/streaming_dbn_loader.rs: Fixed batch dimensions

### Trading Service
- services/trading_service/src/paper_trading_executor.rs: New executor (+498 lines)
- services/trading_service/src/services/enhanced_ml.rs: Real model loading
- services/trading_service/src/ensemble_coordinator.rs: Integration

### Tests
- ml/tests/e2e_mamba2_training.rs: New TDD test suite (+297 lines)

### Trainers
- ml/src/trainers/tft.rs: Fixed CheckpointMetadata signature fields

## Performance Metrics

### TFT Training
- Duration: 7.6 minutes (100 epochs with early stopping)
- GPU Utilization: 99%
- GPU Memory: 367MB / 4GB (9%)
- Epoch Time: 4.4 seconds (vs 43-55s on CPU)
- Speedup: 10x vs CPU
- Status:  PRODUCTION READY

### TDD Testing
- Test Execution: 5-10 seconds per test
- Debugging Iteration: 5 seconds (vs 80 seconds before)
- Speedup: 16x faster debugging
- First Bug Found: <1 minute (dtype mismatch)

## Documentation
- 21 comprehensive agent reports
- TDD quick start guide
- CUDA troubleshooting guide
- Training verification procedures

## Next Steps
1. Fix MAMBA-2 dtype mismatch (F32→F64) - 2 minutes
2. Run MAMBA-2 tests until passing - 5-10 minutes
3. Launch full MAMBA-2 training - 200 epochs
4. Launch Liquid NN training

## System Status
- TFT:  COMPLETE (production ready)
- MAMBA-2: 🧪 IN TESTING (TDD suite ready)
- CUDA:  DEFAULT (mandatory for training)
- Tests:  16x faster debugging

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
2025-10-14 23:13:34 +02:00

184 lines
6.0 KiB
TOML

[package]
name = "ml"
version.workspace = true
edition.workspace = true
rust-version.workspace = true
authors.workspace = true
license.workspace = true
repository.workspace = true
homepage.workspace = true
documentation.workspace = true
publish.workspace = true
keywords.workspace = true
categories.workspace = true
[features]
# MINIMAL features for HFT inference only - ALL HEAVY ML REMOVED
# CUDA is now default for training - GPU acceleration mandatory
default = ["minimal-inference", "cuda"]
# PRODUCTION FEATURES - LIGHTWEIGHT ONLY
minimal-inference = [] # Minimal inference with no optional deps
financial = [] # Basic financial calculations
high-precision = ["rust_decimal/serde-float"]
# PERFORMANCE FEATURES - NO HEAVY ML
simd = [] # SIMD without heavy dependencies
# Storage and memory management features
gc = [] # Garbage collection features
s3-storage = ["aws-config", "aws-sdk-s3", "aws-types", "aws-credential-types", "urlencoding"] # S3 storage backend with AWS SDK
cuda = ["candle-core/cuda", "candle-core/cudnn"] # CUDA support - OPTIONAL for CI/Docker
# ALL HEAVY ML FEATURES REMOVED:
# gpu, pytorch, linfa-ml - MOVED TO ml_training_service
# optimization, graph-models, reinforcement-learning - MOVED TO ml_training_service
# transformers-advanced - MOVED TO ml_training_service
[dependencies]
# Core async and utilities
tokio.workspace = true
futures.workspace = true
async-trait.workspace = true
clap.workspace = true # CLI argument parsing for train_tft binary
# Serialization and error handling
serde.workspace = true
serde_json.workspace = true
uuid.workspace = true
thiserror.workspace = true
anyhow.workspace = true
chrono.workspace = true
rand.workspace = true
# System and I/O
memmap2.workspace = true
tempfile.workspace = true
tracing.workspace = true
tracing-subscriber.workspace = true # For train_tft binary logging
prometheus.workspace = true
reqwest.workspace = true
# Internal workspace crates
trading_engine.workspace = true
config.workspace = true
common.workspace = true
risk = { path = "../risk" }
# Model loading functionality is in storage crate
storage = { path = "../storage" }
# Data crate for test helpers (dev-dependency in tests)
data = { path = "../data" }
# Database for model registry
sqlx.workspace = true
# Essential ML frameworks for HFT inference - CUDA OPTIONAL
# Using specific git rev (671de1db) for cudarc 0.17.3 CUDA 13.0 compatibility
# Rev 671de1db is v0.9.1 + cudarc 0.17.3 upgrade
# CUDA features are optional - controlled by 'cuda' feature flag
candle-core = { git = "https://github.com/huggingface/candle", rev = "671de1db" } # Base without GPU
candle-nn = { git = "https://github.com/huggingface/candle", rev = "671de1db" }
# Use git version of candle-optimisers to match candle version
candle-optimisers = { git = "https://github.com/KGrewal1/optimisers" } # Base without GPU
# HEAVY ML FRAMEWORKS REMOVED - MOVED TO ml_training_service
# ort (ONNX Runtime) - REMOVED (1000+ dependencies alone!)
# tch, torch-sys (PyTorch bindings) - REMOVED (500+ dependencies!)
# Mathematical libraries
# BLAS feature temporarily disabled - requires libopenblas-dev installation
# TODO: Re-enable after running: sudo apt-get install -y libopenblas-dev
ndarray = { version = "0.15", features = ["rayon", "serde"] }
nalgebra = { version = "0.33", features = ["serde-serialize"] }
arrayfire = { version = "3.8", optional = true }
# MINIMAL statistics only - ALL HEAVY ML ALGORITHMS REMOVED
# linfa ecosystem (linfa, linfa-clustering, linfa-linear, linfa-reduction) - REMOVED (200+ deps)
# smartcore - REMOVED (100+ dependencies)
# Basic statistics - always included (not optional)
statrs.workspace = true # Required for statistical computations
rust_decimal.workspace = true
# gymnasium, rerun - REMOVED (RL frameworks moved to ml_training_service)
# cudarc, wgpu - REMOVED (GPU frameworks moved to ml_training_service)
rayon.workspace = true
crossbeam = { version = "0.8", features = ["std"] }
petgraph = { version = "0.6", features = ["serde"] } # Required for TGNN graphs
semver = "1.0"
lru.workspace = true # Required for model caching
# chronoutil, ta, polars - REMOVED or moved to workspace dependencies
# argmin, nlopt, ipopt - REMOVED (optimization frameworks moved to ml_training_service)
half = { version = "2.6.0", features = ["serde"] }
rand_distr.workspace = true
dbn.workspace = true # Databento Binary format for real market data loading
databento = "0.34" # Databento API client for downloading data (includes async by default)
dotenv = "0.15" # Load .env files for API keys
structopt = "0.3" # CLI argument parsing for examples
parking_lot = { version = "0.12", features = ["hardware-lock-elision"] }
dashmap = { version = "6.1", features = ["serde"] }
once_cell = "1.19"
lazy_static.workspace = true
flate2 = "1.0"
sha2 = "0.10"
hmac = "0.12" # HMAC for checkpoint signatures (SEC-001 fix)
hex = "0.4" # Hex encoding for signatures
bincode = "1.3"
fastrand = "2.1"
# wide - REMOVED (SIMD moved to trading_engine)
num-traits = "0.2"
num = "0.4"
libc = "0.2"
fs2 = "0.4"
num_cpus = "1.16"
approx.workspace = true
sysinfo = "0.33" # System information for benchmarks
# AWS SDK dependencies for S3 checkpoint storage (optional, s3-storage feature)
aws-config = { version = "1.1", optional = true }
aws-sdk-s3 = { version = "1.14", optional = true }
aws-types = { version = "1.1", optional = true }
aws-credential-types = { version = "1.1", optional = true }
urlencoding = { version = "2.1", optional = true }
[dev-dependencies]
tokio-test = "0.4"
proptest = "1.5"
tempfile = "3.12"
futures-test = "0.3"
mockall = "0.13"
test-case = "3.0"
rstest = "0.22"
criterion = { version = "0.5", features = ["html_reports", "async_tokio"] }
tokio = { workspace = true, features = ["test-util", "macros"] }
insta = "1.34" # Snapshot testing for ML outputs
serial_test = "3.0" # Sequential testing for GPU resources
tracing-subscriber = { version = "0.3", features = ["env-filter", "fmt"] }
[[example]]
name = "cuda_test"
path = "examples/cuda_test.rs"
[[example]]
name = "gpu_training_benchmark"
path = "examples/gpu_training_benchmark.rs"
[lints]
workspace = true