- gradient_utils: Add TensorId-based fallback when Var identity mismatch causes 0/N vars to match GradStore (Candle Adam clones Arcs). Fallback computes norm AND clips via insert_id. Throttled warning (1st + every 1000th). 7 unit tests including mismatch-actually-clips. - monitoring: Track full 45-action factored space (5 exposure × 3 order × 3 urgency). Fix validate_rewards false alarm on GPU path where single aggregated mean_reward per epoch gives N=1 → std=0. - trainer: GPU experience collection routes exposure actions through route_action() for factored tracking instead of exposure-only counts. Applied in both per-step and epoch-summary paths. - train_baseline_rl: Auto-detect VRAM <8GB → disable GPU replay buffer to prevent OOM on RTX 3050 Ti class GPUs. - smoke_test_real_data: E2E DQN training test with 6 assertions (epoch completion, loss decrease, finite losses, Q-value divergence, 45-action space, finite gradient norms). Validated: 1642 tests pass (ml=915, ml-core=311, ml-dqn=416), 0 clippy warnings, baseline RL trains 10 epochs on CUDA with Sharpe +5.45. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
ml
10-model ML ensemble for the Foxhunt HFT system, built on Candle v0.9.1.
Models
- DQN (Rainbow) — deep Q-network with prioritized replay, dueling heads, noisy nets
- PPO — proximal policy optimization with GAE, LSTM policies, clip-higher
- TFT — temporal fusion transformer for multi-horizon forecasting
- Mamba2 — state space model for sequence prediction
- Liquid Networks — biologically inspired networks for non-stationary data
- TLOB — transformer-based limit order book analysis
- KAN — Kolmogorov-Arnold networks
- xLSTM — extended LSTM architecture
- TGGN — temporal graph neural network
- Diffusion — diffusion-based generative model
Key Modules
ensemble— model ensemble coordination and confidence aggregationhyperopt— PSO-based hyperparameter optimization with per-model adapterstrainers— unified training loops (DQN, PPO, supervised)inference—InferenceAdaptertrait for predictioncheckpoint— model checkpointing and restorationevaluation— walk-forward evaluation pipeline
Usage
use ml::dqn::DQN;
use ml::ppo::PpoTrainer;