Three bugs causing the baseline RL training to hang on first epoch while smoketest completes in 1.2s: 1. VRAM oversubscription (hang root cause): detect_gpu_hardware auto-scaling computed n_episodes without accounting for 2x counterfactual doubling or dtod_clone allocations, causing 4x actual memory vs budget. Replaced with configurable gpu_n_episodes field (smoketest=32, localdev=128, prod=4096). 2. Counterfactual experiences silently dropped: build_next_states_f32 received n_episodes instead of n_episodes*2, and PER insert used base count instead of doubled count — ~50% of augmented training data was generated then lost. 3. CudaEvent leak in hot loop: record_event(None) created+destroyed 8000 events per epoch in forward_online_raw/f32. Pre-allocated 4 events in CublasForward struct, eliminating driver overhead and handle leaks during CUDA Graph capture. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
ml
10-model ML ensemble for the Foxhunt HFT system, built on Candle v0.9.1.
Models
- DQN (Rainbow) — deep Q-network with prioritized replay, dueling heads, noisy nets
- PPO — proximal policy optimization with GAE, LSTM policies, clip-higher
- TFT — temporal fusion transformer for multi-horizon forecasting
- Mamba2 — state space model for sequence prediction
- Liquid Networks — biologically inspired networks for non-stationary data
- TLOB — transformer-based limit order book analysis
- KAN — Kolmogorov-Arnold networks
- xLSTM — extended LSTM architecture
- TGGN — temporal graph neural network
- Diffusion — diffusion-based generative model
Key Modules
ensemble— model ensemble coordination and confidence aggregationhyperopt— PSO-based hyperparameter optimization with per-model adapterstrainers— unified training loops (DQN, PPO, supervised)inference—InferenceAdaptertrait for predictioncheckpoint— model checkpointing and restorationevaluation— walk-forward evaluation pipeline
Usage
use ml::dqn::DQN;
use ml::ppo::PpoTrainer;