Root cause: each trial created a fresh MlDevice (new CudaStream) but the primary CudaContext was shared. cudaFreeAsync returns memory to the allocating STREAM's pool, not the context's global heap. New trial's new stream couldn't reclaim the old stream's freed blocks. Evidence: Trial 0 leaked 2063MB, Trial 2 leaked 3212MB. Available RAM dropped 10.4GB→5.2GB over 3 trials. Later trials ran on a memory-starved system, explaining systematic performance degradation. Fix: fork a new stream from the shared device instead of creating a fresh MlDevice. Forked streams share the same context and async memory pool — free_async blocks are immediately available to the next trial's allocations. Also fixed: TRIAL_SUMMARY best_epoch now shows actual best, not epochs_completed. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
ml
10-model ML ensemble for the Foxhunt HFT system, built on Candle v0.9.1.
Models
- DQN (Rainbow) — deep Q-network with prioritized replay, dueling heads, noisy nets
- PPO — proximal policy optimization with GAE, LSTM policies, clip-higher
- TFT — temporal fusion transformer for multi-horizon forecasting
- Mamba2 — state space model for sequence prediction
- Liquid Networks — biologically inspired networks for non-stationary data
- TLOB — transformer-based limit order book analysis
- KAN — Kolmogorov-Arnold networks
- xLSTM — extended LSTM architecture
- TGGN — temporal graph neural network
- Diffusion — diffusion-based generative model
Key Modules
ensemble— model ensemble coordination and confidence aggregationhyperopt— PSO-based hyperparameter optimization with per-model adapterstrainers— unified training loops (DQN, PPO, supervised)inference—InferenceAdaptertrait for predictioncheckpoint— model checkpointing and restorationevaluation— walk-forward evaluation pipeline
Usage
use ml::dqn::DQN;
use ml::ppo::PpoTrainer;