Root cause: C51 loss kernel launched with grid=batch_size (8192) blocks and a spin-wait inter-block barrier for manifold mixup. On H100 only ~500 blocks can be co-resident → deadlock (blocks 500-8191 wait for scheduling while blocks 0-499 spin-wait for block 8191). Fix (NVIDIA two-phase approach): - Grid capped to min(batch_size, 512) co-resident blocks - Grid-strided sample loop: each block processes multiple samples - Phase 1: all blocks write save_projected for ALL samples (no barrier) - Barrier: waits for gridDim.x (all co-resident), not batch_size - Phase 2: new grid-strided loop reads partner projections, mixes, computes CE Non-mixup path (mixup_alpha=0): CE loss computed inline, no barrier. Mixup path: two separate sample loops with single barrier between them. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
ml
10-model ML ensemble for the Foxhunt HFT system, built on Candle v0.9.1.
Models
- DQN (Rainbow) — deep Q-network with prioritized replay, dueling heads, noisy nets
- PPO — proximal policy optimization with GAE, LSTM policies, clip-higher
- TFT — temporal fusion transformer for multi-horizon forecasting
- Mamba2 — state space model for sequence prediction
- Liquid Networks — biologically inspired networks for non-stationary data
- TLOB — transformer-based limit order book analysis
- KAN — Kolmogorov-Arnold networks
- xLSTM — extended LSTM architecture
- TGGN — temporal graph neural network
- Diffusion — diffusion-based generative model
Key Modules
ensemble— model ensemble coordination and confidence aggregationhyperopt— PSO-based hyperparameter optimization with per-model adapterstrainers— unified training loops (DQN, PPO, supervised)inference—InferenceAdaptertrait for predictioncheckpoint— model checkpointing and restorationevaluation— walk-forward evaluation pipeline
Usage
use ml::dqn::DQN;
use ml::ppo::PpoTrainer;