Mixed-precision loss/grad kernels: - MSE + C51 loss: float softmax/projection/TD-error (prevents bf16 exp overflow) - MSE + C51 grad: float arithmetic + bf16 range clamp before atomicAdd - Shared memory: float (4 bytes/elem) for numerically stable reductions - Bias kernels: float add+clamp ±500 (prevents bf16 Inf cascade between layers) - Noisy bias kernel: same float clamping fast_isnan/fast_isinf (ROOT CAUSE FIX): - nvcc --use_fast_math implies --no-nans → isnan()/isinf() compiled to false - ALL NaN guards across ALL kernels were dead code - Added bit-pattern IEEE 754 checks to common_device_functions.cuh - Replaced isnan/isinf in 7 kernel files (21 occurrences) - ml-dqn build.rs: all kernels now get common header (no more standalone) f32 PER IS-weights: - GpuBatchSlices.weights: CudaSlice<u16> → CudaSlice<f32> - GpuBatch.weights: GpuTensor → CudaSlice<f32> - Loss/grad kernel signatures: const __nv_bfloat16* → const float* - Upload path: separate f32 memcpy instead of bf16 staging - Eliminates bf16 overflow in IS-weight storage CUTLASS padding: - pad32() helper: round up to next multiple of 32 - 6 value-logit buffers: pad32(num_atoms) (51 → 64) - 6 branch-logit buffers: +32*3 padding per branch 895/895 unit tests, 8/9 smoke tests pass. 50-epoch convergence: NaN at step ~100-200 — backward pass produces NaN gradients within the CUDA graph replay (same atomic execution as Adam). Root cause: bf16 backward GemmEx inputs can overflow. Needs mixed-precision backward pass (same pattern as loss kernels) or f32 gradient output buffers. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
ml
10-model ML ensemble for the Foxhunt HFT system, built on Candle v0.9.1.
Models
- DQN (Rainbow) — deep Q-network with prioritized replay, dueling heads, noisy nets
- PPO — proximal policy optimization with GAE, LSTM policies, clip-higher
- TFT — temporal fusion transformer for multi-horizon forecasting
- Mamba2 — state space model for sequence prediction
- Liquid Networks — biologically inspired networks for non-stationary data
- TLOB — transformer-based limit order book analysis
- KAN — Kolmogorov-Arnold networks
- xLSTM — extended LSTM architecture
- TGGN — temporal graph neural network
- Diffusion — diffusion-based generative model
Key Modules
ensemble— model ensemble coordination and confidence aggregationhyperopt— PSO-based hyperparameter optimization with per-model adapterstrainers— unified training loops (DQN, PPO, supervised)inference—InferenceAdaptertrait for predictioncheckpoint— model checkpointing and restorationevaluation— walk-forward evaluation pipeline
Usage
use ml::dqn::DQN;
use ml::ppo::PpoTrainer;