Existing nan_flags_buf [16] covers 13 buffers with per-step NaN checks inside the captured training graph. But read_nan_flags() was only called when training guard's halt_nan fired — which checks pinned- readback grad_norm. NaN-clamped-to-zero gradients reach the pinned scalar as 0, not NaN, so halt_nan never fires for the explosion case. The collapse path (halt_grad_collapse, grad < 0.01) was firing instead without reading the flags. Add flag readback in that path when gr.raw_grad_norm < 1e-6 (suspicious zero) — logs NaN-CLAMPED-TO-ZERO with flagged buffer names. Genuine near-zero gradients get a separate "no NaN flags set" log so we can distinguish the two cases. This should pinpoint which kernel produces the F1 ep2 NaN first (C51 KL projection? IQN aux? CQL? aux heads?). Diagnostic only — keep after fix lands; correct gate for future regressions.
ml
10-model ML ensemble for the Foxhunt HFT system, built on Candle v0.9.1.
Models
- DQN (Rainbow) — deep Q-network with prioritized replay, dueling heads, noisy nets
- PPO — proximal policy optimization with GAE, LSTM policies, clip-higher
- TFT — temporal fusion transformer for multi-horizon forecasting
- Mamba2 — state space model for sequence prediction
- Liquid Networks — biologically inspired networks for non-stationary data
- TLOB — transformer-based limit order book analysis
- KAN — Kolmogorov-Arnold networks
- xLSTM — extended LSTM architecture
- TGGN — temporal graph neural network
- Diffusion — diffusion-based generative model
Key Modules
ensemble— model ensemble coordination and confidence aggregationhyperopt— PSO-based hyperparameter optimization with per-model adapterstrainers— unified training loops (DQN, PPO, supervised)inference—InferenceAdaptertrait for predictioncheckpoint— model checkpointing and restorationevaluation— walk-forward evaluation pipeline
Usage
use ml::dqn::DQN;
use ml::ppo::PpoTrainer;