Replaces half = max(10*q_gap, 3*std_ema, floor) tuned constants with quantile-based: half = max(|q_p95 - v_center|, |v_center - q_p5|) clamped to [min_half_floor, abs_half]. Covers observed Q distribution directly. New GPU kernel q_quantile_reduce reads per-sample Q-values from q_out_buf, sorts per branch (bitonic for power-of-2 N, quickselect otherwise), writes ISV[Q_P05_*=47..51) and ISV[Q_P95_*=51..55). Cold-path per-epoch (4 blocks, 1 thread each). No atomicAdd. ISV slots 47-54 added at tail (8 total). Fingerprint shifted to 55-56. ISV_TOTAL_DIM grows 49 -> 57. Fingerprint seed updated; recomputes hash at compile time (new value: 0xbf6c400c026d77e3). Launch order: reduce_current_q_stats (populates q_out_buf) -> q_quantile_reduce (writes ISV P5/P95) -> reduce_current_q_stats_per_branch -> update_eval_v_range (reads ISV P5/P95 for half-width). Bootstrap: q_p05 = v_min, q_p95 = v_max (matches current atom range at cold start). FoldReset: isv_q_quantiles dispatch arm resets to bootstrap values at fold boundary. StateResetRegistry: isv_q_quantiles as FoldReset; dispatch arm added in training_loop.rs::reset_named_state. Docs: isv-slots.md rows [47..57) updated, fingerprint tail reference corrected to [55..57). dqn-wire-up-audit.md: q_quantile_kernel.cu added. Plan 2 Task 1. Spec §4.C.1. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
ml
10-model ML ensemble for the Foxhunt HFT system, built on Candle v0.9.1.
Models
- DQN (Rainbow) — deep Q-network with prioritized replay, dueling heads, noisy nets
- PPO — proximal policy optimization with GAE, LSTM policies, clip-higher
- TFT — temporal fusion transformer for multi-horizon forecasting
- Mamba2 — state space model for sequence prediction
- Liquid Networks — biologically inspired networks for non-stationary data
- TLOB — transformer-based limit order book analysis
- KAN — Kolmogorov-Arnold networks
- xLSTM — extended LSTM architecture
- TGGN — temporal graph neural network
- Diffusion — diffusion-based generative model
Key Modules
ensemble— model ensemble coordination and confidence aggregationhyperopt— PSO-based hyperparameter optimization with per-model adapterstrainers— unified training loops (DQN, PPO, supervised)inference—InferenceAdaptertrait for predictioncheckpoint— model checkpointing and restorationevaluation— walk-forward evaluation pipeline
Usage
use ml::dqn::DQN;
use ml::ppo::PpoTrainer;