Add warp-cooperative CUDA kernel for DQN experience collection that exploits H100's improved warp shuffle throughput. On sm_90+, 32 threads (1 warp) cooperate on each dot product via __shfl_xor_sync butterfly reduction, cutting per-thread register pressure from ~5.5KB to ~200B and enabling full occupancy on 132 SMs. Key additions: - 6 warp-cooperative device functions in common_device_functions.cuh: distributed/broadcast matvec (clean + NoisyNet), warp RMSNorm, warp_reduce_sum_all - 4 TILE_LAYER_WARP_* macros using __syncwarp() for warp-level sync - 3 warp forward passes: standard dueling, NoisyNet, C51 distributional (distributed heavy layers + broadcast atom layers for softmax) - Full warp kernel: dqn_full_experience_kernel_warp with lane-0 simulation, strided state scatter, warp-shuffle curiosity gather - Host-side compute capability detection (sm_major >= 9) with automatic fallback to standard 256-thread kernel on older GPUs - cuCtxSetLimit(STACK_SIZE, 16KB) for standard kernel safety 874/874 tests pass, 0 clippy warnings. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
ml
10-model ML ensemble for the Foxhunt HFT system, built on Candle v0.9.1.
Models
- DQN (Rainbow) — deep Q-network with prioritized replay, dueling heads, noisy nets
- PPO — proximal policy optimization with GAE, LSTM policies, clip-higher
- TFT — temporal fusion transformer for multi-horizon forecasting
- Mamba2 — state space model for sequence prediction
- Liquid Networks — biologically inspired networks for non-stationary data
- TLOB — transformer-based limit order book analysis
- KAN — Kolmogorov-Arnold networks
- xLSTM — extended LSTM architecture
- TGGN — temporal graph neural network
- Diffusion — diffusion-based generative model
Key Modules
ensemble— model ensemble coordination and confidence aggregationhyperopt— PSO-based hyperparameter optimization with per-model adapterstrainers— unified training loops (DQN, PPO, supervised)inference—InferenceAdaptertrait for predictioncheckpoint— model checkpointing and restorationevaluation— walk-forward evaluation pipeline
Usage
use ml::dqn::DQN;
use ml::ppo::PpoTrainer;