Phase 2: Replace 1-warp/sample fused kernels with cuBLAS SGEMM batched forward/backward. - batched_forward.rs: cuBLAS SGEMM forward (10 GEMM + bias/ReLU per pass) - batched_backward.rs: cuBLAS SGEMM backward (chain rule via GEMM, no atomicAdd) - c51_loss_kernel.cu: standalone C51 distributional loss (256 threads, 2KB shmem) - c51_grad_kernel: dL/d_logits with dueling routing for cuBLAS backward - BF16 alignment fix: pad offsets to even for short2 vectorized loads - Training step: 10.7ms → 0.7ms (15x) on RTX 3050 Phase 3: Unified cuBLAS Q-forward + dead code elimination (-4,400 lines net). - Rewrite experience collector: timestep loop + cuBLAS replaces monolithic 3,272-line kernel - Delete dqn_training_kernel.cu (1,385 lines) — replaced by dqn_utility_kernels.cu (118 lines) - Delete dqn_experience_kernel.cu (3,272 lines) — replaced by experience_kernels.cu (656 lines) - Remove BF16 warp-matvec helpers from common_device_functions.cuh (-159 lines) - Remove dead methods/fields from GpuDqnTrainer (-500 lines) - Experience collection: 348ms → 12ms (29x) on RTX 3050 - No fallback paths — cuBLAS is the only Q-forward implementation - All 1,514 tests pass, GPU smoke test verified with real data Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
ml
10-model ML ensemble for the Foxhunt HFT system, built on Candle v0.9.1.
Models
- DQN (Rainbow) — deep Q-network with prioritized replay, dueling heads, noisy nets
- PPO — proximal policy optimization with GAE, LSTM policies, clip-higher
- TFT — temporal fusion transformer for multi-horizon forecasting
- Mamba2 — state space model for sequence prediction
- Liquid Networks — biologically inspired networks for non-stationary data
- TLOB — transformer-based limit order book analysis
- KAN — Kolmogorov-Arnold networks
- xLSTM — extended LSTM architecture
- TGGN — temporal graph neural network
- Diffusion — diffusion-based generative model
Key Modules
ensemble— model ensemble coordination and confidence aggregationhyperopt— PSO-based hyperparameter optimization with per-model adapterstrainers— unified training loops (DQN, PPO, supervised)inference—InferenceAdaptertrait for predictioncheckpoint— model checkpointing and restorationevaluation— walk-forward evaluation pipeline
Usage
use ml::dqn::DQN;
use ml::ppo::PpoTrainer;