Replace the evaluator's independent CublasForward with QValueProvider trait that routes Q-value computation through the trainer's graphed CublasForward. This eliminates the separate cuBLAS handle that caused evaluation non-determinism. Architecture: - QValueProvider trait in q_value_provider.rs (compute_q_values_to) - FusedTrainingCtx implements it (chunks through trainer batch_size=64) - GpuDqnTrainer::compute_q_values_graphed captures eval forward in CUDA Graph (cuBLAS forward + compute_expected_q) on first call - Pre-captured at deterministic point (right after mega graph) - evaluate_dqn_graphed takes &mut dyn QValueProvider (mandatory) Cleanup (-398 lines): - Deleted evaluator's internal CublasForward + all chunked scratch buffers - Deleted compute_q_values (non-graphed), flatten_weights_for_cublas, ensure_cublas_ready, compute_backtest_param_sizes - Deleted evaluate_dqn (fallback path) - Removed dead imports and stale comments Training: fully bit-identical across runs (CUDA Graph replay). Evaluation: deterministic within process (graph replay), minor variation across processes (cublasLtMatmul capture non-determinism). Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
ml
10-model ML ensemble for the Foxhunt HFT system, built on Candle v0.9.1.
Models
- DQN (Rainbow) — deep Q-network with prioritized replay, dueling heads, noisy nets
- PPO — proximal policy optimization with GAE, LSTM policies, clip-higher
- TFT — temporal fusion transformer for multi-horizon forecasting
- Mamba2 — state space model for sequence prediction
- Liquid Networks — biologically inspired networks for non-stationary data
- TLOB — transformer-based limit order book analysis
- KAN — Kolmogorov-Arnold networks
- xLSTM — extended LSTM architecture
- TGGN — temporal graph neural network
- Diffusion — diffusion-based generative model
Key Modules
ensemble— model ensemble coordination and confidence aggregationhyperopt— PSO-based hyperparameter optimization with per-model adapterstrainers— unified training loops (DQN, PPO, supervised)inference—InferenceAdaptertrait for predictioncheckpoint— model checkpointing and restorationevaluation— walk-forward evaluation pipeline
Usage
use ml::dqn::DQN;
use ml::ppo::PpoTrainer;