Adds evaluate_baseline CLI args: --surrogate-mode=off|random --surrogate-seed=<u64> --surrogate-marginals=<path.json> --emit-action-marginals --emit-pooled-sharpe Random-action surrogate samples actions from marginal distribution (or uniform fallback) using a seeded RNG, bypassing the model entirely. Surrogate mode forces the CPU DQN path so action selection can be overridden (GPU path would need invasive kernel changes and defeats the bypass-the-model sanity check). Flat action-index counts are tracked in the CPU DQN path only; the ACTION_MARGINALS: line emits those as a JSON distribution, or emits an honest stub marker when only the GPU path was exercised. Pooled Sharpe is computed from concatenated per-fold returns (CPU path) or falls back to mean-of-fold-Sharpes with a warning. Smoke test runs 30 surrogate seeds + 1 trained run via evaluate_baseline subprocess, asserts trained pooled Sharpe exceeds the 95th percentile of surrogate Sharpes. Test is #[ignore]'d and requires a trained checkpoint at /workspace/output/dqn_fold0_best.safetensors (Phase 3 deliverable); FOXHUNT_SURROGATE_CKPT env var overrides for local testing. Uses the compiled target/release/examples/evaluate_baseline if present, otherwise falls back to cargo run. Will pass once Phase 3 produces a checkpoint.
ml
10-model ML ensemble for the Foxhunt HFT system, built on Candle v0.9.1.
Models
- DQN (Rainbow) — deep Q-network with prioritized replay, dueling heads, noisy nets
- PPO — proximal policy optimization with GAE, LSTM policies, clip-higher
- TFT — temporal fusion transformer for multi-horizon forecasting
- Mamba2 — state space model for sequence prediction
- Liquid Networks — biologically inspired networks for non-stationary data
- TLOB — transformer-based limit order book analysis
- KAN — Kolmogorov-Arnold networks
- xLSTM — extended LSTM architecture
- TGGN — temporal graph neural network
- Diffusion — diffusion-based generative model
Key Modules
ensemble— model ensemble coordination and confidence aggregationhyperopt— PSO-based hyperparameter optimization with per-model adapterstrainers— unified training loops (DQN, PPO, supervised)inference—InferenceAdaptertrait for predictioncheckpoint— model checkpointing and restorationevaluation— walk-forward evaluation pipeline
Usage
use ml::dqn::DQN;
use ml::ppo::PpoTrainer;