Files
foxhunt/crates/ml
jgrusewski 4189da563f feat(dqn-v2): C.6 migrate 4 more controllers to AdaptiveController (Tasks 13-16)
Tasks 13-16 of Plan 1 §4.C.6 — Batch B controller migrations:

- Task 13 (TauController): cosine-annealed Polyak EMA coefficient with
  health-coupled floor.  Mirrors compute_cosine_annealed_tau +
  apply_health_coupled_tau_floor from fused_training.rs/gpu_dqn_trainer.rs.
  Wired at epoch-boundary B2/G3 block in training_loop.rs; per-step
  actuators (set_tau_host, target_ema_update) remain in fused_training.rs.
  8 unit tests.

- Task 14 (EpsilonController): ISV-adaptive exploration epsilon.
  Replaces inline base_floor*(0.5+volatility) block (~10 lines) in
  initialize_epoch_state.  Volatility from ISV[2]; agent.set_epsilon()
  actuator call preserved.  8 unit tests.

- Task 15 (ConvictionFloorController): IQL branch_scales per-sample floor
  scaffolding.  Static 0.1 (SchemaContract slot ISV[36]).  write_output
  writes to ISV[36].  Never fires.  Plan 3 adds ramp logic.  7 unit tests.

- Task 16 (PlanThresholdController): plan activation threshold scaffolding.
  Static 0.5, mirrors the > 0.5f literal in experience_kernels.cu and
  backtest_plan_kernel.cu.  Never fires.  Plan 3 §B.4 wires ISV slot.
  6 unit tests.

All 29 new tests pass.  Cargo check: 8 warnings (unchanged from baseline).
Audit doc updated per Invariant 7.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-24 15:03:19 +02:00
..

ml

10-model ML ensemble for the Foxhunt HFT system, built on Candle v0.9.1.

Models

  • DQN (Rainbow) — deep Q-network with prioritized replay, dueling heads, noisy nets
  • PPO — proximal policy optimization with GAE, LSTM policies, clip-higher
  • TFT — temporal fusion transformer for multi-horizon forecasting
  • Mamba2 — state space model for sequence prediction
  • Liquid Networks — biologically inspired networks for non-stationary data
  • TLOB — transformer-based limit order book analysis
  • KAN — Kolmogorov-Arnold networks
  • xLSTM — extended LSTM architecture
  • TGGN — temporal graph neural network
  • Diffusion — diffusion-based generative model

Key Modules

  • ensemble — model ensemble coordination and confidence aggregation
  • hyperopt — PSO-based hyperparameter optimization with per-model adapters
  • trainers — unified training loops (DQN, PPO, supervised)
  • inferenceInferenceAdapter trait for prediction
  • checkpoint — model checkpointing and restoration
  • evaluation — walk-forward evaluation pipeline

Usage

use ml::dqn::DQN;
use ml::ppo::PpoTrainer;