Files
foxhunt/crates/ml
jgrusewski 387335e2b9 fix(dqn): SP1 Phase B foundation — stale-doc cleanup + Task 4 prep
Quality-review follow-ups to commit 53bc0bc50 (nan_flags_buf 24->48):

1. Update run_nan_checks_post_forward docstring (gpu_dqn_trainer.rs:14842-14873):
   replace 24-slot map with 48-slot range summary + audit-doc pointer.
   Was actively misleading after the 24->48 expansion; partial-refactor
   residue per feedback_no_partial_refactor.

2. Update '[24] system' comment in training_loop.rs (around line 2044):
   reflect the 48-slot post-expansion state (slots 0-23 fwd, 24-35 bwd,
   36-47 reserved). Also fix stale '0..11' tracing message to '0..47'.

3. Slot 31 (ensemble_d_logits_buf) annotation: flag DEFERRED + owner
   on FusedDqnTraining (different struct than 24-30, 32-35). Prevents
   Task 4 from blanket-launching check_nan_f32 on slot 31's null
   accessor.

4. Both name-table header comments now reference the future
   run_nan_checks_post_backward method (Task 4) plus the audit's
   per-slot table — pre-empts contract drift when Task 4 lands a
   3-way name-table dependency.

Audit-doc entry appended to docs/dqn-wire-up-audit.md SP1 Phase B
section. No behavioral change. Both name tables remain byte-identical.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-30 00:10:37 +02:00
..

ml

10-model ML ensemble for the Foxhunt HFT system, built on Candle v0.9.1.

Models

  • DQN (Rainbow) — deep Q-network with prioritized replay, dueling heads, noisy nets
  • PPO — proximal policy optimization with GAE, LSTM policies, clip-higher
  • TFT — temporal fusion transformer for multi-horizon forecasting
  • Mamba2 — state space model for sequence prediction
  • Liquid Networks — biologically inspired networks for non-stationary data
  • TLOB — transformer-based limit order book analysis
  • KAN — Kolmogorov-Arnold networks
  • xLSTM — extended LSTM architecture
  • TGGN — temporal graph neural network
  • Diffusion — diffusion-based generative model

Key Modules

  • ensemble — model ensemble coordination and confidence aggregation
  • hyperopt — PSO-based hyperparameter optimization with per-model adapters
  • trainers — unified training loops (DQN, PPO, supervised)
  • inferenceInferenceAdapter trait for prediction
  • checkpoint — model checkpointing and restoration
  • evaluation — walk-forward evaluation pipeline

Usage

use ml::dqn::DQN;
use ml::ppo::PpoTrainer;