Files
foxhunt/crates/ml
jgrusewski 05eb574d0c fix(dqn): register NoisyLinear mu vars in VarMap + multi-thread runtime
Three fixes for GPU-accelerated Branching DQN training:

1. **GPU experience collector**: NoisyLinear creates standalone Vars via
   Var::from_tensor(), bypassing VarMap registration. The GPU collector
   looks up weights by name ("value_fc.weight") from VarMap and falls
   back to CPU (~5x slower) when missing. Fix: register mu vars in
   VarMap at construction, keep sigma vars standalone.

2. **Optimizer device mismatch**: Using only vars().all_vars() left
   NoisyLinear head params frozen. backward() produces gradients the
   optimizer doesn't know about → device mismatch in clip_grad_norm.
   Fix: all_trainable_vars() = VarMap (shared+mu) + sigma.

3. **Single-threaded CPU bottleneck**: Runtime::new() creates a
   current-thread scheduler → 1 OS thread → all async work serialized.
   Fix: multi-thread runtime (4 workers) created once in DQNTrainer::new(),
   shared across preload/training/backtest phases. Eliminates 3 fallback
   Runtime::new() callsites.

Also: polyak_update_var_pairs with debug_assert_eq, two-phase target
network sync (VarMap Polyak + sigma var_pairs Polyak), copy_weights_from
handles NoisyLinear heads.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-11 02:05:54 +01:00
..

ml

10-model ML ensemble for the Foxhunt HFT system, built on Candle v0.9.1.

Models

  • DQN (Rainbow) — deep Q-network with prioritized replay, dueling heads, noisy nets
  • PPO — proximal policy optimization with GAE, LSTM policies, clip-higher
  • TFT — temporal fusion transformer for multi-horizon forecasting
  • Mamba2 — state space model for sequence prediction
  • Liquid Networks — biologically inspired networks for non-stationary data
  • TLOB — transformer-based limit order book analysis
  • KAN — Kolmogorov-Arnold networks
  • xLSTM — extended LSTM architecture
  • TGGN — temporal graph neural network
  • Diffusion — diffusion-based generative model

Key Modules

  • ensemble — model ensemble coordination and confidence aggregation
  • hyperopt — PSO-based hyperparameter optimization with per-model adapters
  • trainers — unified training loops (DQN, PPO, supervised)
  • inferenceInferenceAdapter trait for prediction
  • checkpoint — model checkpointing and restoration
  • evaluation — walk-forward evaluation pipeline

Usage

use ml::dqn::DQN;
use ml::ppo::PpoTrainer;