Files
foxhunt/crates/ml
jgrusewski a51345a141 fix: Layer 1+2 — C51 loss clamp + label smoothing (prevents 242K explosion)
Layer 1: Per-sample CE clamped to MAX_PER_SAMPLE_CE=50 before IS weight.
Breaks PER feedback loop: high loss → high priority → high IS weight → repeat.

Layer 2: Bellman target smoothed with ε=0.01 uniform mix after projection.
Prevents any atom from having zero probability → no log(near-zero) in CE.

Root cause: C51 cross-entropy is unbounded when target and predicted
distributions are maximally misaligned. PER amplifies pathological samples.
These two layers cap the maximum possible CE and prevent the extreme
misalignment from occurring.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-23 20:45:11 +01:00
..

ml

10-model ML ensemble for the Foxhunt HFT system, built on Candle v0.9.1.

Models

  • DQN (Rainbow) — deep Q-network with prioritized replay, dueling heads, noisy nets
  • PPO — proximal policy optimization with GAE, LSTM policies, clip-higher
  • TFT — temporal fusion transformer for multi-horizon forecasting
  • Mamba2 — state space model for sequence prediction
  • Liquid Networks — biologically inspired networks for non-stationary data
  • TLOB — transformer-based limit order book analysis
  • KAN — Kolmogorov-Arnold networks
  • xLSTM — extended LSTM architecture
  • TGGN — temporal graph neural network
  • Diffusion — diffusion-based generative model

Key Modules

  • ensemble — model ensemble coordination and confidence aggregation
  • hyperopt — PSO-based hyperparameter optimization with per-model adapters
  • trainers — unified training loops (DQN, PPO, supervised)
  • inferenceInferenceAdapter trait for prediction
  • checkpoint — model checkpointing and restoration
  • evaluation — walk-forward evaluation pipeline

Usage

use ml::dqn::DQN;
use ml::ppo::PpoTrainer;