First half of Phase C. Lands the producer side of the K=3 trade- outcome aux head's state bridge: new kernel populates a per-env 3-slot cache from the K=3 softmax tile every rollout step. The consumer side (state gather reading from this cache → state slots [121..124)) lands in Phase C-2. Mirrors the K=2 head's existing aux_softmax_to_per_env_kernel exactly at K=3: - K=2: prev_aux_dir_prob[env] = 2*softmax[env, 1] - 1 (recentered) - K=3: prev_aux_outcome_probs[env, k] = softmax[env, k] for k in [0, 3) Changes: - state_layout.rs: 3 new constants AUX_OUTCOME_PROFIT_INDEX = 121, AUX_OUTCOME_STOP_INDEX = 122, AUX_OUTCOME_TIMEOUT_INDEX = 123. PROFIT_INDEX aliases AUX_DIR_PROB_INDEX (same value, different semantic). Phase C-2 flips slot 121's meaning from K=2's recentered p_up to K=3's p_Profit. - aux_outcome_softmax_to_per_env_kernel.cu: new kernel + cubin. - gpu_dqn_trainer.rs: new SP22_AUX_OUTCOME_SOFTMAX_TO_PER_ENV_CUBIN embed. - gpu_experience_collector.rs: 2 new struct fields (cache buffer + kernel handle); cubin load + alloc in constructor; struct-init; per-step launch in rollout loop after K=3 forward. - build.rs: kernel registered. Encoding shift K=2 → K=3: K=2 used recentered [-1, +1] to match "no signal = 0" baseline of every other slot. K=3 keeps raw softmax probabilities [0, 1]. Cold-start sentinel 0.0 for all 3 slots = "no prediction yet" (mask). The 3-slot natural distribution is more informative than a scalar. Dead-code status: producer populates cache every step but experience_state_gather doesn't read from it yet — state slot 121 still receives K=2's prev_aux_dir_prob write. Phase C-2 swaps the state gather's source from K=2 cache to K=3 cache (3-slot write). Why split C into C-1 + C-2: experience_state_gather is a hot-path kernel with many consumers. Updating it touches training collector, eval-side backtest evaluator, Rust launcher arg list. C-2 lands that as an atomic state-semantic flip; C-1 lands the GPU-side scaffolding independently so the producer chain can be validated first. Verification: - cargo check -p ml clean. - cargo test -p ml --lib → 1016/0 green. Audit: docs/dqn-wire-up-audit.md Phase C-1 section. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
ml
10-model ML ensemble for the Foxhunt HFT system, built on Candle v0.9.1.
Models
- DQN (Rainbow) — deep Q-network with prioritized replay, dueling heads, noisy nets
- PPO — proximal policy optimization with GAE, LSTM policies, clip-higher
- TFT — temporal fusion transformer for multi-horizon forecasting
- Mamba2 — state space model for sequence prediction
- Liquid Networks — biologically inspired networks for non-stationary data
- TLOB — transformer-based limit order book analysis
- KAN — Kolmogorov-Arnold networks
- xLSTM — extended LSTM architecture
- TGGN — temporal graph neural network
- Diffusion — diffusion-based generative model
Key Modules
ensemble— model ensemble coordination and confidence aggregationhyperopt— PSO-based hyperparameter optimization with per-model adapterstrainers— unified training loops (DQN, PPO, supervised)inference—InferenceAdaptertrait for predictioncheckpoint— model checkpointing and restorationevaluation— walk-forward evaluation pipeline
Usage
use ml::dqn::DQN;
use ml::ppo::PpoTrainer;