First half of Phase B4b (replay-buffer label scatter chain). Amends the Phase A2 trade_outcome_label_kernel to emit a per-(env, t) output column alongside the existing per-env tile, and adds the collector-side buffer + per-step launch arg. Kernel amendment (trade_outcome_label_kernel.cu): - Added NULL-tolerant `out_labels_per_sample` arg after existing `out_labels`. When non-NULL, writes `out_labels_per_sample[env*L + t] = label` at the same offset as the `trade_close_per_sample[env*L + t]` read. NULL = no-op (preserves Phase A2/A3 contract for callers passing old signature). - Pattern mirrors the K=2 head's `aux_sign_labels` per-(i, t) ring column that threads through the replay buffer. Collector field + alloc + launch: - New struct field `exp_aux_to_label_per_sample: CudaSlice<i32>` sized `[alloc_episodes × alloc_timesteps]`. Sentinel -1 (mask) populated by alloc_zeros + per-step kernel writes — survives until a trade-close event overwrites the env's slot at that t. - Updated Phase B3 launcher in collect_experiences_gpu to pass `self.exp_aux_to_label_per_sample.raw_ptr()` as the new arg. Why split B4b into B4b-1 + B4b-2: the full replay-buffer wireup mirrors the K=2 head's aux_sign_labels pattern across ~8 distinct code sites (replay-buffer struct field, sample destination buffer, direct-to-trainer pointer, setter method, scatter on insert, gather direct, gather fallback, GpuBatchPtrs field). Splitting lets us validate the per-(i, t) producer in isolation before touching the consumer pipeline. B4b-1 (this commit) = producer chain complete. Per-(env, t) column populated correctly every rollout step. Consumer wiring (replay- buffer scatter + trainer setter) is B4b-2's scope. Verification: - cargo check -p ml clean (21 warnings, none new). - cargo test -p ml --lib → 1016/0 on clean runs; pre-existing test_dqn_checkpoint_round_trip NoisyLinear flake still surfaces ~50-70% of full-suite runs (flake predates Phase B4b, unrelated to trade-outcome head — disable_noise() zeros ε but leaves some other randomness source intact). Audit: docs/dqn-wire-up-audit.md Phase B4b-1 section. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
ml
10-model ML ensemble for the Foxhunt HFT system, built on Candle v0.9.1.
Models
- DQN (Rainbow) — deep Q-network with prioritized replay, dueling heads, noisy nets
- PPO — proximal policy optimization with GAE, LSTM policies, clip-higher
- TFT — temporal fusion transformer for multi-horizon forecasting
- Mamba2 — state space model for sequence prediction
- Liquid Networks — biologically inspired networks for non-stationary data
- TLOB — transformer-based limit order book analysis
- KAN — Kolmogorov-Arnold networks
- xLSTM — extended LSTM architecture
- TGGN — temporal graph neural network
- Diffusion — diffusion-based generative model
Key Modules
ensemble— model ensemble coordination and confidence aggregationhyperopt— PSO-based hyperparameter optimization with per-model adapterstrainers— unified training loops (DQN, PPO, supervised)inference—InferenceAdaptertrait for predictioncheckpoint— model checkpointing and restorationevaluation— walk-forward evaluation pipeline
Usage
use ml::dqn::DQN;
use ml::ppo::PpoTrainer;