The `val_picked_dir_dist` reader added in 89ece2e36 read `chunked_actions_buf`
which is overwritten each chunk, so the diagnostic only saw the last ~64 bars
of the 214 654-bar val window. The cluster's run-after-run identical
distribution (`s=0.0001 h=0.20 l=0.0000 f=0.80` post-physics; `s=0.18 h=0.18
l=0.41 f=0.22` last-chunk picks) was therefore inconclusive: the post-physics
flatness covers the full window, but the pick-side diversity may only hold for
the last 64 bars while the first 99.97% of bars produce a different
distribution.
Add `picked_action_history_buf` — a window-major `[n_windows, max_len]` i32
buffer parallel to `actions_history_buf`. Filled by an additional
`scatter_intent_chunk` launch immediately after the existing intent-mag
scatter (same kernel handle, different src/dst — `chunked_actions_buf` →
`picked_action_history_buf`). One extra kernel launch per chunk; the launch
config and grid sizing are identical to the existing intent scatter.
The reader `read_chunked_actions_direction_distribution` now reads this
buffer instead of the chunk-local one. The HEALTH_DIAG line stays at
`val_picked_dir_dist [short=... hold=... long=... flat=...]` but now
reflects the full val window.
Decision matrix once the cluster reports the new diagnostic:
both `val_dir_dist` and `val_picked_dir_dist` show ~80% Flat
→ kernel itself produces collapsed picks across the window;
Q-values must be near-uniform for most bars; the eval-collapse is
a learning problem (network can't differentiate states).
`val_picked_dir_dist` diverse but `val_dir_dist` ~80% Flat
→ kernel is diverse, env_step drains active picks; the eval-collapse
is a physics problem (Kelly cap or another gate I haven't found).
Build clean at 11-warning baseline. No new kernel source, no determinism
contract change.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
ml
10-model ML ensemble for the Foxhunt HFT system, built on Candle v0.9.1.
Models
- DQN (Rainbow) — deep Q-network with prioritized replay, dueling heads, noisy nets
- PPO — proximal policy optimization with GAE, LSTM policies, clip-higher
- TFT — temporal fusion transformer for multi-horizon forecasting
- Mamba2 — state space model for sequence prediction
- Liquid Networks — biologically inspired networks for non-stationary data
- TLOB — transformer-based limit order book analysis
- KAN — Kolmogorov-Arnold networks
- xLSTM — extended LSTM architecture
- TGGN — temporal graph neural network
- Diffusion — diffusion-based generative model
Key Modules
ensemble— model ensemble coordination and confidence aggregationhyperopt— PSO-based hyperparameter optimization with per-model adapterstrainers— unified training loops (DQN, PPO, supervised)inference—InferenceAdaptertrait for predictioncheckpoint— model checkpointing and restorationevaluation— walk-forward evaluation pipeline
Usage
use ml::dqn::DQN;
use ml::ppo::PpoTrainer;