The SP7 controller's CQL reference signal was reading cql_sx (post-SAXPY, budget-scaled) which created a self-perpetuating deadlock at small budget values: cql_sx_norm = budget × raw_grad → small budget → small cql_sx → controller can't update → budget stays small. The earlier offset 6 → 3 attempt failed because grad_decomp_launch_cql was never effectively populating slot 3 — the snapshot pattern measures ‖grad_buf − snapshot‖, but apply_cql_gradient writes to cql_grad_scratch (separate buffer), not grad_buf, so the snapshot delta is always 0. This commit adds a real producer kernel cql_raw_norm_compute that reads cql_grad_scratch directly and computes ‖raw_cql‖ over mag/dir/trunk slices. Wired to fire AFTER apply_cql_gradient and BEFORE apply_cql_saxpy, populating grad_decomp_result_pinned[3..6] with the raw norm independent of cql_budget. SP7 launcher updated to read from offset 3 (now: real raw CQL norm, not the never-populated cql snapshot delta). HEALTH_DIAG label renamed cql_sx → cql_raw to reflect the new contract; component index in the cached grad_component_norms_* arrays switched 2 → 1. The historical grad_decomp_launch_cql() call is removed — keeping it would overwrite the slot with 0 after cql_raw_norm_compute fires. The paired grad_decomp_snapshot_cql snapshot is left in place to scope the diff to Path A; buffer cleanup (grad_snapshot_cql allocation + grad_decomp_launch_cql definition) belongs in a follow-up commit per feedback_no_partial_refactor. Files: cql_raw_norm_kernel.cu (new, 97 LOC), build.rs, gpu_dqn_trainer.rs (struct field + cubin static + load + launcher + SP7 read offset 6→3), loss_balance_controller_kernel.cu (docstring + arg comment), fused_training.rs (4 launch_cql_raw_norm call sites, 1 dead grad_decomp_launch_cql call removed), training_loop.rs (HEALTH_DIAG label + index), audit doc Fix 31 SP7 Path A entry. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
ml
10-model ML ensemble for the Foxhunt HFT system, built on Candle v0.9.1.
Models
- DQN (Rainbow) — deep Q-network with prioritized replay, dueling heads, noisy nets
- PPO — proximal policy optimization with GAE, LSTM policies, clip-higher
- TFT — temporal fusion transformer for multi-horizon forecasting
- Mamba2 — state space model for sequence prediction
- Liquid Networks — biologically inspired networks for non-stationary data
- TLOB — transformer-based limit order book analysis
- KAN — Kolmogorov-Arnold networks
- xLSTM — extended LSTM architecture
- TGGN — temporal graph neural network
- Diffusion — diffusion-based generative model
Key Modules
ensemble— model ensemble coordination and confidence aggregationhyperopt— PSO-based hyperparameter optimization with per-model adapterstrainers— unified training loops (DQN, PPO, supervised)inference—InferenceAdaptertrait for predictioncheckpoint— model checkpointing and restorationevaluation— walk-forward evaluation pipeline
Usage
use ml::dqn::DQN;
use ml::ppo::PpoTrainer;