Lands the plan-conditioning concat kernel for the K=3 trade-outcome forward as reusable scaffolding. Phase B5b (integration) is deferred with rationale: Phase C (state slots) is more critical for testing the K=3 head's effect on policy behavior, and can land independently of the plan-conditioning refinement. NEW kernel aux_to_input_concat_kernel.cu: - Writes [B, SH2+P] from h_s2_aux [B, 256] || plan_params [B, 6] - Pure GPU map; one thread per output element, no atomicAdd - NULL-tolerant on plan_params (zeros trailing P cols when source unavailable, e.g., collector cold-start where the trade plan head doesn't run) - Registered in build.rs; cubin compiles (5.7 KB). Dead code at this commit — no Rust launcher yet. Why B5 is split + B5b deferred: Full Phase B5 (integration) requires three coordinated changes: 1. Forward path: bump aux_to_fwd.forward() to SH2=262 + 262-dim input 2. Backward stride mismatch: backward emits dh_s2_aux_to_buf [B, 262], but dh_s2_aux_accum (input to aux trunk backward) is [B, 256]. A direct SAXPY mismatches row strides (262 vs 256) and corrupts the trunk's upstream gradient. Needs a strided-SAXPY kernel. 3. Collector-path plan_params unavailability: trade plan head only runs trainer-side. Workarounds: zero-fill, add trade plan to collector, or skip K=3 forward in collector. All have trade-offs. Phase B5b would need (1) strided-SAXPY kernel and (2) collector plan_params decision. Real work but NOT on the critical path for testing the K=3 head's effect on WR. Why Phase C should land first: The K=3 head currently trains on real labels (post-B4b) but doesn't influence policy behavior. Phase C wires the head's softmax into state slots [121..124) = (p_Profit, p_Stop, p_Timeout), replacing the K=2 single-slot 121 = 2*p_up - 1. WITH Phase C the policy reads aux's outcome predictions as state features → behavior changes → testable. Without Phase C, validation runs would show "K=3 head trains and converges" but predictions don't reach the policy → WR signal isn't a function of K=3 at all. We'd be testing nothing. Recommendation: skip the full B5 for now, do Phase C next, then Phase D (atom-shift). Phase B5b (plan-conditioning) is a refinement we add IF Phase C/D's no-plan-params version shows promise but plateaus below the WR ≥ 0.55 target. Audit: docs/dqn-wire-up-audit.md Phase B5a section. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
ml
10-model ML ensemble for the Foxhunt HFT system, built on Candle v0.9.1.
Models
- DQN (Rainbow) — deep Q-network with prioritized replay, dueling heads, noisy nets
- PPO — proximal policy optimization with GAE, LSTM policies, clip-higher
- TFT — temporal fusion transformer for multi-horizon forecasting
- Mamba2 — state space model for sequence prediction
- Liquid Networks — biologically inspired networks for non-stationary data
- TLOB — transformer-based limit order book analysis
- KAN — Kolmogorov-Arnold networks
- xLSTM — extended LSTM architecture
- TGGN — temporal graph neural network
- Diffusion — diffusion-based generative model
Key Modules
ensemble— model ensemble coordination and confidence aggregationhyperopt— PSO-based hyperparameter optimization with per-model adapterstrainers— unified training loops (DQN, PPO, supervised)inference—InferenceAdaptertrait for predictioncheckpoint— model checkpointing and restorationevaluation— walk-forward evaluation pipeline
Usage
use ml::dqn::DQN;
use ml::ppo::PpoTrainer;