Task 7 — Counterfactual Experience Augmentation: Every timestep writes BOTH original AND mirror experience to output. Mirror action: exposure_idx → (b0_size-1) - exposure_idx (L100↔S100). Mirror reward: -reward. Output buffers are 2x for counterfactual. Always active — no enable_ flag. Doubles effective data diversity. Task 20 — Lottery Ticket Pruning: lottery_ticket_mask kernel: zeros pruned weights in f32 + bf16 after Adam. lottery_ticket_compute_mask kernel: magnitude threshold → binary mask. At pruning epoch (50): read all weights, sort by |w|, bottom 70% → mask=0. After that: mask applied every training step (one kernel launch). Invalidates CUDA graphs on mask creation (new step structure). Enable_ flag removal: ALL enable_ conditionals removed from hot paths. One production path. - enable_domain_randomization → removed (always randomize) - enable_mirror_universe → removed (always mirror odd epochs) - enable_vol_normalization → removed (always normalize) - enable_anti_lr → removed (always anti-intuitive LR) - enable_gradient_vaccine → removed (always project gradients) - enable_causal_intervention → removed (always run interventions) - enable_adversarial_self_play → removed (always saboteur active) - enable_counterfactual → never added (always counterfactual) Domain randomization GPU kernels: removed enable_jitter and enable_dr parameters from CUDA kernel signatures. Kernels always randomize. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
ml
10-model ML ensemble for the Foxhunt HFT system, built on Candle v0.9.1.
Models
- DQN (Rainbow) — deep Q-network with prioritized replay, dueling heads, noisy nets
- PPO — proximal policy optimization with GAE, LSTM policies, clip-higher
- TFT — temporal fusion transformer for multi-horizon forecasting
- Mamba2 — state space model for sequence prediction
- Liquid Networks — biologically inspired networks for non-stationary data
- TLOB — transformer-based limit order book analysis
- KAN — Kolmogorov-Arnold networks
- xLSTM — extended LSTM architecture
- TGGN — temporal graph neural network
- Diffusion — diffusion-based generative model
Key Modules
ensemble— model ensemble coordination and confidence aggregationhyperopt— PSO-based hyperparameter optimization with per-model adapterstrainers— unified training loops (DQN, PPO, supervised)inference—InferenceAdaptertrait for predictioncheckpoint— model checkpointing and restorationevaluation— walk-forward evaluation pipeline
Usage
use ml::dqn::DQN;
use ml::ppo::PpoTrainer;