32 tasks across 5 milestones (E.0 foundation → E.4 shadow-mode), with locked design decisions from three rounds of focused research memos: - Q1 (fill sim): medium-tier Poisson regression from 5.2M trade tape - Q2 (reward): terminal-only, n-step credit (consumes ISV slot 517) - Q3 (alpha trust): implicit calibrated trust via state features - Q4 (state window): current snapshot + 2 short-horizon scalars - Q5 (sizing): hybrid decoupled fractional Kelly × Phase E attenuation Trainer choice: DQN primary (Rainbow + Munchausen target), PPO control on H=600 truncated only if kill criteria fire. Exploration: ε-greedy with kill-criteria gate at end of week 2; NoisyNet escalation (4-6 days due to dead scaffolding in our codebase) if criteria fail; RND beyond that. ISV consumption: 5 existing slots (n_step=517, γ=43-46, ε=41, Kelly=280, reward_caps=452-453); new block 539..550 reserved for Phase E (10 in active use, 2 spare). One new controller (stacker-threshold engagement-rate-self-correction at slot 543). Hardcoded by design: Kelly contract cap (Category-1 safety), kill-criteria thresholds (circuit breakers). All other knobs are ISV-driven per pearl_controller_anchors_isv_driven. Decisive gates at week 1 (H=600 kill criteria), week 4 (composition backtest Sharpe at half-tick > 0), and week 5 (shadow-vs-backtest PnL within 30%). Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
79 KiB
FoxhuntQ-Δ Phase E: Execution-Layer RL Policy — Implementation Plan
For agentic workers: REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (
- [ ]) syntax for tracking.
Goal: Build a learned execution-layer DQN policy that takes Mamba+stacker alpha output + current LOB state and decides order placement (passive L1/L2 vs market vs wait), reducing effective per-trade cost from "kills the strategy at half-tick" to "deployable at quarter-tick or better."
Architecture: GPU-pure DQN (Rainbow + Munchausen target + ε-greedy with kill criteria) over a fill-simulating environment, with state = [alpha_logit, alpha_confidence, spread_bps, L1_imbalance, current_position, time_in_position, recent_OFI_sum, recent_mid_drift, time_since_trade, book_event_rate] and 9-12 discrete actions. Position size is decoupled: fractional Kelly (consuming existing ISV slot KELLY_F_SMOOTH_INDEX=280 × new attenuation slot) determines target contracts; the execution policy only chooses placement. ISV-driven everywhere a controller anchor exists; safety caps and kill criteria hardcoded by design.
Tech Stack:
- Rust 1.85+, Edition 2021
- CUDA 12.4 via cudarc 0.19
- Existing infra:
crates/ml/src/cuda_pipeline/gpu_dqn_trainer.rs(~25K LOC GPU-pure DQN),crates/ml/src/trainers/dqn/,crates/ml/src/cuda_pipeline/sp{4-22}_isv_slots.rs(ISV registry),crates/ml-alpha(Mamba + stacker) - New code: ~3000 LOC in
crates/ml/src/env/,crates/ml/src/cuda_pipeline/phase_e_*,crates/ml/examples/phase_e_*
Anchor commits:
- Alpha foundation:
db874b184(Phase 1c snapshot pipeline) and4cf9499b5(Phase 1d.2 multi-minute validation) - Stacker:
ab4b7c64c - Backtest infra:
20c361a30
Design decisions (LOCKED)
These are the locked-in answers from three rounds of focused research. Tasks below assume these decisions; do not relitigate inline.
| Question | Decision | Source |
|---|---|---|
| Fill simulation | Medium tier: Poisson regression λ_l(spread, imbalance, ofi, time_since_trade) fitted once from 5.2M trade tape |
First memo |
| Reward shaping | Terminal-only at trade close (PnL − fees − implementation shortfall); n-step credit assignment via existing trainer | First memo + pearl_event_driven_reward_density_alignment |
| Alpha trust state | Raw alpha_logit + `alpha_confidence = |
p − 0.5 |
| State window | Current snapshot + 2 short-horizon scalars (5-snap OFI sum, 5-snap mid drift). NO trajectory windows — Mamba already carries history | First memo |
| Position sizing | Hybrid decoupled: closed-form fractional Kelly (consuming existing KELLY_F_SMOOTH_INDEX=280) × NEW attenuation slot; execution RL chooses placement only |
First memo + pearl_kelly_cap_signal_driven_floors |
| Trainer (primary) | DQN with existing Rainbow stack (n-step + dueling + C51 + PER) + add Munchausen target term | Second memo |
| Trainer (secondary, optional) | PPO on H=600 truncated as a sanity check ONLY if DQN's kill criteria fire | Second memo |
| Exploration | ε-greedy with linear schedule 1.0→0.05 over 500K steps + week-1 kill criteria; escalate to NoisyNet if criteria fire (4-6 days due to dead scaffolding); RND if NoisyNet also fails | Third memo |
| ISV consumption (existing slots) | n_step ← slot 517, γ ← slots 43-46, ε ← slot 41, Kelly_smooth ← slot 280, reward_caps ← slots 452-453 | ISV memo |
| ISV new slots | Block 539..550 reserved for Phase E (12 slots, contiguous, single fingerprint bump) | ISV memo |
| Hardcoded by design | Kelly contract cap = 10 (Category-1 safety), kill criteria thresholds (circuit-breakers, NOT controllers) | ISV memo |
File Structure
| File | Responsibility |
|---|---|
crates/ml/src/cuda_pipeline/phase_e_isv_slots.rs (new) |
Define indices 539..550, accessors, sentinel-bootstrap reads, layout fingerprint asserts |
crates/ml/src/trainers/dqn/state_reset_registry.rs (modify) |
Add 12 Phase E entries with dispatch arms per feedback_registry_entries_need_dispatch_arms |
crates/ml/src/env/execution_env.rs (new) |
DQN environment: state assembly, action execution, reward computation |
crates/ml/src/env/fill_model.rs (new) |
Poisson-regression fill simulator |
crates/ml/src/env/action_space.rs (new) |
9-action discrete enum + (level, aggressiveness, size) decoding |
crates/ml/src/cuda_pipeline/munchausen_target_kernel.cu (new) |
One-line addition to standard DQN target: adds `τ·log_π(a |
crates/ml/src/cuda_pipeline/stacker_threshold_controller.cu (new) |
Engagement-rate self-correction kernel: drives slot 543 from observed-rate vs target |
crates/ml/src/cuda_pipeline/phase_e_kill_criteria.cu (new) |
Computes Q-spread, action entropy, return-vs-random, early-Q-movement into diagnostic slots |
crates/ml/examples/phase_e_smoke.rs (new) |
Truncated H=600 DQN smoke for week-1 kill-criteria gate |
crates/ml/examples/phase_e_full.rs (new) |
Full H=6000 DQN training + backtest |
crates/ml/examples/phase_e_compose_backtest.rs (new) |
Alpha-policy + execution-policy composition backtest with realistic costs |
crates/ml/src/cuda_pipeline/gpu_dqn_trainer.rs (modify) |
Bump ISV_TOTAL_DIM; consume Phase E slots; wire kill criteria reads |
services/trading/src/phase_e_shadow.rs (new) |
Shadow-mode wiring: read live ticks, emit orders to log, compare to backtest expectation |
Pre-Flight
Task 0: Verify we're on the right branch and worktree
Files:
-
Read-only:
.git/HEAD,git status -
Step 1: Confirm worktree and branch
cd /home/jgrusewski/Work/foxhunt/.worktrees/sp19-20-wr-first
git rev-parse --abbrev-ref HEAD
Expected: sp20-aux-h-fixed
- Step 2: Confirm anchor commits exist
git log --oneline | grep -E '4cf9499b5|ab4b7c64c|20c361a30' | head -3
Expected: three lines matching the three anchor commits.
- Step 3: Confirm fxcache is present and ≥1GB
ls -lh /home/jgrusewski/Work/foxhunt/test_data/feature-cache/*.fxcache | head -1
Expected: a file ≥ 1G (the snapshot fxcache from 4cf9499b5).
Milestone E.0 — Foundation (~Week 1, Tasks 1-8)
Goal: Reserve ISV block, build environment scaffold + fill simulator + kill-criteria instrumentation, compute random-uniform baseline.
Decisive gate at end of week 1: All four kill criteria are computable on the ε-greedy DQN-untrained baseline. Numerical values printed and saved as ground truth. If we can't even compute them, Phase E.1 can't start.
Task 1: Reserve ISV slot block 539..550
Files:
-
Create:
crates/ml/src/cuda_pipeline/phase_e_isv_slots.rs -
Modify:
crates/ml/src/cuda_pipeline/gpu_dqn_trainer.rs(around theISV_TOTAL_DIMconstant, ~line 2860) -
Step 1: Find current
ISV_TOTAL_DIM
grep -n 'pub const ISV_TOTAL_DIM' crates/ml/src/cuda_pipeline/gpu_dqn_trainer.rs
Expected: a line pub const ISV_TOTAL_DIM: usize = N; where N is the current dim. Note N.
- Step 2: Write the Phase E slot file
Create crates/ml/src/cuda_pipeline/phase_e_isv_slots.rs:
//! Phase E (execution-layer RL) ISV slot block.
//!
//! Reserves indices 539..550 (12 slots, contiguous). All Phase E controller
//! anchors and diagnostic EMAs live here. Consumers read via the standard
//! sentinel-bootstrap pattern: `if isv[idx] <= sentinel + EPS { fallback }`.
//!
//! Slot layout:
//!
//! | Index | Name | Role | Update cadence | Producer kernel |
//! |-------|------------------------------|--------------|-----------------|------------------------------------------|
//! | 539 | Q_SPREAD_EMA_INDEX | Diagnostic | per-rollout-step| phase_e_kill_criteria.cu |
//! | 540 | ACTION_ENTROPY_EMA_INDEX | Diagnostic | per-rollout-step| phase_e_kill_criteria.cu |
//! | 541 | RETURN_VS_RANDOM_EMA_INDEX | Diagnostic | per-trade-close | phase_e_kill_criteria.cu |
//! | 542 | EARLY_Q_MOVEMENT_EMA_INDEX | Diagnostic | per-epoch | phase_e_kill_criteria.cu |
//! | 543 | STACKER_THRESHOLD_INDEX | Controller | per-rollout-step| stacker_threshold_controller.cu |
//! | 544 | TRADE_RATE_TARGET_INDEX | Anchor | static | host init |
//! | 545 | TRADE_RATE_OBSERVED_EMA | Diagnostic | per-rollout-step| stacker_threshold_controller.cu |
//! | 546 | STACKER_KELLY_ATTENUATION | Controller | per-trade-close | stacker_threshold_controller.cu |
//! | 547 | RANDOM_BASELINE_REWARD_MEAN | Anchor | per-fold | host computes once per fold |
//! | 548 | RANDOM_BASELINE_REWARD_STD | Anchor | per-fold | host computes once per fold |
//! | 549 | (reserved) | Spare | - | - |
//! | 550 | (reserved) | Spare | - | - |
//!
//! Reservation discipline (per ISV memo): slots 549..550 are intentionally
//! unused at first-commit to absorb growth without bumping
//! `ISV_TOTAL_DIM` again. Add new Phase E slots HERE first, not by
//! extending past 550 ad-hoc.
pub const PHASE_E_ISV_BLOCK_LO: usize = 539;
pub const PHASE_E_ISV_BLOCK_HI: usize = 550; // inclusive upper
// Diagnostics (read-only for circuit breakers; not controllers)
pub const Q_SPREAD_EMA_INDEX: usize = 539;
pub const ACTION_ENTROPY_EMA_INDEX: usize = 540;
pub const RETURN_VS_RANDOM_EMA_INDEX: usize = 541;
pub const EARLY_Q_MOVEMENT_EMA_INDEX: usize = 542;
// Stacker-threshold controller (engagement-rate self-correction)
pub const STACKER_THRESHOLD_INDEX: usize = 543;
pub const TRADE_RATE_TARGET_INDEX: usize = 544;
pub const TRADE_RATE_OBSERVED_EMA_INDEX: usize = 545;
pub const STACKER_KELLY_ATTENUATION_INDEX: usize = 546;
// Random-uniform baseline (computed once per fold; used by kill criterion)
pub const RANDOM_BASELINE_REWARD_MEAN_INDEX: usize = 547;
pub const RANDOM_BASELINE_REWARD_STD_INDEX: usize = 548;
// Sentinels: zero by convention. Fallbacks use the first-observation
// bootstrap pattern (pearl_first_observation_bootstrap).
pub const PHASE_E_SENTINEL: f32 = 0.0;
pub const PHASE_E_SENTINEL_EPS: f32 = 1e-9;
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn test_block_bounds_consistent() {
assert!(PHASE_E_ISV_BLOCK_HI >= PHASE_E_ISV_BLOCK_LO);
let block_size = PHASE_E_ISV_BLOCK_HI - PHASE_E_ISV_BLOCK_LO + 1;
assert_eq!(block_size, 12, "Phase E reserved 12 contiguous slots");
}
#[test]
fn test_all_indices_inside_block() {
let indices = [
Q_SPREAD_EMA_INDEX,
ACTION_ENTROPY_EMA_INDEX,
RETURN_VS_RANDOM_EMA_INDEX,
EARLY_Q_MOVEMENT_EMA_INDEX,
STACKER_THRESHOLD_INDEX,
TRADE_RATE_TARGET_INDEX,
TRADE_RATE_OBSERVED_EMA_INDEX,
STACKER_KELLY_ATTENUATION_INDEX,
RANDOM_BASELINE_REWARD_MEAN_INDEX,
RANDOM_BASELINE_REWARD_STD_INDEX,
];
for &i in &indices {
assert!(i >= PHASE_E_ISV_BLOCK_LO && i <= PHASE_E_ISV_BLOCK_HI,
"slot {i} out of Phase E block");
}
}
#[test]
fn test_total_dim_covers_block() {
// The parent ISV_TOTAL_DIM must be at least PHASE_E_ISV_BLOCK_HI + 1.
use crate::cuda_pipeline::gpu_dqn_trainer::ISV_TOTAL_DIM;
assert!(ISV_TOTAL_DIM > PHASE_E_ISV_BLOCK_HI,
"ISV_TOTAL_DIM ({ISV_TOTAL_DIM}) must exceed PHASE_E_ISV_BLOCK_HI ({PHASE_E_ISV_BLOCK_HI})");
}
}
- Step 3: Bump
ISV_TOTAL_DIMand add module declaration
Open crates/ml/src/cuda_pipeline/gpu_dqn_trainer.rs. Find the existing ISV_TOTAL_DIM constant (use the line number from Step 1). Set it to 551 (so indices 0..550 fit). Also find mod sp22_isv_slots; (or the highest SP slot module declaration) and add pub mod phase_e_isv_slots; right after.
- Step 4: Compile + test
SQLX_OFFLINE=true cargo test -p ml --lib phase_e_isv_slots -- --nocapture
Expected: 3 passed.
- Step 5: Commit
git add crates/ml/src/cuda_pipeline/phase_e_isv_slots.rs crates/ml/src/cuda_pipeline/gpu_dqn_trainer.rs
git commit -m "$(cat <<'EOF'
feat(phase-e): reserve ISV slot block 539..550 for execution-layer RL
12 contiguous slots reserved for Phase E controllers + diagnostics, per
the ISV research memo recommendation:
- 539-542: kill-criteria diagnostics (Q-spread, entropy, return-vs-random,
early-Q-movement EMAs)
- 543-546: stacker-threshold controller (threshold, target, observed-rate
EMA, Kelly attenuation)
- 547-548: random-uniform baseline (mean, std) — computed once per fold
- 549-550: reserved spare (absorb growth without bumping ISV_TOTAL_DIM)
ISV_TOTAL_DIM bumped to 551. All indices validated in unit tests per
the existing slot-registry pattern (sp{4,5,11,13,14,15,21,22}_isv_slots.rs).
EOF
)"
Task 2: Wire Phase E slots into state-reset registry
Files:
- Modify:
crates/ml/src/trainers/dqn/state_reset_registry.rs
Per feedback_registry_entries_need_dispatch_arms: every new ISV slot needs a registry entry AND a matching dispatch arm in reset_named_state in the same commit. If not, smoke crashes at fold 0 setup with "unknown name" error.
- Step 1: Find the existing registry section
grep -n 'pub fn build_registry' crates/ml/src/trainers/dqn/state_reset_registry.rs
grep -n 'pub fn reset_named_state' crates/ml/src/trainers/dqn/state_reset_registry.rs
Note the line numbers.
- Step 2: Add Phase E registry entries
Add 10 entries to the registry (one per non-spare Phase E slot). Each entry follows the existing pattern. Find an existing SP block (e.g., SP14 entries near crates/ml/src/cuda_pipeline/sp14_isv_slots.rs consumers) and add after it:
// ── Phase E (execution-layer RL) — slot block 539..550 ────────────
// Per phase_e_isv_slots.rs. Lifecycle category rationale:
// - Diagnostics (539-542) are SoftReset { decay_bars: 1000 } — slow EMAs
// that should fade after a fold to avoid contaminating across regimes.
// - Controllers (543, 545, 546) are FoldReset — every fold should learn
// its own threshold and attenuation from scratch.
// - Anchors (544, 547, 548) are TrainingPersist — set once per fold by
// the host, read all fold long.
entries.push(StateResetEntry {
name: "phase_e.q_spread_ema",
index: phase_e_isv_slots::Q_SPREAD_EMA_INDEX,
category: ResetCategory::SoftReset { decay_bars: 1000 },
source: "phase_e_kill_criteria.cu",
});
entries.push(StateResetEntry {
name: "phase_e.action_entropy_ema",
index: phase_e_isv_slots::ACTION_ENTROPY_EMA_INDEX,
category: ResetCategory::SoftReset { decay_bars: 1000 },
source: "phase_e_kill_criteria.cu",
});
entries.push(StateResetEntry {
name: "phase_e.return_vs_random_ema",
index: phase_e_isv_slots::RETURN_VS_RANDOM_EMA_INDEX,
category: ResetCategory::SoftReset { decay_bars: 1000 },
source: "phase_e_kill_criteria.cu",
});
entries.push(StateResetEntry {
name: "phase_e.early_q_movement_ema",
index: phase_e_isv_slots::EARLY_Q_MOVEMENT_EMA_INDEX,
category: ResetCategory::SoftReset { decay_bars: 1000 },
source: "phase_e_kill_criteria.cu",
});
entries.push(StateResetEntry {
name: "phase_e.stacker_threshold",
index: phase_e_isv_slots::STACKER_THRESHOLD_INDEX,
category: ResetCategory::FoldReset,
source: "stacker_threshold_controller.cu",
});
entries.push(StateResetEntry {
name: "phase_e.trade_rate_target",
index: phase_e_isv_slots::TRADE_RATE_TARGET_INDEX,
category: ResetCategory::TrainingPersist,
source: "host (Phase E config)",
});
entries.push(StateResetEntry {
name: "phase_e.trade_rate_observed_ema",
index: phase_e_isv_slots::TRADE_RATE_OBSERVED_EMA_INDEX,
category: ResetCategory::FoldReset,
source: "stacker_threshold_controller.cu",
});
entries.push(StateResetEntry {
name: "phase_e.stacker_kelly_attenuation",
index: phase_e_isv_slots::STACKER_KELLY_ATTENUATION_INDEX,
category: ResetCategory::FoldReset,
source: "stacker_threshold_controller.cu",
});
entries.push(StateResetEntry {
name: "phase_e.random_baseline_reward_mean",
index: phase_e_isv_slots::RANDOM_BASELINE_REWARD_MEAN_INDEX,
category: ResetCategory::TrainingPersist,
source: "host (Phase E.0 baseline computation)",
});
entries.push(StateResetEntry {
name: "phase_e.random_baseline_reward_std",
index: phase_e_isv_slots::RANDOM_BASELINE_REWARD_STD_INDEX,
category: ResetCategory::TrainingPersist,
source: "host (Phase E.0 baseline computation)",
});
- Step 3: Add dispatch arms in
reset_named_state
Find the match block in reset_named_state. Add 10 arms (one per entry above). The arms write the sentinel (0.0) and any Wiener state. Pattern from existing SP arms:
"phase_e.q_spread_ema" => {
isv[phase_e_isv_slots::Q_SPREAD_EMA_INDEX] = phase_e_isv_slots::PHASE_E_SENTINEL;
// Wiener state (3 floats per slot in wiener_state_buf) — reset to 0
let wiener_off = phase_e_isv_slots::Q_SPREAD_EMA_INDEX * 3;
wiener_state_buf[wiener_off] = 0.0; // sample_var
wiener_state_buf[wiener_off + 1] = 0.0; // diff_var
wiener_state_buf[wiener_off + 2] = 0.0; // x_lag
}
// ... same pattern for the 9 other slots
- Step 4: Run the registry coverage test
SQLX_OFFLINE=true cargo test -p ml --lib state_reset_registry -- --nocapture
Expected: existing tests pass, plus the new "every entry has a dispatch arm" coverage test passes.
- Step 5: Commit
git add crates/ml/src/trainers/dqn/state_reset_registry.rs
git commit -m "feat(phase-e): wire 10 Phase E ISV slots into state_reset_registry with dispatch arms per feedback_registry_entries_need_dispatch_arms"
Task 3: Build the action space enum
Files:
-
Create:
crates/ml/src/env/mod.rs(new module) -
Create:
crates/ml/src/env/action_space.rs -
Modify:
crates/ml/src/lib.rs(addpub mod env;) -
Step 1: Define the 9-action space
Create crates/ml/src/env/action_space.rs:
//! Phase E execution-policy action space.
//!
//! 9 discrete actions per the design memo. Size is decoupled (handled by
//! the Kelly layer); the action decides ONLY placement and side. WAIT
//! emits no order; the others emit one limit/market order at the
//! specified level with sign = side.
//!
//! Mapping rationale: 3 sides × 3 placements = 9, plus 1 wait = 10? No —
//! `WAIT` doesn't have a placement, so:
//!
//! 0 WAIT
//! 1 BUY @ market (aggressive, crosses spread)
//! 2 BUY @ L1 bid (passive, joins quote)
//! 3 BUY @ L2 bid (passive, behind quote)
//! 4 SELL @ market (aggressive, crosses spread)
//! 5 SELL @ L1 ask (passive, joins quote)
//! 6 SELL @ L2 ask (passive, behind quote)
//! 7 FLAT @ market (close current position aggressively)
//! 8 FLAT @ L1 (close current position passively at L1)
//!
//! Action gating: actions 1-3 are only legal when not already long;
//! actions 4-6 are only legal when not already short; 7-8 are only legal
//! when holding a position. Illegal actions degrade to WAIT; the gating
//! is enforced in the environment, not the network.
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
#[repr(u8)]
pub enum ExecAction {
Wait = 0,
BuyMarket = 1,
BuyL1 = 2,
BuyL2 = 3,
SellMarket = 4,
SellL1 = 5,
SellL2 = 6,
FlatMarket = 7,
FlatL1 = 8,
}
pub const N_ACTIONS: usize = 9;
#[derive(Debug, Clone, Copy)]
pub struct DecodedAction {
pub side: Side,
pub placement: Placement,
pub closing: bool,
}
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub enum Side { None, Buy, Sell }
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub enum Placement { None, Market, L1, L2 }
impl ExecAction {
pub fn from_u8(i: u8) -> Option<Self> {
match i {
0 => Some(Self::Wait),
1 => Some(Self::BuyMarket), 2 => Some(Self::BuyL1), 3 => Some(Self::BuyL2),
4 => Some(Self::SellMarket), 5 => Some(Self::SellL1), 6 => Some(Self::SellL2),
7 => Some(Self::FlatMarket), 8 => Some(Self::FlatL1),
_ => None,
}
}
pub fn decode(self) -> DecodedAction {
match self {
Self::Wait => DecodedAction { side: Side::None, placement: Placement::None, closing: false },
Self::BuyMarket => DecodedAction { side: Side::Buy, placement: Placement::Market, closing: false },
Self::BuyL1 => DecodedAction { side: Side::Buy, placement: Placement::L1, closing: false },
Self::BuyL2 => DecodedAction { side: Side::Buy, placement: Placement::L2, closing: false },
Self::SellMarket => DecodedAction { side: Side::Sell, placement: Placement::Market, closing: false },
Self::SellL1 => DecodedAction { side: Side::Sell, placement: Placement::L1, closing: false },
Self::SellL2 => DecodedAction { side: Side::Sell, placement: Placement::L2, closing: false },
Self::FlatMarket => DecodedAction { side: Side::None, placement: Placement::Market, closing: true },
Self::FlatL1 => DecodedAction { side: Side::None, placement: Placement::L1, closing: true },
}
}
/// Returns true if the action is legal given current position.
/// `position ∈ {-1, 0, +1}` (sign only; size is decoupled).
pub fn is_legal(self, position: i32) -> bool {
match self {
Self::Wait => true,
Self::BuyMarket | Self::BuyL1 | Self::BuyL2 => position <= 0,
Self::SellMarket | Self::SellL1 | Self::SellL2 => position >= 0,
Self::FlatMarket | Self::FlatL1 => position != 0,
}
}
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn test_n_actions_matches_enum() {
assert_eq!(N_ACTIONS, 9);
for i in 0..N_ACTIONS as u8 {
assert!(ExecAction::from_u8(i).is_some(), "action {i} missing");
}
assert!(ExecAction::from_u8(9).is_none());
}
#[test]
fn test_legality_gating() {
// Flat → can buy or sell, can't close
assert!(ExecAction::BuyMarket.is_legal(0));
assert!(ExecAction::SellMarket.is_legal(0));
assert!(!ExecAction::FlatMarket.is_legal(0));
// Long → can sell or close, can't buy more
assert!(!ExecAction::BuyMarket.is_legal(1));
assert!(ExecAction::SellMarket.is_legal(1));
assert!(ExecAction::FlatMarket.is_legal(1));
// Short → can buy or close, can't sell more
assert!(ExecAction::BuyMarket.is_legal(-1));
assert!(!ExecAction::SellMarket.is_legal(-1));
assert!(ExecAction::FlatMarket.is_legal(-1));
}
}
- Step 2: Wire up mod.rs
Create crates/ml/src/env/mod.rs:
//! Phase E execution-policy environment.
//!
//! Snapshot-stream replay with a Poisson-regression fill simulator and
//! discrete action space. Trains the existing GPU DQN trainer on order-
//! placement decisions for a given alpha signal + LOB state.
pub mod action_space;
pub mod fill_model; // lands in Task 4
pub mod execution_env; // lands in Task 6
pub use action_space::{ExecAction, DecodedAction, Side, Placement, N_ACTIONS};
- Step 3: Register module in lib.rs
In crates/ml/src/lib.rs, add pub mod env; near the other top-level mods.
- Step 4: Run tests
SQLX_OFFLINE=true cargo test -p ml --lib env::action_space -- --nocapture
Expected: 2 tests pass.
- Step 5: Commit
git add crates/ml/src/env/ crates/ml/src/lib.rs
git commit -m "feat(phase-e): 9-action discrete execution action space + legality gating"
Task 4: Build the fill model (Poisson regression scaffolding)
Files:
- Create:
crates/ml/src/env/fill_model.rs
Important context: This task creates the SCAFFOLDING. The actual coefficients get fitted in Task 5 from the historical trade tape. The scaffolding defines the data structure, sampling logic, and config.
- Step 1: Write the failing test for the scaffold
In crates/ml/src/env/fill_model.rs:
//! Poisson-regression fill simulator (medium tier per the design memo).
//!
//! For each level `l ∈ {L1, L2, L3}` and side, we fit a Poisson regression:
//!
//! λ_l(features) = exp(β_l · features)
//!
//! where features = [1, spread_bps, L1_imbalance, OFI_sum_5, log(time_since_trade + 1)].
//! At backtest time, the per-snapshot fill *probability* (Bernoulli) for a
//! posted limit order at level `l` is `1 − exp(−λ_l(features))` (the standard
//! Poisson-to-Bernoulli conversion for a single time bucket). For market
//! orders, fill is always immediate; the *price* is the L1 quote on the
//! opposite side (no slippage modeled at medium tier).
//!
//! The 5 coefficients per (level, side) = 6 distributions × 5 = 30 floats
//! total. Tiny. Fits in registers.
#[derive(Debug, Clone, Copy)]
pub struct FillCoeffs {
/// β_0 + β_1·spread + β_2·imbalance + β_3·ofi_5 + β_4·log(tau+1)
pub beta: [f32; 5],
}
#[derive(Debug, Clone)]
pub struct FillModel {
/// 6 distributions: (level=L1/L2/L3) × (side=BID/ASK). Buy passive uses
/// BID coefficients (we want to be filled by an aggressor crossing
/// down); sell passive uses ASK coefficients.
pub bid_coeffs: [FillCoeffs; 3],
pub ask_coeffs: [FillCoeffs; 3],
}
#[derive(Debug, Clone, Copy)]
pub struct FillFeatures {
pub spread_bps: f32,
pub l1_imbalance: f32,
pub ofi_sum_5: f32,
pub time_since_trade_s: f32,
}
impl FillModel {
/// Untrained-skeleton constructor (coefficients all zero → λ=1 → ~63%
/// fill, which is reasonable as a sanity default before Task 5 fits
/// real coefficients).
pub fn skeleton() -> Self {
let zero = FillCoeffs { beta: [0.0; 5] };
Self {
bid_coeffs: [zero; 3],
ask_coeffs: [zero; 3],
}
}
/// Fill rate λ for a (side, level) at the given features.
pub fn lambda(&self, side: crate::env::action_space::Side, level: usize,
feat: &FillFeatures) -> f32 {
let coeffs = match side {
crate::env::action_space::Side::Buy => &self.bid_coeffs,
crate::env::action_space::Side::Sell => &self.ask_coeffs,
crate::env::action_space::Side::None => return 0.0, // closing actions handled separately
};
let c = &coeffs[level.min(2)].beta;
let x = [1.0, feat.spread_bps, feat.l1_imbalance, feat.ofi_sum_5,
(feat.time_since_trade_s + 1.0).ln()];
let linear = c.iter().zip(x.iter()).map(|(a, b)| a * b).sum::<f32>();
// λ is a rate, must be non-negative. Use exp (canonical Poisson link).
linear.exp().min(50.0) // cap to avoid numerical blow-up; λ=50 is 1-second saturation
}
/// Per-snapshot Bernoulli fill probability for a posted order at level `level`.
pub fn fill_prob(&self, side: crate::env::action_space::Side, level: usize,
feat: &FillFeatures) -> f32 {
let lam = self.lambda(side, level, feat);
1.0 - (-lam).exp()
}
}
#[cfg(test)]
mod tests {
use super::*;
use crate::env::action_space::Side;
#[test]
fn test_skeleton_has_uniform_fill_prob() {
// β=0 → λ=exp(0)=1 → fill_prob = 1 − exp(−1) ≈ 0.632
let m = FillModel::skeleton();
let feat = FillFeatures { spread_bps: 0.5, l1_imbalance: 0.5,
ofi_sum_5: 0.0, time_since_trade_s: 0.5 };
let p = m.fill_prob(Side::Buy, 0, &feat);
assert!((p - 0.632).abs() < 0.01, "skeleton λ=1 → p≈0.632, got {p}");
}
#[test]
fn test_fill_prob_bounded() {
// Even with huge features and zero coefficients, prob is in [0, 1].
let m = FillModel::skeleton();
let feat = FillFeatures { spread_bps: 1000.0, l1_imbalance: 1.0,
ofi_sum_5: 100.0, time_since_trade_s: 1000.0 };
let p = m.fill_prob(Side::Buy, 0, &feat);
assert!((0.0..=1.0).contains(&p));
}
}
- Step 2: Run tests
SQLX_OFFLINE=true cargo test -p ml --lib env::fill_model -- --nocapture
Expected: 2 passed.
- Step 3: Commit
git add crates/ml/src/env/fill_model.rs
git commit -m "feat(phase-e): Poisson regression fill model scaffold (coefficients fitted in Task 5)"
Task 5: Fit the fill model from historical trade tape
Files:
-
Modify:
crates/ml/src/env/fill_model.rs(add fitting routine) -
Create:
crates/ml/examples/phase_e_fit_fill_model.rs -
Step 1: Append fitting routine to
fill_model.rs
Add this section:
/// Fit FillModel coefficients from historical (snapshot, trade) pairs.
///
/// For each (side, level) we treat "was there a trade at price ≤ posted
/// price within the next `dt_seconds` window" as a Bernoulli observation,
/// then transform to Poisson via the canonical link. With ~5M trades over
/// ~16 days, even per-level sample counts are 6 figures — plenty for
/// 5-coefficient Poisson regression. IRLS or gradient-descent both work.
///
/// We use simple gradient descent on the canonical Poisson NLL for
/// implementation simplicity (no external linalg dependency required).
pub fn fit_poisson(
features: &[FillFeatures],
observed_fills: &[bool],
max_iters: usize,
lr: f32,
) -> Result<FillCoeffs, &'static str> {
if features.len() != observed_fills.len() || features.is_empty() {
return Err("fit_poisson: length mismatch or empty");
}
let mut beta = [0.0_f32; 5];
let n = features.len() as f32;
for _ in 0..max_iters {
let mut grad = [0.0_f32; 5];
for (feat, &y) in features.iter().zip(observed_fills) {
let x = [1.0, feat.spread_bps, feat.l1_imbalance, feat.ofi_sum_5,
(feat.time_since_trade_s + 1.0).ln()];
let linear = beta.iter().zip(x.iter()).map(|(a, b)| a * b).sum::<f32>();
let mu = linear.exp().min(50.0); // Poisson mean
let target = if y { 1.0 } else { 0.0 };
// Gradient of -log L per sample: (mu - target) * x
let resid = mu - target;
for k in 0..5 {
grad[k] += resid * x[k];
}
}
for k in 0..5 {
beta[k] -= lr * grad[k] / n;
}
}
Ok(FillCoeffs { beta })
}
Add a test:
#[test]
fn test_fit_recovers_baseline_when_features_uninformative() {
// 1000 random fill events at uniform p=0.4 — fitting should give
// β_0 ≈ ln(λ) where 1−exp(−λ)=0.4 → λ≈0.51 → β_0≈−0.67.
let n = 1000;
let features: Vec<FillFeatures> = (0..n).map(|_| FillFeatures {
spread_bps: 0.0, l1_imbalance: 0.0, ofi_sum_5: 0.0, time_since_trade_s: 0.0,
}).collect();
// Deterministic 40% fill rate
let observed: Vec<bool> = (0..n).map(|i| (i * 401 % 1000) < 400).collect();
let coeffs = fit_poisson(&features, &observed, 1000, 1e-2).expect("fit");
let lam = coeffs.beta[0].exp();
let p = 1.0 - (-lam).exp();
assert!((p - 0.4).abs() < 0.05, "recovered fill rate {p} should be near 0.4");
}
- Step 2: Run the test
SQLX_OFFLINE=true cargo test -p ml --lib env::fill_model::tests::test_fit_recovers -- --nocapture
Expected: passes.
- Step 3: Author the calibration example
Create crates/ml/examples/phase_e_fit_fill_model.rs:
//! Phase E.0 — fit FillModel coefficients from the historical trade tape.
//!
//! Iterates the snapshot fxcache, for each snapshot computes the
//! per-level Bernoulli fill outcome ("was there a contra-side trade at
//! ≤ posted price within the next 60 seconds?"), then fits the Poisson
//! regression. Writes coefficients to `phase_e_fill_coeffs.json` for
//! later consumption by the runtime FillModel.
use anyhow::{Context, Result};
use clap::Parser;
use std::path::PathBuf;
use std::fs::File;
use std::io::Write;
use ml::env::fill_model::{fit_poisson, FillCoeffs, FillFeatures};
use ml::env::action_space::Side;
#[derive(Debug, Parser)]
struct Cli {
#[arg(long)] fxcache_path: String,
#[arg(long)] trades_dir: String,
#[arg(long, default_value_t = 60.0)] window_seconds: f32,
#[arg(long, default_value = "phase_e_fill_coeffs.json")] out_path: PathBuf,
}
fn main() -> Result<()> {
tracing_subscriber::fmt().init();
let cli = Cli::parse();
// (1) Load fxcache snapshots + raw trade tape from MBP-10 files.
// (2) For each snapshot, build FillFeatures.
// (3) For each (side, level), determine: was there an opposing trade
// that would have filled a posted order at this level within
// `window_seconds`?
// (4) Call fit_poisson per (side, level).
// (5) Serialize the 6 FillCoeffs to JSON.
//
// Implementation pattern: copy the existing snapshot + trade
// loading from `crates/ml/examples/precompute_features.rs` (the v10
// imbalance-bar precompute), reuse the (snapshot, trades_in_window)
// iteration loop, replace the feature computation with the binary
// fill-or-not determination.
//
// Concrete code (paste-and-run skeleton; loaders identical to
// precompute_features.rs Stage 1):
let n_snapshots = 1_000_000; // placeholder; load actual count from fxcache
eprintln!("Loaded {n_snapshots} snapshots; trades in window={}s", cli.window_seconds);
// Per (side, level), collect (features, fills).
let mut bid_data: Vec<(Vec<FillFeatures>, Vec<bool>)> =
(0..3).map(|_| (Vec::new(), Vec::new())).collect();
let mut ask_data: Vec<(Vec<FillFeatures>, Vec<bool>)> =
(0..3).map(|_| (Vec::new(), Vec::new())).collect();
// ... iteration loop (omitted: lift from precompute_features.rs) ...
// For each snapshot s, for each level l ∈ {0, 1, 2}:
// feat = FillFeatures { spread_bps: s.spread_bps, l1_imbalance: ...,
// ofi_sum_5: ..., time_since_trade_s: ... };
// bid_fill = trades_in_window.iter().any(|t| t.is_sell && t.price <= s.bid_l[l]);
// ask_fill = trades_in_window.iter().any(|t| !t.is_sell && t.price >= s.ask_l[l]);
// bid_data[l].0.push(feat); bid_data[l].1.push(bid_fill);
// ask_data[l].0.push(feat); ask_data[l].1.push(ask_fill);
let bid_coeffs: Vec<FillCoeffs> = bid_data.iter()
.map(|(f, y)| fit_poisson(f, y, 500, 1e-3).expect("fit"))
.collect();
let ask_coeffs: Vec<FillCoeffs> = ask_data.iter()
.map(|(f, y)| fit_poisson(f, y, 500, 1e-3).expect("fit"))
.collect();
// Serialize to JSON.
let json = serde_json::json!({
"bid_coeffs": bid_coeffs.iter().map(|c| c.beta).collect::<Vec<_>>(),
"ask_coeffs": ask_coeffs.iter().map(|c| c.beta).collect::<Vec<_>>(),
});
let mut f = File::create(&cli.out_path).context("create out file")?;
write!(f, "{}", serde_json::to_string_pretty(&json)?)?;
eprintln!("Wrote {}", cli.out_path.display());
Ok(())
}
Add serde_json to crates/ml/Cargo.toml [dependencies] if not already present (it almost certainly is).
Also make FillCoeffs and FillFeatures derive serde::Serialize, serde::Deserialize for JSON round-trip:
use serde::{Deserialize, Serialize};
#[derive(Debug, Clone, Copy, Serialize, Deserialize)]
pub struct FillCoeffs {
pub beta: [f32; 5],
}
#[derive(Debug, Clone, Copy, Serialize, Deserialize)]
pub struct FillFeatures {
pub spread_bps: f32,
pub l1_imbalance: f32,
pub ofi_sum_5: f32,
pub time_since_trade_s: f32,
}
- Step 4: Build the example
SQLX_OFFLINE=true cargo build -p ml --release --example phase_e_fit_fill_model
Expected: clean compile.
- Step 5: Run on the snapshot fxcache + trades dir
FXC=$(ls -t /home/jgrusewski/Work/foxhunt/test_data/feature-cache/*.fxcache | head -1)
target/release/examples/phase_e_fit_fill_model \
--fxcache-path "$FXC" \
--trades-dir /home/jgrusewski/Work/foxhunt/test_data/futures-baseline-trades \
--window-seconds 60 \
--out-path phase_e_fill_coeffs.json
Expected: prints "Loaded N snapshots", "Wrote phase_e_fill_coeffs.json". File should be ~1KB with 6 arrays of 5 floats.
- Step 6: Sanity-check the fitted coefficients
cat phase_e_fill_coeffs.json | head -20
Expected: β_0 (intercept) values should be negative (since baseline fill rate at 60s is well below 100%). β values on spread_bps should be negative for passive bids (wider spread → fewer fills) and positive for L1_imbalance on bid side (more bids → bid-side fills less likely as price stays up). If signs are reversed, the fitting loop has a bug; debug before committing.
- Step 7: Commit
git add crates/ml/src/env/fill_model.rs crates/ml/examples/phase_e_fit_fill_model.rs phase_e_fill_coeffs.json
git commit -m "feat(phase-e): Poisson regression fit of FillModel coefficients from 5.2M trade tape"
Task 6: Build the ExecutionEnv
Files:
- Create:
crates/ml/src/env/execution_env.rs
This is the core environment. State assembly, step semantics, reward computation, termination.
- Step 1: Write the env skeleton
//! Phase E execution-policy environment.
//!
//! Each EPISODE is a single trade lifecycle: from initial position-open
//! signal (alpha confidence above threshold) through K_horizon snapshots
//! later (forced close). Within the episode, the policy can:
//! - WAIT (no order)
//! - Place a buy/sell at market/L1/L2 (open or add)
//! - FLAT at market/L1 (close)
//!
//! Reward at episode close: realized PnL minus fees minus implementation
//! shortfall vs mid-at-decision. No dense per-snapshot reward.
use crate::env::action_space::{ExecAction, N_ACTIONS, Side, Placement};
use crate::env::fill_model::{FillModel, FillFeatures};
pub const STATE_DIM: usize = 10;
#[derive(Debug, Clone)]
pub struct ExecutionEnvConfig {
pub horizon_snapshots: usize, // K = 6000 (or 600 for E.1 smoke)
pub trade_size_contracts: i32, // From Kelly layer; passed in per-episode
pub cost_per_contract: f32, // Maker rebate / taker fee schedule; in price units
pub max_step_per_episode: usize, // Defaults to horizon_snapshots
}
#[derive(Debug, Clone)]
pub struct EpisodeState {
pub step: usize,
pub position: i32, // signed contracts held
pub entry_price: f32, // 0 if flat
pub mid_at_decision: f32, // Mid price at the original alpha-signal snapshot
pub cumulative_pnl: f32,
pub n_fills: u32,
}
pub struct ExecutionEnv {
pub config: ExecutionEnvConfig,
pub fill_model: FillModel,
// Iterator state over the snapshot fxcache replay
snapshots: Vec<SnapshotRow>,
cursor: usize,
}
#[derive(Debug, Clone)]
pub struct SnapshotRow {
pub mid_price: f32,
pub bid_l: [f32; 3],
pub ask_l: [f32; 3],
pub alpha_logit: f32, // From the alpha layer (Mamba+stacker)
pub alpha_confidence: f32, // |sigmoid(alpha_logit) - 0.5|
pub spread_bps: f32,
pub l1_imbalance: f32,
pub ofi_sum_5: f32, // Rolling 5-snapshot OFI sum
pub mid_drift_5: f32, // Rolling 5-snapshot mid drift (bps)
pub time_since_trade_s: f32,
pub book_event_rate: f32,
}
impl ExecutionEnv {
pub fn new(config: ExecutionEnvConfig, fill_model: FillModel,
snapshots: Vec<SnapshotRow>) -> Self {
Self { config, fill_model, snapshots, cursor: 0 }
}
/// Assemble the state vector at the current cursor.
pub fn state(&self, ep: &EpisodeState) -> [f32; STATE_DIM] {
let s = &self.snapshots[self.cursor];
[
s.alpha_logit,
s.alpha_confidence,
s.spread_bps,
s.l1_imbalance,
s.ofi_sum_5,
s.mid_drift_5,
ep.position as f32,
ep.step as f32 / self.config.horizon_snapshots as f32, // normalised
s.time_since_trade_s.ln_1p(),
s.book_event_rate.ln_1p(),
]
}
/// One env step. Returns (next_state, reward, done).
pub fn step(&mut self, action_id: u8, ep: &mut EpisodeState)
-> Option<([f32; STATE_DIM], f32, bool)>
{
let action = ExecAction::from_u8(action_id)?;
let decoded = action.decode();
// Legality gate: illegal actions degrade to WAIT
let legal = action.is_legal(ep.position);
let effective = if legal { action } else { ExecAction::Wait };
let effective_decoded = effective.decode();
let s = &self.snapshots[self.cursor];
let feat = FillFeatures {
spread_bps: s.spread_bps,
l1_imbalance: s.l1_imbalance,
ofi_sum_5: s.ofi_sum_5,
time_since_trade_s: s.time_since_trade_s,
};
// Execute the action: try to fill, update position + entry_price.
let mut step_pnl = 0.0;
let mut filled = false;
match (effective_decoded.placement, effective_decoded.side, effective_decoded.closing) {
(Placement::None, _, _) => {} // Wait
(Placement::Market, Side::Buy, _) => {
// Immediate fill at ask
let fill_px = s.ask_l[0];
ep.position += self.config.trade_size_contracts;
if ep.entry_price == 0.0 { ep.entry_price = fill_px; }
step_pnl -= self.config.cost_per_contract; // taker fee
filled = true;
}
(Placement::Market, Side::Sell, _) => {
let fill_px = s.bid_l[0];
ep.position -= self.config.trade_size_contracts;
if ep.entry_price == 0.0 { ep.entry_price = fill_px; }
step_pnl -= self.config.cost_per_contract;
filled = true;
}
(Placement::Market, Side::None, true) => {
// Flat at market
let fill_px = if ep.position > 0 { s.bid_l[0] } else { s.ask_l[0] };
let direction = ep.position.signum() as f32;
step_pnl = direction * (fill_px - ep.entry_price)
* (ep.position.abs() as f32);
step_pnl -= self.config.cost_per_contract * ep.position.abs() as f32;
ep.position = 0;
ep.entry_price = 0.0;
filled = true;
}
(Placement::L1, side, closing) | (Placement::L2, side, closing) => {
let level = if effective_decoded.placement == Placement::L1 { 0 } else { 1 };
let fill_p = self.fill_model.fill_prob(
if closing { match side { _ => Side::None } } else { side },
level, &feat,
);
// Deterministic LCG-style "random" for replay reproducibility — in
// real implementation use ChaCha8Rng seeded with episode ID
let rand_v = ((self.cursor.wrapping_mul(2654435761)) as f32) / u32::MAX as f32;
if rand_v < fill_p {
// Filled
let fill_px = match (effective_decoded.side, closing) {
(Side::Buy, _) => s.bid_l[level], // we posted at bid; got hit
(Side::Sell, _) => s.ask_l[level],
(Side::None, true) => if ep.position > 0 { s.ask_l[level] }
else { s.bid_l[level] },
_ => return None,
};
if closing {
let direction = ep.position.signum() as f32;
step_pnl = direction * (fill_px - ep.entry_price)
* (ep.position.abs() as f32);
ep.position = 0;
ep.entry_price = 0.0;
} else {
match side {
Side::Buy => ep.position += self.config.trade_size_contracts,
Side::Sell => ep.position -= self.config.trade_size_contracts,
_ => {}
}
if ep.entry_price == 0.0 { ep.entry_price = fill_px; }
}
// Maker rebate: no taker fee; in some markets a positive rebate
step_pnl -= 0.0; // No fee for makers
filled = true;
}
// else: no fill this step
}
_ => {} // Other combinations shouldn't happen
}
ep.cumulative_pnl += step_pnl;
if filled { ep.n_fills += 1; }
// Advance cursor and step counter
self.cursor = (self.cursor + 1).min(self.snapshots.len() - 1);
ep.step += 1;
// Termination: horizon reached or out of snapshots
let done = ep.step >= self.config.horizon_snapshots
|| self.cursor + 1 >= self.snapshots.len();
let reward = if done {
// Force-close any open position at mid
if ep.position != 0 {
let close_px = self.snapshots[self.cursor].mid_price;
let direction = ep.position.signum() as f32;
let close_pnl = direction * (close_px - ep.entry_price)
* (ep.position.abs() as f32);
ep.cumulative_pnl += close_pnl;
}
ep.cumulative_pnl
} else { 0.0 };
Some((self.state(ep), reward, done))
}
}
#[cfg(test)]
mod tests {
use super::*;
fn synthetic_snapshots(n: usize, slope: f32) -> Vec<SnapshotRow> {
(0..n).map(|i| SnapshotRow {
mid_price: 4500.0 + i as f32 * slope,
bid_l: [4499.875, 4499.625, 4499.375],
ask_l: [4500.125, 4500.375, 4500.625],
alpha_logit: 1.0,
alpha_confidence: 0.3,
spread_bps: 0.25,
l1_imbalance: 0.55,
ofi_sum_5: 0.0,
mid_drift_5: 0.0,
time_since_trade_s: 0.5,
book_event_rate: 5.0,
}).collect()
}
#[test]
fn test_state_dim() {
let snaps = synthetic_snapshots(100, 0.1);
let env = ExecutionEnv::new(
ExecutionEnvConfig {
horizon_snapshots: 100,
trade_size_contracts: 1,
cost_per_contract: 0.0625, // quarter-tick
max_step_per_episode: 100,
},
FillModel::skeleton(),
snaps,
);
let ep = EpisodeState { step: 0, position: 0, entry_price: 0.0,
mid_at_decision: 4500.0, cumulative_pnl: 0.0,
n_fills: 0 };
let s = env.state(&ep);
assert_eq!(s.len(), STATE_DIM);
for v in &s { assert!(v.is_finite()); }
}
#[test]
fn test_market_buy_then_market_sell_on_uptrend_profits() {
let snaps = synthetic_snapshots(20, 1.0); // strong uptrend
let mut env = ExecutionEnv::new(
ExecutionEnvConfig {
horizon_snapshots: 20,
trade_size_contracts: 1,
cost_per_contract: 0.0625,
max_step_per_episode: 20,
},
FillModel::skeleton(),
snaps,
);
let mut ep = EpisodeState { step: 0, position: 0, entry_price: 0.0,
mid_at_decision: 4500.0, cumulative_pnl: 0.0,
n_fills: 0 };
// BUY at market
env.step(ExecAction::BuyMarket as u8, &mut ep);
assert_eq!(ep.position, 1);
// Wait 18 snapshots
for _ in 0..18 {
env.step(ExecAction::Wait as u8, &mut ep);
}
// FLAT at market — episode will end on this step
let (_, reward, done) = env.step(ExecAction::FlatMarket as u8, &mut ep).unwrap();
assert!(done);
// With slope=1.0 over 20 steps, entry at ask=4500.125 and exit at bid
// somewhere ≈ 4519.875 → profit ≈ 19.75 − 2·0.0625 = 19.625
assert!(reward > 15.0, "uptrend buy-sell should profit, got {reward}");
}
}
- Step 2: Wire into mod.rs
In crates/ml/src/env/mod.rs, add pub mod execution_env; and the re-exports:
pub use execution_env::{ExecutionEnv, ExecutionEnvConfig, EpisodeState, SnapshotRow, STATE_DIM};
- Step 3: Run tests
SQLX_OFFLINE=true cargo test -p ml --lib env::execution_env -- --nocapture
Expected: 2 tests pass.
- Step 4: Commit
git add crates/ml/src/env/execution_env.rs crates/ml/src/env/mod.rs
git commit -m "feat(phase-e): ExecutionEnv with 10-dim state, 9-action gating, terminal-only reward"
Task 7: Compute random-uniform-policy baseline reward
Files:
- Create:
crates/ml/examples/phase_e_baseline.rs
The kill-criteria depends on knowing the reward distribution of a random policy. This task computes that ONCE per fold and writes it to ISV slots 547 (mean) and 548 (std).
- Step 1: Write the baseline example
//! Phase E.0 — compute random-uniform-policy baseline reward.
//!
//! Runs N=10,000 episodes with a uniform random action policy, records
//! the terminal reward distribution. Writes mean and std to the Phase E
//! ISV slots so the DQN trainer's kill-criteria can fire on
//! `rollout_reward < random_baseline_mean + 2σ_random`.
use anyhow::{Context, Result};
use clap::Parser;
use rand::{Rng, SeedableRng};
use rand_chacha::ChaCha8Rng;
use ml::env::{ExecutionEnv, ExecutionEnvConfig, EpisodeState, FillModel, N_ACTIONS};
#[derive(Parser)]
struct Cli {
#[arg(long)] fxcache_path: String,
#[arg(long, default_value = "phase_e_fill_coeffs.json")] fill_coeffs_path: String,
#[arg(long, default_value_t = 10_000)] n_episodes: usize,
#[arg(long, default_value_t = 6000)] horizon: usize,
#[arg(long, default_value_t = 42)] seed: u64,
}
fn main() -> Result<()> {
tracing_subscriber::fmt().init();
let cli = Cli::parse();
// 1. Load snapshots from fxcache.
// 2. Load fitted fill coefficients.
// 3. Build env.
// 4. Run N episodes with uniform random actions; record terminal reward.
// 5. Compute mean + std, print, save.
let snapshots: Vec<ml::env::SnapshotRow> = vec![]; // TODO: lift from phase1d_long_horizon.rs
// ... loading logic omitted; reuse fxcache reader + alpha logit retrieval pattern from
// `crates/ml-alpha/examples/phase1d_long_horizon.rs`
let fill_model: FillModel = serde_json::from_reader(
std::fs::File::open(&cli.fill_coeffs_path)?
)?;
let mut rng = ChaCha8Rng::seed_from_u64(cli.seed);
let mut env = ExecutionEnv::new(
ExecutionEnvConfig {
horizon_snapshots: cli.horizon,
trade_size_contracts: 1,
cost_per_contract: 0.0625,
max_step_per_episode: cli.horizon,
},
fill_model,
snapshots,
);
let mut rewards = Vec::with_capacity(cli.n_episodes);
for ep_id in 0..cli.n_episodes {
let mut ep = EpisodeState { step: 0, position: 0, entry_price: 0.0,
mid_at_decision: 4500.0, cumulative_pnl: 0.0,
n_fills: 0 };
let mut total_reward = 0.0_f32;
loop {
let a = rng.gen_range(0..N_ACTIONS as u8);
match env.step(a, &mut ep) {
Some((_, r, done)) => {
total_reward += r;
if done { break; }
}
None => break,
}
}
rewards.push(total_reward);
if ep_id % 1000 == 0 {
eprintln!("ep {ep_id}: reward={total_reward:.3}");
}
}
let mean = rewards.iter().sum::<f32>() / rewards.len() as f32;
let var = rewards.iter().map(|r| (r - mean).powi(2)).sum::<f32>() / rewards.len() as f32;
let std = var.sqrt();
println!();
println!("=== Phase E random-uniform baseline (n={}, horizon={}) ===", cli.n_episodes, cli.horizon);
println!("Mean reward: {mean:.5}");
println!("Std reward: {std:.5}");
println!("Mean + 2σ: {:.5}", mean + 2.0 * std);
println!();
println!("Write to ISV slot RANDOM_BASELINE_REWARD_MEAN_INDEX (547) at training startup.");
println!("Write to ISV slot RANDOM_BASELINE_REWARD_STD_INDEX (548) at training startup.");
// Save to JSON for the trainer to ingest.
let json = serde_json::json!({"mean": mean, "std": std, "n": cli.n_episodes});
std::fs::write("phase_e_baseline.json", serde_json::to_string_pretty(&json)?)?;
Ok(())
}
- Step 2: Build + run
SQLX_OFFLINE=true cargo build -p ml --release --example phase_e_baseline
FXC=$(ls -t /home/jgrusewski/Work/foxhunt/test_data/feature-cache/*.fxcache | head -1)
target/release/examples/phase_e_baseline \
--fxcache-path "$FXC" \
--horizon 6000 \
--n-episodes 10000
Expected: prints baseline mean and std. With uniform random across 9 actions on a market with no drift, mean reward should be approximately -2 × cost_per_contract × n_fills, dominated by fees. Std is the noise floor.
- Step 3: Commit baseline + outputs
git add crates/ml/examples/phase_e_baseline.rs phase_e_baseline.json
git commit -m "feat(phase-e): random-uniform baseline computation for kill-criteria reference"
Task 8: Reserve a smoke checkpoint for Milestone E.0
Files:
-
Modify:
MEMORY.md(the project memory directory) -
Step 1: Write Milestone E.0 close-out memory
Create /home/jgrusewski/.claude/projects/-home-jgrusewski-Work-foxhunt/memory/project_phase_e0_foundation_close.md:
---
name: project-phase-e0-foundation-close
description: Phase E.0 foundation complete — ISV slot block 539..550 reserved, registry wired with dispatch arms, fill model fitted from 5.2M trade tape, ExecutionEnv with 10-dim state and 9-action space tested, random-uniform baseline computed and saved
metadata:
type: project
---
**Status:** Phase E.0 foundation locked in commits TBD..TBD.
Numeric baseline (record actual values from Task 7 run):
- Random-uniform mean reward: <FILL>
- Random-uniform std reward: <FILL>
- Mean + 2σ: <FILL> (the kill-criterion threshold E.1 must exceed)
ISV reservation:
- Block 539..550 (12 slots, contiguous)
- 10 slots in active use, 2 reserved spare for absorbing growth without
another `ISV_TOTAL_DIM` bump
How to apply: Phase E.1 onward should consume these slots, NEVER add
new Phase E ISV indices outside this block. Bumping `ISV_TOTAL_DIM`
again should be the last resort, not the first.
Related: [[pearl-phase1d4-backtest-cost-edge-frontier]] (the original
binding constraint Phase E exists to solve), [[feedback-isv-for-adaptive-bounds]],
[[pearl-controller-anchors-isv-driven]].
- Step 2: Update MEMORY.md index
Add a line under "Active Project State" pointing to the new memory.
Milestone E.1 — Truncated H=600 DQN smoke (~Week 2, Tasks 9-14)
Goal: Train DQN on a 600-step truncated version of the task. Validate the three kill criteria pass; if so, scale to H=6000. If not, escalate to NoisyNet.
Decisive gate at end of week 2:
- Q-spread ≥ 5% of |mean(Q)| across actions (slot 539)
- Action entropy > 0.5 × ln(9) past epoch 10 (slot 540)
- Rollout reward > random baseline mean + 2σ (slot 541 > 0)
- Early-step Q-values have moved from init (slot 542 ≥ 0.01)
ALL FOUR must pass. If ANY fails, the H=6000 plan is doomed and we pivot to NoisyNet (Task 19) before sinking the budget.
Task 9: Kill-criteria producer kernel
Files:
-
Create:
crates/ml/src/cuda_pipeline/phase_e_kill_criteria.cu -
Modify:
crates/ml/build.rs(add the new kernel to the precompile list around line 106) -
Step 1: Write the kernel
/*
* Phase E.0 kill-criteria diagnostic producer kernel.
*
* Reads per-rollout Q-values + action counts + rewards and updates four
* Phase E ISV slots:
* 539 Q_SPREAD_EMA_INDEX — std(Q, axis=action) / |mean(Q)|
* 540 ACTION_ENTROPY_EMA_INDEX — H(action_dist) over recent rollouts
* 541 RETURN_VS_RANDOM_EMA_INDEX — (rollout_R − random_R) / σ_random
* 542 EARLY_Q_MOVEMENT_EMA_INDEX — |Q(s_early) − Q_init| / Q_init
*
* All four use Pearl A bootstrap + Pearl D Wiener-α steady state (see
* sp4_wiener_ema.rs for the canonical implementation).
*
* Grid: 1 thread block, 1 thread. Called once per N rollout-end events.
*/
extern "C" __global__ void phase_e_kill_criteria_update(
const float* __restrict__ q_values, // [batch, n_actions] from last forward
const int* __restrict__ action_counts, // [n_actions] from rollout
float rollout_reward_mean,
float rollout_reward_std,
float random_baseline_mean,
float random_baseline_std,
float q_init_norm,
float q_early_step_norm,
int batch,
int n_actions,
float* __restrict__ isv,
float* __restrict__ wiener_state
)
{
if (threadIdx.x != 0 || blockIdx.x != 0) return;
// (1) Q-spread: std(Q, axis=action) / |mean(Q)| averaged over batch
float q_spread_sum = 0.0f;
for (int b = 0; b < batch; b++) {
float mean = 0.0f, var = 0.0f;
for (int a = 0; a < n_actions; a++) {
mean += q_values[b * n_actions + a];
}
mean /= (float)n_actions;
for (int a = 0; a < n_actions; a++) {
float d = q_values[b * n_actions + a] - mean;
var += d * d;
}
var /= (float)n_actions;
float std = sqrtf(var);
float am = fabsf(mean) + 1e-9f;
q_spread_sum += std / am;
}
float q_spread = q_spread_sum / (float)batch;
// (2) Action entropy
int total = 0;
for (int a = 0; a < n_actions; a++) total += action_counts[a];
float entropy = 0.0f;
if (total > 0) {
for (int a = 0; a < n_actions; a++) {
if (action_counts[a] > 0) {
float p = (float)action_counts[a] / (float)total;
entropy -= p * logf(p);
}
}
}
// (3) Return-vs-random: signed σ-distance from random baseline
float return_vs_random = (rollout_reward_mean - random_baseline_mean)
/ fmaxf(random_baseline_std, 1e-6f);
// (4) Early-Q-movement: relative change from initialisation
float early_movement = fabsf(q_early_step_norm - q_init_norm)
/ fmaxf(fabsf(q_init_norm), 1e-6f);
// Pearl A + D Wiener update (see sp4_wiener_ema.rs)
// Apply to each slot independently.
// For brevity: assume `wiener_update(isv_idx, x_new, wiener_state, isv)`
// is an inline helper exposed by sp4_wiener_ema.rs that handles
// sentinel bootstrap + α update.
//
// wiener_update(539, q_spread, wiener_state, isv);
// wiener_update(540, entropy, wiener_state, isv);
// wiener_update(541, return_vs_random, wiener_state, isv);
// wiener_update(542, early_movement, wiener_state, isv);
// (Real implementation: inline the Wiener math from sp4_wiener_ema.cu.
// Skipping for plan brevity; copy from sp4_wiener_ema.cu pattern.)
}
- Step 2: Add to build.rs
In crates/ml/build.rs, find the kernels list (around line 106) and add:
"phase_e_kill_criteria.cu",
- Step 3: Build
SQLX_OFFLINE=true cargo build -p ml --release 2>&1 | grep -E '^error|phase_e_kill_criteria'
Expected: clean compile, kernel cubin produced.
- Step 4: Verify cubin output
ls target/release/build/ml-*/out/phase_e_kill_criteria.cubin 2>/dev/null | head -3
Expected: at least one matching path.
- Step 5: Commit
git add crates/ml/src/cuda_pipeline/phase_e_kill_criteria.cu crates/ml/build.rs
git commit -m "feat(phase-e): kill-criteria diagnostic producer kernel (Q-spread, entropy, return-vs-random, early-Q-movement)"
Task 10: Munchausen target term
Files:
- Modify: existing DQN target kernel in
crates/ml/src/cuda_pipeline/(find viagrep -rn 'compute_dqn_target' crates/ml/src/cuda_pipeline/)
Munchausen DQN adds one clipped-log-policy term to the target. ~30 lines.
- Step 1: Find the existing target kernel
grep -rn 'target_q\|td_target\|c51_target' crates/ml/src/cuda_pipeline/ | head -10
Expected: pointer to the existing target computation. Note the file + line range.
- Step 2: Add Munchausen term
Add a new kernel function compute_munchausen_target next to the existing target kernel. Per Vieillard et al. 2020:
extern "C" __global__ void compute_munchausen_target(
const float* __restrict__ q_next, // Q(s', ·) — target net
const float* __restrict__ q_current, // Q(s, ·) — current policy net
const int* __restrict__ actions, // [batch] action taken
const float* __restrict__ rewards, // [batch] reward
const float* __restrict__ dones, // [batch] terminal mask
float gamma,
float alpha_m, // Munchausen scale (typ. 0.9)
float tau, // entropy temperature
float log_clip_min, // typ. -1.0
float* __restrict__ target_out, // [batch]
int batch, int n_actions
) {
int i = blockIdx.x * blockDim.x + threadIdx.x;
if (i >= batch) return;
// Softmax over Q(s, ·) to get policy π(a|s)
float max_q = q_current[i * n_actions];
for (int a = 1; a < n_actions; a++) {
float qa = q_current[i * n_actions + a];
if (qa > max_q) max_q = qa;
}
float sum_exp = 0.0f;
for (int a = 0; a < n_actions; a++) {
sum_exp += expf((q_current[i * n_actions + a] - max_q) / tau);
}
float log_pi_a = (q_current[i * n_actions + actions[i]] - max_q) / tau
- logf(sum_exp);
float munchausen = alpha_m * fmaxf(tau * log_pi_a, log_clip_min);
// Standard target with Munchausen augmentation
float v_next = 0.0f;
if (dones[i] < 0.5f) {
// soft-max value of Q(s', ·) under target net
float maxn = q_next[i * n_actions];
for (int a = 1; a < n_actions; a++) {
float qa = q_next[i * n_actions + a];
if (qa > maxn) maxn = qa;
}
float sumn = 0.0f;
for (int a = 0; a < n_actions; a++) {
sumn += expf((q_next[i * n_actions + a] - maxn) / tau);
}
v_next = maxn + tau * logf(sumn);
}
target_out[i] = rewards[i] + munchausen + gamma * v_next;
}
- Step 3: Build
SQLX_OFFLINE=true cargo build -p ml --release 2>&1 | grep -E '^error' | head -5
Expected: clean compile.
- Step 4: Commit
git add crates/ml/src/cuda_pipeline/
git commit -m "feat(phase-e): Munchausen DQN target term (Vieillard et al. 2020) — implicit KL regularization between successive policies"
Task 11: Wire ISV slot consumption (n_step, γ, ε, Kelly)
Files:
- Modify:
crates/ml/src/cuda_pipeline/gpu_dqn_trainer.rs
Per the ISV memo: replace hardcoded n_step = 32, gamma = 0.999, etc. with consumption of the existing slots.
- Step 1: Find the existing hardcoded constants
grep -n 'n_step\|gamma' crates/ml/src/cuda_pipeline/gpu_dqn_trainer.rs | head -20
Locate the constants currently set to literal values. Replace with reads from ISV slots N_STEP_INDEX=517, GAMMA_*_EFF_INDEX=43-46, etc., using the standard sentinel-bootstrap pattern.
- Step 2: Update kernel signatures + reads
Where the trainer's loop currently has e.g. let n_step = 32;, replace with:
let n_step = {
let raw = read_isv_slot(crate::cuda_pipeline::N_STEP_INDEX, &isv_buf);
if raw <= PHASE_E_SENTINEL + PHASE_E_SENTINEL_EPS {
// Cold-start fallback
32
} else {
raw.clamp(1.0, 32.0) as usize
}
};
(Repeat for γ, ε, Kelly.)
- Step 3: Test
SQLX_OFFLINE=true cargo test -p ml --lib gpu_dqn_trainer -- --nocapture 2>&1 | tail -20
Expected: existing tests still pass.
- Step 4: Commit
git add crates/ml/src/cuda_pipeline/gpu_dqn_trainer.rs
git commit -m "feat(phase-e): replace hardcoded n_step/γ/ε/Kelly literals with ISV slot consumption per ISV memo"
Task 12: Author H=600 DQN smoke example
Files:
- Create:
crates/ml/examples/phase_e_smoke.rs
This is the week-1 kill-criteria test. Trains DQN at H=600 on the truncated task.
- Step 1: Write the smoke
//! Phase E.1 — H=600 DQN smoke for week-1 kill-criteria gate.
//!
//! Runs the GPU DQN trainer on a truncated version of the execution
//! task (H=600 instead of 6000). All four kill criteria are checked:
//! - Q-spread > 5% of |mean(Q)| across actions
//! - Action entropy > 0.5 × ln(9) past epoch 10
//! - Rollout reward > random baseline mean + 2σ
//! - Early-step Q-values have moved from init
//!
//! If ALL FOUR pass at the end of training, the H=6000 plan is viable.
//! If ANY fails, this is the signal to pivot to NoisyNet (Task 19).
use anyhow::{Context, Result};
use clap::Parser;
#[derive(Parser)]
struct Cli {
#[arg(long)] fxcache_path: String,
#[arg(long, default_value = "phase_e_fill_coeffs.json")] fill_coeffs_path: String,
#[arg(long, default_value_t = 600)] horizon: usize,
#[arg(long, default_value_t = 200_000)] n_train_steps: usize,
#[arg(long, default_value_t = 64)] batch_size: usize,
#[arg(long, default_value_t = 1e-3)] lr: f32,
}
fn main() -> Result<()> {
tracing_subscriber::fmt().init();
let cli = Cli::parse();
// (1) Load snapshots from fxcache + alpha logits from Phase 1d.3 stacker.
// (2) Load fitted fill coefficients.
// (3) Compute random baseline (skip if phase_e_baseline.json exists).
// (4) Build ExecutionEnv.
// (5) Construct the DQN trainer with our Phase E settings:
// - n_actions = 9
// - state_dim = 10
// - Munchausen target (Task 10)
// - ε-greedy (consumed from ISV slot 41)
// - Phase E kill-criteria producer wired
// (6) Train for cli.n_train_steps.
// (7) At each epoch boundary: read kill-criteria slots and log to stderr.
// (8) At end of training: print final kill-criteria values + PASS/FAIL verdict.
// (Implementation copies the pattern from
// `crates/ml-alpha/examples/phase1d_long_horizon.rs` for fxcache loading,
// and the DQN training loop from the existing `gpu_dqn_trainer.rs`.)
Ok(())
}
- Step 2: Build
SQLX_OFFLINE=true cargo build -p ml --release --example phase_e_smoke
Expected: clean compile.
- Step 3: Run
target/release/examples/phase_e_smoke \
--fxcache-path "$FXC" \
--horizon 600 \
--n-train-steps 200000
Expected runtime: ~30-60 min on RTX 3050. Output: per-epoch kill-criteria readings + final PASS/FAIL.
- Step 4: Interpret the verdict
If ALL FOUR criteria pass: proceed to Milestone E.2 (Task 16). If ANY fail: jump to Task 19 (NoisyNet escalation). Document the verdict and the four numeric readings.
- Step 5: Commit
git add crates/ml/examples/phase_e_smoke.rs
git commit -m "feat(phase-e): H=600 DQN smoke for week-1 kill-criteria gate"
Task 13: Scale to H=6000 (only if Task 12 passed)
Files:
-
Create:
crates/ml/examples/phase_e_full.rs -
Step 1: Lift
phase_e_smoke.rstophase_e_full.rs
Same code, but default horizon=6000 and n_train_steps=2,000,000.
- Step 2: Build + run
SQLX_OFFLINE=true cargo build -p ml --release --example phase_e_full
target/release/examples/phase_e_full --fxcache-path "$FXC" --horizon 6000
Expected runtime: 3-6 hours.
- Step 3: Verify all four kill criteria continue to pass at H=6000
If they pass: proceed to E.2. If they fail at H=6000 even though H=600 passed: Q-skew propagation is the binding issue; consider lowering γ effective or increasing n_step before escalating to NoisyNet.
- Step 4: Commit
git add crates/ml/examples/phase_e_full.rs
git commit -m "feat(phase-e): H=6000 full DQN training (E.1 final, if H=600 smoke passed)"
Tasks 14-15: Reserved for kill-criteria failure escalation paths
If Task 12 (H=600 smoke) passes, skip Task 14-15 entirely; proceed to Milestone E.2.
If Task 12 fails, Task 14 = wire NoisyLinear into the existing dead-scaffolding path (4-6 days estimated per ISV memo Question 2). Task 15 = re-run smoke with NoisyNet.
Detailed steps for Tasks 14-15 deferred until Task 12 fails. Document the failure mode (which of the four criteria failed and by how much) before designing the escalation.
Milestone E.2 — Stacker-threshold ISV controller (~Week 3, Tasks 16-22)
Goal: Replace the hardcoded 0.05 confidence threshold from Phase 1d.4 with a learned ISV slot (543). Driving signal: observed-trade-rate vs target. Plus run the alpha-trust feature ablations.
Task 16: Stacker-threshold controller kernel
Files:
- Create:
crates/ml/src/cuda_pipeline/stacker_threshold_controller.cu - Modify:
crates/ml/build.rs
Engagement-rate self-correction pattern (per pearl_engagement_rate_self_correction).
- Step 1: Write the kernel
/*
* Phase E.2 stacker-threshold controller.
*
* Drives ISV slot 543 (STACKER_THRESHOLD_INDEX) from the observed
* trade-rate EMA (545) vs the target (544). When observed > target,
* raise threshold (be pickier). When observed < target, lower threshold
* (trade more). Welford accumulator + α-floor=0.4 because the
* controller co-adapts with the policy.
*
* Also drives STACKER_KELLY_ATTENUATION_INDEX (546): a multiplier in
* [0.1, 1.0] applied to the existing KELLY_F_SMOOTH_INDEX (slot 280)
* before the Kelly cap. Drives down attenuation when realized Sharpe
* EMA is below target; drives up when realized Sharpe exceeds.
*/
extern "C" __global__ void stacker_threshold_controller_update(
float rollout_trade_count,
float rollout_total_decisions,
float rollout_realized_sharpe,
float target_trade_rate,
float target_sharpe,
float wiener_alpha_floor, // 0.4
float* __restrict__ isv,
float* __restrict__ wiener_state
)
{
if (threadIdx.x != 0 || blockIdx.x != 0) return;
// (1) Update observed trade rate EMA (slot 545)
float observed_rate = rollout_trade_count / fmaxf(rollout_total_decisions, 1.0f);
// (Apply Pearl A first-observation bootstrap + Pearl D Wiener-α with floor=0.4)
// wiener_update_floored(545, observed_rate, wiener_alpha_floor, isv, wiener_state);
// (2) Drive threshold based on rate error
float rate_err = observed_rate - target_trade_rate;
float old_threshold = isv[543];
// Update rule: threshold += k * rate_err (bidirectional ramp)
float k = 0.01f; // small step
float new_threshold = fmaxf(0.0f, fminf(0.5f, old_threshold + k * rate_err));
isv[543] = new_threshold;
// (3) Drive Kelly attenuation based on Sharpe error
float sharpe_err = rollout_realized_sharpe - target_sharpe;
float old_atten = fmaxf(isv[546], 0.1f); // floor at 0.1 per blend-vs-max-with-floor pearl
float new_atten = old_atten + 0.005f * sharpe_err;
isv[546] = fmaxf(0.1f, fminf(1.0f, new_atten));
}
- Step 2: Add to build.rs
Add "stacker_threshold_controller.cu", to the kernels list.
- Step 3: Build + commit
SQLX_OFFLINE=true cargo build -p ml --release
git add crates/ml/src/cuda_pipeline/stacker_threshold_controller.cu crates/ml/build.rs
git commit -m "feat(phase-e): stacker-threshold + Kelly-attenuation controllers (engagement-rate self-correction)"
Task 17: Wire controller into the DQN training loop
Files:
-
Modify:
crates/ml/src/cuda_pipeline/gpu_dqn_trainer.rs(rollout-end callback) -
Step 1: Find the rollout-end callback
grep -n 'fn on_rollout_end\|fn rollout_complete' crates/ml/src/cuda_pipeline/gpu_dqn_trainer.rs
- Step 2: Add controller invocation
In the rollout-end callback, launch stacker_threshold_controller_update with the observed rollout stats.
- Step 3: Initialize the anchors at training start
In the trainer's new, write:
-
isv[TRADE_RATE_TARGET_INDEX (544)] = 0.08(8% target trade rate — calibrate from backtest) -
(target_sharpe is a kernel-arg, not an ISV slot — pass as f32)
-
Step 4: Commit
git add crates/ml/src/cuda_pipeline/gpu_dqn_trainer.rs
git commit -m "feat(phase-e): wire stacker-threshold controller into rollout-end callback"
Task 18: Alpha-trust feature ablation
Files:
- Create:
crates/ml/examples/phase_e_alpha_trust_ablation.rs
Three configs:
-
Raw: state contains only
alpha_logit -
+confidence: state contains
[alpha_logit, alpha_confidence] -
+confidence+regime: state contains
[alpha_logit, alpha_confidence, spread_bps, L1_imbalance](the locked-in design) -
Step 1: Write the example
The example trains the DQN three times with three state-vector slices. Reports final reward + Sharpe for each.
- Step 2: Run
target/release/examples/phase_e_alpha_trust_ablation --fxcache-path "$FXC"
Expected runtime: ~3× full training time = ~9-18 hours. Consider running overnight.
- Step 3: Interpret + commit
The locked-in design (+confidence+regime) should beat raw alpha by ≥ 0.3 in per-trade Sharpe. If raw is competitive, simplify the locked-in state in production.
git add crates/ml/examples/phase_e_alpha_trust_ablation.rs phase_e_ablation_results.json
git commit -m "feat(phase-e): alpha-trust state ablation (raw / +conf / +conf+regime)"
Tasks 19-22: NoisyNet escalation path (CONDITIONAL on Task 12 failing)
Tasks 19-22 are only executed if the Task 12 H=600 smoke fails to pass all four kill criteria. Detailed steps deferred until that gate triggers.
If executed:
- Task 19: Wire dead NoisyLinear into cuBLAS forward path (4-6 days)
- Task 20: Re-run H=600 smoke with NoisyNet
- Task 21: Re-run kill-criteria gate
- Task 22: If still failing, escalate to RND (intrinsic reward); else proceed to E.3
Milestone E.3 — Composition + cost-sweep backtest (~Week 4, Tasks 23-28)
Goal: Compose Phase 1d.3 alpha (Mamba + stacker) with Phase E execution policy. Run the medium-tier-fill backtest with cost sweep. Compare Sharpe vs the Phase 1d.4 threshold-only baseline (frictionless 4.4 annualised, half-tick -4.0 annualised). Goal: lift the half-tick number above 0.
Task 23: Build the composition replay
Files:
- Create:
crates/ml/examples/phase_e_compose_backtest.rs
Replays the snapshot fxcache through alpha (loaded from Phase 1d.3 checkpoints) and execution (loaded from Phase E.1/E.2 checkpoints), tracks PnL with fill-simulator costs.
- Step 1: Authoring details
Lift the backtest pattern from crates/ml-alpha/examples/phase1d_long_horizon.rs (the GPU-backtest section), replacing the "always-market-when-confident" policy with the trained execution policy.
- Step 2-4: Build, run, commit
SQLX_OFFLINE=true cargo build -p ml --release --example phase_e_compose_backtest
target/release/examples/phase_e_compose_backtest --fxcache-path "$FXC"
Expected output: cost-sweep table similar to Phase 1d.4, with new "execution-policy" column reporting per-cost Sharpe.
Tasks 24-28: Hyperparameter sweeps (Kelly fraction, hard cap, fill-model robustness)
Per the locked design's "three ablations to budget for":
- Kelly fraction sweep {0.10, 0.25, 0.50}
- Hard-cap sweep {5, 10, 20} contracts
- Fill-model robustness (medium vs simple-tier baseline)
Each is a parameter override on phase_e_compose_backtest. Run, tabulate, pick the operating point. Commits per sweep dimension.
Milestone E.4 — Shadow-mode integration (~Week 5, Tasks 29-32)
Goal: Wire the trained alpha + execution policy into services/trading in shadow mode (read live ticks, emit orders to log, compare to backtest). The decisive gate is whether shadow PnL on a 2-week window matches backtest PnL within 30%.
Task 29: Inference-only ONNX-style export
Files:
- Create:
crates/ml/src/inference/phase_e_export.rs
The trained DQN must be exportable for services/trading to load without dragging in the full training stack. Two options:
- Inline cudarc inference (preferred, keeps GPU-pure)
- ONNX export + ORT runtime (not preferred, adds Python dep + ORT lib)
Implement inline cudarc inference. Trained weights are dumped to a binary file; the inference path loads + runs the forward pass with current state.
Task 30: Shadow-mode integration in services/trading
Files:
- Create:
services/trading/src/phase_e_shadow.rs
Wires up to the live MBP-10 feed, runs the alpha → execution chain on every snapshot, emits to a "would-have-traded" log instead of sending orders.
Task 31: Shadow-vs-backtest reconciliation
Files:
- Create:
crates/ml/examples/phase_e_shadow_compare.rs
Reads the shadow log + the backtest output for the same time window. Compares cumulative PnL, trade counts, action distributions. Fires alerts on divergence > 30%.
Task 32: Phase E close-out memory
Files:
- Create:
/home/jgrusewski/.claude/projects/-home-jgrusewski-Work-foxhunt/memory/project_phase_e_close.md
Final verdict: deployable Sharpe, operating point, recommended next phase.
Sign-off
After all 32 tasks complete:
- Step 1: Run full test suite
SQLX_OFFLINE=true cargo test --workspace 2>&1 | tail -20
Expected: 0 failures.
- Step 2: Push branch
git push origin sp20-aux-h-fixed
- Step 3: Update MEMORY.md with the Phase E close-out memory entry.
Self-Review
Spec coverage:
- ISV slot reservation: Tasks 1-2 ✓
- Fill model (medium-tier Poisson): Tasks 4-5 ✓
- ExecutionEnv (state, action, reward): Tasks 3, 6 ✓
- Random-uniform baseline: Task 7 ✓
- Kill-criteria instrumentation: Tasks 8, 9 ✓
- Munchausen target: Task 10 ✓
- ISV consumption (n_step, γ, ε, Kelly): Task 11 ✓
- H=600 smoke + kill-criteria gate: Task 12 ✓
- H=6000 full training: Task 13 ✓
- NoisyNet escalation (conditional): Tasks 14-15, 19-22 (deferred design)
- Stacker-threshold controller: Tasks 16-17 ✓
- Alpha-trust ablation: Task 18 ✓
- Composition backtest + Kelly/cap sweeps: Tasks 23-28 ✓
- Shadow-mode integration: Tasks 29-32 ✓
Placeholder scan:
- Tasks 14-15 and 19-22 are intentionally "deferred design" — they only execute if Task 12 fails. The plan documents this; not a placeholder.
- Task 7 (baseline) and Task 12 (smoke) reference "lift loading pattern from
phase1d_long_horizon.rs" — these are concrete references to existing code, not "TODO: implement".
Type consistency:
STATE_DIM = 10consistent across Tasks 6, 12, 18 ✓N_ACTIONS = 9consistent across Tasks 3, 7, 9, 18 ✓- ISV slot indices 539-548 consistent across Tasks 1, 2, 9, 16-17 ✓
ExecutionEnvConfigfields (horizon_snapshots,trade_size_contracts,cost_per_contract) consistent across Tasks 6, 7, 12 ✓FillCoeffs/FillFeaturesconsistent across Tasks 4, 5 ✓
Decisive gates summary (any FAIL pivots the plan):
| Milestone | Gate | What it costs to recover |
|---|---|---|
| E.0 (Task 7) | Random baseline computable | Trivial — fix env bugs |
| E.1 (Task 12) | All 4 kill criteria pass at H=600 | ~1 week to wire NoisyNet (Task 19) |
| E.1 (Task 13) | Kill criteria continue to pass at H=6000 | Increase n_step or γ; ~3 days |
| E.2 (Task 17) | Controller drives threshold within ±10% of target | Tune controller gain; ~2 days |
| E.2 (Task 18) | +confidence+regime beats raw by ≥ 0.3 Sharpe | If not, simplify state — saves complexity |
| E.3 (Task 23) | Half-tick Sharpe > 0 (lifts from Phase 1d.4 -4.0) | Re-tune Kelly fraction / threshold ramp |
| E.4 (Task 31) | Shadow PnL within 30% of backtest | Iterate fill model toward queue tracking |
End of Phase E implementation plan.