docs(phase-e): implementation plan for execution-layer RL policy
32 tasks across 5 milestones (E.0 foundation → E.4 shadow-mode), with locked design decisions from three rounds of focused research memos: - Q1 (fill sim): medium-tier Poisson regression from 5.2M trade tape - Q2 (reward): terminal-only, n-step credit (consumes ISV slot 517) - Q3 (alpha trust): implicit calibrated trust via state features - Q4 (state window): current snapshot + 2 short-horizon scalars - Q5 (sizing): hybrid decoupled fractional Kelly × Phase E attenuation Trainer choice: DQN primary (Rainbow + Munchausen target), PPO control on H=600 truncated only if kill criteria fire. Exploration: ε-greedy with kill-criteria gate at end of week 2; NoisyNet escalation (4-6 days due to dead scaffolding in our codebase) if criteria fail; RND beyond that. ISV consumption: 5 existing slots (n_step=517, γ=43-46, ε=41, Kelly=280, reward_caps=452-453); new block 539..550 reserved for Phase E (10 in active use, 2 spare). One new controller (stacker-threshold engagement-rate-self-correction at slot 543). Hardcoded by design: Kelly contract cap (Category-1 safety), kill-criteria thresholds (circuit breakers). All other knobs are ISV-driven per pearl_controller_anchors_isv_driven. Decisive gates at week 1 (H=600 kill criteria), week 4 (composition backtest Sharpe at half-tick > 0), and week 5 (shadow-vs-backtest PnL within 30%). Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
This commit is contained in:
2112
docs/superpowers/plans/2026-05-15-phase-e-execution-rl.md
Normal file
2112
docs/superpowers/plans/2026-05-15-phase-e-execution-rl.md
Normal file
File diff suppressed because it is too large
Load Diff
Reference in New Issue
Block a user