plan(dqn-v2): Plan 3 — behavioural core implementation plan
Third of five sequential plans decomposing the DQN v2 unified spec
(docs/superpowers/specs/2026-04-24-dqn-v2-unified-design.md).
Covers spec sections:
- §4.C.2 Reward-component attribution (6 ISV slots [52..58), reward_split[..]
in HEALTH_DIAG) — lands first as diagnostic substrate for the rest
- §4.B.1 ISV-driven continuous Flat opp-cost (scales by isv[Q_DIR_ABS_REF])
- §4.B.2 Novelty-weighted trade-attempt bonus (2 ISV slots [50, 51])
- §4.B.4 Adaptive plan-MLP activation threshold (1 ISV slot [49] +
PlanMlpThresholdController via the AdaptiveController trait)
- §4.C.4 Temporal timing bonus (trade-exit, bars_early / hold_time)
- §4.D.4 Generalised temporal reward coupling (3 ISV slots [59..62):
persistence credit, regime-shift penalty, conviction consistency)
- §4.C.3 State-distribution divergence signal (3 ISV slots [58, 63, 64])
- §4.B.3 Replay buffer seeded warm-start (4 scripted policies,
≥100K experiences, CQL α decoupled from seed phase)
- §4.C.5 CQL alpha schedule coupled to seed-phase decay via
CqlAlphaSeedCoupledController
Plan structure: 9 implementation tasks + pre-plan verification + exit-gate
validation task (Task 10). Each task TDD-disciplined with bite-sized
steps. Task 1 (C.2 diagnostic audit) ordered first so subsequent
reward-shape tasks can verify contributions rather than log-archaeology.
Plan 3 exit gate: multi-seed (N=5) × multi-fold (K=6) run passes Tier 1
(convergence) + Tier 2 behavioural subset (val trades/bar ≥ 0.005,
val_active_frac > 0.2). Tier 3 profitability deferred to Plan 5.
Preserves all 9 invariants. Zero stubs. Zero TODO/FIXME. Every new module
wired to its consumer in the same task it lands.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>