H1 result (commit 9adbca826, smoke train-5zmkr, reverted at e8814079d):
aux_dir_acc HD[2] = 78% with H=200 — aux head LEARNS direction at
longer horizon. But policy WR stayed pinned at 50.1-50.2% — the
learned aux signal does NOT propagate to policy (per
pearl_separate_aux_trunk_when_shared_starves: aux on separate
trunk with stop-grad to policy).
H6 design (synthesized from all v5-v11 + H1 evidence):
Wire the aux head's directional probability into the policy STATE
as an input feature.
- Policy gradient flows THROUGH the feature (uses it)
- Stop-grad blocks gradient BACK (aux trunk unaffected, pearl
preserved)
- Uses existing padding slot [121..128) in STATE_DIM=128 (no
layout growth)
Why this is the structural fix:
- Aux PROVED directional signal is in the features (78% at H=200)
- Policy PROVED it can't extract direction (WR=50% across all
v5-v11 conditions)
- Bridge connects the two without violating trunk separation
- Information-theoretic: gives policy a feature it provably
can't compute itself
Test outcome interpretation:
WR > 50.5% → Mechanism 1 was binding (trunk separation gap)
WR pinned → Mechanism 2 (reward density) or Mechanism 3 (V/A
unidentifiability) dominates → H3 or V/A fix next
Files changed:
- docs/plans/2026-05-12-sp22-wr-plateau-investigation.md:
H6 added as new primary hypothesis after H1; experiment order
revised
- docs/dqn-wire-up-audit.md: H6 design entry with three-mechanism
synthesis
Implementation scope (separate commit):
1. State layout: claim slot in padding [121..128) for aux_dir_prob
2. Aux trunk export: pull "up" probability per bar from aux forward
3. Experience collection: write aux_dir_prob into per-bar state
4. Stop-grad verification: confirm policy gradient blocked
5. Trade-open persistence: latch aux_dir_prob for trade duration
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>