Resolves Task 13 — the synthetic-overfit divergence I thought was a
wiring bug was actually init-sensitivity on the n_hid=32 toy. With
seed=0x4242 + lr=3e-2 + constant +1 direction + 200 steps + reset_
hidden_state per sample, the trainer converges loss 0.5669 -> 0.0665
(88% drop, well under the 60% gate threshold).
The 200-step weight trajectory (debug_long_horizon_weight_trajectory)
shows monotone descent:
step 0: loss=0.6932 hb[0]=0.030 hw[0,0]=-0.124
step 50: loss=0.2332 hb[0]=1.164 hw[0,0]= 0.997
step 100: loss=0.1301 hb[0]=1.599 hw[0,0]= 1.412
step 190: loss=0.0747 hb[0]=1.999 hw[0,0]= 1.777
Heads weights drive monotonically into the correct sigmoid tail.
The chain is sound:
- heads_backward finite-diff at 1% relative
- cfc_step_backward finite-diff at 5% relative
- BCE forward+backward at 5% relative finite-diff
- AdamW invariants (zero-grad + wd, descent on g=theta)
- Graph A capture bit-identical to sequential
- end-to-end overfit on constant +1 = 88% loss drop in 200 steps
Removed the #[ignore] + the speculative "wiring bug" doc comment.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>