fix(iqn): hard-sync target from online at fold boundary — kills geometric Q-drift in fold 1

Add `GpuIqnHead::sync_target_from_online` (DtoD copy of online_params →
target_params) and call it from `FusedTrainingCtx::reset_for_fold`
alongside the existing IQN Adam reset.

Root cause: at fold boundary the main DQN does shrink-and-perturb +
hard target sync (gpu_dqn_trainer.rs:12495 — the latter explicitly added
to prevent fold-1 grad explosion). The IQN head was added later but
only had Polyak EMA `target_ema_update` — no hard-sync existed. Polyak
at τ=0.005/step closes ~0.5%/step, far too slow to close the
fold-boundary online↔target gap before inflated Bellman TD-error
compounds geometrically through the IQN backward pass.

Bug signature this resolves:
  F0 ep5 Q=+0.79  (healthy)
  F1 ep1 Q=+0.82  (boundary fine; c51_alpha warmup blends IQN low)
  F1 ep2 Q=+2.23  (drift starts — c51_alpha ramp + IQN online↔target gap widens)
  F1 ep3 Q=+4.05
  F1 ep4 Q=+10.66 (geometric ~2.3×/epoch)
  F1 ep5 NaN at step 5 (fp32 overflow past atom support)

Per `feedback_no_partial_refactor.md`: when a shared contract
(fold-boundary state migration) changes, every consumer must migrate
together. The main DQN sync was added explicitly — the IQN, TLOB, IQL,
and aux Adam states were added later but never extended to that
contract. This commit closes the IQN gap (primary). Aux Adam + IQN
readiness gaps follow in subsequent commits.

Touched:
- gpu_iqn_head.rs: new `sync_target_from_online` (mirrors gpu_dqn_trainer.rs:12495 exactly)
- fused_training.rs: call the new sync inside the existing IQN reset block
- docs/dqn-wire-up-audit.md: Invariant 7 entry covering Fix 1 of 3

cargo check -p ml --lib clean at 13 warnings (workspace baseline). No
fingerprint change. DtoD memcpy only — no HtoD/HtoH/DtoH paths.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
This commit is contained in:
jgrusewski
2026-04-28 19:08:51 +02:00
parent 299981f9b0
commit 7c19b59033
3 changed files with 70 additions and 0 deletions

View File

@@ -2,6 +2,39 @@
**Status:** Populated during Plan 1 Task 6 (A.5 orphan audit). Updated on every commit per Invariant 7.
Fold-boundary reset gap — IQN target hard-sync (2026-04-28, Fix 1 of 3):
the main DQN's `sync_target_from_online` (DtoD `online_params →
target_params` at `gpu_dqn_trainer.rs:12495`) is wired into
`FusedTrainingCtx::reset_for_fold` (`fused_training.rs:887`) explicitly
to prevent fold-1 grad explosion. The IQN head, added later, only
exposed `target_ema_update` (Polyak τ=0.005/step) — no hard-sync. After
fold boundary the IQN online side advances through fold N+1's first
gradient steps while the IQN target buffer still holds end-of-prior-
fold weights, so the Bellman TD error gap widens by a factor that
Polyak takes hundreds of steps to close. Empirically that drives a
geometric ~2.3×/epoch Q-drift inside fold 1
(`F0 ep5 Q=+0.79 → F1 ep1 +0.82 → ep2 +2.23 → ep3 +4.05 → ep4 +10.66
→ ep5 NaN @ step 5` on fp32 overflow past atom support). The boundary
handoff at F1 ep1 looks fine because the c51_alpha warmup ramp blends
IQN low; once c51_alpha ramps up the IQN online↔target gap dominates.
**Fix**: add `GpuIqnHead::sync_target_from_online` (DtoD copy mirroring
the main DQN's helper exactly — same `memcpy_dtod_async` pattern, same
stream, same total_params) and invoke it from
`FusedTrainingCtx::reset_for_fold` immediately before the existing IQN
Adam reset. Per `feedback_no_partial_refactor.md`: the
fold-boundary state-migration contract has multiple consumers (main
DQN trunk, IQN head, aux Adam states, IQN readiness EMAs) and
extending it must migrate every consumer in lockstep. The IQN target
sync is the primary root cause; the aux Adam state resets and IQN
readiness reset land in the next two commits.
Touched: `gpu_iqn_head.rs` (new `sync_target_from_online` method
between `reset_adam_state` and `target_ema_update`), `fused_training.rs`
(invoke the new sync inside the existing `if let Some(iqn) = …` block
in `reset_for_fold`, immediately before `iqn.reset_adam_state()`).
cargo check -p ml --lib clean at 13 warnings (workspace baseline). No
fingerprint change. DtoD only — no HtoD/HtoH/DtoH paths introduced.
MSE warmup loss target clamp to C51 atom support (2026-04-28): the
fold-boundary PopArt carry-forward + iqr floor (`d710f9d50`) reduced
fold-1 epoch-1 train loss readings 16× (318k → 19k) but ep-2/ep-3