fix(popart): carry fold stats forward + raise iqr floor — kills fold-1 loss explosion
L40S 15-epoch repro #2 (train-multi-seed-kdkdv) revealed train_loss exploded 1000-2000× in fold 1 (3.1 → 318k-694k) while val Sharpe stayed healthy and CLIMBING (73 → 114). Train/val disconnect = pure measurement bug, not real instability. Mechanism: reset_for_fold() zeroed cached_iqr/cached_median at every fold boundary. The `if cached_iqr > 0.0` guard at fused_training.rs:1220 then forced fold N+1's first epoch to the Welford GPU path. Welford running stats inherited from fold N, combined with fold N+1's slightly different reward distribution post-S&P + adversarial regime, produced normalized rewards far outside C51 atom support [-50, 50] — categorical loss readings 10^5x inflated. On rare timing-sensitive paths the near-zero divide overflowed to Inf → NaN (the run-1 flagged=[2=on_b_logits, 3=mse_loss, 6=grad_buf, 7=save_current_lp, 8=save_projected] diagnostic — what we caught was downstream of THIS root cause). The diagnostic infrastructure from756b1ef31+32e5375acworked perfectly: its precise signal of "on_v_logits clean, on_b_logits NaN, but loss is also NaN, params clean pre-forward" surfaced the train/val disconnect that pointed at the reward normalization bug. Fix A (carry-forward, fused_training.rs:919-922): Stop resetting cached_iqr / cached_median at fold boundary. Carry fold N's final-epoch median/IQR forward as fold N+1's epoch-1 default. Same instrument, similar reward distribution between adjacent walk-forward folds, so the carry is safe and gets replaced by fresh stats at the end of fold N+1's epoch 1. prev_popart_var still resets (sole consumer is tau-change detection; a fresh fold counts as a change point regardless). Fix B (permanent floor, dqn_utility_kernels.cu:1657): Raise iqr fmaxf floor 1e-6 → 1e-4. With iqr=1e-6 a trade-exit reward of 5.0 normalizes to 5e6 (vs post-fix 5e4) — much harder to hit fp32 overflow. Defensive bound for genuinely-pathological iqr paths (e.g. genuinely degenerate data quantiles), not a tuned knob — Invariant 1 carve-out for numerical-stability bounds. Per pearl_blend_formulas_must_have_permanent_floor.md (the same recipe that resolved Kelly cap warmup + var_scale collapse before). Resolves task #84 ("Fold-boundary state reset gap causes fold 1 grad explosion").
This commit is contained in:
@@ -2,6 +2,50 @@
|
||||
|
||||
**Status:** Populated during Plan 1 Task 6 (A.5 orphan audit). Updated on every commit per Invariant 7.
|
||||
|
||||
Fold-boundary PopArt carry-forward + iqr permanent floor (2026-04-28):
|
||||
the L40S 15-epoch repro #2 (`train-multi-seed-kdkdv`, commit
|
||||
`32e5375ac`) showed a **massive train/val disconnect** in fold 1:
|
||||
mean_loss climbed from 3.1 (fold 0 healthy) to 318,898 → 590,136 →
|
||||
694,323 across fold-1 epochs while val Sharpe stayed healthy and
|
||||
*climbing* 73 → 114. The 10⁵× loss readings turned out to be a
|
||||
fold-boundary PopArt reset bug, not a real instability. Mechanism
|
||||
traced via the diagnostic data (the same diagnostic infrastructure
|
||||
landed in `756b1ef31` + `32e5375ac` worked perfectly to surface the
|
||||
real signal): `reset_for_fold` zeroed `cached_iqr` / `cached_median`
|
||||
at every fold boundary, forcing fold N+1's first epoch to fall
|
||||
through the `cached_iqr > 0.0` guard at fused_training.rs:1220 to
|
||||
the Welford GPU path; combined with the Welford running stats
|
||||
inherited from fold N — and fold N+1's slightly different reward
|
||||
distribution post-S&P + adversarial — the normalised rewards landed
|
||||
far outside the C51 atom support [-50, 50], producing astronomical
|
||||
categorical loss readings. The val path doesn't use PopArt so val
|
||||
Sharpe was unaffected. On rare timing-sensitive paths the
|
||||
near-zero-divide overflowed to Inf → NaN (the run-1 `flagged=[2,3,
|
||||
6,7,8]` diagnostic firing). **Fix A**: stop resetting
|
||||
`cached_iqr` / `cached_median` at fold boundary in
|
||||
`FusedTrainingCtx::reset_for_fold` (fused_training.rs:919) — carry
|
||||
fold N's final-epoch median/IQR forward as fold N+1's epoch-1
|
||||
default; same instrument, similar reward distribution between
|
||||
adjacent walk-forward folds, so the carry is safe and gets replaced
|
||||
by fresh stats at the end of fold N+1's epoch 1. `prev_popart_var`
|
||||
still resets (its sole consumer is tau-change detection — a fresh
|
||||
fold counts as a "change point" anyway). **Fix B**: raise the iqr
|
||||
permanent floor in `popart_normalize_robust`
|
||||
(`dqn_utility_kernels.cu:1657`) from 1e-6 → 1e-4. The previous
|
||||
1e-6 was insufficient defence: a trade-exit reward of 5.0 with
|
||||
iqr=1e-6 normalises to 5e6, vs the post-fix 5e4 (still very large
|
||||
but much harder to hit fp32 overflow). Per
|
||||
`pearl_blend_formulas_must_have_permanent_floor.md` and
|
||||
`feedback_isv_for_adaptive_bounds.md` (numerical-stability bound
|
||||
carve-out under Invariant 1). Touched files: `fused_training.rs`
|
||||
(remove 2 of 3 lines from the `reset_for_fold` PopArt block, expand
|
||||
docstring); `dqn_utility_kernels.cu` (raise the iqr fmaxf floor,
|
||||
expand kernel comment). cargo check clean at 13 warnings (workspace
|
||||
baseline). No fingerprint change. Resolves task #84 ("Fold-boundary
|
||||
state reset gap causes fold 1 grad explosion"). Diagnostic kernels
|
||||
from `756b1ef31` + `32e5375ac` remain in place to catch any future
|
||||
NaN paths through the loss-component buffers.
|
||||
|
||||
NaN diagnostic flag 12 = `save_h_s2` (2026-04-28): the previous wire-up
|
||||
commit (`756b1ef31`) extended NaN coverage to flags 0-11 and proved that
|
||||
the L40S fold-1 NaN entered via `flagged=[2=on_b_logits, 3=mse_loss_scalar,
|
||||
|
||||
Reference in New Issue
Block a user