execute_training_pipeline consumes cached_target_h_s2_ptr via .take(), so
the SP6 Pearl 5 per-branch loop (4 sequential iqn calls per step) only
worked for branch 0. Branches 1, 2, 3 saw cached_ptr=None, returned a
"non-fatal" Err logged as a warning, and silently skipped apply_iqn_
trunk_gradient — only the direction branch's IQN quantile signal ever
reached the trunk, for months.
Surfaced in T10-retry train-multi-seed-ksjcm logs as repeated:
WARN: IQN parallel branch 1 step failed (non-fatal):
WARN: IQN parallel branch 2 step failed (non-fatal):
WARN: IQN parallel branch 3 step failed (non-fatal):
cached_target_h_s2_ptr is None.
Pre-existing since SP6 Pearl 5 (P4.T3-era multi-quantile IQN per-branch
τ schedules) — masked by the non-fatal warning classification.
Likely explains:
- cql_mag persistence at SP7 smoke (cql_mag=0.07 vs cql_dir=1.0):
CQL was the only learning signal mag had because IQN-mag was dead.
- SP7 c51 1000:1 controller suppression on mag/ord/urg may stabilize
at non-collapse values once IQN signal returns to those branches.
Fix is local: re-set cached ptr inside both 4-branch loops (parallel arm
near line 2123, sequential arm near line 2308). The ptr is the same
value (target_h_s2 is per-step, not per-branch), but the .take()
contract requires set-per-call. Defensive single-shot semantic preserved
— if a future bug skips the set, IQN errors loud rather than reusing
stale ptr.
Per feedback_no_partial_refactor: SP6 changed the call pattern 1→4 per
step but didn't migrate the cache contract. Per feedback_no_hiding: the
non-fatal classification hid the structural bug; escalation to hard
error deferred to next commit after post-fix T10 validates the warnings
disappear.
Audit doc updated with Fix 34 entry.