run_nan_checks_post_backward now fires a single fused kernel launch
covering slots 24-35 in 12 blocks. Reduces graph-capture overhead from
8 launches to 1 per step. Slots 33-35 inline checks at backward
orchestration sites (gpu_dqn_trainer.rs:18580/:18601/:6939) unchanged
— multi-point localization preserved.
Method signature simplified — IQN pointers no longer per-step args.
populate_nan_check_meta (called once at construction in fused_training.rs)
baked Option<u64> nullity into the metadata buffer entries; the fused
kernel's null-pointer guard handles slots 27/28 when IQN inactive
without per-step branching at the Rust caller.
Both call sites in fused_training.rs (ungraphed + graph-captured paths)
updated together per feedback_no_partial_refactor.
This is the F0 regression fix — Phase B's per-step kernel-launch
overhead was the F0 cause across 3 SP1 smokes (F0 ~ 35 vs baseline
55.87). Gate 1 smoke (Task A6) validates F0 >= 53.08.
Replaces A2's CudaSlice<u64>/<i32> field types with MappedU64Buffer/
MappedI32Buffer per feedback_no_htod_htoh_only_mapped_pinned. Mapped-
pinned eliminates the HtoD copy entirely — the kernel reads via the
device-mapped pointer (cuMemHostGetDevicePointer_v2) while the trainer
writes through the same mapped pages on the host side.
Adds populate_nan_check_meta on GpuDqnTrainer (one-shot construction-
time write of 12 (ptr, len) tuples for slots 24-35). Slot 31 deferred
(null entry); slots 27/28 nullable on Option<u64> (None when IQN
inactive); slots 33-35 null (inline checks fire separately at backward
orchestration phases — kept individual for entry-point localization).
Adds launch_nan_check_fused_f32 (per-step kernel launch wrapper with
grid_dim=12, block_dim=256, base_flag_idx=24). Registers
dqn_nan_check_fused_f32_kernel in compile_training_kernels (tuple
43→44, info log 38→39 utility kernels) — same module as the per-buffer
dqn_nan_check_f32 to share the captured replay group.
Constructor-time wire-up lands in FusedTrainingCtx::new after gpu_iqn
construction (gpu_iqn is owned by FusedTrainingCtx, not GpuDqnTrainer
— mirrors the same Option<u64> arg pattern used by
apply_iqn_trunk_gradient and run_nan_checks_post_backward).
Wrapper unused yet — call-site replacement (8 individual check_nan_f32
calls in run_nan_checks_post_backward → single fused launch) lands in
A4.
Audit doc updated (Invariant 7).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Adds 12-entry u64 ptr table + 12-entry i32 len table on GpuDqnTrainer
for the fused NaN-check kernel. Allocated as zeros; populated once
in next commit (slot pointers stable across trainer lifetime).
Slot 31 (deferred ensemble) entry will be 0 — kernel skips that block.
Two separate buffers chosen over packed-stride layout for direct
ABI match with the kernel signature (const float* const* buf_ptrs,
const int* buf_lens). Avoids stride-alignment concerns.
Unused yet — populate + wrapper + call-site replacement in next commits.
SP1 (F1 NaN root-cause investigation) closes at commit ab2133463 on
plan-c-phase-2-thompson with the surgical-fix-scope deliverables. Three
L40S validation smokes (smoke-test-xvzgk, dr2bn, ckgv8) characterize
both what SP1 CAN address (cold-start clamp pathology) and what it
CANNOT (structural Adam weight pathology — SP3 scope).
What SP1 produced
-----------------
1. Permanent diagnostic infrastructure verified working:
- nan_flags_buf 24→48 with backward-path coverage at slots 24-35.
- run_nan_checks_post_backward + 3 inline checks during backward
orchestration (slots 33/34/35 multi-point bw_d_h_s2 snapshots).
- Sentinel correctly fires on cuBLAS GEMM accumulator overflow.
2. F1 NaN entry kernels pinpointed (slots 26 + 32):
- apply_iqn_trunk_gradient cuBLAS sgemm (slot 26 = iqn_trunk_m)
- Bottleneck Linear backward sgemm (slot 32 = bn_d_concat_buf)
- IQN-internal buffers (slots 27, 28, 35) stayed CLEAN — the IQN
backward chain is NOT the seed; the cuBLAS consumer of clean
inputs produces NaN via accumulator overflow.
3. Cold-start clamp pathology fixed (ε floor formula):
- Original `(1e6 × isv).max(1e3)` clipped F1-startup gradients to
±1e3 when fold-boundary ISV reset put H_S2_RMS_EMA ≈ 0.
- New `1e6 × isv.max(1.0)` guarantees max_abs ≥ 1e6 in all states.
- Validated: F1 NaN moved from step 240 → 3540 (15× later).
What is SP3 territory (NOT fixed in SP1)
-----------------------------------------
F1 NaN at step 3540 has same `[6, 12, 26, 32]` signature with clean
inputs (slots 27, 35). The cuBLAS GEMM produces NaN/Inf because the
WEIGHTS being multiplied have accumulated NaN/Inf via Adam updates
over 3540 training steps. Surgical clamps on gradient inputs cannot
prevent this; root cause is Q-target inflation drives Adam EMA
saturation drives weight pathology drives GEMM overflow.
SP3 scope per spec:
- Target-Q clipping or pessimistic ensemble
- Atom-position growth bounds (C51)
- Conservative target sync at fold boundary
- Adam EMA reset frequency for fold transitions
What is SP2 territory (NOT fixed in SP1)
-----------------------------------------
F0 regression: 35.30 (Phase B+C+ε-fix) vs 55.87 baseline. Three
consecutive smokes show consistent F0 ≈ 35 across post-Phase-B
commits, indicating Phase B instrumentation has real F0 cost.
SP2 framework codification opportunity: consolidate the 11
per-step check_nan_f32 launches into a single fused reduce
kernel scanning all 12 backward-path slots in parallel.
Expected: F0 returns to ~55 baseline, regression sentinel preserved.
Pearls for memory
-----------------
1. ε-floor placement: when an ISV-driven adaptive bound has a cold-start
safety floor, the ε goes on the ISV MULTIPLIER (`isv.max(1.0)`),
not the BOUND (`bound.max(1e3)`). Putting ε on the bound clips
legitimate signal at cold-start; putting it on the multiplier
preserves the wide guard-band intent.
2. NaN diagnostic ordering: when adding sanitization (clamp,
isfinite-or-zero) to a buffer that's also NaN-checked by a
post-backward sentinel, the inline NaN check MUST fire BEFORE
the sanitization. Otherwise the sentinel sees zero'd buffers
and silently dies.
References
----------
- Spec: docs/superpowers/specs/2026-04-29-numerical-stability-investigation-design.md
- Plan: docs/superpowers/plans/2026-04-29-numerical-stability-sp1-f1-nan-root-cause.md
- Audit (durable for SP2/SP3): docs/dqn-backward-nan-audit.md
- Memory entry: ~/.claude/projects/-home-jgrusewski-Work-foxhunt/memory/project_sp1_f1_nan_root_cause_resolved.md
Smoke smoke-test-dr2bn (commit 19b008e1c) F1-NaN'd at step 240 — earlier
than pre-fix smoke smoke-test-xvzgk (step 890). The fix made things
worse, indicating the ε floor `(1e6 × isv).max(1e3)` is actively
destabilizing F1 startup.
Diagnosis: at fold boundary, ISV[H_S2_RMS_EMA_INDEX=96] and
ISV[Q_DIR_ABS_REF_INDEX=21] reset to 0 (per StateResetRegistry).
Formula `(1e6 × 0).max(1e3) = 1e3` makes max_abs aggressively narrow.
F1 startup gradients can have natural magnitudes > 1e3 (post-fold
Bellman-target shift); clamping them to ±1e3 destabilizes Adam EMAs,
which then drive cuBLAS GEMM accumulators into pathological inputs
that overflow → slot 26 + 32 NaN.
New formula: `1e6 × isv.max(1.0)` guarantees max_abs ≥ 1e6 regardless
of ISV state:
- ISV = 0 → max_abs = 1e6 × max(0, 1.0) = 1e6
- ISV = 0.5 → max_abs = 1e6 × max(0.5, 1.0) = 1e6
- ISV = 2.0 → max_abs = 1e6 × max(2.0, 1.0) = 2e6
- ISV = 100 → max_abs = 1e8
F0 no-op intent preserved (F0 inputs ≪ 1e6 in all states; ISV[96]≈1.0
for converged F0 → max_abs = 1e6, well above F0-typical |iqn_d_h_s2|
≤ ~10²). F1 startup gradients ≤ 1e6 are un-clipped.
The 1.0 ε floor is on the ISV multiplier (Invariant 1 carve-out per
`feedback_isv_for_adaptive_bounds`), not on the bound itself — bound
is still ISV-driven when ISV is meaningful (≥ 1.0).
Two edit sites:
- gpu_dqn_trainer.rs:6967 (apply_iqn_trunk_gradient, slot 26 path)
- gpu_dqn_trainer.rs:18631 (launch_cublas_backward_to, slot 32 path)
F0 regression to 35.24 still unsolved — separate investigation thread
within SP1 (no deferral; Phase B instrumentation timing impact suspected).
Smoke smoke-test-dr2bn (commit 19b008e1c) failed all 7 SP1 pass criteria.
F1 first-fire signature at step 240: flagged=[6=grad_buf, 12=save_h_s2,
26=iqn_trunk_m, 32=bn_d_concat_buf] — IDENTICAL to pre-fix smoke-xvzgk
but at step 240 vs 890. The surgical fix at 19b008e1c made F1 NaN
EARLIER, not prevented it.
Diagnosis: the ε floor `(1e6 * isv).max(1e3)` clips legitimate F1
startup gradients to ±1e3 when fold-boundary ISV reset puts ISV[96]
or ISV[21] near 0. Clipped gradients destabilize Adam EMAs → cuBLAS
GEMM accumulator overflow → slot 26 + 32 NaN.
F0 regression unchanged across Phase C (35.24 vs pre-fix 34.55).
Confirms the F0 regression is a Phase B instrumentation side effect,
not Phase C. Separate investigation thread within SP1.
Next iteration: Task 6 fix-up #2 changes ε floor to `(1e6 * isv.max(1.0))`
guaranteeing max_abs ≥ 1e6 regardless of ISV state. Removes cold-start
clamp pathology while preserving F0 paper-review intent.
Quality-review follow-ups to commit 97f1d25f5:
1. CRITICAL: preserve regression sentinel for slots 24, 25, 26, 32.
The post-clamps in apply_iqn_trunk_gradient + launch_cublas_backward_to
ran BEFORE run_nan_checks_post_backward, silently zeroing the buffers
the diagnostic was supposed to observe. Fix: inline check_nan_f32 calls
IMMEDIATELY BEFORE each launch_clamp_finite_f32 — same pattern as the
existing slot 35 check at line 6949. Now the diagnostic fires on NaN
entry; the clamp then sanitizes (preventing propagation but preserving
the flag — flags are sticky/OR'd in dqn_nan_check_f32, never cleared).
2. CRITICAL: correct F0 paper-review line reference + reasoning. The
commit body cited "5.0 norm-clip at line 16747" — that's a memset, not
a clip. Real norm-clip lives at lines 16827-16868. Also tightened the
L2-vs-per-element reasoning: L2 norm <= 5.0 worst-case bounds per-element
|x| <= 5.0 (single-element edge), still several orders of magnitude
below the 1e6 max_abs guard.
3. Inline comment at line 6951 mislabeled the clamp target as 'slot 27
input' — actually clamps the slot 35 buffer (bw_d_h_s2 post-DtoD). Fixed.
4. Articulated pre-clamp rationale: defense-in-depth regression protection
for input buffers (slots 24, 25, 35) against future pathologies that
could create extreme-but-finite inputs. The F1 ep2 NaN was outputs
overflowing finite inputs, but the same fix template covers both
failure modes per feedback_no_partial_refactor.
5. Renamed clamp_finite_f32_kernel -> dqn_clamp_finite_f32_kernel for
consistency with sibling utility kernels (dqn_nan_check_f32,
dqn_zero_kernel, dqn_grad_norm_kernel).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Patches two cuBLAS GEMM backward operations identified by SP1 Phase B
smoke smoke-test-xvzgk (commit f139a63ee) as the F1 ep2 NaN sources:
1. apply_iqn_trunk_gradient (gpu_dqn_trainer.rs:6885+):
- Pre-call: clamp bw_d_h_s2 (the IQN DtoD-overwrite target that feeds
the GRN trunk encoder's cuBLAS Linear_a/Linear_b dW/dX/dB GEMMs;
symmetric with slot 27 source-side check) to ±1e6×ISV[96].
- Post-call: sanitize iqn_trunk_m (slot 26 output) — isfinite-or-zero
+ magnitude clamp; defence-in-depth.
2. launch_cublas_backward_to main backward path:
- Pre-call: clamp d_value_logits_buf + d_adv_logits_buf (slots 24/25
inputs that feed bw.backward_full's dueling/branch backward GEMMs)
to ±1e6×ISV[21].
- Post-call: sanitize bn_d_concat_buf (slot 32 output, gated on
bottleneck_dim > 0).
Both fixes use the new clamp_finite_f32_kernel utility (CUDA — replaces
NaN/Inf with 0 + clamps finite values to ±max_abs) + launch_clamp_finite_f32
(Rust wrapper). Kernel lives in dqn_utility_kernels.cu (already in
build.rs); loaded into the same module as the NaN check kernels in
compile_training_kernels — both call sites land inside captured children
(forward_child for slot 32, aux_child for slot 26) and use only stream-
bound launch_builder, so the kernel is graph-safe.
ISV-driven max_abs bound = 1e6 × isv[<slot>] with 1e3 ε floor for
uninitialized state (Invariant 1 carve-out per
feedback_isv_for_adaptive_bounds). Wide guard band — F0 inputs sit
several orders of magnitude below the threshold (F0 ISV[96] ≈ 1.0
LayerNorm RMS, F0 ISV[21] ≈ 1.0 Q-value scale; F0 |iqn_d_h_s2| ≤ ~10²,
F0 |d_value_logits| ≤ ~10³ post the existing 5.0 norm-clip at
line 16747), so guards are no-ops on F0; F1 overflows trigger clamp.
F0 paper-review pre-smoke confirmed.
Combined RELATED commit per feedback_no_partial_refactor — both kernels
share the same unsafe-pattern (cuBLAS GEMM extreme-intermediate-product
overflow) + the same fix template, so they ship together.
Phase B slots 24-35 remain as permanent diagnostic infrastructure;
will catch any future regression of this NaN class.
SP1 Phase D validation smoke pending.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Wires per-step NaN checks on backward-path kernel outputs. Coverage
per audit per-slot accessor table (docs/dqn-backward-nan-audit.md
:530-548).
GpuDqnTrainer::run_nan_checks_post_backward (NEW) — fire-once-at-end:
- 24 d_value_logits_buf (post-c51_grad value gradient)
- 25 d_adv_logits_buf (post-c51_grad branch advantage)
- 26 iqn_trunk_m (apply_iqn_trunk_gradient cuBLAS bwd output)
- 27 iqn_d_h_s2_buf (IQN backward dh_s2; arg from FusedTrainingCtx)
- 28 d_branch_logits_buf (IQN production backward; arg from caller)
- 29 cql_d_value_logits (CQL gradient output)
- 30 aux_dh_s2_nb_buf (Aux next-bar backward dh_s2)
- 32 bn_d_concat_buf (Bottleneck Linear backward dy)
Inline checks during backward orchestration:
- 33 bw_d_h_s2 post-main (in launch_cublas_backward_to, after
backward_full + branch concat accum,
BEFORE aux_heads_backward SAXPY)
- 34 bw_d_h_s2 post-aux (in launch_cublas_backward_to, after
aux_heads_backward SAXPY)
- 35 bw_d_h_s2 post-iqn (in apply_iqn_trunk_gradient, after
graph_safe_copy_f32 DtoD overwrite,
BEFORE encoder_backward_chain consumes)
Sequential 33→34→35 fire pattern localises the NaN entry point:
- 33 alone fires → main backward chain (c51 + MSE + branch concat)
- 34 fires after 33-clean → aux SAXPY (aux_heads_backward)
- 35 fires after 34-clean → IQN DtoD or per-sample IQN backward
Slot 31 (ensemble_d_logits_buf) cleanly deferred per Task 2 commit
387335e2b — owner is FusedDqnTraining (different struct); will be
un-deferred when ensemble Phase B saxpy guards are verified.
IQN pointers (slots 27, 28) passed as Option<u64> from FusedTrainingCtx
because GpuDqnTrainer does NOT own GpuIqnHead (Task 3 deviation
finding, audit lines 535-536). None case (when iqn_lambda == 0.0
and gpu_iqn = None) cleanly skips slots 27/28 — semantically honest
"buffer doesn't exist this run" rather than false-clean signal.
Pattern matches apply_iqn_trunk_gradient(iqn_d_h_s2_ptr: u64, ...)
at gpu_dqn_trainer.rs:6843.
Call sites (atomic — feedback_no_partial_refactor):
- fused_training.rs:1519 ungraphed step path
- fused_training.rs:2324 capture_training_graph closure (post_aux child)
- gpu_dqn_trainer.rs::launch_cublas_backward_to (slots 33, 34 — captured
in forward child via submit_forward_ops_main)
- gpu_dqn_trainer.rs::apply_iqn_trunk_gradient (slot 35 — captured in
aux child via submit_aux_ops)
Each check uses existing check_nan_f32(buf_ptr, len, flag_idx) —
single-block GPU reduce, no atomicAdd, no DtoD/HtoD/HtoH copies, no
per-step DtoH. Lengths use CudaSlice::len() where possible (auto-syncs
with allocator padding); inline arithmetic for slot 28 since
d_branch_logits_buf.len() is private to GpuIqnHead. Flags accumulate
within fold; reset at fold boundary via reset_nan_flags(). Readback
flow (commit d1808df14) consumes them in BOTH halt_nan and
halt_grad_collapse paths via name tables annotated in commit 387335e2b.
Permanent diagnostic infrastructure — stays as production-grade
regression sentinel after the surgical fix lands.
Adds 5 new pub(crate) accessor methods exposing backward-path buffer
device pointers for Task 4's per-step NaN checks (slots 24, 25, 28,
29, 30 per audit per-slot table at docs/dqn-backward-nan-audit.md
:530-548):
GpuDqnTrainer (4 new):
- d_value_logits_buf_ptr (slot 24 — post-c51_grad value gradient)
- d_adv_logits_buf_ptr (slot 25 — post-c51_grad branch advantage)
- cql_d_value_logits_ptr (slot 29 — CQL gradient output)
- aux_dh_s2_nb_buf_ptr (slot 30 — aux next-bar backward dh_s2)
GpuIqnHead (1 new):
- d_branch_logits_buf_ptr (slot 28 — production IQN backward output,
iqn_quantile_huber_loss)
Slot 27 (iqn_d_h_s2_buf) reuses existing GpuIqnHead::d_h_s2_raw_ptr()
at gpu_iqn_head.rs:1660 — no new method per feedback_no_legacy_aliases:
the existing accessor is already public and sufficient; renaming +
chasing the single call site adds churn without value.
Slots 26, 32, 33-35 reuse pre-existing handles:
- 26: self.ptrs.iqn_trunk_m
- 32: self.bn_d_concat_buf() (existing, returns &CudaSlice<f32>)
- 33-35: self.ptrs.bw_d_h_s2 (3 different Task 4 call sites)
Slot 31 (ensemble_d_logits_buf) deferred per Task 2 commit 387335e2b
(cross-struct on FusedDqnTraining).
DEVIATION FROM PLAN: the plan called for "delegate accessors on
GpuDqnTrainer for slots 27/28" — structurally invalid because
GpuDqnTrainer does NOT own GpuIqnHead. The IQN head is owned by
FusedTrainingCtx (fused_training.rs:289) alongside the trainer at
line 234. The audit's per-slot accessor table (lines 535-536) is
correct: accessors land on GpuIqnHead. Task 4's
run_nan_checks_post_backward will receive IQN pointers as u64
arguments from the FusedTrainingCtx call site — same pattern already
in use at gpu_dqn_trainer.rs:6843
(apply_iqn_trunk_gradient(&mut self, iqn_d_h_s2_ptr: u64, ...)).
NO new scratch buffer added — the plan's bw_d_h_s2_pre_saxpy scratch
+ DtoD-copy approach was superseded by the audit's 3-call-site
reformulation. Slots 33/34/35 are post-main / post-aux / post-iqn
snapshots of the same bw_d_h_s2 (one buffer, three Task 4 invocations).
Pattern follows commit e9096c7be's GRN-block accessors (concise
pub(crate) fn name_ptr(&self) -> u64 with doc-comment referencing
slot number + audit doc + buffer semantics). Additive — no behavioral
change; new accessors consumed by Task 4's NaN check call sites.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Quality-review follow-ups to commit 53bc0bc50 (nan_flags_buf 24->48):
1. Update run_nan_checks_post_forward docstring (gpu_dqn_trainer.rs:14842-14873):
replace 24-slot map with 48-slot range summary + audit-doc pointer.
Was actively misleading after the 24->48 expansion; partial-refactor
residue per feedback_no_partial_refactor.
2. Update '[24] system' comment in training_loop.rs (around line 2044):
reflect the 48-slot post-expansion state (slots 0-23 fwd, 24-35 bwd,
36-47 reserved). Also fix stale '0..11' tracing message to '0..47'.
3. Slot 31 (ensemble_d_logits_buf) annotation: flag DEFERRED + owner
on FusedDqnTraining (different struct than 24-30, 32-35). Prevents
Task 4 from blanket-launching check_nan_f32 on slot 31's null
accessor.
4. Both name-table header comments now reference the future
run_nan_checks_post_backward method (Task 4) plus the audit's
per-slot table — pre-empts contract drift when Task 4 lands a
3-way name-table dependency.
Audit-doc entry appended to docs/dqn-wire-up-audit.md SP1 Phase B
section. No behavioral change. Both name tables remain byte-identical.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Expands the NaN flag buffer from 24 to 48 slots to make room for
backward-path NaN checks (slots 24-35 per audit doc per-slot table)
plus 12 reserved headroom slots (36-47) for SP2 framework + SP3
observer hooks.
Touches:
- gpu_dqn_trainer.rs: alloc size 24→48 (line ~11178); read_nan_flags
signature [i32; 24]→[i32; 48] (line ~14982); field docstring updated
(line ~2658) to reflect 48-slot layout
- fused_training.rs: pub(crate) read_nan_flags signature [i32; 24]→[i32; 48]
- training_loop.rs: BOTH name table sites (halt_nan + halt_grad_collapse
block from commit d1808df14) updated to 48 entries
Slot names use audit allocation (docs/dqn-backward-nan-audit.md per-slot
table), which supersedes the plan's placeholder names per the audit's
plan-supersession note. Slots 24-35 cover production buffers:
d_value_logits_buf, d_adv_logits_buf, iqn_trunk_m, iqn_d_h_s2_ptr,
d_branch_logits_buf, cql_d_value_logits, aux_dh_s2_nb_buf,
ensemble_d_logits_buf, bn_d_concat_buf, bw_d_h_s2 (×3 call sites).
Slots 36-47 are rsv36-rsv47 (headroom).
No behavioral change — new slots stay at zero until Task 4 wires the
kernel-output NaN checks. Buffer size reviewable by SP2.
Three review-driven fixes:
1. Slot 24/25 cite the actual production buffers d_value_logits_buf
(gpu_dqn_trainer.rs:3123) and d_adv_logits_buf (:3125) — the f32
atomicAdd buffers that launch_c51_grad writes — not the staging
buffers at :3127/:3129. Pointer expressions exist at the launch
site (:17707/:17708), making new accessors optional. Task 3 summary
and priority-list entries updated to match.
2. Explicit plan-supersession note inserted immediately above the
per-slot table: the audit's per-slot allocation supersedes the plan
Task 2 step 6 name table. The plan's allocation was placeholder;
the audit's is grounded in per-kernel inspection. Task 2 should use
the audit's names (d_value_logits_buf, d_adv_logits_buf, iqn_trunk_m,
iqn_d_h_s2_ptr, d_branch_logits_buf, etc.).
3. Summary suspicion-ranking row #3 reconciled with slot-28 dormancy:
iqn_backward_per_sample has no Rust caller (verified via
grep crates/ml/src/), so row #3 is re-pointed at the production
kernel iqn_quantile_huber_loss (iqn_dual_head_kernel.cu:1346-1413,
loaded at gpu_iqn_head.rs:2114). Same unsafe-write pattern (no
isfinite guard at line 1410's d_q_online[idx] = qw*d_huber/Q),
now attributed to the live path. apply_iqn_trunk_gradient at #1
stands — its reasoning (orchestrator consuming iqn_d_h_s2_ptr) is
unchanged by which specific kernel writes that buffer.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Quality-review fixes for the SP1 Phase A audit:
1. nan_flags_buf [i32;24] declaration: quad-cite (field 2659, alloc
11178, constructor 12563, read_nan_flags signature 14982). Task 2's
24->48 expansion must touch all four atomically per
feedback_no_partial_refactor; consumer site in
trainers/dqn/fused_training.rs and the two name-table sites in
trainers/dqn/trainer/training_loop.rs are still listed in Task 2's
plan body.
2. Per-slot Rust buffer-pointer expression added for slots 24-35.
Already-accessible (no new accessor): slots 26, 32, 33, 34, 35
(4 via self.ptrs, 1 via existing bn_d_concat_buf accessor). Need
new accessor on GpuDqnTrainer: slots 24 (d_value_logits), 25
(d_adv_logits), 30 (aux_dh_s2_nb_buf), and 29 (cql_d_value_logits,
only if un-deferred). Need new accessor on GpuIqnHead: slot 28
(d_branch_logits_buf — note: production IQN backward uses
iqn_quantile_huber_loss, NOT iqn_backward_per_sample which is
declared but never loaded). Optional: slot 27 (d_h_s2_buf_ptr) —
can be inlined inside apply_iqn_trunk_gradient instead. Need new
accessor on FusedDqnTraining: slot 31 (only if un-deferred).
Becomes input for Task 3.
3. Citation typo: c51_loss_kernel.cu line 274 reference removed —
line 274 is __syncthreads() in the projection-reduction warp loop;
the a_std=sqrtf reference belongs to c51_grad_kernel.cu:274.
Loss-kernel sqrtf sites are 779 and 804.
Verified by reading each cited line; SQLX_OFFLINE=true cargo check
--workspace passes (docs-only).
Read-only γ audit of every backward-path kernel writing to bw_d_h_s2 /
grad_buf / save_h_s2 accumulators. Per-kernel inventory against
unsafe-pattern checklist (sqrtf-neg, 1/0, logf-≤0, expf-large, EMA
variance, atomicAdd/saxpy NaN-propagation). Each section includes:
identified pattern + line citation, severity, proposed guard form,
ISV bound option, F0 risk, Phase B flag-slot allocation.
Cross-referenced session_2026-04-05_nan_investigation.md residual 8%
step-2 NaN in apply_iqn_trunk_gradient — never closed; ranked top
suspect for SP1.
Output drives SP1 Phase B (12 new flag slots in nan_flags_buf 24→48)
and Phase C surgical fix(es). Durable artifact for SP2 framework
codification + SP3 structural-fix scoping.
Implementation plan for SP1 (Sub-project 1 of 3) of the numerical
stability investigation. Follows the γ + β methodology from the spec
at docs/superpowers/specs/2026-04-29-numerical-stability-investigation-design.md.
8 tasks across 4 phases:
- Phase A (Task 1): γ read-only audit producing docs/dqn-backward-nan-audit.md
- Phase B (Tasks 2-5): always-landing β instrumentation expanding
nan_flags_buf 24→48 with 12 new backward-kernel NaN check slots
+ 12 reserved slots for future coverage
- Phase C (Task 6): surgical fix(es), content-driven by audit + smoke
topology, ISV-driven for any dynamic bound (mandatory)
- Phase D (Task 7): multi-fold L40S smoke validation against 7 pass
criteria (F0 ≥ 95% baseline, F1+F2 monotone improvement, zero
NaN-CLAMPED-TO-ZERO, all 48 NaN flag slots remain at zero)
- Closure (Task 8): audit doc closure, memory entry, SP2/SP3 handoff
Operating principles (mandatory per spec):
- No deferrals — anomalies discovered during investigation get fixed
within SP1, not punted to SP2/SP3
- Combined RELATED fixes ship as rich commits (per
feedback_no_partial_refactor)
- ISV-driven design for any dynamic bound
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
9 substantive issues addressed inline:
1. ISV-driven design elevated from 'if applicable' to MANDATORY for
all dynamic bounds in SP1 fixes. Numerical-stability ε bounds are
the only carve-out (Invariant 1). Hardcoded tuning constants for
dynamic ranges explicitly rejected.
2. F0 Sharpe regression criterion changed from absolute (≥55) to
ratio-based (≥95% of latest baseline; floor 53.08 currently).
Prevents iterative erosion across multiple fix commits.
3. 'F1 trending positive' replaced with concrete monotone-improvement
test: Best Sharpe at last epoch ≥ Best Sharpe at first epoch of
the same fold.
4. Pass criterion distinguishes NaN-CLAMPED-TO-ZERO (failure) from
'Genuine grad collapse' (legitimate observation, permitted) per
the existing infrastructure from commit d1808df14.
5. Multi-source NaN scenarios explicitly supported — γ + β may
identify multiple kernels; SP1 fixes ALL within the same cycle.
6. F0-safety paper-review gate added BEFORE smoke validation. Audit
doc carries 'F0 risk' (low/medium/high) per proposed fix; high-
risk fixes get math-on-paper inspection before consuming L40S.
7. Audit doc structure now requires 'ISV bound option' column —
forces ISV-first thinking at audit stage, not as afterthought.
8. Anti-patterns expanded: micro-clamping (per-op clamping that hides
upstream causes); combining unrelated fixes (anti-pattern of the
rich-commit principle); hardcoded constants for dynamic bounds.
9. 48-slot allocation explicitly justified (24 used + 12 new + 12
headroom) AND marked reviewable by SP2 if right-size differs.
All 5 design sections preserved structurally; revisions integrated
inline. Per brainstorming skill: spec self-review fixes applied
without re-review cycle.
R1: K's hardcoded shrink-and-perturb (m×0.1, v×0.01) at fold boundary
violated feedback_adaptive_not_tuned (untracked tunable knobs) AND
created a downstream pathology: tiny v_hat denominator → oversized
Adam updates 50+ steps post-reset → trunk param overshoot → save_h_s2
NaN at F1 ~step 1745 (smoke-test-bkdx5 diagnostic).
K was introduced (commit 4ef1d8ebb) BEFORE fold_warmup_factor existed
in the same commit's "K + adaptive warmup" pair. With warmup_factor
in place — ISV-driven, dampens lr+clip via lr_eff = lr_base ×
max(MIN_WARMUP_LR_FRAC, fold_warmup_factor) — K is redundant. Single
mechanism, ISV-driven, no hardcoded constants. Eliminating K leaves
m,v reset to 0 at fold boundary; warmup_factor handles cold-start.
P: expanded nan_flags_buf 16→24 with 5 GRN-stage checks
(linear_a_out, elu_out, linear_b_out, glu_sigmoid_out,
layernorm_var/out) for finer-grained source identification if R1
alone doesn't fix F1.
Predicted outcomes:
- If K's tiny-v_hat was the cause: F1 trains successfully (R1 alone)
- If different mechanism: new GRN-stage flags pinpoint which sub-
stage produces NaN, enabling layer-level fix
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Existing nan_flags_buf [16] covers 13 buffers with per-step NaN checks
inside the captured training graph. But read_nan_flags() was only
called when training guard's halt_nan fired — which checks pinned-
readback grad_norm. NaN-clamped-to-zero gradients reach the pinned
scalar as 0, not NaN, so halt_nan never fires for the explosion case.
The collapse path (halt_grad_collapse, grad < 0.01) was firing instead
without reading the flags. Add flag readback in that path when
gr.raw_grad_norm < 1e-6 (suspicious zero) — logs NaN-CLAMPED-TO-ZERO
with flagged buffer names. Genuine near-zero gradients get a separate
"no NaN flags set" log so we can distinguish the two cases.
This should pinpoint which kernel produces the F1 ep2 NaN first
(C51 KL projection? IQN aux? CQL? aux heads?). Diagnostic only —
keep after fix lands; correct gate for future regressions.
After 6+ intervention layers (A.1+A.2+A.3+F+H+K+warmup+N) F1 ep2 still
explodes with grad_norm=2.66T while F1 ep1 trains cleanly. None of our
defenses catch the explosion path. Need data to pinpoint WHICH signal
explodes first.
Per-step FOLD_EXPLOSION_DIAG in fold >= 1: when grad_norm > 1000 OR
jumps 5x from prior guard step, fire diagnostic warn with:
- loss decomposition: total / c51 / mse / iqn (pinned readback, no DtoH)
- Q range: q_min / q_mean / q_max via cold-path reduce
- atom positions: per_sample_support[0,d0] + [0,d2] (v_min, v_max, dz)
- prev grad norm + ratio for context
- trigger flag + post-trigger trajectory countdown
After trigger, continues emitting for 5 guard steps so the explosion
trajectory is captured (not just the first crossing).
Plus an unconditional FOLD_EXPLOSION_DIAG[F1_END_EP1] baseline log at
the end of fold 1 epoch 0 — healthy state immediately preceding the
explosion. Compare-and-contrast with per-step explosion frames pins
the runaway driver.
CQL and ensemble losses are not pinned-readback (transient device
buffers consumed inside the training graph) and are explicitly absent
from the decomposition. If none of {c51, iqn, mse} is the runaway
driver but total_loss still explodes, that implicates the unexposed
CQL/ens path — the absence is itself diagnostic information.
Implementation:
- Three diagnostic-only fields on DQNTrainer: current_fold,
last_logged_grad_norm, explosion_diag_steps_remaining. Reset
last_logged + remaining in reset_for_fold; current_fold
overwritten unconditionally at fold-loop entry.
- Three accessors on FusedTrainingCtx: explosion_diag_loss_components
(pinned), explosion_diag_atom_range (12-elem cold DtoH),
explosion_diag_q_range (reduce + 28B DtoH).
- Three accessors on GpuDqnTrainer: c51_loss_pinned_value,
mse_loss_pinned_value, stream_for_diag.
Diagnostic-only — remove once root cause identified. Not for prod.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Smoke smoke-test-xw4c6 showed F1 ep2 grad_norm=8.4B polluting the
adaptive_clip EMA -> next adaptive_clip threshold became huge ->
subsequent extreme grads passed unclipped -> NaN propagation -> grad
collapse -> fold 1 fails despite K+warmup fixing the first-step problem.
Add Winsorized clamp on raw_grad_norm before the adaptive_clip EMA
update: clamped_sample = min(raw, K * previous_adaptive_clip) where
K = 100 (numerical-stability bound, not a tuned constant — explicit
"single sample can grow EMA at most 100x in one step" semantics).
Companion observability log emits GRAD_CLIP_OUTLIER warning per clamp
event so we can see when this fires in subsequent smokes.
Fast/slow grad_norm EMAs (driving warmup factor) intentionally NOT
winsorized — they're a stability signal that SHOULD respond to
outliers, providing extra warmup damping when the system is unstable.
Predicted impact on Plan C smoke F1 ep2: 8.4B grad -> clamped to
~3000 for EMA -> next adaptive_clip caps at ~1000 -> subsequent F1
batches get bounded gradients passed to Adam -> no NaN propagation.
F1+F2 should now train through, exposing whether the underlying
Q-target inflation requires further structural work or whether
clipping alone suffices.
Composes with K+warmup (4ef1d8ebb), A.2, A.1/A.3, F+H — fourth layer
of the adaptive gradient-stability scaffolding (now five total:
A.2 target-sync drift dampening, K+warmup post-reset transient
dampening, N persistent clip-EMA outlier defense).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Closes the F1 ep1 catastrophic-overshoot gap exposed by smoke-test-s9h4h
(F0 succeeded with Best Sharpe = 36.03; F1 ep1 grad_norm = 355,009 — 5
orders of magnitude larger than F0 steady-state ~10 — leading to NaN
propagation and grad-clamp-to-zero early stop). Combined fix: (1) Adam
shrink-and-perturb at fold boundary (replace m=0/v=0 with m*=0.1, v*=0.01
to preserve direction while damping magnitude), and (2) a single
adaptive ISV signal driving BOTH lr_eff and clip_eff dampening over
the fold's first ~50 steps.
K — Adam shrink-and-perturb in `GpuDqnTrainer::reset_adam_state`:
- m *= 0.1, v *= 0.01 via existing `dqn_scale_f32_kernel` (loaded as
`scale_f32_ungraphed`); t_pinned still zeroed so bias correction
restarts. Architectural constants (preserve direction / lose magnitude
history) per `feedback_isv_for_adaptive_bounds.md` Invariant 1
carve-out — not tuned. Composes with existing param shrink-and-perturb
(`alpha=0.8`) in `FusedTrainingCtx::reset_for_fold`. Root cause for
F1 overshoot: m=0,v=0 → first Adam step ≈ lr × g / ε → 6 OoM
amplification.
New CUDA kernel `fold_warmup_factor_kernel.cu`:
- Single-block single-thread cold-path producer mirroring
`q_drift_rate_ema_kernel.cu` / `moe_lambda_eff_kernel.cu` shape.
- Reads two grad-norm EMAs (fast α=0.1, slow α=0.001) plus host-passed
step counter; writes ISV[FOLD_WARMUP_FACTOR_INDEX=130] = clamp(fast/slow, 0, 1).
- Bootstrap branches (steps_observed < 200, slow EMA < 1e-6) emit
factor=1.0 (no damping during cold-start). No atomicAdd; no DtoH.
New ISV slot Q_DRIFT_RATE_INDEX → FOLD_WARMUP_FACTOR_INDEX = 130:
- ISV_TOTAL_DIM 130 → 131; layout fingerprint shifts (checkpoint-
incompatible per `feedback_no_legacy_aliases.md`, expected for a
real architecture change).
- FoldReset entries: `isv_fold_warmup_factor` → 0.0 and companion
`isv_grad_norm_fast_ema` → 0.0 (lockstep reset per
`feedback_no_partial_refactor.md`); slow EMA persists across folds
as the cross-fold steady-state baseline.
- Two new mapped-pinned scalars on GpuDqnTrainer (grad_norm_fast_ema_pinned,
grad_norm_slow_ema_pinned) fed by `update_adaptive_clip` from the
same `gr.raw_grad_norm` observation source as the existing adaptive
clip EMA.
Two consumers, both monotone (only dampen, never excite):
- lr_eff = cosine_effective_lr × max(MIN_WARMUP_LR_FRAC=0.05, factor)
via `set_lr` per-step. New `cosine_effective_lr_base` field
on DQNTrainer composes the cosine schedule's per-epoch
baseline with the warmup factor's per-step damping (rather
than overriding the cosine schedule).
- clip_eff = clip_base × (MIN_CLIP_FRAC=0.1 + 0.9 × factor) via new
`set_active_clip` setter on FusedTrainingCtx + GpuDqnTrainer.
Composes with the EMA-derived `clip_base = grad_norm_ema × 2`
that `update_adaptive_clip` just wrote to the pinned slot.
Numerical-stability bounds 0.05 / 0.1 are Invariant 1
carve-outs.
Steady-state behaviour unchanged: factor=1 → lr_eff=lr_base,
clip_eff=clip_base. Fold-boundary behaviour: factor starts at 0 →
lr_eff = 0.05 × lr_base, clip_eff ≈ 0.1 × clip_base; rises to 1 over
~50 steps as the fast EMA catches up to the slow steady-state EMA.
Predicted impact on Plan C smoke F1: 355,009-magnitude transient grad
clipped to ~clip_base × 0.1 ≈ 1.0 (vs 10), Adam state shrunk instead
of zeroed → first step update bounded; grad recovers normally over
~50 steps. Companion to A.1 (prev_epoch_q_mean reset), A.2 (adaptive
Polyak-tau), A.3 (gradient_collapse_counter reset), F+H (kill-criterion
robustness) — completes the fold-boundary state-reset family.
Per `pearl_adaptive_moe_lambda.md` (kernel + ISV slot + bootstrap +
reset + observability template), `pearl_cold_path_no_exception_to_gpu_drives.md`
(GPU-stays-on-GPU even at cold-path cadence),
`pearl_blend_formulas_must_have_permanent_floor.md` (lr/clip floors
are permanent minimums), `feedback_adaptive_not_tuned.md` (lr+clip
ISV-driven), `feedback_isv_for_adaptive_bounds.md` (factor IS the
bound; consumers compose at runtime), `feedback_no_atomicadd.md`
(single-thread reduce), `feedback_cudarc_f64_f32_abi.md` (slot index
passed as i32), `feedback_no_partial_refactor.md` (kernel + slot +
reset + producer + 2 consumers all land together),
`feedback_no_quickfixes.md` (replaces brittle full-reset with
adaptive damping; not threshold relaxation).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Smoke smoke-test-vh9bj revealed: after F+H Q-drift kill fix, fold 0
trains successfully (5 epochs, Best Sharpe 36.03), but fold 1 fails
at epoch 1 with "Gradient collapse detected for 5 consecutive epochs"
despite epoch 1 grad_norm=296,417 (healthy). Root cause: the per-step
gradient_collapse_counter on DQN was already near patience (5) from
fold 0's late-epoch near-zero-grad steps, and DQNTrainer::reset_for_fold
didn't clear it. Plus, training_steps accumulating across folds made
the `past_warmup` gate always-true in fold 1+, removing the warmup
grace period for data distribution shifts.
Add DQN::reset_for_fold zeroing both. Wire from DQNTrainer::reset_for_fold
alongside the A.1 prev_epoch_q_mean reset. Pure additive — same
fold-boundary state-reset gap pattern as A.1, A.2, F.
Predicted impact: fold 1+ now starts with counter=0 and warmup window
restored; gradient collapse check has its full per-fold grace period.
Current criterion uses |q_mean| / |prev_q_mean| which explodes near
zero crossings (legitimate cold-start has |prev_q_mean| ≈ 0.01-0.1,
making any non-tiny current trigger the ratio threshold).
smoke-test-n9xzr fired with prev=−0.0795, curr=0.7647, ratio=9.62×
despite this being natural cold-start growth, not geometric runaway.
Replace with rolling 5-epoch window of q_means. Compute median +
MAD (Median Absolute Deviation, a 50%-breakdown estimator robust
to single outliers); kill condition becomes
|q_mean − median(window)| > 4.0 × max(MAD(window), 0.01)
AND the existing adaptive floor.
Constants are declared `const` near the kill block:
- Q_DRIFT_WINDOW_SIZE=5 (matches smoke fold length, gives MAD a
meaningful estimator without averaging across regime shifts)
- Q_DRIFT_WARMUP_SAMPLES=3 (skip until ≥ 3 priors — smaller window
degenerates to half-range MAD that trips on monotonic
trajectories)
- Q_DRIFT_DEVIATION_THRESHOLD=4.0 (4 MADs ≈ 2.7σ Gaussian-
equivalent; clear outlier without firing on every legitimate
dip; literal MAD count rather than σ because q_mean is non-
Gaussian during cold-start)
- Q_DRIFT_MIN_DEVIATION=0.01 (numerical-stability floor when
window is constant)
Window resets at fold boundary alongside prev_epoch_q_mean /
adaptive_tau (A.1 pattern — cross-fold q-stats are independent
training runs, mixing them would inflate MAD or shift the
median).
Predicted impact on smoke-test-n9xzr:
- Plan C smoke fold 0 ep2: window has only 2 priors, warmup gate
not yet satisfied → kill stays silent (was firing on ratio=9.62×)
- Genuine geometric runaway (q_mean → 62.5 from baseline 0.5 over
4 epochs): ep3 deviation = |62.5 − 2.5|/MAD=2.0 = 30 > 4 AND
|62.5| > floor=3×12=36 → kill fires correctly
Both conditions ANDed (floor + deviation), preserving the
production-safety semantics of the original criterion. The floor
check is unchanged; only the divergence detector is replaced.
Per feedback_no_quickfixes.md: this is a principled robust-
statistics replacement, not a threshold relaxation.
Per feedback_no_partial_refactor.md: window field, constants,
criterion site, and fold-boundary reset land in lockstep — the
kill criterion's contract is internally consistent across
trainer/mod.rs, constructor.rs, and training_loop.rs.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
The Q-drift kill criterion's adaptive floor formula is
max(0.5, 3 × max(ISV[Q_ABS_REF=16], ISV[Q_DIR_ABS_REF=21]))
but smoke-test-n9xzr showed both ISVs stuck at bootstrap 0.05 even
when q_mean reached 0.76 — the producers (q_mag_bin_means_reduce
and q_dir_bin_means_reduce) only fired inside reduce_current_q_stats
at epoch boundary, while the per-step captured update_isv_signals
consumed the resulting scratch buffers each training step.
Concrete failure mode at α=0.05 EMA with raw ≈ 1.0:
- Epoch 0: scratch is zero-initialized, all per-step EMA pulls
toward zero, ISV[16,21] stay 0.0
- Epoch 0 boundary: reduce_current_q_stats fires, scratch becomes
q_abs_ref ≈ 1.0, single update_isv_signals writes ISV[16] = 0.05
- Epoch 1+: with stale scratch held constant between boundaries,
per-step EMA over hundreds of steps would saturate — but smoke
step counts are small enough (small batch=64, buffer=256) that
the slot stays near 0.05
- kill_floor degenerates to max(0.5, 3 × 0.05) = 0.5 forever, so
the criterion fires on legitimate cold-start growth
Fix: launch q_mag_bin_means_reduce + q_dir_bin_means_reduce in
both fused_training paths (captured adam_update_child at ~line 2228
AND ungraphed step-0 fallback at ~line 1465) immediately before
update_isv_signals. Per-step graph replays now read q_out_buf that
forward_child populated this same step, write fresh scratch, and
the captured update_isv_signals consumes it on the same stream.
Per pearl_cold_path_no_exception_to_gpu_drives.md: cold-path EMAs
fire alongside their consumers. Per feedback_no_partial_refactor.md:
graphed and ungraphed paths migrate together — same producer-
consumer contract.
After this fix, the floor will adapt: ISV[16,21] saturate to the
policy's actual q_abs_ref scale within ~60 steps at α=0.05, so by
epoch 2 the floor becomes 3 × ~0.5 = ~1.5, no longer firing on
legitimate cold-start growth.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
C.4 timing bonus (experience_kernels.cu) and D.4b regime penalty
multiplied by the trade-cumulative |reward| / |final_pnl| as an
unbounded multiplicand. Across trades this compounds — bonus values
inflate as Q values inflate, then bonus rewards inflate Q values
further (Bellman bootstraps off shaped reward).
Per pearl_one_unbounded_signal_per_reward + feedback_isv_for_adaptive_bounds:
replace |reward|/|final_pnl| with ISV[Q_DIR_ABS_REF_INDEX]-bounded
variant. The unbounded multiplicand becomes the direction-branch
Q-scale EMA (already-tracked, gradient-decoupled), not the
trade-cumulative shaping output. Cap is adaptive (matches Q magnitude
as it evolves) and breaks the multi-trade compounding loop.
Discovered during Plan C Phase 2 smoke diagnosis (researcher report
2026-04-29 a25f669e9df953174). a52d99613 baseline also hits
bonus=235 — pre-existing structural pathology, not Plan-C-specific.
Other shaping sites (D.4a persist, B.2 novelty, D.4c stable) already
have all factors bounded — no fix needed.
Surgical change set:
- experience_env_step gains final `int q_dir_abs_ref_idx` param.
- C.4 (line 2521): pnl_unit = |final_pnl|/max(ISV[21], eps);
pnl_capped = min(pnl_unit, 1) * ISV[21].
- D.4b (line 2581): same pattern, reward_unit/reward_capped.
- gpu_experience_collector.rs:3800 wires Q_DIR_ABS_REF_INDEX as i32.
No new ISV slot, no new producer kernel, no layout-fingerprint shift —
ISV[21] already produced by q_stats_kernel.cu since Plan 1.
Predicted impact:
- bonus EMA drops O(100-256) -> O(0.05-2.0)
- rc[5] -> Bellman-target optimism loop broken
- Plan C smoke: Q-drift kill at F0 ep2 likely no longer fires
- a52d99613 baseline: same effect; validates pre-existing fix
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Wires the q_drift_rate_ema kernel + ISV slot Q_DRIFT_RATE_INDEX=129
(landed in fc8dbb0a8) into the production training path:
- StateResetRegistry FoldReset entry isv_q_drift_rate -> 0.0
(companion to A.1's prev_epoch_q_mean reset; both ensure no
cross-fold leakage of drift state).
- reset_named_state dispatch arm in training_loop.rs.
- Per-epoch launch_q_drift_rate_ema call after the Q-drift kill
check, gated on the same `prev_epoch_q_mean.abs() > 1e-6`
cold-start guard. Both q_mean_curr and q_mean_prev are in scope
BEFORE the `prev_epoch_q_mean = q_mean` update, so the producer
sees the correct delta.
- tau_update_kernel.cu multiplies the cosine-scheduled,
health-coupled tau_eff by `1 / (1 + clip(ISV[129], 0, 4))` so
tau ranges [tau_base/5, tau_base] — monotone dampening under
drift; healthy runs (drift ≈ 0) unaffected.
- reset_for_fold comment in trainer/mod.rs notes the registry
entry handles the new ISV slot.
Predicted impact: Plan C fold 0 ep2 with q_mean(t)=0.82,
q_mean(t-1)=-0.018, ISV[16]+ISV[21]≈0.1 -> drift_rate ≈ 8.4 ->
clipped to 4 -> tau_eff = tau_base × 0.2 (5× slower target sync,
dampens Q-target optimism through the inflation spike).
a52d99613 fold 0 with q_mean staying ~0.21 -> drift_rate ≈ 0 ->
tau_eff = tau_base (unchanged).
Per pearl_adaptive_moe_lambda.md — canonical "EMA-tracked
diagnostic drives a controller" pattern. Per
pearl_cold_path_no_exception_to_gpu_drives.md — cold-path scalar
arithmetic stays on GPU. Per feedback_no_partial_refactor.md —
consumer (tau_update + state_reset + producer launch) all migrate
together in this commit.
Layout fingerprint already shifted by fc8dbb0a8 (slot + ISV_TOTAL_DIM
bump); no additional fingerprint shift in this commit.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Single-block single-thread ISV producer mirrors h_s2_rms_ema /
moe_lambda_eff pattern. Computes per-epoch
drift_rate = |q_mean(t) - q_mean(t-1)| /
max(|q_mean(t-1)|, ISV[Q_ABS_REF] + ISV[Q_DIR_ABS_REF], 1e-6)
clipped to [0, 4] and writes to ISV[Q_DRIFT_RATE_INDEX=129].
Includes the full ISV-contract shift required for the kernel to load:
- ISV_TOTAL_DIM 129 -> 130
- Q_DRIFT_RATE_INDEX = 129 (tail-appended after MOE_LAMBDA_EFF=128)
- Cold-start ISV[129] = 0.0 in constructor (no-op dampening factor)
- layout_fingerprint_seed entry Q_DRIFT_RATE=129 + ISV_TOTAL_DIM=130
- Cubin static, kernel field, kernel load, struct assignment
Per feedback_no_partial_refactor.md the ISV slot + dim + fingerprint
all migrate together (the wrapper references Q_DRIFT_RATE_INDEX so they
cannot be split). Tau consumer + state reset registry + per-epoch
producer launch land in the next commit.
Audit doc dqn-wire-up-audit.md updated with the kernel + ISV slot
description per Invariant 7.
No callers in this commit; layout fingerprint shifts so existing
checkpoints will fail-fast at load per feedback_no_legacy_aliases.md
(expected for a real architecture change).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Plan C Phase 2 T9. Verifies the production kernel's eval-mode argmax
exactly matches a Rust ground-truth E[Q] argmax on a 256-sample batch
with realistic per-direction non-degenerate C51 distributions, and
confirms Thompson explores Long+Short ≥ 40% in training mode.
The plan-prescribed approach (reuse Phase 0 Test 0.F's converged
checkpoint loader) was deferred — Test 0.F itself already exercises
the safetensors load + branching forward path. Replacement strategy
from the dispatch brief: synthetic batch via direct buffer write.
Setup:
- 256 samples, 21 atoms, per-(sample, direction, atom) C51 logits
drawn from a deterministic hash → uniform [-1, 1] (post-Xavier-init
scale of fresh C51 head outputs)
- Per-sample adaptive support [v_min ∈ [-1.0, -0.2], v_max ∈ [0.2, 1.0]]
matching the layout produced by `update_per_sample_support`
- Uniform q_values (mag/ord/urg fall through Boltzmann uniformly)
Rust ground-truth: softmax(b_logits[i, d]) · atom_vals[i, d] computed
with the numerically-stable subtract-max softmax matching the kernel's
softmax_c51_inline.
Assertions:
- Eval mode: per-sample dir_idx EXACTLY matches ground-truth argmax E[Q]
- Train mode: count(d ∈ {Long, Short}) / batch > 0.40
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Plan C Phase 2 T8. The plan-prescribed pre-T2 snapshot approach was
impossible (T2 had already landed); replacement strategy (ii) from the
dispatch brief — behavioral parity vs an analytical Boltzmann reference
computed in Rust — is used.
Setup forces direction = Long (d=2) deterministically via peaked C51
logits (Long peaked at v=+0.8 atom; other directions at v=-0.5), so
the kernel's Hold/Flat → mag_idx=0 short-circuit doesn't mask the
magnitude branch's Boltzmann sampling. q_values are crafted with each
branch (mag/ord/urg) peaked at a single bin with magnitude 1.0:
Mag Q = [0.0, 0.0, 1.0] peak at Full (mag=2)
Order Q = [1.0, 0.0, 0.0] peak at Market (ord=0)
Urgency Q = [0.0, 1.0, 0.0] peak at urg=1
With q_range=1.0 in all three branches, tau collapses to 1.0 and the
analytical Boltzmann probabilities are:
P(best) = 1/(1 + 2/e) ≈ 0.5767
P(other) = 1/e/(1 + 2/e) ≈ 0.2117
Tolerance: at batch=8192 the 1-σ Bernoulli noise is ~0.0055 for p≈0.58;
±5% absolute tolerance covers ~9σ. Algorithmic divergence (e.g. an
inadvertent strict-argmax substitution) would shift P(best) to 1.0 —
trivially detected by the ±5% tolerance.
Assertions:
- dir_idx == Long for every sample (eval argmax E[Q] over peaked C51)
- mag/ord/urg histograms each within ±5% of the Boltzmann reference
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Plan C Phase 2 T7. #[ignore]-gated GPU test that runs BOTH the
production experience_action_select kernel AND the standalone
direction_thompson_v2_test (added in commit 5de5e546a) on identical
Phase 0 Test 0.D failure-mode inputs and asserts the resulting
direction histograms are statistically indistinguishable.
Path (a) — bit-identical comparison via a shared pre-computed uniform
array — would require editing the standalone v2 kernel's signature to
accept a `float* uniforms` instead of generating its own LCG draws.
Out of scope for the 1-2-hour dispatch. Path (b) — KS-style histogram
comparison — is used.
Setup:
- Single per-direction logits tile (Phase 0 Test 0.D failure mode)
- Production: tile replicated across batch=100 000 samples; Philox
keyed on (i, timestep=0, ctr) per thread
- Standalone v2: same tile fed to single-tile launcher with n_seeds=100 000;
LCG keyed on (base_seed=0 + seed_idx) per thread
Assertion:
- KS distance over the {Short, Hold, Long, Flat} histograms ≤ 0.02
(n=100k empirical-CDF noise floor for matching distributions is
~4·sqrt(2/n) ≈ 0.018; algorithm divergences would shift mass by
O(10%) → KS ≈ O(0.1), easily detected)
- Sanity: both histograms have P(Long+Short) ≥ 0.20 (rules out
coincidental shared-mode collapse with KS ≈ 0)
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
reset_for_fold cleared q_value_history (the q_mean source) but missed
the kill criterion's prev_epoch_q_mean baseline AND the adaptive_tau
modulation state. So the kill ratio at fold-N epoch 0 was computed
against fold-(N-1) epoch 5's q_mean — a stale cross-fold baseline.
Discovered while diagnosing why Plan C Phase 2 Thompson smoke triggers
Q-drift kill at fold 0 epoch 2 (researcher report 2026-04-29).
Pairs with the existing q_value_history clear; matches the
project_fold_boundary_q_drift_resolved.md pattern (kill criterion's
companion resets were partially missed).
Independent of Plan C — this is a bug in the kill criterion's
fold-boundary contract regardless of which exploration mechanism is
active.
Plan C Phase 2 T6. Adds an #[ignore]-gated GPU test that exercises
the production experience_action_select kernel directly through a
minimum-viable inline cudarc fixture (ProdActionSelectFixture). The
Plan B audit fixture (GpuExperienceCollector::new_for_test) was never
built, so each test in this Phase 2 batch builds its own buffers,
launches the kernel, and reads back action picks for assertions.
Setup mirrors Phase 0 Test 0.D failure-mode:
- Flat/Hold C51 logits = peaked at v=0 atom (δ(v=0) shape)
- Long/Short C51 logits = log-Gaussian centred at v=-0.001, σ=0.05
- Linear adaptive support [v_min=-0.5, v_max=+0.5, delta_z=0.05]
- Uniform q_values (mag/ord/urg fall through Boltzmann uniformly)
Assertions:
- Eval mode (eps=0): 10 launches at distinct timesteps produce
IDENTICAL per-sample dir_idx (deterministic argmax E[Q]).
Every sample picks Hold or Flat (E[Q] argmax bias reproduces).
- Training mode (eps_start=0.5): >=30% of 10 000 samples pick
Long or Short (Thompson explores directional alternatives).
Shared helpers (PROD_B0..3, decode_dir/mag/ord/urg, fill_peaked_logits,
fill_linear_support, ProdActionSelectFixture) added at the top of the
Phase 2 section to be reused by Tests 2.B-2.D.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
The prior `test_eval_action_select_boltzmann_bounded` asserted
Boltzmann theory (P(best)≈0.366 at tau=q_range) on the direction
branch. Plan C T2 (commit 52a2663a2) replaced the direction-branch
Boltzmann path with single-distribution C51 Thompson + argmax-E[Q]
eval; the T2 amendment (5de5e546a) added 4 args to the kernel ABI
(b_logits_dir, per_sample_support, atom_positions, n_atoms). The old
test was launching with an outdated arg count and asserting a
distribution shape that no longer applies.
Migration (per feedback_no_partial_refactor):
- Rename to `test_eval_action_select_eval_argmax_picks_best`
- Build per-direction C51 logits PEAKED at distinct E[Q] values:
Short=-0.5, Hold=-0.1, Long=+0.5 (best), Flat=+0.1
- Linear adaptive support [v_min=-1, v_max=+1, delta_z=0.1] uniform
across all (sample, direction) pairs
- Run kernel in eval mode (eps_start=eps_end=0): assert P(Long)>=0.99
i.e. the kernel deterministically picks argmax(E[Q]) per sample
This validates the T2 amended ABI runs end-to-end and exercises the
production kernel's argmax-E[Q] eval branch with a controlled
distinct-E[Q] setup. Companion to Phase 2 Tests 2.A-2.D in
distributional_q_tests.rs which add training-mode assertions and
production-vs-standalone parity checks.
Verification: `SQLX_OFFLINE=true cargo check -p ml --lib --tests` clean.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Documents the project-wide invariant: any new distributional Q-head
(atoms, quantiles, ensembles) must expose sampleable interface and
wire into Thompson (training) + argmax (eval) action selection.
Pre-commit hook (Invariant 7) will enforce.
Architecture reflects Plan C Phase 2 amendment (commit 5de5e546a):
single-distribution C51 Thompson is the production pattern; joint
C51+IQN forward-looking under the same contract.
Counterpart to existing val_active_frac. Reports the fraction of
TRAINING rollout actions that were Long or Short. Required for L3
verification gate of Plan D (Phase 3) which checks Thompson is
generating directional exploration during training.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Threads b_logits_dir / per_sample_support / atom_positions / n_atoms
through experience_action_select kernel launch in
gpu_backtest_evaluator.rs's chunked val pipeline (line ~1722, inside
submit_dqn_step_loop_cublas).
The evaluator sources Q-values via the QValueProvider trait
(delegates forward to the trainer's CUDA-Graphed cuBLAS), so this
change extends the trait surface rather than duplicating the
forward:
- New trait method compute_q_and_b_logits_to(states_ptr, batch,
q_out_ptr, b_logits_out_ptr) — DtoD-copies trainer's
on_b_logits_buf into caller's chunked buffer per sub-batch
iteration alongside the existing q_out copy.
- New trait accessors per_sample_support_ptr / atom_positions_ptr /
num_atoms / total_branch_atoms (stable trainer-owned buffers).
- FusedTrainingCtx implements all of the above; trainer gains
on_b_logits_buf_ptr / atom_positions_buf_ptr /
per_sample_support_ptr_get pub accessors.
- Evaluator gains chunked_b_logits_buf field (sized
[n_windows * CHUNK_SIZE, total_actions * num_atoms] + 32*3 tail
safety), allocated in ensure_action_select_ready.
Phase 6's last-step plan_params forward keeps using the original
compute_q_values_to (b_logits not consumed there).
Plan C Task 4 — evaluator companion to T3 collector wire-up. After
this commit, all production callers of experience_action_select use
the amended kernel ABI (T2 amendment 5de5e546a) end-to-end
(feedback_no_partial_refactor).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Threads b_logits_dir (= exp_b_logits, first N*b0*NA region for direction
branch), per_sample_support, atom_positions (NULL = linear), and n_atoms
through experience_action_select kernel launch in gpu_experience_collector.
After amendment commit 5de5e546a, the kernel uses single-distribution C51
Thompson with production's adaptive per-direction support — no IQN
inference required at rollout.
Plan C Task 3 (corrected file path: collector, not trainer per
gpu_dqn_trainer.rs/gpu_experience_collector.rs distinction).
Plan C as authored assumed:
- c51_probs_dir = post-softmax probabilities (didn't exist on collector;
only raw exp_b_logits is materialised)
- atom_values = single global linear support (production uses per-sample
per-direction adaptive [v_min, v_max, delta_z] + optional atom_positions)
- iqn_quantiles_dir = rollout-side IQN inference (doesn't exist; IQN is
training-only)
Amended kernel signature uses the production architecture:
- b_logits_dir [N, b0_size, n_atoms] — raw direction-branch logits
- per_sample_support [N, b0_size, 3] — adaptive per-direction support
- atom_positions [b0_size, n_atoms] — non-linear positions (NULL = linear)
- n_atoms
Direction-branch Thompson is now single-distribution over C51 (with adaptive
support), not joint C51+IQN. Single-distribution still provides the
principled posterior sample that fixes the UCB selector/target asymmetry —
the goal of Plan C is preserved, just sourced from the existing rollout-time
distribution instead of a non-existent IQN inference path.
New device-inline helpers: softmax_c51_inline, compute_atom_values_inline.
Plan + audit doc updated. Phase 0 standalone test kernel gained two new
entry points (direction_thompson_v2_test, argmax_eq_v2_test) matching the
amended production API; original Plan A entry points retained for tests
0.B-0.F. Tasks 3+4 (buffer wiring) unblocked — collector/evaluator already
have per_sample_support_buf and exp_b_logits in production.
Plan C T2 commit 52a2663a2 migrated out_conviction to read from
e_dir[] (joint E[C51 + IQN]) per spec "Conviction stays E[Q]-based"
but left the out_q_gaps writer reading raw q_b0[]. Both signals feed
env_step's Kelly-adjacent position sizing; mixing q_b0-scale gaps
with e_dir-scale conviction inside the same downstream consumer
violates feedback_no_partial_refactor.
Migrate out_q_gaps to read from e_dir[] (already hoisted to outer
scope by T2). Contrarian-mode sign flip preserved symmetrically.
Comment at ~line 1289 updated to reflect joint E[Q] source.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Replaces eps-greedy + Boltzmann on direction with:
- Training: argmax of (sample_C51 + sample_IQN) per direction (Thompson)
- Eval: argmax of (E_C51 + E_IQN) per direction (no exploration)
Conviction calculation now reads e_dir[] (joint E[Q] computed inline)
preserving the spec's requirement that conviction uses raw E[Q] not
samples (avoids Kelly cap jitter).
Magnitude/order/urgency branches: unchanged.
Per-function purpose comments + parameter-contract notes were present
in the Phase 0 reference (thompson_test_kernel.cu) but dropped during
the port to experience_kernels.cu. Restoring them addresses code
review feedback. No functional change.
The original PAUSED state was motivated by measurement-artefact hunt
exit. The bug-hunt cycle is complete (commits a86fba2b1, b8788511c).
A new pathology has surfaced: ff00af68a's UCB count bonus activation
causes selector/target asymmetry → F0 Q-drift kill at epoch 2 →
F1+F2 cascade. Verified by paired DIAG smokes (smoke-test-qlz7t fail,
smoke-test-wmsht pass).
Thompson sampling on C51+IQN distributions eliminates the asymmetry
by construction (sample from learned distribution; no augment-then-
argmax step). Net code-surface decrease — replaces eps-greedy +
Boltzmann + UCB with one principled mechanism.
Plan C Phase 2 execution begins on branch plan-c-phase-2-thompson.
Phase 2A of the HEALTH_DIAG GPU port. Lands the simplest member of the
kernel family (single-block single-thread ISV scalar copies) plus the
orchestrator scaffolding (cubin loader, mapped-pinned snapshot owner,
launcher) that subsequent phases will reuse.
Why simplest first: 5 kernels each with different masking semantics
landing in one commit risks silent numerical mismatches that downstream
parsers (aggregate-multi-seed-metrics.py, smoke summarisers) would
mis-parse. Per feedback_no_quickfixes.md, kernel-by-kernel with
parallel-shadow validation. Phase 4 deletes CPU paths in one commit
once all kernels are bit-identical.
What this kernel does: copies 22 ISV signal-bus slots into the matching
HealthDiagSnapshot fields — reward-component EMAs (slots 63-68), VSN
attention focus EMAs (87, 88), target-drift EMAs (92, 93), aux-head
loss EMAs (113, 114, 117), MoE expert utilisation (118-126), gate
entropy (126), and λ_eff (128). 22 of 147 snapshot words populated
(reward_split[6] + noisy_vsn[2] + noisy_drift[2] + aux[3] + aux_moe[10]).
Producer-only — no consumer reads health_diag.snapshot() yet; Phase 3
wires the call site, Phase 4 makes the snapshot the sole emit source.
Architecture rules upheld:
- No HtoD / DtoH (writes through mapped-pinned device pointer).
- No atomicAdd (single-thread kernel; family-wide rule for 2B-2E too).
- No DtoD memcpy on the diag path (kernel pointer-load + store).
- Kernel WORD-index table inline; static_assert pins WORD_TOTAL=147 in
lockstep with snapshot_size_is_stable test.
- ISV indices passed as kernel args (not #define'd) so the kernel
never embeds the numbering — single source of truth lives in
gpu_dqn_trainer.rs constants.
Changes:
- new crates/ml/src/cuda_pipeline/health_diag_kernel.cu (311 lines).
- new crates/ml/src/cuda_pipeline/gpu_health_diag.rs (190 lines)
holds GpuHealthDiag orchestrator (cubin module + isv_mirror_kernel
handle + MappedHealthDiagSnapshot).
- mod.rs: pub mod gpu_health_diag.
- build.rs: registers health_diag_kernel.cu in kernels_with_common.
- gpu_dqn_trainer.rs:
- imports super::gpu_health_diag::GpuHealthDiag;
- adds health_diag: Option<GpuHealthDiag> field next to aux_heads_fwd;
- constructor calls GpuHealthDiag::new() after the cubin loads;
- adds pub fn launch_health_diag_isv_mirror() (uses compile-time
ISV index constants, single source of truth);
- adds pub fn health_diag_snapshot() accessor.
- docs/health_diag_inventory.md: Phase 2A entry + Phase 4 deletion
list pinned per Invariant 7.
- docs/dqn-wire-up-audit.md: new row for health_diag_kernel.cu under
"CUDA Kernels (.cu files)" — classified Pending wire-up (Phase 3).
Verification:
- cargo check -p ml: 0 new warnings (12 lib warnings pre-existing,
none touch health_diag/gpu_health_diag).
- cargo test -p ml --no-run: clean.
- cargo test -p ml --lib cuda_pipeline::health_diag::tests: 3/3 pass
(snapshot_size_is_stable, default_is_zeroed, alignment_is_4_bytes
— the WORD_TOTAL=147 static_assert in the kernel matches the host
size_of test bit-for-bit).
Scope discipline (per prompt's "if kernel #3 takes too long, stop"):
This commit lands the orchestrator scaffolding + simplest kernel only.
Phases 2B-2E (eval-histogram, q-mag-reduce, per-sample-reduce,
finalise) intentionally deferred — masking semantics for q_mag_reduce
in particular need careful per-CPU-getter inspection before kernel-side
implementation. Producer-only by design; no production caller of
launch_health_diag_isv_mirror in this commit (matches AuxHeadsForwardOps
landing pattern from Plan 4 Task 6 Commit A).