8956c2fb777972bb588478248f660ef8616a8ade
4476 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
8956c2fb77 |
fix(dqn): SP3 close-out — Mech 9 post-Adam weight clamp + Mech 8 revert
Coordinated close-out of SP3 Q-learning numerical-stability per
feedback_no_partial_refactor.
REVERT Mech 8 (slow_ema fold-boundary reset). smoke-test-rxhjh on
|
||
|
|
b8a7ac6f70 |
fix(dqn): SP3 Mech 8 — slow_ema fold-boundary reset (real root cause)
User-identified root cause: grad_norm_slow_ema (α=0.001, half-life
~693 steps) persists across fold boundaries while every other
distribution-tracking signal (Adam m/v, target nets, atom positions)
gets reset. Mech 6's upper_bound formula (100 × slow_ema × isv) was
anchored to the WRONG scale during F1 ramp-up, driving the
non-monotonic multiplier-tuning dance across smokes:
- smoke-test-fxvkk (mult=100×): F0=44, F1 NaN @ 3720
- smoke-test-ftdjz (mult=100×, +Mech7): F0=38, F1 collapse @ 2040
- smoke-test-d25vq (mult=5×): F0=21, F1 collapse @ 2820
The multiplier was searching the wrong dimension — the anchor itself
was stale.
Three coordinated changes (per feedback_no_partial_refactor):
1. RESTORE Mech 6 multiplier 5× → 100×. The original 100× was correct
for steady-state; the F1 saturation was driven by anchor staleness,
not multiplier looseness. Tightening it harmed F0 (over-clip during
ramp-up when slow_ema lagged grad_norm_ema).
2. ADD reset_grad_norm_slow_ema method on GpuDqnTrainer. Zeroes the
mapped-pinned scalar. First step of new fold builds up slow_ema
fresh, with MIN_CLIP=1.0 floor active during the brief transient.
3. WIRE Mech 8 call in fused_training.rs::reset_for_fold, alongside
Mech 4's existing Adam resets. Mech 6's anchor now aligns with
the new fold's grad scale from step 1 — same philosophy as Mech 3
(target net hard-sync) and Mech 4 (Adam EMA reset).
Net SP3 design: Mech 6 stays at 100× multiplier (broad, principled
headroom), Mech 8 keeps the anchor honest. The pair is more robust
than either change alone:
- Without Mech 8: anchor is stale, multiplier tuning has no winning
setting (5× hurts F0 ramp, 100× lets F1 saturate at slot 36).
- With Mech 8: anchor is fresh per fold; 100× multiplier provides
legitimate per-step headroom over CURRENT fold's grad norm.
Mech 7 stays reverted (per-element clip was misdiagnosis — over-clipped
legitimate gradient outliers without addressing the saturation root
cause).
F0 risk: low — F0 starts with slow_ema=0 anyway (cold start), so Mech 8
is a no-op on F0. Only changes F1+F2 fold-boundary behavior.
F1+F2 expectation: Mech 6's upper_bound now scales with the CURRENT
fold's grad norm, providing legitimate ~10× headroom per step without
allowing Adam EMA saturation. Slot 36-42 should stay quiet.
|
||
|
|
f67ede94fb |
fix(dqn): SP3 Mech 6 v2 — tighten clip multiplier 100x -> 5x; revert Mech 7
Smoke smoke-test-ftdjz (commit |
||
|
|
d9a4d98a3d |
feat(dqn): SP3 Mech 7 — per-element gradient clip in Adam kernel
Smoke smoke-test-fxvkk (commit
|
||
|
|
48c25d9997 |
feat(dqn): SP3 Mech 6 — anchored upper bound on adaptive grad clip
Adds an upper bound to update_adaptive_clip's new_clip formula to
prevent clip-drift pathology: consecutive elevated samples were
ratcheting the EMA-driven clip threshold upward without bound,
eventually rendering clipping a no-op against in-distribution drift.
Smoke smoke-test-5rqzs (commit
|
||
|
|
b9edccfc19 | docs(dqn): SP3 Phase F — audit summary of B1-B7 + Gate 2 prep | ||
|
|
5f7db7d8d8 |
feat(dqn): SP3 — name table entries for slots 36-47 diagnostic
Replaces rsv36-rsv47 placeholders with the SP3 Mech 5 diagnostic slot names. Both name table sites (halt_nan + halt_grad_collapse) updated identically per shared-contract migration principle (feedback_no_partial_refactor). Slot names mirror the audit doc table: - 36-39: Adam m max-abs (trunk/value/branch/IQN) - 40-43: Adam v max-abs (trunk/value/branch/IQN) - 44-45: Weight max-abs (trunk/heads) - 46: target_q post-clip - 47: atom_span_max When the fused kernel sets a slot bit, the readback log line names the buffer that exceeded its ISV-derived threshold — providing direct observability for SP3 mechanism effectiveness. |
||
|
|
45e077188f |
feat(dqn): SP3 Mech 5 — fused kernel extended for slots 36-47 threshold checks
Extends dqn_nan_check_fused_f32_kernel to handle 24 slots (12 NaN-only + 12 ISV-threshold). Per-slot thresholds computed inline in the kernel from a single q_abs_ref_eff arg (no HtoD per step; no separate thresholds buffer needed). Slot index → threshold mapping: - 12-15 (slots 36-39): Adam m ≥ 100 × q_abs_ref_eff - 16-19 (slots 40-43): Adam v ≥ 1e6 × q_abs_ref_eff² - 20-21 (slots 44-45): Weight max ≥ 1e3 × q_abs_ref_eff - 22 (slot 46): target_q ≥ 95 × q_abs_ref_eff (= 9.5 × max_abs_target_q) - 23 (slot 47): atom span ≥ 190 × q_abs_ref_eff (= 9.5 × max_atom_abs × 2) nan_check_buf_ptrs/lens resized 12→24 entries (mapped-pinned host write at construction). populate_nan_check_meta extended with 4 new args for IQN Adam m/v ptr+len (Option<u64>/Option<usize> — null when IQN inactive). Grid: 12 → 24 blocks. Single launch covers all 24 backward-path diagnostic slots. Slot 31 (deferred) + slots 33-35 (inline elsewhere) + slots with null IQN entries no-op via the kernel's null-pointer guard. q_abs_ref_eff = max(isv[Q_ABS_REF=16], 1.0) — ε on multiplier per SP1 pearl; cold-start ISV (~0) gives q_abs_ref_eff = 1, so all thresholds floor at their ε-floor multiplier × 1. SP3 Mech 5 closes the diagnostic instrumentation loop — slots 36-47 fire when their threshold is exceeded, providing observability for SP3's other 4 mechanisms' effectiveness. Fail-safe: if SP3 fix doesn't fully resolve F1 NaN, slot 36-47 firing pattern guides the next iteration. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
ef429c25d8 |
feat(dqn): SP3 Mech 4 — comprehensive Adam EMA reset (audit + GpuAttention wire)
Audit of all Adam optimizer state in crates/ml/src/ against the existing reset_adam_state calls at the fold-boundary in fused_training.rs::reset_for_fold: - DQN main (line 971) — covers m_buf/v_buf, IQN-trunk, q_attn, sel, denoise, mamba2, ofi_embed, PopArt, Q-div EMA via GpuDqnTrainer::reset_adam_state. Also covers VSN params (in main params buffer slots) and aux next_bar/regime heads (slots 119-126). - IQN head (line 1011) - TLOB (line 1026) - IQL-high (line 1031) - IQL-low (line 1036) Missing wire identified and added: - GpuAttention (4-head feature attention on h_s2): owns a separate (attn_m, attn_v, attn_adam_step) tuple in gpu_attention.rs, called every training step via attn.adam_step at fused_training.rs:1962/1995, but was never zeroed at fold boundary. Same pathology as IQN/TLOB/IQL — fold N momentum oversizes fold N+1's first-epoch SDP + output projection updates, compounding through the trunk gradient. Implementation: - Add GpuAttention::reset_adam_state mirroring the gpu_iqn_head and gpu_tlob pattern (memset_zeros m/v + zero step counter + write 0 through pinned t_pinned for next adam_step launch). - Wire it under `if let Some(ref mut attn) = self.gpu_attention` in reset_for_fold, immediately after the IQL-low reset, with the same warn-on-failure / info-on-success log pattern. - Append SP3 Task B5 entry to docs/dqn-wire-up-audit.md documenting the audit + fix (Invariant 7 contract). Out-of-scope optimizers (documented in audit, not changed): - DecisionTransformer: scoped within DT pretrain block, dropped at end of pretrain (no fold leak possible). - GpuCuriosityTrainer: owned by GpuExperienceCollector, external to FusedTrainingCtx. Reset belongs at the collector layer if needed. - MetaQNetwork: host-side observability MLP for collapse prediction (does not feed LearningHealth or trunk gradients). - GpuMoeHead: forward-only, no Adam state of its own. Comprehensive coverage now per feedback_no_partial_refactor — every in-context Adam state buffer zeros at fold transition. SP3 Mech 4 closes the "Adam EMA accumulation drives weight pathology" pathway. |
||
|
|
96b77043e0 |
feat(dqn): SP3 Mech 2 — C51 atom-position growth bounds
Clamps C51 atom positions during all dynamic write paths to ±10 × ISV[Q_ABS_REF=16].max(1.0) via inline fminf(fmaxf(...)): - atoms_update_kernel.cu — active GPU-driven shared atom grid (all 4 branches via grid.x = branch_id) - iql_value_kernel.cu::iql_compute_per_sample_support — per-sample [v_min, v_max, delta_z] tile (delta_z recomputed from clamped span) - experience_kernels.cu::adaptive_atom_positions — legacy single-branch entry kept in lockstep per feedback_no_partial_refactor Same ISV bound as Mech 1 (target_q clip) — atoms and target_q share the magnitude scale. ε on the multiplier (`isv.max(1.0)`) per the SP1 ε-floor pearl; bound ≥ 10 even with cold-start ISV. Closes the second of two Q-target inflation pathways: with Mech 1 capping the projection target and Mech 2 capping the atom support, the C51 expected_q (mean of atom × prob) is bounded by ±10 × ISV[16].max(1.0) regardless of probability mass distribution. Prevents the Adam EMA saturation → weight pathology → cuBLAS overflow chain. Reuses isv_signals already in scope at all three kernels — no new kernel arg, no new launch site, no new ISV slot, no new buffer, no new kernel. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
27ed7daa6a |
feat(dqn): SP3 Mech 1 — target_q clipping at single-point source
Clips denoise_target_q_buf to ±10 × ISV[Q_ABS_REF=16].max(1.0) after compute_denoise_target_q. Reuses dqn_clamp_finite_f32_kernel from SP1 (no new kernel). Single-point clip covers all downstream consumers (C51 / IQN / MSE). ε on multiplier per SP1 pearl; bound at least 10 even with cold-start ISV. F0 paper-review: F0-typical |target_q| ≤ 10; with Q_ABS_REF EMA ≈ 1-5 at F0, max_abs = 10-50 — F0 guard no-op. F1 inflation (thousands+) gets clamped, breaking the Q-target inflation pathway that drives Adam EMA saturation → weight pathology → cuBLAS overflow. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
aee11f3210 |
feat(dqn): SP3 — slot 36-47 diagnostic accessors (Task B1)
Adds pub(crate) accessor methods on GpuDqnTrainer and pub accessors
on GpuIqnHead exposing device pointers for the SP3 Mech 5 threshold
checks at slots 36-47:
GpuDqnTrainer (pub(crate)):
36: trunk_adam_m_ptr/_len — m_buf trunk slice (tensors [0..13))
37: value_adam_m_ptr/_len — m_buf value-head slice (tensors [13..17))
38: branch_adam_m_ptr/_len — m_buf branch-head slice (tensors [17..33))
40: trunk_adam_v_ptr/_len — v_buf trunk slice
41: value_adam_v_ptr/_len — v_buf value-head slice
42: branch_adam_v_ptr/_len — v_buf branch-head slice
44: trunk_params_ptr/_len — params_buf trunk slice
45: heads_params_ptr/_len — params_buf value+branch concat
46: target_q_ptr/_len — denoise_target_q_buf [B, 12]
47: atom_positions_ptr/_len — atom_positions_buf [4, num_atoms]
GpuIqnHead (pub):
39: adam_m_ptr/_len — IQN Adam first-moment buffer
43: adam_v_ptr/_len — IQN Adam second-moment buffer
DQN Adam state is UNIFIED in m_buf/v_buf (single TOTAL_PARAMS-sized
buffer covering trunk + heads at the same offsets as params_buf).
Slot 36/37/38 (and 40/41/42) accessors return pointers into the same
buffer at offsets computed via padded_byte_offset over the existing
compute_param_sizes layout. The kernel discriminates by slot index
for threshold checks. IQN Adam state lives separately on GpuIqnHead.
Trunk/heads param-buf split (slots 44/45) uses the same offset helper
— no new buffers, no new ISV slots, no DtoD/HtoD copies. Each ptr
accessor returns self.<field>.raw_ptr() (or +offset for slices); each
len accessor returns self.<field>.len() or computed from
compute_param_sizes.
Audit doc updated (docs/dqn-wire-up-audit.md SP3 Task B1 entry).
Used by populate_nan_check_meta_v2 (Task B6) to feed the fused
kernel's threshold-check entries when extended to 24 slots. Unused
yet — wired in B6.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
05048d03eb | docs(dqn): SP2 Gate 1 partial pass — proceed to Phase B per Option B | ||
|
|
e270b5cc42 | docs(dqn): SP2 Phase E — audit doc summary of A1-A4 + Gate 1 prep | ||
|
|
b5a064f6ca |
feat(dqn): SP2 — replace 8 check_nan_f32 calls with fused launch
run_nan_checks_post_backward now fires a single fused kernel launch covering slots 24-35 in 12 blocks. Reduces graph-capture overhead from 8 launches to 1 per step. Slots 33-35 inline checks at backward orchestration sites (gpu_dqn_trainer.rs:18580/:18601/:6939) unchanged — multi-point localization preserved. Method signature simplified — IQN pointers no longer per-step args. populate_nan_check_meta (called once at construction in fused_training.rs) baked Option<u64> nullity into the metadata buffer entries; the fused kernel's null-pointer guard handles slots 27/28 when IQN inactive without per-step branching at the Rust caller. Both call sites in fused_training.rs (ungraphed + graph-captured paths) updated together per feedback_no_partial_refactor. This is the F0 regression fix — Phase B's per-step kernel-launch overhead was the F0 cause across 3 SP1 smokes (F0 ~ 35 vs baseline 55.87). Gate 1 smoke (Task A6) validates F0 >= 53.08. |
||
|
|
fbf48df9de |
feat(dqn): SP2 A3 — fused NaN check populate + launch wrapper
Replaces A2's CudaSlice<u64>/<i32> field types with MappedU64Buffer/ MappedI32Buffer per feedback_no_htod_htoh_only_mapped_pinned. Mapped- pinned eliminates the HtoD copy entirely — the kernel reads via the device-mapped pointer (cuMemHostGetDevicePointer_v2) while the trainer writes through the same mapped pages on the host side. Adds populate_nan_check_meta on GpuDqnTrainer (one-shot construction- time write of 12 (ptr, len) tuples for slots 24-35). Slot 31 deferred (null entry); slots 27/28 nullable on Option<u64> (None when IQN inactive); slots 33-35 null (inline checks fire separately at backward orchestration phases — kept individual for entry-point localization). Adds launch_nan_check_fused_f32 (per-step kernel launch wrapper with grid_dim=12, block_dim=256, base_flag_idx=24). Registers dqn_nan_check_fused_f32_kernel in compile_training_kernels (tuple 43→44, info log 38→39 utility kernels) — same module as the per-buffer dqn_nan_check_f32 to share the captured replay group. Constructor-time wire-up lands in FusedTrainingCtx::new after gpu_iqn construction (gpu_iqn is owned by FusedTrainingCtx, not GpuDqnTrainer — mirrors the same Option<u64> arg pattern used by apply_iqn_trunk_gradient and run_nan_checks_post_backward). Wrapper unused yet — call-site replacement (8 individual check_nan_f32 calls in run_nan_checks_post_backward → single fused launch) lands in A4. Audit doc updated (Invariant 7). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
82b6bd369e |
feat(dqn): SP2 — nan_check_buf_ptrs + nan_check_buf_lens device buffers
Adds 12-entry u64 ptr table + 12-entry i32 len table on GpuDqnTrainer for the fused NaN-check kernel. Allocated as zeros; populated once in next commit (slot pointers stable across trainer lifetime). Slot 31 (deferred ensemble) entry will be 0 — kernel skips that block. Two separate buffers chosen over packed-stride layout for direct ABI match with the kernel signature (const float* const* buf_ptrs, const int* buf_lens). Avoids stride-alignment concerns. Unused yet — populate + wrapper + call-site replacement in next commits. |
||
|
|
0250722a0d |
feat(dqn): SP2 — fused multi-buffer NaN check kernel
Adds dqn_nan_check_fused_f32_kernel: single-launch replacement for 8 per-step dqn_nan_check_f32 calls in run_nan_checks_post_backward. Block N processes slot base_flag_idx + N; sticky-flag semantics preserved; graph-capture safe; null buf_ptr = no-op (deferred slot 31). Unused yet — Rust wrapper + call-site replacement in next commits. |
||
|
|
9b35fe9d45 |
plan(dqn): SP2 + SP3 implementation — fused NaN kernel + 5-mechanism Q-stability
Implementation plan for the approved combined design spec
(commit
|
||
|
|
ed727b51c0 |
spec(dqn): SP2 + SP3 combined — F0 regression fix + Q-learning structural stability
Designs the fixes for the two SP1 leftover issues identified at SP1 closure
(commit
|
||
|
|
047f175fbf |
docs(dqn): SP1 closure — surgical-fix scope delivered; SP2/SP3 handoff
SP1 (F1 NaN root-cause investigation) closes at commit
|
||
|
|
ab2133463e |
fix(dqn): SP1 Phase C — ε floor fix-up #2 (cold-start clamp pathology)
Smoke smoke-test-dr2bn (commit
|
||
|
|
0c99e08002 |
docs(dqn): SP1 Phase D — multi-fold validation FAIL (criterion 7/7)
Smoke smoke-test-dr2bn (commit |
||
|
|
19b008e1cb |
fix(dqn): SP1 Phase C — preserve NaN diagnostics + correct F0 paper-review
Quality-review follow-ups to commit
|
||
|
|
97f1d25f54 |
fix(dqn): SP1 Phase C — F1 NaN cuBLAS-overflow surgical fix (slots 26 + 32)
Patches two cuBLAS GEMM backward operations identified by SP1 Phase B
smoke smoke-test-xvzgk (commit
|
||
|
|
2ba0eef718 |
docs(dqn): SP1 Phase B smoke result — F1 NaN topology captured (slots 26 + 32)
Smoke smoke-test-xvzgk (commit
|
||
|
|
f139a63eea |
feat(dqn): SP1 Phase B instrumentation — backward NaN checks (slots 24-30, 32-35)
Wires per-step NaN checks on backward-path kernel outputs. Coverage
per audit per-slot accessor table (docs/dqn-backward-nan-audit.md
:530-548).
GpuDqnTrainer::run_nan_checks_post_backward (NEW) — fire-once-at-end:
- 24 d_value_logits_buf (post-c51_grad value gradient)
- 25 d_adv_logits_buf (post-c51_grad branch advantage)
- 26 iqn_trunk_m (apply_iqn_trunk_gradient cuBLAS bwd output)
- 27 iqn_d_h_s2_buf (IQN backward dh_s2; arg from FusedTrainingCtx)
- 28 d_branch_logits_buf (IQN production backward; arg from caller)
- 29 cql_d_value_logits (CQL gradient output)
- 30 aux_dh_s2_nb_buf (Aux next-bar backward dh_s2)
- 32 bn_d_concat_buf (Bottleneck Linear backward dy)
Inline checks during backward orchestration:
- 33 bw_d_h_s2 post-main (in launch_cublas_backward_to, after
backward_full + branch concat accum,
BEFORE aux_heads_backward SAXPY)
- 34 bw_d_h_s2 post-aux (in launch_cublas_backward_to, after
aux_heads_backward SAXPY)
- 35 bw_d_h_s2 post-iqn (in apply_iqn_trunk_gradient, after
graph_safe_copy_f32 DtoD overwrite,
BEFORE encoder_backward_chain consumes)
Sequential 33→34→35 fire pattern localises the NaN entry point:
- 33 alone fires → main backward chain (c51 + MSE + branch concat)
- 34 fires after 33-clean → aux SAXPY (aux_heads_backward)
- 35 fires after 34-clean → IQN DtoD or per-sample IQN backward
Slot 31 (ensemble_d_logits_buf) cleanly deferred per Task 2 commit
|
||
|
|
797f8bf326 |
feat(dqn): SP1 Phase B accessors — backward-buffer NaN check pointers
Adds 5 new pub(crate) accessor methods exposing backward-path buffer
device pointers for Task 4's per-step NaN checks (slots 24, 25, 28,
29, 30 per audit per-slot table at docs/dqn-backward-nan-audit.md
:530-548):
GpuDqnTrainer (4 new):
- d_value_logits_buf_ptr (slot 24 — post-c51_grad value gradient)
- d_adv_logits_buf_ptr (slot 25 — post-c51_grad branch advantage)
- cql_d_value_logits_ptr (slot 29 — CQL gradient output)
- aux_dh_s2_nb_buf_ptr (slot 30 — aux next-bar backward dh_s2)
GpuIqnHead (1 new):
- d_branch_logits_buf_ptr (slot 28 — production IQN backward output,
iqn_quantile_huber_loss)
Slot 27 (iqn_d_h_s2_buf) reuses existing GpuIqnHead::d_h_s2_raw_ptr()
at gpu_iqn_head.rs:1660 — no new method per feedback_no_legacy_aliases:
the existing accessor is already public and sufficient; renaming +
chasing the single call site adds churn without value.
Slots 26, 32, 33-35 reuse pre-existing handles:
- 26: self.ptrs.iqn_trunk_m
- 32: self.bn_d_concat_buf() (existing, returns &CudaSlice<f32>)
- 33-35: self.ptrs.bw_d_h_s2 (3 different Task 4 call sites)
Slot 31 (ensemble_d_logits_buf) deferred per Task 2 commit
|
||
|
|
387335e2b9 |
fix(dqn): SP1 Phase B foundation — stale-doc cleanup + Task 4 prep
Quality-review follow-ups to commit
|
||
|
|
53bc0bc505 |
feat(dqn): SP1 Phase B foundation — nan_flags_buf 24→48
Expands the NaN flag buffer from 24 to 48 slots to make room for
backward-path NaN checks (slots 24-35 per audit doc per-slot table)
plus 12 reserved headroom slots (36-47) for SP2 framework + SP3
observer hooks.
Touches:
- gpu_dqn_trainer.rs: alloc size 24→48 (line ~11178); read_nan_flags
signature [i32; 24]→[i32; 48] (line ~14982); field docstring updated
(line ~2658) to reflect 48-slot layout
- fused_training.rs: pub(crate) read_nan_flags signature [i32; 24]→[i32; 48]
- training_loop.rs: BOTH name table sites (halt_nan + halt_grad_collapse
block from commit
|
||
|
|
2406ebf2f7 |
docs(dqn): SP1 Phase A — slot 24/25 buffer correction + dormancy reconciliation
Three review-driven fixes: 1. Slot 24/25 cite the actual production buffers d_value_logits_buf (gpu_dqn_trainer.rs:3123) and d_adv_logits_buf (:3125) — the f32 atomicAdd buffers that launch_c51_grad writes — not the staging buffers at :3127/:3129. Pointer expressions exist at the launch site (:17707/:17708), making new accessors optional. Task 3 summary and priority-list entries updated to match. 2. Explicit plan-supersession note inserted immediately above the per-slot table: the audit's per-slot allocation supersedes the plan Task 2 step 6 name table. The plan's allocation was placeholder; the audit's is grounded in per-kernel inspection. Task 2 should use the audit's names (d_value_logits_buf, d_adv_logits_buf, iqn_trunk_m, iqn_d_h_s2_ptr, d_branch_logits_buf, etc.). 3. Summary suspicion-ranking row #3 reconciled with slot-28 dormancy: iqn_backward_per_sample has no Rust caller (verified via grep crates/ml/src/), so row #3 is re-pointed at the production kernel iqn_quantile_huber_loss (iqn_dual_head_kernel.cu:1346-1413, loaded at gpu_iqn_head.rs:2114). Same unsafe-write pattern (no isfinite guard at line 1410's d_q_online[idx] = qw*d_huber/Q), now attributed to the live path. apply_iqn_trunk_gradient at #1 stands — its reasoning (orchestrator consuming iqn_d_h_s2_ptr) is unchanged by which specific kernel writes that buffer. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
8fd9cc0603 |
docs(dqn): SP1 Phase A — audit citation + accessor precision
Quality-review fixes for the SP1 Phase A audit: 1. nan_flags_buf [i32;24] declaration: quad-cite (field 2659, alloc 11178, constructor 12563, read_nan_flags signature 14982). Task 2's 24->48 expansion must touch all four atomically per feedback_no_partial_refactor; consumer site in trainers/dqn/fused_training.rs and the two name-table sites in trainers/dqn/trainer/training_loop.rs are still listed in Task 2's plan body. 2. Per-slot Rust buffer-pointer expression added for slots 24-35. Already-accessible (no new accessor): slots 26, 32, 33, 34, 35 (4 via self.ptrs, 1 via existing bn_d_concat_buf accessor). Need new accessor on GpuDqnTrainer: slots 24 (d_value_logits), 25 (d_adv_logits), 30 (aux_dh_s2_nb_buf), and 29 (cql_d_value_logits, only if un-deferred). Need new accessor on GpuIqnHead: slot 28 (d_branch_logits_buf — note: production IQN backward uses iqn_quantile_huber_loss, NOT iqn_backward_per_sample which is declared but never loaded). Optional: slot 27 (d_h_s2_buf_ptr) — can be inlined inside apply_iqn_trunk_gradient instead. Need new accessor on FusedDqnTraining: slot 31 (only if un-deferred). Becomes input for Task 3. 3. Citation typo: c51_loss_kernel.cu line 274 reference removed — line 274 is __syncthreads() in the projection-reduction warp loop; the a_std=sqrtf reference belongs to c51_grad_kernel.cu:274. Loss-kernel sqrtf sites are 779 and 804. Verified by reading each cited line; SQLX_OFFLINE=true cargo check --workspace passes (docs-only). |
||
|
|
6994b9cf97 |
docs(dqn): SP1 Phase A — backward NaN audit
Read-only γ audit of every backward-path kernel writing to bw_d_h_s2 / grad_buf / save_h_s2 accumulators. Per-kernel inventory against unsafe-pattern checklist (sqrtf-neg, 1/0, logf-≤0, expf-large, EMA variance, atomicAdd/saxpy NaN-propagation). Each section includes: identified pattern + line citation, severity, proposed guard form, ISV bound option, F0 risk, Phase B flag-slot allocation. Cross-referenced session_2026-04-05_nan_investigation.md residual 8% step-2 NaN in apply_iqn_trunk_gradient — never closed; ranked top suspect for SP1. Output drives SP1 Phase B (12 new flag slots in nan_flags_buf 24→48) and Phase C surgical fix(es). Durable artifact for SP2 framework codification + SP3 structural-fix scoping. |
||
|
|
05958d3a0c |
plan(dqn): SP1 numerical stability — F1 NaN root-cause investigation
Implementation plan for SP1 (Sub-project 1 of 3) of the numerical stability investigation. Follows the γ + β methodology from the spec at docs/superpowers/specs/2026-04-29-numerical-stability-investigation-design.md. 8 tasks across 4 phases: - Phase A (Task 1): γ read-only audit producing docs/dqn-backward-nan-audit.md - Phase B (Tasks 2-5): always-landing β instrumentation expanding nan_flags_buf 24→48 with 12 new backward-kernel NaN check slots + 12 reserved slots for future coverage - Phase C (Task 6): surgical fix(es), content-driven by audit + smoke topology, ISV-driven for any dynamic bound (mandatory) - Phase D (Task 7): multi-fold L40S smoke validation against 7 pass criteria (F0 ≥ 95% baseline, F1+F2 monotone improvement, zero NaN-CLAMPED-TO-ZERO, all 48 NaN flag slots remain at zero) - Closure (Task 8): audit doc closure, memory entry, SP2/SP3 handoff Operating principles (mandatory per spec): - No deferrals — anomalies discovered during investigation get fixed within SP1, not punted to SP2/SP3 - Combined RELATED fixes ship as rich commits (per feedback_no_partial_refactor) - ISV-driven design for any dynamic bound Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
792812baa1 |
spec(dqn): SP1 numerical stability — revise per critical review
9 substantive issues addressed inline:
1. ISV-driven design elevated from 'if applicable' to MANDATORY for
all dynamic bounds in SP1 fixes. Numerical-stability ε bounds are
the only carve-out (Invariant 1). Hardcoded tuning constants for
dynamic ranges explicitly rejected.
2. F0 Sharpe regression criterion changed from absolute (≥55) to
ratio-based (≥95% of latest baseline; floor 53.08 currently).
Prevents iterative erosion across multiple fix commits.
3. 'F1 trending positive' replaced with concrete monotone-improvement
test: Best Sharpe at last epoch ≥ Best Sharpe at first epoch of
the same fold.
4. Pass criterion distinguishes NaN-CLAMPED-TO-ZERO (failure) from
'Genuine grad collapse' (legitimate observation, permitted) per
the existing infrastructure from commit
|
||
|
|
dab5990287 |
spec(dqn): SP1 numerical stability investigation — F1 NaN root-cause
Brainstorm session 2026-04-29 produced this spec scoping Sub-project 1 of a three-sub-project decomposition addressing the F1 ep2 NaN explosion that persists across 9+ defensive layers in Plan C Phase 2 (commits e445d07a..e9096c7be). Decomposition: - SP1 (this spec): F1 NaN root-cause; γ audit + β always-landing instrumentation + surgical fix + multi-fold validation - SP2 (future): numerical stability framework — codify guards - SP3 (future): Q-learning structural stability — target-Q clip, pessimistic ensemble, atom-range governance Operating principles: - No deferrals (anomalies fixed within SP1, not punted) - Combined fixes — rich commits (per feedback_no_partial_refactor) - Always-landing diagnostic instrumentation (24-slot nan_flags_buf expands to 48; permanent regression sentinel) - F0 Best Sharpe ≥ 55 preserved (no regression on the working path) Pass criterion: all 3 folds train 5 epochs, F0+F1+F2 Sharpe ≥ 0, zero NaN-CLAMPED-TO-ZERO log lines, all 48 flag slots remain zero. 5 design sections approved iteratively: Architecture, Components, Data Flow, Error Handling, Testing. Next: writing-plans produces implementation plan. |
||
|
|
e9096c7be1 |
fix(dqn): R1 — eliminate K Adam shrink + P — GRN-stage NaN checks
R1: K's hardcoded shrink-and-perturb (m×0.1, v×0.01) at fold boundary
violated feedback_adaptive_not_tuned (untracked tunable knobs) AND
created a downstream pathology: tiny v_hat denominator → oversized
Adam updates 50+ steps post-reset → trunk param overshoot → save_h_s2
NaN at F1 ~step 1745 (smoke-test-bkdx5 diagnostic).
K was introduced (commit
|
||
|
|
d1808df14c |
diag(dqn): read NaN flags on grad-collapse path too (find F1 ep2 explosion source)
Existing nan_flags_buf [16] covers 13 buffers with per-step NaN checks inside the captured training graph. But read_nan_flags() was only called when training guard's halt_nan fired — which checks pinned- readback grad_norm. NaN-clamped-to-zero gradients reach the pinned scalar as 0, not NaN, so halt_nan never fires for the explosion case. The collapse path (halt_grad_collapse, grad < 0.01) was firing instead without reading the flags. Add flag readback in that path when gr.raw_grad_norm < 1e-6 (suspicious zero) — logs NaN-CLAMPED-TO-ZERO with flagged buffer names. Genuine near-zero gradients get a separate "no NaN flags set" log so we can distinguish the two cases. This should pinpoint which kernel produces the F1 ep2 NaN first (C51 KL projection? IQN aux? CQL? aux heads?). Diagnostic only — keep after fix lands; correct gate for future regressions. |
||
|
|
7640a681c5 |
diag(dqn): per-step F1 explosion diagnostic — pinpoint which signal blows first
After 6+ intervention layers (A.1+A.2+A.3+F+H+K+warmup+N) F1 ep2 still
explodes with grad_norm=2.66T while F1 ep1 trains cleanly. None of our
defenses catch the explosion path. Need data to pinpoint WHICH signal
explodes first.
Per-step FOLD_EXPLOSION_DIAG in fold >= 1: when grad_norm > 1000 OR
jumps 5x from prior guard step, fire diagnostic warn with:
- loss decomposition: total / c51 / mse / iqn (pinned readback, no DtoH)
- Q range: q_min / q_mean / q_max via cold-path reduce
- atom positions: per_sample_support[0,d0] + [0,d2] (v_min, v_max, dz)
- prev grad norm + ratio for context
- trigger flag + post-trigger trajectory countdown
After trigger, continues emitting for 5 guard steps so the explosion
trajectory is captured (not just the first crossing).
Plus an unconditional FOLD_EXPLOSION_DIAG[F1_END_EP1] baseline log at
the end of fold 1 epoch 0 — healthy state immediately preceding the
explosion. Compare-and-contrast with per-step explosion frames pins
the runaway driver.
CQL and ensemble losses are not pinned-readback (transient device
buffers consumed inside the training graph) and are explicitly absent
from the decomposition. If none of {c51, iqn, mse} is the runaway
driver but total_loss still explodes, that implicates the unexposed
CQL/ens path — the absence is itself diagnostic information.
Implementation:
- Three diagnostic-only fields on DQNTrainer: current_fold,
last_logged_grad_norm, explosion_diag_steps_remaining. Reset
last_logged + remaining in reset_for_fold; current_fold
overwritten unconditionally at fold-loop entry.
- Three accessors on FusedTrainingCtx: explosion_diag_loss_components
(pinned), explosion_diag_atom_range (12-elem cold DtoH),
explosion_diag_q_range (reduce + 28B DtoH).
- Three accessors on GpuDqnTrainer: c51_loss_pinned_value,
mse_loss_pinned_value, stream_for_diag.
Diagnostic-only — remove once root cause identified. Not for prod.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
250cee124e |
fix(dqn): Winsorize adaptive_clip EMA input against single-sample outliers (N)
Smoke smoke-test-xw4c6 showed F1 ep2 grad_norm=8.4B polluting the
adaptive_clip EMA -> next adaptive_clip threshold became huge ->
subsequent extreme grads passed unclipped -> NaN propagation -> grad
collapse -> fold 1 fails despite K+warmup fixing the first-step problem.
Add Winsorized clamp on raw_grad_norm before the adaptive_clip EMA
update: clamped_sample = min(raw, K * previous_adaptive_clip) where
K = 100 (numerical-stability bound, not a tuned constant — explicit
"single sample can grow EMA at most 100x in one step" semantics).
Companion observability log emits GRAD_CLIP_OUTLIER warning per clamp
event so we can see when this fires in subsequent smokes.
Fast/slow grad_norm EMAs (driving warmup factor) intentionally NOT
winsorized — they're a stability signal that SHOULD respond to
outliers, providing extra warmup damping when the system is unstable.
Predicted impact on Plan C smoke F1 ep2: 8.4B grad -> clamped to
~3000 for EMA -> next adaptive_clip caps at ~1000 -> subsequent F1
batches get bounded gradients passed to Adam -> no NaN propagation.
F1+F2 should now train through, exposing whether the underlying
Q-target inflation requires further structural work or whether
clipping alone suffices.
Composes with K+warmup (
|
||
|
|
4ef1d8ebb7 |
fix(dqn): Plan C K — Adam shrink-and-perturb + adaptive fold-warmup ISV
Closes the F1 ep1 catastrophic-overshoot gap exposed by smoke-test-s9h4h
(F0 succeeded with Best Sharpe = 36.03; F1 ep1 grad_norm = 355,009 — 5
orders of magnitude larger than F0 steady-state ~10 — leading to NaN
propagation and grad-clamp-to-zero early stop). Combined fix: (1) Adam
shrink-and-perturb at fold boundary (replace m=0/v=0 with m*=0.1, v*=0.01
to preserve direction while damping magnitude), and (2) a single
adaptive ISV signal driving BOTH lr_eff and clip_eff dampening over
the fold's first ~50 steps.
K — Adam shrink-and-perturb in `GpuDqnTrainer::reset_adam_state`:
- m *= 0.1, v *= 0.01 via existing `dqn_scale_f32_kernel` (loaded as
`scale_f32_ungraphed`); t_pinned still zeroed so bias correction
restarts. Architectural constants (preserve direction / lose magnitude
history) per `feedback_isv_for_adaptive_bounds.md` Invariant 1
carve-out — not tuned. Composes with existing param shrink-and-perturb
(`alpha=0.8`) in `FusedTrainingCtx::reset_for_fold`. Root cause for
F1 overshoot: m=0,v=0 → first Adam step ≈ lr × g / ε → 6 OoM
amplification.
New CUDA kernel `fold_warmup_factor_kernel.cu`:
- Single-block single-thread cold-path producer mirroring
`q_drift_rate_ema_kernel.cu` / `moe_lambda_eff_kernel.cu` shape.
- Reads two grad-norm EMAs (fast α=0.1, slow α=0.001) plus host-passed
step counter; writes ISV[FOLD_WARMUP_FACTOR_INDEX=130] = clamp(fast/slow, 0, 1).
- Bootstrap branches (steps_observed < 200, slow EMA < 1e-6) emit
factor=1.0 (no damping during cold-start). No atomicAdd; no DtoH.
New ISV slot Q_DRIFT_RATE_INDEX → FOLD_WARMUP_FACTOR_INDEX = 130:
- ISV_TOTAL_DIM 130 → 131; layout fingerprint shifts (checkpoint-
incompatible per `feedback_no_legacy_aliases.md`, expected for a
real architecture change).
- FoldReset entries: `isv_fold_warmup_factor` → 0.0 and companion
`isv_grad_norm_fast_ema` → 0.0 (lockstep reset per
`feedback_no_partial_refactor.md`); slow EMA persists across folds
as the cross-fold steady-state baseline.
- Two new mapped-pinned scalars on GpuDqnTrainer (grad_norm_fast_ema_pinned,
grad_norm_slow_ema_pinned) fed by `update_adaptive_clip` from the
same `gr.raw_grad_norm` observation source as the existing adaptive
clip EMA.
Two consumers, both monotone (only dampen, never excite):
- lr_eff = cosine_effective_lr × max(MIN_WARMUP_LR_FRAC=0.05, factor)
via `set_lr` per-step. New `cosine_effective_lr_base` field
on DQNTrainer composes the cosine schedule's per-epoch
baseline with the warmup factor's per-step damping (rather
than overriding the cosine schedule).
- clip_eff = clip_base × (MIN_CLIP_FRAC=0.1 + 0.9 × factor) via new
`set_active_clip` setter on FusedTrainingCtx + GpuDqnTrainer.
Composes with the EMA-derived `clip_base = grad_norm_ema × 2`
that `update_adaptive_clip` just wrote to the pinned slot.
Numerical-stability bounds 0.05 / 0.1 are Invariant 1
carve-outs.
Steady-state behaviour unchanged: factor=1 → lr_eff=lr_base,
clip_eff=clip_base. Fold-boundary behaviour: factor starts at 0 →
lr_eff = 0.05 × lr_base, clip_eff ≈ 0.1 × clip_base; rises to 1 over
~50 steps as the fast EMA catches up to the slow steady-state EMA.
Predicted impact on Plan C smoke F1: 355,009-magnitude transient grad
clipped to ~clip_base × 0.1 ≈ 1.0 (vs 10), Adam state shrunk instead
of zeroed → first step update bounded; grad recovers normally over
~50 steps. Companion to A.1 (prev_epoch_q_mean reset), A.2 (adaptive
Polyak-tau), A.3 (gradient_collapse_counter reset), F+H (kill-criterion
robustness) — completes the fold-boundary state-reset family.
Per `pearl_adaptive_moe_lambda.md` (kernel + ISV slot + bootstrap +
reset + observability template), `pearl_cold_path_no_exception_to_gpu_drives.md`
(GPU-stays-on-GPU even at cold-path cadence),
`pearl_blend_formulas_must_have_permanent_floor.md` (lr/clip floors
are permanent minimums), `feedback_adaptive_not_tuned.md` (lr+clip
ISV-driven), `feedback_isv_for_adaptive_bounds.md` (factor IS the
bound; consumers compose at runtime), `feedback_no_atomicadd.md`
(single-thread reduce), `feedback_cudarc_f64_f32_abi.md` (slot index
passed as i32), `feedback_no_partial_refactor.md` (kernel + slot +
reset + producer + 2 consumers all land together),
`feedback_no_quickfixes.md` (replaces brittle full-reset with
adaptive damping; not threshold relaxation).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
ef243e771d |
fix(dqn): reset gradient_collapse_counter + training_steps at fold boundary (A.3)
Smoke smoke-test-vh9bj revealed: after F+H Q-drift kill fix, fold 0 trains successfully (5 epochs, Best Sharpe 36.03), but fold 1 fails at epoch 1 with "Gradient collapse detected for 5 consecutive epochs" despite epoch 1 grad_norm=296,417 (healthy). Root cause: the per-step gradient_collapse_counter on DQN was already near patience (5) from fold 0's late-epoch near-zero-grad steps, and DQNTrainer::reset_for_fold didn't clear it. Plus, training_steps accumulating across folds made the `past_warmup` gate always-true in fold 1+, removing the warmup grace period for data distribution shifts. Add DQN::reset_for_fold zeroing both. Wire from DQNTrainer::reset_for_fold alongside the A.1 prev_epoch_q_mean reset. Pure additive — same fold-boundary state-reset gap pattern as A.1, A.2, F. Predicted impact: fold 1+ now starts with counter=0 and warmup window restored; gradient collapse check has its full per-fold grace period. |
||
|
|
cca9dd36ae |
fix(dqn): replace q_mean ratio with rolling-window MAD deviation (H — kill criterion robustness)
Current criterion uses |q_mean| / |prev_q_mean| which explodes near zero crossings (legitimate cold-start has |prev_q_mean| ≈ 0.01-0.1, making any non-tiny current trigger the ratio threshold). smoke-test-n9xzr fired with prev=−0.0795, curr=0.7647, ratio=9.62× despite this being natural cold-start growth, not geometric runaway. Replace with rolling 5-epoch window of q_means. Compute median + MAD (Median Absolute Deviation, a 50%-breakdown estimator robust to single outliers); kill condition becomes |q_mean − median(window)| > 4.0 × max(MAD(window), 0.01) AND the existing adaptive floor. Constants are declared `const` near the kill block: - Q_DRIFT_WINDOW_SIZE=5 (matches smoke fold length, gives MAD a meaningful estimator without averaging across regime shifts) - Q_DRIFT_WARMUP_SAMPLES=3 (skip until ≥ 3 priors — smaller window degenerates to half-range MAD that trips on monotonic trajectories) - Q_DRIFT_DEVIATION_THRESHOLD=4.0 (4 MADs ≈ 2.7σ Gaussian- equivalent; clear outlier without firing on every legitimate dip; literal MAD count rather than σ because q_mean is non- Gaussian during cold-start) - Q_DRIFT_MIN_DEVIATION=0.01 (numerical-stability floor when window is constant) Window resets at fold boundary alongside prev_epoch_q_mean / adaptive_tau (A.1 pattern — cross-fold q-stats are independent training runs, mixing them would inflate MAD or shift the median). Predicted impact on smoke-test-n9xzr: - Plan C smoke fold 0 ep2: window has only 2 priors, warmup gate not yet satisfied → kill stays silent (was firing on ratio=9.62×) - Genuine geometric runaway (q_mean → 62.5 from baseline 0.5 over 4 epochs): ep3 deviation = |62.5 − 2.5|/MAD=2.0 = 30 > 4 AND |62.5| > floor=3×12=36 → kill fires correctly Both conditions ANDed (floor + deviation), preserving the production-safety semantics of the original criterion. The floor check is unchanged; only the divergence detector is replaced. Per feedback_no_quickfixes.md: this is a principled robust- statistics replacement, not a threshold relaxation. Per feedback_no_partial_refactor.md: window field, constants, criterion site, and fold-boundary reset land in lockstep — the kill criterion's contract is internally consistent across trainer/mod.rs, constructor.rs, and training_loop.rs. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
790c100719 |
fix(dqn): wire Q_ABS_REF / Q_DIR_ABS_REF EMA producers per-step (F — kill criterion adaptive floor)
The Q-drift kill criterion's adaptive floor formula is max(0.5, 3 × max(ISV[Q_ABS_REF=16], ISV[Q_DIR_ABS_REF=21])) but smoke-test-n9xzr showed both ISVs stuck at bootstrap 0.05 even when q_mean reached 0.76 — the producers (q_mag_bin_means_reduce and q_dir_bin_means_reduce) only fired inside reduce_current_q_stats at epoch boundary, while the per-step captured update_isv_signals consumed the resulting scratch buffers each training step. Concrete failure mode at α=0.05 EMA with raw ≈ 1.0: - Epoch 0: scratch is zero-initialized, all per-step EMA pulls toward zero, ISV[16,21] stay 0.0 - Epoch 0 boundary: reduce_current_q_stats fires, scratch becomes q_abs_ref ≈ 1.0, single update_isv_signals writes ISV[16] = 0.05 - Epoch 1+: with stale scratch held constant between boundaries, per-step EMA over hundreds of steps would saturate — but smoke step counts are small enough (small batch=64, buffer=256) that the slot stays near 0.05 - kill_floor degenerates to max(0.5, 3 × 0.05) = 0.5 forever, so the criterion fires on legitimate cold-start growth Fix: launch q_mag_bin_means_reduce + q_dir_bin_means_reduce in both fused_training paths (captured adam_update_child at ~line 2228 AND ungraphed step-0 fallback at ~line 1465) immediately before update_isv_signals. Per-step graph replays now read q_out_buf that forward_child populated this same step, write fresh scratch, and the captured update_isv_signals consumes it on the same stream. Per pearl_cold_path_no_exception_to_gpu_drives.md: cold-path EMAs fire alongside their consumers. Per feedback_no_partial_refactor.md: graphed and ungraphed paths migrate together — same producer- consumer contract. After this fix, the floor will adapt: ISV[16,21] saturate to the policy's actual q_abs_ref scale within ~60 steps at α=0.05, so by epoch 2 the floor becomes 3 × ~0.5 = ~1.5, no longer firing on legitimate cold-start growth. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
f64eeb64da |
fix(reward): bound bonus shaping by ISV[Q_DIR_ABS_REF] to break optimism loop
C.4 timing bonus (experience_kernels.cu) and D.4b regime penalty multiplied by the trade-cumulative |reward| / |final_pnl| as an unbounded multiplicand. Across trades this compounds — bonus values inflate as Q values inflate, then bonus rewards inflate Q values further (Bellman bootstraps off shaped reward). Per pearl_one_unbounded_signal_per_reward + feedback_isv_for_adaptive_bounds: replace |reward|/|final_pnl| with ISV[Q_DIR_ABS_REF_INDEX]-bounded variant. The unbounded multiplicand becomes the direction-branch Q-scale EMA (already-tracked, gradient-decoupled), not the trade-cumulative shaping output. Cap is adaptive (matches Q magnitude as it evolves) and breaks the multi-trade compounding loop. Discovered during Plan C Phase 2 smoke diagnosis (researcher report 2026-04-29 a25f669e9df953174). |
||
|
|
b4d4a8d046 |
feat(dqn): Plan C A.2 — adaptive Polyak tau coupled to q-drift rate
Wires the q_drift_rate_ema kernel + ISV slot Q_DRIFT_RATE_INDEX=129 (landed in |
||
|
|
fc8dbb0a85 |
feat(dqn): q_drift_rate_ema kernel + Rust wrapper (Plan C A.2 scaffolding)
Single-block single-thread ISV producer mirrors h_s2_rms_ema /
moe_lambda_eff pattern. Computes per-epoch
drift_rate = |q_mean(t) - q_mean(t-1)| /
max(|q_mean(t-1)|, ISV[Q_ABS_REF] + ISV[Q_DIR_ABS_REF], 1e-6)
clipped to [0, 4] and writes to ISV[Q_DRIFT_RATE_INDEX=129].
Includes the full ISV-contract shift required for the kernel to load:
- ISV_TOTAL_DIM 129 -> 130
- Q_DRIFT_RATE_INDEX = 129 (tail-appended after MOE_LAMBDA_EFF=128)
- Cold-start ISV[129] = 0.0 in constructor (no-op dampening factor)
- layout_fingerprint_seed entry Q_DRIFT_RATE=129 + ISV_TOTAL_DIM=130
- Cubin static, kernel field, kernel load, struct assignment
Per feedback_no_partial_refactor.md the ISV slot + dim + fingerprint
all migrate together (the wrapper references Q_DRIFT_RATE_INDEX so they
cannot be split). Tau consumer + state reset registry + per-epoch
producer launch land in the next commit.
Audit doc dqn-wire-up-audit.md updated with the kernel + ISV slot
description per Invariant 7.
No callers in this commit; layout fingerprint shifts so existing
checkpoints will fail-fast at load per feedback_no_legacy_aliases.md
(expected for a real architecture change).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
dfbadf2f3f |
test(dqn): Phase 2 Test 2.D — real-batch e2e on synthetic non-degenerate logits
Plan C Phase 2 T9. Verifies the production kernel's eval-mode argmax
exactly matches a Rust ground-truth E[Q] argmax on a 256-sample batch
with realistic per-direction non-degenerate C51 distributions, and
confirms Thompson explores Long+Short ≥ 40% in training mode.
The plan-prescribed approach (reuse Phase 0 Test 0.F's converged
checkpoint loader) was deferred — Test 0.F itself already exercises
the safetensors load + branching forward path. Replacement strategy
from the dispatch brief: synthetic batch via direct buffer write.
Setup:
- 256 samples, 21 atoms, per-(sample, direction, atom) C51 logits
drawn from a deterministic hash → uniform [-1, 1] (post-Xavier-init
scale of fresh C51 head outputs)
- Per-sample adaptive support [v_min ∈ [-1.0, -0.2], v_max ∈ [0.2, 1.0]]
matching the layout produced by `update_per_sample_support`
- Uniform q_values (mag/ord/urg fall through Boltzmann uniformly)
Rust ground-truth: softmax(b_logits[i, d]) · atom_vals[i, d] computed
with the numerically-stable subtract-max softmax matching the kernel's
softmax_c51_inline.
Assertions:
- Eval mode: per-sample dir_idx EXACTLY matches ground-truth argmax E[Q]
- Train mode: count(d ∈ {Long, Short}) / batch > 0.40
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
cc55c8a25c |
test(dqn): Phase 2 Test 2.C — mag/ord/urg branches unaffected (parity vs Boltzmann ref)
Plan C Phase 2 T8. The plan-prescribed pre-T2 snapshot approach was impossible (T2 had already landed); replacement strategy (ii) from the dispatch brief — behavioral parity vs an analytical Boltzmann reference computed in Rust — is used. Setup forces direction = Long (d=2) deterministically via peaked C51 logits (Long peaked at v=+0.8 atom; other directions at v=-0.5), so the kernel's Hold/Flat → mag_idx=0 short-circuit doesn't mask the magnitude branch's Boltzmann sampling. q_values are crafted with each branch (mag/ord/urg) peaked at a single bin with magnitude 1.0: Mag Q = [0.0, 0.0, 1.0] peak at Full (mag=2) Order Q = [1.0, 0.0, 0.0] peak at Market (ord=0) Urgency Q = [0.0, 1.0, 0.0] peak at urg=1 With q_range=1.0 in all three branches, tau collapses to 1.0 and the analytical Boltzmann probabilities are: P(best) = 1/(1 + 2/e) ≈ 0.5767 P(other) = 1/e/(1 + 2/e) ≈ 0.2117 Tolerance: at batch=8192 the 1-σ Bernoulli noise is ~0.0055 for p≈0.58; ±5% absolute tolerance covers ~9σ. Algorithmic divergence (e.g. an inadvertent strict-argmax substitution) would shift P(best) to 1.0 — trivially detected by the ±5% tolerance. Assertions: - dir_idx == Long for every sample (eval argmax E[Q] over peaked C51) - mag/ord/urg histograms each within ±5% of the Boltzmann reference Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
1a80fcab10 |
test(dqn): Phase 2 Test 2.B — production vs standalone (KS-fallback)
Plan C Phase 2 T7. #[ignore]-gated GPU test that runs BOTH the
production experience_action_select kernel AND the standalone
direction_thompson_v2_test (added in commit
|