Three review-driven fixes:
1. Slot 24/25 cite the actual production buffers d_value_logits_buf
(gpu_dqn_trainer.rs:3123) and d_adv_logits_buf (:3125) — the f32
atomicAdd buffers that launch_c51_grad writes — not the staging
buffers at :3127/:3129. Pointer expressions exist at the launch
site (:17707/:17708), making new accessors optional. Task 3 summary
and priority-list entries updated to match.
2. Explicit plan-supersession note inserted immediately above the
per-slot table: the audit's per-slot allocation supersedes the plan
Task 2 step 6 name table. The plan's allocation was placeholder;
the audit's is grounded in per-kernel inspection. Task 2 should use
the audit's names (d_value_logits_buf, d_adv_logits_buf, iqn_trunk_m,
iqn_d_h_s2_ptr, d_branch_logits_buf, etc.).
3. Summary suspicion-ranking row #3 reconciled with slot-28 dormancy:
iqn_backward_per_sample has no Rust caller (verified via
grep crates/ml/src/), so row #3 is re-pointed at the production
kernel iqn_quantile_huber_loss (iqn_dual_head_kernel.cu:1346-1413,
loaded at gpu_iqn_head.rs:2114). Same unsafe-write pattern (no
isfinite guard at line 1410's d_q_online[idx] = qw*d_huber/Q),
now attributed to the live path. apply_iqn_trunk_gradient at #1
stands — its reasoning (orchestrator consuming iqn_d_h_s2_ptr) is
unchanged by which specific kernel writes that buffer.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>