Replaces the stale `eq / fmaxf(dz, 1e-6f)` post-ReLU scale calibration in
`mag_concat_qdir` with adaptive RMS-match using ISV[H_S2_RMS_EMA_INDEX=96]
(populated each step by the producer kernel landed in 2c.3c.5).
Kernel signature gains two trailing args:
const float* __restrict__ isv,
int isv_h_s2_rms_index
Per-sample formula:
q_rms = sqrt(sum_a(eq_a^2) / b0_size)
scale = (q_rms > 1e-6) ? (h_s2_rms_ema / q_rms) : (1 / max(dz, 1e-6))
concat_out[…, SH2+a] = eq_a * scale
Q_dir's per-sample RMS now strictly tracks the trunk-output RMS regardless
of GRN drift across training. The legacy `~ 1-10` calibration (carried
over from the pre-GRN post-ReLU trunk) is removed; the comment at the
old normalization site is deleted, kernel docstring updated. The fallback
to `1 / max(dz, 1e-6)` activates only when q_rms ≈ 0 (uniform Q across
actions) — a domain-mathematical encoding, not a stub return. A static
`MAG_CONCAT_MAX_DIR=4` register-array bound matches the project's
4-direction (S/H/L/F) layout invariant; `b0_size` stays a runtime arg
for signature stability and production callers always pass 4.
Launch site `launch_mag_concat_from` extended with `isv_signals_dev_ptr`
+ `H_S2_RMS_EMA_INDEX as i32` args. `debug_assert!` mirrors 2c.3c.5's
invariant on the ISV device pointer.
Backward path unchanged: `strided_accumulate` extracts `d_h_s2` from the
first SH2 columns of `d_mag_concat` as before; `h_s2_rms_ema` and `q_rms`
are treated as fixed scalars at this batch's launch (same convention as
`dz`/`v_min` from `per_sample_support`).
Smoke (`cargo test … multi_fold_convergence --ignored --release`,
649.37s, 3 folds × 5 epochs):
fold 0 best train Sharpe 8.06 at epoch 5
fold 1 best train Sharpe 43.06 at epoch 1
fold 2 best train Sharpe 19.16 at epoch 2
geom-mean: 18.80 (vs 2c.3c.5: 20.03, -6.1%; well within 30% band)
All 3 dqn_fold{N}_best.safetensors checkpoints written. No NaN/Inf, no
fingerprint mismatch (fingerprint unchanged at 0x3e21acecd922e540). 0
panic gates added/removed. Closes the 2c.3c chain — H_S2_RMS_EMA producer
(2c.3c.5) + consumer (this commit) both wired.
cargo check clean at 11 warnings (baseline preserved); 81 cubins unchanged
(kernel-signature edit, no new .cu). +69/-6 LOC across experience_kernels.cu
and gpu_dqn_trainer.rs; audit doc appended (Invariant 7).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>