feat(sp14-c6): h_s2_aux_rms_ema producer — ISV[449] per-collector-step

Single-block 256-thread CUDA kernel computing RMS(h_s2_aux [B, SH2])
and EMA-blending the step observation into ISV[H_S2_AUX_RMS_EMA_INDEX=449]
directly. Pearl-A first-observation bootstrap embedded in kernel body
(sentinel 0.0 → replace); fixed α=0.05 EMA blend thereafter.

ISV slot 449 is outside the SP4/SP5 wiener buffer linear span so the
scratch+apply_pearls_ad_kernel path is not available — self-contained
Pearl-A logic mirrors the avg_win_hold_time_update_kernel precedent
(slot 451). No atomicAdd; shmem block-tree-reduce only. Launched after
aux_trunk_forward in the collector per-step hot path.

- h_s2_aux_rms_ema_kernel.cu — new CUDA kernel (81 lines)
- build.rs — cubin manifest entry
- gpu_dqn_trainer.rs — H_S2_AUX_RMS_EMA_CUBIN static
- gpu_aux_trunk.rs — HS2AuxRmsEmaOps struct + launch()
- gpu_experience_collector.rs — field + constructor + hot-path launch
- aux_trunk_oracle_tests.rs — h_s2_aux_rms_ema_pearl_a_bootstrap test
- dqn-wire-up-audit.md — Phase C.6 audit entry

cargo check -p ml --tests: clean (only pre-existing warnings)
Oracle test: 1 new test added (requires GPU to run)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
This commit is contained in:
jgrusewski
2026-05-08 03:17:30 +02:00
parent b26b189925
commit 0e61de408f
7 changed files with 414 additions and 2 deletions

View File

@@ -786,6 +786,18 @@ fn main() {
// (no atomicAdd) per `feedback_no_atomicadd.md`. Pearl-A
// first-observation bootstrap; α=0.05 EMA thereafter. Plan: §C.4b.
"avg_win_hold_time_update_kernel.cu",
// SP14 Layer C Phase C.6 (2026-05-08): h_s2_aux RMS EMA producer.
// Single-block 256-thread kernel computing RMS = sqrt(mean(x²))
// over `h_s2_aux [B, SH2]` (aux trunk final output, no activation)
// and EMA-blending into `ISV[H_S2_AUX_RMS_EMA_INDEX=449]`. Pearl-A
// first-observation bootstrap embedded in kernel (sentinel 0.0 →
// replace directly); fixed α=0.05 EMA blend thereafter. ISV slot 449
// is outside the SP4/SP5 wiener buffer linear span so the shared
// scratch+apply_pearls_ad_kernel path is unavailable — bootstrap
// logic is self-contained per `avg_win_hold_time_update_kernel`
// precedent. Block-tree-reduce (no atomicAdd) per
// `feedback_no_atomicadd.md`. Per-collector-step launch cadence.
"h_s2_aux_rms_ema_kernel.cu",
// SP14 Layer B Task B.3 (2026-05-05): Earned Gradient Flow
// producer kernel. Per-step computes the K=4↔K=2 mapped argmax
// mismatch between the Q-head's 4-way direction action