Files
foxhunt/docs
jgrusewski 68197c2c25 feat(dqn-v2): Plan 4 Task 2c.1 — GRN kernel module (forward + backward, additive)
Plan 4 Task 2c.1. Additive module — kernels compile and load via the
existing kernel-loading infrastructure but have NO production callers
in this commit. Task 2c.3+4 wires them into the trunk encoder.

Eight kernels in grn_kernel.cu (Linear_a/Linear_b/Linear_residual are
cuBLAS GEMMs, not new kernels):
- grn_elu_inplace: element-wise ELU (canonical α=1.0)
- grn_glu_forward: GLU split + sigmoid (saves sigmoid for backward)
- grn_residual_layernorm_forward: residual add + LN, saves mean/rstd/normed
- grn_layernorm_backward_dx: FULL JACOBIAN (not the simplified
  attn_layer_norm_bwd_dx which would silently propagate
  approximation into trunk gradients)
- grn_layernorm_backward_dgamma_dbeta_p1: per-block partial reduction
  (no atomicAdd per feedback_no_atomicadd.md)
- grn_layernorm_backward_dgamma_dbeta_p2: final reduce across blocks
- grn_glu_backward
- grn_elu_backward

LN backward formula (full Jacobian per pearl_cold_path):
  d_out_g = d_out * gamma
  s1 = sum_d(d_out_g)
  s2 = sum_d(d_out_g * normed)
  d_x = (1/H) * rstd * (H * d_out_g - s1 - normed * s2)

Layout convention: row-major [B, H] (sample-major, matches spec §4.E.2
and dt_layernorm_kernel) — intentionally different from
attention_kernel.cu's [D, B] col-major; the trunk encoder downstream
of GRN uses row-major buffers, so per-row LN avoids a transpose.

build.rs registers grn_kernel.cu (kernel count: 57 → 58). Verified
nvcc compiles cleanly via `cargo build -p ml --lib`. Module is dead
code until 2c.3+4. Wire-up audit updated with the additive entry per
Invariant 7.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-25 12:02:36 +02:00
..