Per docs/superpowers/specs/2026-05-17-kloop-parallelization-design.md.
cfc_step_batched (fwd + bwd) refactored from grid=(1,1,1) with internal
n_batch loop to grid=(B,1,1) — each block handles one batch. Removes
the single-SM bottleneck on the K-loop's most-called kernel (64×/step).
Param-grad accumulation moves to per-batch scratch:
cfc_grad_w_in_scratch_d [B, n_hid, n_in]
cfc_grad_w_rec_scratch_d [B, n_hid, n_hid]
cfc_grad_b_scratch_d [B, n_hid]
cfc_grad_tau_scratch_d [B, n_hid]
Zeroed once per training step, K-loop's 64 bwd calls += into them, then
4 reduce_axis0 launches collapse B → final grad buffers (OVERWRITE)
before AdamW. New AdamW-after-reducer invariant: final grads are
meaningful only after the reducer has run in the current step.
New reduce_axis0 kernel: single parameterised reducer [B, N] → [N] via
block tree-reduce (no atomicAdd per feedback_no_atomicadd.md). Same
pattern as layer_norm_reduce_param_grads — CUDA-Graph-safe.
cfc_step_backward_batched shared-mem dropped from (B+1)*n_hid*4 to
2*n_hid*4 bytes per block (only one row of sd_pre needed per block bi).
Tests:
- New stacked_trainer_loss_shrinks_at_batch_32: FIRST test that
actually exercises the cross-batch reduction code path; existing
perception_overfit suite was all B=1. Initial 0.24 → final 0.00.
- Scratch-clears test removed (explanatory comment kept): structurally
hard to assert directly due to begin_capture/end_capture not
executing kernels; the B=32 convergence smoke implicitly validates
scratch zeroing since divergence would otherwise be immediate.
All 9 perception_overfit smokes + 4 backward_finite_diff tests pass.
build.rs:
- KERNELS list adds "reduce_axis0"
- Cache-bust → v11
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>