Files
foxhunt/crates
jgrusewski 7e38e46e65 fix(cuda): h_mag_per_bucket_kernel multi-warp reduction (32→128 lanes)
Latent bug surfaced by f11bab542 test audit. `h_mag_per_bucket_kernel`
launched with block_dim=32 (single warp) but BUCKET_DIMS terciles are
[43, 43, 42] — channels 32..bdim were silently dropped from the sum,
producing under-counted Controller D dead-bucket signal in production
(perception.rs::h_mag_cfg).

Fix: grow block_dim to 128 (smallest pow2 ≥ MAX_BUCKET_DIM=96), expand
shared mem from `__shared__ float sdata[32]` to `sdata[128]`, change
reduction stride from 16→32→64 (single-warp shuffle equivalent) to
64→32→16→8→4→2→1 (multi-warp block-tree). No atomicAdd, no warp
divergence in tail lanes (uniform predicate tid < bdim).

Changes:
- bucket_transition_kernels.cu: kernel block-tree reduction over 128
  lanes, shared mem 128 floats, launch comment updated
- perception.rs::h_mag_cfg: block_dim 32→128, shared 32*4→128*4
- bucket_transition_kernels.rs test: launch config matches new contract
  + bug-surfaced comment replaced with fix-landed reference

Test: tests/bucket_transition_kernels::h_mag_per_bucket_kernel_computes
_mean_abs_per_bucket now passes (was deliberately failing in f11bab542
to surface this bug per feedback_no_todo_fixme — tests assert
invariants not observed behavior).

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-29 00:30:14 +02:00
..