Root cause: SHMEM_MAX_IN_DIM was hardcoded to 256, but on H100 with
hidden_dim_base=2048, SHARED_H1=SHARED_H2=512. The cooperative tile
loader wrote 64×512=32768 floats into shmem allocated for 64×256=16384,
causing CUDA_ERROR_INVALID_VALUE on kernel launch.
Fixes:
- Make SHMEM_MAX_IN_DIM injectable via NVRTC #define (was hardcoded)
- Compute max(state_dim, shared_h1, shared_h2) and inject at compile time
- Use dynamic value for host-side shmem_bytes allocation
- Shrink eps_out[NOISY_MAX_DIM] → eps_out[SHMEM_TILE_ROWS] (saves 768B/thread)
- Force smoke tests to Device::Cpu (CUDA driver sensitivity to binary layout)
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>