Capture the batched spectral norm kernel (10 matrices) + bf16→f32 sync
into a CUDA graph on step 1, replayed on steps 2+ with zero launch
overhead. Spectral norm uses fixed device pointers (params_bf16,
spec_u/v descriptors) that never reallocate, making it a perfect graph
capture candidate. Graph is invalidated on fold reset.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>