variable_selection_bwd refactored from grid=(1,1,1) to grid=(B*K,1,1).
VSN's n_rows = B*K positions (one row per (batch, K-position) pair);
block-per-row matches the existing fwd kernel's layout.
Adds 2 per-row grad scratch buffers + 2 reduce_axis0 launches:
vsn_grad_w_scratch_d [B*K, FEATURE_DIM, FEATURE_DIM]
vsn_grad_b_scratch_d [B*K, FEATURE_DIM]
~210 KB scratch at B=32, K=64.
VSN bwd runs 1×/step (not K×) so the absolute wall-time win here is
small versus commits 1+2. Done for pattern uniformity — every per-batch
or per-row bwd in the trainer now uses scratch+reducer.
All 9 perception_overfit smokes pass.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>