variable_selection_bwd refactored from grid=(1,1,1) to grid=(B*K,1,1). VSN's n_rows = B*K positions (one row per (batch, K-position) pair); block-per-row matches the existing fwd kernel's layout. Adds 2 per-row grad scratch buffers + 2 reduce_axis0 launches: vsn_grad_w_scratch_d [B*K, FEATURE_DIM, FEATURE_DIM] vsn_grad_b_scratch_d [B*K, FEATURE_DIM] ~210 KB scratch at B=32, K=64. VSN bwd runs 1×/step (not K×) so the absolute wall-time win here is small versus commits 1+2. Done for pattern uniformity — every per-batch or per-row bwd in the trainer now uses scratch+reducer. All 9 perception_overfit smokes pass. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
8.6 KiB
8.6 KiB