q_denoise_backward (63.5% GPU time, 169ms/call): decomposed into cuBLAS forward replay + backward GEMMs. 8 cuBLAS GEMMs + elementwise kernels replace 1800-thread serial loop. Target: <1ms/call. attn_bias_grad_reduce + iql_bias_grad_reduce: converted from serial batch loops to 2-phase shared-memory reduction (same pattern as IQN). Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>