Four code paths in PerHorizonTrainState violated CUDA Graph capture invariants and tripped CUDA_ERROR_STREAM_CAPTURE_UNSUPPORTED on smoke alpha-perception-qk2p9: 1. stream.synchronize() inside forward_with_blend (illegal in capture) 2. stream.synchronize() inside backward_through_blend (idem) 3. reduce_per_batch_scratches_to_shared used host-side vec allocations + memcpy_dtoh + CPU summation + memcpy_htod (forbidden during capture per pearl_no_host_branches_in_captured_graph) 4. zero_grads allocated host zero vectors + memcpy_htod each step (idem) Fix: - Remove both synchronizes; same-stream issue order is sufficient. - Replace host-side reduction with three reduce_axis0 GPU kernel launches (q_h, w_res, bias_res). PerHorizonTrainState now owns its own reduce_axis0 cubin handle. - Replace host-zero memcpy with stream.memset_zeros for all nine gradient buffers plus d_alpha_reduced. Verified locally: per_horizon_full_pipeline_smoke (2/2) and both numgrad parity tests (attention pool + residual head) pass on RTX 3050. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
285 KiB
285 KiB