CUDA Graph capture disables cudarc event tracking, so device_ptr()
cannot rely on write events for ordering. Without explicit sync,
the DtoH memcpy can get CUDA_ERROR_INVALID_VALUE because the
kernel writes haven't completed yet.
Add stream.synchronize() between graph launch and scalar readback.
This ensures kernel output buffers (total_loss, grad_norm) are valid
before the host reads them. The sync cost is negligible — it's
1 sync per training step (already bottlenecked by kernel execution).
Fixes H100 CI failure: all DQN pipeline tests returned
"DtoH grad_norm: CUDA_ERROR_INVALID_VALUE" consistently.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>