Forward-port from worktree-gpu-hotpath-audit (2 commits):
1. Zero-roundtrip DQN experience collection via cuMemcpyDtoDAsync:
- GpuExperienceCollector uses device-to-device copies for state tensors
- Shared memory weight caching in GPU replay buffer
- NaN priority clamping on GPU (no CPU readback)
2. Eliminate all GPU→CPU roundtrips from training loop:
- GpuTrainResult: loss + grad_norm stay as GPU scalar tensors
- Single to_scalar() readback at epoch boundary (not per step)
- GPU-resident loss accumulation across training steps
- RegimeConditional head selection via GPU tensor ops
- Removed per-step NaN diagnostic checks (now epoch-level)
Also removes dead `states_tensor` field from ComputeLossResult and
fixes redundant field name clippy warning.
Expected: ~15-20% training throughput improvement on H100 by
eliminating synchronous GPU→CPU transfers in the inner loop.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>