clip_grad_norm: accumulate squared norms on GPU-resident scalar, single
to_scalar() at end (was 16-20 per-param syncs per step × 2917 steps/epoch).
check_gradients_finite: same GPU-accumulation pattern, single sync.
gpu_replay_buffer sample_proportional/rank_based: generate random targets
via rand(0,1)*total_sum on GPU, normalize weights via broadcast_div
(eliminates 2 to_vec0 syncs per sample call).
gpu_replay_buffer update_priorities_gpu: replace CPU loop of 50 individual
slice_scatter calls with single batched index_add delta trick.
dqn NaN detection (every 500 steps): accumulate 3 NaN counts on GPU,
single to_scalar for total; detailed breakdown only if NaN found.
dqn dead neuron detection (every 1000 steps): accumulate dead count on
GPU-resident scalar (was to_vec0 per parameter tensor, ~16-20 syncs).
Net: ~22 GPU→CPU syncs + 50 micro-kernels per training step → 2 syncs.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>