Adds 4 GPU-native kernels replacing host-side operations on the training hot path: gather_f32_rows_padded (row gather + 128-byte zero-pad), gather_f32_scalar, gather_i32_scalar, and increment_step_counters (atomic counter bumps + cosine-annealed tau — zero CPU sync per step). Wired into build.rs (nvcc cubin compile) and gpu_dqn_trainer.rs (GRAPH_UTILITY_CUBIN include_bytes!). Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>