- Migrate from NVRTC JIT to cached nvcc -O3 for all CUDA kernels - Fuse guard kernels, increase prefetch chunk, eliminate per-step GPU alloc - H100-specific: fused Adam, warp reductions, shmem tiling, PPO occupancy - Vectorize gather_states with __ldg() and 4x unroll - sincosf() Box-Muller + paired Gaussian generation in noisy nets - Shared-memory tiled branching DQN forward pass for sm_<90 - GPU-resident training guard kernel replacing Candle tensor ops - Eliminate all to_vec1/to_vec2 CPU roundtrips, DtoD weight copy - GPU PER mandatory everywhere — kill CPU replay path on CUDA - Full GPU action masking — eliminate CPU fallback path - Fix cuBLAS handle sharing via OnceLock (root cause of 49 cascade failures) - Fix ILLEGAL_ADDRESS: scratch1_dist buffer overflow, stack sizing, curand determinism - Fix CudaStream lifetime: bind before .context() to extend lifetime - Keep raw cudarc buffers alive across epochs - Add gpu-hotpath-guard.sh (37 patterns) and ptx-cache-invalidate.sh Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
1.9 KiB
Executable File
1.9 KiB
Executable File