- GpuBatch.indices: GpuTensor → CudaSlice<u32> (fixes 2402 compute-sanitizer memory errors from per_update_priorities_kernel reading u32 from bf16 buffer) - fused_training PER: eliminate bf16→host→u32→GPU roundtrip, pass u32 directly - train_step accumulation: GPU DtoD concat for CudaSlice<u32> indices - grad_norm kernel: float accumulator via separate CudaSlice<f32> buffer (bf16 sum-of-squares overflows at 147K params; atomicAdd on native float) - grad_norm finalize kernel: runs OUTSIDE CUDA graph, converts float→bf16 L2 norm - Adam + clip_grad + clipped_saxpy kernels: read float sum-of-squares directly - training guard: read loss/grad from fused trainer's GPU buffers (not GpuTrainResult's hardcoded zeros), raw_ptr() for kernel args (no event tracking) - guard accumulator: reset between epochs for per-epoch metrics - Q-stats padding: pad input to config.batch_size for CUTLASS tile alignment - training_profile tests: update BF16-tuned values (spectral_norm 1.5, noisy_sigma 0.3) 895/895 unit tests pass, 5/9 smoke tests pass (remaining 4 need loss kernel float arithmetic — C51/MSE softmax overflows bf16 after ~100 training steps). Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
13 KiB
13 KiB