Three issues in one spec: 1. H100 hang: cuBLAS set_stream() inside CUDA graph capture → deadlock 2. Bug: train_walk_forward calls non-stratified generate_folds() 3. GPU regime: classification kernel + prefix sum for O(1) range queries Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>