Single cuGraphExecLaunch per step. PER sampling, gather, training, priority update ALL as child graph nodes in one parent. Direct-to-trainer gather eliminates DtoD copies. GPU-side counters eliminate host writes. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>