Single cuGraphExecLaunch per step. PER sampling, gather, training,
priority update ALL as child graph nodes in one parent. Direct-to-trainer
gather eliminates DtoD copies. GPU-side counters eliminate host writes.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>