Captures 50+ kernel launches as 3 CUDA Graphs (pre-fill, post-fill, replay-step) for single-launch replay. Eliminates ~300μs/step of CPU launch overhead at b=16. Key challenges: device-resident step counter (no scalar arg changes in graph), lobsim fill breaks the graph (split into pre/post), ISV write ordering. Estimated 1.5-3× throughput improvement. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>