Replace per-param gradient norm extraction loop with batched Tensor::stack pattern. Single GPU→CPU sync instead of one per parameter tensor. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Replace per-param gradient norm extraction loop with batched Tensor::stack pattern. Single GPU→CPU sync instead of one per parameter tensor. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>