Replace per-batch loss.to_vec0() in TFT training/validation loops with GPU tensor accumulation. Single extraction per epoch + NaN guard every 100 batches. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Replace per-batch loss.to_vec0() in TFT training/validation loops with GPU tensor accumulation. Single extraction per epoch + NaN guard every 100 batches. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>