jgrusewski
|
169821b3da
|
docs: GPU max performance Phase 2 design (L4 → H100)
8 optimizations: dynamic batch size, tensor core alignment, no-grad
inference, CUDA stream double-buffering, epoch prefetch, INT8
quantized inference, memory pinning, CUDA kernel optimization.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
|
2026-03-01 18:37:03 +01:00 |
|