Root cause: weight tensors packed sequentially in the flat params buffer had non-aligned start offsets when preceding tensors had odd element counts (e.g. bias of 51 atoms = 204 bytes, 204 % 16 = 12). cublasLtMatmul with CUBLAS_COMPUTE_32F_FAST_TF32 requires 16-byte aligned buffer pointers. Fix: pad each tensor to 4-element boundary (16 bytes) in both f32_weight_ptrs_from_base (pointer computation) and compute_total_params (buffer allocation). Added align4() and padded_byte_offset() helpers, fixed shrink_perturb skip range and bottleneck gradient offset. Switched compute type: CUBLAS_COMPUTE_32F → CUBLAS_COMPUTE_32F_FAST_TF32 (forward + backward). Explicit TF32 tensor core path, required by cuBLAS 13.0 on H100 SM90. Deleted dead bf16_weight_ptrs function. 19/19 smoke tests pass on RTX 3050. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
912 B
912 B