Root cause of cublasLtMatmul CUBLAS_STATUS_NOT_SUPPORTED: - states_buf allocated [batch * state_dim] (unpadded) - cuBLAS forward GEMM reads with stride state_dim_padded (pad128) - Buffer overflow: GEMM reads past buffer end cublasLtMatmul validates buffer sizes against layout descriptors and returns NOT_SUPPORTED for undersized buffers. cublasSgemm silently read garbage — this was the source of the "parallel test congestion errors" seen previously. Fix: - Allocate states_buf with state_dim_padded stride (3 allocation sites) - gather_states kernel: add padded_sd parameter, write with padded stride - DtoD copy: use padded row bytes - Both gather_states call sites updated with padded_sd arg Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
33 KiB
Executable File
33 KiB
Executable File