CUDA kernels: 36 of 37 compiled with __nv_bfloat16* (dt_kernels.cu remaining) Rust types: ALL CudaSlice<f32> → CudaSlice<half::bf16> across ml + ml-core Build: all kernels now compiled with common_device_functions.cuh (BF16 helpers) BF16 math wrappers: bf16_sqrt, bf16_log, bf16_exp, bf16_pow, bf16_fabs, bf16_fmax, bf16_fmin, bf16_cos, bf16_zero, bf16_one, bf16(), atomicAddBF16 NOT YET COMPILING — ml-core boundary errors (Vec<f32> → Vec<half::bf16>) and dt_kernels.cu float*__nv_bfloat16 ambiguity remain. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2.1 KiB
2.1 KiB