Complete bf16 elimination across all crates (ml, ml-core, ml-dqn, ml-ppo, ml-supervised). Zero half::bf16, __nv_bfloat16, or CudaSlice<half::bf16> references remain (verified by grep). CUDA: All 60+ .cu kernels and .cuh headers converted to native float. - Half-precision intrinsics (__hmul, __hadd, __hdiv) → float operators - atomicAddBF16 → native atomicAdd(float) - bf16 wrapper functions → f32 identity passthroughs Rust: All CudaSlice<half::bf16> → CudaSlice<f32> across 90+ files. - htod_f32_to_bf16/dtoh_bf16_to_f32 → htod_f32/dtoh_f32 (direct, no conversion) - Deleted bf16 mirror infrastructure (DuelingWeightSetBf16, alloc_bf16_mirror, etc.) - Renamed params_bf16→params_flat, d_value_logits_bf16→d_value_logits, etc. - Fixed .to_f32() sed damage on Decimal::to_f32() and rng.f32() FxCache: Single f32 disk format (was bf16/f64 dual-version). - Deleted --bf16 CLI flag from precompute_features - PVC cache files need regeneration via precompute_features TF32 tensor cores activated via cublasLtMatmul CUBLAS_COMPUTE_32F — no explicit TF32 types needed. Storage is pure f32 everywhere. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
15 KiB
15 KiB