refactor(cuda): eliminate candle from ml-core, ml-ppo, and 4 thin crates
Hard refactor — no shims, no compat layers. Candle removed from Cargo.toml
and all source files in 6 crates:
- ml-core: MlDevice enum, checkpoint.rs (safetensors direct), cudarc imports
fixed from candle re-export to direct, AdamWConfig lr_decay, cuda_compat
gutted. Net -7,341 lines.
- ml-ppo: All 16 files rewritten. LSTM→CudaLSTM, VarMap→GpuVarStore,
PPOAgent 2306→700 lines, checkpoint→binary format.
- ml-ensemble: GPU-resident sigmoid via custom CUDA kernel.
- ml-explainability: Integrated gradients via GPU finite-difference kernels.
- ml-labeling: Device→MlDevice.
- ml-hyperopt: Cargo.toml only.
Remaining: ml-dqn (24 files), ml-supervised (4 files), ml crate (104 files).
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>