Cargo.toml: drops gbdt; adds memmap2 + approx; keeps ml-core only (cannot depend on ml: would cycle since ml depends on ml-alpha for the Mamba2 gate baseline). build.rs: compiles 7 cubins (mamba2_alpha + 6 new placeholders) with -O3 --use_fast_math --ftz --fmad. Skips kernels whose source isn't present yet so partial check-ins work. Every env::var paired with rerun-if-env-changed per the canonical build pearl. src/pinned_mem.rs: local copy of MappedF32Buffer (mirrors ml::cuda_pipeline::mapped_pinned::MappedF32Buffer). Drives the only permitted CPU<->GPU path per feedback_no_htod_htoh_only_mapped_pinned. Eventually the move-to-ml-core refactor will deduplicate; out of scope for the Phase A branch. Addendum: updates the import path to ml_alpha::pinned_mem. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
3 lines
116 B
Plaintext
3 lines
116 B
Plaintext
// placeholder — real implementation in the task adding this kernel
|
|
extern "C" __global__ void cfc_step_stub() {}
|