Forward inference for the supervised snapshot stream — no ISV, no
temporal_weight, no NULL-pointer dispatch. Clean rewrite of the DQN
mamba2 kernel into a purpose-built alpha kernel.
New kernel `crates/ml-alpha/cuda/mamba2_alpha_kernel.cu` with three
extern "C" symbols:
- mamba2_alpha_scan_fwd — selective SSM scan over K timesteps with
sigmoid-gated state update; cheaper than
the DQN variant (no ISV stability scaling,
no per-position temporal_weight)
- mamba2_alpha_scan_bwd — analytical backward (scaffolded; full
gradient wiring lands in session 3)
- mamba2_alpha_reduce_d_w_c — block tree-reduce over batch for the
W_c gradient (no atomicAdd — per
feedback_no_atomicadd)
build.rs swapped from ../ml/src/cuda_pipeline/mamba2_temporal_kernel.cu
to the local cuda/mamba2_alpha_kernel.cu. ml-alpha no longer depends
on ml's CUDA source — fully self-contained alpha-stack.
Forward pipeline:
1. cuBLAS sgemm: input [B,K,in] @ W_in.T + b_in → x [B,K,hidden]
2. cuBLAS sgemm: x @ W_a.T + b_a → a_proj [B,K,state]
3. cuBLAS sgemm: x @ W_b.T + b_b → b_proj [B,K,state]
4. zero-init h_s2, h_enriched [B, hidden]
5. scan kernel: (a_proj, b_proj, W_c, h_s2) → h_enriched
6. cuBLAS sgemm: h_enriched @ W_out.T + b_out → logit [B, 1]
All on GPU; output is a [N] CudaSlice<f32> of raw logits. Caller
sigmoids + thresholds (or feeds directly into BCE-with-logits).
Tests (5 passing on real GPU):
- forward [4, 16, 81] → logit [4, 1], all finite
- reject wrong in_dim
- reject wrong seq_len
- reject state_dim > 16
- reject zero dims
- + parameter-count sanity
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>