Realizes the 2026-05-16 spec amendment merging Mamba2 + CfC into one
stacked architecture (vs the original "compete via gate" framing).
Forward chain:
snap_features × seq_len
-> window pack [1, seq_len, FEATURE_DIM]
-> Mamba2Block.forward_train -> (logit, cache.h_enriched [1, hidden_dim])
-> cfc_step(x=h_enriched, h_old=0) -> h_new
-> heads -> probs [5]
-> BCE(probs, labels)
Backward chain:
BCE -> grad_probs
-> heads_backward -> grad_h_new + grad_W_heads, grad_b_heads
-> cfc_step_backward -> grad_W_in, grad_W_rec, grad_b + grad_x (=grad_h_enriched)
-> Mamba2.backward_from_h_enriched(&cache, &grad_h_enriched_tensor)
-> Mamba2BackwardGrads (full 9-tensor gradient set)
Optimizers (6 total):
- 5 CfC AdamWs (W_in, W_rec, b, heads_w, heads_b) — reused from
PerceptionTrainer's per-param-group pattern
- 1 Mamba2AdamW for all 9 Mamba2 parameter tensors (existing
implementation in mamba2_block.rs)
Synthetic-overfit on constant +1 direction (seq_len=16, state_dim=8,
lr_cfc=3e-3, lr_mamba2=1e-3, 250 steps):
initial_avg=0.5951 → final_avg=0.1917 (68% drop, well past 40% gate).
Monotone descent at all 5 progress checkpoints.
Architectural note (v1): CfC runs with h_old=0 each step (no inter-
step recurrence). With h_old=0, the CfC layer is effectively per-cell
tau-scaled tanh FC. Inter-step CfC state (h_old carrying between
calls) is a v2 extension once the cluster gate validates the v1
foundation.
The cluster gate (Task 18) now has the actual stacked production
trainer to deploy, not a CfC-alone-vs-Mamba2-alone bench.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>