1cccb4e40ebe5bf9bde8d9ec63b13d6ffe77b387
60 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
e004d6c217 |
Revert "arch(ml-alpha): restore count-delta redundancy at feature slot [26]"
This reverts commit
|
||
|
|
008f65d894 |
arch(ml-alpha): restore count-delta redundancy at feature slot [26]
The C16 tick-rule swap (
|
||
|
|
17eb825113 |
arch(ml-alpha): σ becomes sidecar — BCE drops Kendall damping (upgraded path)
Empirical record at this data + this architecture:
Phase 1+2+3 (no σ in loss): mean_auc 0.7749 ± 0.024
v2 (σ + axes B/C/D/E): mean_auc 0.7541 ± 0.005
σ-only (kept σ, dropped B/C/D/E): mean_auc 0.7506 ± 0.008
Perf-fix (σ-only math): mean_auc 0.7499 ± 0.010
ISV-σ (closed-form σ + adaptive λ): mean_auc ~0.75 (2/3 folds)
Every architecture with σ-in-loss lands at 0.75. Removing σ is the
only thing that hits 0.77. That is a framework mismatch, not a tuning
problem.
Kendall+Gal+Cipolla 2018 frames σ as TASK NOISE level — damp the noisy
task, trust the clean one. Our horizons don't have different label
noise; they have different intrinsic difficulty (longer horizon = more
price-walk uncertainty = lower achievable AUC). σ-Kendall sees "high
BCE on h6000" and interprets it as "h6000 is unreliable, back off" —
precisely the opposite of what we want. h6000 is the deployment
target; damping it is a self-inflicted wound. With mean_bce ~ 0.7,
σ ≈ √0.7 ≈ 0.84 ⇒ w_h ≈ 0.71, uniformly attenuating gradient by
~30% across every horizon. λ's z-score boost (max 2×) can rebalance
relative-per-horizon but cannot recover the absolute magnitude.
Per-horizon prioritization remains via the ISV-driven λ controller
(grad scaler in heads_grn_bwd, per pearl_adam_normalizes_loss_weights).
That controller IS appropriate for our problem: it boosts hard
horizons rather than damping them, and it operates on the gradient
into the trunk rather than on the loss aggregate (Adam-cancellation
safe).
Changes to bce_loss_multi_horizon.cu (six lines):
w_h = bw (was: 0.5 * bw * exp(-2 * log_sigma_h))
d_log_sigma_h[h] = 0.0 (was: 1 - 2 * w_h * mean_bce)
total_loss += bw * mean_bce (was: + w_h * mean_bce + log_sigma_h)
σ infrastructure preserved unchanged:
- horizon_ema_and_lambda still computes log_sigma_h closed-form from
loss_ema (Kendall equilibrium) for telemetry / future label-noise
estimation use cases
- log_sigma_h kernel arg still in BCE signature (zero churn at
callsite); ignored inside
All 9 perception_overfit tests pass — including
horizon_ema_and_lambda_track_after_training which validates the
per-horizon controller end-to-end through 64 K-loop iterations of
capture/replay.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
|
||
|
|
410ab6b0ea |
arch(ml-alpha): ISV-driven σ + adaptive Z_SCALE — both controllers anchor on loss_ema
Per pearl_controller_anchors_isv_driven, every controller anchor/target/
cap derives from a tracked signal, not hardcoded constants. The σ-only
revert kept Kendall σ as a free Adam-learned scalar — that violated ISV
discipline and fought Adam's m/√v normalization
(pearl_adam_normalizes_loss_weights).
Single source of truth for both per-horizon controllers:
log_sigma_h[h] ← max(log(0.5), 0.5 * log(loss_ema[h]))
Kendall equilibrium (∂L/∂log σ = 0 ⟹ σ_h² = mean_bce_h)
in closed form. Asymmetric floor at log(0.5) prevents
collapse. No Adam state, no gradient delay.
lambda[h] ← clamp(1.0, 2.0, 1.0 + Z_SCALE_ISV * z_h)
Z_SCALE_ISV = (LAMBDA_CEILING - LAMBDA_FLOOR) / z_max_ema
Adaptive scale auto-uses the full clamp envelope:
the historical-max-z horizon maps exactly to
LAMBDA_CEILING. Replaces hardcoded Z_SCALE=0.5 which
rarely engaged on real data (max observed λ ~1.04).
Both anchor on the same ISV (loss_ema). z_max_ema is a new single-scalar
EMA state tracking max |z| across horizons, with first-obs bootstrap.
Removes:
- opt_log_sigma AdamW optimizer (σ no longer learned)
- grad_log_sigma_h_d memset (BCE kernel writes; output ignored — kept
only to preserve BCE kernel signature)
Kernel signature change (horizon_ema_and_lambda):
+z_max_ema [1] (read+write EMA state)
+log_sigma_h [5] (closed-form output, overwrite)
Discipline:
- First-obs bootstrap (sentinel <= 0) per pearl_first_observation_bootstrap
- Permanent floor (max(real, floor)) per pearl_blend_formulas_must_have_permanent_floor
- Asymmetric clamp per pearl_audit_unboundedness_for_implicit_asymmetry
- Z-score normalisation per pearl_zscore_normalization_for_magnitude_asymmetric_signals
- No nvrtc, no atomicAdd, no host branches in graph capture
All 9 perception_overfit tests pass — including
horizon_ema_and_lambda_track_after_training which validates the kernel
end-to-end through 64 K-loop iterations of capture/replay.
Submit local smoke; cluster A/B vs σ-only baseline (0.7506/0.7519) and
vs Phase 1+2+3 (0.7749/0.7591) follows once the perf-only 3-fold A/B
confirms no regression at
|
||
|
|
b23f8f2efa |
perf(ml-alpha): NVIDIA-grade rewrite of CfC K-loop hot kernels — 2.15× faster
Local L40S profile (perception_overfit smoke) GPU kernel time:
1589ms → 739ms (53.5% reduction, 2.15× speedup).
Wall-clock smoke: 9.6s → 4.94s (1.94× faster).
Per-kernel deltas (nsys --cuda-graph-trace=node):
reduce_axis0: 362ms → 10ms (36× faster)
Block layout: per-column (1 block / output) → 32-wide column tile
(block_dim = 32 × 8). Cross-thread reads were strided by n_tail
(~40K floats = 160KB stride) — one cache line per thread, 8× HBM
bandwidth wasted. New tile gives coalesced 128B transactions per
warp. Block tree-reduce kept (no atomicAdd, per feedback_no_atomicadd).
+1 shared-mem pad to eliminate 32-way bank conflict on the ty reduce.
multi_horizon_heads_grn_bwd_batched: 540ms → 113ms (4.8× faster)
1. Stage h_row[HIDDEN] and a1[5,HEAD_MID] in shared at block entry.
Eliminates ~28K redundant DRAM reads/block across Pass 3 + Pass 5.
2. Pass 5 reorder: k outer / i inner with d_z1[k,m] pinned in
register; writes to grad_w1_scratch are sequential per-thread.
3. Block size 64 → 128 threads. Pass 5/6 now partition over i
(output column): cross-thread writes become COALESCED 128B/warp
(was stride-128 = 512B). Passes 2/3/4 gate on (tid < HEAD_MID).
4. Pass 3 thread role: m_out → m_in/n. Same coalescing fix on
grad_w2 writes AND w2 reads in the d_eta_2 sum.
cfc_step_backward_batched: 351ms → 271ms (1.3× faster)
1. Stage x_b[n_in] and h_old_b[n_hid] in shared (was 128× redundant
DRAM reads per block; now 1× cooperative load).
2. Pass 1 thread role: i (output row) → k (output col). For each
i loop iteration, the warp writes grad_w_in[..., tid] /
grad_w_rec[..., tid] — COALESCED 128B/warp (was stride-128
non-coalesced).
multi_horizon_heads_grn_fwd_batched: 211ms → 204ms
Stage h_row[HIDDEN] in shared — Pass 1 and Pass 3 both consume.
cfc_step_batched (fwd): 95ms → 94ms
Stage x_b and h_old_b in shared.
Shared-mem budgets fit comfortably under the 48KB SM cap (~6KB / ~2KB
respectively). All 9 perception_overfit tests pass — gradient
correctness validated end-to-end (constant-signal overfit, K-loop
capture/replay, stride-4 path, evaluate-only paths).
Discipline:
- Block tree-reduce only, never atomicAdd
- No nvrtc; pre-compiled cubins via build.rs
- Mapped-pinned-only is unaffected (CPU↔GPU contract untouched)
- Single source of truth: replaced kernels in place, no v2 suffixes
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
|
||
|
|
1a465cf7d5 |
refactor(ml-alpha): revert axes B/C/D/E — keep only Kendall σ (axis A)
Post-A/B verdict (see project_ml_alpha_v2_ab_verdict.md): v2 with all
5 axes was marginally tied on h6000 (+0.0013 vs 0.7591 baseline mean,
fails the +0.01 win threshold) and slightly below on mean_auc
(−0.0208 vs 0.7749 baseline mean, within 1σ) at ~5× the wall-time
cost. Per `feedback_v7_gem_methodology` (measure before delete or
wire), the architecture has been measured — it doesn't earn its
compute cost. This commit reverts the axes that didn't lift:
- axis B (L2 anchor + Wiener-α controller) — DROPPED
- axis C (horizon-token attention pool) — DROPPED
- axis D (regime-MoE gate + experts) — DROPPED
- axis E (inverted cross-variate attn) — DROPPED
- axis A (Kendall σ-weighted BCE) — KEPT
Files deleted (kernels, host bindings, numgrad tests, trainer state):
- cuda/{horizon_token_attention_pool, inverted_attention_pool,
inv_pooled_merge, regime_moe_gate, anchor_l2,
horizon_mean_collapse}.cu
- src/{horizon_token_attention_pool, inverted_attention_pool,
inv_pooled_merge, regime_moe_gate, anchor_l2,
horizon_mean_collapse}.rs
- src/trainer/{multi_horizon_attention, anchor_controller}.rs
- tests/{horizon_token_attention_pool_numgrad,
inverted_attention_pool_numgrad,
regime_moe_gate_numgrad,
anchor_l2_numgrad}.rs
Files restored (from V1 commit
|
||
|
|
4ae9a27f48 |
fix(ml-alpha): wire axis E into loss — real add_inv_broadcast kernel
CRITICAL ARCHITECTURAL FIX discovered while planning fused kernel:
The previous Stage 2 `add_inv_broadcast` helper in
`multi_horizon_attention.rs` was a STUB that returned Ok(()) without
doing anything. Practical consequences:
- Forward: `inv_pooled_d` (output of inverted_attention_pool.forward)
was never added into `ctx_h_d`. Downstream MoE + heads never saw
the inverted-attention signal. Axis E contributed ZERO to the
forward output and the loss.
- Backward: `inv_pool.backward` was being fed `grad_ctx_mean` as
its "upstream gradient", but that's the gradient at the CHAIN
TERMINUS — not the gradient w.r.t. inv_pooled_d (which is zero
by construction since inv_pooled wasn't in the loss). The bwd
was injecting incorrect noise into `grad_ln_out`.
Net: paying inverted_attention compute for no gain, plus polluting
ln_b's gradient. Two perf rewrite rounds earlier today showed no
wall-time movement precisely because the slow path was wired into
training while the optimized one was dead.
FIX:
(a) New kernel `cuda/inv_pooled_merge.cu`:
fwd: ctx_h[b, h, d] += inv_pooled[b, d] (broadcast over h)
bwd: grad_inv_pooled[b, d] = Σ_h grad_ctx_h[b, h, d]
Tiny — single block-per-batch, no syncthreads, coalesced reads.
(b) Host binding `src/inv_pooled_merge.rs` (InvPooledMerge).
(c) `MultiHorizonAttention` adds:
- `merge: InvPooledMerge` field.
- `grad_inv_pooled_d [B, H]` buffer for the real upstream of inv_pool.bwd.
(d) `MultiHorizonAttention.forward` now calls `merge.forward(...)`
between inv_pool.forward and moe.forward. Axis E is now actually
in the model's forward output.
(e) `MultiHorizonAttention.backward` now calls `merge.backward` after
moe.bwd writes `grad_ctx_h_d`, producing `grad_inv_pooled_d`.
`inv_pool.backward` consumes the REAL upstream gradient
(`grad_inv_pooled_d`) instead of the prior `grad_ctx_mean` fake.
(f) Old stub `add_inv_broadcast` deleted.
CORRECTNESS:
- perception_overfit 9/9 PASS after the wiring. Loss still shrinks
on the constant-signal test (0.67 → -0.99 over 250 steps), now
with axis E actually contributing.
- All numgrad kernel tests still pass (kernels themselves unchanged
in this commit; only the wiring).
PERF IMPACT (expected):
- +2 tiny launches per step (merge fwd + bwd). Negligible.
- The inverted-attention compute that was previously dead now
actually feeds the loss → same wall-time, but it's earning the
cost. This unblocks meaningful axis-E perf measurement on next
smoke.
NEXT: re-run cluster smoke to confirm wall-time + verify axis E
gradient flow is healthy. After that, consider full MHA forward
fusion as a follow-up.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
|
||
|
|
11b964359b |
perf(ml-alpha): MoE bwd compact scratch + scatter kernel + inv-attn loop interchange
Two perf optimizations bundled:
(1) MoE backward scratch compaction
Old: `grad_w_scratch_d [B, N_H, N_E, H, H]` = 1.3 MB per step (B=1)
New: `grad_w_scratch_d [B, N_H, H, H]` = 320 KB per step
4× memory reduction. Since each batch's top-1 router selects ONE
expert, only that expert's slot was ever non-zero in the prior
layout — the N_E axis was entirely wasteful.
The shape change required:
- Updated `regime_moe_gate_bwd` to write the compact layout.
- New `regime_moe_gate_scatter` kernel scatters per-(b, h)
rank-1 contributions into `grad_experts_W[e]` / `grad_experts_b[e]`
based on `top_e[b]`. Grid (N_E, H, ceil(H/32)) × block (32) —
one warp per (e, d_out, d_in_chunk). 65536 → 16384 grid cells
(4× fewer blocks dispatched).
- Dropped the previously-naive 65536-block `reduce_axis0` for
`grad_experts_w` from `perception.rs` (the scatter kernel
produces the final per-expert grad directly).
- `tests/regime_moe_gate_numgrad.rs` reads `grad_experts_w` from
the scatter output instead of host-side reducing the 5D scratch.
(2) inverted_attention bwd loop interchange
Phase 2's tight loop:
for k:
for j:
ds_myh_j = d_scores[my_h, j] // doesn't depend on k!
ds_j_myh = d_scores[j, my_h] // doesn't depend on k!
...
Hoisted d_scores reads out of the K-loop into J-outer with
per-thread `q_arr[K_MAX]` / `k_arr[K_MAX]` register accumulators.
Net: 32× fewer DRAM reads of d_scores per thread per bwd.
CORRECTNESS:
- regime_moe_gate numgrad PASSES (1/1, 11 numgrad checks).
- inverted_attention numgrad PASSES (1/1, 6 numgrad checks).
- perception_overfit 9/9 PASS — including loss-shrinks tests.
NEXT: re-run cluster smoke to measure the new wall-time vs the 17 s
baseline. Prior smoke at
|
||
|
|
a263cd5446 |
perf(ml-alpha): inverted_attention_pool — cache mean_k, pre-compute pooled
The cluster smoke at
|
||
|
|
9170d24fe3 |
refactor(ml-alpha): replace legacy attention_pool with MultiHorizonAttention [Stage 2]
Single source of truth for the attention path. Deletes the legacy
single-Q `attention_pool.cu` and all `attn_*` fields from
`PerceptionTrainer`; wires `MultiHorizonAttention` (the bundle
introduced in Stage 1) into `step_batched` + `evaluate_batched` as
THE attention summary that seeds CfC's `h_old` at k=0.
Deletions:
cuda/attention_pool.cu (244 lines)
perception.rs::attn_q_d/attn_context_d/
attn_weights_d/grad_attn_q_d/opt_attn_q/
attn_fwd_fn/attn_bwd_fn/_attn_module/
attn_grad_q_scratch_d (all struct fields)
perception.rs::ATTENTION_POOL_CUBIN (include_bytes constant)
Their corresponding init + struct-construction lines.
build.rs::KERNELS (drops "attention_pool")
New kernel + binding:
cuda/horizon_mean_collapse.cu (53 lines)
- `horizon_mean_collapse_fwd/_bwd`: collapses [B, N_H, H] → [B, H]
by averaging over the horizon axis. Single-pass, no reductions.
src/horizon_mean_collapse.rs (host binding)
MHA additions:
- `collapse` field + `ctx_mean_d` + `grad_ctx_mean_d` for the seed.
- `grad_ctx_h_d` scratch (split from grad_horizon_tokens_scratch to
avoid aliasing when MoE bwd writes d_ctx_h while pool bwd writes
d_horizon_tokens).
- `forward(ln_b_out)`: horizon-token pool → inverted pool → MoE
dispatch → mean-collapse → ctx_mean_d.
- `backward(ln_b_out, grad_ctx_mean, grad_ln_out)`: full reverse
chain.
- `apply_anchor()`: launches anchor_l2 on horizon_tokens, Q,
experts_w.
- `adamw_step()`: steps all 6 owned optimizer groups.
PerceptionTrainer integration:
- Section 2d (forward): `self.mha.forward(&self.ln_out_d)` replaces
the legacy attention_pool launch. CfC's h_old at k=0 now reads
`self.mha.ctx_mean_d.device_ptr` (was `self.attn_context_d`).
- Section 7c-pre (backward): `self.mha.backward(ln_out, grad_h_carry,
grad_h_enriched_seq)` replaces the legacy attn_bwd_fn launch.
- Four `reduce_axis0` launches collapse MHA's per-batch scratches
into shared gradient buffers: grad_horizon_tokens, grad_q,
grad_experts_w, grad_experts_b.
- `self.mha.apply_anchor()` adds L2 anchor grad contributions.
- Section 9 (AdamW): `self.mha.adamw_step()` replaces opt_attn_q.
- `evaluate_batched`: `self.mha.forward` replaces the legacy fwd
launch; h_old at k=0 reads `mha.ctx_mean_d`.
- `self.mha.zero_grads()` at step start (capture-safe memset_zeros).
BUG CAUGHT DURING WIRING (NVIDIA-grade discipline): first wiring
attempt mis-sized the reduce_axis0 launches for the MoE
`grad_w_scratch_d` ([B, N_H, N_E, H, H]). Initial `n_tail = N_H * N_E
* H * H = 327680` would have made reduce_axis0 read 5× past the end
of the buffer → CUDA_ERROR_ILLEGAL_ADDRESS. Fix: `n_tail = N_E * H *
H = 65536` with `n_batch = B * N_H`, treating the leading two axes
together as the reduction dimension. Caught by stacked_trainer test
on RTX 3050; would have caused silent corruption then a hard fault
on L40S/H100 later.
LOCAL VERIFICATION (RTX 3050 sm_86):
- ml-alpha builds clean (cuda feature).
- All 38+ tests PASS serially with --test-threads=1:
perception_overfit (8 tests incl. loss-shrinks)
trunk_forward (5)
stacked_loss_shrinks (multiple)
bce_grad_finite_diff (4)
snap_feature_assemble (9)
... (full suite green)
- Numgrad parity for the 4 new MHA kernels (horizon_token, inv_attn,
regime_moe_gate, anchor_l2) PASSES at 5e-2 rel.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
|
||
|
|
6a3f45d872 |
refactor(ml-alpha): consolidate BCE — Kendall σ is THE BCE, wire MHA into trainer
Single source of truth for the multi-horizon BCE per the new
feedback_single_source_of_truth_no_duplicates pearl. Eliminates the
`bce_loss_multi_horizon_sigma.cu` / `loss_sigma.rs` duplication
introduced earlier today and folds the Kendall σ-weighting into the
canonical kernel.
Deletions:
cuda/bce_loss_multi_horizon_sigma.cu (folded into legacy)
src/trainer/loss_sigma.rs (helper merged)
tests/bce_sigma_numgrad.rs (subsumed)
Rewrites:
cuda/bce_loss_multi_horizon.cu — replaced with σ-aware
implementation; kernel function name kept as
`bce_multi_horizon_forward_backward` so cubin symbol stays stable.
NVIDIA-grade warp-shuffle reduce (4 warps, 1 cross-warp barrier);
new args `log_sigma_h` (per-horizon Kendall σ logarithm) and
`d_log_sigma_h` (its gradient).
src/trainer/loss.rs — standalone helper updated to new 11-arg
kernel signature. Exposes optional `log_sigma_h` in `BceInput`
(None → zeros / passthrough Kendall init) and returns
`mean_bce_per_h` + `d_log_sigma_h` in `BceOutput`.
tests/bce_grad_finite_diff.rs — adds `log_sigma_h: None` to the
test inputs; all 4 numgrad tests PASS unchanged.
build.rs — drops `bce_loss_multi_horizon_sigma` entry from
KERNELS. The single canonical `bce_loss_multi_horizon` cubin
now contains the σ-aware kernel.
Wiring:
PerceptionTrainer gains a single `pub mha: MultiHorizonAttention`
field. Owns `log_sigma_h_d` + `grad_log_sigma_h_d` (and the rest
of the multi-horizon attention path, anchored on Stage 2 to fully
replace the legacy `attention_pool` path).
step_batched + evaluate_batched BCE callsites now thread
`mha.log_sigma_h_d` and `mha.grad_log_sigma_h_d` through the
11-arg kernel signature.
This commit keeps the legacy `attention_pool` callsite in place; the
Stage 2 commit will replace it with `mha.pool` + `mha.inv_pool` +
`mha.moe` and delete the `attn_*` fields entirely.
Verified locally: ml-alpha builds clean with the cuda feature,
bce_grad_finite_diff (4/4) PASS on RTX 3050.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
|
||
|
|
55aeddaebd |
feat(ml-alpha): anchor_l2 kernel + Wiener-α controller (v2 B) [V8]
L2 anchor regularization toward initialization (axis B). Anchors
horizon_tokens + Q + MoE experts toward their init values to prevent
the calibration drift observed in v1 (where val_loss climbed as α
opened past epoch 1 in 2 of 3 folds).
KERNEL (`anchor_l2_fwd_bwd`):
loss_out = λ · Σ_i (p[i] − p_init[i])²
grad_p[i] += 2λ · (p[i] − p_init[i])
- Warp-shuffle reduce; one block per parameter group; strided thread
loop over n. Cross-warp reduce uses one __syncthreads.
- Coalesced grad write via stride loop.
- λ passed as device-side [1]-buffer (host writes scalar before launch
— capture-safe).
CONTROLLER (`trainer::anchor_controller::AnchorController`):
- Signal-driven λ floor: λ_floor = ‖p_init‖ / (100 · √numel).
Cross-fold-persistent per pearl_kelly_cap_signal_driven_floors.
- Wiener-α smoother (α = diff_var / (diff_var + sample_var + ε))
on val_loss change; α floored at 0.4 per
pearl_wiener_alpha_floor_for_nonstationary.
- λ blends toward target = |ema_change|·scale with α; floored at
λ_floor per pearl_blend_formulas_must_have_permanent_floor.
- First-observation bootstrap (sentinel state replaced directly on
first step) per pearl_first_observation_bootstrap.
- 4/4 unit tests PASS: signal-floor init, bootstrap returns floor,
floor protection across 1000 steps, λ_max cap.
NUMGRAD VERIFICATION (RTX 3050 sm_86):
anchor_l2_numgrad PASSES with closed-form parity (machine precision)
and central-difference parity (4 random positions) within 5e-2 rel.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
|
||
|
|
11b96dac6a |
feat(ml-alpha): regime_moe_gate kernel + numgrad (v2 D) [V6]
Top-1 Mixture-of-Experts gate per Switch Transformer. Three kernels
in one .cu file:
- regime_moe_gate_fwd: select top-1 expert from gate_logits, apply
its [H, H] linear to fused_ctx → routed_ctx [B, N_H, H].
- regime_moe_gate_bwd: chain-rule grads through the SELECTED
expert's W and bias (sparse-by-expert scratches), plus
d_fused_ctx accumulator. Inactive experts get zero contribution.
- regime_moe_gate_aux: softmax of gate_logits + load-balancing
auxiliary loss (frac · prob_mean × N_EXPERTS).
ARCHITECTURE:
- N_EXPERTS = 4. Each expert is a [H, H] linear with bias.
- Total expert params: 4 · 128 · 128 + 4 · 128 = 66 KB. Cheap.
- STE on gate: gate logit grad is zero from the expert path (top-1
is non-differentiable); the load-balance aux loss provides the
differentiable signal that pushes routing toward balanced usage.
PERFORMANCE:
- Grid (B, N_H, 1), block (HIDDEN_DIM). One block per (b, h).
- Forward: each thread computes one output channel via a dot
product over HIDDEN_DIM input dims (#pragma unroll 8).
- Backward d_fused_ctx: each thread accumulates over d_out
sequentially (HIDDEN_DIM iterations) since the weight matrix
column is naturally aligned to the thread's d_in index.
- Backward d_W/d_b scratches are sparse-by-expert; downstream
reduce_axis0 collapses over (B, N_H).
- Top-1 chosen by thread 0 per block, broadcast via shared mem.
NUMGRAD VERIFICATION (RTX 3050 sm_86):
forward_matches_host_reference_and_backward_matches_numgrad
PASSES 11 checks (4 on d_W, 3 on d_b, 4 on d_fused_ctx) within
5e-2 rel / 5e-3 abs.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
|
||
|
|
d94696620d |
feat(ml-alpha): inverted_attention_pool kernel + numgrad (v2 E) [V3]
iTransformer-style cross-variate attention pass: each of HIDDEN_DIM
features becomes a "variate token" with its K-trajectory as embedding.
FORWARD:
X_inv[h, k] = ln_out[k, h] # transpose
scores[h, j] = inv_scale · Σ_k X_inv[h, k] · X_inv[j, k]
attn[h, j] = softmax_j(scores)
pooled[h] = Σ_j attn[h, j] · mean_k_X_inv[j] # mean-K commutes out
BACKWARD: three independent chains into ln_out:
- value path: (1/K) · attn[h', my_h] · d_pooled[h'] (per-k constant)
- query path: inv_scale · Σ_j d_scores[my_h, j] · X_inv[j, k]
- key path: inv_scale · Σ_h' d_scores[h', my_h] · X_inv[h', k]
Softmax bwd: d_scores[h, j] = attn · (d_attn - Σ_l attn · d_attn)
IMPLEMENTATION NOTES:
- First attempt cached attn [H, H] = 64 KB in shared mem → tripped
the 48 KB dynamic-shared limit on sm_86 (CUDA_ERROR_INVALID_VALUE).
- Fixed by moving d_scores to a DRAM scratch buffer; shared mem
holds only X_inv [H, K] (≤ 16 KB at K = 32). One block-wide barrier
between the d_scores write and the value/query/key accumulation.
- All per-batch slice writes; no atomicAdd, no cross-block race.
- Pooled computation uses the mean-K commute (Σ_k attn · X_inv =
attn · mean_k_X_inv), saving an entire H×K accumulation pass.
LOCAL VERIFICATION (RTX 3050 sm_86):
forward_then_backward_matches_central_difference PASSES 6 numgrad
checks on ln_out positions within 5e-2 rel / 5e-3 abs.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
|
||
|
|
e8f6c4721f |
feat(ml-alpha): horizon_token_attention_pool kernel + numgrad (v2 C) [V2]
Replaces the falsified per-horizon Q_h pool with a single shared
query Q over an extended key sequence [horizon_tokens; LN_b_out],
producing per-horizon outputs via TFT-style horizon-token mixing.
FORWARD:
scores[i] = Σ_d Q[d] · ext[i, d] (i ∈ [0, N_H + K))
attn = softmax_i(scores)
S[d] = Σ_k attn[N_H + k] · ln_out[k, d] (shared time agg)
ctx[h, d] = attn[h] · horizon_tokens[h, d] + S[d] (per-horizon out)
BACKWARD: full chain rule with softmax-centring; gradients to
horizon_tokens, Q, and ln_out via the saved attn weights.
NVIDIA-grade implementation per feedback_nvidia_grade_perf_for_kernels:
- Warp-shuffle reduce (block_reduce_sum / block_reduce_max helpers)
for all per-d dot products and softmax aggregates.
- Cross-warp reduce uses exactly one __syncthreads.
- Non-divergent shuffles: inactive lanes contribute 0 via ternary.
- Block-per-batch + horizon-loop inside block → grad_ln_out += is
race-free without atomicAdd.
- Smem layout computed at launch: [s_attn(N_H+K); s_warp(N_WARPS);
s_d_S(H) on bwd]. No over-allocation.
LOCAL VERIFICATION (RTX 3050 sm_86):
forward_then_backward_matches_central_difference PASSES 12 numgrad
checks (4 each on horizon_tokens / Q / ln_out) at 5e-2 rel / 5e-3
abs envelope. First-try pass.
NOTE: .gitignore adjusted with narrow allow-rules for crates/ml-alpha/{
cuda,src,tests}/horizon_token_* paths — the broad "*token*" rule
intended for auth tokens was hiding these source files. Explicit
allow keeps the security rule intact while exempting these specific
files.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
|
||
|
|
679ab3f5eb |
feat(ml-alpha): Kendall sigma-weighted BCE kernel + numgrad (v2 A) [V7]
New `bce_multi_horizon_sigma_forward_backward` kernel implementing the
Kendall homoscedastic uncertainty weighting per spec axis A:
raw_bce_h = Σ_{i in h, m_i=1} L_i
count_h = #{i in h : m_i = 1}
mean_bce_h = raw_bce_h / count_h
w_h = base_weight_h / (2 · exp(2 · log_sigma_h))
total_loss = Σ_h [ w_h · mean_bce_h + log_sigma_h ]
d L / d p_i = m_i · (w_h / count_h) · (p − y) / (p (1 − p))
d L / d log_sigma_h = 1 − 2 · w_h · mean_bce_h
NVIDIA-grade implementation per feedback_nvidia_grade_perf_for_kernels:
- Warp-shuffle reduction (`__shfl_xor_sync`) for both per-horizon
sums and the global valid count, replacing block tree-reduce.
- One `__syncthreads` for the cross-warp aggregate; no inner-loop
barriers.
- Non-divergent shuffles: inactive lanes contribute 0 via ternary,
never via `if (tid < N) shuffle`.
- Coalesced strided access in both forward and gradient passes.
- Pre-compiled cubin via build.rs; no nvrtc.
Independent of the legacy `bce_loss_multi_horizon` kernel — that one
stays untouched so eval/smoke paths are unaffected. The v2 trainer
wires this kernel in via commit V10.
Standalone helper `bce_sigma_loss_and_grad_gpu` in `trainer::loss_sigma`
for numgrad parity tests. Three numgrad tests all PASS on RTX 3050
(sm_86) within 5e-2 rel / 5e-3 abs:
- d_log_sigma_h ↔ central-difference (numgrad on log_sigma)
- grad_probs ↔ central-difference (8 random positions)
- total_loss ↔ closed-form reconstructed from mean_bce_per_h
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
|
||
|
|
41292303dc |
refactor(ml-alpha): remove per-horizon Q_h path (C21-C25 falsified) [V1]
3-fold A/B sweep 2026-05-18 at commit
|
||
|
|
286ea26e2a |
perf(ml-alpha): warp-shuffle reduce in per-horizon kernels
Cluster A/B sweep with C25 wiring showed 86 s/epoch vs 17 s/epoch baseline = ~5x regression. Root cause: per-horizon attention pool + residual head used block tree-reduce with 8 __syncthreads per K-step in a serialised K-loop, repeated H=5 times in both fwd and bwd = ~3200 barriers/step. Plus the prob_blend bwd reduce kernel ran with a single thread per block, fully serialising over K*B. Replacements: - per_horizon_attention_pool fwd/bwd: introduce block_reduce_sum / block_reduce_max helpers using intra-warp __shfl_xor_sync + cross-warp shuffle (1 syncthread per K instead of 8). Smem shrinks to [K + N_WARPS] / [2K + N_WARPS]. - per_horizon_residual_head fwd: same warp-shuffle reduce pattern. - per_horizon_prob_blend_reduce_alpha_residual: 1 thread → 1 warp per horizon, lane-strided reduction over K*B via shfl_xor_sync. Launch config updated to block_dim=(32,1,1). Tricky bug found while implementing: the cross-warp reduce in the residual head originally guarded `__shfl_xor_sync(0xffffffff, ...)` with `if (tid < PHR_N_WARPS)`, leaving 28 of 32 lanes in warp 0 outside the call. Mask 0xffffffff requires all 32 lanes to participate — divergence is UB and hung the full-pipeline smoke on Ampere/Ada. Fix: read s_warp via ternary into all 32 lanes, then shuffle inside `if (tid < 32)`. Matches the pattern used in block_reduce_sum. Verified locally on RTX 3050 (sm_86): per_horizon_attention_pool numgrad, per_horizon_residual_head numgrad, and per_horizon_full_pipeline_smoke (zero-init identity + non-zero end-to-end) all PASS. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> |
||
|
|
f50b466f77 |
feat(ml-alpha): PerceptionTrainer wires C21+C22 fwd+bwd+AdamW (C25)
The per-horizon attention pool from C21+C22 is now fully integrated
into step_batched's hot loop. Forward and backward both flow; AdamW
updates all four new parameter groups (Q_h, w_res, bias_res, α) every
training step. At α=0 init the contribution is byte-identical to
baseline; training discovers whether α should grow.
New kernel: cuda/per_horizon_prob_blend.cu
per_horizon_prob_blend_fwd
Reads logit_per_k_d (already stored by GRN forward) +
sigmoid(logit_baseline + tanh(α[h]) * residual[b, h]) →
overwrites probs_per_k_d in place. At α=0: r_contrib=0, output
== sigmoid(logit_baseline) == probs_baseline → bit-identical.
per_horizon_prob_blend_reduce_alpha_residual
Reads probs_per_k (= p_final, post-blend) + grad_probs_per_k
(= ∂L/∂p_final from BCE) and computes:
d_logit[k,b,h] = grad_probs[k,b,h] * p_final * (1 - p_final)
d_residual[b,h] = tanh(α[h]) * Σ_k d_logit[k,b,h]
d_alpha[h] = sech²(α[h]) * Σ_{k,b} d_logit[k,b,h] * residual[b,h]
No separate prob_blend_bwd needed — chain-rule equivalence
∂L/∂logit_baseline = ∂L/∂r_contrib (both flow through the same
sigmoid derivative) means the existing GRN backward is UNCHANGED.
trainer/per_horizon_state.rs extensions:
forward_with_blend(ln_b_out, logit_per_k, probs_per_k)
Pool fwd → context_h; head fwd → residual; prob_blend fwd
in-place rewrites probs_per_k.
backward_through_blend(probs_per_k, grad_probs_per_k, ln_b_out,
grad_ln_b_out_target)
Reduce kernel → d_residual + d_alpha. Then:
head bwd → d_w_res_scratch, d_bias_res_scratch, d_context.
pool bwd → d_q_h_scratch, += grad_h_enriched_seq_d.
Per-batch scratches reduced to shared grads host-side
(n_batch ≤ 64 → sub-millisecond on host).
adamw_step()
Steps the four optimizers using the shared grad buffers.
zero_grads()
Called once per step before forward to clear scratch.
trainer/perception.rs step_batched integration:
── 4.5 (after GRN K-loop, before BCE): zero_grads + forward_with_blend
overwrites probs_per_k_d with p_final.
── 5 (existing BCE consumes probs_per_k_d as today; grad_probs is
now ∂L/∂p_final automatically).
── 5a (after BCE, before ISV-lambda + heads bwd): backward_through_blend.
Existing GRN bwd path is UNTOUCHED — the chain rule absorbs
the bias.
── 9 (after existing 17 AdamW group steps): per_horizon.adamw_step
updates Q_h, w_res, bias_res, α.
Verification:
- 34 ml-alpha lib tests still green.
- Per-horizon kernel numgrad parity (C21, C22) still green.
- Per-horizon end-to-end pipeline smoke (C23, including the
alpha=0 byte-identity invariant) still green.
- Full workspace builds clean.
Closes the kernels+wiring portion of #203 (per-horizon attention pool
kernels + wiring). What remains (#204): 30-epoch × 3-fold A/B vs
single-Q baseline. The branch is ready for that sweep when GPU time
is budgeted; the implementation is structurally adoption-safe
(α=0 → identity to baseline) so it can be merged before the A/B if
desired.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
|
||
|
|
8c335caef7 |
feat(ml-alpha): per-horizon residual head kernel + numgrad parity (C22)
Companion kernel to C21's per_horizon_attention_pool. Computes a
per-horizon scalar residual from each horizon's context vector:
residual[b, h] = Σ_d w_res[h, d] * context_h[b, h, d] + bias_res[h]
Designed to be added (behind a learnable α-gate) to the existing
multi_horizon_heads logit output — keeps the existing GRN head kernel
completely unchanged. The per-horizon attention pool's contribution
flows through this lightweight projection without weight-shape
changes elsewhere or checkpoint-V2-bumping.
Path A integration sketch (deferred to follow-up commit C23):
alpha_logit_per_horizon = existing_head(h_K)[h] # from current path
+ tanh(α[h]) * residual_kernel(context_h)[h]
where α[h] is a learnable 5-vector init'd to 0 (no effect at start).
Training discovers per-horizon whether the residual contributes.
This is a strict superset of the existing path — α=0 → bit-identical
to today.
Backward kernel produces:
d_w_res — per-block scratch [B, N_HORIZONS, HIDDEN_DIM]
for host reduce_axis0 → shared [N_HORIZONS, HIDDEN_DIM]
d_bias_res — per-block scratch [B, N_HORIZONS], same reduction
d_context_h — per-batch indexed; += chained with attention bwd
Single-writer discipline preserved (no atomicAdd per
feedback_no_atomicadd.md); horizon loop inside the per-batch block.
Numgrad parity test:
- B=3, N_HORIZONS=5, HIDDEN_DIM=128 fixture.
- Loss = Σ residual_out (so d_residual = 1).
- Probes 8 random w_res indices, all 5 bias_res entries, 8 random
context_h indices via central-difference at ±eps=1e-2.
- All within 5e-2 rel-tol or 5e-3 abs-floor.
- Passes on RTX 3050.
Same scope discipline as C21: kernel + binding + numgrad first;
trainer wiring + α-gate + smoke training + A/B sweep follow once
both kernels are individually validated (now done).
Closes the second kernel-correctness portion of #203. Remaining:
C23: trainer wiring (capture attn_pool fwd into the graph; sum
residual into existing head output with α-gate)
C24: CheckpointV1 → V2 bump (add q_h, w_res, bias_res, alpha fields)
C25: 1-epoch smoke (assert no NaN, loss decreases vs baseline)
C26: 30-epoch × 3-fold A/B (#204) — decision gate per spec §0
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
|
||
|
|
edc449eecf |
feat(ml-alpha): per-horizon attention pool kernel + numgrad parity (C21)
First implementation slice of the per-horizon attention pool design
(docs/superpowers/specs/2026-05-18-per-horizon-attention-pool-design.md).
Lands the kernel + Rust binding + numgrad verification; downstream
wiring into PerceptionTrainer's captured graph + CheckpointV2 bump +
A/B sweep are follow-up commits gated on this proving correctness.
Kernel (cuda/per_horizon_attention_pool.cu):
per_horizon_attention_pool_fwd
Q_h[N_HORIZONS, HIDDEN_DIM] × LNb[B, K, HIDDEN_DIM]
→ context_h[B, N_HORIZONS, HIDDEN_DIM]
attn_h_weights[B, N_HORIZONS, K]
Per-block math identical to the single-Q variant, looped over
N_HORIZONS sequentially within each batch's block. Grid stays
(B, 1, 1) so backward grad_ln_out writes are race-free
(per feedback_no_atomicadd.md — no cross-block contention).
Per-batch shared mem ~k_seq + BLOCK + HIDDEN_DIM floats.
per_horizon_attention_pool_bwd
Same chain-rule pattern as attention_pool_bwd but with the horizon
loop inside the block: each (b, h) slice updates grad_ln_out in
place (sequential horizon accumulation), grad_Q_h is written as
per-block scratch [B, N_HORIZONS, HIDDEN_DIM] for host reduce.
Single-writer discipline preserved.
Rust binding (src/per_horizon_attention_pool.rs):
PerHorizonAttentionPool::{new, forward, backward}. Self-contained;
doesn't yet touch PerceptionTrainer or CfcTrunk. Loads the cubin
via the standard env!("OUT_DIR") path. Dynamic shared-mem byte
count computed per launch from k_seq.
Numgrad parity test (tests/per_horizon_attention_pool_numgrad.rs):
- B=2, K=8, HIDDEN_DIM=128, N_HORIZONS=5 fixture.
- Loss = Σ context_h (so d_context = 1 everywhere — clean analytical).
- Backward kernel produces analytical grads; central-difference of
forward kernel at ±eps=1e-2 across 8 random Q_h indices + 8 random
LNb indices verifies analytical matches CD within 5e-2 rel-tol or
5e-3 abs-floor.
- Passes on RTX 3050.
build.rs picks up the new .cu file automatically via the existing
KERNELS list; cubin compiles cleanly at sm_86 + sm_89.
Same scope discipline as Phase 2D.2 (VSN numgrad) — kernel correctness
first, integration second. The follow-up commit set per the spec §3
appendix:
C22: extend multi_horizon_heads.cu signature to accept per-horizon
context input + bump head_w shape to [N_HORIZONS, 2*HIDDEN_DIM]
C23: wire PerHorizonAttentionPool into PerceptionTrainer + CfcTrunk
captured graph behind AttentionPoolVariant config flag
C24: CheckpointV1 → V2 bump with discriminant + optional q_h field
C25: smoke training (one epoch, no NaN, loss decreases)
C26: 30-epoch × 3-fold A/B sweep (#204) — decision gate per spec §0
Closes the kernel-correctness portion of #203.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
|
||
|
|
a478ba3d84 |
perf(ml-alpha): block-per-batch attention pool bwd refactor (Phase B commit 4)
attention_pool_bwd refactored from grid=(1,1,1) to grid=(B,1,1). The existing per-batch grad_ln_out writes were already uniquely indexed; only grad_Q needed scratch+reducer (1 scratch, 1 reducer launch). Adds 1 per-batch grad scratch buffer + 1 reduce_axis0 launch: attn_grad_q_scratch_d [B, HIDDEN_DIM] ~16 KB scratch at B=32 — trivially small. attn_pool bwd runs 1×/step (not in K-loop) so the absolute wall-time win here is tiny. With this commit every single-SM bwd kernel in the trainer has been refactored to block-per-batch + scratch+reducer. Phase B kernel work complete. Next: local + cluster A/B perf benchmark to verify acceptance gates 6, 7, 8 from the spec. All 9 perception_overfit smokes pass. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> |
||
|
|
9607f33518 |
perf(ml-alpha): block-per-row VSN bwd refactor (Phase B commit 3)
variable_selection_bwd refactored from grid=(1,1,1) to grid=(B*K,1,1). VSN's n_rows = B*K positions (one row per (batch, K-position) pair); block-per-row matches the existing fwd kernel's layout. Adds 2 per-row grad scratch buffers + 2 reduce_axis0 launches: vsn_grad_w_scratch_d [B*K, FEATURE_DIM, FEATURE_DIM] vsn_grad_b_scratch_d [B*K, FEATURE_DIM] ~210 KB scratch at B=32, K=64. VSN bwd runs 1×/step (not K×) so the absolute wall-time win here is small versus commits 1+2. Done for pattern uniformity — every per-batch or per-row bwd in the trainer now uses scratch+reducer. All 9 perception_overfit smokes pass. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> |
||
|
|
5c2c3b65a8 |
perf(ml-alpha): block-per-batch GRN bwd refactor (Phase B commit 2)
multi_horizon_heads_grn_bwd_batched refactored from grid=(1,1,1) to grid=(B,1,1). Removes the single-SM bottleneck on the second-most-called K-loop kernel (64×/step like cfc_bwd). Adds 10 per-batch grad scratch buffers (one per GRN param tensor) + 10 reduce_axis0 launches collapsing B → final grad after the K-loop: grn_grad_w1_scratch_d [B, 5, HEAD_MID, HIDDEN] grn_grad_b1_scratch_d [B, 5, HEAD_MID] grn_grad_w2_scratch_d [B, 5, HEAD_MID, HEAD_MID] grn_grad_b2_scratch_d [B, 5, HEAD_MID] grn_grad_w_gate_scratch_d [B, 5, HEAD_MID] grn_grad_b_gate_scratch_d [B, 5] grn_grad_w_main_scratch_d [B, 5, HEAD_MID] grn_grad_b_main_scratch_d [B, 5] grn_grad_w_skip_scratch_d [B, 5, HIDDEN] grn_grad_b_skip_scratch_d [B, 5] Total: ~8 MB scratch at B=32. All 9 perception_overfit smokes pass (including stacked_trainer_loss_ shrinks_at_batch_32 which exercises the cross-batch reducer path on both cfc and GRN grads). Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> |
||
|
|
494a2e4827 |
perf(ml-alpha): block-per-batch cfc_step + reduce_axis0 reducer (Phase B commit 1)
Per docs/superpowers/specs/2026-05-17-kloop-parallelization-design.md. cfc_step_batched (fwd + bwd) refactored from grid=(1,1,1) with internal n_batch loop to grid=(B,1,1) — each block handles one batch. Removes the single-SM bottleneck on the K-loop's most-called kernel (64×/step). Param-grad accumulation moves to per-batch scratch: cfc_grad_w_in_scratch_d [B, n_hid, n_in] cfc_grad_w_rec_scratch_d [B, n_hid, n_hid] cfc_grad_b_scratch_d [B, n_hid] cfc_grad_tau_scratch_d [B, n_hid] Zeroed once per training step, K-loop's 64 bwd calls += into them, then 4 reduce_axis0 launches collapse B → final grad buffers (OVERWRITE) before AdamW. New AdamW-after-reducer invariant: final grads are meaningful only after the reducer has run in the current step. New reduce_axis0 kernel: single parameterised reducer [B, N] → [N] via block tree-reduce (no atomicAdd per feedback_no_atomicadd.md). Same pattern as layer_norm_reduce_param_grads — CUDA-Graph-safe. cfc_step_backward_batched shared-mem dropped from (B+1)*n_hid*4 to 2*n_hid*4 bytes per block (only one row of sd_pre needed per block bi). Tests: - New stacked_trainer_loss_shrinks_at_batch_32: FIRST test that actually exercises the cross-batch reduction code path; existing perception_overfit suite was all B=1. Initial 0.24 → final 0.00. - Scratch-clears test removed (explanatory comment kept): structurally hard to assert directly due to begin_capture/end_capture not executing kernels; the B=32 convergence smoke implicitly validates scratch zeroing since divergence would otherwise be immediate. All 9 perception_overfit smokes + 4 backward_finite_diff tests pass. build.rs: - KERNELS list adds "reduce_axis0" - Cache-bust → v11 Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> |
||
|
|
f76437c0c7 |
feat(ml-alpha): attention pool forward+backward CUDA kernels (Phase 3.1)
Single-head attention pool over Mamba2 K-positions, designed to replace
the CfC's zero-initialised `h_old` at k=0 with a learned content-
addressable summary over all K LN_b output positions. Forward math:
scores[k] = Q · keys[b, k, :] # [K]
attn[k] = softmax_k(scores) # [K]
context[h] = sum_k attn[k] * values[b, k, h] # [HIDDEN_DIM]
For our attention pool, keys == values == LN_b output [B, K, HIDDEN_DIM].
Single learned param: Q [HIDDEN_DIM]. Tiny (128 params).
Forward layout: grid = (B, 1, 1), block = HIDDEN_DIM=128 threads. Three
passes: (1) K dot-products with tree-reduce over HIDDEN_DIM, (2)
softmax over K with max-subtract+sum, (3) weighted sum into context.
Backward chain rule:
d_attn[k] = sum_h grad_context[h] * values[b, k, h]
d_scores[k] = attn[k] * (d_attn[k] - sum_kp attn[kp] * d_attn[kp])
d_Q[h] += sum_{b, k} d_scores[k] * values[b, k, h]
d_values[b, k, h] += grad_context[h] * attn[k] + d_scores[k] * Q[h]
Both `d_Q` and `d_values` use += semantics:
- d_Q: accumulates across batch (single block, internal n_batch loop).
- d_values: writes ADD onto whatever grad_ln_out already holds, so the
trainer can chain it on top of the K-loop's contribution to the LN_b
output gradient (no separate add-kernel needed).
Single-writer (no atomicAdd): one block per launch, thread h owns
column h of grad_ln_out for ALL (b, k). Internal n_batch loop matches
the GRN / VSN bwd pattern.
build.rs:
- "attention_pool" added to KERNELS
- Cache bust → v10
Wiring into PerceptionTrainer (Phase 3.2) is the follow-up commit:
add attn_q_d learned param + per-batch context + attn_weights buffers,
run attn_pool_fwd between LN_b fwd and the K-loop, use attn_context as
the K-loop's k=0 h_old (instead of zero_h_d), and chain attn_pool_bwd
after the K-loop reverse pass.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
|
||
|
|
c363a7e94c |
feat(ml-alpha): TFT VSN forward+backward CUDA kernels (Phase 2D.1+2D.2)
Per-position softmax-normalised feature gating for the trunk entry. Per (b, k) sample: gate_logit[i] = sum_j W_vsn[i, j] * x[j] + b_vsn[i] gates = softmax(gate_logit) # [FEATURE_DIM] y[i] = x[i] * gates[i] Backward chain rule (cleanly factored from the softmax Jacobian): d_gates[i] = grad_y[i] * x[i] d_logit[i] = gates[i] * (d_gates[i] - sum_j gates[j] * d_gates[j]) grad_W[i,j] += d_logit[i] * x[j] grad_b[i] += d_logit[i] grad_x[j] = grad_y[j] * gates[j] + sum_i d_logit[i] * W[i,j] Single-writer (no atomicAdd): thread tid owns row tid of grad_W and column tid of d_x_via_W. ONE block per launch (loops n_rows internally), same pattern as 2-layer / GRN bwd kernels. Softmax uses standard max-subtract + sum trick for numerical stability. Block dim = 64 (one warp + 24 idle threads at FEATURE_DIM=40). Wiring blocked on: Mamba2 backward needs to emit `d_input` (currently dropped at line 1413 of mamba2_block.rs via `_d_input`). Next commit exposes that so VSN bwd has the right grad_y signal — and the same refactor unblocks Phase 2B (2-stack Mamba2 needs the inter-stack LN to backprop through the 2nd stack's d_input). build.rs: - "variable_selection" added to KERNELS - Cache bust → v9 Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> |
||
|
|
005ded9722 |
feat(ml-alpha): TGN Δt Fourier features in snap_feature_assemble (Phase 2C)
Bumps FEATURE_DIM 32→40. Slots [32..40] now carry 8 TGN-style Fourier
features encoding the elapsed time Δt = ts_ns - prev_ts_ns:
(cos(ω_k · Δt_ns), sin(ω_k · Δt_ns))_{k=0..3}
at log-spaced periods [60s, 6s, 600ms, 60ms].
This gives Mamba2's input vector explicit Δt encoding that's
discriminative across temporal scales — particularly important once
decision-stride > 1 (Phase 2A) lands and the gap between K-positions
becomes irregular. Without these features the model has no way to
distinguish "1ms gap" from "1s gap" between consecutive K-positions.
Slots [0..32] unchanged (bit-equivalent for the first 32 features).
Reserved-zero slots [26..32] kept for future macro context. Slots
[20..26] still hold the loader-precomputed EMA regime cascade.
Frequencies stored in __constant__ memory (SNAP_DT_OMEGAS[4]) — small
table, broadcast read pattern, no register pressure. Frequency
selection rationale (one per log-decade):
60s — minute-scale macro session context
6s — 10s-scale liquidity windows
600ms — sub-second microstructure
60ms — tick-cluster spacing
Δt clamped to >= 0 so the rare out-of-order timestamp doesn't produce
nonsense angles. Each (cos, sin) pair satisfies cos²+sin² = 1
(verified by new test `dt_fourier_features_are_bounded`).
New tests in snap_feature_bit_equiv.rs:
- dt_fourier_features_are_bounded: |slot| <= 1 + cos²+sin² == 1
- dt_fourier_discriminates_scales: Δt=1ms vs Δt=1s produce L2-distinct
Fourier vectors (>0.1)
- reserved_slots_are_zero updated to check [20..32] (regime + reserved)
instead of [20..FEATURE_DIM]
All 8 perception_overfit smokes still pass (synthetic stride=1 and
stride=4 both converge 0.32 → 0.0000) — proves the wider FEATURE_DIM=40
input doesn't break the Mamba2+LN+GRN chain.
build.rs cache-bust → v8.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
|
||
|
|
5e23005dea |
feat(ml-alpha): TFT GRN forward+backward kernels for multi-horizon heads (Phase 1.7a)
Per-horizon GRN structure (Lim et al. 2021 §3.3 adapted to scalar output): eta_2[k, m] = GELU(W1[k, m, :] @ h + b1[k, m]) # [HIDDEN] → [HEAD_MID] eta_1[k, m] = W2[k, m, :] @ eta_2[k, :] + b2[k, m] # [HEAD_MID] → [HEAD_MID] gate_lin[k] = W_gate[k, :] @ eta_1[k, :] + b_gate[k] # → scalar main[k] = W_main[k, :] @ eta_1[k, :] + b_main[k] # → scalar skip[k] = W_skip[k, :] @ h + b_skip[k] # [HIDDEN] → scalar logit[k] = skip[k] + sigmoid(gate_lin[k]) * main[k] p[k] = sigmoid(logit[k]) Gated residual lets each per-horizon head learn "linear vs deeper-transform" gating, matching the regime-conditional alpha pattern from pearl_snapshot_alpha_is_regime_conditional (~20% of book states carry the edge; spread-Q4 hits 75% acc, middle quintiles below chance). Backward chain rule covers all 10 parameter tensors + the trunk gradient (skip-path direct + main-path through W2→GELU→W1, lambda-scaled). Single-writer discipline (no atomicAdd per feedback_no_atomicadd.md): - Thread m owns row m of grad_w1 (col i in 0..HIDDEN), row m of grad_w2 (col m_in in 0..HEAD_MID), and column m of d_eta_2. - Threads 0..4 own per-horizon scalar grads (skip/gate/main biases). - Trunk grad_h tiles i over 2 strides of HEAD_MID for HIDDEN=128 coverage. Shared mem: ~6.5KB (s_a1 + s_z2 + s_d_eta1 + s_d_eta2 + s_d_z1 + scalars), well within 48KB limit. Existing 2-layer MLP kernels (Tasks 1.3/1.4) stay in the cubin as ablation baseline; the wired path becomes GRN once perception.rs lands. build.rs cache-bust → v7. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> |
||
|
|
59e236b4e8 | feat(ml-alpha): 2-layer heads backward kernel with ISV lambda (Phase 1.4) | ||
|
|
e0a497da9e | feat(ml-alpha): 2-layer GELU MLP heads — forward kernel (Phase 1.3) | ||
|
|
167f065647 | feat(ml-alpha): LayerNorm backward + per-row param-grad reducer kernels | ||
|
|
d8cd90c130 | feat(ml-alpha): LayerNorm forward kernel for trunk-pre-CfC normalisation | ||
|
|
00da163078 |
feat(ml-alpha): h6000-aligned ISV — uniform BCE + z-score lambda + K=64
Three correlated fixes addressing the architectural inconsistency
surfaced by the 3-fold ISV CV: we built a horizon-aware gradient
controller (ISV) but suppressed its target horizon (h6000) to 0.36%
of the loss via auto-horizon-weights, then used a ratio formula
that never approached its own clamp ceiling. ISV's lambda was
operating on rounding error.
(1) Uniform BCE weights as auto-default
trainer/perception.rs: `auto_horizon_weights` now returns
[1.0; 5] regardless of seq_len. Prior schedule `min(1, K/h)`
gave h6000 weight 0.0053 at K=32 — combined with lambda ~1.04,
h6000's effective loss contribution was ~0.37%, indistinguishable
from zero. With uniform weights, each horizon contributes 20% and
ISV's lambda actually has something to scale.
(2) Z-score lambda derivation
cuda/horizon_lambda.cu: replace `ratio = ema_h / mean(ema)` with
`z_h = (ema_h - mean) / std(ema); lambda = clamp(1.0 + 0.5*z, 1.0, 2.0)`.
Per `pearl_zscore_normalization_for_magnitude_asymmetric_signals.md`
z-score makes lambda spread scale-invariant of the absolute EMA
level. The ratio formula gave lambdas ≤ 1.04 in our data because
per-horizon BCE clusters tightly (range ~0.04) while mean is
~0.65. Z-score fills the [1.0, 2.0] envelope: 1σ → 1.5, 2σ →
ceiling. Boost-only asymmetric clamp preserved.
Test verification on the existing smoke (after 5 steps):
ema = [0.526, 0.522, 0.608, 0.641, 0.553]
lambda = [1.00, 1.00, 1.40, 1.76, 1.00]
Previously with ratio formula, max lambda on the same data
would have been ~1.05. h1000 (1.5σ above mean BCE here) now
gets a 76% trunk-gradient boost vs uniform.
(3) Default --seq-len 32 → 64
examples/alpha_train.rs: K=32 gives the model 0.5% of the
h6000 prediction window as in-window context. K=64 doubles
that, giving Mamba2's SSM state more material to build
long-horizon predictions. Within the kernel's MAMBA2_KERNEL_SEQ_MAX
cap of 96. Per-epoch wall scales ~K (more K-loop launches in
the captured graph, ~2× wall at K=64 vs K=32 for the K-loop
portion of dispatch).
Cache-bust v5 in build.rs to force nvcc recompile against the new
horizon_lambda.cu formula on the cluster's /cargo-target PVC. Old
cubins compute a numerically different lambda; running them against
the new Rust loop would silently apply the wrong gradient scaler.
Validation: 7 perception_overfit tests + 26 lib + 23 integration
ml-alpha tests pass. Synthetic overfit still converges to 0.0006.
horizon_ema_and_lambda_track_after_training observes the new
lambda spread (1.0-1.76) and asserts the asymmetric clamp envelope.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
|
||
|
|
a45fd85986 |
feat(ml-alpha): raise Mamba2 state cap 16→32 + add auc_h6000 early-stop
Three correlated changes for the next CV round: 1. Mamba2 state_dim cap: 16 → 32 cuda/mamba2_alpha_kernel.cu: MAMBA2_ALPHA_MAX_STATE_D 16 → 32. Per-thread state register `float x[32]` (128 B/thread) and per-thread x_hist replay cache `float x_hist[K*32]` (up to 12 KiB/thread of local memory at K=96). L40S/H100 register file (256 KiB/SM) absorbs this without occupancy collapse for our block dims (32-128 threads). Update Rust-side MAMBA2_KERNEL_STATE_MAX + validation message + test name. Kernel header doc updated. 2. New early-stop option: auc_h6000 examples/alpha_train.rs: add the long-horizon AUC as a third early-stop metric. The ISV CV ( |
||
|
|
0171c8c0ea |
fix(ml-alpha): asymmetric lambda clamp [1.0, 2.0] — boost-only ISV
3-fold A/B CV ( |
||
|
|
5d42ab0e98 |
feat(ml-alpha): ISV horizon weighting Phase 3 — wire lambda into trunk grad
Connects the per-horizon lambda computed by horizon_ema_and_lambda (landed in |
||
|
|
37c3a8f4d7 |
feat(ml-alpha): ISV-driven per-horizon EMA + lambda (Phase 1+2)
Foundation for replacing the static `--auto-horizon-weights` formula (`min(1, K/h)`) with a signal-driven per-horizon gradient scaler. Per `feedback_isv_for_adaptive_bounds.md`: adaptive bounds live in ISV, not hardcoded constants. Per `pearl_adam_normalizes_loss_weights.md`: Adam normalizes per-loss weight lifts (SP13 saw 13× aux_w produce only 0.6%/epoch divergence), so the effective lever is scaling the GRADIENT into the shared trunk, not the BCE coefficient. This commit sets up the EMA + lambda infrastructure; Phase 3 (wiring lambda into heads_bwd to actually scale the trunk gradient) is gated on the 3-fold CV results from |
||
|
|
ab94ce2a49 |
perf(ml-alpha): device-resident AdamW step counter (capture prep)
Stage 1+2 of #162 (CUDA Graph capture of training step). The AdamW kernels previously took the step counter as a host scalar arg, which gets baked into kernel args at CUDA Graph capture time — replays would freeze the counter and produce wrong bias-correction values. Both AdamW variants now read the step from a device pointer, advanced by a tiny 1-thread `increment_counter` kernel that goes inside the captured region. Each replay correctly increments and observes the new step value. Kernel changes: adamw_step.cu: - adamw_step: int step → const int* step_ptr - adamw_increment_counter: new, +=1 on step_ptr[0] mamba2_alpha_kernel.cu: - mamba2_alpha_adamw_step_devscale: int t → const int* step_ptr - mamba2_alpha_increment_step_counter: new Rust changes: trainer/optim.rs (AdamW): - host `step: i32` → device `step_count_d: CudaSlice<i32>` - step(): launch increment kernel BEFORE adamw kernel; both read device counter via pointer arg. - step_count(): test-only accessor, mapped-pinned readback (sync). mamba2_block.rs (Mamba2AdamW): - kept host `step_count: i32` for legacy paths (`step`, `step_from_buffers`) which aren't capture-compatible anyway (host grad-norm dtoh, host scalar grad_scale). - added device `step_count_d: CudaSlice<i32>` for the production gpu_clip path; advances via `kernel_increment_step` kernel inside the captured region. - adamw_apply_devscale: `t: i32` → `step_d: &CudaSlice<i32>`. Validation: - 4 adamw_invariants tests pass (step_count_increments specifically exercises the device counter). - 10 mamba2_block lib tests pass (training_loop_decreases_loss exercises legacy host-counter path). - Synthetic overfit smoke: initial=0.25 → final=0.0006 (matches pre-refactor trajectory bit-for-bit-equivalent). Stage 3+4 (capture brackets + first-call-capture / subsequent-replay state machine in step_batched) follows in the next commit. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> |
||
|
|
b6fb720acd |
perf(ml-alpha): eliminate all dtoh from training hot path
GPU-resident grad-norm + clip-scale; mapped-pinned loss readback.
Replaces 9× memcpy_dtoh per Mamba2 AdamW step (grad-norm host roundtrip)
+ 1× per-step download() (loss). Saves ~10 stream-sync barriers/step.
New kernels (cuda/grad_norm.cu):
- grad_norm_sq_phase1: per-block tree-reduce of x[i]^2 (no atomicAdd)
- grad_norm_sq_phase2: cross-tensor accumulator (sequential stream-ordered)
- grad_clip_scale: writes min(1, max_norm/sqrt(norm_sq)) to device ptr
- mamba2_alpha_adamw_step_devscale: reads grad_scale from device pointer
instead of host scalar, allowing AdamW kernels to launch async without
waiting for a CPU-side norm computation.
Trainer changes (perception.rs):
- loss_d kept device-side; mapped-pinned MappedF32Buffer shadow.
- Single stream.synchronize() at end of step (was 2: post-bwd + download).
- DtoD copy loss_d → loss_host_d queued, then sync flushes both kernels
+ copy in one barrier. Loss read via host_ptr (no dtoh).
Mamba2 AdamW (mamba2_block.rs):
- step_from_buffers_gpu_clip(): all grad-norm tensors processed via
phase1+phase2 chain, scale computed on-device, AdamW launches with
devscale variant. Zero host roundtrips.
- Pre-allocated block_partials_d, grad_norm_sq_d, grad_scale_d.
Optimizer (optim.rs): removed redundant stream.synchronize() per AdamW step.
Each per-tensor AdamW kernel is stream-ordered; sync only needed before
host reads, which the trainer handles centrally.
Synthetic overfit smoke: initial=0.30 → final=0.0006 (matches pre-refactor
trajectory). Full ml-alpha test suite passes (45 tests across lib +
integration).
Honors:
- feedback_no_htod_htoh_only_mapped_pinned.md (MappedF32Buffer only)
- feedback_no_atomicadd.md (block tree-reduce only)
- feedback_no_legacy_aliases.md (step_from_buffers replaced, not aliased)
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
|
||
|
|
c70c5cdf21 |
perf(ml-alpha): fused batched snap_feature_assemble kernel (#4)
Previously the per-step snap_feature path did B*K = 768 single-snapshot kernel launches (at B=8, K=96) + 768 DtoD copies into the window tensor. New `snap_feature_assemble_batched` processes all B*K snapshots in a SINGLE launch and writes outputs directly into the window tensor's storage. Per-step CPU work: pack 12 mapped-pinned staging buffers (~150 KB total host writes), then 10 DtoD copies of the staging → device buffers. Per-step GPU work: 1 batched kernel launch with B*K threads (each writes 32 floats to its output row). Mapped-pinned staging buffers cover the full B*K capacity at trainer init — no per-step allocation. New `MappedI32Buffer` and `MappedI64Buffer` types parallel `MappedF32Buffer` to stage `trade_count` (i32) and `ts_ns` / `prev_ts_ns` (i64) without violating the no-htod rule (`feedback_no_htod_htoh_only_mapped_pinned.md`). Dead per-snapshot scratch + helpers (`bid_px_d`, `snap_feat_d`, `stg_bid_px`, `snap_fn`, `upload_into`, etc.) removed per `feedback_no_legacy_aliases.md` — the only callers were the per-snapshot path, gone. Expected per-step savings: ~5-10 ms launch + DtoD overhead at B=8, K=96. Over 2000 steps/epoch = 10-20 sec/epoch. 77 ml-alpha tests pass. Synthetic overfit unchanged. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> |
||
|
|
737f8e72fa |
fix(ml-alpha): commit transpose_3d_swap_01 kernel (was uncommitted)
This kernel was referenced by PerceptionTrainer.step_batched / evaluate_batched (commit |
||
|
|
829ddfa62c |
feat(ml-alpha): add batched cfc + heads CUDA kernels (foundation for #8)
Adds 4 new kernel symbols alongside the existing single-sample ones —
zero changes to current call sites, so the in-flight qf5mj baseline is
unaffected. The next commit wires these into PerceptionTrainer's K
loop and exposes --batch-size in the CLI.
cfc_step_batched processes [n_batch, n_in/n_hid] tensors
cfc_step_backward_batched same; shared mem holds sd_pre[B, n_hid]
+ sdecay[n_hid] (~16 KiB at B=32, well
under L40S 48 KiB block limit). Param
grads (grad_b/grad_w_in/grad_w_rec/
grad_tau) accumulated via += — thread i
is sole writer to its row across all
samples, so no atomicAdd and no per-
batch scratch buffer.
multi_horizon_heads_batched [n_batch, 5] sigmoid outputs from
[n_batch, 128] hidden inputs.
multi_horizon_heads_backward_batched
shared mem holds sd_z[B, 5]. grad_w
/ grad_b += across batch (thread tid
sole writer). grad_h carries the
optional per-sample grad_h_carry
(cfc recurrence chain).
Design notes:
- Threading: one block of n_hid threads. Each thread loops over
b ∈ 0..B internally. This avoids cross-block races on grad_*
buffers and keeps the existing "no atomicAdd" discipline. Cost:
less raw parallelism than grid-batching, but the bottleneck is
Mamba2 (already batch-parallel via its own kernel grid).
- Per-thread accumulators: grad_b / grad_tau land in registers,
flushed once at end. grad_w_in / grad_w_rec written += per-b
(thread sole writer to its row, safe).
- All B samples processed in stream order inside one kernel launch
— saves K * (B-1) launches per sequence vs serialising B
independent calls.
77 ml-alpha tests pass (kernels not yet exercised — wiring is the
next commit).
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
|
||
|
|
85ce295773 |
feat(ml-alpha): per-horizon BCE weighting (fixes label-correlation inflation)
For seq_len K and horizon h with h ≫ K, the K position-supervised labels in a single sequence are near-identical (sequential positions' forward windows overlap by ~(h-1)/h). Per-position BCE therefore treats ~K highly-correlated labels as independent samples, inflating gradient pressure on long horizons by a factor of K. Concretely at K=96: h=30 → ~3 effective samples per seq (forward windows overlap ~97%) h=100 → ~1 (~99%) h=6000 → ~1 (~99.98%) Per-position supervision was paying 96× the natural signal density on h=6000, pulling the model toward fitting noise at long horizons. Fix: the fused BCE kernel now accepts an optional `loss_weights[N_HORIZONS]` (nullptr → uniform = no-op). Each (k, h) loss + grad contribution is multiplied by w_h; the normaliser is the sum of weighted valid entries instead of the raw valid count. `auto_horizon_weights(K, horizons)` computes `w_h = min(1.0, K/h)` so short horizons stay at full weight and long horizons collapse to their independent-sample density. Exposed via CLI: --auto-horizon-weights # K/h auto-derived --horizon-weights "1,1,0.5,0.1,0.02" # explicit floats Default behaviour is uniform (1.0) — apples-to-apples with the in-flight qf5mj baseline. Synthetic overfit still 0.6268 → 0.1144 in 250 steps (82% drop). 77 tests pass. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> |
||
|
|
248d8fe510 |
feat(ml-alpha): full architectural pass — recurrent CfC, GPU BCE, regime features, training discipline
Comprehensive fix for the issues identified after the BPTT-unroll cluster
run plateaued at val_loss ~0.692 with oscillating AUCs:
ARCHITECTURE
- CfC h_old is now RECURRENT across positions. Previously reset to
zero every step → CfC degenerated to a per-cell tanh-FC layer.
New: h_old at step k IS h_new at step k-1. Heads still operate
on h_new_k, but now the CfC actually carries state. Reverse-order
backward through the K positions accumulates grad_h_old → grad_h_new
via the new optional `grad_h_carry` arg on multi_horizon_heads_backward.
- tau is TRAINED. cfc_step_backward now writes grad_tau (per-cell decay
constant derivative), trainer gets a 7th AdamW group at 0.1× cfc lr.
- 6 NEW regime features (EMA cascade computed loader-side per file)
fill slots out[20..26] of snap_features. Gives the model multi-minute
trend / volatility / liquidity context that is structurally unreachable
inside the K-snapshot BPTT window. Slots: mid-z (med/slow), trend
signal, log-vol slow, log-spread med, log-trade-rate med. All bounded
via log1p / signed-log so no tuned constants leak in.
PERFORMANCE (NVIDIA-style)
- GPU-fused multi-horizon BCE for the entire [K, N_HORIZONS] grid in
ONE launch (was K host roundtrips). Native NaN-label masking.
- K-loop is fully GPU-resident: pre-allocated per-K scratch
(h_new_per_k, probs_per_k, labels_per_k, grad_probs_per_k), zero
device allocs inside step(). Only TWO syncs per sequence (after
forward, after backward) vs previously 2K+1.
- Stream-ordered kernel launches with pointer-offset addressing into
per-K buffers — host doesn't wait between K iterations.
- cfc_step_backward / multi_horizon_heads_backward both use += grad
semantics; trainer pre-zeroes accumulators once per step().
- MAMBA2_ALPHA_MAX_K capped at 96 (was temporarily at 256). 96 covers
h=30/100/300 with room; regime features handle h=1000/h=6000.
TRAINING DISCIPLINE
- LR schedule: linear warmup (default 200 steps) + cosine decay to
lr * lr_min_factor (default 0.1). Applied per training step to both
CfC and Mamba2 AdamW groups via new set_lr_cfc/set_lr_mamba2.
- Best-checkpoint tracking by val_loss; recorded in summary
(best_epoch, best_val_loss, best_val_auc).
- Early stopping on val_loss plateau (default patience = 3).
- CRITICAL BUG FIX: validation now uses new `evaluate()` method
(forward-only) instead of `step()`. Previous CLI called step()
on val data, which ran the full backward + AdamW update on the
validation set. With per-step BPTT that's ~K× more pressure than
the old comment ("statistically negligible") assumed.
Synthetic overfit: 0.6442 → 0.1233 in 250 steps (81% drop, sharper
than the previous 70%). 77 ml-alpha tests pass.
Local 2Q smoke (seq_len=64, 600 train seqs/epoch, 4 epochs):
val_loss 0.7011 → 0.6990, best epoch=1, h300 AUC 0.565 in epoch 0.
Phase E.3 callers (ml/examples/alpha_baseline.rs,
alpha_dqn_h600_smoke.rs) use the LEGACY Mamba2 forward_train +
backward_from_h_enriched path — unaffected by these changes (their
kernels are pre-zeroed via alloc_zeros, so the += grad semantics
remain correct in single-call mode).
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
|
||
|
|
485150c7b7 |
feat(ml-alpha): lift Mamba2 kernel seq_len cap from 32 → 256
Previous BPTT-unroll run had val_loss trending (0.6941→0.6922 over 5
epochs) but AUCs oscillating around 0.50 — the architecture lacked
context for medium/long horizons (h300, h1000, h6000 ≫ seq_len=32).
Phase 1d.2 validated the SSM at seq_len=6000; this is a step toward
restoring useful sequence depth.
Bumps:
- MAMBA2_ALPHA_MAX_K constant: 32 → 256
- x_hist per-thread replay buffer: 2KB → 16KB (spills to
DRAM-backed per-thread local memory; L2-cached, acceptable
perf cost vs the 8x context gain)
- Mamba2BlockConfig::validate updates the cap
Backward compat: legacy Phase E.3 callers (alpha_baseline,
alpha_dqn_h600_smoke) only use K=12 / K=32 — unaffected by the
larger compile-time max.
Synthetic overfit still converges 0.7135 → 0.2079 in 250 steps.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
|
||
|
|
16f5febf27 |
feat(ml-alpha): per-step supervision unrolls BPTT through full sequence
The final-step-only trainer (one BCE prediction per 32-snapshot
window) trained flat at chance on real ES data despite working on
synthetic overfit: train_loss=0.6953, val_loss=0.6943 across 40k
gradient steps. Gradient density was the bottleneck — one supervised
position per sequence × ~8K seqs/epoch isn't enough signal for the
SSM to find the alpha.
This commit supervises the model at EVERY position in the sequence:
mamba2_alpha_scan_fwd_seq — emits h_enriched at every t step
([N, K, sh2] instead of [N, sh2])
mamba2_alpha_scan_bwd_seq — accepts d_h_enriched_seq, injects
gradient at each t before propagating
d_state through the gate chain.
d_w_c and d_h_s2 accumulate across t.
PerceptionTrainer.step() — loop k=0..K; cfc + heads + BCE at
each valid label; cfc/heads grads
accumulate via += in kernel writes.
One Mamba2 backward call consumes the
full grad_h_enriched_seq.
cfc_step_backward — grad_w_in/w_rec/b writes changed
to += (callers MUST pre-zero).
multi_horizon_heads_backward — grad_w/grad_b writes changed to +=.
alpha_train.rs — passes per-position label rows to
step(); AUC still scored from
last-position predictions.
Phase E.3 callers (alpha_baseline.rs, alpha_dqn_h600_smoke.rs) use
the LEGACY Mamba2 forward_train + backward path with `alloc_zeros`
grad buffers — unaffected.
Synthetic overfit still converges 0.6664 → 0.1976 in 250 steps.
Local 2-quarter ES.FUT smoke shows the val AUC at h300 climbing
0.513 → 0.566 over 3 epochs (was flat-at-chance before). First
gradient signal we've gotten through the new architecture.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
|
||
|
|
2289fa062a |
feat(ml-alpha): normalize snap features to O(1)-O(10) scale
First L40S run (alpha-perception-s6hqv, commit
|
||
|
|
deed15b34e |
feat(ml-alpha): cfc_step_backward emits grad_x for upstream chain
Adds grad_x[k] = sum_i d_pre[i] * W_in[i,k] computed by thread 0 of the cfc_step_backward kernel (after the existing __syncthreads in the shared-mem sd_pre relay). Required by the stacked Mamba2 -> CfC design: Mamba2.backward_from_h_enriched needs grad on h_enriched, which is the CfC's "x" input in the stacked topology. For the existing CfC-only PerceptionTrainer (x = snap_features, no upstream learnable layer), grad_x is computed but discarded into a preallocated buffer. backward_finite_diff tests still pass (4/4) — the new arg is the 14th positional kernel arg; existing callers updated. perception_ overfit smoke still passes (loss 0.5669 -> 0.0665 in 200 steps). Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> |
||
|
|
753284f2ae |
feat(ml-alpha): backward kernels — heads + cfc_step (K=1 BPTT)
multi_horizon_heads_backward: sigmoid + linear chain rule. One block, HIDDEN_DIM=128 threads. Computes grad_w, grad_b, grad_h_in. No atomicAdd; per-thread accumulation only. cfc_step_backward: truncated K=1 BPTT through one CfC time step. Forward pre/decay/tanh recomputed inside the kernel; emits grad_w_in, grad_w_rec, grad_b, grad_h_old. tau is held frozen (structural log-uniform init per Hasani 2022; backprop through tau deferred to Phase A v2 if the gate needs it). Uses dynamic shared memory for the d_pre relay between threads (size = 2 * n_hid * 4 bytes). Tests (4/4 on sm_86) validate via on-GPU finite-difference: - heads grad_h vs forward(h±eps) → matches at eps=1e-3, rel<=1% - heads grad_b vs forward(b±eps) → matches at eps=1e-3, rel<=1% - cfc grad_b vs forward(b±eps) → matches at eps=1e-3, rel<=5% - cfc grad_h_old vs forward(h_old±eps) → matches at eps=1e-3, rel<=5% CPU is not the reference (per feedback_no_cpu_test_fallbacks.md). The kernel is the truth; numerical perturbation validates the analytic gradient against the kernel's own forward. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> |