7abfd3b0b6eb4e447f9cf7ea2ce4e0cf6bad742e
85 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
842e90aeb0 |
feat(aux-heads): rewrite for A+B paired (BCE + conditional-Huber) (CB3+CB4)
CB1+CB2 swapped labels D→A+B; this swaps the kernels to match. aux_heads.cu — 4-output structure: - Forward: 12 outputs per snapshot = 4 per direction × N_HORIZONS (prof_long_logit, size_long_pred, prof_short_logit, size_short_pred, each [N_HORIZONS]). Linear projections; sigmoid applied in BCE kernel. - Backward: accepts 4 grad_y inputs, produces 8 grad_W + 8 grad_b + grad_h_aux. Cooperative h_aux staging in shmem once per block. - 8 weight matrices total, Xavier × 0.1 init under scoped_init_seed. aux_loss.cu — 2 kernels: - aux_bce_loss_fwd_bwd: class-weighted BCE+sigmoid fused. pos_weight in shared mem; scales positive-class gradient. NaN-mask y_true. - aux_huber_masked_fwd_bwd: Huber w/ NaN-mask. CB1's y_size=NaN at y_prof=0 provides the conditional-Huber semantics naturally — no separate mask buffer needed. aux_heads.rs: - AuxHeads + AuxHeadsWeights: 8 buffer fields (4 W + 4 b) - AuxBceLoss + AuxMaskedHuberLoss wrappers replace AuxHuberLoss - POS_WEIGHT_MIN/MAX = [1.0, 50.0] clamps per E3 - aux_heads_fwd_gpu/aux_heads_bwd_gpu/aux_bce_loss_gpu/aux_huber_masked_loss_gpu perception.rs (minimal compile-keeping signature updates only): - Renamed/added buffers: 4 prediction (prof/size × long/short), 4 label staging, 4 grad_y per-K, 8 head grad scratches, 8 head Adam optimizers, 2 pos_weight buffers (device + staging) - HOLDING PATTERN: bwd zeroes the 4 grad_y_per_K buffers each step so the head Adam updates are no-ops on grad=0 (no aux gradient signal this commit). CB5 wires the actual aux_bce + aux_huber_masked calls. - BCE direction signal + dir_acc readouts updated to use the new prof_long/prof_short prediction buffers (so existing perception_overfit aux test still passes). 7 GPU oracle tests on RTX 3050 sm_86, all pass in 2.43s: - fwd_matches_naive_reference, bwd_finite_diff_matches_bias_sample, aux_bce_loss_matches_naive_reference, aux_bce_pos_weight_scales_positive_gradient (verified: pos_weight=10 → 10× gradient ratio within 1e-4), aux_huber_masked_does_not_propagate, aux_huber_masked_covers_both_branches, aux_bce_nan_mask_does_not_propagate Cubins rebuilt: aux_heads (8736→10528 bytes, +20% for 4-head fwd/bwd), aux_loss (6944→13344 bytes, +92% for 2 kernels). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
21e7dfd63c |
feat(aux-supervision): wire AuxTrunk + AuxHeads + Huber loss into PerceptionTrainer (B5)
Wires the aux supervision path parallel to BCE: - AuxTrunk (64-hidden single-bucket CfC) consumes the same encoder output as the main BCE trunk - AuxHeads (linear regression long/short) maps aux_trunk output to per-(direction, horizon) predicted outcomes - AuxHuberLoss supervises against D-style labels from MultiHorizonLoader Backward path with asymmetric stop-grad at encoder boundary: - aux_trunk gets gradient signal into its OWN params at all times - aux_trunk's encoder-boundary gradient is INITIALLY blocked (stop_grad_aux_to_encoder = true) - Conditional lift per E3 design: if aux_huber_ema < 0.4 AND aux_dir_acc_ema > 0.85 within 200 steps, lift the stop-grad - When lifted, aux_vec_add kernel folds aux's grad_x into the main grad_h_enriched_seq slot (element-wise += per feedback_no_atomicadd) ISV signals added: aux_huber_per_h, aux_dir_acc_per_h (per pearl). Per-trunk scratch + reduced grad buffers (no Adam state sharing per pearl_adam_normalizes_loss_weights — opt_aux is its own Adam group). New helper kernel cuda/aux_vec_add.cu: position-local dst += src for the asymmetric stop-grad lift accumulation. New synthetic test stacked_trainer_aux_supervision_converges_on_constant_signal validates end-to-end: aux_huber_ema_per_h = [0.087, 0.087, 0.087] (converged) aux_dir_acc_ema_per_h = [1.0, 1.0, 1.0] (perfect on constant) stop_grad_aux_to_encoder = false (lift fired) All 5 stacked_trainer tests pass on RTX 3050 (lib still converges, no regression from parallel aux wiring). Not yet consumed by decision policy (B7) — aux output flows through training only. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
ff70734802 |
feat(aux-heads): linear regression heads + Huber loss for aux supervision (B4)
Adds the aux heads that map AuxTrunk's output (64-dim) to per-(direction, horizon) predicted outcomes, plus the Huber loss kernel that supervises them against the D-style labels generated in B1. Files (1212 LOC): - cuda/aux_heads.cu (197 LOC): fused fwd+bwd for linear projection h_aux[B, 64] → y_long_hat[B, N_HORIZONS] + y_short_hat[B, N_HORIZONS]. Single launch for both directions, per-batch grad scratch + caller-side reduce_axis0, cooperative h_aux staging in shmem. - cuda/aux_loss.cu (143 LOC): Huber loss fwd+bwd with NaN-masking. Per direction call; returns Σ Huber + valid_count separately so caller picks reduction policy. - src/aux_heads.rs (412 LOC): AuxHeads + AuxHuberLoss wrappers, weight structs, Xavier init under scoped_init_seed. - tests/aux_heads.rs (460 LOC): 4 #[ignore]'d GPU oracle tests (fwd_matches_naive, bwd_finite_diff_matches_bias, huber_loss_matches_ naive, huber_loss_nan_mask_does_not_propagate). 4/4 PASS on RTX 3050. - build.rs: KERNELS += aux_heads, aux_loss; cache-bust bumped to v17. - src/lib.rs: pub mod aux_heads. Design decisions (full notes in subagent report): - Linear heads (no GRN) — D-labels already encode asymmetric loss-aversion - Per-direction Huber launches (cleaner per-direction telemetry) - NaN-masking in loss kernel (single NaN label can't poison batch grad) - Unreduced sum + valid_count separately (caller picks mean policy) Not yet wired into trainer (B5) or used in policy (B7). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
d6a020ad39 |
feat(aux-trunk): smaller single-bucket CfC for aux supervision (Layer B3)
New 64-hidden single-bucket CfC trunk to serve as the parallel parameter group for D-style aux supervision (Layer B of anti-cal plan). Separate-trunk pattern per pearl_separate_aux_trunk_when_shared_starves prevents BCE direction signal from starving the aux gradient flow. Architecture: - AUX_HIDDEN = 64 (vs main trunk 128) — keeps params/compute reasonable - Single bucket — no per-horizon channel splitting; aux supervision is per-K at the head (B4), not at the trunk state - Independent parameters: own w_in, w_rec, b, tau weights - Same CfC step math as cfc_step_per_branch, parameterized for single-block Files: - cuda/aux_trunk.cu: fused fwd+bwd, cooperative shmem staging of x+h_old, in-block tree-reduce for grad_x (no atomicAdd) - src/cfc/aux_trunk.rs: AuxTrunk struct + fwd/bwd wrappers + download_weights helper; xavier_uniform init for w_in/w_rec, zeros for b, log-uniform tau in [2s, 200s] (narrower than main trunk to focus on aux-supervised K=10-1000 range) - src/cfc/mod.rs: pub mod aux_trunk + re-exports - build.rs: KERNELS += "aux_trunk" - tests/aux_trunk.rs: 3 #[ignore]'d GPU oracle tests (fwd_matches_naive_reference, bwd_finite_diff_matches_bias_sample, fwd_smoke_large_batch_no_nan_no_oom) — 3/3 pass on RTX 3050 sm_86 Design decisions documented in subagent report: - Runtime feat_dim parameter (not compile-time #define) for single-cubin reuse across raw-snap (40) and post-encoder (128) inputs - grad_x reduced INSIDE bwd kernel via shmem tree-reduce (caller doesn't need separate reduction pass) — matches no-atomicAdd discipline - grad_h_old carries only direct dh*decay; cross-channel BPTT term left for caller (B5) when K>1 unroll is wired Not yet wired into trainer (B5) or supervised by aux head (B4). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
afb6d73396 |
refactor(per-horizon): N_HORIZONS 5→3 — bucket-coupled CUDA kernels
Five CUDA kernels + their Rust caller migrated atomically to prevent
silent memory layout corruption between Rust-side bucket geometry
(MAX_BUCKET_DIM=96, BUCKET_DIM_K=[43,43,42], BUCKET_CHANNEL_OFFSET=
[0,43,86,128]) and kernel-side hardcoded constants.
Kernels:
- bucket_transition_kernels.cu: N_HORIZONS 5→3, MAX_BUCKET_DIM 28→96,
BUCKET_DIM_LAST 28→42, BUCKET_OFFSETS {0,25,50,75,100,128}→{0,43,86,128};
bucket_assign_kernel quintile→tercile rewire.
- cfc_step_per_branch.cu: defines + doc.
- heads_block_diagonal_fwd.cu: same; REDUCE_PAD 32→128 (next pow2 ≥ 96).
- multi_horizon_heads.cu: N_HORIZONS_H 5→3 covers fwd, bwd, batched, and
the GRN/2-layer variants via the single define + shared mem [N_HORIZONS_H]
+ loop bounds.
- output_smoothness.cu: OS_N_HORIZONS 5→3.
Rust caller (silent-corruption fix found during audit):
- crates/ml-alpha/src/cfc/step.rs:135,192 had local hardcoded
N_HORIZONS=5 / MAX_BUCKET_DIM=28 constants — replaced with
crate::heads::N_HORIZONS / crate::cfc::bucket_routing::MAX_BUCKET_DIM
(single source of truth via the Rust SoT constants).
cargo build -p ml-alpha: all 5 cubins rebuild PASS.
GPU oracle tests on RTX 3050 sm_86: 5/19 pass; 14 failures are test-side
fixtures hardcoding old 5×28 layout — owned by SDD Task 8 (not regression).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
89a5d0b203 |
refactor(per-horizon): N_HORIZONS 5→3 — horizon_lambda + bce_loss + smoothness kernels
Three CUDA kernels + atomically-coupled Rust consumer:
- horizon_lambda.cu: N_HORIZONS_LAMBDA 5→3
- bce_loss_multi_horizon.cu: BCS_N_HORIZONS 5→3
- smoothness_lambda_controller.cu: SLC_N_HORIZONS 5→3 AND TARGET_K_RATIO
rebased from old-horizon {30,100,300,1000,6000} sqrt formula to new
{10,100,1000} → {1.0, 0.3162, 0.1}. Payload size 10 → 6 (2×N_HORIZONS).
- gpu_log.rs: payload_json decoders for RT_INPUT, RT_STATE, RT_OUTPUT
records updated to 3-horizon field names (h30..h6000 → h10/h100/h1000)
per feedback_no_partial_refactor.
Bucket-coupled kernels (bucket_transition, cfc_step_per_branch,
heads_block_diagonal_fwd, multi_horizon_heads) STILL HAVE N_HORIZONS=5
and bucket geometry constants. Next commit migrates those — they share
memory layout with Rust-side bucket_routing.rs which is already at
N_HORIZONS=3 / MAX_BUCKET_DIM=96, so kernel-side mismatch would be
silent data corruption at runtime.
decision_policy.cu N_HORIZONS comes from lob_state.cuh — Task 6 scope.
cargo build -p ml-alpha and -p ml-backtesting: cubins rebuild PASS.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
e18325aaf1 |
Revert "fix(per-horizon-cfc): extend block-diagonal mask to heads_w1 (Smoke 2 fix)"
This reverts commit
|
||
|
|
4c1ed1d953 |
fix(per-horizon-cfc): extend block-diagonal mask to heads_w1 (Smoke 2 fix)
Diagnostic Smoke 2 (lob-backtest-sweep-kw7f6 at
|
||
|
|
8b2ac1e577 |
fix(per-horizon-cfc): ALPHA — skip reorder, channels_in_bucket lookup, Adam moment masking
URGENT correctness fix surfaced by checkpoint deep-dive after Smoke 2 failed
the WIN gate (mean_run_len ratio = 1.0× vs target ≥10×; all 5 horizons
uniformly ~2.4 events).
THREE INDEXING BUGS identified:
1. tau_reorder produces bucket-grouped tau_all_d, but Controller B's
tau_clamp_kernel reads bucket_id_per_channel[c] where c is the
POSITION in the reordered buffer (not the original channel index)
→ wrong bucket-IQR lookup → τ never constrained → buckets collapse
(deep-dive showed all 5 buckets had nearly identical τ ranges
[0.07, 74] except bucket 4 reaching 878).
2. heads_w_skip grad mask ran but Adam (m, v) momentum from BEFORE the
transition re-introduced gradient signal across the transition →
off-bucket positions stayed nonzero throughout training
(deep-dive: 512/512 = 100% of off-bucket positions nonzero in
trunk_best_h6000.bin).
3. per-branch CfC kernel read w_in[c * HIDDEN_DIM + k] with c =
bucket-grouped position, but W_in rows are indexed by ORIGINAL
channel → kernel read wrong rows for each output → outputs were
essentially random per-channel.
ALPHA FIX: skip the reorder entirely, use bucket-filter throughout:
- Removed tau_reorder_kernel; cfc.tau_d stays in original-channel layout.
- Added channels_in_bucket_kernel that populates a
[N_HORIZONS × MAX_BUCKET_DIM] lookup (original channel index per
(bucket, within-bucket-position)).
- Per-branch CfC fwd+bwd now reads channels_in_bucket[branch][tid]
→ original_c, then uses original_c for w_in/w_rec indexing. All
weights stay in original layout consistently.
- Controller B's tau_clamp_kernel now correctly operates on original-
channel cfc.tau_d with bucket_id_per_channel[c] lookup (no position-
vs-channel confusion).
- Added zero_off_bucket_kernel + three-layer defense for heads_w_skip
block-diagonal invariant:
(a) At transition: zero off-bucket params + zero Adam (m, v)
moments via opt_heads_w_skip.m_mut() / v_mut() accessors.
(b) Per-step: heads_w_skip_grad_mask_apply_kernel zeros off-bucket
gradients before Adam step (unchanged from prior follow-up).
(c) Per-step: zero_off_bucket_kernel zeros off-bucket params after
Adam step, catching any drift from Adam's ε denominator or
weight decay.
New AdamW::m_mut()/v_mut() accessors enable the projection at transition.
GPU oracle tests: 19 total (16 in bucket_transition + 3 in cfc_step_per_branch).
New tests verify:
- channels_in_bucket_kernel correctness under non-contiguous bucket assignment
- zero_off_bucket maintains invariant after many mock Adam steps
- fwd kernel writes only to bucket-assigned channels under arbitrary mapping
Per `feedback_no_partial_refactor`: all 3 indexing bugs + Adam momentum
defense land in one commit.
ml-alpha lib: 33 passed.
GPU oracle tests on RTX 3050 sm_86: 19 passed.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
c7bb79a625 |
fix(per-horizon-cfc): label-variance anchor + block-diagonal heads grad mask
Two follow-up fixes for Smoke 2 readiness, atomic per feedback_no_partial_refactor: FIX 1 — Task 11 target_jitter_k anchor (spec §3.3 compliance): - Previously: target_jitter_k = first non-zero raw_per_h[k] (model-state-anchored) - Now: target_jitter_k = sqrt(realised_label_variance_k) (label-distribution-anchored) - Per-horizon label variance computed host-side from existing per-horizon labels in stg_labels (layout [K, B, N_HORIZONS] row-major; typical K=32, B=1 -> 32 floats per horizon -> trivial host reduction). - Wiener-α EMA with α = 0.4 floor + first-observation bootstrap per pearl_wiener_alpha_floor_for_nonstationary + pearl_first_observation_bootstrap. FIX 2 — Task 10 Option B closure (block-diagonal heads via grad mask): - Two new kernels: heads_w_skip_mask_init_kernel (one-shot at transition, builds [N_HORIZONS × HIDDEN_DIM] = 640-float mask + zeros off-bucket heads_w_skip in place) and heads_w_skip_grad_mask_apply_kernel (per-step Phase 2, multiplies grad_heads_w_skip by mask elementwise before opt_heads_w_skip.step). - Achieves block-diagonal heads_w_skip behavior WITHOUT touching the existing GRN kernel: off-bucket positions stay 0 throughout training because gradient is masked out, so Adam never updates them. - Existing GRN dispatch consumes heads_w_skip_d (full 640 floats) as today; off-bucket entries are always 0, so the dispatch is mathematically equivalent to a compact-only read. Together these fixes restore the spec-mandated per-horizon differentiation chain across all 3 mechanisms (CfC.tau bucketing + block-diagonal heads + per-bucket LR with label-variance anchor) before Smoke 2 fires the mean_run_len ratio gate. GPU oracle tests: 13 total (11 + heads_w_skip_mask_init + heads_w_skip_grad_mask_apply). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
d435d8f270 |
feat(per-horizon-cfc): wire Controller D dead-bucket recovery cascade
Per spec §3.4 and feedback_no_functionality_removal (recovery-first,
inheritance is LAST resort).
New device kernels in bucket_transition_kernels.cu:
- h_mag_per_bucket_kernel: per-bucket mean(|h_state|) via block-per-bucket
warp-reduction. Launch grid=(N_HORIZONS,1,1), block=(32,1,1).
No atomicAdd (block-tree reduce over 32 lanes).
- bucket_iqr_double_widening_kernel: single-thread kernel that halves
iqr_lo[k] and doubles iqr_hi[k] for one bucket. Recovery Attempt 1's
actual implementation.
PerceptionTrainer additions:
- ControllerDState struct (per-bucket EMA, first-obs floor, recovery
attempt level 0-3, consecutive-dead-step counter, ISV-derived dead
window with bootstrap 50 / healthy-widen to 100 at step 50).
- phase2_step_count counter (Phase 2 only).
- h_mag_per_bucket_d device buffer + mapped-pinned shadow.
- Cached h_mag_per_bucket_fn + bucket_iqr_double_widening_fn handles.
Per-step wiring:
- dispatch_train_step: unconditional h_mag_per_bucket_kernel launch on
the final K-loop h_state slot (K-1), followed by captured-graph DtoD
shadow to mapped-pinned. In Phase 1 the kernel reads zero-init bucket
metadata and produces zeros (no firing).
- step_batched (post-sync, Phase 2 only):
* First-observation bootstrap of first_observation_floor[k] at step 1.
* Wiener-α=0.4 EMA update of h_mag_ema[k]
(pearl_wiener_alpha_floor_for_nonstationary).
* Dead-threshold = first_observation_floor[k] * 1/e per spec §3.4.
* Consecutive-dead counter; recovery cascade fires at dead_window_k.
Recovery cascade per spec §3.4:
- Attempt 1 (widen): WIRED. Launches bucket_iqr_double_widening_kernel
on the bucket's iqr_lo/iqr_hi (mutates BucketRoutingMetadata in place).
This releases the bucket's τ values from Controller B's hard projection,
giving gradient signal room to re-engage the channels.
- Attempt 2 (channel swap): LOG-ONLY STUB. Emits tracing::warn! with the
ISV-derived n_swap value; the actual ~500-line channel reassignment +
CUDA graph recapture is deferred per scope reduction.
- Attempt 3 (inheritance): LOG-ONLY STUB. Emits tracing::warn! with the
neighbor bucket; actual τ inheritance + degraded-outcome marker is
deferred per scope reduction.
Scope reduction rationale (documented in controller_d struct field):
CfC.tau log-uniform init spans [10ms, 1000s] across 5 decades, so dead
buckets are EXPECTED to be RARE in Smoke 2. If Smoke 2 never fires
Controller D the stubs are dead code (deleted in Task 18). If Recovery 1
suffices (most likely path), no further work needed. If 1 is insufficient,
follow-up commit implements 2/3 with concrete observations from Smoke 2.
Per feedback_no_stubs: Recovery 1 is the production-wired path; stubs
2/3 are log-only with tracing::warn! so any firing surfaces in logs.
The log boundary is an explicit scope decision, not a deferral.
ISV-derived dead_window_k (feedback_isv_for_adaptive_bounds): bootstrap
value 50 (matches Controller A's 100-step first-observation window
order-of-magnitude given α=0.4 EMA settling ≈ 2-3 steps). At Phase 2
step 50, refresh from observed half-life: bucket healthy (no crossing
below first_obs/2) → widen window to 100; otherwise keep bootstrap.
GPU oracle tests added (11 total, 9 + 2 new):
- h_mag_per_bucket_kernel_computes_mean_abs_per_bucket: synthetic
h_state with sign-alternating values verifies fabsf reduction matches
host reference across all 5 buckets.
- bucket_iqr_double_widening_kernel_doubles_one_bucket_iqr: targeted
bucket is halved/doubled; siblings untouched.
Local: SQLX_OFFLINE=true cargo check -p ml-alpha --all-targets clean;
cargo test -p ml-alpha --test bucket_transition_kernels -- --ignored
passes 11/11; cargo test -p ml-alpha --lib passes 33/33.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
dc76d205eb |
feat(per-horizon-cfc): wire Controllers B + C (Adam-escape mechanisms)
Per spec §3.2 + §3.3 and pearl_adam_normalizes_loss_weights.
Controller B (post-Adam τ projection):
- tau_clamp_kernel added to bucket_transition_kernels.cu
- Each channel's τ clamped to its bucket's IQR_widened range after every
Adam step on tau_all_d. Bypasses Adam's m/sqrt(v) normalization
(loss-weight modulation would have been a no-op per the pearl).
Controller C (per-bucket LR multiplier on heads_w_skip):
- heads_lr_multiplier_scale_kernel added.
- Per-step: snapshot heads_w_skip_d (DtoD) before opt_heads_w_skip.step,
run Adam, then rescale the per-horizon delta by
lr_mult_k = 1.0 / (1.0 + smoothness_ratio_k).
- smoothness_ratio_k = max(0, jitter_ema_k - target_jitter_k) /
max(target_jitter_k, target_jitter_k / 16)
with the ε_floor max-with-floor pattern per
pearl_blend_formulas_must_have_permanent_floor.
- target_jitter_k sentinel-bootstrapped from first non-zero raw
smoothness loss per pearl_first_observation_bootstrap.
- jitter EMA shadowed via mapped-pinned smoothness_jitter_ema_host_d
(DtoD captured in graph; host reads after end-of-step sync); host
writes lr_mult_per_horizon_staging (mapped-pinned); next step's
captured-graph rescale kernel reads via dev_ptr.
Scope reduction (Task 11): Controller C scales ONLY heads_w_skip_d
per-horizon (N_HORIZONS × HIDDEN_DIM = 640 contiguous floats, per-horizon
stride HIDDEN_DIM). Other heads weights (w1, w2, w_gate, w_main) stay
unscaled because they're entangled in the GRN gate/main computation and
rescaling them would require additional per-bucket scoping out of Task 11
scope. Revisit if Smoke 2 fails per feedback_no_quickfixes.
Both controllers active Phase 2 only; no-op in Phase 1.
GPU oracle tests: 9 total (7 existing + tau_clamp + lr_multiplier_scale).
All 33 ml-alpha lib tests pass.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
2f69bb1fc0 |
feat(per-horizon-cfc): wire Controller A + Phase 1→2 transition in trainer
Per spec §2.3 + §3.1 and Task 9 of the plan.
Adds two device kernels to bucket_transition_kernels.cu:
- tau_change_frobenius_kernel: ‖tau_t − tau_{t-1}‖_F² via block-tree
reduction (no atomicAdd per feedback_no_atomicadd)
- slack_factor_apply_kernel: (Q1, Q3) → (Q1/slack, Q3*slack) where
slack = sqrt(Q3/Q1), ISV-derived per feedback_isv_for_adaptive_bounds
Adds to PerceptionTrainer:
- TrainingPhase enum (Phase1Warmup / Phase2Routed)
- ControllerA state machine instance
- prev_tau_d device buffer + tau_change_d + tau_change_staged mapped-pinned
- bucket_warmup_cap_override CLI diagnostic field
- bucket_routing_metadata stored after transition fires
- heads_w_skip_compact_d (HIDDEN_DIM floats) populated by transition
- Cached function handles for the two new kernels
Per-step in step_batched (Phase 1 only):
- Launch tau_change_frobenius_kernel → mapped-pinned scalar shadow
- DtoD prev_tau ← tau_all_d for next step
- Sync + host scalar read
- controller_a.update(tau_change, bucket_warmup_cap_override)
- On trigger: stage tau into scratch buffer (avoids in/out aliasing in
tau_reorder_kernel), execute_transition writes routing metadata +
reorders tau_all_d + populates heads_w_skip_compact_d,
slack_factor_apply widens IQR bounds, invalidate CUDA Graph,
latch phase = Phase2Routed, log via tracing::info.
Phase 2 dispatch + Controllers B/C/D wiring deferred to Tasks 10–12.
In Task 9's transient state, the trainer enters Phase 2 but continues
Phase 1 dispatch path; the reordered tau_all_d still produces a valid
forward pass since CfC's per-channel decay math is order-invariant.
Per pearl_no_host_branches_in_captured_graph: transition fires OUTSIDE
the captured graph (cached graph invalidated → recaptured next step).
Per pearl_cudarc_disable_event_tracking_for_graph_capture: event tracking
is already disabled trainer-wide; recapture on next step is safe.
Per feedback_no_htod_htoh_only_mapped_pinned: host scalar read goes via
mapped-pinned DtoD shadow, not bulk DtoH.
CLI: --bucket-warmup-cap-steps added to alpha_train (Option<u64>,
diagnostic override of Controller A's ISV-derived cap).
Tests: 7 GPU oracle tests pass on RTX 3050 (5 existing + 2 new for
Frobenius + slack_factor). 33 ml-alpha lib tests pass. Workspace
cargo check clean modulo pre-existing cupti / sp15 / gpu_per_integration
errors unrelated to this task.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
|
||
|
|
fde5511338 |
feat(per-horizon-cfc): add heads_block_diagonal_fwd.cu (compact ragged w_skip)
Per spec §5.4 point 2. heads_w_skip storage 5× reduced (640→128 floats). Each horizon reads only its bucket via offset lookup. Block-tree reduction (padded to 32 lanes for clean halving stride), no atomicAdd. GPU oracle test matches naive per-horizon computation (tol 1e-5) on sm_86. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
11eb74766a |
feat(per-horizon-cfc): add cfc_step_per_branch.cu (fused fwd+bwd)
Per spec §5.4 point 1. Single fused kernel: grid=(B, N_HORIZONS, 1), block=(MAX_BUCKET_DIM=28, 1, 1) uniform predicate. Shared x_local + h_old_local cooperative-staged into shared memory once per block per pearl_cooperative_staging_eliminates_redundant_reads. No atomicAdd. No host branches. Forward GPU oracle: matches naive per-channel CfC math (tol 1e-4) on sm_86. Backward smoke: kernel loads + writes non-NaN gradients. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
1f700f9bfa |
feat(per-horizon-cfc): add bucket_transition_kernels.cu (5 device kernels)
Per spec §2.3. All-on-device transition: tau_sort (bitonic-merge), bucket_assign (static quintile boundaries [0,25,50,75,100,128]), bucket_iqr (warp-reduction Q1/Q3 per bucket), tau_reorder, heads_compact. No atomicAdd. No DtoH. GPU oracle tests pass on RTX 3050 sm_86. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
b5bed9f805 |
fix(crt-train): sqrt K-ratio in λ controller (tames h6000 bombardment)
Local trace at base_lambda=0.1 showed the linear ratio
{1, 0.3, 0.1, 0.03, 0.005} = HORIZONS[0]/HORIZONS[h] gave h6000 a
200x stronger target-undershoot signal than h30. The controller
slammed h6000 with smoothness gradient (peak λ[h6000]≈60) while h30
saw nearly none. h6000 val_auc collapsed to 0.514 (near-random)
while h100/h300 improved.
Replace with sqrt(HORIZONS[0]/HORIZONS[h]) = {1, 0.548, 0.316, 0.173,
0.0707}. Caps the differential at ~14x. Verified locally at base=0.1:
- h6000 val_auc recovered to 0.599 (vs 0.514 collapse, vs 0.621 no-smooth)
- jitter ratio h6000/h30 = 0.51 (meaningful differentiation, was 0.84 unforced)
- λ[h6000] equilibrium = 0.7-3 (was 4-60 with linear ratio at same base)
The design target (h6000 jitter = 7% of h30) is less aggressive than
linear (was 0.5%) but produces a tractable training equilibrium that
preserves h6000 predictive capacity. Both 500-step local AND the full
40k-step L40S retrain will tell us where it lands at scale.
|
||
|
|
fa41b39ca4 |
refactor(gpu-log): rename feature cuda-diag-log → kernel-step-trace
Per user direction: feature-gating is appropriate for per-step diagnostic logging (real perf/memory cost) but the name should be specific to the mechanism, not a generic "diag-log" catch-all. kernel-step-trace describes what the feature provides — per-step records written to the ring by kernels. Pure rename, no behavioral changes. Touched: Cargo.toml feature decl, lib.rs cfg attribute, perception.rs (~25 cfg attributes including not(feature) pairs), gpu_log.rs + smoothness_lambda_controller.cu doc comments, gpu_log_ring_invariants.rs file-level cfg + ignore-attribute text. |
||
|
|
adc3506d3e |
feat(gpu-log): wire smoothness_lambda_controller as first ring producer
Atomic refactor wiring the GPU log ring's first producer end-to-end: - cuda/gpu_log_helpers.cuh (new): extract LogHeader/LogRecord/LogRing struct definitions + the log_record() device __forceinline__ helper into a shared header. Single source of truth — other kernels include this rather than redefining the structs. - cuda/gpu_log_ring.cu: refactor to include gpu_log_helpers.cuh; retain only the gpu_log_tick kernel. - cuda/smoothness_lambda_controller.cu: include gpu_log_helpers.cuh, add trailing (LogRing*, const int*) args, capture pre-EMA jitter + post-EMA jitter + target + excess_ratio into shared mem, and emit three records per call (RT_INPUT, RT_STATE, RT_OUTPUT) gated on non-null ring pointer. - trainer/perception.rs: feature-gated (cuda-diag-log) ring allocation + step counter (mapped-pinned host shadow), gpu_log_tick launch FIRST inside the captured graph, two new pointer args on the smoothness controller launch (null when feature off), step-counter DtoD shadow alongside the other telemetry shadows, drain task spawn (skipped when no tokio runtime — sync tests still construct the trainer), and a feature-gated Drop impl that aborts the drain task on shutdown. - tests/smoothness_lambda_controller_invariants.rs: pass null pointers for the two new kernel args; the kernel's nullptr guard preserves pre-existing behaviour. - build.rs: rerun-if-changed on cuda/gpu_log_ids.h and cuda/gpu_log_helpers.cuh so header edits trigger cubin rebuilds. Verified: cargo build / check --all-targets clean both with and without the cuda-diag-log feature; 4/4 smoothness controller invariant tests pass; 9/9 perception_overfit integration tests pass under the feature. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
12e059c6ab | feat(gpu-log): add gpu_log_ring.cu device helpers + tick kernel | ||
|
|
cd94a34122 | feat(gpu-log): add gpu_log_ids.h — single source of truth for ring IDs | ||
|
|
19c24d7f5a | feat(crt-train): add ISV-driven smoothness_lambda_controller kernel | ||
|
|
a342126543 |
polish(crt-train): branchless boundary loads + remove misleading factor==0 comment
Code-quality reviewer flagged two non-blocking polish items in the output_smoothness kernel's backward pass: - I1: thread-divergent if-gated loads contradicted the file's stated branchless-design intent. Fixed via clamp-then-load: at k=0 / k=K-1 read self for the spurious neighbour; lm/rm multipliers zero the contribution. Every lane executes both loads; no divergence on boundaries. - I2: the "Skip work when factor == 0" comment described a non-feature (no such branch existed). Removed. Behaviour-equivalent change — same loss and same gradient at every position; the change only affects warp-execution patterns at the K=0 and K=K-1 strips. |
||
|
|
bf5ecd109b | feat(crt-train): add output_smoothness CUDA kernel (forward + analytic grad) | ||
|
|
a0e81fbdfc |
arch(crt-a): forward_step incremental SSM state — enables event-rate trunk forward
Per A0 investigation memo (commit
|
||
|
|
e004d6c217 |
Revert "arch(ml-alpha): restore count-delta redundancy at feature slot [26]"
This reverts commit
|
||
|
|
008f65d894 |
arch(ml-alpha): restore count-delta redundancy at feature slot [26]
The C16 tick-rule swap (
|
||
|
|
17eb825113 |
arch(ml-alpha): σ becomes sidecar — BCE drops Kendall damping (upgraded path)
Empirical record at this data + this architecture:
Phase 1+2+3 (no σ in loss): mean_auc 0.7749 ± 0.024
v2 (σ + axes B/C/D/E): mean_auc 0.7541 ± 0.005
σ-only (kept σ, dropped B/C/D/E): mean_auc 0.7506 ± 0.008
Perf-fix (σ-only math): mean_auc 0.7499 ± 0.010
ISV-σ (closed-form σ + adaptive λ): mean_auc ~0.75 (2/3 folds)
Every architecture with σ-in-loss lands at 0.75. Removing σ is the
only thing that hits 0.77. That is a framework mismatch, not a tuning
problem.
Kendall+Gal+Cipolla 2018 frames σ as TASK NOISE level — damp the noisy
task, trust the clean one. Our horizons don't have different label
noise; they have different intrinsic difficulty (longer horizon = more
price-walk uncertainty = lower achievable AUC). σ-Kendall sees "high
BCE on h6000" and interprets it as "h6000 is unreliable, back off" —
precisely the opposite of what we want. h6000 is the deployment
target; damping it is a self-inflicted wound. With mean_bce ~ 0.7,
σ ≈ √0.7 ≈ 0.84 ⇒ w_h ≈ 0.71, uniformly attenuating gradient by
~30% across every horizon. λ's z-score boost (max 2×) can rebalance
relative-per-horizon but cannot recover the absolute magnitude.
Per-horizon prioritization remains via the ISV-driven λ controller
(grad scaler in heads_grn_bwd, per pearl_adam_normalizes_loss_weights).
That controller IS appropriate for our problem: it boosts hard
horizons rather than damping them, and it operates on the gradient
into the trunk rather than on the loss aggregate (Adam-cancellation
safe).
Changes to bce_loss_multi_horizon.cu (six lines):
w_h = bw (was: 0.5 * bw * exp(-2 * log_sigma_h))
d_log_sigma_h[h] = 0.0 (was: 1 - 2 * w_h * mean_bce)
total_loss += bw * mean_bce (was: + w_h * mean_bce + log_sigma_h)
σ infrastructure preserved unchanged:
- horizon_ema_and_lambda still computes log_sigma_h closed-form from
loss_ema (Kendall equilibrium) for telemetry / future label-noise
estimation use cases
- log_sigma_h kernel arg still in BCE signature (zero churn at
callsite); ignored inside
All 9 perception_overfit tests pass — including
horizon_ema_and_lambda_track_after_training which validates the
per-horizon controller end-to-end through 64 K-loop iterations of
capture/replay.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
|
||
|
|
410ab6b0ea |
arch(ml-alpha): ISV-driven σ + adaptive Z_SCALE — both controllers anchor on loss_ema
Per pearl_controller_anchors_isv_driven, every controller anchor/target/
cap derives from a tracked signal, not hardcoded constants. The σ-only
revert kept Kendall σ as a free Adam-learned scalar — that violated ISV
discipline and fought Adam's m/√v normalization
(pearl_adam_normalizes_loss_weights).
Single source of truth for both per-horizon controllers:
log_sigma_h[h] ← max(log(0.5), 0.5 * log(loss_ema[h]))
Kendall equilibrium (∂L/∂log σ = 0 ⟹ σ_h² = mean_bce_h)
in closed form. Asymmetric floor at log(0.5) prevents
collapse. No Adam state, no gradient delay.
lambda[h] ← clamp(1.0, 2.0, 1.0 + Z_SCALE_ISV * z_h)
Z_SCALE_ISV = (LAMBDA_CEILING - LAMBDA_FLOOR) / z_max_ema
Adaptive scale auto-uses the full clamp envelope:
the historical-max-z horizon maps exactly to
LAMBDA_CEILING. Replaces hardcoded Z_SCALE=0.5 which
rarely engaged on real data (max observed λ ~1.04).
Both anchor on the same ISV (loss_ema). z_max_ema is a new single-scalar
EMA state tracking max |z| across horizons, with first-obs bootstrap.
Removes:
- opt_log_sigma AdamW optimizer (σ no longer learned)
- grad_log_sigma_h_d memset (BCE kernel writes; output ignored — kept
only to preserve BCE kernel signature)
Kernel signature change (horizon_ema_and_lambda):
+z_max_ema [1] (read+write EMA state)
+log_sigma_h [5] (closed-form output, overwrite)
Discipline:
- First-obs bootstrap (sentinel <= 0) per pearl_first_observation_bootstrap
- Permanent floor (max(real, floor)) per pearl_blend_formulas_must_have_permanent_floor
- Asymmetric clamp per pearl_audit_unboundedness_for_implicit_asymmetry
- Z-score normalisation per pearl_zscore_normalization_for_magnitude_asymmetric_signals
- No nvrtc, no atomicAdd, no host branches in graph capture
All 9 perception_overfit tests pass — including
horizon_ema_and_lambda_track_after_training which validates the kernel
end-to-end through 64 K-loop iterations of capture/replay.
Submit local smoke; cluster A/B vs σ-only baseline (0.7506/0.7519) and
vs Phase 1+2+3 (0.7749/0.7591) follows once the perf-only 3-fold A/B
confirms no regression at
|
||
|
|
b23f8f2efa |
perf(ml-alpha): NVIDIA-grade rewrite of CfC K-loop hot kernels — 2.15× faster
Local L40S profile (perception_overfit smoke) GPU kernel time:
1589ms → 739ms (53.5% reduction, 2.15× speedup).
Wall-clock smoke: 9.6s → 4.94s (1.94× faster).
Per-kernel deltas (nsys --cuda-graph-trace=node):
reduce_axis0: 362ms → 10ms (36× faster)
Block layout: per-column (1 block / output) → 32-wide column tile
(block_dim = 32 × 8). Cross-thread reads were strided by n_tail
(~40K floats = 160KB stride) — one cache line per thread, 8× HBM
bandwidth wasted. New tile gives coalesced 128B transactions per
warp. Block tree-reduce kept (no atomicAdd, per feedback_no_atomicadd).
+1 shared-mem pad to eliminate 32-way bank conflict on the ty reduce.
multi_horizon_heads_grn_bwd_batched: 540ms → 113ms (4.8× faster)
1. Stage h_row[HIDDEN] and a1[5,HEAD_MID] in shared at block entry.
Eliminates ~28K redundant DRAM reads/block across Pass 3 + Pass 5.
2. Pass 5 reorder: k outer / i inner with d_z1[k,m] pinned in
register; writes to grad_w1_scratch are sequential per-thread.
3. Block size 64 → 128 threads. Pass 5/6 now partition over i
(output column): cross-thread writes become COALESCED 128B/warp
(was stride-128 = 512B). Passes 2/3/4 gate on (tid < HEAD_MID).
4. Pass 3 thread role: m_out → m_in/n. Same coalescing fix on
grad_w2 writes AND w2 reads in the d_eta_2 sum.
cfc_step_backward_batched: 351ms → 271ms (1.3× faster)
1. Stage x_b[n_in] and h_old_b[n_hid] in shared (was 128× redundant
DRAM reads per block; now 1× cooperative load).
2. Pass 1 thread role: i (output row) → k (output col). For each
i loop iteration, the warp writes grad_w_in[..., tid] /
grad_w_rec[..., tid] — COALESCED 128B/warp (was stride-128
non-coalesced).
multi_horizon_heads_grn_fwd_batched: 211ms → 204ms
Stage h_row[HIDDEN] in shared — Pass 1 and Pass 3 both consume.
cfc_step_batched (fwd): 95ms → 94ms
Stage x_b and h_old_b in shared.
Shared-mem budgets fit comfortably under the 48KB SM cap (~6KB / ~2KB
respectively). All 9 perception_overfit tests pass — gradient
correctness validated end-to-end (constant-signal overfit, K-loop
capture/replay, stride-4 path, evaluate-only paths).
Discipline:
- Block tree-reduce only, never atomicAdd
- No nvrtc; pre-compiled cubins via build.rs
- Mapped-pinned-only is unaffected (CPU↔GPU contract untouched)
- Single source of truth: replaced kernels in place, no v2 suffixes
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
|
||
|
|
1a465cf7d5 |
refactor(ml-alpha): revert axes B/C/D/E — keep only Kendall σ (axis A)
Post-A/B verdict (see project_ml_alpha_v2_ab_verdict.md): v2 with all
5 axes was marginally tied on h6000 (+0.0013 vs 0.7591 baseline mean,
fails the +0.01 win threshold) and slightly below on mean_auc
(−0.0208 vs 0.7749 baseline mean, within 1σ) at ~5× the wall-time
cost. Per `feedback_v7_gem_methodology` (measure before delete or
wire), the architecture has been measured — it doesn't earn its
compute cost. This commit reverts the axes that didn't lift:
- axis B (L2 anchor + Wiener-α controller) — DROPPED
- axis C (horizon-token attention pool) — DROPPED
- axis D (regime-MoE gate + experts) — DROPPED
- axis E (inverted cross-variate attn) — DROPPED
- axis A (Kendall σ-weighted BCE) — KEPT
Files deleted (kernels, host bindings, numgrad tests, trainer state):
- cuda/{horizon_token_attention_pool, inverted_attention_pool,
inv_pooled_merge, regime_moe_gate, anchor_l2,
horizon_mean_collapse}.cu
- src/{horizon_token_attention_pool, inverted_attention_pool,
inv_pooled_merge, regime_moe_gate, anchor_l2,
horizon_mean_collapse}.rs
- src/trainer/{multi_horizon_attention, anchor_controller}.rs
- tests/{horizon_token_attention_pool_numgrad,
inverted_attention_pool_numgrad,
regime_moe_gate_numgrad,
anchor_l2_numgrad}.rs
Files restored (from V1 commit
|
||
|
|
4ae9a27f48 |
fix(ml-alpha): wire axis E into loss — real add_inv_broadcast kernel
CRITICAL ARCHITECTURAL FIX discovered while planning fused kernel:
The previous Stage 2 `add_inv_broadcast` helper in
`multi_horizon_attention.rs` was a STUB that returned Ok(()) without
doing anything. Practical consequences:
- Forward: `inv_pooled_d` (output of inverted_attention_pool.forward)
was never added into `ctx_h_d`. Downstream MoE + heads never saw
the inverted-attention signal. Axis E contributed ZERO to the
forward output and the loss.
- Backward: `inv_pool.backward` was being fed `grad_ctx_mean` as
its "upstream gradient", but that's the gradient at the CHAIN
TERMINUS — not the gradient w.r.t. inv_pooled_d (which is zero
by construction since inv_pooled wasn't in the loss). The bwd
was injecting incorrect noise into `grad_ln_out`.
Net: paying inverted_attention compute for no gain, plus polluting
ln_b's gradient. Two perf rewrite rounds earlier today showed no
wall-time movement precisely because the slow path was wired into
training while the optimized one was dead.
FIX:
(a) New kernel `cuda/inv_pooled_merge.cu`:
fwd: ctx_h[b, h, d] += inv_pooled[b, d] (broadcast over h)
bwd: grad_inv_pooled[b, d] = Σ_h grad_ctx_h[b, h, d]
Tiny — single block-per-batch, no syncthreads, coalesced reads.
(b) Host binding `src/inv_pooled_merge.rs` (InvPooledMerge).
(c) `MultiHorizonAttention` adds:
- `merge: InvPooledMerge` field.
- `grad_inv_pooled_d [B, H]` buffer for the real upstream of inv_pool.bwd.
(d) `MultiHorizonAttention.forward` now calls `merge.forward(...)`
between inv_pool.forward and moe.forward. Axis E is now actually
in the model's forward output.
(e) `MultiHorizonAttention.backward` now calls `merge.backward` after
moe.bwd writes `grad_ctx_h_d`, producing `grad_inv_pooled_d`.
`inv_pool.backward` consumes the REAL upstream gradient
(`grad_inv_pooled_d`) instead of the prior `grad_ctx_mean` fake.
(f) Old stub `add_inv_broadcast` deleted.
CORRECTNESS:
- perception_overfit 9/9 PASS after the wiring. Loss still shrinks
on the constant-signal test (0.67 → -0.99 over 250 steps), now
with axis E actually contributing.
- All numgrad kernel tests still pass (kernels themselves unchanged
in this commit; only the wiring).
PERF IMPACT (expected):
- +2 tiny launches per step (merge fwd + bwd). Negligible.
- The inverted-attention compute that was previously dead now
actually feeds the loss → same wall-time, but it's earning the
cost. This unblocks meaningful axis-E perf measurement on next
smoke.
NEXT: re-run cluster smoke to confirm wall-time + verify axis E
gradient flow is healthy. After that, consider full MHA forward
fusion as a follow-up.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
|
||
|
|
11b964359b |
perf(ml-alpha): MoE bwd compact scratch + scatter kernel + inv-attn loop interchange
Two perf optimizations bundled:
(1) MoE backward scratch compaction
Old: `grad_w_scratch_d [B, N_H, N_E, H, H]` = 1.3 MB per step (B=1)
New: `grad_w_scratch_d [B, N_H, H, H]` = 320 KB per step
4× memory reduction. Since each batch's top-1 router selects ONE
expert, only that expert's slot was ever non-zero in the prior
layout — the N_E axis was entirely wasteful.
The shape change required:
- Updated `regime_moe_gate_bwd` to write the compact layout.
- New `regime_moe_gate_scatter` kernel scatters per-(b, h)
rank-1 contributions into `grad_experts_W[e]` / `grad_experts_b[e]`
based on `top_e[b]`. Grid (N_E, H, ceil(H/32)) × block (32) —
one warp per (e, d_out, d_in_chunk). 65536 → 16384 grid cells
(4× fewer blocks dispatched).
- Dropped the previously-naive 65536-block `reduce_axis0` for
`grad_experts_w` from `perception.rs` (the scatter kernel
produces the final per-expert grad directly).
- `tests/regime_moe_gate_numgrad.rs` reads `grad_experts_w` from
the scatter output instead of host-side reducing the 5D scratch.
(2) inverted_attention bwd loop interchange
Phase 2's tight loop:
for k:
for j:
ds_myh_j = d_scores[my_h, j] // doesn't depend on k!
ds_j_myh = d_scores[j, my_h] // doesn't depend on k!
...
Hoisted d_scores reads out of the K-loop into J-outer with
per-thread `q_arr[K_MAX]` / `k_arr[K_MAX]` register accumulators.
Net: 32× fewer DRAM reads of d_scores per thread per bwd.
CORRECTNESS:
- regime_moe_gate numgrad PASSES (1/1, 11 numgrad checks).
- inverted_attention numgrad PASSES (1/1, 6 numgrad checks).
- perception_overfit 9/9 PASS — including loss-shrinks tests.
NEXT: re-run cluster smoke to measure the new wall-time vs the 17 s
baseline. Prior smoke at
|
||
|
|
a263cd5446 |
perf(ml-alpha): inverted_attention_pool — cache mean_k, pre-compute pooled
The cluster smoke at
|
||
|
|
9170d24fe3 |
refactor(ml-alpha): replace legacy attention_pool with MultiHorizonAttention [Stage 2]
Single source of truth for the attention path. Deletes the legacy
single-Q `attention_pool.cu` and all `attn_*` fields from
`PerceptionTrainer`; wires `MultiHorizonAttention` (the bundle
introduced in Stage 1) into `step_batched` + `evaluate_batched` as
THE attention summary that seeds CfC's `h_old` at k=0.
Deletions:
cuda/attention_pool.cu (244 lines)
perception.rs::attn_q_d/attn_context_d/
attn_weights_d/grad_attn_q_d/opt_attn_q/
attn_fwd_fn/attn_bwd_fn/_attn_module/
attn_grad_q_scratch_d (all struct fields)
perception.rs::ATTENTION_POOL_CUBIN (include_bytes constant)
Their corresponding init + struct-construction lines.
build.rs::KERNELS (drops "attention_pool")
New kernel + binding:
cuda/horizon_mean_collapse.cu (53 lines)
- `horizon_mean_collapse_fwd/_bwd`: collapses [B, N_H, H] → [B, H]
by averaging over the horizon axis. Single-pass, no reductions.
src/horizon_mean_collapse.rs (host binding)
MHA additions:
- `collapse` field + `ctx_mean_d` + `grad_ctx_mean_d` for the seed.
- `grad_ctx_h_d` scratch (split from grad_horizon_tokens_scratch to
avoid aliasing when MoE bwd writes d_ctx_h while pool bwd writes
d_horizon_tokens).
- `forward(ln_b_out)`: horizon-token pool → inverted pool → MoE
dispatch → mean-collapse → ctx_mean_d.
- `backward(ln_b_out, grad_ctx_mean, grad_ln_out)`: full reverse
chain.
- `apply_anchor()`: launches anchor_l2 on horizon_tokens, Q,
experts_w.
- `adamw_step()`: steps all 6 owned optimizer groups.
PerceptionTrainer integration:
- Section 2d (forward): `self.mha.forward(&self.ln_out_d)` replaces
the legacy attention_pool launch. CfC's h_old at k=0 now reads
`self.mha.ctx_mean_d.device_ptr` (was `self.attn_context_d`).
- Section 7c-pre (backward): `self.mha.backward(ln_out, grad_h_carry,
grad_h_enriched_seq)` replaces the legacy attn_bwd_fn launch.
- Four `reduce_axis0` launches collapse MHA's per-batch scratches
into shared gradient buffers: grad_horizon_tokens, grad_q,
grad_experts_w, grad_experts_b.
- `self.mha.apply_anchor()` adds L2 anchor grad contributions.
- Section 9 (AdamW): `self.mha.adamw_step()` replaces opt_attn_q.
- `evaluate_batched`: `self.mha.forward` replaces the legacy fwd
launch; h_old at k=0 reads `mha.ctx_mean_d`.
- `self.mha.zero_grads()` at step start (capture-safe memset_zeros).
BUG CAUGHT DURING WIRING (NVIDIA-grade discipline): first wiring
attempt mis-sized the reduce_axis0 launches for the MoE
`grad_w_scratch_d` ([B, N_H, N_E, H, H]). Initial `n_tail = N_H * N_E
* H * H = 327680` would have made reduce_axis0 read 5× past the end
of the buffer → CUDA_ERROR_ILLEGAL_ADDRESS. Fix: `n_tail = N_E * H *
H = 65536` with `n_batch = B * N_H`, treating the leading two axes
together as the reduction dimension. Caught by stacked_trainer test
on RTX 3050; would have caused silent corruption then a hard fault
on L40S/H100 later.
LOCAL VERIFICATION (RTX 3050 sm_86):
- ml-alpha builds clean (cuda feature).
- All 38+ tests PASS serially with --test-threads=1:
perception_overfit (8 tests incl. loss-shrinks)
trunk_forward (5)
stacked_loss_shrinks (multiple)
bce_grad_finite_diff (4)
snap_feature_assemble (9)
... (full suite green)
- Numgrad parity for the 4 new MHA kernels (horizon_token, inv_attn,
regime_moe_gate, anchor_l2) PASSES at 5e-2 rel.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
|
||
|
|
6a3f45d872 |
refactor(ml-alpha): consolidate BCE — Kendall σ is THE BCE, wire MHA into trainer
Single source of truth for the multi-horizon BCE per the new
feedback_single_source_of_truth_no_duplicates pearl. Eliminates the
`bce_loss_multi_horizon_sigma.cu` / `loss_sigma.rs` duplication
introduced earlier today and folds the Kendall σ-weighting into the
canonical kernel.
Deletions:
cuda/bce_loss_multi_horizon_sigma.cu (folded into legacy)
src/trainer/loss_sigma.rs (helper merged)
tests/bce_sigma_numgrad.rs (subsumed)
Rewrites:
cuda/bce_loss_multi_horizon.cu — replaced with σ-aware
implementation; kernel function name kept as
`bce_multi_horizon_forward_backward` so cubin symbol stays stable.
NVIDIA-grade warp-shuffle reduce (4 warps, 1 cross-warp barrier);
new args `log_sigma_h` (per-horizon Kendall σ logarithm) and
`d_log_sigma_h` (its gradient).
src/trainer/loss.rs — standalone helper updated to new 11-arg
kernel signature. Exposes optional `log_sigma_h` in `BceInput`
(None → zeros / passthrough Kendall init) and returns
`mean_bce_per_h` + `d_log_sigma_h` in `BceOutput`.
tests/bce_grad_finite_diff.rs — adds `log_sigma_h: None` to the
test inputs; all 4 numgrad tests PASS unchanged.
build.rs — drops `bce_loss_multi_horizon_sigma` entry from
KERNELS. The single canonical `bce_loss_multi_horizon` cubin
now contains the σ-aware kernel.
Wiring:
PerceptionTrainer gains a single `pub mha: MultiHorizonAttention`
field. Owns `log_sigma_h_d` + `grad_log_sigma_h_d` (and the rest
of the multi-horizon attention path, anchored on Stage 2 to fully
replace the legacy `attention_pool` path).
step_batched + evaluate_batched BCE callsites now thread
`mha.log_sigma_h_d` and `mha.grad_log_sigma_h_d` through the
11-arg kernel signature.
This commit keeps the legacy `attention_pool` callsite in place; the
Stage 2 commit will replace it with `mha.pool` + `mha.inv_pool` +
`mha.moe` and delete the `attn_*` fields entirely.
Verified locally: ml-alpha builds clean with the cuda feature,
bce_grad_finite_diff (4/4) PASS on RTX 3050.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
|
||
|
|
55aeddaebd |
feat(ml-alpha): anchor_l2 kernel + Wiener-α controller (v2 B) [V8]
L2 anchor regularization toward initialization (axis B). Anchors
horizon_tokens + Q + MoE experts toward their init values to prevent
the calibration drift observed in v1 (where val_loss climbed as α
opened past epoch 1 in 2 of 3 folds).
KERNEL (`anchor_l2_fwd_bwd`):
loss_out = λ · Σ_i (p[i] − p_init[i])²
grad_p[i] += 2λ · (p[i] − p_init[i])
- Warp-shuffle reduce; one block per parameter group; strided thread
loop over n. Cross-warp reduce uses one __syncthreads.
- Coalesced grad write via stride loop.
- λ passed as device-side [1]-buffer (host writes scalar before launch
— capture-safe).
CONTROLLER (`trainer::anchor_controller::AnchorController`):
- Signal-driven λ floor: λ_floor = ‖p_init‖ / (100 · √numel).
Cross-fold-persistent per pearl_kelly_cap_signal_driven_floors.
- Wiener-α smoother (α = diff_var / (diff_var + sample_var + ε))
on val_loss change; α floored at 0.4 per
pearl_wiener_alpha_floor_for_nonstationary.
- λ blends toward target = |ema_change|·scale with α; floored at
λ_floor per pearl_blend_formulas_must_have_permanent_floor.
- First-observation bootstrap (sentinel state replaced directly on
first step) per pearl_first_observation_bootstrap.
- 4/4 unit tests PASS: signal-floor init, bootstrap returns floor,
floor protection across 1000 steps, λ_max cap.
NUMGRAD VERIFICATION (RTX 3050 sm_86):
anchor_l2_numgrad PASSES with closed-form parity (machine precision)
and central-difference parity (4 random positions) within 5e-2 rel.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
|
||
|
|
11b96dac6a |
feat(ml-alpha): regime_moe_gate kernel + numgrad (v2 D) [V6]
Top-1 Mixture-of-Experts gate per Switch Transformer. Three kernels
in one .cu file:
- regime_moe_gate_fwd: select top-1 expert from gate_logits, apply
its [H, H] linear to fused_ctx → routed_ctx [B, N_H, H].
- regime_moe_gate_bwd: chain-rule grads through the SELECTED
expert's W and bias (sparse-by-expert scratches), plus
d_fused_ctx accumulator. Inactive experts get zero contribution.
- regime_moe_gate_aux: softmax of gate_logits + load-balancing
auxiliary loss (frac · prob_mean × N_EXPERTS).
ARCHITECTURE:
- N_EXPERTS = 4. Each expert is a [H, H] linear with bias.
- Total expert params: 4 · 128 · 128 + 4 · 128 = 66 KB. Cheap.
- STE on gate: gate logit grad is zero from the expert path (top-1
is non-differentiable); the load-balance aux loss provides the
differentiable signal that pushes routing toward balanced usage.
PERFORMANCE:
- Grid (B, N_H, 1), block (HIDDEN_DIM). One block per (b, h).
- Forward: each thread computes one output channel via a dot
product over HIDDEN_DIM input dims (#pragma unroll 8).
- Backward d_fused_ctx: each thread accumulates over d_out
sequentially (HIDDEN_DIM iterations) since the weight matrix
column is naturally aligned to the thread's d_in index.
- Backward d_W/d_b scratches are sparse-by-expert; downstream
reduce_axis0 collapses over (B, N_H).
- Top-1 chosen by thread 0 per block, broadcast via shared mem.
NUMGRAD VERIFICATION (RTX 3050 sm_86):
forward_matches_host_reference_and_backward_matches_numgrad
PASSES 11 checks (4 on d_W, 3 on d_b, 4 on d_fused_ctx) within
5e-2 rel / 5e-3 abs.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
|
||
|
|
d94696620d |
feat(ml-alpha): inverted_attention_pool kernel + numgrad (v2 E) [V3]
iTransformer-style cross-variate attention pass: each of HIDDEN_DIM
features becomes a "variate token" with its K-trajectory as embedding.
FORWARD:
X_inv[h, k] = ln_out[k, h] # transpose
scores[h, j] = inv_scale · Σ_k X_inv[h, k] · X_inv[j, k]
attn[h, j] = softmax_j(scores)
pooled[h] = Σ_j attn[h, j] · mean_k_X_inv[j] # mean-K commutes out
BACKWARD: three independent chains into ln_out:
- value path: (1/K) · attn[h', my_h] · d_pooled[h'] (per-k constant)
- query path: inv_scale · Σ_j d_scores[my_h, j] · X_inv[j, k]
- key path: inv_scale · Σ_h' d_scores[h', my_h] · X_inv[h', k]
Softmax bwd: d_scores[h, j] = attn · (d_attn - Σ_l attn · d_attn)
IMPLEMENTATION NOTES:
- First attempt cached attn [H, H] = 64 KB in shared mem → tripped
the 48 KB dynamic-shared limit on sm_86 (CUDA_ERROR_INVALID_VALUE).
- Fixed by moving d_scores to a DRAM scratch buffer; shared mem
holds only X_inv [H, K] (≤ 16 KB at K = 32). One block-wide barrier
between the d_scores write and the value/query/key accumulation.
- All per-batch slice writes; no atomicAdd, no cross-block race.
- Pooled computation uses the mean-K commute (Σ_k attn · X_inv =
attn · mean_k_X_inv), saving an entire H×K accumulation pass.
LOCAL VERIFICATION (RTX 3050 sm_86):
forward_then_backward_matches_central_difference PASSES 6 numgrad
checks on ln_out positions within 5e-2 rel / 5e-3 abs.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
|
||
|
|
e8f6c4721f |
feat(ml-alpha): horizon_token_attention_pool kernel + numgrad (v2 C) [V2]
Replaces the falsified per-horizon Q_h pool with a single shared
query Q over an extended key sequence [horizon_tokens; LN_b_out],
producing per-horizon outputs via TFT-style horizon-token mixing.
FORWARD:
scores[i] = Σ_d Q[d] · ext[i, d] (i ∈ [0, N_H + K))
attn = softmax_i(scores)
S[d] = Σ_k attn[N_H + k] · ln_out[k, d] (shared time agg)
ctx[h, d] = attn[h] · horizon_tokens[h, d] + S[d] (per-horizon out)
BACKWARD: full chain rule with softmax-centring; gradients to
horizon_tokens, Q, and ln_out via the saved attn weights.
NVIDIA-grade implementation per feedback_nvidia_grade_perf_for_kernels:
- Warp-shuffle reduce (block_reduce_sum / block_reduce_max helpers)
for all per-d dot products and softmax aggregates.
- Cross-warp reduce uses exactly one __syncthreads.
- Non-divergent shuffles: inactive lanes contribute 0 via ternary.
- Block-per-batch + horizon-loop inside block → grad_ln_out += is
race-free without atomicAdd.
- Smem layout computed at launch: [s_attn(N_H+K); s_warp(N_WARPS);
s_d_S(H) on bwd]. No over-allocation.
LOCAL VERIFICATION (RTX 3050 sm_86):
forward_then_backward_matches_central_difference PASSES 12 numgrad
checks (4 each on horizon_tokens / Q / ln_out) at 5e-2 rel / 5e-3
abs envelope. First-try pass.
NOTE: .gitignore adjusted with narrow allow-rules for crates/ml-alpha/{
cuda,src,tests}/horizon_token_* paths — the broad "*token*" rule
intended for auth tokens was hiding these source files. Explicit
allow keeps the security rule intact while exempting these specific
files.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
|
||
|
|
679ab3f5eb |
feat(ml-alpha): Kendall sigma-weighted BCE kernel + numgrad (v2 A) [V7]
New `bce_multi_horizon_sigma_forward_backward` kernel implementing the
Kendall homoscedastic uncertainty weighting per spec axis A:
raw_bce_h = Σ_{i in h, m_i=1} L_i
count_h = #{i in h : m_i = 1}
mean_bce_h = raw_bce_h / count_h
w_h = base_weight_h / (2 · exp(2 · log_sigma_h))
total_loss = Σ_h [ w_h · mean_bce_h + log_sigma_h ]
d L / d p_i = m_i · (w_h / count_h) · (p − y) / (p (1 − p))
d L / d log_sigma_h = 1 − 2 · w_h · mean_bce_h
NVIDIA-grade implementation per feedback_nvidia_grade_perf_for_kernels:
- Warp-shuffle reduction (`__shfl_xor_sync`) for both per-horizon
sums and the global valid count, replacing block tree-reduce.
- One `__syncthreads` for the cross-warp aggregate; no inner-loop
barriers.
- Non-divergent shuffles: inactive lanes contribute 0 via ternary,
never via `if (tid < N) shuffle`.
- Coalesced strided access in both forward and gradient passes.
- Pre-compiled cubin via build.rs; no nvrtc.
Independent of the legacy `bce_loss_multi_horizon` kernel — that one
stays untouched so eval/smoke paths are unaffected. The v2 trainer
wires this kernel in via commit V10.
Standalone helper `bce_sigma_loss_and_grad_gpu` in `trainer::loss_sigma`
for numgrad parity tests. Three numgrad tests all PASS on RTX 3050
(sm_86) within 5e-2 rel / 5e-3 abs:
- d_log_sigma_h ↔ central-difference (numgrad on log_sigma)
- grad_probs ↔ central-difference (8 random positions)
- total_loss ↔ closed-form reconstructed from mean_bce_per_h
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
|
||
|
|
41292303dc |
refactor(ml-alpha): remove per-horizon Q_h path (C21-C25 falsified) [V1]
3-fold A/B sweep 2026-05-18 at commit
|
||
|
|
286ea26e2a |
perf(ml-alpha): warp-shuffle reduce in per-horizon kernels
Cluster A/B sweep with C25 wiring showed 86 s/epoch vs 17 s/epoch baseline = ~5x regression. Root cause: per-horizon attention pool + residual head used block tree-reduce with 8 __syncthreads per K-step in a serialised K-loop, repeated H=5 times in both fwd and bwd = ~3200 barriers/step. Plus the prob_blend bwd reduce kernel ran with a single thread per block, fully serialising over K*B. Replacements: - per_horizon_attention_pool fwd/bwd: introduce block_reduce_sum / block_reduce_max helpers using intra-warp __shfl_xor_sync + cross-warp shuffle (1 syncthread per K instead of 8). Smem shrinks to [K + N_WARPS] / [2K + N_WARPS]. - per_horizon_residual_head fwd: same warp-shuffle reduce pattern. - per_horizon_prob_blend_reduce_alpha_residual: 1 thread → 1 warp per horizon, lane-strided reduction over K*B via shfl_xor_sync. Launch config updated to block_dim=(32,1,1). Tricky bug found while implementing: the cross-warp reduce in the residual head originally guarded `__shfl_xor_sync(0xffffffff, ...)` with `if (tid < PHR_N_WARPS)`, leaving 28 of 32 lanes in warp 0 outside the call. Mask 0xffffffff requires all 32 lanes to participate — divergence is UB and hung the full-pipeline smoke on Ampere/Ada. Fix: read s_warp via ternary into all 32 lanes, then shuffle inside `if (tid < 32)`. Matches the pattern used in block_reduce_sum. Verified locally on RTX 3050 (sm_86): per_horizon_attention_pool numgrad, per_horizon_residual_head numgrad, and per_horizon_full_pipeline_smoke (zero-init identity + non-zero end-to-end) all PASS. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> |
||
|
|
f50b466f77 |
feat(ml-alpha): PerceptionTrainer wires C21+C22 fwd+bwd+AdamW (C25)
The per-horizon attention pool from C21+C22 is now fully integrated
into step_batched's hot loop. Forward and backward both flow; AdamW
updates all four new parameter groups (Q_h, w_res, bias_res, α) every
training step. At α=0 init the contribution is byte-identical to
baseline; training discovers whether α should grow.
New kernel: cuda/per_horizon_prob_blend.cu
per_horizon_prob_blend_fwd
Reads logit_per_k_d (already stored by GRN forward) +
sigmoid(logit_baseline + tanh(α[h]) * residual[b, h]) →
overwrites probs_per_k_d in place. At α=0: r_contrib=0, output
== sigmoid(logit_baseline) == probs_baseline → bit-identical.
per_horizon_prob_blend_reduce_alpha_residual
Reads probs_per_k (= p_final, post-blend) + grad_probs_per_k
(= ∂L/∂p_final from BCE) and computes:
d_logit[k,b,h] = grad_probs[k,b,h] * p_final * (1 - p_final)
d_residual[b,h] = tanh(α[h]) * Σ_k d_logit[k,b,h]
d_alpha[h] = sech²(α[h]) * Σ_{k,b} d_logit[k,b,h] * residual[b,h]
No separate prob_blend_bwd needed — chain-rule equivalence
∂L/∂logit_baseline = ∂L/∂r_contrib (both flow through the same
sigmoid derivative) means the existing GRN backward is UNCHANGED.
trainer/per_horizon_state.rs extensions:
forward_with_blend(ln_b_out, logit_per_k, probs_per_k)
Pool fwd → context_h; head fwd → residual; prob_blend fwd
in-place rewrites probs_per_k.
backward_through_blend(probs_per_k, grad_probs_per_k, ln_b_out,
grad_ln_b_out_target)
Reduce kernel → d_residual + d_alpha. Then:
head bwd → d_w_res_scratch, d_bias_res_scratch, d_context.
pool bwd → d_q_h_scratch, += grad_h_enriched_seq_d.
Per-batch scratches reduced to shared grads host-side
(n_batch ≤ 64 → sub-millisecond on host).
adamw_step()
Steps the four optimizers using the shared grad buffers.
zero_grads()
Called once per step before forward to clear scratch.
trainer/perception.rs step_batched integration:
── 4.5 (after GRN K-loop, before BCE): zero_grads + forward_with_blend
overwrites probs_per_k_d with p_final.
── 5 (existing BCE consumes probs_per_k_d as today; grad_probs is
now ∂L/∂p_final automatically).
── 5a (after BCE, before ISV-lambda + heads bwd): backward_through_blend.
Existing GRN bwd path is UNTOUCHED — the chain rule absorbs
the bias.
── 9 (after existing 17 AdamW group steps): per_horizon.adamw_step
updates Q_h, w_res, bias_res, α.
Verification:
- 34 ml-alpha lib tests still green.
- Per-horizon kernel numgrad parity (C21, C22) still green.
- Per-horizon end-to-end pipeline smoke (C23, including the
alpha=0 byte-identity invariant) still green.
- Full workspace builds clean.
Closes the kernels+wiring portion of #203 (per-horizon attention pool
kernels + wiring). What remains (#204): 30-epoch × 3-fold A/B vs
single-Q baseline. The branch is ready for that sweep when GPU time
is budgeted; the implementation is structurally adoption-safe
(α=0 → identity to baseline) so it can be merged before the A/B if
desired.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
|
||
|
|
8c335caef7 |
feat(ml-alpha): per-horizon residual head kernel + numgrad parity (C22)
Companion kernel to C21's per_horizon_attention_pool. Computes a
per-horizon scalar residual from each horizon's context vector:
residual[b, h] = Σ_d w_res[h, d] * context_h[b, h, d] + bias_res[h]
Designed to be added (behind a learnable α-gate) to the existing
multi_horizon_heads logit output — keeps the existing GRN head kernel
completely unchanged. The per-horizon attention pool's contribution
flows through this lightweight projection without weight-shape
changes elsewhere or checkpoint-V2-bumping.
Path A integration sketch (deferred to follow-up commit C23):
alpha_logit_per_horizon = existing_head(h_K)[h] # from current path
+ tanh(α[h]) * residual_kernel(context_h)[h]
where α[h] is a learnable 5-vector init'd to 0 (no effect at start).
Training discovers per-horizon whether the residual contributes.
This is a strict superset of the existing path — α=0 → bit-identical
to today.
Backward kernel produces:
d_w_res — per-block scratch [B, N_HORIZONS, HIDDEN_DIM]
for host reduce_axis0 → shared [N_HORIZONS, HIDDEN_DIM]
d_bias_res — per-block scratch [B, N_HORIZONS], same reduction
d_context_h — per-batch indexed; += chained with attention bwd
Single-writer discipline preserved (no atomicAdd per
feedback_no_atomicadd.md); horizon loop inside the per-batch block.
Numgrad parity test:
- B=3, N_HORIZONS=5, HIDDEN_DIM=128 fixture.
- Loss = Σ residual_out (so d_residual = 1).
- Probes 8 random w_res indices, all 5 bias_res entries, 8 random
context_h indices via central-difference at ±eps=1e-2.
- All within 5e-2 rel-tol or 5e-3 abs-floor.
- Passes on RTX 3050.
Same scope discipline as C21: kernel + binding + numgrad first;
trainer wiring + α-gate + smoke training + A/B sweep follow once
both kernels are individually validated (now done).
Closes the second kernel-correctness portion of #203. Remaining:
C23: trainer wiring (capture attn_pool fwd into the graph; sum
residual into existing head output with α-gate)
C24: CheckpointV1 → V2 bump (add q_h, w_res, bias_res, alpha fields)
C25: 1-epoch smoke (assert no NaN, loss decreases vs baseline)
C26: 30-epoch × 3-fold A/B (#204) — decision gate per spec §0
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
|
||
|
|
edc449eecf |
feat(ml-alpha): per-horizon attention pool kernel + numgrad parity (C21)
First implementation slice of the per-horizon attention pool design
(docs/superpowers/specs/2026-05-18-per-horizon-attention-pool-design.md).
Lands the kernel + Rust binding + numgrad verification; downstream
wiring into PerceptionTrainer's captured graph + CheckpointV2 bump +
A/B sweep are follow-up commits gated on this proving correctness.
Kernel (cuda/per_horizon_attention_pool.cu):
per_horizon_attention_pool_fwd
Q_h[N_HORIZONS, HIDDEN_DIM] × LNb[B, K, HIDDEN_DIM]
→ context_h[B, N_HORIZONS, HIDDEN_DIM]
attn_h_weights[B, N_HORIZONS, K]
Per-block math identical to the single-Q variant, looped over
N_HORIZONS sequentially within each batch's block. Grid stays
(B, 1, 1) so backward grad_ln_out writes are race-free
(per feedback_no_atomicadd.md — no cross-block contention).
Per-batch shared mem ~k_seq + BLOCK + HIDDEN_DIM floats.
per_horizon_attention_pool_bwd
Same chain-rule pattern as attention_pool_bwd but with the horizon
loop inside the block: each (b, h) slice updates grad_ln_out in
place (sequential horizon accumulation), grad_Q_h is written as
per-block scratch [B, N_HORIZONS, HIDDEN_DIM] for host reduce.
Single-writer discipline preserved.
Rust binding (src/per_horizon_attention_pool.rs):
PerHorizonAttentionPool::{new, forward, backward}. Self-contained;
doesn't yet touch PerceptionTrainer or CfcTrunk. Loads the cubin
via the standard env!("OUT_DIR") path. Dynamic shared-mem byte
count computed per launch from k_seq.
Numgrad parity test (tests/per_horizon_attention_pool_numgrad.rs):
- B=2, K=8, HIDDEN_DIM=128, N_HORIZONS=5 fixture.
- Loss = Σ context_h (so d_context = 1 everywhere — clean analytical).
- Backward kernel produces analytical grads; central-difference of
forward kernel at ±eps=1e-2 across 8 random Q_h indices + 8 random
LNb indices verifies analytical matches CD within 5e-2 rel-tol or
5e-3 abs-floor.
- Passes on RTX 3050.
build.rs picks up the new .cu file automatically via the existing
KERNELS list; cubin compiles cleanly at sm_86 + sm_89.
Same scope discipline as Phase 2D.2 (VSN numgrad) — kernel correctness
first, integration second. The follow-up commit set per the spec §3
appendix:
C22: extend multi_horizon_heads.cu signature to accept per-horizon
context input + bump head_w shape to [N_HORIZONS, 2*HIDDEN_DIM]
C23: wire PerHorizonAttentionPool into PerceptionTrainer + CfcTrunk
captured graph behind AttentionPoolVariant config flag
C24: CheckpointV1 → V2 bump with discriminant + optional q_h field
C25: smoke training (one epoch, no NaN, loss decreases)
C26: 30-epoch × 3-fold A/B sweep (#204) — decision gate per spec §0
Closes the kernel-correctness portion of #203.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
|
||
|
|
a478ba3d84 |
perf(ml-alpha): block-per-batch attention pool bwd refactor (Phase B commit 4)
attention_pool_bwd refactored from grid=(1,1,1) to grid=(B,1,1). The existing per-batch grad_ln_out writes were already uniquely indexed; only grad_Q needed scratch+reducer (1 scratch, 1 reducer launch). Adds 1 per-batch grad scratch buffer + 1 reduce_axis0 launch: attn_grad_q_scratch_d [B, HIDDEN_DIM] ~16 KB scratch at B=32 — trivially small. attn_pool bwd runs 1×/step (not in K-loop) so the absolute wall-time win here is tiny. With this commit every single-SM bwd kernel in the trainer has been refactored to block-per-batch + scratch+reducer. Phase B kernel work complete. Next: local + cluster A/B perf benchmark to verify acceptance gates 6, 7, 8 from the spec. All 9 perception_overfit smokes pass. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> |
||
|
|
9607f33518 |
perf(ml-alpha): block-per-row VSN bwd refactor (Phase B commit 3)
variable_selection_bwd refactored from grid=(1,1,1) to grid=(B*K,1,1). VSN's n_rows = B*K positions (one row per (batch, K-position) pair); block-per-row matches the existing fwd kernel's layout. Adds 2 per-row grad scratch buffers + 2 reduce_axis0 launches: vsn_grad_w_scratch_d [B*K, FEATURE_DIM, FEATURE_DIM] vsn_grad_b_scratch_d [B*K, FEATURE_DIM] ~210 KB scratch at B=32, K=64. VSN bwd runs 1×/step (not K×) so the absolute wall-time win here is small versus commits 1+2. Done for pattern uniformity — every per-batch or per-row bwd in the trainer now uses scratch+reducer. All 9 perception_overfit smokes pass. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> |
||
|
|
5c2c3b65a8 |
perf(ml-alpha): block-per-batch GRN bwd refactor (Phase B commit 2)
multi_horizon_heads_grn_bwd_batched refactored from grid=(1,1,1) to grid=(B,1,1). Removes the single-SM bottleneck on the second-most-called K-loop kernel (64×/step like cfc_bwd). Adds 10 per-batch grad scratch buffers (one per GRN param tensor) + 10 reduce_axis0 launches collapsing B → final grad after the K-loop: grn_grad_w1_scratch_d [B, 5, HEAD_MID, HIDDEN] grn_grad_b1_scratch_d [B, 5, HEAD_MID] grn_grad_w2_scratch_d [B, 5, HEAD_MID, HEAD_MID] grn_grad_b2_scratch_d [B, 5, HEAD_MID] grn_grad_w_gate_scratch_d [B, 5, HEAD_MID] grn_grad_b_gate_scratch_d [B, 5] grn_grad_w_main_scratch_d [B, 5, HEAD_MID] grn_grad_b_main_scratch_d [B, 5] grn_grad_w_skip_scratch_d [B, 5, HIDDEN] grn_grad_b_skip_scratch_d [B, 5] Total: ~8 MB scratch at B=32. All 9 perception_overfit smokes pass (including stacked_trainer_loss_ shrinks_at_batch_32 which exercises the cross-batch reducer path on both cfc and GRN grads). Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> |
||
|
|
494a2e4827 |
perf(ml-alpha): block-per-batch cfc_step + reduce_axis0 reducer (Phase B commit 1)
Per docs/superpowers/specs/2026-05-17-kloop-parallelization-design.md. cfc_step_batched (fwd + bwd) refactored from grid=(1,1,1) with internal n_batch loop to grid=(B,1,1) — each block handles one batch. Removes the single-SM bottleneck on the K-loop's most-called kernel (64×/step). Param-grad accumulation moves to per-batch scratch: cfc_grad_w_in_scratch_d [B, n_hid, n_in] cfc_grad_w_rec_scratch_d [B, n_hid, n_hid] cfc_grad_b_scratch_d [B, n_hid] cfc_grad_tau_scratch_d [B, n_hid] Zeroed once per training step, K-loop's 64 bwd calls += into them, then 4 reduce_axis0 launches collapse B → final grad buffers (OVERWRITE) before AdamW. New AdamW-after-reducer invariant: final grads are meaningful only after the reducer has run in the current step. New reduce_axis0 kernel: single parameterised reducer [B, N] → [N] via block tree-reduce (no atomicAdd per feedback_no_atomicadd.md). Same pattern as layer_norm_reduce_param_grads — CUDA-Graph-safe. cfc_step_backward_batched shared-mem dropped from (B+1)*n_hid*4 to 2*n_hid*4 bytes per block (only one row of sd_pre needed per block bi). Tests: - New stacked_trainer_loss_shrinks_at_batch_32: FIRST test that actually exercises the cross-batch reduction code path; existing perception_overfit suite was all B=1. Initial 0.24 → final 0.00. - Scratch-clears test removed (explanatory comment kept): structurally hard to assert directly due to begin_capture/end_capture not executing kernels; the B=32 convergence smoke implicitly validates scratch zeroing since divergence would otherwise be immediate. All 9 perception_overfit smokes + 4 backward_finite_diff tests pass. build.rs: - KERNELS list adds "reduce_axis0" - Cache-bust → v11 Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> |