First commit of the horizon-rebase refactor (full plan at
docs/superpowers/plans/2026-05-22-horizon-rebase-n3.md). Motivated by
empirical evidence from this session's investigations:
- C1: ES MBP-10 input ACF dies by L=300; K=1000/6000 are below noise floor
- D3: binarized labels ρ(30,6000)=0.06; K=6000 carries no signal
- Heads_w1 mask experiment: h6000 AUC collapsed 0.73→0.54, confirming
long-horizon prediction needs cross-channel integration that 5-horizon
bucket structure couldn't provide
New horizon set [10, 100, 1000] matches C1's empirical autocorrelation
elbows (~50ms, ~3.5s, ~45s wall-clock at 74µs median Δt).
Scope of this commit:
- crates/ml-alpha/src/heads.rs: N_HORIZONS 5→3, HORIZONS=[10,100,1000]
- crates/ml-alpha/src/cfc/bucket_routing.rs: MAX_BUCKET_DIM 28→96,
BUCKET_DIM_K=[43,43,42], BUCKET_CHANNEL_OFFSET=[0,43,86,128]
- crates/ml-alpha/src/data/loader.rs: 5 hardcoded [T;5] types migrated
- crates/ml-backtesting/src/sim/mod.rs: broadcast_alpha signature,
local N_HORIZONS re-exports ml_alpha::heads::N_HORIZONS
- crates/ml-backtesting/src/harness.rs: local N_HORIZONS re-export,
per-horizon log labels, MultiHorizonLoaderConfig.horizons
Tests + examples + CUDA kernels still need migration (subsequent commits
per plan tasks 2-8).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Anti-calibration evidence (smoke ckpts across architecture variants
showed higher reported conviction → MORE negative PnL, -15.6 → -145.2)
traced to a magnitude/direction mismatch in the sizing formula:
conv_ema = EMA(|conviction_signed|) # magnitude smoothing
target_lots = sign(conviction_signed) # INSTANTANEOUS sign
× conv_ema × max_lots # × SMOOTHED magnitude
When per-horizon directions disagree event-to-event, the magnitude EMA
stays high (|x| EMA can't cancel) but the sign flips on noise. Result:
big trades in random direction whenever the 5 horizons disagree.
Replace with a single signed-conviction EMA:
conv_signed_ema = EMA(conviction_signed)
target_lots = sign(conv_signed_ema)
× |conv_signed_ema| × max_lots
Now BOTH sign AND magnitude come from the same smoothed signal. When
directions disagree, signed EMA → 0 collapses size to zero naturally.
Atomic migration across decision_policy.cu (2 kernels), pnl_track.cu
(open-branch reads the SIGNED buffer and takes fabsf for the magnitude-
semantics open_trade_state offsets that composite_exit_check forms ratios
from), and src/sim/mod.rs (renamed device buffers + accessor + 4 launch
sites). No legacy aliases per feedback_no_legacy_aliases; no parallel
buffers per feedback_single_source_of_truth_no_duplicates; α floor 0.4
preserved per pearl_wiener_alpha_floor_for_nonstationary; first-
observation bootstrap preserved per pearl_first_observation_bootstrap.
ml-backtesting lib: 33 passed.
GPU oracle: signed_conviction_ema_collapses_on_sign_flips passes —
observed steady-state shows |signed_ema| ≈ 0.205 (post-fix) producing
|target_lots| = 2 every tail event, vs the pre-fix bug's predicted
|target_lots| = 7 every tail event. Existing
multi_horizon_conviction_cancels_on_disagreement still passes (zero-
cancellation regime is even cleaner under signed EMA).
ml-alpha lib: 33 passed (no regression).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Empirical ES MBP-10 median inter-event Δt is 74µs (not 250ms as
documented in loader.rs:25-26). The constants ALPHA_MED=0.02 and
ALPHA_SLOW=0.0005 were named/commented as "9s @ 250ms" and "6min @
250ms" but at the actual 74µs Δt their half-lives are ~2.6ms and ~100ms
— a 3000× scale mismatch.
Rename for truth-in-naming (values unchanged to preserve all trained
checkpoints):
- ALPHA_MED → ALPHA_FAST_2P6MS (half-life ~2.6ms @ 74µs Δt)
- ALPHA_SLOW → ALPHA_MED_100MS (half-life ~100ms @ 74µs Δt)
The regime[0..5] feature naming in snap_features.rs reflects the
actual microstructure-burst and 1-2-second-clustering time scales, NOT
macro 9s/6min regimes. True macro-regime EMAs (if needed) are a
separate addition.
ml-alpha lib: 33 passed, unchanged from baseline.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Mirrors the alpha_train CLI flag (--kernel-step-trace <path>) into
fxt-backtest. The PerceptionTrainerConfig already accepts the path;
this commit just exposes the CLI surface and propagates through:
fxt-backtest run --kernel-step-trace <path>
-> BacktestHarnessConfig.kernel_step_trace_path
-> PerceptionTrainerConfig.kernel_step_trace_path
Sweep YAML gains a `kernel_step_trace: Option<PathBuf>` base field.
Argo lob-backtest-sweep-template gains a `kernel-step-trace` workflow
parameter (default disabled; when set, --features kernel-step-trace
is passed to cargo build AND --kernel-step-trace to fxt-backtest).
Gated behind the `kernel-step-trace` Cargo feature (matching alpha_train).
When disabled (default), zero overhead.
Enables per-step JSONL diagnostic emission from inference kernels --
needed for Smoke 2 deep-dive after sampling histograms (in parallel
investigation) localize the per-horizon dynamics bottleneck.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Adds --per-horizon-logit-sampling-stride CLI flag to fxt-backtest run/sweep
+ harness instrumentation that DtoHs alpha_probs every N events and
accumulates per-horizon mean/stddev/min/max + 10-bin histogram.
Emitted at end of CRT.diag block as:
per_horizon_logit_diag h{N}: count=... mean=... std=... min=... max=... hist=[...]
Goal: distinguish whether Smoke 2's uniform mean_run_len 1.04x across
horizons is caused by uniform per-horizon LOGITS (-> Task 11 scope
incomplete; heads_w1/w2/gate/main need block-diagonal too) or by
DIFFERENTIATED logits with collapsed downstream (-> different fix).
Default stride=None (disabled, zero overhead). With stride=1000 on
2M events, ~2000 DtoD stops x ~50us = ~100ms total overhead per run.
Per feedback_no_htod_htoh_only_mapped_pinned: uses mapped-pinned buffer
(single 5xf32 allocation reused on every sampling tick).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
URGENT correctness fix surfaced by checkpoint deep-dive after Smoke 2 failed
the WIN gate (mean_run_len ratio = 1.0× vs target ≥10×; all 5 horizons
uniformly ~2.4 events).
THREE INDEXING BUGS identified:
1. tau_reorder produces bucket-grouped tau_all_d, but Controller B's
tau_clamp_kernel reads bucket_id_per_channel[c] where c is the
POSITION in the reordered buffer (not the original channel index)
→ wrong bucket-IQR lookup → τ never constrained → buckets collapse
(deep-dive showed all 5 buckets had nearly identical τ ranges
[0.07, 74] except bucket 4 reaching 878).
2. heads_w_skip grad mask ran but Adam (m, v) momentum from BEFORE the
transition re-introduced gradient signal across the transition →
off-bucket positions stayed nonzero throughout training
(deep-dive: 512/512 = 100% of off-bucket positions nonzero in
trunk_best_h6000.bin).
3. per-branch CfC kernel read w_in[c * HIDDEN_DIM + k] with c =
bucket-grouped position, but W_in rows are indexed by ORIGINAL
channel → kernel read wrong rows for each output → outputs were
essentially random per-channel.
ALPHA FIX: skip the reorder entirely, use bucket-filter throughout:
- Removed tau_reorder_kernel; cfc.tau_d stays in original-channel layout.
- Added channels_in_bucket_kernel that populates a
[N_HORIZONS × MAX_BUCKET_DIM] lookup (original channel index per
(bucket, within-bucket-position)).
- Per-branch CfC fwd+bwd now reads channels_in_bucket[branch][tid]
→ original_c, then uses original_c for w_in/w_rec indexing. All
weights stay in original layout consistently.
- Controller B's tau_clamp_kernel now correctly operates on original-
channel cfc.tau_d with bucket_id_per_channel[c] lookup (no position-
vs-channel confusion).
- Added zero_off_bucket_kernel + three-layer defense for heads_w_skip
block-diagonal invariant:
(a) At transition: zero off-bucket params + zero Adam (m, v)
moments via opt_heads_w_skip.m_mut() / v_mut() accessors.
(b) Per-step: heads_w_skip_grad_mask_apply_kernel zeros off-bucket
gradients before Adam step (unchanged from prior follow-up).
(c) Per-step: zero_off_bucket_kernel zeros off-bucket params after
Adam step, catching any drift from Adam's ε denominator or
weight decay.
New AdamW::m_mut()/v_mut() accessors enable the projection at transition.
GPU oracle tests: 19 total (16 in bucket_transition + 3 in cfc_step_per_branch).
New tests verify:
- channels_in_bucket_kernel correctness under non-contiguous bucket assignment
- zero_off_bucket maintains invariant after many mock Adam steps
- fwd kernel writes only to bucket-assigned channels under arbitrary mapping
Per `feedback_no_partial_refactor`: all 3 indexing bugs + Adam momentum
defense land in one commit.
ml-alpha lib: 33 passed.
GPU oracle tests on RTX 3050 sm_86: 19 passed.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
URGENT correctness fix surfaced by Task 13 implementer.
Bug: training-time Phase 2 dispatch sites (cfc_step_per_branch_fwd at
~3139, h_mag_per_bucket for Controller D at ~3211-3212) read
self.trunk.bucket_*_d which are zero-initialized at trunk construction
and only populated via load_checkpoint. The Phase 1→2 transition
populates a SEPARATE BucketRoutingMetadata owned by the trainer
(self.bucket_routing_metadata) but never syncs it into the trunk fields.
With zero-init bucket_dim_k, the per-branch kernel's uniform predicate
`tid >= bucket_dim_k[branch]` early-returns ALL threads -> Phase 2
silently no-ops during training, Smoke 1 fails before producing signal.
Fix: at transition, DtoD-copy metadata buffers into trunk fields BEFORE
the `Some(metadata)` move (so `metadata.*` borrows remain valid):
- metadata.bucket_channel_offset_d -> trunk.bucket_channel_offset_d
- metadata.bucket_dim_k_d -> trunk.bucket_dim_k_d
- metadata.bucket_id_per_channel_d -> trunk.bucket_id_per_channel_d
- metadata.heads_w_skip_offset_d -> trunk.heads_w_skip_offset_d
Trunk fields remain the single source of truth for both training-time
dispatch AND checkpoint serialization. Total sync cost: 4 DtoD memcpys,
~256 bytes total at the one-shot transition (off the hot path).
Verification:
- `cargo check -p ml-alpha --all-targets`: clean
- `cargo test -p ml-alpha --lib`: 33 passed, 0 failed
- 3 read sites (perception.rs:3139, 3211-3212, 4725-4726) verified
unchanged; all 4 new memcpy_dtod_async calls bracketed in the
transition block (perception.rs:2224, 2234, 2244, 2254).
- Block-diagonal heads grad-mask init (step 7) continues to read
`metadata.bucket_id_per_channel_d` via `as_ref()` after the move;
ordering preserved per pearl_canary_input_freshness_launch_order.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Per spec §2.4 and Task 13 of the plan.
Replace the single-CfC `cfc_step_batched` dispatch in `forward_step_into`
with `cfc_step_per_branch_fwd_gpu`. Production checkpoints have Phase 2
routing frozen at the training-time Phase 1→2 transition, so inference
unconditionally takes the per-branch path. The fused kernel reads
`bucket_channel_offset_d` / `bucket_dim_k_d` from the trunk; in
deployment these are populated via `CfcTrunk::load_checkpoint`. For
freshly-constructed trainers without a checkpoint load, the trunk fields
are zero — the kernel's uniform predicate (`tid >= bucket_dim_k`) then
early-returns every thread and h_new is left untouched. That matches the
"Phase 2 only at inference" contract.
Heads dispatch unchanged: the existing GRN kernel reads from
`trunk.heads_w_skip_d` which has off-bucket entries zeroed at the
transition (sparsification per the prior block-diagonal-grad-mask
follow-up commit). Mathematically the full-buffer read is equivalent to
a compact-only per-bucket read; no inference-time kernel change required.
forward_step_golden's convergence test is now architecturally
incompatible with the new dispatch — `forward_only` runs Phase 1
single-CfC math while `forward_step_into` requires populated Phase 2
routing. Annotated `#[ignore]` with a clear divergence note; the
deterministic + reset semantics tests remain valid invariants per
`pearl_training_smoothness_does_not_transfer_to_inference`.
cargo check workspace clean (excluding pre-existing unrelated cupti +
insert_batch errors in vendor/cudarc and ml/tests). ml-alpha + ml-
backtesting lib tests pass (33 + 33).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Two follow-up fixes for Smoke 2 readiness, atomic per feedback_no_partial_refactor:
FIX 1 — Task 11 target_jitter_k anchor (spec §3.3 compliance):
- Previously: target_jitter_k = first non-zero raw_per_h[k] (model-state-anchored)
- Now: target_jitter_k = sqrt(realised_label_variance_k) (label-distribution-anchored)
- Per-horizon label variance computed host-side from existing per-horizon
labels in stg_labels (layout [K, B, N_HORIZONS] row-major; typical K=32,
B=1 -> 32 floats per horizon -> trivial host reduction).
- Wiener-α EMA with α = 0.4 floor + first-observation bootstrap per
pearl_wiener_alpha_floor_for_nonstationary +
pearl_first_observation_bootstrap.
FIX 2 — Task 10 Option B closure (block-diagonal heads via grad mask):
- Two new kernels: heads_w_skip_mask_init_kernel (one-shot at transition,
builds [N_HORIZONS × HIDDEN_DIM] = 640-float mask + zeros off-bucket
heads_w_skip in place) and heads_w_skip_grad_mask_apply_kernel
(per-step Phase 2, multiplies grad_heads_w_skip by mask elementwise
before opt_heads_w_skip.step).
- Achieves block-diagonal heads_w_skip behavior WITHOUT touching the
existing GRN kernel: off-bucket positions stay 0 throughout training
because gradient is masked out, so Adam never updates them.
- Existing GRN dispatch consumes heads_w_skip_d (full 640 floats) as
today; off-bucket entries are always 0, so the dispatch is
mathematically equivalent to a compact-only read.
Together these fixes restore the spec-mandated per-horizon differentiation
chain across all 3 mechanisms (CfC.tau bucketing + block-diagonal heads
+ per-bucket LR with label-variance anchor) before Smoke 2 fires the
mean_run_len ratio gate.
GPU oracle tests: 13 total (11 + heads_w_skip_mask_init + heads_w_skip_grad_mask_apply).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Per spec §3.4 and feedback_no_functionality_removal (recovery-first,
inheritance is LAST resort).
New device kernels in bucket_transition_kernels.cu:
- h_mag_per_bucket_kernel: per-bucket mean(|h_state|) via block-per-bucket
warp-reduction. Launch grid=(N_HORIZONS,1,1), block=(32,1,1).
No atomicAdd (block-tree reduce over 32 lanes).
- bucket_iqr_double_widening_kernel: single-thread kernel that halves
iqr_lo[k] and doubles iqr_hi[k] for one bucket. Recovery Attempt 1's
actual implementation.
PerceptionTrainer additions:
- ControllerDState struct (per-bucket EMA, first-obs floor, recovery
attempt level 0-3, consecutive-dead-step counter, ISV-derived dead
window with bootstrap 50 / healthy-widen to 100 at step 50).
- phase2_step_count counter (Phase 2 only).
- h_mag_per_bucket_d device buffer + mapped-pinned shadow.
- Cached h_mag_per_bucket_fn + bucket_iqr_double_widening_fn handles.
Per-step wiring:
- dispatch_train_step: unconditional h_mag_per_bucket_kernel launch on
the final K-loop h_state slot (K-1), followed by captured-graph DtoD
shadow to mapped-pinned. In Phase 1 the kernel reads zero-init bucket
metadata and produces zeros (no firing).
- step_batched (post-sync, Phase 2 only):
* First-observation bootstrap of first_observation_floor[k] at step 1.
* Wiener-α=0.4 EMA update of h_mag_ema[k]
(pearl_wiener_alpha_floor_for_nonstationary).
* Dead-threshold = first_observation_floor[k] * 1/e per spec §3.4.
* Consecutive-dead counter; recovery cascade fires at dead_window_k.
Recovery cascade per spec §3.4:
- Attempt 1 (widen): WIRED. Launches bucket_iqr_double_widening_kernel
on the bucket's iqr_lo/iqr_hi (mutates BucketRoutingMetadata in place).
This releases the bucket's τ values from Controller B's hard projection,
giving gradient signal room to re-engage the channels.
- Attempt 2 (channel swap): LOG-ONLY STUB. Emits tracing::warn! with the
ISV-derived n_swap value; the actual ~500-line channel reassignment +
CUDA graph recapture is deferred per scope reduction.
- Attempt 3 (inheritance): LOG-ONLY STUB. Emits tracing::warn! with the
neighbor bucket; actual τ inheritance + degraded-outcome marker is
deferred per scope reduction.
Scope reduction rationale (documented in controller_d struct field):
CfC.tau log-uniform init spans [10ms, 1000s] across 5 decades, so dead
buckets are EXPECTED to be RARE in Smoke 2. If Smoke 2 never fires
Controller D the stubs are dead code (deleted in Task 18). If Recovery 1
suffices (most likely path), no further work needed. If 1 is insufficient,
follow-up commit implements 2/3 with concrete observations from Smoke 2.
Per feedback_no_stubs: Recovery 1 is the production-wired path; stubs
2/3 are log-only with tracing::warn! so any firing surfaces in logs.
The log boundary is an explicit scope decision, not a deferral.
ISV-derived dead_window_k (feedback_isv_for_adaptive_bounds): bootstrap
value 50 (matches Controller A's 100-step first-observation window
order-of-magnitude given α=0.4 EMA settling ≈ 2-3 steps). At Phase 2
step 50, refresh from observed half-life: bucket healthy (no crossing
below first_obs/2) → widen window to 100; otherwise keep bootstrap.
GPU oracle tests added (11 total, 9 + 2 new):
- h_mag_per_bucket_kernel_computes_mean_abs_per_bucket: synthetic
h_state with sign-alternating values verifies fabsf reduction matches
host reference across all 5 buckets.
- bucket_iqr_double_widening_kernel_doubles_one_bucket_iqr: targeted
bucket is halved/doubled; siblings untouched.
Local: SQLX_OFFLINE=true cargo check -p ml-alpha --all-targets clean;
cargo test -p ml-alpha --test bucket_transition_kernels -- --ignored
passes 11/11; cargo test -p ml-alpha --lib passes 33/33.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Per spec §3.2 + §3.3 and pearl_adam_normalizes_loss_weights.
Controller B (post-Adam τ projection):
- tau_clamp_kernel added to bucket_transition_kernels.cu
- Each channel's τ clamped to its bucket's IQR_widened range after every
Adam step on tau_all_d. Bypasses Adam's m/sqrt(v) normalization
(loss-weight modulation would have been a no-op per the pearl).
Controller C (per-bucket LR multiplier on heads_w_skip):
- heads_lr_multiplier_scale_kernel added.
- Per-step: snapshot heads_w_skip_d (DtoD) before opt_heads_w_skip.step,
run Adam, then rescale the per-horizon delta by
lr_mult_k = 1.0 / (1.0 + smoothness_ratio_k).
- smoothness_ratio_k = max(0, jitter_ema_k - target_jitter_k) /
max(target_jitter_k, target_jitter_k / 16)
with the ε_floor max-with-floor pattern per
pearl_blend_formulas_must_have_permanent_floor.
- target_jitter_k sentinel-bootstrapped from first non-zero raw
smoothness loss per pearl_first_observation_bootstrap.
- jitter EMA shadowed via mapped-pinned smoothness_jitter_ema_host_d
(DtoD captured in graph; host reads after end-of-step sync); host
writes lr_mult_per_horizon_staging (mapped-pinned); next step's
captured-graph rescale kernel reads via dev_ptr.
Scope reduction (Task 11): Controller C scales ONLY heads_w_skip_d
per-horizon (N_HORIZONS × HIDDEN_DIM = 640 contiguous floats, per-horizon
stride HIDDEN_DIM). Other heads weights (w1, w2, w_gate, w_main) stay
unscaled because they're entangled in the GRN gate/main computation and
rescaling them would require additional per-bucket scoping out of Task 11
scope. Revisit if Smoke 2 fails per feedback_no_quickfixes.
Both controllers active Phase 2 only; no-op in Phase 1.
GPU oracle tests: 9 total (7 existing + tau_clamp + lr_multiplier_scale).
All 33 ml-alpha lib tests pass.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Per spec §2.1 + §5.4 point 1 and Task 10 of the plan.
Adds binding `cfc_step_per_branch_fwd_gpu` in cfc/step.rs that wraps
the fused per-branch kernel with the spec-mandated launch config:
grid=(B, N_HORIZONS=5, 1), block=(MAX_BUCKET_DIM=28, 1, 1) uniform
predicate, cooperative staging of x_local + h_old_local (2 × HIDDEN_DIM
floats). Also adds `cfc_step_per_branch_bwd_gpu` binding for Task 11.
CfcTrunk loads two new cubins (cfc_step_per_branch + heads_block_diagonal)
and caches three function handles. `heads_w_skip_compact_d` (from Task 9)
remains allocated as Phase-2-prepared metadata; full heads consumption
deferred to a follow-up task (Phase 2 heads dispatch needs GRN
integration, which is out of Task 10 scope per
`feedback_no_partial_refactor`).
In `dispatch_train_step`, the CfC dispatch branches on TrainingPhase:
- Phase1Warmup: existing `cfc_step_batched` (unchanged)
- Phase2Routed: `cfc_step_per_branch_fwd` consuming bucket-routed
`tau_all_d` + `bucket_channel_offset_d` + `bucket_dim_k_d`
GRN heads dispatch stays unchanged for both phases per the Task 10
scope adjustment (Option B in the task brief). The `heads_w_skip` storage
remains 640 floats; the compact 128-float `heads_w_skip_compact_d` is
metadata-ready but not yet consumed.
Workspace cargo check clean (ml-alpha lib + trainer compile; only
pre-existing cupti/sp15/gpu_per_integration test errors remain).
ml-alpha lib tests: 33 passed.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Per spec §3.1, §2.3. Controller A computes tau-change EMA with
first-observation bootstrap + Wiener-α floor 0.4 (co-adapting closed
loop); trigger when EMA < max(noise_floor/4, noise_floor/16); ISV-derived
hard cap from observed MAD/median dispersion in first 100 steps.
execute_transition() dispatches the 5 device kernels in order:
tau_sort → bucket_assign → bucket_iqr → tau_reorder → heads_compact.
All-on-device per feedback_no_htod_htoh_only_mapped_pinned (one-time
static-constant upload at transition is allowed off-hot-path).
Returns raw Q1/Q3 in BucketRoutingMetadata.bucket_tau_iqr_{lo,hi}_d;
slack_factor=sqrt(Q3/Q1) widening kernel applied by trainer in Task 9.
3 CPU-only unit tests verify Controller A's first-obs bootstrap,
threshold-trigger, and ISV-cap-trigger semantics.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Per spec §4.1 + Q-3-seed-baseline. Both baseline Argo workflows
succeeded:
- alpha-perception-baseline-seed16963 (39min)
- alpha-perception-baseline-seed16964 (26min)
Precise AUC values deferred: train pod stdout is in MinIO archive
(xl.meta format, needs mc auth). The on-disk alpha_train_summary.json
on /feature-cache/alpha-perception-runs/f22f3f948/ corresponds to
seed 16963 (last to finish). Task 15 implementer retrieves via debug
pod mount of feature-cache PVC.
Known data point: seed 16962 = 0.7454 best_auc_h6000.
The Smoke 1 invariant gate threshold is signal-derived per
pearl_tests_must_prove_not_lock_observations — formulas captured in
the JSON for Task 15 to apply once values are retrieved.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Per spec §5.4 point 2. heads_w_skip storage 5× reduced (640→128 floats).
Each horizon reads only its bucket via offset lookup. Block-tree reduction
(padded to 32 lanes for clean halving stride), no atomicAdd.
GPU oracle test matches naive per-horizon computation (tol 1e-5) on sm_86.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Per spec §5.4 point 1. Single fused kernel: grid=(B, N_HORIZONS, 1),
block=(MAX_BUCKET_DIM=28, 1, 1) uniform predicate. Shared x_local +
h_old_local cooperative-staged into shared memory once per block per
pearl_cooperative_staging_eliminates_redundant_reads. No atomicAdd.
No host branches.
Forward GPU oracle: matches naive per-channel CfC math (tol 1e-4) on
sm_86. Backward smoke: kernel loads + writes non-NaN gradients.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Per spec §2.5 and feedback_no_partial_refactor. Removes:
- LoaderMode enum (single-variant after Sequential removal → deleted entirely)
- next_sequence_sequential + sequential_file_idx + sequential_anchor_idx + last_call_was_file_boundary + is_file_boundary
- last_seen_file_boundary, train_graph_boundary_state fields
- notify_file_boundary method + boundary-state graph recapture logic
- need_attn_pool_bootstrap match — always true now (random mode always bootstraps)
- --loader-mode CLI flag + notify_file_boundary() call
- loader-mode Argo template param
Additional consumers migrated atomically (beyond the 4 files listed in
the plan): ml-alpha/tests/perception_overfit.rs,
ml-alpha/tests/multi_horizon_loader.rs, ml-backtesting/src/harness.rs,
ml-backtesting/tests/trainer_parity.rs,
ml-backtesting/tests/ring3_replay.rs — all referenced LoaderMode or the
removed config fields.
Workspace builds clean at this commit (pre-existing cudarc-cupti example
and ml-crate test errors are unrelated to MTER removal — they fail at
HEAD too). New per-horizon CfC arch lands in subsequent tasks of the
same plan.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Spec: docs/superpowers/specs/2026-05-21-per-horizon-cfc-inference-design.md
Plan: docs/superpowers/plans/2026-05-21-per-horizon-cfc-inference-plan.md
Spec went through 2 critical-review passes (32 total findings, all resolved).
Bucketing source: CfC.tau (per-channel, trained, log-uniform init at HIDDEN_DIM=128).
Atomic refactor reverts MTER scaffolding. 5 ISV-driven controllers. All-on-device
transition (no bulk DtoH). Single fused per-branch kernel. Compact ragged heads_w_skip.
Validated via 3 Argo smokes (training stability, CRT.diag inference differentiation,
fxt-backtest end-to-end).
Also committing historical record: superseded MTER spec/plan, intervention A
plan, ISV λ controller spec/plan, GPU log ring spec/plan. Per spec §5.3 these
stay as audit trail; not in build path.
Plan has 18 tasks, full TDD with bite-sized steps, ready for
subagent-driven-development execution.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Task 1's attn_pool gate (commit 7dd953aa) was inside the captured-graph
region, so the host branch's decision was frozen at step-2 capture time.
In Sequential mode this would lock the captured graph to whichever
boundary state happened to hold at step 2, breaking subsequent file-
boundary detection — violating pearl_no_host_branches_in_captured_graph.
Fix: track the boundary state at capture time as a new
`train_graph_boundary_state: Option<bool>` field. On each step_batched,
if the current `last_seen_file_boundary` diverges from the captured
state, drop the cached graph and re-capture on the next step.
- Random mode: gate is unconditionally true; state never diverges;
zero recaptures.
- Sequential mode: ~8 recaptures per epoch (one per file boundary).
Each recapture costs the same as the original step-2 capture (~50ms).
Negligible.
Verified clean cargo check -p ml-alpha --all-targets.
Loader can now serve sequences in temporal order (Sequential mode);
when active, the trainer skips attn_pool reset at non-file-boundary
sequence boundaries (next commit will use this for stateful CfC + MTER).
Random mode preserved as default; ZERO behavioral change for existing
runs. Phased commit 1/8 of intervention B per
docs/superpowers/specs/2026-05-21-crt-train-intervention-b-multi-timescale-readout.md
Local trace at base_lambda=0.1 showed the linear ratio
{1, 0.3, 0.1, 0.03, 0.005} = HORIZONS[0]/HORIZONS[h] gave h6000 a
200x stronger target-undershoot signal than h30. The controller
slammed h6000 with smoothness gradient (peak λ[h6000]≈60) while h30
saw nearly none. h6000 val_auc collapsed to 0.514 (near-random)
while h100/h300 improved.
Replace with sqrt(HORIZONS[0]/HORIZONS[h]) = {1, 0.548, 0.316, 0.173,
0.0707}. Caps the differential at ~14x. Verified locally at base=0.1:
- h6000 val_auc recovered to 0.599 (vs 0.514 collapse, vs 0.621 no-smooth)
- jitter ratio h6000/h30 = 0.51 (meaningful differentiation, was 0.84 unforced)
- λ[h6000] equilibrium = 0.7-3 (was 4-60 with linear ratio at same base)
The design target (h6000 jitter = 7% of h30) is less aggressive than
linear (was 0.5%) but produces a tractable training equilibrium that
preserves h6000 predictive capacity. Both 500-step local AND the full
40k-step L40S retrain will tell us where it lands at scale.
The drain task previously used tokio::spawn + tokio::time::sleep,
gated by Handle::try_current().is_ok(). alpha_train's main() is
synchronous (not #[tokio::main]) so Handle::try_current() returned
Err and the drain task was silently skipped — the trace file was
never created.
Refactor to std::thread::spawn + std::thread::sleep. No runtime
required; spawn is unconditional; Drop joins synchronously instead
of aborting; BufWriter flushes on every drain cycle so even short
runs (< 1 sec) capture their tail. Existing gpu_log_ring_invariants
tests migrated from #[tokio::test] back to plain #[test].
Also resolves an exit-1 issue on alpha_train shutdown — likely a
tokio task abort panicking through the synchronous Drop path.
Local smoke (alpha_train --kernel-step-trace ...):
- EXIT=0
- step_trace.jsonl: 1497 lines (~500 steps × 3 smoothness records)
- valid JSONL, schema matches smoothness_controller emission contract
The compile-time feature kernel-step-trace gates inclusion of the ring
code. NEW: --kernel-step-trace <PATH> CLI flag on alpha_train gates
RUNTIME activation:
- Path provided -> ring allocated, drain spawns, JSONL records written
to <PATH> (truncated). One record per kernel emission.
- Path omitted -> ring not allocated, drain not spawned, zero overhead
even with the feature compiled in.
JSONL schema: {"step": u32, "kid": u8, "kname": str, "rt": u8,
"rt_name": str, "payload": {field: f32, ...}}. The payload field names
come from the existing per-(kid, rt) decoders in gpu_log.rs (ported to
serde_json::json! in decode_to_json).
Trainer ring fields are now Option<_>; populated only when the runtime
trace path is Some. Tick kernel, step-counter shadow, smoothness
controller pointer-passing, and Drop all become path-conditional.
The tracing::info!-based drain variant is removed (file-writer is
strictly more useful; tests migrated to assert JSONL file contents).
Argo template:
- New parameter kernel-step-trace-enable (build-time feature opt-in)
- New parameter kernel-step-trace-path (runtime CLI value)
- ensure-binary cache-busts when feature toggles
- training step passes --kernel-step-trace flag conditionally
Per feedback_no_feature_flags: compile-time gate retained because the
ring carries real memory cost (~2 MiB pinned) and per-step write overhead;
the specific name kernel-step-trace narrows scope to this mechanism.
Runtime gate is Option<PathBuf>, not a boolean enable_*.
Per user direction: feature-gating is appropriate for per-step diagnostic
logging (real perf/memory cost) but the name should be specific to the
mechanism, not a generic "diag-log" catch-all. kernel-step-trace
describes what the feature provides — per-step records written to the
ring by kernels.
Pure rename, no behavioral changes. Touched: Cargo.toml feature decl,
lib.rs cfg attribute, perception.rs (~25 cfg attributes including
not(feature) pairs), gpu_log.rs + smoothness_lambda_controller.cu doc
comments, gpu_log_ring_invariants.rs file-level cfg + ignore-attribute
text.
Atomic refactor wiring the GPU log ring's first producer end-to-end:
- cuda/gpu_log_helpers.cuh (new): extract LogHeader/LogRecord/LogRing
struct definitions + the log_record() device __forceinline__ helper
into a shared header. Single source of truth — other kernels include
this rather than redefining the structs.
- cuda/gpu_log_ring.cu: refactor to include gpu_log_helpers.cuh; retain
only the gpu_log_tick kernel.
- cuda/smoothness_lambda_controller.cu: include gpu_log_helpers.cuh, add
trailing (LogRing*, const int*) args, capture pre-EMA jitter +
post-EMA jitter + target + excess_ratio into shared mem, and emit
three records per call (RT_INPUT, RT_STATE, RT_OUTPUT) gated on
non-null ring pointer.
- trainer/perception.rs: feature-gated (cuda-diag-log) ring allocation +
step counter (mapped-pinned host shadow), gpu_log_tick launch FIRST
inside the captured graph, two new pointer args on the smoothness
controller launch (null when feature off), step-counter DtoD shadow
alongside the other telemetry shadows, drain task spawn (skipped when
no tokio runtime — sync tests still construct the trainer), and a
feature-gated Drop impl that aborts the drain task on shutdown.
- tests/smoothness_lambda_controller_invariants.rs: pass null pointers
for the two new kernel args; the kernel's nullptr guard preserves
pre-existing behaviour.
- build.rs: rerun-if-changed on cuda/gpu_log_ids.h and
cuda/gpu_log_helpers.cuh so header edits trigger cubin rebuilds.
Verified: cargo build / check --all-targets clean both with and without
the cuda-diag-log feature; 4/4 smoothness controller invariant tests
pass; 9/9 perception_overfit integration tests pass under the feature.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Task 5 of the GPU log ring plan. Adds `gpu_log` module behind the
`cuda-diag-log` feature gate:
- `LogRecord` (#[repr(C)], 64 B) mirroring the C-side layout in
`gpu_log_ring.cu`; const compile-time size assertion catches drift.
- `LogRing::alloc()` allocates and zero-initialises a 32k-slot
mapped-pinned ring (2 MiB pinned).
- `alloc_step_counter()` allocates the device-side i32 step counter.
- Decoder dispatch table with per-(kid, rt) decoders for
`KID_SMOOTHNESS_CONTROLLER` × {RT_INPUT, RT_STATE, RT_OUTPUT}
emitting structured tracing::info! events.
- `spawn_drain_task()` polls a mapped-pinned step-counter shadow at
500 ms cadence, walks new slots, validates magic + step, dispatches
to decoders. Emits tracing::warn! when the drainer falls behind the
ring window (no silent record drops).
Kernel ID + record-type constants are a manual mirror of
`crates/ml-alpha/cuda/gpu_log_ids.h`; drift is caught by the const
size assertion at compile time and (subsequently) by build.rs.
Task 5 added smoothness_base_lambda to PerceptionTrainerConfig but only
updated the struct definition + Default impl. Eight test sites in
perception_overfit.rs and one site in alpha_train.rs construct the
config via explicit field list (no ..Default::default() spread), so
they broke with E0063 missing-field errors under cargo check
--all-targets.
Per feedback_no_partial_refactor.md: when a contract changes, every
consumer migrates atomically. This commit adds smoothness_base_lambda:
0.0 to all 9 sites so the workspace compiles cleanly. The
alpha_train.rs value of 0.0 is a placeholder — Task 7 will replace it
with cli.smoothness_base_lambda once the CLI flag is added.
Code-quality reviewer flagged two non-blocking polish items in the
output_smoothness kernel's backward pass:
- I1: thread-divergent if-gated loads contradicted the file's stated
branchless-design intent. Fixed via clamp-then-load: at k=0 / k=K-1
read self for the spurious neighbour; lm/rm multipliers zero the
contribution. Every lane executes both loads; no divergence on
boundaries.
- I2: the "Skip work when factor == 0" comment described a non-feature
(no such branch existed). Removed.
Behaviour-equivalent change — same loss and same gradient at every
position; the change only affects warp-execution patterns at the K=0
and K=K-1 strips.