Commit Graph

5401 Commits

Author SHA1 Message Date
jgrusewski
ff9207e7bb feat(per-horizon-cfc): plumb kernel-step-trace into fxt-backtest inference path
Mirrors the alpha_train CLI flag (--kernel-step-trace <path>) into
fxt-backtest. The PerceptionTrainerConfig already accepts the path;
this commit just exposes the CLI surface and propagates through:

  fxt-backtest run --kernel-step-trace <path>
  -> BacktestHarnessConfig.kernel_step_trace_path
  -> PerceptionTrainerConfig.kernel_step_trace_path

Sweep YAML gains a `kernel_step_trace: Option<PathBuf>` base field.
Argo lob-backtest-sweep-template gains a `kernel-step-trace` workflow
parameter (default disabled; when set, --features kernel-step-trace
is passed to cargo build AND --kernel-step-trace to fxt-backtest).

Gated behind the `kernel-step-trace` Cargo feature (matching alpha_train).
When disabled (default), zero overhead.

Enables per-step JSONL diagnostic emission from inference kernels --
needed for Smoke 2 deep-dive after sampling histograms (in parallel
investigation) localize the per-horizon dynamics bottleneck.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-21 22:33:44 +02:00
jgrusewski
d2d34b6d4e diag(per-horizon-cfc): per-horizon alpha_probs sampling for Smoke 2 investigation
Adds --per-horizon-logit-sampling-stride CLI flag to fxt-backtest run/sweep
+ harness instrumentation that DtoHs alpha_probs every N events and
accumulates per-horizon mean/stddev/min/max + 10-bin histogram.

Emitted at end of CRT.diag block as:
  per_horizon_logit_diag h{N}: count=... mean=... std=... min=... max=... hist=[...]

Goal: distinguish whether Smoke 2's uniform mean_run_len 1.04x across
horizons is caused by uniform per-horizon LOGITS (-> Task 11 scope
incomplete; heads_w1/w2/gate/main need block-diagonal too) or by
DIFFERENTIATED logits with collapsed downstream (-> different fix).

Default stride=None (disabled, zero overhead). With stride=1000 on
2M events, ~2000 DtoD stops x ~50us = ~100ms total overhead per run.
Per feedback_no_htod_htoh_only_mapped_pinned: uses mapped-pinned buffer
(single 5xf32 allocation reused on every sampling tick).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-21 22:20:42 +02:00
jgrusewski
8b2ac1e577 fix(per-horizon-cfc): ALPHA — skip reorder, channels_in_bucket lookup, Adam moment masking
URGENT correctness fix surfaced by checkpoint deep-dive after Smoke 2 failed
the WIN gate (mean_run_len ratio = 1.0× vs target ≥10×; all 5 horizons
uniformly ~2.4 events).

THREE INDEXING BUGS identified:

1. tau_reorder produces bucket-grouped tau_all_d, but Controller B's
   tau_clamp_kernel reads bucket_id_per_channel[c] where c is the
   POSITION in the reordered buffer (not the original channel index)
   → wrong bucket-IQR lookup → τ never constrained → buckets collapse
   (deep-dive showed all 5 buckets had nearly identical τ ranges
   [0.07, 74] except bucket 4 reaching 878).

2. heads_w_skip grad mask ran but Adam (m, v) momentum from BEFORE the
   transition re-introduced gradient signal across the transition →
   off-bucket positions stayed nonzero throughout training
   (deep-dive: 512/512 = 100% of off-bucket positions nonzero in
   trunk_best_h6000.bin).

3. per-branch CfC kernel read w_in[c * HIDDEN_DIM + k] with c =
   bucket-grouped position, but W_in rows are indexed by ORIGINAL
   channel → kernel read wrong rows for each output → outputs were
   essentially random per-channel.

ALPHA FIX: skip the reorder entirely, use bucket-filter throughout:
- Removed tau_reorder_kernel; cfc.tau_d stays in original-channel layout.
- Added channels_in_bucket_kernel that populates a
  [N_HORIZONS × MAX_BUCKET_DIM] lookup (original channel index per
  (bucket, within-bucket-position)).
- Per-branch CfC fwd+bwd now reads channels_in_bucket[branch][tid]
  → original_c, then uses original_c for w_in/w_rec indexing. All
  weights stay in original layout consistently.
- Controller B's tau_clamp_kernel now correctly operates on original-
  channel cfc.tau_d with bucket_id_per_channel[c] lookup (no position-
  vs-channel confusion).
- Added zero_off_bucket_kernel + three-layer defense for heads_w_skip
  block-diagonal invariant:
    (a) At transition: zero off-bucket params + zero Adam (m, v)
        moments via opt_heads_w_skip.m_mut() / v_mut() accessors.
    (b) Per-step: heads_w_skip_grad_mask_apply_kernel zeros off-bucket
        gradients before Adam step (unchanged from prior follow-up).
    (c) Per-step: zero_off_bucket_kernel zeros off-bucket params after
        Adam step, catching any drift from Adam's ε denominator or
        weight decay.

New AdamW::m_mut()/v_mut() accessors enable the projection at transition.

GPU oracle tests: 19 total (16 in bucket_transition + 3 in cfc_step_per_branch).
New tests verify:
- channels_in_bucket_kernel correctness under non-contiguous bucket assignment
- zero_off_bucket maintains invariant after many mock Adam steps
- fwd kernel writes only to bucket-assigned channels under arbitrary mapping

Per `feedback_no_partial_refactor`: all 3 indexing bugs + Adam momentum
defense land in one commit.

ml-alpha lib: 33 passed.
GPU oracle tests on RTX 3050 sm_86: 19 passed.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-21 20:43:51 +02:00
jgrusewski
bf4d0c899f fix(per-horizon-cfc): sync trunk bucket metadata from transition output
URGENT correctness fix surfaced by Task 13 implementer.

Bug: training-time Phase 2 dispatch sites (cfc_step_per_branch_fwd at
~3139, h_mag_per_bucket for Controller D at ~3211-3212) read
self.trunk.bucket_*_d which are zero-initialized at trunk construction
and only populated via load_checkpoint. The Phase 1→2 transition
populates a SEPARATE BucketRoutingMetadata owned by the trainer
(self.bucket_routing_metadata) but never syncs it into the trunk fields.
With zero-init bucket_dim_k, the per-branch kernel's uniform predicate
`tid >= bucket_dim_k[branch]` early-returns ALL threads -> Phase 2
silently no-ops during training, Smoke 1 fails before producing signal.

Fix: at transition, DtoD-copy metadata buffers into trunk fields BEFORE
the `Some(metadata)` move (so `metadata.*` borrows remain valid):
- metadata.bucket_channel_offset_d -> trunk.bucket_channel_offset_d
- metadata.bucket_dim_k_d          -> trunk.bucket_dim_k_d
- metadata.bucket_id_per_channel_d -> trunk.bucket_id_per_channel_d
- metadata.heads_w_skip_offset_d   -> trunk.heads_w_skip_offset_d

Trunk fields remain the single source of truth for both training-time
dispatch AND checkpoint serialization. Total sync cost: 4 DtoD memcpys,
~256 bytes total at the one-shot transition (off the hot path).

Verification:
- `cargo check -p ml-alpha --all-targets`: clean
- `cargo test -p ml-alpha --lib`: 33 passed, 0 failed
- 3 read sites (perception.rs:3139, 3211-3212, 4725-4726) verified
  unchanged; all 4 new memcpy_dtod_async calls bracketed in the
  transition block (perception.rs:2224, 2234, 2244, 2254).
- Block-diagonal heads grad-mask init (step 7) continues to read
  `metadata.bucket_id_per_channel_d` via `as_ref()` after the move;
  ordering preserved per pearl_canary_input_freshness_launch_order.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-21 19:16:27 +02:00
jgrusewski
ed8b53b474 feat(per-horizon-cfc): forward_step_into uses Phase 2 fused dispatch
Per spec §2.4 and Task 13 of the plan.

Replace the single-CfC `cfc_step_batched` dispatch in `forward_step_into`
with `cfc_step_per_branch_fwd_gpu`. Production checkpoints have Phase 2
routing frozen at the training-time Phase 1→2 transition, so inference
unconditionally takes the per-branch path. The fused kernel reads
`bucket_channel_offset_d` / `bucket_dim_k_d` from the trunk; in
deployment these are populated via `CfcTrunk::load_checkpoint`. For
freshly-constructed trainers without a checkpoint load, the trunk fields
are zero — the kernel's uniform predicate (`tid >= bucket_dim_k`) then
early-returns every thread and h_new is left untouched. That matches the
"Phase 2 only at inference" contract.

Heads dispatch unchanged: the existing GRN kernel reads from
`trunk.heads_w_skip_d` which has off-bucket entries zeroed at the
transition (sparsification per the prior block-diagonal-grad-mask
follow-up commit). Mathematically the full-buffer read is equivalent to
a compact-only per-bucket read; no inference-time kernel change required.

forward_step_golden's convergence test is now architecturally
incompatible with the new dispatch — `forward_only` runs Phase 1
single-CfC math while `forward_step_into` requires populated Phase 2
routing. Annotated `#[ignore]` with a clear divergence note; the
deterministic + reset semantics tests remain valid invariants per
`pearl_training_smoothness_does_not_transfer_to_inference`.

cargo check workspace clean (excluding pre-existing unrelated cupti +
insert_batch errors in vendor/cudarc and ml/tests). ml-alpha + ml-
backtesting lib tests pass (33 + 33).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-21 19:10:22 +02:00
jgrusewski
c7bb79a625 fix(per-horizon-cfc): label-variance anchor + block-diagonal heads grad mask
Two follow-up fixes for Smoke 2 readiness, atomic per feedback_no_partial_refactor:

FIX 1 — Task 11 target_jitter_k anchor (spec §3.3 compliance):
- Previously: target_jitter_k = first non-zero raw_per_h[k] (model-state-anchored)
- Now: target_jitter_k = sqrt(realised_label_variance_k) (label-distribution-anchored)
- Per-horizon label variance computed host-side from existing per-horizon
  labels in stg_labels (layout [K, B, N_HORIZONS] row-major; typical K=32,
  B=1 -> 32 floats per horizon -> trivial host reduction).
- Wiener-α EMA with α = 0.4 floor + first-observation bootstrap per
  pearl_wiener_alpha_floor_for_nonstationary +
  pearl_first_observation_bootstrap.

FIX 2 — Task 10 Option B closure (block-diagonal heads via grad mask):
- Two new kernels: heads_w_skip_mask_init_kernel (one-shot at transition,
  builds [N_HORIZONS × HIDDEN_DIM] = 640-float mask + zeros off-bucket
  heads_w_skip in place) and heads_w_skip_grad_mask_apply_kernel
  (per-step Phase 2, multiplies grad_heads_w_skip by mask elementwise
  before opt_heads_w_skip.step).
- Achieves block-diagonal heads_w_skip behavior WITHOUT touching the
  existing GRN kernel: off-bucket positions stay 0 throughout training
  because gradient is masked out, so Adam never updates them.
- Existing GRN dispatch consumes heads_w_skip_d (full 640 floats) as
  today; off-bucket entries are always 0, so the dispatch is
  mathematically equivalent to a compact-only read.

Together these fixes restore the spec-mandated per-horizon differentiation
chain across all 3 mechanisms (CfC.tau bucketing + block-diagonal heads
+ per-bucket LR with label-variance anchor) before Smoke 2 fires the
mean_run_len ratio gate.

GPU oracle tests: 13 total (11 + heads_w_skip_mask_init + heads_w_skip_grad_mask_apply).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-21 18:59:28 +02:00
jgrusewski
d435d8f270 feat(per-horizon-cfc): wire Controller D dead-bucket recovery cascade
Per spec §3.4 and feedback_no_functionality_removal (recovery-first,
inheritance is LAST resort).

New device kernels in bucket_transition_kernels.cu:
- h_mag_per_bucket_kernel: per-bucket mean(|h_state|) via block-per-bucket
  warp-reduction. Launch grid=(N_HORIZONS,1,1), block=(32,1,1).
  No atomicAdd (block-tree reduce over 32 lanes).
- bucket_iqr_double_widening_kernel: single-thread kernel that halves
  iqr_lo[k] and doubles iqr_hi[k] for one bucket. Recovery Attempt 1's
  actual implementation.

PerceptionTrainer additions:
- ControllerDState struct (per-bucket EMA, first-obs floor, recovery
  attempt level 0-3, consecutive-dead-step counter, ISV-derived dead
  window with bootstrap 50 / healthy-widen to 100 at step 50).
- phase2_step_count counter (Phase 2 only).
- h_mag_per_bucket_d device buffer + mapped-pinned shadow.
- Cached h_mag_per_bucket_fn + bucket_iqr_double_widening_fn handles.

Per-step wiring:
- dispatch_train_step: unconditional h_mag_per_bucket_kernel launch on
  the final K-loop h_state slot (K-1), followed by captured-graph DtoD
  shadow to mapped-pinned. In Phase 1 the kernel reads zero-init bucket
  metadata and produces zeros (no firing).
- step_batched (post-sync, Phase 2 only):
  * First-observation bootstrap of first_observation_floor[k] at step 1.
  * Wiener-α=0.4 EMA update of h_mag_ema[k]
    (pearl_wiener_alpha_floor_for_nonstationary).
  * Dead-threshold = first_observation_floor[k] * 1/e per spec §3.4.
  * Consecutive-dead counter; recovery cascade fires at dead_window_k.

Recovery cascade per spec §3.4:
- Attempt 1 (widen): WIRED. Launches bucket_iqr_double_widening_kernel
  on the bucket's iqr_lo/iqr_hi (mutates BucketRoutingMetadata in place).
  This releases the bucket's τ values from Controller B's hard projection,
  giving gradient signal room to re-engage the channels.
- Attempt 2 (channel swap): LOG-ONLY STUB. Emits tracing::warn! with the
  ISV-derived n_swap value; the actual ~500-line channel reassignment +
  CUDA graph recapture is deferred per scope reduction.
- Attempt 3 (inheritance): LOG-ONLY STUB. Emits tracing::warn! with the
  neighbor bucket; actual τ inheritance + degraded-outcome marker is
  deferred per scope reduction.

Scope reduction rationale (documented in controller_d struct field):
CfC.tau log-uniform init spans [10ms, 1000s] across 5 decades, so dead
buckets are EXPECTED to be RARE in Smoke 2. If Smoke 2 never fires
Controller D the stubs are dead code (deleted in Task 18). If Recovery 1
suffices (most likely path), no further work needed. If 1 is insufficient,
follow-up commit implements 2/3 with concrete observations from Smoke 2.

Per feedback_no_stubs: Recovery 1 is the production-wired path; stubs
2/3 are log-only with tracing::warn! so any firing surfaces in logs.
The log boundary is an explicit scope decision, not a deferral.

ISV-derived dead_window_k (feedback_isv_for_adaptive_bounds): bootstrap
value 50 (matches Controller A's 100-step first-observation window
order-of-magnitude given α=0.4 EMA settling ≈ 2-3 steps). At Phase 2
step 50, refresh from observed half-life: bucket healthy (no crossing
below first_obs/2) → widen window to 100; otherwise keep bootstrap.

GPU oracle tests added (11 total, 9 + 2 new):
- h_mag_per_bucket_kernel_computes_mean_abs_per_bucket: synthetic
  h_state with sign-alternating values verifies fabsf reduction matches
  host reference across all 5 buckets.
- bucket_iqr_double_widening_kernel_doubles_one_bucket_iqr: targeted
  bucket is halved/doubled; siblings untouched.

Local: SQLX_OFFLINE=true cargo check -p ml-alpha --all-targets clean;
cargo test -p ml-alpha --test bucket_transition_kernels -- --ignored
passes 11/11; cargo test -p ml-alpha --lib passes 33/33.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-21 17:53:52 +02:00
jgrusewski
dc76d205eb feat(per-horizon-cfc): wire Controllers B + C (Adam-escape mechanisms)
Per spec §3.2 + §3.3 and pearl_adam_normalizes_loss_weights.

Controller B (post-Adam τ projection):
- tau_clamp_kernel added to bucket_transition_kernels.cu
- Each channel's τ clamped to its bucket's IQR_widened range after every
  Adam step on tau_all_d. Bypasses Adam's m/sqrt(v) normalization
  (loss-weight modulation would have been a no-op per the pearl).

Controller C (per-bucket LR multiplier on heads_w_skip):
- heads_lr_multiplier_scale_kernel added.
- Per-step: snapshot heads_w_skip_d (DtoD) before opt_heads_w_skip.step,
  run Adam, then rescale the per-horizon delta by
  lr_mult_k = 1.0 / (1.0 + smoothness_ratio_k).
- smoothness_ratio_k = max(0, jitter_ema_k - target_jitter_k) /
                       max(target_jitter_k, target_jitter_k / 16)
  with the ε_floor max-with-floor pattern per
  pearl_blend_formulas_must_have_permanent_floor.
- target_jitter_k sentinel-bootstrapped from first non-zero raw
  smoothness loss per pearl_first_observation_bootstrap.
- jitter EMA shadowed via mapped-pinned smoothness_jitter_ema_host_d
  (DtoD captured in graph; host reads after end-of-step sync); host
  writes lr_mult_per_horizon_staging (mapped-pinned); next step's
  captured-graph rescale kernel reads via dev_ptr.

Scope reduction (Task 11): Controller C scales ONLY heads_w_skip_d
per-horizon (N_HORIZONS × HIDDEN_DIM = 640 contiguous floats, per-horizon
stride HIDDEN_DIM). Other heads weights (w1, w2, w_gate, w_main) stay
unscaled because they're entangled in the GRN gate/main computation and
rescaling them would require additional per-bucket scoping out of Task 11
scope. Revisit if Smoke 2 fails per feedback_no_quickfixes.

Both controllers active Phase 2 only; no-op in Phase 1.

GPU oracle tests: 9 total (7 existing + tau_clamp + lr_multiplier_scale).
All 33 ml-alpha lib tests pass.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-21 17:37:23 +02:00
jgrusewski
d16cab18bb feat(per-horizon-cfc): wire Phase 2 fused CfC forward dispatch in trainer
Per spec §2.1 + §5.4 point 1 and Task 10 of the plan.

Adds binding `cfc_step_per_branch_fwd_gpu` in cfc/step.rs that wraps
the fused per-branch kernel with the spec-mandated launch config:
grid=(B, N_HORIZONS=5, 1), block=(MAX_BUCKET_DIM=28, 1, 1) uniform
predicate, cooperative staging of x_local + h_old_local (2 × HIDDEN_DIM
floats). Also adds `cfc_step_per_branch_bwd_gpu` binding for Task 11.

CfcTrunk loads two new cubins (cfc_step_per_branch + heads_block_diagonal)
and caches three function handles. `heads_w_skip_compact_d` (from Task 9)
remains allocated as Phase-2-prepared metadata; full heads consumption
deferred to a follow-up task (Phase 2 heads dispatch needs GRN
integration, which is out of Task 10 scope per
`feedback_no_partial_refactor`).

In `dispatch_train_step`, the CfC dispatch branches on TrainingPhase:
- Phase1Warmup: existing `cfc_step_batched` (unchanged)
- Phase2Routed: `cfc_step_per_branch_fwd` consuming bucket-routed
  `tau_all_d` + `bucket_channel_offset_d` + `bucket_dim_k_d`

GRN heads dispatch stays unchanged for both phases per the Task 10
scope adjustment (Option B in the task brief). The `heads_w_skip` storage
remains 640 floats; the compact 128-float `heads_w_skip_compact_d` is
metadata-ready but not yet consumed.

Workspace cargo check clean (ml-alpha lib + trainer compile; only
pre-existing cupti/sp15/gpu_per_integration test errors remain).
ml-alpha lib tests: 33 passed.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-21 17:22:21 +02:00
jgrusewski
2f69bb1fc0 feat(per-horizon-cfc): wire Controller A + Phase 1→2 transition in trainer
Per spec §2.3 + §3.1 and Task 9 of the plan.

Adds two device kernels to bucket_transition_kernels.cu:
- tau_change_frobenius_kernel: ‖tau_t − tau_{t-1}‖_F² via block-tree
  reduction (no atomicAdd per feedback_no_atomicadd)
- slack_factor_apply_kernel: (Q1, Q3) → (Q1/slack, Q3*slack) where
  slack = sqrt(Q3/Q1), ISV-derived per feedback_isv_for_adaptive_bounds

Adds to PerceptionTrainer:
- TrainingPhase enum (Phase1Warmup / Phase2Routed)
- ControllerA state machine instance
- prev_tau_d device buffer + tau_change_d + tau_change_staged mapped-pinned
- bucket_warmup_cap_override CLI diagnostic field
- bucket_routing_metadata stored after transition fires
- heads_w_skip_compact_d (HIDDEN_DIM floats) populated by transition
- Cached function handles for the two new kernels

Per-step in step_batched (Phase 1 only):
- Launch tau_change_frobenius_kernel → mapped-pinned scalar shadow
- DtoD prev_tau ← tau_all_d for next step
- Sync + host scalar read
- controller_a.update(tau_change, bucket_warmup_cap_override)
- On trigger: stage tau into scratch buffer (avoids in/out aliasing in
  tau_reorder_kernel), execute_transition writes routing metadata +
  reorders tau_all_d + populates heads_w_skip_compact_d,
  slack_factor_apply widens IQR bounds, invalidate CUDA Graph,
  latch phase = Phase2Routed, log via tracing::info.

Phase 2 dispatch + Controllers B/C/D wiring deferred to Tasks 10–12.
In Task 9's transient state, the trainer enters Phase 2 but continues
Phase 1 dispatch path; the reordered tau_all_d still produces a valid
forward pass since CfC's per-channel decay math is order-invariant.

Per pearl_no_host_branches_in_captured_graph: transition fires OUTSIDE
the captured graph (cached graph invalidated → recaptured next step).
Per pearl_cudarc_disable_event_tracking_for_graph_capture: event tracking
is already disabled trainer-wide; recapture on next step is safe.
Per feedback_no_htod_htoh_only_mapped_pinned: host scalar read goes via
mapped-pinned DtoD shadow, not bulk DtoH.

CLI: --bucket-warmup-cap-steps added to alpha_train (Option<u64>,
diagnostic override of Controller A's ISV-derived cap).

Tests: 7 GPU oracle tests pass on RTX 3050 (5 existing + 2 new for
Frobenius + slack_factor). 33 ml-alpha lib tests pass. Workspace
cargo check clean modulo pre-existing cupti / sp15 / gpu_per_integration
errors unrelated to this task.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-21 17:09:01 +02:00
jgrusewski
fc0311de77 feat(per-horizon-cfc): atomic Checkpoint envelope migration (3 consumers)
Per spec §5.2 and feedback_no_partial_refactor. Adds bucket-routing
metadata to the Checkpoint envelope:
- Renames tau -> tau_all (same HIDDEN_DIM shape)
- Adds bucket_id_per_channel: Vec<u8>
- Adds bucket_channel_offset: Vec<u32>
- Adds bucket_dim_k: Vec<u32>
- Adds heads_w_skip_offset: Vec<u32>

heads_w_skip stays at [N_HORIZONS x HIDDEN_DIM]=640 floats; compaction
to [HIDDEN_DIM]=128 deferred to Task 10's forward dispatch update.

new_random zero-initializes bucket metadata; Phase 1->2 transition
(Task 9) populates them.

CfcTrunk field `tau_d` renamed to `tau_all_d` across:
- crates/ml-alpha/src/cfc/trunk.rs (struct + save_checkpoint
  + load_checkpoint + new_random + save_load_roundtrip bit-equality test
  with 4 new u8/u32 assertions on the bucket metadata fields)
- crates/ml-alpha/src/trainer/perception.rs (6 self.trunk.tau_d
  references + 1 trunk.tau_d upload + 1 comment line)

ml-backtesting/tests/checkpoint_smoke.rs needed no edit: its only
consumer surface is CfcTrunk::load_checkpoint(&dev, &cfg, &ckpt),
whose signature is unchanged. The 3-consumer atomic-migration audit
is satisfied because consumer #3's API contract is preserved.

Greenfield envelope (per existing cfc/trunk.rs:25-28 convention -
no backward compat).

Verified:
- cargo check --workspace --all-targets: clean (pre-existing
  cupti/insert_batch errors in crates/ml/tests excluded per task scope)
- cargo test -p ml-alpha --lib: 33 passed
- cargo test -p ml-backtesting --lib: 33 passed
- cargo test -p ml-alpha --lib save_load_roundtrip -- --ignored: PASS
  (bit-equality on tau_all_d + 4 new bucket metadata fields, RTX 3050)
- cargo test -p ml-backtesting --test checkpoint_smoke: compiles clean

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-21 16:35:48 +02:00
jgrusewski
1a7c4fc66a feat(per-horizon-cfc): add bucket_routing.rs (Controller A + transition dispatch)
Per spec §3.1, §2.3. Controller A computes tau-change EMA with
first-observation bootstrap + Wiener-α floor 0.4 (co-adapting closed
loop); trigger when EMA < max(noise_floor/4, noise_floor/16); ISV-derived
hard cap from observed MAD/median dispersion in first 100 steps.

execute_transition() dispatches the 5 device kernels in order:
tau_sort → bucket_assign → bucket_iqr → tau_reorder → heads_compact.
All-on-device per feedback_no_htod_htoh_only_mapped_pinned (one-time
static-constant upload at transition is allowed off-hot-path).

Returns raw Q1/Q3 in BucketRoutingMetadata.bucket_tau_iqr_{lo,hi}_d;
slack_factor=sqrt(Q3/Q1) widening kernel applied by trainer in Task 9.

3 CPU-only unit tests verify Controller A's first-obs bootstrap,
threshold-trigger, and ISV-cap-trigger semantics.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-21 16:18:04 +02:00
jgrusewski
c535be046f baseline(per-horizon-cfc): 3-seed h6000 baseline JSON skeleton + retrieval gap
Per spec §4.1 + Q-3-seed-baseline. Both baseline Argo workflows
succeeded:
- alpha-perception-baseline-seed16963 (39min)
- alpha-perception-baseline-seed16964 (26min)

Precise AUC values deferred: train pod stdout is in MinIO archive
(xl.meta format, needs mc auth). The on-disk alpha_train_summary.json
on /feature-cache/alpha-perception-runs/f22f3f948/ corresponds to
seed 16963 (last to finish). Task 15 implementer retrieves via debug
pod mount of feature-cache PVC.

Known data point: seed 16962 = 0.7454 best_auc_h6000.

The Smoke 1 invariant gate threshold is signal-derived per
pearl_tests_must_prove_not_lock_observations — formulas captured in
the JSON for Task 15 to apply once values are retrieved.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-21 16:10:48 +02:00
jgrusewski
fde5511338 feat(per-horizon-cfc): add heads_block_diagonal_fwd.cu (compact ragged w_skip)
Per spec §5.4 point 2. heads_w_skip storage 5× reduced (640→128 floats).
Each horizon reads only its bucket via offset lookup. Block-tree reduction
(padded to 32 lanes for clean halving stride), no atomicAdd.

GPU oracle test matches naive per-horizon computation (tol 1e-5) on sm_86.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-21 16:04:29 +02:00
jgrusewski
11eb74766a feat(per-horizon-cfc): add cfc_step_per_branch.cu (fused fwd+bwd)
Per spec §5.4 point 1. Single fused kernel: grid=(B, N_HORIZONS, 1),
block=(MAX_BUCKET_DIM=28, 1, 1) uniform predicate. Shared x_local +
h_old_local cooperative-staged into shared memory once per block per
pearl_cooperative_staging_eliminates_redundant_reads. No atomicAdd.
No host branches.

Forward GPU oracle: matches naive per-channel CfC math (tol 1e-4) on
sm_86. Backward smoke: kernel loads + writes non-NaN gradients.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-21 15:56:20 +02:00
jgrusewski
1f700f9bfa feat(per-horizon-cfc): add bucket_transition_kernels.cu (5 device kernels)
Per spec §2.3. All-on-device transition: tau_sort (bitonic-merge),
bucket_assign (static quintile boundaries [0,25,50,75,100,128]),
bucket_iqr (warp-reduction Q1/Q3 per bucket), tau_reorder, heads_compact.
No atomicAdd. No DtoH.

GPU oracle tests pass on RTX 3050 sm_86.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-21 15:46:56 +02:00
jgrusewski
f13d6c6dcf refactor(per-horizon-cfc): atomically remove MTER scaffolding
Per spec §2.5 and feedback_no_partial_refactor. Removes:
- LoaderMode enum (single-variant after Sequential removal → deleted entirely)
- next_sequence_sequential + sequential_file_idx + sequential_anchor_idx + last_call_was_file_boundary + is_file_boundary
- last_seen_file_boundary, train_graph_boundary_state fields
- notify_file_boundary method + boundary-state graph recapture logic
- need_attn_pool_bootstrap match — always true now (random mode always bootstraps)
- --loader-mode CLI flag + notify_file_boundary() call
- loader-mode Argo template param

Additional consumers migrated atomically (beyond the 4 files listed in
the plan): ml-alpha/tests/perception_overfit.rs,
ml-alpha/tests/multi_horizon_loader.rs, ml-backtesting/src/harness.rs,
ml-backtesting/tests/trainer_parity.rs,
ml-backtesting/tests/ring3_replay.rs — all referenced LoaderMode or the
removed config fields.

Workspace builds clean at this commit (pre-existing cudarc-cupti example
and ml-crate test errors are unrelated to MTER removal — they fail at
HEAD too). New per-horizon CfC arch lands in subsequent tasks of the
same plan.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-21 15:36:28 +02:00
jgrusewski
ed34d356a8 docs(per-horizon-cfc): kickoff — spec + plan + historical record
Spec: docs/superpowers/specs/2026-05-21-per-horizon-cfc-inference-design.md
Plan: docs/superpowers/plans/2026-05-21-per-horizon-cfc-inference-plan.md

Spec went through 2 critical-review passes (32 total findings, all resolved).
Bucketing source: CfC.tau (per-channel, trained, log-uniform init at HIDDEN_DIM=128).
Atomic refactor reverts MTER scaffolding. 5 ISV-driven controllers. All-on-device
transition (no bulk DtoH). Single fused per-branch kernel. Compact ragged heads_w_skip.
Validated via 3 Argo smokes (training stability, CRT.diag inference differentiation,
fxt-backtest end-to-end).

Also committing historical record: superseded MTER spec/plan, intervention A
plan, ISV λ controller spec/plan, GPU log ring spec/plan. Per spec §5.3 these
stay as audit trail; not in build path.

Plan has 18 tasks, full TDD with bite-sized steps, ready for
subagent-driven-development execution.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-21 15:05:28 +02:00
jgrusewski
f22f3f9488 fix(intervention-b): recapture train graph on boundary-state change
Task 1's attn_pool gate (commit 7dd953aa) was inside the captured-graph
region, so the host branch's decision was frozen at step-2 capture time.
In Sequential mode this would lock the captured graph to whichever
boundary state happened to hold at step 2, breaking subsequent file-
boundary detection — violating pearl_no_host_branches_in_captured_graph.

Fix: track the boundary state at capture time as a new
`train_graph_boundary_state: Option<bool>` field. On each step_batched,
if the current `last_seen_file_boundary` diverges from the captured
state, drop the cached graph and re-capture on the next step.

- Random mode: gate is unconditionally true; state never diverges;
  zero recaptures.
- Sequential mode: ~8 recaptures per epoch (one per file boundary).
  Each recapture costs the same as the original step-2 capture (~50ms).
  Negligible.

Verified clean cargo check -p ml-alpha --all-targets.
2026-05-21 11:39:29 +02:00
jgrusewski
7dd953aa2b feat(intervention-b): add LoaderMode::Sequential + attn_pool gate
Loader can now serve sequences in temporal order (Sequential mode);
when active, the trainer skips attn_pool reset at non-file-boundary
sequence boundaries (next commit will use this for stateful CfC + MTER).

Random mode preserved as default; ZERO behavioral change for existing
runs. Phased commit 1/8 of intervention B per
docs/superpowers/specs/2026-05-21-crt-train-intervention-b-multi-timescale-readout.md
2026-05-21 11:30:50 +02:00
jgrusewski
6d65423e32 fix(argo): drop --decision-stride from alpha-perception template (removed from CLI per A1) 2026-05-21 09:30:25 +02:00
jgrusewski
b5bed9f805 fix(crt-train): sqrt K-ratio in λ controller (tames h6000 bombardment)
Local trace at base_lambda=0.1 showed the linear ratio
{1, 0.3, 0.1, 0.03, 0.005} = HORIZONS[0]/HORIZONS[h] gave h6000 a
200x stronger target-undershoot signal than h30. The controller
slammed h6000 with smoothness gradient (peak λ[h6000]≈60) while h30
saw nearly none. h6000 val_auc collapsed to 0.514 (near-random)
while h100/h300 improved.

Replace with sqrt(HORIZONS[0]/HORIZONS[h]) = {1, 0.548, 0.316, 0.173,
0.0707}. Caps the differential at ~14x. Verified locally at base=0.1:
- h6000 val_auc recovered to 0.599 (vs 0.514 collapse, vs 0.621 no-smooth)
- jitter ratio h6000/h30 = 0.51 (meaningful differentiation, was 0.84 unforced)
- λ[h6000] equilibrium = 0.7-3 (was 4-60 with linear ratio at same base)

The design target (h6000 jitter = 7% of h30) is less aggressive than
linear (was 0.5%) but produces a tractable training equilibrium that
preserves h6000 predictive capacity. Both 500-step local AND the full
40k-step L40S retrain will tell us where it lands at scale.
2026-05-21 09:14:49 +02:00
jgrusewski
cd0fe8549c fix(gpu-log): use std::thread for drain task (no tokio runtime needed)
The drain task previously used tokio::spawn + tokio::time::sleep,
gated by Handle::try_current().is_ok(). alpha_train's main() is
synchronous (not #[tokio::main]) so Handle::try_current() returned
Err and the drain task was silently skipped — the trace file was
never created.

Refactor to std::thread::spawn + std::thread::sleep. No runtime
required; spawn is unconditional; Drop joins synchronously instead
of aborting; BufWriter flushes on every drain cycle so even short
runs (< 1 sec) capture their tail. Existing gpu_log_ring_invariants
tests migrated from #[tokio::test] back to plain #[test].

Also resolves an exit-1 issue on alpha_train shutdown — likely a
tokio task abort panicking through the synchronous Drop path.

Local smoke (alpha_train --kernel-step-trace ...):
- EXIT=0
- step_trace.jsonl: 1497 lines (~500 steps × 3 smoothness records)
- valid JSONL, schema matches smoothness_controller emission contract
2026-05-21 02:08:55 +02:00
jgrusewski
2d5a66f6e4 feat(kernel-step-trace): runtime --kernel-step-trace CLI flag + JSONL drain
The compile-time feature kernel-step-trace gates inclusion of the ring
code. NEW: --kernel-step-trace <PATH> CLI flag on alpha_train gates
RUNTIME activation:

- Path provided   -> ring allocated, drain spawns, JSONL records written
                    to <PATH> (truncated). One record per kernel emission.
- Path omitted    -> ring not allocated, drain not spawned, zero overhead
                    even with the feature compiled in.

JSONL schema: {"step": u32, "kid": u8, "kname": str, "rt": u8,
"rt_name": str, "payload": {field: f32, ...}}. The payload field names
come from the existing per-(kid, rt) decoders in gpu_log.rs (ported to
serde_json::json! in decode_to_json).

Trainer ring fields are now Option<_>; populated only when the runtime
trace path is Some. Tick kernel, step-counter shadow, smoothness
controller pointer-passing, and Drop all become path-conditional.

The tracing::info!-based drain variant is removed (file-writer is
strictly more useful; tests migrated to assert JSONL file contents).

Argo template:
- New parameter kernel-step-trace-enable (build-time feature opt-in)
- New parameter kernel-step-trace-path (runtime CLI value)
- ensure-binary cache-busts when feature toggles
- training step passes --kernel-step-trace flag conditionally

Per feedback_no_feature_flags: compile-time gate retained because the
ring carries real memory cost (~2 MiB pinned) and per-step write overhead;
the specific name kernel-step-trace narrows scope to this mechanism.
Runtime gate is Option<PathBuf>, not a boolean enable_*.
2026-05-21 01:55:57 +02:00
jgrusewski
fa41b39ca4 refactor(gpu-log): rename feature cuda-diag-log → kernel-step-trace
Per user direction: feature-gating is appropriate for per-step diagnostic
logging (real perf/memory cost) but the name should be specific to the
mechanism, not a generic "diag-log" catch-all. kernel-step-trace
describes what the feature provides — per-step records written to the
ring by kernels.

Pure rename, no behavioral changes. Touched: Cargo.toml feature decl,
lib.rs cfg attribute, perception.rs (~25 cfg attributes including
not(feature) pairs), gpu_log.rs + smoothness_lambda_controller.cu doc
comments, gpu_log_ring_invariants.rs file-level cfg + ignore-attribute
text.
2026-05-21 01:44:29 +02:00
jgrusewski
adc3506d3e feat(gpu-log): wire smoothness_lambda_controller as first ring producer
Atomic refactor wiring the GPU log ring's first producer end-to-end:

- cuda/gpu_log_helpers.cuh (new): extract LogHeader/LogRecord/LogRing
  struct definitions + the log_record() device __forceinline__ helper
  into a shared header. Single source of truth — other kernels include
  this rather than redefining the structs.
- cuda/gpu_log_ring.cu: refactor to include gpu_log_helpers.cuh; retain
  only the gpu_log_tick kernel.
- cuda/smoothness_lambda_controller.cu: include gpu_log_helpers.cuh, add
  trailing (LogRing*, const int*) args, capture pre-EMA jitter +
  post-EMA jitter + target + excess_ratio into shared mem, and emit
  three records per call (RT_INPUT, RT_STATE, RT_OUTPUT) gated on
  non-null ring pointer.
- trainer/perception.rs: feature-gated (cuda-diag-log) ring allocation +
  step counter (mapped-pinned host shadow), gpu_log_tick launch FIRST
  inside the captured graph, two new pointer args on the smoothness
  controller launch (null when feature off), step-counter DtoD shadow
  alongside the other telemetry shadows, drain task spawn (skipped when
  no tokio runtime — sync tests still construct the trainer), and a
  feature-gated Drop impl that aborts the drain task on shutdown.
- tests/smoothness_lambda_controller_invariants.rs: pass null pointers
  for the two new kernel args; the kernel's nullptr guard preserves
  pre-existing behaviour.
- build.rs: rerun-if-changed on cuda/gpu_log_ids.h and
  cuda/gpu_log_helpers.cuh so header edits trigger cubin rebuilds.

Verified: cargo build / check --all-targets clean both with and without
the cuda-diag-log feature; 4/4 smoothness controller invariant tests
pass; 9/9 perception_overfit integration tests pass under the feature.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-21 01:36:18 +02:00
jgrusewski
17cddf7f4a test(gpu-log): invariant tests for ring allocator + drainer 2026-05-21 01:21:55 +02:00
jgrusewski
339eaca983 feat(gpu-log): Rust-side LogRing, drain task, decoder registry
Task 5 of the GPU log ring plan. Adds `gpu_log` module behind the
`cuda-diag-log` feature gate:

- `LogRecord` (#[repr(C)], 64 B) mirroring the C-side layout in
  `gpu_log_ring.cu`; const compile-time size assertion catches drift.
- `LogRing::alloc()` allocates and zero-initialises a 32k-slot
  mapped-pinned ring (2 MiB pinned).
- `alloc_step_counter()` allocates the device-side i32 step counter.
- Decoder dispatch table with per-(kid, rt) decoders for
  `KID_SMOOTHNESS_CONTROLLER` × {RT_INPUT, RT_STATE, RT_OUTPUT}
  emitting structured tracing::info! events.
- `spawn_drain_task()` polls a mapped-pinned step-counter shadow at
  500 ms cadence, walks new slots, validates magic + step, dispatches
  to decoders. Emits tracing::warn! when the drainer falls behind the
  ring window (no silent record drops).

Kernel ID + record-type constants are a manual mirror of
`crates/ml-alpha/cuda/gpu_log_ids.h`; drift is caught by the const
size assertion at compile time and (subsequently) by build.rs.
2026-05-21 01:16:55 +02:00
jgrusewski
b1eaef5db2 refactor(gpu-log): generic MappedRecordBuffer<T>; MappedF32Buffer as alias 2026-05-21 01:11:22 +02:00
jgrusewski
92841bc5aa feat(gpu-log): add cuda-diag-log Cargo feature flag 2026-05-21 01:07:30 +02:00
jgrusewski
12e059c6ab feat(gpu-log): add gpu_log_ring.cu device helpers + tick kernel 2026-05-21 01:06:54 +02:00
jgrusewski
cd94a34122 feat(gpu-log): add gpu_log_ids.h — single source of truth for ring IDs 2026-05-21 01:04:50 +02:00
jgrusewski
0f664cf361 test(crt-train): invariant tests for smoothness_lambda_controller 2026-05-21 00:34:06 +02:00
jgrusewski
29ab397689 feat(crt-train): wire ISV-driven λ controller; remove SMOOTHNESS_LAMBDA_RATIO 2026-05-21 00:30:47 +02:00
jgrusewski
011a6b3cf2 build(crt-train): register smoothness_lambda_controller.cu in build.rs 2026-05-21 00:25:26 +02:00
jgrusewski
19c24d7f5a feat(crt-train): add ISV-driven smoothness_lambda_controller kernel 2026-05-21 00:23:45 +02:00
jgrusewski
70ecfc0fe6 feat(crt-train): plumb --smoothness-base-lambda through alpha-perception template 2026-05-20 23:55:15 +02:00
jgrusewski
60d0b62b8d feat(crt-train): add --smoothness-base-lambda CLI flag + telemetry 2026-05-20 23:51:20 +02:00
jgrusewski
c9aa0c2a98 feat(crt-train): launch output_smoothness kernel after BCE in step_batched 2026-05-20 23:45:51 +02:00
jgrusewski
4a32186a5f fix(crt-train): migrate all PerceptionTrainerConfig literals to new smoothness_base_lambda field
Task 5 added smoothness_base_lambda to PerceptionTrainerConfig but only
updated the struct definition + Default impl. Eight test sites in
perception_overfit.rs and one site in alpha_train.rs construct the
config via explicit field list (no ..Default::default() spread), so
they broke with E0063 missing-field errors under cargo check
--all-targets.

Per feedback_no_partial_refactor.md: when a contract changes, every
consumer migrates atomically. This commit adds smoothness_base_lambda:
0.0 to all 9 sites so the workspace compiles cleanly. The
alpha_train.rs value of 0.0 is a placeholder — Task 7 will replace it
with cli.smoothness_base_lambda once the CLI flag is added.
2026-05-20 23:40:45 +02:00
jgrusewski
b83c14cf44 feat(crt-train): allocate smoothness buffers + upload λ in PerceptionTrainer 2026-05-20 23:38:33 +02:00
jgrusewski
3e07bf41b8 test(crt-train): finite-diff validate output_smoothness kernel grad 2026-05-20 23:31:40 +02:00
jgrusewski
ac7c52b5be feat(crt-train): add SMOOTHNESS_LAMBDA_RATIO const in heads.rs 2026-05-20 23:27:18 +02:00
jgrusewski
a342126543 polish(crt-train): branchless boundary loads + remove misleading factor==0 comment
Code-quality reviewer flagged two non-blocking polish items in the
output_smoothness kernel's backward pass:

- I1: thread-divergent if-gated loads contradicted the file's stated
  branchless-design intent. Fixed via clamp-then-load: at k=0 / k=K-1
  read self for the spurious neighbour; lm/rm multipliers zero the
  contribution. Every lane executes both loads; no divergence on
  boundaries.
- I2: the "Skip work when factor == 0" comment described a non-feature
  (no such branch existed). Removed.

Behaviour-equivalent change — same loss and same gradient at every
position; the change only affects warp-execution patterns at the K=0
and K=K-1 strips.
2026-05-20 23:24:33 +02:00
jgrusewski
3e1d66f865 build(crt-train): register output_smoothness.cu in ml-alpha build.rs KERNELS 2026-05-20 23:22:45 +02:00
jgrusewski
bf5ecd109b feat(crt-train): add output_smoothness CUDA kernel (forward + analytic grad) 2026-05-20 23:16:17 +02:00
jgrusewski
faa9a73c10 spec(crt-train): output smoothness retraining intervention (intervention A)
Per project_crt_diag_findings.md — CRT.diag + diag.2 empirically
falsified all three hypotheses about why the model's per-event output
doesn't track horizon-specific dynamics. Labels ARE differentiated per
horizon, AUC=0.66 IS per-event measured, but per-event predictions
behave identically across all 5 horizons (2.5-event mean run length
raw, 3.3-event after aggressive Wiener-α EMA smoothing). h6000
predictions should change every thousands of events, not every 3.3.

Diagnosis: training dynamics issue. The BCE loss provides no signal
pushing toward horizon-coherent predictions. Adjacent-event labels at
horizon K share K-1/K of the forward window so SHOULD produce similar
predictions, but the loss doesn't require this.

Intervention A (this spec, minimum-scope): output smoothness
regularizer
  L_smooth[h] = λ[h] × mean_t (p[h](t) - p[h](t-1))²

With horizon-weighted λ (stronger for h6000 than h30): forces h6000
predictions to change slowly while leaving h30 responsive.

Implementation surface:
  - multi_horizon_heads.cu backward: extend grad_probs with smoothness
    gradient
  - perception.rs trainer step: prev_probs_d buffer, per-horizon (p[t]
    - p[t-1])² accumulation
  - heads.rs: SMOOTHNESS_LAMBDA constant per horizon
  - alpha_train.rs: --smoothness-base-lambda CLI flag

Validation gate (CRT.train Gate):
  MUST: per-horizon AUC ≥ baseline - 0.02 (no signal destruction)
  WIN: h6000 mean_run_len ≥ 100 events; h6000/h30 ratio ≥ 10× (horizons
       actually differentiated post-training)
  STRETCH: CRT.1 controller smoke shows mean PnL CORRELATED with
       conviction (vs anti-correlated in lnfwd)

Future work if intervention A doesn't pass WIN:
  B: horizon-conditional output structure (architecture change)
  C: curriculum on horizon (multi-day training)
  D: different label generation (majority-vote vs single-point)

Status: Design — awaiting user review before plan.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-20 22:59:49 +02:00
jgrusewski
0e17fd4f2d diag(crt-2): per-horizon alpha-input EMA test — hypothesis investigation
Driven by ffr59 (commit b44a97ff9) findings: all 5 horizons flip
direction every 2.5 events; win rate flat 24% across conviction; mean
PnL anti-correlated with conviction. Hypothesis: per-event alpha output
is high-frequency noise on top of a slower signal. If true, smoothing
input alpha BEFORE the conviction formula should reduce direction flips
and recover signal.

Adds Wiener-α adaptive EMA on raw alpha_probs[h] for each horizon,
applied BEFORE the multi-horizon conviction formula. Floor at 0.1
(stronger than the 0.4 floor on the output-side conviction EMA — this
tests whether INPUT smoothing has different impact than OUTPUT smoothing).

Three new device slots:
  - alpha_ema_per_b_per_h (per-horizon EMA state)
  - alpha_diff_var_per_b_per_h (variance of changes)
  - alpha_sample_var_per_b_per_h (variance of value)

The CRT.diag Group A direction-flip counter still reads RAW alpha_probs
so we have a head-to-head comparison: raw flip rate vs smoothed flip rate.
Group E adds the smoothed-direction counter + mean run length.

End-of-run log adds one line per horizon:
  crt_diag h<X> smoothed: flips=Y mean_run_len=Z events (vs raw F / M)

If smoothed mean_run_len >> raw mean_run_len: hypothesis is RIGHT, the
input had signal under the noise. Next step would be to make this an
operational EMA in the controller.

If smoothed and raw are similar: hypothesis is WRONG, per-event output
is genuinely noisy. Next step would be to investigate model training
(horizon collapse) OR the AUC=0.66 measurement definition.

Either way, definitive result from one smoke run.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-20 22:34:35 +02:00
jgrusewski
b44a97ff94 diag(crt): empirical measurement battery for signal × controller × cost analysis
Pure-instrumentation commit. Zero behavior change. The Gate CRT.1 smoke
p9cnk (commit 1e656948b) produced 222k trades and 29688% drawdown,
worse than the failed CRT.A. The structural fixes (multi-horizon
conviction, no-trade band, composite exit) all work as designed — but
the trade cadence is fundamentally too fast for the signal/cost ratio.
Before another round of threshold tuning, measure the structural truth.

Adds device-side counters dumped at smoke close (single end-of-run
memcpy_dtoh batch in read_diagnostics(), NOT per-event):

Group A — per-horizon signal persistence:
  - flip_count[h]:        total direction-sign flips
  - sum_run_length[h]:    cumulative events between flips (u64)
  - run_length_hist[h]:   5-bucket log-scale histogram (1-9, 10-99,
                          100-999, 1k-9.9k, 10k+)
  Output: mean_run_len per horizon = signal half-life proxy

Group B — conviction smoothed-EMA distribution:
  - conv_hist[10]:        10-bucket histogram over [0, 1)
  Output: shape of conviction distribution at decision time

Group C — hold-time distribution:
  - hold_hist[6]:         <1s, 1-10s, 10-60s, 1-10m, 10-60m, >1h
  Output: empirical distribution of trade hold times

Group D — outcome by entry-conviction:
  - outcome_n[10]:        trades per conviction bucket
  - outcome_sum_pnl[10]:  cumulative realized PnL per bucket (price-units)
  - outcome_n_wins[10]:   wins per bucket
  Output: validates "high conviction = better trades" hypothesis OR
          surfaces signal miscalibration

Per-backtest memory: ~330 bytes. Negligible vs existing state.

All counters single-writer-per-block (threadIdx.x == 0 convention; no
atomicAdd per feedback_no_atomicadd). All updates kernel-internal (no
host roundtrip per event). The end-of-run dump uses memcpy_dtoh in
read_diagnostics() — called once at harness termination from the
post-stop_ctrl_counters block, never per-event. Per
feedback_no_htod_htoh_only_mapped_pinned, the hot path is untouched.

N_HORIZONS = 5 (heads.rs canonical {30, 100, 300, 1000, 6000}); the
plan text says 4 but the workspace constant is 5 — used the code value.

Output at end of cluster smoke as a `crt_diag` log block:
  crt_diag h30: flips=X mean_run_len=Y events
  crt_diag h30 run_length_hist: 1-9:N 10-99:N 100-999:N ...
  ...
  crt_diag conv_ema_hist: [0.0-0.1]:N [0.1-0.2]:N ...
  crt_diag hold_time_hist: <1s:N 1-10s:N ...
  crt_diag outcome_by_entry_conv[0.0-0.1]: n=N win_rate=X.X% mean_pnl_pu=Y.YYY
  ...

The output drives the next round of CRT.1 threshold design (delta_floor,
composite exit factors, min-hold cooldown) from data instead of guesses.

Per pearl_adaptive_not_tuned and pearl_controller_anchors_isv_driven:
thresholds will become ISV-derived in CRT.2; CRT.1 must first establish
defensible defaults from the empirical signal characteristics.

Verified: stop_controller (22/22 pass), decision_floor_coldstart (3/3
pass), threshold_and_cost (3/3 pass), parallel_sim_correctness (1/1
pass), lob_sim_integrated_fuzz (3/3 pass). Two lob_sim_fixtures
failures (fix_decision_alpha_buy_close, fix_decision_program_h4_only)
were pre-existing at HEAD 1e656948b — not caused by this commit.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-20 21:56:07 +02:00
jgrusewski
1e656948b3 arch(crt-1): composite exit_signal safety circuit-breaker (C1.4)
Per spec §4.3 + plan C1.4. Safety backup for the primary exit path
(target → 0 from collapsed multi-horizon conviction, C1.2). Catches
the edge case where conviction stays bounded above zero but the trade
is deep in unrealized loss, or short/long horizons disagree for many
events while threshold-gate logic still permits sizing.

Three terms, max-of-three composite:
  degradation_term  = (1 − conv_ema_now / conviction_at_entry) / 2.0
  disagreement_term = disagree_consecutive_events / 32.0
  pnl_decay_term    = max(−pnl_adjusted_conv_ema, 0) / 1.0
  exit_signal       = max(deg, dis, pnl_decay)

Fires force-flat (market_targets = (3, 0)) when exit_signal >= 1.0.

State-update flow:
  - pnl_track open branch: write conviction_at_entry (offset 20),
    conviction_ema (offset 24, initialised to entry value),
    pnl_adjusted_conviction_ema (offset 44, = 0),
    peak_unrealized_pnl (offset 48, = 0). Takes new conviction_ema_per_b
    read-only arg to snapshot at open.
  - decision kernel intra-event (when pos open): update offsets
    24/44/48/56 (disagreement counter) via update_open_trade_trajectory.
  - composite_exit_check device helper: reads offsets 20/24/44/56,
    writes (3, 0) on fire. Sibling of stop_check_isv with same return
    semantics; called after trajectory update so all four reads see
    the freshly-written values.

Fields not read in CRT.1 (offsets 28-43 per-horizon EMA, 52 degradation
counter, 60 horizon_idx_dominant) are NOT written by the open branch —
alloc_zeros covers init. Per feedback_no_hiding: don't populate unused
fields. Also fixes a latent C1.1 bug: pnl_track close branch read
horizon_idx from offset 20 (now conviction_at_entry, f32); moved the
read to offset 60 per the documented layout.

Threshold-bootstrap defaults (will be ISV-derived in CRT.2 alongside
the rest of the controller anchors per pearl_controller_anchors_isv_driven):
  degradation_factor  = 2.0  (50% conviction decay vs entry → ≤0.5 term)
  disagreement_factor = 32.0 (32 events of short/long disagreement)
  pnl_decay_factor    = 1.0  (pnl_adj_conv structurally ∈ [-1, +1])

Under these defaults disagreement_term is the only term that can fire
on its own — degradation_term caps at 0.5 (conv_ema/conviction_at_entry
≥ 0) and pnl_decay_term caps at 1.0 in the limit. The three terms still
add evidence in superposition; CRT.2 will retune factors against ISV
percentiles. Unit-test coverage validates the disagreement path
(composite_exit_signal_fires_on_disagreement).

Per pearl_no_host_branches_in_captured_graph — all composite logic
lives inside the decision kernel; no host roundtrip in the per-event
hot path. Per feedback_no_htod_htoh_only_mapped_pinned — no new
memcpy or stream.synchronize() in step_pnl_track or
dispatch_latent_market_orders.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-20 21:15:40 +02:00