Commit Graph

2220 Commits

Author SHA1 Message Date
jgrusewski
e968f4ded9 feat(sp15-wave3a): kernel-side foundation — baseline output buffers + position_history derivation
Wave 3a half of the val-cost-streams refactor (3b host-side wire-up
follows). Atomically migrates the kernel-side contracts; 5 launchers
remain orphan transiently awaiting 3b production callers.

Baseline kernels (1.4):
  - 4 baseline_*_kernel signatures gain 'out: float*' parameter writing
    per-window [mean, std, raw_sharpe] (matches 1.1.b sharpe_per_bar shape)
  - ISV writes to slots 409, 410, 412, 416 removed entirely
    (per-window output is correct for WindowMetrics consumption;
    ISV-scalar writes were spec scaffolding for a single-fold-aggregate
    version that 1.4.b's per-window contract supersedes)
  - 4 ISV slot constants removed from sp15_isv_slots.rs
  - state_reset_registry: NO entries to remove (verified via grep —
    the 4 slots never had registry entries / dispatch arms in the first
    place; they were single-fold-aggregate scalars defaulted at every
    fold start by the constructor-write that initialises the ISV bus).
    Task 4 from the dispatch is a no-op; the
    every_fold_and_soft_reset_entry_has_dispatch_arm regression test
    continues to pass unchanged.
  - 4 oracle tests migrated to output-buffer assertion
  - layout_fingerprint_seed string updated (4 retired entries removed,
    4 trunk-shared entries retained; layout-break-class change)

New action_decoding_helpers.cuh:
  - Extracts factored_action_to_dir_idx + factored_action_to_position
    __device__ helpers (the latter is a higher-level position state-
    machine helper not previously available)
  - Mirrors trade_physics.cuh::decode_direction_4b semantics exactly so
    on-policy and counterfactual paths agree on factored-action meaning
  - Single source of truth for action→direction→position mapping;
    consumers #include the header

New position_history_derivation_kernel.cu (post-loop derivation for
cost_net_sharpe consumer in Wave 3b):
  - Reads actions_history_buf, reconstructs per-bar position_history
    (-1/0/+1), side_ind (1.0 on position change), rt_ind (1.0 on
    transition-to-flat from non-flat) via sequential walk (single
    block per window, no atomicAdd per feedback_no_atomicadd)
  - New launcher launch_sp15_position_history_derivation in
    gpu_dqn_trainer.rs
  - New cubin manifest entry in build.rs
  - 1 oracle test covering 8-bar Short→Hold→Long→Hold→Flat→Long→Flat→
    Short sequence; expected position/side_ind/rt_ind triples match
    hand-computed values

Atomic per feedback_no_partial_refactor for the ISV-contract change
(every consumer of slots 409/410/412/416 migrated in this commit; their
consumers were the 4 oracle tests, all migrated). The orphan launcher
transient state for the 5 baselines + derivation kernel is explicitly
the 3a/3b split point — production callers land in 3b's
GpuBacktestEvaluator::new constructor signature change.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-06 23:03:30 +02:00
jgrusewski
334b496647 feat(sp15-wave2): fused post-SP11 reward-axis composer (layered architecture)
Closes the deferred-consumer gap left by Phase 3.1
(r_quality_discipline_split_kernel, commit 5d36f3238), Phase 3.3
(dd_penalty_kernel), and Phase 3.5.2 (dd_asymmetric_reward_kernel) —
three SP15 reward-axis kernels that landed only as standalone scalar
producers awaiting deferred consumer wiring per
feedback_no_partial_refactor.md.

Wave 2 chooses Option β (layered, composable post-modifier) over
Option α (replace SP11 entirely): SP11 B1b stays canonical
"trader-quality" composer that writes out_rewards; the SP15 reward-axis
composition becomes a fused PER-(i,t) post-modifier read-modify-writing
the same buffer in place. This preserves SP11's z-score mag-ratio
contract (canary tests untouched) while landing all three deferred SP15
consumers atomically with zero parallel paths.

Architecture (Q1-Q4 user resolutions):
  Q1: SP12 caps stay as state_layout.cuh macros (REWARD_NEG_CAP=-10,
      REWARD_POS_CAP=+5), NOT lifted to ISV — spec'd constants per the
      SP12 v3 design, not adaptive bounds.
  Q2: Both on-policy + CF slots get the same DD-aware shaping. CF reward
      = w_cf × r_cf (SP11 controller weight already applied) is composed
      via the same helpers as on-policy. DD context is per-step, not
      per-action.
  Q3: New slot_completed_normally[N*L] flag preserves SP11's early-return
      semantics (data-end at experience_kernels.cu:2142 → reward=0.0;
      blown-account at :2203 → reward=-10.0). Fused kernel skips slots
      where flag==0.
  Q4: alpha_split_producer_kernel OWNS the warm-count increment (per-step
      scalar producer; the fused kernel has 2*N*L threads and would
      over-tick by that factor if it owned the increment).

Per-step launch sequence in gpu_experience_collector.rs (after Phase
1.3.b's dd_state launch): experience_env_step → alpha_split_producer
(reads grad-norm slots 418/419, writes ALPHA_SPLIT slot 417, increments
warm-count) → compute_sp15_final_reward (fused per-(i,t) over [N*2*L]
slots, applies α-blend → DD asymmetric → DD penalty → SP12 cap, writes
back to out_rewards in place).

experience_env_step signature change — 2 new output params:
  r_discipline_out: float* [N*L] — per-step REGRET_EMA mirror, written
    at end of normal reward composition.
  slot_completed_normally_out: int* [N*L] — 0 default at entry, set to
    1 only on the path that reaches out_rewards[out_off] = reward.

3 deleted kernel files:
  - r_quality_discipline_split_kernel.cu (composer + producer; replaced
    by renamed alpha_split_producer_kernel.cu keeping ONLY the producer
    with the moved warm-count increment).
  - dd_penalty_kernel.cu (replaced by sp15_dd_penalty __device__ helper).
  - dd_asymmetric_reward_kernel.cu (replaced by sp15_dd_asymmetric_reward
    helper).

3 new files:
  - alpha_split_producer_kernel.cu (per-step scalar producer of α from
    grad-norm ratio, with warm-count increment moved here per Q4).
  - sp15_reward_axis_helpers.cuh (4 __device__ inline helpers:
    sp15_alpha_blend, sp15_dd_asymmetric_reward, sp15_dd_penalty,
    sp15_apply_sp12_cap).
  - compute_sp15_final_reward_kernel.cu (fused per-(i,t) parallel over
    [N*2*L] slots — α-blend + DD-asymmetric + DD-penalty + SP12 cap,
    skips early-return sentinel slots).

3 deleted launchers + 3 deleted CUBIN statics in gpu_dqn_trainer.rs:
  - launch_sp15_r_quality_discipline_split (composer scalar variant).
  - launch_sp15_dd_penalty.
  - launch_sp15_dd_asymmetric_reward.
  - SP15_R_QUALITY_DISCIPLINE_SPLIT_CUBIN.
  - SP15_DD_PENALTY_CUBIN.
  - SP15_DD_ASYMMETRIC_REWARD_CUBIN.

2 new launchers + 2 new CUBIN statics:
  - launch_sp15_final_reward + SP15_FINAL_REWARD_CUBIN.
  - SP15_ALPHA_SPLIT_PRODUCER_CUBIN (the retained launch_sp15_alpha_split_producer
    now loads this).

GpuExperienceCollector: 2 new CudaSlice fields
(r_discipline_per_sample, slot_completed_normally_per_sample) +
sp15_alpha_warm_count_dev_ptr field + set_sp15_alpha_warm_count_ptr
setter + Step 5b launch block.

State reset registry — no new entries: r_discipline +
slot_completed_normally are per-step ephemeral (defaulted at every
kernel entry); cross-fold leakage impossible. Existing Phase 3.1/3.3/
3.5.2 ISV slot entries cover the rest.

6 new oracle tests in sp15_phase1_oracle_tests.rs::mod gpu drive
compute_sp15_final_reward_kernel directly with hand-crafted buffers:
  - final_reward_alpha_blend_at_cold_start (Stage 1)
  - final_reward_dd_penalty_above_threshold (Stage 3)
  - final_reward_dd_asymmetric_gain (Stage 2 + R_GAIN_DD_BOOST diag)
  - final_reward_dd_asymmetric_loss (Stage 2 asymmetric guard)
  - final_reward_sp12_cap_clamps_both_directions (Stage 4)
  - final_reward_skips_early_return_slots (Q3 sentinel preservation)

9 deleted scalar-kernel oracle tests:
  - r_split_uses_sentinel_alpha_at_cold_start
  - r_quality_subtracts_explicit_cost
  - dd_penalty_quadratic_above_threshold + dd_penalty_zero_below_threshold
  - dd_asymmetric_reward_gain_amplified_by_dd_pct +
    dd_asymmetric_reward_loss_unchanged + dd_asymmetric_reward_no_op_at_ath
  (the new fused-kernel tests cover the same behavioral surface
  end-to-end through the in-place RMW path).

Verified: SQLX_OFFLINE=true cargo check -p ml --features cuda clean (18
unrelated warnings); all 29 SP15 phase1 oracle tests pass on RTX 3050 Ti
(includes the 6 new fused-kernel tests + 17 retained tests + 6 old);
SP11 mag-ratio canary tests still untouched (no canary-test renames or
deletions); ml lib suite holds 946 pass / 13 fail = baseline.

Hard rules: feedback_no_partial_refactor (3 phases' deferred consumers +
3 deletions + 3 new files + 6 new tests + audit doc all in this commit;
no parallel paths, no feature flags), feedback_wire_everything_up
(closes 3 SP15 phase orphan launchers atomically), feedback_no_legacy_aliases
(deletions land in same commit as replacement; no compatibility shim),
feedback_no_atomicadd (fused kernel is per-(i,t) parallel, pure scalar
arithmetic), pearl_audit_unboundedness_for_implicit_asymmetry (gain-only
DD multiplier + asymmetric NEG/POS caps preserve loss aversion),
pearl_symmetric_clamp_audit (SP12 cap is bilateral via fmaxf/fminf even
though bounds are intentionally asymmetric per spec),
pearl_no_host_branches_in_captured_graph (new launches happen inside
collect_experiences_gpu::launch_timestep_loop per-step, OUTSIDE the
experience-fwd CUDA Graph capture region — same precedent as Phase
1.3.b's dd_state launch).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-06 22:22:33 +02:00
jgrusewski
f01a292f6f feat(sp15-p3.5.b+3.5.3.b): wire hold_floor (inline) + cooldown mask into experience_action_select
Phase 3.5 (hold_floor_kernel) + Phase 3.5.3 (cooldown_kernel) landed
the producer + state machinery; both deferred the action-selection
consumer wiring. This task wires both atomically.

Architectural decision: hold_floor is now an INLINE __device__
computation inside experience_action_select reading ISV slots
426/427/428/429 directly. The standalone hold_floor_kernel.cu +
launch_sp15_hold_floor + HOLD_FLOOR_CUBIN are deleted — launching a
kernel to write one f32 just to read it back was unnecessary. ISV
slots + state_reset_registry entries remain; only the launch path is
removed per feedback_wire_everything_up + feedback_no_legacy_aliases.

Entropy source: per-step Shannon entropy of softmax(e_dir) computed
inline from the 4 e_dir floats already in registers (Pass 1 of the
Thompson direction selector). High entropy = uncertain policy → Hold
gets the floor lift; low entropy = confident policy → floor ≈ 0.

q_eff_dir scratch preserves e_dir for downstream consumers
(out_conviction, out_q_gaps, out_magnitude_conviction) — adding
hold_floor there would corrupt the Kelly-cap warmup floor with a
meta-confidence mask.

cooldown mask: when ISV[COOLDOWN_BARS_REMAINING=435] > 0,
action_select hard short-circuits to dir_idx = DIR_HOLD before
Pass 2 — sidesteps the temperature-blend numerics where a
finite-sentinel-on-non-Hold approach would let pure-Thompson (τ=1)
samples dominate the masked direction. Cooldown supersedes
hold_floor — when forcing Hold the floor is moot.

3 new oracle tests:
  - action_select_applies_hold_floor_inline (no cooldown)
  - action_select_forces_hold_during_cooldown
  - action_select_no_force_hold_when_cooldown_zero

Atomic per feedback_no_partial_refactor: action_select changes +
hold_floor_kernel deletion + cubin manifest update + 3 oracle tests +
audit doc all in this commit. No parallel paths, no feature flags.

Eliminates Phase 3.5 + Phase 3.5.3 deferred consumers. The
cooldown_kernel itself remains (it maintains the consecutive_losses
streak + decrements COOLDOWN_BARS_REMAINING per bar); only its
consumer is now wired.

Verified: cargo check -p ml --features cuda clean; ml lib suite
holds 946 pass / 13 fail = baseline; all 6 oracle tests pass on
RTX 3050 Ti (3 pre-existing cooldown + 3 new action_select).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-06 21:18:31 +02:00
jgrusewski
d7f60d4dd7 feat(sp15-p1.6.b+1.7.b): wire dev-eval (Q8 final-fold) + test-eval (per-fold) into trainer
Phase 1.6 (CLI flags + dev_features/holdout_features stash) and Phase
1.7 (set_test_data_from_slices observer + test_features stash) landed
the data-flow scaffolding; both deferred the actual eval consumer.

This task wires both atomically as parallel evaluator instances:
  - dev_evaluator: Option<GpuBacktestEvaluator> -- lazy-init after final
    fold, fires once against Q8 dev_features (when dev_quarters > 0)
  - test_evaluator: Option<GpuBacktestEvaluator> -- lazy-init per fold,
    fires inside the fold loop against the WF test slice (when
    fold.test_end > fold.test_start)

Architectural choice: parallel evaluator instances (NOT window-swap on
val_evaluator). Window-swap would require invalidating the CUDA graph
between val and dev/test runs -- fragile, and a direct violation of
pearl_no_host_branches_in_captured_graph. Parallel instances mirror
val_evaluator's lazy-init pattern (TLOB sync, ISV signal pointer,
training_mode = false toggle). Implementation lives behind a single
shared helper `launch_extra_eval` keyed on an `ExtraEvalKind` enum so
Dev / Test share TLOB / ISV / config setup verbatim.

HEALTH_DIAG additions:
  HEALTH_DIAG[N]: dev_eval dev_sharpe_net=... dev_calmar=... dev_max_dd=... dev_trades=...
  HEALTH_DIAG[N]: test_slice fold=K test_sharpe_net=... test_calmar=... test_max_dd=... test_trades=...

The *_sharpe_net key uses the fused-metrics-kernel cost-aware Sharpe
(post-Phase-1.1.b split -- already includes tx_cost_bps + spread_cost
via the env-step PnL feed); when Phase 1.2.b cost-net sharpe lands, the
key name is preserved so the aggregator-script contract holds.

Atomic per feedback_no_partial_refactor: both eval calls + both
evaluator fields + both HEALTH_DIAG lines + audit doc all in this
commit. No parallel paths, no feature flags. Dev_eval runs synchronously
via evaluate_dqn_graphed (one-shot, no async pipelining benefit since
it doesn't fire per-epoch); val path stays async.

Sealed Q9 holdout remains untouched -- Phase 4.3 will load Q9 via a
separate eval-only entry point (NOT train_walk_forward). The Phase 1.6
debug_assert sealed-slice guard catches accidental future refactors.

Verified: cargo check -p ml --features cuda clean; ml lib suite holds
946 pass / 13 fail baseline; the existing Phase 1.7 oracle test
set_test_data_from_slices_fires_observer_and_stashes still passes.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-06 20:55:23 +02:00
jgrusewski
132609724e feat(sp15-p1.3.b): wire dd_state per-step launch + drop equity-recompute bug + env-0 canonical observable
Path A of the blocked 1.3.b investigation: fixes two architectural
issues atomically and wires the launcher.

(1) Bug fix: dd_state_kernel.cu was recomputing new_equity =
PS_PREV_EQUITY + pnl_step and writing it back, but experience_env_step
already maintains PS_PREV_EQUITY (experience_kernels.cu:3473-3475) —
wiring as-is would silently double-accumulate equity every step.
Kernel now READS PS_PREV_EQUITY / PS_PEAK_EQUITY only; does not
modify them. pnl_step parameter dropped from both kernel and
launcher signatures.

(2) Per-env shape decision: kernel is single-thread/single-block;
production has N envs but DD ISV slots [401..407) are scalars.
Picks 'env 0 as canonical observable' — kernel reads
pos_state[0 * PS_STRIDE + ...]. Per-env redesign (per-env tiles +
reduction kernel) deferred to Phase 1.3.b-followup if L40S smoke
shows single-env DD aggregation is insufficient.

(3) Wire-up: launch added at gpu_experience_collector.rs step 5b in
launch_timestep_loop, immediately after env_step writes PS_PREV_EQUITY,
outside the exp-fwd graph capture region (which ends at line ~3829,
well before env_step). Atomic per feedback_no_partial_refactor:
kernel signature change + oracle test update + launcher call site
update all in this commit.

Eliminates the Phase 1.3 orphan launcher per feedback_wire_everything_up.
Downstream Phase 3.3 / 3.5.2 / 3.5.4 / 3.5.5 readers will receive live
DD values when their consumer wiring lands.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-06 20:06:57 +02:00
jgrusewski
eda1eccb1a merge(sp15): bring phase2a (LobBar + behavioral test scaffold + 17 tests) into phase1
Brings in Phase 2A.1 (LobBar canonical ABI + 4 synthetic market generators),
Phase 2A.2 (oracle + harness + pre-commit hook), and Phase 2B (17 #[ignore]
behavioral test contracts) so the phase1 honest-numbers branch has access
to the LobBar (price, half_spread, ofi) ABI needed for Phase 1.2.b cost-net
sharpe consumer wiring.

Path 2 of the BLOCKED 1.2.b investigation: the cost-net kernel needs
GPU-resident streams (half_spread, ofi, rt_ind, side_ind, position) that
do not exist on phase1; Phase 2A.1's LobBar provides the canonical ABI.

Conflict resolution: docs/dqn-wire-up-audit.md — both branches prepended
entries; merged by keeping all three (Phase 2A.1 from phase2a, Phase 1.6,
Phase 1.7 from phase1). Phase 2A.1 entry placed above Phase 1.6 / 1.7 to
keep this region's audit ordering consistent (newer-first locally).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-06 19:43:46 +02:00
jgrusewski
b791bc8f7f refactor(sp15-p1.1.b): split sharpe out of fused backtest_metrics kernel; wire dedicated launch_sp15_sharpe_per_bar
Phase 1.1 landed sharpe_per_bar_kernel.cu + launch_sp15_sharpe_per_bar
as orphan scaffolding because the val-side sharpe was inline in
backtest_metrics_kernel's 8-metric fusion (lines 208-211, 277), not
a host-side loop the spec sketch had assumed.

This refactor splits sharpe out:
  - backtest_metrics_kernel computes 7 metrics now (sortino, win_rate,
    max_dd, calmar, omega, VaR, CVaR; remaining counters unchanged).
    Output stride drops 14 -> 13; shmem 6 -> 5 reduction arrays.
  - gpu_backtest_evaluator calls launch_sp15_sharpe_per_bar against
    the same GPU-resident per-bar returns buffer, once per window
    (kernel is single-block by design; n_windows is small).
  - Annualization moves host-side: WindowMetrics.sharpe =
    raw_sharpe * annualization_factor.

Atomic per feedback_no_partial_refactor: kernel split + offset
rebase (every metric below sharpe shifted down by 1) + the lone
WindowMetrics.sharpe consumer (consume_metrics_after_event)
migrated in one commit. No parallel paths.

Output value of WindowMetrics.sharpe is preserved (verified to
1e-5 relative error against f64 closed-form via new oracle test
unified_sharpe_kernel_equivalence_under_annualization). All
existing Phase 1.1 oracle tests still pass; ml lib test suite
holds at the 945/13 baseline (no new regressions).

Eliminates the Phase 1.1 orphan launcher per feedback_wire_everything_up.
Sets up Phase 1.2.b cost-net sharpe to also use launch_sp15_cost_net_sharpe
on the cost-net returns buffer (separate task, separate commit).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-06 19:29:41 +02:00
jgrusewski
69b8fdb61a feat(sp15-p3.5.5): recovery curriculum — per-step DD_TRAJECTORY_DECREASING proxy
Per spec §9.2 (3.5.5) post-amendment-2: replaces non-existent
episode-level metadata with per-bar signal that fires when
dd_pct(t) < dd_pct(t-1) AND dd_pct(t-1) > DD_TRAJECTORY_FLOOR — i.e.
transition is part of a recovery from non-trivial DD.

PER sampler (Phase 3.5.5.b follow-up) will read this and weight:
  sampling_weight = base × (1 + RECOVERY_OVERSAMPLE_WEIGHT × signal)
so recovery transitions get amplified gradient signal, completing the
downward-spiral break-out chain (3.5.2 reward asymmetry → 3.5.3
cooldown gate → 3.5.4 plasticity → 3.5.5 PER recovery curriculum).

3 ISV slots: 439 DD_TRAJECTORY_DECREASING, 440 RECOVERY_OVERSAMPLE_
WEIGHT (2.0 sentinel; ISV-driven from current dd_pct in follow-up),
441 DD_TRAJECTORY_FLOOR (0.02 sentinel; ISV-driven 25th percentile
of running dd_pct distribution in Phase 3.5.5.c follow-up per
feedback_isv_for_adaptive_bounds).

New sp15_dd_trajectory_prev_dd MappedF32Buffer (size 1) tracks
prev_dd across kernel calls — mirrors Task 3.5.3 sp15_cooldown_
consecutive_losses non-ISV mapped-pinned scratch pattern.

4 fold-reset registry entries + dispatch arms (3 ISV + 1 scratch).

Per established Phase precedent: kernel + launcher land first; PER
sampler integration and 25th-percentile floor producer are purely
additive follow-ups per feedback_no_partial_refactor.

Anchor test 2.10 recovery_after_streak (Phase 2C / Phase 3.5 paired) —
fully green via 3.5.2 + 3.5.4 + 3.5.5 combined.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-06 18:19:53 +02:00
jgrusewski
e0e0abfb28 feat(sp15-p3.5.4): plasticity injection trigger + warm-up tracker (weight-reset deferred)
Per spec §9.2 (3.5.4) post-amendment-2 fix. TWO-STEP recovery:
  1. Fire when DD_PERSISTENCE > PLASTICITY_PERSISTENCE_THRESHOLD AND
     PLASTICITY_FIRED_THIS_FOLD == 0 → set fired flag, set warm-bars
     counter to M_warm (default 200). [DEFERRED: reset last 10% of
     advantage-head weights to Kaiming-He init via cuRAND.]
  2. Per-bar warm-bars decrement; action-selection layer (consumer wiring
     follow-up) reads max(COOLDOWN_BARS_REMAINING, PLASTICITY_WARM_BARS_
     REMAINING) and forces Hold while > 0.

Weight-reset DEFERRED to Phase 3.5.4.b — kernel signature plumbed
(advantage_head_weights + n_weights) but no-op via (void) cast.
Documented in audit doc.

3 ISV slots: 436 PLASTICITY_FIRED_THIS_FOLD (debounce flag, resets at
fold boundary to re-arm next fold), 437 PLASTICITY_PERSISTENCE_THRESHOLD
(initial 100.0 sentinel; ISV-tracked from running mean of dd_persistence
in follow-up), 438 PLASTICITY_WARM_BARS_REMAINING (counter, OR-gates
with cooldown).

3 fold-reset registry entries + dispatch arms.

Three GPU oracle tests pass: fires-when-persistence-exceeds-threshold
(warm_bars [198, 200] post-fire-and-decrement), debounced-within-fold
(no re-fire when fired=1; warm decrements 50→49), no-fire-below-threshold.

Anchor test 2.22 plasticity_cooldown_interlock (Phase 2C / Phase 3.5
paired) — green via 3.5.4.b follow-up + action-selection consumer wiring.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-06 17:37:56 +02:00
jgrusewski
649128e739 feat(sp15-p3.5.3): cooldown gate — K=5 force Hold for M=20 bars (initial sentinels)
Per spec §9.2 (3.5.3). After K consecutive losing trades, force Hold for
M bars. Counter at COOLDOWN_BARS_REMAINING (slot 435) part of state —
model can reason about it.

Initial K=5, M=20 hardcoded sentinels. ISV-driven K via MEDIAN_STREAK_
LENGTH (slot 442) producer using two-heap median tracking is documented
Phase 3.5.3 follow-up. ISV-driven M from vol_normalizer time-to-mean-
reversion is also follow-up.

4 ISV slots: 433 K_THRESHOLD, 434 M_BARS, 435 BARS_REMAINING,
442 MEDIAN_STREAK_LENGTH. New sp15_cooldown_consecutive_losses
MappedF32Buffer tracks streak counter persistent across kernel calls.

5 fold-reset registry entries + dispatch arms (4 ISV slots + scratch
buffer reset).

Per spec post-amendment-2 fix: streak counter only updates on trade-close
events; per-bar non-close calls just decrement the cooldown counter. The
trigger gate fires only when a trade-close lands during an inactive
cooldown — re-arming mid-cooldown would extend the gate every closed
trade during the freeze, which is not the spec.

Per established Phase precedent: kernel + launcher land first; action-
selection wiring (force Hold while cooldown_remaining > 0) deferred to
follow-up commit per feedback_no_partial_refactor.

Anchor test 2.12 cooldown_engagement (Phase 2C / Phase 3.5 paired) —
green via this commit.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-06 17:14:53 +02:00
jgrusewski
aa91ce4d82 feat(sp15-p3.5.2): asymmetric reward under DD — gain × (1 + λ × dd_pct), pre-SP12-cap
Per spec §9.2 (3.5.2). For gains: r_adjusted = r × (1 + λ × dd_pct).
For losses: unchanged. Multiplier applies BEFORE SP12 NEG/POS cap clamp;
saturating the cap for big recovery trades is behaviorally correct
per spec (encourages frequent small recoveries).

3 ISV slots: 430 DD_ASYMMETRY_LAMBDA (initial 0.5; ISV-tracked from
running DD variance via DD_DIST_VAR is follow-up), 431 R_GAIN_DD_BOOST
(diagnostic — most recent boost), 432 DD_DIST_VAR (running variance,
producer follow-up).

3 fold-reset registry entries + dispatch arms.

Per established Phase precedent: kernel + launcher land first; reward
composer site (apply multiplier BEFORE SP12 cap) deferred to follow-up
commit per feedback_no_partial_refactor.

Anchor test 2.10 recovery_after_streak (Phase 2C / Phase 3.5 paired) —
green via 3.5.2 + 3.5.3 + 3.5.4 + 3.5.5 combined.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-06 16:58:39 +02:00
jgrusewski
5d36f3238c feat(sp15-p3.5): confidence-aware Hold floor — bounded sigmoid
Per spec §8.2 (3.5) post-amendment-2 fix. hold_floor = α × σ(k × (entropy − ε₀))
added to Q_hold pre-argmax/Thompson selection. Hold becomes uncertainty
expression, not distributional default.

α (HOLD_FLOOR_ALPHA slot 426) initial 0.5 sentinel — producer kernel
updating from rolling 95th percentile of |Q_dir| (NOT running max —
outlier-ratchet vulnerable per spec second-review #6) is documented
Phase 3.5 follow-up.
k (HOLD_FLOOR_K slot 427) initial 10.0 — producer from running variance
of entropy is follow-up.
ε₀ (HOLD_FLOOR_EPS0 slot 428) initial 1.0 — producer from 75th percentile
of entropy distribution (ENTROPY_DIST_REF slot 429) is follow-up.

4 fold-reset registry entries + dispatch arms.

Per established Phase precedent: kernel + launcher land first; action-
selection wiring (add hold_floor to Q_hold pre-argmax/Thompson) deferred
to follow-up commit per feedback_no_partial_refactor.

Anchor tests: 2.1 flat_market_holds + 2.6 regime_silences (Phase 2B
contracts) — green via this teaching's action-selection wiring.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-06 16:12:17 +02:00
jgrusewski
8bfc480d92 feat(sp15-p3.4): regret signal = r_discipline content + first-obs bootstrap EMA
Per spec §8.2 (3.4). When policy holds AND a trade would have been
profitable past cost+trail, regret_bar = ideal_pnl_missed - cost_t.
EMA tracked at REGRET_EMA (slot 423) with first-observation bootstrap
(per pearl_first_observation_bootstrap) + simple exponential decay
(α=0.05; Wiener-α adaptive swap is a documented follow-up).

r_discipline_per_bar = -LAMBDA_REGRET × REGRET_EMA composed at the
reward-split site in the follow-up consumer commit per
feedback_no_partial_refactor.

3 ISV slots: 423 REGRET_EMA, 424 LAMBDA_REGRET (initial 1.0), 425
REGRET_GRAD_NORM. 3 fold-reset registry entries + dispatch arms.

Anchor test 2.6 regime_silences (Phase 2B contract) — green via
3.4 + 3.5 combined.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-06 16:02:59 +02:00
jgrusewski
7753cbef1b feat(sp15-p3.3): quadratic DD penalty + ISV-driven λ + threshold
Per spec §8.2 (3.3). penalty = λ_dd × max(0, dd_current − dd_threshold)²

Asymmetric: zero below threshold, quadratic growth above. Encodes loss
aversion per pearl_audit_unboundedness_for_implicit_asymmetry.

3 ISV slots: 420 LAMBDA_DD (initial 1.0; ISV-tracked from grad-balance
in follow-up), 421 DD_THRESHOLD (initial 0.05 = 5% drawdown trigger),
422 DD_PENALTY_GRAD_NORM (initial 0.0).

3 fold-reset registry entries + dispatch arms.

Per established Phase precedent: kernel + launcher land first; reward
composition site (subtract penalty from r_total) deferred to follow-up
commit per feedback_no_partial_refactor.

Anchor test 2.5 drawdown_de_risks (Phase 2C / Phase 3.5 paired) — green
via Phase 3.5 mechanisms; this commit lands the penalty primitive.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-06 15:49:45 +02:00
jgrusewski
1eac41d644 feat(sp15-p3.2): explicit cost in r_quality on trade-close events
Per spec §8.2 (3.2). Extends the Phase 3.1 composer kernel
`r_quality_discipline_split_kernel` with two new args (`float cost_t`,
`unsigned int trade_close_indicator`) and subtracts
`cost_t * (float)trade_close_indicator` from `r_quality` BEFORE the α
blend. Same `cost_t` scalar shape as the Phase 1.2 cost_net_sharpe
accumulator (commission + per-side spread + OFI-impact); the gate
`trade_close_indicator=1` on round-trip-close bars (rt_ind=1) and 0
otherwise so non-close bars receive a structural no-op identical to
the pre-Phase-3.2 behaviour. Model SEES the bill in the gradient
signal during training, not just in eval-time metrics.

Approach: kernel-signature-extension (NOT wrapper-kernel). The only
existing call site in the tree is the Phase 3.1 oracle test
(`training_loop.rs` has dispatch-arm reset wiring but does NOT yet
invoke the launcher per the Phase 3.1 commit's deferred-consumer
note), so the cascade is bounded to that single test — extending the
existing kernel is cleaner than a parallel wrapper that would have to
be retired the moment the Phase 3.1 deferred consumer migration
lands.

Anchor test 2.4 cost_sensitivity (Phase 2B contract) — green via this
commit + Phase 3.1 split structure (already landed in 2d226e6e7).

Per established Phase precedent: kernel/launcher signature change +
existing-test migration land atomically per
feedback_no_partial_refactor; production reward-composition wire-up
that feeds real `cost_t` from cost_net_sharpe is the same deferred
follow-up Phase 3.1 declared (no new debt added — both share one
follow-up commit).

cargo check -p ml --features cuda: clean (18 pre-existing warnings).
cargo test -p ml --test sp15_phase1_oracle_tests --features cuda
  -- --ignored r_quality_subtracts_explicit_cost: 1 passed.
cargo test -p ml --test sp15_phase1_oracle_tests --features cuda
  -- --ignored r_split_uses_sentinel_alpha_at_cold_start: 1 passed
  (Phase 3.1 sentinel test post-migration).
cargo test -p ml --features cuda: 946 passed / 13 failed
  (same 13 failures as parent 2d226e6e7; zero introduced).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-06 15:36:58 +02:00
jgrusewski
2d226e6e76 feat(sp15-p3.1): r_quality + r_discipline split with ISV-driven α + sentinel cold-start
Per spec §8.2 (3.1) post-amendment-2 fix: ALPHA_SPLIT slot initialized
DIRECTLY to 0.5 in trainer constructor. Formula α = grad_norm_q /
(grad_norm_q + grad_norm_d + ε) takes over only after BOTH grad-norm
EMAs accumulate ≥ N_WARM=100 non-zero observations.

Two kernels in r_quality_discipline_split_kernel.cu (single cubin per
established 1:1-source-to-cubin pattern with multiple kernels):
  - r_quality_discipline_split_kernel: per-step composition + warm count
  - alpha_split_producer_kernel: per-step ALPHA_SPLIT update from grad ratio
    (gated on warm count to prevent premature formula activation)

3 ISV slots (417 ALPHA_SPLIT, 418 GRAD_NORM_QUALITY, 419 GRAD_NORM_DISCIPLINE)
+ sp15_alpha_warm_count [1] mapped-pinned scratch buffer on the trainer
struct. 4 fold-reset registry entries + dispatch arms (one for the
non-ISV warm-count buffer mirrors the sp11_novelty_hash host_slice_mut
pattern).

Per established Phase precedent: kernels + launchers land first; consumer
migration (per-step launches in training_loop.rs reward composition site)
deferred to a follow-up commit per feedback_no_partial_refactor.

Anchor tests: 2.4 cost_sensitivity + 2.6 regime_silences (Phase 2B
contracts) — green via Phase 3.4 regret + 3.2 cost; this commit lands
the split structure they depend on.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-06 15:26:56 +02:00
jgrusewski
ef373c34d7 feat(sp15-p1.7): consume the abandoned walk-forward test slice (stash + observer; eval invocation deferred)
Per spec §6.7. The walk-forward generator emits `test_start..test_end`
per fold but the trainer at `mod.rs:1294` only consumed train+val — the
12.5% test slice was silently dropped, the model was never measured on
held-out data the train/val pipeline didn't see.

This commit lands the foundation: `set_test_data_from_slices` stashes
the per-fold range immediately after `set_val_data_from_slices`, gated
on `fold.test_end > fold.test_start` for defensively-empty slices. A
`set_test_data_observer` hook lets unit tests verify the wiring without
spinning up a full GPU eval pipeline.

The actual `evaluate_dqn_graphed` invocation against the stashed slice
plus the per-fold `HEALTH_DIAG test_slice fold=K test_sharpe_net=...`
emit is deferred to a follow-up commit per `feedback_no_partial_refactor`.
Wiring it through requires either standing up a second
`GpuBacktestEvaluator` instance (parallel to the val one at
`metrics.rs:550`) or refactoring the existing val evaluator to swap
window data between val and test eval — the val evaluator's lazy-init
path is fundamentally tied to the window passed at construction. Plus
TLOB weight sync, ISV signal wiring, and a `training_mode` toggle (no
such field exists yet on `DQNTrainer`).

This deferral matches the Phase 1.5 (kernel + launcher first, trunk
consumer follow-up) and Phase 1.6 (stash dev/holdout slices, eval
consumer follow-up) precedents on this branch. The stash + observer
surface is the analogous foundation; the L40S smoke once Task 1.7.b
lands will surface the per-fold `test_sharpe_net` HEALTH_DIAG line as
the canonical end-to-end verifier.

New oracle test `set_test_data_from_slices_fires_observer_and_stashes`
in `sp15_phase1_oracle_tests.rs` constructs a real trainer (sync init,
no GPU forward), registers an observer, exercises the API with a
synthetic [5000..6000) range, asserts the observer fires once with the
right bounds. Passes locally on RTX 3050 Ti.

`docs/dqn-wire-up-audit.md` extended with a Phase 1.7 entry documenting
what landed, what's deferred, the wire-up locations, and the rationale.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-06 14:53:05 +02:00
jgrusewski
ce019c72d2 feat(sp15-p1.6): --holdout-quarters + --dev-quarters CLI flags + sealed Q1-Q7/Q8/Q9 split
Per spec §6.6 / Q6 train/dev/test split (defaults Q1-Q7 train, Q8 dev,
Q9 sealed final test):

- DQNHyperparameters: holdout_quarters + dev_quarters (default 1+1)
- crates/ml/examples/train_baseline_rl.rs: --holdout-quarters /
  --dev-quarters CLI flags forwarded to hyperparams (this is the actual
  training binary; bin/fxt/src/commands/train.rs is a gRPC client and
  services/ml_training_service/src/main.rs accepts training params via
  proto not CLI — see audit doc note).
- DQNTrainer::train_walk_forward slices training_data BEFORE fold
  generation; folds run on Q1..Q(9 - holdout - dev) only.
- DQNTrainer struct: dev_features/dev_targets/holdout_features/
  holdout_targets fields stash trailing slices for end-of-training dev
  eval and the Phase 4.3 separate eval-only workflow.
- debug_assert sealed-slice guard catches future refactors that
  re-introduce holdout into the training path.

Per established Phase 1 precedent (1.1-1.5: kernel/state lands first,
consumer wiring deferred to follow-up commit per
feedback_no_partial_refactor): CLI plumbing + slicing + dev/holdout
storage land in this commit. The post-final-fold dev evaluation call
(consumer of dev_features) is deferred to a follow-up commit and will
mirror Task 1.7's evaluate_dqn_graphed integration pattern. Phase 4.3
argo-eval-final.sh is the sole legitimate consumer of holdout_features
(separate eval-only workflow that does NOT call train_walk_forward).

cargo check -p ml --features cuda --example train_baseline_rl: clean
cargo check -p fxt: clean (no fxt changes needed; gRPC client only)

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-06 14:39:37 +02:00
jgrusewski
5309d4bee5 feat(sp15-p1.5): dd_pct foundational state input concat kernel — LAYOUT FINGERPRINT BREAK
LAYOUT FINGERPRINT BREAK: pre-SP15 checkpoints WILL NOT LOAD after this
commit. Greenfield OK per spec Q1.

Per spec §6.5: dd_pct (slot 406, written by Task 1.3 dd_state_kernel)
gets concatenated to the trunk forward input as the last dim. Eval-time
policy SEES drawdown context on every forward pass; Phase 3 teachings
can condition on dd_pct directly via state, not just reward modulation.

New kernel `dd_pct_concat_kernel.cu` produces a [B, state_dim_padded + 1]
buffer whose leading state_dim_padded columns equal the input states_buf
and whose +1 last column equals isv[DD_PCT_INDEX=406] broadcast across
batch. Pure scatter-copy (no atomicAdd per feedback_no_atomicadd).

layout_fingerprint_seed extended with 'TRUNK_INPUT_DD_PCT=sp15_phase_1_5'
marker — FNV1a hash changes; old checkpoints fail to load with the
existing layout-mismatch error path (same fail-fast that fired on the
SP4 / SP14 layout breaks).

Phase 1.5 lands kernel + launcher + layout fingerprint marker + GPU
oracle test only. Trunk consumer migration (re-pointing forward_online
to consume the concat buffer + bumping s1_input_dim from 48 → 49 +
propagating through GRN encoder, VSN gate input, bottleneck path, and
backward dx scratch) is deferred to a follow-up atomic commit per
feedback_no_partial_refactor — matches the established Phase 1.1-1.4
precedent (kernels + launchers verify in isolation first; consumer
migration is a load-bearing change touching the GRN encoder, VSN
partition boundaries, fxcache schema, and backward gradient flow).

GPU oracle test `dd_pct_concat_kernel_writes_last_column` validates B=4,
raw state_dim=48, state_dim_padded=128: leading 128 columns of each
output row match the input states (including pad zeros), and the 129th
column equals isv[DD_PCT_INDEX]=0.42 broadcast across all 4 rows. Test
passes on local RTX 3050 Ti (sm_86) in 1.62s.

cargo test -p ml --lib --features cuda: 945 passed / 14 failed — same
14 failures pre-existing on the parent commit `c6fd4b4b2` (Task 1.4
partial baseline); zero introduced by this commit.

Per spec §6.5 step 7 (trunk-grounding behavioral test): Phase 4 L40S
smoke verifies dd_pct propagation at production scale via the existing
layout-fingerprint-mismatch fail-fast on cold-start of any pre-SP15
checkpoint. The follow-up Phase 1.5.b commit that lands the consumer
migration adds the explicit trunk-grounding KL test alongside the
forward_online wiring.

Touched:
- crates/ml/src/cuda_pipeline/dd_pct_concat_kernel.cu (new)
- crates/ml/build.rs (+1 cubin manifest entry)
- crates/ml/src/cuda_pipeline/gpu_dqn_trainer.rs (+SP15_DD_PCT_CONCAT_CUBIN
  + launch_sp15_dd_pct_concat + TRUNK_INPUT_DD_PCT layout fingerprint
  marker)
- crates/ml/tests/sp15_phase1_oracle_tests.rs (+1 GPU oracle test)
- docs/dqn-wire-up-audit.md (+1 Phase 1.5 entry)

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-06 14:31:16 +02:00
jgrusewski
c6fd4b4b2a feat(sp15-p1.4-partial): 4 constant-policy baselines (buyhold, hold_only, momentum, reversion)
Per spec §6.4. Pure-CUDA kernels in baseline_kernels.cu — single cubin,
4 extern "C" __global__ functions sharing a templated compute_baseline_sharpe
helper. Trunk-shared baselines (random_dir_kelly slot 411, aux_only slot 413,
mag_quarter_fixed slot 414, trail_only slot 415) deferred to Task 1.4.b
follow-up — they need partial-policy-forward access from the main eval pass.

ISV slots written: 409 (buyhold), 410 (hold_only), 412 (naive_momentum),
416 (naive_reversion). Slots 411/413/414/415 stay at sentinel 0.0 until
follow-up commit.

Per established Phase 1 precedent: kernels + launchers land first;
per-eval-pass launches + HEALTH_DIAG baseline_deltas emit deferred to
follow-up commit per feedback_no_partial_refactor.

Anchor tests: buyhold positive on +drift, hold_only emits 0,
momentum + reversion sum near zero on mean-reverting series.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-06 14:10:20 +02:00
jgrusewski
9e84602486 feat(sp15-p1.3): drawdown state kernel — DD_CURRENT/MAX/RECOVERY/PERSISTENCE/CALMAR/PCT
Per spec §6.3. Per-step kernel reads PS_PEAK_EQUITY (slot 7) and
PS_PREV_EQUITY (slot 9) from existing position state buffer (no new
equity slot needed). Writes 6 ISV slots (401-406). Calmar uses
max(dd_max, 1e-4) floor — eliminates the saturation-at-100 artifact
seen in train-dd4xl HEALTH_DIAG.

6 fold-reset registry entries + dispatch arms. 5 sentinel-0 stateful
outputs; calmar uses sentinel 1e-4 (same value as the kernel's floor)
so cold-start division uses the floor rather than ±inf.

Atomic split per feedback_no_partial_refactor.md (mirrors Task 1.1 +
1.2 precedent): kernel + launcher + registry land here; per-step
production wire-up + HEALTH_DIAG composer deferred to a follow-up
commit. Test 1.3 oracle on 6-step synthetic equity curve passes
locally on RTX 3050 Ti (sm_86) in 1.61s. cargo test -p ml --lib
--features cuda: 946 passed / 13 failed — same 13 pre-existing
failures as Task 1.2 baseline (a92ff28a9); zero introduced.
State-reset registry tests: 4/4 pass.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-06 13:56:57 +02:00
jgrusewski
a92ff28a98 feat(sp15-p1.2): cost-net sharpe kernel — commission + spread + OFI-impact
Per spec §6.2 per-side semantics:
  cost_t = commission_per_rt × rt_ind[t]
         + half_spread[t] × |pos[t]| × side_ind[t]
         + ofi_lambda × |pos[t]| × |ofi[t]| × side_ind[t]

Commission charged at close; half-spread × |pos| at entry AND exit
(sums to one full spread per RT); OFI impact same per-side. Initial
λ=2.0e-4 in ISV[OFI_IMPACT_LAMBDA_INDEX=407] as Invariant-1 anchor
(constructor-write + FoldReset rewrite per feedback_isv_for_adaptive_bounds);
per-fold ISV refit may overwrite in later phases. Mean cost-per-bar
emitted to ISV[COST_PER_BAR_AVG_INDEX=408] (stateful kernel output;
FoldReset sentinel 0 + Pearl A first-observation bootstrap).

Same kernel reads LobBar fields from synthetic markets (Phase 2A) and
real fxcache LOB (prod) — dev/prod parity per Q3.

Phase 1.2 lands kernel + launcher + anchor seed + 2 registry entries
+ 2 dispatch arms only; consumer migration deferred to a follow-up
commit per feedback_no_partial_refactor (mirrors Phase 1.1 atomic
pattern). Oracle test cost_net_sharpe_round_trip_charges_full_spread
validates single round-trip cost = 2.00 / 10 bars = 0.20/bar with
mean_pnl_net = 0.80 on 1.69s RTX 3050 Ti (sm_86). Zero regressions
introduced (946 passed / 13 pre-existing failures, same as Task 1.1).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-06 13:46:03 +02:00
jgrusewski
3667cd1b04 feat(sp15-p1.1): unified sharpe kernel — single formula for train and val
Replaces sharpe_ema (per-batch EMA, train) vs sharpe_annualised (val ×
sqrt(525600)) split. Single GPU kernel computes mean/std/sharpe via
2-pass block-tree-reduce; annualisation at the call site.

Per feedback_no_partial_refactor: consumer migration in metrics.rs /
training_loop.rs deferred to a follow-up commit; kernel + launcher land
in isolation first. Verified via 2 GPU oracle tests
(unified_sharpe_kernel_zero_mean, unified_sharpe_kernel_positive_drift)
passing on local RTX 3050 Ti (sm_86) in 1.78s.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-06 11:41:00 +02:00
jgrusewski
7d0a29dced feat(sp15-p2a.1): LobBar canonical ABI + 4 synthetic market generators
Phase 2A scaffolding lands FIRST per spec §4.4 ABI contract. Phase 1.2
cost kernel reads LobBar; both dev synthetic and prod fxcache produce
LobBar — dev/prod parity per Q3.

Generators: flat_market, drift_market, ou_market, regime_switch_market
(seeded RNG for reproducible tests). Regime-switch test uses sticky
0.99/0.01 transitions (true regime persistence; spec's 50/50 was a
random walk, not a regime switch — corrected with code comment).

behavioral_suite test target wired into Cargo.toml; will run all
22 Phase 2 tests once they land in Phase 2B/2C.

Audit doc: SP15 Phase 2A.1 entry appended to docs/dqn-wire-up-audit.md
per Invariant 7 (component changes require audit-doc update).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-06 10:57:51 +02:00
jgrusewski
c146c4fffd feat(sp15): scaffold sp15_isv_slots.rs with 46 slots [397..443) — ISV_TOTAL_DIM 396→443
Per spec §4.3 allocation map. Pre-allocates disjoint slot ranges to
enable Approach B parallel sub-worktrees without index collisions:
  - Phase 0.B EGF retune: [397..401)
  - Phase 1.3 drawdown: [401..407)
  - Phase 1.2 cost: [407..409)
  - Phase 1.4 baselines: [409..417)
  - Phase 3.X-3.5.X teachings + recovery: [417..441)
  - Phase 3.5 deferred anchors: [441..443)

Layout fingerprint extended with all 46 slot names. Pre-SP15 checkpoints
will be incompatible (greenfield OK per Q1).

Two regression tests verify: (1) every slot < ISV_TOTAL_DIM, (2) layout
fingerprint locked at named indices. docs/isv-slots.md gets the SP15
section documenting the allocation map + greenfield sub-worktree plan.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-06 10:49:33 +02:00
jgrusewski
c0fc28e455 fix(sp14): delete warmup_gate — let variance-driven k_aux/k_q handle warmup (ISV-driven)
Per `feedback_isv_for_adaptive_bounds`, the hardcoded
`warmup_gate = (fold_step_counter / WARMUP_STEPS_FALLBACK).min(1.0)`
ramp violated the rule: adaptive bounds in ISV, never hardcoded
constants. The variance-driven k_aux/k_q sigmoid steepness already
provides warmup behavior intrinsically:

- High variance (cold-start, EMAs still moving) → k → K_MIN → flat
  sigmoid → gate ≈ 0.5 regardless of input. That IS the warmup.
- Low variance (settled) → k → K_BASE → sharp sigmoid → gates
  respond correctly to driver signals.

Adding a separate hardcoded step-counter multiplier on top was
double-counting + tuning-driven (the 1000-step threshold had no
principled basis). Removed entirely.

Removed (per `feedback_no_partial_refactor`, all atomically):
- `WARMUP_STEPS_FALLBACK` constant in `sp14_isv_slots.rs`
- `warmup_gate: f32` parameter in `alpha_grad_compute_kernel.cu`
- `gate1 * gate2 * warmup_gate` → `gate1 * gate2` in kernel
- `warmup_gate` arg from `launch_sp14_alpha_grad_compute`
- `fold_step_counter: usize` field on the trainer struct
- `fold_step_counter = 0` reset in `reset_for_fold`
- `fold_step_counter` init in trainer constructor
- `let warmup_gate: f32 = 1.0;` and `.arg(&warmup_gate)` from B.4
  oracle tests (4 launches: 2 in alpha_grad_schmitt_hysteresis,
  20 in alpha_grad_adaptive_beta loop)

Build: clean, 18 warnings (pre-existing baseline).
Tests: cargo test --no-run on sp14_oracle_tests succeeds.

Net result: EGF gate's warmup behavior now lives entirely in the
variance-driven k_aux/k_q sigmoid steepness controller (ISV slots
388/var_aux, 389/var_q). No hardcoded step counter. Honors
`feedback_isv_for_adaptive_bounds` and `pearl_controller_anchors_isv_driven`.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-05 22:07:55 +02:00
jgrusewski
60ad42676e fix(sp14): bump ISV_TOTAL_DIM 383 → 396 to cover SP14 EGF slots
ROOT CAUSE of L1+L2 from Smoke A2-B: the ISV bus was sized for top
of SP13 (ISV_TOTAL_DIM=383) but B.1 allocated SP14 slots at 383-395.
Every SP14 read/write was OUT-OF-BOUNDS memory access. That's why:

- gate1 (slot 391) read as 0 always (OOB zero-init memory)
- post_open_min (slot 394) accumulated garbage values 9.5 → 28 → 46
- α_smoothed/α_raw values appeared to work but were undefined behavior

SP4/SP5 had a regression test (`all_sp4/5_slots_fit_within_isv_total_dim`)
that catches this exact failure mode at unit-test time. SP14 was missing
it — that gap let the bug ship across all 16 commits without being caught.

Changes:
- ISV_TOTAL_DIM: 383 → 396 (covers SP14 slots 383-395)
- layout_fingerprint_seed: extended with SP14 slot names + new
  ISV_TOTAL_DIM=396 marker (forces fingerprint hash bump per
  Invariant 8 — old checkpoints invalidated correctly)
- sp14_isv_slots.rs: 2 regression tests (mirror SP4/SP5 patterns)

Both tests pass. After this fix, SP14 EGF kernels will read/write
the correct slots; gate1 should actually flip open when aux_dir_acc
crosses target+0.03; gradient_hack_detect post_open_min stays bounded
in [0, 1] as designed.

NOT yet addressed (separate follow-up):
- warmup_gate hardcoded WARMUP_STEPS_FALLBACK=1000 violates
  feedback_isv_for_adaptive_bounds. Should be ISV-signal-driven OR
  removed entirely (k_aux/k_q already provide variance-driven warmup).
  Redesign post-re-smoke once bus-size fix is verified.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-05 21:59:47 +02:00
jgrusewski
e41dbb7d8a diag(sp14 B.12): per-epoch pearl_egf_diag HEALTH_DIAG emit
Adds a new HEALTH_DIAG[{epoch}]: pearl_egf_diag line immediately after
the aux_moe block in the per-epoch metrics section of training_loop.rs.
Reads all 13 SP14 ISV slots [383..396) — α_smoothed, α_raw, β, k_aux,
k_q, var_aux, var_q, var_α, q_dis_short, q_dis_long, gate1 state,
post_open_min, lockout — via the established read_isv_signal_at pattern,
giving forensic visibility into EGF pearl state each epoch.

gate1/gate2 sigmoid outputs are intentionally omitted: recomputing them
host-side would violate feedback_no_cpu_compute_strict; the sigmoid
inputs are sufficient for a reader to infer the output values.

docs/isv-slots.md updated (Invariant 7): records B.12 HEALTH_DIAG wire-up.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-05 21:14:58 +02:00
jgrusewski
857722e774 feat(sp14): B.11 — orchestrator wire-up for 3 EGF producer kernels + var_aux gap closure
Per-step launches (in graph capture order):
  1. Forward (existing)
  2. Action select (existing) → q_dir_logits available
  3. launch_sp14_q_disagreement_update → ISV[383, 384, 389]
  4. launch_sp14_alpha_grad_compute → ISV[385..395] (consumes q_disagreement)
  5. Backward (existing) — wire-col scale at B.10 reads ISV[393]

Per-epoch launch (end of epoch):
  6. launch_sp14_gradient_hack_detect → circuit breaker

α_short=0.3, α_long=0.05, α_var=0.05 per spec; warmup_gate derived from
steps_in_fold / WARMUP_STEPS_FALLBACK.

Var_aux producer gap closed (option C from B.4): alpha_grad_compute_kernel
now also writes ISV[VAR_AUX_INDEX=388] via Welford EMA against
(aux_dir_acc_short - aux_dir_acc_long). Adaptive k_aux is now functional
(was degenerate at K_BASE_AUX=20.0 constant pre-B.11). Closes the
"adaptive_k_aux currently degenerate" concern flagged in B.4 commit.

After this commit, the EGF pearl is FULLY ACTIVE end-to-end:
- Forward: aux signal feeds direction Q-head input (B.8/B.9)
- Backward: wire-col gradient gated by α_grad_smoothed (B.10)
- Producers: α_grad computed every step from real driver signals (B.11)
- Pre-B.11 force-closed gate (sentinel 0.0) → post-B.11 responsive gate

Build clean: 18 warnings pre-existing baseline, 0 new.
Tests: 4/4 P0b aux_w tests pass (no regression).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-05 21:11:26 +02:00
jgrusewski
dc3f948ee9 feat(sp14): B.10 — backward wire gradient gating by ALPHA_GRAD_SMOOTHED
Critical safety mechanism that completes the EGF pearl: scales the wire
column of `dL/dx_concat [B, SH2 + 1]` (the gradient flowing FROM the
direction Q-head's first FC SGEMM TO `aux_softmax_diff`) by
`ISV[ALPHA_GRAD_SMOOTHED_INDEX = 393]`, computed by B.4's
`alpha_grad_compute_kernel` and orchestrated per-step in B.11.

`dL/dW[wire_col]` (Q-head's own weight gradient for the appended column)
is NOT scaled — the dW SGEMM `dY^T × x_concat` and the dX SGEMM
`dY × W^T` are independent, so scaling `dx[:, SH2]` AFTER both have
completed leaves dW unaffected. Q-head learns to USE the wire freely;
only the gradient PROPAGATING BACK to aux is gated.

Pre-B.11 (no producer wired) `ISV[393]` holds sentinel `0.0` →
wire force-closed (gradient zeroed) — the conservative safety state.
Post-B.11, B.4 writes the live gate output ∈ [0, 1] each step.

Closes the latent K-mismatch B.8/B.9 left in backward
============================================================

B.8 grew `w_b0fc` to `[adv_h, SH2 + 1]` end-to-end (Adam m/v +
spectral-norm vector + smoke fixtures); B.9 closed the forward dispatch
K-mismatch. The backward dW/dX SGEMMs for `d == 0` still used `K = SH2`
against the new `LDA = SH2 + 1` weight tensor — silently dropping the
last column of dW and zeroing the wire-col gradient. B.10 closes that
gap atomically with the wire-col scale per `feedback_no_partial_refactor`:

  * `backward_branch_dw` for `d == 0` now uses `(dir_qaux_concat_ptr, SH2 + 1)`
    instead of `(save_h_s2, SH2)` — matching the forward consumer
    pattern from B.9.
  * `backward_branch_dx` for `d == 0` now writes to
    `d_dir_qaux_concat [B, SH2 + 1]` with `K = SH2 + 1` instead of
    `scratch_d_h_s2 [B, SH2]` with `K = SH2`. Mirrors the magnitude
    branch's wider-buffer pattern.

New artifacts
=============

  * `sp14_scale_wire_col_kernel.cu`: one thread per batch row, scales
    `dx_concat[b, SH2]` by `isv[393]` IN-PLACE. NaN-safe per the
    `dqn_scale_f32_kernel` precedent (explicit `α==0 ⇒ 0` branch).
    Pure per-thread map, no atomicAdd, no shared memory.
  * `sp14_d_dir_qaux_concat: CudaSlice<f32>` `[B, SH2 + 1]` trainer-
    struct field. Dx SGEMM destination; the wire-col scale acts on
    this buffer; the strided accumulator copies the first SH2 columns
    into `bw_d_h_s2` after the scale.
  * `launch_sp14_scale_wire_col` launcher reads `self.isv_signals_dev_ptr`
    and the new buffer's raw_ptr.
  * `backward_full` signature grows two trailing `u64` args
    (`dir_qaux_concat_ptr`, `d_dir_qaux_concat_ptr`); both
    `backward_full` call sites (CQL aux + main online) wired
    atomically per `feedback_no_partial_refactor`.

Post-call orchestration at trainer level
========================================

  1. `launch_sp14_dir_concat_qaux(save_h_s2)` rebuilds the ONLINE
     concat in `sp14_dir_qaux_concat_scratch` (the forward pass had
     overwritten it with the TARGET concat at line ~25817). Same
     one-step-lag semantic preserved — `aux_nb_softmax_buf` is
     unchanged between forward and backward.
  2. `cuMemsetD32Async` zero of `d_h_s2` — pre-B.10 the direction
     branch (d==0) wrote it with beta=0; post-B.10 the dir-Q dX
     lives in `d_dir_qaux_concat` and is gated + accumulated AFTER
     `backward_full` returns, so the value-FC dx accumulator inside
     `backward_full` (beta=1) needs an explicit zero baseline.
  3. `backward_full` runs: dir branch → `d_dir_qaux_concat`,
     mag/ord/urg branches → their concat dX buffers, value-FC →
     `d_h_s2` (beta=1, on top of zeroed buffer).
  4. `launch_sp14_scale_wire_col` gates col SH2 of `d_dir_qaux_concat`.
  5. `accumulate_d_h_s2_from_concat` (beta=1) copies first SH2 cols
     of `d_dir_qaux_concat` into `d_h_s2`. Wire col stays in
     `d_dir_qaux_concat[:, SH2]`, untouched by this accumulator (its
     destination range is [0, SH2)). Pre-B.11 the wire is already
     zeroed by the sentinel-α gate; the orchestrator that propagates
     the gated wire-col gradient back to the aux head's softmax CE
     backward chain lives in B.11.
  6. mag/ord/urg accumulators continue with beta=1 (comments updated).

Wire status
===========

  * Forward dispatch: unchanged (B.9-complete).
  * Backward dispatch: GATED on both call sites (CQL aux + main online).
  * dW unchanged: the `dW = dY^T × x_concat` SGEMM writes
    `grad_buf[goff_w_b0fc..]` BEFORE the scale-wire-col launches;
    the scale operates ONLY on `d_dir_qaux_concat` (the dx buffer)
    AFTER both dW and dX SGEMMs complete.
  * Target net unaffected: Polyak EMA-only, no backward.
  * CudaSlice wrapper path: passes `0u64` for both new args, falls
    back to the legacy K=SH2 path. Consistent with the forward
    wrapper's diagnostic-only residual.

Verified
========

  * `SQLX_OFFLINE=true cargo check -p ml` clean, 18 warnings (baseline)
  * `cargo test -p ml --test sp14_oracle_tests` 2 passed, 6 ignored (GPU)
  * Audit doc `docs/dqn-wire-up-audit.md` updated per Invariant 7.

After this commit, the EGF pearl is architecturally complete; the
orchestration of when/how the alpha_grad gates fire happens in B.11
(producer chain orchestrator).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-05 20:49:27 +02:00
jgrusewski
ecf4757c0d feat(sp14): B.9 — wire forward concat into direction Q-head SGEMM
Closes the latent SGEMM K-mismatch left by B.8 (6715ab4ea):
`w_b0fc` had grown from `[adv_h, SH2]` to `[adv_h, SH2 + 1]` end-to-end,
but every direction-Q-head consumer's SGEMM still used `K = shared_h2`
against the new `LDA = SH2 + 1` weight tensor — safe ONLY because the
new column was zero-init in B.8 and Adam had not yet updated it. After
this commit the forward wire is FULLY ACTIVE; the SGEMM consumes
`sp14_dir_qaux_concat_scratch [B, SH2 + 1]` with `K = shared_h2 + 1`.

Direction Q-head input pointer: `h_s2_buf` → `sp14_dir_qaux_concat_scratch`.
K dim: `shared_h2` → `shared_h2 + 1`.

Concat kernel runs immediately before the direction Q-head SGEMM in
the same stream, enforcing `pearl_canary_input_freshness_launch_order`.
Mirrors the `launch_mag_concat_from` precedent: the aux head forward
that writes `aux_nb_softmax_buf` runs AFTER the per-step online
forward (line ~25599 in the new layout), so each forward consumes
the PREVIOUS step's aux predictions — same one-step-lag semantic as
mag_concat. Step 0 sees alloc_zeros (uniform 0.5/0.5 → diff = 0),
step 1+ sees the prior step's aux next-bar softmax.

Atomic-migration consumers (`feedback_no_partial_refactor`):

- `gpu_dqn_trainer.rs` — new `launch_sp14_dir_concat_qaux` method;
  online forward (line ~25583) and target forward (line ~25758)
  each precede their `forward_*_raw` call with a concat launch and
  pass `sp14_dir_qaux_concat_scratch.raw_ptr()`. Both replay paths
  (`replay_forward_ungraphed`, `replay_forward_for_q_values`
  ungraphed fallback) get the same wire — they use online weights
  and produce direction Q-values consumed by training/eval. Causal
  intervention sites (×2) and DDQN argmax pass `0u64` per spec
  (their direction Q outputs are either unread by the consumer or
  the spec accepts the K=SH2 fallback's residual one-step bias).

- `batched_forward.rs` — five `forward_*_raw` / `launch_vsn_glu_branch`
  signatures grow a trailing `dir_qaux_concat_ptr: u64`; new
  `d == 0 && dir_qaux_concat_ptr != 0` branch in every legacy
  ReLU-FC FC dispatch (multi-stream / sequential × online / target /
  F32-output) returning `(dir_qaux_concat_ptr, self.shared_h2 + 1)`.
  VSN-GLU branch path scatters `vsn_masked` into the first SH2 cols
  of the scratch, identical to the `d == 1/2/3` scatter pattern
  (the trailing aux_softmax_diff column was already written by the
  pre-VSN concat-kernel launch and survives the scatter). The
  `CublasGemmSet::new` heuristic-cache shape table grows by one
  unique tuple `(adv_h, batch, SH2 + 1, SH2 + 1)` so the first-call
  cublasLt heuristic search hits a fresh cache slot instead of the
  pre-B.8 `(adv_h, batch, SH2, SH2)` entry.

- `gpu_experience_collector.rs` / `value_decoder.rs` — pass `0u64`
  for the new arg (no aux-head dependency on those forwards;
  documented inline with rationale).

- `docs/dqn-wire-up-audit.md` — new SP14 Layer B B.9 entry per
  Invariant 7, documenting every new dispatch site, the
  diagnostic-path residual, and the launch-order constraint.

After this commit the forward wire is FULLY ACTIVE: aux-head
gradients flow back through the kernel's `s1 - s0` derivative into
`aux_nb_softmax_buf`'s logits, co-training the aux head with Q-loss.
Backward gradient flow is INTENTIONALLY UNGATED in this commit —
the EGF pearl gating (scale `dL/dx[wire_col]` by `α_grad_smoothed`
to prevent gradient-hacking) lands in B.10. Per
`feedback_no_partial_refactor`, this intermediate state is
functional (the model trains; aux gets co-trained by Q-loss) but
not yet behavior-protected by the gate.

Diagnostic-path residual (causal intervention, DDQN argmax, exp
collector, value decoder): the cuBLAS heuristic for `K=SH2, LDA=SH2`
against the underlying `[adv_h, SH2 + 1]` weight tensor reads the
first `adv_h * SH2` floats with stride SH2 — within bounds (no
OOB), produces stable-but-incorrect outputs for the residual paths.
Their direction Q outputs feed either (a) only-value-logit consumers
(causal sensitivity) or (b) downstream argmax-only consumers with
one-step-bias acknowledged by the spec (DDQN). The train-time wire
(online + target + replay) is fully closed.

Test: `SQLX_OFFLINE=true cargo check -p ml` clean (18 warnings,
pre-existing baseline). The smoke validation that the model
converges with the active forward wire happens in B.11 alongside
the captured-graph integration (B.10 gates backward first).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-05 20:26:23 +02:00
jgrusewski
6715ab4ea1 feat(sp14): B.8 — direction Q-head input dim SH2 → SH2+1 + fingerprint
Forward wire (weight-side only — dispatch consumer lands in B.9):
direction Q-head's first FC weight tensor `w_b0fc` (param-table index
17) input dim grows by 1 to accept aux_softmax_diff. The (rows, cols)
shape table at the trainer-init Xavier site is updated in lock-step
with `compute_param_sizes()`. Output dim (adv_h) and the bias `b_b0fc`
are unchanged — bias is per-output, not per-input.

`layout_fingerprint_seed()` entry renamed `PARAM_W_B0FC` →
`PARAM_W_B0FC_AUX1`, forcing the FNV-1a hash to bump per Invariant 8
(old checkpoints fail-fast at load via `check_layout_fingerprint`
instead of silently aliasing onto the new architecture). Per
`feedback_no_legacy_aliases`, no `_DEPRECATED` shim — straight
in-place rename.

New weight column zero-initialised in trainer construction (mirrors
the OFI column zeroing for w_b2fc/w_b3fc): the model starts ignoring
the new input and learns to use it through gradient descent. Xavier-
init on this column would inject day-0 noise the trunk would have to
denoise — let the EGF gate decide when the aux signal is trustworthy.

Atomic-migration consumers (`feedback_no_partial_refactor`) updated:

- `gpu_dqn_trainer.rs` — Xavier (rows, cols) table at index 17 grows
  to (adv_h, SH2+1); spectral-norm descriptor entry [4] for W_a1 grows
  in_dim from sh2 to sh2+1; spec_v_a1 power-iteration vector grows
  from sh2 to sh2+1 floats.
- `dqn/smoke_tests/gradient_budget.rs` — both `alloc_dueling`
  fixtures' slot 8 grow `cfg.adv_h * cfg.shared_h2` →
  `cfg.adv_h * (cfg.shared_h2 + 1)`.
- `docs/dqn-wire-up-audit.md` — new SP14 Layer B B.8 entry per
  Invariant 7.

Note: forward GEMM dispatch still uses `K = shared_h2` until B.9 lands
the concat-then-SGEMM consumer (per plan §2358). Until then the new
column reads as ignored padding; this is safe because (a) it's zero-
initialised, (b) GPU-only smoke tests are skipped on this CPU CI, (c)
the fingerprint bump invalidates any pre-SP14 checkpoint that would
attempt to load.

Test: `layout_fingerprint_bumps_after_sp14_wire` (CPU-only, in
`sp14_oracle_tests.rs`) hashes the pre-B.8 seed verbatim with the
single difference `PARAM_W_B0FC` (vs post-B.8 `PARAM_W_B0FC_AUX1`)
and asserts `LAYOUT_FINGERPRINT_CURRENT` differs — any silent revert
of the rename trips this test. Mirrors the `fingerprint_bumped_from_
pre_b1_1a` pattern from `sp13_layer_b_oracle_tests.rs`.

Verified:
- `SQLX_OFFLINE=true cargo check -p ml` clean, 18 warnings (baseline)
- `cargo test -p ml --test sp14_oracle_tests` 2 passed, 6 ignored (GPU)
- `cargo test -p ml --test sp13_layer_b_oracle_tests
   fingerprint_bumped_from_pre_b1_1a` still passes (sister test)

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-05 19:58:18 +02:00
jgrusewski
9843de5e3d feat(sp14): B.7 — trainer struct fields for EGF kernels + concat scratch
Adds 4 CudaFunction handles + 1 scratch buffer to GpuDqnTrainer:
- sp14_q_disagreement_update_kernel
- sp14_alpha_grad_compute_kernel
- sp14_gradient_hack_detect_kernel
- sp14_dir_concat_qaux_kernel
- sp14_dir_qaux_concat_scratch: CudaSlice<f32> [B * (SH2 + 1)]

All loaded from precompiled cubins in trainer construction, mirroring
the SP13 aux_pred_to_isv_tanh / aux_sign_label kernel-loading pattern.
Per feedback_no_partial_refactor: handles + scratch are held but not
yet wired in. Subsequent tasks (B.9+) launch them.

Build: cargo check --workspace clean.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-05 19:41:12 +02:00
jgrusewski
4527d8c852 feat(sp14): B.6 — dir_concat_qaux_kernel (pre-SGEMM forward-wire concat)
Mirrors mag_concat_qdir precedent (experience_kernels.cu:4560).
Concats [h_s2 ; aux_softmax_diff] into scratch buffer [B, SH2+1] for
the direction Q-head's first FC SGEMM.

aux_softmax_diff = softmax[b, 1] - softmax[b, 0] in [-1, +1] computed
inline; structurally bounded by softmax components per
pearl_bounded_modifier_outputs_require_structural_activation.

Pure per-thread map; no reduction; no atomicAdd. Does not read or write
ISV slots — purely data-movement. Launch order constraint: aux head
forward MUST complete before this concat reads aux_nb_softmax_buf
(enforced by orchestrator in B.10/B.11).

Test: dir_concat_qaux_correct verifies row-wise contiguous concat
+ correct softmax diff values for both 'down' (-0.8) and 'up' (+0.8)
synthetic aux predictions. B.3+B.4+B.5 regression: 5 GPU tests
unchanged (6 total pass).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-05 19:36:03 +02:00
jgrusewski
82fe6cea66 feat(sp14): B.5 — gradient_hack_detect_kernel (anti-mesa-opt circuit breaker)
Detects suspected gradient hacking: when gate1 is open AND aux_dir_acc
post-open-minimum drops > LOCKOUT_TRIGGER_DROP (0.05) below the Schmitt
open-threshold (target + SCHMITT_BAND = target + 0.03) AND q_disagreement
rises > LOCKOUT_TRIGGER_DIS_RISE (0.10) above the analytic random-alignment
baseline 0.5, simultaneously.

Action: force gate1_open_state = 0 (ISV[391]); set lockout_remaining = 2.0
epochs (LOCKOUT_EPOCHS). During lockout, gate1 stays force-closed
regardless of alpha_grad_compute_kernel output.

Tracks AUX_DIR_ACC_POST_OPEN_MIN (ISV[394]): running minimum of aux_dir_acc
since gate1 last opened; resets to 1.0 sentinel when gate closes naturally
or when circuit breaker fires.

Slot indices shifted +2 from original plan (SP13 closeout added
HOLD_RATE_TARGET=381 + HOLD_RATE_OBSERVED_EMA=382): Q_DIS_SHORT=383,
GATE1=391, POST_OPEN_MIN=394, LOCKOUT=395. Matches sp14_isv_slots.rs.

Single-thread state-machine kernel (threadIdx.x==0 guard); runs at end of
each epoch after alpha_grad_compute_kernel. No atomicAdd per
feedback_no_atomicadd.md.

1 oracle test: gradient_hack_circuit_breaker_fires verifies trigger
conditions (aux_drop=0.08 > 0.05, q_rise=0.15 > 0.10) cause lockout=2.0
and gate1 force-close=0.0.

B.3+B.4 regression: 4 GPU tests unchanged (5 total GPU pass, 1 host pass).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-05 19:29:22 +02:00
jgrusewski
49cdf90ecc feat(sp14): B.4 — alpha_grad_compute_kernel (EGF heart)
Single-thread state-machine kernel that is the heart of the Earned
Gradient Flow pearl. Reads driver signals from the global ISV bus,
runs Schmitt-trigger Gate 1, computes adaptive k_aux/k_q/β, evaluates
two sigmoids, multiplies with a host-supplied warmup gate, applies a
β-rate-limiter, and writes 7 outputs back to ISV.

Per-step pipeline:

1. Read aux_dir_acc (slot 373), q_disagreement (slot 383), Welford
   variance EMAs (388, 389, 390), persistent Schmitt state (391),
   alpha_smoothed_prev (393).
2. Compute adaptive k_aux = K_BASE_AUX/(1 + var_aux/VARIANCE_REF_AUX)
   and k_q analogously (B.2.5; floor at K_MIN = 1.0).
3. Run Schmitt-trigger Gate 1 state update (open at target+0.03,
   close at target-0.03; intentional discontinuity at transition is
   smoothed by the β rate-limiter downstream).
4. Evaluate Gate 1 sigmoid (aux competence, distance from threshold)
   and Gate 2 sigmoid (Q-aux disagreement vs analytic 0.5 baseline).
5. alpha_grad_raw = gate1 × gate2 × warmup_gate (structurally bounded
   to [0, 1] per pearl_bounded_modifier_outputs_require_structural_
   activation; no runtime clamp).
6. Update Welford variance of alpha_grad_raw → adaptive β (B.2.8;
   floor BETA_BASE = 0.5, ceiling BETA_MAX = 0.95).
7. alpha_grad_smoothed = β × prev + (1-β) × raw (rate-limited).
8. Write back 7 outputs: k_aux (385), k_q (386), β (387), var_alpha
   (390), gate1_state (391), alpha_raw (392), alpha_smoothed (393).

Sigmoid arguments clipped to [-30, 30] before __expf for fp32
overflow guard (precision-neutral; sigmoid saturates bit-equal at
those bounds).

Per pearl_bounded_modifier_outputs_require_structural_activation:
sigmoid composition produces structurally-bounded [0, 1] output.

KNOWN LIMITATION: as of B.4 landing, NO upstream kernel writes
ISV[388] (AUX_DIR_ACC_VARIANCE_EMA). The grep at status-report time
finds only the sp14_isv_slots.rs declaration. Effect: var_aux stays
at sentinel 0.0 forever, so k_aux is degenerate-but-non-fatal at
K_BASE_AUX (constant). Gate 1 still works, the sigmoid just doesn't
soften under noisy aux_dir_acc. To be resolved in B.11 producer-
chain orchestrator OR a separate fix-up task that adds a Welford-
variance update next to the existing AUX_DIR_ACC_SHORT_EMA producer.
var_q (389) IS written by q_disagreement_update_kernel (B.3), so
adaptive k_q is fully functional from B.4 onward.

Slot indices hardcoded inside the kernel via const int locals — must
match crates/ml/src/cuda_pipeline/sp14_isv_slots.rs (and 372/373
from sp13_isv_slots.rs). The plan originally documented 381/383/
384/385/386/387/388/389/390/391 for SP14 slots; the actual values
are +2 because SP13 closeout added HOLD_RATE_TARGET=381 +
HOLD_RATE_OBSERVED_EMA=382 after the plan was written.

Tests (RTX 3050 Ti pass; B.3's 2 tests still pass — no regression):

- alpha_grad_schmitt_hysteresis: 4-step trajectory verifies the
  closed→open→open→closed transition. Closed at aux=0.55 (below
  open=0.58); opens at aux=0.60; stays open at aux=0.54 (in
  hysteresis band [close=0.52, open=0.58]); finally closes at
  aux=0.50 (below close=0.52).
- alpha_grad_adaptive_beta: 20-oscillation regime verifies β grows
  above β_base=0.5 and remains bounded by β_max=0.95.

docs/dqn-wire-up-audit.md updated per Invariant 7 with full B.4
behaviour contract, per-step pipeline, single-thread launch
convention, sigmoid clip rationale, Schmitt discontinuity note,
and the var_aux Known Limitation.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-05 19:14:48 +02:00
jgrusewski
d3a35cc6e9 feat(sp14): B.3 — q_disagreement_update_kernel with K=4↔K=2 mapping
First of four producer kernels in the Earned Gradient Flow pearl chain.
Per-step computes the K=4↔K=2 mapped argmax mismatch between the Q-head's
4-way direction action (DIR_SHORT/DIR_HOLD/DIR_LONG/DIR_FLAT) and the
aux head's 2-way next-bar prediction (down/up), updates fast + slow
EMAs of the disagreement rate, and updates a Welford-style variance
EMA on the same signal in a single launch.

Mapping (per state_layout.cuh):
  DIR_SHORT (=0)  → aux down (=0)        contributes
  DIR_HOLD  (=1)  → masked                no contribution
  DIR_LONG  (=2)  → aux up   (=1)        contributes
  DIR_FLAT  (=3)  → masked                no contribution

Hold and Flat are masked because they represent "no NEW directional
commitment" (Hold = keep prior position; Flat = close all positions).
Penalising the aux head for not matching them would conflate
position-management actions with directional predictions.

Block tree-reduce on numerator + count separately, then divide and
update EMAs in a single thread (tid == 0). No atomicAdd per
feedback_no_atomicadd.md. Pure GPU compute per
feedback_no_cpu_compute_strict.md.

Pearl-A first-observation: when ISV[Q_DISAGREEMENT_SHORT_EMA=383] AND
ISV[Q_DISAGREEMENT_LONG_EMA=384] both equal sentinel 0.5f AND batch
has at least one valid contribution (total_cnt > 0), both EMAs are
replaced directly with the first observation per
pearl_first_observation_bootstrap.md. The 0.5f exact-match is safe
because 0.5 is exactly representable in IEEE 754 single precision
(mantissa = 1.0, exponent = -1). The Welford variance EMA at slot
389 drives the adaptive k_q sigmoid steepness consumed by the
alpha_grad_compute_kernel in B.4.

Slot indices are hardcoded inside the kernel via #define — must match
crates/ml/src/cuda_pipeline/sp14_isv_slots.rs (currently 383, 384, 389).
The kernel header documents the coupling explicitly.

Launch contract:
  grid_dim = (1, 1, 1)
  block_dim = (256, 1, 1)
  shared_mem_bytes = 2 * 256 * sizeof(float) = 2048

A shared_mem_bytes = 0 launch reads garbage and corrupts the EMA — the
kernel header documents the launcher requirement; oracle tests pass
2048 explicitly.

Files:
  - crates/ml/src/cuda_pipeline/q_disagreement_update_kernel.cu (NEW)
  - crates/ml/build.rs — register cubin in kernels_with_common
  - crates/ml/tests/sp14_oracle_tests.rs — append #[cfg(feature="cuda")]
    mod gpu with cubin handle + 2 GPU oracle tests
  - docs/dqn-wire-up-audit.md — SP14 B.3 section (Invariant 7)

Tests (both pass on RTX 3050 Ti):
  - q_disagreement_k4_k2_mapping: 8-row batch with 2 agreements,
    2 disagreements, 4 masked Hold/Flat → first-obs Pearl-A replaces
    both EMAs with batch_mean = 0.5 (asserted within 1e-4)
  - q_disagreement_all_hold_no_contribution: all-Hold edge case;
    total_cnt = 0 so batch_mean = 0/1 = 0; sentinel-bootstrap guard
    (total_cnt > 0) keeps EMA from collapsing to 0; current kernel
    blends to 0.35, test bound [0.0, 0.5] tolerant of either current
    blend or future skip-update refinement, asserts is_finite()

Wire-up status: producer kernel exists and is exercised only by the
oracle tests. The Rust launcher and graph-capture integration land in
B.7+ alongside the consumer (alpha_grad_compute in B.4 reads slots
383/384/389). This is one producer kernel of four (B.3 q_disagreement,
B.4 alpha_grad_compute, B.5 gradient_hack_detect, B.6 dir_concat_qaux);
known-orphan for the duration of the producer chain per
feedback_wire_everything_up.md (the same-commit wire-up rule is
relaxed for atomic chained-producer-consumer landings, with each
orphan documented in the audit doc; this orphan is acknowledged in
docs/dqn-wire-up-audit.md SP14 B.3 section).

Build + test verification:
  - SQLX_OFFLINE=true CUDA_COMPUTE_CAP=86 cargo check -p ml --features cuda
    → clean (only the 18 pre-existing warnings)
  - SQLX_OFFLINE=true cargo test -p ml --test sp14_oracle_tests
    set_aux_weight → host-only A.2 still passes (no shared dependency)
  - SQLX_OFFLINE=true CUDA_COMPUTE_CAP=86 cargo test -p ml
    --test sp14_oracle_tests --features cuda q_disagreement
    -- --ignored --nocapture → 2 PASS

Slot indices reflect the SP13-closeout +2 shift documented in
sp14_isv_slots.rs (the original plan's 381/382/387 became 383/384/389
because HOLD_RATE_TARGET=381 and HOLD_RATE_OBSERVED_EMA=382 landed
between when the plan was written and B.1 was implemented).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-05 19:04:18 +02:00
jgrusewski
84de278dfe feat(sp14): B.2 — register 11 SP14 ISV slots for fold-boundary reset
Each EGF pearl EMA / state slot resets to its Pearl-A sentinel at fold
boundary, mirroring sp13_aux_dir_acc_short_ema / long_ema entries.
Atomic refactor (feedback_no_partial_refactor): both halves land
together — registry entry + reset_named_state dispatch arm.

Reset slots (11 total, sentinel in parens):
  - Q_DISAGREEMENT_SHORT/LONG_EMA (slots 383, 384) → 0.5
  - K_AUX_ADAPTIVE (385) → K_BASE_AUX = 20.0
  - K_Q_ADAPTIVE (386) → K_BASE_Q = 15.0
  - BETA_RATE_LIMITER_ADAPTIVE (387) → BETA_BASE = 0.5
  - AUX_DIR_ACC_VARIANCE_EMA, Q_DISAGREEMENT_VARIANCE_EMA,
    ALPHA_GRAD_RAW_VARIANCE_EMA (388, 389, 390) → 0.0
    (initial k = k_base, β = β_base via ISV-driven controllers)
  - GATE1_OPEN_STATE (391) → 0.0 (closed)
  - ALPHA_GRAD_SMOOTHED (393) → 0.0
  - AUX_DIR_ACC_POST_OPEN_MIN (394) → 1.0 (no min observed)

ALPHA_GRAD_RAW (slot 392, recomputed every step from variance EMAs)
and GRADIENT_HACK_LOCKOUT_REMAINING (slot 395, decays at epoch
boundary) are NOT in the fold-reset registry; both naturally
re-initialise without explicit reset.

Also corrects the isv-slots.md SP14 table: slots 392 and 395 were
incorrectly marked FoldReset in the B.1 entry; corrected to reflect
their actual reset semantics (NOT reset / epoch-boundary decay).

Producer + consumer wiring lands in subsequent tasks (B.3-B.12);
this commit is additive infrastructure only — no behavior change.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-05 18:51:00 +02:00
jgrusewski
d63cb7992e feat(sp14): B.1 — sp14_isv_slots.rs with 13 new ISV slot constants
Allocates ISV slots [383..396) for the Aux→Q Wire + Earned Gradient
Flow pearl (Layer B of SP14). Mirrors sp13_isv_slots.rs pattern.
The plan originally documented [381..394), but Phase 0 verification
found SP13 closeout added HOLD_RATE_TARGET_INDEX=381 and
HOLD_RATE_OBSERVED_EMA_INDEX=382 after the plan was written, so the
range shifts by +2.

Slots fall into 4 functional groups:
- Q-disagreement EMAs (short, long; K=4↔K=2 mapping with Hold/Flat masked)
- Adaptive controllers (k_aux, k_q, β; variance-driven)
- Welford variance EMAs (3, one per adaptive scalar)
- Schmitt state + α_grad outputs + circuit breaker

Plus 14 structural constants for numerical-stability anchors:
K_BASE_*, K_MIN, VARIANCE_REF_*, BETA_BASE, BETA_MAX, SCHMITT_BAND,
WARMUP_STEPS_FALLBACK, LOCKOUT_*, Q_DISAGREEMENT_BASELINE.

Per feedback_isv_for_adaptive_bounds: adaptive bounds (k_*, β,
post_open_min, lockout) live in ISV; numerical anchors live as
structural constants. Per pearl_first_observation_bootstrap: all
EMAs reset to sentinels and Pearl-A bootstraps on first observation.

Producer + consumer wiring lands in subsequent tasks (B.2-B.12);
this commit is additive infrastructure only — no behavior change.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-05 18:45:49 +02:00
jgrusewski
f75786fc5a fix(sp14): A.1 — C51 inv_a_std floor lift (1e-6 → 1e-3)
c51_grad_kernel.cu line 275: lift floor from 1e-6 to 1e-3 in
\`inv_a_std = 1.0f / (a_std + 1e-3f)\`, capping the magnitude-branch
gradient amplifier at 1000 instead of ~1e6 in the degenerate case.

Why: Smoke A produced 1109 GRAD_CLIP_OUTLIER events with C51 grad
reaching 9.5e6 — the SP7 budget controller saturated at the EPS_DIV
floor instead of rebalancing proportionally. Phase-0 verification
against the actual kernel found the amplifier is NOT the spec's
claimed −log(p)/p divide (which does not exist; the kernel uses
the CE-stable expf(lp) - proj form at line 81). The actual
amplifier is inv_a_std = 1/(a_std + 1e-6) at line 275, gated by
\`if (d == 1) grad_val *= inv_a_std\` at line 282. When magnitude
advantage logits collapse near-uniform (Smoke A: var_q=9e-10),
a_std → ~1e-9, so inv_a_std → ~1e6.

Per feedback_isv_for_adaptive_bounds, this is a numerical-stability
anchor (Invariant 1: prevent division-by-near-zero amplification),
not a behavioural bound. ISV-driven bounds govern behavior; the
existing 1e-12f floor on a_std at line 274 is also a structural
anchor — same class of fix.

Validation gate: Smoke A2-A GRAD_CLIP_OUTLIER count <100 in fold 2
(was 1109 pre-fix). The 3-order-of-magnitude reduction in worst-
case amplification should bring C51 grad spikes back under SP7
budget controller authority.

Audit doc: Fix 40 added (parallel to Fix 39 for A.2 and Fix 41 for
A.3); the stale "A.1 deferred" note was removed in the A.3 commit.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-05 18:09:21 +02:00
jgrusewski
1420383212 fix(sp14): A.3 — stagnation warmup gate at fold boundary
compute_aux_w_p0b's stagnation term was firing inappropriately at
fold reset because Pearl-A first-observation bootstrap forces both
EMAs equal:
  short_ema = sentinel 0.5 → first observation X
  long_ema  = sentinel 0.5 → first observation X (same update)
  improvement = max(0, short - long) = 0
  stagnation = (1 - 0/max(deficit, 0.005)) = 1.0
  aux_w *= (1 - 0.7) = 0.3 → spurious decay on a non-stagnation

Fix: add epochs_in_fold parameter; skip stagnation when < 1. Wait
one full epoch for the α=0.3 vs α=0.05 short/long EMA timescale
split to produce real improvement signal.

Atomic refactor (feedback_no_partial_refactor): all 5 callers
migrated — 1 production site (training_loop.rs:4257) plus 4
existing unit tests at trainers/dqn/trainer/tests.rs.

Test: aux_w_stagnation_warmup_gate_epoch_0 verifies:
  - Epoch 0: stagnation = 0, aux_w = base × deficit_amp = 0.625
  - Epoch 1: stagnation = 1.0, aux_w = floor 0.15

Combined with A.2 (clamp lift), Fold 2's aux_w should now hit
0.625 in epoch 0 (the controller's intended post-deficit-amp
value) instead of being collapsed to 0.164.

Audit doc: Fix 41 added; stale "A.1 deferred" note from a prior
killed agent was also removed (A.1 lands in the next commit, not
deferred per user direction).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-05 18:07:12 +02:00
jgrusewski
731cae4c80 fix(sp14): A.2 — lift set_aux_weight clamp to SP13 P0b range [0.15, 1.5]
The pre-fix clamp at gpu_dqn_trainer.rs:14722 was [0.05, 0.3] — the
SP11-era cap. P0b's controller computes aux_w in [0.15, 1.5] (base
0.5 × [0.3, 3.0]) but the setter silently chopped everything above
0.3, masking the deficit-amplification term entirely.

Smoke A trace confirmed:
  Fold 0/1: raw aux_w = 0.66-0.80 → clamped to 0.30 (deficit invisible)
  Fold 2:   raw aux_w = 0.164 (stagnation; below clamp) → 45% deficit

Post-fix: deficit-amp term `(1 + 5 × deficit)` actually expresses
through to the trainer. Fold 2 stagnation will get the designed
floor 0.15 instead of being capped at 0.3 — but the upper range
1.5 also opens, so deficit-amp can pull aux_w up when accuracy is
below target.

Constants imported from sp13_isv_slots.rs (AUX_W_BASE=0.5,
AUX_W_HARD_FLOOR_RATIO=0.3, AUX_W_HARD_CEIL_RATIO=3.0). No new
slots; existing constants exposed as the clamp bounds.

Test: set_aux_weight_clamp_range verifies constants resolve correctly.

A.1 (C51 atom-probability floor) deferred — Phase 0 verification
found the spec's stated `−log(p)/p · ∂p/∂z` divide does NOT exist
in c51_grad_kernel.cu at HEAD 037c24116; actual kernel uses the
numerically-stable `expf(lp) - proj`. The 1109 GRAD_CLIP_OUTLIER
events in Smoke A are real but their mechanism is different. Will
be re-spec'd as a separate task post-SP14.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-05 17:52:18 +02:00
jgrusewski
6657e56265 feat(sp13): B1.1b — producer kernel + replay direct path + experience collector
Final piece of the SP13 Layer B chain. B1.1a flipped the aux head from K=1
MSE regression to K=2 softmax CE classification but aux_nb_label_buf was
zero-init — model was training on "all bars are class 0 (down)". B1.1b
lands the producer kernel that fills i32 -1/0/1 labels from the 30-bar
price trajectory, the replay direct-path 8th gather that carries those
labels into the trainer, and the experience collector hoist that ensures
bar_indices_pinned is always populated (producer + hindsight relabel both
consume it). Aux head finally trains on real classification signal.

Recovery commit: this completes a B1.1b agent dispatch that crashed
mid-edit. The implementer had landed ~95% of the cascade (kernel file,
build.rs, replay buffer signature + direct path, fused_training getter,
trainer accessor, both training_loop.rs callers, kernel field + cubin
loader on the experience collector) before being killed. The missing
pieces (experience collector launch + bar_indices_pinned hoist + 6
producer tests + audit doc) were completed manually post-crash and
verified end-to-end.

Three contracts (atomic single commit per feedback_no_partial_refactor):

  1. NEW aux_sign_label_kernel.cu producer — pure per-thread O(1) map
     reading targets[bar*6+2] (raw_close column) at bar and bar+lookahead,
     writing -1 (skip if bar+lookahead >= total_bars), 0 (down/flat under
     strict greater-than tie-break), or 1 (up). Replaces B0 alloc_zeros.

  2. Replay direct-path 8th gather — set_trainer_buffers gains 8th arg
     trainer_aux_sign_labels_ptr; direct branch in sample_proportional
     adds gather_i32_scalar into the trainer ptr; fallback gather wrapped
     in if !direct_to_trainer (avoids wasted DtoD). Direct-mode
     GpuBatchPtrs return points aux_sign_labels_ptr at trainer ptr.

  3. Experience collector bar_indices_pinned cpu-fill hoist — moved out
     of if hindsight_fraction > 0.0 so producer + hindsight share it.

Files (9 total):
  - crates/ml/src/cuda_pipeline/aux_sign_label_kernel.cu (NEW)
  - crates/ml/build.rs (cubin registration)
  - crates/ml/src/cuda_pipeline/gpu_experience_collector.rs (kernel
    field + cubin loader + struct init + hoist + producer launch)
  - crates/ml-dqn/src/gpu_replay_buffer.rs (8th arg + direct gather +
    fallback skip + GpuBatchPtrs return)
  - crates/ml/src/cuda_pipeline/gpu_dqn_trainer.rs (aux_nb_label_buf_ptr
    accessor mirrors 6 existing trainer-buf accessors)
  - crates/ml/src/trainers/dqn/fused_training.rs
    (trainer_aux_sign_labels_buf_ptr getter)
  - crates/ml/src/trainers/dqn/trainer/training_loop.rs (both
    set_trainer_buffers callers updated)
  - crates/ml/tests/sp13_layer_b_oracle_tests.rs (6 NEW producer tests)
  - docs/dqn-wire-up-audit.md (B1.1b section)

Hard rules upheld:
  - feedback_no_partial_refactor: every consumer of the 3 contracts
    migrates atomically
  - feedback_no_atomicadd: producer is pure map; no reductions
  - feedback_cpu_is_read_only: producer GPU-only; only host work is
    pre-existing bar_indices_pinned cpu-fill (hoisted unchanged)
  - feedback_no_stubs: kernel output flows through real chain — ring
    buffer → direct gather → aux_nb_label_buf → CE consumer
  - feedback_no_legacy_aliases: 8-arg setter gets
    #[allow(clippy::too_many_arguments)] not an alias shim
  - feedback_no_htod_htoh_only_mapped_pinned: targets_buf and
    bar_indices_pinned both pre-existing mapped-pinned

Build + test:
  - cargo check --workspace clean (only pre-existing warnings)
  - cargo check --workspace --tests clean
  - 17 tests in sp13_layer_b_oracle_tests.rs:
    - 2 CPU-only (fingerprint bump + HEALTH_DIAG snap stable) pass
    - 15 GPU on RTX 3050 Ti pass (9 B1.1a + 6 new B1.1b producer):
        aux_sign_label_monotone_up_all_ones
        aux_sign_label_monotone_down_all_zeros
        aux_sign_label_flat_all_zeros_strict_gt
        aux_sign_label_last_30_bars_skip
        aux_sign_label_boundary_first_valid_last_skip
        aux_sign_label_multi_episode_per_episode_skip

Producer tests cover every edge case in the kernel:
  - Monotone trajectories (label=1 / label=0 across all valid bars)
  - Flat tie-break (strict greater-than means flat → 0)
  - Skip sentinel for last lookahead bars
  - First/last bar boundary (bar=0 valid, bar=L-1 skip)
  - Multi-episode global skip semantics

Next: Smoke A — L40S 5-epoch validation of full SP13 stack
(P0a + P0b + B0 + B0.1 + B1.0 + B1.1a + B1.1b). Expected: aux head
trains on real K=2 softmax CE labels; aux_dir_acc_short_ema rises above
0.5 within first epoch (vs B1.1a degraded baseline at 0.5);
HEALTH_DIAG aux_b1_diag emits per-epoch with n_down/n_up/n_skip/mask_frac.
If aux_dir_acc_short_ema > 0.55 by epoch 5, B1.1b is validated and the
chain merges to main.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-05 14:45:47 +02:00
jgrusewski
7d10ea8b3e feat(sp13): B1.1a — K=1→2 + softmax CE kernel rewrites + struct flips
Flips the aux next-bar head from K=1 MSE regression to K=2 softmax
cross-entropy classification. Kernel ABIs, struct fields, partial-buf
shapes, and per-step ISV producers all migrate atomically; the producer
that fills `aux_sign_labels` with real -1/0/1 from the price trajectory
lands separately in B1.1b.

Why split B1.1a from B1.1b: the original B1 brief was decomposed (B1.0
+ B1.1) after six full-B1 dispatches confirmed agent-session capacity
is the bottleneck, not cascade understanding. B1.1 itself is now further
split into B1.1a (kernel/struct/test cascade — this commit) and B1.1b
(producer kernel + replay direct path + experience collector hoist +
remaining tests + Smoke A). B1.1a is contract-consistent atomic
kernel-side; B1.1b lands the producer + smoke.

Why labels stay zero-init: B1.1a is producer-less by design. The
`aux_nb_label_buf: CudaSlice<i32>` is `alloc_zeros<i32>` so every
sample receives label 0 ("down"). The model converges on "predict
class 0 (down) everywhere" until B1.1b lands the producer kernel that
fills real -1/0/1 from the 30-bar price trajectory. This degraded
training behavior is intentional and known — the cascade is internally
consistent (every consumer of K-flip / softmax tile / CE / i32 label
migrates atomically per `feedback_no_partial_refactor`); the labels are
placeholder. Local unit tests (CE correctness, dir_acc correctness,
isv_tanh correctness, fingerprint bump) validate B1.1a in isolation;
no L40S smoke runs between B1.1a and B1.1b.

Four contracts (atomic in this commit):
  1. K_NB flip 1 → 2: AUX_NEXT_BAR_K constant, compute_param_sizes
     ([121]/[122] grow), fingerprint seed rename
     (PARAM_AUX_NB_W2/B2 → PARAM_AUX_NB_W2_K2/B2_K2 — bumps the hash),
     forward + backward kernels, partial-buf allocs (nb_w2 [B,H]→[B,K,H],
     nb_b2 [B]→[B,K]), saxpy spec table, max_aux_tensor_len,
     aux_nb_pred_buf renamed to aux_nb_logits_buf per
     feedback_no_legacy_aliases.
  2. Softmax tile: aux_next_bar_forward writes [B, K] softmax via
     in-kernel stable softmax (max-shift form, K=2 single-thread fanout
     mirrors regime kernel); 4 consumers read the tile (loss, backward,
     dir-acc, isv-tanh). NEW field aux_nb_softmax_buf [B, K].
  3. MSE → CE: aux_next_bar_loss_reduce reads softmax + i32 labels,
     masks -1, divides by B_valid (mean-over-valid-rows), writes loss +
     B_valid scalar. aux_next_bar_backward reads B_valid so loss + grad
     share the same divisor — derivatives of the same scalar function.
     NEW field aux_nb_valid_count_buf [1]. All-skip batch produces
     loss = 0 (no NaN; fmaxf(valid, 1) divisor) and zero gradients.
     Numerical floor 1e-30 prevents -log(0) = +inf in extreme-logit path.
  4. i32 label dtype: aux_nb_label_buf flipped f32 → i32; -1 mask
     sentinel handled across loss + backward + dir-acc + isv-tanh.
     The strided_gather of next_states[:, 0] retired entirely.

Cascade (atomic per feedback_no_partial_refactor):
- aux_heads_kernel.cu: aux_next_bar_forward gains K + softmax tile
  output (via in-kernel stable softmax); aux_next_bar_loss_reduce
  ABI flipped (softmax + i32 labels, mean-over-valid CE,
  valid_count_out); aux_next_bar_backward ABI flipped (softmax +
  i32 labels + valid_count, K-fanout d_logits, masked rows zero
  across the K-vector)
- aux_dir_acc_reduce_kernel.cu: read softmax + i32 labels, argmax
  over K, output grew 3 → 6 floats (added n_down/n_up/n_skip);
  shmem 4 → 6 int arrays
- aux_pred_to_isv_tanh_kernel.cu: read softmax tile, compute
  mean(softmax[:, 1] - softmax[:, 0]); tanh transcend retired
  (structural [-1, +1] bound via softmax components per
  pearl_bounded_modifier_outputs_require_structural_activation)
- gpu_aux_heads.rs: AUX_NEXT_BAR_K 1 → 2; forward_next_bar gains K +
  logits_out + softmax_out args; next_bar_loss_reduce gains K +
  valid_count_out; backward_next_bar gains softmax_in + labels_i32_in
  + valid_count_in + K
- gpu_dqn_trainer.rs: compute_param_sizes ([121]/[122]); fingerprint
  seed rename (W2/B2 → W2_K2/B2_K2); struct fields (logits, softmax,
  i32 label, valid_count); aux_dir_acc_buf 3 → 6 floats; partial-buf
  allocs grow; max_aux_tensor_len extended; saxpy spec table updated;
  orchestrator launchers (launch_aux_dir_acc_reduce,
  launch_aux_pred_to_isv_tanh, launch_sp13_aux_dir_metrics) gain K
  arg; strided_gather block deleted entirely
- training_loop.rs: aux_b1_diag HEALTH_DIAG line reads
  aux_dir_acc_buf [3..6] for n_down / n_up / n_skip + mask_frac;
  doc comment update for the per-step aux dir-metrics block
- tests/sp13_phase0_oracle_tests.rs: 6 dir_acc + 3 isv_tanh tests
  rewritten in-place to new ABI (no shadow tests, per
  feedback_no_legacy_aliases)
- tests/sp13_layer_b_oracle_tests.rs (NEW): 11 B1.1a tests — 5 CE
  loss/backward (single-row, batch-mixed, all-skip-NaN, backward
  single-row, backward batch-mixed); 2 dir_acc (handcrafted argmax,
  all-skip NaN-safe); 2 isv_tanh (bounded fuzz, mean handcrafted);
  2 layout regression (fingerprint bump, HEALTH_DIAG snap stable)
- docs/dqn-wire-up-audit.md: B1.1a section

Hard rules upheld:
- feedback_no_partial_refactor: every consumer of K-flip / softmax tile
  / CE / i32 labels migrates atomically — kernels + orchestrators +
  struct + diag + existing oracle tests
- feedback_no_atomicadd: block tree-reduce only; CE loss reduce uses
  2 parallel partial-reduction strips; CE backward uses existing
  per-sample partial → final aux_param_grad_reduce pattern
- feedback_cpu_is_read_only: aux_nb_label_buf is GPU-resident
  CudaSlice<i32>; HEALTH_DIAG aux_b1_diag reads via mapped-pinned
  aux_dir_acc_buf (no DtoH)
- feedback_no_stubs: every new buffer + kernel arg wired through to
  a real consumer; CE forward / loss / backward chain executes
  end-to-end against placeholder labels (degraded behavior, not stub)
- feedback_no_legacy_aliases: aux_nb_pred_buf renamed in-place
  (no shim); PARAM_AUX_NB_W2/B2 renamed to _W2_K2/_B2_K2 in seed
  (no _DEPRECATED alias); 9 existing oracle tests rewritten in-place
- feedback_no_cpu_test_fallbacks: 9 GPU tests gated #[ignore]; 2
  layout-regression tests are CPU-only (pub const + size_of)
- feedback_no_htod_htoh_only_mapped_pinned: every CPU↔GPU buffer
  in tests + production is MappedF32Buffer / MappedI32Buffer
- feedback_isv_for_adaptive_bounds: no hardcoded thresholds added
  (1e-30 numerical floor on log is stability epsilon, not tunable)
- feedback_trust_code_not_docs: 8/8 Phase 0 anchors verified at
  HEAD 75e94858c before editing
- pearl_bounded_modifier_outputs_require_structural_activation:
  softmax IS the structural activation; runtime tanh squash retired

Build: cargo check --workspace --tests clean.
Tests:
  - snapshot_size_is_stable passes at 149*4=596 bytes (B1.1a
    doesn't touch HEALTH_DIAG snap-words)
  - fingerprint_bumped_from_pre_b1_1a passes (proves W2/B2 rename
    flipped seed hash)
  - health_diag_snap_size_stable_at_149_floats passes
  - 9 GPU oracle tests in sp13_layer_b_oracle_tests pass on local
    RTX 3050 Ti (B=1/4)
  - 14 tests in sp13_phase0_oracle_tests pass (6 dir_acc + 3
    isv_tanh rewrites + 5 unchanged hold/EMA tests)

Pre-existing test failures (14) on `cargo test -p ml --lib` are
environmental (OFI test data missing, etc.) and reproduce identically
at HEAD 75e94858c (B1.0) — verified via git stash round-trip.

Net delta: 9 files (8 source + 1 audit doc), +1039 / −539 LOC.

Next: B1.1b — producer kernel + replay direct-path 8th gather +
experience collector hoist + remaining 7 tests (6 producer + 1
end-to-end round-trip) + Smoke A on L40S.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-05 14:09:48 +02:00
jgrusewski
75e94858c5 feat(sp13): B1.0 — ISV[117] retirement + scale-free MSE bridge
Retires ISV[117]=AUX_LABEL_SCALE_EMA_INDEX together with its producer
kernel (aux_label_scale_ema_update), launch site, backward pass-through,
StateResetRegistry entry, HEALTH_DIAG snapshot field, and unit test.

Why: labels at the data layer are z-normalised, so the
mean(|label|) EMA tracked by ISV[117] sits at ~1.0 empirically.
Dividing by max(scale, 1e-6) before the residual `(pred - label)`
reduces to `(pred - label)` within rounding. The divisor was a
defensive scaffold from when the data layer carried mixed-scale
labels (1e-3 log returns vs 5000 raw prices); z-normalisation made
that scaffold redundant.

This is a numerical bridge, NOT the final fix. B1.1 lands on top:
- Aux head 1→2 dim (next-bar regression → 2-class direction logit)
- MSE → CE loss flip
- aux_dir_acc reads softmax over the 2 logits
- aux_pred_to_isv_tanh rewrite as logit-diff
- Producer kernel that fills aux_sign_labels with real -1/0/1 from
  the 30-bar price trajectory (B0 plumbing currently zero-init)
- dqn_param_layout fingerprint bump (head dim changes)
- aux_b1_diag HEALTH_DIAG metric
- 17+ GPU oracle unit tests

Cascade (atomic per feedback_no_partial_refactor):
- aux_heads_kernel.cu: aux_next_bar_loss_reduce + aux_next_bar_backward
  drop `isv` + `isv_label_scale_index` params; residual is (pred - label)
- aux_heads_loss_ema_kernel.cu: aux_label_scale_ema_update kernel deleted
- gpu_aux_heads.rs: kernel field/loader + launch_label_scale_ema +
  isv_* args from next_bar_loss_reduce / backward_next_bar all dropped
- gpu_dqn_trainer.rs: Step 2b producer launch + ISV slot uses dropped;
  AUX_LABEL_SCALE_EMA=117 line retained in fingerprint seed
  (no fingerprint bump in B1.0; B1.1 will bump on head-dim flip)
- gpu_health_diag.rs + health_diag.rs: aux_label_scale snapshot field
  dropped; aux block 4→3 floats, downstream offsets shift down by 1,
  WORD_TOTAL 150→149, snapshot_size_is_stable test 150*4 → 149*4
- health_diag_kernel.cu: WORD_AUX_LABEL_SCALE removed, downstream
  offsets shift, static_assert(WORD_TOTAL == 149)
- state_reset_registry.rs: isv_aux_label_scale_ema FoldReset dropped
- training_loop.rs: reset_named_state arm + HEALTH_DIAG read +
  aux line label_scale field all dropped
- sp4_producer_unit_tests.rs: load_aux_label_scale_ema_kernel helper +
  sp4_aux_label_scale_ema_writes_step_obs_via_pearl_a_then_converges_pearl_d
  test dropped

Hard rules upheld:
- feedback_no_partial_refactor: every consumer of ISV[117] migrates
  atomically — kernel + Rust orchestrator + producer launch + backward
  + HEALTH_DIAG + reset registry + unit test all in this commit
- feedback_no_stubs: not a stub — divisor is removed at every site,
  not aliased through a 1.0_const shim
- feedback_no_legacy_aliases: no legacy AUX_LABEL_SCALE_EMA_INDEX → 1.0
  alias function
- feedback_no_hiding: doc comments forward to B1.1 explicitly; no
  underscore suppression or #[allow(dead_code)]

Build: cargo check --workspace --tests clean.
Tests: snapshot_size_is_stable passes at 149*4=596 bytes.
       cargo test -p ml --lib + cargo test -p ml-dqn --lib compile.

Net delta: 10 files, −288 LOC.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-05 10:52:13 +02:00
jgrusewski
6a869ad366 fix(sp13): B0 cascade gap — 5 unaudited insert_batch call sites
The B0 audit (commit 62ab8ed85) under-counted insert_batch test
callers as 2 (1 production + 1 in-file unit test). Surfaced during
B1.0 implementation when cargo check --workspace --tests failed
with 5 arity-mismatch errors after B0's signature change.

Root cause: B0 audit's grep filter was `grep -v test` and didn't
enumerate crates/ml/src/trainers/dqn/smoke_tests/ (compiled as
part of the lib's test binary, not behind #[cfg(test)]) nor
crates/ml/tests/.

Sites fixed (zero-init i32 alloc, threaded through):
  - crates/ml/src/trainers/dqn/smoke_tests/training_stability.rs:152, 196
  - crates/ml/src/trainers/dqn/smoke_tests/performance.rs:142
  - crates/ml/src/trainers/dqn/smoke_tests/gpu_residency.rs:75
  - crates/ml/tests/gpu_per_integration_test.rs:125

No behavior change — the column carries zero data and no consumer
reads it pre-B1.1. B1.1 lands the producer kernel that fills with
-1/0/1 from the 30-bar price trajectory.

Process correction documented in docs/dqn-wire-up-audit.md
"B0.1 cascade-gap fix-up" subsection: future B-series audits must
run cargo check --workspace --tests before claiming cardinality
completeness.

Build: cargo check --workspace --tests clean.
Tests: cargo test -p ml --lib compiles + passes (GPU tests
#[ignore]-gated).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-05 10:46:20 +02:00
jgrusewski
62ab8ed850 feat(sp13): B0 — replay buffer i32 ring + GpuBatchPtrs plumbing
Pure plumbing: threads a new i32 column (aux_sign_labels) through every
layer of the replay path so B1 can wire the aux head's CrossEntropy
classification target without touching any aggregator or batch-shape
contract on its own. No consumer reads the column yet — all labels are
zero-initialized; smoke between B0 and B1 should be bit-identical to
parent 0ad5b6fa4 modulo allocator entropy.

Why split B0 from B1: original Layer B Commit 1 bundled replay plumbing
with the aux head contract change (1->2 dim, MSE->CE, slot 117 retire,
fingerprint bump). Four implementer subagents in a row hit
NEEDS_CONTEXT on brief-accuracy issues — too large for a single
dispatch. Split keeps B0 mechanical and isolates B1 for fresh
implementer with audit-derived brief.

Wiring (data flow):
  1. scatter_insert_i32 kernel mirrors scatter_insert_u32 but preserves
     -1 sentinels (u32 would alias to 4294967295). gather_i32_scalar at
     #18c is symmetric.
  2. ReplayKernels.scatter_insert_i32 field + ld() loader.
  3. GpuReplayBuffer.aux_sign_labels: CudaSlice<i32> ring [capacity] +
     sample_aux_sign_labels: CudaSlice<i32> per-batch gather dst.
  4. insert_batch signature gains aux_sign: &CudaSlice<i32>. Body
     scatter-launches alongside actions/rewards/dones using same
     (ci, cpi, bsi) tuple.
  5. sample_proportional Step 3c launches gather_i32_scalar from ring
     into sample buffer. Both GpuBatchPtrs return paths set
     aux_sign_labels_ptr to sample buffer raw_ptr.
  6. GpuBatchPtrs.aux_sign_labels_ptr: u64 publicly exposed.
  7. GpuExperienceBatch.aux_sign_labels field — collect_experiences_gpu
     alloc_zeros total-sized i32 slice (B1 replaces with producer kernel).
  8. Single production caller training_loop.rs:2032 threads through.

Audit-verified call site cardinality (per implementer 4 finding):
  - 1 production insert_batch caller (training_loop.rs:2032)
  - 1 in-file unit test caller (signature updated)
  - 2 GpuBatchPtrs construction sites (both inside sample_proportional)
  - 1 GpuExperienceBatch construction site (collect_experiences_gpu)

Implementer 4 caught the over-counted 4 sites claim from the original
brief — the 2 extras were doc-comment references to
insert_batch_with_episode_ids (no separate function exists).

Hard rules upheld:
  - feedback_no_partial_refactor: every insert_batch consumer migrates
    atomically (1 prod + 1 test, both updated)
  - feedback_no_stubs: column carries real data through real kernels;
    zero values are valid i32 that B1 overwrites with -1/0/1
  - feedback_isv_for_adaptive_bounds: no new ISV slots; slot 117
    AUX_LABEL_SCALE_EMA retirement deferred to B1 atomic
  - feedback_no_atomicadd: no new producers/reductions
  - feedback_cpu_is_read_only: alloc_zeros + scatter + gather all GPU

Build: cargo check --workspace clean in 21s.

Audit: docs/dqn-wire-up-audit.md SP13 Layer B B0 section added with
wiring diagram, call site cardinality table, and B1 next-steps for
fresh-implementer dispatch.

Files: 92 LOC net across 4 source files + audit doc:
  - crates/ml-dqn/src/replay_buffer_kernels.cu (+21)
  - crates/ml-dqn/src/gpu_replay_buffer.rs (+54)
  - crates/ml/src/cuda_pipeline/gpu_experience_collector.rs (+17)
  - crates/ml/src/trainers/dqn/trainer/training_loop.rs (+1)
  - docs/dqn-wire-up-audit.md (B0 section)

Next: B1 — aux head 1->2 dim, MSE->CE loss, aux_dir_acc reads softmax,
aux_pred_to_isv_tanh logit-diff rewrite, slot 117 retirement,
dqn_param_layout fingerprint bump, kernel populates aux_sign_labels
with real -1/0/1 labels, aux_b1_diag HEALTH_DIAG, 17+ unit tests.
B1 dispatched to fresh implementer with this audit doc as truth-source.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-05 09:43:00 +02:00
jgrusewski
bdc5cb8bb2 feat(sp13): P0b — aux_w deficit+stagnation controller + Hold cost lift
P0a smoke (train-67gqb on f934ea171) returned PARTIAL: aux_dir_acc climbed
0.149 → 0.483 in 10 epochs (signal exists), but val_win_rate stuck at 0.4638
and observed_hold_rate climbed to 0.479 despite controller saturating at
2.4× base. Two findings drive P0b:

(1) aux_w=0.05 (SP11-era clamp) starves the aux head of gradient. Replace
    inverted formula at training_loop.rs SP11 site with deficit+stagnation:
        deficit  = max(0, target - short_ema)
        improve  = max(0, short_ema - long_ema)
        stag     = (deficit > 0.005) ? clamp(1 - improve/deficit, 0, 1) : 0
        aux_w    = base × (1 + 5 × deficit) × (1 - 0.7 × stag),
                   clamped [0.3×base, 3.0×base]
    base 0.05 → 0.5 (10× lift). Stagnation decay prevents permanent
    destabilisation in data-limited case. Formula extracted as host helper
    `compute_aux_w_p0b(target, short, long)` for unit-test coverage.

(2) HOLD_COST_BASE=0.001 was too weak — max cost 0.005/bar × 30-bar hold =
    0.15 cumulative vs ±5-10 reward range = <3% of magnitude. Lift to 0.005
    (max 0.025/bar × 30 bars = 0.75 cumulative ≈ 10-15% of capped reward).
    Genuinely deters lazy Hold without crippling MFT use. Constructor
    static-init unchanged: it still writes the (now-lifted) HOLD_COST_BASE
    constant to slot 380 at fold boundary.

Per pearl_event_driven_reward_density_alignment tension already addressed
in P0a spec; the lift doesn't change the architecture, just the calibration.
Per feedback_isv_for_adaptive_bounds: base/gain/decay/floor/ceil are
numerical anchors; target/short/long EMAs read from ISV.

Tests: 3 new unit tests for controller formula (aux_w_at_target_returns_base,
aux_w_stagnation_decays_to_floor, aux_w_improving_amplifies_above_base).
14 SP13 GPU oracle tests + 14 SP12 reward-math tests still green; no kernel
changes.

Audit-doc updated with P0b section per Invariant 7.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-05 08:10:35 +02:00
jgrusewski
f934ea1719 feat(sp13): P0a atomic — Hold-pricing + dir_acc instrumentation (additive)
Tests user's hypothesis (Hold being FREE is the bug, not Hold itself) by
pricing Hold via ISV-driven adaptive controller targeting 20% Hold-rate.

11 new SP13 ISV slots [372..383). 5 new GPU kernels:
- aux_dir_acc_reduce_kernel.cu (correct/pos_pred/pos_label/valid → 3 scalars)
- hold_rate_observer_kernel.cu (packed batch_actions decode, count(Hold)/B)
- apply_fixed_alpha_ema_kernel.cu (preserves short/long timescale split that
  Wiener-optimal apply_pearls_ad_kernel would collapse)
- aux_pred_to_isv_tanh_kernel.cu (mean(tanh(aux_pred)) → ISV[375])
- 3 reward-composition sites in experience_kernels.cu subtract isv[HOLD_COST]
  on Hold actions (segment_complete pre-asymmetric-cap, positioned-non-event
  per-bar, flat per-bar)

Host-side controller in training_loop.rs:
  excess = max(0, observed - target)
  hold_cost = HOLD_COST_BASE × (1 + 5 × excess), clamped [0.5×, 5.0×base]

Per-step observer + EMA chain in gpu_experience_collector.rs after
experience_action_select. Per-epoch HEALTH_DIAG emit:
  aux_dir_acc target/short/long/pred_tanh
  hold_pricing observed_rate/target/cost

4-way action space stays (ExposureLevel::Hold preserved). Replay buffer /
fxcache compatibility preserved. SP11 (11/11) + SP12 (14/14) tests no
regression. SP13 P0a oracle tests: 14/14 on RTX 3050 Ti.

Spec/plan: docs/superpowers/{specs,plans}/2026-05-04-sp13-redefine-success-for-predictive-skill.md (v3)
Audit: docs/dqn-wire-up-audit.md (SP13 P0a section appended)

v2 → v3 reframe: P0a.T3 v2 implementer's audit found DirectionAction enum
doesn't exist (codebase uses 8-variant fused ExposureLevel cascading through
77 files). v3 reframes from "eliminate Hold" (250 LOC + 32-test cascade) to
"price Hold" (additive, no contract change, no cross-crate cascade).

Tension with pearl_event_driven_reward_density_alignment acknowledged in spec
— per-bar Hold cost is exposure-NEGATIVE (pulls policy AWAY from Hold-default,
inverse of the pearl's failure mode), models real economic carry, ISV-bounded
by controller. Faithful reward modeling, not artificial shaping.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-05 00:52:54 +02:00
jgrusewski
a1681abc46 Revert "exp(sp13): aux_w=1.0 override + directional accuracy metric for data investigation"
This reverts commit d2a27a0042.
2026-05-04 22:23:53 +02:00