Commit Graph

35 Commits

Author SHA1 Message Date
jgrusewski
c146c4fffd feat(sp15): scaffold sp15_isv_slots.rs with 46 slots [397..443) — ISV_TOTAL_DIM 396→443
Per spec §4.3 allocation map. Pre-allocates disjoint slot ranges to
enable Approach B parallel sub-worktrees without index collisions:
  - Phase 0.B EGF retune: [397..401)
  - Phase 1.3 drawdown: [401..407)
  - Phase 1.2 cost: [407..409)
  - Phase 1.4 baselines: [409..417)
  - Phase 3.X-3.5.X teachings + recovery: [417..441)
  - Phase 3.5 deferred anchors: [441..443)

Layout fingerprint extended with all 46 slot names. Pre-SP15 checkpoints
will be incompatible (greenfield OK per Q1).

Two regression tests verify: (1) every slot < ISV_TOTAL_DIM, (2) layout
fingerprint locked at named indices. docs/isv-slots.md gets the SP15
section documenting the allocation map + greenfield sub-worktree plan.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-06 10:49:33 +02:00
jgrusewski
c0fc28e455 fix(sp14): delete warmup_gate — let variance-driven k_aux/k_q handle warmup (ISV-driven)
Per `feedback_isv_for_adaptive_bounds`, the hardcoded
`warmup_gate = (fold_step_counter / WARMUP_STEPS_FALLBACK).min(1.0)`
ramp violated the rule: adaptive bounds in ISV, never hardcoded
constants. The variance-driven k_aux/k_q sigmoid steepness already
provides warmup behavior intrinsically:

- High variance (cold-start, EMAs still moving) → k → K_MIN → flat
  sigmoid → gate ≈ 0.5 regardless of input. That IS the warmup.
- Low variance (settled) → k → K_BASE → sharp sigmoid → gates
  respond correctly to driver signals.

Adding a separate hardcoded step-counter multiplier on top was
double-counting + tuning-driven (the 1000-step threshold had no
principled basis). Removed entirely.

Removed (per `feedback_no_partial_refactor`, all atomically):
- `WARMUP_STEPS_FALLBACK` constant in `sp14_isv_slots.rs`
- `warmup_gate: f32` parameter in `alpha_grad_compute_kernel.cu`
- `gate1 * gate2 * warmup_gate` → `gate1 * gate2` in kernel
- `warmup_gate` arg from `launch_sp14_alpha_grad_compute`
- `fold_step_counter: usize` field on the trainer struct
- `fold_step_counter = 0` reset in `reset_for_fold`
- `fold_step_counter` init in trainer constructor
- `let warmup_gate: f32 = 1.0;` and `.arg(&warmup_gate)` from B.4
  oracle tests (4 launches: 2 in alpha_grad_schmitt_hysteresis,
  20 in alpha_grad_adaptive_beta loop)

Build: clean, 18 warnings (pre-existing baseline).
Tests: cargo test --no-run on sp14_oracle_tests succeeds.

Net result: EGF gate's warmup behavior now lives entirely in the
variance-driven k_aux/k_q sigmoid steepness controller (ISV slots
388/var_aux, 389/var_q). No hardcoded step counter. Honors
`feedback_isv_for_adaptive_bounds` and `pearl_controller_anchors_isv_driven`.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-05 22:07:55 +02:00
jgrusewski
60ad42676e fix(sp14): bump ISV_TOTAL_DIM 383 → 396 to cover SP14 EGF slots
ROOT CAUSE of L1+L2 from Smoke A2-B: the ISV bus was sized for top
of SP13 (ISV_TOTAL_DIM=383) but B.1 allocated SP14 slots at 383-395.
Every SP14 read/write was OUT-OF-BOUNDS memory access. That's why:

- gate1 (slot 391) read as 0 always (OOB zero-init memory)
- post_open_min (slot 394) accumulated garbage values 9.5 → 28 → 46
- α_smoothed/α_raw values appeared to work but were undefined behavior

SP4/SP5 had a regression test (`all_sp4/5_slots_fit_within_isv_total_dim`)
that catches this exact failure mode at unit-test time. SP14 was missing
it — that gap let the bug ship across all 16 commits without being caught.

Changes:
- ISV_TOTAL_DIM: 383 → 396 (covers SP14 slots 383-395)
- layout_fingerprint_seed: extended with SP14 slot names + new
  ISV_TOTAL_DIM=396 marker (forces fingerprint hash bump per
  Invariant 8 — old checkpoints invalidated correctly)
- sp14_isv_slots.rs: 2 regression tests (mirror SP4/SP5 patterns)

Both tests pass. After this fix, SP14 EGF kernels will read/write
the correct slots; gate1 should actually flip open when aux_dir_acc
crosses target+0.03; gradient_hack_detect post_open_min stays bounded
in [0, 1] as designed.

NOT yet addressed (separate follow-up):
- warmup_gate hardcoded WARMUP_STEPS_FALLBACK=1000 violates
  feedback_isv_for_adaptive_bounds. Should be ISV-signal-driven OR
  removed entirely (k_aux/k_q already provide variance-driven warmup).
  Redesign post-re-smoke once bus-size fix is verified.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-05 21:59:47 +02:00
jgrusewski
e41dbb7d8a diag(sp14 B.12): per-epoch pearl_egf_diag HEALTH_DIAG emit
Adds a new HEALTH_DIAG[{epoch}]: pearl_egf_diag line immediately after
the aux_moe block in the per-epoch metrics section of training_loop.rs.
Reads all 13 SP14 ISV slots [383..396) — α_smoothed, α_raw, β, k_aux,
k_q, var_aux, var_q, var_α, q_dis_short, q_dis_long, gate1 state,
post_open_min, lockout — via the established read_isv_signal_at pattern,
giving forensic visibility into EGF pearl state each epoch.

gate1/gate2 sigmoid outputs are intentionally omitted: recomputing them
host-side would violate feedback_no_cpu_compute_strict; the sigmoid
inputs are sufficient for a reader to infer the output values.

docs/isv-slots.md updated (Invariant 7): records B.12 HEALTH_DIAG wire-up.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-05 21:14:58 +02:00
jgrusewski
84de278dfe feat(sp14): B.2 — register 11 SP14 ISV slots for fold-boundary reset
Each EGF pearl EMA / state slot resets to its Pearl-A sentinel at fold
boundary, mirroring sp13_aux_dir_acc_short_ema / long_ema entries.
Atomic refactor (feedback_no_partial_refactor): both halves land
together — registry entry + reset_named_state dispatch arm.

Reset slots (11 total, sentinel in parens):
  - Q_DISAGREEMENT_SHORT/LONG_EMA (slots 383, 384) → 0.5
  - K_AUX_ADAPTIVE (385) → K_BASE_AUX = 20.0
  - K_Q_ADAPTIVE (386) → K_BASE_Q = 15.0
  - BETA_RATE_LIMITER_ADAPTIVE (387) → BETA_BASE = 0.5
  - AUX_DIR_ACC_VARIANCE_EMA, Q_DISAGREEMENT_VARIANCE_EMA,
    ALPHA_GRAD_RAW_VARIANCE_EMA (388, 389, 390) → 0.0
    (initial k = k_base, β = β_base via ISV-driven controllers)
  - GATE1_OPEN_STATE (391) → 0.0 (closed)
  - ALPHA_GRAD_SMOOTHED (393) → 0.0
  - AUX_DIR_ACC_POST_OPEN_MIN (394) → 1.0 (no min observed)

ALPHA_GRAD_RAW (slot 392, recomputed every step from variance EMAs)
and GRADIENT_HACK_LOCKOUT_REMAINING (slot 395, decays at epoch
boundary) are NOT in the fold-reset registry; both naturally
re-initialise without explicit reset.

Also corrects the isv-slots.md SP14 table: slots 392 and 395 were
incorrectly marked FoldReset in the B.1 entry; corrected to reflect
their actual reset semantics (NOT reset / epoch-boundary decay).

Producer + consumer wiring lands in subsequent tasks (B.3-B.12);
this commit is additive infrastructure only — no behavior change.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-05 18:51:00 +02:00
jgrusewski
d63cb7992e feat(sp14): B.1 — sp14_isv_slots.rs with 13 new ISV slot constants
Allocates ISV slots [383..396) for the Aux→Q Wire + Earned Gradient
Flow pearl (Layer B of SP14). Mirrors sp13_isv_slots.rs pattern.
The plan originally documented [381..394), but Phase 0 verification
found SP13 closeout added HOLD_RATE_TARGET_INDEX=381 and
HOLD_RATE_OBSERVED_EMA_INDEX=382 after the plan was written, so the
range shifts by +2.

Slots fall into 4 functional groups:
- Q-disagreement EMAs (short, long; K=4↔K=2 mapping with Hold/Flat masked)
- Adaptive controllers (k_aux, k_q, β; variance-driven)
- Welford variance EMAs (3, one per adaptive scalar)
- Schmitt state + α_grad outputs + circuit breaker

Plus 14 structural constants for numerical-stability anchors:
K_BASE_*, K_MIN, VARIANCE_REF_*, BETA_BASE, BETA_MAX, SCHMITT_BAND,
WARMUP_STEPS_FALLBACK, LOCKOUT_*, Q_DISAGREEMENT_BASELINE.

Per feedback_isv_for_adaptive_bounds: adaptive bounds (k_*, β,
post_open_min, lockout) live in ISV; numerical anchors live as
structural constants. Per pearl_first_observation_bootstrap: all
EMAs reset to sentinels and Pearl-A bootstraps on first observation.

Producer + consumer wiring lands in subsequent tasks (B.2-B.12);
this commit is additive infrastructure only — no behavior change.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-05 18:45:49 +02:00
jgrusewski
b3b4d02789 fix(sp11): B1b smoke-recovery — z-score normalization for mag-ratio canary
Linear magnitude ratios in reward_component_mag_ratio_compute_kernel
amplified popart's intrinsic O(100) magnitude over the other 5
components' O(0.1-2) magnitudes, causing controller to saturate
w_pop toward MAX_WEIGHT regardless of actual signal quality.

Replaced with z-score: z[c] = mag[c] / max(sqrt(var[c]), EPS_DIV).
6 new ISV slots [361..367) for per-component variance EMAs computed
via Welford's online algorithm in extended popart_component_ema_kernel
and reward_component_ema_kernel.

Atomic per feedback_no_partial_refactor: slot allocation +
state-reset registry + 2 producer kernels + canary signature +
launcher Pearls A+D + tests.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-04 11:51:30 +02:00
jgrusewski
5e16b67ca6 fix(sp11): B1b follow-up — add slot 360 for popart-component mag EMA
Per spec §4 amendment at 52c0b7521 on main: B1b smoke surfaced that
SP11 mag-ratio canary was reading slot 63 (REWARD_POPART_EMA_INDEX)
which is overloaded — pre-SP11 PopArt's normalization input (total
reward mag EMA) was the same value as popart-component magnitude
because composition was inline accumulation. B1b decomposition exposed
the overload; controller emitted w_pop ≈ 2.0 based on contaminated
ratio → 10× sharpe drop in smoke.

Resolution:
- Allocate ISV slot 360 = POPART_COMPONENT_MAG_EMA_INDEX
- Add popart_component_per_sample mapped-pinned buffer + write site
  in experience_env_step at the r_popart assignment
- New popart_component_ema_kernel.cu writes slot 360 (single-block
  block-tree-reduce per feedback_no_atomicadd)
- mag-ratio canary kernel signature changes from single
  popart_ema_base_slot to (popart_specific_slot, cf_others_base_slot)
  pair so it reads non-contiguous slot 360 + slots 64..68
- Reset registry: sp11_popart_component_mag_ema entry + dispatch arm
- Slot 63 (PopArt's input) UNCHANGED — pre-SP11 invariant preserved

ISV total: 360 → 361. SP5_SLOT_END = 361. SP5_PRODUCER_COUNT = 187.

cargo check + build clean; SP11 GPU oracle tests pass (6/6 including
updated mag_ratio test with 2 slot-index args); sp5_isv_slots layout
tests pass (10/10 with 185 unique slots / 187 linear span); state
reset registry tests pass (4/4 with new sp11_popart_component_mag_ema
entry + dispatch arm). Local multi_fold_convergence smoke gated on
data volume (175k bars on local fxcache vs 10-month walk-forward
requirement); validation deferred to L40S Argo run on PVC data per
the spec's pass criterion.
2026-05-04 10:29:47 +02:00
jgrusewski
034ba16801 feat(sp11): B1b — structural reward composition refactor (production flip)
Per spec §3.5.3 amended at 7ddaf9c51 on main: experience_env_step
reward composition decomposed from 8+ inline accumulation sites into
explicit per-component locals (r_popart, r_cf, r_trail, r_micro,
r_opp_cost, r_bonus), then composed as Σ w_i × r_i with controller
weights from ISV[340..346).

Trail reward extraction (§3.5.4): trail-fire P&L now flows through
r_trail (forced-exit signal) instead of r_popart (voluntary-exit
signal). REWARD_TRAIL_WEIGHT_INDEX has real signal — controller can
weight forced-exit vs voluntary-exit P&L differently. rc[2] (the
prior structural-placeholder slot) now carries trail magnitude.

Universal post-composition modifiers (§3.4.4): drawdown / capital-
floor / inventory / churn / conviction-scale / cf-flip apply AFTER
the weighted Σ, unweighted. They are risk constraints and structural
operators, NOT learning components — agent cannot weigh them away.

Mean=1 normalization (B0, §3.4.3): weights normalize to mean=1 so per-
bar `w_active × r_active` ≈ pre-SP11 absolute scale on average.

Sentinel-defense: experience_env_step runs at start of epoch, SP11
controller runs at end (training_loop.rs ~3475). At fold 0 epoch 0
step 0 the controller has not yet emitted, so ISV[340..346) hold
sentinel 0. Defense: fmaxf(w_raw, 0.01) — same Invariant-1 hard floor
the controller enforces post-renorm. Cold-start scale = 1% of
pre-SP11; Pearl A bootstrap on first emit replaces sentinel.

cf_reward path: out_rewards[cf_off] now writes controller-weighted
cf reward (w_cf × r_cf with sentinel-defense). Loss-kernel
cf_weight=0.3f at mse:318/c51:789 (structural Q-blend, NOT reward
weight) UNTOUCHED per §3.5 amendment.

Mutual exclusivity preserved (popart / trail / micro / opp_cost):
exactly one path fires per bar; others stay 0. The cascade scalar
`reward` mirrors per-component locals so C.4/D.4b bonus blocks that
read in-progress trade reward (Q-cap pattern from
pearl_one_unbounded_signal_per_reward) keep bit-identical compounding.
After cascade, `reward` is overwritten with r_weighted; post-
composition modifiers operate on r_weighted as before.

This is the production-flip commit. Trainer is now on the SP11
controller end-to-end (modulo replay-time curiosity which lands in
B1c after Layer C audit per §3.5.5).

Verification:
  - cargo check + build clean (1m32s release).
  - 6/6 SP11 GPU oracle tests pass (none exercise env_step directly).
  - 14/14 contract tests pass (sp5_isv_slots=10, state_reset_registry=4).
  - Local smoke (RTX 3050 Ti, 20-epoch magnitude_distribution) verifies
    HEALTH_DIAG sp11_reward weights drift epoch-over-epoch:
      epoch 0: w_pop=1.000 w_cf=1.000 w_tr=1.000 ...   (uniform sentinel-defense floor → mean=1)
      epoch 3: w_pop=1.991 w_cf=2.036 w_tr=0.493 ...   (controller redistributes)
      epoch 9: w_pop=1.823 w_cf=1.887 w_tr=0.572 ...   (mean ≈ 1.0 preserved, Σ ≈ 6)
    EVAL_DIST bit-identical to B1a baseline (eq=0.803 eh=0.197 ef=0.000)
    — pre-existing magnitude eval-collapse pathology
    (project_magnitude_eval_collapse_kelly_capped) unchanged by B1b.

Audit doc updated (Invariant 7): docs/isv-slots.md SP11 section now
reflects Layer B status with B0/B1a/B1b/B1c rollout timeline.
2026-05-04 09:46:53 +02:00
jgrusewski
25eba79ad5 feat(sp11): A2 — controller kernel + SimHash novelty buffer
reward_subsystem_controller_kernel: 5 canaries → 10 outputs, true Z-score
(delta_ema/sqrt(var_ema)), sigmoid blending, weight renormalization to Σ=1,
saboteur post-clamp, curiosity permanent floor (0.2 × bound). Pearls A+D
chained on outputs per spec §3.4.1.

novelty_simhash_kernel: 42×16 random projection → 16-bit SimHash code,
1M-slot bucket count table for novelty signal `1/sqrt(1+count)`. Race-
tolerated update per feedback_no_atomicadd (under-counts bias novelty
UPWARD — safe direction).

novelty_simhash_proj_init_kernel: Philox-seeded GPU init for the
projection matrix (CPU is read-only per feedback_no_cpu_forwards).

HEALTH_DIAG `sp11_reward` line emits 10 outputs + improvement_z each
epoch. Reset registry: novelty hash table reset arm wired (closes the
A0 deferral); projection matrix is frozen at trainer init for run
lifetime, not reset.

All 20 SP11 slots populate every step. No consumer reads them yet —
training behavior unchanged from A1. 3 new GPU oracle tests pass on
RTX 3050 Ti (controller midpoint, weight renorm, saboteur clamp).

Spec: docs/superpowers/specs/2026-05-04-sp11-reward-as-controlled-subsystem.md §3.4 §3.5.2

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-04 02:26:05 +02:00
jgrusewski
91b48bc7a5 feat(sp11): A1 — three canary producer kernels (no behavior change)
Adds val_sharpe_delta + saboteur_engagement + reward_component_mag_ratio
GPU producers for the SP11 reward-as-controlled-subsystem chain. Each is
a single-block producer chained with apply_pearls_ad_kernel for Pearls
A+D smoothing per pearl_first_observation_bootstrap.md +
pearl_wiener_optimal_adaptive_alpha.md. All three write to slots in
[350..360) which no consumer reads yet — Layer A is additive; consumer
migration lands atomically in Layer B.

A1.1 — val_sharpe_delta_compute_kernel.cu
  Two-pass: writes raw delta + (delta - prev_delta_ema)^2 to scratch.
  Chained Pearls A+D (n_slots=2) → ISV[VAL_SHARPE_DELTA_EMA_INDEX=350,
  VAL_SHARPE_VAR_EMA_INDEX=351]. Host writes val_sharpe to mapped-pinned
  history[1]; rotation handled in training_loop.rs at val emit boundary
  (a literal already-computed value — no host-side compute, no htod_copy).

A1.2 — saboteur_engagement_compute_kernel.cu
  Per-bar |Δreward| > 0.01 × ISV[PNL_REWARD_MAGNITUDE_EMA_INDEX] check
  with block tree-reduce (no atomicAdd per feedback_no_atomicadd). The
  per-bar Δreward signal is produced by experience_env_step's saboteur
  perturbation site as `traded × |reward| × max(|eff_spread − 1|,
  |eff_slip − 1|)` — a structural proxy for the cost-differential the
  saboteur imposed on bars where the model traded. Single kernel-side
  emit (no parallel reward computation), per spec §3.3.1.
  Chained Pearls A+D → ISV[SABOTEUR_ENGAGEMENT_RATE_INDEX=358].

A1.3 — reward_component_mag_ratio_compute_kernel.cu
  Reads ISV[REWARD_POPART_EMA_INDEX..+6) (the SP4 reward-component
  magnitude EMAs), normalises to ratios, and mirrors popart magnitude
  into scratch[6] as a side-output. ONE non-pointer parameter
  (popart_ema_base_slot) — no _unused param per feedback_no_stubs.
  Two chained Pearls A+D launches:
    n_slots=6 → ISV[REWARD_COMPONENT_MAG_RATIO_BASE..+6)
    n_slots=1 → ISV[PNL_REWARD_MAGNITUDE_EMA_INDEX=359]
  (slots non-contiguous: 352..358 then 359.)

Wire-up (per feedback_wire_everything_up):
- 3 cubin entries appended to crates/ml/build.rs
- 3 kernel handles + val_sharpe_history_pinned (MappedF32Buffer[2]) +
  saboteur_delta_reward dev-ptr cache fields on GpuDqnTrainer
- 3 launchers (launch_sp11_*) + 1 setter (set_sp11_saboteur_delta_reward_buf)
- saboteur_delta_reward_per_sample buffer field on GpuExperienceCollector
- experience_env_step kernel signature extended with the new buffer arg;
  every call site in the same commit per feedback_no_partial_refactor
- training_loop.rs init wires collector→trainer setter; val emit boundary
  invokes launch_sp11_val_sharpe_delta_compute; per-epoch metrics block
  invokes launch_sp11_mag_ratio_compute then
  launch_sp11_saboteur_engagement_compute (mag_ratio first so the
  signal-relative threshold base is populated before the saboteur reader)
- SP5_SCRATCH_TOTAL grown 266 → 276 (10 new scratch slots: 2+1+7)
- docs/isv-slots.md SP11 section updated to reflect A1 producers

3 GPU oracle tests in crates/ml/tests/sp11_producer_unit_tests.rs
pass on RTX 3050 Ti via MappedF32Buffer fixtures (zero htod_copy /
dtoh_sync_copy / alloc_zeros — feedback_no_htod_htoh_only_mapped_pinned
compliant).

Note on Step 8a path: the plan offered two routes for the saboteur
Δreward producer — in-kernel diff emission OR a small dedicated
reader-of-existing-buffers. The existing reward path emits ONE reward
(not both with/without), so the dedicated-reader alternative was
infeasible. The in-kernel emission landed as a small write site at the
END of experience_env_step (after total_reward_per_sample is finalised),
threading saboteur_eff_spread/saboteur_eff_slip from the perturbation
site forward to the END via stack vars. Single new kernel parameter,
single new GPU-only buffer, single existing call site updated.

Spec: docs/superpowers/specs/2026-05-04-sp11-reward-as-controlled-subsystem.md §5
Plan: docs/superpowers/plans/2026-05-04-sp11-reward-as-controlled-subsystem.md (Task A1)

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-04 01:40:38 +02:00
jgrusewski
1e5a65912c fix(sp11): A0 follow-up — update stale wiener-buffer + isv-slots header
Code-quality review on bf3a32d63 found two stale references that need
SP11 numbers:

- training_loop.rs:6672 + state_reset_registry.rs:891 — sp5_wiener_state
  comments referenced the post-SP4/post-SP8 buffer sizes (543, 681);
  post-SP11 is (71 + SP5_PRODUCER_COUNT) × 3 = 771 floats. Replaced the
  literal sizes with formula form citing SP5_PRODUCER_COUNT directly so
  this drifts less in the future.

- docs/isv-slots.md header — "Current ISV_TOTAL_DIM" said 171 (post-SP4
  Task A1) while actual is 360. Updated header; SP11 section already
  appended at the end of the file.

No logic changes. Cargo check + sp5_isv_slots / state_reset_registry
tests still pass.
2026-05-04 01:11:11 +02:00
jgrusewski
bf3a32d63a feat(sp11): A0 — allocate 20 ISV slots [340..360) + 20 reset entries
Pure infrastructure. No producer kernels, no consumer reads. Existing
training paths trace identically because no consumer reads slots [340..360)
yet. Layout-fingerprint bumped to ISV_TOTAL_DIM=360.

Spec: docs/superpowers/specs/2026-05-04-sp11-reward-as-controlled-subsystem.md
2026-05-04 00:55:07 +02:00
jgrusewski
8df28b81c8 merge: SP6 Pearl 5 — IQN τ per-branch (replaces SP5 broken upload pattern)
Brings in worktree-agent-acd65ada (commit 89fadec24): per-branch IQN
τ schedules via 4 forward passes per step, each with one branch's τ
slab swapped in via mem::swap. ÷4 budget normalization at every
per-branch IQN call preserves total gradient magnitude.

CRITICAL FIX: this merge replaces the SP5 Layer B Pearl 5 implementation
that violated feedback_no_cpu_compute_strict. The old refresh_taus_from_isv
called upload_f32_via_pinned every step inside the CUDA Graph capture
region, triggering CUDA_ERROR_STREAM_CAPTURE_UNSUPPORTED and the
'parent graph capture failed (continuing ungraphed)' fallback observed
in smoke-test-qtn7c at SP5 HEAD.

SP6 Pearl 5 uses pre-uploaded τ slabs + mem::swap pointer arithmetic.

Conflict resolution in fused_training.rs:
- Pearl 5 worktree branched from pre-Pearl-2 HEAD; renamed
  iqn_budget → iqn_trunk in 2 sites (parallel + sequential paths) to
  match Pearl 2's compute_adaptive_budgets() new return signature.
- Removed redundant single-call apply_iqn_trunk_gradient(iqn_trunk)
  from post-join block — Pearl 5's per-branch loop above already
  applies iqn_budget_per_branch = iqn_trunk/4 to the IQN trunk gradient
  4 times, achieving the same total magnitude with per-branch τ.

cargo check + cargo test --lib (sp4 sp5 state_reset_registry: 13/13)
both clean.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-02 09:14:52 +02:00
jgrusewski
76d58406bd merge: SP6 Pearl 3 — NoisyLinear per-branch σ array
Brings in worktree-agent-a4d8a879 (commit ed3fa066b): per-branch σ via
[4]-element mapped-pinned device buffer. add_advantage_noise kernel
indexes σ by branch derived from action_idx % total_actions; Q-value
layout is branch-major contiguous so per-branch σ derivation requires
no forward-pass restructuring.

3 ExperienceCollectorConfig constructors updated.

Resolves Pearl 3 averaging from SP5 Layer B which collapsed 4 per-branch
σ values into a single scalar via training_loop.rs:1747.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-02 09:10:51 +02:00
jgrusewski
f62860ea82 feat(sp6): Pearl 2 — SAXPY per-branch loss budgets (correction-factor pattern)
SP6 sub-project 1 (Pearl 2): converts the 4 SAXPY launchers from a single
scalar budget (mean of 4 ISV branch slots) to per-branch differentiated
scaling via the correction-factor pattern.

Problem: SP5 Layer B read ISV[190..210) per-branch budget slots but
collapsed them to a scalar via sum/4.0, then passed the single scalar to
apply_c51_budget_scale / apply_cql_saxpy / apply_iqn_trunk_gradient.
Branch HEAD parameters received the same budget as trunk, defeating
per-branch differentiation.

Fix: compute_adaptive_budgets() now returns ([f32;4], [f32;4], [f32;4],
[f32;4], f32, f32, f32, f32) — four per-branch arrays + four trunk-mean
scalars. The trunk mean (D3 decision) is used for the full-buffer trunk/value
SAXPY call (preserving SP5 Layer B behavior for shared params). Branch HEAD
parameter slices receive a correction sub-launch:

  correction = branch_budget[b] / trunk_mean    (skip if |correction-1| <= 1e-6)

After both launches, branch HEAD slice is effectively scaled by
branch_budget[b], trunk/value is scaled by trunk_mean. No double-scaling.

New helpers added to GpuDqnTrainer:
  apply_c51_budget_scale_branch(branch_idx, correction): scale_f32_ungraphed
    on branch-slice [f32 elements], offset via padded_byte_offset.
  apply_cql_saxpy_branch(branch_idx, correction): saxpy_f32_aux on
    branch-slice of both grad_buf and cql_grad_scratch.

IQN trunk gradient: uses iqn_trunk (mean of 4 branch IQN budgets) — IQN
backward flows through trunk only; per-branch IQN routing is beyond SP6 scope.

HEALTH_DIAG: three new per-epoch info! lines emit per-branch c51/iqn/cql
budget arrays (dir/mag/ord/urg) after the intent_dist line.

State: per-branch budget arrays cached on GpuDqnTrainer
(last_*_budget_per_branch: [f32;4]) and on DqnTrainer
(last_*_budget_per_branch: Option<[f32;4]>) for diagnostics.

docs/isv-slots.md: updated Pearl 2 slot rows to reflect SP6 consumer wiring.

Verification: cargo check + release build clean (13 warnings, pre-existing).
13 sp5+sp4+state_reset_registry lib tests pass. sp5_producer_unit_tests
--no-run clean. Sanity grep for old sum/4.0 averaging pattern: empty.

Files changed (6): gpu_dqn_trainer.rs, fused_training.rs, constructor.rs,
mod.rs, training_loop.rs, docs/isv-slots.md. No Pearl 3 or Pearl 5 files touched.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-02 02:08:54 +02:00
jgrusewski
89fadec24a feat(sp6): Pearl 5 — IQN τ per-branch schedule (4 forward passes, ÷4 budget normalization)
GpuIqnHead gains 12 new CudaSlice<f32> buffers (online_taus_branch[4],
target_taus_branch[4], cos_features_branch[4]) allocated at construction time
via alloc_f32. Each slab is [B,N] for taus and [D,N] for cos_features — same
sizes as the existing main buffers.

refresh_taus_for_branch(branch_idx, tau5): uploads one branch's 5-quantile
τ schedule from ISV[IQN_TAU_BASE + b*5 .. +5] to per-branch slabs with cold-start
floor (FIXED_TAUS[q] when ISV slot is zero). No cross-branch averaging.

activate_branch_taus(b) / deactivate_branch_taus(b): symmetric mem::swap helpers
install/restore one branch's slab into self.online_taus/target_taus/cos_features
for a per-branch IQN forward pass. activate→deactivate(b) is its own inverse.

fused_training.rs:
- Tau refresh block calls refresh_taus_for_branch(b, tau5) for all 4 branches,
  then refresh_taus_from_isv for the averaged main buffer (CVaR backward compat).
- grad_decomp_snapshot_iqn() moved BEFORE the parallel/sequential fork so the
  snapshot is taken before any of the 4 per-branch apply_iqn_trunk_gradient calls.
- Parallel path: 4 sequential IQN passes on iqn_stream; after each pass, event
  sync to main stream, apply_iqn_trunk_gradient(iqn_budget/4), re-fork so next
  pass starts after main has consumed d_h_s2_buf. iqn_done_event recorded after
  all 4 passes.
- Sequential path: 4 sequential IQN passes on main stream; apply_iqn_trunk_gradient
  (iqn_budget/4) inline after each pass while d_h_s2_buf holds that branch's result.
- Post-join: single apply_iqn_trunk_gradient removed (now inline); target_ema_update
  and PER loss cast remain.

÷4 normalization: both apply_iqn_trunk_gradient call sites use iqn_budget_per_branch
= iqn_budget / 4.0_f32 so 4 × (budget/4) = budget total — matching SP5 Layer B
gradient magnitude contract exactly.

docs/isv-slots.md: add SP6 Pearl 5 consumer wiring section under the SP5 table.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-02 02:08:45 +02:00
jgrusewski
ed3fa066b9 feat(sp6): Pearl 3 — NoisyLinear per-branch σ array
Replace ExperienceCollectorConfig.noise_sigma: f32 with
noise_sigma_per_branch: [f32; 4] (branch order: dir/mag/ord/urg).

add_advantage_noise kernel (experience_kernels.cu) now takes
const float* noise_sigma[4] + b0/b1/b2/b3 branch-size params.
Each thread derives its branch_idx from action offset using cumulative
branch size offsets; applies that branch's sigma. Sigma=0 fast-exits
with no PRNG work.

GpuExperienceCollector gains noise_sigma_dev: MappedF32Buffer[4]
(mapped-pinned, zero HtoD copy per feedback_no_htod). CPU writes the
4 sigma values via write_from_slice before each kernel launch; kernel
reads via dev_ptr.

training_loop.rs reads ISV[NOISY_SIGMA_BASE..+4] = ISV[210..214)
directly — one slot per branch — instead of averaging all 4 into a
scalar. Cold-start floor 0.01 is Invariant 1 (numerical stability).
Falls back to [hyperparams.noise_sigma; 4] when fused_ctx unavailable.

Default::default() supplies [0.1; 4]. mod.rs test (line 895) uses
Default::default() unchanged — no explicit field to update.

docs/isv-slots.md updated to reflect SP6 Pearl 3 consumer wired.

Files changed: 4 (experience_kernels.cu, gpu_experience_collector.rs,
training_loop.rs, docs/isv-slots.md). No Pearl 2 or Pearl 5 files touched.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-02 02:05:42 +02:00
jgrusewski
f91098e456 refactor(sp5): Task A0 quality fix-up — rustdoc + HashSet overlap test
Address two code-quality review items on commit 6dcaf1a1c:

1. Convert file-level + Pearl-section comments from `//` to rustdoc
   (`//!` module-level + `///` on first constant of each section).
   Matches sp4_isv_slots.rs style; SP5 module now `cargo doc`-discoverable
   as peer of SP4.

2. Replace spot-check assertions in slot_layout_no_overlaps_and_total_correct
   with a HashSet enumerating every slot reachable through every accessor.
   Asserts exactly 110 unique slots, min=174, max=285, the 2-slot carve-out
   gap (278, 279) absent, and the set equals {174..278} ∪ {280..286}.
   Test now actually verifies the no-overlaps invariant its name promised.

Also adds SP5 section to docs/isv-slots.md (Invariant 7 audit-doc update
required by pre-commit hook for cuda_pipeline component changes).

No constant values, accessor signatures, or fingerprint string changed.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-01 20:20:53 +02:00
jgrusewski
206ebd3558 fix(sp4): Task A1 review fix-ups — fingerprint seed + branch count + doc refresh
3 important + 1 optional findings from code quality review:

1. layout_fingerprint_seed() now lists all 40 SP4 slots and bumps
   `ISV_TOTAL_DIM=131` -> `ISV_TOTAL_DIM=171`. Without this, the fail-fast
   checkpoint-load guard would not detect the SP4 layout extension —
   a binary built against new code would falsely compare-equal to old
   checkpoints' fingerprints.

2. Added `SP4_BRANCH_COUNT=4` constant and `debug_assert!(branch <
   SP4_BRANCH_COUNT)` in `atom_pos_bound`. Test extended to verify
   atom_pos_bound max does not alias into WEIGHT_BOUND family.

3. Refreshed stale content in docs/isv-slots.md: ISV_TOTAL_DIM 96->171,
   fingerprint location [37..39)/[47..49) -> [115..117), table entries
   [94]/[95] -> [115]/[116].

4. Added `debug_assert!(group < SP4_PARAM_GROUP_COUNT)` to weight_bound,
   adam_m_bound, adam_v_bound, wd_rate accessors (symmetric with the
   atom_pos_bound change).

cargo check clean, cargo test slot_layout_is_contiguous_and_total_40 passes.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-30 22:02:41 +02:00
jgrusewski
c724fa99be feat(sp4): Task A1 — define 40 ISV slot indices for signal-driven bounds
Layer A additive: ISV_TOTAL_DIM 131 → 171. New module sp4_isv_slots.rs
defines TARGET_Q_BOUND, ATOM_POS_BOUND[4], WEIGHT_BOUND[8], ADAM_M_BOUND[8],
ADAM_V_BOUND[8], WD_RATE[8], GRAD_CLIP_BOUND, H_S2_BOUND, L1_LAMBDA_TRUNK.
40 slots total. ParamGroup enum for type-safe per-group launches.

No consumers wired yet — slots are reserved but unused. Behavior unchanged.
docs/isv-slots.md updated with the SP4 reservation table per Invariant 7.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-30 21:54:11 +02:00
jgrusewski
c247d9cab0 refactor(dqn-v2): delete legacy CPU-DtoH per_branch_vsn_mean / per_branch_target_drift
Per pearl_cold_path_no_exception_to_gpu_drives.md and explicit user
direction ("remove the legacy paths!"), the two Plan-1-era host-side
DtoH+CPU-loop helpers are GONE — not just routed around. Replaced
with GPU kernel reductions writing to ISV slots.

Deleted (160 lines):
- GpuDqnTrainer::per_branch_vsn_mean()    — DtoH params slice + host abs+mean loop
- GpuDqnTrainer::per_branch_target_drift() — DtoH (target+online) slices + host RMS loop
- FusedTrainingCtx::per_branch_vsn_mean() and per_branch_target_drift() wrappers

Added (GPU-only producers, ~150 lines):
- target_drift_kernel.cu: 2-block reduction RMS(target − online) for mag/dir
  branches. 256-thread smem tree-reduce per block; thread 0 EMA-updates
  ISV slot via pinned device-mapped (no DtoH).
- ISV slots [92] TARGET_DRIFT_MAG_EMA_INDEX, [93] TARGET_DRIFT_DIR_EMA_INDEX
- Fingerprint shifted [90,91] → [94,95]; ISV_TOTAL_DIM 92 → 96
- Trainer accessors: branch_param_slice_indices(), target_params_buf_device_ptr()
- Collector launcher launch_target_drift_ema_inplace()
- Constructor cold-start writes for the 2 new slots

HEALTH_DIAG site refactor:
- vsn_mag/vsn_dir read from ISV[VSN_MAG_EMA_INDEX=87], ISV[88] (Task 5
  GPU-driven kernel produced these, replacing the legacy VSN scalars)
- drift_mag/drift_dir read from ISV[TARGET_DRIFT_MAG_EMA_INDEX=92], ISV[93]
- All 4 reads via FusedTrainingCtx::read_isv_signal_at — pinned device-mapped
  so host reads are coherent with GPU kernel writes without explicit DtoH

isv-slots.md: rows for [87..89] updated to reflect GPU-only producers
(replaces stale text claiming host-DtoH); 2 new rows for [92, 93];
fingerprint shifted to [94, 95]; ISV_TOTAL_DIM bumped 92 → 96.

Smoke fold-2 best Sharpe 100.45 at ep5; per-fold best_val_metric
3.75/10.09/20.21 (avg 11.35) — within Plan 3 T5 baseline (95-117 range).

Pearl validation: this commit demonstrates the cold-path-no-exception
rule applied retroactively. The legacy methods existed for an entire
plan generation; "cold path is fine" was the rationale. New rule:
if a value is computed (reduction/EMA/RMS), the compute is in a kernel,
period — regardless of frequency.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-25 11:02:23 +02:00
jgrusewski
cfc4ccb72f feat(dqn-v2): E.5 Mode A attention-focus ISV — fully GPU-driven
Plan 4 Task 5 Mode A. Light, ISV-only, no model parameters added,
no checkpoint break.

ISV tail-append:
- [87] VSN_MAG_EMA_INDEX
- [88] VSN_DIR_EMA_INDEX
- [89] MAMBA2_RETENTION_EMA_INDEX
- Fingerprint shifted [85,86] → [90,91]; ISV_TOTAL_DIM 87 → 92

Producer (attention_focus_ema_kernel.cu):
- Single kernel, 3-block multi-reduction, fully GPU-driven
- Block 0 reduces magnitude-branch VSN-weight slice of params buffer
- Block 1 reduces direction-branch VSN-weight slice
- Block 2 reduces mamba2_h_enriched buffer
- Each block: 256-thread smem tree-reduce → mean |x| → EMA-update
  ISV slot via pinned device-mapped (no explicit DtoH)
- Adaptive α matches Plan 3 Task 3 convention

GPU-only correction (per user direction "no dtoh, pinned memory"):
First-pass agent implementation used host-DtoH + CPU loop for the
Mamba2 retention scalar (mirroring the pre-existing
per_branch_vsn_mean() pattern). User flagged this as a violation
of spec §4.C.6 GPU-drives-CPU-reads even on cold path. Rewrote
the kernel to do all 3 reductions on-device via smem block-reduce,
reading params/mamba2 buffers via raw_ptr() and writing to pinned
ISV slots directly. Removed mamba2_retention_mean() from
GpuDqnTrainer + FusedTrainingCtx (was the host-DtoH culprit).

Read-only AttentionMonitor mirrors PlanThresholdMonitor pattern.
StateResetRegistry: all 3 slots FoldReset.

Smoke fold-2 best Sharpe 95.09, avg best_val_metric 10.98 — within
Plan 3 baseline range. ISV slots populate non-zero from epoch 2.

Pearl: cold-path is not an exception to GPU-drives-CPU-reads. If
the producer is computing a value (not just reading a CPU-side
constant), the compute belongs on GPU regardless of frequency.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-25 10:17:29 +02:00
jgrusewski
3cb083f182 feat(dqn-v2): B.3 + C.5 GPU-only replay seed warm-start + CQL α ramp
Plan 3 Tasks 8 + 9. Single commit because Task 9 directly consumes Task 8's
seed-fraction signal; no useful intermediate state.

ISV tail-append:
- [82] SEED_STEPS_TARGET_INDEX — config replay_seed_steps (CPU constructor write)
- [83] SEED_STEPS_DONE_INDEX — GPU-incremented per collect_experiences_gpu
- [84] SEED_FRAC_EMA_INDEX — adaptive EMA of (1 - done/target)
- Fingerprint shifted [80,81] → [85,86]; ISV_TOTAL_DIM 82 → 87

GPU-only design (per user direction "fully gpu driven no cpu involvement"):
- 4 scripted policies as ONE CUDA kernel (scripted_policy_kernel.cu)
- Per-sample policy mix (40% uniform LCG / 20% momentum / 20% mean-rev /
  20% vwap-deviation) deterministic by `i % 5`
- Action source switched at launch boundary (CPU per-epoch read of ISV slot
  decides which kernel to dispatch; the action computation itself is 100% GPU)
- No CPU physics mirror — existing GPU `experience_env_step` runs unchanged

seed_step_counter_update_kernel.cu:
- Single-thread cold-path; increments DONE, computes FRAC = max(0, 1-done/target)
- Adaptive α matches Task 3/4 convention (α_base × (1 + 0.5×|clamp(sharpe,±2)|))

cql_alpha_seed_update_kernel.cu (Task 9):
- target = config.cql_alpha × max(0, 1 - seed_frac)
- During seed phase (frac=1) → target=0 → CQL α decays to 0 (no pessimism on
  exploration data); as frac → 0 → CQL α ramps to config value
- Updates ISV[CQL_ALPHA_INDEX=48]; CQL gradient kernel reads slot 48 via
  pinned device-mapped ISV (Plan 1 Task 12 consumer pattern unchanged)

Registry: SEED_STEPS_DONE + SEED_FRAC_EMA both FoldReset; CQL_ALPHA flipped
SchemaContract → FoldReset; SEED_STEPS_TARGET stays SchemaContract.

Read-only monitors (mirror PlanThresholdMonitor / StateKlMonitor pattern):
- monitors/seed_monitor.rs — surfaces ISV[82..85) for HEALTH_DIAG +
  controller_activity smoke fire-rate
- monitors/cql_alpha_monitor.rs — surfaces ISV[48] + ISV[84] dependency

Smoke (RTX 3050 Ti, 3 folds × 5 epochs, dqn-smoketest profile with
replay_seed_steps=1000 override so seed phase completes mid-fold):
- All 3 folds saved best-checkpoint
- Fold 2 best Sharpe = 92.4938 at epoch 1 (target range 80-120) ✓
- Per-fold val_metric: f0=3.80 / f1=9.73 / f2=20.24 (loss-based)
- HEALTH_DIAG[3..4] cql_alpha=0.0500 with health=0.49 → base ≈ 0.10 from ISV[48],
  consistent with kernel ramping toward final×(1-frac); regime gate
  (1-regime)×health applies on top
- 11 cargo check warnings (matches pre-task baseline; no new warnings)
- 6/6 monitor unit tests pass (read/diagnose/observe×fire_rate)

Smoke override rationale: smoke runs ~200 samples per collect (4 episodes ×
50 timesteps) × 5 epochs × 3 folds ≈ 3000 total. Default 100k target would
keep entire smoke in seed phase. Override to 1000 lets the seed→network
transition complete mid-fold so the CQL α ramp is observable.

Per pearl_one_unbounded_signal_per_reward.md: cql_alpha is bounded (clamped
to config_final × (1 - seed_frac) ∈ [0, config_final]), composes safely with
downstream CQL loss.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-25 09:13:39 +02:00
jgrusewski
673eb66124 feat(dqn-v2): C.3 state-distribution KL divergence + adaptive amplification
Plan 3 Task 7.

ISV tail-append:
- [78] STATE_KL_TRAIN_VAL_EMA_INDEX — moment-match KL on OFI dims, EMA
- [79] STATE_KL_AMPLIFICATION_INDEX — Flat-trap escape multiplier ∈ [1,2]
- Fingerprint shifted [76,77] → [80,81]; ISV_TOTAL_DIM 78 → 82

Producer (state_kl_divergence_kernel.cu):
- Per-dim Gaussian moment-match KL summed over OFI block (32 dims)
- Adaptive α matches Task 3 convention: α_base × (1 + 0.5 × |sharpe|), α_base=0.05
- Amp formula: ratio = new_ema / prev_ema; target = 1 + clamp(ratio−1, 0, 1)
- Kernel-internal trailing-EMA-of-self pattern — no separate threshold ISV slot
- Single block, one thread per OFI dim, smem reduction, no atomicAdd, no DtoH

Val-side launch: GpuBacktestEvaluator exposes val_state_sample() returning
(ptr, n) into the chunked_states_buf the val pass populates. Train side uses
the experience-collector's existing states buffer. Both buffers in same
CUDA context; launch enqueues on collector's stream after compute_validation_loss.

Consumer (experience_kernels.cu):
- B.1 opp_cost ×= kl_amp (penalty intensifies on train/val drift)
- B.2 bonus ×= kl_amp (novelty bonus intensifies on drift)
- fmaxf(1.0, ISV[STATE_KL_AMP]) guard at consumers handles cold-start 0 → no-op

Per pearl_one_unbounded_signal_per_reward.md: amp ∈ [1,2] is bounded; stacks
safely with B.1's q_abs_ref unbounded multiplicand. No formula blow-out risk.

Constructor cold-start: KL=0.0, amp=1.0 (consumer no-op until kernel fires).
Registry: both slots FoldReset with cold-start values restored on fold boundary.
Read-only StateKLMonitor surfaces both slots in diagnose snapshot.

Smoke results: All 3 folds pass cleanly with diverse epoch convergence
patterns (best at ep5/ep2/ep3 — addresses prior concern about always-ep1
in fold 2). Per-fold best Sharpe: 105.13, 83.28, 103.49. Average
best_val_metric 10.64.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-25 02:26:09 +02:00
jgrusewski
8f59c1e3b8 feat(dqn-v2): B.4 adaptive plan threshold — upgrade PLAN_THRESHOLD_INDEX static→GPU-driven
Plan 3 Task 4.

ISV tail-append:
- [75] READINESS_EMA_INDEX — batch-mean readiness EMA (GPU-written)
- [49] PLAN_THRESHOLD_INDEX — producer upgraded from static constructor
  write to GPU kernel (same consumer path unchanged)
- Fingerprint shifted [73,74] → [76,77]; ISV_TOTAL_DIM 75 → 78

Producer (plan_threshold_update_kernel.cu):
- Single-block reduction of readiness_per_sample [N*L]
- Adaptive α = α_base × (1 + 0.5 × |clamp(sharpe, -2, 2)|); α_base=0.05
- Derived: threshold = max(0.1, 0.5 × readiness_ema) → ISV[49]

Consumer sites unchanged (experience_kernels.cu 4 sites + backtest_plan_kernel.cu).
The upgrade is producer-only; consumers keep reading ISV[49] as before but now
receive an adaptive value tracking the policy's actual readiness distribution
rather than a hardcoded 0.5 midpoint.

PlanThresholdMonitor (read-only observer) surfaces plan_threshold.eff +
plan_threshold.readiness_ema for HEALTH_DIAG / controller_activity smoke.

StateResetRegistry: PLAN_THRESHOLD flipped SchemaContract→FoldReset;
READINESS_EMA registered as FoldReset. Both fold-reset arms restore the
cold-start values (0.5 / 1.0) before the first kernel fire on the new fold.

No breaking changes to consumer API.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-25 01:35:07 +02:00
jgrusewski
b5d19c1004 feat(dqn-v2): B.2 ISV-driven trade-attempt bonus — novelty at Flat→Positioned
Plan 3 Task 3.

ISV tail-append:
- [71] TRADE_ATTEMPT_RATE_EMA — Flat→Positioned transition rate EMA
       (GPU-written by trade_attempt_rate_ema_update, adaptive α)
- [72] TRADE_TARGET_RATE — reference rate, CPU-frozen at epoch 5
- Fingerprint shifted [69,70] → [73,74]; ISV_TOTAL_DIM 71 → 75

Producer kernel (trade_rate_ema_kernel.cu):
- Single-block reduction of flat_to_pos_per_sample [N*L] (no atomicAdd)
- Adaptive EMA: α = α_base × (1 + 0.5 × |clamp(sharpe, -2, 2)|)
- α_base = 0.05 (matches reward_component_ema convention)
- Launched from training_loop alongside reward_component_ema

Consumer (experience_env_step):
- Flat→Positioned site: novelty = max(0, 1 - attempt/target)
- bonus = conviction_core × vol_proxy × novelty
- reward += shaping_scale × bonus; rc[5] captures bonus for ISV[68]
- Explicit freeze gate: target_raw > 1e-6f, so bonus is structurally
  inert pre-freeze (prevents spurious novelty=1.0 on epoch-1 when
  attempt_rate is still 0)

Epoch-5 freeze (training_loop):
- measured = ISV[TRADE_ATTEMPT_RATE_EMA]; floor at 0.001
- Prevents novelty from sticking at 1.0 post-freeze

StateResetRegistry: both slots registered as FoldReset with per-slot
reset dispatch arms in training_loop's fold-boundary path.

Smoke: multi_fold_convergence passes (fold 2 best Sharpe 87.55 —
 slightly above Task 2 baseline 85.6, within noise; bonus inert
 in 5-epoch smoke so training matches pre-B.2 baseline as designed).
cargo check clean at 11 warnings baseline.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-24 23:11:20 +02:00
jgrusewski
44539d8f4b feat(dqn-v2): C.2 reward-component attribution — 6 ISV EMAs + HEALTH_DIAG split
New GPU kernel reward_component_ema reads per-sample per-component
rewards [popart, cf, trail, micro, opp_cost, bonus] and updates 6 ISV
EMA slots via adaptive-rate EMA (α=0.05). CPU-side RewardComponentMonitor
is a read-only observer per spec §4.C.6; HEALTH_DIAG emits reward_split
line each epoch.

experience_kernels.cu extended with reward_components_per_sample [N*L,6]
output param; rc[0]=popart, rc[1]=cf, rc[3]=micro populated; rc[2,4,5]
are structural placeholders for Plan 3 B.1/B.2/C.4/D.4.

ISV slots 63-68 added (REWARD_{POPART,CF,TRAIL,MICRO,OPP_COST,BONUS}_EMA);
fingerprint shifted 61-62 → 69-70; ISV_TOTAL_DIM 63 → 71.
Layout fingerprint seed updated; auto-recomputes on next run.

StateResetRegistry: isv_reward_component_emas → FoldReset.
training_loop.rs: reset_named_state arm zeroes slots 63..69 at fold boundary.
docs/isv-slots.md + docs/dqn-wire-up-audit.md updated (+2 Wired rows, 107 total).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-24 22:08:40 +02:00
jgrusewski
3c18ebd637 feat(dqn-v2): D.8 Plan 2 Task 6C — TLOB cuBLAS port, train+val parity
Implement TLOB attention (OFI[32]→TLOB[16]) using existing DQN cuBLAS
infrastructure.  Random Xavier init, trained end-to-end via DQN reward.
Both training (experience_env_step) and val (backtest_plan_kernel) write
TLOB features to SL_TLOB_START=74, preserving state-dist parity.

State layout changes:
  - SL_TLOB_DIM=16 inserted at SL_OFI_START+SL_OFI_DIM (=74)
  - SL_MTF_START 74→90, SL_PORTFOLIO_START 90→106, SL_PLAN_ISV_START 98→114
  - SL_PADDING_START 105→121, SL_STATE_DIM 112→128
  - state_layout.rs mirrored: TLOB_START=74, STATE_DIM=128

New files:
  - cuda_pipeline/gpu_tlob.rs: GpuTlob with 9 cublasLt GEMM descriptors,
    forward/backward/adam_step, copy_params_from for val weight sync
  - cuda_pipeline/tlob_kernel.cu: tlob_sdp_forward, tlob_sdp_backward,
    tlob_write_states, tlob_read_grad_from_states kernels

Wiring:
  - fused_training.rs: TLOB forward before trunk; TLOB backward+Adam Phase 6
  - training_loop.rs: per-epoch ISV[60]=TLOB_REGIME_FOCUS_EMA update
  - gpu_backtest_evaluator.rs: val TLOB forward after each launch_gather,
    set_tlob_from_training syncs weights from training TLOB each epoch
  - metrics.rs: creates eval GpuTlob on evaluator stream, syncs params
  - gpu_dqn_trainer.rs: ISV_TOTAL_DIM 60→63, fingerprint indices 58/59→61/62

ISV slot audit: isv-slots.md updated (slot 60=TLOB_REGIME_FOCUS_EMA,
fingerprint shifted to 61/62, ISV_TOTAL_DIM=63).
Wire-up audit: dqn-wire-up-audit.md updated (gpu_tlob.rs + tlob_kernel.cu).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-24 21:38:42 +02:00
jgrusewski
d071501e44 feat(dqn-v2): D.2 per-branch gamma — GPU kernel + monitor (spec §4.D.2 + §4.C.6)
Replaces scalar GAMMA_EFF_INDEX=43 with 4 per-branch slots at 43-46
(DIR, MAG, ORD, URG). New per_branch_gamma_update_kernel.cu reads
per-branch v-range + health and writes 4 slots. Deletes
gamma_update_kernel.cu per feedback_no_partial_refactor.md.

Per-branch base/max: direction [0.92, 0.99] (trend horizon),
magnitude [0.88, 0.95], order [0.85, 0.93], urgency [0.80, 0.90]
(execution horizon). Kernel diverges each branch's γ from uniform
bootstrap based on its branch's Q-stats.

ISV slots downstream of 43 shift +3: KELLY (44→47), CQL_ALPHA (45→48),
PLAN_THRESHOLD (46→49), Q_P05_* (47→50..54), Q_P95_* (51→54..58),
fingerprint (55-56 → 58-59). ISV_TOTAL_DIM grows 57 → 60.
Layout fingerprint auto-recomputes from updated seed bytes.

c51_loss_kernel reads per-branch γ from ISV[43+d] inside the branch loop;
iql_compute_per_sample_support reads ISV[43+d] per-branch for half_w.
kelly_cap_update_kernel write index updated 44→47.
q_quantile_kernel slot bases updated 47/51 → 50/54.
state_layout.cuh ISV_PLAN_THRESHOLD_IDX updated 46→49.

CPU-side PerBranchGammaMonitor (gamma_monitor.rs) is read-only observer
of all 4 slots; read() returns mean, diagnose() exposes individual γ values
+ spread + health. No CPU-side γ computation (spec §4.C.6 GPU-drives).

Plan 2 Task 3. Spec §4.D.2 + §4.C.6.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-24 20:01:38 +02:00
jgrusewski
2cf9927647 feat(dqn-v2): C.1 quantile-based atom support (Plan 2 Task 1)
Replaces half = max(10*q_gap, 3*std_ema, floor) tuned constants with
quantile-based: half = max(|q_p95 - v_center|, |v_center - q_p5|)
clamped to [min_half_floor, abs_half]. Covers observed Q distribution
directly.

New GPU kernel q_quantile_reduce reads per-sample Q-values from q_out_buf,
sorts per branch (bitonic for power-of-2 N, quickselect otherwise), writes
ISV[Q_P05_*=47..51) and ISV[Q_P95_*=51..55). Cold-path per-epoch (4 blocks,
1 thread each). No atomicAdd.

ISV slots 47-54 added at tail (8 total). Fingerprint shifted to 55-56.
ISV_TOTAL_DIM grows 49 -> 57. Fingerprint seed updated; recomputes hash
at compile time (new value: 0xbf6c400c026d77e3).

Launch order: reduce_current_q_stats (populates q_out_buf) ->
q_quantile_reduce (writes ISV P5/P95) -> reduce_current_q_stats_per_branch
-> update_eval_v_range (reads ISV P5/P95 for half-width).

Bootstrap: q_p05 = v_min, q_p95 = v_max (matches current atom range at
cold start). FoldReset: isv_q_quantiles dispatch arm resets to bootstrap
values at fold boundary.

StateResetRegistry: isv_q_quantiles as FoldReset; dispatch arm added in
training_loop.rs::reset_named_state.

Docs: isv-slots.md rows [47..57) updated, fingerprint tail reference
corrected to [55..57). dqn-wire-up-audit.md: q_quantile_kernel.cu added.

Plan 2 Task 1. Spec §4.C.1.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-24 19:29:26 +02:00
jgrusewski
ac9bcab94a feat(dqn-v2): C.6 static ISV writes (Tasks 12, 15, 16) + slot pre-allocation
Lands the static-configuration branch of Plan 1 C.6: cql_alpha,
conviction_floor (no-op if IQL_BRANCH_SCALE_FLOOR already serves),
plan_threshold written to dedicated ISV slots at constructor. Consumer
kernels (CQL, backtest_plan, experience) read from ISV instead of
config fields / hardcoded literals.

Also pre-allocates ISV slots for the 6 upcoming GPU-kernel tasks
(atoms, gamma, kelly_cap, tau, epsilon outputs + CPU-born epoch inputs).
Those slots start at 0; GPU kernels in follow-up commits fill them.

New ISV slots:
- EPOCH_IDX_INDEX=39, TOTAL_EPOCHS_INDEX=40 (CPU-born inputs)
- EPSILON_EFF_INDEX=41, TAU_EFF_INDEX=42, GAMMA_EFF_INDEX=43,
  KELLY_CAP_EFF_INDEX=44 (GPU-written in follow-up tasks)
- CQL_ALPHA_INDEX=45, PLAN_THRESHOLD_INDEX=46 (static config; this commit)
- Task 15 confirmed no-op: IQL_BRANCH_SCALE_FLOOR_INDEX=36 already
  serves conviction-floor role (constructor + ISV read in iql kernel).

Layout fingerprint auto-updated via seed-byte edits; fingerprint
re-tail at [47..49). ISV_TOTAL_DIM 39 -> 49.

GpuDqnTrainConfig gains total_epochs field; fused_training.rs passes
hyperparams.epochs at construction; default 0 for smoke tests.

write_isv_signal_at bound extended from ISV_DIM(23) to ISV_TOTAL_DIM(49)
so tail slots are writable by CPU.

cql_alpha consumer: compute_cql_logit_gradients reads base from
ISV[CQL_ALPHA_INDEX] instead of config.cql_alpha; adaptive formula
(base x health x (1-regime_stability)) unchanged.

plan_threshold consumers: experience_kernels.cu (experience_state_gather,
experience_action_select, experience_env_step) and backtest_plan_kernel.cu
(backtest_plan_state_isv) read plan_thr from ISV[ISV_PLAN_THRESHOLD_IDX=46]
via isv_signals pointer already present in both kernels; null-guard defaults
to 0.5f for smoke-scale runs without ISV warmup.

ISV_PLAN_THRESHOLD_IDX=46 defined in state_layout.cuh (included by both
plan kernel files); value must match PLAN_THRESHOLD_INDEX in gpu_dqn_trainer.rs.

StateResetRegistry entries added for all new slots: SchemaContract for
TOTAL_EPOCHS/CQL_ALPHA/PLAN_THRESHOLD; FoldReset for EPOCH_IDX and
GPU-written per-fold slots.

No behavioural change: all threshold/base values remain at their prior
defaults; consumers now read adaptive values from ISV instead of baked-in
config or hardcoded literals.

Plan 1 Tasks 12, 15, 16 + pre-allocation for 9, 10, 11, 13, 14.
Spec §4.C.6 (2026-04-24 GPU-drives-CPU-reads revision).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-24 16:45:46 +02:00
jgrusewski
fbb8694a0b feat(dqn-v2): A.2 ISV layout fingerprint at ISV[37..39) (tail placement)
Implements spec §4.A.2 structural layout fingerprint with tail placement
rather than head placement (spec alternative: §4.A.2 Step 5.3 alt).

Head placement (ISV[0..2)) was rejected because isv_signals[0] and [1]
are actively written by the isv_signal_update kernel (Q-drift EMA and
gradient-norm EMA). Shifting those would require updating every literal
reference in experience_kernels.cu — a larger change than warranted for
pure contract enforcement. Tail placement leaves all existing indices
intact, touches zero kernel .cu files, and fulfils the same design contract.

Key changes:
- ISV_LAYOUT_FINGERPRINT_LO_INDEX = 37, HI_INDEX = 38 (u64 across 2×f32).
- LAYOUT_FINGERPRINT_CURRENT: u64 = FNV-1a of slot-list seed bytes.
  Value: 0x85d4d76b578a7c17. Any slot change updates seed bytes,
  which updates the hash automatically.
- Constructor writes fingerprint after zero-init; calls
  check_layout_fingerprint() to self-verify before returning.
- check_layout_fingerprint(): reads pinned slots [37..39), recomposes u64,
  fails-fast on mismatch with "retrain required" message.
- Error message does NOT mention migration as an option.
- Pre-commit hook rejects `fn migrate_isv|upgrade_isv` names — makes the
  no-migration rule structurally enforced (check_no_isv_migrations).
- ISV_TOTAL_DIM: 37 → 39.
- Zero existing index shifts (no kernel literal sites affected).
- StateResetRegistry entry renamed ISV_SCHEMA_VERSION → ISV_LAYOUT_FINGERPRINT.
- ResetCategory::SchemaContract docstring updated to remove "migration" framing.
- docs/isv-slots.md: updated table + design note for tail placement.

Tests: state_reset_registry 3 unit tests pass with renamed entry.
       cargo check -p ml clean (pre-existing warnings only).

Plan 1 Task 5. Spec §4.A.2 (tail-placement alternative).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-24 12:40:59 +02:00
jgrusewski
1353e71ae6 spec+plan(dqn-v2): layout fingerprint replaces schema version (§4.A.2)
User raised a backward-compat concern: integer "schema version"
semantically implies a family of coexisting versions with upgrade
paths between them, inviting the forbidden pattern
(feedback_no_legacy_aliases.md, feedback_no_partial_refactor.md).

Replace with a compile-time structural fingerprint:
- ISV[0..2) stores a u64 FNV-1a hash of the slot layout (split across
  two f32 lanes preserving raw bits).
- The hash is computed by a const fn over the slot list; any slot
  change automatically updates the fingerprint. No human decides
  "what version is this now".
- Checkpoint load is fail-fast only. Error message does NOT mention
  migration as an option.
- Pre-commit hook rejects any `fn migrate_isv|upgrade_isv` to make
  the no-migration rule structurally enforced (landed as part of
  Plan 1 Task 5 hook extension).
- StateResetRegistry entry renamed ISV_SCHEMA_VERSION →
  ISV_LAYOUT_FINGERPRINT (Task 2's landed code touched in Task 5's
  commit per no-partial-refactor).

Updated:
- spec §4.A.2 (fingerprint design + rationale)
- spec §5 landing-order note (fingerprint auto-updates on layout change)
- Plan 1 architecture line
- Plan 1 Task 2 StateResetRegistry test + entry naming
- Plan 1 Task 5 — full rewrite of implementation steps
- Plan 1 exit criteria #6
- Plan 2 pre-plan verification (grep fingerprint constants, not version==0)
- Plan 2 Task 6D.1 (update fingerprint seed, not bump version)
- Plan 2 Task 6D exit criteria #7
- Plan 3 pre-plan verification
- docs/isv-slots.md ISV[0..2) row

No code changed. Plan 1 Task 5 is not yet implemented — this commit
realigns the spec + plans so the implementer subagent works from the
corrected design.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-24 11:47:07 +02:00
jgrusewski
06989cfdf9 infra(dqn-v2): audit doc scaffolding + pre-commit enforcement
Plan 1 Task 1. Creates the five audit docs plus config/metric-bands.toml
that track Invariants 2, 7, 8 per the DQN v2 spec, and extends the
pre-commit hook with two checks:

  - component-adding commits must touch an audit doc (Invariant 7)
  - added code may not contain TODO/FIXME/XXX/HACK/TBD/unimplemented!/
    todo! markers (Invariant 9)

Tests: manually verified by staging a TODO-marked file; commit
rejected with the correct error message.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-24 10:25:27 +02:00