First half of Phase C. Lands the producer side of the K=3 trade-
outcome aux head's state bridge: new kernel populates a per-env
3-slot cache from the K=3 softmax tile every rollout step. The
consumer side (state gather reading from this cache → state slots
[121..124)) lands in Phase C-2.
Mirrors the K=2 head's existing aux_softmax_to_per_env_kernel exactly
at K=3:
- K=2: prev_aux_dir_prob[env] = 2*softmax[env, 1] - 1 (recentered)
- K=3: prev_aux_outcome_probs[env, k] = softmax[env, k] for k in [0, 3)
Changes:
- state_layout.rs: 3 new constants AUX_OUTCOME_PROFIT_INDEX = 121,
AUX_OUTCOME_STOP_INDEX = 122, AUX_OUTCOME_TIMEOUT_INDEX = 123.
PROFIT_INDEX aliases AUX_DIR_PROB_INDEX (same value, different
semantic). Phase C-2 flips slot 121's meaning from K=2's recentered
p_up to K=3's p_Profit.
- aux_outcome_softmax_to_per_env_kernel.cu: new kernel + cubin.
- gpu_dqn_trainer.rs: new SP22_AUX_OUTCOME_SOFTMAX_TO_PER_ENV_CUBIN
embed.
- gpu_experience_collector.rs: 2 new struct fields (cache buffer +
kernel handle); cubin load + alloc in constructor; struct-init;
per-step launch in rollout loop after K=3 forward.
- build.rs: kernel registered.
Encoding shift K=2 → K=3: K=2 used recentered [-1, +1] to match
"no signal = 0" baseline of every other slot. K=3 keeps raw softmax
probabilities [0, 1]. Cold-start sentinel 0.0 for all 3 slots =
"no prediction yet" (mask). The 3-slot natural distribution is more
informative than a scalar.
Dead-code status: producer populates cache every step but
experience_state_gather doesn't read from it yet — state slot 121
still receives K=2's prev_aux_dir_prob write. Phase C-2 swaps the
state gather's source from K=2 cache to K=3 cache (3-slot write).
Why split C into C-1 + C-2: experience_state_gather is a hot-path
kernel with many consumers. Updating it touches training collector,
eval-side backtest evaluator, Rust launcher arg list. C-2 lands that
as an atomic state-semantic flip; C-1 lands the GPU-side scaffolding
independently so the producer chain can be validated first.
Verification:
- cargo check -p ml clean.
- cargo test -p ml --lib → 1016/0 green.
Audit: docs/dqn-wire-up-audit.md Phase C-1 section.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Phase 1 post-mortem traced an actual `pearl_first_observation_bootstrap`
violation in my own H6 implementation: state slot 121 wrote
`aux_softmax[env, 1] = p_up ∈ [0, 1]` with sentinel 0.5, but every
OTHER state slot uses 0 as the "no signal" baseline (zero-padding,
feature_mask, ofi-missing, mtf-missing). The encoder had to learn TWO
things about slot 121 (directional mapping + non-zero bias offset)
instead of one. Phase 2 fixes the encoding to match the project
convention BEFORE declaring H6 fully falsified.
Mechanism change
────────────────
- `aux_softmax_to_per_env_kernel.cu` writes `2*p_up - 1 ∈ [-1, +1]`
instead of `p_up`. Still structurally bounded (softmax components
in [0, 1] sum to 1).
- Cold-start + FoldReset sentinel: 0.5 → 0.0 via the same pure-GPU
`fill_f32` path. No HtoD per
`feedback_no_htod_htoh_only_mapped_pinned`.
- NULL-fallback in 3 state-gather kernels (training +
backtest-per-step + backtest-chunk): 0.5f → 0.0f.
- Constant + device-function comment updates to document the
recentered encoding.
Atomic per `feedback_no_partial_refactor`: the encoding contract
spans 5 source files; partial migration produces inconsistent slot
semantics between training and eval.
Verification gates (all clean)
──────────────────────────────
- cargo check -p ml --features cuda: 0 errors, 21 pre-existing warnings
(parity with Phase 1 baseline)
- gpu_backtest_validation: 4/4 expected-passing tests still pass; 2
pre-existing PnL-assertion failures bit-identical to Phase 1
(confirms recentering does not perturb scripted-policy paths)
- compute-sanitizer --tool=memcheck: ERROR SUMMARY: 0 errors
Smoke dispatch deferred pending an orthogonal investigation into the
2 pre-existing gpu_backtest_validation failures (stale action
constants in the tests; addressed in a follow-up commit, NOT a
Phase 2 regression).
Verdict criteria (per spec, evaluated after smoke)
──────────────────────────────────────────────────
- WR > 50.5% within 3 epochs → recentering binding, H6 + Phase 2
sufficient → justify A2.
- a_var for mag/ord/urg > 1e-3 → sub-branches gradient-coupled under
recentered signal.
- WR pinned at 50.1–50.2% → Phase 2 falsified, pivot to amplitude
scaling or deeper hypothesis.
Refs
────
- docs/plans/2026-05-12-sp22-h6-phase2-recenter.md (spec)
- docs/plans/2026-05-12-sp22-h6-phase2-recenter-runbook.md (this plan)
- pearl_first_observation_bootstrap (sentinel = 0)
- feedback_no_partial_refactor (5-file atomic)
- feedback_no_htod_htoh_only_mapped_pinned (fill_f32, not HtoD)
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Wires the aux head's per-env directional probability into policy STATE
slot 121 (AUX_DIR_PROB_INDEX = PADDING_START + 0), preserving the trunk-
separation invariant from `pearl_separate_aux_trunk_when_shared_starves`.
H1 (label horizon) confirmed aux learns 78% dir-acc within-fold at H=200
but the policy was walled off; this commit conducts that signal through
the state input with a one-step lag.
Mechanism (rollout-time, collector-only)
────────────────────────────────────────
- State[env, t] reads `prev_aux_dir_prob[env]` (= p_up from step t-1).
- After aux forward at step t, the new copy kernel writes
`aux_softmax[env, 1]` → `prev_aux_dir_prob[env]` for step t+1.
- Cold-start + FoldReset seed the buffer to 0.5 (neutral; p_up = 50%)
via the pure-GPU `fill_f32` kernel — no HtoD per
`feedback_no_htod_htoh_only_mapped_pinned.md`.
- Launch sits in the same `isv_signals && trainer_params != 0` gate as
the aux forward, so when aux is skipped the cache keeps its previous
(sentinel or last-good) value instead of copying stale `alloc_zeros`.
Three state-gather kernels updated atomically (per
`feedback_no_partial_refactor.md`):
- `experience_state_gather` — training, reads `aux_dir_prob_per_env[i]`
- `backtest_state_gather` — eval (single-step), NULL → 0.5 (A3 fallback)
- `backtest_state_gather_chunk` — eval (chunked), NULL → 0.5 (A3 fallback)
`assemble_state` gained a 7th param `float aux_dir_prob` written to
`out[SL_PADDING_START + 0]`; the remaining 6 padding slots stay zero
for 8-alignment.
Phase 1 scope = training-side + eval A3 NULL fallback. A2 (aux trunk
forward in eval) is deferred per the runbook — gates on whether the
smoke moves WR off the 50.1–50.2% plateau.
New files
─────────
- crates/ml/src/cuda_pipeline/aux_softmax_to_per_env_kernel.cu
Modified
────────
- crates/ml-core/src/state_layout.rs (+AUX_DIR_PROB_INDEX)
- crates/ml/src/cuda_pipeline/state_layout.cuh (+assemble_state param)
- crates/ml/src/cuda_pipeline/experience_kernels.cu
(3 state-gather kernels + NULL-defensive sentinel)
- crates/ml/src/cuda_pipeline/gpu_experience_collector.rs
(per-env buffer + 2 kernel handles + cold-start fill + copy launch
+ FoldReset re-fill + state-gather arg)
- crates/ml/src/cuda_pipeline/gpu_backtest_evaluator.rs
(NULL aux_dir_prob_per_env at all 3 launchers for A3)
- crates/ml/src/cuda_pipeline/gpu_dqn_trainer.rs
(SP22_AUX_SOFTMAX_TO_PER_ENV_CUBIN static)
- crates/ml/src/cuda_pipeline/gpu_action_selector.rs
(EPSILON_GREEDY_CUBIN → pub(crate) so collector reuses fill_f32)
- crates/ml/build.rs (register new kernel)
- docs/dqn-wire-up-audit.md
(## 2026-05-12 — SP22 H6 implementation: Phase 1 entry)
Verification (all three gates clean, no smoke yet)
──────────────────────────────────────────────────
- cargo check -p ml --features cuda: 0 errors
- gpu_backtest_validation: 4/4 expected-passing tests still pass
(the 2 PnL-assertion failures are pre-existing per the runbook)
- compute-sanitizer --tool=memcheck: 0 CUDA errors
Refs
────
- docs/plans/2026-05-12-sp22-h6-aux-policy-state-bridge.md
- docs/plans/2026-05-12-sp22-h6-next-session-prompt.md
- pearl_separate_aux_trunk_when_shared_starves
- pearl_first_observation_bootstrap (sentinel = 0.5 cold-start)
- feedback_no_htod_htoh_only_mapped_pinned (fill_f32 not HtoD)
- feedback_no_partial_refactor (3 state-gather kernels atomic)
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Adds 12 new `SL_*_GROUP_BEGIN/END` macros to `state_layout.cuh` covering the
6 feature groups (market/ofi/tlob/mtf/portfolio/plan_isv), plus
`SL_NUM_FEATURE_GROUPS=6` and `SL_MAX_FEATURE_GROUP_DIM=42` (= largest group
dim, MARKET_DIM). Six new `static_assert`s anchor each group dim ≤ max and
confirm contiguity / start-at-zero / end-at-padding invariants.
Rust mirror in `crates/ml-core/src/state_layout.rs` exposes:
- `SL_NUM_FEATURE_GROUPS`
- `SL_MAX_FEATURE_GROUP_DIM`
- `FEATURE_GROUP_RANGES: [(usize, usize); 6]` (half-open `[begin, end)`,
`plan_isv` ends at `PADDING_START` — padding is not a group)
- `FEATURE_GROUP_NAMES: [&str; 6]`
Three `const _: () = assert!(...)` blocks validate (1) first range starts at
0, (2) last range ends at `PADDING_START`, (3) every adjacent pair
satisfies `ranges[g].end == ranges[g+1].begin` (no gaps), (4) every group
dim ≤ `SL_MAX_FEATURE_GROUP_DIM`.
Group dims at this commit: market=42, ofi=32, tlob=16, mtf=16, portfolio=8,
plan_isv=7 — total 121 = `PADDING_START`.
Prerequisite for Plan 4 Task 1B (E.1 Variable Selection Network — pre-trunk
per-group softmax-over-groups gating), Task 5 Mode B (per-group ISV
diagnostics), and Task 4 follow-on (group-aware encoder interface).
Additive only — ZERO production callers in this commit. No new module /
kernel / ISV slot / param tensor / Orphan row. No fingerprint change (group
ranges are derivable from existing `SL_*_START` constants and contribute no
new structural-hash invariant). cargo check clean at 11 warnings (workspace
baseline preserved); cargo build compiles all kernel cubins (state_layout.cuh
edit triggers full kernel rebuild — no compile errors).
Audit doc: 1 new entry under "Plan 4 Task 1A".
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Plan 3 Task 6b.
Portfolio-state tail-append:
- PS_REGIME_SHIFT_BAR = 40 (hold_time of first detected regime shift, 0 if none)
- PS_STRIDE 40 → 41
- All 6 hardcoded-stride sites migrated in lockstep
Detector (experience_kernels.cu):
- Adaptive threshold = clamp(0.25 × |clamp(sharpe, -2, 2)|, 0.05, 0.5)
- Fires first bar where |regime_now - PS_PLAN_ENTRY_REGIME| > threshold
- First-shift-only (short-circuits on non-zero PS_REGIME_SHIFT_BAR)
- Uses new ISV_SHARPE_EMA_IDX = 22 macro in state_layout.cuh
Consumer (segment_complete block):
- bars_late_frac = clamp(bars_late / hold_time, 0, 1)
- penalty = shaping × conviction × bars_late_frac × |reward|
- All multiplicands except |reward| in [0,1]; max penalty = |reward|
- reward -= penalty; rc[5] -= penalty (cancels with B.2/C.4/D.4a at
other (i,t) slots; ISV[68] REWARD_BONUS_EMA shows net)
- Consumer resets PS_REGIME_SHIFT_BAR after use
**Iteration history.** First pass multiplied by ISV[Q_DIR_ABS_REF] (~5–50,
an absolute Q-magnitude) AND |reward| — produced penalties 5–50× the
reward, destabilising training (smoke: Return swings ±300–900%, Sharpe
oscillating wildly). Root cause: Q_DIR_ABS_REF is an absolute
magnitude, not a [0,1] coefficient; B.1 uses it as a DENOMINATOR to
normalize q_range, not as a multiplier on an already-unbounded signal.
Fix: drop q_scale, keep |reward| as the only unbounded factor. Smoke
now passes cleanly with fold-2 best Sharpe 117.92 (up from T6a's 100.10).
No new ISV slot.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Plan 3 Task 6a.
Portfolio-state tail-append (shared-contract migration, all in same commit):
- PS_INTRA_TRADE_MIN_PNL = 39 (symmetric to PS_INTRA_TRADE_MAX_PNL = 21)
- PS_STRIDE 39 -> 40
- 6 PORTFOLIO_STRIDE hardcoded copies bumped in lockstep
Producer (experience_kernels.cu):
- MIN_PNL tracked per bar (fminf against pnl_pct) in the same block
as MAX_PNL update
- Reset to 0 at all 5 MAX_PNL reset sites (entry, reverse, 2x fold boundary)
Consumer (experience_kernels.cu segment_complete):
- Fires only on reward > 0 AND drawdown_depth > 1e-6
- persist_bonus = shaping x conviction x |min_pnl| x tanh(reward/|min_pnl|)
- reward += persist_bonus; rc[5] += persist_bonus (accumulates with
B.2 entry bonus + C.4 timing bonus — different (i,t) slots per trade)
Self-scaling via tanh: no tuned coefficients. Saturates when recovery
is large relative to drawdown; near-zero when recovery is trivial.
Attribution lands in ISV[68] REWARD_BONUS_EMA via the Task 1 kernel.
No new ISV slot.
Smoke: multi_fold_convergence PASS (fold-2 best Sharpe 100.10, threshold >=80).
HEALTH_DIAG reward_split bonus=17.21 (post-Task-5 rises with new D.4a credit
firing on profitable drawdown recoveries).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Plan 3 Task 5.
Portfolio-state tail-append (shared-contract migration, all in same commit):
- PS_PEAK_PNL_BAR = 38 (hold_time snapshotted when MAX_PNL updates)
- PS_STRIDE 38 -> 39 in state_layout.cuh and ml-core/state_layout.rs
- PORTFOLIO_STRIDE 38 -> 39 in trade_stats_kernel.cu (hardcoded copy)
- PORTFOLIO_STRIDE 38 -> 39 in gpu_experience_collector.rs allocator
- ps_stride 38 -> 39 in gpu_dqn_trainer.rs launch_kelly_cap_update
Producer (experience_kernels.cu):
- Peak bar snapshotted alongside every MAX_PNL update (uses local
hold_time, not ps[PS_HOLD_TIME], because the portfolio-state commit
block runs later in the kernel).
- Peak bar reset to 0 at every MAX_PNL reset site: plan-entry (1856),
entering_trade (2014), reversing_trade (2019), fold hard-reset (2736),
trade-complete soft-reset (2751).
Consumer (experience_kernels.cu segment_complete block):
- bars_early = max(0, segment_hold_time - PS_PEAK_PNL_BAR)
- timing_bonus = shaping_scale x (bars_early / segment_hold_time)
x |final_pnl| x conviction_core
- reward += timing_bonus; rc[5] += timing_bonus
(accumulates with Task 3 B.2 entry bonus — different (i,t) slots).
No new ISV slot — rc[5] bonus semantics unchanged; B.2 and C.4 share it
via += accumulate semantics (defensively idempotent, but the two sites
fire at distinct (i,t) by construction: entry vs exit).
Self-scaling: shaping_scale x conviction_core x |pnl| keeps the bonus
proportional to trade magnitude, no tuned coefficients.
Smoke multi_fold_convergence (RTX 3050 Ti): all 3 folds complete,
fold-2 best Sharpe 84.44 at epoch 1 (expected ~85 range).
cargo check --workspace clean at 11 warnings baseline.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Adds 7th plan_isv dimension: PLAN_ISV_REMAINING_FRACTION = max(0, min(1,
(plan_target_bars - hold_time) / plan_target_bars)) when plan active, else
0. Exposes temporal pressure to the policy.
State vector grows 104 → 112 (105 + 7 padding for 8-alignment).
SL_PORTFOLIO_PLAN_DIM 6 → 7. SL_PADDING_DIM 0 → 7.
Both training (experience_env_step) and val (backtest_plan_state_isv) write
the new slot identically — preserves train/val state-distribution parity.
All hardcoded stride-6 references in backtest_plan_kernel.cu replaced with
SL_PORTFOLIO_PLAN_DIM. plan_isv_buf allocation updated to n_windows * 7.
Stale offset comments ([86..92)) corrected to [98..105) across all files.
No behavioural change to existing dimensions. New signal is additive.
Plan 2 Task 6A. Spec §4.D.6.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
fix kernel-read gap
Adds 12 features to the DQN input pipeline:
- 10 MicrostructureState::snapshot()[0..10] slots that were previously computed
every bar and then discarded before reaching fxcache: ofi_trajectory,
realized_variance, hawkes_intensity, book_pressure (weighted 10-level),
spread_dynamics, aggression_ratio, queue_depletion_asymmetry,
order_count_flux, intra_bar_momentum, regime_score.
- 2 TLOB-novel slots derived directly from Mbp10Snapshot:
order_count_imbalance = (Σbid_ct − Σask_ct) / Σ(bid_ct + ask_ct),
microprice_residual = (weighted_mid − mid) / mid.
Also fixes a production gap: ofi_acceleration (slot 18) and
toxicity_gradient (slot 19) were persisted to fxcache via OFI_DIM=20
but the OFI embed kernel (experience_kernels.cu:6146-6173) only read
[0..18), silently discarding them every bar. Kernel extended to
consume full SL_OFI_DIM=32.
Dimension bumps (all 8-aligned):
OFI_DIM 20 → 32
FXCACHE_VERSION 4 → 5 (invalidates existing caches; regen via
precompute_features)
STATE_DIM 96 → 104
PADDING_DIM 4 → 0 (OFI expansion consumed padding, still 8-aligned)
STATE_DIM_PADDED 128 (unchanged)
OFI_EMBED_IN 18 → 32 (MLP input width; W/grad/Adam/m/v buffers
resized in lockstep via named constants)
fxcache regen results (175874 bars ES.FUT 2024-Q1):
deltas_nonzero: 175781 / 175874 (99.9 percent)
book_aggression: 102137 / 175874 (58.1 percent)
microstructure[20-30): 175874 / 175874 (100 percent)
tlob_novel[30-32): 133615 / 175874 (76.0 percent)
Compile status: SQLX_OFFLINE=true CARGO_INCREMENTAL=0 cargo check
--workspace --tests passes cleanly (0 errors, pre-existing warnings
only).
Test results:
fxcache roundtrip (unit + integration): PASS (4+6 tests)
magnitude_distribution smoke: ran through epoch 1 successfully
(OFI_DIAG fires, state_dim=104 confirmed, feature_dim=74 in
validation kernel); epoch 2 OOM on local RTX 3050 Ti (4 GB) —
expected hardware limit from state_dim growth. Full 20-epoch run
requires L40S/H100 CI verification.
multi_fold_convergence smoke: not verified locally (same VRAM
ceiling applies). L40S/H100 CI verification required.
The new slots follow the existing OFICalculator/MicrostructureState
pattern and consume signals already computed by ml-features — no new
crate, no ONNX, no stubs. All 12 sources were audited against their
implementation before persistence; every slot traces back to real
Mbp10Snapshot or MicrostructureState math.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>