Files
foxhunt/docs/plans/2026-05-12-sp22-h6-phase2-recenter.md
jgrusewski 7883c5ca1b docs(sp22): H6 Phase 2 spec — recenter state[121] to [-1, +1]
Post-mortem of the H6 Phase 1 smoke (WR = 50.21% verdict in
`docs/dqn-wire-up-audit.md`) traced an actual pearl violation in the
H6 implementation itself: state slot 121 uses the [0, 1] range
(`aux_softmax[env, 1] = p_up`) with sentinel 0.5, while every OTHER
state slot uses 0 as the "no signal" baseline (zero-padding,
feature_mask, ofi-missing, mtf-missing).

Per `pearl_first_observation_bootstrap`: "sentinel = 0; first
observation replaces directly." The H6 design violated this for
slot 121 alone. The encoder must learn TWO things about slot 121:
(1) the directional mapping AND (2) the appropriate bias offset for
the non-zero baseline — every other dim is single-step (mapping only).

Phase 2 is the simplest possible amplification fix: rewrite the bridge
to use the same convention as every other state slot. If the encoder
can't gradient-couple even after this fix, H6 is truly falsified and
we pivot to amplitude scaling or deeper hypothesis.

Change scope (atomic per `feedback_no_partial_refactor`):
- aux_softmax_to_per_env_kernel.cu: write `2*p_up - 1` instead of `p_up`
- gpu_experience_collector.rs: cold-start + FoldReset sentinel 0.5 → 0.0
- experience_kernels.cu: NULL-fallback in 3 state-gather kernels
  0.5f → 0.0f
- state_layout.rs / state_layout.cuh: comment updates to reflect
  [-1, +1] range and 0 sentinel

Estimated effort: ~45 min walltime (edits + verification gates) +
~25–40 min smoke wall-clock.

Verdict criteria (cycle 1 WR + a_var [d/m/o/u] in HEALTH_DIAG[0]):
- WR > 50.5% → recentering binding, H6 + Phase 2 sufficient → justify A2
- a_var for mag/ord/urg > 1e-3 → sub-branches learning under recentered signal
- WR pinned at 50.1–50.2% → Phase 2 falsified, pivot to amplitude
  scaling or deeper hypothesis

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-12 22:23:24 +02:00

8.7 KiB
Raw Blame History

SP22 H6 Phase 2 — Recenter state[121] to [-1, +1]

Author: jgrusewski + assistant pair Date: 2026-05-12 Branch: sp20-aux-h-fixed Predecessor: SP22 H6 Phase 1 (docs/plans/2026-05-12-sp22-h6-aux-policy-state-bridge.md), falsified at cycle 1 (WR = 50.21% in workflow train-cr9hl @ commit 7fc979934; verdict in docs/dqn-wire-up-audit.md "Smoke result (2026-05-12) — H6 FALSIFIED"). Audit pointer: append findings to docs/dqn-wire-up-audit.md under ## 2026-05-12 — SP22 H6 Phase 2 implementation.

Why this is Phase 2 of H6, not H3

The smoke verdict supports the architectural reading "the bridge wires aux into state, but the policy can't extract directional signal in 3 epochs." Before declaring the H6 hypothesis fully falsified and pivoting to H3 (reward density), one pearl-violation in the H6 implementation itself was discovered during the post-mortem trace:

pearl_first_observation_bootstrap.md: "sentinel = 0; first observation replaces directly."

The H6 design used:

  • State slot encoding: state[121] = aux_softmax[env, 1] = p_up ∈ [0, 1]
  • Cold-start / FoldReset sentinel: 0.5 (mid-range of [0, 1])

Every other state slot uses 0 as its "no signal" baseline — zero-padding, feature_mask = 0, missing OFI features, missing MTF features. Slot 121 alone uses 0.5. The encoder must learn two things about state[121] before it can use the signal: (1) the directional mapping AND (2) the appropriate bias offset for this dim's non-zero baseline. Every other dim is single-step (mapping only).

This is the simplest possible amplification fix: rewrite the bridge so slot 121 uses the same convention as every other state slot. If the encoder can't gradient-couple even after this fix, then H6 is truly falsified and we pivot.

Change set

Three points of edit. Atomic per feedback_no_partial_refactor.

Change 1 — Recenter the copy kernel

File: crates/ml/src/cuda_pipeline/aux_softmax_to_per_env_kernel.cu

Current write at line ~73 (inside the kernel body):

prev_aux_dir_prob[env] = aux_softmax[(size_t)env * (size_t)K + 1];

New write:

/* SP22 H6 Phase 2 (2026-05-12): write recentered p_up so the
 * encoder sees the same "no signal" = 0 baseline as every other
 * state slot, per pearl_first_observation_bootstrap. Range
 * [-1, +1]; 0 = neutral / no aux conviction; +1 = "up with full
 * conviction"; -1 = "down with full conviction". The softmax
 * components are non-negative and sum to 1, so 2*p_up - 1 is
 * structurally bounded in [-1, +1]. */
prev_aux_dir_prob[env] =
    2.0f * aux_softmax[(size_t)env * (size_t)K + 1] - 1.0f;

Update the kernel-file header comments to reflect the new range and sentinel value.

Change 2 — Update sentinel value (collector cold-start + FoldReset)

File: crates/ml/src/cuda_pipeline/gpu_experience_collector.rs

Two let sentinel: f32 = 0.5; literals (one in new() cold-start fill, one in reset_episodes() FoldReset) become let sentinel: f32 = 0.0;. Update the surrounding comments and the field-doc on prev_aux_dir_prob.

Change 3 — Update kernel-side NULL fallback

File: crates/ml/src/cuda_pipeline/experience_kernels.cu

Three state-gather kernels currently fall back to 0.5f when aux_dir_prob_per_env == NULL (Phase 1 A3 path):

float aux_dir_prob = (aux_dir_prob_per_env != NULL)
    ? aux_dir_prob_per_env[i]   // or [w] in backtest kernels
    : 0.5f;

becomes:

float aux_dir_prob = (aux_dir_prob_per_env != NULL)
    ? aux_dir_prob_per_env[i]   // or [w] in backtest kernels
    : 0.0f;

Applies to experience_state_gather, backtest_state_gather, and backtest_state_gather_chunk — all three migrate atomically.

Change 4 — Constant + comment update in state_layout.rs

File: crates/ml-core/src/state_layout.rs

The doc-comment on AUX_DIR_PROB_INDEX (line ~5863) currently says "Sentinel = 0.5 (neutral; p_up = 50%)". Update to:

// SP22 H6 Phase 2 (2026-05-12) — recentered encoding [-1, +1] for
// gradient-init parity with other state slots (per
// pearl_first_observation_bootstrap; every other slot uses 0 as the
// "no signal" baseline). Sentinel = 0.0 (neutral; aux contributes
// nothing). +1 = up with full conviction; -1 = down with full
// conviction. The kernel writes `2*p_up - 1` post H6 Phase 2.
pub const AUX_DIR_PROB_INDEX: usize = PADDING_START;  // 121

Change 5 — assemble_state device-function comment

File: crates/ml/src/cuda_pipeline/state_layout.cuh

Update the inline comment on the padding block (around the out[SL_PADDING_START + 0] = aux_dir_prob write) to reflect the new range [-1, +1] and sentinel 0.

Out of scope

  • Reward shaping changes (SP12 + SP18 + SP20 are already pearl-aligned).
  • Per-branch IQN tau / C51 atom span (SP5 already wired; pearls pearl_per_branch_iqn_tau_schedule etc. already implemented).
  • V/A projection (SP17 already landed).
  • A2 (eval-side aux integration) — still deferred. The eval NULL fallback continues to use the kernel-side sentinel (now 0.0).
  • Aux temperature / magnitude scaling — IF Phase 2 also falsifies, THIS becomes the next candidate. Not in scope yet.

Verification gates

Identical to H6 Phase 1, all required clean before smoke dispatch:

  1. SQLX_OFFLINE=true cargo check -p ml --features cuda → 0 errors
  2. cargo test -p ml --test gpu_backtest_validation --features cuda --release → 4 expected-passing tests still pass (the 2 PnL-assertion failures are pre-existing per the H6 Phase 1 baseline)
  3. /usr/local/cuda/bin/compute-sanitizer --tool=memcheck <test_bin> → 0 CUDA errors

Smoke dispatch + verdict criteria

./scripts/argo-train.sh dqn --baseline --branch sp20-aux-h-fixed \
  --epochs 3 --gpu-pool ci-training-l40s

Watch: cycle 1 WR + a_var [d/m/o/u] in HEALTH_DIAG[0].

Decision:

  • WR > 50.5% within 3 epochs → recentering was the binding constraint; H6 + Phase 2 is sufficient. Justify A2 (eval-side aux integration) for production parity. Capture verdict in audit doc.
  • a_var for mag/ord/urg moves off 0 (> 1e-3) → secondary success signal: sub-branches are getting enough gradient to learn under the recentered aux signal.
  • WR pinned at 50.150.2% → Phase 2 falsified. The encoder successfully consumes a recentered signal but still can't extract directional alpha from one state slot in 3 epochs. Pivot to:
    • (a) Aux temperature / magnitude amplification — multiply state[121] write by some scalar > 1 to widen its dynamic range relative to other dims. Smallest-possible follow-up.
    • (b) True H3 (reward density audit, deeper than already done) or per-branch optimization audit.

Kill criteria (per feedback_kill_runs_on_anomaly_quickly):

  • Compile errors that can't be resolved in 30 min → checkpoint
  • compute-sanitizer reports any non-zero CUDA errors → fix before dispatch
  • Cycle 1 WR firmly in 50.150.2% band → Phase 2 falsified, kill and pivot

Effort

Phase Effort Risk
Edits (5 small files) 20 min Low (mechanical)
Verification gates 10 min cargo check + 5 min test + 2 min sanitizer Low
Commit + push + smoke dispatch 5 min Low
Smoke wall-clock ~37 min (with cache miss on first build) or ~20 min (cache hit) (waiting)
Total ~45 min walltime + ~2540 min smoke

Risk + mitigations

Risk Mitigation
Replay-buffer rows captured pre-Phase-1 with sentinel 0.5 still cycling through training Replay buffer was created at trainer start; pre-Phase-1 rows would only exist if we were resuming from a checkpoint. Baseline smoke from scratch has no prior replay. No mitigation needed.
Recentered signal's mean is exactly 0.5 of the original [0, 1] mean (no bias shift) but the variance is 4× wider ((2x-1).var = 4*x.var). Could cause early grad-clip events. Variance shift is the intended effect: wider dynamic range = stronger encoder gradient. Grad-clipper handles outliers (already saw clipping at step ~75 in Phase 1 smoke; same mechanism handles Phase 2).
The kernel-side NULL fallback (0.0f) is used by eval at A3. Eval is currently unused for verdict-determination (training-side WR is the verdict signal), but if A2 ever lands, the 0.0 is the right cold-start. None needed; already a future-compatible choice.

Falsification confidence

Recentering is information-theoretically the minimum-cost intervention. If WR doesn't move, we have very strong evidence the bottleneck is not in the state[121]→encoder coupling, narrowing the next investigation to amplitude scaling or downstream pathology. The kill-fast / learn-fast discipline of feedback_kill_runs_on_anomaly_quickly applies — one smoke (~30 min wall-clock) gives a clean signal either way.