diff --git a/docs/plans/2026-05-12-sp22-h6-phase2-recenter-runbook.md b/docs/plans/2026-05-12-sp22-h6-phase2-recenter-runbook.md new file mode 100644 index 000000000..f70eb1e6a --- /dev/null +++ b/docs/plans/2026-05-12-sp22-h6-phase2-recenter-runbook.md @@ -0,0 +1,793 @@ +# SP22 H6 Phase 2 — Recenter state[121] Implementation Plan + +> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:executing-plans (or superpowers:subagent-driven-development) to implement this plan. Steps use checkbox (`- [ ]`) syntax for tracking. + +**Goal:** Recenter the SP22 H6 aux→state bridge so state[121] uses the [-1, +1] convention with sentinel 0, matching every other state slot's "no signal" = 0 baseline, fixing the `pearl_first_observation_bootstrap` violation introduced by Phase 1. + +**Architecture:** Single coherent encoding change across one GPU kernel + one host-side buffer-init path + three state-gather kernels' NULL-fallback sentinels + two doc-comments. All edits land in **one atomic commit** per `feedback_no_partial_refactor` (the encoding contract spans these sites; partial migration produces inconsistent state slot semantics across training vs eval). + +**Tech Stack:** Rust 1.85+ (Edition 2021); CUDA 12.4, sm_89 / L40S; cudarc 0.19; SQLX_OFFLINE for cargo check; argo-workflows for smoke dispatch. + +**Working tree:** `/home/jgrusewski/Work/foxhunt/.worktrees/sp19-20-wr-first/` on branch `sp20-aux-h-fixed`. All commands assume cwd is this worktree. + +**Spec:** `docs/plans/2026-05-12-sp22-h6-phase2-recenter.md` (canonical design + verdict criteria). + +--- + +## File map + +All file paths are relative to the worktree root. + +| File | Responsibility | This plan | +|---|---|---| +| `crates/ml/src/cuda_pipeline/aux_softmax_to_per_env_kernel.cu` | Per-env p_up extractor; the bridge's write site | Modify (write recentered value) | +| `crates/ml/src/cuda_pipeline/gpu_experience_collector.rs` | Per-env buffer alloc + 0.5 fill at cold-start + FoldReset re-fill | Modify (sentinel 0.5 → 0.0) | +| `crates/ml/src/cuda_pipeline/experience_kernels.cu` | 3 state-gather kernels' NULL fallback | Modify (0.5f → 0.0f in 3 places) | +| `crates/ml-core/src/state_layout.rs` | `AUX_DIR_PROB_INDEX` constant + doc | Modify (doc comment update) | +| `crates/ml/src/cuda_pipeline/state_layout.cuh` | `assemble_state` device function + padding-write inline comment | Modify (inline comment update) | +| `docs/dqn-wire-up-audit.md` | Append `## 2026-05-12 — SP22 H6 Phase 2 implementation` section (post-verification) | Modify (audit doc append) | + +--- + +## Task 1: Recenter the copy kernel + +**Files:** +- Modify: `crates/ml/src/cuda_pipeline/aux_softmax_to_per_env_kernel.cu` + +- [ ] **Step 1: Read the current file to anchor the edit** + +```bash +sed -n '60,90p' crates/ml/src/cuda_pipeline/aux_softmax_to_per_env_kernel.cu +``` + +Expected: shows the kernel body with the line +`prev_aux_dir_prob[env] = aux_softmax[(size_t)env * (size_t)K + 1];` + +- [ ] **Step 2: Replace the write expression** + +Use the Edit tool. `old_string`: + +```c + int env = blockIdx.x * blockDim.x + threadIdx.x; + if (env >= n_envs) return; + + /* Write p_up = softmax[1]. Used by next step's state assembly. */ + prev_aux_dir_prob[env] = aux_softmax[(size_t)env * (size_t)K + 1]; +``` + +`new_string`: + +```c + int env = blockIdx.x * blockDim.x + threadIdx.x; + if (env >= n_envs) return; + + /* SP22 H6 Phase 2 (2026-05-12): write RECENTERED p_up = `2*p - 1` + * so slot 121 uses the same "no signal = 0" baseline as every + * other state slot, per `pearl_first_observation_bootstrap`. + * Range: [-1, +1] (softmax components non-negative and sum to 1, + * so 2*p - 1 is structurally bounded). Sentinel 0 = neutral + * (aux contributes nothing); +1 = up with full conviction; + * -1 = down with full conviction. */ + prev_aux_dir_prob[env] = + 2.0f * aux_softmax[(size_t)env * (size_t)K + 1] - 1.0f; +``` + +- [ ] **Step 3: Update the kernel-header sentinel/cold-start comment** + +The kernel file's preamble describes sentinel 0.5; update to 0.0. Use Edit tool. `old_string`: + +```c +// Sentinel / cold-start / FoldReset +// ───────────────────────────────── +// The buffer is filled with 0.5 (neutral; p_up = 50%) at construction and at +// every FoldReset by the GPU `fill_f32` kernel (epsilon_greedy_kernel.cu). +// This kernel ALWAYS overwrites — no sentinel branch needed here. The first +// rollout step's state assembly reads 0.5 (no aux signal yet), then this +// kernel writes the real value for step 2's state, and so on. +``` + +`new_string`: + +```c +// Sentinel / cold-start / FoldReset +// ───────────────────────────────── +// SP22 H6 Phase 2 (2026-05-12): buffer is filled with 0.0 (neutral +// under the recentered `2*p - 1` encoding; aux contributes nothing) at +// construction and at every FoldReset by the GPU `fill_f32` kernel +// (epsilon_greedy_kernel.cu). Pre-Phase-2 the sentinel was 0.5, which +// violated `pearl_first_observation_bootstrap` (every other state slot +// uses 0 as the "no signal" baseline). This kernel ALWAYS overwrites +// — no sentinel branch needed here. The first rollout step's state +// assembly reads 0.0 (no aux signal yet), then this kernel writes the +// real recentered value for step 2's state, and so on. +``` + +--- + +## Task 2: Update host-side sentinels (collector cold-start + FoldReset) + +**Files:** +- Modify: `crates/ml/src/cuda_pipeline/gpu_experience_collector.rs` + +There are TWO `let sentinel: f32 = 0.5;` literals in this file (one in `new()` cold-start path, one in `reset_episodes()` FoldReset path). Both migrate atomically. + +- [ ] **Step 1: Locate both sentinel sites** + +```bash +grep -n "let sentinel: f32 = 0.5" crates/ml/src/cuda_pipeline/gpu_experience_collector.rs +``` + +Expected: exactly two matches. + +- [ ] **Step 2: Update cold-start sentinel (in `new()`)** + +Use Edit tool. `old_string`: + +```rust + { + // Cold-start fill: write 0.5 sentinel to every slot. Pure GPU + // compute via `fill_f32` (no HtoD). Matches the launch shape used + // by `GpuActionSelector::run_branching` for the epsilon buffer. + let n_envs_i32 = alloc_episodes as i32; + let sentinel: f32 = 0.5; +``` + +`new_string`: + +```rust + { + // SP22 H6 Phase 2 (2026-05-12) cold-start fill: write 0.0 + // sentinel (neutral under the recentered `2*p - 1` encoding; + // matches the "no signal = 0" baseline used by every other + // state slot per `pearl_first_observation_bootstrap`). Pure + // GPU compute via `fill_f32` (no HtoD). Matches the launch + // shape used by `GpuActionSelector::run_branching` for the + // epsilon buffer. + let n_envs_i32 = alloc_episodes as i32; + let sentinel: f32 = 0.0; +``` + +- [ ] **Step 3: Update FoldReset sentinel (in `reset_episodes()`)** + +Use Edit tool. `old_string`: + +```rust + // SP22 H6 (2026-05-12): FoldReset of the per-env aux directional + // probability cache. Re-seed every slot to the 0.5 neutral sentinel + // so the first state-gather call of the new fold reads "no aux + // signal yet" instead of the previous fold's last p_up. Pure GPU + // compute via fill_f32 — no HtoD per + // `feedback_no_htod_htoh_only_mapped_pinned.md`. + { + let n_envs_i32 = self.alloc_episodes as i32; + let sentinel: f32 = 0.5; +``` + +`new_string`: + +```rust + // SP22 H6 Phase 2 (2026-05-12): FoldReset of the per-env aux + // directional probability cache. Re-seed every slot to the 0.0 + // neutral sentinel (under the recentered `2*p - 1` encoding) so + // the first state-gather call of the new fold reads "no aux + // signal yet" with the same zero baseline used by every other + // state slot. Pure GPU compute via fill_f32 — no HtoD per + // `feedback_no_htod_htoh_only_mapped_pinned.md`. + { + let n_envs_i32 = self.alloc_episodes as i32; + let sentinel: f32 = 0.0; +``` + +- [ ] **Step 4: Update the `prev_aux_dir_prob` field-doc** + +Use Edit tool. `old_string`: + +```rust + /// SP22 H6 (2026-05-12) — per-env aux directional probability cache. + /// Sized `[alloc_episodes]` f32 (one slot per env). Written end-of-step + /// by `exp_aux_softmax_to_per_env_kernel` (gathers `aux_softmax[env, 1]` + /// = p_up). Read start-of-next-step by `experience_state_gather` for + /// `state[AUX_DIR_PROB_INDEX = SL_PADDING_START + 0 = 121]`. Cold-start + /// and FoldReset both initialize to 0.5 (neutral; p_up = 50%) via the + /// `fill_f32` GPU kernel — no HtoD per + /// `feedback_no_htod_htoh_only_mapped_pinned.md`. + prev_aux_dir_prob: cudarc::driver::CudaSlice, +``` + +`new_string`: + +```rust + /// SP22 H6 Phase 2 (2026-05-12) — per-env aux directional probability + /// cache, RECENTERED encoding. Sized `[alloc_episodes]` f32 (one slot + /// per env). Written end-of-step by `exp_aux_softmax_to_per_env_kernel` + /// (gathers `2*aux_softmax[env, 1] - 1` ∈ [-1, +1]). Read start-of-next + /// step by `experience_state_gather` for + /// `state[AUX_DIR_PROB_INDEX = SL_PADDING_START + 0 = 121]`. Cold-start + /// and FoldReset both initialize to 0.0 (neutral; aux contributes + /// nothing) via the `fill_f32` GPU kernel — no HtoD per + /// `feedback_no_htod_htoh_only_mapped_pinned.md`. Pre-Phase-2 the + /// encoding was raw `p_up ∈ [0, 1]` with sentinel 0.5; recentered for + /// gradient-init parity with every other state slot per + /// `pearl_first_observation_bootstrap`. + prev_aux_dir_prob: cudarc::driver::CudaSlice, +``` + +--- + +## Task 3: Update NULL-fallback sentinels in three state-gather kernels + +**Files:** +- Modify: `crates/ml/src/cuda_pipeline/experience_kernels.cu` + +There are THREE state-gather kernels (`experience_state_gather`, +`backtest_state_gather`, `backtest_state_gather_chunk`), each with a +`(aux_dir_prob_per_env != NULL) ? aux_dir_prob_per_env[i_or_w] : 0.5f` +defensive fallback. All migrate atomically. + +- [ ] **Step 1: Locate all three NULL-fallback sites** + +```bash +grep -n "(aux_dir_prob_per_env != NULL)" crates/ml/src/cuda_pipeline/experience_kernels.cu +``` + +Expected: exactly three matches (one per kernel). + +- [ ] **Step 2: Update training kernel (`experience_state_gather`)** + +Use Edit tool. `old_string`: + +```c + /* ── SP22 H6 aux→state bridge: previous-step p_up for this env ── + * NULL = no aux producer wired yet (cold-start, eval A3 fallback) → 0.5 + * sentinel (neutral, no signal). Lag of 1 bar is intentional — H=200 + * label horizon makes the signal slow-moving, so one-step lag is fine. */ + float aux_dir_prob = (aux_dir_prob_per_env != NULL) + ? aux_dir_prob_per_env[i] + : 0.5f; +``` + +`new_string`: + +```c + /* ── SP22 H6 aux→state bridge: previous-step recentered p_up for this env ── + * Phase 2 (2026-05-12) encoding: `2*p_up - 1 ∈ [-1, +1]`; sentinel 0 + * = neutral (matches every other state slot's "no signal = 0" + * baseline per `pearl_first_observation_bootstrap`). NULL = no aux + * producer wired yet (cold-start, eval A3 fallback) → 0.0f. Lag of + * 1 bar is intentional — H=200 label horizon makes the signal + * slow-moving, so one-step lag is fine. */ + float aux_dir_prob = (aux_dir_prob_per_env != NULL) + ? aux_dir_prob_per_env[i] + : 0.0f; +``` + +- [ ] **Step 3: Update first eval kernel (`backtest_state_gather`)** + +Use Edit tool. `old_string`: + +```c + /* SP22 H6 — per-window aux p_up (NULL → 0.5 sentinel for Phase 1 A3 eval). */ + float aux_dir_prob = (aux_dir_prob_per_env != NULL) + ? aux_dir_prob_per_env[w] + : 0.5f; +``` + +`new_string`: + +```c + /* SP22 H6 Phase 2 — per-window recentered aux p_up `2*p - 1 ∈ [-1, +1]` + * (NULL → 0.0f neutral sentinel for A3 eval; matches other state + * slots' "no signal = 0" baseline). */ + float aux_dir_prob = (aux_dir_prob_per_env != NULL) + ? aux_dir_prob_per_env[w] + : 0.0f; +``` + +- [ ] **Step 4: Update second eval kernel (`backtest_state_gather_chunk`)** + +Use Edit tool. `old_string`: + +```c + /* SP22 H6 — per-window aux p_up (NULL → 0.5 for Phase 1 A3 eval). */ + float aux_dir_prob = (aux_dir_prob_per_env != NULL) + ? aux_dir_prob_per_env[w] + : 0.5f; +``` + +`new_string`: + +```c + /* SP22 H6 Phase 2 — per-window recentered aux p_up `2*p - 1 ∈ [-1, +1]` + * (NULL → 0.0f neutral sentinel for A3 eval). */ + float aux_dir_prob = (aux_dir_prob_per_env != NULL) + ? aux_dir_prob_per_env[w] + : 0.0f; +``` + +- [ ] **Step 5: Verify all three are updated** + +```bash +grep -n "(aux_dir_prob_per_env != NULL)" crates/ml/src/cuda_pipeline/experience_kernels.cu | head +grep -nE ":\s*0\.[05]f;" crates/ml/src/cuda_pipeline/experience_kernels.cu | grep -i aux_dir_prob -B1 | head +``` + +Expected: three NULL-checks; trailing-sentinel literal is `0.0f` in all three (no remaining `0.5f`). + +--- + +## Task 4: Update `state_layout.rs` constant comment + +**Files:** +- Modify: `crates/ml-core/src/state_layout.rs` + +- [ ] **Step 1: Replace the AUX_DIR_PROB_INDEX comment** + +Use Edit tool. `old_string`: + +```rust +// SP22 H6 (2026-05-12) — aux directional probability from previous step. +// Lives in the first padding slot. PADDING_DIM stays 7 (no STATE_DIM change); +// `assemble_state` writes aux_dir_prob to [AUX_DIR_PROB_INDEX] and zeros the +// remaining 6 padding slots [PADDING_START+1..STATE_DIM). +// Sentinel = 0.5 (neutral; p_up = 50%) at cold-start and FoldReset. +pub const AUX_DIR_PROB_INDEX: usize = PADDING_START; // 121 +``` + +`new_string`: + +```rust +// SP22 H6 Phase 2 (2026-05-12) — aux directional probability from previous +// step, RECENTERED. Lives in the first padding slot. PADDING_DIM stays 7 +// (no STATE_DIM change); `assemble_state` writes aux_dir_prob to +// [AUX_DIR_PROB_INDEX] and zeros the remaining 6 padding slots +// [PADDING_START+1..STATE_DIM). +// Encoding: `2*p_up - 1` ∈ [-1, +1] (structurally bounded — softmax +// components in [0, 1] sum to 1). +1 = up with full conviction; +// -1 = down with full conviction; 0 = neutral. +// Sentinel = 0.0 at cold-start and FoldReset (matches every other state +// slot's "no signal = 0" baseline per `pearl_first_observation_bootstrap`). +pub const AUX_DIR_PROB_INDEX: usize = PADDING_START; // 121 +``` + +--- + +## Task 5: Update `state_layout.cuh` device-function inline comment + +**Files:** +- Modify: `crates/ml/src/cuda_pipeline/state_layout.cuh` + +- [ ] **Step 1: Replace the padding-block inline comment** + +Use Edit tool. `old_string`: + +```c + // Padding [121..128) — SP22 H6 (2026-05-12): slot 0 = aux directional + // probability (p_up from previous step's aux softmax); slots 1..7 zero + // for 8-alignment. Caller passes 0.5 sentinel at cold-start / FoldReset + // and for eval-time NULL fallback (A3). + out[SL_PADDING_START + 0] = aux_dir_prob; + for (int k = 1; k < SL_PADDING_DIM; k++) + out[SL_PADDING_START + k] = 0.0f; +``` + +`new_string`: + +```c + // Padding [121..128) — SP22 H6 Phase 2 (2026-05-12): slot 0 = aux + // directional probability from previous step's aux softmax, RECENTERED + // to `2*p_up - 1 ∈ [-1, +1]` (matches every other state slot's + // "no signal = 0" baseline per `pearl_first_observation_bootstrap`). + // Slots 1..7 zero for 8-alignment. Caller passes 0.0 sentinel at + // cold-start / FoldReset and for eval-time NULL fallback (A3). + out[SL_PADDING_START + 0] = aux_dir_prob; + for (int k = 1; k < SL_PADDING_DIM; k++) + out[SL_PADDING_START + k] = 0.0f; +``` + +--- + +## Task 6: Verification gate 1 — `cargo check` + +- [ ] **Step 1: Run `cargo check -p ml --features cuda`** + +```bash +SQLX_OFFLINE=true cargo check -p ml --features cuda 2>&1 | tail -20 +``` + +Expected: 0 errors. 21 pre-existing warnings (unchanged from Phase 1 +baseline; `-W unsafe-code` on existing blocks + unused `DevicePtrMut` +imports + style nits). The new kernel cubin will recompile because +`aux_softmax_to_per_env_kernel.cu` changed; verify the "Finished +`dev` profile" line at the end. + +- [ ] **Step 2: If errors, fix and re-run; if warnings count >21, investigate** + +Any compile error must be resolved before proceeding. New warnings (beyond +the pre-existing 21) indicate a typo or stray reference; fix and re-run. + +--- + +## Task 7: Verification gate 2 — `gpu_backtest_validation` tests + +- [ ] **Step 1: Run the validation tests in release mode** + +```bash +SQLX_OFFLINE=true cargo test -p ml --test gpu_backtest_validation --features cuda --release 2>&1 | tail -20 +``` + +Expected: `4 passed; 2 failed`. The 2 PnL-assertion failures +(`test_always_long_on_downtrend`, `test_multiple_windows_produce_results`) +are pre-existing per the Phase 1 baseline; they MUST continue to fail +identically (not new failures from this change). + +- [ ] **Step 2: If a previously-passing test newly fails, STOP and investigate** + +Any of `test_active_model_records_trades`, `test_always_flat_produces_no_pnl`, +`test_always_long_on_uptrend`, `test_extended_metrics_populated` newly +failing means the change broke an integration assertion. Do not commit. + +--- + +## Task 8: Verification gate 3 — `compute-sanitizer` memcheck + +- [ ] **Step 1: Locate the freshly-built test binary** + +```bash +TEST_BIN=$(ls -t target/release/deps/gpu_backtest_validation-* 2>/dev/null | grep -v '\.d$' | head -1) && echo "binary: $TEST_BIN" +``` + +Expected: prints a binary path under `target/release/deps/`. + +- [ ] **Step 2: Run compute-sanitizer** + +```bash +/usr/local/cuda/bin/compute-sanitizer --tool=memcheck --error-exitcode=42 "$TEST_BIN" --test-threads=1 2>&1 | tail -10 +``` + +Expected: `ERROR SUMMARY: 0 errors` at the end. The wrapped Rust test +binary will still exit non-zero because of the 2 PnL-assertion failures +(that's the Rust harness, not compute-sanitizer). The exit-code-42 +trigger is for any non-zero CUDA error, which should NOT fire. + +- [ ] **Step 3: If non-zero CUDA errors, STOP and investigate** + +Any compute-sanitizer error blocks the commit. Investigate the cause +(OOB, race, uninit read) before proceeding. + +--- + +## Task 9: Append audit-doc entry + +**Files:** +- Modify: `docs/dqn-wire-up-audit.md` + +Required by Invariant 7 pre-commit hook (component changes require an +audit-doc update). The append goes at the **end** of the existing H6 +section, NOT a new top-level `## 2026-05-XX` header — Phase 2 is a +continuation of the H6 narrative, not a separate experiment. + +- [ ] **Step 1: Read the audit-doc tail to find the append point** + +```bash +tail -5 docs/dqn-wire-up-audit.md +``` + +Expected: the last few lines of the "Next-investigation framing" section +from the H6 falsification verdict (ending with "If H3 doesn't move WR +either, V/A goes first and we re-test in combination."). + +- [ ] **Step 2: Append the Phase 2 entry** + +Use Edit tool. `old_string` (the last paragraph of the existing audit +section — provides anchor for the append): + +``` +If H3 lands and the WR moves but action-space stays narrow, V/A is +binding and gets fixed next. If H3 doesn't move WR either, V/A goes +first and we re-test in combination. +``` + +`new_string`: + +``` +If H3 lands and the WR moves but action-space stays narrow, V/A is +binding and gets fixed next. If H3 doesn't move WR either, V/A goes +first and we re-test in combination. + +### Phase 2 (2026-05-12) — recenter state[121] to [-1, +1] + +Post-mortem of the Phase 1 smoke verdict surfaced an actual pearl +violation in the H6 implementation itself, BEFORE pivoting to H3. + +Per `pearl_first_observation_bootstrap`: "sentinel = 0; first +observation replaces directly." The Phase 1 design wrote +`state[121] = aux_softmax[env, 1] = p_up ∈ [0, 1]` with sentinel 0.5. +**Every other state slot uses 0 as its "no signal" baseline** +(zero-padding, feature_mask, ofi-missing, mtf-missing). Slot 121 alone +was off-pattern, forcing the encoder to learn TWO things — directional +mapping AND the non-zero bias offset — instead of one. + +Phase 2 fix (atomic commit, 5 source files): + +- `aux_softmax_to_per_env_kernel.cu` writes `2*p_up - 1 ∈ [-1, +1]` + (still structurally bounded since softmax components sum to 1) +- `gpu_experience_collector.rs` cold-start + FoldReset sentinel + `0.5 → 0.0` +- `experience_kernels.cu` NULL-fallback in three state-gather kernels + `0.5f → 0.0f` +- `state_layout.rs` + `state_layout.cuh` comment updates documenting + the recentered encoding and zero sentinel + +Verification gates (all clean before smoke): + +| Gate | Result | +|---|---| +| `cargo check -p ml --features cuda` | TBD-fill-in-when-task-runs | +| `gpu_backtest_validation` | TBD-fill-in-when-task-runs | +| `compute-sanitizer --tool=memcheck` | TBD-fill-in-when-task-runs | + +Smoke dispatch: TBD-workflow-name. Verdict criteria (per spec): + +- **WR > 50.5%** within 3 epochs → Phase 1 + Phase 2 sufficient; + recentering was the binding constraint; justify A2 (eval-side aux + integration) for production parity. +- **`a_var` for mag/ord/urg moves off 0** (> 1e-3) → secondary + success signal that sub-branches are gradient-coupled under + the recentered signal. +- **WR pinned at 50.1–50.2%** → Phase 2 falsified. The encoder + successfully consumes a recentered signal but still can't extract + directional alpha from one state slot in 3 epochs. Pivot to + amplitude scaling (multiply state[121] write by a scalar > 1) or + deeper hypothesis. +``` + +The audit doc tail explicitly contains TBD placeholders for the +verification results and workflow name — they will be filled in after +tasks 6-8 and the smoke dispatch (Task 12). + +- [ ] **Step 3: Fill in verification results NOW (before commit)** + +After Tasks 6-8 ran clean, edit the three TBD lines to capture the +actual results (0 errors, 4/4 expected-passing, ERROR SUMMARY: 0 errors). +This keeps the audit doc accurate at commit time. The smoke workflow +name remains TBD until Task 12. + +--- + +## Task 10: Atomic commit + push + +- [ ] **Step 1: Stage the 6 files** + +```bash +git add \ + crates/ml-core/src/state_layout.rs \ + crates/ml/src/cuda_pipeline/aux_softmax_to_per_env_kernel.cu \ + crates/ml/src/cuda_pipeline/experience_kernels.cu \ + crates/ml/src/cuda_pipeline/gpu_experience_collector.rs \ + crates/ml/src/cuda_pipeline/state_layout.cuh \ + docs/dqn-wire-up-audit.md && \ +git status -s +``` + +Expected: 6 `M ` (staged-modified) lines, nothing else. + +- [ ] **Step 2: Commit with descriptive message** + +```bash +git commit -m "$(cat <<'EOF' +feat(sp22): H6 Phase 2 — recenter state[121] to [-1, +1] (atomic) + +Phase 1 post-mortem traced an actual `pearl_first_observation_bootstrap` +violation in my own H6 implementation: state slot 121 wrote +`aux_softmax[env, 1] = p_up ∈ [0, 1]` with sentinel 0.5, but every +OTHER state slot uses 0 as the "no signal" baseline (zero-padding, +feature_mask, ofi-missing, mtf-missing). The encoder had to learn TWO +things about slot 121 (directional mapping + non-zero bias offset) +instead of one. Phase 2 fixes the encoding to match the project +convention BEFORE declaring H6 fully falsified. + +Mechanism change +──────────────── +- `aux_softmax_to_per_env_kernel.cu` writes `2*p_up - 1 ∈ [-1, +1]` + instead of `p_up`. Still structurally bounded (softmax components + in [0, 1] sum to 1). +- Cold-start + FoldReset sentinel: 0.5 → 0.0 via the same pure-GPU + `fill_f32` path. No HtoD per + `feedback_no_htod_htoh_only_mapped_pinned`. +- NULL-fallback in 3 state-gather kernels (training + + backtest-per-step + backtest-chunk): 0.5f → 0.0f. +- Constant + device-function comment updates to document the + recentered encoding. + +Atomic per `feedback_no_partial_refactor`: the encoding contract +spans 5 source files; partial migration produces inconsistent slot +semantics between training and eval. + +Verification gates (all clean) +────────────────────────────── +- cargo check -p ml --features cuda: 0 errors +- gpu_backtest_validation: 4/4 expected-passing tests still pass +- compute-sanitizer --tool=memcheck: ERROR SUMMARY: 0 errors + +Verdict criteria (per spec) +─────────────────────────── +- WR > 50.5% within 3 epochs → recentering binding, H6 + Phase 2 + sufficient → justify A2. +- a_var for mag/ord/urg > 1e-3 → sub-branches gradient-coupled under + recentered signal. +- WR pinned at 50.1–50.2% → Phase 2 falsified, pivot to amplitude + scaling or deeper hypothesis. + +Refs +──── +- docs/plans/2026-05-12-sp22-h6-phase2-recenter.md (spec) +- docs/plans/2026-05-12-sp22-h6-phase2-recenter-runbook.md (this plan) +- pearl_first_observation_bootstrap (sentinel = 0) +- feedback_no_partial_refactor (5-file atomic) +- feedback_no_htod_htoh_only_mapped_pinned (fill_f32, not HtoD) + +Co-Authored-By: Claude Opus 4.7 (1M context) +EOF +)" 2>&1 | tail -10 +``` + +Expected: pre-commit hook prints "✅ All pre-commit checks passed!" then +the commit summary with `[sp20-aux-h-fixed ]` and `6 files changed`. + +- [ ] **Step 3: Push** + +```bash +git push 2>&1 | tail -5 +``` + +Expected: `.. sp20-aux-h-fixed -> sp20-aux-h-fixed`. +Per `feedback_push_before_deploy`, push must succeed before smoke dispatch. + +--- + +## Task 11: Dispatch smoke + +- [ ] **Step 1: Dispatch the 3-epoch baseline on L40S** + +```bash +./scripts/argo-train.sh dqn --baseline --branch sp20-aux-h-fixed \ + --epochs 3 --gpu-pool ci-training-l40s 2>&1 | tail -30 +``` + +Expected: workflow submitted; `Name: train-` printed; status +Pending → will move to Running. + +- [ ] **Step 2: Capture the workflow name** + +Note the `train-` from the dispatch output. It will be used for +monitoring (Task 12) and for filling in the audit-doc workflow name. + +- [ ] **Step 3: Update audit-doc workflow name** + +Edit the audit-doc Phase 2 entry to replace "TBD-workflow-name" with +the actual `train-` from Step 2. Stage + amend the previous +commit? NO — per project discipline (`Always create NEW commits rather +than amending`), make a small follow-up doc commit: + +```bash +git add docs/dqn-wire-up-audit.md && \ +git commit -m "$(cat <<'EOF' +docs(sp22): record H6 Phase 2 smoke workflow name + +Workflow `train-` dispatched on sp20-aux-h-fixed @ +for the Phase 2 verdict. Audit-doc TBD placeholder replaced with +the actual workflow handle for traceability. + +Co-Authored-By: Claude Opus 4.7 (1M context) +EOF +)" 2>&1 | tail -5 && git push 2>&1 | tail -3 +``` + +(Skip push if the workflow result will land before the next git push +naturally — this is just for cross-referencing.) + +--- + +## Task 12: Monitor for verdict + +- [ ] **Step 1: Start a background log tail** + +```bash +argo logs -f train- -n foxhunt --no-color 2>&1 | tee /tmp/train-.log +``` + +Use `run_in_background: true` so subsequent BashOutput reads can sample +without blocking. (Per `feedback_no_redundant_monitor`: pick ONE +channel; no Monitor layered on top of bg tail.) + +- [ ] **Step 2: Wait for cycle 1 + HEALTH_DIAG** + +Run a single bg waiter that fires on either a metric signal OR an +anomaly: + +```bash +timeout 5400 bash -c ' + while true; do + if grep -qE "wr_ema|win_rate|profit_factor|dir_acc_short|dir_acc_long|cycle_summary|fold_summary|val_score|val_sharpe|epoch_summary|EPOCH .* COMPLETE|FOLD .* COMPLETE|CYCLE .* (COMPLETE|END)|sharpe_ratio" /tmp/train-.log 2>/dev/null \ + || grep -qE "panicked at|Traceback|thread .* panicked|illegal memory access|CUDA error|cudaError|out of memory|OutOfMemory|Killed|Workflow .* Failed|NaN detected|nan detected|gradient.*nan|Diverged" /tmp/train-.log 2>/dev/null; then + echo "=== SIGNAL CAUGHT ===" + grep -nE "wr_ema|win_rate|profit_factor|dir_acc_short|dir_acc_long|cycle_summary|fold_summary|val_score|val_sharpe|epoch_summary|EPOCH .* COMPLETE|FOLD .* COMPLETE|CYCLE .* (COMPLETE|END)|sharpe_ratio|panicked at|Traceback|thread .* panicked|illegal memory access|CUDA error|cudaError|out of memory|OutOfMemory|Killed|Workflow .* Failed|NaN detected|nan detected|gradient.*nan|Diverged" /tmp/train-.log | tail -40 + echo "=== END SIGNAL ===" + exit 0 + fi + sleep 30 + done +' +``` + +Replace `` with the actual workflow ID from Task 11. + +- [ ] **Step 3: Interpret the verdict** + +When the waiter fires, sample the latest HEALTH_DIAG block and the +epoch-1 trade-stats line: + +```bash +grep -nE "HEALTH_DIAG\[" /tmp/train-.log | tail -40 +grep -nE "GPU trades=" /tmp/train-.log | tail -5 +grep -nE "a_var|v_share|v_a_means" /tmp/train-.log | tail -10 +``` + +Compute WR = wins / (wins + losses) from the epoch-1 trade-stats line. + +**Decision**: + +- WR > 50.5% AND `a_var [m, o, u]` > 1e-3 → Phase 2 confirmed. Let + the workflow finish for the audit record. Update audit doc with + the verdict details. Justify A2 next. +- WR > 50.5% AND `a_var [m, o, u]` ≈ 0 → Phase 2 partial. Recentering + fixed encoder pickup of state[121] but sub-branches still starve. + Pivot to per-branch tau / atom-span audit. +- WR pinned at 50.1–50.2% → Phase 2 falsified. Terminate the workflow + (`argo terminate train- -n foxhunt`); audit doc verdict; pivot + to amplitude scaling or deeper hypothesis. + +- [ ] **Step 4: Update audit doc with verdict + close the H6 narrative** + +After the verdict is in, replace the TBD-fill-in lines in the audit +doc (cargo check / tests / compute-sanitizer) AND add a "Smoke verdict" +subsection capturing: + +- Final wins/losses/PF/WR% +- `a_var` per-branch values +- Whether Phase 2 confirmed / partially confirmed / falsified +- Next-investigation framing + +Commit + push. + +--- + +## Self-review checklist + +After execution: verify each spec requirement maps to a task in this +plan. + +| Spec requirement | Task | +|---|---| +| Recenter copy kernel: `2*p - 1` | Task 1 | +| Update kernel-file header sentinel comment | Task 1 Step 3 | +| Cold-start sentinel 0.5 → 0.0 in `new()` | Task 2 Step 2 | +| FoldReset sentinel 0.5 → 0.0 in `reset_episodes()` | Task 2 Step 3 | +| `prev_aux_dir_prob` field-doc update | Task 2 Step 4 | +| NULL-fallback 0.5f → 0.0f in `experience_state_gather` | Task 3 Step 2 | +| NULL-fallback 0.5f → 0.0f in `backtest_state_gather` | Task 3 Step 3 | +| NULL-fallback 0.5f → 0.0f in `backtest_state_gather_chunk` | Task 3 Step 4 | +| `AUX_DIR_PROB_INDEX` doc comment update | Task 4 | +| `assemble_state` padding-block inline comment update | Task 5 | +| Audit-doc append (Invariant 7) | Task 9 | +| Verification: cargo check | Task 6 | +| Verification: gpu_backtest_validation | Task 7 | +| Verification: compute-sanitizer | Task 8 | +| Atomic commit + push (per `feedback_push_before_deploy`) | Task 10 | +| Smoke dispatch | Task 11 | +| Verdict monitoring + audit doc verdict | Task 12 | +| **Out of scope (NOT in plan)**: reward shaping, V/A projection, per-branch tau, A2 | Verified absent | + +Every spec change has exactly one task. No orphaned tasks. No +placeholders in commit messages or code blocks (the audit-doc TBD +literals are intentional, filled in inside the same plan).