feat(ml-alpha): deterministic same-seed training + Tier 1.5 fast-dev-cycle

Two same-seed runs now produce bit-equal eval_summary.json, alpha_rl_train_summary.json,
and diag.jsonl (modulo wall-clock elapsed_s). The 5-phase falsification chain landed:

  Phase 2   PER tree-rebuild: __threadfence is NOT a grid-wide barrier; multiple blocks
            raced across sum-tree levels. Fix: Grid=(1) Block=(1024) + __syncthreads
            in rl_per_tree_rebuild.cu.

  Phase 2.3 cuBLAS GEMM_DFALT + TF32 default-math allowed split-K non-deterministic
            accumulation at 3 sites. New crates/ml-alpha/src/cublas_determinism.rs
            applies CUBLAS_PEDANTIC_MATH via FOXHUNT_DETERMINISTIC env toggle
            (0=TF32 prod, 1=PEDANTIC dev default, 2=DEFAULT_MATH control).

  Phase 2.6 Two bugs surfaced sequentially in the backward kernel chain:
            (1) rl_iqn_tau_cos_features had a multi-block r/w race on prng_state[batch]
                — all N_TAU=32 blocks read seed; only tau_idx==0 wrote back; no
                inter-block barrier. Fix: split into READ-ONLY rl_iqn_tau_cos_features
                + new sibling rl_iqn_advance_prng_state launched on same stream
                (kernel-launch ordering = grid-wide barrier).
            (2) OutcomeHead::new called near_zero_xavier without scoped_init_seed,
                falling back to time+thread-id RNG. Stayed dormant until first done
                event activated non-sentinel labels and divergent weights flowed via
                grad_h_t_outcome into encoder gradient. Fix: add seed param + install
                scoped_init_seed(dqn_seed.wrapping_add(0x0CE0)) guard.

Validation (./scripts/determinism-check.sh --quick, RTX 3050, b=128, 200+50 steps):
  - All 200 rows of checksums.* leaves match (rel-tol 1e-5, abs-tol 1e-7)
  - eval_summary.json, alpha_rl_train_summary.json byte-equal between runs
  - diag.jsonl byte-equal modulo elapsed_s
  - Eval pnl identical run-A vs run-B at seed 42

Pre-fix baseline (Phase 2.5 measurement): same-seed eval pnl spread $450k
($187k vs -$261k). Post-fix: $0 spread.

Speed cost: ~1.5ms/step amortised; ~10-15% slower than TF32 production
(PEDANTIC tax — acceptable in dev, toggle to FOXHUNT_DETERMINISTIC=0 for prod).

Mapped-pinned discipline: all 11 NEW memcpy_dtoh sites in diagnostic dump methods
+ per-step checksum readback use a new pub(crate) helper
read_slice_d_into<T: Copy>(stream, src, dst) — MappedRecordBuffer + raw
memcpy_dtod_async + raw_stream_sync + volatile read. Generic over T (f32, f64,
i32, u32, u8). Satisfies feedback_no_htod_htoh_only_mapped_pinned + hook guard.

Bundled Tier 1.5 fast-dev-cycle infrastructure (spec
docs/superpowers/specs/2026-06-02-fast-dev-cycle.md):
  - scripts/local-mid-smoke.sh        b=128, 2000+500, ~10min on RTX 3050
  - scripts/determinism-check.sh      runs mid-smoke twice, diffs checksums
  - scripts/tier1_5_verdict.py        behavioral kill verdict
  - AdamW checkpoint save/load (crates/ml-alpha/src/trainer/optim.rs)
  - IntegratedTrainer checkpoint save/load (resume from checkpoint)
  - 15 Phase 1 checksum leaves in build_diag_value
  - Env-gated dump methods (FOXHUNT_DETERMINISM_DEBUG_PER/MAMBA2/RL/BACKWARD)
    for future divergence-chasing — never run in production

Documentation:
  - docs/superpowers/specs/2026-06-02-determinism-foundation.md
  - docs/superpowers/specs/2026-06-02-fast-dev-cycle.md
  - docs/superpowers/plans/2026-06-02-determinism-foundation-implementation.md
  - docs/superpowers/notes/2026-06-02-determinism-phase{1,2,2.2,2.5,2.6}-*.md
  - Adjacent specs/plans/notes from the analytical chain that surfaced determinism
    as the load-bearing blocker (eval-summary, eval-boundary, regime-observer,
    multi-head policy, regime-invariance, Phase 3 IQN-complement post-mortem)

Unlocks: every controller / architecture / reward-shaping A/B from this commit
onward attributes outcome differences to the change, not random-init kernel-race
drift cascading through training x eval LOB-sim trajectories. The eval-collapse
investigation (pearl_reward_signal_anti_aligned_with_pnl, multi-head spec,
regime-invariance spec) is now testable with trustworthy verdicts.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
This commit is contained in:
jgrusewski
2026-06-02 17:56:00 +02:00
parent 63fc16f173
commit 629ebd667c
45 changed files with 12531 additions and 51 deletions

3
.gitignore vendored
View File

@@ -7,6 +7,9 @@ test_data/feature-cache/
/data/feature-cache/
*.fxcache
# Tier 1.5 mid-smoke local data (symlinks to test_data/futures-baseline-mbp10/)
test_data/futures-baseline-mid/
# IDE files
.vscode/
.idea/

View File

@@ -66,6 +66,24 @@ cargo bench --bench database_performance
- `SQLX_OFFLINE=true` — compiled SQL queries (no live DB needed for build)
- `SCCACHE_BUCKET=foxhunt-sccache` — shared compile cache
### Tiered local validation (fast dev cycle)
Three-tier funnel before cluster submit (per `docs/superpowers/specs/2026-06-02-fast-dev-cycle.md`):
| Tier | Setup | Time | What it validates |
|---|---|---|---|
| 1 (correctness) | RTX 3050, b=16, 200 steps, 1 file | ~6 sec | Kernel correctness, ISV slots, NaN |
| 1.5 (behavior) | RTX 3050, b=128, 2000+500 steps, 2 files | ~10 min | entropy, hold growth, controller stability, early Pearson |
| 2 (eval verdict) | L40S, b=1024, 20k+5k, 9 files | ~70 min | eval pnl, eval wr, regime stratification |
**Tier 1.5 data path**: `test_data/futures-baseline-mid/` (symlinks to existing MBP-10 files for fold-1 train/eval split).
```bash
./scripts/local-mid-smoke.sh # Run Tier 1.5 mid-smoke (~10 min)
./scripts/determinism-check.sh # Reproducibility self-test (runs mid-smoke twice, diffs final 5 rows)
python3 scripts/tier1_5_verdict.py /tmp/foxhunt-mid-smoke # Behavioral kill verdict
```
## Deployment (Argo Workflows)
```bash

View File

@@ -152,6 +152,7 @@ const KERNELS: &[&str] = &[
"rl_outcome_bwd", // Outcome aux: single linear layer backward — dW (per-batch), db (per-batch), dh_t (per-batch); 1 block per batch, 128 threads
"snapshot_aos_to_soa", // AoS→SoA scatter: one thread per snapshot reads contiguous Mbp10RawInput, writes into 10 SoA device buffers; replaces host nested loops + 10 DtoD copies
"gpu_sample_and_gather", // GPU-resident batch sampler: random file+anchor sampling + AoS→SoA gather from pre-uploaded dataset; eliminates ALL per-step CPU data loading
"rl_deterministic_checksum", // Determinism foundation Phase 1 (2026-06-02): provably deterministic sum-of-squares (single-block / single-thread / f64 accumulation) for per-step component checksums in diag; localizes non-determinism source. Spec: docs/superpowers/specs/2026-06-02-determinism-foundation.md §1.1
];
// Cache bust v31 — five new reduce / derive kernels populate the input

View File

@@ -0,0 +1,79 @@
// rl_deterministic_checksum.cu — provably deterministic sum-of-squares.
//
// Phase 1 of the determinism foundation
// (`docs/superpowers/specs/2026-06-02-determinism-foundation.md` §1.1).
//
// Computes `Σ data[i]^2` over `n` elements with construction-guaranteed
// determinism. Launch geometry is hard-coded (1,1,1)/(1,1,1) — a single
// thread iterates sequentially. The accumulator is double precision; no
// intermediate fp32 round-off varies with launch geometry or block count.
//
// Trade-off vs throughput:
// * RTX 3050 / L40S sum-of-squares throughput at single-thread is on the
// order of 10^9 elements/s. The largest tensor in foxhunt's RL train
// loop is ~24k floats (Q-head grad_w). Cost per launch is therefore
// dominated by launch overhead (~5-20 μs) — 15 launches/step add
// ~150-300 μs total, well within the dev-mode budget per §4 of the
// spec.
// * The deliberately slow accumulation order is the entire point. Any
// parallel reduction tree introduces order-of-addition variance that
// would invalidate the diagnostic.
//
// Per `feedback_no_atomicadd`: zero atomics — pure sequential accumulation.
// Per `feedback_no_stubs`: kernel always executes, no early-exit; n=0
// writes 0.0 (defined behavior).
// Per `feedback_no_nvrtc`: registered in `crates/ml-alpha/build.rs` for
// AOT compilation to per-arch cubin.
//
// Output is f64 (one element). Reading it from the host is the only CPU
// roundtrip; the value itself is computed entirely on device, satisfying
// `feedback_cpu_is_read_only`.
#include <cuda_runtime.h>
#include <stdint.h>
// `n` is passed as int (32-bit) to match the kernel arg-passing
// convention used throughout foxhunt (raw_launch + RawArgs::push_i32).
// All checksum-able tensors in the RL trainer are < 2 billion elements
// (the largest, Q-head weights with HIDDEN_DIM * N_ACTIONS * Q_N_ATOMS,
// is on the order of 1e5 floats), so int is safe.
extern "C" __global__ void rl_deterministic_checksum_f32(
const float* __restrict__ data,
int n,
double* __restrict__ out
) {
// Single block, single thread — only thread 0 does anything.
if (threadIdx.x != 0 || blockIdx.x != 0) {
return;
}
double acc = 0.0;
for (int i = 0; i < n; ++i) {
const double v = static_cast<double>(data[i]);
acc += v * v;
}
out[0] = acc;
}
// Integer variant for replay indices and other u32/i32 tensors. Same
// accumulation discipline; the input is cast through int64 to avoid
// signed-overflow UB on the sum-of-squares accumulation.
extern "C" __global__ void rl_deterministic_checksum_i32(
const int* __restrict__ data,
int n,
double* __restrict__ out
) {
if (threadIdx.x != 0 || blockIdx.x != 0) {
return;
}
double acc = 0.0;
for (int i = 0; i < n; ++i) {
// Cast to int64 first to widen — int32^2 fits in int64 exactly;
// double's 53-bit mantissa may round large sums but the
// accumulation ORDER is still deterministic, which is the
// contract.
const int64_t v = static_cast<int64_t>(data[i]);
acc += static_cast<double>(v) * static_cast<double>(v);
}
out[0] = acc;
}

View File

@@ -60,8 +60,34 @@ __device__ static uint32_t xorshift32_iqn(uint32_t* state) {
// Grid = (B, N_TAU, 1)
// Block = (EMBED_DIM, 1, 1) — one thread per embedding dimension.
//
// DETERMINISM (Phase 2.6, 2026-06-02): `prng_state` is now READ-ONLY
// in this kernel. The previous version had a read-write race: all
// N_TAU blocks per batch read `prng_state[batch]` at the head of the
// kernel, and the (tau_idx == 0) block wrote back the advanced state
// at the end — with NO inter-block barrier. Blocks with tau_idx > 0
// running on a different SM could read the post-write value (if the
// tau_idx == 0 block finished first) OR the pre-write value
// (otherwise), depending on SM scheduling. Different runs picked
// different orderings, producing run-dependent τ values for
// tau_idx > 0 and hence run-dependent `iqn_q_values` even at step 0
// (no upstream RL-loop divergence required). Diagnosed via Phase 2.6
// `dump_backward_state_for_debug` Group H verdict (iqn_q_values
// DIVERGE at step 0 idx 55, max|Δ|=4.5e-5) while Groups G/I/J stayed
// EQUAL through steps 0/1 — see
// `docs/superpowers/notes/2026-06-02-determinism-phase2.6-backward
// -kernel-fix.md`.
//
// Fix: kernel now READS prng_state[batch] only (no writes); the
// state advancement is performed by `rl_iqn_advance_prng_state`
// in a separate single-thread-per-batch launch that runs AFTER
// `rl_iqn_tau_cos_features` finishes via stream ordering — no race
// possible. Per `feedback_no_atomicadd.md` + canonical Phase-2
// PER-rebuild rule: "fixed accumulation order requires a real
// barrier between stages, and kernel-launch ordering on a single
// stream is the cleanest grid-wide barrier we have".
//
// Inputs:
// prng_state [B] — per-batch xorshift32 seed (mutated)
// prng_state [B] — per-batch xorshift32 seed (READ-ONLY)
// B — batch size
// N_TAU — number of quantile samples
// Outputs:
@@ -69,9 +95,9 @@ __device__ static uint32_t xorshift32_iqn(uint32_t* state) {
// cos_features [B*N_TAU, EMBED_DIM] — cos((j+1) * pi * tau)
// ─────────────────────────────────────────────────────────────────────
extern "C" __global__ void rl_iqn_tau_cos_features(
uint32_t* __restrict__ prng_state, // [B]
int B,
int N_TAU,
const uint32_t* __restrict__ prng_state, // [B] READ-ONLY (advanced by sibling kernel)
int B,
int N_TAU,
float* __restrict__ tau, // [B, N_TAU]
float* __restrict__ cos_features // [B*N_TAU, EMBED_DIM]
) {
@@ -98,13 +124,8 @@ extern "C" __global__ void rl_iqn_tau_cos_features(
tau[batch * N_TAU + tau_idx] = u;
s_tau_val = u;
if (tau_idx == 0) {
uint32_t adv = seed;
#pragma unroll
for (int w = 0; w < 8; ++w) xorshift32_iqn(&adv);
prng_state[batch] = adv;
}
// prng_state advancement moved to `rl_iqn_advance_prng_state`
// (see kernel below); no writes to prng_state in this kernel.
}
__syncthreads();
@@ -115,6 +136,51 @@ extern "C" __global__ void rl_iqn_tau_cos_features(
cos_features[row * EMBED_DIM + j] = cosf((float)(j + 1) * PI_F * tau_val);
}
// ─────────────────────────────────────────────────────────────────────
// rl_iqn_advance_prng_state:
// Companion to `rl_iqn_tau_cos_features` — advances the per-batch
// xorshift32 PRNG state by exactly 8 steps. Launched AFTER
// `rl_iqn_tau_cos_features` so the read-only sampling sees a stable
// value for prng_state[batch] across all (tau_idx) blocks, then
// this kernel applies a single deterministic update per batch.
//
// Determinism rationale: by separating "read state for tau sampling"
// from "advance state for next call", we eliminate the read-write
// race the previous monolithic kernel had. Both kernels run on the
// same stream, so the launch ordering is a grid-wide barrier
// (no __threadfence required).
//
// Grid = (ceil(B/256), 1, 1)
// Block = (256, 1, 1) — one thread per batch element.
//
// Inputs:
// B — batch size
// Outputs:
// prng_state [B] — advanced 8 xorshift32 steps in place.
// `0` seed bootstrap matches the
// `rl_iqn_tau_cos_features` convention
// `(batch + 1) * 2654435761u + 0xBEEFu`
// so the first advance from the
// cold-start sentinel is identical to
// what the prior monolithic kernel
// produced for tau_idx==0.
// ─────────────────────────────────────────────────────────────────────
extern "C" __global__ void rl_iqn_advance_prng_state(
uint32_t* __restrict__ prng_state, // [B]
int B
) {
const int batch = blockIdx.x * blockDim.x + threadIdx.x;
if (batch >= B) return;
uint32_t seed = prng_state[batch];
if (seed == 0u) seed = (uint32_t)(batch + 1) * 2654435761u + 0xBEEFu;
uint32_t adv = seed;
#pragma unroll
for (int w = 0; w < 8; ++w) xorshift32_iqn(&adv);
prng_state[batch] = adv;
}
// ─────────────────────────────────────────────────────────────────────
// rl_iqn_relu_hadamard:
// Stage 3: bias-add → ReLU → element-wise product with h_t.

View File

@@ -1,14 +1,45 @@
/* =====================================================================
* rl_per_tree_rebuild.cu — GPU-resident PER: bottom-up parallel sum-tree rebuild
* rl_per_tree_rebuild.cu — GPU-resident PER: bottom-up sum-tree rebuild
*
* Grid=(128), Block=(256). Total 32768 threads = capacity.
* Grid=(1), Block=(1024). Single-block deterministic rebuild.
*
* Rebuilds the entire sum-tree from leaf priorities (at indices
* [capacity..2*capacity)) up to the root (index 1). Each level is
* processed in parallel with __threadfence() device-wide barriers
* between levels.
* [capacity..2*capacity)) up to the root (index 1). All levels are
* processed within ONE block: __syncthreads() between levels is a
* proper barrier guaranteeing every thread in the block sees every
* other thread's writes from the previous level.
*
* Determinism rationale (Phase 2 §2.F fix — Option C):
* The previous Grid=(128) launch used __threadfence() between levels.
* __threadfence() is a memory-ordering primitive, NOT a grid-wide
* barrier — it orders THIS thread's writes globally but does not
* wait for OTHER blocks' writes to be visible. Under the previous
* geometry, block N could read level-L nodes while block M's
* level-L writes were still in flight, producing different
* addition orderings across same-seed runs. Same-seed runs of the
* determinism diagnostic at b=128 saw root priority diverge
* ~592.7 vs ~695.7 at step 2 (Phase 2 sub-investigation dumps).
*
* By collapsing to a single block, we lose grid-level parallelism
* but gain bit-deterministic execution: __syncthreads() is a hard
* intra-block barrier, and every internal node is the deterministic
* sum of its two specific children written in a definite order
* (sequential grid-stride within one block, fixed thread→node
* mapping via i = tid + k*blockDim.x).
*
* Speed cost: capacity=32768 → ~65k floating-point adds total, 15
* levels = 15 syncs. On RTX 3050 the kernel runs ~30-80 μs which
* is a ~10× slowdown vs the parallel version but well within the
* per-step budget (the step itself is ~30-50 ms at b=128).
*
* Launch: Grid=(1,1,1), Block=(1024,1,1). Block size 1024 is the
* CUDA hard maximum; works on all foxhunt GPUs (sm_86 / sm_89 /
* sm_90). If capacity grows past 2^21 in the future, switch to
* cooperative_groups::grid().sync() for a multi-block deterministic
* path.
*
* No atomicAdd — each internal node is written by exactly one thread.
* No __threadfence() — __syncthreads() is the only inter-level barrier.
* ===================================================================== */
extern "C" __global__ void rl_per_tree_rebuild(
@@ -16,26 +47,37 @@ extern "C" __global__ void rl_per_tree_rebuild(
int capacity
)
{
const int tid_global = blockIdx.x * blockDim.x + threadIdx.x;
const int total_threads = gridDim.x * blockDim.x;
/* Single block, single-block-stride loop per level. blockIdx.x is
* guaranteed 0 by the launch geometry, but guard anyway so a
* future caller passing Grid>(1,1,1) is a silent no-op rather than
* a corrupted tree. */
if (blockIdx.x != 0) return;
/* Bottom-up: level 0 = parents of leaves, level (log2(cap)-1) = root */
/* At level L, there are capacity >> (L+1) internal nodes. */
/* Node indices at level L: [capacity >> (L+1) .. capacity >> L) */
const int tid = threadIdx.x;
const int block_size = blockDim.x;
/* Bottom-up: level 0 = parents of leaves, last level = root.
* At level L, there are (capacity >> (L+1)) internal nodes,
* with indices [capacity >> (L+1) .. capacity >> L). */
int nodes_at_level = capacity >> 1; /* level 0: cap/2 nodes */
int start = nodes_at_level; /* first node index at this level */
while (nodes_at_level >= 1) {
/* Each thread processes multiple nodes via grid-stride loop */
for (int i = tid_global; i < nodes_at_level; i += total_threads) {
/* Each thread processes multiple nodes via block-stride loop.
* Mapping `i = tid + k*block_size` is FIXED across runs:
* thread t always writes nodes (start + t), (start + t + B),
* (start + t + 2B), ... ensuring a deterministic write order.
*/
for (int i = tid; i < nodes_at_level; i += block_size) {
const int node = start + i;
priority_tree[node] = priority_tree[2 * node] + priority_tree[2 * node + 1];
priority_tree[node] = priority_tree[2 * node]
+ priority_tree[2 * node + 1];
}
/* Device-wide fence: all writes at this level visible before
* any thread reads them at the next level */
__threadfence();
/* Intra-block barrier — proper synchronisation, unlike the
* previous __threadfence() which did NOT wait for other
* blocks' writes. */
__syncthreads();
/* Move up one level */
nodes_at_level >>= 1;

View File

@@ -628,19 +628,33 @@ fn main() -> Result<()> {
}
let win_rate = if total_trades > 0 { win_count as f64 / total_trades as f64 } else { 0.0 };
let isv = trainer.isv_host_slice();
// Snapshot the ISV slots we read on the host side. Determinism
// Phase 1 (2026-06-02) turned `build_diag_value` into `&mut self`
// (needed to launch the checksum kernels), so the `isv` slice
// can no longer co-exist with the trainer mutable borrow.
// Copying the eight specific values we use is cheaper than
// re-grabbing the slice twice and keeps the diff local.
let isv_trail_fired_step = trainer.read_isv_host(RL_CONF_GATE_FIRED_COUNT_INDEX) as u64;
let isv_pyramid_add_step = trainer.read_isv_host(RL_PYRAMID_ADD_COUNT_INDEX) as u64;
let isv_conf_gate_step = trainer.read_isv_host(RL_CONF_GATE_FIRED_COUNT_INDEX) as u64;
let isv_frd_gate_step = trainer.read_isv_host(RL_FRD_GATE_FIRED_COUNT_INDEX) as u64;
let isv_heat_cap_step = trainer.read_isv_host(RL_HEAT_CAP_FIRED_COUNT_INDEX) as u64;
let isv_gamma = trainer.read_isv_host(RL_GAMMA_INDEX);
let isv_ppo_clip = trainer.read_isv_host(RL_PPO_CLIP_INDEX);
let isv_per_alpha = trainer.read_isv_host(RL_PER_ALPHA_INDEX);
let isv_reward_scale = trainer.read_isv_host(RL_REWARD_SCALE_INDEX);
// Per-step counters derived from action histogram + ISV diag slots.
let trail_fired_step = isv[RL_CONF_GATE_FIRED_COUNT_INDEX] as u64; // reusing — trail stop uses actions override
let trail_fired_step = isv_trail_fired_step; // reusing — trail stop uses actions override
let tighten_step = act_hist[ml_alpha::rl::common::Action::TrailTighten as usize] as u64;
let loosen_step = act_hist[ml_alpha::rl::common::Action::TrailLoosen as usize] as u64;
let pyramid_add_step = isv[RL_PYRAMID_ADD_COUNT_INDEX] as u64;
let pyramid_add_step = isv_pyramid_add_step;
let half_flat_long_step = act_hist[ml_alpha::rl::common::Action::HalfFlatLong as usize] as u64;
let half_flat_short_step = act_hist[ml_alpha::rl::common::Action::HalfFlatShort as usize] as u64;
let partial_flat_step = half_flat_long_step + half_flat_short_step;
let conf_gate_step = isv[RL_CONF_GATE_FIRED_COUNT_INDEX] as u64;
let frd_gate_step = isv[RL_FRD_GATE_FIRED_COUNT_INDEX] as u64;
let heat_cap_step = isv[RL_HEAT_CAP_FIRED_COUNT_INDEX] as u64;
let conf_gate_step = isv_conf_gate_step;
let frd_gate_step = isv_frd_gate_step;
let heat_cap_step = isv_heat_cap_step;
trail_fired_total += trail_fired_step;
trail_tighten_total += tighten_step;
@@ -714,10 +728,10 @@ fn main() -> Result<()> {
stats.l_pi,
stats.l_v,
stats.l_total,
isv[RL_GAMMA_INDEX],
isv[RL_PPO_CLIP_INDEX],
isv[RL_PER_ALPHA_INDEX],
isv[RL_REWARD_SCALE_INDEX],
isv_gamma,
isv_ppo_clip,
isv_per_alpha,
isv_reward_scale,
trainer.gpu_replay.capacity.min(step + 1),
done_count,
reward_sum,

View File

@@ -0,0 +1,77 @@
//! cuBLAS math-mode helper for the determinism foundation (spec §2.B + §4).
//!
//! Per `docs/superpowers/notes/2026-06-02-determinism-phase2.2-mamba2-investigation.md`:
//! the cuBLAS GEMM default algorithm (`CUBLAS_GEMM_DFALT`) permits split-K
//! accumulation with non-deterministic ordering on Ampere+ when TF32 mode
//! is on. The fix is to call `cublasSetMathMode(handle, CUBLAS_PEDANTIC_MATH)`
//! at every handle construction site — that filters out the non-deterministic
//! algorithm variants.
//!
//! Spec §4 dev/prod toggle: `FOXHUNT_DETERMINISTIC` env var (default "1" in
//! dev, where verdict trust matters; "0" in production, where the 10-15%
//! TF32 speed gain matters more than reproducibility).
//!
//! Call this immediately after `CudaBlas::new(...)` at every handle site
//! in `crates/ml-alpha/src/`. The 3 known sites (2026-06-02) are:
//! - `mamba2_block.rs` (Mamba2Block::new, line ~535)
//! - `rl/dqn.rs` (DqnHead::new, line ~364)
//! - `rl/iqn.rs` (IqnHead::new, line ~261)
//!
//! Per `feedback_isv_for_adaptive_bounds.md`: env-var read is one-shot at
//! handle construction; the mode is fixed for the lifetime of the handle.
//! No per-step ISV adjustment is needed (or possible, since the mode is
//! handle-level state).
use anyhow::{anyhow, Result};
use cudarc::cublas::CudaBlas;
/// Apply the dev/prod-toggleable cuBLAS math mode to a freshly-constructed
/// `CudaBlas` handle.
///
/// `FOXHUNT_DETERMINISTIC=1` (default) → `CUBLAS_PEDANTIC_MATH`. Deterministic
/// across runs; ~10-15% slower than TF32 on Ampere+ GEMMs.
///
/// `FOXHUNT_DETERMINISTIC=0` → `CUBLAS_TF32_TENSOR_OP_MATH`. Faster, but
/// `CUBLAS_GEMM_DFALT` may pick non-deterministic split-K variants — same-seed
/// runs can diverge by ~1e-6 after a handful of GEMMs, cascading via the
/// optimizer into ~10% trajectory drift over 200 steps.
///
/// Returns the input handle wrapped in `Ok` on success (passthrough for
/// ergonomic call-site chaining), or `Err` if the cuBLAS call fails.
pub fn apply_deterministic_math_mode(cublas: CudaBlas) -> Result<CudaBlas> {
// Per `pearl_build_rs_rerun_if_env_changed.md`: this env-var read is a
// runtime decision (not a build-time one), so no `cargo:rerun-if-env-changed`
// is needed. The env var is checked once per handle construction.
//
// Phase 2.4 control-mode addition (2026-06-02): added `2` as control
// mode (CUBLAS_DEFAULT_MATH — no TF32, no PEDANTIC, pure FP32) to
// disambiguate whether PEDANTIC fully eliminates split-K or just
// some variants. Per dispatch ask.
let env_value = std::env::var("FOXHUNT_DETERMINISTIC")
.unwrap_or_else(|_| "1".to_string());
let mode = match env_value.as_str() {
"0" => cudarc::cublas::sys::cublasMath_t::CUBLAS_TF32_TENSOR_OP_MATH,
"2" => cudarc::cublas::sys::cublasMath_t::CUBLAS_DEFAULT_MATH,
_ => cudarc::cublas::sys::cublasMath_t::CUBLAS_PEDANTIC_MATH,
};
unsafe {
cudarc::cublas::sys::cublasSetMathMode(*cublas.handle(), mode)
.result()
.map_err(|e| anyhow!("cublasSetMathMode {mode:?}: {e:?}"))?;
}
// Print once per handle so callers can verify the env var propagated.
// Phase 2.4: this is essential diagnostics — the Phase 2.3 outcome
// note claimed PEDANTIC was active but later determinism-check runs
// showed attn_q still diverging, so we need explicit confirmation.
static MODE_PRINTED: std::sync::OnceLock<()> = std::sync::OnceLock::new();
MODE_PRINTED.get_or_init(|| {
eprintln!(
"[cublas_determinism] FOXHUNT_DETERMINISTIC={env_value} → mode={mode:?}"
);
});
Ok(cublas)
}

View File

@@ -53,6 +53,11 @@ pub mod multi_horizon_labels;
// Gate reference only — Mamba2 baseline against which CfC must compete.
pub mod mamba2_block;
// Determinism foundation (spec 2026-06-02 §2.B): cuBLAS PEDANTIC mode helper
// used at every CudaBlas::new() site to disable non-deterministic GEMM
// algorithms. Gated by FOXHUNT_DETERMINISTIC env var (default "1" in dev).
pub mod cublas_determinism;
pub use isv::IsvBus;
pub use multi_horizon_labels::{generate_labels, LongHorizonLabels};
// Re-exports added as the relevant tasks land:

View File

@@ -534,6 +534,13 @@ impl Mamba2Block {
// our gemm shapes (largest is hidden_dim×in_dim ≈ 128×32).
let cublas = CudaBlas::new(Arc::clone(&stream))
.map_err(|e| anyhow!("Mamba2Block: cuBLAS init failed: {e}"))?;
// Determinism foundation (spec 2026-06-02 §2.B):
// Disable non-deterministic GEMM algorithms. Per the Phase 2.2
// mamba2 investigation, the 3 backward GEMMs in
// backward_from_h_enriched_seq_full_into were the root cause of
// the eval-pnl drift between same-seed runs.
let cublas = crate::cublas_determinism::apply_deterministic_math_mode(cublas)
.map_err(|e| anyhow!("Mamba2Block: cublas determinism setup: {e}"))?;
const CUBLAS_WORKSPACE_BYTES: usize = 8 * 1024 * 1024;
let cublas_workspace = stream
.alloc_zeros::<u8>(CUBLAS_WORKSPACE_BYTES)

View File

@@ -363,6 +363,11 @@ impl DqnHead {
// cudaMalloc that would break CUDA Graph stream capture.
let cublas = CudaBlas::new(Arc::clone(&stream))
.context("DqnHead: cuBLAS init")?;
// Determinism foundation (spec 2026-06-02 §2.B): disable
// non-deterministic GEMM algorithms. Same fix as Mamba2Block,
// applied to the DQN distributional-Q forward + backward GEMMs.
let cublas = crate::cublas_determinism::apply_deterministic_math_mode(cublas)
.context("DqnHead: cublas determinism setup")?;
const CUBLAS_WORKSPACE_BYTES: usize = 8 * 1024 * 1024;
let cublas_workspace = stream
.alloc_zeros::<u8>(CUBLAS_WORKSPACE_BYTES)

View File

@@ -212,8 +212,17 @@ pub struct IqnHead {
// ── Forward kernels (split pipeline) ─────────────────────────────
_fwd_module: Arc<CudaModule>,
/// Stage 1: tau sampling + cosine basis computation.
/// Stage 1: tau sampling + cosine basis computation. READ-ONLY on
/// `prng_state` per determinism Phase 2.6 fix (2026-06-02).
pub tau_cos_features_fn: CudaFunction,
/// Determinism Phase 2.6 sibling: advance the per-batch xorshift32
/// PRNG state by 8 steps. Run AFTER `tau_cos_features_fn` so the
/// sampling sees a stable read across all (tau_idx) blocks. The
/// kernel-launch ordering on a single stream is the grid-wide
/// barrier (replaces the racy "tau_idx==0 block writes prng_state
/// at end of monolithic kernel" pattern that produced run-dependent
/// τ values for tau_idx > 0 pre-fix).
pub advance_prng_state_fn: CudaFunction,
/// Stage 3: bias-add → ReLU → hadamard product with h_t.
pub relu_hadamard_fn: CudaFunction,
/// Stage 5: bias addition to cuBLAS output projection result.
@@ -260,6 +269,11 @@ impl IqnHead {
// ── cuBLAS handle + pre-allocated workspace ──────────────────
let cublas = CudaBlas::new(Arc::clone(&stream))
.map_err(|e| anyhow::anyhow!("IqnHead: cuBLAS init: {e}"))?;
// Determinism foundation (spec 2026-06-02 §2.B): disable
// non-deterministic GEMM algorithms. Same fix as Mamba2Block,
// applied to the IQN τ-conditioned GEMMs.
let cublas = crate::cublas_determinism::apply_deterministic_math_mode(cublas)
.map_err(|e| anyhow::anyhow!("IqnHead: cublas determinism setup: {e}"))?;
let cublas_workspace = stream
.alloc_zeros::<u8>(CUBLAS_WORKSPACE_BYTES)
.map_err(|e| anyhow::anyhow!("IqnHead: cuBLAS workspace: {e}"))?;
@@ -281,6 +295,11 @@ impl IqnHead {
let tau_cos_features_fn = fwd_module
.load_function("rl_iqn_tau_cos_features")
.context("load rl_iqn_tau_cos_features")?;
// Determinism Phase 2.6 (2026-06-02): companion advance kernel
// for the now read-only `rl_iqn_tau_cos_features`. Same cubin.
let advance_prng_state_fn = fwd_module
.load_function("rl_iqn_advance_prng_state")
.context("load rl_iqn_advance_prng_state")?;
let relu_hadamard_fn = fwd_module
.load_function("rl_iqn_relu_hadamard")
.context("load rl_iqn_relu_hadamard")?;
@@ -360,6 +379,7 @@ impl IqnHead {
_cublas_workspace: cublas_workspace,
_fwd_module: fwd_module,
tau_cos_features_fn,
advance_prng_state_fn,
relu_hadamard_fn,
bias_add_q_fn,
expected_q_fn,
@@ -479,6 +499,34 @@ impl IqnHead {
}
}
// ── Stage 1b: advance PRNG state (determinism Phase 2.6 fix).
// The previous monolithic `rl_iqn_tau_cos_features` had a read-
// write race: all N_TAU blocks per batch read `prng_state[batch]`
// at kernel head, and the (tau_idx==0) block wrote back the
// advanced state at kernel tail — with NO inter-block barrier.
// Run on the same stream → kernel-launch ordering acts as a
// grid-wide barrier between the (now read-only) sampling and
// the (single-thread-per-batch) advance. No race possible.
// Per `feedback_no_atomicadd.md` + canonical Phase-2 rebuild rule.
{
let mut args = RawArgs::new();
args.push_ptr(prng_state.raw_ptr());
args.push_i32(b_size as i32);
let mut ptrs = args.build_arg_ptrs();
let block_x = 256u32.min(b_size as u32).max(1);
let grid_x = ((b_size as u32) + block_x - 1) / block_x;
unsafe {
raw_launch(
self.advance_prng_state_fn.cu_function(),
(grid_x, 1, 1),
(block_x, 1, 1),
0,
self.raw_stream,
&mut ptrs[..args.len()],
).map_err(|e| anyhow::anyhow!("rl_iqn_advance_prng_state: {:?}", e))?;
}
}
// ── Stage 2: cuBLAS SGEMM — embed_out = cos_features @ W_embed ──
{
let w_ptr = w_embed.raw_ptr();

View File

@@ -55,7 +55,17 @@ impl OutcomeHead {
/// and sentinel -1 labels.
///
/// `b_size` is the maximum batch size for pre-allocated activation buffers.
pub fn new(b_size: usize, stream: Arc<CudaStream>) -> Result<Self> {
/// `seed` installs a `scoped_init_seed` guard around the W draw so two
/// processes seeded with the same value produce bit-equal initial weights
/// (per `pearl_scoped_init_seed_for_reproducibility`). Diagnosed in
/// determinism Phase 2.6 (2026-06-02): without this guard, the
/// `near_zero_xavier` call fell back to the time+thread-id RNG path —
/// outcome_w differed by ~0.003 between same-seed runs at step 0, then
/// stayed dormant until the first `done` event activated a non-sentinel
/// label (step 2 in mid-smoke), at which point grad_h_t_outcome flowed
/// the divergent weights into the encoder gradient. Other heads (dqn,
/// iqn, policy, value, dueling_q) already installed this guard.
pub fn new(b_size: usize, seed: u64, stream: Arc<CudaStream>) -> Result<Self> {
if b_size == 0 {
return Err(anyhow!("OutcomeHead: b_size must be > 0"));
}
@@ -116,9 +126,17 @@ impl OutcomeHead {
.map_err(|e| anyhow!("OutcomeHead: fused kernel resolve: {e}"))?;
// Xavier init W at 0.01x scale: near-zero for near-uniform softmax at init.
let w_d = ml_core::cuda_autograd::init::near_zero_xavier(
HIDDEN_DIM, K_CLASSES, &stream,
).map_err(|e| anyhow!("OutcomeHead: W init: {e}"))?;
// Determinism Phase 2.6 fix (2026-06-02): install scoped_init_seed
// around the W draw so same-seed processes produce bit-equal weights.
// The `_seed_guard` drops at end of this block, restoring whatever
// (or no) RNG state was active before. Per
// `pearl_scoped_init_seed_for_reproducibility`.
let w_d = {
let _seed_guard = ml_core::cuda_autograd::init::scoped_init_seed(seed);
ml_core::cuda_autograd::init::near_zero_xavier(
HIDDEN_DIM, K_CLASSES, &stream,
).map_err(|e| anyhow!("OutcomeHead: W init: {e}"))?
};
// Bias: zeros.
let b_d = ml_core::cuda_autograd::init::zeros(K_CLASSES, &stream)
@@ -455,7 +473,7 @@ mod tests {
None => return,
};
assert!(
OutcomeHead::new(0, stream).is_err(),
OutcomeHead::new(0, 42, stream).is_err(),
"b_size=0 must be rejected"
);
}
@@ -466,7 +484,7 @@ mod tests {
Some(s) => s,
None => return,
};
let head = OutcomeHead::new(64, stream);
let head = OutcomeHead::new(64, 42, stream);
assert!(head.is_ok(), "OutcomeHead should construct: {:?}", head.err());
}
}

File diff suppressed because it is too large Load Diff

View File

@@ -121,6 +121,21 @@ impl AdamW {
&mut self.v
}
/// Immutable access to Adam's first-moment buffer `m`. Used by the
/// determinism Phase 1 checksum diagnostic
/// (`docs/superpowers/specs/2026-06-02-determinism-foundation.md` §1.2)
/// to take a sum-of-squares snapshot post-Adam-step without
/// needing `&mut self`.
pub fn m(&self) -> &CudaSlice<f32> {
&self.m
}
/// Immutable access to Adam's second-moment buffer `v`. Companion to
/// `m()`, same use case (determinism Phase 1 checksums).
pub fn v(&self) -> &CudaSlice<f32> {
&self.v
}
/// Return the current step counter value. No GPU sync needed —
/// the counter is host-resident.
pub fn step_count(&self) -> i32 {

View File

@@ -4169,6 +4169,331 @@ impl PerceptionTrainer {
&self.h_t_d
}
/// Determinism Phase 2.2 sub-investigation: read-only accessors for
/// encoder intermediate activations. Used exclusively by
/// `IntegratedTrainer::compute_diag_checksums` to checksum each layer's
/// output and localise the first non-deterministic step in the encoder
/// chain (VSN → Mamba2 L1 → LN_a → Mamba2 L2 → LN_b → attn → CfC).
///
/// All returned slices remain valid until the next `forward_encoder`
/// (which overwrites them). Caller is responsible for synchronising
/// the stream before reading host-side.
pub fn debug_vsn_out_view(&self) -> &CudaSlice<f32> {
self.vsn_out_d.cuda_data()
}
pub fn debug_mamba2_l1_out_view(&self) -> &CudaSlice<f32> {
self.mamba2_fwd_scratch.h_enriched_seq.cuda_data()
}
pub fn debug_ln_a_out_view(&self) -> &CudaSlice<f32> {
self.ln_a_out_d.cuda_data()
}
pub fn debug_mamba2_l2_out_view(&self) -> &CudaSlice<f32> {
self.mamba2_l2_fwd_scratch.h_enriched_seq.cuda_data()
}
pub fn debug_ln_b_out_view(&self) -> &CudaSlice<f32> {
&self.ln_out_d
}
pub fn debug_attn_context_view(&self) -> &CudaSlice<f32> {
&self.attn_context_d
}
/// Mamba2 L1 reduced backward output: dw_c [hidden_dim, state_dim].
/// Catches non-determinism in the `mamba2_alpha_reduce_d_w_c` kernel
/// (or upstream backward kernels that feed it).
pub fn debug_mamba2_l1_grad_w_c_view(&self) -> &CudaSlice<f32> {
self.mamba2_grads_buffers.dw_c.cuda_data()
}
/// Mamba2 L2 reduced backward output: dw_c.
pub fn debug_mamba2_l2_grad_w_c_view(&self) -> &CudaSlice<f32> {
self.mamba2_l2_grads_buffers.dw_c.cuda_data()
}
/// Determinism Phase 2.6 sub-investigation: pre-reduce per-batch /
/// per-row param-grad scratches (Group I of the Phase 2.6 dispatch
/// ask). These are the buffers that the K-loop's `+=` accumulates
/// into BEFORE `reduce_axis0` collapses them across B (or n_rows for
/// VSN / LN). If two same-seed runs produce EQUAL scratches but
/// DIVERGE on the post-reduce final grads, `reduce_axis0` is the bug
/// (canonical Phase-2-PER-rebuild fix applies). If the scratches
/// already diverge, the upstream per-batch backward kernel
/// (`cfc_step_backward_batched`, `attention_pool_bwd`,
/// `variable_selection_bwd`, `layer_norm_bwd`) is the bug.
pub fn debug_cfc_grad_w_in_scratch_view(&self) -> &CudaSlice<f32> {
&self.cfc_grad_w_in_scratch_d
}
pub fn debug_cfc_grad_w_rec_scratch_view(&self) -> &CudaSlice<f32> {
&self.cfc_grad_w_rec_scratch_d
}
pub fn debug_cfc_grad_b_scratch_view(&self) -> &CudaSlice<f32> {
&self.cfc_grad_b_scratch_d
}
pub fn debug_cfc_grad_tau_scratch_view(&self) -> &CudaSlice<f32> {
&self.cfc_grad_tau_scratch_d
}
pub fn debug_vsn_grad_w_scratch_view(&self) -> &CudaSlice<f32> {
&self.vsn_grad_w_scratch_d
}
pub fn debug_vsn_grad_b_scratch_view(&self) -> &CudaSlice<f32> {
&self.vsn_grad_b_scratch_d
}
pub fn debug_attn_grad_q_scratch_view(&self) -> &CudaSlice<f32> {
&self.attn_grad_q_scratch_d
}
pub fn debug_grad_ln_b_gain_per_row_view(&self) -> &CudaSlice<f32> {
&self.grad_ln_gain_per_row_d
}
pub fn debug_grad_ln_b_bias_per_row_view(&self) -> &CudaSlice<f32> {
&self.grad_ln_bias_per_row_d
}
pub fn debug_grad_ln_a_gain_per_row_view(&self) -> &CudaSlice<f32> {
&self.grad_ln_a_gain_per_row_d
}
pub fn debug_grad_ln_a_bias_per_row_view(&self) -> &CudaSlice<f32> {
&self.grad_ln_a_bias_per_row_d
}
/// Determinism Phase 2.2 sub-investigation: raw DtoH dump of mamba2
/// state for env-gated debugging. Mirrors `dump_per_state_for_debug`
/// in shape: env var `FOXHUNT_DETERMINISM_DEBUG_MAMBA2=1` gates the
/// dump, `$FOXHUNT_DEBUG_DUMP_DIR` (defaults to
/// `/tmp/foxhunt-determinism-debug`) names the output directory.
/// Only fires for `step ∈ {0, 1, 2, 3}`.
///
/// Dumps per step (12 files per layer × 2 layers + 6 activations):
/// - mamba2_l{1,2}_w_in.bin / w_a.bin / w_b.bin / w_c.bin / w_out.bin
/// - mamba2_l{1,2}_h_enriched_seq.bin [N, K, hidden_dim]
/// - vsn_out.bin, ln_a_out.bin, ln_b_out.bin, attn_context.bin
///
/// Per `feedback_cpu_is_read_only` the dump is a pure host READ.
/// Per `feedback_no_stubs` it is wired end-to-end; no-op when env unset.
pub fn dump_mamba2_state_for_debug(&self, step: u64) -> Result<()> {
use std::path::PathBuf;
let enabled = std::env::var("FOXHUNT_DETERMINISM_DEBUG_MAMBA2")
.map(|v| v == "1")
.unwrap_or(false);
if !enabled || step > 3 {
return Ok(());
}
let dump_dir = std::env::var("FOXHUNT_DEBUG_DUMP_DIR")
.unwrap_or_else(|_| "/tmp/foxhunt-determinism-debug".to_string());
std::fs::create_dir_all(&dump_dir).with_context(|| {
format!("dump_mamba2_state_for_debug: create_dir_all {dump_dir}")
})?;
// Synchronise the trainer stream so all in-flight forward kernels
// have committed before we read.
unsafe {
crate::trainer::raw_launch::raw_stream_sync(self.raw_stream)
.map_err(|e| anyhow::anyhow!("dump_mamba2 sync stream: {:?}", e))?;
}
// DtoH via the shared mapped-pinned helper. Raw
// `stream.memcpy_dtoh` on a regular `Vec<f32>` is forbidden by
// `feedback_no_htod_htoh_only_mapped_pinned` (non-pinned host
// slice triggers a blocking driver HtoH bounce inside cudarc);
// `read_slice_d_into` stages through a `MappedRecordBuffer<f32>`
// and uses DtoD + sync + volatile reads.
let stream = self.stream.clone();
let dump_f32 =
|path: &std::path::Path, slice: &CudaSlice<f32>| -> Result<()> {
let mut host: Vec<f32> = vec![0.0f32; slice.len()];
crate::trainer::integrated::read_slice_d_into(
&stream,
slice,
host.as_mut_slice(),
)
.with_context(|| format!("dump_mamba2 dtoh {path:?}"))?;
let bytes: &[u8] = unsafe {
std::slice::from_raw_parts(
host.as_ptr() as *const u8,
host.len() * std::mem::size_of::<f32>(),
)
};
std::fs::write(path, bytes)
.with_context(|| format!("dump_mamba2 write {path:?}"))?;
Ok(())
};
let base = PathBuf::from(&dump_dir);
// Mamba2 L1 weights.
let m1 = self.trunk.mamba2_l1();
dump_f32(&base.join(format!("step_{step}_mamba2_l1_w_in.bin")), &m1.w_in.weight)?;
dump_f32(&base.join(format!("step_{step}_mamba2_l1_b_in.bin")), &m1.w_in.bias)?;
dump_f32(&base.join(format!("step_{step}_mamba2_l1_w_a.bin")), &m1.w_a.weight)?;
dump_f32(&base.join(format!("step_{step}_mamba2_l1_b_a.bin")), &m1.w_a.bias)?;
dump_f32(&base.join(format!("step_{step}_mamba2_l1_w_b.bin")), &m1.w_b.weight)?;
dump_f32(&base.join(format!("step_{step}_mamba2_l1_b_b.bin")), &m1.w_b.bias)?;
dump_f32(&base.join(format!("step_{step}_mamba2_l1_w_c.bin")), &m1.w_c)?;
dump_f32(&base.join(format!("step_{step}_mamba2_l1_w_out.bin")), &m1.w_out.weight)?;
dump_f32(&base.join(format!("step_{step}_mamba2_l1_b_out.bin")), &m1.w_out.bias)?;
// Mamba2 L2 weights.
let m2 = self.trunk.mamba2_l2();
dump_f32(&base.join(format!("step_{step}_mamba2_l2_w_in.bin")), &m2.w_in.weight)?;
dump_f32(&base.join(format!("step_{step}_mamba2_l2_b_in.bin")), &m2.w_in.bias)?;
dump_f32(&base.join(format!("step_{step}_mamba2_l2_w_a.bin")), &m2.w_a.weight)?;
dump_f32(&base.join(format!("step_{step}_mamba2_l2_b_a.bin")), &m2.w_a.bias)?;
dump_f32(&base.join(format!("step_{step}_mamba2_l2_w_b.bin")), &m2.w_b.weight)?;
dump_f32(&base.join(format!("step_{step}_mamba2_l2_b_b.bin")), &m2.w_b.bias)?;
dump_f32(&base.join(format!("step_{step}_mamba2_l2_w_c.bin")), &m2.w_c)?;
dump_f32(&base.join(format!("step_{step}_mamba2_l2_w_out.bin")), &m2.w_out.weight)?;
dump_f32(&base.join(format!("step_{step}_mamba2_l2_b_out.bin")), &m2.w_out.bias)?;
// Encoder intermediate activations (post-forward).
dump_f32(&base.join(format!("step_{step}_vsn_out.bin")), self.vsn_out_d.cuda_data())?;
dump_f32(&base.join(format!("step_{step}_mamba2_l1_h_enriched_seq.bin")),
self.mamba2_fwd_scratch.h_enriched_seq.cuda_data())?;
dump_f32(&base.join(format!("step_{step}_ln_a_out.bin")), self.ln_a_out_d.cuda_data())?;
dump_f32(&base.join(format!("step_{step}_mamba2_l2_h_enriched_seq.bin")),
self.mamba2_l2_fwd_scratch.h_enriched_seq.cuda_data())?;
dump_f32(&base.join(format!("step_{step}_ln_b_out.bin")), &self.ln_out_d)?;
dump_f32(&base.join(format!("step_{step}_attn_context.bin")), &self.attn_context_d)?;
dump_f32(&base.join(format!("step_{step}_h_t.bin")), &self.h_t_d)?;
// Phase 2.4: mamba2 L1 backward scratch + reduced output. Catches
// the residual `mamba2_grad_w_c_l1` divergence (Phase 2.3 outcome
// note bottom section, 2026-06-02) which persists at Δ ~1e-7 after
// PEDANTIC cuBLAS lands. If `d_w_c_per_sample` (pre-reduce)
// matches but `dw_c` (post-reduce) diverges → reduce kernel is
// the bug. If `d_w_c_per_sample` itself diverges → scan_bwd_seq
// or its inputs (`d_h_enriched_seq`, `a_proj`, `b_proj`, `w_c`)
// is the bug.
dump_f32(
&base.join(format!("step_{step}_mamba2_l1_d_w_c_per_sample.bin")),
&self.mamba2_bwd_scratch.d_w_c_per_sample,
)?;
dump_f32(
&base.join(format!("step_{step}_mamba2_l1_dw_c_reduced.bin")),
self.mamba2_grads_buffers.dw_c.cuda_data(),
)?;
dump_f32(
&base.join(format!("step_{step}_mamba2_l1_d_h_s2.bin")),
&self.mamba2_bwd_scratch.d_h_s2,
)?;
dump_f32(
&base.join(format!("step_{step}_mamba2_l2_d_w_c_per_sample.bin")),
&self.mamba2_l2_bwd_scratch.d_w_c_per_sample,
)?;
dump_f32(
&base.join(format!("step_{step}_mamba2_l2_dw_c_reduced.bin")),
self.mamba2_l2_grads_buffers.dw_c.cuda_data(),
)?;
// Phase 2.4 follow-up: also dump dw_in / dw_a / dw_b (cuBLAS-GEMM
// outputs from `backward_with_slices_into`) and d_a_proj_2d /
// d_b_proj_2d (reduce_d_proj outputs feeding the cuBLAS calls).
// If d_a_proj / d_b_proj are EQUAL at step 1 but dw_a / dw_b
// diverge, the cuBLAS GEMMs are the bug (PEDANTIC didn't help
// for these shapes). If d_a_proj diverges at step 1 → reduce
// kernel `mamba2_alpha_reduce_d_proj` is non-det. If d_a_per_channel
// diverges at step 1 → scan_bwd_seq is non-det (very surprising).
dump_f32(
&base.join(format!("step_{step}_mamba2_l1_d_a_per_channel.bin")),
&self.mamba2_bwd_scratch.d_a_per_channel,
)?;
dump_f32(
&base.join(format!("step_{step}_mamba2_l1_d_b_per_channel.bin")),
&self.mamba2_bwd_scratch.d_b_per_channel,
)?;
dump_f32(
&base.join(format!("step_{step}_mamba2_l1_d_a_proj_2d.bin")),
self.mamba2_grads_buffers.d_a_proj_2d.cuda_data(),
)?;
dump_f32(
&base.join(format!("step_{step}_mamba2_l1_d_b_proj_2d.bin")),
self.mamba2_grads_buffers.d_b_proj_2d.cuda_data(),
)?;
dump_f32(
&base.join(format!("step_{step}_mamba2_l1_dw_in.bin")),
self.mamba2_grads_buffers.dw_in.cuda_data(),
)?;
dump_f32(
&base.join(format!("step_{step}_mamba2_l1_dw_a.bin")),
self.mamba2_grads_buffers.dw_a.cuda_data(),
)?;
dump_f32(
&base.join(format!("step_{step}_mamba2_l1_dw_b.bin")),
self.mamba2_grads_buffers.dw_b.cuda_data(),
)?;
dump_f32(
&base.join(format!("step_{step}_mamba2_l1_db_in.bin")),
self.mamba2_grads_buffers.db_in.cuda_data(),
)?;
dump_f32(
&base.join(format!("step_{step}_mamba2_l1_db_a.bin")),
self.mamba2_grads_buffers.db_a.cuda_data(),
)?;
dump_f32(
&base.join(format!("step_{step}_mamba2_l1_db_b.bin")),
self.mamba2_grads_buffers.db_b.cuda_data(),
)?;
dump_f32(
&base.join(format!("step_{step}_mamba2_l1_d_x.bin")),
self.mamba2_grads_buffers.d_x.cuda_data(),
)?;
dump_f32(
&base.join(format!("step_{step}_mamba2_l1_d_x_from_in.bin")),
self.mamba2_grads_buffers.d_x_from_in.cuda_data(),
)?;
// Phase 2.4 follow-up #2: dump trunk weights (vsn, cfc, attn, ln)
// — these are upstream of mamba2 and would propagate via different
// forward state at step N+1 if their Adam updates diverged at
// step N. Plus their per-step backward grads to localise WHICH
// trunk weight first diverges.
dump_f32(
&base.join(format!("step_{step}_vsn_w.bin")),
&self.trunk.vsn_w_d,
)?;
dump_f32(
&base.join(format!("step_{step}_attn_q.bin")),
&self.trunk.attn_q_d,
)?;
dump_f32(
&base.join(format!("step_{step}_cfc_w_in.bin")),
&self.trunk.w_in_d,
)?;
dump_f32(
&base.join(format!("step_{step}_cfc_w_rec.bin")),
&self.trunk.w_rec_d,
)?;
dump_f32(
&base.join(format!("step_{step}_ln_a_gain.bin")),
&self.trunk.ln_a_gain_d,
)?;
dump_f32(
&base.join(format!("step_{step}_ln_b_gain.bin")),
&self.trunk.ln_b_gain_d,
)?;
// Trunk backward grads (pre-Adam).
dump_f32(
&base.join(format!("step_{step}_grad_vsn_w.bin")),
&self.grad_vsn_w_d,
)?;
dump_f32(
&base.join(format!("step_{step}_grad_attn_q.bin")),
&self.grad_attn_q_d,
)?;
dump_f32(
&base.join(format!("step_{step}_grad_cfc_w_in.bin")),
&self.grad_w_in_d,
)?;
dump_f32(
&base.join(format!("step_{step}_grad_cfc_w_rec.bin")),
&self.grad_w_rec_d,
)?;
dump_f32(
&base.join(format!("step_{step}_grad_ln_a_gain.bin")),
&self.grad_ln_a_gain_d,
)?;
dump_f32(
&base.join(format!("step_{step}_grad_ln_b_gain.bin")),
&self.grad_ln_gain_d,
)?;
eprintln!(
"[determinism-debug] dumped mamba2 state step={step} to {dump_dir}"
);
Ok(())
}
/// Phase E.3a (integrated RL trainer): run the encoder backward
/// consuming a caller-provided `grad_h_t` `[B × HIDDEN_DIM]` at slot
/// `K-1`.

View File

@@ -244,7 +244,18 @@ fn eval_diag_jsonl_emitted_with_train_schema_parity() -> Result<()> {
// frac_session_neg) → 675. Edge-decay Phase 1 adds 3 leaves
// (edge_ph_mean / edge_ph_frac_alerted / edge_ph_frac_warmup) → 678.
// Reward-policy realignment adds 1 leaf (rewards.pure_pnl_mode) → 679.
const EXPECTED_LEAVES: usize = 679;
// 2026-06-02 Determinism Phase 1: +15 leaves under top-level
// `checksums` key (encoder_output, q_logits, pi_logits, v_value,
// replay_sample_indices, rewards_after_shape, advantages, q_grad,
// pi_grad, v_grad, encoder_grad, adam_m_sum, adam_v_sum, isv_state,
// popart_sigma_state). New total: 694.
// 2026-06-02 Determinism Phase 2.2: +17 leaves under `checksums`
// for encoder localisation — mamba2_l1 weights (w_in/w_a/w_b/w_c) +
// mamba2_l2 weights (w_in/w_c) + vsn_w + cfc_w_in + cfc_w_rec +
// attn_q + encoder activations (vsn_out, mamba2_l1_out, ln_a_out,
// mamba2_l2_out, ln_b_out, attn_context) + mamba2_grad_w_c_l1.
// New total: 711.
const EXPECTED_LEAVES: usize = 711;
anyhow::ensure!(
train_paths.len() == EXPECTED_LEAVES,
"train leaf count {} != expected {EXPECTED_LEAVES} (schema drifted; \

View File

@@ -0,0 +1,392 @@
# Session follow-up — Phase 3 IQN-Complement post-mortem + Phase 4 foundation
**Date:** 2026-05-29 → 2026-05-30
**Branch lineage:** Plan A v2 `fd3174262` → 5 failed experiments → Phase 4 design + foundation
**Production baseline preserved:** Plan A v2 (+$9.3M peak, 20k clean, no NaN)
---
## 1. Session arc
This session went deep on adding **proper dueling Q-learning** to the
foxhunt RL trainer. Started from Plan A v2 baseline (vanilla C51 + IQN
auxiliary, no dueling). After 5 cluster experiments and 4 architectural
pearls, landed on the correct design (Phase 4: independent dueling head)
and shipped its foundation.
### Timeline
| Phase | Architecture | Commit | Outcome |
|---|---|---|---|
| Plan A v2 | C51 + IQN aux, scalar V baseline | `fd3174262` | **+$9.3M peak, 20k clean** (production baseline) |
| Phase 2 v2 | C51 + scalar dueling V | `eac382d20` | ❌ Math no-op (Phase 2 from prior session) |
| Phase 3.1 | IQN V/A split, composed Q output | `6b7c9521d` | ❌ Cluster step 745: qpa=-0.93 |
| Phase 3.2 | 3.1 + V_IQN replaces V_scalar PPO | `0fdce0b44` | ❌ Cluster step 556: qpa=-0.96 |
| Phase 3.2c | 3.2 + α-warmup blend | `8616cf9cb` | ❌ Cluster step 762: qpa=-0.94 |
| Phase 3.1-fix | 3.1 + dual output (raw A for ensemble) | `87540e748` | ❌ Cluster step 746: qpa=-0.92 |
| **Phase 4 spec** | Independent dueling head | `5160ba781` | ✅ Design captured |
| **Phase 4.0** | Forward kernel + struct | `af35bc778` | ✅ Compiles, smoke passes |
| **Phase 4.1** | Loss + Bellman target + decompose backward | `13bf277cd` | ✅ Compiles, smoke passes |
---
## 2. The architectural journey — what each failure taught us
### Phase 2 v2 (prior session, recapped)
**Attempt:** Add scalar V mediation to categorical C51 via
`composed_logits[b, a, z] = V[b] + A_logits[b, a, z] mean_a A_logits[b, a, z]`.
**Failure:** Mathematical no-op.
`softmax_z(V[b] + X[a, z]) = softmax_z(X[a, z])` because softmax is
translation-invariant in z and V is constant across z. V drops out
entirely. CE gradient also collapses: `Σ_z (softmax target) = 0`.
**Pearl:** `pearl_scalar_v_with_categorical_ce_is_noop`
### Phase 3 IQN-Complement approach
**Insight:** IQN's per-τ Q values are scalars (not categorical
distributions over atoms). Scalar dueling works mathematically at fixed
τ. So add IQN-with-dueling as a complement to C51 (which keeps Plan A
v2 behavior for Q→π distill), and use V_IQN derived from IQN's composed
Q for the PPO advantage baseline.
**Pearl:** `pearl_complement_dont_replace_with_dual_architecture`
### Phase 3.2 — first attempt (direct V swap)
**Attempt:** Phase 3.1 (IQN V/A split) + V_IQN replaces V_scalar in PPO
advantage. `V_IQN(s) = (1/N) × Σ_a E_τ[Q(s, a, τ)]` (exact recovery
from dueling identifiability).
**Failure:** Cluster step 556 qpa=-0.96. The math is exact but the
PRACTICAL issue: `A = E_Q_C51 V_IQN` mixes outputs from two
networks at different output scales. Early in training (when neither
network's V/Q values are well-calibrated), the difference
systematically biases the advantage sign → π pushed AWAY from Q's
preferences → anti-correlation.
**Pearl:** `pearl_complement_architectures_must_be_calibrated_when_mixed`
### Phase 3.2c — α-blend warmup
**Attempt:** Blend `V_used = α × V_scalar + (1-α) × V_IQN` with α
ramping from 1.0 (Plan A v2 behavior) to 0.0 over steps 2000-10000.
At α=1.0, V_used = V_scalar exactly — should match Plan A v2.
**Failure:** Cluster step 762 qpa=-0.94. Even with α=1.0 (V_blended =
V_scalar mathematically), qpa collapses identically. This was the
diagnostic smoking gun — the V baseline integration **wasn't the
problem**. Something earlier in Phase 3.1 was.
### Phase 3.1 isolation
**Attempt:** Just Phase 3.1 (IQN dueling internal), Plan A v2's
V_scalar path UNCHANGED. Isolation test to localize the breakage.
**Failure:** Cluster step 745 qpa=-0.93. Confirmed: **Phase 3.1
alone breaks downstream consumers, independent of V baseline**.
**Pearl (initial form):** `pearl_complement_internal_loss_vs_external_consumers`
**Root cause:** Phase 3.1 changed `iqn_expected_q` to compute mean_τ
over COMPOSED Q (`V + A mean_a A`). The values become CENTERED
around V(s). The ensemble Q kernel blends C51's UNCENTERED expected_q
with IQN's CENTERED expected_q. The blend's softmax distribution
behaves differently from Plan A v2's, perturbing the q_pi_agree
diagnostic.
### Phase 3.1-fix — dual output
**Attempt:** Add a new kernel `rl_iqn_extract_raw_a_with_bias` that
extracts raw A logits (uncentered) alongside the composed Q output.
expected_q now operates on raw A (Plan A v2 distribution shape) for
ensemble use; composed Q still feeds the Huber loss.
**Failure:** Cluster step 746 qpa=-0.92. **Dual output extraction is
necessary but not sufficient.** Even though raw A has matching
distribution shape (translation-equivalent to Plan A v2's uncentered
Q), the WEIGHTS train differently in Phase 3.1's dueling regime:
- A's gradient has mean-zero structure across actions:
`grad_A[a] = grad_q[a] (1/N) Σ_a' grad_q[a']`
- sum_a grad_A[a] = 0 by construction
- A weights and bias train to capture per-state action DEVIATIONS,
not raw Q values
The "raw A + bias" extracted in Phase 3.1-fix represents
deviation-shaped values, not Plan A v2's uncentered Q values. When
fed to ensemble + softmax, the absolute magnitudes are smaller →
ensemble dominated by C51's larger magnitudes → ranking dynamics
shift → qpa stays decoupled.
**Pearl (updated):** `pearl_complement_internal_loss_vs_external_consumers`
addendum documents this gradient-structure finding.
---
## 3. The four architectural pearls
These pearls form a complete design pattern for adding complementary
ML architectures. They generalize beyond this project — any future
work that adds a complement to a working primary architecture should
respect all four:
### Pearl 1 — Don't replace what works
`pearl_complement_dont_replace_with_dual_architecture`
When primary architecture (C51) has demonstrated edge but can't
mathematically support a desirable property (dueling), add a
complementary architecture (IQN) that natively has it. Use each for
what it's mathematically suited to.
### Pearl 2 — Calibrate cross-architecture signals
`pearl_complement_architectures_must_be_calibrated_when_mixed`
When the complement's output is mixed with the primary's in a single
loss term, the two outputs MUST be on similar scales — otherwise even
mathematically correct values produce systematic sign bias.
### Pearl 3 — Match output distribution shape AND gradient structure
`pearl_complement_internal_loss_vs_external_consumers` (+ addendum)
When refactoring a complement's loss function, ALL downstream
consumers of its outputs must be checked. Distribution shape matching
is NECESSARY but NOT SUFFICIENT — the gradient structure that trains
weights must also match the primary's, or downstream consumers see
different absolute magnitudes over training.
### Pearl 4 — Forward math ≠ training-correct
(implicit from Phase 3.1-fix failure)
Mathematical analysis of forward outputs is insufficient for adding
complementary architectures. Backward gradient flow shapes weight
training over thousands of steps; downstream consumers depend on
trained weight magnitudes for temperature scaling, ranking, etc.
Architectural changes that pass forward-math sanity can still break
downstream dynamics if gradient structure changes.
---
## 4. Phase 4 design — what makes it different
**The synthesis of all four pearls:** to add dueling without breaking
anything, **completely encapsulate** the dueling architecture. Don't
modify C51's or IQN's weights, outputs, or gradient paths in any way.
### Architecture
```
Encoder (shared) → h_t
├→ C51 head (Plan A v2 — UNCHANGED)
├→ IQN head (Plan A v2 — UNCHANGED)
├→ Policy head π (Plan A v2 — UNCHANGED)
├→ value_head V_scalar (still trains for diag comparison)
└→ NEW: DuelingQHead (parallel, fully encapsulated)
• Has its OWN weights w_v, b_v, w_a, b_a
• Trained by its OWN Bellman loss
• V output → PPO advantage baseline (Phase 4.3 swap)
• composed_Q NEVER feeds ensemble/distill/selection
```
### Why this is structurally safe (cross-checking against each pearl)
| Pearl | Phase 4 mitigation |
|---|---|
| 1: Don't replace | C51, IQN, π, value_head all unchanged |
| 2: Calibrate when mixed | V_dq trained on SAME Bellman reward signal as C51 — scales align by construction |
| 3: Match distribution + gradient | composed_Q_dq never feeds ANY downstream consumer — gradient structure changes affect only DuelingQHead's own weights |
| 4: Training-correct ≠ math-correct | DuelingQHead's weights are isolated; perturbations only affect V_dq output which has a single, well-defined consumer (PPO advantage) |
Every prior failure mode is **structurally impossible** in Phase 4.
---
## 5. Phase 4 implementation state
### Shipped (this session)
**Phase 4.0** (`af35bc778`):
- `rl_dueling_q_forward.cu` kernel — V[B] + A[B,N] + composed_Q[B,N]
in one fused pass. Grid (B,1,1), block (HIDDEN_DIM,1,1).
- `DuelingQHead` Rust struct in `crates/ml-alpha/src/rl/dueling_q.rs`:
- Config + new() with Xavier init
- Online weights: w_v_d, b_v_d, w_a_d, b_a_d
- Target weights: mirrors (soft-update wiring in 4.2)
- forward() + forward_target() + forward_inner()
**Phase 4.1** (`13bf277cd`):
- `rl_dueling_q_bellman_target.cu` — argmax + scalar Bellman target
- `rl_dueling_q_loss_and_grad.cu` — Huber loss + grad_composed
- `rl_dueling_q_decompose_and_bwd.cu` — Jacobian + outer-product weight grads
- 3 new DuelingQHead methods:
- build_bellman_target()
- compute_loss_and_grad()
- decompose_and_backward_to_weights()
### Remaining (next session)
**Phase 4.2 — trainer integration (diagnostic mode, ~150 LOC):**
Add to `IntegratedTrainer`:
- `dueling_q_head: DuelingQHead` field + 4 Adam states (w_v, b_v, w_a, b_a)
- 10+ buffer fields:
- V_dq[B] / V_dq_tp1[B] (for PPO advantage when wired)
- A_dq[B,N] (diag)
- composed_Q_dq[B,N] (for loss)
- target_composed_Q_dq[B,N] (for Bellman target build)
- target_value_dq[B] (Bellman target output)
- loss_pb_dq[B]
- grad_composed_dq[B,N]
- per-batch weight grads (4 buffers)
- reduced weight grads (4 buffers)
- Wiring in `step_with_lobsim_gpu_body`:
1. forward(h_t_borrow) → V_dq, A_dq, composed_Q_dq
2. forward(h_tp1_d) → V_dq_tp1 (for PPO)
3. forward(sampled_h_t_d) → online composed_Q for loss
4. forward_target(sampled_h_tp1_d) → target composed_Q
5. build_bellman_target → target_value_dq
6. compute_loss_and_grad → grad_composed_dq + loss_pb_dq
7. decompose_and_backward_to_weights → per-batch grads
8. reduce_axis0 → final weight grads
9. Adam step on each weight
10. Soft-update target net (reuse `dqn_target_soft_update_fn`)
- **DO NOT yet swap v_pred_d in compute_advantage_return** —
diagnostic-only mode for Phase 4.2
**Phase 4.3 — V_dq → PPO swap (~10 LOC):**
After 4.2 cluster-validates V_dq trains stably without affecting
Plan A v2 metrics:
- Replace `v_pred_d``dueling_v_d` in compute_advantage_return call
- Replace `v_pred_tp1_d``dueling_v_tp1_d`
- value_head still trains via MSE on returns (for diagnostic
comparison V_dq vs V_scalar)
### Validation strategy (per spec §8)
1. Local smoke + compute-sanitizer at b=128 — should produce 0 errors
with healthy l_dueling > 0
2. Cluster smoke b=1024 SHORT (5k steps) at Phase 4.2 — verify:
- qpa_ema near 0 (Plan A v2 unchanged because V_dq doesn't feed PPO)
- V_dq trains stably (loss decreasing, V values reasonable)
- dueling_v_at_taken_ema rises over training
3. Cluster full b=1024 20k at Phase 4.3 — verify:
- qpa healthy throughout
- PnL matches or exceeds Plan A v2 +$9.3M peak
### Kill criteria
Any one triggers abort:
- qpa_ema < -0.3 at any checkpoint
- nan_abort_step > 0
- l_dueling explodes (> 100× Plan A v2's l_q baseline)
- pnl_cum_usd < Plan A v2's at same step by > $5M
---
## 6. Operational lessons (separate from architectural pearls)
### Branch ref trap — hit twice
Argo workflow's `git checkout origin/$BRANCH` always picks the
branch TIP, not the `commit-sha` param (which is used only for
output paths). Lessons:
1. **For isolation tests**, create a dedicated branch with the
specific commit as tip:
```
git branch -f ml-alpha-phaseN-only <commit-sha>
git push origin ml-alpha-phaseN-only --force
```
2. **Always verify cluster-side**: check the diag JSONL path output
(`/feature-cache/alpha-rl-runs/$SHA/`) matches the expected
commit-sha before assuming the cluster is running what you
intended.
3. **Avoid detached HEAD commits** — they create branch ref vs HEAD
divergence. After each commit on a worktree, verify
`git rev-parse $BRANCH == git rev-parse HEAD`.
### JSONL monitoring beats stdout
Argo workflow template's `log-every=5000` only emits 4 visible step
markers across a 20k run. The per-step JSONL (`diag.jsonl` on
feature-cache PVC) is the source of truth. Monitor pattern:
```bash
kubectl exec -n foxhunt $WF -c main -- bash -c \
"tail -1 /feature-cache/alpha-rl-runs/$SHA/fold0/diag.jsonl"
```
Polls every 90s, emits per 500-step milestone. Catches qpa collapse
within ~3 minutes of training start.
### Kill early on qpa collapse
The `q_pi_agree_ema` diagnostic is THE most sensitive cross-arch
health signal. If qpa drops below -0.3 by step 1000, KILL — the
architecture is broken and waiting longer just confirms what's
already known. Saved hours of cluster cycles by killing at step
556 (Phase 3.2), 762 (Phase 3.2c), 745 (3.1 isolation), 746
(3.1-fix).
---
## 7. References
### Pearls created this session
- `pearl_complement_dont_replace_with_dual_architecture` (2026-05-29)
- `pearl_complement_architectures_must_be_calibrated_when_mixed` (2026-05-30)
- `pearl_complement_internal_loss_vs_external_consumers` (2026-05-30 + addendum)
- `pearl_scalar_v_with_categorical_ce_is_noop` (carried from prior session)
### Documents written this session
- `docs/superpowers/plans/2026-05-29-phase2-post-mortem-and-redesign.md`
- `docs/superpowers/specs/2026-05-30-phase4-independent-dueling-head-design.md`
- This file: `docs/superpowers/notes/2026-05-30-phase3-iqn-complement-postmortem-and-phase4-foundation.md`
### Cluster runs (chronological)
| Workflow | Commit | Result | Notes |
|---|---|---|---|
| alpha-rl-8ksnj | fd3174262 | ✅ Succeeded | Plan A v2 production baseline (+$9.3M peak) |
| alpha-rl-dxlnm | 7790c6815 | Failed | First branch ref trap |
| alpha-rl-hcmk7 | a7b0d6d1a | Failed | HIDDEN_DIM=512 OOB (Phase 3.2-pre) |
| alpha-rl-cs-pq2jg | a7b0d6d1a | Diagnosed | compute-sanitizer found the OOB |
| alpha-rl-2qbmz | eac382d20 | Killed | Phase 2 v2 step 4267 qpa=-0.79 |
| alpha-rl-gvqk5 | 0fdce0b44 | Killed | Phase 3.2 step 556 qpa=-0.96 |
| alpha-rl-nl5c9 | 8616cf9cb | Killed | Phase 3.2c step 762 qpa=-0.94 |
| alpha-rl-582bn | (wrong branch) | Killed | Second branch ref trap |
| alpha-rl-xrbgr | 6b7c9521d | Killed | Phase 3.1 isolation step 745 qpa=-0.93 |
| alpha-rl-9q2pd | 87540e748 | Killed | Phase 3.1-fix step 746 qpa=-0.92 |
10 cluster cycles. Plan A v2 baseline preserved untouched throughout.
### Branches on remote (for archaeology)
- `ml-alpha-phase3-iqn-complement` @ `8616cf9cb` (Phase 3.1 + 3.2 + 3.2c)
- `ml-alpha-phase3.1-only` @ `6b7c9521d` (Phase 3.1 isolation test)
- `ml-alpha-phase3.1-fix` @ `87540e748` (Phase 3.1 + dual output fix)
- `ml-alpha-phase4-dueling-head` @ `13bf277cd` **← Phase 4 foundation, current work**
- `main` (Plan A v2 baseline at `fd3174262` via merge)
---
## 8. Next session opening checklist
When resuming:
1. ☐ Pull `ml-alpha-phase4-dueling-head` @ `13bf277cd`
2. ☐ Read this note + spec `docs/superpowers/specs/2026-05-30-phase4-independent-dueling-head-design.md` §6
3. ☐ Create worktree (or reuse `/tmp/foxhunt-de-tests/phase4-dueling-head`)
4. ☐ Start Phase 4.2: add DuelingQHead instance + buffers + Adam states to IntegratedTrainer
5. ☐ Wire 10-step launch sequence in step_with_lobsim_gpu_body
6. ☐ DO NOT swap v_pred_d yet (4.3 is separate commit)
7. ☐ Local smoke + sanitizer at b=128 — expect 0 errors, l_dueling > 0
8. ☐ Cluster smoke b=1024 5k — verify qpa unchanged from Plan A v2
9. ☐ If ✅, Phase 4.3: swap V_dq into compute_advantage_return
10. ☐ Cluster b=1024 20k full validation

View File

@@ -0,0 +1,205 @@
# Determinism Phase 1 — Diagnosis
**Date**: 2026-06-02
**Branch**: ml-alpha-adaptive-controller-floors
**Spec**: `docs/superpowers/specs/2026-06-02-determinism-foundation.md` §1
**Status**: Phase 1 SHIPPED — diagnosis complete, fix is Phase 2 (separate dispatch).
## TL;DR
`scripts/determinism-check.sh --quick` (200 train + 50 eval at b=128, seed=42) on
`ml-alpha-adaptive-controller-floors` produces:
```
FIRST_DIVERGENCE: step=2 leaf=checksums.replay_sample_indices
run_A=709847.0 run_B=2851599.0 delta=2.14175e+06
```
**Culprit**: PER (Prioritized Experience Replay) sampling. Spec §2.F is the fix.
Steps 0-1 are bit-equivalent across runs. Divergence appears at step 2 the
moment the replay buffer transitions from "trivially deterministic" (small,
all entries equally probable) into "PER-sampled with priority weights".
## Phase 1 deliverables
1. **Kernel**`crates/ml-alpha/cuda/rl_deterministic_checksum.cu`
- Two entry points: `rl_deterministic_checksum_f32` + `rl_deterministic_checksum_i32`
- Both are `(1,1,1)/(1,1,1)` launches — a single thread sums `data[i]²`
into a double-precision accumulator. NO atomics, NO multi-block reduce,
NO non-deterministic patterns by construction. Per `feedback_no_atomicadd`.
- Cost per launch: ~5-20 μs launch overhead + sequential f64 accumulation
(~1ns per element on RTX 3050). Largest tensor we checksum is the
Q-head weight `ss_q_grad_w_d` at ~24k floats — 24 μs sequential.
- Registered in `crates/ml-alpha/build.rs` (AOT cubin compile per
`feedback_no_nvrtc`).
2. **Diag emission**`crates/ml-alpha/src/trainer/integrated.rs`
- New method `IntegratedTrainer::compute_diag_checksums(&mut self) → Result<[f64; 15]>`
- Allocates a 28-f64 scratch buffer (`checksum_scratch_d`) once at
trainer construction; 15 final-leaf slots + 12 Adam staging slots
(the optimizer's m and v live in six separate sub-buffers for the
DQN/π/V w+b parameter groups; host-folded into adam_m_sum / adam_v_sum
after the dtoh).
- `build_diag_value` is now `&mut self`; it calls `compute_diag_checksums`
up-front and inserts the 15 leaves under a new top-level `checksums`
JSON object.
- 15 leaves emitted per spec §1.2: `encoder_output`, `q_logits`,
`pi_logits`, `v_value`, `replay_sample_indices`, `rewards_after_shape`,
`advantages`, `q_grad`, `pi_grad`, `v_grad`, `encoder_grad`,
`adam_m_sum`, `adam_v_sum`, `isv_state`, `popart_sigma_state`.
- `tests/eval_diag_emission.rs` `EXPECTED_LEAVES` bumped 679 → 694.
- Call-site fix in `crates/ml-alpha/examples/alpha_rl_train.rs`: the
train-loop `let isv = trainer.isv_host_slice()` borrow conflicted with
the new `&mut self` on `build_diag_value` because the slice was
referenced past the call in an `eprintln!`. Fix is local — pre-extract
the 4 ISV values into `f32` locals before the diag call. Eval loop
was unaffected (NLL drops the borrow before the diag call there).
3. **Diagnostic**`scripts/determinism-check.sh` rewritten
- Iterates over EVERY row of `diag.jsonl` (was: last 5 rows)
- Compares only the `checksums.*` subtree leaf-by-leaf (other diag
fields are derived quantities; localising to component checksums
gives a cleaner signal)
- Reports first `(step, leaf)` divergence at rel-tol=1e-5 + abs-tol=1e-7
- Also lists ALL leaves diverging at the first divergent step (multiple
components frequently co-diverge — the script ranks them by Δ
magnitude so the top entry is the most-likely root cause)
- Maps each known leaf to its spec §2 fix recommendation
- Exit 0 = deterministic / 1 = drift detected / 2 = inputs missing
## Diagnostic findings — step-by-step
Output of `./scripts/determinism-check.sh --quick`:
| Step | Diverging leaves | Top Δ |
|------|------------------|-------|
| 0 | 0 (bit-equivalent) | — |
| 1 | 0 (bit-equivalent) | — |
| 2 | 5 simultaneously: **`replay_sample_indices`**, `q_grad`, `adam_m_sum`, `q_logits`, `adam_v_sum` | 2.14e6 on `replay_sample_indices` |
| 3 | 11 (drift cascades through forward + grads + ISV) | — |
| 4-199 | 10-15 leaves drift on every step | — |
Total divergent rows: 198 / 200.
## Interpretation
The cascade ordering is structurally clean:
1. **Step 0-1**: PER buffer has not accumulated enough transitions for
priority-weighted sampling to actually choose different indices (or
the priorities are still degenerate). Both runs produce identical
indices, identical sampled `h_t`/`h_tp1`, identical Q forward, identical
gradients, identical Adam state. **Bit-equivalent**.
2. **Step 2 — divergence enters via `replay_sample_indices`**. The two
runs choose different replay transitions to sample. From this moment
the Q gradient (which depends on the sampled batch) diverges, which
immediately diverges Q-head Adam m+v, which on the same step diverges
`q_logits` (via the immediately-following Q forward in the next step's
prefill). The 5-simultaneous divergence at step 2 is consistent with
a single root cause (PER sampling) propagating to all downstream
buffers in the same training step.
3. **Step 3+**: now divergent gradients have been Adam-applied to
parameters → encoder forward diverges (`encoder_output`), π and V
forward diverge, advantages diverge, controllers' EMA inputs diverge,
ISV state diverges. Cascade complete.
The fact that **`replay_sample_indices` is the FIRST and LARGEST-magnitude
divergence at step 2**, while every other field is bit-equivalent up
through step 1, is decisive: the bug is in the PER sampling kernel chain.
## Candidate root causes (Phase 2 investigation scope)
Spec §2.F lists the entry points to audit:
1. **`rl_per_sample` kernel** — stratified proportional sampling via
sum-tree walk + gather. The walk uses a device-side xorshift32 PRNG
(`gpu_replay.sample_prng_d`). Audit whether:
- The PRNG state is seeded deterministically and not reset between
runs (must use `cli.seed.wrapping_add(...)` per the existing CPU-side
audit in items 1-5 impl log).
- The kernel's launch configuration (Grid=B) reads state in a
deterministic order (each block owns one batch slot → fixed).
- The sum-tree walk converges deterministically — IEEE-754 sum-tree
comparisons can produce different paths if the tree values were
accumulated in a non-deterministic order (PER push_flush uses
prefix-sum allocation, which is non-deterministic on parallel
warps if the input order varies).
2. **`rl_per_push_flush` and `rl_per_tree_rebuild`** — coordinated
coalesced replay writes + bottom-up parallel sum-tree rebuild. Both
"have no atomics" per their build.rs comments, but the rebuild's
bottom-up parallelism may produce different floating-point sums
if the tree leaf order differs across runs. Audit: does the tree
rebuild deterministically reduce children before parents?
3. **`rl_per_update_priority`** — writes `|TD|^α` to tree leaves + block-
wide max reduction. The block-wide max should be deterministic, but
any precision-sensitive aggregation could be the culprit. Less likely.
4. **Priority update sequencing**`pearl_rl_sample_tau_xorshift32` notes
the PRNG seeding pattern. Verify the seed reaches `sample_prng_d` AND
is not perturbed by any unrelated state at construction time.
## Recommendation for Phase 2
Apply **spec §2.F** (PER sampling determinism audit). Concrete steps:
1. Add per-batch dumps of `sample_prng_d`, `priority_tree_d`, and
`sample_indices_d` at step 0, 1, 2 of two same-seed runs. Compare
PRNG state at step 1's exit / step 2's entry — if PRNG already
diverges → fix the seeding; if PRNG matches but indices differ →
the priority-tree walk is the culprit.
2. If priority tree is suspect, audit `rl_per_tree_rebuild` for any
parallel reduction whose addition order depends on warp scheduling.
The "no atomics" claim doesn't preclude order-of-addition variance
in a parallel scan.
3. Apply the deterministic-by-construction fix (single-block reduction
for small trees, or fixed-warp ordering for large trees).
4. Re-run `scripts/determinism-check.sh --quick` after the fix. Success
= "DETERMINISTIC: all `checksums.*` leaves match across all 200 rows".
Once PER is fixed, the next divergence (if any) localises automatically.
Phase 1 infrastructure is reusable — Phase 2 just runs the same script.
## Files touched (Phase 1, no commits per dispatch instruction)
| Path | Change |
|------|--------|
| `crates/ml-alpha/cuda/rl_deterministic_checksum.cu` (NEW) | Sum-of-squares kernel, two entry points (f32 + i32) |
| `crates/ml-alpha/build.rs` | Registered the new kernel in `KERNELS` |
| `crates/ml-alpha/src/trainer/optim.rs` | Added immutable `m()` / `v()` accessors on `AdamW` |
| `crates/ml-alpha/src/trainer/integrated.rs` | New cubin const + struct fields, module loading in `new()`, `compute_diag_checksums()` method, `build_diag_value` switched to `&mut self` + emits `checksums` JSON object |
| `crates/ml-alpha/examples/alpha_rl_train.rs` | Pre-extracted ISV values into locals to release the `isv_host_slice()` borrow before calling `build_diag_value` (now `&mut self`) |
| `crates/ml-alpha/tests/eval_diag_emission.rs` | `EXPECTED_LEAVES` 679 → 694 |
| `scripts/determinism-check.sh` | Rewritten to compare all rows × `checksums.*` leaves; reports FIRST_DIVERGENCE + spec §2 fix recommendation |
| `docs/superpowers/notes/2026-06-02-determinism-phase1-diagnosis.md` (NEW, this file) | Diagnosis log |
## Discipline checklist
- [x] No new files in repo root
- [x] No new docs outside `docs/superpowers/notes/`
- [x] CUDA kernel registered in `build.rs` per `feedback_no_nvrtc`
- [x] Kernel uses NO atomics per `feedback_no_atomicadd`
- [x] Result read from CPU only; computation entirely on device per `feedback_cpu_is_read_only`
- [x] All 15 checksums end-to-end wired — no stubs per `feedback_no_stubs`
- [x] Local RTX 3050 smoke validation only — NO cluster submission
- [x] Read-before-Edit on all touched existing files
- [x] No fix applied — diagnosis only, per dispatch instruction (Phase 2 is separate)
- [x] No commit — left uncommitted for human review
## Reproduction
```bash
# Build
SQLX_OFFLINE=true cargo build --release --example alpha_rl_train -p ml-alpha
# Run
./scripts/determinism-check.sh --quick
# Expected output (current branch ml-alpha-adaptive-controller-floors):
# FIRST_DIVERGENCE: step=2 leaf=checksums.replay_sample_indices
# run_A=709847.0 run_B=2851599.0 delta=2.14175e+06
```

View File

@@ -0,0 +1,204 @@
# Determinism Phase 2 — PER Sub-investigation + Tree-Rebuild Fix
**Date**: 2026-06-02
**Branch**: ml-alpha-regime-observer
**Spec**: `docs/superpowers/specs/2026-06-02-determinism-foundation.md` §2.F
**Plan**: `docs/superpowers/plans/2026-06-02-determinism-foundation-implementation.md` §2.F
**Phase 1 input**: `docs/superpowers/notes/2026-06-02-determinism-phase1-diagnosis.md`
## TL;DR
Phase 1 identified `checksums.replay_sample_indices` as the first
divergent leaf at step 2 with Δ=2.14e6 on the sum-of-squares. Because
the checksum is permutation-invariant, that result alone could not
distinguish between three candidates: (A) PRNG seeding drift, (B)
tree-walk fp comparison flip, or (C) tree-rebuild non-determinism.
This sub-investigation added env-gated raw-state dumps for
`sample_prng_d`, `priority_tree_d`, and `sample_indices_d` at steps
0/1/2, then ran the determinism check twice with the dumps enabled
and exact-bytewise-compared the resulting binaries.
**Finding — HYPOTHESIS C wins**: at step 2, `prng` matches
bit-exactly, `sample_indices` matches bit-exactly (steps 0/1/2 all
match), but `priority_tree` **DIVERGES at idx 1 (the root)**:
```
step 2:
prng: EQUAL (128 elements)
indices: EQUAL (128 elements)
tree: DIVERGE at idx 1
A=592.7081298828125 B=695.7353515625
```
This is the root of the priority sum-tree. A divergence of ~100 at
the root cannot come from sub-eps fp noise: it is a genuine
order-of-addition difference in the tree rebuild.
## Cause — `__threadfence()` is not a grid-wide barrier
`crates/ml-alpha/cuda/rl_per_tree_rebuild.cu` previously launched as
`Grid=(128), Block=(256)` and synchronised between levels with
`__threadfence()`. `__threadfence()` is a memory-ordering primitive
(it orders THIS thread's prior writes globally), NOT a grid-wide
barrier (it does NOT wait for OTHER blocks' writes). Once any block
finished its level-0 writes and proceeded to level 1, it could read
INTER-MEDIATE values of level-0 nodes that other blocks were still
writing. The resulting addition order varied across runs even at the
same seed, producing different `priority_tree[1]` (root) values.
The kernel did NOT use `atomicAdd`, but determinism requires more
than that — it requires a **fixed addition order** across runs. The
parallel grid-stride loop with `__threadfence()` did not provide it.
## Fix — Option C: single-block deterministic rebuild
Files changed (uncommitted):
- `crates/ml-alpha/cuda/rl_per_tree_rebuild.cu` (~50 LOC, full rewrite
of the kernel body): launch geometry collapsed to `Grid=(1)`,
`Block=(1024)`; inter-level barrier switched from `__threadfence()`
to `__syncthreads()` (a proper intra-block barrier). Each thread
processes nodes via a block-stride loop with a FIXED thread→node
mapping (`i = tid + k*blockDim.x`), so every same-seed run uses
identical addition orderings. Header comment documents the
determinism rationale, the speed analysis, and the upgrade path to
`cooperative_groups::grid().sync()` if capacity ever grows past 2^21.
- `crates/ml-alpha/src/trainer/integrated.rs` (~8 LOC in the launch
site at the K-loop body): updated the `raw_launch` config from
`(128,1,1)/(256,1,1)` to `(1,1,1)/(1024,1,1)` with an inline
comment pointing at the kernel header for the determinism rationale.
Diagnostic infrastructure (new files):
- `crates/ml-alpha/src/trainer/integrated.rs`
`IntegratedTrainer::dump_per_state_for_debug` — env-gated method
triggered from `build_diag_value` at steps 0/1/2. Reads
`FOXHUNT_DETERMINISM_DEBUG=1` to enable, dumps to
`$FOXHUNT_DEBUG_DUMP_DIR` (default `/tmp/foxhunt-determinism-debug`).
No-op in normal runs.
- `scripts/determinism-check.sh` `--debug-dump` flag: runs both A/B
with the env vars set, then invokes `compare-per-state.py`.
- `scripts/compare-per-state.py` (NEW, 200 LOC): reads the raw dumps,
exact-byte-compares prng / tree / indices, and prints a verdict
identifying which hypothesis (A / B / C) is in play.
## Validation
### Task 3 — `scripts/determinism-check.sh --quick`
Post-fix output (full 200-step quick run):
```
FIRST_DIVERGENCE: step=3 leaf=checksums.encoder_output
run_A=6896.334218005304 run_B=6896.194631539084
delta=0.139586
All leaves diverging at step=3:
checksums.encoder_output A=6896.334... B=6896.194... Δ=0.139586
checksums.v_value A=0.23311... B=0.23318... Δ=6.59e-05
checksums.v_grad A=0.09886... B=0.09888... Δ=1.74e-05
```
PER is **no longer the first-divergent component**. The next bug
surfaces at step 3 in `encoder_output` (mamba2 selective scan, per
spec §2.A). Per dispatch instructions: documenting and stopping —
this is a separate Phase 2.2 dispatch.
The debug-dump run also confirms PER state is bit-equal at steps
0/1/2:
```
step 2:
prng: EQUAL (128 elements)
indices: EQUAL (128 elements)
tree: EQUAL (65536 elements)
VERDICT: all three buffers match across all dumped steps.
```
### Speed cost
Quick mid-smoke timing on RTX 3050 (200 train + 50 eval, b=128):
| Phase | Run A | Run B |
|-------|-------|-------|
| Pre-fix (Grid=128, Block=256, __threadfence) | 0m 45s | 0m 44s |
| Post-fix (Grid=1, Block=1024, __syncthreads) | 0m 43s | 0m 43s |
**No measurable slowdown.** The single-block rebuild touches ~65k
floats (15 tree levels), each thread doing ~64 ops, completing in a
few hundred μs. The kernel launch overhead (one launch per K-loop
iter) is unchanged. If anything, the single-block variant slightly
reduces launch-config overhead.
If the post-fix run is faster, the difference is within run-to-run
variance. Confirmed safe to land.
### Acceptance criteria from §2.F
- [x] Phase 2.F.1 — PER RNG seeding check: confirmed deterministic
(prng-equal dumps).
- [x] Phase 2.F.2 — Fix applied (tree-rebuild kernel + launch site).
- [x] Phase 2.F.3 — `determinism-check.sh` no longer reports
`replay_sample_indices` as the first divergent leaf. The next
first-divergent leaf is `encoder_output` (separate Phase 2.A
issue, mamba2 selective scan).
## Files touched (uncommitted)
| Path | Change | LOC |
|------|--------|-----|
| `crates/ml-alpha/cuda/rl_per_tree_rebuild.cu` | Single-block deterministic rebuild | +60 / -34 |
| `crates/ml-alpha/src/trainer/integrated.rs` | (a) launch geometry change (b) `dump_per_state_for_debug` method (c) call from `build_diag_value` | +120 / -2 |
| `scripts/determinism-check.sh` | `--debug-dump` flag + env-var plumbing | +35 / -5 |
| `scripts/compare-per-state.py` (NEW) | Diff PRNG / tree / indices, hypothesis verdict | +200 |
| `docs/superpowers/notes/2026-06-02-determinism-phase2-per-investigation.md` (NEW) | this note | new |
## Discipline checklist
- [x] No new files in repo root
- [x] No new docs outside `docs/superpowers/notes/`
- [x] Per `feedback_no_atomicadd`: kernel uses no atomics (block-wide
barrier + fixed-order writes).
- [x] Per `feedback_no_nvrtc`: kernel registered in `build.rs` (no
change required — already AOT compiled, just the source body
changed).
- [x] Per `feedback_cpu_is_read_only`: the debug-dump method does a
DtoH copy after a stream sync; this is the standard host-read
pattern, no host-side compute.
- [x] Per `feedback_no_stubs`: dump is wired end-to-end, runs only
when env-gate is set, otherwise zero overhead.
- [x] Per `feedback_local_smoke_before_cluster`: validated on RTX
3050 with the quick determinism check. No cluster submission.
- [x] Per `feedback_no_partial_refactor`: there is only ONE launch
site for `rl_per_tree_rebuild` (verified via grep); updated in
lockstep with the kernel.
- [x] Per `feedback_no_quickfixes`: chose the principled fix
(deterministic-by-construction single-block) over a cheap epsilon
shift; this is the same pattern the spec §1.1 already
recommends for the deterministic checksum kernel.
- [x] No commit — left uncommitted for human review per dispatch.
## Next step — Phase 2.2 dispatch for `encoder_output`
The `encoder_output` divergence at step 3 (Δ=0.14 on a sum-of-squares
of ~6896) corresponds to mamba2 selective scan, per spec §2.A. This
is a separate sub-investigation; the Phase 1 infrastructure (15-leaf
checksums + scripts/determinism-check.sh) already supports localising
it. Recommend dispatching a Phase 2.2 worker on §2.A.
## Reproduction
```bash
# Build (idempotent, sccache helps)
SQLX_OFFLINE=true cargo build --release --example alpha_rl_train -p ml-alpha
# Validate the fix (no debug dumps)
./scripts/determinism-check.sh --quick
# Re-run the sub-investigation (env-gated dumps at steps 0/1/2)
./scripts/determinism-check.sh --quick --debug-dump
```

View File

@@ -0,0 +1,654 @@
# Determinism Phase 2.2 — Mamba2 Sub-Investigation
**Date**: 2026-06-02
**Branch**: ml-alpha-regime-observer
**Spec**: `docs/superpowers/specs/2026-06-02-determinism-foundation.md` §2.A (hypothesised) / §2.B (actual)
**Plan**: `docs/superpowers/plans/2026-06-02-determinism-foundation-implementation.md`
**Phase 2 input**: `docs/superpowers/notes/2026-06-02-determinism-phase2-per-investigation.md`
## TL;DR — STOP, the bug is NOT mamba2; it is cuBLAS GEMM non-determinism (§2.B)
Phase 2 fixed the PER tree-rebuild and surfaced
`checksums.encoder_output` at step 3 (Δ=0.14) as the new
first-divergent leaf, with §2.A (mamba2 selective scan) listed
as the most-likely culprit per the spec lookup.
Phase 2.2 added 17 encoder-localisation checksums plus an
env-gated raw-state dumper. The dumps tell a different and
**decisive** story:
1. At **step 2** (one step earlier than `encoder_output`), the
first divergent leaf is **`checksums.attn_q`** — but the raw
mamba2 dumps show that ALL encoder weights diverge
simultaneously at step 2: vsn_w, cfc_w_in, cfc_w_rec, attn_q,
mamba2 L1 w_in/b_in/w_a/b_a/w_b/b_b/w_c, mamba2 L2 same
set — but **mamba2 w_out / b_out STAY EQUAL** (those weights
aren't touched by the seq backward path).
2. Step 1's `encoder_grad` (= `grad_h_t_combined_d`) checksum
**MATCHES bit-exactly** between runs A and B at step 1's diag
emission (1.3292713308910526e-10 in both). All per-head Q / π /
V grads at step 1 also match.
3. So the encoder backward at step 1 sees IDENTICAL inputs in
both runs, yet produces DIVERGENT gradient updates that diverge
ALL trainable encoder weights simultaneously at step 2.
4. The only operation in the encoder backward chain that runs
**before** the first non-deterministic weight (mamba2 L2) but
AFTER the deterministic upstream chain is the mamba2 backward
itself — specifically, the **cuBLAS GEMM calls inside
`Mamba2Block::backward_from_h_enriched_seq_full_into`** at
`crates/ml-alpha/src/mamba2_block.rs:1918-1944` (3 GEMMs:
w_b.backward, w_a.backward, w_in.backward).
5. The cuBLAS path uses `cublasGemmEx` with algo
`CUBLAS_GEMM_DFALT` (see
`crates/ml-core/src/cuda_autograd/linear.rs:88`). Per NVIDIA
docs, default GEMM algorithms are NOT guaranteed deterministic
— split-K with non-deterministic accumulation order is
permitted. No call to `cublasSetMathMode(CUBLAS_PEDANTIC_MATH)`
exists anywhere in the codebase
(`grep -rn "CublasMathMode\|setMathMode" crates/` returns
nothing).
**Conclusion — this is spec §2.B, not §2.A.** The mamba2 kernels
themselves are bit-deterministic by construction (one thread per
(i, j), unique slot writes, no atomics; reductions iterate a
single thread over a fixed range). The non-determinism enters
through the **cuBLAS GEMMs** that surround the scan kernel
(W_in, W_a, W_b forward + backward dw/db/dx).
Per dispatch instruction "If Task 1 finds something completely
unexpected (e.g., the bug is in the mamba2_block.rs Rust glue
rather than the CUDA kernel), document and STOP — separate
dispatch": this finding lands in the Rust-glue + cuBLAS-handle
category. Fix is a different scope (every cuBLAS handle in the
trainer must enter PEDANTIC mode, not a mamba2-local edit).
## Investigation infrastructure
Three artefacts were added (uncommitted) to localise the bug:
### 1. 17 new encoder-localisation checksum leaves
Slot indices 28-44 in `IntegratedTrainer::checksum_scratch_d`, output
indices 15-31 in `compute_diag_checksums`. Wired into the existing
`build_diag_value` JSON `checksums` map:
| idx | leaf | source tensor |
|-----|------|---------------|
| 15 | `mamba2_l1_w_in` | `perception.trunk.mamba2_l1().w_in.weight` |
| 16 | `mamba2_l1_w_a` | `mamba2_l1.w_a.weight` |
| 17 | `mamba2_l1_w_b` | `mamba2_l1.w_b.weight` |
| 18 | `mamba2_l1_w_c` | `mamba2_l1.w_c` |
| 19 | `mamba2_l2_w_in` | `mamba2_l2.w_in.weight` |
| 20 | `mamba2_l2_w_c` | `mamba2_l2.w_c` |
| 21 | `vsn_w` | `trunk.vsn_w_d` |
| 22 | `cfc_w_in` | `trunk.w_in_d` |
| 23 | `cfc_w_rec` | `trunk.w_rec_d` |
| 24 | `attn_q` | `trunk.attn_q_d` |
| 25 | `vsn_out` | encoder intermediate activation |
| 26 | `mamba2_l1_out` | `h_enriched_seq` L1 |
| 27 | `ln_a_out` | LN_a output |
| 28 | `mamba2_l2_out` | `h_enriched_seq` L2 |
| 29 | `ln_b_out` | LN_b output |
| 30 | `attn_context` | attention pool context |
| 31 | `mamba2_grad_w_c_l1` | reduced backward `dw_c` for L1 |
`EXPECTED_LEAVES` in `eval_diag_emission.rs` bumped 694 → 711.
### 2. Env-gated raw-state dumper
`PerceptionTrainer::dump_mamba2_state_for_debug(step)` triggered
from `IntegratedTrainer::build_diag_value` at steps 0/1/2/3 when
`FOXHUNT_DETERMINISM_DEBUG_MAMBA2=1`. Writes 26 raw f32 binaries
per step (18 weight tensors + 7 activations + 1 reduced grad) to
`$FOXHUNT_DEBUG_DUMP_DIR` (default `/tmp/foxhunt-determinism-debug`).
Mirrors the pattern in Phase 2's `dump_per_state_for_debug`.
### 3. `scripts/compare-mamba2-state.py`
Exact-bytewise diff with verdict logic for sub-cases 2.A.1
through 2.A.5 (weights vs intermediates vs reduced grads).
### 4. `scripts/determinism-check.sh --debug-dump-mamba2`
New flag plus `mamba2_l1_w_in` / `vsn_out` / etc. entries in the
per-leaf "fix recommendation" table.
## Empirical findings — checksums
`./scripts/determinism-check.sh --quick` after Phase 2.2 wiring:
```
FIRST_DIVERGENCE: step=2 leaf=checksums.attn_q
run_A=0.31439592314009085 run_B=0.31440285748760444
delta=6.93435e-06
All leaves diverging at step=2:
checksums.attn_q Δ=6.93e-06
checksums.mamba2_grad_w_c_l1 Δ=3.57e-07
```
Following the trace via JSONL inspection:
| Step | run_A vsn_w | run_B vsn_w | run_A mamba2_l1_w_in | run_B mamba2_l1_w_in |
|------|-------------|-------------|----------------------|----------------------|
| 0 | 18.66958 | 18.66958 | 79.54464 | 79.54464 |
| 1 | 18.66958 | 18.66958 | 79.54464 | 79.54464 |
| 2 | **18.75015** | **18.75025** | **79.57704** | **79.57719** |
| 3 | 18.86302 | 18.86328 | 79.59899 | 79.59909 |
Steps 0 and 1 are bit-identical between runs (l_q = l_pi = l_v = 0
at step 0; step 1 fires the first non-zero backward). At step 2's
diag, EVERY encoder weight has diverged.
`encoder_grad` (= `grad_h_t_combined_d`) at step 1:
- run A: 1.3292713308910526e-10
- run B: 1.3292713308910526e-10 (MATCH)
`q_grad`, `pi_grad`, `v_grad` at step 2 also all MATCH between
runs (sum-of-squares 3636.114, 0.0003699, 4.770977 in both).
## Empirical findings — raw dumps
`./scripts/determinism-check.sh --quick --debug-dump-mamba2`:
```
=== step 0 ===
ALL 26 files: EQUAL
=== step 1 ===
ALL 26 files: EQUAL
=== step 2 ===
mamba2_l1_w_in: DIVERGE at idx 0
mamba2_l1_b_in: DIVERGE at idx 0
mamba2_l1_w_a: DIVERGE
mamba2_l1_b_a: DIVERGE
mamba2_l1_w_b: DIVERGE
mamba2_l1_b_b: DIVERGE
mamba2_l1_w_c: DIVERGE
mamba2_l1_w_out: EQUAL
mamba2_l1_b_out: EQUAL
(same pattern for L2)
vsn_out: EQUAL (forward output at step 2)
mamba2_l1_h_enriched_seq: EQUAL
ln_a_out: EQUAL
mamba2_l2_h_enriched_seq: EQUAL
ln_b_out: EQUAL
attn_context: EQUAL
h_t: EQUAL
=== step 3 ===
ALL diverge (cascade)
```
**This is decisive sub-case 2.A.4 (weight drift) — but with the
critical refinement that `w_out` / `b_out` STAY EQUAL.** Per
`crates/ml-alpha/src/mamba2_block.rs:305`:
> "W_out is unused in the seq path — dw_out / db_out are zero-init
> shells of the right shape so the Mamba2AdamW step is a no-op for
> those parameters."
The diverging weights are exactly the set of params with non-zero
gradients in the seq path. Their gradients flow through
`backward_from_h_enriched_seq_full_into` which uses three
cuBLAS GEMMs (W_b backward, W_a backward, W_in backward).
The cuBLAS handle is initialised via `CudaBlas::new(stream)` and
never has `cublasSetMathMode` called on it. Per NVIDIA docs:
- `CUBLAS_DEFAULT_MATH` (and `CUBLAS_GEMM_DFALT` algorithm) allows
TF32 + split-K + heuristic-selected variants on Ampere+.
- Split-K with atomic accumulation across blocks produces
**non-deterministic** floating-point sums across runs even with
identical inputs and identical algorithm selection.
- The only way to disable this is `cublasSetMathMode(handle,
CUBLAS_PEDANTIC_MATH)` PLUS explicit algorithm selection (or
rely on `CUBLAS_PEDANTIC_MATH` to filter out non-deterministic
algos).
## Why the OUTPUTS at step 2 match but the WEIGHTS diverge
The dump shows at step 2:
- `mamba2_l1_h_enriched_seq`: **EQUAL** (forward output bit-exact)
- `mamba2_l1_w_in`: **DIVERGE** (weights bit-different)
This is consistent with: at step 2's diag emission, the encoder
forward has run with the POST-Adam weights from step 1's backward.
But our dumps for the activations show pre-step-2 state — i.e.,
the forward dump was taken AFTER the diag forward fired but the
ACTIVATIONS were captured into stable scratch buffers that haven't
been re-overwritten by step 2's backward yet.
Wait — actually the dump fires inside `build_diag_value` AFTER the
step has completed. Step 2's full step has run: forward (with
post-step-1 weights, so the forward used DIVERGED weights) +
backward (writing new grads). Then the dump captures BOTH the
current weights AND the activations.
So either:
(a) The forward at step 2 used divergent weights but happened to
produce equal h_enriched_seq within fp32 noise floor, or
(b) The forward output match is genuine (the divergence at step 2
in mamba2 weights is so tiny — Δ ~1e-5 — that the forward
output difference is below the exact-bytewise threshold of
the dump, which compares bit-equality).
Spot-check at run A vs B w_in[0]: 0.17429381608963013 vs
0.17429661750793457 → diff ~2.8e-6. Multiplied through a 128-wide
GEMM with input magnitudes ~1, the output deviation is on the
order of 1e-6 * sqrt(128) ~ 1e-5. This IS below f32 epsilon-level
in some sums but should still show up as bit-different in most.
So (a) is more likely — the propagated forward errors at step 2
ARE below fp32 last-bit thresholds for SOME tensors but should
manifest at step 3 (which they do — all activations DIVERGE at
step 3).
This is consistent with the cuBLAS GEMM hypothesis: the BACKWARD
GEMM is split-K non-deterministic (running over a B*K reduction
dimension where bits can shift across runs), but the forward GEMM
uses different shapes where the same algo IS deterministic in
practice (no split-K because the reduction dim is small enough).
## What needs to happen next — Phase 2.3 dispatch
**Per dispatch instruction: STOP.** The fix is not a single-file
mamba2 edit; it is a cross-cutting cuBLAS handle reconfiguration
spec §2.B requires:
1. Every cuBLAS handle in the trainer must set
`cublasSetMathMode(CUBLAS_PEDANTIC_MATH)` at construction.
At minimum: `Mamba2Block.cublas`,
`dqn.cublas`, `iqn.cublas` (per the grep at the start of this
investigation). Possibly others.
2. Document the performance cost — NVIDIA quotes ~10-15% on H100
for general GEMMs. Acceptable per spec §4 (`FOXHUNT_DETERMINISTIC`
env toggle).
3. Re-run `./scripts/determinism-check.sh --quick` on seeds
42 / 43 / 44 — if the fix is correct, all three should pass
exit-0 with no first-divergence reported.
4. If a residual divergence surfaces in some other layer (e.g.,
a custom kernel we haven't yet checksummed), Phase 2.3 dumps
will pinpoint it.
## Out of scope for this dispatch
Per `feedback_extending_existing_code_audits_for_existing_violations`,
flagging two pre-existing items observed during this audit but
NOT fixing (out of scope):
1. **`CUBLAS_GEMM_DFALT` comment is misleading**:
`crates/ml-core/src/cuda_autograd/linear.rs:82-88` calls
the algorithm "Deterministic default algo" but per NVIDIA
docs `CUBLAS_GEMM_DFALT` is the heuristic-selected default
and is NOT guaranteed deterministic. Comment should be
corrected when §2.B is applied.
2. **`Mamba2BlockForwardScratch` exposes only `pub`-fields** —
public access by the perception trainer dispatch is via
`mamba2_fwd_scratch.h_enriched_seq.cuda_data()`. When the
dump method here borrows it via a new public accessor, it
crosses module boundaries that were previously informal. No
structural issue; just noted.
## Speed cost (post-instrumentation)
Tier 1.5 mid-smoke (200 train + 50 eval, b=128) on RTX 3050:
| Version | Run A | Run B |
|---------|-------|-------|
| Phase 2 (15 checksums) | 0m 43s | 0m 43s |
| Phase 2.2 (32 checksums) | 0m 47s | 0m 46s |
The +4s per 250-step run is ~16 ms/step extra, consistent with
17 additional 1-thread checksum kernels each ~1 ms on the largest
tensor (524288 f32s = 524288 ns at ~1 GFLOP/s sequential f64).
Negligible vs the 200-ms/step training step.
## Files touched (Phase 2.2, no commits per dispatch)
| Path | Change | LOC |
|------|--------|-----|
| `crates/ml-alpha/src/trainer/perception.rs` | 7 new `debug_*_view` accessors + `dump_mamba2_state_for_debug` method | +130 |
| `crates/ml-alpha/src/trainer/integrated.rs` | (a) scratch+host buffer 28→45 slots (b) 17 new checksum launches (c) returns `[f64; 32]` (d) JSON emission +17 leaves (e) `dump_mamba2_state_for_debug` call from `build_diag_value` | +90 / -10 |
| `crates/ml-alpha/tests/eval_diag_emission.rs` | `EXPECTED_LEAVES` 694 → 711 | +5 |
| `scripts/determinism-check.sh` | `--debug-dump-mamba2` flag + env plumbing + 17 new candidate-leaf entries | +30 / -5 |
| `scripts/compare-mamba2-state.py` (NEW) | Mamba2 dump comparison + sub-case verdict | +240 |
| `docs/superpowers/notes/2026-06-02-determinism-phase2.2-mamba2-investigation.md` (NEW) | this note | new |
## Discipline checklist
- [x] No new files in repo root
- [x] No new docs outside `docs/superpowers/notes/`
- [x] No `.cu` files added/changed (the bug is NOT in mamba2 kernels)
- [x] Per `feedback_no_atomicadd.md`: not relevant (this is cuBLAS)
- [x] Per `feedback_no_nvrtc.md`: not relevant
- [x] Per `feedback_cpu_is_read_only.md`: the dump method only does
DtoH copies for off-line analysis, no host compute
- [x] Per `feedback_no_stubs.md`: dumps + checksums wired end-to-end
- [x] Per `feedback_local_smoke_before_cluster.md`: validated on
RTX 3050 with the quick determinism check; no cluster submit
- [x] No fix applied — diagnosis only, per dispatch STOP instruction
- [x] No commit — left uncommitted for human review
## Reproduction
```bash
# Build (idempotent, sccache helps)
SQLX_OFFLINE=true cargo build --release --example alpha_rl_train -p ml-alpha
# Run the quick determinism check WITH 17 new encoder leaves
./scripts/determinism-check.sh --quick
# Output:
# FIRST_DIVERGENCE: step=2 leaf=checksums.attn_q
# delta=6.93e-06
# Run again with raw mamba2 dumps for the sub-case verdict
./scripts/determinism-check.sh --quick --debug-dump-mamba2
# Output (key bit):
# VERDICT (first divergent step = 2):
# CASE 2.A.4 — WEIGHTS DIVERGE upstream of forward.
# Mamba2 L1/L2 weights diverged at step 2.
```
## Recommended next-dispatch ask (Phase 2.3)
**One-liner ask**: Apply spec §2.B (cuBLAS PEDANTIC_MATH) to every
`CudaBlas::new(...)` site in `crates/ml-alpha/src/`. Verify
`determinism-check.sh --quick` exits 0 on seeds 42/43/44.
**Sites to patch** (3 known cuBLAS handles, grep
`"CudaBlas::new"`):
- `crates/ml-alpha/src/mamba2_block.rs:535` (Mamba2Block.cublas)
- `crates/ml-alpha/src/rl/iqn.rs:261` (IqnHead.cublas)
- `crates/ml-alpha/src/rl/dqn.rs:364` (DqnHead.cublas)
Required call: immediately after `CudaBlas::new`, call
```rust
unsafe {
cudarc::cublas::sys::cublasSetMathMode(
*cublas.handle(),
cudarc::cublas::sys::cublasMath_t::CUBLAS_PEDANTIC_MATH,
).result()
.map_err(|e| anyhow!("cublasSetMathMode PEDANTIC: {e:?}"))?;
}
```
Plus an env toggle `FOXHUNT_DETERMINISTIC` per spec §4 (default
"1" in dev → PEDANTIC, "0" in production → TF32_TENSOR_OP).
Once §2.B lands, re-run the same `determinism-check.sh --quick`
infrastructure (no further new instrumentation needed) — should
go straight to exit 0 with all 32 checksums matching.
---
## Phase 2.3 outcome — cuBLAS PEDANTIC fix lands MAJOR bug; residual remains
**Date appended**: 2026-06-02 (continuation session)
**Fix shipped**: `crates/ml-alpha/src/cublas_determinism.rs` (NEW helper) + 3 site insertions at the post-`CudaBlas::new` lines in `mamba2_block.rs:535`, `rl/dqn.rs:364`, `rl/iqn.rs:261`. Env-gated by `FOXHUNT_DETERMINISTIC=1` (default ON in dev).
### Empirical post-fix state
`./scripts/determinism-check.sh --quick` after Phase 2.3:
```
FIRST_DIVERGENCE: step=2 leaf=checksums.mamba2_grad_w_c_l1
run_A=2.04e-07 run_B=9.09e-07 delta=7.05e-07
All leaves diverging at step=2:
checksums.mamba2_grad_w_c_l1 Δ=7.05e-07
```
**Comparison to pre-fix**:
| Leaf | Pre-Phase-2.3 Δ | Post-Phase-2.3 Δ | Status |
|------|-----------------|------------------|--------|
| `checksums.attn_q` | 6.93e-06 (FIRST) | (not divergent) | ✅ FIXED |
| `checksums.mamba2_grad_w_c_l1` | 3.57e-07 (cascade) | 7.05e-07 (FIRST) | ⚠️ residual exposed |
| All others (q_grad, encoder weights cascade) | divergent | not divergent | ✅ FIXED |
The PEDANTIC fix landed the major cuBLAS-driven non-determinism. The attn_q cascade is GONE. What remains is a small residual in mamba2's L1 d_w_c gradient.
### Trajectory-level impact of residual
Measured at step 197 of the 200-step quick determinism check:
| Signal | Run A | Run B | Δ |
|--------|-------|-------|---|
| train pnl | $33,064.97 | $32,674.07 | $391 |
| wr | 0.304264 | 0.303969 | 2.95e-4 |
| entropy | 1.289456 | 1.289806 | 3.5e-4 |
Eval summary (50 steps):
| Signal | Run A | Run B | Δ |
|--------|-------|-------|---|
| max_drawdown | $254,825 | $438,637 | **$184k** |
| eval trades seen | 392 | 431 | 39 |
The 7e-7 step-2 gradient checksum residual cascades to a **$184k eval max_drawdown** difference. Sub-eps gradient drift × non-linear RL dynamics × 247 steps = real trajectory divergence. **This residual must be eliminated.**
### What's been ruled out (post-PEDANTIC)
- **Mamba2 forward scan**: per-thread sequential, deterministic by construction
- **`mamba2_alpha_scan_bwd_seq`**: per-thread sequential with unique slot writes
- **`mamba2_alpha_reduce_d_proj`**: per-thread sequential over fixed j=0..sh2-1
- **`mamba2_alpha_reduce_d_w_c`**: per-thread sequential over fixed i=0..N-1
- **The 3 mamba2 cuBLAS backward GEMMs**: now PEDANTIC, attn_q proves they're deterministic
- **All foreseen mamba2 paths read on inspection**
### What's NOT ruled out (Phase 2.4 scope)
1. **PEDANTIC may not be sufficient for some GEMM shapes**: NVIDIA docs say PEDANTIC disables non-deterministic algorithms; but for the specific shapes the mamba2 backward uses (e.g., K-dim is the time axis K=32, so K-split could still occur), some variants might persist. Try also `CUBLAS_DEFAULT_MATH` (no TF32 at all) as a control.
2. **AdamW kernel**: `mamba2_alpha_adamw_step` (line 525 of build.rs reference) — uses per-thread state per parameter; should be deterministic but worth dumping.
3. **Reductions in the trunk**: vsn / cfc / attn / ln_a / ln_b have their own custom kernels we haven't dump-compared at step 1 backward; one of them may have a parallel reduce.
4. **f32 catastrophic cancellation**: at the tiny gradient magnitudes (~3e-4) where d_w_c lives, subtraction patterns in different orders can produce different fp results. May need f64 accumulation.
### Phase 2.4 dispatch ask
Reuse the existing dump infrastructure (`FOXHUNT_DETERMINISTIC_DEBUG_MAMBA2=1` + `scripts/compare-mamba2-state.py`) to capture step-1 backward outputs:
- `d_a_per_channel`, `d_b_per_channel`, `d_w_c_per_sample` (the scan_bwd_seq outputs)
- `d_a_proj`, `d_b_proj`, `d_w_c` (the reduce outputs)
- The upstream `d_h_enriched_seq` input to mamba2 L1 backward
If `d_w_c_per_sample` diverges with deterministic inputs → bug is in scan_bwd_seq (despite inspection).
If `d_h_enriched_seq` diverges → bug is upstream of mamba2 L1 (likely in some trunk reduction).
Also worth trying `CUBLAS_DEFAULT_MATH` (no TF32, no PEDANTIC) on the 3 sites as a control — see if the residual changes.
The fix is likely small (~30 LOC kernel modification or a different cuBLAS mode) but requires the dump evidence to localize.
---
## Phase 2.4 outcome — residual is NOT in mamba2 or cuBLAS; bug is upstream of mamba2 in the RL step loop
**Date appended**: 2026-06-02 (Phase 2.4 dispatch)
**Investigation**: dump-compare per dispatch Task 1, plus DEFAULT_MATH control per Task 2.
### Task 1 — dump-compare at step 1 backward outputs
Phase 2.4 extended the existing `dump_mamba2_state_for_debug` method to capture (per L1 backward; L2 selectively):
- Per-channel scan_bwd_seq outputs: `d_a_per_channel`, `d_b_per_channel`, `d_w_c_per_sample`, `d_h_s2`
- Reduce kernel outputs: `d_a_proj_2d`, `d_b_proj_2d`, `dw_c_reduced`
- cuBLAS GEMM outputs: `dw_in`, `dw_a`, `dw_b`, `db_in`, `db_a`, `db_b`, `d_x`, `d_x_from_in`
- Trunk weights (upstream): `vsn_w`, `attn_q`, `cfc_w_in`, `cfc_w_rec`, `ln_a_gain`, `ln_b_gain`
- Trunk gradients (pre-Adam): `grad_vsn_w`, `grad_attn_q`, `grad_cfc_w_in`, `grad_cfc_w_rec`, `grad_ln_a_gain`, `grad_ln_b_gain`
**Result at step 1 dump (`FOXHUNT_DETERMINISTIC=1`, PEDANTIC)**:
```
=== step 1 ===
mamba2_l1_w_in: EQUAL (7168 elements)
... [all 18 mamba2 weights] ALL EQUAL
[all 7 forward activations] ALL EQUAL
[all 5 mamba2 backward scratches] ALL EQUAL
[all 13 mamba2 backward outputs] ALL EQUAL ← INCLUDES dw_in, dw_a, dw_b, dw_c
[all 6 trunk weights] ALL EQUAL
[all 6 trunk gradients] ALL EQUAL
=== step 2 ===
mamba2_l1_w_in: DIVERGE ← divergence enters HERE
```
**Verdict**: at step 1's full post-state dump (post-forward, post-backward, post-Adam), **EVERYTHING is bit-identical** between run A and run B. This includes:
- All mamba2 backward kernels' outputs (scan_bwd_seq scratch, reduce outputs, cuBLAS GEMM outputs)
- All trunk weight gradients (vsn / attn / cfc / LN)
- All weights after Adam step 1 applied
Yet at step 2's dump, mamba2 weights diverge. **This means the bug is in something we have NOT yet captured** — between step 1's diag emission and step 2's diag emission.
### What's between step 1 dump and step 2 dump
Per `crates/ml-alpha/examples/alpha_rl_train.rs:496-714` the per-step sequence is:
1. `step_with_lobsim_gpu()` — full training step (data load, forward, backward, Adam, RL env update)
2. `diag_staging.sync_and_swap()` + `snapshot_async()`
3. `build_diag_value()` ← **where the dump fires**
So at step 2's `build_diag_value`, the state captured is **post-step-2 full cycle**.
Between step 1 dump and step 2 dump, step 2's `step_with_lobsim_gpu` runs. Its inputs are:
- Step 1's POST-Adam weights (EQUAL per dump)
- Step 2's training data batch
- Step 2's RL environment state (positions, cfc carry, ISV controllers)
- Step 2's PER replay sample
- Step 2's action sampling RNG state
The not-yet-instrumented candidates are:
1. **CfC recurrent carry** (`cfc_h_state_step_d`) — crosses steps, could diverge
2. **PER buffer state** — sample priorities, tree state (Phase 2 fixed the rebuild — but sampling itself?)
3. **Action sampling RNG / categorical sampling** at step 2
4. **LOB simulator state** (positions, pyramid, balance) — RL env evolves across steps
5. **ISV controllers** — cross-step state (popart, win_rate EMA, kelly state, etc.)
6. **Reward shaping kernels** (rl_fused_reward_pipeline)
### Task 2 — `CUBLAS_DEFAULT_MATH` control
Tested `FOXHUNT_DETERMINISTIC=2` (DEFAULT_MATH, no TF32, no PEDANTIC).
**Result**: same residual `mamba2_grad_w_c_l1` Δ ≈ 1e-7 between runs. **DEFAULT_MATH does NOT eliminate the residual**, meaning the residual is **NOT a cuBLAS algorithm-selection artefact**. PEDANTIC alone is sufficient for the cuBLAS GEMM determinism.
Side note: across multiple runs, the first-divergent leaf alternates between `attn_q`, `vsn_w`, and `mamba2_grad_w_c_l1` (the divergence magnitudes differ run-to-run from 1e-7 to 5e-5). This suggests the divergence enters from a non-deterministic source whose magnitude varies — consistent with **action sampling RNG** or **PER sampling** (random-mediated non-determinism), NOT a deterministic-but-non-bit-exact kernel.
### Why the previous "attn_q FIXED" claim was wrong
The Phase 2.3 outcome section claimed "attn_q cascade is GONE" — but my repeat runs show attn_q can still emerge as first-divergent. The reason: at b=128 the absolute checksum delta is tiny (~1e-5 to 1e-7), and the random seed of the dump infrastructure interacts with the very-small divergence to produce different first-divergent leaves on different runs. The PEDANTIC fix DOES eliminate large cuBLAS-driven divergence (post-fix Δ never exceeds 1e-4 vs pre-fix 6.93e-6 → cascade to 0.14 by step 3); it does NOT achieve bit-exactness.
### What WAS proven
- **§2.B (cuBLAS split-K non-determinism) is FALSIFIED at PEDANTIC**: confirmed PEDANTIC is being applied (added print statement to `cublas_determinism.rs`). Both PEDANTIC and DEFAULT_MATH leave the same residual.
- **mamba2 forward+backward kernels are bit-deterministic** (all dumps match at step 1 with identical inputs).
- **mamba2 reduce kernels (`reduce_d_w_c`, `reduce_d_proj`) are bit-deterministic** (dw_c_reduced matches at step 1).
- **Trunk backward chain (CfC/LN/VSN/attn) is bit-deterministic at step 1** (all grad outputs match).
### Phase 2.4 STOPS per dispatch instruction
Per dispatch: *"If Task 1 finds the bug is in a kernel I haven't suspected, document and STOP — separate dispatch."*
This applies: the bug is **NOT in any mamba2 backward kernel, NOT in any reduce kernel, NOT in any cuBLAS call, NOT in any trunk backward kernel**. The bug is in step 2's `step_with_lobsim_gpu` execution path — most likely in the RL env state evolution, PER sampling, action sampling RNG, or reward shaping kernels.
Per `feedback_systematic_debugging` + `feedback_no_quickfixes`: applying Task 3's proposed f64-accumulation fix to `mamba2_alpha_reduce_d_w_c` would be a quick-fix masking the real bug elsewhere. The reduce kernel produces identical outputs given identical inputs — its f32 precision is sufficient for the divergence we observe (Δ=1e-7).
### Phase 2.5 dispatch ask
Extend the dump infrastructure to capture step-N PRE-STEP state (before `step_with_lobsim_gpu` runs), focused on the not-yet-instrumented candidates:
1. **CfC recurrent carry** — `cfc_h_state_step_d`: dump at step 1 dump-point (= post-step-1 state, fed into step 2 forward). If divergent → cfc cross-step carry is non-det (very surprising since Phase 2 confirmed PER was fixed).
2. **PER replay state** — `replay_buffer.priorities_d`, `replay_buffer.priority_tree`, the next-sample indices for step 2. If divergent → PER sampling at step 2 produces different indices despite Phase 2 fix.
3. **Action sampling RNG state** — wherever step 2's categorical action sampling reads from a curandState. If RNG state is per-batch but shared across steps, and one batch's RNG advance depends on conditional logic, this could diverge.
4. **LOB simulator positions/balance** — `pyramid_units_count`, `unit_entry_step`, etc. — if these diverge at step 2 dump but were equal at step 1 dump, the LOB sim is the source.
5. **ISV controllers** — popart_sigma, kelly state, wr_ema. Cross-step state. If divergent at step 1 dump (despite Phase 2 fixing PER), one of these has hidden non-determinism.
Reuse the same env-gated dump pattern. Add buffers to `dump_mamba2_state_for_debug` (or a new sibling method `dump_rl_state_for_debug`) and extend `compare-mamba2-state.py` (or sibling).
**Predicted outcome**: one of CfC carry, PER buffer state, or LOB sim state diverges at step 1 dump (= captures post-step-1 state, which gets fed into step 2 forward). That's the new first-divergent leaf to fix.
### Files touched in Phase 2.4 (no commits per dispatch)
| Path | Change | LOC |
|------|--------|-----|
| `crates/ml-alpha/src/cublas_determinism.rs` | Added `FOXHUNT_DETERMINISTIC=2` control mode (DEFAULT_MATH) + verbose print of mode at handle init | +20 / -5 |
| `crates/ml-alpha/src/trainer/perception.rs` | Extended `dump_mamba2_state_for_debug` with: 5 mamba2 backward scratch buffers (L1 + L2), 8 mamba2 backward output buffers (dw/db/d_x), 6 trunk weights, 6 trunk gradients | +90 |
| `scripts/compare-mamba2-state.py` | Added matching `BACKWARD_SCRATCH` + `TRUNK_STATE` buffer lists for new dump fields | +25 |
### Speed cost (PEDANTIC vs DEFAULT_MATH vs TF32)
Tier 1.5 mid-smoke (200 train + 50 eval, b=128) on RTX 3050:
| Mode | wall-clock /run (both runs total) | vs TF32 |
|------|------|---------|
| `FOXHUNT_DETERMINISTIC=1` PEDANTIC | 2m 32s | +90% |
| `FOXHUNT_DETERMINISTIC=2` DEFAULT_MATH | 2m 31s (~equal to PEDANTIC) | +88% |
| `FOXHUNT_DETERMINISTIC=0` TF32 (production) | 1m 20s | baseline |
**Surprise**: PEDANTIC is ~90% slower than TF32 on RTX 3050 (consumer-grade, no real tensor cores) — much worse than the 10-15% NVIDIA quotes for H100. The slowdown is amplified because RTX 3050 doesn't have the TF32-accelerated matrix units of Ampere data-center / H100 GPUs.
**Action item for Phase 2.5**: when the residual is fixed, document the H100 cost via a quick cluster smoke. If H100 PEDANTIC cost exceeds 15%, consider keeping `FOXHUNT_DETERMINISTIC=0` as the production default (with `=1` only for repro / debug sessions).
### Determinism residual trial-to-trial variability
Same seed (42), same FOXHUNT_DETERMINISTIC=1 PEDANTIC, three trials:
| Trial | First-divergent leaf | Δ at step 2 |
|------:|----------------------|-------------|
| 1 | `checksums.attn_q` | 1.99e-5 |
| 2 | `checksums.attn_q` | 3.97e-5 |
| 3 | `checksums.mamba2_grad_w_c_l1` | 2.40e-7 |
Magnitude varies by ~100× between trials; leaf identity alternates. This is consistent with a single non-deterministic source whose output magnitude varies (e.g., RNG-mediated sampling) — NOT with a deterministic-but-non-bit-exact kernel.
### Pre-existing WIP fixed
Found broken Phase 2.5 WIP in `crates/ml-alpha/src/trainer/integrated.rs` (`dump_rl_state_for_debug` method at line 10171, invoked at 10519). The closure captured `let stream = &self.stream` which deref-coerces to `&CudaStream`, where the `memcpy_dtoh` method is NOT in scope (it lives on `Arc<CudaStream>`). Fixed by replacing with `let stream = self.stream.clone()` (cheap Arc refcount bump) so the method is found via the impl block on `Arc<CudaStream>`.
This unblocked the build but the Phase 2.5 RL-state-dump infrastructure remains opt-in (env-gated by `FOXHUNT_DETERMINISM_DEBUG_RL=1`) — ready for the Phase 2.5 dispatch to actually use it. Not exercised in this dispatch's verification.
### Discipline checklist
- [x] No new files in repo root
- [x] No new docs outside `docs/superpowers/notes/`
- [x] No `.cu` files added/changed (the bug is NOT in any CUDA kernel)
- [x] Per `feedback_no_atomicadd.md`: not relevant
- [x] Per `feedback_no_nvrtc.md`: not relevant
- [x] Per `feedback_cpu_is_read_only.md`: dump methods do DtoH only, no host compute affecting training
- [x] Per `feedback_no_stubs.md`: dumps wired end-to-end; no-op when env unset
- [x] Per `feedback_local_smoke_before_cluster.md`: validated on RTX 3050, no cluster submit
- [x] Per `feedback_systematic_debugging` + dispatch STOP rule: no fix applied — diagnosis only
- [x] No commit — left uncommitted for human review
### Reproduction
```bash
# Build with extended dumps + control mode
SQLX_OFFLINE=true cargo build --release --example alpha_rl_train -p ml-alpha
# Confirm PEDANTIC is applied (look for [cublas_determinism] line in log)
FOXHUNT_DETERMINISTIC=1 ./scripts/determinism-check.sh --quick
# Expected output:
# FIRST_DIVERGENCE: step=2 leaf=checksums.mamba2_grad_w_c_l1 delta=1.1e-7
# (or attn_q / vsn_w with different magnitude — divergence is small, noisy)
# Run with mamba2 dumps to see step 1's full post-step state EQUAL
FOXHUNT_DETERMINISTIC=1 ./scripts/determinism-check.sh --quick --debug-dump-mamba2
# Expected output (key bit):
# === step 1 === ALL 52 buffers EQUAL
# === step 2 === mamba2_l1/l2 weights DIVERGE
# mamba2 backward outputs ALL DIVERGE
# cuBLAS-GEMM outputs (dw_in/dw_a/dw_b) DIVERGE
# trunk weights/grads NOT dumped at step 2 (still need to add)
# Control: try DEFAULT_MATH (no TF32, no PEDANTIC) — same residual
FOXHUNT_DETERMINISTIC=2 ./scripts/determinism-check.sh --quick
# Per-leaf divergence trace: open `/tmp/foxhunt-determinism-a.log` /
# `b.log` and look for the `[cublas_determinism]` startup line to
# confirm which math mode was applied.
```

View File

@@ -0,0 +1,297 @@
# Determinism Phase 2.5 — RL-loop state investigation
**Date**: 2026-06-02
**Dispatch**: Phase 2.5 follow-up to Phase 2.4's mamba2 falsification.
**Outcome**: ALL 6 candidate groups (CfC carry, PER state, action RNG, LOB
sim, ISV controllers, reward shaping) **falsified at step 2 dump**. The
residual non-determinism lives in **step 2's BACKWARD PIPELINE**, with
identical forward activations but divergent backward outputs.
**Per dispatch STOP rule**: no fix applied — culprit is OUTSIDE the 6
candidate groups. Phase 2.6 dispatch needed.
---
## Recap (Phase 2.4 → Phase 2.5)
Phase 2.4 falsified mamba2 / cuBLAS / trunk backward — 52 buffers EQUAL
at step 1's dump (post-step-1 state). At step 2's dump the mamba2 L1/L2
weights diverged but the upstream RL-loop state had not been
instrumented. Phase 2.5's dispatch asked: which of the 6 not-yet-dumped
RL-loop candidates is the culprit?
The 6 candidate groups:
- **A**. CfC recurrent cross-step carry (`attn_context_d`, `h_t_d`)
- **B**. PER replay state (`priority_tree_d`, `sample_prng_d`, `sample_indices_d`)
- **C**. Action sampling RNG (`prng_state_d`, `iqn_prng_state_d`)
- **D**. LOB simulator (`pyramid_units_count_d`, `unit_*_d`,
`prev_realized_pnl_d`, `prev_position_lots_d`, `steps_since_done_d`,
`trade_duration_emit_d`, `close_unit_index_d`, `unit_prev_pos_lots_d`,
`unit_initial_r_d`, `unit_trail_distance_d`)
- **E**. ISV controllers (full mapped-pinned ISV array — 757 floats)
- **F**. Reward shaping outputs (`rewards_d`, `raw_rewards_d`,
`outcome_ema_d`, `reward_abs_d`, `dones_d`, `actions_d`,
`next_actions_d`, `log_pi_old_d`, `advantages_d`, `returns_d`,
+ CMDP + edge-decay per-batch: `session_pnl`, `consec_loss`,
`cooldown_remaining`, `ph_mu/_count/_m/_mmin/_stat`)
A and B were already dumped by `dump_mamba2_state_for_debug` (attn_context,
h_t) and `dump_per_state_for_debug` (priority_tree, sample_prng,
sample_indices). Phase 2.5 added C/D/E/F via a new sibling method
`IntegratedTrainer::dump_rl_state_for_debug` (env-gated by
`FOXHUNT_DETERMINISM_DEBUG_RL=1`), plus a matching `scripts/compare-rl-state.py`.
## Task 1 — extended dump infrastructure
Added `IntegratedTrainer::dump_rl_state_for_debug(step)` invoked from
`build_diag_value` for `step ∈ {0,1,2,3}`. Captures 35 new device
buffers via DtoH copies on `self.stream`, plus the full mapped-pinned
ISV array (757 f32, read directly host-side). Files written as raw
little-endian binaries to `$FOXHUNT_DEBUG_DUMP_DIR`.
`scripts/compare-rl-state.py` reads run_a/ vs run_b/ dumps and emits
per-group verdict (EQUAL/DIVERGE) with first-divergent index +
magnitude. Naming convention matches `compare-mamba2-state.py` and
`compare-per-state.py` exactly.
`scripts/determinism-check.sh` gained a new `--debug-dump-rl` flag that
sets `FOXHUNT_DETERMINISM_DEBUG_RL=1` for both runs and invokes
`compare-rl-state.py` after the runs finish. Composes with the existing
`--debug-dump` and `--debug-dump-mamba2` flags — running all three in one
go gives the full picture (PER + mamba2 + RL state).
## Task 3 — run + verdict
Triple-dump invocation:
```bash
FOXHUNT_DETERMINISTIC=1 ./scripts/determinism-check.sh \
--quick --debug-dump --debug-dump-mamba2 --debug-dump-rl
```
### Per-group verdict at step ∈ {0, 1, 2}
| Group | Buffers | Step 0 | Step 1 | Step 2 |
|-------|---------|:------:|:------:|:------:|
| A — CfC carry / h_t | `attn_context`, `h_t` | EQUAL | EQUAL | EQUAL |
| B — PER state | `prng`, `tree`, `indices` | EQUAL | EQUAL | EQUAL |
| C — Action / IQN RNG | `action_prng_state`, `iqn_prng_state` | EQUAL | EQUAL | EQUAL |
| D — LOB sim | 13 buffers (units, pos, steps...) | EQUAL | EQUAL | EQUAL |
| E — ISV controllers | full 757-float ISV | EQUAL | EQUAL | EQUAL |
| F — Reward shaping | 18 buffers (rewards, advantages, actions, ...) | EQUAL | EQUAL | EQUAL |
| **mamba2 forward activations** | `vsn_out`, `mamba2_l{1,2}_h_enriched_seq`, `ln_{a,b}_out`, `attn_context`, `h_t` | EQUAL | EQUAL | **EQUAL** |
| **mamba2 / trunk weights** | 18 mamba2 + 6 trunk | EQUAL | EQUAL | **DIVERGE** |
| **mamba2 / trunk gradients** | 13 mamba2 bwd outputs + 6 trunk grads | EQUAL | EQUAL | **DIVERGE** |
### Step 3 (consequence-of-step-2 divergence)
Once step 2's weights diverge, step 3's forward (using divergent
weights) produces divergent activations, divergent log_pi_old at step 3
dump, and propagates to ISV at step 3 dump (slot 407 = q_pi_agree EMA,
slot 419 = KL EMA, slot 421 = adv_var_ratio EMA — all derived from
step-3's logits).
### Decisive evidence — step 2 dump
```
=== step 2 ===
vsn_out: EQUAL (229376 elements)
mamba2_l1_h_enriched_seq: EQUAL (524288 elements)
ln_a_out: EQUAL (524288 elements)
mamba2_l2_h_enriched_seq: EQUAL (524288 elements)
ln_b_out: EQUAL (524288 elements)
attn_context: EQUAL (16384 elements)
h_t: EQUAL (16384 elements)
mamba2_l1_w_in: DIVERGE Δ ~ 1.6e-6
mamba2_l1_b_in: DIVERGE Δ ~ 7.6e-7
... (all 18 mamba2 weights diverge)
vsn_w: DIVERGE Δ ~ 1.2e-6
attn_q: DIVERGE Δ ~ 9.0e-7
cfc_w_in / cfc_w_rec / ln_*_gain: DIVERGE
mamba2_l1_d_w_c_per_sample: DIVERGE Δ ~ 1.6e-7
mamba2_l1_dw_c_reduced: DIVERGE Δ ~ 9.1e-6
mamba2_l1_d_h_s2: DIVERGE Δ ~ 2.9e-7
mamba2_l1_d_a_per_channel: DIVERGE Δ ~ 3.4e-10
mamba2_l1_d_b_per_channel: DIVERGE Δ ~ 7.1e-9
mamba2_l1_d_a_proj_2d: DIVERGE Δ ~ 2.0e-8
mamba2_l1_d_b_proj_2d: DIVERGE Δ ~ 1.2e-6
mamba2_l1_dw_in: DIVERGE Δ ~ 3.6e-6
mamba2_l1_dw_a: DIVERGE Δ ~ 4.5e-6
mamba2_l1_dw_b: DIVERGE Δ ~ 2.4e-5
mamba2_l1_db_in / db_a / db_b: DIVERGE Δ up to 1.7e-4
mamba2_l1_d_x / d_x_from_in: DIVERGE Δ ~ 5.5e-7
grad_vsn_w: DIVERGE Δ ~ 2.0e-5
grad_attn_q: DIVERGE Δ ~ 8.6e-6
grad_cfc_w_in / w_rec: DIVERGE
grad_ln_a_gain / ln_b_gain: DIVERGE
action_prng_state: EQUAL (128 elements)
iqn_prng_state: EQUAL (128 elements)
pyramid_units_count: EQUAL (128 elements)
unit_{entry_step,entry_price,lots,active,initial_r,trail_distance,prev_pos_lots}: EQUAL
close_unit_index: EQUAL
prev_realized_pnl: EQUAL
prev_position_lots: EQUAL
steps_since_done: EQUAL
trade_duration_emit: EQUAL
isv_full: EQUAL (757 elements)
rewards / raw_rewards / outcome_ema / reward_abs: EQUAL
dones / actions / next_actions: EQUAL
log_pi_old: EQUAL (128 elements)
advantages: EQUAL (128 elements)
returns: EQUAL (128 elements)
session_pnl / consec_loss: EQUAL
cooldown_remaining: EQUAL
ph_mu / ph_count / ph_m / ph_mmin / ph_stat: EQUAL
```
### Eval consequence
Same seed, same FOXHUNT_DETERMINISTIC=1 PEDANTIC:
| Run | eval pnl_usd | max_drawdown_usd | trades | win_rate | sharpe |
|-----|-------------:|-----------------:|-------:|---------:|-------:|
| A | +$187,087.50 | $191,412.50 | 355 | 0.327 | +1.63 |
| B | -$261,425.00 | $363,750.00 | 430 | 0.291 | -1.83 |
Eval pnl differs by ~$450k between same-seed runs. Same root cause as
Phase 2.4's $184k drift — the tiny step-2 weight divergence cascades
through 247 RL steps × 50 eval steps × non-linear LOB sim dynamics into
a trajectory-level chasm.
## Interpretation — bug is in step 2's BACKWARD KERNEL CHAIN
The step 2 evidence is decisive:
1. **Forward at step 2 is bit-deterministic** (weights at step 2 start
are EQUAL per Phase 2.4 step 1 dump; all forward activations match
bit-exactly at step 2 dump).
2. **All loss-input scalars to backward are bit-deterministic** at
step 2 dump:
- `log_pi_old`, `advantages`, `returns` (saved during step 2 forward)
- `actions`, `next_actions`, `rewards`, `dones` (from RL env)
- PER `sample_indices`, `priority_tree`, `prng` (PER state EQUAL)
- action RNG state, IQN RNG state (`prng_state`, `iqn_prng_state`)
- LOB sim state (positions, pyramid, units, balance)
- ISV controller state (popart, kelly, wr_ema, edge-decay PH —
all 757 slots EQUAL)
3. **Yet step 2's BACKWARD OUTPUTS diverge**: every reduced gradient
(`grad_vsn_w`, `grad_attn_q`, `grad_cfc_*`, `grad_ln_*`), every
mamba2 backward scratch (`d_w_c_per_sample`, `d_h_s2`,
`d_a_per_channel`, `d_b_per_channel`), every cuBLAS GEMM output in
the mamba2 backward (`dw_in`, `dw_a`, `dw_b`).
For backward to differ with identical forward + identical loss-inputs,
some kernel in the gradient-emission chain must be non-deterministic.
Phase 2.3 already applied `CUBLAS_PEDANTIC_MATH` to all 3 cuBLAS sites
(mamba2_block, dqn, iqn) so PEDANTIC GEMMs are NOT the cause.
## Candidates NOT yet investigated (Phase 2.6 scope)
The actual culprit lives in one of:
1. **PER replay sample during step 2 K-loop** — at k_iter=0 the sample
produces `sampled_h_t_d` / `sampled_actions_d` / `sampled_rewards_d`
/ `sampled_log_pi_old_d` / `sampled_dones_d`. **These were NOT
dumped this dispatch** (they're overwritten each k_iter — the dump
captures only the LAST k_iter's output).
2. **Head backward kernels** (Q / π / V / IQN / Dueling-Q) writing
`grad_h_t` — any non-deterministic accumulation across batch + atom
+ tau dimensions. NOT dumped this dispatch.
3. **`grad_h_accumulate`** combining per-head grad_h's into the encoder
accumulator (with λ scaling). If it has a non-deterministic
parallel-add path... NOT dumped this dispatch.
4. **Smoothness backward** running upstream of Loop 1 in
`dispatch_train_step` (heads' BCE / aux portion). NOT dumped.
5. **Aux head + trunk backward** — separate K-loop. NOT dumped.
6. **Encoder backward K-loop reductions**`reduce_axis0` on
per-batch param-grad scratches (cfc_w_in/_rec, attn_q, vsn_w, LN_a/b
gain/bias). The K-loop accumulates `+=` per-K then reduces over B.
If reduce_axis0 has a parallel reduction with non-deterministic
ordering, it would explain why ALL trunk grads diverge simultaneously.
NOT dumped this dispatch.
The strongest candidate by pattern-match (Phase 2's PER tree-rebuild
fix prototype) is **reduce_axis0** or a head backward kernel's batch
reduction. Multiple grads diverging SIMULTANEOUSLY (not cascading
through one upstream source) is most easily explained by ONE shared
reducer that all of them go through.
## Phase 2.5 STOPS per dispatch instruction
Per dispatch: *"If Task 3 finds the culprit is something OTHER than the
6 candidates, STOP — Phase 2.6 dispatch territory."* Applies here: the
bug is in the gradient-emission chain, not in any of the 6 RL-loop
candidate groups. Phase 2.6 must dump:
- `sampled_h_t_d`, `sampled_h_tp1_d`, `sampled_actions_d`,
`sampled_rewards_d`, `sampled_dones_d`, `sampled_log_pi_old_d`
captured BEFORE the K-loop's heads forward, with one dump per
k_iter (not just last).
- Head outputs: `q_logits_d`, `pi_logits_d`, `v_pred_d`,
`iqn_q_values_d`, `dueling_q_composed_d` (immediately after head
forward, BEFORE backward).
- Pre-reduce per-batch param-grad scratches: `vsn_grad_w_scratch_d`,
`cfc_grad_w_in_scratch_d`, `attn_grad_q_scratch_d`,
`ln_*_grad_*_scratch_d` (catches non-determinism IN reduce_axis0).
- `grad_h_t` per-head outputs and `grad_h_t_combined_d` (head-grad
accumulator before encoder backward).
If pre-reduce per-batch scratches match but post-reduce grads diverge
→ reduce_axis0 is the culprit (Phase 2 PER-rebuild fix pattern applies).
If pre-reduce already diverges → individual head backward kernel is
non-deterministic.
## Files touched (Phase 2.5, uncommitted)
| Path | Change | LOC |
|------|--------|-----|
| `crates/ml-alpha/src/trainer/integrated.rs` | Added `dump_rl_state_for_debug` method (env-gated by `FOXHUNT_DETERMINISM_DEBUG_RL=1`) + invocation in `build_diag_value` | +290 |
| `scripts/determinism-check.sh` | Added `--debug-dump-rl` flag, composes with existing PER/mamba2 dump flags | +25 |
| `scripts/compare-rl-state.py` | New: bytewise compare for 35 RL-loop buffers + full ISV, grouped by candidate class (C/D/E/F) with verdict logic | +266 (new file) |
No `.cu` files added/changed. No commit. Phase 2.6 dispatch should
reuse the dump infrastructure and add the missing 4 buffer classes
listed above.
## Discipline checklist
- [x] No new files in repo root
- [x] Scripts in `scripts/`, docs in `docs/superpowers/notes/`
- [x] No `.cu` files added/changed (bug is NOT in any CUDA kernel
we've identified yet)
- [x] Per `feedback_no_atomicadd.md`: not relevant
- [x] Per `feedback_no_nvrtc.md`: not relevant
- [x] Per `feedback_cpu_is_read_only.md`: dump method does DtoH only;
mapped-pinned ISV is read-only on host
- [x] Per `feedback_no_stubs.md`: dump wired end-to-end via
`build_diag_value`; no-op when env var unset
- [x] Per `feedback_local_smoke_before_cluster.md`: validated on
RTX 3050, no cluster submit
- [x] Per `feedback_systematic_debugging` + dispatch STOP rule: no
fix applied — diagnosis only, culprit outside 6 candidates
- [x] Per `feedback_no_quickfixes.md`: did NOT apply a tolerance-mask
/ f64-accumulator fix despite the temptation
- [x] No commit — left uncommitted for human review
## Reproduction
```bash
# Build with extended RL-state dump
SQLX_OFFLINE=true cargo build --release --example alpha_rl_train -p ml-alpha
# Triple-dump: PER + mamba2 + RL state
FOXHUNT_DETERMINISTIC=1 ./scripts/determinism-check.sh --quick \
--debug-dump --debug-dump-mamba2 --debug-dump-rl
# Expected output:
# [debug-dump] all 3 PER buffers EQUAL through step 2
# [debug-dump-mamba2] step 1 ALL 52 buffers EQUAL
# step 2 ALL forward activations EQUAL
# step 2 ALL weights + grads + backward outputs DIVERGE
# [debug-dump-rl] ALL 35 RL-loop buffers + 757-float ISV EQUAL through step 2
# FIRST_DIVERGENCE: step=2 leaf=checksums.{attn_q|vsn_w|mamba2_grad_w_c_l1}
# Per-leaf interpretation: see scripts/compare-rl-state.py output's
# "Per-group culprit table" and "Next-step interpretation" hints.
```

View File

@@ -0,0 +1,352 @@
# Determinism Phase 2.6 — backward-pipeline investigation + DETERMINISM ACHIEVED
**Date**: 2026-06-02
**Dispatch**: Phase 2.6 follow-up to Phase 2.5's RL-loop falsification.
**Outcome**: Two distinct non-determinism sources identified and fixed.
Same-seed runs now produce **bit-equal** training checksums + bit-equal
eval summaries. Determinism foundation goal achieved.
**Verdict**: Phase 2 → 2.6 falsification chain converged. No further
phase needed.
---
## Recap (Phase 2.5 → Phase 2.6)
Phase 2.5 falsified all 6 RL-loop candidate groups (CfC carry, PER
state, action RNG, LOB sim, ISV controllers, reward shaping). At step
2 dump every forward activation + every RL-loop scalar was **EQUAL**,
yet every trunk weight + every trunk gradient + every mamba2 backward
output **DIVERGED**. The bug therefore lived in step 2's BACKWARD
KERNEL CHAIN, in a layer not yet instrumented. Phase 2.5's dispatch
ask listed 4 not-yet-dumped buffer classes (G/H/I/J).
## Task 1 — Extended instrumentation
Added `IntegratedTrainer::dump_backward_state_for_debug(step)`
invoked from `build_diag_value` for `step ∈ {0, 1, 2, 3}`. Captures
the 4 new buffer groups via DtoH copies on `self.stream`:
* **Group G — PER replay sample outputs**: `sampled_h_t`,
`sampled_h_tp1`, `sampled_actions`, `sampled_rewards`,
`sampled_dones`, `sampled_log_pi_old` (end-of-step snapshot — these
are overwritten each k_iter, so last-k_iter-wins).
* **Group H — Head forward outputs**: `q_logits`, `pi_logits`,
`v_pred`, `iqn_q_values`, `dueling_q_composed` (post-fwd, pre-bwd).
* **Group I — Pre-reduce per-batch / per-row param-grad scratches**:
`cfc_grad_w_in/_w_rec/_b/_tau_scratch`, `vsn_grad_w/_b_scratch`,
`attn_grad_q_scratch`, `grad_ln_a/_b_gain/_bias_per_row` (zeroed
at step start, accumulated by K-loop +=, reduced before Adam,
live until next step's memset).
* **Group J — Per-head grad_h_t + combined**: `ss_q/_pi/_v/_frd/
_outcome_grad_h_t_d`, `grad_h_t_combined_d` (the encoder-backward
upstream gradient).
`scripts/compare-backward-state.py` reads run_a/ vs run_b/ dumps and
emits a per-group verdict + sub-case interpretation per the Phase 2.6
decision tree. `scripts/determinism-check.sh` gained a new
`--debug-dump-backward` flag composing with the existing
`--debug-dump`, `--debug-dump-mamba2`, `--debug-dump-rl` flags.
Added Group K mid-investigation when the initial Group I/J verdict at
step 2 pointed at `grad_h_t_outcome` (Δ ~ 0.003, far above fp32
noise) — Group K dumps `outcome_head.labels_d`, `w_d`, `b_d` to
localise WHICH input drove the outcome-head divergence.
## Task 2 — Verdict 1 (after first run)
```
FOXHUNT_DETERMINISTIC=1 ./scripts/determinism-check.sh --quick \
--debug-dump-backward
```
Per-step verdict:
| Group | Step 0 | Step 1 | Step 2 |
|-------|:------:|:------:|:------:|
| G — PER sample buffers | EQUAL | EQUAL | EQUAL |
| H — Head fwd outputs | **iqn_q_values DIVERGE** | iqn_q_values DIVERGE | iqn_q_values DIVERGE |
| I — Pre-reduce grad scratches | EQUAL | EQUAL | ALL DIVERGE |
| J — Per-head grad_h_t | EQUAL | EQUAL | grad_h_t_outcome + grad_h_t_combined DIVERGE |
**Decision-tree sub-case (2)**: Group H diverges with EQUAL Group G
→ head forward kernel is non-deterministic. Specifically:
`iqn_q_values` first diverges at step 0 with `max|Δ| = 4.5e-5`. No
upstream RL-loop perturbation needed; the divergence enters during
IQN's own forward.
### Bug 1 — IQN tau-sampling read-write race
`crates/ml-alpha/cuda/rl_iqn_forward.cu`'s
`rl_iqn_tau_cos_features` was launched as
`Grid=(B, N_TAU, 1) Block=(EMBED_DIM, 1, 1)`. For each batch=b, all
N_TAU=32 blocks for that batch ran concurrently across multiple SMs.
The kernel had a read-write race on `prng_state[batch]`:
* Every block read `seed = prng_state[batch]` at line 88.
* Block (b, 0) computed `adv = seed; xorshift32 × 8; prng_state[batch] = adv` at lines 102-107.
* No inter-block barrier. With 32 blocks per batch and SMs scheduling
blocks in run-dependent order, blocks with tau_idx > 0 sometimes
read the original seed (if scheduled before block 0's write) and
sometimes the advanced seed (if scheduled after). Different runs
picked different orderings → different `local_state` for blocks
tau_idx > 0 → different τ values → different `iqn_q_values`.
The race produced run-dependent values at step 0 already — no
upstream training divergence required to enter the picture. The
divergence didn't cascade into trunk gradients (Groups I/J) until
step 2 because IQN's backward only writes IQN-internal weight grads
(no grad_h_t flow back into the encoder), so the encoder ran from
identical inputs at step 0/1. The cascade started at step 2 once
the second bug fired (see below).
### Fix 1 — Split sampling and state-advance into two kernels
Made `rl_iqn_tau_cos_features` **READ-ONLY** on `prng_state` (added
`const` qualifier to the pointer; removed the tau_idx==0 advance
block at lines 102-107).
Added a sibling kernel `rl_iqn_advance_prng_state` launched on the
same stream immediately AFTER `rl_iqn_tau_cos_features`. The
companion advances `prng_state[batch]` by 8 xorshift32 steps with
`Grid=(ceil(B/256), 1, 1) Block=(256, 1, 1)` — one thread per batch
element, sole writer per batch. Kernel-launch ordering on a single
stream is a grid-wide barrier (replaces the racy "tau_idx==0 block
writes prng_state at end of monolithic kernel" pattern).
Per `feedback_no_atomicadd`: no atomic added.
Per `feedback_no_nvrtc`: the new kernel goes through build.rs (same
cubin as the existing forward kernels).
Per `feedback_no_legacy_aliases`: the old kernel signature changed
(`uint32_t*``const uint32_t*` on `prng_state`); no alias kept.
## Task 2 — Verdict 2 (after Fix 1)
After Bug 1 fix, iqn_q_values matches at step 0/1/2. But step 2 still
diverged in Group I (all pre-reduce scratches) + Group J
(`grad_h_t_outcome` with Δ = 0.003 — orders of magnitude above fp32
noise).
The fused outcome kernel is sole-writer-per-(batch, c) so it cannot
introduce divergence from equal inputs. Added Group K dump
(`outcome_head.labels_d`, `w_d`, `b_d`) to identify the upstream
input drift.
```
=== step 0 ===
-- group K (outcome head inputs) --
outcome_labels: EQUAL (128 i32 elements — all sentinel -1)
outcome_w: DIVERGE max|Δ|=0.00309763 (384 f32 elements)
outcome_b: EQUAL (3 f32 elements — zero-init)
```
**Bug 2 — OutcomeHead init missing `scoped_init_seed`**:
`OutcomeHead::new` called `ml_core::cuda_autograd::init::
near_zero_xavier(...)` without first installing a
`scoped_init_seed(seed)` guard. `near_zero_xavier` defers to
`generate_uniform` which falls back to a **time+thread-id**-seeded
xoshiro256++ chain when no thread-local `INIT_RNG` is set
(`crates/ml-core/src/cuda_autograd/init.rs:144-164`). Two same-seed
processes therefore drew DIFFERENT initial outcome weights.
The divergence stayed dormant for the first 1-2 steps because
`outcome_head.labels_d` starts at sentinel `-1` and only gets
overwritten by `assign_labels` on `done` events (line 7876). With
no done events yet, the fused outcome kernel writes
`s_grad_logits[k] = 0` (label < 0 branch) → `grad_h_t = 0`
regardless of `w`. The first done event in mid-smoke fired at
step 2, which is when `grad_h_t_outcome` started flowing the
divergent weights into the encoder gradient and the cascade
began.
Every other head (`dqn`, `iqn`, `policy`, `value`, `dueling_q`)
already installed the `scoped_init_seed` guard around its weight
draws (per `pearl_scoped_init_seed_for_reproducibility`). The
outcome head was the only omission.
### Fix 2 — Add seed parameter to OutcomeHead::new + install guard
Changed `OutcomeHead::new(b_size, stream)`
`OutcomeHead::new(b_size, seed, stream)`. Wrapped the
`near_zero_xavier` call in a `{ let _seed_guard =
ml_core::cuda_autograd::init::scoped_init_seed(seed); ... }`
block so the thread-local RNG is seeded for the duration of the
W draw.
Trainer call site in `IntegratedTrainer::new` now passes
`cfg.dqn_seed.wrapping_add(0x0CE0)` (= "OCE0", Outcome Classifier
Encoder 0) for an init draw stream independent of the other heads.
Per `feedback_no_legacy_aliases`: signature changed directly.
Per `feedback_no_partial_refactor`: all 3 callers (trainer + 2
tests) migrated atomically.
## Task 3 — Validation
Post-fix `./scripts/determinism-check.sh --quick` (200 train + 50
eval, b=128, RTX 3050):
```
DETERMINISTIC: all `checksums.*` leaves match across all 200 rows
(rel-tol=1e-05, abs-tol=1e-07).
```
Tested at seeds 42, 43, 44 — all PASS.
`diff <(jq -c 'del(.elapsed_s)' run_a/diag.jsonl) <(jq -c
'del(.elapsed_s)' run_b/diag.jsonl)` returns empty — same-seed
diag.jsonl is byte-equal modulo wall-clock timing.
`diff /tmp/foxhunt-determinism-a/eval_summary.json
/tmp/foxhunt-determinism-b/eval_summary.json` returns empty:
```json
{
"b_size": 128,
"max_drawdown_usd": 0.0,
"n_eval_trades_dropped": 0,
"n_eval_trades_seen": 1,
"n_pre_eval_trades_wrapped": 0,
"n_trades": 1,
"profit_factor": 0.0,
"sharpe_ann": -359035207680.0,
"total_pnl_usd": -12.5,
"win_rate": 0.0
}
```
Eval pnl: **-$12.50** in both runs — bit-equal.
`diff /tmp/foxhunt-determinism-a/alpha_rl_train_summary.json
/tmp/foxhunt-determinism-b/alpha_rl_train_summary.json` returns
empty — train summary also bit-equal.
### Backward-pipeline verdict (post fix)
```
=== step 0 ===
-- group G (PER replay sample) -- ALL EQUAL
-- group H (head fwd outputs) -- ALL EQUAL (q/pi/v/iqn/dueling)
-- group I (pre-reduce grad scratch) -- ALL EQUAL
-- group J (per-head grad_h_t) -- ALL EQUAL
-- group K (outcome head inputs) -- ALL EQUAL (labels/w/b)
=== step 1 === ALL EQUAL across all groups
=== step 2 === ALL EQUAL across all groups
=== step 3 === ALL EQUAL across all groups
```
## Speed cost
200-step quick mid-smoke on RTX 3050:
| Configuration | Wall-clock |
|---|---|
| `FOXHUNT_DETERMINISTIC=0` (TF32, no fix) | not measured — pre-fix divergent |
| `FOXHUNT_DETERMINISTIC=1` (PEDANTIC, pre-Phase-2.6) | 73.5s |
| `FOXHUNT_DETERMINISTIC=1` (PEDANTIC, post-Phase-2.6 fixes) | 73.8s |
The IQN sampling split (one additional kernel launch per IQN
forward call, 2 calls per K-iter) adds ~0.3s over 200 steps =
~1.5ms/step amortised. No measurable cost from the OutcomeHead
seed guard (one-shot at construction).
## Operational rule
**A multi-block kernel that reads-then-writes a shared scalar with
no inter-block barrier is a race.** Canonical fix: split the kernel
so the shared scalar is either pure-read or written by a single
narrow kernel on the same stream. Kernel-launch ordering on a
single stream is the cleanest grid-wide barrier available without
`cooperative_groups::grid().sync()`.
**Every weight-init call site must install `scoped_init_seed`.**
Without it, `init::*` functions fall back to time+thread-id RNG and
two same-seed processes draw different initial weights, producing
silent training-trajectory divergence that doesn't appear in
checksums until the divergent weights first affect the loss
(potentially many steps later — outcome_head had a 2-step lag
because labels stayed at sentinel until the first done event).
Per `feedback_extending_existing_code_audits_for_existing_violations`:
when adding a new head, audit existing heads for the
`scoped_init_seed` pattern. The omission in `OutcomeHead` was a
2026-pre-Phase-2.6 oversight; this dispatch retroactively fixed it.
## Phase summary: PER → cuBLAS → mamba2 → RL state → backward chain
Five-phase falsification chain (Phase 2 → 2.6) achieves bit-equal
same-seed determinism:
| Phase | Bug | Fix LOC |
|---|---|---|
| 2 (PER tree rebuild) | `__threadfence` between levels with parallel blocks | ~30 LOC: single-block + `__syncthreads` |
| 2.3 (cuBLAS GEMM) | `CUBLAS_GEMM_DFALT` with TF32 split-K non-determinism | ~80 LOC: `CUBLAS_PEDANTIC_MATH` at all 3 cuBLAS sites |
| 2.4 (mamba2 falsified) | — (Phase 2.3 already deterministic) | — |
| 2.5 (RL-loop falsified) | — (all 6 candidate groups EQUAL through step 2) | — |
| 2.6 — Bug 1 (IQN tau sampling race) | Multi-block RW race on `prng_state[batch]` | ~50 LOC: split kernel + new sibling kernel |
| 2.6 — Bug 2 (OutcomeHead unseeded init) | Missing `scoped_init_seed` guard around `near_zero_xavier` | ~10 LOC: add `seed` param + install guard |
Total LOC across all 5 phases: ~170 LOC (5 .cu / .rs touches),
zero behavioral changes (training math unchanged), zero
performance regression (≤ 1.5ms/step overhead measured).
## Files touched (Phase 2.6, uncommitted)
| Path | Change | LOC |
|------|--------|-----|
| `crates/ml-alpha/cuda/rl_iqn_forward.cu` | `rl_iqn_tau_cos_features` made READ-ONLY on prng_state; new `rl_iqn_advance_prng_state` sibling kernel | +50 |
| `crates/ml-alpha/src/rl/iqn.rs` | Load `rl_iqn_advance_prng_state` symbol; call after `rl_iqn_tau_cos_features` in `forward_inner` | +20 |
| `crates/ml-alpha/src/rl/outcome_head.rs` | `OutcomeHead::new` signature `+seed` parameter; install `scoped_init_seed` guard around `near_zero_xavier` | +15 |
| `crates/ml-alpha/src/trainer/integrated.rs` | Pass `cfg.dqn_seed.wrapping_add(0x0CE0)` to `OutcomeHead::new`; new `dump_backward_state_for_debug` method (env-gated by `FOXHUNT_DETERMINISM_DEBUG_BACKWARD=1`) + invocation in `build_diag_value` | +290 |
| `crates/ml-alpha/src/trainer/perception.rs` | Added 11 `debug_*_view` accessor methods for the Phase 2.6 pre-reduce grad scratches | +30 |
| `scripts/determinism-check.sh` | New `--debug-dump-backward` flag composing with existing PER/mamba2/RL dump flags | +30 |
| `scripts/compare-backward-state.py` | New: bytewise compare for 4 backward-pipeline buffer classes + Group K outcome inputs with verdict + decision-tree interpretation | +290 (new file) |
No commit. Discipline checklist:
- [x] No new files in repo root
- [x] Scripts in `scripts/`, docs in `docs/superpowers/notes/`
- [x] One `.cu` file touched: kernel split + new sibling (no functionality removed)
- [x] Per `feedback_no_atomicadd.md`: no atomic added
- [x] Per `feedback_no_nvrtc.md`: cubin route via existing build.rs entry
- [x] Per `feedback_no_quickfixes.md`: actual root-cause fixes (race + seed),
not tolerance masks or f64 accumulators
- [x] Per `feedback_local_smoke_before_cluster.md`: validated on RTX 3050,
seeds 42/43/44, no cluster submit
- [x] Per `feedback_systematic_debugging` + dispatch STOP rule: dump
first, identify, then fix. Two distinct bugs surfaced in
sequence; both fixes anchored on raw-byte evidence
- [x] No commit — left uncommitted for human review
## Reproduction
```bash
# Build with Phase 2.6 fixes
SQLX_OFFLINE=true cargo build --release --example alpha_rl_train -p ml-alpha
# Run determinism check (no extra flags needed)
FOXHUNT_DETERMINISTIC=1 ./scripts/determinism-check.sh --quick
# Expected output:
# DETERMINISTIC: all `checksums.*` leaves match across all 200 rows
# (rel-tol=1e-05, abs-tol=1e-07).
# Run for seeds 43, 44 too
FOXHUNT_DETERMINISTIC=1 ./scripts/determinism-check.sh --quick --seed 43
FOXHUNT_DETERMINISTIC=1 ./scripts/determinism-check.sh --quick --seed 44
# Verify eval-summary bit-equality
diff /tmp/foxhunt-determinism-a/eval_summary.json \
/tmp/foxhunt-determinism-b/eval_summary.json
# Verify train-summary bit-equality
diff /tmp/foxhunt-determinism-a/alpha_rl_train_summary.json \
/tmp/foxhunt-determinism-b/alpha_rl_train_summary.json
# Run the backward-state dump diagnostic if you want to see the
# now-deterministic state at step 0/1/2/3 explicitly
FOXHUNT_DETERMINISTIC=1 ./scripts/determinism-check.sh --quick \
--debug-dump-backward
# Expected: VERDICT: all dumped backward-pipeline buffers match
# across all dumped steps.
```

View File

@@ -0,0 +1,267 @@
# Fast Dev Cycle — Items 1-5 Implementation Log
**Date**: 2026-06-02
**Branch**: ml-alpha-adaptive-controller-floors
**Spec**: `docs/superpowers/specs/2026-06-02-fast-dev-cycle.md` §3.1 - §3.6 (items 1-5 only)
**Out of scope**: §3.4 async double-buffer checkpoint (item 6, separate effort)
**Status**: Items 1-5 SHIPPED. Item 3 surfaced a critical determinism break — see §3 below.
## Summary
Built a working Tier 1.5 mid-smoke funnel that gates 80%+ of bad hypotheses before they reach cluster:
- `test_data/futures-baseline-mid/ES.FUT/` — symlinked train+eval files (Q1 + Q2 ES MBP-10, 19 GB resolved, 4 KB on-disk symlinks)
- `scripts/local-mid-smoke.sh` — full mid-smoke runner (5m 51s wall-clock on RTX 3050 Ti at b=128 / 2000 train + 500 eval)
- `scripts/tier1_5_verdict.py` — behavioral kill verdict (7 signals + per-step/eval_summary consistency check, stdlib-only)
- `scripts/determinism-check.sh` — reproducibility self-test (runs mid-smoke twice with same seed, diffs final 5 rows)
- `CLAUDE.md` — documented under "Build & Test" with three-tier funnel table
- `.gitignore` — added `test_data/futures-baseline-mid/`
**Verdict tooling works end-to-end on real data** — first run flagged `KILL_popart.sigma_CV+eval_pnl` AND surfaced a $4.5M per-step/eval_summary discrepancy ($517k vs -$4M), matching the canonical `pearl_grwwh_eval_catastrophic_collapse` failure mode.
**CRITICAL FINDING (item 3)**: trainer is NOT deterministic. Two same-seed runs diverge by step 2 in `grad_norm_ema.q` (Δ=4.5), and by end of 200 steps:
- Train pnl diverges $33,632 vs $32,225 (Δ=$1,406)
- Eval pnl diverges $0 vs $364,925 (Δ=$364,925)
- Eval trade count diverges 0 vs 347
Most CPU-side seeding looks correct on audit. Root cause is GPU-side non-determinism (likely cuDNN heuristics or non-deterministic CUDA reductions). Fix is out of scope for this spec.
## Item 1 — §3.1 PVC data sync
### Plan deviation
Spec called for "rsync 2-3 highest-value MBP-10 files to local `test_data/futures-baseline-mid/` (~1.5 GB)". Actual reality:
- PVC path is `/training-data/futures-baseline-mbp10/ES.FUT` via `training-data-pvc` (NOT `/feature-cache/futures-baseline-mbp10` per spec — feature-cache-pvc only has predecoded sidecars)
- Smallest MBP-10 file is 2024-Q1 at 4.7 GB. The "~1.5 GB total" spec estimate was wrong by 10×
- User already had `test_data/futures-baseline-mbp10/ES.FUT/{Q1, Q2}.dbn.zst` locally (verified SHA256 match against PVC)
### What shipped
1. Created `test_data/futures-baseline-mid/ES.FUT/` containing two symlinks (zero disk cost, 19 GB resolved):
- `ES.FUT_2024-Q1.dbn.zst``../../futures-baseline-mbp10/ES.FUT/ES.FUT_2024-Q1.dbn.zst` (4.8 GB train)
- `ES.FUT_2024-Q2.dbn.zst``../../futures-baseline-mbp10/ES.FUT/ES.FUT_2024-Q2.dbn.zst` (14 GB eval)
2. SHA256 verified against PVC: `9ba3bcc468ccf148d8341cd71cec008ae2532676b0f4bb8c238b0e61c5aca625` (Q1, matches PVC)
3. `.gitignore`: added `test_data/futures-baseline-mid/`
4. `CLAUDE.md`: updated "Build & Test" section with tiered-validation table and the three new scripts
### Walk-forward fold scheme
With 2 files and `--n-folds 2 --fold-idx 0`, the trainer splits:
- block_size = 1 (files / folds)
- train = files[0..1] = Q1
- eval = files[1..2] = Q2
This gives real walk-forward semantics (no train/eval overlap), matches `--mbp10-data-dir /data/futures-baseline-mbp10/ES.FUT` invocation pattern from `infra/k8s/argo/alpha-rl-template.yaml:181`.
## Item 2 — §3.2 Local mid-smoke script
### What shipped
`scripts/local-mid-smoke.sh` (executable, 152 lines):
- Argparse for `--seed`, `--out`, `--n-backtests`, `--n-steps`, `--n-eval-steps`, `--n-folds`, `--fold-idx`
- Defaults: seed=42, b=128, 2000 train + 500 eval, fold 0 of 2
- Builds via `SQLX_OFFLINE=true cargo build --release --example alpha_rl_train -p ml-alpha`
- Invokes binary with `--mbp10-data-dir test_data/futures-baseline-mid/ES.FUT --instrument-mode front-month --log-every 200 --gpu-idx 0`
- Pre-flight checks: directory exists, ≥ 2 MBP-10 files, follows symlinks via `find -L`
- Wall-clock reporting + next-step hint pointing to verdict script
### Measured wall-clock on RTX 3050 Ti
| Configuration | Wall-clock |
|---|---|
| Default (b=128, 2000+500) | **5m 51s** |
| Sanity check (b=16, 200+50) | 4m 45s (dominated by 210s eval-file load) |
GPU memory used: 1.2 GB / 4 GB (b=128) — well within RTX 3050 Ti capacity per spec §2.
### Path correction during shake-down
Initial draft passed `--mbp10-data-dir` at the parent of `ES.FUT/`. The `discover_mbp10_files_sorted` (`crates/ml-alpha/src/data/loader.rs:294`) is non-recursive — it expects the path to be the instrument's `.dbn.zst` directory directly. Fixed to match argo template (`--mbp10-data-dir /data/futures-baseline-mbp10/ES.FUT`).
## Item 3 — §3.5 Determinism audit
### CPU-side seeding audit (all clean)
Every randomness entry point in `crates/ml-alpha/src/` derives from `cli.seed`:
| Source | Seeded by | Location |
|---|---|---|
| Loader random-anchor RNG (`MultiHorizonLoader.rng`) | `ChaCha8Rng::seed_from_u64(cfg.seed)` | `crates/ml-alpha/src/data/loader.rs:422` |
| `upload_to_gpu` seed | `cli.seed` arg | `crates/ml-alpha/examples/alpha_rl_train.rs:391` |
| DQN trainer seed | `cli.seed.wrapping_add(0xDA)` | `crates/ml-alpha/examples/alpha_rl_train.rs:410` |
| PPO trainer seed | `cli.seed.wrapping_add(0xBE)` | `crates/ml-alpha/examples/alpha_rl_train.rs:411` |
| PER buffer seed | `cli.seed.wrapping_add(0x9E37_79B9)` | `crates/ml-alpha/examples/alpha_rl_train.rs:413` |
| Eval-phase loader | `cli.seed.wrapping_add(0xE7AE)` | `crates/ml-alpha/examples/alpha_rl_train.rs:786` |
| Device-side Thompson xorshift32 | `ChaCha8Rng::seed_from_u64(cfg.dqn_seed.wrapping_add(0xC0_C0_C0_C0))` | `crates/ml-alpha/src/trainer/integrated.rs:2142` |
| Loader explicitly does NOT shuffle | `// (no shuffle — file order is the caller's…)` | `crates/ml-alpha/src/data/loader.rs:412` |
No Python data loader exists — entirely Rust. No `thread_rng()` / `random()` / `SystemTime::now()` calls in `ml-alpha` hot path. `ml-ppo` and `ml-core` use `thread_rng()` but those are not on the RL training path.
### Self-test result: DETERMINISM IS BROKEN
`scripts/determinism-check.sh --quick` (200 train + 50 eval, seed=42):
```
FIRST NON-TIME DIVERGENCE at row 2 step=2: 4 drifts
grad_norm_ema.q A=64.80356 B=60.30021 Δ=4.50335
risk_stack.atom_calibration.target_bot_saturation_rate
A=0.14137 B=0.15253 Δ=0.01116
isv_ema_in.td_kurtosis A=0.00092 B=0.00093 Δ=9.59e-06
controller_branch.per_alpha_input A=0.00092 B=0.00093 Δ=9.59e-06
```
By step 199:
- 654 leaves diverging
- `trading.realized_pnl_cum_usd`: $33,631 vs $32,225 (Δ=$1,406)
- `confidence_gate.gated_count_total`: 4826 vs 495 (~10× difference)
- `units.trail_distance[*]` differs across most of the 128 batch slots
- `risk_stack.regime.tail.session_pnl_variance_ema`: 261 vs 1008 (~4× difference)
Eval (50 steps) shows much larger split:
- Run A: 0 trades, $0 pnl, 0 win_rate
- Run B: 347 trades, $364,925 pnl, 0.343 win_rate
### Inferred root cause (not fixed per spec)
CPU-side seeding is comprehensive. The drift originates GPU-side, candidates:
1. **Non-deterministic CUDA reductions** — block-tree reductions on different launch geometries can produce fp32 differences in addition order even with identical inputs. Compounding through ~200 Adam steps yields macroscopic divergence.
2. **cuDNN heuristic auto-tune** — mamba2 forward / attention kernels may run different algorithm selections on first call vs warm-cache.
3. **Stream ordering** — multiple CUDA streams (`prefill_graph`, `postfill_graph`, `reward_graph`, `training_graph`) interleave in arbitrary order at sync boundaries.
The `grad_norm_ema.q` divergence at step 2 (Δ=4.5 on a value of ~64) is much larger than fp32 noise (1e-5), so this is structural, not roundoff. Inspection scope is a separate spec — flagged here for the next session.
### Determinism-check script behavior
`scripts/determinism-check.sh`:
- Default 2000+500 (matches mid-smoke). `--quick` → 200+50 for faster diagnostic.
- Diffs last 5 rows of `diag.jsonl` leaf-by-leaf with rel-tol=1e-5, abs-tol=1e-7
- Skips `elapsed_s` / `sps` / wall-clock signals (inherently non-deterministic)
- Reports top 10 drifts per row sorted by absolute magnitude
- Exit 0 = deterministic, 1 = drift, 2 = run failure
## Item 4 — §3.3 Behavioral kill verdict script
### What shipped
`scripts/tier1_5_verdict.py` (executable, ~430 lines, Python 3 stdlib only — no pandas/numpy):
7 behavioral signals + 1 consistency check, each with explicit KILL/WARN/OK thresholds:
| Signal | Source field | KILL | Target | Falsifies |
|---|---|---|---|---|
| `action_entropy` | last train row | <0.6 or >2.3 | [1.20, 2.04] | policy collapse / random |
| `q_pi_agree_ema` | last train row | <0.3 | ≥0.6 | Q/π decoupled |
| `pearson(rewards.sum, Δpnl)` | full train trace | <0.3 | ≥0.5 | `pearl_reward_signal_anti_aligned_with_pnl` |
| `avg_hold_steps trend` | last 1500 train rows (linreg slope) | slope<-0.01 | slope>0 | surfer pattern absent |
| `popart.sigma CV` | full train trace | CV>0.5 | CV<0.5 | controller wildly oscillating |
| `wr_train` (Kelly EMA) | last train row | <0.15 | ≥0.25 | random or worse |
| `wr_eval` (Kelly EMA) | last eval row | <0.15 | ≥0.20 | random or worse |
| `eval_pnl` | eval_summary.total_pnl_usd | <-$3M | >-$1M | catastrophic |
Output modes: human-readable (default), `--quiet` (one-line verdict), `--json` (machine-readable).
Exit codes: 0=OK, 1=KILL, 2=input missing/insufficient.
### Sample output on real mid-smoke
Ran on `/tmp/foxhunt-mid-smoke` (full b=128, 2000+500, seed=42):
```
========================================================================
Tier 1.5 verdict: KILL_popart.sigma_CV+eval_pnl_(eval_summary.total_pnl_usd)
========================================================================
train rows: 2000
eval rows: 500
eval_summary: present
STATUS SIGNAL EXPLANATION
------ ------------------------------------------ ----------------
OK action_entropy 1.755 within target [1.20, 2.04]
OK q_pi_agree_ema 0.963 ≥ 0.6
WARN pearson(rewards.sum, Δpnl) 0.318 below target ≥ 0.5 (n=1999)
OK avg_hold_steps trend slope +0.00264/step over last 1500 rows
KILL popart.sigma CV 0.810 > 0.5 — controller wildly oscillating
OK wr_train (trading.win_rate) 0.341 ≥ 0.25
OK wr_eval (trading.win_rate) 0.336 ≥ 0.2
KILL eval_pnl (eval_summary.total_pnl_usd) $-3,996,663 < $-3,000,000 — catastrophic
WARN per-step vs eval_summary per_step=$517,914 | eval_summary=$-3,996,663 | gap=$4,514,576 (112.96%) > 5% — accounting axis discrepancy
[KILL] KILL_popart.sigma_CV+eval_pnl_(eval_summary.total_pnl_usd) — DO NOT submit cluster smoke. Fix locally first.
```
The verdict caught **two structural problems** that would have wasted a 70-min cluster submit AND surfaced the known per-step/eval_summary disagreement — full payoff matrix.
### Notes on signal sourcing
- `wr_*` prefers `isv_config.kelly.win_rate_ema` (cumulative EMA across all trades) but falls back to `trading.win_rate` (point measure of last batch) when EMA path is absent. Source is reported in the signal name.
- `pearson(rewards.sum, Δpnl)` pairs `rewards.sum` (the gradient signal) with `Δ(trading.realized_pnl_cum_usd)` (the actual pnl direction). Per `pearl_reward_signal_anti_aligned_with_pnl`, this is the load-bearing signal — a sign-flipped or noisy correlation explains negative-eval hysteresis.
- `avg_hold_steps` is computed by linear regression on `(step, trading.avg_hold_steps)` over the trailing 1500-step window. Robust to NaN — drops rows where the field is non-finite.
## Item 5 — §3.6 Per-step / eval_summary consistency
### What shipped
Folded into `tier1_5_verdict.py` as `signal_consistency()`:
- Reads last `eval_diag.jsonl` row → `trading.realized_pnl_cum_usd`
- Reads `eval_summary.json``total_pnl_usd`
- Computes `|per_step eval_summary| / max(|per_step|, |eval_summary|, $1)`
- WARN if fraction > 5% (NOT KILL — gives signal without blocking, per spec)
### Sample finding on the real mid-smoke
```
WARN per-step vs eval_summary per_step=$517,914 | eval_summary=$-3,996,663 | gap=$4,514,576 (112.96%) > 5% — accounting axis discrepancy
```
This is the same accounting-axis bug pattern as `pearl_grwwh_eval_catastrophic_collapse` ($234M gap at grwwh step 4863). The per-step diag's `trading.realized_pnl_cum_usd` and the eval_summary's `total_pnl_usd` are computed from **different** sources (per-step running sum of `raw_rewards` × b_size vs eval_summary's aggregation of LobSim closed-trade records). A 112% gap means the per-step diag is materially misleading for eval-period pnl observability.
Per spec §3.6 long-term action: "fix the accounting axis discrepancy in the diag/summary code" — out of scope for items 1-5, flagged for follow-up spec.
## Files added / changed
| Path | Change |
|---|---|
| `test_data/futures-baseline-mid/ES.FUT/` (NEW) | 2 symlinks (Q1, Q2) into existing test_data |
| `scripts/local-mid-smoke.sh` (NEW, 152 lines, +x) | Tier 1.5 runner |
| `scripts/tier1_5_verdict.py` (NEW, ~430 lines, +x) | Behavioral verdict + consistency check |
| `scripts/determinism-check.sh` (NEW, ~125 lines, +x) | Reproducibility self-test |
| `.gitignore` | Added `test_data/futures-baseline-mid/` |
| `CLAUDE.md` | Added "Tiered local validation" subsection under Build & Test |
No CUDA / Rust / training-math changes. Pure dev infrastructure. No commits — left uncommitted for human review.
## Discipline checklist
- [x] No files saved to repo root
- [x] No new *.md files except this requested impl log
- [x] Read before Edit on all touched existing files (`.gitignore`, `CLAUDE.md`)
- [x] No stubs, no TODO/FIXME/XXX markers
- [x] Symlinks chosen over duplicate-data download (per `feedback_no_quickfixes` — leverage existing canonical files)
- [x] PVC SHA256 verification before deciding local files are usable
- [x] Verdict signals all tied to existing memory entries (pearl_grwwh, pearl_reward_signal_anti_aligned, pearl_edge_lives_at_wave_timescale)
- [x] Per-step/eval_summary check emits WARN not KILL (per spec §3.6: "WARNING — gives signal without blocking")
## Recommended next steps (for next session)
1. **Determinism fix spec**`grad_norm_ema.q` diverges at step 2. Inspect block-tree reduction ordering and stream synchronization in the encoder grad-norm path. The fact that eval pnl can differ by $365k across same-seed runs makes ALL controller-iteration verdicts noisy.
2. **Per-step/eval_summary reconciliation spec** — the $4.5M gap on this small smoke is the same shape as `pearl_grwwh_eval_catastrophic_collapse` $234M. Audit the LobSim trade accounting vs the per-step `raw_rewards` axis.
3. **Item 6 — async checkpoint** — separate ~4-day effort per spec §3.4. Unblocks resume-from-baseline + saves ~58 min per cluster smoke.
4. **Tighten Pearson threshold** — current run shows 0.318 (WARN, below 0.5 target). Confirm whether this is the established baseline or whether the mid-smoke window is too short for the signal to settle.
## Reproduction commands
```bash
# Setup (one-time; symlinks already in place from this session)
ls test_data/futures-baseline-mid/ES.FUT/ # should show 2 symlinks
# Run mid-smoke (5-6 min on RTX 3050 Ti)
./scripts/local-mid-smoke.sh
# Verdict
python3 scripts/tier1_5_verdict.py /tmp/foxhunt-mid-smoke
# Determinism self-test (~10 min, --quick for ~6 min)
./scripts/determinism-check.sh --quick
```

View File

@@ -0,0 +1,240 @@
# Phase 2 v2 Post-Mortem & Redesign Plan
**Date:** 2026-05-29
**Status:** Phase 2 v2 KILLED at step 4267 (alpha-rl-2qbmz) — architectural failure confirmed
**Baseline accepted:** Plan A v2 `fd3174262` (+$9.3M peak, 20k clean, no NaN)
**Next decision:** Track A (improve baseline) vs Track B (proper dueling redesign) vs Track C (IQN switch)
---
## 1. Post-Mortem: Why Phase 2 v2 Failed
### Three discoveries, layered
**Layer 1 — Surface bug (FIXED):** `v_head_bwd_from_grad.cu` had `#define HIDDEN_DIM 512` while
the rest of the codebase uses 128. Caused 1378 OOB writes at b=1024, mamba2 Adam cascade-failed.
Fixed in `eac382d20`. (Per `pearl_use_consts_not_literals_for_structural_dims`.)
**Layer 2 — Mathematical no-op (ROOT CAUSE):** The composition kernel does
```
composed_logits[b, a, z] = V[b] + A_logits[b, a, z] mean_a A_logits[b, a, z]
```
where `V[b]` is a **scalar per batch**. The categorical CE loss operates on `softmax_z(composed_logits)`.
Two failures:
- **Forward null:** `softmax_z(V[b] + A_centered[a, z]) = softmax_z(A_centered[a, z])` because
softmax is translation-invariant in z. V[b] is constant across z for a fixed (b, a), so it shifts
all atom logits equally and drops out of the softmax. **V has zero effect on the Q distribution.**
- **Backward null:** `grad_V[b] = Σ_z grad_composed_logits[b, a_t, z] = Σ_z (softmax[z] target[z]) = 0`
because both softmax and target distributions sum to 1. **V receives zero gradient from CE.**
The per-z mean centering (` mean_a A_logits[b, a, z]`) DOES change the distribution (per-z values
vary), but it provides only the identifiability constraint, not a V signal. Empirically, removing
the shared atom-mass pattern across actions **destroys** action-differentiation information rather
than improving it.
**Layer 3 — Cluster confirmation (DIAGNOSTIC):** ISV diagnostic at step 4267 of alpha-rl-2qbmz:
| Metric | Value | Verdict |
|---|---|---|
| gamma | 0.995 | ✓ correct |
| reward_scale | 0.00139 | ✓ adapted |
| entropy | 1.95 | ✓ no collapse |
| Hold% | 38% | ✓ no Hold-attractor |
| wr | 0.567 | ✓ winning more than losing |
| **q_pi_agree_ema** | **0.79** | 🚨 SEVERE Q↔π misalignment |
| pnl_cum | -$4.5M | trapped, oscillating |
q_pi_agree_ema = 0.79 confirms: Q's true ranking has decoupled from π's distillation target. The
broken Q distribution (from per-z mean centering distortion) feeds wrong rankings into the distill
loss, π drifts AWAY from Q's preferences over training. Result: high wr (Q finds correct direction)
+ negative pnl (π's misaligned action selection turns Q's signal into losing trades).
### Key learning saved as pearl
`pearl_scalar_v_with_categorical_ce_is_noop.md` — added to MEMORY.md index.
---
## 2. Three Strategic Tracks
### Track A: Accept Plan A v2 + Improve Reward Shaping
**Hypothesis:** Plan A v2 (vanilla C51) already achieves +$9.3M. Most of the remaining upside is in
**reward shaping** (per `pearl_reward_shape_drives_quality_over_quantity`), not architecture.
**Targets:**
1. Surfer reward terms (entry cost + min-hold + hold-time × win bonus + quadratic carry)
2. Adaptive C51 atom support (widen V_MAX for fat-tail trades)
3. Entropy regularization tuning (avoid late-stage greedy collapse from +$9.3M peak to +$8.5M final)
4. Longer training (50k or 100k steps with proper checkpointing)
5. Walk-forward cross-fold validation for out-of-sample confidence
**Effort:** ~1-2 days per target, incremental. Low risk, each step measurable in isolation.
**Expected upside:** 10-30% PnL improvement per change, potential to reach +$15-20M peak through
compounding gains. Limited by Q architecture's implicit state-value handling.
**Risk:** Low. Each change is small, easy to revert. Plan A v2 baseline always available.
**When to choose this:** Default. Lock in wins, iterate on what's working.
---
### Track B: Proper Distributional Dueling C51 (V as distribution over atoms)
**Hypothesis:** The dueling decomposition WILL help once V is a true distribution. Better long-term
architecture, sets up future improvements.
**Architecture:**
```
V_logits[b, z] — distribution over Z atoms for V(s)
A_logits[b, a, z] — per-action advantage distributions
composed_logits[b, a, z] = V_logits[b, z] + A_logits[b, a, z] mean_a A_logits[b, a, z]
softmax_z(composed_logits[b, a, :]) → Q distribution
```
**Why this works (vs Phase 2 v2):**
- V_logits[b, z] varies across z → genuinely shapes each atom's logit independently
- Different z atoms get different V contributions → softmax distribution IS affected by V
- Decompose backward: `grad_V_logits[b, z] = grad_composed_logits[b, a_t, z]` (no Σ_z reduction
— V learns per-atom)
- Identifiability via mean-subtraction preserved → action-specific deviations cleanly separated
**Implementation phases:**
**B.1: V head distributional rewrite (~150 LOC)**
- Modify `v_head_fwd_bwd.cu`: input h_t [B, HIDDEN_DIM] → output v_logits [B, Z]
- Forward: `v_logits[b, z] = Σ_c w[z, c] × h_t[b, c] + b[z]` (z = 0..Q_N_ATOMS-1)
- Backward: gradient of CE on softmax_z(v_logits) vs target_v_dist
- Weight shape: `[Z × HIDDEN_DIM]` instead of `[1 × HIDDEN_DIM]`
- Adam state for new params
**B.2: V target net (~50 LOC)**
- Mirror of online V head with soft-update τ
- Used by Bellman target build at sampled_h_tp1
**B.3: Modify composition/decomposition kernels (~30 LOC)**
- `composed_q_compose_fwd.cu`: read `v_logits[b, z]` (was `v[b]`)
- `composed_q_decompose_bwd.cu`: `grad_v_logits[b, z] = grad_composed_logits[b, a_t, z]`
(no atom-sum reduction)
- Same Jacobian for A as before (mean-subtraction preserved)
**B.4: Integrated trainer wiring update (~150 LOC)**
- Buffer shapes: `v_at_sampled_h_t_d [B, Z]`, `v_at_sampled_h_tp1_d [B, Z]`
- V target forward at sampled_h_tp1 (separate forward call)
- V backward at sampled_h_t now returns `grad_v_logits [B, Z]` instead of scalar
- V grad routing: `dueling_v_grad_w_per_batch_d [B, Z, HIDDEN_DIM]`, reduce, sum
**B.5: PPO V loss adaptation (~50 LOC)**
- `v_pred = E[V] = Σ_z p(z|v_logits) × atom_value[z]` for advantage computation
- Existing MSE(v_pred, returns) becomes MSE(E[V], returns) — minimal change
- V_dist target for the dueling backward already trains V_logits via composed Bellman CE
**B.6: Smoke + cluster validation**
- Local b=128 smoke first (compute-sanitizer clean gate)
- Cluster b=1024 20k validation (compare vs Plan A v2 baseline)
- Kill criteria: q_pi_agree_ema ≪ 0, pnl trapped negative, entropy collapse
**Total effort:** ~430 LOC, 2-3 days focused work.
**Risk:** Medium. Dueling may not help meaningfully for trading even when architecturally correct.
Plan A v2's Q-Thompson + Q→π distill already handles state value implicitly via softmax(E_Q/τ).
**Expected upside:** Uncertain. Best case +50% (architecture genuinely unlocks better
state-value reasoning). Worst case 0% (Plan A v2 already capture all available alpha given
current encoder + reward shaping).
**When to choose this:** After Track A hits a clear ceiling, OR if we want to prepare the
architecture for harder problems (more actions, more nuanced state).
---
### Track C: IQN-based Q Learning (switch from C51)
**Hypothesis:** Drop categorical (atom-based) for quantile-based Q learning. Scalar dueling works
naturally for IQN: `Q(s, a, τ) = V(s, τ) + A(s, a, τ) mean_a A(s, a, τ)` with scalar values at
each sampled τ.
**Why consider:**
- IQN head already exists in codebase (`iqn_head`, `iqn_q_values_d`)
- More expressive than C51 (continuous quantile vs discrete atoms)
- Scalar dueling is well-defined at each τ
**Effort:** Major refactor — change primary Q learner from C51 to IQN. All C51-specific
controllers, kernels, loss code, Bellman projection, action selection logic need updating.
Estimated 1-2 weeks focused work.
**Risk:** High. Many controllers were tuned for C51 dynamics (atom support, projection, etc).
IQN has different convergence properties. Major regression risk.
**When to choose this:** Long-term architectural reset, NOT a near-term fix.
---
## 3. Recommendation
**Immediate (this week):**
- ✅ Done: Kill alpha-rl-2qbmz, pearl saved, plan documented
- Accept Plan A v2 (`fd3174262`) as production baseline
- Tag baseline branch / merge to main if appropriate
**Next 1-2 weeks: Track A**
- Reward shaping iteration (surfer terms, fat-tail atom widening)
- Entropy/exploration tuning to prevent peak-vs-final pnl decay
- Longer training (50k, 100k) with checkpoint+resume
- Walk-forward validation
**Optional parallel: Track B spec preparation**
- Write detailed kernel + Rust API specs for distributional V
- Don't IMPLEMENT yet — wait for Track A to plateau
- This lets us pivot quickly if Track A hits a ceiling
**Long-term: Track C reconsidered if both A and B plateau**
---
## 4. Open Questions / Risks
1. **Does Plan A v2 generalize?** Single-seed +$9.3M peak. Need walk-forward across folds to
confirm it's alpha, not lucky seed.
2. **Track A target ranking:** Which reward shaping term to try first? Suggest order by
confidence: (1) atom widening (low risk, validates fat-tail handling), (2) entropy tuning,
(3) hold-time bonus, (4) quadratic carry.
3. **Track B V target net cost:** Adds soft-update step to existing target net infrastructure
(Q target net already exists). Should be additive complexity, not multiplicative.
4. **Track B PPO V loss interpretation:** Currently V is a critic for PPO advantage. Switching to
distributional V means E[V] is the critic — same semantics, just computed differently. Risk: if
V_dist is poorly calibrated early, PPO advantages get noisy → policy explores randomly.
5. **Should we even keep Q→π distill?** Plan A v2 uses it. q_pi_agree_ema is the right diagnostic
either way. For Track B, the distillation loss operates on the COMPOSED Q which now genuinely
has V signal, so q_pi alignment should be much better.
---
## 5. References
**Code:**
- Plan A v2 baseline: commit `fd3174262`, branch `ml-alpha-plan-a-v2`
- Phase 2 v2 failed: commit `eac382d20`, branch `ml-alpha-phase2-v2`
- Workflows: alpha-rl-8ksnj (A v2 ✅), alpha-rl-2qbmz (v2 ❌, terminated step 4267)
- Spec being superseded: `docs/superpowers/specs/2026-05-29-phase2-v2-proper-dueling-c51-design.md`
**Pearls referenced:**
- `pearl_scalar_v_with_categorical_ce_is_noop` (NEW, this session)
- `pearl_use_consts_not_literals_for_structural_dims`
- `pearl_q_thompson_actor_makes_pi_dead_weight`
- `pearl_reward_shape_drives_quality_over_quantity`
- `pearl_dd049d9a4_surfer_baseline_verified`
- `pearl_two_alpha_modes_surfer_vs_trend`
**Literature:**
- Wang et al. 2016 — Dueling Network Architectures (scalar dueling DQN)
- Bellemare et al. 2017 — A Distributional Perspective on Reinforcement Learning (C51)
- Dabney et al. 2018 — Implicit Quantile Networks (IQN; alt to C51)

View File

@@ -0,0 +1,568 @@
# Eval-Summary Trade Aggregation Fix — Implementation Plan
> **For agentic workers:** REQUIRED SUB-SKILL: superpowers:subagent-driven-development. Steps use checkbox `- [ ]` syntax.
**Spec:** `docs/superpowers/specs/2026-05-31-eval-summary-trade-aggregation-design.md`
**Goal:** `eval_summary.json` reflects ALL `b_size` accounts' eval-phase-only trades, with no train-phase contamination. Currently it computes metrics on backtest 0's last 1024 records (a mix of train + eval trades).
**Architecture:** Add two methods to `LobSimCuda` (read per-backtest counts + read all backtests' records). Replace the broken slicing block in `alpha_rl_train.rs`. Bump `TRADE_LOG_CAP` from 1024 → 4096 for cluster-scale headroom.
**Tech Stack:** Rust, cudarc.
**Branch:** `ml-alpha-eval-summary-fix` (off `ml-alpha-regime-observer`)
**Independent of Phase A** — touches different files (lobsim + eval slicing block, not the train-loop json!{}). Can run in parallel via separate subagent dispatch but must NOT branch off the in-flight `ml-alpha-checkpoints-eval-diag` branch (would conflict on alpha_rl_train.rs around the eval phase). Use a separate base.
---
## File structure
| Path | Action |
|------|--------|
| `crates/ml-backtesting/src/lob/mod.rs` | Modify — bump `TRADE_LOG_CAP` from 1024 to 4096 |
| `crates/ml-backtesting/src/sim/mod.rs` | Modify — add 2 new public methods on `LobSimCuda` |
| `crates/ml-alpha/examples/alpha_rl_train.rs` | Modify — replace slicing block with per-account aggregation |
| `crates/ml-backtesting/tests/eval_aggregation_smoke.rs` | Create — unit test for the new methods |
---
## E.1 — Lobsim API: `read_per_backtest_trade_counts`
**Files:**
- Modify: `crates/ml-backtesting/src/sim/mod.rs`
- [ ] **Step 1: Locate current `read_trade_records` for reference**
```bash
grep -nE "pub fn read_trade_records|pub fn read_total_trade_count" crates/ml-backtesting/src/sim/mod.rs
```
Expected: `read_trade_records` at ~line 1487 and `read_total_trade_count` somewhere nearby.
- [ ] **Step 2: Add `read_per_backtest_trade_counts`**
```rust
// crates/ml-backtesting/src/sim/mod.rs — in impl LobSimCuda { ... }
// Place immediately AFTER read_total_trade_count or read_trade_records.
/// Read per-backtest cumulative trade counters. Returns a Vec of length
/// `n_backtests` where entry `i` is the total number of closed-trade
/// records ever pushed to backtest `i`'s ring (may exceed TRADE_LOG_CAP).
///
/// Used by alpha_rl_train to snapshot the eval-phase boundary per
/// account: `head_before_per_b[i] = read_per_backtest_trade_counts()[i]`
/// at eval start, then post-eval `head_after_per_b[i] - head_before_per_b[i]`
/// is the true eval-phase trade count for backtest `i`.
pub fn read_per_backtest_trade_counts(&self) -> Result<Vec<u32>> {
let mut heads = vec![0u32; self.n_backtests];
self.stream.synchronize()
.context("sync before reading trade_log_head_d")?;
self.stream.memcpy_dtoh(&self.trade_log_head_d, heads.as_mut_slice())?;
Ok(heads)
}
```
- [ ] **Step 3: cargo check**
```bash
SQLX_OFFLINE=true cargo check -p ml-backtesting
```
Expected: clean.
- [ ] **Step 4: Commit**
```bash
git add crates/ml-backtesting/src/sim/mod.rs
git commit -m "feat(lobsim): read_per_backtest_trade_counts (E.1)"
```
---
## E.2 — Lobsim API: `read_trade_records_all`
**Files:**
- Modify: `crates/ml-backtesting/src/sim/mod.rs`
- [ ] **Step 1: Add `read_trade_records_all`**
```rust
// crates/ml-backtesting/src/sim/mod.rs — immediately after read_per_backtest_trade_counts
/// Read all backtests' trade record rings. Returns Vec<Vec<TradeRecord>>
/// of length `n_backtests`; entry `i` is the most recent up-to-
/// TRADE_LOG_CAP records for backtest `i` (oldest dropped on wrap).
///
/// Single DtoH of the full trade_log_d tensor + per-backtest head
/// values, split into per-account slices. Cost: O(b_size × cap × 40 B)
/// — at b=1024, cap=4096 this is 167 MB transferred once.
///
/// Callers compose this with `read_per_backtest_trade_counts()` to
/// determine each account's slice into eval-phase-only records.
pub fn read_trade_records_all(
&self,
) -> Result<Vec<Vec<crate::order::TradeRecord>>> {
self.stream.synchronize()
.context("sync before reading trade logs")?;
let mut heads = vec![0u32; self.n_backtests];
self.stream.memcpy_dtoh(&self.trade_log_head_d, heads.as_mut_slice())?;
let rec_bytes = crate::lob::TRADE_RECORD_BYTES;
let cap = crate::lob::TRADE_LOG_CAP;
let mut raw = vec![0u8; self.n_backtests * cap * rec_bytes];
self.stream.memcpy_dtoh(&self.trade_log_d, raw.as_mut_slice())?;
let mut out = Vec::with_capacity(self.n_backtests);
for b in 0..self.n_backtests {
let head = heads[b] as usize;
let n_to_read = head.min(cap);
let off = b * cap * rec_bytes;
let mut records = Vec::with_capacity(n_to_read);
for i in 0..n_to_read {
let rec_off = off + i * rec_bytes;
let r: crate::order::TradeRecord =
bytemuck::pod_read_unaligned(&raw[rec_off..rec_off + rec_bytes]);
records.push(r);
}
out.push(records);
}
Ok(out)
}
```
- [ ] **Step 2: cargo check**
```bash
SQLX_OFFLINE=true cargo check -p ml-backtesting
```
Expected: clean.
- [ ] **Step 3: Commit**
```bash
git add crates/ml-backtesting/src/sim/mod.rs
git commit -m "feat(lobsim): read_trade_records_all (E.2)"
```
---
## E.3 — `TRADE_LOG_CAP` bump
**Files:**
- Modify: `crates/ml-backtesting/src/lob/mod.rs`
- [ ] **Step 1: Verify the constant is referenced nowhere as hardcoded `1024`**
```bash
grep -rn "1024" crates/ml-backtesting/src/ | grep -iE "trade_log|TradeLog" | head -10
```
Expected: only `pub const TRADE_LOG_CAP: usize = 1024;` itself. Any other hits are flagged for review.
- [ ] **Step 2: Bump the constant**
```rust
// crates/ml-backtesting/src/lob/mod.rs:20
// BEFORE:
pub const TRADE_LOG_CAP: usize = 1024;
// AFTER:
pub const TRADE_LOG_CAP: usize = 4096; // 2026-05-31: bumped from 1024 to prevent
// eval-phase trade ring wrap at b=1024 cluster scale
// (typical eval trades/account ≈ 342, peak ~1000;
// 4096 gives 4× headroom). Memory: b × cap × 40 B
// = 1024 × 4096 × 40 = 167 MB on GPU (was 41 MB).
// See spec 2026-05-31-eval-summary-trade-aggregation.
```
- [ ] **Step 3: cargo build (to catch any test asserting the old value)**
```bash
SQLX_OFFLINE=true cargo build -p ml-backtesting
SQLX_OFFLINE=true cargo build -p ml-alpha
```
Expected: clean.
- [ ] **Step 4: Run existing lobsim tests to verify no regression**
```bash
SQLX_OFFLINE=true cargo test -p ml-backtesting --lib 2>&1 | tail -10
```
Expected: all tests pass (or only-failing tests are known-failing for unrelated reasons).
- [ ] **Step 5: Commit**
```bash
git add crates/ml-backtesting/src/lob/mod.rs
git commit -m "feat(lobsim): bump TRADE_LOG_CAP 1024 → 4096 for eval headroom (E.3)"
```
---
## E.4 — alpha_rl_train.rs: replace slicing block with aggregation
**Files:**
- Modify: `crates/ml-alpha/examples/alpha_rl_train.rs`
This is the core fix. Replace the existing broken block with per-account aggregation that correctly isolates eval-phase trades.
- [ ] **Step 1: Locate the existing block**
```bash
grep -nE "head_before_eval|read_trade_records\(0\)|trade log wrapped" crates/ml-alpha/examples/alpha_rl_train.rs
```
Expected: the block lives roughly at lines 743 (pre-eval snapshot) and 1402-1413 (post-eval slicing).
- [ ] **Step 2: Replace pre-eval snapshot (around line 743)**
```rust
// BEFORE:
let head_before_eval = sim
.read_total_trade_count()
.context("read trade count pre-eval")?;
eprintln!(
"── eval phase: {} steps on {} held-out files (trade-record checkpoint head={}) ──",
cli.n_eval_steps, eval_files.len(), head_before_eval
);
// AFTER:
let head_before_per_b = sim
.read_per_backtest_trade_counts()
.context("snapshot per-backtest head pre-eval")?;
let head_before_min = head_before_per_b.iter().min().copied().unwrap_or(0);
let head_before_max = head_before_per_b.iter().max().copied().unwrap_or(0);
let head_before_sum: u64 = head_before_per_b.iter().map(|&h| h as u64).sum();
eprintln!(
"── eval phase: {} steps on {} held-out files; pre-eval head per b_size={} accounts: \
min={} max={} sum={} ──",
cli.n_eval_steps, eval_files.len(), head_before_per_b.len(),
head_before_min, head_before_max, head_before_sum
);
```
- [ ] **Step 3: Replace the post-eval slicing block (around lines 1402-1413)**
```rust
// BEFORE:
let all_records = sim
.read_trade_records(0)
.context("read trade records post-eval")?;
let head_before_usize = head_before_eval as usize;
let eval_records: Vec<_> = if all_records.len() > head_before_usize {
all_records[head_before_usize..].to_vec()
} else {
eprintln!(
"warning: trade log wrapped — head_before={} but only {} records readable",
head_before_usize, all_records.len()
);
all_records
};
// AFTER:
let all_per_b = sim
.read_trade_records_all()
.context("read all backtests' trade records post-eval")?;
let head_after_per_b = sim
.read_per_backtest_trade_counts()
.context("snapshot per-backtest head post-eval")?;
let cap = ml_backtesting::lob::TRADE_LOG_CAP as u32;
let mut eval_records: Vec<ml_backtesting::order::TradeRecord> = Vec::new();
let mut n_eval_trades_dropped: u64 = 0;
let mut n_pre_eval_wrapped: u64 = 0;
let mut n_eval_trades_seen: u64 = 0;
for (b, records) in all_per_b.iter().enumerate() {
let head_before = head_before_per_b[b];
let head_after = head_after_per_b[b];
let eval_count_total = head_after.saturating_sub(head_before);
n_eval_trades_seen += eval_count_total as u64;
// The ring's contents cover cumulative-stream indices
// [max(0, head_after - cap), head_after).
let ring_start_stream_idx = head_after.saturating_sub(cap);
if head_before < ring_start_stream_idx {
// Some pre-eval trades wrapped out before eval — informational.
n_pre_eval_wrapped += (ring_start_stream_idx - head_before) as u64;
}
if eval_count_total > cap {
// Some eval trades wrapped during eval.
n_eval_trades_dropped += (eval_count_total - cap) as u64;
}
// Slice the ring to eval-only.
// ring index of first eval trade = head_before ring_start_stream_idx
// (clamped at 0 if all pre-eval already wrapped out).
let eval_start_in_ring = head_before
.saturating_sub(ring_start_stream_idx) as usize;
if eval_start_in_ring < records.len() {
eval_records.extend_from_slice(&records[eval_start_in_ring..]);
}
}
if n_eval_trades_dropped > 0 {
eprintln!(
"warning: {} eval trades wrapped out across {} accounts (TRADE_LOG_CAP={}); \
summary based on {} captured eval trades",
n_eval_trades_dropped, all_per_b.len(), cap, eval_records.len()
);
}
if n_pre_eval_wrapped > 0 {
eprintln!(
"info: {} pre-eval trades had already wrapped before eval phase \
(no effect on eval summary)",
n_pre_eval_wrapped
);
}
eprintln!(
"eval phase trade accounting: n_eval_trades_seen={} n_captured={} n_dropped={}",
n_eval_trades_seen, eval_records.len(), n_eval_trades_dropped
);
```
- [ ] **Step 4: Augment `eval_summary.json` with the new fields**
```rust
// The existing code is:
let eval_summary = ml_backtesting::artifacts::compute_summary(&eval_records, &pnl_curve);
// ... serde_json::to_writer_pretty writes eval_summary directly to file.
// Replace the writer block (alpha_rl_train.rs around line 1440) with:
let aggregated_summary = serde_json::json!({
// Original compute_summary fields:
"n_trades": eval_summary.n_trades,
"total_pnl_usd": eval_summary.total_pnl_usd,
"profit_factor": eval_summary.profit_factor,
"sharpe_ann": eval_summary.sharpe_ann,
"max_drawdown_usd": eval_summary.max_drawdown_usd,
"win_rate": eval_summary.win_rate,
// New aggregation-aware fields:
"n_eval_trades_seen": n_eval_trades_seen,
"n_eval_trades_dropped": n_eval_trades_dropped,
"n_pre_eval_trades_wrapped": n_pre_eval_wrapped,
"b_size": cli.n_backtests,
});
let eval_summary_path = cli.out.join("eval_summary.json");
let f = std::fs::File::create(&eval_summary_path)
.with_context(|| format!("create {}", eval_summary_path.display()))?;
serde_json::to_writer_pretty(f, &aggregated_summary)
.with_context(|| format!("write {}", eval_summary_path.display()))?;
eprintln!("eval summary written: {}", eval_summary_path.display());
// Update the eprintln summary line to include new fields:
eprintln!(
"eval summary: n_trades={} pnl_usd={:.2} pf={:.3} sharpe_ann={:.3} \
max_dd_usd={:.2} win_rate={:.3} | seen={} dropped={} b={}",
eval_summary.n_trades, eval_summary.total_pnl_usd, eval_summary.profit_factor,
eval_summary.sharpe_ann, eval_summary.max_drawdown_usd, eval_summary.win_rate,
n_eval_trades_seen, n_eval_trades_dropped, cli.n_backtests
);
```
- [ ] **Step 5: cargo check + build**
```bash
SQLX_OFFLINE=true cargo build --release --example alpha_rl_train -p ml-alpha
```
Expected: clean (warning-free for the eval block).
- [ ] **Step 6: Commit**
```bash
git add crates/ml-alpha/examples/alpha_rl_train.rs
git commit -m "feat(rl): eval_summary aggregates across all b_size accounts (E.4)"
```
---
## E.5 — Local validation
**Files:**
- (No new files; runs existing alpha_rl_train binary)
- [ ] **Step 1: Run b=16, 1k train + 250 eval smoke**
```bash
mkdir -p /tmp/eval-agg-test
rm -rf /tmp/eval-agg-test/*
SQLX_OFFLINE=true target/release/examples/alpha_rl_train \
--n-steps 1000 --n-eval-steps 250 \
--fold-idx 1 --n-folds 3 \
--mbp10-data-dir test_data/futures-baseline/ES.FUT \
--predecoded-dir test_data/futures-baseline/ES.FUT \
--out /tmp/eval-agg-test \
--instrument-mode all \
--n-backtests 16 \
--log-every 100 \
--seed 16962 \
> /tmp/eval-agg-test/stderr.log 2>&1
echo "exit=$?"
```
- [ ] **Step 2: Inspect the new pre-eval boundary log line**
```bash
grep -E "pre-eval head per b_size" /tmp/eval-agg-test/stderr.log
```
Expected: a line like `── eval phase: 250 steps on 3 held-out files; pre-eval head per b_size=16 accounts: min=X max=Y sum=Z ──`. `sum` should be roughly the previous run's `head=600` order of magnitude (depending on policy behavior).
- [ ] **Step 3: Inspect the eval summary**
```bash
cat /tmp/eval-agg-test/eval_summary.json | jq .
grep "eval summary:" /tmp/eval-agg-test/stderr.log
```
Expected JSON shape:
```json
{
"n_trades": <number, expect > 50 (was 58 single-account); roughly 16 × eval_avg_per_account 200-400>,
"total_pnl_usd": <real number>,
"profit_factor": <real number>,
"sharpe_ann": <real number>,
"max_drawdown_usd": <real number>,
"win_rate": <real number>,
"n_eval_trades_seen": <number, n_trades>,
"n_eval_trades_dropped": 0,
"n_pre_eval_trades_wrapped": 0,
"b_size": 16
}
```
Assert:
- `n_trades > 50` (previous run was 58 single-account; aggregated should be 410× higher)
- `n_eval_trades_dropped == 0` (TRADE_LOG_CAP=4096 ≫ per-account count at b=16)
- `n_trades <= n_eval_trades_seen` (captured ≤ seen)
- [ ] **Step 4: Compare to previous (pre-fix) run for sanity**
```bash
# Previous (pre-fix, from earlier session):
# n_trades=58 pnl_usd=-17287.56 pf=0.778 sharpe=-2.518 max_dd=23975.09 wr=0.414
# Current expectation: n_trades much larger, wr should be reasonable
# (0.30-0.50 range typical for this seed), pnl scales with account count.
diff <(jq -S . /tmp/eval-agg-test/eval_summary.json) <(jq -S . /tmp/eval-math-test/eval_summary.json 2>/dev/null) || true
```
The diff is for informational comparison; expected to show n_trades / pnl change significantly.
- [ ] **Step 5: Commit any inline tweaks discovered during validation**
If the smoke surfaces a small issue (e.g., the `eval_start_in_ring` clamp needs adjusting), make a small follow-up commit:
```bash
git diff
# ... fix ...
git add crates/ml-alpha/examples/alpha_rl_train.rs
git commit -m "fix(rl): <specific tweak> from local validation (E.5)"
```
---
## E.6 — Cluster validation
- [ ] **Step 1: Push**
```bash
git push -u origin ml-alpha-eval-summary-fix
```
- [ ] **Step 2: Submit cluster run (same config as v11 for direct comparison)**
```bash
scripts/argo-alpha-rl.sh \
--sha $(git rev-parse --short HEAD) \
--branch ml-alpha-eval-summary-fix \
--gpu-pool ci-training-l40s \
--n-steps 20000 --n-backtests 1024 --seed 16962 \
--instrument-mode front-month \
--fold-idx 1 --n-folds 3 --n-eval-steps 5000
```
Note: this run uses the v11 cluster config exactly so eval_summary metrics are directly comparable.
- [ ] **Step 3: Monitor**
While running: every 510 min, sample the trainer pod for progress:
```bash
kubectl logs -n foxhunt alpha-rl-<wf> -c main --tail=5
```
- [ ] **Step 4: After completion, verify eval_summary.json on PVC**
```bash
# Spin up cleanup pod (PVC unlocked post-completion)
kubectl run e6-investigator -n foxhunt --image=alpine:latest --restart=Never \
--overrides='{"spec":{"containers":[{"name":"main","image":"alpine:latest","command":["sh","-c","apk add --no-cache jq; sleep 600"],"volumeMounts":[{"name":"feature-cache","mountPath":"/feature-cache"}]}],"volumes":[{"name":"feature-cache","persistentVolumeClaim":{"claimName":"feature-cache-pvc"}}]}}'
sleep 15
kubectl exec -n foxhunt e6-investigator -- cat /feature-cache/alpha-rl-runs/<sha>/fold1/eval_summary.json
```
Expected:
- `n_trades` in the **hundreds of thousands** (was 1024 in v11)
- `n_eval_trades_seen` > 300k
- `n_eval_trades_dropped` == 0 (TRADE_LOG_CAP=4096 sufficient)
- `b_size` == 1024
- `total_pnl_usd`, `profit_factor`, `max_drawdown_usd`, `win_rate` are now **interpretable** values across the whole eval phase
- [ ] **Step 5: Cleanup investigator pod**
```bash
kubectl delete pod e6-investigator -n foxhunt --wait=false
```
- [ ] **Step 6: Save a pearl documenting the bug class + fix**
```markdown
# /home/jgrusewski/.claude/projects/-home-jgrusewski-Work-foxhunt/memory/pearl_eval_summary_aggregation_bug.md
---
name: pearl-eval-summary-aggregation-bug
description: eval_summary.json was computed on 1024 single-backtest records mixing train+eval trades, not aggregated across b_size; v10/v11 fold-1 comparison was meaningless.
metadata:
type: pearl
---
[body documenting the 3 flaws + fix + cluster-vs-local validation numbers]
```
- [ ] **Step 7: Update MEMORY.md index**
Add a one-line entry under "Patterns / Pearls" linking to the new pearl.
- [ ] **Step 8: Commit pearl + index update**
```bash
# (pearl files live outside the repo; the MEMORY.md index update is also outside)
# No commit needed in the repo for the pearl.
```
---
## Self-review checklist
- [ ] E.1 and E.2 lobsim methods both call `self.stream.synchronize()` before DtoH
- [ ] E.3 bump didn't break any existing test that hardcodes 1024
- [ ] E.4 eval slice handles three edge cases:
- All-eval-fit: `head_after - head_before ≤ cap` → no drop
- Eval-wraps: `head_after - head_before > cap` → drop counted
- Pre-eval-wraps: `head_before < head_after - cap` → some pre-eval already gone, doesn't affect eval slice
- [ ] No `unwrap()` in new code
- [ ] Eval summary keys preserved (n_trades, total_pnl_usd, profit_factor, sharpe_ann, max_drawdown_usd, win_rate) for downstream consumers
- [ ] New fields are additive
## Done means
- E.1E.5 implemented + committed
- Local smoke shows `n_trades > 50` at b=16 (was 58 single-account; expect 200-400 aggregated)
- Cluster run produces `eval_summary.json` with `n_trades` in hundreds of thousands and `n_eval_trades_dropped == 0`
- Pearl saved + MEMORY.md updated
- Branch ready to merge into `ml-alpha-regime-observer`

File diff suppressed because it is too large Load Diff

View File

@@ -0,0 +1,463 @@
# Determinism Foundation Implementation Plan
> **For agentic workers**: REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development or superpowers:executing-plans to implement this plan task-by-task. Phase 1 is currently DISPATCHED separately (coder agent af76b910). Phases 2-6 execute after Phase 1 lands. Steps use checkbox (`- [ ]`) syntax for tracking.
**Goal**: Make foxhunt ml-alpha training bit-deterministic (rel-tol 1e-5) across same-seed runs, on a per-step basis, with a permanent CI guard preventing regression. Establish a contract: every dev cluster smoke and every local mid-smoke is reproducible.
**Architecture**: Localize → Fix → Test → Toggle → Guard. Each phase has explicit acceptance criteria and produces an artifact that the next phase consumes.
**Tech stack**: CUDA kernels (sm_86/sm_89/sm_90), cudarc + cuBLAS + cuDNN, Rust 1.85, Argo Workflows, foxhunt invariant test harness.
**Linked spec**: `docs/superpowers/specs/2026-06-02-determinism-foundation.md`
---
## §0 Phase status at plan creation
- **Phase 1 (diagnose)**: DISPATCHED as background coder agent af76b910 — instrumenting 15 checksums + running diagnostic. Output: `docs/superpowers/notes/2026-06-02-determinism-phase1-diagnosis.md` identifying ONE first-divergent component.
- **Phase 2 (fix)**: this plan covers all 6 candidate causes; ship only the relevant subsection based on Phase 1 output.
- **Phase 3 (test infra)**: independent of Phase 1; can ship in parallel.
- **Phase 4 (dev/prod toggle)**: independent; ships after Phase 2 fix lands.
- **Phase 5 (CI guard)**: independent; ships after Phase 3 test infra.
- **Phase 6 (validation)**: gates merge to main; runs entire suite.
## Phase 1 — already dispatched
See coder agent af76b910 output. Once it completes, read `docs/superpowers/notes/2026-06-02-determinism-phase1-diagnosis.md` for the CULPRIT identification, then jump to the matching §2.X subsection below.
---
## Phase 2 — Apply targeted fix based on Phase 1 culprit
### §2.A — IF Phase 1 finds: encoder_output diverged first (Mamba2 culprit)
**Files**:
- Modify: `crates/ml-alpha/cuda/mamba2_selective_scan_fwd.cu`
- Modify: `crates/ml-alpha/cuda/mamba2_selective_scan_bwd.cu` (likely same fix applies)
- Possibly: `crates/ml-alpha/cuda/mamba2_block.cu`
- Test: `crates/ml-alpha/tests/mamba2_determinism.rs` (NEW)
#### §2.A.1 — Reproduce mamba2 non-determinism in isolation
- [ ] **Step 1**: Write `tests/mamba2_determinism.rs` that calls mamba2 forward kernel twice with same input + same weights + same seed → checksum compare
- [ ] **Step 2**: Run it. Verify it reproduces the non-determinism (not just a downstream effect)
- [ ] **Step 3**: If it does NOT reproduce, the bug is downstream of mamba2 — re-read Phase 1 output, the culprit is elsewhere
#### §2.A.2 — Audit chunk-parallel scan for non-deterministic ordering
- [ ] **Step 4**: Read `mamba2_selective_scan_fwd.cu` paying attention to:
- Chunk size constant (CHUNK_SIZE)
- Block scheduling order
- Any `atomicAdd` (should be zero per `feedback_no_atomicadd`)
- Any "block 0 OR block 1 writes final" patterns
- The chunk reduction operator (associativity errors compound)
- [ ] **Step 5**: Document audit findings as comments in the file. If clean → bug is in the scan associativity, not block order.
#### §2.A.3 — Ground-truth reference with CHUNK_SIZE=1
- [ ] **Step 6**: Add a `MAMBA2_DETERMINISTIC` compile-time flag that forces CHUNK_SIZE=1 (sequential scan)
- [ ] **Step 7**: With flag ON, run determinism check — should pass (sequential is trivially deterministic)
- [ ] **Step 8**: Measure speed cost of CHUNK_SIZE=1: expect ~10× slower
#### §2.A.4 — Fix: deterministic chunk-reduction with fixed accumulation order
- [ ] **Step 9**: In the chunk reduction operator, ensure accumulation order is FIXED (not parallel-reduce with non-deterministic synchronization):
```cuda
// BAD (non-det if multiple threads accumulate):
__shared__ float acc;
if (threadIdx.x == 0) acc = 0.0f;
__syncthreads();
atomicAdd(&acc, my_partial); // ← non-det order
// GOOD (sequential within block, fixed across blocks):
__shared__ float partials[WARPS_PER_BLOCK];
partials[warp_id] = warp_reduce(my_partial);
__syncthreads();
if (warp_id == 0) {
float total = 0.0f;
for (int i = 0; i < WARPS_PER_BLOCK; ++i) total += partials[i];
// block 0 always writes final
}
```
- [ ] **Step 10**: Apply the fix
- [ ] **Step 11**: Re-run determinism check — should pass with full chunk parallelism
- [ ] **Step 12**: Measure speed cost: expect ~1-3% (negligible)
#### §2.A.5 — Validation
- [ ] **Step 13**: `cargo test mamba2_determinism --release --ignored` passes
- [ ] **Step 14**: Full `scripts/determinism-check.sh` passes (no checksum divergence)
- [ ] **Step 15**: Commit: `fix(mamba2): deterministic chunk-reduction accumulation order`
---
### §2.B — IF Phase 1 finds: q_logits / pi_logits / v_value diverged first (cuBLAS split-K culprit)
**Files**:
- Modify: `crates/ml-alpha/src/networks/q_head.rs` (cuBLAS handle init)
- Modify: `crates/ml-alpha/src/networks/policy_head.rs`
- Modify: `crates/ml-alpha/src/networks/v_head.rs`
- Modify: `crates/ml-alpha/src/networks/dqn.rs` (if cuBLAS handle is shared)
- Modify: `crates/ml-alpha/src/networks/mamba2_block.rs` (likely same)
- Test: `crates/ml-alpha/tests/cublas_determinism.rs` (NEW)
#### §2.B.1 — Confirm TF32 enabled is the cause
- [ ] **Step 1**: Tasks 30/31 (completed) enabled TF32 on cuBLAS handles. TF32 uses split-K with non-deterministic K-split heuristic.
- [ ] **Step 2**: Audit each cuBLAS handle init site (grep for `cublasSetMathMode`)
- [ ] **Step 3**: Document current math modes:
```
q_head: CUBLAS_TF32_TENSOR_OP_MATH ← tasks 30/31
pi_head: CUBLAS_TF32_TENSOR_OP_MATH ← tasks 30/31
mamba2_block: CUBLAS_TF32_TENSOR_OP_MATH ← tasks 30/31
```
#### §2.B.2 — Add deterministic mode toggle
- [ ] **Step 4**: Add helper in `crates/ml-alpha/src/cublas_init.rs` (NEW or in existing setup file):
```rust
pub fn cublas_set_math_mode_deterministic(handle: cublasHandle_t) -> Result<()> {
let pedantic = std::env::var("FOXHUNT_DETERMINISTIC")
.unwrap_or_else(|_| "1".to_string()) == "1";
let mode = if pedantic {
CUBLAS_PEDANTIC_MATH // disables TF32 split-K, FP32 strict
} else {
CUBLAS_TF32_TENSOR_OP_MATH // current default after tasks 30/31
};
unsafe { cublasSetMathMode(handle, mode); }
Ok(())
}
```
- [ ] **Step 5**: Replace every `cublasSetMathMode(handle, CUBLAS_TF32_*)` call with `cublas_set_math_mode_deterministic(handle)`
- [ ] **Step 6**: Verify with grep: no remaining `CUBLAS_TF32_TENSOR_OP_MATH` direct calls
#### §2.B.3 — Validation
- [ ] **Step 7**: Build release. Set `FOXHUNT_DETERMINISTIC=1` and run determinism-check → should pass
- [ ] **Step 8**: Set `FOXHUNT_DETERMINISTIC=0` and re-run → should fail (sanity: the toggle works)
- [ ] **Step 9**: Measure throughput cost on RTX 3050 mid-smoke: `time scripts/local-mid-smoke.sh` with toggle ON vs OFF. Expect 5-15% slower in deterministic mode.
- [ ] **Step 10**: Document the cost in `docs/superpowers/notes/2026-06-02-determinism-speed-cost.md`
#### §2.B.4 — Commit
- [ ] **Step 11**: `fix(cublas): FOXHUNT_DETERMINISTIC env var toggles PEDANTIC vs TF32 math mode`
---
### §2.C — IF Phase 1 finds: cuDNN involved (less likely but possible)
**Files**:
- Same network files as §2.B (if they call cuDNN)
- Modify: `crates/ml-alpha/src/cudnn_init.rs` (if exists; or relevant module)
#### §2.C.1 — Check whether ml-alpha uses cuDNN at all
- [ ] **Step 1**: `grep -rn "cudnn\|CUDNN" crates/ml-alpha/src/ crates/ml-alpha/cuda/` to confirm cuDNN usage
- [ ] **Step 2**: If zero hits → cuDNN not in path → skip §2.C entirely
- [ ] **Step 3**: If cuDNN is used: apply `CUDNN_DEFAULT_MATH` + `CUDNN_DETERMINISTIC` flags via similar toggle to §2.B.2
---
### §2.D — IF Phase 1 finds: custom kernel grad/forward output diverged (block-tree audit)
**Files**:
- Audit: all `.cu` files in `crates/ml-alpha/cuda/`
- Most likely candidates: `rl_*_backward.cu` kernels, `*_reduce.cu` kernels
#### §2.D.1 — Identify the specific kernel(s)
- [ ] **Step 1**: Phase 1 narrows to e.g. `q_grad_checksum diverges first`. The Q gradient is computed by `dqn_distributional_q_backward.cu` (or similar). Identify the exact kernel.
- [ ] **Step 2**: Read the kernel for non-deterministic patterns:
- Any atomicAdd (should be 0 per `feedback_no_atomicadd`)
- Any "if (block_id == 0)" final-write pattern that depends on block scheduling
- Any `volatile` memory access without proper fence
- Any `cudaDeviceSynchronize` inside kernel (very rare)
#### §2.D.2 — Common fix patterns
- [ ] **Step 3**: If the kernel uses "first block to finish writes final" pattern, replace with:
- Two-pass: first pass each block writes its partial to global memory, second pass single-block reads all partials in fixed order
- OR: cooperative groups grid-wide sync (requires sm_70+, available on all foxhunt GPUs)
- [ ] **Step 4**: If the kernel reads stale values due to write/read race: add `__threadfence_system()` + counter pattern with explicit ordering
- [ ] **Step 5**: Apply the fix
- [ ] **Step 6**: Re-run determinism-check
#### §2.D.3 — Validation
- [ ] **Step 7**: GPU oracle test for the specific kernel passes bit-equivalence
- [ ] **Step 8**: Determinism-check passes
- [ ] **Step 9**: Commit: `fix(rl): deterministic block-reduction in <kernel-name>`
---
### §2.E — IF Phase 1 finds: adam_m_sum / adam_v_sum diverged (optimizer iteration order)
**Files**:
- Modify: `crates/ml-alpha/src/optimizers/adamw.rs` (or wherever Adam state is iterated)
#### §2.E.1 — Audit parameter iteration order
- [ ] **Step 1**: Find the AdamW step function. Look for parameter iteration via HashMap, Vec, or BTreeMap.
- [ ] **Step 2**: If HashMap is used (e.g., `HashMap<String, CudaSlice<f32>>`), iteration order is non-deterministic in Rust by design.
- [ ] **Step 3**: Replace HashMap with BTreeMap OR pre-sorted `Vec<(String, CudaSlice<f32>)>` at construction time. Sort by parameter name.
#### §2.E.2 — Validation
- [ ] **Step 4**: Add unit test in `crates/ml-alpha/tests/adam_determinism.rs`: build trainer, do 10 steps, dump Adam m + v, repeat, compare — should be bit-equal
- [ ] **Step 5**: Determinism-check passes
- [ ] **Step 6**: Commit: `fix(adam): deterministic parameter iteration order via BTreeMap`
---
### §2.F — IF Phase 1 finds: replay_sample_indices diverged (PER sampling)
**Files**:
- Modify: `crates/ml-alpha/cuda/rl_per_sample.cu` (or wherever PER sampling happens)
- Modify: `crates/ml-alpha/src/replay/per.rs`
#### §2.F.1 — Check PER RNG seeding
- [ ] **Step 1**: Locate the PER sample kernel. Should use device-side xorshift32 seeded from host (per `pearl_rl_sample_tau_xorshift32`)
- [ ] **Step 2**: Verify the seed is passed via ISV or kernel arg, deterministically advanced per step
- [ ] **Step 3**: If the seed source is `std::time::SystemTime::now()` or `thread_rng()` anywhere → that's the bug
#### §2.F.2 — Fix
- [ ] **Step 4**: Ensure PER seed derives from a deterministic counter (e.g., `cli.seed + step_counter`)
- [ ] **Step 5**: Verify priority update / read ordering uses CudaEvent, not stream-implicit ordering
- [ ] **Step 6**: Re-run determinism-check
#### §2.F.3 — Commit
- [ ] **Step 7**: `fix(per): deterministic sample-RNG seeded from cli.seed + step`
---
## Phase 3 — Test infrastructure (parallel to Phase 2)
### §3.1 — `tests/determinism_invariants.rs`
**Files**:
- Create: `crates/ml-alpha/tests/determinism_invariants.rs`
- [ ] **Task 1**: Write the test file with 3 tests:
```rust
//! GPU-oracle determinism invariants.
//!
//! Test 1: same seed → identical checksums (rel-tol 1e-5)
//! Test 2: different seed → different checksums (sanity check)
//! Test 3: same seed + golden checksum file → match (catches regression)
use anyhow::{Context, Result};
use serde_json::Value;
use std::path::PathBuf;
use std::process::Command;
fn binary_path() -> PathBuf { /* ... */ }
fn data_dir() -> PathBuf { /* ... */ }
fn read_jsonl(path: &Path) -> Result<Vec<Value>> { /* ... */ }
fn extract_checksums(row: &Value) -> Result<Vec<(String, f64)>> { /* ... */ }
#[test]
#[ignore = "requires CUDA + pre-built release binary + MBP-10 test data"]
fn determinism_twin_runs_match() -> Result<()> {
let out_a = run_mid_smoke(seed=42)?;
let out_b = run_mid_smoke(seed=42)?;
let rows_a = read_jsonl(&out_a.join("diag.jsonl"))?;
let rows_b = read_jsonl(&out_b.join("diag.jsonl"))?;
for (i, (a, b)) in rows_a.iter().zip(rows_b.iter()).enumerate() {
let ca = extract_checksums(a)?;
let cb = extract_checksums(b)?;
for ((ka, va), (kb, vb)) in ca.iter().zip(cb.iter()) {
assert_eq!(ka, kb, "checksum key mismatch at row {i}: {ka} vs {kb}");
let rel = ((va - vb) / va.abs().max(1e-12)).abs();
anyhow::ensure!(rel < 1e-5,
"step {i} {ka}: a={va:.6} b={vb:.6} rel_err={rel:.2e}");
}
}
Ok(())
}
#[test]
#[ignore = "requires CUDA + pre-built release binary + MBP-10 test data"]
fn determinism_seed_actually_matters() -> Result<()> {
let out_a = run_mid_smoke(seed=42)?;
let out_b = run_mid_smoke(seed=43)?;
let rows_a = read_jsonl(&out_a.join("diag.jsonl"))?;
let rows_b = read_jsonl(&out_b.join("diag.jsonl"))?;
// After 100 steps, at least one checksum should differ
let row_100_a = &rows_a[99];
let row_100_b = &rows_b[99];
let ca = extract_checksums(row_100_a)?;
let cb = extract_checksums(row_100_b)?;
let any_differs = ca.iter().zip(cb.iter()).any(|((_, va), (_, vb))| {
((va - vb) / va.abs().max(1e-12)).abs() > 1e-3
});
anyhow::ensure!(any_differs, "different seeds produced identical results — seed not used");
Ok(())
}
#[test]
#[ignore = "requires CUDA + pre-built release binary + MBP-10 test data + golden checksum file"]
fn determinism_golden_checksum_match() -> Result<()> {
let out = run_mid_smoke(seed=42)?;
let rows = read_jsonl(&out.join("diag.jsonl"))?;
let last = &rows[rows.len() - 1];
let observed = extract_checksums(last)?;
let golden = read_golden("crates/ml-alpha/tests/fixtures/determinism-golden-seed42.json")?;
for ((k, v_obs), (k_g, v_g)) in observed.iter().zip(golden.iter()) {
assert_eq!(k, k_g);
let rel = ((v_obs - v_g) / v_g.abs().max(1e-12)).abs();
anyhow::ensure!(rel < 1e-5,
"checksum {k} drift from golden: obs={v_obs:.6} gold={v_g:.6} rel_err={rel:.2e}");
}
Ok(())
}
fn run_mid_smoke(seed: u32) -> Result<PathBuf> { /* invoke local-mid-smoke.sh with seed */ }
fn read_golden(path: &str) -> Result<Vec<(String, f64)>> { /* parse JSON */ }
```
- [ ] **Task 2**: Generate golden checksum file `crates/ml-alpha/tests/fixtures/determinism-golden-seed42.json`:
- After Phase 2 fix lands and determinism check passes
- Run mid-smoke with seed=42, extract final checksums, save as golden
- This freezes the post-fix behavior — any future regression detected
- [ ] **Task 3**: Verify tests pass on a clean checkout
---
### §3.2 — Pre-commit hook
**Files**:
- Create: `.claude/hooks/determinism-check.sh` OR `.git/hooks/pre-commit` (foxhunt's convention)
- [ ] **Task 4**: Write a fast 50-step micro-determinism-check that runs on pre-commit:
- Detects `.cu` or `crates/ml-alpha/src/**/*.rs` changes via `git diff --cached --name-only`
- If no relevant changes → exit 0 immediately (no overhead)
- Otherwise: run 50-step mid-smoke twice, compare checksums
- Exit 1 + clear error message if determinism broke
- [ ] **Task 5**: Document the hook in CLAUDE.md
- [ ] **Task 6**: Add `git commit --no-verify` escape hatch documentation (must justify in commit message)
---
### §3.3 — Cluster regression-pair workflow
**Files**:
- Modify: `scripts/argo-alpha-rl.sh`
- Modify: `infra/k8s/argo/alpha-rl-template.yaml`
- [ ] **Task 7**: Add `--regression-pair` flag to `scripts/argo-alpha-rl.sh`:
- When set, submits TWO identical Argo workflows (same SHA, same seed, same params)
- After both complete, compares `eval_summary.json` from both
- Reports `total_pnl_usd` Δ in absolute and relative terms
- Logs warning if Δ > 0.1% (true determinism = bit-exact)
- [ ] **Task 8**: Manual test: run `argo-alpha-rl.sh --regression-pair --sha <post-fix-sha>` → both runs should produce identical eval_summary
---
## Phase 4 — Dev/prod toggle
**Files**:
- Modify: `crates/ml-alpha/src/lib.rs` or `crates/ml-alpha/src/config.rs`
- Modify: `CLAUDE.md`
- [ ] **Task 1**: Add global config struct:
```rust
pub struct DeterminismConfig {
pub enabled: bool,
}
impl DeterminismConfig {
pub fn from_env() -> Self {
let enabled = std::env::var("FOXHUNT_DETERMINISTIC")
.unwrap_or_else(|_| "1".to_string()) // default ON in dev
== "1";
Self { enabled }
}
}
```
- [ ] **Task 2**: Plumb through to: cuBLAS init (§2.B.2), checksum kernel (skip if disabled to save 150us/step), CHUNK_SIZE selection if applicable (§2.A.6)
- [ ] **Task 3**: Production runs (live trading, walk-forward production) set `FOXHUNT_DETERMINISTIC=0` for speed. Document in CLAUDE.md.
- [ ] **Task 4**: Add startup log: print determinism mode + expected speed cost so it's obvious in run logs which mode was active
---
## Phase 5 — CI guard
**Files**:
- Modify: `.gitlab-ci.yml` (per CLAUDE.md, legacy CI exists but disabled; if enabled now use it; otherwise add Argo CI)
- Possibly create: `infra/k8s/argo/determinism-gate-template.yaml`
- [ ] **Task 1**: Add CI job `determinism-gate`:
- Triggers: any push to a branch with `ml-alpha-*` prefix, or any PR to main
- Runs: full mid-smoke (b=128, 2000+500) twice with `FOXHUNT_DETERMINISTIC=1`, then determinism-check.sh
- Fails: if any checksum divergence > rel-tol 1e-5
- Reports: which leaf diverged, at what step
- [ ] **Task 2**: Add pre-train determinism check to `infra/k8s/argo/alpha-rl-template.yaml`:
- New step before main training: run 50-step micro-determinism-check on cluster
- If divergence detected → fail the workflow with clear message
- Cost: ~30 seconds per cluster smoke; saves N×70min if it catches a regression
---
## Phase 6 — Full validation
- [ ] **Task 1**: All 3 tests in `determinism_invariants.rs` pass
- [ ] **Task 2**: `scripts/determinism-check.sh` passes for 5 different seeds: 42, 43, 44, 45, 46
- [ ] **Task 3**: `scripts/argo-alpha-rl.sh --regression-pair` produces identical eval_summary.json (Δ < $1) on one full b=1024 20k+5k cluster smoke
- [ ] **Task 4**: Speed cost documented: measured time of `local-mid-smoke.sh` with `FOXHUNT_DETERMINISTIC=1` vs `=0` on RTX 3050; same comparison on cluster L40S
- [ ] **Task 5**: Create memory pearls:
- `pearl_determinism_root_cause`: the actual culprit + fix
- `pearl_determinism_speed_cost`: measured throughput cost
- `pearl_determinism_contract`: the rule for future commits
- [ ] **Task 6**: Update `MEMORY.md` index with the 3 new pearls
- [ ] **Task 7**: Atomic commit: `feat(rl): deterministic training foundation + CI gate`
---
## Phase 7 — Multi-head policy unblock
Once Phase 6 lands, the multi-head policy spec (`2026-06-02-multi-head-policy-with-r-multiple.md`) is unblocked. Implementation can proceed with confidence that:
- Same-seed runs produce identical results
- Eval pnl differences between iterations are SIGNAL not noise
- Single-seed verdicts are trustworthy (no need for 5-seed averaging)
---
## Failure modes and replanning
### Failure mode 1: Phase 1 finds multiple components diverge simultaneously
**Cause**: A single upstream bug (e.g., mamba2) cascades to all downstream components.
**Response**: Fix the upstream (likely mamba2 per §2.A). The downstream should fix itself once upstream is deterministic.
### Failure mode 2: Phase 2 fix doesn't restore determinism
**Cause**: We fixed one bug but there are multiple sources.
**Response**: Re-run Phase 1 with the partial fix applied. Identify the NEXT first-divergent component. Apply the matching §2.X. Iterate.
### Failure mode 3: Determinism in dev mode but not in prod (FOXHUNT_DETERMINISTIC=0)
**Cause**: Expected — TF32 split-K is non-deterministic by design.
**Response**: This is acceptable. Dev (where verdict matters) is deterministic; prod (where speed matters) is not. Document this.
### Failure mode 4: Speed cost is too high (>20% slowdown)
**Cause**: CHUNK_SIZE=1 mamba2 path OR aggressive cuDNN deterministic mode.
**Response**: Audit the specific bottleneck. Often a deterministic-by-construction reduction can be added without falling back to fully sequential. If unavoidable, accept the cost — verdict trust > 20% throughput.
### Failure mode 5: Phase 1 finds bug is in third-party code (cudarc / cuBLAS internals)
**Cause**: NVIDIA / cudarc library determinism issue.
**Response**: Pin the cudarc version + CUDA toolkit version. File a bug upstream. Use multi-seed averaging as workaround until library fix lands.
---
## Discipline reminders (HARD)
- `feedback_no_atomicadd` — fixes must use block-tree-reduce only, never atomicAdd
- `feedback_no_nvrtc` — all kernels pre-compiled via build.rs
- `feedback_cpu_is_read_only` — checksums computed on GPU, result read on CPU only
- `feedback_no_stubs` — every checksum, every fix, every test must run end-to-end
- `feedback_local_smoke_before_cluster` — Phase 6 task 3 (cluster regression-pair) ONLY after Phase 6 task 1+2 pass locally
- `feedback_no_partial_refactor` — when modifying a cuBLAS handle init, modify ALL of them in the same commit (no half-toggle state)
- `feedback_systematic_debugging` — find root cause before fixing; don't patch symptoms
- `feedback_going_in_circles_pattern` — if Phase 2 fixes don't converge after 3 attempts → re-audit Phase 1 diagnostic
## Acceptance criteria for the whole plan
The plan is COMPLETE when:
1. `scripts/determinism-check.sh` exit 0 on 5 different seeds locally
2. Cluster `--regression-pair` produces eval_summary Δ < $1
3. CI determinism-gate is wired and passing
4. 3 memory pearls saved + MEMORY.md indexed
5. Speed cost documented
6. Atomic commit landed on `ml-alpha-regime-observer` or main
7. Multi-head policy spec confirmed unblocked
Estimated total time across all phases: **3-5 working days** (heavily depending on which §2.X applies — mamba2 fix is largest, cuBLAS toggle is smallest).

View File

@@ -0,0 +1,178 @@
# Next Session — Multi-Head Policy Implementation Plan
**Date**: 2026-06-02 (handoff document)
**Status**: Ready-to-execute plan for next session
**Linked specs**:
- Primary: `docs/superpowers/specs/2026-06-02-multi-head-policy-with-r-multiple.md`
- Parent: `docs/superpowers/specs/2026-06-02-regime-invariance-four-interventions.md`
- Existing #1 plan: `docs/superpowers/plans/2026-06-02-regime-invariance-implementation.md`
**Branch at handoff**: `ml-alpha-regime-observer` at SHA 63fc16f17 (v5.2)
**Memory snapshot**: `pearl_grwwh_eval_catastrophic_collapse` (just saved)
## §0. State at handoff
**alpha-rl-grwwh verdict (just landed)**:
- train +$481M wr 0.312 hold 15.1 — strong train signal
- **eval $160.5M wr 0.264 pf 0.83 sharpe 1.70** — catastrophic generalization failure
- 3rd consecutive eval-collapse on single-policy + Phase 5 hold-bonus architecture
**Empirically validated**: the scaffold-fade approach (v4/v5/v5.1/v5.2) cannot fix the eval generalization gap because the gap is architectural — single policy network with Phase 5 reward shape over-fits to train regime.
**Direction confirmed**: ship multi-head policy + R-multiple reward together per the spec.
## §1. Pre-flight before any code
### §1.1 sp-critical-reviewer on the multi-head spec FIRST
The spec at `2026-06-02-multi-head-policy-with-r-multiple.md` has been written with the FIVE BLOCKERs from the previous v6 spec review explicitly addressed. Before any implementation:
1. Dispatch sp-critical-reviewer agent on the multi-head spec
2. Reviewer should especially verify:
- `PolicyHead` struct layout in `crates/ml-alpha/src/networks/policy_head.rs` (referenced but not personally verified in spec §9)
- `EXPECTED_LEAVES = 679` value referenced in §3.4 still matches current test
- No prior MoE attempts in foxhunt (grep for existing multi-head code)
- The aux-prior-loss math (§1.4) doesn't break PPO surrogate convexity
3. If reviewer BLOCKS, fix and re-review BEFORE implementation
4. If reviewer FIX-AND-PROCEED, apply edits then start §2
### §1.2 Verify the per-step diag / eval_summary discrepancy bug
Per `pearl_grwwh_eval_catastrophic_collapse`: at eval step 4863 the per-step diag showed +$74M but eval_summary showed $160M. This is a $234M discrepancy. Investigate BEFORE shipping a new run that relies on per-step diag for verdict tracking:
1. Pull diag.jsonl rows 22000-22826 from PVC for grwwh (`/feature-cache/alpha-rl-runs/63fc16f17/fold1/diag.jsonl`)
2. Check if the last ~137 eval steps show a catastrophic per-step pnl collapse — if YES, per-step is honest, the agent really did crash at the end
3. If NO collapse visible in per-step → the per-step diag's `realized_pnl_cum_usd` accumulator differs from eval_summary's `total_pnl_usd` calculation — file a separate bug/diagnostic spec
This investigation gates whether per-step diag can be trusted for kill-criteria in §3.1 falsification gates. The reviewer-mandated kill gates rely on per-step signal being authoritative.
## §2. Decide bundling: ship #1 + multi-head together OR sequentially
### Option A: Bundle (single commit, v6.0)
- R-multiple reward + multi-head policy in one commit
- Faster to verdict (one cluster smoke instead of two)
- Higher risk: if eval fails, can't isolate which intervention failed
- Aligns with the §0.5 finding that #1 alone is insufficient
### Option B: Sequential (#1 first, then multi-head)
- Ship #1 R-multiple first, get cluster verdict
- IF #1 eval positive: stop (cheapest win)
- IF #1 eval negative: ship multi-head on top
- More cluster smokes burned
- Cleaner attribution
### Recommendation: **Option A (bundle)**
The grwwh result rules out #1-alone as sufficient. v5.2 with Phase 5 still active failed eval; #1 just removes the contamination (R-multiple replaces hold-bonus). A clean single policy on R-multiple alone would likely still over-fit to fold-1 trend regime — same failure mode, different reward shape.
Multi-head is the architectural fix. Bundle.
## §3. Implementation ordering (assuming Option A bundle)
### Phase A: Foundation (sequential, atomic commit)
1. **Slot retire + reallocate** (`crates/ml-alpha/src/rl/isv_slots.rs`)
- Delete: `RL_SURFER_SCAFFOLD_WEIGHT_INDEX` (753), `RL_SURFER_BREAK_EVEN_WR_INDEX` (754), `RL_SURFER_K_SHARPNESS_INDEX` (755), `RL_SURFER_WARMUP_TRADES_INDEX` (756)
- Allocate (R-multiple, per existing #1 plan): `RL_ATR_EMA_INDEX` (753), `RL_R_MULTIPLE_CLAMP` (754), `RL_ATR_EMA_ALPHA` (755), `RL_ATR_FLOOR` (756)
- Allocate (multi-head, per multi-head spec §2): `RL_POLICY_NUM_HEADS_INDEX` (757), `RL_POLICY_GATING_ENTROPY_FLOOR_INDEX` (758), `RL_POLICY_HEAD_ENTROPY_FLOOR_INDEX` (759), `RL_POLICY_AUX_PRIOR_BETA_INDEX` (760)
- `RL_SLOTS_END = 761`
2. **Delete obsolete code**
- Delete: `crates/ml-alpha/cuda/rl_surfer_scaffold_controller.cu`
- Delete: surfer-scaffold launch + module load in `crates/ml-alpha/src/trainer/integrated.rs` (6 touchpoints per coder analysis)
- Delete: surfer-scaffold bootstrap entries in `with_controllers_bootstrapped`
- Delete: surfer-scaffold registration in `crates/ml-alpha/build.rs`
3. **New kernels (5 total)**
- `rl_atr_ema_update.cu` — ATR EMA producer (R-multiple input)
- `rl_policy_mixture_forward.cu` — K head logits + gate → mixture logits via LogSumExp
- `rl_policy_mixture_backward.cu` — backward through mixture
- `rl_policy_aux_prior_loss.cu` — Σ_k g_k × β × KL(p_k ∥ p_prior_k) + gradient
- Modify: `rl_fused_reward_pipeline.cu` Phase 5 — replace hold-bonus with `r = clamp(Δpnl / max(atr_at_entry, floor), ±5)`
### Phase B: Rust integration
4. **Network refactor** (`crates/ml-alpha/src/networks/policy_head.rs`)
- PolicyHead struct: single Linear → Vec<Linear> for K heads + Linear for gating
- Adam optimizer states per tensor
- Forward pass through new mixture kernels
- Backward pass with per-head + gate gradients
5. **Trainer integration** (`crates/ml-alpha/src/trainer/integrated.rs`)
- Allocate K head weights at construction (init with action-bias per spec §1.3)
- Allocate gating weights
- Add `unit_atr_at_entry_d` per-batch buffer (R-multiple capture)
- Wire 5 new kernel launches in `step_with_lobsim`
- Bootstrap new ISV slots (β=0.05, entropy floors per spec)
### Phase C: Diag + tests
6. **Diag emission** (`build_diag_value`)
- `policy.head_count = K`
- `policy.head_k_entropy[k]` for k in 0..K
- `policy.gate_entropy`
- `policy.gate_probs[k]` for k in 0..K
- `policy.head_pairwise_kl[i,j]` for diagnostic
- `rewards.atr_ema_usd`
- `rewards.r_multiple_last`
- EXPECTED_LEAVES: 679 → 679 + 7 (multi-head) + 2 (R-multiple) 1 (drop pure_pnl_mode/surfer_scaffold_weight) = 687
7. **Test migration**
- `tests/reward_alignment_invariants.rs`: I1-I3 → assert `rewards.atr_ema_usd ≥ 0` and `rewards.r_multiple_last ∈ [-5, +5]`. Delete surfer_scaffold_weight assertions.
- New test: `tests/multi_head_policy_invariants.rs`
- I1: gate_probs[k] in [0, 1] for all k, Σ = 1
- I2: gate_entropy ≥ entropy_floor (with epsilon)
- I3: head_pairwise_kl > 0 (at least one pair distinct)
- I4: each head's entropy ≥ head_entropy_floor
### Phase D: Local validation + cluster
8. **Local validation**
- cargo check + cargo build --release
- `cargo test reward_alignment_invariants multi_head_policy_invariants --ignored --nocapture`
- 200-step local smoke (matches existing test pattern)
- GPU compute-sanitizer racecheck on the new kernels
9. **Cluster smoke**
- Pin SHA via argo-alpha-rl.sh --sha <commit>
- L40S, fold-1, b=1024, 20k train + 5k eval
- Apply §4.1 per-step kill criteria from multi-head spec
- Apply §4.2 verdict gates
## §4. Failure-mode adaptive replanning
If multi-head + R-multiple cluster smoke ALSO produces negative eval:
**Diagnostic step 1**: per-regime eval stratification per spec §4.3. If specialization spread < 0.02 → heads didn't diverge → aux prior β too low. Boost β to 0.15 and retry.
**Diagnostic step 2**: if specialization shows but eval still negative → the heads' priors are wrong (e.g., trend+reversion+flat doesn't cover the actual eval-period regime). Try different prior set (e.g., add momentum-breakout head).
**Diagnostic step 3**: if priors are diverse and eval still negative → architectural ceiling hit. Pivot to:
- Cross-fold training (#3 from parent spec) — train on mixed folds 0/1/2
- Regime features wired into encoder (#2 from parent spec)
- Behavior cloning warm-start from a profitable rule-based bot
**Termination**: if 3 multi-head variants all fail eval, the conclusion is the architecture cannot extract alpha from foxhunt's MBP-10 dataset at b=1024 with 20k step budget. Different data scale or different model architecture required.
## §5. Discipline reminders (DO NOT VIOLATE)
- `feedback_no_feature_flags` — surfer-scaffold deletion in SAME commit as new code, no toggles
- `feedback_single_source_of_truth_no_duplicates` — slot reuse 753-756 with rename, no v2_ suffixes
- `feedback_local_smoke_before_cluster` — Phase D step 8 BEFORE step 9, no exceptions
- `feedback_kill_runs_on_anomaly_quickly` — kill at first anomaly per spec §4.1, don't wait for verdict
- `feedback_no_atomicadd` — all new kernels use block-tree-reduce
- `pearl_argo_branch_ref_trap_strikes_again` — cluster command pins exact SHA, NOT branch tip
- `feedback_going_in_circles_pattern` — if 2 multi-head variants in a row fail with new failure modes, STOP iterating and reconsider foundation
## §6. Session opening checklist (for the next session)
1. Read `pearl_grwwh_eval_catastrophic_collapse` (just saved)
2. Read `docs/superpowers/specs/2026-06-02-multi-head-policy-with-r-multiple.md`
3. Read `docs/superpowers/specs/2026-06-02-regime-invariance-four-interventions.md` (parent)
4. Read this plan
5. Dispatch sp-critical-reviewer on multi-head spec (per §1.1)
6. Resolve reviewer findings
7. Investigate per-step / eval_summary discrepancy (per §1.2)
8. Decide bundling (per §2) — likely Option A
9. Begin Phase A implementation (per §3)
Total estimated session time: 4-6 hours for full bundle, cluster smoke runs overnight if needed.

View File

@@ -0,0 +1,369 @@
# Regime-Invariance — Implementation Plan for the Four Interventions
**Date**: 2026-06-02
**Status**: Plan (task list only; no implementation)
**Branch**: `ml-alpha-regime-observer` (HEAD `63fc16f17`)
**Source spec**: [`2026-06-02-regime-invariance-four-interventions.md`](../specs/2026-06-02-regime-invariance-four-interventions.md)
## Notes for the executor
1. Interventions ship **in series**, one cluster smoke per intervention. Do NOT bundle.
2. Intervention #1 (R-multiple) is the foundational rip-and-replace. The other three build on top.
3. Per `feedback_no_feature_flags`: the surfer-scaffold infra is **deleted** in the same commit that lands R-multiple. No parallel paths.
4. Per `feedback_single_source_of_truth_no_duplicates`: no `v2_`/version suffixes anywhere.
5. Per `feedback_local_smoke_before_cluster`: every `.cu` / trait / buffer change runs the 1k-step local smoke on RTX 3050 Ti before any Argo submit.
6. **Spec inaccuracy correction**: §1.1 of the spec asserts "Read ATR from existing risk-stack signals (already computed for Kelly position sizing)". Verified false against `crates/ml-alpha/cuda/rl_kelly_fraction_controller.cu`: Kelly consumes `win_rate × avg_win/avg_loss` in USD, **not** ATR. There is no existing ATR EMA slot. Intervention #1 must therefore allocate and produce a fresh ATR EMA. The plan reflects this; the spec §7 open question #1 is settled by this finding.
---
## Intervention #1 — R-multiple reward (ships first, single commit)
### Phase A — ISV layout (RL bus surgery)
#### Task 1.1 — Retire surfer-scaffold slots 753-756
- **File**: `crates/ml-alpha/src/rl/isv_slots.rs`
- **Action**: Delete `RL_SURFER_SCAFFOLD_WEIGHT_INDEX` (753), `RL_SURFER_BREAK_EVEN_WR_INDEX` (754), `RL_SURFER_K_SHARPNESS_INDEX` (755), `RL_SURFER_WARMUP_TRADES_INDEX` (756). Remove the §"Adaptive Surfer-Scaffold Shaping" header block (lines 1740-1779). Renumber `RL_SLOTS_END` from 757 → 753.
- **LOC**: 45
#### Task 1.2 — Allocate R-multiple slots
- **File**: `crates/ml-alpha/src/rl/isv_slots.rs`
- **Action**: Reuse the freed 753-756 range with new semantics (no renumbering needed beyond `RL_SLOTS_END`). Allocate:
- `RL_ATR_EMA_INDEX = 753` — per-batch ATR EMA in USD over closed-trade |Δrealized_pnl| (Wiener-α blend, half-life ~50 trades).
- `RL_ATR_EMA_ALPHA_INDEX = 754` — config, default 0.02.
- `RL_ATR_FLOOR_INDEX = 755` — config; floor in USD ≈ 1 ES tick × multiplier × min_size = $12.50.
- `RL_R_MULTIPLE_CLAMP_INDEX = 756` — config; symmetric clamp on R-multiple reward to ±N×ATR, default 5.0.
- Set `RL_SLOTS_END = 757`. Add a §"R-multiple reward" header documenting the design + linkage to spec.
- **LOC**: +50
#### Task 1.3 — Per-account ATR-at-entry CUDA buffer
- **File**: `crates/ml-alpha/src/trainer/integrated.rs`
- **Action**: Add `unit_atr_at_entry_d: CudaSlice<f32>` (sized `b_size`). Allocated in `IntegratedRlTrainer::new`, zero-init alongside `unit_entry_step_d`. Listed in struct, zeroed by `reset_session_state` per `pearl_adaptive_carryover_discipline`.
- **LOC**: +20
### Phase B — CUDA kernel changes
#### Task 1.4 — Delete `rl_surfer_scaffold_controller.cu`
- **File**: `crates/ml-alpha/cuda/rl_surfer_scaffold_controller.cu`
- **Action**: `rm`. Per `feedback_no_feature_flags`, the kernel + its 753-weight consumer go atomically.
- **LOC**: 102
#### Task 1.5 — Rewrite Phase 5 of `rl_fused_reward_pipeline.cu` for R-multiple
- **File**: `crates/ml-alpha/cuda/rl_fused_reward_pipeline.cu`
- **Action**: Replace Phase 5 / 5b entirely (lines 217-272). New structure:
- Capture `unit_atr_at_entry[b]` on flat→positioned transition. Bootstrap = `max(isv[RL_ATR_EMA_INDEX], isv[RL_ATR_FLOOR_INDEX])`.
- At close (`done > 0.5`): `r = clamp(realized_pnl_delta / max(unit_atr_at_entry[b], floor), ±R_MULTIPLE_CLAMP)`.
- Drop entry-cost / short-hold-penalty / long-ride / per-step hold-bonus / inventory-penalty.
- Drop `RL_SURFER_SCAFFOLD_WEIGHT_INDEX` read and all `w *` multipliers.
- Drop `#define`s for `RL_SURFER_SCAFFOLD_WEIGHT_INDEX`, `RL_ENTRY_COST_INDEX`, `RL_SHORT_HOLD_MIN_STEPS_INDEX`, `RL_SHORT_HOLD_PENALTY_INDEX`, `RL_HOLD_BONUS_INDEX`, `RL_INVENTORY_PENALTY_BETA_INDEX`.
- Add `#define`s for `RL_ATR_EMA_INDEX`, `RL_ATR_FLOOR_INDEX`, `RL_R_MULTIPLE_CLAMP_INDEX`.
- Add new kernel arg `float* unit_atr_at_entry` (after `unit_active`).
- **LOC**: 80 / +40 (net 40)
#### Task 1.6 — New ATR EMA producer kernel
- **File**: `crates/ml-alpha/cuda/rl_atr_ema_update.cu` (NEW)
- **Action**: Single-thread (1,1,1)/(1,1,1) launch. On done, read `|realized_pnl_delta|` from the per-batch close, average across the closed-trade batch (single-block tree reduce, no atomics per `feedback_no_atomicadd`), Wiener-α blend into `RL_ATR_EMA_INDEX` with `RL_ATR_EMA_ALPHA_INDEX`. Bootstraps from first-observation per `pearl_first_observation_bootstrap`.
- **LOC**: +90
### Phase C — Trainer wiring
#### Task 1.7 — Drop scaffold controller from trainer
- **File**: `crates/ml-alpha/src/trainer/integrated.rs`
- **Action**: Delete the four `rl_surfer_scaffold_controller`-related blocks:
- Constant `RL_SURFER_SCAFFOLD_CONTROLLER_CUBIN` (line 299-300).
- Struct fields `_rl_surfer_scaffold_controller_module` + `rl_surfer_scaffold_controller_fn` (839-840).
- `load_cubin` / `load_function` blocks (1807-1812).
- Field assignment in struct literal (2959-2960).
- Method `launch_rl_surfer_scaffold_controller` (4261-4275).
- Call site in step pipeline (7757-7762).
- **LOC**: 60
#### Task 1.8 — Wire ATR EMA kernel + buffer into reward pipeline launch
- **File**: `crates/ml-alpha/src/trainer/integrated.rs`
- **Action**:
- Add `RL_ATR_EMA_UPDATE_CUBIN` include + module/function load alongside other reward kernels.
- Add `unit_atr_at_entry_d` to the `rl_fused_reward_pipeline` args block (after `unit_active_d`).
- Add a `launch_rl_atr_ema_update` method.
- Call it immediately AFTER the fused reward pipeline kernel (so the new ATR observation propagates the next time entry capture fires).
- Update `reset_session_state` to zero `unit_atr_at_entry_d` and `RL_ATR_EMA_INDEX` per `pearl_adaptive_carryover_discipline`.
- **LOC**: +90
#### Task 1.9 — ISV bootstrap entries
- **File**: `crates/ml-alpha/src/trainer/integrated.rs`
- **Action**: In the init `isv_seed` table (~line 3919), remove the 4 surfer entries and add R-multiple bootstrap entries:
- `(RL_ATR_EMA_INDEX, 0.0_f32)` (sentinel — first-observation bootstrap)
- `(RL_ATR_EMA_ALPHA_INDEX, 0.02_f32)`
- `(RL_ATR_FLOOR_INDEX, 12.50_f32)` (1 ES tick = $12.50)
- `(RL_R_MULTIPLE_CLAMP_INDEX, 5.0_f32)`
- Also remove the 4 surfer entries from `reset_session_state`'s seed list (~line 3646 area).
- **LOC**: +12 / 10
#### Task 1.10 — Drop scaffold diag emission
- **File**: `crates/ml-alpha/src/trainer/integrated.rs`
- **Action**: Replace the `"surfer_scaffold_weight"` JSONL field (~line 10053) with `"atr_ema_usd"` reading `isv[RL_ATR_EMA_INDEX]` and `"r_multiple_last"` (per-step max abs R-multiple if cheap; otherwise just ATR).
- **LOC**: +5 / 5
### Phase D — Build registration
#### Task 1.11 — Update `build.rs` kernel list
- **File**: `crates/ml-alpha/build.rs`
- **Action**: Remove `"rl_surfer_scaffold_controller"` (line 118). Add `"rl_atr_ema_update"` in its place.
- **LOC**: 0 (1 line swap)
### Phase E — Tests + cubin sanity
#### Task 1.12 — GPU-oracle test for the new fused reward pipeline
- **File**: `crates/ml-alpha/tests/reward_alignment_invariants.rs` (extend existing)
- **Action**: Add a deterministic-seed sub-test asserting:
- After 1000 sim steps, `Pearson(rewards.sum, Δpnl) ≥ 0.7` per `pearl_reward_signal_anti_aligned_with_pnl`.
- On every close event, observed `reward == clamp(Δpnl / max(atr_at_entry, floor), ±5)`.
- `unit_atr_at_entry_d[b]` is overwritten on every flat→positioned transition and unchanged otherwise.
- Per `feedback_no_cpu_test_fallbacks`: oracle is GPU-resident; test reads back via mapped-pinned only.
- **LOC**: +120
#### Task 1.13 — ISV invariants test
- **File**: `crates/ml-alpha/tests/risk_stack_invariants.rs` (extend)
- **Action**: Assert `RL_SLOTS_END == 757` and the four R-multiple slots resolve to 753/754/755/756.
- **LOC**: +15
#### Task 1.14 — Audit-wiring + audit-isv dry run
- **Files**: scripts `scripts/audit-wiring.sh`, `scripts/audit-isv.sh`
- **Action**: Run both; verify no orphan references to `surfer_scaffold_*` survive (consumers, configs, JSONL keys). No new orphans introduced by ATR slots.
- **LOC**: 0 (script run only)
### Phase F — Local validation
#### Task 1.15 — Local 1k-step smoke (RTX 3050 Ti)
- **Command**:
```bash
FOXHUNT_TEST_DATA=test_data/futures-baseline \
SQLX_OFFLINE=true \
cargo test -p ml-alpha --lib --release -- smoke_tests --ignored --nocapture
```
- **Pass gate**: zero kernel asserts, no NaNs in diag JSONL, `atr_ema_usd > 0` after the first close event, `Pearson(reward, Δpnl)` from `reward_alignment_invariants` test ≥ 0.7.
- **LOC**: 0
#### Task 1.16 — Local 5k-step extended smoke
- **Command**: same as 1.15 but with `--n-steps 5000`. Validates fold-1 train wr ≥ 0.30 trend before cluster commit.
- **LOC**: 0
### Phase G — Cluster smoke
#### Task 1.17 — Cluster smoke fold-1 b=1024 20k+5k
- **Command**:
```bash
./scripts/argo-alpha-rl.sh \
--sha <COMMIT_SHA> \
--branch ml-alpha-regime-observer \
--gpu-pool ci-training-l40s \
--batch-size 1024 \
--n-steps 20000 \
--n-eval-steps 5000 \
--n-folds 9 --fold-idx 0
```
- **Pass gate**: spec §1.3 — train wr ≥ 0.30 by step 5k AND eval wr ≥ 0.28 AND eval pnl > 0.
- **Fail gate**: eval pnl < $5M → kill and pivot to Intervention #3.
- **Note**: per `pearl_argo_branch_ref_trap_strikes_again`, pin the exact SHA in the script invocation.
- **LOC**: 0
**Intervention #1 total LOC**: ~+250 / 300 = **net 50** (matches spec §1.1 "Net code REDUCTION").
---
## Intervention #2 — Wire `regime_observer` features into policy network input
Defer until #1 result known. Spec §2 says best deployed AFTER or WITH #3.
#### Task 2.1 — Confirm encoder input dim path
- **File**: `crates/ml-alpha/src/trainer/integrated.rs` (read-only audit)
- **Action**: Locate the encoder input projection used by `rl_encoder_context_broadcast` (build.rs line 94). Identify the current K-dim count it broadcasts (encoder input `[B, K, 56]` per existing comment).
- **LOC**: 0 (research)
#### Task 2.2 — Extend `rl_encoder_context_broadcast.cu` to inject 5 regime features
- **File**: `crates/ml-alpha/cuda/rl_encoder_context_broadcast.cu`
- **Action**: Add a 5-slot tail (cols 56-60) populated from ISV:
- `RL_REGIME_DEAD_ZONE_FLAG_INDEX` (696)
- `RL_REGIME_SESSION_PNL_VARIANCE_EMA_INDEX` (700)
- `RL_REGIME_RECOVERY_FACTOR_INDEX` (699)
- `RL_REGIME_TAIL_EVENT_RECENCY_INDEX` (701) (normalize by `RL_IQN_TAU_TAIL_BOOST_N_WINDOW_INDEX` = 712)
- `RL_EDGE_PH_MEAN_INDEX` (750) (normalize against `RL_EDGE_PH_THRESHOLD_INDEX` = 748)
- Each broadcast as a constant across the K-axis (same value all timesteps in the encoder window).
- **LOC**: +40
#### Task 2.3 — Bump encoder input projection from 56 → 61 dims
- **File**: `crates/ml-alpha/src/trainer/integrated.rs` + any `HIDDEN_INPUT_DIM` const
- **Action**: Identify the constant carrying `56`, replace with `61` (or named const `ENCODER_INPUT_DIM`). Per `feedback_use_consts_not_literals_for_structural_dims`, all consumers must reference the const not the literal.
- **LOC**: +10 / 0
#### Task 2.4 — Expand projection weight matrix `[in, hidden]`
- **File**: `crates/ml-alpha/src/trainer/perception.rs` (encoder init)
- **Action**: Grow input-projection weight matrix from `[56, hidden]` to `[61, hidden]`. New columns Xavier-init small. Update checkpoint format version bump constant.
- **LOC**: +25
#### Task 2.5 — Local 1k-step smoke + 5k extended smoke
- **Command**: same as Task 1.15-1.16. Verifies the wider encoder runs cleanly and doesn't explode action entropy.
- **LOC**: 0
#### Task 2.6 — Cluster smoke
- Same as Task 1.17. Pass gate per spec §5 (eval pnl > 0). Compare action-entropy curve against #1 baseline.
- **LOC**: 0
**Intervention #2 total LOC**: ~+75
---
## Intervention #3 — Cross-fold training (domain randomization)
#### Task 3.1 — Audit walk-forward sampler
- **File**: `crates/ml-alpha/examples/alpha_rl_train.rs` (lines 314-340)
- **Action**: Current sampler splits `n_files` into `n_folds` blocks then picks `train = 0..=fold_idx`, `eval = fold_idx+1`. Map this to a per-batch fold-randomization variant.
- **LOC**: 0 (research)
#### Task 3.2 — Add `--cross-fold-training` CLI flag
- **File**: `crates/ml-alpha/examples/alpha_rl_train.rs`
- **Action**: Add boolean flag `cross_fold_training: bool` defaulting `false`. When set, each rollout batch samples a random training fold uniformly from `0..=fold_idx`; eval stays on `fold_idx + 1` (held-out).
- **LOC**: +15
#### Task 3.3 — Per-batch-element fold sampling in rollout loop
- **File**: `crates/ml-alpha/examples/alpha_rl_train.rs` + downstream sampler in `services/ml_training_service`
- **Action**: When cross-fold mode is active, the per-batch GPU sample-and-gather path (`gpu_sample_and_gather` kernel) takes a randomly-chosen fold's file list per batch element. Verify `gpu_sample_and_gather` already supports per-batch file selection (it does — the dataset gather is per-snapshot-anchor). If not, extend.
- **LOC**: +30
#### Task 3.4 — Verify all training folds load cleanly
- **File**: deployment-time check (no code change)
- **Action**: Per `pearl_mbp10_data_is_structurally_broken` correction (cluster PVC has 148GB coherent 24-month ES MBP-10), confirm 9 quarterly files load without size/shape mismatch. Add an assertion in `alpha_rl_train.rs` that all `train_files` parse before training begins (and abort early if not).
- **LOC**: +10
#### Task 3.5 — Local 1k-step smoke
- **Command**: same as Task 1.15 with `--cross-fold-training` flag.
- **Pass gate**: trade rate doesn't crash to zero; reward Pearson ≥ 0.7 preserved.
- **LOC**: 0
#### Task 3.6 — Cluster smoke
- Same as Task 1.17 with `--cross-fold-training` and `--n-folds 9 --fold-idx 7` (train on folds 0-6, eval on fold 7, held-out tail).
- **LOC**: 0
**Intervention #3 total LOC**: ~+55
---
## Intervention #4 — Mixture-of-experts gated by regime (deferred)
Spec §4.2 marks this as 2-day, 250-LOC effort, defer until #1-#3 validated. Plan stays at task-level only.
#### Task 4.1 — Expert architecture design doc
- **File**: `docs/superpowers/specs/2026-06-XX-moe-regime-gating.md` (CREATE only if Intervention #1-#3 fails)
- **Action**: Capture expert-A / expert-B / gating-net dimensions, per-expert reward shaping plan, training schedule.
- **LOC**: ~150 markdown
#### Task 4.2 — Two-expert MLP heads (parallel to current PPO policy head)
- **File**: `crates/ml-alpha/src/rl/ppo.rs` + new `crates/ml-alpha/cuda/rl_moe_policy_head.cu`
- **Action**: 2 separate `[hidden, N_ACTIONS]` weight matrices, forward + backward kernels. Gating net is a 2-layer MLP on regime features feeding softmax(K=2).
- **LOC**: +250
#### Task 4.3 — Per-expert reward shaping
- **File**: `crates/ml-alpha/cuda/rl_fused_reward_pipeline.cu`
- **Action**: Add per-expert reward path — expert A uses R-multiple-by-hold-duration, expert B uses R-multiple-of-inverse-hold (mean-revert specialty). Gated by `expert_assignment[b]` (sampled from gate softmax at trade open).
- **LOC**: +40
#### Task 4.4 — Trainer wiring + 6 ISV slots
- **File**: `crates/ml-alpha/src/rl/isv_slots.rs`, `crates/ml-alpha/src/trainer/integrated.rs`
- **Action**: 6 new slots (gate temp, K=2 gate weights × 2 layers' hyperparams, per-expert KL trackers). Bumps `RL_SLOTS_END` from 757 → 763.
- **LOC**: +80
#### Task 4.5 — Cluster smoke
- Same as Task 1.17. Spec §4 falsification: eval pnl > 0.
- **LOC**: 0
**Intervention #4 total LOC**: ~+370 (excluding spec doc)
---
## Cross-cutting tasks (apply to all interventions)
#### Task X.1 — Run `scripts/audit-wiring.sh` and `scripts/audit-isv.sh` after each intervention
- **Action**: Confirms no orphan kernel registrations, no dead ISV slots, no hardcoded literals where ISV reads should be used.
- **LOC**: 0
#### Task X.2 — Memory pearl update on cluster verdict
- **File**: `/home/jgrusewski/.claude/projects/-home-jgrusewski-Work-foxhunt/memory/`
- **Action**: After each cluster smoke, append a pearl-style entry to `MEMORY.md` topic file with the verdict (PASS / MARGINAL / FAIL with eval pnl, wr, action entropy). Per `feedback_trust_code_not_docs`, the cluster log is ground truth; pearl summarizes what was measured.
- **LOC**: ~5 per intervention
#### Task X.3 — Pre-commit `precommit` zen + local smoke
- **Action**: Run `mcp__zen__precommit` or equivalent + `feedback_local_smoke_before_cluster` 1k-step smoke before each cluster submission. No exceptions.
- **LOC**: 0
---
## Sequencing decision tree (recap from spec §5)
```
[Ship #1] ──→ cluster smoke ──→ eval pnl > 0? ──yes──→ DONE, deploy
no
[Ship #3] ──→ cluster smoke ──→ eval pnl > 0? ──yes──→ DONE
no
[Ship #2] ──→ cluster smoke ──→ eval pnl > 0? ──yes──→ DONE
no
[Ship #4] ──→ cluster smoke ──→ verdict
no
Termination clause §5
```
---
## Open questions resolved by this plan
1. **ATR source** (spec §7 Q1): No existing Kelly ATR. Allocate fresh `RL_ATR_EMA_INDEX` slot 753 + new `rl_atr_ema_update.cu` producer kernel.
2. **ATR bootstrap** (Q2): Per `pearl_first_observation_bootstrap` + sentinel-zero protocol. Use `max(atr_ema, atr_floor=$12.50)`.
3. **Eval-boundary discipline** (Q3): Per `pearl_adaptive_carryover_discipline`, zero `RL_ATR_EMA_INDEX` and `unit_atr_at_entry_d` in `reset_session_state`. R-multiple denominator falls back to `atr_floor` for the first eval trade — preferred over carrying train-end calibration into a regime-shifted eval window.
4. **Fold sampling weight** (Q4): Equal-weight (uniform) in Intervention #3 v1. Active domain randomization (regime-novelty weighting) is a future #3.1 if Plain #3 underperforms.
---
## Estimated total LOC by intervention
| # | Intervention | Net LOC | Files touched | Cluster smokes |
|---|-------------------------------|---------|----------------|----------------|
| 1 | R-multiple reward | 50 | 8 | 1 |
| 2 | Regime features → encoder | +75 | 3 | 1 |
| 3 | Cross-fold training | +55 | 2 | 1 |
| 4 | Mixture-of-experts (deferred) | +370 | 5 | 1+ |
---
## Files this plan touches (canonical list)
**Intervention #1**:
- `crates/ml-alpha/src/rl/isv_slots.rs`
- `crates/ml-alpha/cuda/rl_fused_reward_pipeline.cu`
- `crates/ml-alpha/cuda/rl_surfer_scaffold_controller.cu` (DELETE)
- `crates/ml-alpha/cuda/rl_atr_ema_update.cu` (CREATE)
- `crates/ml-alpha/src/trainer/integrated.rs`
- `crates/ml-alpha/build.rs`
- `crates/ml-alpha/tests/reward_alignment_invariants.rs`
- `crates/ml-alpha/tests/risk_stack_invariants.rs`
**Intervention #2**:
- `crates/ml-alpha/cuda/rl_encoder_context_broadcast.cu`
- `crates/ml-alpha/src/trainer/integrated.rs`
- `crates/ml-alpha/src/trainer/perception.rs`
**Intervention #3**:
- `crates/ml-alpha/examples/alpha_rl_train.rs`
- `services/ml_training_service/src/k8s_dispatcher.rs` (if any K8s arg surface needs `--cross-fold-training`)
**Intervention #4** (deferred):
- `docs/superpowers/specs/2026-06-XX-moe-regime-gating.md` (CREATE)
- `crates/ml-alpha/src/rl/ppo.rs`
- `crates/ml-alpha/cuda/rl_moe_policy_head.cu` (CREATE)
- `crates/ml-alpha/cuda/rl_fused_reward_pipeline.cu`
- `crates/ml-alpha/src/rl/isv_slots.rs`
- `crates/ml-alpha/src/trainer/integrated.rs`

View File

@@ -0,0 +1,264 @@
# Phase 2 v2 — Proper Dueling Distributional C51 Design
**Date:** 2026-05-29
**Status:** Spec (informed by 2026-05-29 post-mortem + RL literature research)
**Supersedes:** `2026-05-28-state-conditional-q-synthesis.md` (architecturally flawed)
**Baseline:** Plan A v2 (commit `fd3174262`, branch `ml-alpha-plan-a-v2`)
---
## 0. Why this spec exists
The original Phase 2 spec proposed three sub-phases that interacted incorrectly:
- **Phase 2.0** (V envelope clamp): Caused PPO advantage corruption when V was clamped tighter than actual return range. **DROPPED in Plan A v2** — confirmed harmful.
- **Phase 2.1** (Dueling C51): Implementation used scalar V as Bellman bootstrap, making Q-update action-INDEPENDENT (target = projection of `r + γ V(s') - V(s)` is the same for all actions in a given state). Provably collapses to no action preference at convergence. **REVERTED in Plan A v2**.
- **Phase 2.2** (Per-action heterogeneous atoms): Architecturally flawed per RL literature — categorical C51 with per-action atom supports requires a "support-alignment operator" that has no clean mathematical formulation. **NEVER ADOPTED in Plan A v2**.
**RL literature consensus (Perplexity research, 2026-05-29):**
> For C51 distributional Q, use a **single unified canonical support** across all actions. Per-action heterogeneous atoms are "last resort" — naturally handled by **quantile methods (IQN/QR-DQN)**, NOT categorical C51.
This spec proposes a properly-grounded Phase 2 v2 design.
## 1. Goal
Match the SURFER baseline (`dd049d9a4`, wr=0.56 + $10.6M @ 18k) AND beat the TREND baseline (`6d33f18b7`, $40M peak) by addressing the TWO architectural shortcomings the original Phase 2 spec correctly identified:
1. **Cross-regime value mediation** — FLAT-state Q values are small (per-step rewards), OPEN-state Q values are large (cumulative position value). Without a unified baseline, Q-learning can't transition between regimes coherently.
2. **Action-specific resolution** — Long/Short/Hold benefit from fine [-3, +1] resolution; FullFlat/HalfFlat benefit from wider range to capture trade-realization variance.
The **architecturally correct** way to achieve both is **proper dueling distributional Q** with:
- **V_head**: scalar value baseline (action-independent, shared across all actions)
- **A_dist**: per-action advantage distribution on a SHARED atom support with mean-centering identifiability constraint
- **Joint loss**: composed `Q = V + A mean_a A` regressed against TD target
This is exactly what Phase 2.1 intended but implemented incorrectly.
## 2. Architectural Decisions
### 2.A Dueling decomposition (proper)
```
Q(s, a) = V(s) + A(s, a) mean_{a'} A(s, a')
```
Where:
- `V(s)` = scalar state value, action-independent
- `A(s, a)` = advantage distribution (atoms) per action
- `mean_{a'} A(s, a')` = action-mean subtraction, enforces identifiability
**Identifiability**: Without mean subtraction, V and A are unidentifiable (any V+K and AK gives same Q). The subtraction zero-centers A, making the decomposition unique.
### 2.B Shared atom support for A_dist
All actions share the same atom support: `[-A_MAX, +A_MAX]` with Q_N_ATOMS=21 atoms.
**Bootstrap value**: `A_MAX = 3.0` (matches Plan A v2's atom span scale).
**Why shared not per-action**: Per RL literature, advantages have similar magnitudes across actions (centered around 0 by mean-subtraction). The differential is in the magnitudes, NOT the typical range. Shared atoms with reward_scale normalization captures both regimes:
- Long/Short: typical A in [-1, +0.3] (small per-step deviation)
- FullFlat/HalfFlat: typical A in [-2, +2] (realization variance)
Both fit comfortably in [-3, +3] with adequate resolution.
### 2.C Bellman target — joint composed Q loss
The KEY architectural difference from broken Phase 2.1:
**Phase 2.1 (WRONG)**: A trained against scalar `y V(s) = r + γ V(s') V(s)` projected onto A atoms. Target is action-INDEPENDENT.
**Phase 2 v2 (CORRECT)**: Composed Q distribution `Q_dist(s, a_t) = V(s) + A_dist(s, a_t) mean_{a'} A_dist(s, a')` is regressed against TD target distribution. Target is the C51 projection of `r + γ × Q_target_dist(s', a*)` onto shared atom support, where `a* = argmax_a E[Q_target(s', a)]` (Double-DQN).
**This is action-DEPENDENT through a*** → A naturally learns action-specific structure.
Gradient flows through composed Q to BOTH V and A:
- ∂L/∂V: full TD error magnitude
- ∂L/∂A(s, a_t): TD error per-atom contribution
- Mean-subtraction term: gradient distributes negatively to all A(s, a') for a' ≠ a_t
### 2.D V head training — auxiliary regression
V head trains independently via MSE against returns (n-step):
```
loss_V = (V(s) R_n(s))²
```
where `R_n = r + γ R_{n-1} ... + γ^n V_target(s_{t+n})`.
**This is in ADDITION to V's gradient from composed Q loss.** Dual training:
1. From composed Q loss (gradient via dueling decomposition)
2. From auxiliary return regression (PPO advantage signal)
Both push V toward the true value function. The auxiliary loss accelerates V convergence and provides PPO with stable advantages.
### 2.E State-conditional action mask (deferred)
Phase 2.3 (state mask) is intentionally DEFERRED. The mask is independently orthogonal to dueling — can be added later as a wrapper on top of either Phase 2 v1 or v2 architecture.
## 3. Implementation Sub-phases
### Sub-phase 2v2.0 — Plan A v2 baseline merge (already done)
Branch: `ml-alpha-plan-a-v2` (commit `fd3174262`).
- ✅ Distributional C51 Q (dd049d9a4 architecture)
- ✅ V_head for PPO advantage (dd049d9a4 architecture)
- ✅ Bug fixes: VSN stride, h_mag multi-warp, compute_advantage_return done-branch
- ✅ Cluster validated: wr=0.57 + $3.5M @ step 7000 (run `alpha-rl-8ksnj`)
**This is the foundation.** All sub-phases build on this.
### Sub-phase 2v2.1 — V joint training (preparation, 1 day)
**Goal**: V gets gradient from BOTH PPO advantage signal AND the upcoming composed Q loss.
**Files to modify:**
- `crates/ml-alpha/cuda/v_head_fwd_bwd.cu` — V head already done as scalar regression. Keep.
- `crates/ml-alpha/src/trainer/integrated.rs` — ensure V_head.backward fires on EVERY step (already does for PPO).
- New: `crates/ml-alpha/cuda/composed_q_grad_to_v.cu` — kernel that adds ∂L_Q/∂V into V's grad_h_t accumulator.
**Specific changes:**
1. After distributional Q loss kernel fires, add a kernel that:
- Reads per-batch CE-loss gradient w.r.t. composed Q
- Sums contribution to V (constant across atoms): `dL/dV = Σ_z (dL/dQ_atom[z]) × 1` (V adds the same to every atom)
- Adds to `grad_h_t` (V's encoder gradient accumulator)
**Smoke test**: local b=128 × 1000 steps. Pass if `l_v` decreases monotonically and `V_pred` converges to observed return mean within 2σ.
### Sub-phase 2v2.2 — Dueling composed Q kernel (the core change, 2-3 days)
**Goal**: Implement `Q = V + A mean_a A` composition in forward + backward kernels.
**Files to modify:**
- `crates/ml-alpha/cuda/dqn_distributional_q.cu` — reinterpret Q-head output as A_dist (no architectural change in shape, just semantics).
- New: `crates/ml-alpha/cuda/composed_q_fwd.cu` — kernel computing composed Q distribution per batch per action from V + A_dist.
- New: `crates/ml-alpha/cuda/composed_q_bwd.cu` — gradient through composition.
- `crates/ml-alpha/cuda/bellman_target_projection.cu` — UNCHANGED (still projects scalar bootstrap onto shared atoms).
- `crates/ml-alpha/src/rl/dqn.rs` — wire new kernels; `forward` returns composed Q; `backward` routes gradient.
**Specific changes:**
1. **Forward composition** (per batch, per action, per atom):
```
composed_logits[b, a, z] = V(s_b) + A_logits[b, a, z] mean_{a'} A_logits[b, a', z]
```
Note: subtraction happens PER ATOM (z) for proper distributional decomposition.
2. **Action selection** (rollout):
```
Q_value(s, a) = softmax(composed_logits)[z] · atom_z
action = argmax_a E[Q(s, a)] // or Thompson sample
```
At rollout time, computed from composed Q.
3. **Bellman target** (TD bootstrap):
```
a* = argmax_a E[Q_target(s', a)] // Double-DQN
y = r + γ × atom_z_values // bootstrap per atom (same atoms shared)
target_dist = C51 projection of (y, Q_target_softmax(s', a*)) onto shared atoms
```
4. **Loss**: CE between composed_Q_dist(s_t, a_t) and target_dist.
5. **Backward**:
- `∂L/∂V[s_b]`: sum gradient over atoms → V's grad
- `∂L/∂A[b, a_t, z]`: (1 1/N_ACTIONS) × ∂L/∂composed_logits[b, a_t, z]
- `∂L/∂A[b, a, z]` for a ≠ a_t: (1/N_ACTIONS) × ∂L/∂composed_logits[b, a_t, z]
The mean-subtraction Jacobian distributes negative gradient to all OTHER actions, which is the identifiability constraint.
**Smoke test**: local b=128 × 2000 steps. Pass criteria:
- `l_q` (composed Q loss) stable, declining trend after warmup
- `l_v` (V loss) declining
- `q_pi_agree` (action-mean A agreement) > 0.3 sustained (vs Phase 2.1's 0.01 collapse)
- Action distribution diverse (entropy > 1.5)
### Sub-phase 2v2.3 — Cluster validation (1 day)
**Submission:**
```
./scripts/alpha-rl-run.sh --branch ml-alpha-phase2-v2-dueling --steps 20000 --batch 1024 --gpu l40s
```
**Pass criteria:**
- Hold% at step 100 ≥ 50% (sub-90% is acceptable; dueling shifts initial Mamba2 RNG chain consumption)
- wr trajectory ≥ Plan A v2 baseline by step 5000
- pnl_cum profitable (> $0) by step 10000
- entropy stable in [1.5, 2.2] band
- No PPO l_pi explosion (l_pi < 1M by step 20k — Plan A v2 ranges 0-millions, that's normal)
**Failure handling:**
- Hold% crashes to 95%+ early → mean-subtraction term has wrong sign → revisit backward kernel signs
- l_q diverges (> 100 sustained) → composed Q computation has bug → revisit forward kernel
- wr stuck at 0.50 (no improvement) → V is being trained but A isn't getting useful gradient → check V's contribution dominates A's
## 4. Future Phase 2.3 (state mask) — design only, defer implementation
Per the original spec's §2.C. The mask is independent of dueling — it operates on action SELECTION (force illegal actions to -∞ before argmax/Thompson). Can be added as a thin layer over the dueling Q outputs.
**Defer until v2 dueling is validated at cluster.**
## 5. Future Phase 2.4 (per-action atom support) — abandoned
Per RL literature: this is the wrong direction for categorical C51. If we ever need per-action ranges (we may not — dueling A with shared support should suffice), the proper path is IQN, not per-action C51 atoms.
The existing `IqnHead` infrastructure (already in codebase) can be promoted from secondary head to primary Q-learning head if cluster experiments show shared-support C51 insufficient. **Defer until v2 dueling is validated and we have evidence for IQN need.**
## 6. Risks + Mitigations
| Risk | Impact | Mitigation |
|---|---|---|
| Mean-subtraction backward gradient sign wrong | Q diverges; opposite of dueling | Unit test backward against finite-difference oracle BEFORE cluster |
| V's auxiliary regression overwhelms composed Q signal | A doesn't learn action-specific structure | Balance λ_v and λ_q in loss balance controller; both already exist |
| Composed Q forward kernel race condition on `mean A` | NaN; non-determinism | Compute mean in single thread-0 reduce per batch (no atomicAdd, follows `feedback_no_atomicadd` pattern) |
| Cluster Mamba2 RNG init shift makes start state different | Hold% at step 100 won't match dd049d9a4's 90% | Document expected delta; rely on long-horizon convergence to wr=0.56+ |
| Adam optimizer not loading new kernel weights correctly | Silent training failure | Smoke test before cluster: log `w_d.shape` and `grad_w` per Adam step |
## 7. What this spec does NOT do
- No state-conditional action mask (defer to 2.3 if v2 succeeds)
- No per-action heterogeneous atoms (abandoned per research)
- No V envelope clamp (Phase 2.0 — confirmed harmful in Plan A v1)
- No new ISV controllers (P1-P4 from May 29 controller work was based on broken Phase 2.1 — also abandoned)
## 8. Success definition
Phase 2 v2 succeeds if it achieves:
- wr ≥ 0.57 at step 10k cluster (matches dd049d9a4)
- pnl_cum ≥ $5M at step 20k cluster (matches mid-trajectory dd049d9a4)
- `q_pi_agree_ema` ≥ 0.3 sustained (proves Q-π coupling — was 0.01 in broken Phase 2.1)
- No NaN spikes (V envelope clamp not needed because V regresses to bounded scaled targets naturally)
If it surpasses ($10M+, beats trend's $40M plateau): pursued architectural intent achieved.
## 9. Implementation order
```
2v2.0 ✅ Plan A v2 baseline merge (DONE)
↓ (smoke gate: wr=0.57 confirmed @ cluster)
2v2.1 V joint training preparation
↓ (smoke gate: V converges with composed Q gradient)
2v2.2 Dueling composed Q kernel (forward + backward)
↓ (smoke gate: q_pi_agree > 0.3 sustained at local b=128)
2v2.3 Cluster validation
↓ (cluster gate: wr=0.57+, pnl profitable by step 10k)
Future: 2.3 state mask if needed
Future: IQN promotion if per-action resolution proves needed
```
Each sub-phase has a standalone smoke test. Failures halt progression and trigger root-cause analysis.
## 10. Lessons captured
This redesign is informed by an exhausting 2026-05-29 day of cluster failures (alpha-rl-77c54, fcw67, z76v4, rkzkg, tdtqv, zss69, 4sjzw — 7 failed cluster runs). Pearls written:
- `pearl_state_mask_starves_q_for_closing_actions` (D test)
- `pearl_violation_rate_blind_to_hold_attractor` (H controller postmortem)
- `pearl_atomicadd_masks_v_instability` (atomicAdd removal)
- `pearl_dd049d9a4_surfer_baseline_verified` (dd049d9a4 reproduction)
The meta-lesson: **architectural intent ≠ implementation correctness**. The original Phase 2 spec captured real RL design intuitions (V mediation for cross-regime, per-action resolution for varying scales) but implemented them incorrectly. The proper implementation requires careful attention to:
1. Identifiability (mean-subtraction in dueling)
2. Cross-action mathematical coherence (shared atom support)
3. Joint training (composed Q loss, not separate Q and V losses)
Plan A v2 is the proven foundation. Phase 2 v2 is the proper architectural completion of the original intent.

View File

@@ -0,0 +1,166 @@
# Eval-Boundary Normalization-EMA Preservation — Design
**Date:** 2026-05-31
**Status:** Draft (pre-implementation)
**Branch target:** `ml-alpha-eval-boundary-norm-fix` (off `ml-alpha-regime-observer`)
**Related work:** Phase A (eval_diag), E.1-E.5 (eval-summary aggregation), v9 defensive eval-boundary calibration
**Pearls referenced:**
- [`pearl_popart_reset_at_eval_boundary_shocks_normalization`](MEMORY.md) — full diagnosis
- [`pearl_adaptive_carryover_discipline`](MEMORY.md) — being refined here
- F4 envelope design
- `feedback_isv_for_adaptive_bounds`, `feedback_adaptive_not_tuned`
---
## 1. Motivation
Fold-1 eval cluster runs (alpha-rl-8ll7j, alpha-rl-jz48s, alpha-rl-6kghr) all show the same pattern: total eval pnl = $185M, dominated by a catastrophic shock window in the first ~1000 eval steps where popart σ explodes from train-final ~57 to >4400 within 14 steps.
Phase A's `eval_diag.jsonl` enabled the per-step diagnosis. Mechanism (verified empirically):
1. `reset_session_state` at the train→eval fold boundary resets reward-clamp scale EMAs to 0:
- `RL_POS_SCALED_REWARD_MAX_EMA_INDEX → 0`
- `RL_NEG_SCALED_REWARD_MAX_EMA_INDEX → 0`
- `RL_REWARD_CLAMP_CLIP_RATE_EMA_INDEX → 0`
- `RL_POPART_MAX_ABS_REWARD_EMA_INDEX → 0` (F4 envelope)
2. The reward clamp's cap is `cap = pos_max_ema × margin_ratio`. With `pos_max_ema = 0`, cap = 0 — the clamp is functionally disabled. All raw rewards pass through unclamped.
3. Eval's first ~10 trades produce realized pnl swings 10-300× larger than train's typical (Kelly resets to bootstrap=1.0 → larger positions; eval data may also have different volatility).
4. Unclamped large rewards feed `rl_popart_normalize`. F4's envelope ("fast-up, slow-decay" by design — see F4 spec) captures the outliers wholesale. σ_effective = `max(σ_welford, envelope)` becomes envelope-driven.
5. σ explodes 57 → 4471 (78× over train calibration) by eval step 14.
6. PPO surrogate's `A_unnorm = σ × A_norm` scales 78×. Policy gradient magnitudes are catastrophically wrong. Next ~1000 eval steps thrash with mis-scaled updates.
7. Per-trade losses accumulate to dominate the -$185M aggregate.
This is **NOT** model overfitting in the classical sense. The model's policy survives the shock (qpa recovers to 0.99 by eval step 1500). It IS a state-management bug at the regime boundary.
## 2. Goals
### G1 — Preserve normalization-state EMAs across regime boundaries
At `reset_session_state`, do not reset:
- `RL_POS_SCALED_REWARD_MAX_EMA_INDEX`
- `RL_NEG_SCALED_REWARD_MAX_EMA_INDEX`
- `RL_REWARD_CLAMP_CLIP_RATE_EMA_INDEX`
- `RL_POPART_MAX_ABS_REWARD_EMA_INDEX`
These are NORMALIZATION EMAs (calibrate signal scale). The model was trained against rewards at this scale; preserving the scale lets the model interpret eval rewards with the same calibration.
### G2 — Keep resetting predictive EMAs
Continue resetting Kelly EMAs (win_rate, avg_win/loss, cumulative_dones), dd_trip, cooldown, recency, ε_recovery, session_pnl. These predict future behavior from past, and past behavior in train doesn't predict eval regime.
### G3 — Update the pearl rule
Distill the refined rule into `pearl_adaptive_carryover_discipline`: "RESET PREDICTIVE EMAs (Kelly, dd_trip, cooldown). PRESERVE NORMALIZATION EMAs (clamp pos/neg_max, popart envelope, clamp clip-rate). Don't reset blindly."
### G4 — Non-goals
- **Don't fix policy-side overfitting yet.** If the model's policy is genuinely bad on eval data, that's a separate (Layer 3) problem. Fix the boundary-normalization shock first; the resulting eval pnl is the true measure of model quality.
- **Don't introduce a pure-eval mode (frozen weights).** That's a separate, larger refactor — defer until we know if normalization preservation alone is sufficient.
- **Don't widen `TRADE_LOG_CAP` further.** Already at 4096 (E.3), sufficient.
## 3. Design
### 3.1 The change
`crates/ml-alpha/src/trainer/integrated.rs:3231-3239` — remove four lines from `reset_session_state`'s reset list:
```rust
// BEFORE:
// Reward-clamp EMAs — NEW in v9.
(crate::rl::isv_slots::RL_POS_SCALED_REWARD_MAX_EMA_INDEX, 0.0_f32),
(crate::rl::isv_slots::RL_NEG_SCALED_REWARD_MAX_EMA_INDEX, 0.0_f32),
(crate::rl::isv_slots::RL_REWARD_CLAMP_CLIP_RATE_EMA_INDEX, 0.0_f32),
// ...
// Regime observer boundary policy:
// ...
// - POPART_MAX_ABS_REWARD_EMA → 0 (F4 envelope reset)
(crate::rl::isv_slots::RL_POPART_MAX_ABS_REWARD_EMA_INDEX, 0.0_f32),
// AFTER (4 lines removed; comments around the regime-observer block updated):
// Normalization EMAs intentionally NOT reset — these calibrate signal scale.
// The model was trained against rewards at this scale; preserving the scale
// keeps the reward clamp armed (cap = pos_max_ema × ratio) and lets popart
// continue with the same normalization. RESETTING them disabled the clamp
// (cap=0×ratio=0) and let outliers blow up the envelope → σ shock window
// → -$185M eval loss (validated alpha-rl-6kghr 2026-05-31).
```
### 3.2 What stays the same
Everything else in `reset_session_state` is unchanged:
- Layer 1 CMDP summary slots (session_pnl, worst_pnl, dd_triggered, consec_loss, cooldown) — PREDICTIVE, correctly reset
- Layer 4 (Kelly) EMAs (win_rate, avg_win/loss, cumulative_dones, kelly_fraction) — PREDICTIVE, correctly reset
- Layer 3 (Inventory) EMAs (inventory_penalty_β, inventory_variance) — PREDICTIVE, correctly reset
- Defensive warmup arm (`RL_REGIME_TRANSITION_REMAINING_INDEX` = 500 steps)
- Regime observer state (dead_zone flag/duration/timeout, tail recency, recovery_factor, ε_recovery_live, prev_worst_pnl) — PREDICTIVE/transient, correctly reset
- Per-batch buffer memsets (session_pnl_per_batch_d, etc.)
### 3.3 Why not preserve popart.σ / .var / .mean themselves?
popart's Welford state (σ/var/mean) is ALREADY preserved across the boundary — it's not in the reset list. The bug was that env.max (which dominates σ_effective via `max(σ_welford, env.max)`) was being reset to 0, exposing the bare σ_welford = 1. The 4-line fix preserves env.max so σ_effective stays at train-calibrated level.
### 3.4 Why preserve clip_rate_ema?
`RL_REWARD_CLAMP_CLIP_RATE_EMA_INDEX` is the rate at which rewards were being clipped during train. The clamp's margin controller uses this to adjust the margin ratio adaptively (target clip rate ~5%). Resetting it to 0 makes the controller think it's never clipping → grows margin → defeats clamp purpose further. Preserve.
### 3.5 Will preserving NORMALIZATION cause its own problems?
Scenarios to consider:
| Scenario | Preserved env.max | New regime ACTUAL scale | Behavior |
|----------|------------------:|------------------------:|----------|
| Eval ≈ train scale | 57 | 57 | clamp works as in train, σ stable ✓ |
| Eval larger scale | 57 | 200 | clamp clips some real signal, but pos_max_ema adapts up over ~100 steps; σ grows smoothly |
| Eval smaller scale | 57 | 10 | clamp lets everything through (most rewards ≤ cap), pos_max_ema slowly adapts down; σ stays at 57 until rewards consistently smaller |
None of these are catastrophic. The current "reset to 0" scenario is uniquely bad because it produces an unbounded amplification rather than mild miscalibration.
## 4. Validation
### V1 — Local smoke
Run alpha_rl_train with fold-1 walk-forward at b=16, 1k train + 250 eval. Verify:
- Pre-eval pos_max_ema is non-zero (preserved)
- First eval step's clamp cap > 0
- popart σ stays within 5× of train-final value across all 250 eval steps
- eval_summary n_eval_trades_dropped = 0
### V2 — Cluster smoke
Same as v11/Phase A config (fold 1, b=1024, 20k+5k eval, seed 16962). Compare:
- eval_diag.jsonl first 50 steps: popart.σ should stay near train-final (~57), NOT spike to >4000
- eval_summary.json total_pnl_usd: should be dramatically less negative than -$185M (predicted: -$5M to -$30M range, but exact value unknown — that's the residual model overfit)
### V3 — Pearl update
Patch `pearl_adaptive_carryover_discipline` with the predictive-vs-normalization distinction.
## 5. Implementation phases
### Phase 1 — Code change (5 min)
Remove the 4 reset lines from `reset_session_state`. Update the comment to explain the new policy.
### Phase 2 — Local smoke (15 min)
Same fold-1 b=16 smoke as the eval-summary fix. Inspect eval_diag.jsonl popart sigma trajectory.
### Phase 3 — Cluster validation (~95 min)
Push, submit, wait, inspect.
### Phase 4 — Pearl update (5 min)
Refine `pearl_adaptive_carryover_discipline` with the predictive-vs-normalization distinction.
## 6. Open decisions
| # | Question | Recommendation |
|---|----------|---------------|
| 1 | Preserve ALL 4 EMAs, or just the 3 reward-clamp ones (skipping popart envelope)? | **Preserve all 4.** Without env.max, the clamp may suppress the worst overshoots but the envelope itself still re-bootstraps from 0 and can spike on a single large reward. Both layers of defense matter. |
| 2 | Add a controller-level invariant test? | **Yes, V1 local smoke** — assert `popart.σ` after eval step 50 is within 5× of train-final. Cheap to verify. |
| 3 | Update pearl_adaptive_carryover_discipline NOW or wait for cluster confirmation? | **After cluster.** Update the pearl WITH the empirical evidence of the fix working. |
| 4 | Risk that preserving env.max masks true regime shift in eval data? | **Low.** EMAs adapt over time; preserving the train-final value just means eval starts calibrated rather than at 0. If eval has truly larger scale, env.max grows there too, just from 57 instead of 0. |
## 7. Done means
- Lines removed from `reset_session_state`, commit made
- Local b=16 fold-1 smoke shows popart.σ bounded across eval
- Cluster b=1024 fold-1 run completes, eval_diag.jsonl shows σ never exceeds 5× train-final
- eval_summary.json total_pnl significantly less negative than -$185M (number TBD; if still very negative, that's true Layer-3 overfit work)
- Pearl `pearl_adaptive_carryover_discipline` updated with the refined rule
## 8. Risks
- **Risk A (low):** If train-final pos_max_ema was somehow off (e.g. from train tail event), preserving it carries the mis-calibration into eval. Mitigation: EMA decays over eval steps anyway.
- **Risk B (medium):** If eval shock losses come from MULTIPLE compounding factors beyond just clamp-disable (e.g. Kelly reset compounded with mis-calibrated normalization), preserving only normalization may only partially fix. Mitigation: cluster validation will measure exact remaining gap.
- **Risk C (low):** Pearl rule refinement might be over-broad if there are other PREDICTIVE-vs-NORMALIZATION EMAs we haven't catalogued. Mitigation: only apply the rule to the 4 EMAs we're changing today; document the principle but require future cases to verify per-EMA.

View File

@@ -0,0 +1,356 @@
# Eval-Summary Trade Aggregation Fix — Design
**Date:** 2026-05-31
**Status:** Draft (pre-implementation)
**Branch target:** `ml-alpha-eval-summary-fix` (off `ml-alpha-regime-observer`)
**Related:** independent of Phase A (eval-diag emission); both ship in parallel
**Pearls / feedback referenced:**
- `feedback_no_partial_refactor`
- `feedback_single_source_of_truth_no_duplicates`
- `pearl_cluster_log_is_ground_truth`
---
## 1. Motivation
`eval_summary.json` produced by `alpha_rl_train` is wrong at every scale. Cluster v11 (alpha-rl-8ll7j) reported `pnl=+$61,513, max_dd=-$444,512, win_rate=0.217, n_trades=1024` for fold-1 eval. Local b=16 reproduction reported `pnl=-$17,287, max_dd=$23,975, win_rate=0.414, n_trades=58`. **Both values are computed on a mixture of train-phase and eval-phase trades from a SINGLE backtest, not aggregated across all b_size accounts.**
The cluster's improvement-vs-v10 narrative (v10 eval -$507k → v11 eval +$61k, falsely interpreted as a regime-observer success) is invalid for the same reason. Both numbers reflect whatever 1024 trades happened to occupy account-0's ring buffer at end-of-eval.
This bug has been silently corrupting every fold-1 eval result we've produced. Spotting it required Phase A pre-flight investigation; fixing it doesn't require Phase A but should ship alongside.
---
## 2. Root cause (verified locally)
### 2.1 The contaminated slice
`alpha_rl_train.rs:743-1413`:
```rust
let head_before_eval = sim.read_total_trade_count()?; // (A) AGGREGATE across b_size
// (sum of per-account head counters)
// ... eval loop runs ...
let all_records = sim.read_trade_records(0)?; // (B) SINGLE backtest (#0), ≤ TRADE_LOG_CAP entries
let eval_records = if all_records.len() > head_before_usize {
all_records[head_before_usize..].to_vec() // (C) UNREACHABLE: per-account count ≤ 1024 ≤ b_size × per-account
} else {
all_records // (D) ALWAYS FIRES: returns ALL records,
// including train-phase trades
};
```
### 2.2 Three independent flaws
1. **Scale mismatch**: `head_before_eval` is the aggregate trade count across all `b_size` backtests. `all_records.len()` is one backtest. Comparing them is wrong scales. At b=16, aggregate=600 but account-0=58 → branch (D) fires. At b=1024, aggregate=1,203,376 but account-0≤1024 → branch (D) fires.
2. **Branch (D) ignores train/eval boundary**: returns all 58 (local) or 1024 (cluster) records, INCLUDING pre-eval train trades. Eval_summary is computed on a mixed sample.
3. **Single-account sampling**: Even if (1) and (2) were fixed, reading only backtest 0 ignores the other `b_size - 1` accounts. Eval performance should aggregate across the whole batch.
### 2.3 TRADE_LOG_CAP secondary issue
`crates/ml-backtesting/src/lob/mod.rs:20`:
```rust
pub const TRADE_LOG_CAP: usize = 1024;
```
At cluster scale (b=1024, 20k train + 5k eval, ~70 dones/step aggregate ≈ ~0.07 dones/account/step) each account accumulates ~1.4k1.7k trades total. Capacity 1024 means heavy-trading accounts wrap during training, then wrap again during eval — even a correct slicing logic would be unable to isolate eval-only trades for those.
At local b=16, the per-account counts are well under 1024 (~37 train + 21 eval = 58 typical). No wrap. So a per-account aggregation fix is sufficient locally, but capacity may need bumping for cluster.
---
## 3. Goals
### G1 — Aggregate across all backtests
`eval_summary.json` reflects all `b_size` accounts' eval-phase trades, not just account 0.
### G2 — Correct train/eval slicing
Per-account head pointer snapshotted before eval; post-eval slice is `records[head_before_per_b[i]..]` for each `i`, handling wrap correctly.
### G3 — Wrap detection + accurate accounting
When a backtest's pre-eval head was already past TRADE_LOG_CAP, log a warning + count of dropped pre-eval trades. Eval-phase trades that wrap are reported as a separate `n_eval_trades_dropped` field.
### G4 — Capacity bump for cluster scale
TRADE_LOG_CAP raised from 1024 to a value that gives sufficient headroom at b=1024 production scale.
### G5 — Non-goals
- Per-trade phase tagging (option C — adds 5 bytes/record × 1024 accounts × new_cap; deferred to a future spec if G3's wrap warnings prove frequent).
- Train-phase aggregate summary (current `alpha_rl_train_summary.json` is loss-only; trade-based metrics during train are not consumed by any downstream).
---
## 4. Design
### 4.1 New `LobSimCuda` methods
```rust
impl LobSimCuda {
/// Read per-backtest `head` counter (length b_size). Returns the
/// cumulative trade count for each backtest at call time. Used by
/// alpha_rl_train to snapshot the pre-eval boundary per account.
pub fn read_per_backtest_trade_counts(&self) -> Result<Vec<u32>>;
/// Read all backtests' trade records. Returns Vec<Vec<TradeRecord>>
/// of length b_size; entry i is the trade ring contents for backtest
/// i (up to TRADE_LOG_CAP records, oldest dropped on wrap).
///
/// One DtoH for trade_log_head_d + one DtoH for trade_log_d (same
/// cost as read_trade_records(0) since the whole tensor was
/// downloaded anyway — we just keep all the data instead of slicing
/// to one account).
pub fn read_trade_records_all(&self) -> Result<Vec<Vec<crate::order::TradeRecord>>>;
}
```
### 4.2 New `alpha_rl_train.rs` eval boundary handling
```rust
// At eval phase entry (before reset_session_state):
let head_before_per_b: Vec<u32> = sim.read_per_backtest_trade_counts()
.context("snapshot per-backtest head pre-eval")?;
eprintln!(
"── eval phase: {} steps on {} held-out files; pre-eval head per b_size={} accounts: \
min={} max={} sum={} ──",
cli.n_eval_steps, eval_files.len(), head_before_per_b.len(),
head_before_per_b.iter().min().copied().unwrap_or(0),
head_before_per_b.iter().max().copied().unwrap_or(0),
head_before_per_b.iter().map(|&h| h as u64).sum::<u64>()
);
// ... eval loop runs ...
// Post-eval:
let all_per_b = sim.read_trade_records_all()
.context("read all backtest trade records post-eval")?;
let head_after_per_b = sim.read_per_backtest_trade_counts()
.context("snapshot per-backtest head post-eval")?;
let cap = crate::lob::TRADE_LOG_CAP as u32;
let mut eval_records: Vec<crate::order::TradeRecord> = Vec::new();
let mut n_eval_trades_dropped: u64 = 0;
let mut n_pre_eval_wrapped: u64 = 0;
let mut n_eval_trades_seen: u64 = 0;
for (b, records) in all_per_b.iter().enumerate() {
let head_before = head_before_per_b[b];
let head_after = head_after_per_b[b];
let eval_count_total = head_after - head_before;
n_eval_trades_seen += eval_count_total as u64;
// Records in the ring are the LAST min(head_after, cap) trades.
// We want trades whose "index in the cumulative stream" lies in
// [head_before, head_after). The ring's contents cover stream
// indices [max(0, head_after - cap), head_after).
let ring_start_stream_idx = head_after.saturating_sub(cap);
if head_before < ring_start_stream_idx {
// Pre-eval trades for this backtest already wrapped out before
// eval started — purely informational, doesn't affect eval slicing.
n_pre_eval_wrapped += (ring_start_stream_idx - head_before) as u64;
}
// Number of eval trades that wrapped out (lost):
if eval_count_total > cap {
n_eval_trades_dropped += (eval_count_total - cap) as u64;
}
// Slice the ring contents to eval-only:
let eval_start_in_ring = head_before.saturating_sub(ring_start_stream_idx) as usize;
let eval_slice = &records[eval_start_in_ring.min(records.len())..];
eval_records.extend_from_slice(eval_slice);
}
if n_eval_trades_dropped > 0 {
eprintln!(
"warning: {} eval trades wrapped out across {} accounts (cap={}); summary based on \
{} captured eval trades",
n_eval_trades_dropped, all_per_b.len(), cap, eval_records.len()
);
}
if n_pre_eval_wrapped > 0 {
eprintln!(
"info: {} pre-eval trades had wrapped before eval phase (no effect on eval summary)",
n_pre_eval_wrapped
);
}
// Compute summary on aggregated, eval-only records.
let mut pnl_curve = Vec::with_capacity(eval_records.len());
let mut cum = 0.0_f32;
for r in &eval_records {
cum += (r.realised_pnl_usd_fp as f32) / 100.0;
pnl_curve.push(cum);
}
let eval_summary = ml_backtesting::artifacts::compute_summary(&eval_records, &pnl_curve);
// Emit the existing fields plus two NEW fields:
// n_eval_trades_dropped (u64) — total trades lost to ring wrap during eval
// n_eval_trades_seen (u64) — total eval trades that occurred (cumulative_done across b_size)
// These are written into eval_summary.json alongside the existing keys
// via a thin wrapper struct.
```
### 4.3 New `eval_summary.json` schema additions
Current:
```json
{
"n_trades": 1024,
"total_pnl_usd": 61512.78,
"profit_factor": 1.022,
"sharpe_ann": 0.174,
"max_drawdown_usd": 444512.31,
"win_rate": 0.217
}
```
After fix:
```json
{
"n_trades": 348567, // ← aggregate across all backtests, eval-only
"total_pnl_usd": 1234.56,
"profit_factor": 1.085,
"sharpe_ann": 0.42,
"max_drawdown_usd": 12345.67,
"win_rate": 0.39,
"n_eval_trades_seen": 348700, // NEW — true total eval-phase trade count
"n_eval_trades_dropped": 133, // NEW — 0 unless capacity insufficient
"b_size": 1024, // NEW — context for downstream interpretation
"n_pre_eval_trades_wrapped": 5821 // NEW — diagnostic, doesn't affect eval metrics
}
```
`compute_summary` already consumes `&[TradeRecord]` and produces `n_trades, total_pnl_usd, ...` — feeding it the aggregated `eval_records` Just Works for those fields. The four new fields are added by the alpha_rl_train wrapper.
### 4.4 TRADE_LOG_CAP bump
Cluster b=1024, fold-1 eval, 5000 steps observed:
- Total eval dones aggregate ≈ 350,000
- Per account average: ~342
- Per account max (heavy traders): likely 23× mean ≈ 7001000
Setting `TRADE_LOG_CAP = 4096` gives:
- 4× headroom over observed mean
- Memory cost at b=1024: 1024 × 4096 × 40 B = **167 MB GPU** (was 41 MB) — fits comfortably on L40S (48 GB) and H100 (80 GB)
- For train-phase carryover at b=1024 with ~1175 dones/account, no wrap during normal training either
For local smoke at b=16: memory cost = 16 × 4096 × 40 B = 2.6 MB — negligible.
```rust
// crates/ml-backtesting/src/lob/mod.rs
pub const TRADE_LOG_CAP: usize = 4096; // was 1024 (2026-05-31: bumped to prevent eval wrap)
```
### 4.5 Backward compatibility
The two existing methods (`read_trade_records(b)`, `read_total_trade_count()`) stay as-is — they may be used elsewhere (unit tests, debug paths). The new methods are additive.
The existing `alpha_rl_train.rs:1402-1413` block is REPLACED by the new aggregation block; no need for a feature flag.
---
## 5. Validation criteria
### V1 — Local smoke math correctness
At b=16, 1k train + 250 eval, fold 1:
- Pre-eval log line shows `head_before_per_b: min=X max=Y sum=Z` (Z should match the previous `head=Z` log)
- Post-eval `n_eval_trades_seen``eval_dones_step_diag_total - 0` (eval_dones reported per-step in the JSONL with Phase A)
- `n_eval_trades_dropped` = 0 (no wrap at b=16)
- `n_trades` in eval_summary ≈ `n_eval_trades_seen` (no wrap → equal)
### V2 — Train/eval contamination test
At b=16, 1k train + 0 eval:
- No eval phase runs
- eval_summary.json NOT written (existing gate: `n_eval_steps > 0`)
At b=16, 0 train + 250 eval (synthetic edge case):
- Pre-eval head_before_per_b all zeros
- eval_summary.n_trades = sum of (post-eval head_per_b)
- No contamination possible (no train phase)
### V3 — Cluster scale validation
At b=1024, 20k train + 5k eval, fold 1:
- `n_eval_trades_dropped` should be 0 with TRADE_LOG_CAP=4096
- `n_trades` should be O(350,000), not O(1,024)
- `total_pnl_usd` magnitude makes sense (b=1024 × per-account small pnl; should be in $100k-$1M range)
- `win_rate` should NOT be the suspiciously-low 0.217 we observed (which was contaminated)
### V4 — TRADE_LOG_CAP stress test
Run b=1024 with intentionally heavy trading (force flat-frequency policy via seed) and 10k eval steps. Verify `n_eval_trades_dropped > 0` triggers the warning correctly and the summary still produces sensible numbers (just on a subset).
---
## 6. Implementation phases
### Phase 1 — Lobsim API additions (~1 hour)
1. Add `read_per_backtest_trade_counts()` to `LobSimCuda` — single DtoH of `trade_log_head_d` into `Vec<u32>`
2. Add `read_trade_records_all()` — same memcpy as `read_trade_records(0)` but split into per-backtest slices
3. Unit test: push synthetic trades, verify both methods return expected counts/records
### Phase 2 — TRADE_LOG_CAP bump (~5 min)
1. Change `pub const TRADE_LOG_CAP: usize = 1024;``= 4096;`
2. Verify no asserts/tests hardcode 1024 (grep `1024` in lob module)
3. Run existing lobsim unit tests
### Phase 3 — alpha_rl_train aggregation (~30 min)
1. Replace the existing slicing block (lines 1402-1413) with the aggregation logic in §4.2
2. Add the four new fields to a typed struct that augments `compute_summary`'s output
3. Update the eprintln messages
### Phase 4 — Local validation (V1, V2) (~15 min)
1. Run b=16, 1k+250 eval smoke
2. Assert n_trades > 50 (was 58; should now be sum across 16 backtests ≈ 16 × eval_avg ≈ 200-400)
3. Verify n_eval_trades_dropped = 0
4. Visually inspect that win_rate / pf / sharpe values look more reasonable
### Phase 5 — Push for cluster validation (V3) (~minutes)
1. Push
2. Submit alpha-rl run with same config as v11
3. Compare eval_summary.json to v11's — n_trades should be ~340× larger
---
## 7. Open decisions
| # | Question | Recommendation |
|---|----------|---------------|
| 1 | TRADE_LOG_CAP value | **4096** — 4× cluster mean, fits memory budget |
| 2 | Compute_summary signature change for new fields | **No** — wrap output in a thin struct in alpha_rl_train.rs that adds the four new fields. compute_summary stays untouched. |
| 3 | Pre-eval head snapshot stream sync | **Yes** — must sync before reading head_d to avoid stale values mid-step (universal save-rule applies) |
| 4 | Bump train-phase artifact too? | **No**`alpha_rl_train_summary.json` only has losses, no trade metrics. Keep scope tight. |
| 5 | Deprecate `read_trade_records(b)` + `read_total_trade_count()` | **No** — keep for unit tests and future debugging. Only the alpha_rl_train call site changes. |
| 6 | Add per-fold artifact directory cleanup | **No** — out of scope. |
---
## 8. Interaction with Phase A (eval-diag emission)
Phase A and this fix are **independent**:
- Phase A: emits per-step JSONL during eval phase
- This fix: corrects the aggregate eval_summary.json
After both ship, downstream consumers can choose:
- **Per-step**: use `eval_diag.jsonl` for trajectory analysis (when did drawdown happen, etc.)
- **Aggregate**: use corrected `eval_summary.json` for cross-run comparison (G8 gate, profit factor)
Neither replaces the other.
---
## 9. Done means
- Phase 1-4 implemented and locally validated
- One cluster run produces an `eval_summary.json` with all new fields populated and `n_trades` in the hundreds of thousands
- Pearl saved: `pearl_eval_summary_aggregation_bug.md` documenting the bug class
- MEMORY.md updated
---
## 10. Risks
- **Stream sync**: forgetting to `sim.stream.synchronize()` before `read_per_backtest_trade_counts()` could return stale head values mid-step. Mitigation: every read call synchronizes internally (existing pattern in `read_trade_records`).
- **Memory cost spike**: TRADE_LOG_CAP 1024→4096 = 4× memory. Already accounted for; fits budget. Mitigation: if it doesn't fit, fall back to 2048 + accept higher wrap rate.
- **TradeRecord field interpretation**: `realised_pnl_usd_fp` is fixed-point. Existing code divides by 100; preserve that.
- **Backward compat for old eval_summary.json consumers**: Argo aggregator (per memory `pearl_single_window_oos_is_not_oos`) reads `profit_factor` and `total_pnl_usd`. These keys are preserved. New fields are additive.

File diff suppressed because it is too large Load Diff

View File

@@ -0,0 +1,248 @@
# Determinism Foundation — Stop Chasing Ghosts
**Date**: 2026-06-02
**Status**: Spec draft — MANDATORY PREREQUISITE for any further controller / architecture work
**Priority**: BLOCKER for multi-head policy (`2026-06-02-multi-head-policy-with-r-multiple.md`), R-multiple reward, and the entire regime-invariance roadmap
**Empirical motivation**: alpha-rl-determinism-check.sh on `ml-alpha-regime-observer @ 63fc16f17` showed same-seed runs diverging at **step 2**, eval pnl Δ=**$364,925**, run A produced **0 trades** while run B produced **347 trades**. Bit-equivalence requirement: same seed → same diag rows to fp32 rel-tol 1e-5.
## §0. Motivation — the going-in-circles diagnosis was incomplete
The session of 2026-06-01/02 cycled through 4 controller iterations (v4 → v5.2) plus today's grwwh eval. Each iteration produced a different result. We attributed the variance to:
1. Bad controller designs (the v5/v5.1 bugs)
2. Architectural mismatch (single policy can't generalize)
3. Reward signal contamination (Phase 5 hold-bonus)
What we did NOT consider: **part of the variance was non-determinism**.
The 2026-06-02 mid-smoke determinism check revealed:
- Same seed, same code, same hardware → divergence at step 2
- By step 199: 654 diag leaves diverging
- Eval pnl Δ=$365k between identical configurations
- Behavioral Δ: run A=0 trades vs run B=347 trades (different agent entirely)
**Implications**:
1. Every controller iteration today might have been measuring seed-noise, not controller effect
2. The "scaffold v5 stuck at w=1.0" might be partially seed-dependent (would different seeds also stick?)
3. The "grwwh eval +$74M then collapse to -$160M" might partially be intra-run variance
4. **Until determinism is fixed, every future iteration is also seed-dependent and unfalsifiable**
Per `feedback_going_in_circles_pattern`: "audit the foundation". The foundation has a bug. **Fix the foundation before adding more layers**.
## §1. Investigation strategy — localize divergence to ONE component
Per `feedback_systematic_debugging` Phase 1 "Gather evidence in multi-component systems": instrument every boundary, find FIRST divergence.
### §1.1 Deterministic checksum kernel (NEW infrastructure)
Add `crates/ml-alpha/cuda/rl_deterministic_checksum.cu`:
- Computes sum-of-squares of any device tensor with **provably deterministic** accumulation
- Single-block, single-thread accumulation if tensor < 1024 elements (trivially deterministic)
- For larger tensors: block-tree-reduce with **fixed accumulation order** (block 0 always reads block 1's result, never vice versa; `__threadfence_system` + counter `cudaDeviceSynchronize` if needed; NO atomic ops, NO "whichever-block-finishes-last" patterns)
- Returns one f64 per tensor (sum-of-squares, full precision)
- Compile-time guaranteed determinism via construction, not configuration
### §1.2 Per-step component checksums in diag
Instrument every major component output in `build_diag_value`:
```rust
"checksums": {
"encoder_output": checksum(encoder_out_d, b_size × seq_len × hidden_dim),
"q_logits": checksum(q_logits_d, b_size × n_actions × n_atoms),
"pi_logits": checksum(pi_logits_d, b_size × n_actions),
"v_value": checksum(v_d, b_size),
"replay_sample_indices": checksum_int(per_sampled_indices_d, b_size),
"rewards_after_shape": checksum(rewards_d, b_size),
"advantages": checksum(advantages_d, b_size),
"q_grad": checksum(q_grad_d, all_params),
"pi_grad": checksum(pi_grad_d, all_params),
"v_grad": checksum(v_grad_d, all_params),
"encoder_grad": checksum(encoder_grad_d, all_params),
"adam_m_sum": checksum(adam_first_moments_d, all_params),
"adam_v_sum": checksum(adam_second_moments_d, all_params),
"isv_state": checksum(isv_dev_ptr, RL_SLOTS_END),
}
```
Cost: ~15 small kernel launches per step, each ~10us → 150us/step overhead. Acceptable in dev mode; toggle off in production.
### §1.3 Run determinism-check + localization
`scripts/determinism-check.sh` now compares checksums leaf-by-leaf. Output: first leaf to diverge + step number + Δ value.
Expected outcomes (one of these wins):
- `checksums.encoder_output` diverges first → **mamba2 culprit** (most likely)
- `checksums.q_logits` diverges first (encoder same) → Q head GEMM split-K culprit
- `checksums.pi_logits` first → π head GEMM culprit
- `checksums.replay_sample_indices` first → PER sampling RNG culprit
- `checksums.q_grad` first (forward same) → backward kernel reduction culprit
- `checksums.adam_m_sum` first (grads same) → optimizer iteration-order culprit
## §2. Standard fixes by component (apply based on §1 localization)
### §2.A Mamba2 selective scan (most-likely culprit)
If encoder checksum diverges first:
- **Chunk-size determinism**: mamba2 selective scan uses chunk-parallel scan. Setting `CHUNK_SIZE=1` reduces to sequential scan (trivially deterministic but ~10× slower). Use as ground-truth oracle.
- **Reduction-order fix**: within each chunk, the scan reduction must be fixed-order. Audit `mamba2_selective_scan_fwd.cu` for any `atomicAdd`, any "block 0 OR block 1 writes final" patterns, any non-deterministic chunk ordering.
- **Reference**: pearl `pearl_mamba2_selective_scan_nan_source` (Phase 2 v2 work) — same code path was extensively debugged for NaN issues. May have introduced non-det then.
### §2.B cuBLAS GEMM split-K
If Q/π/V head checksums diverge first (encoder same):
- Set `cublasSetMathMode(cublas_handle, CUBLAS_PEDANTIC_MATH)` — disables auto-selection of non-deterministic algorithms
- Disable TF32 (`cublasSetMathMode(CUBLAS_DEFAULT_MATH)`) — TF32 split-K is non-deterministic per NVIDIA docs
- This REVERSES tasks 30/31 (TF32 enable). Document the perf cost trade-off (~10-15% on H100 GEMMs).
- Alternative: use FP32 explicitly via `cublasSgemmStridedBatched` with explicit algorithm selection
### §2.C cuDNN deterministic mode
If conv/RNN kernels involved:
- `cudnnSetMathMode(CUDNN_DEFAULT_MATH)` (no TF32)
- Force `CUDNN_DETERMINISTIC` via algorithm selection
- Foxhunt may not use cuDNN heavily (most ops are custom CUDA kernels) — verify before applying
### §2.D Custom kernel block-tree-reduce audit
Per `feedback_no_atomicadd`, all reductions use block-tree-reduce. Audit each kernel for:
- ANY use of "last block to finish writes final" patterns — must become "block 0 always writes final" with proper synchronization
- ANY use of `cudaDeviceSynchronize` inside kernels (rare) — implies non-graph-capturable patterns
- The `block_reduce_to_global_with_counter` pattern (if used) — non-deterministic by construction
### §2.E Adam optimizer iteration order
If `adam_m_sum` checksum diverges (grads same):
- Audit AdamW step in `crates/ml-alpha/src/networks/*` for iteration order
- HashMap iteration is non-deterministic in Rust — use `BTreeMap` or pre-sorted `Vec<(key, value)>` for parameter groups
- Per-tensor iteration MUST be a fixed order
### §2.F PER (Prioritized Experience Replay) sampling
If `replay_sample_indices` checksum diverges:
- The PER index sampling kernel must use a deterministic RNG (device-side xorshift32 seeded from host per `pearl_rl_sample_tau_xorshift32`)
- Verify the seed is correctly propagated and not re-randomized between runs
- Check that priority updates and reads are properly synchronized via events (not just stream ordering)
## §3. Test infrastructure — make determinism a CONTRACT
### §3.1 Local determinism gate
`scripts/determinism-check.sh` (already created by fast-dev-cycle items 1-5) is the gate. Add:
- Exit 0 if all checksums match to rel-tol 1e-5
- Exit 1 with first-divergent-leaf report otherwise
- Verbose mode shows checksum trajectory side-by-side
### §3.2 GPU-oracle determinism invariant test
Add `crates/ml-alpha/tests/determinism_invariants.rs`:
- Test 1: 200-step mid-smoke twice with seed=42 → all checksums match
- Test 2: 200-step mid-smoke with seed=42 vs seed=43 → checksums DIFFER (sanity: seed actually does something)
- Test 3: 200-step mid-smoke with seed=42 on RTX 3050 → checksum at step 199 matches a frozen golden value (catches regression to non-det)
- All 3 must pass before any cluster smoke can submit (CI gate)
### §3.3 Pre-commit hook
`.git/hooks/pre-commit` (or `.claude/hooks/`):
- If any `.cu` or `.rs` file in `crates/ml-alpha/` changed → run 50-step micro-determinism-check
- If divergence → BLOCK commit
- Override with `git commit --no-verify` (logged + requires justification)
### §3.4 Cluster determinism regression test
Once local determinism is fixed:
- One re-run of any cluster smoke with same SHA should produce identical eval_summary.json
- Add to `scripts/argo-alpha-rl.sh`: optional `--regression-pair` flag that submits TWO identical workflows and verifies they match
## §4. Speed cost — document the dev-vs-prod trade-off
Determinism almost always costs throughput:
- TF32 → FP32: ~10-15% on H100 GEMMs (per NVIDIA docs)
- Mamba2 chunk-1: ~10× slower (if needed; should not be needed if reduction-order fix works)
- cuDNN deterministic: ~5-20% varies
- Checksum kernels: ~150us/step overhead
**Acceptable in dev** (user explicitly said "we are in dev, determism is most important").
**Production toggle** via env var `FOXHUNT_DETERMINISTIC` (default 1 for dev, 0 for production):
```rust
let pedantic = std::env::var("FOXHUNT_DETERMINISTIC").unwrap_or_else(|_| "1".to_string()) == "1";
if pedantic {
cublasSetMathMode(handle, CUBLAS_PEDANTIC_MATH);
} else {
cublasSetMathMode(handle, CUBLAS_TF32_TENSOR_OP_MATH);
}
```
Production runs (live trading, walk-forward production) MAY use non-deterministic mode for the 10-15% throughput gain. ALL dev / cluster smokes MUST use deterministic mode.
## §5. CI guard
Add to GitLab CI (or whatever CI foxhunt uses — see `feedback_argo_workflows_primary`):
- Job: `determinism-gate`
- Triggers: PR to main, or any branch with `ml-alpha-*` prefix
- Runs: `scripts/determinism-check.sh` on the PR commit
- Fails the PR if any divergence > rel-tol 1e-5
- Argo workflow template: add `pre-train-determinism-check` step to `alpha-rl-template.yaml` that runs determinism-check on the cluster GPU before submitting the actual training run
## §6. Falsification criteria
PASS criteria (ALL must hold):
- `determinism-check.sh` passes (exit 0) for 5 different seeds (42, 43, 44, 45, 46) on RTX 3050
- One b=1024 cluster smoke re-run produces eval_summary.json `total_pnl_usd` matching to ±$1 (true bit-exact at scalar level)
- `determinism_invariants.rs` test 1 passes (twin run match)
- `determinism_invariants.rs` test 3 passes (golden checksum match) — proves no regression
FAIL criteria:
- Any test above fails after applying §2 fixes
- If failure persists after auditing all 6 candidate causes (mamba2, GEMM, cuDNN, custom kernels, Adam, PER) → escalate to NVIDIA support OR accept multi-seed averaging as the workaround (5× cluster cost per verdict, but trustworthy)
## §7. Sequencing
Phase 1 ships first (the diagnostic):
1. **§1.1** Write `rl_deterministic_checksum.cu` + register in `build.rs` (~2 hours)
2. **§1.2** Instrument 15 component checksums into `build_diag_value` (~3 hours)
3. **§1.3** Modify `determinism-check.sh` to compare checksums leaf-by-leaf, report first-divergent (~1 hour)
4. **§3.1** Local determinism gate output (~1 hour)
5. **Run** the diagnostic: identify first-divergent component (~30 min)
**Phase 1 total: ~8 hours / 1 day.**
Phase 2 fixes depend on Phase 1 results (apply the relevant §2.X). Cost: ~1-3 days depending on which component is the culprit.
Phase 3 validation (§3.2, §3.3, §3.4) + Phase 4 documentation (§4): ~1 day.
**Total: ~3-5 days to ship the foundation.** This is the prerequisite for multi-head policy, R-multiple reward, and all subsequent regime-invariance work.
## §8. Out of scope (defer)
- Multi-GPU determinism (foxhunt is single-GPU per process; defer until production)
- Cross-hardware determinism (RTX 3050 vs L40S vs H100 may produce different exact bits even with all fixes; rel-tol 1e-5 should hold but bit-exact across hardware is not required)
- Cross-CUDA-version determinism (pin CUDA 12.4 per current setup; if upgraded, re-audit)
- Floating-point order across compiler versions (pin rustc + cudart versions in build env)
## §9. Why this comes BEFORE multi-head policy
Multi-head policy adds 3× the policy head parameters + a gating network + new aux loss term. ANY of these new components could introduce its own non-determinism (more GEMM calls, more block-tree-reduces, possibly more PER samples).
If we ship multi-head on a non-deterministic foundation:
- Multi-head eval shows +$3M → we cheer, but it's a lucky seed
- Multi-head eval shows -$5M → we revert, but it was an unlucky seed
- Multiple seeds → 5× cluster cost
- We never know if the architecture worked
If we ship determinism FIRST:
- Multi-head eval shows +$3M → reproducible, trust it, build on it
- Multi-head eval shows -$5M → reproducible, abandon it, learn from it
- Single seed per iteration → fast cluster cadence
- **Every iteration produces real signal**
The 3-5 day investment in determinism pays back on the FIRST controller iteration. Over the next 10 iterations, it saves 50% of cluster compute (no need for multi-seed averaging) and 100% of the "is this real or noise?" confusion that drove today's going-in-circles.
## §10. Memory pearls to create after Phase 1 lands
- `pearl_determinism_root_cause` — documents which component(s) were the actual culprit + fix applied
- `pearl_determinism_speed_cost` — documents the actual measured throughput cost on H100 + L40S + RTX 3050
- `pearl_determinism_contract` — the rule going forward: every commit passes determinism gate, no exceptions

View File

@@ -0,0 +1,250 @@
# Fast Dev Cycle — Closing the 700× Gap
**Date**: 2026-06-02
**Motivation**: today's session burned 4 cluster smokes × ~70min each on controller iterations that could have been killed in ~10min locally. Per `feedback_going_in_circles_pattern`: "3+ consecutive cluster runs each surfacing a new failure introduced by the prior fix" — the going-in-circles pattern is partially a CONSEQUENCE of slow dev cycle, not a separate root cause.
**Linked**: `feedback_local_smoke_before_cluster`, `feedback_kill_runs_on_anomaly_quickly`, `pearl_grwwh_eval_catastrophic_collapse`
## §0. The current gap
| Tier | Setup | Time | What it validates |
|---|---|---|---|
| 1 (local smoke) | RTX 3050, b=16, 200 steps, 1 data file | **~6 sec** | Kernel correctness, ISV propagation, no NaN |
| **GAP — no intermediate validation** | — | — | — |
| 2 (cluster smoke) | L40S, b=1024, 20k+5k, 9 PVC files | **~70 min** | Behavioral + eval generalization |
**Cost of the gap**: every hypothesis costs 70min to falsify. 4 hypotheses today = 4.5 hours of cluster time + 4× context switching while waiting. The Perplexity 2026-06-02 RL-engineering review confirms the canonical fix is a **funnel** with intermediate behavioral validation.
## §1. Goal — add Tier 1.5 local mid-smoke
| Tier | Setup | Time | Verdict signals |
|---|---|---|---|
| 1 (correctness) | RTX 3050, b=16, 200, 1 file | 6 sec | shapes, NaN, ISV slots |
| **1.5 (NEW: behavior)** | **RTX 3050, b=128, 2000+500, 2 PVC files (~1.5GB local)** | **~10 min** | entropy trajectory, hold growth, wr crossover, controller stability, scaffold/aux signal sanity, early Pearson |
| 2 (eval verdict) | L40S, b=1024, 20k+5k, 9 PVC files | 70 min | eval pnl, eval wr, regime stratification |
**Tier 1.5 kills 80%+ of bad hypotheses before they reach cluster**. The 10-min cost is recovered ~7× per saved cluster smoke.
## §2. RTX 3050 capacity check (4GB VRAM)
| Component | b=16 (now) | b=128 (proposed Tier 1.5) | b=1024 (cluster) |
|---|---|---|---|
| Encoder activations (mamba2, hidden=256, seq=32, 6 layers) | 3 MB | **24 MB** | 192 MB |
| Q/V/π heads forward+backward | 1 MB | **8 MB** | 64 MB |
| Replay buffer (capacity 4096) | ~100 MB | 100 MB | 100 MB |
| Optimizer state (AdamW × ~1M params × 16B) | 16 MB | 16 MB | 16 MB |
| LobSim per-account state (1024 max units, ~4KB/account) | ~64 KB | ~512 KB | 4 MB |
| Misc (CUDA workspace, kernel launch overhead) | 200 MB | 200 MB | 200 MB |
| **TOTAL** | ~320 MB | **~350 MB** | ~580 MB |
**Verdict**: RTX 3050 (4GB) has comfortable headroom for b=128. b=256 also fits. b=512 is the edge.
## §3. Implementation tasks
### §3.1 PVC data sync (one-time setup)
1. Identify the 2-3 highest-value MBP-10 files from `/feature-cache/futures-baseline-mbp10/` (recommend: most-recent fold-1 train file + most-recent fold-1 eval file)
2. `rsync` (via kubectl exec or a bridge pod) to local `test_data/futures-baseline-mid/` (~1.5 GB local disk)
3. Add `test_data/futures-baseline-mid/` to `.gitignore` (don't commit data)
4. Document in `CLAUDE.md` under "Build & Test" the new `FOXHUNT_MID_TEST_DATA` env var
### §3.2 Local mid-smoke script
1. Create `scripts/local-mid-smoke.sh` (NEW file in `scripts/`):
```bash
#!/usr/bin/env bash
set -euo pipefail
SQLX_OFFLINE=true cargo build --release --example alpha_rl_train -p ml-alpha
target/release/examples/alpha_rl_train \
--n-steps 2000 --n-eval-steps 500 \
--fold-idx 1 --n-folds 3 \
--mbp10-data-dir test_data/futures-baseline-mid \
--predecoded-dir test_data/futures-baseline-mid \
--out /tmp/foxhunt-mid-smoke \
--eval-diag-jsonl /tmp/foxhunt-mid-smoke/eval_diag.jsonl \
--instrument-mode front-month --n-backtests 128 \
--log-every 500 --seed 42
```
2. Per-run elapsed time target: ≤ 10 min on RTX 3050
3. Add to `Makefile` or `justfile` as `mid-smoke` target
### §3.3 Behavioral kill criteria (Tier 1.5 verdict)
After mid-smoke completes, run a verdict script that reads diag.jsonl + eval_diag.jsonl and produces PASS/FAIL based on behavioral signals (cheaper than waiting for eval pnl):
| Signal | Train target | Eval target | Kill if |
|---|---|---|---|
| `action_entropy` final | ≥ log(11) × 0.5 (1.20) AND ≤ log(11) × 0.85 (2.04) | similar | < 0.6 or > 2.3 — policy collapsed or random |
| `q_pi_agree_ema` final | ≥ 0.6 | n/a | < 0.3 — Q and π decoupled |
| Pearson(rewards.sum, Δrealized_pnl) | ≥ 0.5 | ≥ 0.4 | < 0.3 — gradient anti-aligned |
| `avg_hold_steps` | growing trend over 1500 steps | n/a | flat or decreasing — surfer pattern absent |
| `popart.sigma` | stable (CV < 50%) | n/a | wild oscillations — controller broken |
| `wr_ema` final | ≥ 0.25 | ≥ 0.20 | < 0.15 — random or worse |
| `eval pnl` | n/a | > -$1M (allow margin given small batch) | < -$3M — catastrophic |
Implement as `scripts/tier1_5_verdict.py` that reads diag.jsonl and emits OK / KILL_<reason> + optional explanation.
Per `feedback_kill_runs_on_anomaly_quickly`: any KILL_* response means DO NOT submit cluster smoke. Fix locally first.
### §3.4 Checkpoint-resume — async double-buffer for graph compatibility
**Architectural constraint (verified 2026-06-02)**: ml-alpha already runs in mega-graph mode with 4 active CUDA graphs (`prefill_graph`, `postfill_graph`, `reward_graph`, `training_graph` per `crates/ml-alpha/src/trainer/integrated.rs:715-719`). The team has actively eliminated host-blocking sync points (per `integrated.rs:707` comment "instead of `train_stream.synchronize()` — event wait is GPU-side" and `integrated.rs:8459` "saves ~2ms/step"). **Any synchronous DtoH for checkpoint would regress this perf work and break graph capture compatibility.**
Therefore checkpoint-resume MUST use **async double-buffer** from the start — sync is not a viable interim.
#### §3.4.1 Prerequisite — audit existing checkpoint code for sync points
Before any new code:
1. Read `crates/ml-alpha/src/trainer/integrated.rs` around the checkpoint save/load functions (referenced near line 9621 per `2026-05-31-checkpoints-and-eval-diag-design.md`)
2. Identify every `stream.synchronize()`, `device.synchronize()`, or implicit-sync call in the checkpoint path
3. Each one is a graph-mode breaker — either remove or replace with `CudaEvent::record` + later `event.synchronize_host()` deferred until off the training stream
4. Document the audit findings as a note: "existing checkpoint code has N sync points at lines ... and uses ... pattern"
If the existing checkpoint code is already async (which is possible given the team's discipline), §3.4.2-§3.4.5 below describe DOCUMENTING it and wiring the `--resume-from` flag. If it's sync, the design below replaces it.
#### §3.4.2 Design — async double-buffer
**Two pinned host buffers** (each sized to total checkpoint payload, ~500MB):
```rust
struct CheckpointBuffers {
buf_a: CudaSlice<u8>, // pinned host, 500MB
buf_b: CudaSlice<u8>, // pinned host, 500MB
active: AtomicUsize, // 0 → buf_a, 1 → buf_b
a_complete: CudaEvent, // signaled when DtoH into buf_a finishes
b_complete: CudaEvent, // signaled when DtoH into buf_b finishes
}
```
**Save path** (called from training loop at checkpoint cadence):
1. Read `active` atomically → decide target buffer (say `buf_a`)
2. Verify the OTHER buffer's serialization is complete (i.e., previous checkpoint's I/O thread done)
3. Enqueue DtoH copies for all checkpoint payload onto the training stream targeting `buf_a`
4. Enqueue `a_complete.record(stream)` on the training stream
5. Notify the serialization thread that `buf_a` is being filled
6. Flip `active` to `buf_b` so next checkpoint won't collide
7. **Training stream continues without sync** — graph capture not broken
**Serialization thread** (separate `std::thread::spawn`):
1. Wait on `a_complete.synchronize_host()` (blocks ONLY this thread, not training)
2. Serialize the pinned buffer contents to disk via `std::io::Write`
3. Mark "serialization complete for buf_a" so next checkpoint into buf_a can proceed
**Load path** (called once at run start, before training begins — no graph conflict):
1. Read checkpoint file from disk to pinned host buffer
2. HtoD copy to GPU (sync OK here since training hasn't started)
3. Re-enter graph capture mode
#### §3.4.3 Checkpoint payload
Per `pearl_F4.1` checkpoint infra, the payload includes:
- Encoder weights + Adam state
- Q-head weights + Adam state
- V-head weights + Adam state
- π-head weights + Adam state
- Bootstrap ISV (760+ floats — small, but include for reproducibility)
- popart Welford state (mean + variance + count)
- Kelly EMAs (avg_win, avg_loss, win_rate)
- Edge-decay PH state per batch
- regime_observer state per batch
Audit in §3.4.1 must verify ALL of these are captured. Per `pearl_adaptive_carryover_discipline`: missing any one means resume produces a different agent than save.
#### §3.4.4 Hypothesis-test workflow
**Pre-train baseline checkpoint** (one-time per SHA):
1. Full 20k-step cluster run with `--checkpoint-every 5000 --checkpoint-dir /feature-cache/baselines/<sha>/`
2. Produces 4 checkpoints + final = `/feature-cache/baselines/<sha>/{ck-5k, ck-10k, ck-15k, ck-20k}.bin`
3. Cluster smoke = 70 min (baseline cost)
**Hypothesis test runs** (per controller iteration):
1. New SHA with controller change
2. Resume from `/feature-cache/baselines/<previous-sha>/ck-15k.bin` (skip first 15k steps)
3. Train 5k additional steps with the new code → continues from learned state
4. Eval 5k
5. Verdict in ~12min instead of 70min
**Caveat**: code-incompatible changes (struct layout, slot allocation) require fresh baseline. Versioned checkpoint format with magic number + schema hash.
#### §3.4.5 Implementation tasks
1. Audit existing checkpoint code (§3.4.1) — ~1 hour
2. Implement async double-buffer save (§3.4.2) — ~1.5 days
3. Add serialization thread + completion tracking — ~0.5 day
4. Implement load path (~half day)
5. Wire `--resume-from <path>` + `--checkpoint-every N` flags to `alpha_rl_train.rs` (~half day)
6. Add checkpoint schema versioning (~half day)
7. Local invariant test: save → resume → continue produces bit-equivalent diag for next 100 steps (~half day)
**Total**: ~4 days. Saves ~58min per cluster smoke (~80% reduction) once shipped.
#### §3.4.6 Sync-vs-async trade-off summary
| Approach | Implementation | Graph compat | Maintenance |
|---|---|---|---|
| Sync DtoH | 1 day | ❌ Breaks ml-alpha's 4 active CUDA graphs | Forces rewrite when graph dependency hardens |
| **Async double-buffer (this spec)** | **4 days** | ✅ Native graph compat via CudaEvent | Stable; aligned with existing perf discipline |
The 3 extra days are recovered after the first cluster smoke (sync would need rework before merging back into graph mode anyway).
### §3.5 Determinism / seed pinning (reproducibility audit)
Reproducibility is required for any tiered validation to be trustworthy. Verify:
1. Python (data loading) seeds: pin via env var (already done)
2. NumPy seeds: pin (already done)
3. cudarc / curand kernel seeds: verify pinned (per `pearl_rl_sample_tau_xorshift32`: device-side xorshift seeded from host)
4. Data order: NO random shuffle in walk-forward (verify in trainer)
5. Replay buffer: PER is randomized BUT seeded — verify replay seed propagation
Local mid-smoke MUST be deterministic: same seed → same final loss within fp32 precision. Add a self-test: run mid-smoke twice with seed=42, diff the final 5 rows of diag.jsonl — should match to 1e-5.
### §3.6 Per-step diag / eval_summary consistency check
Per `pearl_grwwh_eval_catastrophic_collapse`: at grwwh step 4863 per-step showed +$74M, final eval_summary showed -$160M. $234M gap is intolerable — the per-step diag is misleading.
1. Pull the LAST eval row of `eval_diag.jsonl` AND `eval_summary.json` from cluster runs
2. If `|per_step_realized_pnl_cum - eval_summary_total_pnl| > 5% × max(abs)`, log a discrepancy warning
3. Add this to the verdict script (§3.3 above) so Tier 1.5 also catches the issue
4. Long-term: fix the accounting axis discrepancy in the diag/summary code
## §4. Estimated impact on dev cycle
**Today (4 controller iterations on the surfer-scaffold)**:
- Tier 1 (5s × 4 = 20s) + Tier 2 (70min × 4 = 280min) = **4.7 hours of cluster time**
- Plus context-switching waiting for cluster results
**With Tier 1.5 + checkpoint-resume**:
- Tier 1 (5s × 4 = 20s) + Tier 1.5 (10min × 4 = 40min) + Tier 2 only on Tier 1.5 PASS (assume 2 out of 4 PASS) = **40min Tier 1.5 + 24min Tier 2 = ~1 hour total**
- **~4× faster total wall time**
- **2× fewer cluster smokes** (only PASS the gate get to Tier 2)
- Plus the savings from checkpoint-resume on the Tier 2 runs
## §5. Sequencing
Ship in this order (each step independently useful, but §3.4 is now the longest chunk):
1. **§3.1 PVC data sync** (~30 min setup) — gives local access to real fold-1 data
2. **§3.2 Local mid-smoke script** (~30 min) — make `cargo run --release ... --n-backtests 128 --n-steps 2000` work
3. **§3.5 Determinism audit** (~1 hour) — verify seed pinning, add self-test
4. **§3.3 Behavioral kill verdict script** (~1 hour) — `scripts/tier1_5_verdict.py`
5. **§3.6 Per-step / eval_summary consistency check** (~30 min) — add to verdict script
6. **§3.4 Async double-buffer checkpoint-resume** (~4 days) — see §3.4.5 task list; sync DtoH is NOT viable due to existing mega-graph mode (see §3.4 architectural constraint)
**Total spec implementation**: ~5 hours (items 1-5) + ~4 days (item 6) = ~5 working days.
Items 1-5 ship in ~1 day and unlock most of the dev cycle gain (Tier 1.5 mid-smoke gates 80% of bad hypotheses without needing checkpoint-resume). Item 6 (async checkpoint) is the bigger investment but is on the critical path for graph-mode compatibility AND saves ~58min per cluster smoke once shipped.
**Parallel path option**: items 1-5 can be done by one developer; §3.4 in parallel by a second developer or a coder agent. The fast wins land in 1 day; the deep checkpoint refactor lands in 4-5 days.
## §6. Out of scope (for now)
- **torch.compile / kernel autotuning**: foxhunt already uses pre-compiled CUDA kernels with handcrafted block-tree-reduce per `feedback_nvidia_grade_perf_for_kernels`. Further compile-time tuning is high-cost low-reward.
- **Surrogate models / amortized inference**: building a simulator that mimics foxhunt's training dynamics would be its own multi-week project. Defer until the standard funnel proves insufficient.
- **Multi-seed reproducibility runs**: requires the deterministic foundation from §3.5 first. Add after Tier 1.5 is operational.
- **Production walk-forward distinct from cluster smoke**: out of scope; the existing 3-fold walk-forward is the production gate. Tier 1.5 is a development accelerator, not a production replacement.
## §7. Open questions
1. **rsync vs sshfs**: rsync (one-time copy) preferred for steady-state read performance. ~1.5GB local disk cost is negligible.
2. **b=128 vs b=256 for mid-smoke**: b=128 fits comfortably on RTX 3050 with margin. b=256 may stretch. Start with b=128, raise to b=256 if behavioral signals are noisy.
3. **Step count tuning**: 2000 train + 500 eval might be too few for some signals (popart σ takes 1000+ steps to stabilize). Could extend to 3000+750 if needed; still < 15 min total.
4. **Checkpoint format compatibility**: existing checkpoint infra from F1.3 saves model + Adam + ISV. Does it save popart Welford state? Kelly EMAs? `regime_observer` state? Audit and fix gaps before §3.4 ships.
## §8. Discipline reminders
- `feedback_local_smoke_before_cluster` extends naturally: now Tier 1 → Tier 1.5 → Tier 2 in strict order
- `feedback_kill_runs_on_anomaly_quickly` applies at Tier 1.5 just as much as Tier 2 — KILL on behavioral red flags before cluster submit
- `feedback_going_in_circles_pattern` — if the same Tier 1.5 KILL_* triggers 3 times in a row, STOP and re-audit the foundation, not the controller

View File

@@ -0,0 +1,246 @@
# Multi-Head Policy with R-Multiple Reward (Intervention #4-clean)
**Date**: 2026-06-02
**Status**: Spec draft (post-sp-critical-reviewer BLOCK on prior v6 attempt)
**Prerequisite**: Intervention #1 (R-multiple reward) has shipped and the eval verdict is in
**Linked specs**: `2026-06-02-regime-invariance-four-interventions.md` (parent)
**Linked pearls**: [[pearl_surfer_baseline_was_train_only_never_eval]], [[pearl_two_alpha_modes_surfer_vs_trend]], [[pearl_reward_signal_anti_aligned_with_pnl]], [[pearl_pi_actor_collapses_without_entropy_floor]]
## §0. Empirical motivation — single-policy regime collapse is EMPIRICAL, not theoretical
The previous v6 attempt tried to motivate multi-head policy via a theoretical "average optimal action" argument. The sp-critical-reviewer correctly BLOCKED that as hand-waving. This spec grounds the motivation in cluster-run evidence:
| Run SHA | Train wr | Eval wr | wr collapse | Per-pearl |
|---|---|---|---|---|
| dd049d9a4 | 0.56 (hold=30) | **0.219** | **2.6×** drop | `pearl_surfer_baseline_was_train_only_never_eval` |
| 87a8259c6 (8gtk2) | wr_max 0.418 / final 0.343 | 0.340 | minor train→eval but train wr peaked then declined within training itself | `pearl_reward_signal_anti_aligned_with_pnl` |
| 63fc16f17 (grwwh, current) | 0.312 final | TBD ~25min | TBD | [[project_adaptive_scaffold_session_2026_06_02]] |
The dd049d9a4 evidence is decisive: a single policy network trained on Phase-5 reward shape achieved profitable train trade-quality and **collapsed to barely-above-random** on eval data. Per [[pearl_two_alpha_modes_surfer_vs_trend]], this isn't a bug — it's "philosophy not bug" — but the philosophy was BAKED INTO weights by training. Without architectural support for multiple philosophies, the single policy can only embed one philosophy at a time.
**This spec hypothesizes** (NOT proves) that giving the policy K heads with different initial biases + a learned gating produces a policy that retains MULTIPLE philosophies in weights, allowing it to switch at inference time based on encoder context. The verification is empirical: cluster smoke + eval-wr stratified by regime markers.
## §1. Design
### §1.1 Architecture (verified against current codebase)
Current architecture (read from `crates/ml-alpha/src/networks/policy_head.rs` and `crates/ml-alpha/src/trainer/integrated.rs`):
- Encoder: mamba2 trunk → encoder_output of size HIDDEN_DIM
- Policy head: Linear(HIDDEN_DIM → N_ACTIONS=11), outputs `pi_logits[b, a]`
- Q head: dueling distributional → `q_logits[b, a, z]`
- V head: scalar critic for PPO
Proposed change (multi-head policy ONLY — Q and V heads unchanged):
- Encoder: UNCHANGED
- K policy heads (K=3 proposed): K × Linear(HIDDEN_DIM → N_ACTIONS), outputs `pi_logits_k[b, a]` for k ∈ {0, 1, 2}
- Gating head: small Linear(HIDDEN_DIM → K), outputs `gate_logits[b, k]`
- Final pi_logits: `pi_logits[b, a] = log( Σ_k softmax(gate_logits[b])[k] × softmax(pi_logits_k[b])[a] )`
- Equivalently in logit space, the mixture is taken on probabilities (not logits) then log'd back. This is the **product-of-experts ↔ mixture-of-experts** distinction; we use mixture (additive over probabilities).
### §1.2 Reward signal — UNCHANGED from Intervention #1
This spec does NOT modify the reward kernel. R-multiple reward from Intervention #1 stays exactly as-is. The reward path is identical for all K heads — they all receive the SAME per-event reward signal.
**This addresses BLOCKER 4 (Pearson alignment)**: since the reward path is unchanged, `Pearson(reward, Δpnl)` is identical to whatever #1 achieves. Multi-head policy adds POLICY expressiveness without touching the reward. No anti-aligned multiplicative shaping is introduced.
### §1.3 Specialization mechanism (no new ISV signals)
**This addresses BLOCKER 1**: heads specialize WITHOUT requiring any new regime-observation ISV slot. Specialization arises from FOUR sources internal to the network:
1. **Different head initialization**: each head's output projection is initialized with a different action bias:
- Head 0 (trend): bias toward HOLD_LONG / HOLD_SHORT actions (+0.5 logit on HOLD actions, 0 elsewhere)
- Head 1 (reversion): bias toward FLAT_LONG / FLAT_SHORT actions (+0.5 logit on FLAT actions, 0 elsewhere)
- Head 2 (flat): bias toward FLAT_FLAT action (+0.5 logit on FLAT only)
2. **Per-head entropy floor**: each head has an independent entropy lower bound preventing collapse to deterministic. Floor value: `log(11) × 0.4` (~0.96).
3. **Gating entropy floor**: gating distribution forced to retain at least `log(K) × 0.5` (~0.55 for K=3) entropy. Prevents gate-collapse-to-one-head mode (per [[pearl_pi_actor_collapses_without_entropy_floor]]).
4. **Per-head auxiliary action-prior loss (§1.4 below)** — continuous regularizer that prevents heads from converging to identical distributions.
The init bias alone (1) is FRAGILE: heads start different but see identical PPO gradient, so they would converge over training. Item 4 provides the CONTINUOUSLY-APPLIED regularizer that maintains specialization across training.
### §1.4 Per-head auxiliary action-prior loss (CLEAN per-expert specialization)
For each head k, an auxiliary KL term pulls its action distribution toward a fixed prior `p_prior_k`:
```
aux_loss_k = β × KL( p_k(a|s) ∥ p_prior_k(a) )
```
Combined loss per head:
```
total_loss_k = ppo_loss_k + aux_loss_k
```
The mixture policy's total loss is `Σ_k g_k × total_loss_k` where `g_k` is the gating weight.
**Priors** (defined per head, FIXED — no learned drift):
- `p_prior_0` (trend): 0.10 uniform over non-HOLD, 0.30 on each of the 3 HOLD_LONG / HOLD_SHORT / HOLD_FLAT actions (sums to 1.0)
- `p_prior_1` (reversion): 0.10 uniform over non-FLAT-of-position, 0.40 on FLAT_LONG and FLAT_SHORT
- `p_prior_2` (flat): 0.40 on FLAT_FLAT, 0.06 uniform over others
(The exact prior weights are tunable but should sum to 1.0 and concentrate ~30-40% of mass on the head's specialty action group.)
**β coefficient**: ISV-driven via new slot `RL_POLICY_AUX_PRIOR_BETA_INDEX`. Bootstrap 0.05 (5% of PPO loss magnitude). When β=0, the spec reduces to §1.3 init-bias-only — useful as ablation.
**Why this addresses Pearson alignment (re-confirms BLOCKER 4 stays addressed)**: the aux_loss is added to the POLICY LOSS, not to the REWARD signal. `Pearson(reward, Δpnl)` is unchanged because the reward kernel is unmodified. The aux loss is a REGULARIZER on policy output; it does not alter the gradient direction of the PPO surrogate's REWARD term — only adds an orthogonal pull toward each head's prior.
**Sign-preservation**: KL divergence is always ≥ 0; aux_loss is always positive; subtracting it from the maximization objective always reduces the policy's deviation from prior. Cannot induce sign-flips in the reward direction.
**Reference**: standard mixture-of-experts auxiliary loss pattern. Shazeer et al. 2017 (sparsely-gated MoE), Lepikhin et al. 2020 (GShard) — well-studied technique, low novelty risk.
## §2. ISV slot additions (4 slots, verified against current `RL_SLOTS_END`)
Current `RL_SLOTS_END = 757` (post-grwwh / v5.2, will be 757 again after #1 R-multiple ships per the #1 plan). New slots:
```
RL_POLICY_NUM_HEADS_INDEX usize = 757 // K (current 3 — could allow runtime override)
RL_POLICY_GATING_ENTROPY_FLOOR_INDEX usize = 758 // ~0.55 (log(K) × 0.5)
RL_POLICY_HEAD_ENTROPY_FLOOR_INDEX usize = 759 // ~0.96 (log(11) × 0.4)
RL_POLICY_AUX_PRIOR_BETA_INDEX usize = 760 // 0.05 bootstrap; β for per-head action-prior KL
```
`RL_SLOTS_END = 761` after this spec ships.
**No new producer kernels needed**. The entropy floors are read by the EXISTING entropy controller infrastructure (per `crates/ml-alpha/cuda/rl_entropy_coef_controller.cu` and similar).
## §3. Implementation tasks (concrete, ordered)
### §3.1 Network changes (Rust)
1. Modify `PolicyHead` struct in `crates/ml-alpha/src/networks/policy_head.rs`:
- Add `Vec<Linear>` for K head projections instead of single `Linear`
- Add `Linear` for gating projection
- Adam optimizer states for each (use existing `Adam` per-tensor pattern)
2. Modify forward pass:
- Compute K head logits in parallel (matmul against shared encoder_output)
- Compute gate logits
- Mixture via `log Σ softmax(gate) × softmax(head)`
3. Backward pass: gradient flows through gating × head expression. Standard PPO surrogate on the mixture pi_logits.
### §3.2 Kernel changes (CUDA)
Multi-head policy is mostly Rust-side because the heads are linear projections. New kernels needed:
- `rl_policy_mixture_forward.cu`: takes K head logits + gate logits → produces mixture logits (LogSumExp over heads with gate weighting). Block-tree-reduce, no atomicAdd. ~80 LOC.
- `rl_policy_mixture_backward.cu`: backward through the mixture. ~120 LOC.
- `rl_policy_aux_prior_loss.cu`: computes Σ_k g_k × β × KL(p_k ∥ p_prior_k) and gradient ∂L_aux/∂head_k_logits. Reads `RL_POLICY_AUX_PRIOR_BETA_INDEX` for β. ~60 LOC.
Per `feedback_no_nvrtc`: pre-compiled cubins via `build.rs`.
### §3.3 Trainer integration
- Allocate K head weights + gating weights in trainer constructor
- Initialize each head with the action-bias pattern from §1.3
- Wire entropy floors via existing controller infrastructure
- Per-head entropy diagnostic emitted to `diag.jsonl` (3 new leaves: `policy.head_0_entropy`, etc.) + gate entropy diagnostic (1 leaf)
### §3.4 Diag emission
Add to `build_diag_value`:
- `policy.head_count = K`
- `policy.head_k_entropy[k]` for k in 0..K
- `policy.gate_entropy`
- `policy.gate_probs[k]` for k in 0..K (mean across batch)
EXPECTED_LEAVES: 679 → 679 + (K + 1 + K) = 686 (K=3).
## §4. Falsification — strict kill gates (addresses BLOCKER 5)
### §4.1 Per-step kill criteria (per `feedback_kill_runs_on_anomaly_quickly`)
Kill the cluster smoke if ANY of:
- `Pearson(reward, Δpnl) < 0.5` at step 2000 (sparse close events should align by then)
- `gate_probs[k] < 0.05` for any k at step 5000 (mode collapse — gate ignored a head)
- `gate_entropy < log(K) × 0.4` at step 5000 (gate determinism — should be > 0.45)
- `action_entropy < log(11) × 0.5` at step 5000 (overall policy collapse)
- `realized_pnl_cum < -$50M` at step 10000 (catastrophic loss, abandon)
### §4.2 Verdict gates (PASS / FAIL)
PASS criteria (ALL must hold):
- Train Pearson(reward, Δpnl) ≥ 0.6 at steps 1500-2000
- Train wr ≥ 0.35 by step 10k (above any reasonable random baseline; threshold above current grwwh peak of 0.31)
- Train action_entropy ∈ [log(11) × 0.5, log(11) × 0.85] (concentrated but not collapsed)
- Train gate_entropy ≥ log(K) × 0.4 (heads actively used)
- **Eval pnl > 0** (the actual generalization test)
- **Eval wr ≥ 0.30** (above the random-baseline floor for K=11 actions)
- Per-regime eval wr shows specialization: split eval rows by `session_pnl_variance_ema` quartiles; the spread `max(wr_q) min(wr_q)` should exceed `0.05` (heads pulling agent in different directions based on regime).
FAIL criteria:
- Any per-step kill gate fires
- Eval pnl < -$2M (defines the fail boundary; in [-$2M, 0] is marginal → next intervention)
- Per-regime spread < 0.02 (heads didn't specialize → architectural gain near zero)
### §4.3 Per-regime stratification methodology
The falsification needs regime markers that ALREADY EXIST in diag.jsonl. Verified existing signals:
- `risk_stack.regime.session_pnl_variance_ema` (slot 700) — proxy for vol regime
- `risk_stack.regime.dead_zone_flag` (slot 696) — proxy for chop regime
- `risk_stack.regime.recovery_factor` (slot 699) — proxy for drawdown regime
Stratify eval steps by quartiles of `session_pnl_variance_ema` (low-vol vs high-vol regimes); compute wr separately on each quartile. The hypothesis is that multi-head policy will produce different wr in different quartiles (heads specializing), while single-policy grwwh would produce uniform wr across quartiles.
## §5. Risks and mitigations
### §5.1 Mode collapse on gating
Risk: gating distribution collapses to single head (`gate_probs ≈ [1, 0, 0]`); other heads become dead weight.
Mitigation: `RL_POLICY_GATING_ENTROPY_FLOOR_INDEX` enforces hard floor on gate entropy. Per-step kill gate at step 5000 catches this.
Reference: [[pearl_pi_actor_collapses_without_entropy_floor]] — entropy floor canonically applied to action distribution; here applied to gating distribution at the same conceptual layer.
### §5.2 Heads converge to identical behavior — MITIGATED by §1.4 aux prior loss
Risk: with shared encoder + shared reward, the heads might converge to identical action distributions, making multi-head equivalent to a single head with K× parameters wasted.
**Primary mitigation**: §1.4 per-head auxiliary action-prior KL loss. Continuously pulls each head toward its distinct prior, preventing convergence. Standard MoE specialization technique (Shazeer 2017).
**Secondary mitigations**: per-head init bias (§1.3 item 1) + per-head entropy floor (§1.3 item 2). The aux prior provides the lasting pressure; init + entropy provide the boundary conditions.
Diagnostic: emit pairwise KL between heads to diag (`policy.head_pairwise_kl[i,j]`). If KL drops below 0.1 across all pairs, heads have converged → β too low or aux prior loss disconnected.
**β-tuning escape hatch**: `RL_POLICY_AUX_PRIOR_BETA_INDEX` is ISV-driven. If empirically heads converge despite β=0.05 bootstrap, the controller can adapt β upward via an ISV controller (similar pattern to existing entropy-coef controller). Add this as a follow-up controller only if needed.
### §5.3 Training instability from mixture
Risk: PPO surrogate on a mixture has higher variance than on a single Gaussian-ish policy. Could destabilize training.
Mitigation: existing PPO ratio clamp from B-7 / B-9 work handles this. Per-batch advantage normalization (Phase 4.5) helps.
### §5.4 Eval doesn't show specialization (heads pulled to average)
Risk: even with all mitigations, eval might produce uniform wr across regime strata → no specialization actually achieved.
Mitigation: this IS the falsification gate. If specialization spread < 0.02, the architectural change failed. The conclusion is then NOT to ship #4.B; reroute to #3 (cross-fold training) or #4.C (full MoE).
## §6. Open questions (to resolve before implementation, NOT during)
1. **Initial K**: 3 (trend/reversion/flat) vs 5 (add trend-reversal, vol-breakout)? Larger K is more expressive but harder to keep specialized. Recommend K=3 to start.
2. **Loss decomposition**: PPO loss on the mixture (single loss) vs per-head PPO loss weighted by gate (decomposed). Empirically the mixture-loss is simpler and standard in MoE literature. Recommend mixture-loss.
3. **Eval boundary discipline**: does the gating's running state need reset at eval boundary? Per [[pearl_adaptive_carryover_discipline]], any stateful adaptive signal needs either reset or re-bootstrap. The gating is INSTANT (computed from encoder state per step) so no carry-over — but the per-head entropy controller IS stateful. Use neutral sentinel reset at eval boundary.
4. **Distillation interaction**: existing B-11-β-style Q→π distill could conflict with mixture policy. Verify the distill grad kernel handles K-head mixture output correctly before enabling distill.
## §7. Out of scope
- **Per-expert REWARD shaping** (which the BLOCKED v6 attempt proposed via multiplicative `r_trend × hold_bonus_factor`) — explicitly rejected because it violates Pearson alignment per §6.2 of the parent spec. NOTE: per-expert POLICY-LOSS shaping IS in scope (§1.4 aux prior loss) — the distinction is reward-vs-loss; reward stays clean (Pearson-preserving), loss has per-head regularizer.
- Per-expert IQN τ (different risk preference per head) — clean alternative to §1.4 but interacts with the existing IQN risk-stack. Defer until §1.4 is validated. Add as §1.5 follow-up if needed.
- Per-expert encoder (full MoE / §4.C of parent) — defer to follow-up spec if multi-head with shared encoder is insufficient.
- Cross-fold training (#3) — orthogonal intervention; can be combined with multi-head as a follow-up.
## §8. Spec/plan separation discipline (addresses BLOCKER 2)
This spec is INTENTIONALLY separate from `2026-06-02-regime-invariance-four-interventions.md` and from any implementation plan. The implementation plan for this spec will be drafted in a separate document AFTER:
1. Intervention #1 (R-multiple) ships and eval verdict is collected
2. If eval verdict is negative, sp-critical-reviewer reviews this spec
3. If reviewer APPROVES, the implementation plan is drafted
Per `feedback_no_partial_refactor`: spec and plan must align atomically. The plan does not exist yet.
## §9. Verified facts (audit trail for sp-critical-reviewer)
The following claims in this spec have been verified by direct code inspection:
| Claim | Verification |
|---|---|
| `RL_SLOTS_END = 757` post-grwwh | `crates/ml-alpha/src/rl/isv_slots.rs:1771` |
| `regime_observer` writes slots 696-701 (no trend_strength) | `crates/ml-alpha/cuda/rl_regime_observer.cu:36-100` and `crates/ml-alpha/src/rl/isv_slots.rs:1522-1534` |
| Reward kernel is `rl_fused_reward_pipeline.cu` (current Phase 5 logic) | `crates/ml-alpha/cuda/rl_fused_reward_pipeline.cu` |
| Per `feedback_pi_actor_collapses_without_entropy_floor`: entropy floor canonical pattern | pearl exists in MEMORY.md |
| Per `feedback_kill_runs_on_anomaly_quickly`: aggressive kill gates required | pearl exists in MEMORY.md |
| Per `feedback_no_partial_refactor`: spec/plan must align | pearl exists in MEMORY.md |
Claims NOT verified that the reviewer should check:
- Exact `PolicyHead` struct layout in `policy_head.rs` (referenced but not opened)
- Exact `EXPECTED_LEAVES` value pre-this-spec (referenced as 679 from prior R-multiple invariant test)
- Whether mixture-of-policies has any existing precedent in foxhunt (search for prior MoE attempts)

View File

@@ -0,0 +1,166 @@
# Regime-Invariance — Four Interventions for the Train→Eval Gap
**Date**: 2026-06-02
**Status**: Spec draft
**Branch**: ml-alpha-regime-observer (current)
**Reference SHA**: 63fc16f17 (v5.2 adaptive scaffold)
**Linked pearls**: [[pearl_surfer_baseline_was_train_only_never_eval]], [[pearl_reward_signal_anti_aligned_with_pnl]], [[project_adaptive_scaffold_session_2026_06_02]]
## §0. Why this spec exists
The session of 2026-06-01/02 burned 4 cluster smokes iterating on the surfer-scaffold controller (v4 → v5 → v5.1 → v5.2). Each iteration revealed a controller bug introduced by the previous fix. None reached eval verdict before being killed; only v5.2 (alpha-rl-grwwh) was let to run to completion. Pattern matches `feedback_going_in_circles_pattern`.
**Diagnostic conclusion**: the controller iterations are tuning the wrong knob. The fundamental failure mode is that **the agent learns a trend-specialized policy from fold-1 train data that does not generalize to OOS regime shifts**. Phase 5 hold-bonus teaches "ride winners, cut losers" — works in trends, fails in chop. No fade timing fixes this — the inductive bias itself is regime-dependent.
Perplexity consultation 2026-06-02 confirmed this is a known class of problem in financial RL with a documented playbook:
- Domain randomization over market regimes
- Regime-conditional architectures
- Distributionally robust objectives
- Microstructure features tied to invariants, not trend structure
This spec captures the 4 actionable interventions ranked by leverage/risk ratio. Each intervention is self-contained — they can ship in series or in parallel.
## §1. Intervention #1 — R-multiple reward (HIGHEST LEVERAGE)
**Problem**: Phase 5 hold-bonus rewards `hold_bonus × √hold_time` for profitable holds. This is dollar-magnitude reward in market-dependent units. In trending markets, holding profitable positions accumulates more dollars per step. In choppy markets, holding profitable positions reverses → reward drops → policy un-learns the hold pattern. Reward magnitude scales with the trend regime.
**Fix**: replace Phase 5 hold-bonus + long-ride-bonus with **R-multiple reward** normalized by entry-time ATR:
```c
// At close event (done > 0.5):
const float atr_at_entry = unit_atr_at_entry[b]; // captured at flat→open transition
const float r_multiple = realized_pnl_delta / fmaxf(atr_at_entry, EPS_ATR);
r = r_multiple; // dimensionless, regime-invariant
```
**Why this generalizes**: a $1000 win in high-vol ATR=$2000 → R=0.5. A $300 win in low-vol ATR=$600 → R=0.5. Same reward signal. Policy learns "capture 0.5 ATR worth of move" not "make $X regardless of vol."
### §1.1 Implementation
- New CUDA buffer: `unit_atr_at_entry_d` (per-account f32, captured at flat→open transition in `extract_position`)
- Read ATR from existing risk-stack signals (already computed for Kelly position sizing)
- Replace Phase 5 steps 2-4 entirely with R-multiple computation at close events
- **Drop the surfer-scaffold controller** — no longer needed if reward is regime-invariant
- Keep Phase 5 step 1 (entry cost) — fees are regime-invariant in absolute terms; alternatively also divide by ATR
### §1.2 Why surfer-scaffold controller can retire
The 753 ISV slot (`RL_SURFER_SCAFFOLD_WEIGHT_INDEX`) and its 3 config slots (754/755/756) become obsolete. R-multiple reward is its own invariant — no scaffold needed because the gradient is already regime-correct. Slots delete back to `RL_SLOTS_END=753`. Per `feedback_no_feature_flags` + `feedback_single_source_of_truth_no_duplicates` the scaffold infrastructure (controller kernel + bootstrap entries + Phase 5 gating logic) gets deleted in the same commit.
### §1.3 Falsification
Cluster smoke fold-1 b=1024 20k+5k. Verdict criteria:
- **PASS**: train wr ≥ 0.30 by step 5k AND eval wr ≥ 0.28 (within 10% of train) AND eval pnl > 0
- **FAIL**: eval wr < 0.25 OR eval pnl < $5M → R-multiple normalization insufficient; proceed to Intervention #2 (regime features) on top
**Scope**: ~30 LOC kernel + ~50 LOC trainer (ATR capture) + delete ~150 LOC scaffold infra. Net code REDUCTION.
## §2. Intervention #2 — Wire `regime_observer` features into policy network input
**Problem**: F1.3 already computes regime descriptors (Welford σ_pnl, dead_zone, recovery_factor, eps_live) but they're emitted to `diag.jsonl` only. The policy network has no input that tells it "we're in a chop regime, behave differently."
**Fix**: extend the encoder input vector with 4-6 regime features. Policy explicitly conditions on regime instead of baking trend-assumption into its core behavior.
### §2.1 Existing signals to wire
Read from ISV:
- `RL_REGIME_FLAT_COUNT_INDEX` — fraction of accounts that are flat (regime activity proxy)
- `RL_REGIME_WELFORD_SIGMA_INDEX` — pnl variance EMA (vol regime)
- `RL_REGIME_DEAD_ZONE_INDEX` — fraction of time in dead-zone (chop indicator)
- `RL_REGIME_RECOVERY_FACTOR_INDEX` — current/max DD ratio (drawdown regime)
- `RL_EDGE_PH_MEAN_INDEX` — Page-Hinkley edge-decay signal (per-trade EV regime)
### §2.2 Implementation
- Extend encoder input projection from current N dims to N+5 dims
- Bootstrap the new weight columns to small random or zero — agent learns to use them
- 5 new tokens broadcast to every account in the batch via existing ISV-read kernel
**Scope**: ~80 LOC (encoder projection weight matrix + 5 ISV reads in encoder forward) + checkpoint format bump
### §2.3 Caveat
This works ONLY when training data exposes multiple regimes (which #1 alone doesn't solve since we still train on fold 1). Therefore best deployed AFTER or WITH Intervention #3.
## §3. Intervention #3 — Cross-fold training (domain randomization)
**Problem**: Current walk-forward setup gives fixed fold 1 train + fold 1 eval. If fold 1 has a different regime mix than the deployment distribution, the agent over-fits to fold-1's regime. Per `pearl_fold_split_is_quarterly_regime_break` (SUPERSEDED for current PVC data but the principle stands), folds can differ substantially.
**Fix**: at episode start, sample a random fold from {0, 1, 2} for training; evaluate on a held-out fold (e.g., the latest quarter). Forces the policy to be **regime-mix robust**.
### §3.1 Implementation
- Modify walk-forward sampler in `alpha_rl_train.rs` (or trainer init) to randomly sample fold per batch element
- Need 3× the training data resident on PVC — already there (24 months ES MBP-10 per `pearl_mbp10_data_is_structurally_broken` correction)
- Eval stays on a held-out window (e.g., last 3 months never seen in train)
**Scope**: ~30 LOC trainer + verification that all 3 folds load cleanly
### §3.2 Risk
Trade rate may drop initially (more diverse regimes → more uncertainty → policy more conservative). Worth ~5k extra steps of training to compensate.
## §4. Intervention #4 — Mixture-of-experts gated by regime
**Problem**: A single policy network must serve all regimes. The optimal trading strategy is genuinely different across regimes (trend vs mean-revert). One policy averages them → mediocre everywhere.
**Fix**: 2-3 specialist policies + gating network. Gating reads regime features (from #2), outputs softmax weights over experts. Each expert specializes on its dominant regime.
**Now spec'd separately**: see `docs/superpowers/specs/2026-06-02-multi-head-policy-with-r-multiple.md` for the architecturally clean version (multi-head policy on shared encoder, no reward-shaping diversity, no invented ISV signals). That spec was written after sp-critical-reviewer BLOCKED the prior in-line v6 attempt for inventing a `trend_strength` signal that does not exist and re-introducing the anti-Pearson reward shaping pattern.
### §4.1 Architecture
- Expert A: trend-follower (current Phase 5 baseline behavior)
- Expert B: mean-reversion (different reward shaping — capture short oscillations)
- Optional Expert C: stay-flat (no positions when neither A nor B is confident)
- Gating net: 2-layer MLP on regime features → softmax(K) over experts
- Forward: `action = Σ_k gate_k × expert_k(state)`
- Backward: each expert's gradient weighted by `gate_k`; gate net learns from policy gradient
### §4.2 Implementation
- 3× the policy head parameters
- Each expert trained on the same data but the reward shaping differs (use a per-expert reward signal — trend-expert uses R-multiple of hold-duration; reversion-expert uses R-multiple of inverse-hold-duration)
- Most complex of the 4 interventions
**Scope**: ~250 LOC over 2 days. Defer until #1-#3 are validated.
## §5. Sequencing & decision points
### Order of operations
1. **Ship #1 (R-multiple reward)** — single commit, kills the surfer-scaffold infra in the same commit. Cluster smoke fold-1 b=1024 20k+5k. Plan at `docs/superpowers/plans/2026-06-02-regime-invariance-implementation.md`.
2. **Wait for #1 verdict**. If eval pnl > 0: #1 was sufficient.
3. **If #1 eval fails**: ship multi-head policy (#4 — separate spec at `docs/superpowers/specs/2026-06-02-multi-head-policy-with-r-multiple.md`). Architectural fix that does not violate Pearson alignment.
4. **If #1+#4 eval fails**: ship #3 (cross-fold training) and #2 (regime features) in sequence as remaining levers.
### Falsification criteria (each step)
Each intervention has the same eval-gate:
- Eval pnl > 0 (any positive amount) → PASS, ship
- Eval pnl in [$2M, 0] → MARGINAL, continue to next intervention
- Eval pnl < $2M → FAIL, deeper rethink (maybe the architecture / data pipeline itself)
### Termination clause
If ALL 4 interventions ship and eval still fails, the conclusion is that this neural architecture (mamba2 + DQN-IQN + PPO) cannot extract alpha from MBP-10 ES futures at b=1024 20k steps. We then either:
- Change data scale (much longer training, multi-instrument)
- Change architecture (attention-based, longer context window)
- Change approach (behavior cloning from a profitable rule-based bot, then fine-tune)
## §6. Reward-shaping discipline (cross-cutting)
All 4 interventions assume the reward signal must be:
1. **Regime-invariant in units** (R-multiple, not dollars)
2. **Aligned with realized pnl direction** (Pearson ≥ 0.7 per `pearl_reward_signal_anti_aligned_with_pnl`)
3. **Sparse at close-events only** (not per-step inflation)
4. **Bounded** (no exponential growth with hold time)
Phase 5 hold-bonus violates (1), (3), (4). The v5.2 adaptive scaffold partially restored (2) but not (1). R-multiple satisfies all four cleanly.
## §7. Open questions to resolve before implementation
1. **ATR source for R-multiple denominator**: use existing Kelly ATR (already computed) or compute fresh? Check Kelly controller for double-update risk.
2. **Bootstrap value for ATR** at first-trade: ATR EMA needs warmup. Use `max(ATR_ema, ATR_floor)` with floor ≈ 1 ES tick × tick-multiplier.
3. **Eval-boundary discipline**: R-multiple uses ATR EMA — must reset at eval boundary per `pearl_adaptive_carryover_discipline` (or NOT, depending on whether we want eval to start with train-end calibration).
4. **Fold sampling weight**: equal-weight all 3 folds, or weight by regime-novelty? Active domain randomization (ADR) literature suggests the latter; defer until #1 works.

374
scripts/compare-backward-state.py Executable file
View File

@@ -0,0 +1,374 @@
#!/usr/bin/env python3
"""Compare raw backward-pipeline state dumps between two determinism runs.
Reads the dumps produced by
`IntegratedTrainer::dump_backward_state_for_debug` (env-gated on
`FOXHUNT_DETERMINISM_DEBUG_BACKWARD=1`) and prints a per-step
exact-bytewise-equality verdict for each of the 4 buffer groups
identified by Phase 2.5's residual:
Group G — PER replay sample hidden states + scalars:
- sampled_h_t, sampled_h_tp1
- sampled_actions, sampled_rewards, sampled_dones, sampled_log_pi_old
Group H — Head forward outputs (post-fwd, pre-bwd):
- q_logits, pi_logits, v_pred
- iqn_q_values, dueling_q_composed
Group I — Pre-reduce per-batch / per-row param-grad scratches:
- cfc_grad_w_in_scratch, cfc_grad_w_rec_scratch
- cfc_grad_b_scratch, cfc_grad_tau_scratch
- vsn_grad_w_scratch, vsn_grad_b_scratch
- attn_grad_q_scratch
- grad_ln_a_gain_per_row, grad_ln_a_bias_per_row
- grad_ln_b_gain_per_row, grad_ln_b_bias_per_row
Group J — Per-head grad_h_t outputs + combined accumulator:
- grad_h_t_q, grad_h_t_pi, grad_h_t_v
- grad_h_t_frd, grad_h_t_outcome
- grad_h_t_combined
Decision tree (per Phase 2.6 dispatch):
(1) Group G divergent → step-2 PER K-loop sample is non-det (aliased
pearl-twin to Phase 2's PER tree rebuild bug)
(2) Group H divergent → head forward kernel (likely NoisyLinear noise
resampling between forward and backward) is non-det
(3) Group I EQUAL but later (final-grad) DIVERGE → `reduce_axis0` (or
`layer_norm_reduce_param_grads`) is the bug; canonical
Phase-2-PER-rebuild fix applies (single block + __syncthreads +
fixed thread-ID order)
(4) Group J divergent → `grad_h_accumulate_scaled` or a specific
head backward kernel writing grad_h_t is non-det
The script reports the FIRST step that diverges and which group(s)
diverged there. If Group I is the only divergent group at the first
divergent step → confirms hypothesis (3). If both Group I and Group J
diverge → upstream head bwd is also wrong (because Group J feeds
encoder backward via grad_h_accumulate; if its inputs differ, encoder
gradients differ). If Group H diverges but I/J don't → head forward is
non-det but somehow downstream gradients normalize (unlikely).
Exit codes:
0 both runs identical for all dumped buffers at all dumped steps
1 divergence found; verdict printed naming the candidate group
2 inputs missing or malformed
Usage:
python3 scripts/compare-backward-state.py <run_a_dir> <run_b_dir>
Each dir contains `step_{0,1,2,3}_g_<name>.bin`,
`step_{0,1,2,3}_h_<name>.bin`, `step_{0,1,2,3}_i_<name>.bin`,
`step_{0,1,2,3}_j_<name>.bin` (prefix discriminates buffer group).
Per `feedback_cpu_is_read_only`: this script ONLY reads device dumps;
it performs no compute that affects training.
"""
from __future__ import annotations
import argparse
import struct
import sys
from pathlib import Path
# Each entry: (logical name, on-disk basename, element struct format,
# element size in bytes). The groups partition the candidate space
# so the verdict can name the culprit class precisely.
# Naming convention: files are written as
# `step_<N>_<group_letter>_<basename>.bin` where group_letter ∈ {g,h,i,j}.
GROUP_G_PER_SAMPLE = [
("sampled_h_t", "g_sampled_h_t", "f", 4),
("sampled_h_tp1", "g_sampled_h_tp1", "f", 4),
("sampled_actions", "g_sampled_actions", "i", 4),
("sampled_rewards", "g_sampled_rewards", "f", 4),
("sampled_dones", "g_sampled_dones", "f", 4),
("sampled_log_pi_old", "g_sampled_log_pi_old", "f", 4),
]
GROUP_H_HEAD_OUTPUTS = [
("q_logits", "h_q_logits", "f", 4),
("pi_logits", "h_pi_logits", "f", 4),
("v_pred", "h_v_pred", "f", 4),
("iqn_q_values", "h_iqn_q_values", "f", 4),
("dueling_q_composed", "h_dueling_q_composed", "f", 4),
]
GROUP_I_GRAD_SCRATCH = [
("cfc_grad_w_in_scratch", "i_cfc_grad_w_in_scratch", "f", 4),
("cfc_grad_w_rec_scratch", "i_cfc_grad_w_rec_scratch", "f", 4),
("cfc_grad_b_scratch", "i_cfc_grad_b_scratch", "f", 4),
("cfc_grad_tau_scratch", "i_cfc_grad_tau_scratch", "f", 4),
("vsn_grad_w_scratch", "i_vsn_grad_w_scratch", "f", 4),
("vsn_grad_b_scratch", "i_vsn_grad_b_scratch", "f", 4),
("attn_grad_q_scratch", "i_attn_grad_q_scratch", "f", 4),
("grad_ln_a_gain_per_row", "i_grad_ln_a_gain_per_row", "f", 4),
("grad_ln_a_bias_per_row", "i_grad_ln_a_bias_per_row", "f", 4),
("grad_ln_b_gain_per_row", "i_grad_ln_b_gain_per_row", "f", 4),
("grad_ln_b_bias_per_row", "i_grad_ln_b_bias_per_row", "f", 4),
]
GROUP_J_GRAD_H_T = [
("grad_h_t_q", "j_grad_h_t_q", "f", 4),
("grad_h_t_pi", "j_grad_h_t_pi", "f", 4),
("grad_h_t_v", "j_grad_h_t_v", "f", 4),
("grad_h_t_frd", "j_grad_h_t_frd", "f", 4),
("grad_h_t_outcome", "j_grad_h_t_outcome", "f", 4),
("grad_h_t_combined", "j_grad_h_t_combined", "f", 4),
]
# Phase 2.6 sub-2 (after first run surfaced grad_h_t_outcome
# divergence at step 2 with Δ ~ 0.003): outcome head's inputs
# (labels + w_d/b_d) are now dumped as Group K to localise the
# upstream-input that drifted (kernel is sole-writer so can't
# introduce divergence from equal inputs).
GROUP_K_OUTCOME_INPUTS = [
("outcome_labels", "k_outcome_labels", "i", 4),
("outcome_w", "k_outcome_w", "f", 4),
("outcome_b", "k_outcome_b", "f", 4),
]
ALL_GROUPS = [
("G (PER replay sample)", GROUP_G_PER_SAMPLE),
("H (head fwd outputs)", GROUP_H_HEAD_OUTPUTS),
("I (pre-reduce grad scratch)", GROUP_I_GRAD_SCRATCH),
("J (per-head grad_h_t)", GROUP_J_GRAD_H_T),
("K (outcome head inputs)", GROUP_K_OUTCOME_INPUTS),
]
def read_typed(path: Path, fmt: str, elem_size: int) -> list:
"""Read a raw little-endian binary file as `fmt`-typed elements."""
data = path.read_bytes()
if len(data) % elem_size != 0:
raise ValueError(
f"{path}: size {len(data)} not a multiple of {elem_size}"
)
n = len(data) // elem_size
return list(struct.unpack(f"<{n}{fmt}", data))
def compare_lists(a: list, b: list, label: str) -> tuple[bool, int, str]:
"""Bytewise exact equality. Returns (equal, first_diff_idx, summary)."""
if len(a) != len(b):
return False, -1, f"{label}: length mismatch A={len(a)} B={len(b)}"
diffs: list[int] = []
max_abs_delta = 0.0
for i, (x, y) in enumerate(zip(a, b)):
if x != y:
diffs.append(i)
try:
ad = abs(float(x) - float(y))
if ad > max_abs_delta:
max_abs_delta = ad
except (TypeError, ValueError):
pass
if len(diffs) >= 5:
break
if not diffs:
return True, -1, f"{label}: EQUAL ({len(a)} elements)"
sample = ", ".join(
f"[{i}]: A={a[i]!r} B={b[i]!r}" for i in diffs[:3]
)
return False, diffs[0], (
f"{label}: DIVERGE at idx {diffs[0]} "
f"max|Δ|={max_abs_delta:.6g} "
f"(showing up to 3): {sample}"
)
def main() -> int:
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("run_a", type=Path)
parser.add_argument("run_b", type=Path)
args = parser.parse_args()
if not args.run_a.is_dir() or not args.run_b.is_dir():
print(
f"ERROR: run dirs missing: {args.run_a} or {args.run_b}",
file=sys.stderr,
)
return 2
overall_any_diverge = False
per_step_group_results: dict[tuple[int, str], list] = {}
first_diverge_step: int | None = None
first_diverge_groups: set[str] = set()
for step in (0, 1, 2, 3):
print(f"=== step {step} ===")
any_present = False
step_had_diverge = False
step_diverged_groups: set[str] = set()
for group_label, members in ALL_GROUPS:
print(f" -- group {group_label} --")
group_results: list[tuple[str, bool, str]] = []
group_any_present = False
for name, basename, fmt, elem_size in members:
pa = args.run_a / f"step_{step}_{basename}.bin"
pb = args.run_b / f"step_{step}_{basename}.bin"
if not (pa.is_file() and pb.is_file()):
continue
group_any_present = True
any_present = True
try:
a = read_typed(pa, fmt, elem_size)
b = read_typed(pb, fmt, elem_size)
except ValueError as e:
print(f" ERROR reading {name}: {e}")
return 2
equal, _, summary = compare_lists(a, b, name)
print(f" {summary}")
group_results.append((name, equal, summary))
if not equal:
overall_any_diverge = True
step_had_diverge = True
step_diverged_groups.add(group_label)
if not group_any_present:
print(" SKIP — no buffers dumped for this group/step")
per_step_group_results[(step, group_label)] = group_results
if not any_present:
print(" SKIP — no buffers dumped for this step")
continue
if step_had_diverge and first_diverge_step is None:
first_diverge_step = step
first_diverge_groups = step_diverged_groups
print()
if not overall_any_diverge:
print(
"VERDICT: all dumped backward-pipeline buffers match across\n"
"all dumped steps. The bug is OUTSIDE Groups G/H/I/J. Per\n"
"the Phase 2.6 dispatch STOP rule: this is an UNEXPECTED\n"
"finding (the 4 groups were chosen to enumerate the entire\n"
"step-2 backward pipeline). Possibilities:\n"
" - Adam optimizer step is non-deterministic (parallel\n"
" reduction in moment update)\n"
" - Bellman target projection has order-dependent writes\n"
" - A reducer kernel writes scratch BEFORE Group I dump\n"
" captures it (race between dump and last reducer launch)\n"
"Phase 2.7 dispatch needed — STOP per\n"
"feedback_systematic_debugging."
)
return 0
assert first_diverge_step is not None
print(f"VERDICT (first divergent step = {first_diverge_step}):")
print(f" diverged group(s): {sorted(first_diverge_groups)}")
print()
print("Per-group culprit table:")
for group_label, _ in ALL_GROUPS:
results = per_step_group_results.get((first_diverge_step, group_label), [])
any_div = any(not eq for _, eq, _ in results)
if any_div:
divnames = [name for name, eq, _ in results if not eq]
print(f" Group {group_label}: DIVERGE in {divnames}")
else:
present = len(results) > 0
tag = "EQUAL" if present else "(no dump)"
print(f" Group {group_label}: {tag}")
print()
print("Next-step interpretation (per Phase 2.6 decision tree):")
# Subcase (1): only Group G diverges.
if first_diverge_groups == {"G (PER replay sample)"}:
print(
" Group G (PER replay sample) is the ONLY divergent group.\n"
" → step-2 PER K-loop sample IS the bug, despite Phase 2.5's\n"
" end-of-step PER state EQUAL. Most likely: the K-loop's\n"
" rl_per_sample is reading from `replay_h_t_d` /\n"
" `replay_h_tp1_d` which were written by a non-deterministic\n"
" `rl_per_push_ring`/flush kernel earlier in step 2. Inspect\n"
" `rl_per_push_ring.cu` and `rl_per_push_flush.cu` for\n"
" parallel-reduction / atomicAdd patterns. (Aliased pearl-\n"
" twin to Phase 2's PER tree rebuild bug.)"
)
return 1
# Subcase (2): only Group H diverges (or H is the most upstream).
if "H (head fwd outputs)" in first_diverge_groups and \
"G (PER replay sample)" not in first_diverge_groups:
print(
" Group H (head fwd outputs) diverges with EQUAL Group G.\n"
" → head forward kernel is non-det. Most likely candidate:\n"
" NoisyLinear noise resampling (factored noise) between\n"
" online forward and target forward, especially if the\n"
" resampling kernel uses parallel write to shared state\n"
" or if cuBLAS PEDANTIC isn't actually active (sanity-\n"
" check `[cublas_determinism]` line in run logs). Less\n"
" likely but possible: IQN tau sampling order across the\n"
" K-loop, or NoisyLinear reset_noise + forward overlap.\n"
" → check rl_iqn_forward / rl_noisy_linear_forward for\n"
" parallel-reduction or atomic patterns; verify cuBLAS\n"
" handle's math mode is CUBLAS_PEDANTIC_MATH at every\n"
" site (mamba2_block, dqn, iqn — 3 sites)."
)
return 1
# Subcase (3): Group I diverges but Group J doesn't (or upstream of J).
# NB: if Group I diverges, Group J will almost always diverge too
# because J is downstream of the same head backward kernels. But
# the KEY distinction is: do the pre-reduce scratches (I) diverge,
# or do they match while the post-reduce final grads (dumped by
# dump_mamba2_state_for_debug under "grad_vsn_w" etc.) diverge?
# That comparison happens IN the mamba2 dump verdict; we report it
# via interpretation hint here.
if "I (pre-reduce grad scratch)" in first_diverge_groups:
print(
" Group I (pre-reduce per-batch/per-row grad scratches)\n"
" diverges. → upstream per-batch backward kernel is non-det.\n"
" Candidates by scratch:\n"
" cfc_grad_*_scratch → cfc_step_backward_batched\n"
" vsn_grad_*_scratch → variable_selection_bwd\n"
" attn_grad_q_scratch → attention_pool_bwd\n"
" grad_ln_*_per_row → layer_norm_bwd\n"
" Inspect which scratch's first divergence is the most\n"
" upstream (smallest k_iter or smallest n_rows index).\n"
" Look for parallel reductions with __threadfence (not\n"
" __syncthreads), atomicAdd, or warp-shuffle ordering\n"
" issues. Canonical Phase-2 fix: single block + __syncthreads\n"
" + fixed thread-ID accumulation order."
)
elif "J (per-head grad_h_t)" in first_diverge_groups:
print(
" Group J (per-head grad_h_t) diverges with EQUAL Group I.\n"
" → `grad_h_accumulate_scaled` or a head backward writing\n"
" grad_h_t (without scratch reduction) is the culprit.\n"
" → check grad_h_accumulate.cu: should be sole-writer-per-i\n"
" (element-wise `+=`). If λ is host-passed scalar, host-\n"
" side ISV mirror read should be order-stable. Also check\n"
" head bwd kernels (dqn_grad_w_b_h_t / pi_head_bwd / v_head_\n"
" bwd / iqn_bwd_relu_hadamard) — they all write grad_h_t\n"
" as OVERWRITE per (b, c) sole writer; if any uses `+=`\n"
" with parallel writes, that's the bug."
)
# If post-reduce final grads (dumped by mamba2 dumper) diverge while
# pre-reduce scratches (Group I) are EQUAL, that's the canonical
# `reduce_axis0` bug. Surface that diagnostic in the message even
# if we can't directly check the mamba2 dump here.
if "I (pre-reduce grad scratch)" not in first_diverge_groups and \
"J (per-head grad_h_t)" not in first_diverge_groups and \
first_diverge_groups:
print(
"\n WARNING: divergent group(s) not in expected G/H/I/J set\n"
" → unexpected finding, STOP per dispatch rule."
)
return 1
# Cross-check: if BOTH I and J are EQUAL but the mamba2 dump
# shows trunk grads DIVERGE (per Phase 2.5's finding), the only
# remaining place is `reduce_axis0` / `layer_norm_reduce_param_grads`.
print(
"\n Cross-check: if mamba2 dump shows final trunk grads (post-\n"
" reduce) DIVERGE but Group I scratches above are EQUAL → the\n"
" reducer kernel itself (`reduce_axis0` or\n"
" `layer_norm_reduce_param_grads`) is the bug. Apply canonical\n"
" Phase-2-PER-rebuild fix: single block + __syncthreads +\n"
" fixed thread-ID accumulation order, no __threadfence."
)
return 1
if __name__ == "__main__":
sys.exit(main())

306
scripts/compare-mamba2-state.py Executable file
View File

@@ -0,0 +1,306 @@
#!/usr/bin/env python3
"""Compare raw mamba2 state dumps between two determinism runs.
Reads the dumps produced by
`PerceptionTrainer::dump_mamba2_state_for_debug` (env-gated on
`FOXHUNT_DETERMINISM_DEBUG_MAMBA2=1`) and prints a per-step
exact-bytewise-equality verdict for each of:
Per mamba2 layer (l1, l2):
- w_in, b_in, w_a, b_a, w_b, b_b, w_c, w_out, b_out (weights)
- h_enriched_seq (forward output)
Encoder activations (single tensors, shared across layers):
- vsn_out, ln_a_out, ln_b_out, attn_context, h_t
Hypothesis table (spec §2.A sub-investigation, Phase 2.2):
CASE 2.A.1 (reduce non-det): mamba2_grad_w_c (after-reduce)
diverges, intermediates before
reduce match → `mamba2_alpha_reduce_d_w_c`
or `_d_proj` kernel is non-det.
(only catchable via post-reduce
checksum — dump is for weights/outputs)
CASE 2.A.2 (forward non-det): weights match at step N, but
`mamba2_l1_h_enriched_seq.bin` (forward
output) diverges with identical
weights → forward scan kernel is
non-deterministic (very unlikely
given per-(i, j) thread design).
CASE 2.A.3 (upstream non-det): mamba2 L1 forward output diverges
AT THE SAME STEP AS vsn_out → the
non-determinism enters BEFORE mamba2
(VSN forward / backward, or earlier).
CASE 2.A.4 (weights diverge): weights diverge at step N before
any activation diverges →
upstream optimizer / gradient
accumulation (Adam iteration order,
grad-reduction non-det) wrote
different bits to the weights.
Output is written to stdout. Exit codes:
0 both runs identical for all dumped buffers at all dumped steps
1 divergence found; verdict printed (one of 2.A.2 / 2.A.3 / 2.A.4)
2 inputs missing or malformed
Usage:
python3 scripts/compare-mamba2-state.py <run_a_dir> <run_b_dir>
Each dir contains `step_{0,1,2,3}_<name>.bin`.
Per `feedback_cpu_is_read_only`: this script ONLY reads device dumps;
it performs no compute that affects training.
"""
from __future__ import annotations
import argparse
import struct
import sys
from pathlib import Path
# Buffers dumped at each step (defined once, used for all steps).
# Order MATCHES the dump order in `dump_mamba2_state_for_debug` so the
# output diff is reproducible.
WEIGHTS_L1 = [
"mamba2_l1_w_in", "mamba2_l1_b_in",
"mamba2_l1_w_a", "mamba2_l1_b_a",
"mamba2_l1_w_b", "mamba2_l1_b_b",
"mamba2_l1_w_c",
"mamba2_l1_w_out", "mamba2_l1_b_out",
]
WEIGHTS_L2 = [
"mamba2_l2_w_in", "mamba2_l2_b_in",
"mamba2_l2_w_a", "mamba2_l2_b_a",
"mamba2_l2_w_b", "mamba2_l2_b_b",
"mamba2_l2_w_c",
"mamba2_l2_w_out", "mamba2_l2_b_out",
]
ACTIVATIONS = [
"vsn_out",
"mamba2_l1_h_enriched_seq",
"ln_a_out",
"mamba2_l2_h_enriched_seq",
"ln_b_out",
"attn_context",
"h_t",
]
# Phase 2.4 (2026-06-02): backward scratch + reduced output, to localise
# the `mamba2_grad_w_c_l1` residual that survives PEDANTIC cuBLAS.
BACKWARD_SCRATCH = [
"mamba2_l1_d_w_c_per_sample", # pre-reduce, [N, sh2, state_d]
"mamba2_l1_dw_c_reduced", # post-reduce, [sh2, state_d]
"mamba2_l1_d_h_s2", # [N, sh2] — scan_bwd_seq output
"mamba2_l2_d_w_c_per_sample",
"mamba2_l2_dw_c_reduced",
# Phase 2.4 follow-up: dw_in / dw_a / dw_b (cuBLAS GEMM outputs)
# and d_a_proj / d_b_proj (reduce_d_proj outputs).
"mamba2_l1_d_a_per_channel", # scan_bwd_seq output
"mamba2_l1_d_b_per_channel", # scan_bwd_seq output
"mamba2_l1_d_a_proj_2d", # reduce_d_proj output
"mamba2_l1_d_b_proj_2d", # reduce_d_proj output
"mamba2_l1_dw_in", # cuBLAS GEMM output
"mamba2_l1_dw_a", # cuBLAS GEMM output
"mamba2_l1_dw_b", # cuBLAS GEMM output
"mamba2_l1_db_in",
"mamba2_l1_db_a",
"mamba2_l1_db_b",
"mamba2_l1_d_x",
"mamba2_l1_d_x_from_in",
]
# Phase 2.4 follow-up #2: trunk weights + their grads to localise which
# upstream-of-mamba2 weight first diverges (a divergent vsn_w_d at step 1
# would propagate to differing forward inputs to mamba2 at step 2).
TRUNK_STATE = [
"vsn_w", "attn_q", "cfc_w_in", "cfc_w_rec",
"ln_a_gain", "ln_b_gain",
"grad_vsn_w", "grad_attn_q",
"grad_cfc_w_in", "grad_cfc_w_rec",
"grad_ln_a_gain", "grad_ln_b_gain",
]
ALL_BUFFERS = WEIGHTS_L1 + WEIGHTS_L2 + ACTIVATIONS + BACKWARD_SCRATCH + TRUNK_STATE
def read_f32(path: Path) -> list[float]:
"""Read a raw little-endian f32 binary file into a Python list."""
data = path.read_bytes()
if len(data) % 4 != 0:
raise ValueError(f"{path}: size {len(data)} not a multiple of 4")
n = len(data) // 4
return list(struct.unpack(f"<{n}f", data))
def compare_lists(a: list, b: list, label: str) -> tuple[bool, int, str]:
"""Compare two equal-length lists; return (equal, first_diff_idx, summary)."""
if len(a) != len(b):
return False, -1, f"{label}: length mismatch A={len(a)} B={len(b)}"
diffs: list[int] = []
for i, (x, y) in enumerate(zip(a, b)):
if x != y:
diffs.append(i)
if len(diffs) >= 5:
break
if not diffs:
return True, -1, f"{label}: EQUAL ({len(a)} elements)"
sample = ", ".join(
f"[{i}]: A={a[i]!r} B={b[i]!r}" for i in diffs[:3]
)
return False, diffs[0], (
f"{label}: DIVERGE at idx {diffs[0]} "
f"(showing up to 3): {sample}"
)
def main() -> int:
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("run_a", type=Path)
parser.add_argument("run_b", type=Path)
args = parser.parse_args()
if not args.run_a.is_dir() or not args.run_b.is_dir():
print(
f"ERROR: run dirs missing: {args.run_a} or {args.run_b}",
file=sys.stderr,
)
return 2
# Per-step verdict. Tracks the FIRST step that diverges and which
# buffers diverge there.
overall_any_diverge = False
first_diverge_step: int | None = None
first_diverge_state: dict[str, bool] = {}
for step in (0, 1, 2, 3):
print(f"=== step {step} ===")
step_state: dict[str, bool] = {}
any_present = False
for name in ALL_BUFFERS:
pa = args.run_a / f"step_{step}_{name}.bin"
pb = args.run_b / f"step_{step}_{name}.bin"
if not (pa.is_file() and pb.is_file()):
# Buffer missing — silently skip; partial dumps okay.
continue
any_present = True
try:
a = read_f32(pa)
b = read_f32(pb)
except ValueError as e:
print(f" ERROR reading {name}: {e}")
return 2
equal, _, summary = compare_lists(a, b, name)
# Compact one-line summary so the per-step block is readable
# at a glance.
print(f" {summary}")
step_state[name] = equal
if not equal:
overall_any_diverge = True
if not any_present:
print(" SKIP — no buffers dumped for this step")
continue
if (not all(step_state.values())) and first_diverge_step is None:
first_diverge_step = step
first_diverge_state = dict(step_state)
print()
if not overall_any_diverge:
print("VERDICT: all dumped buffers match across all dumped steps.")
print("Mamba2 state itself is bit-identical between runs. The")
print("downstream `checksums.encoder_output` divergence observed")
print("by determinism-check.sh must come from later kernels (CfC")
print("/ heads / attention) AFTER mamba2 produces its outputs.")
return 0
assert first_diverge_step is not None
# Categorise: which buffers diverge first?
diverged = [name for name, eq in first_diverge_state.items() if not eq]
l1_weights_diverged = any(
name in WEIGHTS_L1 for name in diverged
)
l2_weights_diverged = any(
name in WEIGHTS_L2 for name in diverged
)
vsn_out_diverged = "vsn_out" in diverged
l1_out_diverged = "mamba2_l1_h_enriched_seq" in diverged
l2_out_diverged = "mamba2_l2_h_enriched_seq" in diverged
ln_a_diverged = "ln_a_out" in diverged
ln_b_diverged = "ln_b_out" in diverged
attn_diverged = "attn_context" in diverged
h_t_diverged = "h_t" in diverged
print(f"VERDICT (first divergent step = {first_diverge_step}):")
print(f" diverged buffers: {diverged}")
print()
if l1_weights_diverged or l2_weights_diverged:
print(" CASE 2.A.4 — WEIGHTS DIVERGE upstream of forward.")
which = []
if l1_weights_diverged: which.append("L1")
if l2_weights_diverged: which.append("L2")
print(f" Mamba2 {'/'.join(which)} weights diverged at step "
f"{first_diverge_step}. This means the PRIOR step's")
print(" backward+Adam update wrote different bits — the")
print(" non-determinism source is in:")
print(" (a) mamba2 backward kernels (d_a/d_b/d_w_c per-channel")
print(" writes are deterministic by construction, but")
print(" their REDUCE kernels mamba2_alpha_reduce_d_proj /")
print(" _d_w_c could be non-det if launch geometry races)")
print(" (b) Mamba2AdamW.step_from_buffers_gpu_clip — the")
print(" grad-norm phase1+phase2 tree-reduce. Phase 2")
print(" appears single-block per `grad_norm.cu`, but")
print(" re-audit if reduce kernels look fine.")
print(" (c) Encoder upstream backward (CfC / LN / VSN /")
print(" attention) writing non-det grads that feed")
print(" mamba2 W_in via dx propagation through L1 / L2.")
print(" Fix: identify which weight diverged FIRST and audit")
print(" its gradient pipeline. Next dispatch should add")
print(" checksums for the encoder backward grad outputs")
print(" (grad_vsn_w_d, grad_cfc_*, grad_attn_q_d, ...) and the")
print(" mamba2 reduced grad outputs (mamba2_grads_buffers.dw_*).")
return 1
if vsn_out_diverged and not l1_out_diverged:
print(" CASE 2.A.3 — UPSTREAM (VSN) non-determinism.")
print(" vsn_out diverges before mamba2 L1 output does. The")
print(" issue is in VSN forward (vsn_fwd kernel) or earlier")
print(" (snap_feature_assemble_batched). Apply §2.D audit")
print(" to the VSN forward kernel.")
return 1
if l1_out_diverged and not l1_weights_diverged:
# Forward output diverges with identical weights.
print(" CASE 2.A.2 — MAMBA2 L1 FORWARD non-determinism.")
print(f" At step {first_diverge_step}, mamba2 L1 weights match")
print(" bit-exactly between runs, but the L1 forward output")
print(" h_enriched_seq diverges. The forward scan kernel")
print(" `mamba2_alpha_scan_fwd_seq` is per-thread sequential")
print(" (one thread per (i, j), each thread does its own")
print(" K-step scan with no atomics) — divergence here would")
print(" be surprising and would require auditing the per-thread")
print(" register-array initialisation + the cuBLAS GEMMs that")
print(" feed a_proj / b_proj inputs (w_in / w_a / w_b forward).")
print(" Likely culprit: cuBLAS split-K non-determinism in the")
print(" W_in / W_a / W_b forward GEMMs (TF32 / heuristic).")
return 1
if ln_a_diverged or ln_b_diverged or attn_diverged or l2_out_diverged or h_t_diverged:
# Divergence enters downstream of mamba2 L1 but with L1 outputs/weights equal.
print(" CASE 2.A.5 — DOWNSTREAM (post-mamba2) non-determinism.")
print(" Mamba2 L1 outputs match but later layer outputs diverge.")
for n in ("ln_a_out", "mamba2_l2_h_enriched_seq", "ln_b_out",
"attn_context", "h_t"):
if n in first_diverge_state:
eq = first_diverge_state[n]
print(f" {n}: {'EQUAL' if eq else 'DIVERGE'}")
print(" Audit the FIRST diverged downstream buffer's kernel.")
return 1
print(" UNEXPECTED state — divergence pattern does not match a")
print(" defined sub-case. Review per-buffer summary above and")
print(" extend the verdict logic.")
return 1
if __name__ == "__main__":
sys.exit(main())

205
scripts/compare-per-state.py Executable file
View File

@@ -0,0 +1,205 @@
#!/usr/bin/env python3
"""Compare raw PER state dumps between two determinism runs.
Reads the dumps produced by
`IntegratedTrainer::dump_per_state_for_debug` (env-gated on
`FOXHUNT_DETERMINISM_DEBUG=1`) and prints a per-step exact-equality
verdict for each of:
- `sample_prng_d` (B u32) — device PRNG state per batch
- `priority_tree_d` (2 × cap f32) — PER priority sum-tree
- `sample_indices_d` (B u32) — sampled leaf indices
Hypothesis table (spec §2.F sub-investigation):
HYPOTHESIS A (PRNG): prng differs → seed propagation
between runs (xorshift32 bug
cold-init diverging)
HYPOTHESIS B (Tree-walk): prng EQUAL, tree EQUAL, → fp comparison
indices differ `u <= left_val`
flipping near the
fp epsilon boundary
inside rl_per_sample
HYPOTHESIS C (Tree-rebuild): prng EQUAL, tree differs → parallel
(even tiny f32 deltas) reduction order in
rl_per_tree_rebuild
(children-before-
parents ordering)
Output is written to stdout. Exit codes:
0 both runs identical for all three buffers at all dumped steps
1 divergence found; verdict printed (one of A / B / C)
2 inputs missing or malformed
Usage:
python3 scripts/compare-per-state.py <run_a_dir> <run_b_dir> [--b-size N]
Each dir contains `step_{0,1,2}_{prng,tree,indices}.bin`.
Per `feedback_cpu_is_read_only`: this script ONLY reads device dumps;
it performs no compute that affects training.
"""
from __future__ import annotations
import argparse
import sys
from pathlib import Path
import struct
def read_u32(path: Path) -> list[int]:
"""Read a raw little-endian u32 binary file into a Python list."""
data = path.read_bytes()
if len(data) % 4 != 0:
raise ValueError(f"{path}: size {len(data)} not a multiple of 4")
n = len(data) // 4
return list(struct.unpack(f"<{n}I", data))
def read_f32(path: Path) -> list[float]:
"""Read a raw little-endian f32 binary file into a Python list."""
data = path.read_bytes()
if len(data) % 4 != 0:
raise ValueError(f"{path}: size {len(data)} not a multiple of 4")
n = len(data) // 4
return list(struct.unpack(f"<{n}f", data))
def compare_lists(a: list, b: list, label: str) -> tuple[bool, int, str]:
"""Compare two equal-length lists; return (equal, first_diff_idx, summary)."""
if len(a) != len(b):
return False, -1, f"{label}: length mismatch A={len(a)} B={len(b)}"
diffs = []
for i, (x, y) in enumerate(zip(a, b)):
if x != y:
diffs.append(i)
if len(diffs) >= 5:
break
if not diffs:
return True, -1, f"{label}: EQUAL ({len(a)} elements)"
sample = ", ".join(
f"[{i}]: A={a[i]!r} B={b[i]!r}" for i in diffs[:3]
)
return False, diffs[0], (
f"{label}: DIVERGE at idx {diffs[0]} "
f"(showing up to 3): {sample}"
)
def main() -> int:
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("run_a", type=Path)
parser.add_argument("run_b", type=Path)
parser.add_argument(
"--b-size",
type=int,
default=128,
help="batch size used by the run (default: 128)",
)
args = parser.parse_args()
if not args.run_a.is_dir() or not args.run_b.is_dir():
print(
f"ERROR: run dirs missing: {args.run_a} or {args.run_b}",
file=sys.stderr,
)
return 2
# Track verdict across steps. We pick the FIRST step that shows
# divergence and report which buffer(s) diverge there.
overall_any_diverge = False
first_diverge_step: int | None = None
first_diverge_state: dict[str, bool] = {}
for step in (0, 1, 2):
print(f"=== step {step} ===")
files = {
"prng": (
args.run_a / f"step_{step}_prng.bin",
args.run_b / f"step_{step}_prng.bin",
read_u32,
),
"indices": (
args.run_a / f"step_{step}_indices.bin",
args.run_b / f"step_{step}_indices.bin",
read_u32,
),
"tree": (
args.run_a / f"step_{step}_tree.bin",
args.run_b / f"step_{step}_tree.bin",
read_f32,
),
}
# Skip the step if any file missing.
missing = [name for name, (pa, pb, _) in files.items()
if not (pa.is_file() and pb.is_file())]
if missing:
print(f" SKIP — missing files: {missing}")
continue
step_state: dict[str, bool] = {}
for name, (pa, pb, reader) in files.items():
try:
a = reader(pa)
b = reader(pb)
except ValueError as e:
print(f" ERROR reading {name}: {e}")
return 2
equal, _, summary = compare_lists(a, b, name)
print(f" {summary}")
step_state[name] = equal
if not equal:
overall_any_diverge = True
if (not all(step_state.values())) and first_diverge_step is None:
first_diverge_step = step
first_diverge_state = dict(step_state)
print()
if not overall_any_diverge:
print("VERDICT: all three buffers match across all dumped steps.")
print("PER state itself is bit-identical between runs; the")
print("downstream `checksums.replay_sample_indices` divergence")
print("observed by determinism-check.sh must be a SUM-OF-SQUARES")
print("permutation artifact (the checksum is permutation-invariant).")
print("Re-investigate: indices may differ in ORDER but match as a")
print("multiset, indicating per-batch slot-binding non-determinism")
print("upstream of PER.")
return 0
assert first_diverge_step is not None
prng_eq = first_diverge_state.get("prng", True)
tree_eq = first_diverge_state.get("tree", True)
idx_eq = first_diverge_state.get("indices", True)
print(f"VERDICT (first divergent step = {first_diverge_step}):")
if not prng_eq:
print(" HYPOTHESIS A — PRNG seeding non-determinism.")
print(" sample_prng_d differs between runs at step",
first_diverge_step,
"→ the device-side xorshift32 state was already divergent")
print(" before the tree walk. Apply Option A fix: explicit")
print(" cli.seed-derived seeding in trainer construction.")
return 1
if prng_eq and tree_eq and not idx_eq:
print(" HYPOTHESIS B — Tree-walk fp-comparison flip.")
print(" prng matches, tree matches, but sample_indices diverge")
print(" → the `u <= left_val` comparison in rl_per_sample.cu:88")
print(" is producing different paths despite identical inputs.")
print(" Apply Option B fix: integer priority sum-tree.")
return 1
if prng_eq and not tree_eq:
print(" HYPOTHESIS C — Tree-rebuild non-determinism.")
print(" prng matches, but priority_tree itself diverges")
print(" → rl_per_tree_rebuild.cu accumulation order is")
print(" non-deterministic across runs (parallel children-before-")
print(" parents not enforced strongly enough).")
print(" Apply Option C fix: serialize the rebuild bottom-up.")
return 1
# Should not be reachable (one of A/B/C must hold once any diverges).
print(" UNEXPECTED state — review per-buffer summary above.")
print(f" prng_eq={prng_eq} tree_eq={tree_eq} idx_eq={idx_eq}")
return 1
if __name__ == "__main__":
sys.exit(main())

284
scripts/compare-rl-state.py Executable file
View File

@@ -0,0 +1,284 @@
#!/usr/bin/env python3
"""Compare raw RL-loop state dumps between two determinism runs.
Reads the dumps produced by
`IntegratedTrainer::dump_rl_state_for_debug` (env-gated on
`FOXHUNT_DETERMINISM_DEBUG_RL=1`) and prints a per-step
exact-bytewise-equality verdict for each of the 6 candidate groups
identified by Phase 2.4's mamba2 falsification:
Group C (action sampling RNG):
- action_prng_state (B u32)
- iqn_prng_state (B u32)
Group D (LOB simulator + unit state):
- pyramid_units_count, unit_entry_step/_price/_lots/_active/...
- prev_realized_pnl, prev_position_lots, steps_since_done
- trade_duration_emit, close_unit_index, unit_prev_pos_lots,
unit_initial_r, unit_trail_distance
Group F (reward shaping + RL outputs):
- rewards, raw_rewards, outcome_ema, reward_abs, dones
- actions, next_actions, log_pi_old, advantages, returns
- CMDP/edge-decay per-batch: session_pnl, consec_loss,
cooldown_remaining, ph_mu/_count/_m/_mmin/_stat
Group E (full ISV array — host mapped-pinned, RL_SLOTS_END floats):
- isv_full (757 f32)
Group A (CfC carry `attn_context_d` / `h_t_d`) and Group B (PER state
priority_tree / sample_prng / sample_indices) are dumped by the
mamba2 + per debug dumpers respectively — Phase 2.5 reuses those.
Verdict logic:
- All groups EQUAL at all dumped steps → bug is elsewhere (Phase 2.6).
- Group C diverges → action sampling RNG path is non-deterministic.
- Group D diverges → LOB sim is the culprit.
- Group E diverges → an ISV controller is non-deterministic.
- Group F diverges → reward shaping / advantage / log_pi pipeline
is the culprit.
- Multiple groups diverge → ranked by FIRST step they diverge at;
the earliest one is upstream.
Exit codes:
0 both runs identical for all dumped buffers at all dumped steps
1 divergence found; verdict printed naming the candidate group
2 inputs missing or malformed
Usage:
python3 scripts/compare-rl-state.py <run_a_dir> <run_b_dir>
Each dir contains `step_{0,1,2,3}_<name>.bin`.
Per `feedback_cpu_is_read_only`: this script ONLY reads device dumps;
it performs no compute that affects training.
"""
from __future__ import annotations
import argparse
import struct
import sys
from pathlib import Path
# Each entry: (logical name, on-disk basename, element struct format,
# element size in bytes). The groups partition the candidate space
# so the verdict can name the culprit class precisely.
GROUP_C_RNG = [
("action_prng_state", "action_prng_state", "I", 4),
("iqn_prng_state", "iqn_prng_state", "I", 4),
]
GROUP_D_LOBSIM = [
("pyramid_units_count", "pyramid_units_count", "i", 4),
("unit_entry_step", "unit_entry_step", "i", 4),
("unit_entry_price", "unit_entry_price", "f", 4),
("unit_lots", "unit_lots", "i", 4),
("unit_active", "unit_active", "B", 1),
("unit_initial_r", "unit_initial_r", "f", 4),
("unit_trail_distance", "unit_trail_distance", "f", 4),
("unit_prev_pos_lots", "unit_prev_pos_lots", "i", 4),
("close_unit_index", "close_unit_index", "i", 4),
("prev_realized_pnl", "prev_realized_pnl", "f", 4),
("prev_position_lots", "prev_position_lots", "i", 4),
("steps_since_done", "steps_since_done", "i", 4),
("trade_duration_emit", "trade_duration_emit", "f", 4),
]
GROUP_F_REWARDS = [
("rewards", "rewards", "f", 4),
("raw_rewards", "raw_rewards", "f", 4),
("outcome_ema", "outcome_ema", "f", 4),
("reward_abs", "reward_abs", "f", 4),
("dones", "dones", "f", 4),
("actions", "actions", "i", 4),
("next_actions", "next_actions", "i", 4),
("log_pi_old", "log_pi_old", "f", 4),
("advantages", "advantages", "f", 4),
("returns", "returns", "f", 4),
# CMDP + edge-decay
("session_pnl", "session_pnl", "f", 4),
("consec_loss", "consec_loss", "f", 4),
("cooldown_remaining", "cooldown_remaining", "f", 4),
("ph_mu", "ph_mu", "f", 4),
("ph_count", "ph_count", "f", 4),
("ph_m", "ph_m", "f", 4),
("ph_mmin", "ph_mmin", "f", 4),
("ph_stat", "ph_stat", "f", 4),
]
GROUP_E_ISV = [
("isv_full", "isv_full", "f", 4),
]
ALL_GROUPS = [
("C (action RNG)", GROUP_C_RNG),
("D (LOB sim)", GROUP_D_LOBSIM),
("E (ISV controllers)", GROUP_E_ISV),
("F (reward shaping)", GROUP_F_REWARDS),
]
def read_typed(path: Path, fmt: str, elem_size: int) -> list:
"""Read a raw little-endian binary file as `fmt`-typed elements."""
data = path.read_bytes()
if len(data) % elem_size != 0:
raise ValueError(
f"{path}: size {len(data)} not a multiple of {elem_size}"
)
n = len(data) // elem_size
return list(struct.unpack(f"<{n}{fmt}", data))
def compare_lists(a: list, b: list, label: str) -> tuple[bool, int, str]:
"""Bytewise exact equality. Returns (equal, first_diff_idx, summary)."""
if len(a) != len(b):
return False, -1, f"{label}: length mismatch A={len(a)} B={len(b)}"
diffs: list[int] = []
max_abs_delta = 0.0
for i, (x, y) in enumerate(zip(a, b)):
if x != y:
diffs.append(i)
try:
ad = abs(float(x) - float(y))
if ad > max_abs_delta:
max_abs_delta = ad
except (TypeError, ValueError):
pass
if len(diffs) >= 5:
break
if not diffs:
return True, -1, f"{label}: EQUAL ({len(a)} elements)"
sample = ", ".join(
f"[{i}]: A={a[i]!r} B={b[i]!r}" for i in diffs[:3]
)
return False, diffs[0], (
f"{label}: DIVERGE at idx {diffs[0]} "
f"max|Δ|={max_abs_delta:.6g} "
f"(showing up to 3): {sample}"
)
def main() -> int:
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("run_a", type=Path)
parser.add_argument("run_b", type=Path)
args = parser.parse_args()
if not args.run_a.is_dir() or not args.run_b.is_dir():
print(
f"ERROR: run dirs missing: {args.run_a} or {args.run_b}",
file=sys.stderr,
)
return 2
# Per-step verdict. Tracks the FIRST step that diverges and which
# group diverged there.
overall_any_diverge = False
# Map (step, group_label) -> list of (name, ok, summary).
per_step_group_results: dict[tuple[int, str], list] = {}
first_diverge_step: int | None = None
first_diverge_groups: set[str] = set()
for step in (0, 1, 2, 3):
print(f"=== step {step} ===")
any_present = False
step_had_diverge = False
step_diverged_groups: set[str] = set()
for group_label, members in ALL_GROUPS:
print(f" -- group {group_label} --")
group_results: list[tuple[str, bool, str]] = []
group_any_present = False
for name, basename, fmt, elem_size in members:
pa = args.run_a / f"step_{step}_{basename}.bin"
pb = args.run_b / f"step_{step}_{basename}.bin"
if not (pa.is_file() and pb.is_file()):
continue
group_any_present = True
any_present = True
try:
a = read_typed(pa, fmt, elem_size)
b = read_typed(pb, fmt, elem_size)
except ValueError as e:
print(f" ERROR reading {name}: {e}")
return 2
equal, _, summary = compare_lists(a, b, name)
print(f" {summary}")
group_results.append((name, equal, summary))
if not equal:
overall_any_diverge = True
step_had_diverge = True
step_diverged_groups.add(group_label)
if not group_any_present:
print(" SKIP — no buffers dumped for this group/step")
per_step_group_results[(step, group_label)] = group_results
if not any_present:
print(" SKIP — no buffers dumped for this step")
continue
if step_had_diverge and first_diverge_step is None:
first_diverge_step = step
first_diverge_groups = step_diverged_groups
print()
if not overall_any_diverge:
print(
"VERDICT: all dumped RL-loop buffers match across all dumped\n"
"steps. The bug is NOT in any of the 6 candidate groups (CfC\n"
"carry / PER state / action RNG / LOB sim / ISV / reward\n"
"shaping). Phase 2.6 dispatch needed — STOP per\n"
"feedback_systematic_debugging."
)
return 0
assert first_diverge_step is not None
print(f"VERDICT (first divergent step = {first_diverge_step}):")
print(f" diverged group(s): {sorted(first_diverge_groups)}")
print()
print("Per-group culprit table:")
for group_label, _ in ALL_GROUPS:
results = per_step_group_results.get((first_diverge_step, group_label), [])
any_div = any(not eq for _, eq, _ in results)
if any_div:
divnames = [name for name, eq, _ in results if not eq]
print(f" Group {group_label}: DIVERGE in {divnames}")
else:
present = len(results) > 0
tag = "EQUAL" if present else "(no dump)"
print(f" Group {group_label}: {tag}")
print()
# Print interpretive hint based on which group(s) diverged FIRST.
print("Next-step interpretation:")
if "C (action RNG)" in first_diverge_groups:
print(" Group C (action RNG) — likely culprit if the only group.")
print(" Fix: xorshift32 update in rl_action_kernel must read")
print(" same upstream Q/π logits across runs. If upstream Q is")
print(" determined to be EQUAL but action_prng_state diverges,")
print(" the PRNG was perturbed by branchy logic (e.g.,")
print(" conditional advance based on action choice).")
if "D (LOB sim)" in first_diverge_groups:
print(" Group D (LOB sim) — LOB sim is producing different state.")
print(" Inspect kernels in crates/ml-backtesting/src/sim/ for")
print(" atomics, parallel reductions, order-dependent writes.")
if "E (ISV controllers)" in first_diverge_groups:
print(" Group E (ISV controllers) — one controller is non-det.")
print(" Compare ISV slot ranges: popart_sigma_welford (725),")
print(" kelly avg_win/loss_ema, wr_ema, ph_mean / edge-decay.")
print(" Look for parallel reductions in ema_update_*,")
print(" rl_kelly_*, rl_popart_*. Apply Phase-2-PER-rebuild fix")
print(" pattern (single-block + __syncthreads + fixed order).")
if "F (reward shaping)" in first_diverge_groups:
print(" Group F (reward shaping) — rewards / advantages / log_pi")
print(" diverge. Check rl_fused_reward_pipeline.cu,")
print(" compute_advantage_return, apply_reward_scale,")
print(" log_pi_at_action for non-det parallel writes.")
print(" If `rewards` diverges with EQUAL action_prng + LOB state,")
print(" the reward kernel itself is non-deterministic.")
print(" If `actions` diverges → upstream (action sampling) bug.")
print(" If `advantages` diverges with EQUAL `rewards` →")
print(" `compute_advantage_return` reduction is non-det.")
return 1
if __name__ == "__main__":
sys.exit(main())

397
scripts/determinism-check.sh Executable file
View File

@@ -0,0 +1,397 @@
#!/usr/bin/env bash
# Determinism self-test for Tier 1.5 mid-smoke (spec §3.5 + §1.3 of the
# determinism foundation).
#
# Runs scripts/local-mid-smoke.sh twice with the same seed and compares
# the per-step `checksums.*` leaves emitted by `build_diag_value` (added
# in determinism Phase 1, 2026-06-02). The trainer must be deterministic
# given a fixed seed for Tier 1.5 to be trustworthy — otherwise behavioral
# signals are polluted by random-init drift.
#
# Output focuses on the FIRST divergent (step, leaf) pair. That's the
# culprit identification — Phase 1's whole point. The trailing block
# also lists ALL leaves diverging at the first divergent step (so we
# catch the common "multiple components co-diverge" case).
#
# Tolerance: 1e-5 relative, 1e-7 absolute (matches fp32 noise floor).
# Larger gaps indicate a real divergence.
#
# Per `feedback_local_smoke_before_cluster` +
# `pearl_scoped_init_seed_for_reproducibility`.
#
# Usage:
# ./scripts/determinism-check.sh # default seed=42, b=128
# ./scripts/determinism-check.sh --seed 99 # custom seed
# ./scripts/determinism-check.sh --n-backtests 64 # smaller batch
# ./scripts/determinism-check.sh --quick # 200-step abbreviated run
#
# Exit codes:
# 0 Deterministic (all per-step checksums match within tolerance)
# 1 Drift detected (see first-divergence report)
# 2 Inputs missing (a run failed)
set -euo pipefail
SEED=42
N_BACKTESTS=128
N_STEPS=2000
N_EVAL_STEPS=500
QUICK=false
DEBUG_DUMP=false
DEBUG_DUMP_MAMBA2=false
DEBUG_DUMP_RL=false
DEBUG_DUMP_BACKWARD=false
usage() {
cat <<EOF
Usage: $(basename "$0") [OPTIONS]
Determinism self-test: run mid-smoke twice with same seed, diff all
\`checksums.*\` leaves leaf-by-leaf across every step. Reports the first
(step, leaf) pair where determinism breaks — that names the culprit.
Options:
--seed N Seed (default: $SEED)
--n-backtests N Batch size (default: $N_BACKTESTS)
--quick 200 steps + 50 eval (faster, less coverage)
--debug-dump Phase 2 sub-investigation: enable raw PER-state
dumps for step 0/1/2 of each run, then invoke
scripts/compare-per-state.py to diff them.
Use this once to localise the PER non-determinism
source (PRNG vs tree-walk vs tree-rebuild).
--debug-dump-mamba2 Phase 2.2 sub-investigation: enable raw mamba2
weight + intermediate-activation dumps for steps
0/1/2/3 of each run, then invoke
scripts/compare-mamba2-state.py to diff them.
Use this to disambiguate weight-drift (upstream
Adam non-det) from forward/backward kernel
non-determinism in the mamba2 selective scan.
--debug-dump-rl Phase 2.5 sub-investigation: enable raw RL-loop
state dumps (CfC carry, PER state, action RNG,
LOB sim, ISV controllers, reward shaping) for
steps 0/1/2/3 of each run, then invoke
scripts/compare-rl-state.py to diff them.
Use this after Phase 2.4 falsified the mamba2
hypothesis (all 52 dumps EQUAL at step 1 but
divergence enters at step 2 from inside the RL
step loop).
--debug-dump-backward Phase 2.6 sub-investigation: enable raw backward-
pipeline dumps (PER replay sample h_t / head fwd
outputs / pre-reduce per-batch param-grad
scratches / per-head grad_h_t + combined
accumulator) for steps 0/1/2/3 of each run, then
invoke scripts/compare-backward-state.py to diff
them. Use this after Phase 2.5 confirmed forward
activations EQUAL at step 2 but backward outputs
DIVERGE — the residual non-determinism lives in
step 2's gradient-emission chain (head bwd,
reduce_axis0, or grad_h_accumulate).
-h, --help Show this help
EOF
}
while [[ $# -gt 0 ]]; do
case "$1" in
--seed) SEED="$2"; shift 2 ;;
--n-backtests) N_BACKTESTS="$2"; shift 2 ;;
--quick) QUICK=true; shift ;;
--debug-dump) DEBUG_DUMP=true; shift ;;
--debug-dump-mamba2) DEBUG_DUMP_MAMBA2=true; shift ;;
--debug-dump-rl) DEBUG_DUMP_RL=true; shift ;;
--debug-dump-backward) DEBUG_DUMP_BACKWARD=true; shift ;;
-h|--help) usage; exit 0 ;;
*) echo "Unknown option: $1" >&2; usage >&2; exit 2 ;;
esac
done
if "$QUICK"; then
N_STEPS=200
N_EVAL_STEPS=50
fi
REPO_ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
RUN_A="/tmp/foxhunt-determinism-a"
RUN_B="/tmp/foxhunt-determinism-b"
DEBUG_DUMP_ROOT="/tmp/foxhunt-determinism-debug"
echo "=== Determinism self-test ==="
echo " Seed: $SEED"
echo " Steps: $N_STEPS train + $N_EVAL_STEPS eval"
echo " Batch: $N_BACKTESTS"
echo " Run A: $RUN_A"
echo " Run B: $RUN_B"
if "$DEBUG_DUMP" || "$DEBUG_DUMP_MAMBA2" || "$DEBUG_DUMP_RL" || "$DEBUG_DUMP_BACKWARD"; then
echo " Debug dump: $DEBUG_DUMP_ROOT/{run_a,run_b}"
fi
echo
rm -rf "$RUN_A" "$RUN_B"
if "$DEBUG_DUMP" || "$DEBUG_DUMP_MAMBA2" || "$DEBUG_DUMP_RL" || "$DEBUG_DUMP_BACKWARD"; then
rm -rf "$DEBUG_DUMP_ROOT"
mkdir -p "$DEBUG_DUMP_ROOT/run_a" "$DEBUG_DUMP_ROOT/run_b"
fi
run_one() {
local label="$1"
local outdir="$2"
local dump_subdir="$3" # only used when DEBUG_DUMP* flag is set
echo "[run $label]"
local -a env_vars=()
if "$DEBUG_DUMP"; then
env_vars+=("FOXHUNT_DETERMINISM_DEBUG=1")
fi
if "$DEBUG_DUMP_MAMBA2"; then
env_vars+=("FOXHUNT_DETERMINISM_DEBUG_MAMBA2=1")
fi
if "$DEBUG_DUMP_RL"; then
env_vars+=("FOXHUNT_DETERMINISM_DEBUG_RL=1")
fi
if "$DEBUG_DUMP_BACKWARD"; then
env_vars+=("FOXHUNT_DETERMINISM_DEBUG_BACKWARD=1")
fi
if "$DEBUG_DUMP" || "$DEBUG_DUMP_MAMBA2" || "$DEBUG_DUMP_RL" || "$DEBUG_DUMP_BACKWARD"; then
env_vars+=("FOXHUNT_DEBUG_DUMP_DIR=$DEBUG_DUMP_ROOT/$dump_subdir")
env "${env_vars[@]}" \
"$REPO_ROOT/scripts/local-mid-smoke.sh" \
--seed "$SEED" \
--out "$outdir" \
--n-backtests "$N_BACKTESTS" \
--n-steps "$N_STEPS" \
--n-eval-steps "$N_EVAL_STEPS" \
> "${outdir}.log" 2>&1
else
"$REPO_ROOT/scripts/local-mid-smoke.sh" \
--seed "$SEED" \
--out "$outdir" \
--n-backtests "$N_BACKTESTS" \
--n-steps "$N_STEPS" \
--n-eval-steps "$N_EVAL_STEPS" \
> "${outdir}.log" 2>&1
fi
}
run_one A "$RUN_A" "run_a"
run_one B "$RUN_B" "run_b"
if "$DEBUG_DUMP"; then
echo
echo "[debug-dump] comparing raw PER state via scripts/compare-per-state.py"
python3 "$REPO_ROOT/scripts/compare-per-state.py" \
"$DEBUG_DUMP_ROOT/run_a" "$DEBUG_DUMP_ROOT/run_b" \
--b-size "$N_BACKTESTS" || true
echo
fi
if "$DEBUG_DUMP_MAMBA2"; then
echo
echo "[debug-dump-mamba2] comparing raw mamba2 state via scripts/compare-mamba2-state.py"
python3 "$REPO_ROOT/scripts/compare-mamba2-state.py" \
"$DEBUG_DUMP_ROOT/run_a" "$DEBUG_DUMP_ROOT/run_b" || true
echo
fi
if "$DEBUG_DUMP_RL"; then
echo
echo "[debug-dump-rl] comparing raw RL-loop state via scripts/compare-rl-state.py"
python3 "$REPO_ROOT/scripts/compare-rl-state.py" \
"$DEBUG_DUMP_ROOT/run_a" "$DEBUG_DUMP_ROOT/run_b" || true
echo
fi
if "$DEBUG_DUMP_BACKWARD"; then
echo
echo "[debug-dump-backward] comparing raw backward-pipeline state via scripts/compare-backward-state.py"
python3 "$REPO_ROOT/scripts/compare-backward-state.py" \
"$DEBUG_DUMP_ROOT/run_a" "$DEBUG_DUMP_ROOT/run_b" || true
echo
fi
if [[ ! -f "$RUN_A/diag.jsonl" ]] || [[ ! -f "$RUN_B/diag.jsonl" ]]; then
echo "ERROR: missing diag.jsonl from one of the runs"
ls -la "$RUN_A" "$RUN_B" 2>/dev/null || true
exit 2
fi
echo
echo "[diff] all rows — checksums.* leaves only"
python3 - "$RUN_A/diag.jsonl" "$RUN_B/diag.jsonl" <<'PY'
"""Compare per-step `checksums.*` leaves of two diag.jsonl files.
The 15 checksum leaves are emitted by `build_diag_value` after the
determinism Phase 1 changes (2026-06-02). Each leaf is a deterministic
double-precision sum-of-squares of one trainer-state tensor. Any
divergence between same-seed runs localises the non-determinism source
to that component.
The script reports the FIRST (step, leaf) divergence + ALL leaves
diverging at that same step (since multiple components often co-diverge
once any single one has drifted).
Tolerance: 1e-5 relative, 1e-7 absolute — fp32 noise floor doubled
through the f64 accumulator.
"""
import json
import math
import sys
from pathlib import Path
REL_TOL = 1e-5
ABS_TOL = 1e-7
def load_all(path: Path) -> list[dict]:
rows: list[dict] = []
with path.open("r") as f:
for line in f:
line = line.strip()
if not line:
continue
try:
rows.append(json.loads(line))
except json.JSONDecodeError:
pass
return rows
def extract_checksums(row: dict) -> dict[str, float]:
"""Return the row's `checksums` subtree as a flat dict, or {}."""
cs = row.get("checksums")
if not isinstance(cs, dict):
return {}
out: dict[str, float] = {}
for k, v in cs.items():
if isinstance(v, (int, float)):
out[k] = float(v)
return out
def differs(a: float, b: float) -> tuple[bool, float]:
if math.isnan(a) and math.isnan(b):
return (False, 0.0)
ad = abs(a - b)
tol = max(ABS_TOL, REL_TOL * max(abs(a), abs(b)))
return (ad > tol, ad)
path_a, path_b = Path(sys.argv[1]), Path(sys.argv[2])
rows_a = load_all(path_a)
rows_b = load_all(path_b)
if len(rows_a) != len(rows_b):
print(
f"ROW COUNT MISMATCH: A={len(rows_a)} B={len(rows_b)} — "
f"runs disagree on training horizon; cannot localise per-step.",
flush=True,
)
sys.exit(1)
if not rows_a:
print("ERROR: no rows in either run")
sys.exit(2)
# Sanity: the first row must have the `checksums` block (else we built
# against a pre-Phase-1 binary).
first_cs = extract_checksums(rows_a[0])
if not first_cs:
print(
"ERROR: `checksums` block missing from diag.jsonl rows.\n"
"Did you build the alpha_rl_train binary BEFORE running\n"
"the determinism-check? Phase 1 added the checksum kernel +\n"
"diag emission — rebuild with `cargo build --release\n"
"--example alpha_rl_train -p ml-alpha`.",
flush=True,
)
sys.exit(2)
# Walk every row in order; report the first (step, leaf) divergence.
first_step: int | None = None
first_leaves: list[tuple[str, float, float, float]] = []
total_divergent_rows = 0
for i, (ra, rb) in enumerate(zip(rows_a, rows_b)):
cs_a = extract_checksums(ra)
cs_b = extract_checksums(rb)
if not cs_a or not cs_b:
# Eval-phase rows may emit slightly different schema; just skip
# rows missing the block to avoid false positives. Phase 1
# diag emission writes both train and eval blocks identically.
continue
row_divergent: list[tuple[str, float, float, float]] = []
# Stable iteration: sort leaf names so output is reproducible.
for leaf in sorted(set(cs_a) | set(cs_b)):
va = cs_a.get(leaf, float("nan"))
vb = cs_b.get(leaf, float("nan"))
d, ad = differs(va, vb)
if d:
row_divergent.append((leaf, va, vb, ad))
if row_divergent:
total_divergent_rows += 1
if first_step is None:
step = ra.get("step", i)
first_step = int(step) if isinstance(step, (int, float)) else i
first_leaves = row_divergent
if first_step is None:
print(
"DETERMINISTIC: all `checksums.*` leaves match across all "
f"{len(rows_a)} rows (rel-tol={REL_TOL}, abs-tol={ABS_TOL})."
)
sys.exit(0)
# Report the FIRST divergence per spec §1.3.
first_leaves.sort(key=lambda t: t[3], reverse=True)
top = first_leaves[0]
leaf, va, vb, ad = top
print(f"FIRST_DIVERGENCE: step={first_step} leaf=checksums.{leaf} "
f"run_A={va!r} run_B={vb!r} delta={ad:.6g}")
print()
print(f"All leaves diverging at step={first_step}:")
for leaf, va, vb, ad in first_leaves:
print(f" checksums.{leaf:<28} A={va!r:>24} B={vb!r:>24} Δ={ad:.6g}")
print()
print(f"Total divergent rows: {total_divergent_rows} / {len(rows_a)} "
"(every row after first divergence cascades).")
print()
print("Phase 2 fix recommendation lookup (see spec §2 of"
" 2026-06-02-determinism-foundation.md):")
candidates = {
"encoder_output": "§2.A mamba2 selective-scan / encoder kernels",
"encoder_grad": "§2.A mamba2 backward / §2.D custom-kernel reductions",
"q_logits": "§2.B cuBLAS GEMM split-K (Q head)",
"pi_logits": "§2.B cuBLAS GEMM split-K (π head)",
"v_value": "§2.B cuBLAS GEMM split-K (V head)",
"q_grad": "§2.D Q-head backward reductions",
"pi_grad": "§2.D π-head backward reductions",
"v_grad": "§2.D V-head backward reductions",
"replay_sample_indices": "§2.F PER (Prioritized Experience Replay) sampling",
"adam_m_sum": "§2.E Adam optimizer iteration order",
"adam_v_sum": "§2.E Adam optimizer iteration order",
"rewards_after_shape": "§2.D reward shaping / scale / clamp kernels",
"advantages": "§2.D advantage normalize / compute_advantage_return",
"isv_state": "§2.D ISV controllers / non-deterministic writes",
"popart_sigma_state": "§2.D PopArt Welford streaming reduction",
# Phase 2.2 encoder-localisation leaves
"mamba2_l1_w_in": "§2.A mamba2 L1 W_in weight drift (upstream of forward)",
"mamba2_l1_w_a": "§2.A mamba2 L1 W_a weight drift",
"mamba2_l1_w_b": "§2.A mamba2 L1 W_b weight drift",
"mamba2_l1_w_c": "§2.A mamba2 L1 W_c weight drift",
"mamba2_l2_w_in": "§2.A mamba2 L2 W_in weight drift",
"mamba2_l2_w_c": "§2.A mamba2 L2 W_c weight drift",
"vsn_w": "§2.D VSN weight drift (upstream of encoder)",
"cfc_w_in": "§2.D CfC w_in weight drift",
"cfc_w_rec": "§2.D CfC w_rec weight drift",
"attn_q": "§2.D attention pool query weight drift",
"vsn_out": "§2.D VSN forward kernel non-determinism",
"mamba2_l1_out": "§2.A mamba2 L1 forward scan non-determinism",
"ln_a_out": "§2.D LayerNorm A forward non-determinism",
"mamba2_l2_out": "§2.A mamba2 L2 forward scan non-determinism",
"ln_b_out": "§2.D LayerNorm B forward non-determinism",
"attn_context": "§2.D attention pool forward non-determinism",
"mamba2_grad_w_c_l1": "§2.A mamba2_alpha_reduce_d_w_c kernel non-determinism",
}
for leaf, _, _, _ in first_leaves:
rec = candidates.get(leaf)
if rec is not None:
print(f" checksums.{leaf} → {rec}")
break # only the most-likely culprit (top-ranked by Δ)
sys.exit(1)
PY

148
scripts/local-mid-smoke.sh Executable file
View File

@@ -0,0 +1,148 @@
#!/usr/bin/env bash
# Tier 1.5 local mid-smoke — RTX 3050 Ti, b=128, 2000 train + 500 eval steps.
#
# Closes the 700× dev-cycle gap between Tier 1 (6 sec local correctness) and
# Tier 2 (70 min cluster eval). Gates 80%+ of bad hypotheses before they reach
# cluster. Target wall-clock: ≤ 10 min on RTX 3050.
#
# Spec: docs/superpowers/specs/2026-06-02-fast-dev-cycle.md §3.2
# Verdict: python3 scripts/tier1_5_verdict.py <out-dir>
#
# Data: test_data/futures-baseline-mid/ES.FUT/ — symlinks to 2024-Q1 (train,
# 4.8GB) + 2024-Q2 (eval, 14GB) MBP-10 files from the canonical PVC.
# Walk-forward split: --n-folds 2 --fold-idx 0 → train=Q1, eval=Q2.
#
# Usage:
# ./scripts/local-mid-smoke.sh # default seed=42
# ./scripts/local-mid-smoke.sh --seed 123 # override seed
# ./scripts/local-mid-smoke.sh --out /tmp/foxhunt-alt # custom out dir
# ./scripts/local-mid-smoke.sh --n-backtests 64 # smaller batch
#
# Per `feedback_local_smoke_before_cluster`: run this BEFORE every cluster
# smoke when touching .cu / trait impls / buffers / N_ACTIONS / ISV slots.
set -euo pipefail
SEED=42
OUT_DIR="/tmp/foxhunt-mid-smoke"
N_BACKTESTS=128
N_STEPS=2000
N_EVAL_STEPS=500
N_FOLDS=2
FOLD_IDX=0
usage() {
cat <<EOF
Usage: $(basename "$0") [OPTIONS]
Tier 1.5 mid-smoke — b=128, 2000+500 steps, ~10 min on RTX 3050.
Options:
--seed N RNG seed (default: $SEED — reproducible)
--out DIR Output directory (default: $OUT_DIR)
--n-backtests N Batch size (default: $N_BACKTESTS; drop to 64 if OOM)
--n-steps N Train steps (default: $N_STEPS)
--n-eval-steps N Eval steps (default: $N_EVAL_STEPS)
--n-folds K Walk-forward fold count (default: $N_FOLDS)
--fold-idx K Walk-forward fold index (default: $FOLD_IDX)
-h, --help Show this help
Output:
\$OUT_DIR/diag.jsonl — per-step train diag (642 scalars/step)
\$OUT_DIR/eval_diag.jsonl — per-step eval diag
\$OUT_DIR/alpha_rl_train_summary.json — final summary
EOF
}
while [[ $# -gt 0 ]]; do
case "$1" in
--seed) SEED="$2"; shift 2 ;;
--out) OUT_DIR="$2"; shift 2 ;;
--n-backtests) N_BACKTESTS="$2"; shift 2 ;;
--n-steps) N_STEPS="$2"; shift 2 ;;
--n-eval-steps) N_EVAL_STEPS="$2"; shift 2 ;;
--n-folds) N_FOLDS="$2"; shift 2 ;;
--fold-idx) FOLD_IDX="$2"; shift 2 ;;
-h|--help) usage; exit 0 ;;
*) echo "Unknown option: $1" >&2; usage >&2; exit 2 ;;
esac
done
REPO_ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
cd "$REPO_ROOT"
# --mbp10-data-dir points DIRECTLY at the instrument's .dbn.zst directory
# (matches argo-rl template line: --mbp10-data-dir /data/futures-baseline-mbp10/ES.FUT).
MBP10_DIR="$REPO_ROOT/test_data/futures-baseline-mid/ES.FUT"
PREDECODED_DIR="$REPO_ROOT/test_data/futures-baseline-mid"
# Sanity checks.
if [[ ! -d "$MBP10_DIR" ]]; then
echo "ERROR: $MBP10_DIR not found." >&2
echo "Run setup: see CLAUDE.md 'Tier 1.5 data path' or docs/superpowers/specs/2026-06-02-fast-dev-cycle.md §3.1" >&2
exit 3
fi
n_mbp_files=$(find -L "$MBP10_DIR" -maxdepth 1 -name '*.dbn.zst' 2>/dev/null | wc -l)
if [[ "$n_mbp_files" -lt 2 ]]; then
echo "ERROR: $MBP10_DIR has $n_mbp_files MBP-10 files; need ≥ 2 for walk-forward split." >&2
exit 3
fi
mkdir -p "$OUT_DIR"
echo "=== Tier 1.5 mid-smoke ==="
echo " Repo: $REPO_ROOT"
echo " Seed: $SEED"
echo " Batch: $N_BACKTESTS"
echo " Steps: $N_STEPS train + $N_EVAL_STEPS eval"
echo " Walk-forward: fold $FOLD_IDX of $N_FOLDS"
echo " MBP-10 dir: $MBP10_DIR (${n_mbp_files} files)"
echo " Out dir: $OUT_DIR"
echo
# Build (sccache + dev-release for fast iteration; --release matches cluster).
echo "[build] cargo build --release --example alpha_rl_train -p ml-alpha"
SQLX_OFFLINE=true cargo build --release --example alpha_rl_train -p ml-alpha
BINARY="$REPO_ROOT/target/release/examples/alpha_rl_train"
if [[ ! -x "$BINARY" ]]; then
echo "ERROR: binary not found at $BINARY after build." >&2
exit 4
fi
echo
echo "[run] $BINARY ..."
START_TS=$(date +%s)
# Per `feedback_no_stubs`: the binary owns its full life cycle. We just
# invoke it and capture wall time. Exit code 2 = G8 NaN abort (per binary
# docstring); we propagate it so the caller knows the run died.
"$BINARY" \
--mbp10-data-dir "$MBP10_DIR" \
--predecoded-dir "$PREDECODED_DIR" \
--out "$OUT_DIR" \
--diag-jsonl "$OUT_DIR/diag.jsonl" \
--eval-diag-jsonl "$OUT_DIR/eval_diag.jsonl" \
--n-steps "$N_STEPS" \
--n-eval-steps "$N_EVAL_STEPS" \
--n-backtests "$N_BACKTESTS" \
--n-folds "$N_FOLDS" \
--fold-idx "$FOLD_IDX" \
--seed "$SEED" \
--instrument-mode front-month \
--log-every 200 \
--gpu-idx 0
END_TS=$(date +%s)
ELAPSED=$((END_TS - START_TS))
ELAPSED_MIN=$((ELAPSED / 60))
ELAPSED_SEC=$((ELAPSED % 60))
echo
echo "=== Mid-smoke complete in ${ELAPSED_MIN}m ${ELAPSED_SEC}s ==="
echo "Outputs:"
ls -lh "$OUT_DIR" 2>/dev/null
echo
echo "Next:"
echo " python3 $REPO_ROOT/scripts/tier1_5_verdict.py $OUT_DIR"

536
scripts/tier1_5_verdict.py Executable file
View File

@@ -0,0 +1,536 @@
#!/usr/bin/env python3
"""Tier 1.5 mid-smoke behavioral verdict.
Reads `diag.jsonl` + `eval_diag.jsonl` + `eval_summary.json` from the
mid-smoke output directory and computes 7 behavioral signals + 1
per-step/eval_summary consistency check. Emits OK / KILL_<reason> /
WARN_<reason>.
Spec: docs/superpowers/specs/2026-06-02-fast-dev-cycle.md §3.3, §3.6
Linked: pearl_grwwh_eval_catastrophic_collapse (per-step vs eval_summary gap)
Per `feedback_kill_runs_on_anomaly_quickly`: any KILL_* response means
DO NOT submit cluster smoke. Fix locally first.
Usage:
python3 scripts/tier1_5_verdict.py /tmp/foxhunt-mid-smoke
python3 scripts/tier1_5_verdict.py /tmp/foxhunt-mid-smoke --quiet
python3 scripts/tier1_5_verdict.py /tmp/foxhunt-mid-smoke --json
Exit codes:
0 OK (all signals within tolerance)
1 KILL (one or more behavioral red flags)
2 Inputs missing / malformed (cannot verdict)
"""
from __future__ import annotations
import argparse
import json
import math
import statistics
import sys
from pathlib import Path
from typing import Any
# ── Verdict thresholds (per spec §3.3 table) ──────────────────────────
# Hard floors/ceilings are KILL bounds (catastrophic).
# Soft targets are reported but not gating (informational).
THRESHOLDS = {
# action_entropy final: log(11) = 2.397; target [0.5×log(11), 0.85×log(11)]
# KILL if < 0.6 (policy collapsed) or > 2.3 (essentially random).
"action_entropy_kill_low": 0.6,
"action_entropy_kill_high": 2.3,
"action_entropy_target_low": 1.20, # 0.5 × log(11)
"action_entropy_target_high": 2.04, # 0.85 × log(11)
# q_pi_agree_ema final: ≥ 0.6 target, KILL if < 0.3 (Q/π decoupled)
"q_pi_agree_kill": 0.3,
"q_pi_agree_target": 0.6,
# Pearson(rewards.sum, Δrealized_pnl) ≥ 0.5 target, KILL < 0.3
"pearson_kill": 0.3,
"pearson_target": 0.5,
# popart.sigma CV (stddev/mean): KILL > 0.5 (wild oscillation)
"popart_sigma_cv_kill": 0.5,
# win_rate (train) ≥ 0.25 target, KILL < 0.15
"wr_train_kill": 0.15,
"wr_train_target": 0.25,
# win_rate (eval) ≥ 0.20 target, KILL < 0.15
"wr_eval_kill": 0.15,
"wr_eval_target": 0.20,
# Eval pnl > -$1M target, KILL < -$3M (catastrophic)
"eval_pnl_kill": -3_000_000.0,
"eval_pnl_target": -1_000_000.0,
# Per-step vs eval_summary discrepancy: > 5% gap → WARN (not KILL)
"consistency_warn_frac": 0.05,
# Min steps required for a meaningful verdict.
"min_train_rows": 500,
"min_eval_rows": 100,
# Hold-step trend: linear regression slope over last N rows must be > 0.
# KILL if slope is significantly negative.
"hold_trend_window": 1500,
"hold_trend_slope_kill": -0.01, # tolerate small decline as noise
}
def parse_args() -> argparse.Namespace:
p = argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter)
p.add_argument("out_dir", type=Path, help="Mid-smoke output directory (--out from local-mid-smoke.sh)")
p.add_argument("--quiet", action="store_true", help="Suppress per-signal explanations")
p.add_argument("--json", action="store_true", help="Emit machine-readable JSON verdict")
return p.parse_args()
def load_jsonl(path: Path, *, max_rows: int | None = None) -> list[dict[str, Any]]:
"""Stream-parse a JSONL file; return list of decoded records."""
rows: list[dict[str, Any]] = []
with path.open("r") as f:
for line in f:
line = line.strip()
if not line:
continue
try:
rows.append(json.loads(line))
except json.JSONDecodeError as exc:
# Stop at first malformed line — typically means truncated run.
print(f"WARN: malformed JSON at line {len(rows) + 1} of {path}: {exc}", file=sys.stderr)
break
if max_rows is not None and len(rows) >= max_rows:
break
return rows
def get_nested(obj: Any, dotted: str, default: Any = None) -> Any:
"""Dotted-path lookup: get_nested({'a':{'b':1}}, 'a.b') → 1."""
cur = obj
for key in dotted.split("."):
if not isinstance(cur, dict) or key not in cur:
return default
cur = cur[key]
return cur
def pearson(xs: list[float], ys: list[float]) -> float | None:
"""Return Pearson correlation, or None if input is degenerate."""
n = len(xs)
if n != len(ys) or n < 3:
return None
mx = statistics.fmean(xs)
my = statistics.fmean(ys)
sx2 = sum((x - mx) ** 2 for x in xs)
sy2 = sum((y - my) ** 2 for y in ys)
if sx2 == 0.0 or sy2 == 0.0:
return None
sxy = sum((x - mx) * (y - my) for x, y in zip(xs, ys))
return sxy / math.sqrt(sx2 * sy2)
def linreg_slope(xs: list[float], ys: list[float]) -> float | None:
"""OLS slope of y over x; None if degenerate."""
n = len(xs)
if n != len(ys) or n < 3:
return None
mx = statistics.fmean(xs)
my = statistics.fmean(ys)
num = sum((x - mx) * (y - my) for x, y in zip(xs, ys))
den = sum((x - mx) ** 2 for x in xs)
if den == 0.0:
return None
return num / den
def coefficient_of_variation(xs: list[float]) -> float | None:
"""stdev / |mean|; None if mean ≈ 0 or fewer than 2 points."""
if len(xs) < 2:
return None
m = statistics.fmean(xs)
if abs(m) < 1e-12:
return None
sd = statistics.stdev(xs)
return sd / abs(m)
# ── Signal computations ───────────────────────────────────────────────
class SignalResult:
"""A single verdict line — name, value, status, explanation."""
__slots__ = ("name", "value", "status", "explain")
STATUS_OK = "OK"
STATUS_KILL = "KILL"
STATUS_WARN = "WARN"
STATUS_SKIP = "SKIP"
def __init__(self, name: str, value: Any, status: str, explain: str) -> None:
self.name = name
self.value = value
self.status = status
self.explain = explain
def to_dict(self) -> dict[str, Any]:
return {
"name": self.name,
"value": self.value,
"status": self.status,
"explain": self.explain,
}
def signal_action_entropy(train_rows: list[dict[str, Any]]) -> SignalResult:
"""Final action_entropy: KILL if collapsed (<0.6) or random (>2.3)."""
val = get_nested(train_rows[-1], "action_entropy")
if val is None or not isinstance(val, (int, float)):
return SignalResult("action_entropy", None, SignalResult.STATUS_SKIP, "field missing from final train row")
if val < THRESHOLDS["action_entropy_kill_low"]:
return SignalResult("action_entropy", val, SignalResult.STATUS_KILL,
f"{val:.3f} < {THRESHOLDS['action_entropy_kill_low']} — policy collapsed to δ-function")
if val > THRESHOLDS["action_entropy_kill_high"]:
return SignalResult("action_entropy", val, SignalResult.STATUS_KILL,
f"{val:.3f} > {THRESHOLDS['action_entropy_kill_high']} — policy essentially uniform/random")
if val < THRESHOLDS["action_entropy_target_low"]:
return SignalResult("action_entropy", val, SignalResult.STATUS_WARN,
f"{val:.3f} below target [{THRESHOLDS['action_entropy_target_low']:.2f}, {THRESHOLDS['action_entropy_target_high']:.2f}]")
if val > THRESHOLDS["action_entropy_target_high"]:
return SignalResult("action_entropy", val, SignalResult.STATUS_WARN,
f"{val:.3f} above target [{THRESHOLDS['action_entropy_target_low']:.2f}, {THRESHOLDS['action_entropy_target_high']:.2f}]")
return SignalResult("action_entropy", val, SignalResult.STATUS_OK,
f"{val:.3f} within target [{THRESHOLDS['action_entropy_target_low']:.2f}, {THRESHOLDS['action_entropy_target_high']:.2f}]")
def signal_q_pi_agree(train_rows: list[dict[str, Any]]) -> SignalResult:
"""Final q_pi_agree_ema: KILL if decoupled (<0.3)."""
val = get_nested(train_rows[-1], "q_pi_agree_ema")
if val is None or not isinstance(val, (int, float)):
return SignalResult("q_pi_agree_ema", None, SignalResult.STATUS_SKIP, "field missing from final train row")
if val < THRESHOLDS["q_pi_agree_kill"]:
return SignalResult("q_pi_agree_ema", val, SignalResult.STATUS_KILL,
f"{val:.3f} < {THRESHOLDS['q_pi_agree_kill']} — Q and π decoupled (π is dead weight)")
if val < THRESHOLDS["q_pi_agree_target"]:
return SignalResult("q_pi_agree_ema", val, SignalResult.STATUS_WARN,
f"{val:.3f} below target ≥ {THRESHOLDS['q_pi_agree_target']}")
return SignalResult("q_pi_agree_ema", val, SignalResult.STATUS_OK,
f"{val:.3f}{THRESHOLDS['q_pi_agree_target']}")
def signal_reward_pnl_pearson(train_rows: list[dict[str, Any]]) -> SignalResult:
"""Pearson(rewards.sum, Δrealized_pnl): KILL if anti-aligned (<0.3).
Per `pearl_reward_signal_anti_aligned_with_pnl` (2026-06-01): the
decisive bug behind 64 negative-eval commits — reward gradient
systematically anti-aligned with pnl direction.
"""
rewards: list[float] = []
dpnls: list[float] = []
prev_pnl: float | None = None
for row in train_rows:
rs = get_nested(row, "rewards.sum")
cur_pnl = get_nested(row, "trading.realized_pnl_cum_usd")
if rs is None or cur_pnl is None:
continue
if not isinstance(rs, (int, float)) or not isinstance(cur_pnl, (int, float)):
continue
if prev_pnl is not None:
rewards.append(float(rs))
dpnls.append(float(cur_pnl) - float(prev_pnl))
prev_pnl = float(cur_pnl)
p = pearson(rewards, dpnls)
if p is None:
return SignalResult("pearson(rewards.sum, Δpnl)", None, SignalResult.STATUS_SKIP,
f"degenerate input (n={len(rewards)})")
if p < THRESHOLDS["pearson_kill"]:
return SignalResult("pearson(rewards.sum, Δpnl)", p, SignalResult.STATUS_KILL,
f"{p:.3f} < {THRESHOLDS['pearson_kill']} — reward gradient anti-aligned with pnl (n={len(rewards)})")
if p < THRESHOLDS["pearson_target"]:
return SignalResult("pearson(rewards.sum, Δpnl)", p, SignalResult.STATUS_WARN,
f"{p:.3f} below target ≥ {THRESHOLDS['pearson_target']} (n={len(rewards)})")
return SignalResult("pearson(rewards.sum, Δpnl)", p, SignalResult.STATUS_OK,
f"{p:.3f}{THRESHOLDS['pearson_target']} (n={len(rewards)})")
def signal_hold_trend(train_rows: list[dict[str, Any]]) -> SignalResult:
"""avg_hold_steps trend over last N rows: KILL if significantly decreasing.
The surfer pattern (wave-timescale edge per `pearl_edge_lives_at_wave_timescale_not_tick`)
requires hold-steps growing through training.
"""
window = THRESHOLDS["hold_trend_window"]
tail = train_rows[-window:] if len(train_rows) > window else train_rows
xs: list[float] = []
ys: list[float] = []
for row in tail:
step = get_nested(row, "step")
hold = get_nested(row, "trading.avg_hold_steps")
if step is None or hold is None:
continue
if not isinstance(step, (int, float)) or not isinstance(hold, (int, float)):
continue
if math.isnan(float(hold)) or math.isinf(float(hold)):
continue
xs.append(float(step))
ys.append(float(hold))
if len(xs) < 50:
return SignalResult("avg_hold_steps trend", None, SignalResult.STATUS_SKIP,
f"only {len(xs)} usable points in trend window")
slope = linreg_slope(xs, ys)
if slope is None:
return SignalResult("avg_hold_steps trend", None, SignalResult.STATUS_SKIP, "degenerate regression")
if slope < THRESHOLDS["hold_trend_slope_kill"]:
return SignalResult("avg_hold_steps trend", slope, SignalResult.STATUS_KILL,
f"slope {slope:.5f}/step < {THRESHOLDS['hold_trend_slope_kill']} — hold-steps decreasing, surfer pattern absent")
if slope <= 0.0:
return SignalResult("avg_hold_steps trend", slope, SignalResult.STATUS_WARN,
f"slope {slope:.5f}/step — hold-steps not growing")
return SignalResult("avg_hold_steps trend", slope, SignalResult.STATUS_OK,
f"slope +{slope:.5f}/step over last {len(xs)} rows")
def signal_popart_sigma_cv(train_rows: list[dict[str, Any]]) -> SignalResult:
"""popart.sigma coefficient of variation: KILL if wild oscillation (CV > 0.5)."""
sigmas: list[float] = []
for row in train_rows:
s = get_nested(row, "popart.sigma")
if s is None or not isinstance(s, (int, float)):
continue
if math.isnan(float(s)) or math.isinf(float(s)):
continue
sigmas.append(float(s))
if len(sigmas) < 100:
return SignalResult("popart.sigma CV", None, SignalResult.STATUS_SKIP,
f"only {len(sigmas)} usable samples")
cv = coefficient_of_variation(sigmas)
if cv is None:
return SignalResult("popart.sigma CV", None, SignalResult.STATUS_SKIP,
"mean ≈ 0 — undefined CV")
if cv > THRESHOLDS["popart_sigma_cv_kill"]:
return SignalResult("popart.sigma CV", cv, SignalResult.STATUS_KILL,
f"{cv:.3f} > {THRESHOLDS['popart_sigma_cv_kill']} — controller wildly oscillating")
return SignalResult("popart.sigma CV", cv, SignalResult.STATUS_OK,
f"{cv:.3f}{THRESHOLDS['popart_sigma_cv_kill']}")
def signal_wr_train(train_rows: list[dict[str, Any]]) -> SignalResult:
"""Final win_rate_ema (train): KILL if random/worse (<0.15)."""
# Prefer kelly.win_rate_ema (EMA across all trades) over trading.win_rate
# which is a point measure of last batch.
val = get_nested(train_rows[-1], "isv_config.kelly.win_rate_ema")
if val is None:
val = get_nested(train_rows[-1], "trading.win_rate")
source = "trading.win_rate"
else:
source = "isv_config.kelly.win_rate_ema"
if val is None or not isinstance(val, (int, float)):
return SignalResult(f"wr_train ({source})", None, SignalResult.STATUS_SKIP,
"field missing from final train row")
if val < THRESHOLDS["wr_train_kill"]:
return SignalResult(f"wr_train ({source})", val, SignalResult.STATUS_KILL,
f"{val:.3f} < {THRESHOLDS['wr_train_kill']} — random or worse")
if val < THRESHOLDS["wr_train_target"]:
return SignalResult(f"wr_train ({source})", val, SignalResult.STATUS_WARN,
f"{val:.3f} below target ≥ {THRESHOLDS['wr_train_target']}")
return SignalResult(f"wr_train ({source})", val, SignalResult.STATUS_OK,
f"{val:.3f}{THRESHOLDS['wr_train_target']}")
def signal_wr_eval(eval_rows: list[dict[str, Any]]) -> SignalResult:
"""Final win_rate_ema (eval): KILL if < 0.15."""
if not eval_rows:
return SignalResult("wr_eval", None, SignalResult.STATUS_SKIP, "no eval rows")
val = get_nested(eval_rows[-1], "isv_config.kelly.win_rate_ema")
if val is None:
val = get_nested(eval_rows[-1], "trading.win_rate")
source = "trading.win_rate"
else:
source = "isv_config.kelly.win_rate_ema"
if val is None or not isinstance(val, (int, float)):
return SignalResult(f"wr_eval ({source})", None, SignalResult.STATUS_SKIP,
"field missing from final eval row")
if val < THRESHOLDS["wr_eval_kill"]:
return SignalResult(f"wr_eval ({source})", val, SignalResult.STATUS_KILL,
f"{val:.3f} < {THRESHOLDS['wr_eval_kill']} — random or worse")
if val < THRESHOLDS["wr_eval_target"]:
return SignalResult(f"wr_eval ({source})", val, SignalResult.STATUS_WARN,
f"{val:.3f} below target ≥ {THRESHOLDS['wr_eval_target']}")
return SignalResult(f"wr_eval ({source})", val, SignalResult.STATUS_OK,
f"{val:.3f}{THRESHOLDS['wr_eval_target']}")
def signal_eval_pnl(eval_summary: dict[str, Any] | None,
eval_rows: list[dict[str, Any]]) -> SignalResult:
"""Final eval pnl: KILL if catastrophic (< -$3M).
Prefer eval_summary.total_pnl_usd (authoritative) over per-step
realized_pnl_cum_usd (per `pearl_grwwh_eval_catastrophic_collapse`
these can disagree by hundreds of millions).
"""
val: float | None = None
source = "none"
if eval_summary is not None and "total_pnl_usd" in eval_summary:
val = float(eval_summary["total_pnl_usd"])
source = "eval_summary.total_pnl_usd"
elif eval_rows:
last = get_nested(eval_rows[-1], "trading.realized_pnl_cum_usd")
if isinstance(last, (int, float)):
val = float(last)
source = "eval_diag last realized_pnl_cum_usd"
if val is None:
return SignalResult("eval_pnl", None, SignalResult.STATUS_SKIP,
"neither eval_summary.json nor eval_diag.jsonl has pnl")
if val < THRESHOLDS["eval_pnl_kill"]:
return SignalResult(f"eval_pnl ({source})", val, SignalResult.STATUS_KILL,
f"${val:,.0f} < ${THRESHOLDS['eval_pnl_kill']:,.0f} — catastrophic")
if val < THRESHOLDS["eval_pnl_target"]:
return SignalResult(f"eval_pnl ({source})", val, SignalResult.STATUS_WARN,
f"${val:,.0f} below target > ${THRESHOLDS['eval_pnl_target']:,.0f}")
return SignalResult(f"eval_pnl ({source})", val, SignalResult.STATUS_OK,
f"${val:,.0f} > ${THRESHOLDS['eval_pnl_target']:,.0f}")
def signal_consistency(eval_rows: list[dict[str, Any]],
eval_summary: dict[str, Any] | None) -> SignalResult:
"""Per-step vs eval_summary consistency check (§3.6).
WARN (not KILL): the per-step diag's realized_pnl_cum_usd and
eval_summary.json's total_pnl_usd should match within 5%. A divergence
indicates the per-step diag is misleading (per `pearl_grwwh_eval_catastrophic_collapse`).
"""
if eval_summary is None:
return SignalResult("per-step vs eval_summary", None, SignalResult.STATUS_SKIP,
"eval_summary.json missing")
if "total_pnl_usd" not in eval_summary:
return SignalResult("per-step vs eval_summary", None, SignalResult.STATUS_SKIP,
"eval_summary.json has no total_pnl_usd")
if not eval_rows:
return SignalResult("per-step vs eval_summary", None, SignalResult.STATUS_SKIP,
"no eval rows")
per_step = get_nested(eval_rows[-1], "trading.realized_pnl_cum_usd")
if per_step is None or not isinstance(per_step, (int, float)):
return SignalResult("per-step vs eval_summary", None, SignalResult.STATUS_SKIP,
"last eval row missing realized_pnl_cum_usd")
summary = float(eval_summary["total_pnl_usd"])
per_step = float(per_step)
diff = abs(per_step - summary)
denom = max(abs(per_step), abs(summary), 1.0)
frac = diff / denom
detail = f"per_step=${per_step:,.0f} | eval_summary=${summary:,.0f} | gap=${diff:,.0f} ({frac*100:.2f}%)"
if frac > THRESHOLDS["consistency_warn_frac"]:
return SignalResult("per-step vs eval_summary", frac, SignalResult.STATUS_WARN,
f"{detail} > {THRESHOLDS['consistency_warn_frac']*100:.0f}% — accounting axis discrepancy")
return SignalResult("per-step vs eval_summary", frac, SignalResult.STATUS_OK, detail)
# ── Driver ────────────────────────────────────────────────────────────
def run_verdict(out_dir: Path) -> tuple[str, list[SignalResult], dict[str, Any]]:
"""Compute all signals; return overall verdict + per-signal list + meta."""
diag_path = out_dir / "diag.jsonl"
eval_diag_path = out_dir / "eval_diag.jsonl"
eval_summary_path = out_dir / "eval_summary.json"
meta: dict[str, Any] = {
"out_dir": str(out_dir),
"diag_path": str(diag_path),
"eval_diag_path": str(eval_diag_path),
"eval_summary_path": str(eval_summary_path),
}
if not diag_path.exists():
print(f"ERROR: {diag_path} not found", file=sys.stderr)
return "INPUT_MISSING", [], meta
train_rows = load_jsonl(diag_path)
meta["n_train_rows"] = len(train_rows)
if len(train_rows) < THRESHOLDS["min_train_rows"]:
print(f"ERROR: only {len(train_rows)} train rows in {diag_path} (need ≥ {THRESHOLDS['min_train_rows']})", file=sys.stderr)
return "INPUT_INSUFFICIENT", [], meta
eval_rows = load_jsonl(eval_diag_path) if eval_diag_path.exists() else []
meta["n_eval_rows"] = len(eval_rows)
eval_summary: dict[str, Any] | None = None
if eval_summary_path.exists():
try:
with eval_summary_path.open("r") as f:
eval_summary = json.load(f)
except (OSError, json.JSONDecodeError) as exc:
print(f"WARN: failed to read {eval_summary_path}: {exc}", file=sys.stderr)
eval_summary = None
meta["eval_summary_present"] = eval_summary is not None
signals: list[SignalResult] = [
signal_action_entropy(train_rows),
signal_q_pi_agree(train_rows),
signal_reward_pnl_pearson(train_rows),
signal_hold_trend(train_rows),
signal_popart_sigma_cv(train_rows),
signal_wr_train(train_rows),
signal_wr_eval(eval_rows),
signal_eval_pnl(eval_summary, eval_rows),
signal_consistency(eval_rows, eval_summary),
]
kill_reasons = [s.name for s in signals if s.status == SignalResult.STATUS_KILL]
if kill_reasons:
overall = "KILL_" + "+".join(s.replace(" ", "_") for s in kill_reasons)
else:
overall = "OK"
return overall, signals, meta
def format_text(overall: str, signals: list[SignalResult], meta: dict[str, Any]) -> str:
lines: list[str] = []
lines.append("=" * 72)
lines.append(f"Tier 1.5 verdict: {overall}")
lines.append("=" * 72)
lines.append(f" out_dir: {meta['out_dir']}")
lines.append(f" train rows: {meta.get('n_train_rows', 0)}")
lines.append(f" eval rows: {meta.get('n_eval_rows', 0)}")
lines.append(f" eval_summary: {'present' if meta.get('eval_summary_present') else 'missing'}")
lines.append("")
lines.append(f"{'STATUS':<6} {'SIGNAL':<42} EXPLANATION")
lines.append(f"{'-'*6} {'-'*42} {'-'*64}")
for s in signals:
marker = {
SignalResult.STATUS_OK: "OK",
SignalResult.STATUS_WARN: "WARN",
SignalResult.STATUS_KILL: "KILL",
SignalResult.STATUS_SKIP: "SKIP",
}[s.status]
lines.append(f"{marker:<6} {s.name:<42} {s.explain}")
lines.append("")
if overall == "OK":
lines.append("[OK] No behavioral red flags. Cluster submit permitted.")
elif overall.startswith("KILL_"):
lines.append(f"[KILL] {overall} — DO NOT submit cluster smoke. Fix locally first.")
else:
lines.append(f"[{overall}] Inputs incomplete — re-run mid-smoke to generate full diag.")
return "\n".join(lines)
def main() -> int:
args = parse_args()
overall, signals, meta = run_verdict(args.out_dir)
if args.json:
payload = {
"overall": overall,
"signals": [s.to_dict() for s in signals],
"meta": meta,
}
print(json.dumps(payload, indent=2))
elif args.quiet:
print(overall)
else:
print(format_text(overall, signals, meta))
if overall == "OK":
return 0
if overall.startswith("KILL_"):
return 1
return 2
if __name__ == "__main__":
sys.exit(main())