From ba83fcd1f5fde328bcd6b20950eeb7c5fef88a2f Mon Sep 17 00:00:00 2001 From: jgrusewski Date: Mon, 4 May 2026 23:18:58 +0200 Subject: [PATCH] =?UTF-8?q?docs(sp13):=20v3=20spec=20+=20plan=20=E2=80=94?= =?UTF-8?q?=20Hold-pricing=20replaces=20Hold-elimination?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit P0a.T3 v2 implementer's audit revealed `DirectionAction` enum doesn't exist; the codebase uses an 8-variant fused `ExposureLevel` (ShortSmall/Half/Full, Hold, LongSmall/Half/Full, Flat) with cross-crate consumers across 77 files and 32+ test files pinning the 8-variant invariant. Atomic Hold elimination would cascade massively. User insight (2026-05-04): Hold being FREE is the bug, not Hold itself. MFT trading legitimately needs multi-bar holds; we want the model to use them deliberately, not as a CQL-bias lazy default. Holding isn't free in the real world — broker fees, margin interest, opportunity cost. v3 reframes as Hold-pricing: - 4-way action space stays; ExposureLevel::Hold stays; no cross-crate cascade - 3 new ISV slots (380-382): HOLD_COST_INDEX, HOLD_RATE_TARGET_INDEX, HOLD_RATE_OBSERVED_EMA_INDEX - Hold-rate observer: small GPU kernel + Pearls A+D smoothing - Hold-cost controller: 5-line deficit-driven formula (excess > target → cost rises 1×→5× base; observed ≤ target → relax) - Per-bar reward subtraction at action == DIR_HOLD site - 2 GPU oracle tests for the controller P0a.T3 cuts from ~250 LOC + 32-test cascade → ~120 LOC additive. T1+T2 already-staged work unchanged. T4/T5/Layer B/C/D structure preserved. Tension with pearl_event_driven_reward_density_alignment acknowledged in spec — per-bar Hold cost is exposure-NEGATIVE (away from Hold), models real economic carry, ISV-bounded by controller. Inverse of the pearl's failure mode. Faithful reward modeling, not artificial shaping. Co-Authored-By: Claude Opus 4.7 (1M context) --- ...3-redefine-success-for-predictive-skill.md | 308 +++++++++++------- ...3-redefine-success-for-predictive-skill.md | 176 ++++++---- 2 files changed, 299 insertions(+), 185 deletions(-) diff --git a/docs/superpowers/plans/2026-05-04-sp13-redefine-success-for-predictive-skill.md b/docs/superpowers/plans/2026-05-04-sp13-redefine-success-for-predictive-skill.md index ab2fd5833..1304b4bed 100644 --- a/docs/superpowers/plans/2026-05-04-sp13-redefine-success-for-predictive-skill.md +++ b/docs/superpowers/plans/2026-05-04-sp13-redefine-success-for-predictive-skill.md @@ -124,9 +124,11 @@ --- -## Phase 0a — Hold elimination + dir_acc instrumentation (1 atomic commit) +## Phase 0a — Hold-pricing + dir_acc instrumentation (1 atomic commit, additive) -**Per `feedback_no_partial_refactor`**: every consumer of the 4-way direction action space migrates to 3-way in a single commit. Replay buffer compatibility intentionally broken (greenfield). +> **v3 NOTE**: P0a.T3 was v2's atomic Hold elimination (would have cascaded across 77 files / 200+ sites including 32 test files). After the T3 implementer's audit found `DirectionAction` enum doesn't exist (codebase uses 8-variant fused `ExposureLevel`), v3 reframed as **Hold-PRICING**: ~120 LOC, additive, no contract change, no cross-crate cascade. The File Structure table above was authored for v2 — entries that conflict with v3 (line 5159 direction-bias re-derivation, financials.rs/monitoring.rs action-packing 12→9, val_dir_dist 4→3, SP5 Pearl 8 trail-dist migration, 30+ stale-doc cleanup) are **NO LONGER IN SCOPE**. The authoritative T3 spec is the v3 P0a.T3 section below. + +**P0a is additive, atomic per `feedback_no_partial_refactor`** — but "atomic" here means single commit, NOT contract change. Replay buffer / fxcache compatibility preserved. ### Task P0a.T1: ISV slot constants + state-reset registry @@ -356,163 +358,217 @@ fn dir_acc_empty_batch_returns_sentinel() { --- -### Task P0a.T3: Direction action space refactor — atomic Hold removal +### Task P0a.T3 (v3): Hold-pricing — adaptive cost controller (additive, no contract change) -**Goal:** `NUM_DIRECTIONS: 4 → 3`. Eliminate Hold from `DirectionAction` enum. Update all 7+ consumers in lockstep. Per `feedback_no_partial_refactor` — single commit, no partial migration. +**Goal:** Make Hold action cost something so the policy uses it deliberately. v2 was atomic Hold elimination; v3 prices Hold via ISV-driven adaptive controller targeting a Hold-rate (~20% — MFT-ish). Per `feedback_isv_for_adaptive_bounds` — base cost is the only constant; rate target and observed rate are ISV slots; cost output is ISV-driven. -**Action semantics in 3-way world** (replaces `project_hold_action_design`): +**Architecture:** +- 3 new ISV slots (380-382): `HOLD_COST_INDEX`, `HOLD_RATE_TARGET_INDEX`, `HOLD_RATE_OBSERVED_EMA_INDEX` +- Hold-rate observer: small GPU kernel writing `(action == DIR_HOLD ? 1.0 : 0.0)` to scratch per-bar; Pearls A+D smoothing (α=0.05) into ISV[382] +- Hold-cost controller: 5-line formula in `training_loop.rs` reading observed rate and writing cost to ISV[380] +- Reward composition: per-bar `reward -= ISV[HOLD_COST_INDEX]` when action == Hold at the existing reward composition site in `experience_kernels.cu` +- 2 GPU oracle tests for the controller -| Prior position | Action picked | Result | -|---|---|---| -| Flat | Long | Open Long | -| Long | Long | Keep Long (was old Hold-Long) | -| Short | Long | Flip to Long | -| Long | Short | Flip to Short | -| Long | Flat | Close position | -| Flat | Flat | No-op | +**Per `pearl_event_driven_reward_density_alignment` tension** (acknowledged): per-bar Hold cost is per-bar reward shaping, which the pearl warns against. The cost is justified as: (a) economically realistic (broker fees, margin interest, opportunity cost), (b) ISV-bounded by the controller, (c) inverse-direction from the pearl's failure mode (pulls policy AWAY from Hold-default, doesn't create exposure-positive bias). -`Long` is "have Long position", not "open Long". Same for Short. Flat = "no position". +**Files:** +- Modify: `crates/ml/src/cuda_pipeline/sp13_isv_slots.rs` (add 3 constants + defaults) +- Modify: `crates/ml/src/cuda_pipeline/sp5_isv_slots.rs` (`SP5_SLOT_END = 383`; layout fingerprint extension) +- Modify: `crates/ml/src/cuda_pipeline/state_layout.cuh` (3 new SP13 slot defines + `ISV_TOTAL_DIM = 383`) +- Modify: `crates/ml/src/trainers/dqn/state_reset_registry.rs` (1 new entry — slot 382 per-fold, sentinel 0.0) +- Modify: `crates/ml/src/trainers/dqn/trainer/training_loop.rs` (reset arm for slot 382 + per-step controller formula) +- Modify: `crates/ml/src/cuda_pipeline/gpu_dqn_trainer.rs` (constructor static-init for slots 380, 381; bump `ISV_TOTAL_DIM = 383`; fingerprint extension) +- Create: `crates/ml/src/cuda_pipeline/hold_rate_observer_kernel.cu` (~30 LOC GPU kernel) +- Modify: `crates/ml/build.rs` (cubin registration) +- Modify: `crates/ml/src/cuda_pipeline/experience_kernels.cu` (locate reward composition site near SP12 alpha cap; add per-bar Hold cost subtraction when action == DIR_HOLD) +- Modify: `crates/ml/tests/sp13_phase0_oracle_tests.rs` (add 2 controller tests) -**Encoding (CRITICAL — preserves codebase Short-first canonical anchor):** -``` -OLD: Short=0, Hold=1, Long=2, Flat=3 (4-way) -NEW: Short=0, Long=1, Flat=2 (3-way) — Hold removed; Long & Flat shift down -``` - -This is NOT `Long=0, Short=1, Flat=2` — keeping Short=0 avoids 30+ stale-comment churn across kernel headers and tests. - -- [ ] **Step 1:** Update `state_layout.cuh`: - -```c -// SP13: direction action space reduced 4 → 3. Hold removed; Long shifts 2→1, Flat 3→2. -// Short-first ordering preserved per codebase canonical anchor. -#define NUM_DIRECTIONS 3 -#define DIRECTION_SHORT 0 -#define DIRECTION_LONG 1 -#define DIRECTION_FLAT 2 -// (DIRECTION_HOLD definition deleted — was =1 in 4-way scheme) -``` - -- [ ] **Step 2:** Update canonical Rust enum at **`crates/ml-core/src/common/action.rs:124-132`** (NOT `crates/ml/src/agents/policy.rs` — that path was speculative; Explore agent confirmed the canonical site): +- [ ] **Step 1:** Add 3 ISV slot constants + defaults to `sp13_isv_slots.rs`: ```rust -#[repr(u8)] -#[derive(Clone, Copy, Debug, PartialEq, Eq)] -pub enum DirectionAction { - Short = 0, - Long = 1, - Flat = 2, -} +pub const HOLD_COST_INDEX: usize = 380; +pub const HOLD_RATE_TARGET_INDEX: usize = 381; +pub const HOLD_RATE_OBSERVED_EMA_INDEX: usize = 382; + +/// Per-bar Hold cost base (controller can scale up to 5×, down to 0.5×). +/// Small per-bar (~10 ticks of price for ES.FUT at ~0.25 tick value) but cumulative +/// across long Hold runs becomes meaningful relative to capped trade reward (±5..10). +pub const HOLD_COST_BASE: f32 = 0.001; + +pub const HOLD_RATE_TARGET_DEFAULT: f32 = 0.20; +pub const HOLD_COST_CONTROLLER_GAIN: f32 = 5.0; +pub const HOLD_COST_FLOOR_RATIO: f32 = 0.5; /* never below 0.5 × base */ +pub const HOLD_COST_CEIL_RATIO: f32 = 5.0; /* never above 5.0 × base */ ``` -And `crates/ml-core/src/state_layout.rs:48-49`: +- [ ] **Step 2:** Bump `SP5_SLOT_END = 383` (was 380 from T1) in `sp5_isv_slots.rs`. Extend `SP5_LAYOUT_FINGERPRINT_FRAGMENT` with the 3 new slot names. Update test expectations: `SP5_PRODUCER_COUNT = 209`. -```rust -// (3 actions: Short / Long / Flat; Hold removed in SP13) -pub const NUM_DIRECTIONS: usize = 3; // was 4 -``` +- [ ] **Step 3:** Mirror in `state_layout.cuh`: `#define ISV_TOTAL_DIM 383`; add the 3 new SP13 slot defines. -Compiler will surface every `Hold` match arm. Fix each — do NOT add `_ => unreachable!()` cover-ups. +- [ ] **Step 4:** Add reset registry entry in `state_reset_registry.rs` for slot 382 (per-fold, sentinel `0.0` — no observations at fold start). Slots 380 and 381 are static-init in constructor. -- [ ] **Step 3:** Update `experience_action_select` direction branch at `experience_kernels.cu:1013`: argmax over 3 instead of 4. Thompson sampling on 3 atoms (`pearl_thompson_for_distributional_action_selection`). +- [ ] **Step 5:** Add reset_named_state dispatch arm for slot 382 in `training_loop.rs` (writes 0.0). -- [ ] **Step 4:** **NEW — Direction-bias signal at `experience_kernels.cu:5159`:** +- [ ] **Step 6:** Constructor static-init in `gpu_dqn_trainer.rs` for slots 380 (`HOLD_COST_BASE`) and 381 (`HOLD_RATE_TARGET_DEFAULT`). Extend `layout_fingerprint_seed()` byte-string with 3 new slot names + `ISV_TOTAL_DIM=383`. Bump `ISV_TOTAL_DIM` constant. -The 4-tuple per-direction bias weights `[Short=1.0, Hold=0.5, Long=1.0, Flat=0.0]` encoded a softening gate where Hold attenuated the Q-signal. With Hold removed: +- [ ] **Step 7:** Write the Hold-rate observer kernel `crates/ml/src/cuda_pipeline/hold_rate_observer_kernel.cu`: ```cuda -// SP13: 4-tuple direction-bias weights → 3-tuple (Hold's 0.5 softening gate removed -// since Hold itself is removed; Long and Short stay full-weight, Flat stays soft). -// OLD: const float dir_bias[4] = {1.0f, 0.5f, 1.0f, 0.0f}; // [Short, Hold, Long, Flat] -const float dir_bias[3] = {1.0f, 1.0f, 0.0f}; // [Short, Long, Flat] +// hold_rate_observer_kernel.cu +// Per-step: writes 1.0 if any bar in batch picked Hold, 0.0 otherwise (or aggregated +// fraction). Single-block reduce; mapped-pinned output for Pearls A+D smoothing. + +#include "state_layout.cuh" + +extern "C" __global__ void hold_rate_observer_kernel( + const int* __restrict__ actions_dir, /* [B] direction actions, values ∈ {0..3} */ + int batch_size, + float* __restrict__ hold_rate_out /* [1] = (count(Hold) / B) */ +) { + const int tid = threadIdx.x; + const int bdim = blockDim.x; + extern __shared__ int sh_hold[]; + + int local_hold = 0; + for (int i = tid; i < batch_size; i += bdim) { + if (actions_dir[i] == DIR_HOLD) local_hold += 1; + } + sh_hold[tid] = local_hold; + __syncthreads(); + + for (int s = bdim / 2; s > 0; s >>= 1) { + if (tid < s) sh_hold[tid] += sh_hold[tid + s]; + __syncthreads(); + } + + if (tid == 0) { + hold_rate_out[0] = (batch_size > 0) ? ((float)sh_hold[0] / (float)batch_size) : 0.0f; + } +} ``` -Verify the loop bounds at `:5159` use `NUM_DIRECTIONS` (now 3). If hardcoded `<4`, change. +(Single-block, 256 threads, 1 KB shared mem. Mirror the launch infrastructure of `aux_dir_acc_reduce` from T2.) -- [ ] **Step 5:** Update Q-head output dim in `dqn_inference.cu` and `dqn_param_layout.rs`: 4 → 3. Bump layout fingerprint seed (invalidates cached checkpoints — greenfield acceptable). +- [ ] **Step 8:** Register cubin in `build.rs` matching the existing pattern. -- [ ] **Step 6:** Update C51/IQN per-direction-bucket loops at the specific sites Explore agent located: - - `c51_loss_kernel.cu:394` — doc comment update - - `c51_loss_kernel.cu:460` — `for (int k=1; k<4; k++)` → `< NUM_DIRECTIONS` - - `c51_loss_kernel.cu:1088` — argmax across direction bins → 3 - - `c51_grad_kernel.cu` — **AUDIT**: grep for `for.*<.*4` direction-axis loops, update each - - `experience_kernels.cu:4059` — direction accumulation loop - - `experience_kernels.cu:5306` — conviction EMA loop - - `experience_kernels.cu:6371-6397` — 4× per-atom direction loops, all need bound change - - Per-direction ISV slot ranges (SP4/SP5/SP6 atoms_pos/tau/...) KEEP 4-entry length; slot index 3 dead-space: - - ```rust - // SP13: NUM_DIRECTIONS reduced 4→3 (Hold removed). Slot index 3 dead-space until - // SP13 close-out recompaction; consumers read by direction index ∈ 0..2 so dead - // slot is never touched. - ``` - -- [ ] **Step 7:** SP5 Pearl 8 trail-distance migration (per-direction ISV slot range stays 4 entries; producer kernel rewrites for new direction indexing): - - `sp5_isv_slots.rs:272` — `TRAIL_DIST_PER_DIR_BASE=270..274` doc-comment update from `[4 dirs: Short, Hold, Long, Flat]` to `[3 dirs at Short=0, Long=1, Flat=2; slot 273 dead-space]` - - `pearl_8_trail_kernel.cu` — producer writes at new indices: Short→slot 270, Long→slot 271 (was Hold), Flat→slot 272 (was Long); slot 273 stays zero - - `state_reset_registry.rs:974` — comment update; reset values stay 4-entry (slot 273 stays sentinel) - - `training_loop.rs:4046-4048` — Pearl 8 `apply_pearls_ad_kernel` calls: 4 → 3 (skip slot 273 — no smoothing of dead slot) - -- [ ] **Step 8:** **Action packing recompacts 12 → 9 slots** (real shape change, not slot-3 dead-space): - - `crates/ml/src/trainers/dqn/financials.rs:27,197` — `exp_idx = dir*3 + mag` semantics: was 4 dirs × 3 mags = 12 indices `[Short[0..3], Hold[3..6], Long[6..9], Flat[9..12]]`; now 3 dirs × 3 mags = 9 indices `[Short[0..3], Long[3..6], Flat[6..9]]` - - `crates/ml/src/trainers/dqn/monitoring.rs:341` — `action_counts[9..12]` Flat-mag indexing: now `action_counts[6..9]` - - `crates/ml/src/cuda_pipeline/iqn_dual_head_kernel.cu:200,245,1199` — action layout `[direction, magnitude, order, urgency]` is 4 SUB-ACTIONS per sample (NOT 4-way direction); only the direction VALUE range shrinks 0..3 → 0..2; `actions[sample * 4 + 0]` indexing stays. Verify consumers expect dir < 3. - - Action-counts buffer allocation: shrink from `12 * batch_size` → `9 * batch_size` (search for allocation site) - -- [ ] **Step 9:** Update HEALTH_DIAG 4-bucket → 3-bucket emit at all sites: - - `metrics.rs:949` — `"val_dir_dist [short={:.4} hold={:.4} long={:.4} flat={:.4}]"` → `"val_dir_dist [short={:.4} long={:.4} flat={:.4}]"` - - `magnitude_distribution.rs:88,217` — `EVAL_DIR_DIST` 4-tuple → 3-tuple - - `distributional_q_tests.rs:790` — `σ_C51 per direction` test output - -- [ ] **Step 10:** Update KL divergence calc if present. Explore agent did NOT find a `kl_divergence_kernel.cu` — verify with `find crates -name "*kl*" -type f`. If absent, this step is a no-op (record in commit message). - -- [ ] **Step 11:** Update tests + build comments: - - `crates/ml/src/trainers/dqn/trainer/tests.rs:542` — `assert_eq!(config.num_actions, 4, …)` → `3` - - `crates/ml/src/trainers/dqn/distributional_q_tests.rs:888` — `const PROD_B0: i32 = 4` → `3` - - `crates/ml/src/cuda_pipeline/thompson_test_kernel.cu` — direction count constant 4 → 3 - - `crates/ml/build.rs:490,496` — comment `"4-thread single-block kernel (one per direction: Short=0, Hold=1, Long=2, Flat=3)"` → `"3-thread … Short=0, Long=1, Flat=2"` - -- [ ] **Step 12:** Add 5 GPU oracle tests for 3-direction action selection (NEW Short-first indexing) in `sp13_phase0_oracle_tests.rs`: +- [ ] **Step 9:** Add launcher in `gpu_dqn_trainer.rs` (mirroring `launch_aux_dir_acc_reduce`): ```rust -// New encoding: Short=0, Long=1, Flat=2 - -#[test] fn action_select_3way_argmax_short() { /* q=[2.0, 1.0, 0.5] → 0 (Short wins) */ } -#[test] fn action_select_3way_argmax_long() { /* q=[1.0, 2.0, 0.5] → 1 (Long wins) */ } -#[test] fn action_select_3way_argmax_flat() { /* q=[0.5, 0.3, 2.0] → 2 (Flat wins) */ } -#[test] fn action_select_no_hold_leaks() { - // 100 random q-vectors → 0 ≤ action < 3 always; never 3 -} -#[test] fn action_packing_9_slot_invariant() { - // Pack (dir, mag) for all 9 (dir, mag) combinations; verify exp_idx ∈ [0, 9) +pub fn launch_hold_rate_observer( + &self, + actions_dir: &CudaSlice, + batch_size: i32, + out_1: &mut CudaSlice, +) -> Result<(), MLError> { + let bdim: u32 = 256; + let cfg = LaunchConfig { + grid_dim: (1, 1, 1), + block_dim: (bdim, 1, 1), + shared_mem_bytes: bdim * std::mem::size_of::() as u32, + }; + unsafe { + self.hold_rate_observer + .clone() + .launch(cfg, (actions_dir, batch_size, out_1))?; + } + Ok(()) } ``` -- [ ] **Step 13:** Stale doc cleanup — implementer audit: +Add `hold_rate_observer: CudaFunction` and `hold_rate_buf: MappedF32Buffer` (1 float). -```bash -grep -rn "Short=0.*Hold=1.*Long=2.*Flat=3\|S=Short.*H=Hold.*L=Long.*F=Flat\|\[Short, Hold, Long, Flat\]" \ - crates/ --include="*.rs" --include="*.cu" --include="*.cuh" \ - | grep -v "// SP13:" +- [ ] **Step 10:** Wire the per-step launch + Pearls A+D smoothing + controller in `training_loop.rs` (after action selection, before next step): + +```rust +// SP13 v3: Hold-rate observer + Hold-cost controller +self.launch_hold_rate_observer( + &actions_dir, /* [B] direction actions from policy this step */ + batch_size as i32, + self.hold_rate_buf.device_slice_mut(), +)?; + +// Pearls A+D smoothing (α=0.05) into ISV[HOLD_RATE_OBSERVED_EMA_INDEX] +self.launch_apply_pearls_ad( + self.hold_rate_buf.device_slice(), + 1, + HOLD_RATE_OBSERVED_EMA_INDEX, + /* alpha */ 0.05, +)?; + +// Controller (deficit-driven): when observed > target, raise cost +let observed = self.read_isv_slot(HOLD_RATE_OBSERVED_EMA_INDEX); +let target = self.read_isv_slot(HOLD_RATE_TARGET_INDEX); +let excess = (observed - target).max(0.0); +let hold_cost = HOLD_COST_BASE * (1.0 + HOLD_COST_CONTROLLER_GAIN * excess); +let hold_cost = hold_cost.clamp( + HOLD_COST_BASE * HOLD_COST_FLOOR_RATIO, + HOLD_COST_BASE * HOLD_COST_CEIL_RATIO, +); +self.write_isv_slot(HOLD_COST_INDEX, hold_cost)?; ``` -Update narrative comments to `Short=0, Long=1, Flat=2` ordering. Specific known sites: -- `crates/ml/src/trainers/dqn/monitoring.rs:13` -- `crates/ml/src/trainers/dqn/trainer/mod.rs:1966` -- 20+ kernel header comments +(Match existing patterns for ISV reads/writes in this file. The `apply_pearls_ad_kernel` API may differ slightly; consult adjacent SP11/SP12 launches for the exact signature.) -- [ ] **Step 14:** Sanity-check no Hold leaks: +- [ ] **Step 11:** Add per-bar Hold cost subtraction in `experience_kernels.cu`. Locate the reward composition site (SP12 asymmetric cap is around line 2788 per `pearl_symmetric_clamp_audit`): -```bash -grep -rn "DIRECTION_HOLD\|DirectionAction::Hold\|::Hold\b\|hold=\|action == 3\|< 4.*direction\|dir.*== 1.*Hold" \ - crates/ --include="*.rs" --include="*.cu" --include="*.cuh" \ - | grep -v test | grep -v "// SP13" +```cuda +// SP13 v3: per-bar Hold cost (priced action — ISV-driven adaptive) +if (action_dir == DIR_HOLD) { + base_reward -= isv[HOLD_COST_INDEX]; +} + +// (then existing SP12 asymmetric cap, etc.) +float r = fmaxf(-10.0f, fminf(base_reward, 5.0f)); ``` -Expected: zero hits. +(Adapt to actual variable names at the call site. Hold cost subtraction goes BEFORE the SP12 asymmetric cap so the cost participates in the bounded clamp — keeps total reward in the same range.) -- [ ] **Step 15:** Stage all P0a.T3 changes (atomic with T1, T2, T4 in P0a.T5 commit). +- [ ] **Step 12:** Add 2 GPU oracle tests to `sp13_phase0_oracle_tests.rs`: + +```rust +#[test] +fn hold_rate_observer_all_hold_returns_one() { + // Batch where every action == DIR_HOLD (=1) → out = 1.0 + let actions = vec![1_i32; 10]; + let out = run_hold_rate_observer(&actions); + assert!((out - 1.0).abs() < 1e-6); +} + +#[test] +fn hold_rate_observer_no_hold_returns_zero() { + // Batch with mix of Short(0)/Long(2)/Flat(3), no Hold(1) → out = 0.0 + let actions = vec![0_i32, 2, 3, 0, 2, 3]; + let out = run_hold_rate_observer(&actions); + assert!(out.abs() < 1e-6); +} + +// (Optional 3rd: half-hold case verifying out = 0.5 — implementer adds if scaffold permits) +``` + +- [ ] **Step 13:** Verify on RTX 3050 Ti: + +```bash +SQLX_OFFLINE=true cargo check -p ml --lib # clean +SQLX_OFFLINE=true cargo build -p ml --release # cubin compiles +SQLX_OFFLINE=true cargo test -p ml --features cuda --test sp13_phase0_oracle_tests -- --ignored +# Expected: 8 passing (6 from T2 + 2 new) +SQLX_OFFLINE=true cargo test -p ml --features cuda --test sp12_reward_math_tests -- --ignored # SP12 still green +``` + +- [ ] **Step 14:** Stage all P0a.T3 v3 changes (atomic with T1, T2, T4 in P0a.T5 commit): + +```bash +git add crates/ml/src/cuda_pipeline/sp13_isv_slots.rs \ + crates/ml/src/cuda_pipeline/sp5_isv_slots.rs \ + crates/ml/src/cuda_pipeline/state_layout.cuh \ + crates/ml/src/cuda_pipeline/gpu_dqn_trainer.rs \ + crates/ml/src/cuda_pipeline/hold_rate_observer_kernel.cu \ + crates/ml/src/cuda_pipeline/experience_kernels.cu \ + crates/ml/src/trainers/dqn/state_reset_registry.rs \ + crates/ml/src/trainers/dqn/trainer/training_loop.rs \ + crates/ml/build.rs \ + crates/ml/tests/sp13_phase0_oracle_tests.rs +``` --- diff --git a/docs/superpowers/specs/2026-05-04-sp13-redefine-success-for-predictive-skill.md b/docs/superpowers/specs/2026-05-04-sp13-redefine-success-for-predictive-skill.md index e375f4a07..980d427ec 100644 --- a/docs/superpowers/specs/2026-05-04-sp13-redefine-success-for-predictive-skill.md +++ b/docs/superpowers/specs/2026-05-04-sp13-redefine-success-for-predictive-skill.md @@ -1,6 +1,30 @@ -# SP13 — Redefine Success for Predictive Skill (v2) +# SP13 — Redefine Success for Predictive Skill (v3) -**Date**: 2026-05-04 (v2 — supersedes v1 in-place) +**Date**: 2026-05-04 (v3 — supersedes v2 in-place after architectural finding) + +## v2 → v3 change (added during P0a.T3 dispatch) + +P0a.T3 implementer's audit revealed `DirectionAction` enum doesn't exist — the action +space uses an 8-variant `ExposureLevel` (fused dir+mag) with cross-crate consumers +across 77 files / 200+ Hold sites. Eliminating Hold would cascade through 32+ test +files pinning the 8-variant invariant. + +User insight (2026-05-04): "Why is Hold free anyway? MFT-style with Hold for when it's +really needed could be fine. Holding isn't free — infrastructure costs still money even +doing nothing. The model should learn to hold carefully." + +v3 reframes: **Hold being FREE is the bug, not Hold itself**. Pricing Hold: +- Preserves legitimate MFT use cases (multi-bar conviction holds) +- Avoids 200+ site cross-crate cascade +- Tests a sharper hypothesis: "free Hold + CQL bias = lazy default; priced Hold forces deliberate use" +- Cuts P0a.T3 from ~250 LOC + 32-test cascade → ~120 LOC, no cross-crate refactor +- Replay buffer / fxcache compatibility preserved (no contract change) +- 4-way action space stays; HEALTH_DIAG val_dir_dist stays 4-bucket; no stale-doc + cleanup needed + +The data-investigation hypothesis ("data has signal, Hold hides it") is unchanged. The +test mechanism shifts from "remove Hold so policy must commit" to "price Hold so policy +must commit when no edge exists, can still hold deliberately when edge persists". **Branch**: continues `sp11-reward-as-controlled-subsystem` (single coherent architectural lineage) **Builds on**: SP11 (cap stability) + SP12 (event-driven density) + asymmetric loss aversion **Anchor pearls**: @@ -78,48 +102,59 @@ persists because Hold hides it. P0a smoke decides cleanly: ## The five changes -### Change 1: Eliminate Hold action +### Change 1: Price Hold (replaces v2's "Eliminate Hold") **Current**: 4-way direction `{Short=0, Hold=1, Long=2, Flat=3}` × magnitude × ord × urg. -**New**: 3-way `{Short=0, Long=1, Flat=2}` — Hold removed, Short stays at 0 (preserves -canonical anchor used across 30+ kernel comments and tests), Long shifts down 2 → 1, -Flat shifts down 3 → 2. +Hold has zero cost: model can pick it indefinitely with no penalty. Combined with CQL's +conservative Q-bias (pulls toward minimum-variance action), Hold becomes the safe +default ~45% of the time across SP1-SP12. -**Action semantics in 3-way world** (replaces `project_hold_action_design`): -- `Long` at t when prior position = Flat → open Long -- `Long` at t when prior position = Long → keep Long (same as old "Hold-Long") -- `Long` at t when prior position = Short → flip to Long -- `Flat` at t when any position → close all positions -- `Short` symmetric +**New**: same 4-way action space; **Hold action incurs a per-bar cost**. Adaptive, +ISV-driven controller targets a Hold-rate (~20% — MFT-ish, deliberate use). Cost +self-balances: -`Long` is no longer "open Long" specifically; it is "have Long position". `Long` is a -state, not a verb. +``` +At each bar, if action == Hold: + reward[bar] -= ISV[HOLD_COST_INDEX] -**Affected** (atomic per `feedback_no_partial_refactor`, ~50–70 touch-sites mapped by -Explore agent 2026-05-04): -- Canonical Rust enum at `crates/ml-core/src/common/action.rs:124-132` and - `NUM_DIRECTIONS: 4 → 3` at `crates/ml-core/src/state_layout.rs:48-49` -- `state_layout.cuh` C/CUDA mirror: `#define NUM_DIRECTIONS 3` and `DIRECTION_SHORT=0, - DIRECTION_LONG=1, DIRECTION_FLAT=2` (DIRECTION_HOLD removed) -- Q-head output dim 4 → 3, weight/bias re-init, layout fingerprint bump -- Action selection kernel argmax over 3 (`experience_action_select`) -- C51/IQN per-direction-bucket loops (~6 `for d<4` sites become `< NUM_DIRECTIONS`) -- **Action packing recompacts 12 → 9 slots**: `action_counts` buffer and `exp_idx = - dir*3 + mag` packing in `financials.rs:197` shrink from 12 → 9 (3 dirs × 3 mags). - Real shape change, not slot-3 dead-space. Consumers in `monitoring.rs:341` reindex. -- **Direction-bias signal at `experience_kernels.cu:5159`**: 4-tuple per-direction - bias weights `[Short=1.0, Hold=0.5, Long=1.0, Flat=0.0]` re-derives to 3-tuple - `[Short=1.0, Long=1.0, Flat=0.0]`. Hold's 0.5 was a softening gate; with Hold gone - the gate is no longer needed (Long/Short are full-weight, Flat is the soft branch). -- SP5 Pearl 8 per-direction trail-distance ISV range `TRAIL_DIST_PER_DIR_BASE=270..274`: - producer rewrites at new indices `{0=Short, 1=Long, 2=Flat}`; slot 273 dead-space - (per-direction-indexed ISV ranges keep length 4 to avoid cascade-renumber of all - downstream slots — consumers read by direction index ∈ 0..2 so dead slot is never - touched). Recompaction deferred to SP13 close-out. -- HEALTH_DIAG val_dir_dist 3-bucket emit: `metrics.rs:949`, `magnitude_distribution.rs:88,217`, - `distributional_q_tests.rs:790` -- KL divergence calc (if present) -- 30+ stale doc/comment lines across kernel headers — implementer audits in same commit +Controller (per-step, after action selection): + observed_hold_rate = ISV[HOLD_RATE_OBSERVED_EMA_INDEX] /* fast EMA of Hold-pick frequency */ + target_hold_rate = ISV[HOLD_RATE_TARGET_INDEX] /* default 0.20 */ + excess = max(0, observed_hold_rate - target_hold_rate) + hold_cost = HOLD_COST_BASE × (1 + 5 × excess) + hold_cost = clamp(hold_cost, HOLD_COST_BASE × 0.5, HOLD_COST_BASE × 5.0) + ISV[HOLD_COST_INDEX] = hold_cost +``` + +When the model abuses Hold (rate > target), cost rises until Hold becomes uneconomic. +When Hold rate ≤ target, cost relaxes to baseline (just realistic carry). The model +learns to use Hold deliberately for high-conviction multi-bar positions. + +**Action semantics preserved**: +- `Hold` (action 1) — keep current position; pays per-bar cost; legitimate for MFT +- `Flat` (action 3) — close to no position; pays close-slippage once, then near-zero infra +- `Long`/`Short` — open/keep directional position; pays carry implicitly via P&L + +**Affected** (~120 LOC, additive, no contract change): +- 3 new ISV slots: `HOLD_COST_INDEX`, `HOLD_RATE_TARGET_INDEX`, `HOLD_RATE_OBSERVED_EMA_INDEX` +- Hold-rate observer: small GPU kernel writing `(action == Hold ? 1.0 : 0.0)` to scratch, + Pearls A+D smoothing into ISV[HOLD_RATE_OBSERVED_EMA_INDEX] +- Hold-cost controller: 5-line formula in `training_loop.rs` (per-step, after observer) +- Reward composition: per-bar subtraction at `experience_kernels.cu` reward composition + site when action == Hold +- 2 GPU oracle tests for the controller (target tracking, deficit response) + +**No changes to** (vs v2 P0a): action space dimension, `ExposureLevel` enum, Q-head dim, +C51/IQN bucket loops, action packing, replay buffer encoding, HEALTH_DIAG val_dir_dist +buckets, SP5 Pearl 8 trail-dist, 30+ stale-doc audit. The 4-way world stays. + +**Tension with `pearl_event_driven_reward_density_alignment`** (acknowledged): +the pearl warns against per-bar reward shaping creating exposure-positive bias toward +dense-signal optimization. Per-bar Hold cost is the inverse case — exposure-NEGATIVE +shaping that models real economic carry cost (broker fees, margin interest, opportunity +cost). It's faithful reward modeling, not artificial shaping. ISV-bounded by controller +target so cost can't run away. We KEEP the pearl's intent (don't add fake dense rewards +the policy will optimize against the objective) while pricing a real economic friction. ### Change 2: Direction-skill bonus (bounded) @@ -218,7 +253,7 @@ the fix, not a confounder, because the SP11 formula was structurally wrong). ## ISV slot allocation -8 new slots in `[372..380)`: +11 new slots in `[372..383)` (8 from v2 + 3 added in v3 for Hold-pricing): | Slot | Name | Purpose | Reset | |---|---|---|---| @@ -230,8 +265,18 @@ the fix, not a confounder, because the SP11 formula was structurally wrong). | 377 | `DIR_SKILL_BONUS_BETA_INDEX` | Skill bonus penalty (wrong direction), default 1.0 | per-fold | | 378 | `LUCK_WIN_DISCOUNT_INDEX` | Lucky-win discount factor, default 0.3 | per-fold | | 379 | `SKILL_BONUS_CAP_RATIO_INDEX` | Cap ratio: bonus ≤ ratio × \|alpha\|, default 0.3 | per-fold | +| 380 | `HOLD_COST_INDEX` | Per-bar Hold cost (controller output), default `HOLD_COST_BASE` | per-fold | +| 381 | `HOLD_RATE_TARGET_INDEX` | Target Hold-pick rate (default 0.20) | static | +| 382 | `HOLD_RATE_OBSERVED_EMA_INDEX` | Observed Hold-pick rate EMA (α=0.05) | per-fold | -Bump `SP5_SLOT_END = 380`, `ISV_TOTAL_DIM = 380`. +Defaults: `HOLD_COST_BASE = 0.001` (small per-bar; cumulative across long Hold runs +becomes significant). `HOLD_RATE_TARGET_DEFAULT = 0.20`. Sentinel for slot 382: 0.0 +(zero observed Hold-picks at fold start). + +Bump `SP5_SLOT_END = 383`, `ISV_TOTAL_DIM = 383`. + +T1 already staged with `SP5_SLOT_END = 380` (8 slots). T3 v3 work bumps to 383 +additively (slots 380-382 are net-new in T3, no T1 rework). **Slot 371 (was BENCHMARK_PNL_CUMULATIVE_INDEX in v1) — DROPPED.** Per-fold buy-and-hold delta is computed at fold-end in HEALTH_DIAG emit (trivial: `price[fold_end] − @@ -239,31 +284,44 @@ price[fold_start]`). No kernel, no slot, no reward dependency. ## Implementation phases -### Phase 0a — Hold elimination only (1 atomic commit) +### Phase 0a — Hold-pricing + dir_acc instrumentation (1 atomic commit, additive) -**Purpose**: clean test of user's hypothesis. Hold removed; aux_w controller, dir_acc -instrumentation, skill bonus, aux classification all UNCHANGED from current main. +**Purpose**: clean test of user's hypothesis. Hold-pricing forces the policy to commit +when no edge exists (CQL bias can no longer make Hold the lazy default), while +preserving Hold for legitimate MFT use. aux_w controller, skill bonus, aux +classification all UNCHANGED from current main (those land in P0b/B/C). -- ISV slot 372 (target, static), 373 (short EMA), 374 (long EMA), 375 (aux prediction — - but produced by the existing regression aux head, just `tanh(scalar_pred)` for now to - fit the [−1, +1] interface) -- `aux_dir_acc_reduce_kernel` — single-block tree-reduce of (correct_count, pos_pred, - pos_label, valid_count) → 3 output floats per epoch -- HEALTH_DIAG `aux_dir_acc accuracy=… pos_pred=… pos_label=…` per epoch -- Hold elimination atomic refactor (state_layout, Q-head, action select, C51/IQN bucket - state, replay encoding, val_dir_dist, KL, Rust enum) -- ~300 LOC +- 11 ISV slots (372-382): dir_acc tracking (372-374), aux prediction (375), Layer C + reward composition slots (376-379, populated but unused until Layer C), and the v3 + Hold-pricing trio (380-382). +- T1 (already staged): 8 slots scaffolding (372-379), state-reset registry, constructor + static-init. +- T2 (already staged): `aux_dir_acc_reduce_kernel` + 6 GPU oracle tests. +- **T3 v3** (replaces v2's Hold-elimination): adds 3 Hold-pricing slots (380-382), + Hold-rate observer kernel, Hold-cost controller, per-bar reward subtraction site, + 2 GPU oracle tests for the controller. ~120 LOC, NO contract change. +- T4: wire dir_acc launch + HEALTH_DIAG `aux_dir_acc` + `hold_cost` + `observed_hold_rate` + emit per epoch. +- T5: atomic commit + 10-epoch L40S smoke. **Phase 0a smoke** (10-epoch L40S): - ✅ **Pass**: `val_dir_acc[ep9] > val_dir_acc[ep1]` AND `val_dir_acc[ep9] > 0.52` AND - `val_win_rate[ep9] > 0.51` → user's hypothesis confirmed; proceed Layer B. -- ⚠️ **Partial** (one trending up, the other flat): proceed P0b. -- ❌ **Both flat** (`dir_acc < 0.51` AND `win_rate < 0.50` after 10 epochs): proceed P0b - to add aux amplification. If P0b also fails, halt SP13. + `val_win_rate[ep9] > 0.51` AND `observed_hold_rate ≤ 0.30` → hypothesis confirmed; + priced Hold + remaining policy commits to directions; proceed Layer B. +- ⚠️ **Partial** (one trending up, the other flat, OR Hold rate stuck > 0.40 despite + controller escalation): proceed P0b for aux amplification. +- ❌ **All flat** (dir_acc < 0.51 AND win_rate < 0.50 AND Hold-rate stuck high): proceed + P0b. If P0b also fails, halt SP13 — data lacks signal at this timescale, plan SP14. 10 epochs (not 5) because aux head regression has been suppressed for entire SP1-SP12 -lineage — needs time even at unchanged aux_w to show whether signal is recoverable. -Cost: ~€0.60. +lineage — needs time at unchanged aux_w to show whether signal is recoverable. Cost: ~€0.60. + +**Why this is a sharper test than v2 elimination**: v2 forced ALL Hold-equivalent +behavior (the model couldn't pick Hold). v3 lets the model still pick Hold, but only +when the perceived edge justifies the cost. If model converges to a low Hold-rate (≤ +30%) AND WR climbs, both conditions pass — clean confirmation. If model maintains high +Hold-rate despite escalating cost AND WR stays flat, the data lacks signal (the model +prefers paying for Hold over committing because committing has no edge). ### Phase 0b — Add corrected aux_w controller (only if 0a partial/red, atomic commit)