fix(sp22): H6 Phase 3 beta - activate via 0.5 structural prior (B6-minimal)

Bug: state_reset_registry has FoldReset entries for sp22_reward_aux_align_ema
and sp22_aux_align_scale, but reset_named_state dispatch had no matching arms
-> unknown name error at first fold boundary with these registry entries.

Latent in 5106e3b117 - single-fold smoke (current train-qsltr) avoids the
trigger but any --folds 2+ run would fail.

Fix: add dispatch arms for both names. Plus elevate scale_beta cold-start
from 0 to 0.5 - a structural prior matching W [-0.5, 0, +0.5, 0] init
approach. Activates beta reward shaping with fixed magnitude from epoch 0
without waiting for Phase B6 full SP11 controller extension.

Once B6 lands (SP11 controller emits scale_beta adaptively per bounds
[0.05, 2.0]), the 0.5 prior is overwritten on first emit per
pearl_first_observation_bootstrap.

Docs updated: sp22_isv_slots.rs, state_reset_registry.rs, audit doc.
Verification: cargo check -p ml --lib clean.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
This commit is contained in:
jgrusewski
2026-05-13 08:30:31 +02:00
parent 9c0aaacdfb
commit ff98edc774
4 changed files with 55 additions and 10 deletions

View File

@@ -23,15 +23,21 @@
//! //!
//! ## Cold-start sentinel discipline //! ## Cold-start sentinel discipline
//! //!
//! Both slots cold-start at sentinel 0 per `pearl_first_observation_bootstrap`. //! `REWARD_AUX_ALIGN_EMA_INDEX` cold-starts at sentinel 0 per
//! For `SP22_AUX_ALIGN_SCALE_INDEX`, the cold-start zero makes β a no-op //! `pearl_first_observation_bootstrap` — the producer's first observation
//! until the SP11 controller's first emit (sometime in the first epoch //! replaces directly (Pearl A bootstrap).
//! after the EMA producer has accumulated a non-zero signal). For
//! `REWARD_AUX_ALIGN_EMA_INDEX`, the producer's first observation replaces
//! the sentinel directly (Pearl A bootstrap).
//! //!
//! Registered in `state_reset_registry.rs` as FoldReset category so both //! `SP22_AUX_ALIGN_SCALE_INDEX` cold-starts at sentinel **0.5** — a
//! slots reset to 0 at fold boundaries. //! structural prior matching the W [-0.5, 0, +0.5, 0] init approach.
//! Activates β with a fixed magnitude from epoch 0. This deviates from
//! the original spec (sentinel 0) because slot 537's producer (Phase B6
//! SP11 controller extension to 7 components) has NOT yet landed. With
//! sentinel 0, β would be a runtime no-op. The 0.5 prior makes β fire
//! from epoch 0 with a fixed magnitude. Once B6 lands and emits adaptively,
//! the prior is overwritten on first emit per `pearl_first_observation_bootstrap`.
//!
//! Registered in `state_reset_registry.rs` as FoldReset category — slot
//! 536 resets to 0, slot 537 resets to 0.5.
/// SP22 H6 Phase 3 (2026-05-13) — EMA of the 7th reward component /// SP22 H6 Phase 3 (2026-05-13) — EMA of the 7th reward component
/// `r_aux_align`. Produced by the extended (0..7 iteration) /// `r_aux_align`. Produced by the extended (0..7 iteration)
@@ -46,7 +52,7 @@
/// non-zero observation replaces directly. FoldReset. /// non-zero observation replaces directly. FoldReset.
pub const REWARD_AUX_ALIGN_EMA_INDEX: usize = 536; pub const REWARD_AUX_ALIGN_EMA_INDEX: usize = 536;
/// SP22 H6 Phase 3 (2026-05-13) — SP11-controller-emitted adaptive /// SP22 H6 Phase 3 (2026-05-13) — SP11-controller-emitted (eventual) adaptive
/// `scale_β` for the β producer at trade-close. /// `scale_β` for the β producer at trade-close.
/// ///
/// Producer: extended `reward_subsystem_controller_kernel` (7-weight /// Producer: extended `reward_subsystem_controller_kernel` (7-weight

View File

@@ -2164,7 +2164,7 @@ impl StateResetRegistry {
RegistryEntry { RegistryEntry {
name: "sp22_aux_align_scale", name: "sp22_aux_align_scale",
category: ResetCategory::FoldReset, category: ResetCategory::FoldReset,
description: "ISV[SP22_AUX_ALIGN_SCALE_INDEX=537] — SP22 H6 Phase 3 (2026-05-13) SP11-controller-emitted adaptive scale_β for the β producer at trade-close. Producer: extended `reward_subsystem_controller_kernel.cu` (7-weight output, anchor on `REWARD_AUX_ALIGN_EMA_INDEX=536`, target ~10% of total reward magnitude, bounds [0.05, 2.0]). Consumers: `experience_env_step::segment_complete` (training-side β: `r_aux_align = scale_β × alignment × profit_pos`) and `backtest_env_kernel::segment_complete` (eval-side β with the same formula); HEALTH_DIAG `sp11_reward` printer (the `w_aux=…` column). FoldReset sentinel 0 — when scale_β = 0, β is a no-op (no reward signal added), which is the correct behavior pre-first-controller-emit. The controller's first emit replaces the sentinel directly per `pearl_first_observation_bootstrap`; subsequent EMA-driven adaptation tracks the anchor signal within the bounds.", description: "ISV[SP22_AUX_ALIGN_SCALE_INDEX=537] — SP22 H6 Phase 3 (2026-05-13) SP11-controller-emitted adaptive scale_β for the β producer at trade-close. Producer (eventual): extended `reward_subsystem_controller_kernel.cu` (7-weight output, anchor on `REWARD_AUX_ALIGN_EMA_INDEX=536`, target ~10% of total reward magnitude, bounds [0.05, 2.0]) — Phase B6, NOT yet landed. Consumers: `experience_env_step::segment_complete` (training-side β: `r_aux_align = scale_β × alignment × profit_pos`) and `backtest_env_kernel::segment_complete` (eval-side β with the same formula); HEALTH_DIAG `sp11_reward` printer (the `w_aux=…` column). FoldReset sentinel 0.5structural prior matching W's [-0.5, 0, +0.5, 0] init approach; activates β with fixed magnitude from epoch 0. The 0.5 prior was chosen over 0 (the original spec) because slot 537 has no producer yet (Phase B6 deferred); 0 would make β a runtime no-op. Once B6 lands and the controller emits adaptively per the bounds [0.05, 2.0], the prior is overwritten on first emit per `pearl_first_observation_bootstrap`; subsequent EMA-driven adaptation tracks the anchor signal.",
}, },
]; ];
Self { entries } Self { entries }

View File

@@ -10570,6 +10570,31 @@ impl DQNTrainer {
fused.trainer().write_isv_signal_at(LOW_DD_RATIO_INDEX, 0.0); fused.trainer().write_isv_signal_at(LOW_DD_RATIO_INDEX, 0.0);
} }
} }
// SP22 H6 Phase 3 (2026-05-13) — aux→policy bypass routing.
// Both ISV slots reset at fold boundary. The aux-align EMA
// sentinel is the canonical 0 per pearl_first_observation_bootstrap
// (first non-zero β emission replaces directly). The scale_β
// sentinel is 0.5 — a structural prior matching W's
// [-0.5, 0, +0.5, 0] init pattern; activates β with a fixed
// magnitude from epoch 0 until the SP11 controller's Phase B6
// extension to 7 components lands and overwrites this with an
// adaptive emit. Without the 0.5 prior, β would be a runtime
// no-op because slot 537 has no producer yet.
"sp22_reward_aux_align_ema" => {
if let Some(ref fused) = self.fused_ctx {
use crate::cuda_pipeline::sp22_isv_slots::REWARD_AUX_ALIGN_EMA_INDEX;
fused.trainer().write_isv_signal_at(REWARD_AUX_ALIGN_EMA_INDEX, 0.0);
}
}
"sp22_aux_align_scale" => {
if let Some(ref fused) = self.fused_ctx {
use crate::cuda_pipeline::sp22_isv_slots::SP22_AUX_ALIGN_SCALE_INDEX;
// Structural prior 0.5 — activates β with fixed magnitude
// from epoch 0. SP11 controller's Phase B6 extension will
// later overwrite with adaptive emit anchored on slot 536.
fused.trainer().write_isv_signal_at(SP22_AUX_ALIGN_SCALE_INDEX, 0.5);
}
}
_ => { _ => {
return Err(crate::MLError::ModelError(format!( return Err(crate::MLError::ModelError(format!(
"StateResetRegistry reset dispatch: unknown name '{}'. \ "StateResetRegistry reset dispatch: unknown name '{}'. \

View File

@@ -17319,3 +17319,17 @@ Estimated ~1-2 hr engineering (single-kernel edit + launcher update if N_COMPONE
| WR ~46%, dist shifts | Mechanism active but no WR translation | β probably necessary — land B6 next | | WR ~46%, dist shifts | Mechanism active but no WR translation | β probably necessary — land B6 next |
| WR ~46%, no dist shift | α inactive (state_121 not flowing) | Investigate aux head + state_121 wiring | | WR ~46%, no dist shift | α inactive (state_121 not flowing) | Investigate aux head + state_121 wiring |
| Training collapse | Loss landscape unstable | Reduce W init magnitude or revert | | Training collapse | Loss landscape unstable | Reduce W init magnitude or revert |
#### Phase B6-minimal — β activation via 0.5 structural prior (2026-05-13)
**Bug**: `state_reset_registry.rs` has FoldReset entries for `sp22_reward_aux_align_ema` and `sp22_aux_align_scale`, but `reset_named_state` dispatch in `training_loop.rs` had no matching arms → "unknown name" error at first fold boundary that touches these entries. Latent in `5106e3b117` because the current 1-fold smoke (train-qsltr) doesn't trigger fold-reset; any `--folds 2+` run would fail.
**Fix**: Added dispatch arms for both names. Plus elevated `scale_β` cold-start from 0 to 0.5 — a structural prior matching W's `[-0.5, 0, +0.5, 0]` init approach. Activates β reward shaping with fixed magnitude from epoch 0 without waiting for Phase B6's full SP11 controller extension.
**Effect**:
- aux_align_ema (slot 536) stays at 0 sentinel → producer's first observation replaces directly per `pearl_first_observation_bootstrap`.
- scale_β (slot 537) cold-starts at 0.5 → β producer fires `r_aux_align = 0.5 × alignment × profit_pos` from epoch 0 whenever a profitable aux-aligned trade closes.
**Once full B6 lands** (SP11 controller emits adaptive scale_β per bounds [0.05, 2.0]), the 0.5 prior is overwritten on first emit per `pearl_first_observation_bootstrap`.
**Status update**: the previous "α ACTIVE β DEAD" scope clarification is superseded. After this commit, both α and β are active when the smoke re-runs at this commit. The current train-qsltr smoke (at parent commit 5106e3b117) still has β-dead since it predates this fix.