Each EGF pearl EMA / state slot resets to its Pearl-A sentinel at fold
boundary, mirroring sp13_aux_dir_acc_short_ema / long_ema entries.
Atomic refactor (feedback_no_partial_refactor): both halves land
together — registry entry + reset_named_state dispatch arm.
Reset slots (11 total, sentinel in parens):
- Q_DISAGREEMENT_SHORT/LONG_EMA (slots 383, 384) → 0.5
- K_AUX_ADAPTIVE (385) → K_BASE_AUX = 20.0
- K_Q_ADAPTIVE (386) → K_BASE_Q = 15.0
- BETA_RATE_LIMITER_ADAPTIVE (387) → BETA_BASE = 0.5
- AUX_DIR_ACC_VARIANCE_EMA, Q_DISAGREEMENT_VARIANCE_EMA,
ALPHA_GRAD_RAW_VARIANCE_EMA (388, 389, 390) → 0.0
(initial k = k_base, β = β_base via ISV-driven controllers)
- GATE1_OPEN_STATE (391) → 0.0 (closed)
- ALPHA_GRAD_SMOOTHED (393) → 0.0
- AUX_DIR_ACC_POST_OPEN_MIN (394) → 1.0 (no min observed)
ALPHA_GRAD_RAW (slot 392, recomputed every step from variance EMAs)
and GRADIENT_HACK_LOCKOUT_REMAINING (slot 395, decays at epoch
boundary) are NOT in the fold-reset registry; both naturally
re-initialise without explicit reset.
Also corrects the isv-slots.md SP14 table: slots 392 and 395 were
incorrectly marked FoldReset in the B.1 entry; corrected to reflect
their actual reset semantics (NOT reset / epoch-boundary decay).
Producer + consumer wiring lands in subsequent tasks (B.3-B.12);
this commit is additive infrastructure only — no behavior change.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
393 lines
38 KiB
Markdown
393 lines
38 KiB
Markdown
# ISV Slot Allocation Registry
|
||
|
||
**Source of truth for ISV bus slot allocations.** Every slot has a named constant in `gpu_dqn_trainer.rs`. This doc is the cross-reference.
|
||
|
||
**Design: layout fingerprint, not schema version.** ISV[115..117) carries a
|
||
compile-time structural hash of the slot layout. Checkpoint load is
|
||
fail-fast; no migration functions exist. Backward compat is structurally
|
||
unwritable (there is no ordered version space to pair migrations against).
|
||
See spec §4.A.2.
|
||
|
||
**Tail placement rationale:** The fingerprint occupies the current tail (`[115..117)`) rather
|
||
than the head because `isv_signals[0]` and `isv_signals[1]` are actively written
|
||
by `isv_signal_update` (Q-drift EMA and gradient-norm EMA respectively). Inserting
|
||
at the head would displace those live signals and require shifting every upstream
|
||
literal in `experience_kernels.cu`. The fingerprint moves to the new tail each time
|
||
new slots are appended to the bus.
|
||
|
||
**Current `ISV_TOTAL_DIM`:** 360 (post-SP11 Task A0 reservation extending the bus from 340 → 360 with 20 contiguous reward-subsystem-controller slots at `[340..360)`; SP11 A0 slots are reservation-only and zero-initialized until A1/A2/Layer-B land producers and atomic consumer migration — see the SP11 section at the bottom of this file). Earlier expansions: SP4 Task A1 (131 → 171, 40 slots at `[131..171)`); SP5 Task A0 (173 → 286, 110 slots; producer wave landed across A1–A8 + Layer D); SP6/SP7/SP8/SP9/SP10 incremental allocations onto the SP5 block (286 → 340). Producers and consumers for slots `[0..131)` are wired across Plan 1, Plan 2 (Task 1 C.1, Task 3 D.2 per-branch gamma, Task 6C D.8 TLOB), Plan 3 (Task 1 C.2 reward-component EMAs, Task 3 B.2 trade-attempt novelty, Task 4 B.4 readiness-EMA, Task 7 C.3 state-distribution KL, Task 8 B.3 GPU-only seed warm-start, Task 9 C.5 CQL α ramp), Plan 4 (Task 5 Mode A E.5 attention-focus EMAs, follow-up target-drift EMAs, Task 6 C.A multi-task aux head EMAs, Task 6 / Plan 5 follow-up label-scale EMA, MoE expert-utilisation + gate-entropy EMAs + adaptive-lambda controller, Q-drift-rate diagnostic, fold warmup factor).
|
||
|
||
| Index | Name constant | Type | Producer | Consumers | Reset-category | Notes |
|
||
|---|---|---|---|---|---|---|
|
||
| [0] | `SLOT_0_Q_DRIFT` (seed only) | f32 | isv_signal_update | ISV encoder | FoldReset | `isv_signals[0]` = (q_mean − q_ema) / max(|q_ema|, 0.1) |
|
||
| [1] | `SLOT_1_GRAD_NORM_EMA` (seed only) | f32 | isv_signal_update | ISV encoder | FoldReset | EMA of sqrt(grad_norm²) |
|
||
| [2] | `SLOT_2_TD_ERR_EMA` (seed only) | f32 | isv_signal_update | ISV encoder | FoldReset | EMA of TD-error scalar |
|
||
| [3] | `SLOT_3_ENS_VAR_EMA` (seed only) | f32 | isv_signal_update | ISV encoder | FoldReset | C51 Q-distribution variance EMA (batch-mean atom-spread) |
|
||
| [4] | `SLOT_4_ENS_VAR_VEL` (seed only) | f32 | isv_signal_update | ISV encoder | FoldReset | Delta of slot 3 (variance velocity) |
|
||
| [5] | `SLOT_5_REWARD_EMA` (seed only) | f32 | isv_signal_update | ISV encoder | FoldReset | EMA of per-batch mean reward |
|
||
| [6] | `SLOT_6_ATOM_UTIL_EMA` (seed only) | f32 | isv_signal_update | ISV encoder | FoldReset | EMA of atom utilization fraction |
|
||
| [7] | `SLOT_7_LOSS_EMA` (seed only) | f32 | isv_signal_update | ISV encoder | FoldReset | EMA of total training loss |
|
||
| [8] | `SLOT_8_ADX_EMA` (seed only) | f32 | isv_signal_update | ISV encoder | FoldReset | Batch-mean ADX EMA (regime velocity indicator) |
|
||
| [9] | `SLOT_9_REGIME_DISAGREE` (seed only) | f32 | isv_signal_update | ISV encoder | FoldReset | |norm(ADX) − norm(CUSUM)| disagreement signal |
|
||
| [10] | `SLOT_10_REGIME_VEL_EMA` (seed only) | f32 | isv_signal_update | ISV encoder | FoldReset | EMA of regime transition velocity (ADX + CUSUM delta) |
|
||
| [11] | `SLOT_11_REGIME_STABILITY` (seed only) | f32 | isv_signal_update | ISV encoder | FoldReset | 1 − sigmoid(5 × regime_vel); high = stable regime |
|
||
| [12] | `LEARNING_HEALTH_INDEX` | f32 | isv_signal_update | many | FoldReset | health score ∈ [0, 1] |
|
||
| [13..17) | `Q_MAG_MEAN_*_INDEX` | f32 | q_mag_means_reduce | c51 kernels | FoldReset | Quarter/Half/Full mag Q-mean EMAs + \|Q\| ref |
|
||
| [17..22) | `Q_DIR_MEAN_*_INDEX` | f32 | q_dir_means_reduce | c51 kernels | FoldReset | Short/Hold/Long/Flat dir Q-mean EMAs + \|Q\| ref |
|
||
| [22] | `SHARPE_EMA_INDEX` | f32 | training_loop host | isv_signal_update | FoldReset | Training Sharpe EMA |
|
||
| [23..31) | `V_{CENTER,HALF}_{DIR,MAG,ORD,URG}_INDEX` | f32 | update_eval_v_range | adaptive_atoms, warm_start | FoldReset | Per-branch Q-support |
|
||
| [31..35) | `GRAD_NORM_TARGET_*_INDEX` | f32 | grad_balance_isv_update | branch_grad_rescale | SoftReset(decay_bars=500) | Per-branch grad-norm target |
|
||
| [35] | `GRAD_SCALE_LIMIT_INDEX` | f32 | grad_balance_isv_update | branch_grad_rescale | SoftReset(decay_bars=500) | Scale clamp limit |
|
||
| [36] | `IQL_BRANCH_SCALE_FLOOR_INDEX` | f32 | construct (static) | iql_per_branch_advantage | SchemaContract | Per-sample branch_scales floor; safety bound |
|
||
| [37..39) | (gap — fingerprint moved to tail) | — | — | — | — | Previously held the fingerprint pair; promoted across the Plan 1 C.6 / Plan 3 / Plan 4 / Plan 5 expansions to its current location at [115..117). Slots unused; zero-filled. |
|
||
| [39] | `EPOCH_IDX_INDEX` | f32 (int cast) | CPU epoch-loop (follow-up task) | GPU adaptive kernels | FoldReset | Current epoch index; 0 at construction, CPU writes at each epoch boundary |
|
||
| [40] | `TOTAL_EPOCHS_INDEX` | f32 (int cast) | construct (CPU static) | GPU adaptive kernels | SchemaContract | Total epoch count for this run; written at construction from `config.total_epochs` |
|
||
| [41] | `EPSILON_EFF_INDEX` | f32 | GPU epsilon-adaptive kernel (follow-up) | epsilon-greedy action select | FoldReset | Effective epsilon; 0 at construction; GPU fills each step |
|
||
| [42] | `TAU_EFF_INDEX` | f32 | GPU tau-adaptive kernel (follow-up) | target-net Polyak update | FoldReset | Effective tau; 0 at construction; GPU fills each step |
|
||
| [43] | `GAMMA_DIR_EFF_INDEX` | f32 | GPU per_branch_gamma_update kernel (Plan 2 D.2) | c51_loss_batched, iql_compute_per_sample_support | FoldReset | Direction-branch gamma; base=0.92, max=0.99; GPU fills each epoch boundary |
|
||
| [44] | `GAMMA_MAG_EFF_INDEX` | f32 | GPU per_branch_gamma_update kernel (Plan 2 D.2) | c51_loss_batched, iql_compute_per_sample_support | FoldReset | Magnitude-branch gamma; base=0.88, max=0.95 |
|
||
| [45] | `GAMMA_ORD_EFF_INDEX` | f32 | GPU per_branch_gamma_update kernel (Plan 2 D.2) | c51_loss_batched, iql_compute_per_sample_support | FoldReset | Order-branch gamma; base=0.85, max=0.93 |
|
||
| [46] | `GAMMA_URG_EFF_INDEX` | f32 | GPU per_branch_gamma_update kernel (Plan 2 D.2) | c51_loss_batched, iql_compute_per_sample_support | FoldReset | Urgency-branch gamma; base=0.80, max=0.90 |
|
||
| [47] | `KELLY_CAP_EFF_INDEX` | f32 | GPU kelly-cap-adaptive kernel (Plan 1 Task 11) | experience_env_step | FoldReset | Effective Kelly cap; 0 at construction; GPU fills each step |
|
||
| [48] | `CQL_ALPHA_INDEX` | f32 | construct (CPU static) | CQL adaptive formula in Rust | SchemaContract | CQL pessimism base coefficient; written from `config.cql_alpha`; read in `compute_cql_logit_gradients` |
|
||
| [49] | `PLAN_THRESHOLD_INDEX` | f32 | GPU `plan_threshold_update` kernel (Plan 3 B.4) | `experience_kernels.cu`, `backtest_plan_kernel.cu` | FoldReset | Plan-MLP activation threshold; producer upgraded from static constructor write to GPU kernel by Plan 3 Task 4 B.4. Cold-start 0.5f preserved. Kernel writes `max(0.1, 0.5 × ISV[READINESS_EMA_INDEX=75])` each epoch. Consumers (4 sites in `experience_kernels.cu` + 1 in `backtest_plan_kernel.cu`) unchanged — still read via `ISV_PLAN_THRESHOLD_IDX`. |
|
||
| [50] | `Q_P05_DIR_INDEX` | f32 | q_quantile_reduce GPU kernel (cold-path per-epoch) | update_eval_v_range | FoldReset | Direction-branch 5th-percentile Q EMA. Bootstrap: v_min. |
|
||
| [51] | `Q_P05_MAG_INDEX` | f32 | q_quantile_reduce | update_eval_v_range | FoldReset | Magnitude-branch 5th-percentile Q EMA. Bootstrap: v_min. |
|
||
| [52] | `Q_P05_ORD_INDEX` | f32 | q_quantile_reduce | update_eval_v_range | FoldReset | Order-branch 5th-percentile Q EMA. Bootstrap: v_min. |
|
||
| [53] | `Q_P05_URG_INDEX` | f32 | q_quantile_reduce | update_eval_v_range | FoldReset | Urgency-branch 5th-percentile Q EMA. Bootstrap: v_min. |
|
||
| [54] | `Q_P95_DIR_INDEX` | f32 | q_quantile_reduce GPU kernel (cold-path per-epoch) | update_eval_v_range | FoldReset | Direction-branch 95th-percentile Q EMA. Bootstrap: v_max. |
|
||
| [55] | `Q_P95_MAG_INDEX` | f32 | q_quantile_reduce | update_eval_v_range | FoldReset | Magnitude-branch 95th-percentile Q EMA. Bootstrap: v_max. |
|
||
| [56] | `Q_P95_ORD_INDEX` | f32 | q_quantile_reduce | update_eval_v_range | FoldReset | Order-branch 95th-percentile Q EMA. Bootstrap: v_max. |
|
||
| [57] | `Q_P95_URG_INDEX` | f32 | q_quantile_reduce | update_eval_v_range | FoldReset | Urgency-branch 95th-percentile Q EMA. Bootstrap: v_max. |
|
||
| [58..60) | (gap — fingerprint shifted to tail) | — | — | — | — | Previously [58..60); fingerprint promoted to [61..63) by Plan 2 Task 6C D.8 TLOB expansion. Slots unused; zero-filled. |
|
||
| [60] | `TLOB_REGIME_FOCUS_EMA_INDEX` | f32 | CPU training_loop (per-epoch) | ISV consumers (diagnostic) | FoldReset | Mean-max TLOB attention weight EMA (α=0.05); written at epoch boundary from `GpuTlob::mean_max_attention_weight`. Indicates OFI feature focus sharpness. |
|
||
| [61] | (gap — fingerprint shifted to tail) | — | — | — | — | Previously [61]; fingerprint promoted to [69..71) by Plan 3 Task 1 C.2 reward-EMA expansion. Slot unused; zero-filled. |
|
||
| [62] | (gap — fingerprint shifted to tail) | — | — | — | — | Previously [62]; see [61] note above. Slot unused; zero-filled. |
|
||
| [63] | `REWARD_POPART_EMA_INDEX` | f32 | GPU `reward_component_ema` kernel (Plan 3 C.2) | HEALTH_DIAG `reward_split`, `RewardComponentMonitor` | FoldReset | EMA of mean \|popart reward\| across batch (α=0.05). Final on-policy reward component. |
|
||
| [64] | `REWARD_CF_EMA_INDEX` | f32 | GPU `reward_component_ema` kernel (Plan 3 C.2) | HEALTH_DIAG `reward_split`, `RewardComponentMonitor` | FoldReset | EMA of mean \|counterfactual reward\| across batch (α=0.05). Zero until Plan 3 B.1. |
|
||
| [65] | `REWARD_TRAIL_EMA_INDEX` | f32 | GPU `reward_component_ema` kernel (Plan 3 C.2) | HEALTH_DIAG `reward_split`, `RewardComponentMonitor` | FoldReset | EMA of mean \|trail reward\| across batch (α=0.05). Structural placeholder; populated by Plan 3 B.2. |
|
||
| [66] | `REWARD_MICRO_EMA_INDEX` | f32 | GPU `reward_component_ema` kernel (Plan 3 C.2) | HEALTH_DIAG `reward_split`, `RewardComponentMonitor` | FoldReset | EMA of mean \|OFI micro-reward\| across batch (α=0.05). Populated via `rc[3]` in `experience_env_step`. |
|
||
| [67] | `REWARD_OPP_COST_EMA_INDEX` | f32 | GPU `reward_component_ema` kernel (Plan 3 C.2) | HEALTH_DIAG `reward_split`, `RewardComponentMonitor` | FoldReset | EMA of mean \|opportunity-cost reward\| across batch (α=0.05). Populated by Plan 3 B.1 at the Flat-branch opp-cost site. |
|
||
| [68] | `REWARD_BONUS_EMA_INDEX` | f32 | GPU `reward_component_ema` kernel (Plan 3 C.2) | HEALTH_DIAG `reward_split`, `RewardComponentMonitor` | FoldReset | EMA of mean \|bonus reward\| across batch (α=0.05). Populated by Plan 3 B.2 at the Flat→Positioned site (rc[5]). |
|
||
| [71] | `TRADE_ATTEMPT_RATE_EMA_INDEX` | f32 | GPU `trade_attempt_rate_ema_update` kernel (Plan 3 B.2) | `experience_env_step` Flat→Positioned bonus path | FoldReset | EMA of Flat→Positioned transition rate (0..1). Adaptive α = α_base × (1 + 0.5×\|clamp(sharpe, -2, 2)\|). α_base = 0.05. Zero at construction and fold boundary. |
|
||
| [72] | `TRADE_TARGET_RATE_INDEX` | f32 | CPU `training_loop` (epoch 5 freeze) | `experience_env_step` novelty computation | FoldReset | Reference rate for novelty = max(0, 1 - attempt/target). Frozen at epoch 5 from measured TRADE_ATTEMPT_RATE_EMA (min 0.001). Pre-freeze the slot is 0 and the bonus site gates on `target_raw > 1e-6f`, so the reward term is structurally inert until the freeze fires. |
|
||
| [73..76) | (gap — fingerprint shifted to tail) | — | — | — | — | Previously [73..75); fingerprint shifted by Plan 3 Task 4 B.4 to accommodate the new READINESS_EMA slot. Slots unused; zero-filled. |
|
||
| [75] | `READINESS_EMA_INDEX` | f32 | GPU `plan_threshold_update` kernel (Plan 3 B.4) | derived → ISV[PLAN_THRESHOLD_INDEX=49]; `PlanThresholdMonitor` | FoldReset | Per-batch mean readiness EMA (α=α_base × (1+0.5×\|clamp(sharpe,−2,2)\|), α_base=0.05). Cold-start 1.0 so derived threshold = 0.5 matches the previous static default. Drives slot 49 producer-only — consumers unchanged. |
|
||
| [78] | `STATE_KL_TRAIN_VAL_EMA_INDEX` | f32 | GPU `state_kl_moment_match` kernel (Plan 3 C.3) | `StateKLMonitor` (HEALTH_DIAG) | FoldReset | Per-validation-epoch moment-match KL EMA between train-state sample and val-state batch over the OFI block (32 dims). Adaptive α=α_base × (1+0.5×\|clamp(sharpe,−2,2)\|), α_base=0.05. Cold-start 0.0 — first kernel fire EMAs the actual measured KL toward this. |
|
||
| [79] | `STATE_KL_AMPLIFICATION_INDEX` | f32 | GPU `state_kl_moment_match` kernel (Plan 3 C.3) | `experience_kernels.cu` B.1 opp_cost + B.2 bonus consumers (via `ISV_STATE_KL_AMP_IDX` macro) | FoldReset | Bounded amplifier ∈ [1, 2] tracking trailing-EMA-of-self KL ratio: `1 + clamp(new_ema/prev_ema − 1, 0, 1)`. Cold-start 1.0 — neutral element behind consumers' `fmaxf(1.0, …)` guard. When train/val distributions diverge, both Flat opp-cost and trade-attempt bonus scale up together. Per `pearl_one_unbounded_signal_per_reward.md`: amp is bounded so it stacks safely with B.1's `q_abs_ref` unbounded multiplicand. |
|
||
| [48] | `CQL_ALPHA_INDEX` (post-Plan-3-T9 producer upgrade) | f32 | GPU `cql_alpha_seed_update` kernel (Plan 3 C.5) | `compute_cql_logit_gradients` (Rust ISV read) | FoldReset (was SchemaContract) | CQL pessimism base coefficient. Producer upgraded SchemaContract → FoldReset by Plan 3 Task 9 C.5: per-epoch GPU kernel EMAs slot 48 toward `config.cql_alpha × max(0, 1 - ISV[SEED_FRAC_EMA_INDEX=84])`. During seed phase (frac=1), target=0 → CQL α decays to 0 (no pessimism on exploration data); as frac → 0, α ramps to `config.cql_alpha` (full pessimism on network-driven trajectories). Constructor cold-starts to `config.cql_alpha`; fold-boundary reset re-applies. |
|
||
| [82] | `SEED_STEPS_TARGET_INDEX` | f32 | construct (CPU static, from `config.replay_seed_steps`) | `seed_step_counter_update` GPU kernel (Plan 3 B.3) | SchemaContract | Plan 3 Task 8 B.3 — target number of seed-phase experience-collection samples. Default 100_000; smoke configs may override to 10_000 so the seed→network transition is observable inside the smoke run. |
|
||
| [83] | `SEED_STEPS_DONE_INDEX` | f32 | GPU `seed_step_counter_update` kernel (Plan 3 B.3) | training_loop CPU dispatch decision (per-epoch read) + `seed_step_counter_update` self | FoldReset | Plan 3 Task 8 B.3 — cumulative seed-phase steps completed. GPU-incremented by `n_samples_this_step` per `collect_experiences_gpu` call, capped at TARGET. Cold-start 0; fold-boundary reset re-applies. CPU per-epoch read of (DONE, TARGET) decides whether next collect dispatches scripted-policy kernel (DONE < TARGET) or `experience_action_select`. |
|
||
| [84] | `SEED_FRAC_EMA_INDEX` | f32 | GPU `seed_step_counter_update` kernel (Plan 3 B.3) | GPU `cql_alpha_seed_update` kernel (Plan 3 C.5) + `SeedMonitor` | FoldReset | Plan 3 Task 8 B.3 — adaptive EMA of `max(0, 1 - DONE/TARGET) ∈ [0, 1]`. Cold-start 1.0 (fully in seed phase) so Task 9's CQL ramp sees `target = final × max(0, 1 - 1.0) = 0` until the seed phase actually decays. Adaptive α = α_base × (1 + 0.5 × \|clamp(sharpe, ±2)\|), α_base = 0.05. |
|
||
| [87] | `VSN_MAG_EMA_INDEX` | f32 | GPU `attention_focus_ema_update` kernel (Plan 4 E.5 Mode A; rewritten 2026-04-25 for GPU-only reduction) | `AttentionMonitor` (HEALTH_DIAG); HEALTH_DIAG `noisy [vsn_mag=…]` reads via `read_isv_signal_at` | FoldReset | Magnitude-branch VSN-weight magnitude EMA. **GPU-computed** by block 0 of `attention_focus_ema_update` (256-thread smem reduction over the params buffer slice corresponding to VSN mag tensors). Adaptive α = α_base × (1 + 0.5 × \|clamp(sharpe, ±2)\|), α_base = 0.05. Cold-start 0.0. Replaces the legacy CPU-DtoH `per_branch_vsn_mean()` that was deleted 2026-04-25 per `pearl_cold_path_no_exception_to_gpu_drives.md`. |
|
||
| [88] | `VSN_DIR_EMA_INDEX` | f32 | GPU `attention_focus_ema_update` kernel block 1 | same | FoldReset | Direction-branch equivalent of [87]. Same producer/contract. |
|
||
| [89] | `MAMBA2_RETENTION_EMA_INDEX` | f32 | GPU `attention_focus_ema_update` kernel block 2 | same | FoldReset | Mamba2 state-transition magnitude EMA (retention proxy). **GPU-computed** by block 2's smem reduction over the live `mamba2_h_enriched` device buffer (no DtoH). Same EMA convention as [87]. Diagnostic only. |
|
||
| [92] | `TARGET_DRIFT_MAG_EMA_INDEX` | f32 | GPU `target_drift_ema_update` kernel block 0 (Plan 4 follow-up, 2026-04-25) | HEALTH_DIAG `noisy [drift_mag=…]` via `read_isv_signal_at` | FoldReset | RMS(target − online) of magnitude-branch params (tensors 12..16), EMA-tracked. **GPU-computed** by block 0 of `target_drift_ema_update` (256-thread smem reduction of squared diffs, sqrt at thread 0). Adaptive α = α_base × (1 + 0.5 × \|clamp(sharpe, ±2)\|), α_base = 0.05. Cold-start 0.0. Replaces the legacy CPU-DtoH `per_branch_target_drift()` that was deleted 2026-04-25 per `pearl_cold_path_no_exception_to_gpu_drives.md`. |
|
||
| [93] | `TARGET_DRIFT_DIR_EMA_INDEX` | f32 | GPU `target_drift_ema_update` kernel block 1 | HEALTH_DIAG `noisy [drift_dir=…]` | FoldReset | Direction-branch equivalent of [92] (params tensors 8..12). |
|
||
| [115] | `ISV_LAYOUT_FINGERPRINT_LO_INDEX` | u32 bits (in f32) | construct | check_layout_fingerprint | SchemaContract | Low 32 bits of u64 FNV-1a structural hash. Fail-fast on mismatch — NOT a version number, no migration path. Shifted 69→73 by Plan 3 Task 3 B.2, 73→76 by Plan 3 Task 4 B.4, 76→80 by Plan 3 Task 7 C.3, 80→85 by Plan 3 Task 8 B.3, 85→90 by Plan 4 Task 5 Mode A, 90→94 by Plan 4 follow-up (target-drift EMA replacing legacy CPU-DtoH path), and 94→115 over the Plan 4 Task 6 / Plan 5 / MoE expansions that landed slots `[96..115)`. |
|
||
| [116] | `ISV_LAYOUT_FINGERPRINT_HI_INDEX` | u32 bits (in f32) | construct | check_layout_fingerprint | SchemaContract | High 32 bits of u64 FNV-1a structural hash. Tracks the same shift sequence as [115]. |
|
||
| [117..131) | (post-fingerprint slots filled by Plan 4 Task 6 / Plan 5 / MoE) | | | | | Includes `AUX_LABEL_SCALE_EMA_INDEX=117`, MoE expert-utilisation EMAs `[118..126)`, MoE `GATE_ENTROPY_EMA=126`, adaptive `MOE_LAMBDA_EFF=128`, `Q_DRIFT_RATE=129`, `FOLD_WARMUP_FACTOR=130`. Cross-reference the named constants in `gpu_dqn_trainer.rs`. |
|
||
| [131..171) | SP4 reservation (Layer A, no producers wired yet) | | | | | See "SP4: Signal-driven magnitude bounds" section below. |
|
||
|
||
## SP4: Signal-driven magnitude bounds (Layer A reservation)
|
||
|
||
SP4 Task A1 extends `ISV_TOTAL_DIM` from 131 → 171 by reserving 40 contiguous
|
||
slots at indices `[131..171)` for the signal-driven bounds that replace
|
||
hardcoded magnitude multipliers in SP3 mechanisms. Constants live in
|
||
`crates/ml/src/cuda_pipeline/sp4_isv_slots.rs` and are re-exported from
|
||
`cuda_pipeline::mod`. **Layer A is reservation-only — no producer kernels or
|
||
consumers are wired yet. All 40 slots remain zero-initialized at construction
|
||
and behave as no-ops until subsequent SP4 tasks land producers.**
|
||
|
||
| Index | Name constant | Family | Purpose |
|
||
|---|---|---|---|
|
||
| [131] | `TARGET_Q_BOUND_INDEX` | scalar | p99(\|target_q\|) — Mech 1 clamp bound |
|
||
| [132..136) | `ATOM_POS_BOUND_BASE` + branch | per-branch | p99(\|atom_positions[branch]\|) — Mech 2 clamp |
|
||
| [136..144) | `WEIGHT_BOUND_BASE` + group | per-param-group | p99(\|params\|) — Mech 9 clamp |
|
||
| [144..152) | `ADAM_M_BOUND_BASE` + group | per-param-group | p99(\|adam_m\|) — Mech 5 diagnostic |
|
||
| [152..160) | `ADAM_V_BOUND_BASE` + group | per-param-group | p99(\|adam_v\|) — Mech 5 diagnostic |
|
||
| [160..168) | `WD_RATE_BASE` + group | per-param-group | EMA of \|w·g\|/\|\|w\|\|² — AdamW weight_decay arg |
|
||
| [168] | `GRAD_CLIP_BOUND_INDEX` | scalar | p99(grad_norm) — Mech 6 adaptive_clip upper bound |
|
||
| [169] | `H_S2_BOUND_INDEX` | scalar | p99(\|h_s2\|) — Mech 10 clamp bound |
|
||
| [170] | `L1_LAMBDA_TRUNK_INDEX` | scalar (group-0 only) | (mean\|g\|/mean\|w\|) × entropy_deficit — trunk L1 lambda |
|
||
|
||
Param-group ordering (`ParamGroup` enum, all 8 groups `[0..8)`):
|
||
DqnTrunk=0, DqnValue=1, DqnBranches=2, Iqn=3, IqlHigh=4, IqlLow=5, Attn=6,
|
||
Curiosity=7. Convenience `const fn` accessors `weight_bound(g)`,
|
||
`adam_m_bound(g)`, `adam_v_bound(g)`, `wd_rate(g)`, `atom_pos_bound(branch)`
|
||
return the absolute slot index for a given group/branch.
|
||
|
||
Producers and consumers will be wired in subsequent SP4 layer-B and layer-C
|
||
tasks; until then the 40 slots are zero-initialized and reserved.
|
||
|
||
## SP5: Per-branch + per-group adaptation layer (Task A0 reservation)
|
||
|
||
SP5 Task A0 extends `ISV_TOTAL_DIM` from 173 → 286 by reserving 110 slots at
|
||
`[174..278) ∪ [280..286)`. An intentional 2-slot carve-out gap at `[278..280)`
|
||
separates the per-fold block from the cross-fold-persistent Kelly block.
|
||
Constants live in `crates/ml/src/cuda_pipeline/sp5_isv_slots.rs`.
|
||
|
||
| Range | Base constant | Family | Pearl | Purpose |
|
||
|---|---|---|---|---|
|
||
| [174..178) | `ATOM_V_CENTER_BASE` | per-branch [4] | Pearl 1 | C51 atom distribution center |
|
||
| [178..182) | `ATOM_V_HALF_BASE` | per-branch [4] | Pearl 1 | C51 atom half-width |
|
||
| [182..186) | `ATOM_HEADROOM_BASE` | per-branch [4] | Pearl 1 | C51 atom headroom |
|
||
| [186..190) | `ATOM_CLIP_RATE_BASE` | per-branch [4] | Pearl 1 | C51 atom clip rate |
|
||
| [190..194) | `BUDGET_C51_BASE` | per-branch [4] | Pearl 2 | C51 loss budget weight. SP6 Pearl 2: `compute_adaptive_budgets()` reads individually, applies correction-factor sub-launches via `apply_c51_budget_scale_branch`. |
|
||
| [194..198) | `BUDGET_IQN_BASE` | per-branch [4] | Pearl 2 | IQN loss budget weight. SP6 Pearl 2: used as trunk-mean only (`iqn_trunk`) — IQN backward targets trunk params exclusively. |
|
||
| [198..202) | `BUDGET_CQL_BASE` | per-branch [4] | Pearl 2 | CQL loss budget weight. SP6 Pearl 2: `compute_adaptive_budgets()` reads individually, applies correction-factor sub-launches via `apply_cql_saxpy_branch`. |
|
||
| [202..206) | `BUDGET_ENS_BASE` | per-branch [4] | Pearl 2 | Ensemble loss budget weight. SP6 Pearl 2: used as trunk-mean only. |
|
||
| [206..210) | `FLATNESS_BASE` | per-branch [4] | Pearl 2 | Loss flatness diagnostic |
|
||
| [210..214) | `NOISY_SIGMA_BASE` | per-branch [4] | Pearl 3 | NoisyNet σ level — SP6 Pearl 3 consumer wired: `add_advantage_noise` kernel reads per-branch σ via mapped-pinned dev_ptr; `training_loop.rs` reads slots directly (no averaging) |
|
||
| [214..218) | `SIGMA_FRACTION_BASE` | per-branch [4] | Pearl 3 | NoisyNet σ fraction |
|
||
| [218..222) | `BRANCH_ENTROPY_BASE` | per-branch [4] | Pearl 3 | Branch action entropy |
|
||
| [222..226) | `Q_VAR_PER_BRANCH_BASE` | per-branch [4] | shared | Q-value variance |
|
||
| [226..234) | `ADAM_BETA1_BASE` | per-group [8] | Pearl 4 | Adam β1 per param group |
|
||
| [234..242) | `ADAM_BETA2_BASE` | per-group [8] | Pearl 4 | Adam β2 per param group |
|
||
| [242..250) | `ADAM_EPS_BASE` | per-group [8] | Pearl 4 | Adam ε per param group |
|
||
| [250..270) | `IQN_TAU_BASE` | per-branch×quantile [4×5] | Pearl 5 | IQN τ schedule |
|
||
| [270..274) | `TRAIL_DIST_PER_DIR_BASE` | per-direction [4] | Pearl 8 | Trail stop distance |
|
||
| [274..278) | `ATOM_NUM_ATOMS_BASE` | per-branch [4] | Pearl 1-ext | C51 atom count |
|
||
| [278..280) | — | gap | — | Intentional carve-out (not allocated) |
|
||
| [280] | `KELLY_F_SMOOTH_INDEX` | scalar | Pearl 6 | Kelly fraction EMA (cross-fold) |
|
||
| [281] | `CONVICTION_SMOOTH_INDEX` | scalar | Pearl 6 | Conviction EMA (cross-fold) |
|
||
| [282] | `TRADE_VAR_SMOOTH_INDEX` | scalar | Pearl 6 | Trade variance EMA (cross-fold) |
|
||
| [283] | `KELLY_SAMPLE_COUNT_INDEX` | scalar | Pearl 6 | Kelly sample count (cross-fold) |
|
||
| [284] | `WIN_RATE_SMOOTH_INDEX` | scalar | Pearl 6 | Win rate EMA (cross-fold) |
|
||
| [285] | `LOSS_RATE_SMOOTH_INDEX` | scalar | Pearl 6 | Loss rate EMA (cross-fold) |
|
||
|
||
Kelly slots `[280..286)` are NOT in the fold-reset registry. All other SP5 slots
|
||
are per-fold and reset at fold boundaries. Task A0 is reservation-only; producers
|
||
and consumers land in subsequent SP5 tasks.
|
||
|
||
## SP6 Pearl 5: IQN τ per-branch consumer wiring
|
||
|
||
SP6 Pearl 5 (commit in `sp6-pearl-5` worktree) upgrades the consumer side of
|
||
`ISV[250..270)` (`IQN_TAU_BASE`, 4 branches × 5 quantiles).
|
||
|
||
**Before SP6 (SP5 Layer B contract):** `fused_training.rs` read all 20 slots once
|
||
per epoch, averaged 4 branches per quantile to a single `[5]` tau array, and
|
||
uploaded it to the shared `online_taus/target_taus/cos_features` buffers.
|
||
One IQN forward pass per training step used the averaged schedule.
|
||
|
||
**After SP6 Pearl 5:** `GpuIqnHead` holds 12 additional `CudaSlice<f32>` buffers:
|
||
`online_taus_branch[4]`, `target_taus_branch[4]`, `cos_features_branch[4]`.
|
||
`refresh_taus_for_branch(b, tau5)` populates one branch slab per call — no
|
||
averaging. `activate_branch_taus(b)` / `deactivate_branch_taus(b)` swap the main
|
||
buffers for a per-branch pass via `mem::swap` (no allocation).
|
||
|
||
Training step: 4 sequential IQN forward passes, each with one branch's τ active.
|
||
Each pass calls `apply_iqn_trunk_gradient(iqn_budget / 4.0)` so total IQN
|
||
gradient contribution = `iqn_budget` (same as SP5 Layer B). The averaged
|
||
`refresh_taus_from_isv` is kept for CVaR and backward-compat consumers.
|
||
|
||
**Files changed:** `crates/ml/src/cuda_pipeline/gpu_iqn_head.rs`,
|
||
`crates/ml/src/trainers/dqn/fused_training.rs`.
|
||
|
||
## SP11: Reward as controlled subsystem (Tasks A0 + A1 + A2 + B0 + B1a + B1b)
|
||
|
||
SP11 Task A0 (Fix 39, 2026-05-04) extends `ISV_TOTAL_DIM` from 340 → 360 by
|
||
allocating 20 contiguous slots at `[340..360)` for the reward-subsystem
|
||
controller. Constants live in `crates/ml/src/cuda_pipeline/sp11_isv_slots.rs`
|
||
and are re-exported via `cuda_pipeline::sp5_isv_slots::*`.
|
||
|
||
**Task A1 (2026-05-04): three canary producer kernels added**, each
|
||
chained with `apply_pearls_ad_kernel` for Pearls A+D smoothing. Slots
|
||
[350..360) now have producers populating them every epoch (val_sharpe
|
||
emit cadence for the Δ canary; per-epoch metrics block for mag-ratio +
|
||
saboteur engagement).
|
||
|
||
**Task A2 (2026-05-04): controller producer + SimHash novelty buffer**.
|
||
The main `reward_subsystem_controller_kernel` reads all 5 canary slots
|
||
[350..360), runs the spec §3.4 control law (true Z-score → sigmoid blend
|
||
of winner/diversifier weights, post-floor renormalize Σ=1, permanent-
|
||
floor curiosity, post-clamp saboteur), and writes 10 outputs to scratch
|
||
which a chained `apply_pearls_ad_kernel` (n_slots=10) smooths into the
|
||
remaining slots [340..350). After A2 **all 20 SP11 slots populate every
|
||
step**.
|
||
|
||
A2 also lands the SimHash novelty infrastructure for B1's replay-time
|
||
curiosity bonus: `novelty_simhash_proj_init_kernel` (one-shot Philox-
|
||
driven init at trainer construct, populates a 42×16 ±1 projection
|
||
matrix on-device per `feedback_no_cpu_forwards.md`),
|
||
`novelty_simhash_kernel` (lookup + update, race-tolerated update per
|
||
`feedback_no_atomicadd.md` — under-counts bias novelty UPWARD, the safe
|
||
direction). The 1M-slot hash table at `GpuDqnTrainer.novelty_hash_buf`
|
||
gets a fold-reset registry entry (`sp11_novelty_hash`) that closes the
|
||
A0 deferral; the projection matrix is frozen for the run lifetime and
|
||
intentionally has NO registry entry.
|
||
|
||
**Layer A was additive (A0/A1/A2). Layer B atomic consumer migration:**
|
||
|
||
- **B0 (commit `302992f63`):** controller renormalisation flipped from
|
||
`Σ=1` to `mean=1` (Σ=N=6) per spec §3.4.3 amendment so per-bar
|
||
`w_active × r_active` ≈ pre-SP11 absolute scale on average. Per-
|
||
component cap MAX_WEIGHT=3.0 (Invariant-1, prevents winner-take-all).
|
||
- **B1a (commit `d5e1214f2`):** saboteur GPU multiplication via
|
||
`SABOTEUR_INTENSITY_MULT_INDEX` and SimHash novelty-buffer
|
||
`state_stride` correctness for B1c curiosity wiring.
|
||
- **B1b (this commit):** structural reward decomposition in
|
||
`experience_env_step` — 8+ inline accumulation sites replaced with
|
||
explicit per-component locals (`r_popart`, `r_cf`, `r_trail`,
|
||
`r_micro`, `r_opp_cost`, `r_bonus`) composed as `Σ w_i × r_i` with
|
||
controller weights from ISV[340..346). Universal post-composition
|
||
modifiers (drawdown / capital-floor / inventory / churn /
|
||
conviction-scale / cf-flip) apply unweighted per spec §3.4.4. Trail
|
||
P&L extraction (§3.5.4) — forced-exit P&L now flows through
|
||
`r_trail` (rc[2]); voluntary-exit through `r_popart`. cf-tuple
|
||
reward at `out_rewards[cf_off]` is now `w_cf × r_cf` (the
|
||
`cf_weight=0.3f` constants in `mse_loss_kernel.cu:318` /
|
||
`c51_loss_kernel.cu:789` are STRUCTURAL Q-blend weights, NOT reward
|
||
weights — left UNTOUCHED per §3.5 amendment). Sentinel-defense:
|
||
`fmaxf(w_raw, 0.01)` covers the cold-start ordering gap (the
|
||
controller runs at end-of-epoch, so step 0 of fold 0 reads sentinel
|
||
0; defense pins minimum scale at the Invariant-1 hard floor).
|
||
- **B1c (pending):** replay-time curiosity bonus per §3.5.5 (Layer C
|
||
audit gate — `rewards_buf` single-read-site verification).
|
||
|
||
A1 producers:
|
||
- `val_sharpe_delta_compute_kernel.cu` — two-pass: raw Δ + squared
|
||
deviation against prior `VAL_SHARPE_DELTA_EMA_INDEX`. Reads
|
||
mapped-pinned 2-element val_sharpe_history; chain → ISV[350, 351].
|
||
- `saboteur_engagement_compute_kernel.cu` — block tree-reduce over
|
||
per-bar `|Δreward|` populated by `experience_env_step` at the
|
||
saboteur perturbation site (proxy: `traded × |reward| ×
|
||
max(|eff_spread − 1|, |eff_slip − 1|)`); threshold = 0.01 ×
|
||
ISV[PNL_REWARD_MAGNITUDE_EMA_INDEX]. Chain → ISV[358].
|
||
- `reward_component_mag_ratio_compute_kernel.cu` — reads 6 EMAs at
|
||
ISV[REWARD_POPART_EMA_INDEX..+6), normalises to ratios, mirrors
|
||
popart magnitude into scratch[6] as a side-output. Chain (n_slots=6)
|
||
→ ISV[352..358); chain (n_slots=1) → ISV[359].
|
||
|
||
6 GPU oracle tests in `crates/ml/tests/sp11_producer_unit_tests.rs`
|
||
cover (A1) first-observation behaviour, threshold classification, and
|
||
normalisation invariants on the canary kernels; plus (A2) the
|
||
controller midpoint-at-z=0 invariant, weight-renormalization-after-floor
|
||
invariant, and saboteur-post-clamp-at-extreme-regression invariant.
|
||
All MappedF32Buffer fixtures (zero
|
||
htod_copy/dtoh_sync_copy/alloc_zeros per
|
||
`feedback_no_htod_htoh_only_mapped_pinned`).
|
||
|
||
| Range | Name constant | Family | Purpose |
|
||
|---|---|---|---|
|
||
| [340] | `REWARD_POPART_WEIGHT_INDEX` | scalar | Component weight: PopArt reward |
|
||
| [341] | `REWARD_CF_WEIGHT_INDEX` | scalar | Component weight: counterfactual (replaces hardcoded `cf_weight=0.3f`) |
|
||
| [342] | `REWARD_TRAIL_WEIGHT_INDEX` | scalar | Component weight: trail-stop |
|
||
| [343] | `REWARD_MICRO_WEIGHT_INDEX` | scalar | Component weight: microstructure |
|
||
| [344] | `REWARD_OPP_COST_WEIGHT_INDEX` | scalar | Component weight: opportunity cost |
|
||
| [345] | `REWARD_BONUS_WEIGHT_INDEX` | scalar | Component weight: exploration bonus |
|
||
| [346] | `CURIOSITY_PRESSURE_INDEX` | scalar | Controller output: curiosity scale (replay-time bonus) |
|
||
| [347] | `SABOTEUR_INTENSITY_MULT_INDEX` | scalar | Controller output: saboteur intensity multiplier ∈ [0.5, 2.0] |
|
||
| [348] | `REWARD_WEIGHT_FLOOR_INDEX` | scalar | Controller output: minimum-weight floor |
|
||
| [349] | `CURIOSITY_BOUND_INDEX` | scalar | Controller output: curiosity hard cap (= 0.3 × pnl_magnitude_ema) |
|
||
| [350] | `VAL_SHARPE_DELTA_EMA_INDEX` | scalar canary | val-sharpe trend EMA (improvement_z numerator) |
|
||
| [351] | `VAL_SHARPE_VAR_EMA_INDEX` | scalar canary | val-sharpe variance EMA (improvement_z denominator) |
|
||
| [352..358) | `REWARD_COMPONENT_MAG_RATIO_BASE` (+6) | per-component canary | Per-component reward-magnitude ratio (winner/diversifier blend input) |
|
||
| [358] | `SABOTEUR_ENGAGEMENT_RATE_INDEX` | scalar canary | Per-bar saboteur engagement rate EMA |
|
||
| [359] | `PNL_REWARD_MAGNITUDE_EMA_INDEX` | scalar canary | `|reward_total|` EMA (curiosity-bound scale factor) |
|
||
| [360] | `POPART_COMPONENT_MAG_EMA_INDEX` | scalar canary | popart-component-specific magnitude EMA (B1b fix-up; mag-ratio canary popart axis input) |
|
||
| [361..367) | `REWARD_COMPONENT_VAR_EMA_BASE` (+6) | per-component canary | Per-reward-component variance EMA (B1b smoke-recovery z-score normalisation; popart var at 361, then cf/trail/micro/opp_cost/bonus at 362..367) |
|
||
|
||
Invariant-1 fraction constants (rate-limiter anchors per spec §3.4.2):
|
||
`SP11_EPS_DIV=1e-6`, `SP11_WEIGHT_HARD_FLOOR=0.01`,
|
||
`SP11_SABOTEUR_MIN=0.5`, `SP11_SABOTEUR_MAX=2.0`,
|
||
`SP11_CURIOSITY_PERMANENT_FRACTION=0.2`, `SP11_CURIOSITY_BOUND_FRACTION=0.3`,
|
||
`SP11_WEIGHT_FLOOR_FRACTION=0.5`, `SP11_ENGAGEMENT_FLOOR=0.1`. Plus
|
||
`SP11_PROJ_SEED_SALT=0x_5511_0001` for the A2 novelty-SimHash projection-init
|
||
launcher's seed derivation.
|
||
|
||
**B1b fix-up (2026-05-04, spec §4 amendment):** ISV[360] = `POPART_COMPONENT_MAG_EMA_INDEX`
|
||
added to resolve the slot-63 PopArt-input vs popart-component-magnitude
|
||
overload that B1b smoke surfaced. Pre-fix-up the SP11 mag-ratio canary
|
||
read 6 contiguous slots starting at slot 63 — but slot 63 was overloaded
|
||
(PopArt's normalisation input AND popart-component magnitude collapsed to
|
||
the same value because reward composition was inline accumulation). B1b's
|
||
structural decomposition surfaced the contamination — controller
|
||
redistributed weight toward popart (w_pop ≈ 2.0) based on contaminated
|
||
ratio → 10× sharpe drop. Post-fix-up the mag-ratio kernel takes two slot
|
||
indices: `popart_specific_slot=360` for the popart axis,
|
||
`cf_others_base_slot=64` for cf/trail/micro/opp_cost/bonus. Slot 63 is
|
||
unchanged (PopArt's input = total reward magnitude, pre-SP11 invariant
|
||
preserved); slot 360 is the new dedicated popart-component magnitude
|
||
produced by `popart_component_ema_kernel`. ISV total: 360 → 361.
|
||
|
||
**B1b smoke-recovery (2026-05-04, spec §4 amendment "Why z-score" lines
|
||
564-619):** ISV[361..367) = `REWARD_COMPONENT_VAR_EMA_BASE..+6` added
|
||
to drive z-score normalisation in `reward_component_mag_ratio_compute_kernel`.
|
||
Pre-z-score the linear `winner_weight = mag_ratio` formula amplified
|
||
popart's intrinsic O(100) magnitude over the other 5 components'
|
||
O(0.1-2). Result: w_pop saturated toward MAX_WEIGHT, curiosity_b
|
||
exploded, sharpe collapsed in B1b smoke `smoke-test-6wd2c` on commit
|
||
`61b2fa962`. Z-score makes ratios scale-invariant —
|
||
z[c] = mag[c] / max(sqrt(var[c]), EPS_DIV); ratio[c] = z[c] / Σz.
|
||
Variance EMAs are produced via Welford's online algorithm in the
|
||
extended `popart_component_ema_kernel` (slot 361) and
|
||
`reward_component_ema_kernel` (slots 362..367); the canary now takes
|
||
4 slot indices (popart mag/var split + cf-others mag/var bases). The
|
||
mirror at scratch[6] still emits popart RAW magnitude (slot 359 input
|
||
for curiosity bound + saboteur engagement which need physical scale).
|
||
ISV total: 361 → 367.
|
||
|
||
All 21 slots are FoldReset (sentinel 0; Pearl A bootstraps from first
|
||
producer launch on each fold per `pearl_first_observation_bootstrap.md`).
|
||
After A2 all 20 slots have producers; B1b wires the on-policy reward
|
||
composition consumer for [340..346) (the 6 component weights).
|
||
B1c will wire the curiosity consumer for [346] + [349] (replay-time).
|
||
[347] saboteur multiplier already consumed in B1a's
|
||
gpu_experience_collector site. [348] `REWARD_WEIGHT_FLOOR_INDEX` is
|
||
self-consumed by the controller on the next step. [350..360) canaries
|
||
are read by the controller (intra-cycle).
|
||
|
||
The novelty-hash device buffer (`GpuDqnTrainer.novelty_hash_buf`,
|
||
1M slots × f32 mapped-pinned) lands in A2 alongside its registry entry
|
||
(`sp11_novelty_hash`) and dispatch arm. The 42×16 SimHash projection
|
||
matrix (`GpuDqnTrainer.novelty_simhash_proj`) is populated on-device by
|
||
`launch_novelty_simhash_proj_init` at trainer construct time (Philox-
|
||
driven from `SP11_PROJ_SEED_SALT`); it is **frozen for the run
|
||
lifetime** and intentionally has NO registry entry — it is the hash
|
||
function, not state.
|
||
|
||
**Spec:** `docs/superpowers/specs/2026-05-04-sp11-reward-as-controlled-subsystem.md`
|
||
**Plan:** `docs/superpowers/plans/2026-05-04-sp11-reward-as-controlled-subsystem.md`
|
||
|
||
---
|
||
|
||
## SP14 — Aux→Q Wire + Earned Gradient Flow [383..396)
|
||
|
||
**13 slots** allocated in `sp14_isv_slots.rs` (B.1, 2026-05-05).
|
||
Slots [381..383) were already claimed by SP13 closeout
|
||
(`HOLD_RATE_TARGET_INDEX=381`, `HOLD_RATE_OBSERVED_EMA_INDEX=382`), so
|
||
SP14 starts at 383 (shifted +2 from the original plan which documented
|
||
[381..394)).
|
||
|
||
| Slot | Constant | Reset | Notes |
|
||
|------|----------|-------|-------|
|
||
| 383 | `Q_DISAGREEMENT_SHORT_EMA_INDEX` | FoldReset (0.5) | K=4↔K=2 argmax mismatch fast EMA |
|
||
| 384 | `Q_DISAGREEMENT_LONG_EMA_INDEX` | FoldReset (0.5) | K=4↔K=2 argmax mismatch slow EMA |
|
||
| 385 | `K_AUX_ADAPTIVE_INDEX` | FoldReset (0.0→K_BASE_AUX) | Sigmoid steepness for Gate 1 (aux competence) |
|
||
| 386 | `K_Q_ADAPTIVE_INDEX` | FoldReset (0.0→K_BASE_Q) | Sigmoid steepness for Gate 2 (Q-head disagreement) |
|
||
| 387 | `BETA_RATE_LIMITER_ADAPTIVE_INDEX` | FoldReset (0.0→BETA_BASE) | Rate-limiter β (variance-driven) |
|
||
| 388 | `AUX_DIR_ACC_VARIANCE_EMA_INDEX` | FoldReset (0.0) | Welford variance for k_aux adaptation |
|
||
| 389 | `Q_DISAGREEMENT_VARIANCE_EMA_INDEX` | FoldReset (0.0) | Welford variance for k_q adaptation |
|
||
| 390 | `ALPHA_GRAD_RAW_VARIANCE_EMA_INDEX` | FoldReset (0.0) | Welford variance for β adaptation |
|
||
| 391 | `GATE1_OPEN_STATE_INDEX` | FoldReset (0.0) | Schmitt-trigger state (0=closed, 1=open) |
|
||
| 392 | `ALPHA_GRAD_RAW_INDEX` | NOT reset (recomputed every step) | Raw gate output (HEALTH_DIAG visibility) |
|
||
| 393 | `ALPHA_GRAD_SMOOTHED_INDEX` | FoldReset (0.0) | Rate-limited gate consumed by backward path |
|
||
| 394 | `AUX_DIR_ACC_POST_OPEN_MIN_INDEX` | FoldReset (1.0) | Post-open aux_dir_acc minimum (circuit breaker) |
|
||
| 395 | `GRADIENT_HACK_LOCKOUT_REMAINING_INDEX` | NOT reset (epoch-boundary decay) | Lockout epochs remaining (anti-gradient-hacking) |
|
||
|
||
Slots 396-398 are conceptual buffer; no constants allocated.
|
||
|
||
11 of the 13 slots are FoldReset per `pearl_first_observation_bootstrap.md`.
|
||
Slots 392 (ALPHA_GRAD_RAW) and 395 (GRADIENT_HACK_LOCKOUT_REMAINING) are
|
||
intentionally excluded from fold-reset: raw is recomputed every step from
|
||
the three Welford variance EMAs; lockout decays at epoch boundary and must
|
||
not be force-reset mid-fold.
|
||
|
||
Producer + consumer wiring lands in B.3-B.12; B.1 is additive
|
||
infrastructure (slot constants), B.2 wires the 11 fold-reset registry
|
||
entries + dispatch arms — both are additive only (no behavior change).
|
||
|
||
**Spec:** `docs/superpowers/specs/2026-05-05-sp14-aux-q-wire-earned-gradient-flow.md`
|
||
**Plan:** `docs/superpowers/plans/2026-05-05-sp14-aux-q-wire-earned-gradient-flow.md`
|