Files
foxhunt/docs/isv-slots.md
jgrusewski 2feeeda8bb fix(phase-e-4-a): GPU kernel for h_enriched slot copy — eliminate per-step CPU roundtrip
The Phase E.4.A T8 wiring stored Mamba2's per-step cache.h_enriched
into h_enriched_buf_dev via a dtoh+htod sequence:

  let h_host = stream.clone_dtoh(cache.h_enriched.cuda_data())?;
  let mut buf_host = stream.clone_dtoh(&h_enriched_buf_dev)?;  // <- whole buffer
  for j in 0..hidden_dim { buf_host[slot_offset + j] = h_host[j]; }
  stream.memcpy_htod(&buf_host, &mut h_enriched_buf_dev)?;     // <- whole buffer

This violates feedback_cpu_is_read_only AND
feedback_no_htod_htoh_only_mapped_pinned. Worse, the buffer-wide
dtoh+htod every step is ~20K floats × 600 steps × 500 eps × 30 cells
= ~9M roundtrips totaling significant PCIe latency in the backtest.

Fix: new tiny CUDA kernel alpha_h_enriched_store_kernel in
alpha_window_push.cu (one thread per hidden-dim feature, writes
src[j] → buf[slot_offset + j]). Replaces the dtoh/htod sequence
in both smoke and backtest binaries.

Estimated speed-up at backtest scale: 3-6× on the temporal eval
path. Pure-GPU per-step inference restored — no synchronisation
points on the hot path.

docs/isv-slots.md updated per kernel-audit-doc hook.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-15 22:20:28 +02:00

73 KiB
Raw Blame History

ISV Slot Allocation Registry

Source of truth for ISV bus slot allocations. Every slot has a named constant in gpu_dqn_trainer.rs. This doc is the cross-reference.

Design: layout fingerprint, not schema version. ISV[115..117) carries a compile-time structural hash of the slot layout. Checkpoint load is fail-fast; no migration functions exist. Backward compat is structurally unwritable (there is no ordered version space to pair migrations against). See spec §4.A.2.

Tail placement rationale: The fingerprint occupies the current tail ([115..117)) rather than the head because isv_signals[0] and isv_signals[1] are actively written by isv_signal_update (Q-drift EMA and gradient-norm EMA respectively). Inserting at the head would displace those live signals and require shifting every upstream literal in experience_kernels.cu. The fingerprint moves to the new tail each time new slots are appended to the bus.

Current ISV_TOTAL_DIM: 360 (post-SP11 Task A0 reservation extending the bus from 340 → 360 with 20 contiguous reward-subsystem-controller slots at [340..360); SP11 A0 slots are reservation-only and zero-initialized until A1/A2/Layer-B land producers and atomic consumer migration — see the SP11 section at the bottom of this file). Earlier expansions: SP4 Task A1 (131 → 171, 40 slots at [131..171)); SP5 Task A0 (173 → 286, 110 slots; producer wave landed across A1A8 + Layer D); SP6/SP7/SP8/SP9/SP10 incremental allocations onto the SP5 block (286 → 340). Producers and consumers for slots [0..131) are wired across Plan 1, Plan 2 (Task 1 C.1, Task 3 D.2 per-branch gamma, Task 6C D.8 TLOB), Plan 3 (Task 1 C.2 reward-component EMAs, Task 3 B.2 trade-attempt novelty, Task 4 B.4 readiness-EMA, Task 7 C.3 state-distribution KL, Task 8 B.3 GPU-only seed warm-start, Task 9 C.5 CQL α ramp), Plan 4 (Task 5 Mode A E.5 attention-focus EMAs, follow-up target-drift EMAs, Task 6 C.A multi-task aux head EMAs, Task 6 / Plan 5 follow-up label-scale EMA, MoE expert-utilisation + gate-entropy EMAs + adaptive-lambda controller, Q-drift-rate diagnostic, fold warmup factor).

Index Name constant Type Producer Consumers Reset-category Notes
[0] SLOT_0_Q_DRIFT (seed only) f32 isv_signal_update ISV encoder FoldReset isv_signals[0] = (q_mean q_ema) / max(
[1] SLOT_1_GRAD_NORM_EMA (seed only) f32 isv_signal_update ISV encoder FoldReset EMA of sqrt(grad_norm²)
[2] SLOT_2_TD_ERR_EMA (seed only) f32 isv_signal_update ISV encoder FoldReset EMA of TD-error scalar
[3] SLOT_3_ENS_VAR_EMA (seed only) f32 isv_signal_update ISV encoder FoldReset C51 Q-distribution variance EMA (batch-mean atom-spread)
[4] SLOT_4_ENS_VAR_VEL (seed only) f32 isv_signal_update ISV encoder FoldReset Delta of slot 3 (variance velocity)
[5] SLOT_5_REWARD_EMA (seed only) f32 isv_signal_update ISV encoder FoldReset EMA of per-batch mean reward
[6] SLOT_6_ATOM_UTIL_EMA (seed only) f32 isv_signal_update ISV encoder FoldReset EMA of atom utilization fraction
[7] SLOT_7_LOSS_EMA (seed only) f32 isv_signal_update ISV encoder FoldReset EMA of total training loss
[8] SLOT_8_ADX_EMA (seed only) f32 isv_signal_update ISV encoder FoldReset Batch-mean ADX EMA (regime velocity indicator)
[9] SLOT_9_REGIME_DISAGREE (seed only) f32 isv_signal_update ISV encoder FoldReset
[10] SLOT_10_REGIME_VEL_EMA (seed only) f32 isv_signal_update ISV encoder FoldReset EMA of regime transition velocity (ADX + CUSUM delta)
[11] SLOT_11_REGIME_STABILITY (seed only) f32 isv_signal_update ISV encoder FoldReset 1 sigmoid(5 × regime_vel); high = stable regime
[12] LEARNING_HEALTH_INDEX f32 isv_signal_update many FoldReset health score ∈ [0, 1]
[13..17) Q_MAG_MEAN_*_INDEX f32 q_mag_means_reduce c51 kernels FoldReset Quarter/Half/Full mag Q-mean EMAs + |Q| ref
[17..22) Q_DIR_MEAN_*_INDEX f32 q_dir_means_reduce c51 kernels FoldReset Short/Hold/Long/Flat dir Q-mean EMAs + |Q| ref
[22] SHARPE_EMA_INDEX f32 training_loop host isv_signal_update FoldReset Training Sharpe EMA
[23..31) V_{CENTER,HALF}_{DIR,MAG,ORD,URG}_INDEX f32 update_eval_v_range adaptive_atoms, warm_start FoldReset Per-branch Q-support
[31..35) GRAD_NORM_TARGET_*_INDEX f32 grad_balance_isv_update branch_grad_rescale SoftReset(decay_bars=500) Per-branch grad-norm target
[35] GRAD_SCALE_LIMIT_INDEX f32 grad_balance_isv_update branch_grad_rescale SoftReset(decay_bars=500) Scale clamp limit
[36] IQL_BRANCH_SCALE_FLOOR_INDEX f32 construct (static) iql_per_branch_advantage SchemaContract Per-sample branch_scales floor; safety bound
[37..39) (gap — fingerprint moved to tail) Previously held the fingerprint pair; promoted across the Plan 1 C.6 / Plan 3 / Plan 4 / Plan 5 expansions to its current location at [115..117). Slots unused; zero-filled.
[39] EPOCH_IDX_INDEX f32 (int cast) CPU epoch-loop (follow-up task) GPU adaptive kernels FoldReset Current epoch index; 0 at construction, CPU writes at each epoch boundary
[40] TOTAL_EPOCHS_INDEX f32 (int cast) construct (CPU static) GPU adaptive kernels SchemaContract Total epoch count for this run; written at construction from config.total_epochs
[41] EPSILON_EFF_INDEX f32 GPU epsilon-adaptive kernel (follow-up) epsilon-greedy action select FoldReset Effective epsilon; 0 at construction; GPU fills each step
[42] TAU_EFF_INDEX f32 GPU tau-adaptive kernel (follow-up) target-net Polyak update FoldReset Effective tau; 0 at construction; GPU fills each step
[43] GAMMA_DIR_EFF_INDEX f32 GPU per_branch_gamma_update kernel (Plan 2 D.2) c51_loss_batched, iql_compute_per_sample_support FoldReset Direction-branch gamma; base=0.92, max=0.99; GPU fills each epoch boundary
[44] GAMMA_MAG_EFF_INDEX f32 GPU per_branch_gamma_update kernel (Plan 2 D.2) c51_loss_batched, iql_compute_per_sample_support FoldReset Magnitude-branch gamma; base=0.88, max=0.95
[45] GAMMA_ORD_EFF_INDEX f32 GPU per_branch_gamma_update kernel (Plan 2 D.2) c51_loss_batched, iql_compute_per_sample_support FoldReset Order-branch gamma; base=0.85, max=0.93
[46] GAMMA_URG_EFF_INDEX f32 GPU per_branch_gamma_update kernel (Plan 2 D.2) c51_loss_batched, iql_compute_per_sample_support FoldReset Urgency-branch gamma; base=0.80, max=0.90
[47] KELLY_CAP_EFF_INDEX f32 GPU kelly-cap-adaptive kernel (Plan 1 Task 11) experience_env_step FoldReset Effective Kelly cap; 0 at construction; GPU fills each step
[48] CQL_ALPHA_INDEX f32 construct (CPU static) CQL adaptive formula in Rust SchemaContract CQL pessimism base coefficient; written from config.cql_alpha; read in compute_cql_logit_gradients
[49] PLAN_THRESHOLD_INDEX f32 GPU plan_threshold_update kernel (Plan 3 B.4) experience_kernels.cu, backtest_plan_kernel.cu FoldReset Plan-MLP activation threshold; producer upgraded from static constructor write to GPU kernel by Plan 3 Task 4 B.4. Cold-start 0.5f preserved. Kernel writes max(0.1, 0.5 × ISV[READINESS_EMA_INDEX=75]) each epoch. Consumers (4 sites in experience_kernels.cu + 1 in backtest_plan_kernel.cu) unchanged — still read via ISV_PLAN_THRESHOLD_IDX.
[50] Q_P05_DIR_INDEX f32 q_quantile_reduce GPU kernel (cold-path per-epoch) update_eval_v_range FoldReset Direction-branch 5th-percentile Q EMA. Bootstrap: v_min.
[51] Q_P05_MAG_INDEX f32 q_quantile_reduce update_eval_v_range FoldReset Magnitude-branch 5th-percentile Q EMA. Bootstrap: v_min.
[52] Q_P05_ORD_INDEX f32 q_quantile_reduce update_eval_v_range FoldReset Order-branch 5th-percentile Q EMA. Bootstrap: v_min.
[53] Q_P05_URG_INDEX f32 q_quantile_reduce update_eval_v_range FoldReset Urgency-branch 5th-percentile Q EMA. Bootstrap: v_min.
[54] Q_P95_DIR_INDEX f32 q_quantile_reduce GPU kernel (cold-path per-epoch) update_eval_v_range FoldReset Direction-branch 95th-percentile Q EMA. Bootstrap: v_max.
[55] Q_P95_MAG_INDEX f32 q_quantile_reduce update_eval_v_range FoldReset Magnitude-branch 95th-percentile Q EMA. Bootstrap: v_max.
[56] Q_P95_ORD_INDEX f32 q_quantile_reduce update_eval_v_range FoldReset Order-branch 95th-percentile Q EMA. Bootstrap: v_max.
[57] Q_P95_URG_INDEX f32 q_quantile_reduce update_eval_v_range FoldReset Urgency-branch 95th-percentile Q EMA. Bootstrap: v_max.
[58..60) (gap — fingerprint shifted to tail) Previously [58..60); fingerprint promoted to [61..63) by Plan 2 Task 6C D.8 TLOB expansion. Slots unused; zero-filled.
[60] TLOB_REGIME_FOCUS_EMA_INDEX f32 CPU training_loop (per-epoch) ISV consumers (diagnostic) FoldReset Mean-max TLOB attention weight EMA (α=0.05); written at epoch boundary from GpuTlob::mean_max_attention_weight. Indicates OFI feature focus sharpness.
[61] (gap — fingerprint shifted to tail) Previously [61]; fingerprint promoted to [69..71) by Plan 3 Task 1 C.2 reward-EMA expansion. Slot unused; zero-filled.
[62] (gap — fingerprint shifted to tail) Previously [62]; see [61] note above. Slot unused; zero-filled.
[63] REWARD_POPART_EMA_INDEX f32 GPU reward_component_ema kernel (Plan 3 C.2) HEALTH_DIAG reward_split, RewardComponentMonitor FoldReset EMA of mean |popart reward| across batch (α=0.05). Final on-policy reward component.
[64] REWARD_CF_EMA_INDEX f32 GPU reward_component_ema kernel (Plan 3 C.2) HEALTH_DIAG reward_split, RewardComponentMonitor FoldReset EMA of mean |counterfactual reward| across batch (α=0.05). Zero until Plan 3 B.1.
[65] REWARD_TRAIL_EMA_INDEX f32 GPU reward_component_ema kernel (Plan 3 C.2) HEALTH_DIAG reward_split, RewardComponentMonitor FoldReset EMA of mean |trail reward| across batch (α=0.05). Structural placeholder; populated by Plan 3 B.2.
[66] REWARD_MICRO_EMA_INDEX f32 GPU reward_component_ema kernel (Plan 3 C.2) HEALTH_DIAG reward_split, RewardComponentMonitor FoldReset EMA of mean |OFI micro-reward| across batch (α=0.05). Populated via rc[3] in experience_env_step.
[67] REWARD_OPP_COST_EMA_INDEX f32 GPU reward_component_ema kernel (Plan 3 C.2) HEALTH_DIAG reward_split, RewardComponentMonitor FoldReset EMA of mean |opportunity-cost reward| across batch (α=0.05). Populated by Plan 3 B.1 at the Flat-branch opp-cost site.
[68] REWARD_BONUS_EMA_INDEX f32 GPU reward_component_ema kernel (Plan 3 C.2) HEALTH_DIAG reward_split, RewardComponentMonitor FoldReset EMA of mean |bonus reward| across batch (α=0.05). Populated by Plan 3 B.2 at the Flat→Positioned site (rc[5]).
[71] TRADE_ATTEMPT_RATE_EMA_INDEX f32 GPU trade_attempt_rate_ema_update kernel (Plan 3 B.2) experience_env_step Flat→Positioned bonus path FoldReset EMA of Flat→Positioned transition rate (0..1). Adaptive α = α_base × (1 + 0.5×|clamp(sharpe, -2, 2)|). α_base = 0.05. Zero at construction and fold boundary.
[72] TRADE_TARGET_RATE_INDEX f32 CPU training_loop (epoch 5 freeze) experience_env_step novelty computation FoldReset Reference rate for novelty = max(0, 1 - attempt/target). Frozen at epoch 5 from measured TRADE_ATTEMPT_RATE_EMA (min 0.001). Pre-freeze the slot is 0 and the bonus site gates on target_raw > 1e-6f, so the reward term is structurally inert until the freeze fires.
[73..76) (gap — fingerprint shifted to tail) Previously [73..75); fingerprint shifted by Plan 3 Task 4 B.4 to accommodate the new READINESS_EMA slot. Slots unused; zero-filled.
[75] READINESS_EMA_INDEX f32 GPU plan_threshold_update kernel (Plan 3 B.4) derived → ISV[PLAN_THRESHOLD_INDEX=49]; PlanThresholdMonitor FoldReset Per-batch mean readiness EMA (α=α_base × (1+0.5×|clamp(sharpe,2,2)|), α_base=0.05). Cold-start 1.0 so derived threshold = 0.5 matches the previous static default. Drives slot 49 producer-only — consumers unchanged.
[78] STATE_KL_TRAIN_VAL_EMA_INDEX f32 GPU state_kl_moment_match kernel (Plan 3 C.3) StateKLMonitor (HEALTH_DIAG) FoldReset Per-validation-epoch moment-match KL EMA between train-state sample and val-state batch over the OFI block (32 dims). Adaptive α=α_base × (1+0.5×|clamp(sharpe,2,2)|), α_base=0.05. Cold-start 0.0 — first kernel fire EMAs the actual measured KL toward this.
[79] STATE_KL_AMPLIFICATION_INDEX f32 GPU state_kl_moment_match kernel (Plan 3 C.3) experience_kernels.cu B.1 opp_cost + B.2 bonus consumers (via ISV_STATE_KL_AMP_IDX macro) FoldReset Bounded amplifier ∈ [1, 2] tracking trailing-EMA-of-self KL ratio: 1 + clamp(new_ema/prev_ema 1, 0, 1). Cold-start 1.0 — neutral element behind consumers' fmaxf(1.0, …) guard. When train/val distributions diverge, both Flat opp-cost and trade-attempt bonus scale up together. Per pearl_one_unbounded_signal_per_reward.md: amp is bounded so it stacks safely with B.1's q_abs_ref unbounded multiplicand.
[48] CQL_ALPHA_INDEX (post-Plan-3-T9 producer upgrade) f32 GPU cql_alpha_seed_update kernel (Plan 3 C.5) compute_cql_logit_gradients (Rust ISV read) FoldReset (was SchemaContract) CQL pessimism base coefficient. Producer upgraded SchemaContract → FoldReset by Plan 3 Task 9 C.5: per-epoch GPU kernel EMAs slot 48 toward config.cql_alpha × max(0, 1 - ISV[SEED_FRAC_EMA_INDEX=84]). During seed phase (frac=1), target=0 → CQL α decays to 0 (no pessimism on exploration data); as frac → 0, α ramps to config.cql_alpha (full pessimism on network-driven trajectories). Constructor cold-starts to config.cql_alpha; fold-boundary reset re-applies.
[82] SEED_STEPS_TARGET_INDEX f32 construct (CPU static, from config.replay_seed_steps) seed_step_counter_update GPU kernel (Plan 3 B.3) SchemaContract Plan 3 Task 8 B.3 — target number of seed-phase experience-collection samples. Default 100_000; smoke configs may override to 10_000 so the seed→network transition is observable inside the smoke run.
[83] SEED_STEPS_DONE_INDEX f32 GPU seed_step_counter_update kernel (Plan 3 B.3) training_loop CPU dispatch decision (per-epoch read) + seed_step_counter_update self FoldReset Plan 3 Task 8 B.3 — cumulative seed-phase steps completed. GPU-incremented by n_samples_this_step per collect_experiences_gpu call, capped at TARGET. Cold-start 0; fold-boundary reset re-applies. CPU per-epoch read of (DONE, TARGET) decides whether next collect dispatches scripted-policy kernel (DONE < TARGET) or experience_action_select.
[84] SEED_FRAC_EMA_INDEX f32 GPU seed_step_counter_update kernel (Plan 3 B.3) GPU cql_alpha_seed_update kernel (Plan 3 C.5) + SeedMonitor FoldReset Plan 3 Task 8 B.3 — adaptive EMA of max(0, 1 - DONE/TARGET) ∈ [0, 1]. Cold-start 1.0 (fully in seed phase) so Task 9's CQL ramp sees target = final × max(0, 1 - 1.0) = 0 until the seed phase actually decays. Adaptive α = α_base × (1 + 0.5 × |clamp(sharpe, ±2)|), α_base = 0.05.
[87] VSN_MAG_EMA_INDEX f32 GPU attention_focus_ema_update kernel (Plan 4 E.5 Mode A; rewritten 2026-04-25 for GPU-only reduction) AttentionMonitor (HEALTH_DIAG); HEALTH_DIAG noisy [vsn_mag=…] reads via read_isv_signal_at FoldReset Magnitude-branch VSN-weight magnitude EMA. GPU-computed by block 0 of attention_focus_ema_update (256-thread smem reduction over the params buffer slice corresponding to VSN mag tensors). Adaptive α = α_base × (1 + 0.5 × |clamp(sharpe, ±2)|), α_base = 0.05. Cold-start 0.0. Replaces the legacy CPU-DtoH per_branch_vsn_mean() that was deleted 2026-04-25 per pearl_cold_path_no_exception_to_gpu_drives.md.
[88] VSN_DIR_EMA_INDEX f32 GPU attention_focus_ema_update kernel block 1 same FoldReset Direction-branch equivalent of [87]. Same producer/contract.
[89] MAMBA2_RETENTION_EMA_INDEX f32 GPU attention_focus_ema_update kernel block 2 same FoldReset Mamba2 state-transition magnitude EMA (retention proxy). GPU-computed by block 2's smem reduction over the live mamba2_h_enriched device buffer (no DtoH). Same EMA convention as [87]. Diagnostic only.
[92] TARGET_DRIFT_MAG_EMA_INDEX f32 GPU target_drift_ema_update kernel block 0 (Plan 4 follow-up, 2026-04-25) HEALTH_DIAG noisy [drift_mag=…] via read_isv_signal_at FoldReset RMS(target online) of magnitude-branch params (tensors 12..16), EMA-tracked. GPU-computed by block 0 of target_drift_ema_update (256-thread smem reduction of squared diffs, sqrt at thread 0). Adaptive α = α_base × (1 + 0.5 × |clamp(sharpe, ±2)|), α_base = 0.05. Cold-start 0.0. Replaces the legacy CPU-DtoH per_branch_target_drift() that was deleted 2026-04-25 per pearl_cold_path_no_exception_to_gpu_drives.md.
[93] TARGET_DRIFT_DIR_EMA_INDEX f32 GPU target_drift_ema_update kernel block 1 HEALTH_DIAG noisy [drift_dir=…] FoldReset Direction-branch equivalent of [92] (params tensors 8..12).
[115] ISV_LAYOUT_FINGERPRINT_LO_INDEX u32 bits (in f32) construct check_layout_fingerprint SchemaContract Low 32 bits of u64 FNV-1a structural hash. Fail-fast on mismatch — NOT a version number, no migration path. Shifted 69→73 by Plan 3 Task 3 B.2, 73→76 by Plan 3 Task 4 B.4, 76→80 by Plan 3 Task 7 C.3, 80→85 by Plan 3 Task 8 B.3, 85→90 by Plan 4 Task 5 Mode A, 90→94 by Plan 4 follow-up (target-drift EMA replacing legacy CPU-DtoH path), and 94→115 over the Plan 4 Task 6 / Plan 5 / MoE expansions that landed slots [96..115).
[116] ISV_LAYOUT_FINGERPRINT_HI_INDEX u32 bits (in f32) construct check_layout_fingerprint SchemaContract High 32 bits of u64 FNV-1a structural hash. Tracks the same shift sequence as [115].
[117..131) (post-fingerprint slots filled by Plan 4 Task 6 / Plan 5 / MoE) Includes AUX_LABEL_SCALE_EMA_INDEX=117, MoE expert-utilisation EMAs [118..126), MoE GATE_ENTROPY_EMA=126, adaptive MOE_LAMBDA_EFF=128, Q_DRIFT_RATE=129, FOLD_WARMUP_FACTOR=130. Cross-reference the named constants in gpu_dqn_trainer.rs.
[131..171) SP4 reservation (Layer A, no producers wired yet) See "SP4: Signal-driven magnitude bounds" section below.

SP4: Signal-driven magnitude bounds (Layer A reservation)

SP4 Task A1 extends ISV_TOTAL_DIM from 131 → 171 by reserving 40 contiguous slots at indices [131..171) for the signal-driven bounds that replace hardcoded magnitude multipliers in SP3 mechanisms. Constants live in crates/ml/src/cuda_pipeline/sp4_isv_slots.rs and are re-exported from cuda_pipeline::mod. Layer A is reservation-only — no producer kernels or consumers are wired yet. All 40 slots remain zero-initialized at construction and behave as no-ops until subsequent SP4 tasks land producers.

Index Name constant Family Purpose
[131] TARGET_Q_BOUND_INDEX scalar p99(|target_q|) — Mech 1 clamp bound
[132..136) ATOM_POS_BOUND_BASE + branch per-branch p99(|atom_positions[branch]|) — Mech 2 clamp
[136..144) WEIGHT_BOUND_BASE + group per-param-group p99(|params|) — Mech 9 clamp
[144..152) ADAM_M_BOUND_BASE + group per-param-group p99(|adam_m|) — Mech 5 diagnostic
[152..160) ADAM_V_BOUND_BASE + group per-param-group p99(|adam_v|) — Mech 5 diagnostic
[160..168) WD_RATE_BASE + group per-param-group EMA of |w·g|/||w||² — AdamW weight_decay arg
[168] GRAD_CLIP_BOUND_INDEX scalar p99(grad_norm) — Mech 6 adaptive_clip upper bound
[169] H_S2_BOUND_INDEX scalar p99(|h_s2|) — Mech 10 clamp bound
[170] L1_LAMBDA_TRUNK_INDEX scalar (group-0 only) (mean|g|/mean|w|) × entropy_deficit — trunk L1 lambda

Param-group ordering (ParamGroup enum, all 8 groups [0..8)): DqnTrunk=0, DqnValue=1, DqnBranches=2, Iqn=3, IqlHigh=4, IqlLow=5, Attn=6, Curiosity=7. Convenience const fn accessors weight_bound(g), adam_m_bound(g), adam_v_bound(g), wd_rate(g), atom_pos_bound(branch) return the absolute slot index for a given group/branch.

Producers and consumers will be wired in subsequent SP4 layer-B and layer-C tasks; until then the 40 slots are zero-initialized and reserved.

SP5: Per-branch + per-group adaptation layer (Task A0 reservation)

SP5 Task A0 extends ISV_TOTAL_DIM from 173 → 286 by reserving 110 slots at [174..278) [280..286). An intentional 2-slot carve-out gap at [278..280) separates the per-fold block from the cross-fold-persistent Kelly block. Constants live in crates/ml/src/cuda_pipeline/sp5_isv_slots.rs.

Range Base constant Family Pearl Purpose
[174..178) ATOM_V_CENTER_BASE per-branch [4] Pearl 1 C51 atom distribution center
[178..182) ATOM_V_HALF_BASE per-branch [4] Pearl 1 C51 atom half-width
[182..186) ATOM_HEADROOM_BASE per-branch [4] Pearl 1 C51 atom headroom
[186..190) ATOM_CLIP_RATE_BASE per-branch [4] Pearl 1 C51 atom clip rate
[190..194) BUDGET_C51_BASE per-branch [4] Pearl 2 C51 loss budget weight. SP6 Pearl 2: compute_adaptive_budgets() reads individually, applies correction-factor sub-launches via apply_c51_budget_scale_branch.
[194..198) BUDGET_IQN_BASE per-branch [4] Pearl 2 IQN loss budget weight. SP6 Pearl 2: used as trunk-mean only (iqn_trunk) — IQN backward targets trunk params exclusively.
[198..202) BUDGET_CQL_BASE per-branch [4] Pearl 2 CQL loss budget weight. SP6 Pearl 2: compute_adaptive_budgets() reads individually, applies correction-factor sub-launches via apply_cql_saxpy_branch.
[202..206) BUDGET_ENS_BASE per-branch [4] Pearl 2 Ensemble loss budget weight. SP6 Pearl 2: used as trunk-mean only.
[206..210) FLATNESS_BASE per-branch [4] Pearl 2 Loss flatness diagnostic
[210..214) NOISY_SIGMA_BASE per-branch [4] Pearl 3 NoisyNet σ level — SP6 Pearl 3 consumer wired: add_advantage_noise kernel reads per-branch σ via mapped-pinned dev_ptr; training_loop.rs reads slots directly (no averaging)
[214..218) SIGMA_FRACTION_BASE per-branch [4] Pearl 3 NoisyNet σ fraction
[218..222) BRANCH_ENTROPY_BASE per-branch [4] Pearl 3 Branch action entropy
[222..226) Q_VAR_PER_BRANCH_BASE per-branch [4] shared Q-value variance
[226..234) ADAM_BETA1_BASE per-group [8] Pearl 4 Adam β1 per param group
[234..242) ADAM_BETA2_BASE per-group [8] Pearl 4 Adam β2 per param group
[242..250) ADAM_EPS_BASE per-group [8] Pearl 4 Adam ε per param group
[250..270) IQN_TAU_BASE per-branch×quantile [4×5] Pearl 5 IQN τ schedule
[270..274) TRAIL_DIST_PER_DIR_BASE per-direction [4] Pearl 8 Trail stop distance
[274..278) ATOM_NUM_ATOMS_BASE per-branch [4] Pearl 1-ext C51 atom count
[278..280) gap Intentional carve-out (not allocated)
[280] KELLY_F_SMOOTH_INDEX scalar Pearl 6 Kelly fraction EMA (cross-fold)
[281] CONVICTION_SMOOTH_INDEX scalar Pearl 6 Conviction EMA (cross-fold)
[282] TRADE_VAR_SMOOTH_INDEX scalar Pearl 6 Trade variance EMA (cross-fold)
[283] KELLY_SAMPLE_COUNT_INDEX scalar Pearl 6 Kelly sample count (cross-fold)
[284] WIN_RATE_SMOOTH_INDEX scalar Pearl 6 Win rate EMA (cross-fold)
[285] LOSS_RATE_SMOOTH_INDEX scalar Pearl 6 Loss rate EMA (cross-fold)

Kelly slots [280..286) are NOT in the fold-reset registry. All other SP5 slots are per-fold and reset at fold boundaries. Task A0 is reservation-only; producers and consumers land in subsequent SP5 tasks.

SP6 Pearl 5: IQN τ per-branch consumer wiring

SP6 Pearl 5 (commit in sp6-pearl-5 worktree) upgrades the consumer side of ISV[250..270) (IQN_TAU_BASE, 4 branches × 5 quantiles).

Before SP6 (SP5 Layer B contract): fused_training.rs read all 20 slots once per epoch, averaged 4 branches per quantile to a single [5] tau array, and uploaded it to the shared online_taus/target_taus/cos_features buffers. One IQN forward pass per training step used the averaged schedule.

After SP6 Pearl 5: GpuIqnHead holds 12 additional CudaSlice<f32> buffers: online_taus_branch[4], target_taus_branch[4], cos_features_branch[4]. refresh_taus_for_branch(b, tau5) populates one branch slab per call — no averaging. activate_branch_taus(b) / deactivate_branch_taus(b) swap the main buffers for a per-branch pass via mem::swap (no allocation).

Training step: 4 sequential IQN forward passes, each with one branch's τ active. Each pass calls apply_iqn_trunk_gradient(iqn_budget / 4.0) so total IQN gradient contribution = iqn_budget (same as SP5 Layer B). The averaged refresh_taus_from_isv is kept for CVaR and backward-compat consumers.

Files changed: crates/ml/src/cuda_pipeline/gpu_iqn_head.rs, crates/ml/src/trainers/dqn/fused_training.rs.

SP11: Reward as controlled subsystem (Tasks A0 + A1 + A2 + B0 + B1a + B1b)

SP11 Task A0 (Fix 39, 2026-05-04) extends ISV_TOTAL_DIM from 340 → 360 by allocating 20 contiguous slots at [340..360) for the reward-subsystem controller. Constants live in crates/ml/src/cuda_pipeline/sp11_isv_slots.rs and are re-exported via cuda_pipeline::sp5_isv_slots::*.

Task A1 (2026-05-04): three canary producer kernels added, each chained with apply_pearls_ad_kernel for Pearls A+D smoothing. Slots [350..360) now have producers populating them every epoch (val_sharpe emit cadence for the Δ canary; per-epoch metrics block for mag-ratio + saboteur engagement).

Task A2 (2026-05-04): controller producer + SimHash novelty buffer. The main reward_subsystem_controller_kernel reads all 5 canary slots [350..360), runs the spec §3.4 control law (true Z-score → sigmoid blend of winner/diversifier weights, post-floor renormalize Σ=1, permanent- floor curiosity, post-clamp saboteur), and writes 10 outputs to scratch which a chained apply_pearls_ad_kernel (n_slots=10) smooths into the remaining slots [340..350). After A2 all 20 SP11 slots populate every step.

A2 also lands the SimHash novelty infrastructure for B1's replay-time curiosity bonus: novelty_simhash_proj_init_kernel (one-shot Philox- driven init at trainer construct, populates a 42×16 ±1 projection matrix on-device per feedback_no_cpu_forwards.md), novelty_simhash_kernel (lookup + update, race-tolerated update per feedback_no_atomicadd.md — under-counts bias novelty UPWARD, the safe direction). The 1M-slot hash table at GpuDqnTrainer.novelty_hash_buf gets a fold-reset registry entry (sp11_novelty_hash) that closes the A0 deferral; the projection matrix is frozen for the run lifetime and intentionally has NO registry entry.

Layer A was additive (A0/A1/A2). Layer B atomic consumer migration:

  • B0 (commit 302992f63): controller renormalisation flipped from Σ=1 to mean=1 (Σ=N=6) per spec §3.4.3 amendment so per-bar w_active × r_active ≈ pre-SP11 absolute scale on average. Per- component cap MAX_WEIGHT=3.0 (Invariant-1, prevents winner-take-all).
  • B1a (commit d5e1214f2): saboteur GPU multiplication via SABOTEUR_INTENSITY_MULT_INDEX and SimHash novelty-buffer state_stride correctness for B1c curiosity wiring.
  • B1b (this commit): structural reward decomposition in experience_env_step — 8+ inline accumulation sites replaced with explicit per-component locals (r_popart, r_cf, r_trail, r_micro, r_opp_cost, r_bonus) composed as Σ w_i × r_i with controller weights from ISV[340..346). Universal post-composition modifiers (drawdown / capital-floor / inventory / churn / conviction-scale / cf-flip) apply unweighted per spec §3.4.4. Trail P&L extraction (§3.5.4) — forced-exit P&L now flows through r_trail (rc[2]); voluntary-exit through r_popart. cf-tuple reward at out_rewards[cf_off] is now w_cf × r_cf (the cf_weight=0.3f constants in mse_loss_kernel.cu:318 / c51_loss_kernel.cu:789 are STRUCTURAL Q-blend weights, NOT reward weights — left UNTOUCHED per §3.5 amendment). Sentinel-defense: fmaxf(w_raw, 0.01) covers the cold-start ordering gap (the controller runs at end-of-epoch, so step 0 of fold 0 reads sentinel 0; defense pins minimum scale at the Invariant-1 hard floor).
  • B1c (pending): replay-time curiosity bonus per §3.5.5 (Layer C audit gate — rewards_buf single-read-site verification).

A1 producers:

  • val_sharpe_delta_compute_kernel.cu — two-pass: raw Δ + squared deviation against prior VAL_SHARPE_DELTA_EMA_INDEX. Reads mapped-pinned 2-element val_sharpe_history; chain → ISV[350, 351].
  • saboteur_engagement_compute_kernel.cu — block tree-reduce over per-bar |Δreward| populated by experience_env_step at the saboteur perturbation site (proxy: traded × |reward| × max(|eff_spread 1|, |eff_slip 1|)); threshold = 0.01 × ISV[PNL_REWARD_MAGNITUDE_EMA_INDEX]. Chain → ISV[358].
  • reward_component_mag_ratio_compute_kernel.cu — reads 6 EMAs at ISV[REWARD_POPART_EMA_INDEX..+6), normalises to ratios, mirrors popart magnitude into scratch[6] as a side-output. Chain (n_slots=6) → ISV[352..358); chain (n_slots=1) → ISV[359].

6 GPU oracle tests in crates/ml/tests/sp11_producer_unit_tests.rs cover (A1) first-observation behaviour, threshold classification, and normalisation invariants on the canary kernels; plus (A2) the controller midpoint-at-z=0 invariant, weight-renormalization-after-floor invariant, and saboteur-post-clamp-at-extreme-regression invariant. All MappedF32Buffer fixtures (zero htod_copy/dtoh_sync_copy/alloc_zeros per feedback_no_htod_htoh_only_mapped_pinned).

Range Name constant Family Purpose
[340] REWARD_POPART_WEIGHT_INDEX scalar Component weight: PopArt reward
[341] REWARD_CF_WEIGHT_INDEX scalar Component weight: counterfactual (replaces hardcoded cf_weight=0.3f)
[342] REWARD_TRAIL_WEIGHT_INDEX scalar Component weight: trail-stop
[343] REWARD_MICRO_WEIGHT_INDEX scalar Component weight: microstructure
[344] REWARD_OPP_COST_WEIGHT_INDEX scalar Component weight: opportunity cost
[345] REWARD_BONUS_WEIGHT_INDEX scalar Component weight: exploration bonus
[346] CURIOSITY_PRESSURE_INDEX scalar Controller output: curiosity scale (replay-time bonus)
[347] SABOTEUR_INTENSITY_MULT_INDEX scalar Controller output: saboteur intensity multiplier ∈ [0.5, 2.0]
[348] REWARD_WEIGHT_FLOOR_INDEX scalar Controller output: minimum-weight floor
[349] CURIOSITY_BOUND_INDEX scalar Controller output: curiosity hard cap (= 0.3 × pnl_magnitude_ema)
[350] VAL_SHARPE_DELTA_EMA_INDEX scalar canary val-sharpe trend EMA (improvement_z numerator)
[351] VAL_SHARPE_VAR_EMA_INDEX scalar canary val-sharpe variance EMA (improvement_z denominator)
[352..358) REWARD_COMPONENT_MAG_RATIO_BASE (+6) per-component canary Per-component reward-magnitude ratio (winner/diversifier blend input)
[358] SABOTEUR_ENGAGEMENT_RATE_INDEX scalar canary Per-bar saboteur engagement rate EMA
[359] PNL_REWARD_MAGNITUDE_EMA_INDEX scalar canary `
[360] POPART_COMPONENT_MAG_EMA_INDEX scalar canary popart-component-specific magnitude EMA (B1b fix-up; mag-ratio canary popart axis input)
[361..367) REWARD_COMPONENT_VAR_EMA_BASE (+6) per-component canary Per-reward-component variance EMA (B1b smoke-recovery z-score normalisation; popart var at 361, then cf/trail/micro/opp_cost/bonus at 362..367)

Invariant-1 fraction constants (rate-limiter anchors per spec §3.4.2): SP11_EPS_DIV=1e-6, SP11_WEIGHT_HARD_FLOOR=0.01, SP11_SABOTEUR_MIN=0.5, SP11_SABOTEUR_MAX=2.0, SP11_CURIOSITY_PERMANENT_FRACTION=0.2, SP11_CURIOSITY_BOUND_FRACTION=0.3, SP11_WEIGHT_FLOOR_FRACTION=0.5, SP11_ENGAGEMENT_FLOOR=0.1. Plus SP11_PROJ_SEED_SALT=0x_5511_0001 for the A2 novelty-SimHash projection-init launcher's seed derivation.

B1b fix-up (2026-05-04, spec §4 amendment): ISV[360] = POPART_COMPONENT_MAG_EMA_INDEX added to resolve the slot-63 PopArt-input vs popart-component-magnitude overload that B1b smoke surfaced. Pre-fix-up the SP11 mag-ratio canary read 6 contiguous slots starting at slot 63 — but slot 63 was overloaded (PopArt's normalisation input AND popart-component magnitude collapsed to the same value because reward composition was inline accumulation). B1b's structural decomposition surfaced the contamination — controller redistributed weight toward popart (w_pop ≈ 2.0) based on contaminated ratio → 10× sharpe drop. Post-fix-up the mag-ratio kernel takes two slot indices: popart_specific_slot=360 for the popart axis, cf_others_base_slot=64 for cf/trail/micro/opp_cost/bonus. Slot 63 is unchanged (PopArt's input = total reward magnitude, pre-SP11 invariant preserved); slot 360 is the new dedicated popart-component magnitude produced by popart_component_ema_kernel. ISV total: 360 → 361.

B1b smoke-recovery (2026-05-04, spec §4 amendment "Why z-score" lines 564-619): ISV[361..367) = REWARD_COMPONENT_VAR_EMA_BASE..+6 added to drive z-score normalisation in reward_component_mag_ratio_compute_kernel. Pre-z-score the linear winner_weight = mag_ratio formula amplified popart's intrinsic O(100) magnitude over the other 5 components' O(0.1-2). Result: w_pop saturated toward MAX_WEIGHT, curiosity_b exploded, sharpe collapsed in B1b smoke smoke-test-6wd2c on commit 61b2fa962. Z-score makes ratios scale-invariant — z[c] = mag[c] / max(sqrt(var[c]), EPS_DIV); ratio[c] = z[c] / Σz. Variance EMAs are produced via Welford's online algorithm in the extended popart_component_ema_kernel (slot 361) and reward_component_ema_kernel (slots 362..367); the canary now takes 4 slot indices (popart mag/var split + cf-others mag/var bases). The mirror at scratch[6] still emits popart RAW magnitude (slot 359 input for curiosity bound + saboteur engagement which need physical scale). ISV total: 361 → 367.

All 21 slots are FoldReset (sentinel 0; Pearl A bootstraps from first producer launch on each fold per pearl_first_observation_bootstrap.md). After A2 all 20 slots have producers; B1b wires the on-policy reward composition consumer for [340..346) (the 6 component weights). B1c will wire the curiosity consumer for [346] + [349] (replay-time). [347] saboteur multiplier already consumed in B1a's gpu_experience_collector site. [348] REWARD_WEIGHT_FLOOR_INDEX is self-consumed by the controller on the next step. [350..360) canaries are read by the controller (intra-cycle).

The novelty-hash device buffer (GpuDqnTrainer.novelty_hash_buf, 1M slots × f32 mapped-pinned) lands in A2 alongside its registry entry (sp11_novelty_hash) and dispatch arm. The 42×16 SimHash projection matrix (GpuDqnTrainer.novelty_simhash_proj) is populated on-device by launch_novelty_simhash_proj_init at trainer construct time (Philox- driven from SP11_PROJ_SEED_SALT); it is frozen for the run lifetime and intentionally has NO registry entry — it is the hash function, not state.

Spec: docs/superpowers/specs/2026-05-04-sp11-reward-as-controlled-subsystem.md Plan: docs/superpowers/plans/2026-05-04-sp11-reward-as-controlled-subsystem.md


SP14 — Aux→Q Wire + Earned Gradient Flow [383..396)

13 slots allocated in sp14_isv_slots.rs (B.1, 2026-05-05). Slots [381..383) were already claimed by SP13 closeout (HOLD_RATE_TARGET_INDEX=381, HOLD_RATE_OBSERVED_EMA_INDEX=382), so SP14 starts at 383 (shifted +2 from the original plan which documented [381..394)).

Slot Constant Reset Notes
383 Q_DISAGREEMENT_SHORT_EMA_INDEX FoldReset (0.5) K=4↔K=2 argmax mismatch fast EMA
384 Q_DISAGREEMENT_LONG_EMA_INDEX FoldReset (0.5) K=4↔K=2 argmax mismatch slow EMA
385 K_AUX_ADAPTIVE_INDEX FoldReset (0.0→K_BASE_AUX) Sigmoid steepness for Gate 1 (aux competence)
386 K_Q_ADAPTIVE_INDEX FoldReset (0.0→K_BASE_Q) Sigmoid steepness for Gate 2 (Q-head disagreement)
387 BETA_RATE_LIMITER_ADAPTIVE_INDEX FoldReset (0.0→BETA_BASE) Rate-limiter β (variance-driven)
388 AUX_DIR_ACC_VARIANCE_EMA_INDEX FoldReset (0.0) Welford variance for k_aux adaptation
389 Q_DISAGREEMENT_VARIANCE_EMA_INDEX FoldReset (0.0) Welford variance for k_q adaptation
390 ALPHA_GRAD_RAW_VARIANCE_EMA_INDEX FoldReset (0.0) Welford variance for β adaptation
391 GATE1_OPEN_STATE_INDEX FoldReset (0.0) Schmitt-trigger state (0=closed, 1=open)
392 ALPHA_GRAD_RAW_INDEX NOT reset (recomputed every step) Raw gate output (HEALTH_DIAG visibility)
393 ALPHA_GRAD_SMOOTHED_INDEX FoldReset (0.0) Rate-limited gate consumed by backward path
394 AUX_DIR_ACC_POST_OPEN_MIN_INDEX FoldReset (1.0) Post-open aux_dir_acc minimum (circuit breaker)
395 GRADIENT_HACK_LOCKOUT_REMAINING_INDEX NOT reset (epoch-boundary decay) Lockout epochs remaining (anti-gradient-hacking)

Slots 396-398 are conceptual buffer; no constants allocated.

11 of the 13 slots are FoldReset per pearl_first_observation_bootstrap.md. Slots 392 (ALPHA_GRAD_RAW) and 395 (GRADIENT_HACK_LOCKOUT_REMAINING) are intentionally excluded from fold-reset: raw is recomputed every step from the three Welford variance EMAs; lockout decays at epoch boundary and must not be force-reset mid-fold.

Producer + consumer wiring lands in B.3-B.12; B.1 is additive infrastructure (slot constants), B.2 wires the 11 fold-reset registry entries + dispatch arms — both are additive only (no behavior change).

B.12 (2026-05-05) adds the per-epoch HEALTH_DIAG[{epoch}]: pearl_egf_diag emit in training_loop.rs reading all 13 slots. gate1/gate2 sigmoid output values are omitted from the emit (recomputing host-side would violate feedback_no_cpu_compute_strict); the inputs are sufficient for inference.

Spec: docs/superpowers/specs/2026-05-05-sp14-aux-q-wire-earned-gradient-flow.md Plan: docs/superpowers/plans/2026-05-05-sp14-aux-q-wire-earned-gradient-flow.md

SP14 L1+L2 root-cause fix — ISV_TOTAL_DIM bump (2026-05-05)

Smoke A2-B exposed: SP14 slots at 383-395 were OOB of the ISV bus (sized for ISV_TOTAL_DIM=383, top of SP13). All SP14 reads/writes touched memory past the end of the bus → gate1 stuck at 0 (OOB zero-init), post_open_min accumulated garbage (9.5 → 28 → 46), α_smoothed values were undefined behavior.

Fix:

  • ISV_TOTAL_DIM: 383 → 396 in gpu_dqn_trainer.rs:1088
  • Extended layout_fingerprint_seed() with all 13 SP14 slot name=index entries + new ISV_TOTAL_DIM=396 marker → forces fingerprint hash bump per Invariant 8 (old checkpoints invalidated correctly)
  • 2 new regression tests in sp14_isv_slots.rs:
    • all_sp14_slots_fit_within_isv_total_dim: catches the exact bug Smoke A2-B exposed
    • sp14_slot_layout_locked: locks slot layout against silent renumbering (mirrors SP5 pattern)

Why this slipped through: SP4/SP5 had analogous regression tests (all_sp4_slots_fit_within_isv_total_dim); SP14's was missing. The bug was silent because OOB GPU reads return zeros (or adjacent-allocation contents) instead of crashing — sentinel values that look almost-plausible.

Open follow-up (separate from L1/L2): warmup_gate uses hardcoded WARMUP_STEPS_FALLBACK=1000 constant. Per feedback_isv_for_adaptive_bounds, this should be ISV-signal-driven (variance-driven) OR removed entirely since k_aux/k_q already provide variance-driven warmup behavior. Redesign post-re-smoke once bus-size fix is verified.

SP14 L1+L2 follow-up — warmup_gate deletion (2026-05-05, atomic with bus-fix)

Resolution of the open follow-up above: Removed warmup_gate entirely instead of redesigning ISV-driven. The variance-driven k_aux/k_q already provide warmup behavior — when var_aux/var_q are high (cold-start), k = K_BASE / (1 + var/VARIANCE_REF) collapses toward K_MIN, which makes the sigmoid flat → gate ≈ 0.5 regardless of input. Once EMAs settle, k → K_BASE → sharp sigmoid → gates respond correctly. The separate hardcoded WARMUP_STEPS_FALLBACK ramp was double-counting and tuning-driven.

Removed:

  • WARMUP_STEPS_FALLBACK constant in sp14_isv_slots.rs
  • warmup_gate: f32 parameter in alpha_grad_compute_kernel.cu
  • gate1 * gate2 * warmup_gategate1 * gate2 in kernel
  • warmup_gate argument from launch_sp14_alpha_grad_compute launcher
  • fold_step_counter: usize field on the trainer struct (was the ramp driver)
  • fold_step_counter = 0 reset in reset_for_fold
  • fold_step_counter init in trainer constructor
  • let warmup_gate: f32 = 1.0; and .arg(&warmup_gate) from B.4 oracle tests

Net result: EGF gate's warmup behavior now lives entirely in the variance-driven k_aux/k_q sigmoid steepness controller. ISV-signal-driven via slots 388 (var_aux) and 389 (var_q). No hardcoded step counter. Honors feedback_isv_for_adaptive_bounds.

SP15 — Trader Discipline Recovery (Pre-Phase P.2 reservation, 2026-05-06)

SP15 Pre-Phase Task P.2 extends ISV_TOTAL_DIM from 396 → 443 by reserving 46 contiguous slots at [397..443). SP15 Phase 1.3.b-followup (2026-05-07) appends one more slot at [443..444) (DD_PERSISTENCE_MAX_INDEX) for the per-env DD redesign, bringing ISV_TOTAL_DIM to 444 and SP15_SLOT_END to 444 (47 SP15 slots total). The allocation pre-claims disjoint ranges across 3 sub-worktrees (Approach B parallel-dispatch) so Phase 0/1/2A can land producers and consumers independently without index collisions.

Allocation map (per spec §4.3 + Phase 1.3.b-followup audit doc entry):

Range Phase Purpose
[397..401) Phase 0.B EGF retune anchors (SCHMITT_HI/LO + VAR_AUX/Q_REF)
[401..407) Phase 1.3 + 1.3.b-followup Drawdown reporting (current/max/recovery_bars/persistence/calmar/dd_pct). Phase 1.3.b-followup (2026-05-07): now mean-aggregated across envs by dd_state_reduce_kernel from the per-env tile sp15_dd_state_per_env[alloc_episodes * 6]; the per-(env, slot) DD context is read directly by compute_sp15_final_reward_kernel from the tile
[407..409) Phase 1.2 Cost kernel (OFI_IMPACT_LAMBDA + COST_PER_BAR_AVG)
[409..417) Phase 1.4 8 counterfactual baselines (buyhold/hold-only/random-dir-kelly/naive-momentum/aux-only/mag-quarter-fixed/trail-only/naive-reversion sharpe)
[417..420) Phase 3.1 r_quality + r_discipline split (ALPHA_SPLIT + grad-norm quality/discipline)
[420..423) Phase 3.3 Quadratic DD penalty (LAMBDA_DD + DD_THRESHOLD + DD_PENALTY_GRAD_NORM)
[423..426) Phase 3.4 Regret signal (REGRET_EMA + LAMBDA_REGRET + REGRET_GRAD_NORM)
[426..430) Phase 3.5 Confidence-aware Hold floor (HOLD_FLOOR_ALPHA/K/EPS0 + ENTROPY_DIST_REF)
[430..433) Phase 3.5.2 DD asymmetric reward (DD_ASYMMETRY_LAMBDA + R_GAIN_DD_BOOST + DD_DIST_VAR)
[433..436) Phase 3.5.3 Cooldown gate (K_THRESHOLD + M_BARS + BARS_REMAINING)
[436..439) Phase 3.5.4 Plasticity injection (FIRED_THIS_FOLD + PERSISTENCE_THRESHOLD + WARM_BARS_REMAINING)
[439..441) Phase 3.5.5 Recovery curriculum in PER (DD_TRAJECTORY_DECREASING + RECOVERY_OVERSAMPLE_WEIGHT)
[441..443) Phase 3.5 deferred DD_TRAJECTORY_FLOOR + MEDIAN_STREAK_LENGTH ISV anchors
[443..444) Phase 1.3.b-followup DD_PERSISTENCE_MAX (max-aggregate of per-env DD_PERSISTENCE; consumed by plasticity_injection_kernel so the global advantage-head reset fires when ANY env exceeds the threshold — one set of advantage weights → global firing semantics → max-aggregate is the only correct rule)

Files touched (P.2 atomic):

  • crates/ml/src/cuda_pipeline/sp15_isv_slots.rs (new) — 46 pub const *_INDEX constants + SP15_SLOT_BASE/END/COUNT + 2 regression tests (all_sp15_slots_fit_within_isv_total_dim, sp15_slot_layout_locked)
  • crates/ml/src/cuda_pipeline/mod.rspub mod sp15_isv_slots; declaration
  • crates/ml/src/cuda_pipeline/gpu_dqn_trainer.rsISV_TOTAL_DIM: 396 → 443, layout_fingerprint_seed() extended with all 46 slot name=index entries before the ISV_TOTAL_DIM=443 marker (forces fingerprint hash bump → pre-SP15 checkpoints invalidated, greenfield OK per spec Q1)

Greenfield sub-worktrees: .worktrees/sp15-phase0-egf-retune, .worktrees/sp15-phase1-honest-numbers, .worktrees/sp15-phase2a-test-scaffold rebase onto this commit so each phase's producers/consumers see the slot map.

Producers/consumers status: P.2 lands the slot reservation only; all 46 slots are zero-initialized until per-phase tasks land producers (Phase 0/1/3) and atomic consumer migration (Phase 3.X). Mirrors the SP4/SP5/SP11 pre-allocation pattern.

Class A audit-fix Batch 4-A — adaptive DD saturation floor (2026-05-08)

Range Phase Purpose
[458..459) Batch 4-A DD_SATURATION_FLOOR_ADAPTIVE — replaces hardcoded 0.25f in trade_physics.cuh::apply_margin_cap (the upper end of the linear position-size scaling ramp dd_scale = max(0.05, 1.0 dd_frac/saturation_floor)). Producer dd_saturation_floor_update_kernel.cu aggregates per-env DD_MAX from sp15_dd_state_per_env[n_envs * 6] (offset 1) via Welford mean + Z_75 × sigma p75 estimator × 1.5 safety factor. Pearl-A bootstrap (sentinel 0.25 = pre-fix value) + α=0.01 slow EMA (per-epoch boundary). Bounds [0.10, 0.50] = Category-1 dimensional safety. Distinct from SP15_DD_THRESHOLD_INDEX=421 (the SP15 quadratic DD-penalty trigger threshold, a lower bound).

Atomic deletion alongside this addition: legacy compute_drawdown_penalty device function in trade_physics.cuh + its single caller at experience_kernels.cu:3822 + the w_dd Rust config field + the dd_threshold and w_dd kernel args + all w_dd profile/TOML references. Per feedback_no_legacy_aliases + feedback_no_partial_refactor — SP15's quadratic asymmetric DD penalty in compute_sp15_final_reward_kernel.cu:154 is the production-grade replacement; layering both creates double-counting.

ISV_TOTAL_DIM: 458 → 459 (Item 1 adds 1 slot; Item 2 is pure deletion).

SP20 — WR-first Reward Optimization (Pre-Phase reservation, 2026-05-09)

SP20 Pre-Phase Task Pre.2 reserves 10 contiguous slots at [510..520) for the WR-first reward design. This reservation occurs at the start of the SP20 work stream and establishes the slot base for upcoming components:

Index Name constant Purpose Producer Notes
[510] LOSS_CAP_INDEX Target loss ceiling per transaction (tuning anchor for asymmetric cap) Per-epoch ISV seed Bounds the reward penalty for loss scenarios; component 1a
[511] ALPHA_EMA_INDEX Adaptive alpha for transaction-cost EMA decay GPU kernel (planned) Tunes the responsiveness of trading cost adjustments; component 1b
[512] WR_EMA_INDEX Win rate EMA for per-bar reward shaping GPU kernel (planned) Tracks historical WR; component 2 input signal
[513] HOLD_COST_SCALE_INDEX Per-bar Hold-state carrying cost scaling Per-epoch ISV seed Modulates the cost of staying in a position; component 3a
[514] TARGET_HOLD_PCT_INDEX Target percentage of bars spent in Hold state Per-epoch ISV seed Reference for adaptive controller; component 3b
[515] HOLD_PCT_EMA_INDEX EMA of actual Hold percentage across bars GPU kernel (planned) Driven by adaptive controller; component 3c
[516] HOLD_REWARD_EMA_INDEX EMA of per-bar Hold reward component GPU kernel (planned) Tracks reward contribution from Hold state; component 3d
[517] N_STEP_INDEX N-step lookahead window for reward shaping Per-epoch ISV seed Controls temporal credit assignment; component 4
[518] AUX_CONF_THRESHOLD_INDEX Aux confidence threshold for gating auxiliary supervision Per-epoch ISV seed Prevents aux contamination; component 5a
[519] AUX_GATE_TEMP_INDEX Temperature parameter for aux gating sigmoid Per-epoch ISV seed Controls steepness of confidence-gating function; component 5b

Files touched (Pre.2 atomic):

  • crates/ml/src/cuda_pipeline/sp14_isv_slots.rs — 10 pub const *_INDEX constants (at the top level, before the test module) + SP20_SLOT_BASE (510) + SP20_SLOT_END (520) constants + 1 regression test (sp20_isv_slots_reserved_510_to_520) verifying all 10 slot indices and ISV_TOTAL_DIM=520
  • crates/ml/src/cuda_pipeline/gpu_dqn_trainer.rsISV_TOTAL_DIM: 510 → 520, layout_fingerprint_seed() extended with all 10 slot name=index entries + ISV_TOTAL_DIM=520 marker (forces fingerprint hash bump → pre-SP20 checkpoints invalidated, greenfield OK)
  • crates/ml/src/trainers/dqn/state_reset_registry.rs — Task Pre.3 registers all 10 SP20 slots with FoldReset category and sentinel=0.0 per pearl_first_observation_bootstrap; includes unit test sp20_isv_slots_registered_with_sentinel_zero

Producers/consumers status: Pre.2 lands the slot reservation only; Pre.3 registers in state reset machinery; all 10 slots are zero-initialized until subsequent SP20 phases land producers (per-epoch ISV seeds + GPU kernels) and consumer migration. Mirrors the SP4/SP5/SP11/SP15 pre-allocation pattern.

ISV_TOTAL_DIM: 510 → 520 (Pre.2 adds 10 slots).

Alpha trading system — durable infrastructure block (2026-05-15)

Phase E intro commit reserves 12 contiguous slots at [539..551) for system-scoped infrastructure of the alpha trading stack (stacker / execution-policy diagnostics, controller anchors, population baselines). Unlike the per-iteration SP4..SP22 reservations, this block is named for the system — slots are intended to outlive any individual Phase E/F/G milestone. New slots in this range should be added in alpha_isv_slots.rs first; spares at 549..550 absorb growth without bumping ISV_TOTAL_DIM.

Index Name constant Role Producer (initial commit) Update cadence Notes
539 Q_SPREAD_EMA_INDEX Diagnostic (none; Phase E.1 kill-criteria kernel) per-rollout-step `std(Q, axis=action) /
540 ACTION_ENTROPY_EMA_INDEX Diagnostic (none; Phase E.1 kill-criteria kernel) per-rollout-step H(action_dist) over recent rollouts; kill-criterion #2
541 RETURN_VS_RANDOM_EMA_INDEX Diagnostic (none; Phase E.1 kill-criteria kernel) per-trade-close (rollout_R random_R) / σ_random; kill-criterion #3
542 EARLY_Q_MOVEMENT_EMA_INDEX Diagnostic (none; Phase E.1 kill-criteria kernel) per-epoch `
543 STACKER_THRESHOLD_INDEX Controller (none; Phase E.2 stacker controller) per-rollout-step Confidence threshold for trade eligibility; engagement-rate self-correction
544 TRADE_RATE_TARGET_INDEX Anchor host init at training start static Target trade-rate fraction (e.g. 0.08)
545 TRADE_RATE_OBSERVED_EMA_INDEX Diagnostic (none; Phase E.2 stacker controller) per-rollout-step Observed trade-rate EMA
546 STACKER_KELLY_ATTENUATION_INDEX Controller (none; Phase E.2 stacker controller) per-trade-close Multiplied with KELLY_F_SMOOTH_INDEX=280 before contract-cap; bounded [0.1, 1.0] per pearl_blend_formulas_must_have_permanent_floor
547 RANDOM_BASELINE_MEAN_INDEX Anchor host computes once per fold per-fold Mean reward of uniform-random execution policy at the given horizon
548 RANDOM_BASELINE_STD_INDEX Anchor host computes once per fold per-fold Std of same — used to compute RETURN_VS_RANDOM_EMA σ-distance
549 (reserved spare) Absorb growth; do not use ad-hoc
550 (reserved spare) Absorb growth; do not use ad-hoc

Files touched (intro commit, slot reservation only):

  • crates/ml/src/cuda_pipeline/alpha_isv_slots.rs (new) — 10 pub const *_INDEX constants + ALPHA_ISV_BLOCK_LO=539/ALPHA_ISV_BLOCK_HI=550 bookends + 4 unit tests (bounds, membership, uniqueness, total-dim coverage)
  • crates/ml/src/cuda_pipeline/mod.rs — adds pub mod alpha_isv_slots;
  • crates/ml/src/cuda_pipeline/gpu_dqn_trainer.rsISV_TOTAL_DIM: 539 → 551; comment updated to reference alpha_isv_slots.rs

Producers/consumers status: intro commit is reservation-only. Producers land later in Phase E:

  • Phase E.1: alpha_kill_criteria.cu writes slots 539-542
  • Phase E.2: stacker_threshold_controller.cu writes slots 543, 545, 546 (reads 544 as anchor)
  • Phase E.0: host writes 544 (trade-rate target) at training start and 547, 548 (baselines) per fold

State-reset-registry dispatch arms for all 10 active slots land in the immediately-following commit (Phase E Task 2), per feedback_registry_entries_need_dispatch_arms.

Naming rationale: SP4..SP22 reservations are milestone-scoped (a specific SP iteration); this block is system-scoped (alpha trading system across Phase E/F/G/...). The naming difference is intentional — future alpha-system additions go in alpha_isv_slots.rs, not in a per-phase file.

ISV_TOTAL_DIM: 539 → 551 (12 slots reserved, 10 active, 2 spare).

Alpha trading system block — registry + dispatch wiring (2026-05-15, follow-up)

Phase E Task 2 lands the registry entries + 7 dispatch arms (4 SoftReset diagnostics + 3 FoldReset controllers) for the alpha-system slots reserved in Task 1. Per feedback_registry_entries_need_dispatch_arms, both must land in the same commit; the test every_fold_and_soft_reset_entry_has_dispatch_arm walks the source of reset_named_state and verifies every FoldReset/SoftReset entry has a matching match arm. 3 TrainingPersist anchors (544, 547, 548) intentionally do NOT have dispatch arms — the registry test filters them out per the convention at state_reset_registry.rs:2325.

Files touched (atomic with Task 1 if you want to squash; otherwise standalone):

  • crates/ml/src/trainers/dqn/state_reset_registry.rs — 10 RegistryEntry rows appended after the SP22 block, each with full producer/consumer rationale (Pearl A bootstrap, α-update mechanism, reset semantics)
  • crates/ml/src/trainers/dqn/trainer/training_loop.rs — 7 dispatch arms in reset_named_state (the 3 TrainingPersist names are not dispatched)

All 10 registry tests + the global registry coverage test pass: cargo test -p ml --lib state_reset_registry.

Alpha block — Phase E Task 9 producer kernel (2026-05-15)

Phase E.1 Task 9 lands the producer for slots 539-542 (alpha_kill_criteria.cu). Single-block, single-thread kernel; writes 4 raw scalar observations into scratch_out[0..4], then the existing apply_pearls_ad_kernel (chained on the same stream, n_slots=4) applies Pearl A bootstrap + Pearl D Wiener-α smoothing into the four ISV slots:

  • scratch_out[0] → ISV[539] Q_SPREAD_EMA_INDEX — mean over batch of std(Q, axis=action) / |mean(Q)|
  • scratch_out[1] → ISV[540] ACTION_ENTROPY_EMA_INDEXH(empirical action distribution) in nats
  • scratch_out[2] → ISV[541] RETURN_VS_RANDOM_EMA_INDEX(rollout_R ISV[547]) / max(ISV[548], 1e-6)
  • scratch_out[3] → ISV[542] EARLY_Q_MOVEMENT_EMA_INDEX|q_early_norm q_init_norm| / max(|q_init_norm|, 1e-6)

Kernel reads ISV[547]/ISV[548] (random baseline mean/std, populated once by Task 7c, TrainingPersist) via slot indices passed as kernel args. Composability per val_sharpe_delta_compute_kernel — no inline Wiener math here, the canonical apply_pearls_ad_kernel handles it.

Files touched:

  • crates/ml/src/cuda_pipeline/alpha_kill_criteria.cu (new) — producer kernel
  • crates/ml/build.rs — added kernel to kernels_with_common list

Cubin: target/release/build/ml-*/out/alpha_kill_criteria.cubin (~16.8 KB).

Launcher wiring: lands in Phase E.1 Task 11 (the DQN training loop reads alpha_kill_criteria_compute_kernel from the compiled cubin and chains apply_pearls_ad_kernel on the same stream after it).

Phase E.1 Task 10 — Munchausen target kernel (2026-05-15)

Standalone DQN target-augmentation kernel implementing Vieillard et al. 2020 (arXiv:2007.14430). Augments the standard TD target with:

  • Munchausen bonus: m = α_m · max(τ · log π(a|s), log_clip_min) — implicit KL regularisation between successive policies (the policy is softmax(Q_online / τ))
  • Soft V-bootstrap: V_soft(s') = max(Q_next) + τ · log Σ_a' exp((Q_next max) / τ) — replaces the hard max with a temperature-weighted softmax average

Final target: r + m + γ · V_soft(s') (or r + m if terminal).

Both softmaxes use log-sum-exp via the max-trick (essential at τ ≈ 0.03 where raw exp(Q/τ) would overflow f32 for any Q-spread > 25 nats).

Does NOT touch any ISV slot directly — α_m, τ, and log_clip_min are exposed as kernel args so a downstream controller (Phase E.2+) can ISV-drive them.

Files touched:

  • crates/ml/src/cuda_pipeline/alpha_munchausen_target.cu (new)
  • crates/ml/build.rs — kernel added to kernels_with_common

Cubin: target/release/build/ml-*/out/alpha_munchausen_target.cubin (~12.8 KB).

Wiring: lands in Phase E.1 Task 11. The integration point is wherever the existing C51/MSE loss kernels currently consume r + γ · V_next — they'll consume target_out from this kernel instead.

Phase E.1 Task 11 — Rust launchers alpha_kernels.rs (2026-05-15)

crates/ml/src/cuda_pipeline/alpha_kernels.rs exposes two pub(crate) launchers:

  • launch_alpha_kill_criteria(...) — calls alpha_kill_criteria_compute_kernel (cubin from Task 9). Grid (1, 1, 1) × Block (1, 1, 1). Reads RANDOM_BASELINE_MEAN_INDEX (547) and RANDOM_BASELINE_STD_INDEX (548) from alpha_isv_slots and passes them as kernel args. Caller is responsible for chaining apply_pearls_ad_kernel(n_slots=4, isv_idx_base=539) afterward to smooth raw scratch outputs into slots 539..542.

  • launch_alpha_munchausen_target(...) — calls alpha_munchausen_target_kernel (cubin from Task 10). Grid sized ceil(batch/256) × 1 × 1, Block (256, 1, 1). α_m, τ, log_clip_min are scalar args so the smoke / future controller can ISV-drive them.

Both launchers follow launch_apply_pearls (in sp4_wiener_ema.rs) — pre-loaded CudaFunction passed as parameter, u64 device pointers, debug_assert! guards.

Plan deviation rationale: the plan's Task 11 said "replace hardcoded n_step=32, gamma=0.999 literals" but grep across the trainer found zero such literals — the trainer already reads γ via read_isv_signal_at(GAMMA_DIR_EFF_INDEX) and ε via read_isv_signal_at(AUX_TRUNK_EPS_INDEX). The actual missing piece for E.1 was Rust-side launchers for the new cubins, delivered here.

Files touched:

  • crates/ml/src/cuda_pipeline/alpha_kernels.rs (new) — both launchers, pub(crate) visibility
  • crates/ml/src/cuda_pipeline/mod.rs — registers pub(crate) mod alpha_kernels;

Tests: cargo test -p ml --lib alpha_kernels — 1 compile-witness passes. Real GPU integration test lands in Phase E.1 Task 12 (kernel smoke binary).

Phase E.1 — kill-criteria GPU smoke (2026-05-15)

crates/ml/src/cuda_pipeline/alpha_kernels.rs::tests::kill_criteria_smoke_matches_hand_computation exercises the alpha_kill_criteria_compute_kernel end-to-end on the local GPU with synthetic inputs covering all 4 outputs. Hand-math:

  • q_values = [[1, 2, 3], [5, 5, 5]] → row 0 spread / |mean| = 0.4082; row 1 = 0 → batch mean = 0.2041
  • action_counts = [10, 30, 60] → H(p=[0.1, 0.3, 0.6]) = 0.8980 nats
  • rollout_R = 100, ISV[547] = -5185, ISV[548] = 4953(100 + 5185) / 4953 = 1.0670 σ above random
  • q_init = 50, q_early = 555 / 50 = 0.1000

ISV buffer is sized to 552 floats with the production slot indices populated using the committed Task 7c baseline values; this exercises the slot-indexing path through isv[random_baseline_mean_slot] / isv[random_baseline_std_slot] reads.

Companion to the munchausen_target_smoke_matches_hand_computation test (91d1a52b9). Both pass on RTX 3050 Ti.

Phase E.1 Task 12 wiring — chained pipeline smoke (2026-05-15)

crates/ml/src/cuda_pipeline/alpha_kernels.rs::tests::kill_criteria_chained_with_pearls_populates_isv_slots is the gating composition smoke for the H=600 DQN smoke (Task 12 proper). Validates that alpha_kill_criteria_compute_kernel correctly composes with the canonical apply_pearls_ad_kernel on the same stream:

t+0  alpha_kill_criteria_compute_kernel  → scratch[0..4]
t+1  apply_pearls_ad_kernel(n_slots=4)   → ISV[539..543]

Two iterations with stationary synthetic inputs:

  • Iter 1 (Pearl A bootstrap): prev_x_mean=0 AND x_lag=0 → ISV[539..542] = raw observations = [0.2041, 0.8980, 1.0670, 0.1] within 0.01 tolerance. Anchor slots 547/548 must remain at Task 7c values (-5185, 4953) — Pearls only writes 539..542.

  • Iter 2 (Pearl D stationary): dx_mean = dx_step = 0 → α* = 0 → ISV unchanged from iter 1 within 1e-4. Stationary signal stays at the bootstrap value.

Tests it would catch:

  • Producer's scratch write not visible to applicator (stream-ordering bug) → iter 2 produces 0
  • Wrong Pearls isv_idx_base → ISV slots wrong
  • Wrong wiener_offset_base → Wiener state corruption visible at iter 2
  • Pearl A sentinel detection broken → iter 1 produces 0 (formula yields 0 at t=0 without sentinel branch)

Reuses launch_apply_pearls from sp4_wiener_ema.rs (pub(crate)); both kernels load directly from cubins via include_bytes!.

Companion to munchausen_target_smoke_matches_hand_computation and kill_criteria_smoke_matches_hand_computation. All three pass on RTX 3050 Ti in ~2s total.

The full H=600 DQN training smoke (real trainer integration with this pipeline at rollout boundaries) is the remaining Task 12 work, queued for a dedicated session.

Phase E.1 Task 12a — alpha_linear_q kernels (2026-05-15)

Three new CUDA kernels in crates/ml/src/cuda_pipeline/alpha_linear_q.cu for the H=600 DQN smoke (Task 12 proper):

  • alpha_linear_q_forward_kernel: Q = X · W^T + b for batched states. Grid: ceil(batch · n_actions / 256), Block: 256. One thread per (b, i) inner-products over state_dim.
  • alpha_linear_q_grad_kernel: sparse MSE-TD gradients. For each (i, j): dW[i,j] = scale · sum_{b: actions[b]==i} (Q[b,i] target[b]) · X[b,j]. Only the taken action contributes — no atomicAdd needed.
  • alpha_linear_q_sgd_step_kernel: element-wise params -= lr · grad. Reused for both W and b updates with different n args.

All three are wired through pub(crate) launchers in alpha_kernels.rs. Cubin embedded at compile time via include_bytes!(env!("OUT_DIR") + "/alpha_linear_q.cubin").

Test: linear_q_forward_grad_sgd_round_trip_matches_hand_math exercises all three kernels end-to-end on a small (batch=2, state_dim=2, n_actions=3) case with hand-computed expected values for Q, dW, db, W_after, b_after. All within 1e-4 tolerance. Passes on RTX 3050 Ti.

Architecture rationale (linear, no hidden layer): the 10-dim Phase E state has meaningful direct features (alpha_logit, spread_bps, position, ...) so linear Q captures real relations like Q[Buy] ∝ alpha_logit. If linear can't pass the kill-criteria gate, no architecture upgrade will save it.

Next step (Task 12 proper): smoke binary that loads the env from MBP-10 + fitted FillModel, instantiates the linear Q-net via these kernels, runs ε-greedy training for N episodes with Munchausen target, periodically launches the kill-criteria pipeline, and emits PASS/FAIL verdict.

Phase E.1 Task 12 — H=600 DQN smoke binary (2026-05-15)

crates/ml/examples/alpha_dqn_h600_smoke.rs is the runnable kill-criteria-gate test for Milestone E.1. Linear Q-network on GPU (W [9×10] + b [9], no hidden layer) trained with ε-greedy action selection + Munchausen target augmentation on the Phase E ExecutionEnv. All compute on GPU; action selection reads 9 Q-values to CPU per step (read-only).

Reads:

  • config/ml/alpha_fill_coeffs.json (Task 5c artifact)
  • ISV anchors 547/548 = Task 7c baseline (mean=-5185, std=4953)

Writes:

  • config/ml/alpha_dqn_h600_smoke.json — verdict + per-checkpoint KC trajectory

Visibility bumps required for examples (not crate-internal):

  • cuda_pipeline::alpha_kernels module: pub(crate)pub
  • All launch_alpha_* launchers: pub(crate)pub
  • cuda_pipeline::sp4_wiener_ema::launch_apply_pearls: pub(crate)pub
  • ALPHA_LINEAR_Q_CUBIN: pub(crate)pub

These are appropriate pub exports — they're the public API for using the Phase E.1 kernel layer from external trainers / smokes.

Initial micro-smoke (horizon=100, n_episodes=50, lr=1e-6):

  • Q_SPREAD_EMA = 3.12 (≥0.05, PASS)
  • ACTION_ENTROPY_EMA = 2.12 (≥1.0986, PASS)
  • RETURN_VS_RANDOM_EMA = +1.03 (≥0.0, PASS)
  • EARLY_Q_MOVEMENT_EMA = 2268 (≥0.01, PASS but uncalibrated — ||W||/||W_init|| grew 2000×)

All four pass but training stability is poor (early_mvmt ≈ 2000× is unphysical). Known follow-ups: (a) gradient clipping, (b) target network with periodic hard-update, (c) reward normalisation. With lr=1e-4 the network diverges to NaN at H=600 — Munchausen target produces large gradients without clipping. lr=1e-6 keeps it stable but slow.

Reproduction:

cargo run -p ml --release --example alpha_dqn_h600_smoke -- \
  --mbp10-dir /home/jgrusewski/Work/foxhunt/test_data/futures-baseline-mbp10/ES.FUT \
  --horizon 600 --n-episodes 1000

This is the runnable Task 12 deliverable. The full H=6000 scale-up (Task 13) is the next milestone gated on a more stable run at H=600.

Phase E.1 Task 12 stabilization (2026-05-15)

Added three stabilizers to alpha_dqn_h600_smoke.rs after the initial smoke showed unstable training (early_mvmt = 2268× at lr=1e-6, NaN at lr=1e-4):

  1. Reward normalization (--reward-scale, default 1000): rewards divided by scale before being fed to the Munchausen target. Brings TD error from ~1000 (raw reward magnitude) into O(1). Action selection and rollout-R reporting use ORIGINAL rewards — only the bootstrap target uses normalized.

  2. Target network (--target-update-every, default 10 episodes): separate w_target_dev and b_target_dev buffers. Q_next(s') forward uses target weights; SGD updates only the online weights. Hard-update copies online → target every K episodes. Breaks the V_soft(s') chase-its-own-tail divergence loop characteristic of online-only Munchausen.

  3. Gradient clipping (--grad-clip, default 1.0): new alpha_clip_inplace_kernel in alpha_linear_q.cu (element-wise clamp). Applied to dW and db after grad kernel, before SGD step. Safety net for gradient bursts that survive the other stabilizers.

Also fixed the diagnostic: weight_norm was direction-insensitive (orthogonal rotations don't change ||W||_F), so early_mvmt read ≈0 even when training. Switched to weight_distance_from_init: ||W_now W_init||_F + ||b_now b_init||_F. The Rust side then sets q_early = q_init + distance, so the kernel's |q_early q_init| / |q_init| ratio captures relative weight-space distance from init.

Verified at horizon=100, n_episodes=200, lr=1e-4 (with all stabilizers on):

  • Q_SPREAD_EMA = 23.64 (≥ 0.05) PASS
  • ACTION_ENTROPY_EMA = 1.86 (≥ 1.0986) PASS
  • RETURN_VS_RANDOM_EMA = +0.586 (≥ 0.0) PASS
  • EARLY_Q_MOVEMENT_EMA = 0.0212 (≥ 0.01) PASS
  • Overall: PASS (H=6000 scale-up VIABLE)

early_mvmt grew monotonically (0.005 → 0.021) — direction-sensitive diagnostic confirms genuine policy learning across episodes.

Next: tune for the real H=600 / 1000-episode run; if PASS holds, proceed to Task 13 (H=6000 scale-up).

Phase E.2 Task 16 — stacker-threshold controller kernel (2026-05-15)

crates/ml/src/cuda_pipeline/stacker_threshold_controller.cu is the engagement-rate self-correcting controller for the alpha trading system. Single-block, single-thread; runs once per rollout-end boundary.

Drives:

  • ISV[543] STACKER_THRESHOLD_INDEX (clamped [0, 0.5]) — P-controller on rate error
  • ISV[545] TRADE_RATE_OBSERVED_EMA_INDEX — Pearl A bootstrap + Pearl D Wiener-α (floored at 0.4 per pearl_wiener_alpha_floor_for_nonstationary)
  • ISV[546] STACKER_KELLY_ATTENUATION_INDEX (clamped [0.1, 1.0]) — P-controller on Sharpe error

Reads:

  • ISV[544] TRADE_RATE_TARGET_INDEX (TrainingPersist anchor, set once at training start)

Control law:

observed = trade_count / max(decisions, 1)
ISV[545] ← Pearl_A+D_floored(observed, prev_x_mean, x_lag)
err_rate = ISV[545]  ISV[544]
ISV[543] ← clamp(0, 0.5, ISV[543] + k_threshold · err_rate)

err_sharpe = rollout_sharpe  target_sharpe
ISV[546] ← clamp(0.1, 1.0, ISV[546] + k_atten · err_sharpe)  (sentinel 0 → start at 1.0)

Wiener-α is computed INLINE (not via the canonical apply_pearls_ad_kernel chain) because the controller's job IS to drive the slot — the EMA is part of the control loop, not a separate diagnostic concern. Lower latency, fewer kernels per step.

Test: stacker_threshold_controller_smoke_matches_hand_computation in alpha_kernels.rs::tests verifies two iterations:

  • Iter 1 (Pearl A): observed = 0.30 → ISV[545] = 0.30 raw; threshold 0.05 → 0.052
  • Iter 2 (Pearl D): observed = 0.05 → ISV[545] = 0.175 (α* hit floor 0.5); threshold 0.052 → 0.05275

Both within 1e-5 tolerance. Anchor slot 544 unchanged after both iterations.

Cubin: target/release/build/ml-*/out/stacker_threshold_controller.cubin.

Wiring: Task 17 — invoke at each rollout-end in alpha_dqn_h600_smoke.rs; initialize slot 544 = 0.08 (8% target trade rate) at training start.

Phase E.3 follow-up — C51 distributional Q (2026-05-15)

crates/ml/src/cuda_pipeline/alpha_c51.cu adds five new kernels for distributional Q-learning (forward, project, grad, expected_q, thompson_select). None of them WRITE to ISV slots — they're pure Q-network compute kernels. They depend on the existing controller-driven slot 543 (stacker_threshold) for the GPU Thompson selector's confidence gate at inference: the threshold is populated into a kernel scalar arg from a host-cached clone_dtoh of ISV[543] at episode boundaries (same pattern as the linear-Q smoke).

ISV-continual-learning (Phase E.4 Pillar B, designed in specs/2026-05-15-phase-e-temporal-encoder-design.md — not yet implemented): the stacker-threshold controller is intended to fire at BOTH training and inference. Q-net weights stay frozen at inference; effective policy adapts via ISV slot 543 (threshold), 545 (observed-rate EMA), 546 (Kelly attenuation). No new ISV slots allocated for the E.3 follow-up; E.4.B will add MoE-gate-entropy and Pearl-1 atom-headroom slots when those land.

Phase E.4.A.10 — alpha_c51_grad_input (2026-05-15)

crates/ml/src/cuda_pipeline/alpha_c51.cu adds a 6th kernel: alpha_c51_grad_input_kernel. Computes dL/d_input[b, j] for the C51 layer (gradient w.r.t. the upstream encoder's output). Required to chain C51's gradient into Mamba2's backward in --temporal.

NO ISV slot interaction — pure Q-layer gradient compute.

T10 status: kernel landed but Mamba2 backward NOT yet wired in the smoke binary because ml_alpha::Mamba2Block::backward takes d_logit [B, 1] (post-W_out gradient) rather than d_h_enriched [B, hidden_dim]. Two paths to complete T10:

  1. Modify ml-alpha to expose backward_from_h_enriched.
  2. Replicate the post-W_out backward logic inline.

Deferred pending backtest validation — if frozen-Mamba2 already lifts backtest Sharpe, T10 becomes optimisation rather than prerequisite.

Phase E.4.A.6 — alpha_window_push (2026-05-15)

crates/ml/src/cuda_pipeline/alpha_window_push.cu: two kernels.

  1. alpha_window_push_kernel — shift+insert buffer push primitive (matches production mamba2_update_history pattern).
  2. alpha_h_enriched_store_kernel (added 2026-05-15.fix) — GPU-side slot copy from Mamba2 cache.h_enriched into the per-step buffer h_enriched_buf_dev[slot_offset..slot_offset+hidden_dim]. Replaces a dtoh/htod sequence that was violating feedback_cpu_is_read_only and slowing the eval loop ~6× (per-step full-buffer dtoh of ~20K floats).

NO ISV slot interaction — both kernels operate purely on the smoke / backtest binary's window_dev and h_enriched_buf_dev buffers. Documented here only for the kernel-audit-doc hook requirement; truly slot-agnostic.