Lands the static-configuration branch of Plan 1 C.6: cql_alpha, conviction_floor (no-op if IQL_BRANCH_SCALE_FLOOR already serves), plan_threshold written to dedicated ISV slots at constructor. Consumer kernels (CQL, backtest_plan, experience) read from ISV instead of config fields / hardcoded literals. Also pre-allocates ISV slots for the 6 upcoming GPU-kernel tasks (atoms, gamma, kelly_cap, tau, epsilon outputs + CPU-born epoch inputs). Those slots start at 0; GPU kernels in follow-up commits fill them. New ISV slots: - EPOCH_IDX_INDEX=39, TOTAL_EPOCHS_INDEX=40 (CPU-born inputs) - EPSILON_EFF_INDEX=41, TAU_EFF_INDEX=42, GAMMA_EFF_INDEX=43, KELLY_CAP_EFF_INDEX=44 (GPU-written in follow-up tasks) - CQL_ALPHA_INDEX=45, PLAN_THRESHOLD_INDEX=46 (static config; this commit) - Task 15 confirmed no-op: IQL_BRANCH_SCALE_FLOOR_INDEX=36 already serves conviction-floor role (constructor + ISV read in iql kernel). Layout fingerprint auto-updated via seed-byte edits; fingerprint re-tail at [47..49). ISV_TOTAL_DIM 39 -> 49. GpuDqnTrainConfig gains total_epochs field; fused_training.rs passes hyperparams.epochs at construction; default 0 for smoke tests. write_isv_signal_at bound extended from ISV_DIM(23) to ISV_TOTAL_DIM(49) so tail slots are writable by CPU. cql_alpha consumer: compute_cql_logit_gradients reads base from ISV[CQL_ALPHA_INDEX] instead of config.cql_alpha; adaptive formula (base x health x (1-regime_stability)) unchanged. plan_threshold consumers: experience_kernels.cu (experience_state_gather, experience_action_select, experience_env_step) and backtest_plan_kernel.cu (backtest_plan_state_isv) read plan_thr from ISV[ISV_PLAN_THRESHOLD_IDX=46] via isv_signals pointer already present in both kernels; null-guard defaults to 0.5f for smoke-scale runs without ISV warmup. ISV_PLAN_THRESHOLD_IDX=46 defined in state_layout.cuh (included by both plan kernel files); value must match PLAN_THRESHOLD_INDEX in gpu_dqn_trainer.rs. StateResetRegistry entries added for all new slots: SchemaContract for TOTAL_EPOCHS/CQL_ALPHA/PLAN_THRESHOLD; FoldReset for EPOCH_IDX and GPU-written per-fold slots. No behavioural change: all threshold/base values remain at their prior defaults; consumers now read adaptive values from ISV instead of baked-in config or hardcoded literals. Plan 1 Tasks 12, 15, 16 + pre-allocation for 9, 10, 11, 13, 14. Spec §4.C.6 (2026-04-24 GPU-drives-CPU-reads revision). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
6.0 KiB
ISV Slot Allocation Registry
Source of truth for ISV bus slot allocations. Every slot has a named constant in gpu_dqn_trainer.rs. This doc is the cross-reference.
Design: layout fingerprint, not schema version. ISV[37..39) carries a compile-time structural hash of the slot layout. Checkpoint load is fail-fast; no migration functions exist. Backward compat is structurally unwritable (there is no ordered version space to pair migrations against). See spec §4.A.2.
Tail placement rationale: The fingerprint occupies the tail ([47..49)) rather
than the head because isv_signals[0] and isv_signals[1] are actively written
by isv_signal_update (Q-drift EMA and gradient-norm EMA respectively). Inserting
at the head would displace those live signals and require shifting every upstream
literal in experience_kernels.cu. The fingerprint moves to the new tail each time
new slots are appended to the bus.
Current ISV_TOTAL_DIM: 49 (Plan 1 Tasks 12/15/16 + pre-allocation for Tasks 9-11/13/14). Post-full DQN v2 rollout: 72.
| Index | Name constant | Type | Producer | Consumers | Reset-category | Notes |
|---|---|---|---|---|---|---|
| [0] | SLOT_0_Q_DRIFT (seed only) |
f32 | isv_signal_update | ISV encoder | FoldReset | isv_signals[0] = (q_mean − q_ema) / max( |
| [1] | SLOT_1_GRAD_NORM_EMA (seed only) |
f32 | isv_signal_update | ISV encoder | FoldReset | EMA of sqrt(grad_norm²) |
| [2] | SLOT_2_TD_ERR_EMA (seed only) |
f32 | isv_signal_update | ISV encoder | FoldReset | EMA of TD-error scalar |
| [3] | SLOT_3_ENS_VAR_EMA (seed only) |
f32 | isv_signal_update | ISV encoder | FoldReset | C51 Q-distribution variance EMA (batch-mean atom-spread) |
| [4] | SLOT_4_ENS_VAR_VEL (seed only) |
f32 | isv_signal_update | ISV encoder | FoldReset | Delta of slot 3 (variance velocity) |
| [5] | SLOT_5_REWARD_EMA (seed only) |
f32 | isv_signal_update | ISV encoder | FoldReset | EMA of per-batch mean reward |
| [6] | SLOT_6_ATOM_UTIL_EMA (seed only) |
f32 | isv_signal_update | ISV encoder | FoldReset | EMA of atom utilization fraction |
| [7] | SLOT_7_LOSS_EMA (seed only) |
f32 | isv_signal_update | ISV encoder | FoldReset | EMA of total training loss |
| [8] | SLOT_8_ADX_EMA (seed only) |
f32 | isv_signal_update | ISV encoder | FoldReset | Batch-mean ADX EMA (regime velocity indicator) |
| [9] | SLOT_9_REGIME_DISAGREE (seed only) |
f32 | isv_signal_update | ISV encoder | FoldReset | |
| [10] | SLOT_10_REGIME_VEL_EMA (seed only) |
f32 | isv_signal_update | ISV encoder | FoldReset | EMA of regime transition velocity (ADX + CUSUM delta) |
| [11] | SLOT_11_REGIME_STABILITY (seed only) |
f32 | isv_signal_update | ISV encoder | FoldReset | 1 − sigmoid(5 × regime_vel); high = stable regime |
| [12] | LEARNING_HEALTH_INDEX |
f32 | isv_signal_update | many | FoldReset | health score ∈ [0, 1] |
| [13..17) | Q_MAG_MEAN_*_INDEX |
f32 | q_mag_means_reduce | c51 kernels | FoldReset | Quarter/Half/Full mag Q-mean EMAs + |Q| ref |
| [17..22) | Q_DIR_MEAN_*_INDEX |
f32 | q_dir_means_reduce | c51 kernels | FoldReset | Short/Hold/Long/Flat dir Q-mean EMAs + |Q| ref |
| [22] | SHARPE_EMA_INDEX |
f32 | training_loop host | isv_signal_update | FoldReset | Training Sharpe EMA |
| [23..31) | V_{CENTER,HALF}_{DIR,MAG,ORD,URG}_INDEX |
f32 | update_eval_v_range | adaptive_atoms, warm_start | FoldReset | Per-branch Q-support |
| [31..35) | GRAD_NORM_TARGET_*_INDEX |
f32 | grad_balance_isv_update | branch_grad_rescale | SoftReset(decay_bars=500) | Per-branch grad-norm target |
| [35] | GRAD_SCALE_LIMIT_INDEX |
f32 | grad_balance_isv_update | branch_grad_rescale | SoftReset(decay_bars=500) | Scale clamp limit |
| [36] | IQL_BRANCH_SCALE_FLOOR_INDEX |
f32 | construct (static) | iql_per_branch_advantage | SchemaContract | Per-sample branch_scales floor; safety bound |
| [37..39) | (gap — fingerprint moved to tail) | — | — | — | — | Previously [37..39); fingerprint promoted to [47..49) by Plan 1 C.6 expansion. Slots unused; zero-filled. |
| [39] | EPOCH_IDX_INDEX |
f32 (int cast) | CPU epoch-loop (follow-up task) | GPU adaptive kernels | FoldReset | Current epoch index; 0 at construction, CPU writes at each epoch boundary |
| [40] | TOTAL_EPOCHS_INDEX |
f32 (int cast) | construct (CPU static) | GPU adaptive kernels | SchemaContract | Total epoch count for this run; written at construction from config.total_epochs |
| [41] | EPSILON_EFF_INDEX |
f32 | GPU epsilon-adaptive kernel (follow-up) | epsilon-greedy action select | FoldReset | Effective epsilon; 0 at construction; GPU fills each step |
| [42] | TAU_EFF_INDEX |
f32 | GPU tau-adaptive kernel (follow-up) | target-net Polyak update | FoldReset | Effective tau; 0 at construction; GPU fills each step |
| [43] | GAMMA_EFF_INDEX |
f32 | GPU gamma-adaptive kernel (follow-up) | Bellman target kernel | FoldReset | Effective gamma; 0 at construction; GPU fills each step |
| [44] | KELLY_CAP_EFF_INDEX |
f32 | GPU kelly-cap-adaptive kernel (follow-up) | experience_env_step | FoldReset | Effective Kelly cap; 0 at construction; GPU fills each step |
| [45] | CQL_ALPHA_INDEX |
f32 | construct (CPU static) | CQL adaptive formula in Rust | SchemaContract | CQL pessimism base coefficient; written from config.cql_alpha; read in compute_cql_logit_gradients |
| [46] | PLAN_THRESHOLD_INDEX |
f32 | construct (CPU static) | experience_kernels.cu, backtest_plan_kernel.cu |
SchemaContract | Plan-MLP activation threshold; 0.5f at construction; read via ISV_PLAN_THRESHOLD_IDX in both kernels |
| [47] | ISV_LAYOUT_FINGERPRINT_LO_INDEX |
u32 bits (in f32) | construct | check_layout_fingerprint | SchemaContract | Low 32 bits of u64 FNV-1a structural hash. Fail-fast on mismatch — NOT a version number, no migration path. |
| [48] | ISV_LAYOUT_FINGERPRINT_HI_INDEX |
u32 bits (in f32) | construct | check_layout_fingerprint | SchemaContract | High 32 bits of u64 FNV-1a structural hash. |
| [49..72) | (reserved for DQN v2) | Allocated incrementally by Plans 2-5 |