Plan 3 Tasks 8 + 9. Single commit because Task 9 directly consumes Task 8's seed-fraction signal; no useful intermediate state. ISV tail-append: - [82] SEED_STEPS_TARGET_INDEX — config replay_seed_steps (CPU constructor write) - [83] SEED_STEPS_DONE_INDEX — GPU-incremented per collect_experiences_gpu - [84] SEED_FRAC_EMA_INDEX — adaptive EMA of (1 - done/target) - Fingerprint shifted [80,81] → [85,86]; ISV_TOTAL_DIM 82 → 87 GPU-only design (per user direction "fully gpu driven no cpu involvement"): - 4 scripted policies as ONE CUDA kernel (scripted_policy_kernel.cu) - Per-sample policy mix (40% uniform LCG / 20% momentum / 20% mean-rev / 20% vwap-deviation) deterministic by `i % 5` - Action source switched at launch boundary (CPU per-epoch read of ISV slot decides which kernel to dispatch; the action computation itself is 100% GPU) - No CPU physics mirror — existing GPU `experience_env_step` runs unchanged seed_step_counter_update_kernel.cu: - Single-thread cold-path; increments DONE, computes FRAC = max(0, 1-done/target) - Adaptive α matches Task 3/4 convention (α_base × (1 + 0.5×|clamp(sharpe,±2)|)) cql_alpha_seed_update_kernel.cu (Task 9): - target = config.cql_alpha × max(0, 1 - seed_frac) - During seed phase (frac=1) → target=0 → CQL α decays to 0 (no pessimism on exploration data); as frac → 0 → CQL α ramps to config value - Updates ISV[CQL_ALPHA_INDEX=48]; CQL gradient kernel reads slot 48 via pinned device-mapped ISV (Plan 1 Task 12 consumer pattern unchanged) Registry: SEED_STEPS_DONE + SEED_FRAC_EMA both FoldReset; CQL_ALPHA flipped SchemaContract → FoldReset; SEED_STEPS_TARGET stays SchemaContract. Read-only monitors (mirror PlanThresholdMonitor / StateKlMonitor pattern): - monitors/seed_monitor.rs — surfaces ISV[82..85) for HEALTH_DIAG + controller_activity smoke fire-rate - monitors/cql_alpha_monitor.rs — surfaces ISV[48] + ISV[84] dependency Smoke (RTX 3050 Ti, 3 folds × 5 epochs, dqn-smoketest profile with replay_seed_steps=1000 override so seed phase completes mid-fold): - All 3 folds saved best-checkpoint - Fold 2 best Sharpe = 92.4938 at epoch 1 (target range 80-120) ✓ - Per-fold val_metric: f0=3.80 / f1=9.73 / f2=20.24 (loss-based) - HEALTH_DIAG[3..4] cql_alpha=0.0500 with health=0.49 → base ≈ 0.10 from ISV[48], consistent with kernel ramping toward final×(1-frac); regime gate (1-regime)×health applies on top - 11 cargo check warnings (matches pre-task baseline; no new warnings) - 6/6 monitor unit tests pass (read/diagnose/observe×fire_rate) Smoke override rationale: smoke runs ~200 samples per collect (4 episodes × 50 timesteps) × 5 epochs × 3 folds ≈ 3000 total. Default 100k target would keep entire smoke in seed phase. Override to 1000 lets the seed→network transition complete mid-fold so the CQL α ramp is observable. Per pearl_one_unbounded_signal_per_reward.md: cql_alpha is bounded (clamped to config_final × (1 - seed_frac) ∈ [0, config_final]), composes safely with downstream CQL loss. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
15 KiB
ISV Slot Allocation Registry
Source of truth for ISV bus slot allocations. Every slot has a named constant in gpu_dqn_trainer.rs. This doc is the cross-reference.
Design: layout fingerprint, not schema version. ISV[37..39) carries a compile-time structural hash of the slot layout. Checkpoint load is fail-fast; no migration functions exist. Backward compat is structurally unwritable (there is no ordered version space to pair migrations against). See spec §4.A.2.
Tail placement rationale: The fingerprint occupies the tail ([55..57)) rather
than the head because isv_signals[0] and isv_signals[1] are actively written
by isv_signal_update (Q-drift EMA and gradient-norm EMA respectively). Inserting
at the head would displace those live signals and require shifting every upstream
literal in experience_kernels.cu. The fingerprint moves to the new tail each time
new slots are appended to the bus.
Current ISV_TOTAL_DIM: 87 (Plan 1 + Plan 2 Task 1 C.1 + Plan 2 Task 3 D.2 per-branch gamma + Plan 2 Task 6C D.8 TLOB + Plan 3 Task 1 C.2 reward-component EMAs + Plan 3 Task 3 B.2 trade-attempt novelty + Plan 3 Task 4 B.4 readiness-EMA + Plan 3 Task 7 C.3 state-distribution KL + Plan 3 Task 8 B.3 GPU-only seed warm-start [82..85)). Post-full DQN v2 rollout: 87+.
| Index | Name constant | Type | Producer | Consumers | Reset-category | Notes |
|---|---|---|---|---|---|---|
| [0] | SLOT_0_Q_DRIFT (seed only) |
f32 | isv_signal_update | ISV encoder | FoldReset | isv_signals[0] = (q_mean − q_ema) / max( |
| [1] | SLOT_1_GRAD_NORM_EMA (seed only) |
f32 | isv_signal_update | ISV encoder | FoldReset | EMA of sqrt(grad_norm²) |
| [2] | SLOT_2_TD_ERR_EMA (seed only) |
f32 | isv_signal_update | ISV encoder | FoldReset | EMA of TD-error scalar |
| [3] | SLOT_3_ENS_VAR_EMA (seed only) |
f32 | isv_signal_update | ISV encoder | FoldReset | C51 Q-distribution variance EMA (batch-mean atom-spread) |
| [4] | SLOT_4_ENS_VAR_VEL (seed only) |
f32 | isv_signal_update | ISV encoder | FoldReset | Delta of slot 3 (variance velocity) |
| [5] | SLOT_5_REWARD_EMA (seed only) |
f32 | isv_signal_update | ISV encoder | FoldReset | EMA of per-batch mean reward |
| [6] | SLOT_6_ATOM_UTIL_EMA (seed only) |
f32 | isv_signal_update | ISV encoder | FoldReset | EMA of atom utilization fraction |
| [7] | SLOT_7_LOSS_EMA (seed only) |
f32 | isv_signal_update | ISV encoder | FoldReset | EMA of total training loss |
| [8] | SLOT_8_ADX_EMA (seed only) |
f32 | isv_signal_update | ISV encoder | FoldReset | Batch-mean ADX EMA (regime velocity indicator) |
| [9] | SLOT_9_REGIME_DISAGREE (seed only) |
f32 | isv_signal_update | ISV encoder | FoldReset | |
| [10] | SLOT_10_REGIME_VEL_EMA (seed only) |
f32 | isv_signal_update | ISV encoder | FoldReset | EMA of regime transition velocity (ADX + CUSUM delta) |
| [11] | SLOT_11_REGIME_STABILITY (seed only) |
f32 | isv_signal_update | ISV encoder | FoldReset | 1 − sigmoid(5 × regime_vel); high = stable regime |
| [12] | LEARNING_HEALTH_INDEX |
f32 | isv_signal_update | many | FoldReset | health score ∈ [0, 1] |
| [13..17) | Q_MAG_MEAN_*_INDEX |
f32 | q_mag_means_reduce | c51 kernels | FoldReset | Quarter/Half/Full mag Q-mean EMAs + |Q| ref |
| [17..22) | Q_DIR_MEAN_*_INDEX |
f32 | q_dir_means_reduce | c51 kernels | FoldReset | Short/Hold/Long/Flat dir Q-mean EMAs + |Q| ref |
| [22] | SHARPE_EMA_INDEX |
f32 | training_loop host | isv_signal_update | FoldReset | Training Sharpe EMA |
| [23..31) | V_{CENTER,HALF}_{DIR,MAG,ORD,URG}_INDEX |
f32 | update_eval_v_range | adaptive_atoms, warm_start | FoldReset | Per-branch Q-support |
| [31..35) | GRAD_NORM_TARGET_*_INDEX |
f32 | grad_balance_isv_update | branch_grad_rescale | SoftReset(decay_bars=500) | Per-branch grad-norm target |
| [35] | GRAD_SCALE_LIMIT_INDEX |
f32 | grad_balance_isv_update | branch_grad_rescale | SoftReset(decay_bars=500) | Scale clamp limit |
| [36] | IQL_BRANCH_SCALE_FLOOR_INDEX |
f32 | construct (static) | iql_per_branch_advantage | SchemaContract | Per-sample branch_scales floor; safety bound |
| [37..39) | (gap — fingerprint moved to tail) | — | — | — | — | Previously [37..39); fingerprint promoted to [47..49) by Plan 1 C.6 expansion. Slots unused; zero-filled. |
| [39] | EPOCH_IDX_INDEX |
f32 (int cast) | CPU epoch-loop (follow-up task) | GPU adaptive kernels | FoldReset | Current epoch index; 0 at construction, CPU writes at each epoch boundary |
| [40] | TOTAL_EPOCHS_INDEX |
f32 (int cast) | construct (CPU static) | GPU adaptive kernels | SchemaContract | Total epoch count for this run; written at construction from config.total_epochs |
| [41] | EPSILON_EFF_INDEX |
f32 | GPU epsilon-adaptive kernel (follow-up) | epsilon-greedy action select | FoldReset | Effective epsilon; 0 at construction; GPU fills each step |
| [42] | TAU_EFF_INDEX |
f32 | GPU tau-adaptive kernel (follow-up) | target-net Polyak update | FoldReset | Effective tau; 0 at construction; GPU fills each step |
| [43] | GAMMA_DIR_EFF_INDEX |
f32 | GPU per_branch_gamma_update kernel (Plan 2 D.2) | c51_loss_batched, iql_compute_per_sample_support | FoldReset | Direction-branch gamma; base=0.92, max=0.99; GPU fills each epoch boundary |
| [44] | GAMMA_MAG_EFF_INDEX |
f32 | GPU per_branch_gamma_update kernel (Plan 2 D.2) | c51_loss_batched, iql_compute_per_sample_support | FoldReset | Magnitude-branch gamma; base=0.88, max=0.95 |
| [45] | GAMMA_ORD_EFF_INDEX |
f32 | GPU per_branch_gamma_update kernel (Plan 2 D.2) | c51_loss_batched, iql_compute_per_sample_support | FoldReset | Order-branch gamma; base=0.85, max=0.93 |
| [46] | GAMMA_URG_EFF_INDEX |
f32 | GPU per_branch_gamma_update kernel (Plan 2 D.2) | c51_loss_batched, iql_compute_per_sample_support | FoldReset | Urgency-branch gamma; base=0.80, max=0.90 |
| [47] | KELLY_CAP_EFF_INDEX |
f32 | GPU kelly-cap-adaptive kernel (Plan 1 Task 11) | experience_env_step | FoldReset | Effective Kelly cap; 0 at construction; GPU fills each step |
| [48] | CQL_ALPHA_INDEX |
f32 | construct (CPU static) | CQL adaptive formula in Rust | SchemaContract | CQL pessimism base coefficient; written from config.cql_alpha; read in compute_cql_logit_gradients |
| [49] | PLAN_THRESHOLD_INDEX |
f32 | GPU plan_threshold_update kernel (Plan 3 B.4) |
experience_kernels.cu, backtest_plan_kernel.cu |
FoldReset | Plan-MLP activation threshold; producer upgraded from static constructor write to GPU kernel by Plan 3 Task 4 B.4. Cold-start 0.5f preserved. Kernel writes max(0.1, 0.5 × ISV[READINESS_EMA_INDEX=75]) each epoch. Consumers (4 sites in experience_kernels.cu + 1 in backtest_plan_kernel.cu) unchanged — still read via ISV_PLAN_THRESHOLD_IDX. |
| [50] | Q_P05_DIR_INDEX |
f32 | q_quantile_reduce GPU kernel (cold-path per-epoch) | update_eval_v_range | FoldReset | Direction-branch 5th-percentile Q EMA. Bootstrap: v_min. |
| [51] | Q_P05_MAG_INDEX |
f32 | q_quantile_reduce | update_eval_v_range | FoldReset | Magnitude-branch 5th-percentile Q EMA. Bootstrap: v_min. |
| [52] | Q_P05_ORD_INDEX |
f32 | q_quantile_reduce | update_eval_v_range | FoldReset | Order-branch 5th-percentile Q EMA. Bootstrap: v_min. |
| [53] | Q_P05_URG_INDEX |
f32 | q_quantile_reduce | update_eval_v_range | FoldReset | Urgency-branch 5th-percentile Q EMA. Bootstrap: v_min. |
| [54] | Q_P95_DIR_INDEX |
f32 | q_quantile_reduce GPU kernel (cold-path per-epoch) | update_eval_v_range | FoldReset | Direction-branch 95th-percentile Q EMA. Bootstrap: v_max. |
| [55] | Q_P95_MAG_INDEX |
f32 | q_quantile_reduce | update_eval_v_range | FoldReset | Magnitude-branch 95th-percentile Q EMA. Bootstrap: v_max. |
| [56] | Q_P95_ORD_INDEX |
f32 | q_quantile_reduce | update_eval_v_range | FoldReset | Order-branch 95th-percentile Q EMA. Bootstrap: v_max. |
| [57] | Q_P95_URG_INDEX |
f32 | q_quantile_reduce | update_eval_v_range | FoldReset | Urgency-branch 95th-percentile Q EMA. Bootstrap: v_max. |
| [58..60) | (gap — fingerprint shifted to tail) | — | — | — | — | Previously [58..60); fingerprint promoted to [61..63) by Plan 2 Task 6C D.8 TLOB expansion. Slots unused; zero-filled. |
| [60] | TLOB_REGIME_FOCUS_EMA_INDEX |
f32 | CPU training_loop (per-epoch) | ISV consumers (diagnostic) | FoldReset | Mean-max TLOB attention weight EMA (α=0.05); written at epoch boundary from GpuTlob::mean_max_attention_weight. Indicates OFI feature focus sharpness. |
| [61] | (gap — fingerprint shifted to tail) | — | — | — | — | Previously [61]; fingerprint promoted to [69..71) by Plan 3 Task 1 C.2 reward-EMA expansion. Slot unused; zero-filled. |
| [62] | (gap — fingerprint shifted to tail) | — | — | — | — | Previously [62]; see [61] note above. Slot unused; zero-filled. |
| [63] | REWARD_POPART_EMA_INDEX |
f32 | GPU reward_component_ema kernel (Plan 3 C.2) |
HEALTH_DIAG reward_split, RewardComponentMonitor |
FoldReset | EMA of mean |popart reward| across batch (α=0.05). Final on-policy reward component. |
| [64] | REWARD_CF_EMA_INDEX |
f32 | GPU reward_component_ema kernel (Plan 3 C.2) |
HEALTH_DIAG reward_split, RewardComponentMonitor |
FoldReset | EMA of mean |counterfactual reward| across batch (α=0.05). Zero until Plan 3 B.1. |
| [65] | REWARD_TRAIL_EMA_INDEX |
f32 | GPU reward_component_ema kernel (Plan 3 C.2) |
HEALTH_DIAG reward_split, RewardComponentMonitor |
FoldReset | EMA of mean |trail reward| across batch (α=0.05). Structural placeholder; populated by Plan 3 B.2. |
| [66] | REWARD_MICRO_EMA_INDEX |
f32 | GPU reward_component_ema kernel (Plan 3 C.2) |
HEALTH_DIAG reward_split, RewardComponentMonitor |
FoldReset | EMA of mean |OFI micro-reward| across batch (α=0.05). Populated via rc[3] in experience_env_step. |
| [67] | REWARD_OPP_COST_EMA_INDEX |
f32 | GPU reward_component_ema kernel (Plan 3 C.2) |
HEALTH_DIAG reward_split, RewardComponentMonitor |
FoldReset | EMA of mean |opportunity-cost reward| across batch (α=0.05). Populated by Plan 3 B.1 at the Flat-branch opp-cost site. |
| [68] | REWARD_BONUS_EMA_INDEX |
f32 | GPU reward_component_ema kernel (Plan 3 C.2) |
HEALTH_DIAG reward_split, RewardComponentMonitor |
FoldReset | EMA of mean |bonus reward| across batch (α=0.05). Populated by Plan 3 B.2 at the Flat→Positioned site (rc[5]). |
| [71] | TRADE_ATTEMPT_RATE_EMA_INDEX |
f32 | GPU trade_attempt_rate_ema_update kernel (Plan 3 B.2) |
experience_env_step Flat→Positioned bonus path |
FoldReset | EMA of Flat→Positioned transition rate (0..1). Adaptive α = α_base × (1 + 0.5×|clamp(sharpe, -2, 2)|). α_base = 0.05. Zero at construction and fold boundary. |
| [72] | TRADE_TARGET_RATE_INDEX |
f32 | CPU training_loop (epoch 5 freeze) |
experience_env_step novelty computation |
FoldReset | Reference rate for novelty = max(0, 1 - attempt/target). Frozen at epoch 5 from measured TRADE_ATTEMPT_RATE_EMA (min 0.001). Pre-freeze the slot is 0 and the bonus site gates on target_raw > 1e-6f, so the reward term is structurally inert until the freeze fires. |
| [73..76) | (gap — fingerprint shifted to tail) | — | — | — | — | Previously [73..75); fingerprint shifted by Plan 3 Task 4 B.4 to accommodate the new READINESS_EMA slot. Slots unused; zero-filled. |
| [75] | READINESS_EMA_INDEX |
f32 | GPU plan_threshold_update kernel (Plan 3 B.4) |
derived → ISV[PLAN_THRESHOLD_INDEX=49]; PlanThresholdMonitor |
FoldReset | Per-batch mean readiness EMA (α=α_base × (1+0.5×|clamp(sharpe,−2,2)|), α_base=0.05). Cold-start 1.0 so derived threshold = 0.5 matches the previous static default. Drives slot 49 producer-only — consumers unchanged. |
| [78] | STATE_KL_TRAIN_VAL_EMA_INDEX |
f32 | GPU state_kl_moment_match kernel (Plan 3 C.3) |
StateKLMonitor (HEALTH_DIAG) |
FoldReset | Per-validation-epoch moment-match KL EMA between train-state sample and val-state batch over the OFI block (32 dims). Adaptive α=α_base × (1+0.5×|clamp(sharpe,−2,2)|), α_base=0.05. Cold-start 0.0 — first kernel fire EMAs the actual measured KL toward this. |
| [79] | STATE_KL_AMPLIFICATION_INDEX |
f32 | GPU state_kl_moment_match kernel (Plan 3 C.3) |
experience_kernels.cu B.1 opp_cost + B.2 bonus consumers (via ISV_STATE_KL_AMP_IDX macro) |
FoldReset | Bounded amplifier ∈ [1, 2] tracking trailing-EMA-of-self KL ratio: 1 + clamp(new_ema/prev_ema − 1, 0, 1). Cold-start 1.0 — neutral element behind consumers' fmaxf(1.0, …) guard. When train/val distributions diverge, both Flat opp-cost and trade-attempt bonus scale up together. Per pearl_one_unbounded_signal_per_reward.md: amp is bounded so it stacks safely with B.1's q_abs_ref unbounded multiplicand. |
| [48] | CQL_ALPHA_INDEX (post-Plan-3-T9 producer upgrade) |
f32 | GPU cql_alpha_seed_update kernel (Plan 3 C.5) |
compute_cql_logit_gradients (Rust ISV read) |
FoldReset (was SchemaContract) | CQL pessimism base coefficient. Producer upgraded SchemaContract → FoldReset by Plan 3 Task 9 C.5: per-epoch GPU kernel EMAs slot 48 toward config.cql_alpha × max(0, 1 - ISV[SEED_FRAC_EMA_INDEX=84]). During seed phase (frac=1), target=0 → CQL α decays to 0 (no pessimism on exploration data); as frac → 0, α ramps to config.cql_alpha (full pessimism on network-driven trajectories). Constructor cold-starts to config.cql_alpha; fold-boundary reset re-applies. |
| [82] | SEED_STEPS_TARGET_INDEX |
f32 | construct (CPU static, from config.replay_seed_steps) |
seed_step_counter_update GPU kernel (Plan 3 B.3) |
SchemaContract | Plan 3 Task 8 B.3 — target number of seed-phase experience-collection samples. Default 100_000; smoke configs may override to 10_000 so the seed→network transition is observable inside the smoke run. |
| [83] | SEED_STEPS_DONE_INDEX |
f32 | GPU seed_step_counter_update kernel (Plan 3 B.3) |
training_loop CPU dispatch decision (per-epoch read) + seed_step_counter_update self |
FoldReset | Plan 3 Task 8 B.3 — cumulative seed-phase steps completed. GPU-incremented by n_samples_this_step per collect_experiences_gpu call, capped at TARGET. Cold-start 0; fold-boundary reset re-applies. CPU per-epoch read of (DONE, TARGET) decides whether next collect dispatches scripted-policy kernel (DONE < TARGET) or experience_action_select. |
| [84] | SEED_FRAC_EMA_INDEX |
f32 | GPU seed_step_counter_update kernel (Plan 3 B.3) |
GPU cql_alpha_seed_update kernel (Plan 3 C.5) + SeedMonitor |
FoldReset | Plan 3 Task 8 B.3 — adaptive EMA of max(0, 1 - DONE/TARGET) ∈ [0, 1]. Cold-start 1.0 (fully in seed phase) so Task 9's CQL ramp sees target = final × max(0, 1 - 1.0) = 0 until the seed phase actually decays. Adaptive α = α_base × (1 + 0.5 × |clamp(sharpe, ±2)|), α_base = 0.05. |
| [85] | ISV_LAYOUT_FINGERPRINT_LO_INDEX |
u32 bits (in f32) | construct | check_layout_fingerprint | SchemaContract | Low 32 bits of u64 FNV-1a structural hash. Fail-fast on mismatch — NOT a version number, no migration path. Shifted 69→73 by Plan 3 Task 3 B.2, 73→76 by Plan 3 Task 4 B.4, 76→80 by Plan 3 Task 7 C.3, 80→85 by Plan 3 Task 8 B.3. |
| [86] | ISV_LAYOUT_FINGERPRINT_HI_INDEX |
u32 bits (in f32) | construct | check_layout_fingerprint | SchemaContract | High 32 bits of u64 FNV-1a structural hash. Shifted 70→74 by Plan 3 Task 3 B.2, 74→77 by Plan 3 Task 4 B.4, 77→81 by Plan 3 Task 7 C.3, 81→86 by Plan 3 Task 8 B.3. |
| [87) | (reserved for DQN v2) | Allocated incrementally by Plans 3-5 |