Files
foxhunt/docs/isv-slots.md
jgrusewski b5d19c1004 feat(dqn-v2): B.2 ISV-driven trade-attempt bonus — novelty at Flat→Positioned
Plan 3 Task 3.

ISV tail-append:
- [71] TRADE_ATTEMPT_RATE_EMA — Flat→Positioned transition rate EMA
       (GPU-written by trade_attempt_rate_ema_update, adaptive α)
- [72] TRADE_TARGET_RATE — reference rate, CPU-frozen at epoch 5
- Fingerprint shifted [69,70] → [73,74]; ISV_TOTAL_DIM 71 → 75

Producer kernel (trade_rate_ema_kernel.cu):
- Single-block reduction of flat_to_pos_per_sample [N*L] (no atomicAdd)
- Adaptive EMA: α = α_base × (1 + 0.5 × |clamp(sharpe, -2, 2)|)
- α_base = 0.05 (matches reward_component_ema convention)
- Launched from training_loop alongside reward_component_ema

Consumer (experience_env_step):
- Flat→Positioned site: novelty = max(0, 1 - attempt/target)
- bonus = conviction_core × vol_proxy × novelty
- reward += shaping_scale × bonus; rc[5] captures bonus for ISV[68]
- Explicit freeze gate: target_raw > 1e-6f, so bonus is structurally
  inert pre-freeze (prevents spurious novelty=1.0 on epoch-1 when
  attempt_rate is still 0)

Epoch-5 freeze (training_loop):
- measured = ISV[TRADE_ATTEMPT_RATE_EMA]; floor at 0.001
- Prevents novelty from sticking at 1.0 post-freeze

StateResetRegistry: both slots registered as FoldReset with per-slot
reset dispatch arms in training_loop's fold-boundary path.

Smoke: multi_fold_convergence passes (fold 2 best Sharpe 87.55 —
 slightly above Task 2 baseline 85.6, within noise; bonus inert
 in 5-epoch smoke so training matches pre-B.2 baseline as designed).
cargo check clean at 11 warnings baseline.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-24 23:11:20 +02:00

11 KiB
Raw Blame History

ISV Slot Allocation Registry

Source of truth for ISV bus slot allocations. Every slot has a named constant in gpu_dqn_trainer.rs. This doc is the cross-reference.

Design: layout fingerprint, not schema version. ISV[37..39) carries a compile-time structural hash of the slot layout. Checkpoint load is fail-fast; no migration functions exist. Backward compat is structurally unwritable (there is no ordered version space to pair migrations against). See spec §4.A.2.

Tail placement rationale: The fingerprint occupies the tail ([55..57)) rather than the head because isv_signals[0] and isv_signals[1] are actively written by isv_signal_update (Q-drift EMA and gradient-norm EMA respectively). Inserting at the head would displace those live signals and require shifting every upstream literal in experience_kernels.cu. The fingerprint moves to the new tail each time new slots are appended to the bus.

Current ISV_TOTAL_DIM: 75 (Plan 1 + Plan 2 Task 1 C.1 + Plan 2 Task 3 D.2 per-branch gamma + Plan 2 Task 6C D.8 TLOB + Plan 3 Task 1 C.2 reward-component EMAs + Plan 3 Task 3 B.2 trade-attempt novelty). Post-full DQN v2 rollout: 75+.

Index Name constant Type Producer Consumers Reset-category Notes
[0] SLOT_0_Q_DRIFT (seed only) f32 isv_signal_update ISV encoder FoldReset isv_signals[0] = (q_mean q_ema) / max(
[1] SLOT_1_GRAD_NORM_EMA (seed only) f32 isv_signal_update ISV encoder FoldReset EMA of sqrt(grad_norm²)
[2] SLOT_2_TD_ERR_EMA (seed only) f32 isv_signal_update ISV encoder FoldReset EMA of TD-error scalar
[3] SLOT_3_ENS_VAR_EMA (seed only) f32 isv_signal_update ISV encoder FoldReset C51 Q-distribution variance EMA (batch-mean atom-spread)
[4] SLOT_4_ENS_VAR_VEL (seed only) f32 isv_signal_update ISV encoder FoldReset Delta of slot 3 (variance velocity)
[5] SLOT_5_REWARD_EMA (seed only) f32 isv_signal_update ISV encoder FoldReset EMA of per-batch mean reward
[6] SLOT_6_ATOM_UTIL_EMA (seed only) f32 isv_signal_update ISV encoder FoldReset EMA of atom utilization fraction
[7] SLOT_7_LOSS_EMA (seed only) f32 isv_signal_update ISV encoder FoldReset EMA of total training loss
[8] SLOT_8_ADX_EMA (seed only) f32 isv_signal_update ISV encoder FoldReset Batch-mean ADX EMA (regime velocity indicator)
[9] SLOT_9_REGIME_DISAGREE (seed only) f32 isv_signal_update ISV encoder FoldReset
[10] SLOT_10_REGIME_VEL_EMA (seed only) f32 isv_signal_update ISV encoder FoldReset EMA of regime transition velocity (ADX + CUSUM delta)
[11] SLOT_11_REGIME_STABILITY (seed only) f32 isv_signal_update ISV encoder FoldReset 1 sigmoid(5 × regime_vel); high = stable regime
[12] LEARNING_HEALTH_INDEX f32 isv_signal_update many FoldReset health score ∈ [0, 1]
[13..17) Q_MAG_MEAN_*_INDEX f32 q_mag_means_reduce c51 kernels FoldReset Quarter/Half/Full mag Q-mean EMAs + |Q| ref
[17..22) Q_DIR_MEAN_*_INDEX f32 q_dir_means_reduce c51 kernels FoldReset Short/Hold/Long/Flat dir Q-mean EMAs + |Q| ref
[22] SHARPE_EMA_INDEX f32 training_loop host isv_signal_update FoldReset Training Sharpe EMA
[23..31) V_{CENTER,HALF}_{DIR,MAG,ORD,URG}_INDEX f32 update_eval_v_range adaptive_atoms, warm_start FoldReset Per-branch Q-support
[31..35) GRAD_NORM_TARGET_*_INDEX f32 grad_balance_isv_update branch_grad_rescale SoftReset(decay_bars=500) Per-branch grad-norm target
[35] GRAD_SCALE_LIMIT_INDEX f32 grad_balance_isv_update branch_grad_rescale SoftReset(decay_bars=500) Scale clamp limit
[36] IQL_BRANCH_SCALE_FLOOR_INDEX f32 construct (static) iql_per_branch_advantage SchemaContract Per-sample branch_scales floor; safety bound
[37..39) (gap — fingerprint moved to tail) Previously [37..39); fingerprint promoted to [47..49) by Plan 1 C.6 expansion. Slots unused; zero-filled.
[39] EPOCH_IDX_INDEX f32 (int cast) CPU epoch-loop (follow-up task) GPU adaptive kernels FoldReset Current epoch index; 0 at construction, CPU writes at each epoch boundary
[40] TOTAL_EPOCHS_INDEX f32 (int cast) construct (CPU static) GPU adaptive kernels SchemaContract Total epoch count for this run; written at construction from config.total_epochs
[41] EPSILON_EFF_INDEX f32 GPU epsilon-adaptive kernel (follow-up) epsilon-greedy action select FoldReset Effective epsilon; 0 at construction; GPU fills each step
[42] TAU_EFF_INDEX f32 GPU tau-adaptive kernel (follow-up) target-net Polyak update FoldReset Effective tau; 0 at construction; GPU fills each step
[43] GAMMA_DIR_EFF_INDEX f32 GPU per_branch_gamma_update kernel (Plan 2 D.2) c51_loss_batched, iql_compute_per_sample_support FoldReset Direction-branch gamma; base=0.92, max=0.99; GPU fills each epoch boundary
[44] GAMMA_MAG_EFF_INDEX f32 GPU per_branch_gamma_update kernel (Plan 2 D.2) c51_loss_batched, iql_compute_per_sample_support FoldReset Magnitude-branch gamma; base=0.88, max=0.95
[45] GAMMA_ORD_EFF_INDEX f32 GPU per_branch_gamma_update kernel (Plan 2 D.2) c51_loss_batched, iql_compute_per_sample_support FoldReset Order-branch gamma; base=0.85, max=0.93
[46] GAMMA_URG_EFF_INDEX f32 GPU per_branch_gamma_update kernel (Plan 2 D.2) c51_loss_batched, iql_compute_per_sample_support FoldReset Urgency-branch gamma; base=0.80, max=0.90
[47] KELLY_CAP_EFF_INDEX f32 GPU kelly-cap-adaptive kernel (Plan 1 Task 11) experience_env_step FoldReset Effective Kelly cap; 0 at construction; GPU fills each step
[48] CQL_ALPHA_INDEX f32 construct (CPU static) CQL adaptive formula in Rust SchemaContract CQL pessimism base coefficient; written from config.cql_alpha; read in compute_cql_logit_gradients
[49] PLAN_THRESHOLD_INDEX f32 construct (CPU static) experience_kernels.cu, backtest_plan_kernel.cu SchemaContract Plan-MLP activation threshold; 0.5f at construction; read via ISV_PLAN_THRESHOLD_IDX in both kernels
[50] Q_P05_DIR_INDEX f32 q_quantile_reduce GPU kernel (cold-path per-epoch) update_eval_v_range FoldReset Direction-branch 5th-percentile Q EMA. Bootstrap: v_min.
[51] Q_P05_MAG_INDEX f32 q_quantile_reduce update_eval_v_range FoldReset Magnitude-branch 5th-percentile Q EMA. Bootstrap: v_min.
[52] Q_P05_ORD_INDEX f32 q_quantile_reduce update_eval_v_range FoldReset Order-branch 5th-percentile Q EMA. Bootstrap: v_min.
[53] Q_P05_URG_INDEX f32 q_quantile_reduce update_eval_v_range FoldReset Urgency-branch 5th-percentile Q EMA. Bootstrap: v_min.
[54] Q_P95_DIR_INDEX f32 q_quantile_reduce GPU kernel (cold-path per-epoch) update_eval_v_range FoldReset Direction-branch 95th-percentile Q EMA. Bootstrap: v_max.
[55] Q_P95_MAG_INDEX f32 q_quantile_reduce update_eval_v_range FoldReset Magnitude-branch 95th-percentile Q EMA. Bootstrap: v_max.
[56] Q_P95_ORD_INDEX f32 q_quantile_reduce update_eval_v_range FoldReset Order-branch 95th-percentile Q EMA. Bootstrap: v_max.
[57] Q_P95_URG_INDEX f32 q_quantile_reduce update_eval_v_range FoldReset Urgency-branch 95th-percentile Q EMA. Bootstrap: v_max.
[58..60) (gap — fingerprint shifted to tail) Previously [58..60); fingerprint promoted to [61..63) by Plan 2 Task 6C D.8 TLOB expansion. Slots unused; zero-filled.
[60] TLOB_REGIME_FOCUS_EMA_INDEX f32 CPU training_loop (per-epoch) ISV consumers (diagnostic) FoldReset Mean-max TLOB attention weight EMA (α=0.05); written at epoch boundary from GpuTlob::mean_max_attention_weight. Indicates OFI feature focus sharpness.
[61] (gap — fingerprint shifted to tail) Previously [61]; fingerprint promoted to [69..71) by Plan 3 Task 1 C.2 reward-EMA expansion. Slot unused; zero-filled.
[62] (gap — fingerprint shifted to tail) Previously [62]; see [61] note above. Slot unused; zero-filled.
[63] REWARD_POPART_EMA_INDEX f32 GPU reward_component_ema kernel (Plan 3 C.2) HEALTH_DIAG reward_split, RewardComponentMonitor FoldReset EMA of mean |popart reward| across batch (α=0.05). Final on-policy reward component.
[64] REWARD_CF_EMA_INDEX f32 GPU reward_component_ema kernel (Plan 3 C.2) HEALTH_DIAG reward_split, RewardComponentMonitor FoldReset EMA of mean |counterfactual reward| across batch (α=0.05). Zero until Plan 3 B.1.
[65] REWARD_TRAIL_EMA_INDEX f32 GPU reward_component_ema kernel (Plan 3 C.2) HEALTH_DIAG reward_split, RewardComponentMonitor FoldReset EMA of mean |trail reward| across batch (α=0.05). Structural placeholder; populated by Plan 3 B.2.
[66] REWARD_MICRO_EMA_INDEX f32 GPU reward_component_ema kernel (Plan 3 C.2) HEALTH_DIAG reward_split, RewardComponentMonitor FoldReset EMA of mean |OFI micro-reward| across batch (α=0.05). Populated via rc[3] in experience_env_step.
[67] REWARD_OPP_COST_EMA_INDEX f32 GPU reward_component_ema kernel (Plan 3 C.2) HEALTH_DIAG reward_split, RewardComponentMonitor FoldReset EMA of mean |opportunity-cost reward| across batch (α=0.05). Populated by Plan 3 B.1 at the Flat-branch opp-cost site.
[68] REWARD_BONUS_EMA_INDEX f32 GPU reward_component_ema kernel (Plan 3 C.2) HEALTH_DIAG reward_split, RewardComponentMonitor FoldReset EMA of mean |bonus reward| across batch (α=0.05). Populated by Plan 3 B.2 at the Flat→Positioned site (rc[5]).
[71] TRADE_ATTEMPT_RATE_EMA_INDEX f32 GPU trade_attempt_rate_ema_update kernel (Plan 3 B.2) experience_env_step Flat→Positioned bonus path FoldReset EMA of Flat→Positioned transition rate (0..1). Adaptive α = α_base × (1 + 0.5×|clamp(sharpe, -2, 2)|). α_base = 0.05. Zero at construction and fold boundary.
[72] TRADE_TARGET_RATE_INDEX f32 CPU training_loop (epoch 5 freeze) experience_env_step novelty computation FoldReset Reference rate for novelty = max(0, 1 - attempt/target). Frozen at epoch 5 from measured TRADE_ATTEMPT_RATE_EMA (min 0.001). Pre-freeze the slot is 0 and the bonus site gates on target_raw > 1e-6f, so the reward term is structurally inert until the freeze fires.
[73] ISV_LAYOUT_FINGERPRINT_LO_INDEX u32 bits (in f32) construct check_layout_fingerprint SchemaContract Low 32 bits of u64 FNV-1a structural hash. Fail-fast on mismatch — NOT a version number, no migration path. Shifted 69→73 by Plan 3 Task 3 B.2.
[74] ISV_LAYOUT_FINGERPRINT_HI_INDEX u32 bits (in f32) construct check_layout_fingerprint SchemaContract High 32 bits of u64 FNV-1a structural hash. Shifted 70→74 by Plan 3 Task 3 B.2.
[75) (reserved for DQN v2) Allocated incrementally by Plans 3-5