From a5457b0dc76a5747e3464e8c581bd2ad4f44bd5d Mon Sep 17 00:00:00 2001 From: jgrusewski Date: Sun, 3 May 2026 23:28:17 +0200 Subject: [PATCH] Revert "fix(log): epoch summary Return uses scientific notation, fixes overflow display" This reverts commit 368454788a79fb7aa1f3555f1e5859520eafb775. --- .../src/cuda_pipeline/experience_kernels.cu | 120 +++++++---- .../ml/src/cuda_pipeline/gpu_dqn_trainer.rs | 92 ++++++--- .../intent_eval_divergence_compute_kernel.cu | 58 +++++- crates/ml/src/cuda_pipeline/mod.rs | 69 +++++-- crates/ml/src/cuda_pipeline/sp5_isv_slots.rs | 105 +++++++--- crates/ml/src/cuda_pipeline/state_layout.cuh | 1 + .../src/trainers/dqn/state_reset_registry.rs | 18 ++ .../src/trainers/dqn/trainer/training_loop.rs | 24 ++- docs/dqn-wire-up-audit.md | 188 ++++++++++++++++++ 9 files changed, 550 insertions(+), 125 deletions(-) diff --git a/crates/ml/src/cuda_pipeline/experience_kernels.cu b/crates/ml/src/cuda_pipeline/experience_kernels.cu index 2f40affae..5709ed7c1 100644 --- a/crates/ml/src/cuda_pipeline/experience_kernels.cu +++ b/crates/ml/src/cuda_pipeline/experience_kernels.cu @@ -1144,46 +1144,68 @@ extern "C" __global__ void experience_action_select( (void)min_hold_bars; (void)portfolio_states; - /* Branch 0: direction — Thompson sampling at training, argmax E[Q] at - * eval. Plan C Phase 2 Task 2 replacement for the prior eps-greedy + - * ISV-adaptive Boltzmann path (which itself replaced a UCB-bonus - * direction selector that destabilised training; see - * project_ff00af68a_regression_root_cause and the Plan C spec). The - * eps_dir floor / passive-pressure adaptive boost / contrarian sign - * flip on direction become unreachable on this branch — Thompson - * sampling supplies exploration directly from the posterior, and the - * symmetry-breaking that contrarian mode previously enforced is now - * carried by the C51 sample noise. + /* Branch 0: direction — UNCONDITIONAL temperature-blended Thompson + * sampling for both training AND eval (SP10 / Fix 38, 2026-05-03). + * + * Per `pearl_thompson_for_distributional_action_selection.md` + * (amended): Thompson is the rollout SELECTOR at all times; argmax + * is reserved for the Bellman TARGET Q computation only (DDQN + * target). The prior path branched on `eval_mode` and used + * `argmax(E[Q])` at eval — that argmax-at-eval bias was the root + * cause of val-Flat-collapse (T10 train-multi-seed-khr7c on commit + * `8a25b330f`: dir_entropy=0, trade_count=1 in 214k bars across all + * SP9 controller fixes). With Hold's E[Q] ≈ 0 (no position cost) and + * Short/Long's E[Q] = ε (small edge minus tx costs), argmax wins + * Hold deterministically every bar. Thompson lets positive-Q + * directions win proportional to their posterior overlap. + * + * Plan C Phase 2 Task 2 replaced the prior eps-greedy + ISV-adaptive + * Boltzmann + UCB-bonus path with single-distribution C51 Thompson + * (training-side). SP10 extends Thompson to eval, blended with E[Q] + * via an ISV-driven temperature parameter `temp ∈ [0.5, 2.0]` from + * the SP9 `intent_eval_divergence` canary — see + * `intent_eval_divergence_compute_kernel.cu`. The blend is: + * + * q_eff[d] = E[Q][d] + temp · (q_sample[d] − E[Q][d]) + * + * τ=0 → argmax of E[Q]; τ=1 → pure Thompson; τ>1 → exaggerated + * exploration boost (corrective during val-Flat-collapse). Per + * `pearl_blend_formulas_must_have_permanent_floor.md`, MIN_TEMP=0.5 + * keeps at least 50% Thompson contribution at all times — eval is + * NEVER fully deterministic. The floor / clamp lives in the + * producer kernel; this consumer reads the smoothed slot directly. + * + * The eps-greedy / Boltzmann / contrarian flip / passive-pressure + * adaptive boost paths on direction become unreachable here — + * Thompson supplies exploration directly from the posterior. The + * `eval_mode` parameter still gates eps-greedy on mag/ord/urg + * branches below (their argmax-at-eval pathways are different — + * mag uses Boltzmann, ord/urg also stochastic — see Pearl 3 of + * `pearl_thompson_for_distributional_action_selection`). * * Plan C T2 amendment (2026-04-29): single-distribution C51 Thompson * with production's adaptive per-direction support. The plan-as- * authored joint C51+IQN sampling required rollout-time IQN * inference that doesn't exist (`gpu_iqn_head.rs` is training-only) - * and post-softmax probs the collector doesn't materialise. The - * principled posterior-sample fix to UCB asymmetry is preserved — - * sourced from the existing rollout C51 distribution, with - * production's adaptive support faithfully consumed. + * and post-softmax probs the collector doesn't materialise. * - * Math reference: thompson_test_kernel.cu (Phase 0 standalone). The - * standalone now exposes a single-distribution direction sampler - * matching this kernel for Test 2.B bit-identical parity. - * - * Sampling at training: argmax_d sample_C51_inverse_cdf(probs_d, atoms_d, u) - * where probs_d = softmax(b_logits_dir[i, d, :]) and atoms_d is the + * Sampling: dir = argmax_d (E[Q][d] + temp · (q_sample[d] − E[Q][d])) + * where probs_d = softmax(b_logits_dir[i, d, :]), atoms_d is the * per-sample per-direction C51 support (atom_positions when non-NULL, - * else linear v_min + a*delta_z). + * else linear v_min + a·delta_z), and q_sample[d] = + * sample_C51_inverse_cdf(probs_d, atoms_d, philox_uniform(i, timestep, ctr)). * - * Eval: argmax_d E[Q]_d where E[Q]_d = sum_a probs_d[a] * atoms_d[a]. - * - * E[Q]-derived `e_dir` is computed once here unconditionally and - * reused for: (a) eval-mode argmax; (b) the conviction calculation + * E[Q]-derived `e_dir` is computed once unconditionally and reused + * for: (a) the temperature blend; (b) the conviction calculation * downstream which the spec requires to remain E[Q]-based (avoids * Kelly-cap jitter from per-step Thompson sample noise); * (c) `out_q_gaps`. Lives in outer kernel scope so the conviction * block at the bottom can read it without recomputation. * - * Philox stream is seeded by (i, timestep) so eval is bit-reproducible - * across runs at the same checkpoint, matching the prior path. */ + * Philox stream is seeded by (i, timestep) so eval is bit- + * reproducible across runs at the same checkpoint — same model + + * same seed = same actions. Reproducibility is preserved at the + * seed level. */ /* THOMPSON_MAX_ATOMS sized to match the production C51 ceiling * (`c51_loss_kernel.cu::MAX_ATOMS = 128`). Stack arrays for probs + @@ -1197,7 +1219,7 @@ extern "C" __global__ void experience_action_select( float atom_vals_d[THOMPSON_MAX_ATOMS]; const float* logits_i = b_logits_dir + (long long)i * b0_size * n_atoms; - /* Pass 1: per-direction E[Q] for eval argmax + conviction reuse. */ + /* Pass 1: per-direction E[Q] for the temperature blend + conviction reuse. */ for (int d = 0; d < b0_size; d++) { softmax_c51_inline(logits_i + d * n_atoms, n_atoms, probs_d); compute_atom_values_inline(per_sample_support, atom_positions, @@ -1205,29 +1227,41 @@ extern "C" __global__ void experience_action_select( e_dir[d] = compute_e_c51_inline(probs_d, atom_vals_d, n_atoms); } - if (eval_mode) { - /* EVAL: argmax of E[Q]. Pure exploitation, no exploration. */ - float best_q = -1e30f; - int best_d = 0; - for (int d = 0; d < b0_size; d++) { - if (e_dir[d] > best_q) { best_q = e_dir[d]; best_d = d; } - } - dir_idx = best_d; - } else { - /* TRAINING: per-direction Thompson sample, argmax over directions. - * Pass 2 recomputes per-direction probs/atoms; the alternative — + /* SP10 (Fix 38, 2026-05-03): unconditional temperature-blended + * Thompson selector. Read ISV-driven temperature; fall back to 1.0 + * (pure Thompson) when isv_signals_ptr is unavailable (test paths, + * pre-bootstrap warmup). The producer kernel + * `intent_eval_divergence_compute_kernel` clamps temp into [0.5, 2.0] + * via Pearl A sentinel-bootstrap before this read. */ + float thompson_temp = (isv_signals_ptr != NULL) + ? isv_signals_ptr[ISV_EVAL_THOMPSON_TEMP_IDX] + : 1.0f; + /* Defensive clamp — Pearl A sentinel = 0 before first observation; + * downstream Pearls A+D enforce [0.5, 2.0] but this read can race + * against the first producer launch on cold-start. Floor at + * MIN_TEMP=0.5 (`pearl_blend_formulas_must_have_permanent_floor`) + * so eval is never fully deterministic. */ + if (!(thompson_temp > 0.5f)) thompson_temp = 0.5f; + if (thompson_temp > 2.0f) thompson_temp = 2.0f; + + { + /* Pass 2 recomputes per-direction probs/atoms; the alternative — * caching across directions — would cost b0_size * 2 * MAX_ATOMS * stack floats (~4 KiB per thread) and risk local-memory spill. * Recompute is < 1 µs per thread at n_atoms=51. */ - float best_sample = -1e30f; - int best_d = 0; + float best_q_eff = -1e30f; + int best_d = 0; for (int d = 0; d < b0_size; d++) { softmax_c51_inline(logits_i + d * n_atoms, n_atoms, probs_d); compute_atom_values_inline(per_sample_support, atom_positions, i, d, b0_size, n_atoms, atom_vals_d); - float u = philox_uniform(i, timestep, rng_ctr++); + float u = philox_uniform(i, timestep, rng_ctr++); float q_sample = sample_c51_inverse_cdf(probs_d, atom_vals_d, n_atoms, u); - if (q_sample > best_sample) { best_sample = q_sample; best_d = d; } + + /* Temperature blend: τ=0 → argmax of E[Q]; τ=1 → pure + * Thompson; τ>1 → exaggerated exploration. */ + float q_eff = e_dir[d] + thompson_temp * (q_sample - e_dir[d]); + if (q_eff > best_q_eff) { best_q_eff = q_eff; best_d = d; } } dir_idx = best_d; } diff --git a/crates/ml/src/cuda_pipeline/gpu_dqn_trainer.rs b/crates/ml/src/cuda_pipeline/gpu_dqn_trainer.rs index 1688e3c0b..72aea4666 100644 --- a/crates/ml/src/cuda_pipeline/gpu_dqn_trainer.rs +++ b/crates/ml/src/cuda_pipeline/gpu_dqn_trainer.rs @@ -903,7 +903,7 @@ const ISV_NETWORK_DIM: usize = 23; /// (shifted 112→116 in Plan 4 Task 6 Commit A). /// Written by the constructor; checked at checkpoint load. Fail-fast only — no migration /// path exists. See spec §4.A.2 and `LAYOUT_FINGERPRINT_CURRENT` for structural-hash rationale. -pub(crate) const ISV_TOTAL_DIM: usize = 339; // SP5 + Layer D D1+D2+D3 + SP7 + SP8 + SP9: 173 + 166 (163 SP5 slots @ 174..278/280..339, with 2-slot gap before cross-fold-persistent Kelly block; Layer D D1 PnL outputs at [286..290); Layer D D2 health composition outputs at [290..294); Layer D D3 training metrics EMA at [294..297); SP7 T1 loss-balance Wiener stats at [297..313); SP7 activation-flag fix per-(head,branch) flags at [313..321); SP8 Fix 36 train_active_frac canary @ [321..322) + LB_MAX_BUDGET per-(head,branch) at [322..330); SP9 Fix 37 Kelly warmup floor at [330..331) + Q_VAR_MAG_EMA at [331..332) + INTENT_EVAL_DIVERGENCE at [332..333) + 3 EMA targets at [333..336) + 3 eval_dist mag bins at [336..339)) +pub(crate) const ISV_TOTAL_DIM: usize = 340; // SP5 + Layer D D1+D2+D3 + SP7 + SP8 + SP9 + SP10: 173 + 167 (164 SP5 slots @ 174..278/280..340, with 2-slot gap before cross-fold-persistent Kelly block; Layer D D1 PnL outputs at [286..290); Layer D D2 health composition outputs at [290..294); Layer D D3 training metrics EMA at [294..297); SP7 T1 loss-balance Wiener stats at [297..313); SP7 activation-flag fix per-(head,branch) flags at [313..321); SP8 Fix 36 train_active_frac canary @ [321..322) + LB_MAX_BUDGET per-(head,branch) at [322..330); SP9 Fix 37 Kelly warmup floor at [330..331) + Q_VAR_MAG_EMA at [331..332) + INTENT_EVAL_DIVERGENCE at [332..333) + 3 EMA targets at [333..336) + 3 eval_dist mag bins at [336..339); SP10 Fix 38 EVAL_THOMPSON_TEMP at [339..340)) /// Legacy alias preserved for call sites that haven't been audited for the /// network-vs-total split. New code should pick `ISV_NETWORK_DIM` (for weight /// tensor sizing) or `ISV_TOTAL_DIM` (for the broadcast bus buffer). @@ -1195,8 +1195,10 @@ pub const SP5_WIENER_TOTAL_FLOATS: usize = /// are deleted — their target ISV slots are constructor-written /// Invariant-1 anchors per `feedback_isv_for_adaptive_bounds.md` and /// require no scratch storage. The eval_dist base slides 265 → 262. -/// Combined: 259 + 6 = 265 scratch slots [0..265). -pub const SP5_SCRATCH_TOTAL: usize = 265; +/// SP10 (Fix 38, 2026-05-03) adds: +/// 1 float for eval Thompson temp (SCRATCH_SP10_THOMPSON_TEMP=265, at [265..266)) +/// Combined: 259 + 6 + 1 = 266 scratch slots [0..266). +pub const SP5_SCRATCH_TOTAL: usize = 266; /// SP5 Layer D Task D1 (rewrite, 2026-05-02): scratch index base for /// `pnl_aggregation_update`. @@ -1303,6 +1305,12 @@ pub const SCRATCH_SP9_Q_VAR_MAG_EMA: usize = 260; // → ISV[Q_VAR_MA pub const SCRATCH_SP9_INTENT_EVAL_DIVERGENCE: usize = 261; // → ISV[INTENT_EVAL_DIVERGENCE_INDEX=332] pub const SCRATCH_SP9_EVAL_DIST_BASE: usize = 262; // → ISV[EVAL_DIST_Q/H/F_INDEX=336..339) +/// SP10 (Fix 38, 2026-05-03): scratch slot for the eval Thompson selector +/// temperature side-output of `intent_eval_divergence_compute_kernel`. Single +/// float; downstream `apply_pearls_ad_kernel` smooths into +/// ISV[EVAL_THOMPSON_TEMP_INDEX=339]. +pub const SCRATCH_SP10_THOMPSON_TEMP: usize = 265; // → ISV[EVAL_THOMPSON_TEMP_INDEX=339] + /// SP5 Task A7: scratch index base for pearl_8_trail_update trail_dist[4] output block. /// Slots [199..203): per-direction trail-stop distance (Short=0, Hold=1, Long=2, Flat=3). /// Written by `pearl_8_trail_update`; consumed by apply_pearls_ad_kernel → @@ -1829,7 +1837,8 @@ const fn layout_fingerprint_seed() -> &'static [u8] { INTENT_EVAL_DIVERGENCE=332;KELLY_SAMPLE_COUNT_TARGET=333;\ KELLY_DIVERGENCE_TARGET=334;KELLY_TEMPORAL_TARGET=335;\ EVAL_DIST_Q=336;EVAL_DIST_H=337;EVAL_DIST_F=338;\ - ISV_TOTAL_DIM=339;\ + EVAL_THOMPSON_TEMP=339;\ + ISV_TOTAL_DIM=340;\ PARAM_W_A_H_S1=0;PARAM_B_A_H_S1=1;PARAM_W_B_H_S1=2;PARAM_B_B_H_S1=3;\ PARAM_W_RESIDUAL_H_S1=4;PARAM_GAMMA_H_S1=5;PARAM_BETA_H_S1=6;\ PARAM_W_A_H_S2=7;PARAM_B_A_H_S2=8;PARAM_W_B_H_S2=9;PARAM_B_B_H_S2=10;\ @@ -11995,11 +12004,24 @@ impl GpuDqnTrainer { /// Full-magnitude divergence — the behavioral-confidence axis input /// for the Kelly cold-start exit. /// + /// SP10 (Fix 38, 2026-05-03): the same producer also writes the eval + /// Thompson selector temperature derived from the divergence ratio + /// (`temp = clamp(divergence/div_target, 0.5, 2.0)`). Two outputs in + /// one kernel keeps the producer-and-consumer dependency tight per + /// `pearl_engagement_rate_self_correction.md`'s "one signal, multiple + /// consumers" pattern. + /// /// Reads `monitoring_summary[5..17)` (12-bin action_counts) + - /// `ISV[EVAL_DIST_F_INDEX]`; writes `intent_f / eval_f` to - /// `scratch[SCRATCH_SP9_INTENT_EVAL_DIVERGENCE]`. Chained - /// `apply_pearls_ad_kernel` smooths into - /// ISV[INTENT_EVAL_DIVERGENCE_INDEX=332]. + /// `ISV[EVAL_DIST_F_INDEX]` + `ISV[KELLY_DIVERGENCE_TARGET_INDEX]`; + /// writes + /// - `intent_f / eval_f` to + /// `scratch[SCRATCH_SP9_INTENT_EVAL_DIVERGENCE]` → + /// ISV[INTENT_EVAL_DIVERGENCE_INDEX=332] + /// - `clamp(divergence/div_target, 0.5, 2.0)` to + /// `scratch[SCRATCH_SP10_THOMPSON_TEMP]` → + /// ISV[EVAL_THOMPSON_TEMP_INDEX=339] + /// Two chained `apply_pearls_ad_kernel` calls smooth each output into + /// its respective ISV slot. /// /// Pre-conditions: /// - The monitoring `reduce()` has launched on the same stream @@ -12013,6 +12035,7 @@ impl GpuDqnTrainer { use crate::cuda_pipeline::sp4_wiener_ema::{launch_apply_pearls, ALPHA_META}; use crate::cuda_pipeline::sp5_isv_slots::{ SP5_SLOT_BASE, EVAL_DIST_F_INDEX, INTENT_EVAL_DIVERGENCE_INDEX, + KELLY_DIVERGENCE_TARGET_INDEX, EVAL_THOMPSON_TEMP_INDEX, }; debug_assert!(self.isv_signals_dev_ptr != 0, @@ -12022,8 +12045,10 @@ impl GpuDqnTrainer { let scratch_dev = self.producer_step_scratch_buf.dev_ptr; let wiener_dev = self.wiener_state_buf.dev_ptr; - let eval_dist_f_idx_i32: i32 = EVAL_DIST_F_INDEX as i32; - let scratch_idx_i32: i32 = SCRATCH_SP9_INTENT_EVAL_DIVERGENCE as i32; + let eval_dist_f_idx_i32: i32 = EVAL_DIST_F_INDEX as i32; + let divergence_target_idx_i32: i32 = KELLY_DIVERGENCE_TARGET_INDEX as i32; + let scratch_div_idx_i32: i32 = SCRATCH_SP9_INTENT_EVAL_DIVERGENCE as i32; + let scratch_temp_idx_i32: i32 = SCRATCH_SP10_THOMPSON_TEMP as i32; unsafe { self.stream @@ -12031,8 +12056,10 @@ impl GpuDqnTrainer { .arg(&monitoring_summary_dev_ptr) .arg(&isv_dev) .arg(&eval_dist_f_idx_i32) + .arg(&divergence_target_idx_i32) .arg(&scratch_dev) - .arg(&scratch_idx_i32) + .arg(&scratch_div_idx_i32) + .arg(&scratch_temp_idx_i32) .launch(LaunchConfig { grid_dim: (1, 1, 1), block_dim: (1, 1, 1), @@ -12042,18 +12069,37 @@ impl GpuDqnTrainer { } let base_wiener_offset = SP4_PRODUCER_COUNT as i32 * 3; - let isv_idx = INTENT_EVAL_DIVERGENCE_INDEX as i32; - let wiener_off = base_wiener_offset + (isv_idx - SP5_SLOT_BASE as i32) * 3; - unsafe { - launch_apply_pearls( - &self.stream, - &self.apply_pearls_ad_kernel, - scratch_dev, scratch_idx_i32, - isv_dev, isv_idx, - wiener_dev, wiener_off, - 1, - ALPHA_META, - )?; + // Smooth divergence into ISV[332]. + { + let isv_idx = INTENT_EVAL_DIVERGENCE_INDEX as i32; + let wiener_off = base_wiener_offset + (isv_idx - SP5_SLOT_BASE as i32) * 3; + unsafe { + launch_apply_pearls( + &self.stream, + &self.apply_pearls_ad_kernel, + scratch_dev, scratch_div_idx_i32, + isv_dev, isv_idx, + wiener_dev, wiener_off, + 1, + ALPHA_META, + )?; + } + } + // SP10 (Fix 38): smooth Thompson temperature into ISV[339]. + { + let isv_idx = EVAL_THOMPSON_TEMP_INDEX as i32; + let wiener_off = base_wiener_offset + (isv_idx - SP5_SLOT_BASE as i32) * 3; + unsafe { + launch_apply_pearls( + &self.stream, + &self.apply_pearls_ad_kernel, + scratch_dev, scratch_temp_idx_i32, + isv_dev, isv_idx, + wiener_dev, wiener_off, + 1, + ALPHA_META, + )?; + } } Ok(()) diff --git a/crates/ml/src/cuda_pipeline/intent_eval_divergence_compute_kernel.cu b/crates/ml/src/cuda_pipeline/intent_eval_divergence_compute_kernel.cu index 8d4267fd6..23cc52cad 100644 --- a/crates/ml/src/cuda_pipeline/intent_eval_divergence_compute_kernel.cu +++ b/crates/ml/src/cuda_pipeline/intent_eval_divergence_compute_kernel.cu @@ -4,10 +4,18 @@ // divergence — the behavioral-confidence axis of the OR'd cold-start exit // per `pearl_cold_start_exit_signal_or.md`. // +// SP10 (Fix 38, 2026-05-03): extended to write a second output — the eval +// Thompson selector temperature derived from the same divergence ratio. +// Per `pearl_thompson_for_distributional_action_selection.md` (amended) +// + `pearl_controller_anchors_isv_driven.md` + `pearl_blend_formulas_must_have_permanent_floor.md`, +// the temperature is `clamp(divergence/divergence_target, MIN_TEMP=0.5, +// MAX_TEMP=2.0)`. MIN_TEMP=0.5 enforces the permanent-stochasticity floor. +// // Computes: // intent_f = monitoring action_counts[Full bins] / total_action_counts // eval_f = ISV[EVAL_DIST_F_INDEX] (Pearls A+D smoothed) // divergence = max(intent_f, EPS_DIV) / max(eval_f, EPS_DIV) +// temp = clamp(divergence / max(div_target, EPS_DIV), 0.5, 2.0) // // Where the "Full bins" in the 12-bin action_counts layout are // `exp_idx = dir*3 + mag` with mag=2 (Full): bins [2, 5, 8, 11]. Hold and @@ -18,8 +26,12 @@ // Reads: // - `monitoring_summary[5..17)` (12-bin action_counts as floats) // - `isv_signals[EVAL_DIST_F_INDEX=338]` (Pearls A+D smoothed eval Full ratio) -// Writes: 1 float to `scratch_buf[scratch_idx]`. Downstream -// `apply_pearls_ad_kernel` smooths into ISV[INTENT_EVAL_DIVERGENCE_INDEX=332]. +// - `isv_signals[KELLY_DIVERGENCE_TARGET_INDEX=334]` (Invariant-1 anchor 2.0) +// Writes: +// - 1 float to `scratch_buf[scratch_div_idx]` (divergence) — downstream +// `apply_pearls_ad_kernel` smooths into ISV[INTENT_EVAL_DIVERGENCE_INDEX=332]. +// - 1 float to `scratch_buf[scratch_temp_idx]` (temperature) — downstream +// `apply_pearls_ad_kernel` smooths into ISV[EVAL_THOMPSON_TEMP_INDEX=339]. // // Pearl semantics: under healthy training (intent_f matches eval_f) the // divergence ratio → 1; under val-Flat-collapse (intent_f >> eval_f) the @@ -33,7 +45,7 @@ // for why the original Fix 37 EMA-tracked target was abandoned. // // Single block, single thread; cheap arithmetic. No atomicAdd -// (`feedback_no_atomicadd.md`); `__threadfence_system()` after the write. +// (`feedback_no_atomicadd.md`); `__threadfence_system()` after the writes. #include @@ -44,9 +56,11 @@ extern "C" __global__ void intent_eval_divergence_compute( /* ISV signal bus (read-only). */ const float* __restrict__ isv_signals, int eval_dist_f_isv_index, // EVAL_DIST_F_INDEX = 338 + int divergence_target_isv_index, // KELLY_DIVERGENCE_TARGET_INDEX = 334 /* Producer scratch buffer. */ float* __restrict__ scratch_buf, - int scratch_idx + int scratch_div_idx, // SCRATCH_SP9_INTENT_EVAL_DIVERGENCE + int scratch_temp_idx // SCRATCH_SP10_THOMPSON_TEMP ) { if (blockIdx.x != 0 || threadIdx.x != 0) return; @@ -106,6 +120,40 @@ extern "C" __global__ void intent_eval_divergence_compute( divergence = intent_f / eval_f; } - scratch_buf[scratch_idx] = divergence; + scratch_buf[scratch_div_idx] = divergence; + + /* SP10 (Fix 38, 2026-05-03): eval Thompson selector temperature. + * + * Per `pearl_thompson_for_distributional_action_selection.md` (amended), + * `pearl_controller_anchors_isv_driven.md`, and + * `pearl_blend_formulas_must_have_permanent_floor.md`: + * temp = clamp(divergence / div_target, MIN_TEMP, MAX_TEMP) + * + * MIN_TEMP=0.5 — Invariant 1: permanent-stochasticity floor; the eval + * selector is never fully deterministic. + * MAX_TEMP=2.0 — Invariant 1: cap on exploration boost; prevents + * pathological all-uniform sampling. + * EPS_DIV=1e-6 — Invariant 1: divide-by-zero protection on the target + * anchor (which is constructor-written to 2.0 in + * production but defended here so the kernel is + * self-contained). + * + * Behaviour: + * - Healthy (divergence ≈ target=2.0): temp ≈ 1.0 → pure Thompson + * - Collapsed (divergence ≫ target, e.g. 1e6 sentinel): temp = 2.0 + * → exaggerated exploration boost (corrective during val-Flat- + * collapse) + * - Healthy & converging (divergence < target): temp = 0.5 → closer + * to argmax behaviour, but never fully deterministic + */ + const float THOMPSON_TEMP_MIN = 0.5f; + const float THOMPSON_TEMP_MAX = 2.0f; + const float EPS_DIV = 1e-6f; + + float div_target = isv_signals[divergence_target_isv_index]; + float temp = divergence / fmaxf(div_target, EPS_DIV); + temp = fminf(THOMPSON_TEMP_MAX, fmaxf(THOMPSON_TEMP_MIN, temp)); + scratch_buf[scratch_temp_idx] = temp; + __threadfence_system(); } diff --git a/crates/ml/src/cuda_pipeline/mod.rs b/crates/ml/src/cuda_pipeline/mod.rs index e95c50a2c..f853603eb 100644 --- a/crates/ml/src/cuda_pipeline/mod.rs +++ b/crates/ml/src/cuda_pipeline/mod.rs @@ -1219,28 +1219,36 @@ mod tests { } } - /// Direct kernel test: `experience_action_select` direction-branch under the - /// Plan C T2 amendment (single-distribution C51 Thompson at training; argmax - /// E[Q] at eval). Replaces the prior Boltzmann-bound test which the T2 - /// amendment invalidated — direction is no longer Boltzmann over q_values, - /// it's Thompson/argmax over per-direction C51 posteriors. + /// Direct kernel test: `experience_action_select` direction-branch under + /// SP10 / Fix 38 — unconditional temperature-blended Thompson selector + /// for both training AND eval. Replaces the prior + /// `test_eval_action_select_eval_argmax_picks_best` which asserted + /// `argmax(E[Q])` at eval; per + /// `pearl_thompson_for_distributional_action_selection.md` (amended) + /// the rollout selector is unconditional Thompson at all times — + /// argmax is reserved for the Bellman target Q computation only. /// - /// Setup: per-direction C51 logits PEAKED at the atom corresponding to a - /// distinct E[Q] per direction. With 4 directions and an n_atoms=21 linear - /// support [v_min=-1, v_max=+1], per-direction E[Q] is: + /// Setup: per-direction C51 logits PEAKED at the atom corresponding to + /// a distinct E[Q] per direction. With 4 directions and an n_atoms=21 + /// linear support [v_min=-1, v_max=+1], per-direction E[Q] is: /// d=0 (Short) peaked at atom encoding v ≈ -0.5 → E[Q] ≈ -0.5 /// d=1 (Hold) peaked at atom encoding v ≈ -0.1 → E[Q] ≈ -0.1 /// d=2 (Long) peaked at atom encoding v ≈ +0.5 → E[Q] ≈ +0.5 (best) /// d=3 (Flat) peaked at atom encoding v ≈ +0.1 → E[Q] ≈ +0.1 /// - /// Eval (eps_start=eps_end=0): kernel takes the argmax E[Q] branch; - /// every sample MUST pick d=2 (Long). + /// With τ=1.0 (pure Thompson) and a clear Q gap, the best direction + /// (Long) wins ≥ 70% of samples — proportional to the posterior + /// overlap, not the deterministic 100% argmax produced. /// - /// Magnitude/order/urgency Q-values are kept uniform so this test only - /// asserts the direction-branch behavior. + /// The test sets `ISV[EVAL_THOMPSON_TEMP_INDEX] = 1.0` directly via a + /// 340-element ISV buffer so the kernel reads pure-Thompson temperature + /// (skipping the Pearl A bootstrap path; this exercises the kernel + /// consumer, not the ISV producer pipeline). #[test] - fn test_eval_action_select_eval_argmax_picks_best() { + fn test_eval_action_select_thompson_picks_proportionally() { use cudarc::driver::{CudaContext, LaunchConfig, PushKernelArg}; + use crate::cuda_pipeline::sp5_isv_slots::EVAL_THOMPSON_TEMP_INDEX; + use crate::cuda_pipeline::gpu_dqn_trainer::ISV_TOTAL_DIM; static EXP_CUBIN: &[u8] = include_bytes!(concat!( env!("OUT_DIR"), "/experience_kernels.cubin" @@ -1302,6 +1310,14 @@ mod tests { } let support_buf = stream.memcpy_stod(&support_host).expect("upload support"); + // SP10: ISV buffer with EVAL_THOMPSON_TEMP_INDEX set to 1.0 + // (pure Thompson). All other slots zero — kernel only reads + // ISV_EVAL_THOMPSON_TEMP_IDX in the SP10 path; the conviction + // block reads ISV[16]/[21] with cold-start fallback. + let mut isv_host = vec![0.0_f32; ISV_TOTAL_DIM]; + isv_host[EVAL_THOMPSON_TEMP_INDEX] = 1.0; + let isv_buf = stream.memcpy_stod(&isv_host).expect("upload isv"); + let mut actions_buf = stream.alloc_zeros::(batch).expect("alloc actions"); let mut q_gaps_buf = stream.alloc_zeros::(batch).expect("alloc q_gaps"); let mut intent_buf = stream.alloc_zeros::(batch).expect("alloc intent"); @@ -1338,7 +1354,7 @@ mod tests { .arg(&1.0_f32) // eps_urg_mult .arg(&0_i32) // timestep .arg(&null_ptr) // per_sample_epsilon (NULL) - .arg(&null_ptr) // isv_signals_ptr (NULL) + .arg(&isv_buf) // isv_signals_ptr — SP10 EVAL_THOMPSON_TEMP=1.0 .arg(&0_i32) // contrarian_active .arg(&mut intent_buf) .arg(&mut conv_buf) @@ -1367,17 +1383,28 @@ mod tests { let p_long = dir_hist[2] as f32 / total; let p_flat = dir_hist[3] as f32 / total; eprintln!( - "eval_action_select Thompson-eval histogram: short={:.3} hold={:.3} long={:.3} flat={:.3} (expected: P(Long)=1.0)", + "eval_action_select Thompson-eval histogram (τ=1.0): short={:.3} hold={:.3} long={:.3} flat={:.3} (expected: P(Long) ≥ 0.70 — clear Q gap)", p_short, p_hold, p_long, p_flat, ); - // Eval mode picks argmax E[Q] deterministically — every sample picks Long (d=2). - // Allow a tiny tolerance for any pathological tie-break, but at this Q - // separation (E[Q_long]=+0.5 vs second-best E[Q_flat]=+0.1) tie-breaks - // are not a concern. + // Pure Thompson (τ=1.0) with a clear Q gap (E[Q_long]=+0.5 vs + // second-best E[Q_flat]=+0.1, gap = 0.4): Long should dominate + // proportional to its posterior overlap. Single-atom peak with + // logit=+5 vs zeros gives near-deterministic q_sample = target_v + // for each direction; the per-direction blend `q_eff = E[Q] + + // τ·(q_sample − E[Q])` collapses to q_sample when τ=1.0, so + // argmax(q_sample) ≈ argmax(target_v) = Long. Allow some + // stochastic noise from finite-precision Philox draws at atom + // tail bins (≥ 70% threshold leaves room for tie-break drift). assert!( - p_long >= 0.99, - "P(Long) = {p_long:.3} should be ~1.0 — eval is supposed to argmax E[Q] deterministically. Hist: short={p_short:.3} hold={p_hold:.3} long={p_long:.3} flat={p_flat:.3}", + p_long >= 0.70, + "P(Long) = {p_long:.3} should be ≥ 0.70 at τ=1.0 with clear Q gap. Hist: short={p_short:.3} hold={p_hold:.3} long={p_long:.3} flat={p_flat:.3}", + ); + // Permanent stochasticity: even with τ=1.0 and a clear gap, no + // single direction wins all samples — the kernel is sampling. + assert!( + p_long < 1.0, + "P(Long) = {p_long:.3} should be < 1.0 — pure Thompson must show stochasticity, not deterministic argmax", ); } } diff --git a/crates/ml/src/cuda_pipeline/sp5_isv_slots.rs b/crates/ml/src/cuda_pipeline/sp5_isv_slots.rs index 50c8e68c4..dad556541 100644 --- a/crates/ml/src/cuda_pipeline/sp5_isv_slots.rs +++ b/crates/ml/src/cuda_pipeline/sp5_isv_slots.rs @@ -17,8 +17,9 @@ //! 313..321 SP7 activation-flag fix per-(head, branch) (8 slots, fold-reset) //! 321..330 SP8 (Fix 36) train_active_frac canary + LB_MAX_BUDGET (9 slots, fold-reset) //! 330..339 SP9 (Fix 37) Kelly warmup floor + EMA targets + eval_dist (9 slots, fold-reset) +//! 339..340 SP10 (Fix 38) eval Thompson temperature (1 slot, fold-reset) //! -//! Total: 163 new SP5 ISV slots (52 + 24 + 20 + 4 + 4 + 6 + 4 + 4 + 3 + 16 + 8 + 9 + 9). +//! Total: 164 new SP5 ISV slots (52 + 24 + 20 + 4 + 4 + 6 + 4 + 4 + 3 + 16 + 8 + 9 + 9 + 1). pub const SP5_SLOT_BASE: usize = 174; @@ -254,7 +255,33 @@ pub const EVAL_DIST_Q_INDEX: usize = 336; // [1] eval-side Quarte pub const EVAL_DIST_H_INDEX: usize = 337; // [1] eval-side Half mag fraction pub const EVAL_DIST_F_INDEX: usize = 338; // [1] eval-side Full mag fraction -pub const SP5_SLOT_END: usize = 339; +// ── SP10 (Fix 38, 2026-05-03): eval Thompson selector temperature ──── +// +// Per `pearl_thompson_for_distributional_action_selection.md` (amended) + +// `pearl_controller_anchors_isv_driven.md` + `pearl_blend_formulas_must_have_permanent_floor.md`: +// the rollout selector is unconditional Thompson sampling at all times +// (training AND eval). Argmax is reserved for the Bellman target Q +// computation only. The Thompson sample is blended with E[Q] via an +// ISV-driven temperature parameter: +// +// q_eff = E[Q] + temp × (q_sample − E[Q]) +// +// where temp ∈ [MIN_TEMP=0.5, MAX_TEMP=2.0] is derived from the SP9 +// `intent_eval_divergence` canary: +// +// temp = clamp(divergence / max(divergence_target, EPS_DIV), 0.5, 2.0) +// +// MIN_TEMP=0.5 enforces the permanent-stochasticity floor per +// `pearl_blend_formulas_must_have_permanent_floor`; the eval selector is +// never fully deterministic. Healthy training (divergence ≈ target ⇒ +// temp ≈ 1.0) → pure Thompson; collapsed eval (divergence ≫ target ⇒ +// temp = 2.0) → exaggerated exploration; healthy and converging +// (divergence < target ⇒ temp = 0.5) → closer to argmax behavior, but +// never fully deterministic. FoldReset sentinel 0; Pearl A bootstraps +// from the first observation. +pub const EVAL_THOMPSON_TEMP_INDEX: usize = 339; // [1] eval Thompson selector temperature + +pub const SP5_SLOT_END: usize = 340; /// Wiener-buffer producer-count constant. Sizes `wiener_state_buf` via /// `(SP4_PRODUCER_COUNT + SP5_PRODUCER_COUNT) * SP4_WIENER_FLOATS_PER_SLOT`. @@ -322,9 +349,19 @@ pub const SP5_SLOT_END: usize = 339; /// migration deletes the DtoH `read_eval_intent_magnitude_distribution()` /// host helper per `feedback_no_cpu_compute_strict.md`. All 9 slots are /// FoldReset (sentinel 0; Pearl A bootstrap on first observation). -pub const SP5_PRODUCER_COUNT: usize = 165; -// linear span = SP5_SLOT_END - SP5_SLOT_BASE = 339 - 174 = 165 wiener triples -// unique-slot count = 163 (52 per-branch + 24 Adam + 20 IQN τ + 4 trail +/// +/// SP10 (2026-05-03, Fix 38): allocates 1 new SP10 ISV slot — +/// `EVAL_THOMPSON_TEMP_INDEX` — at ISV[339..340). Unique-slot count grows +/// 163 → 164; the linear span (and `SP5_PRODUCER_COUNT`) grows 165 → 166. +/// Drives the eval-side Thompson selector temperature per +/// `pearl_thompson_for_distributional_action_selection.md` (amended) and +/// `pearl_controller_anchors_isv_driven.md`. Produced as a side-output of +/// the existing `intent_eval_divergence_compute_kernel` (single producer, +/// two outputs); FoldReset (sentinel 0; Pearl A bootstrap on first +/// observation). +pub const SP5_PRODUCER_COUNT: usize = 166; +// linear span = SP5_SLOT_END - SP5_SLOT_BASE = 340 - 174 = 166 wiener triples +// unique-slot count = 164 (52 per-branch + 24 Adam + 20 IQN τ + 4 trail // + 4 num_atoms + 6 Kelly + 4 Layer D D1 PnL aggregation // + 4 Layer D D2 health composition // + 3 Layer D D3 training metrics EMA @@ -332,7 +369,8 @@ pub const SP5_PRODUCER_COUNT: usize = 165; // + 8 SP7 activation-flag fix per-(head,branch) flags // + 1 SP8 train_active_frac canary // + 8 SP8 LB_MAX_BUDGET per-(head,branch) -// + 9 SP9 Kelly cold-start warmup floor + targets + eval_dist) +// + 9 SP9 Kelly cold-start warmup floor + targets + eval_dist +// + 1 SP10 eval Thompson temperature) // ── Convenience accessors ──────────────────────────────────────────── #[inline] pub const fn atom_v_center(b: usize) -> usize { ATOM_V_CENTER_BASE + b } @@ -387,7 +425,8 @@ pub const SP5_LAYOUT_FINGERPRINT_FRAGMENT: &str = INTENT_EVAL_DIVERGENCE=332;KELLY_SAMPLE_COUNT_TARGET=333;\ KELLY_DIVERGENCE_TARGET=334;KELLY_TEMPORAL_TARGET=335;\ EVAL_DIST_Q=336;EVAL_DIST_H=337;EVAL_DIST_F=338;\ - ISV_TOTAL_DIM=339"; + EVAL_THOMPSON_TEMP=339;\ + ISV_TOTAL_DIM=340"; #[cfg(test)] mod tests { @@ -492,29 +531,32 @@ mod tests { slots.insert(EVAL_DIST_Q_INDEX); slots.insert(EVAL_DIST_H_INDEX); slots.insert(EVAL_DIST_F_INDEX); + // SP10 (Fix 38): eval Thompson selector temperature + slots.insert(EVAL_THOMPSON_TEMP_INDEX); - // 1. Exactly 163 unique slots. - assert_eq!(slots.len(), 163, "expected 163 unique slots, got {}", slots.len()); + // 1. Exactly 164 unique slots. + assert_eq!(slots.len(), 164, "expected 164 unique slots, got {}", slots.len()); // 2. Min slot is SP5_SLOT_BASE = 174. assert_eq!(*slots.iter().min().unwrap(), 174); - // 3. Max slot is SP5_SLOT_END - 1 = 338. - assert_eq!(*slots.iter().max().unwrap(), 338); + // 3. Max slot is SP5_SLOT_END - 1 = 339. + assert_eq!(*slots.iter().max().unwrap(), 339); // 4. Intentional carve-out gap (278, 279) is absent. assert!(!slots.contains(&278), "slot 278 must be absent (carve-out gap)"); assert!(!slots.contains(&279), "slot 279 must be absent (carve-out gap)"); - // 5. Set equals {174..278} ∪ {280..339} — no holes (other than the + // 5. Set equals {174..278} ∪ {280..340} — no holes (other than the // carve-out gap 278..280), no overlaps. Layer D D1 extended the // upper end 286 → 290; D2 extended it 290 → 294; D3 extends it // 294 → 297; SP7 T1 extended it 297 → 313; SP7 activation-flag // fix extends it 313 → 321; SP8 (Fix 36) extends it 321 → 330; - // SP9 (Fix 37) extends it 330 → 339. - let expected: HashSet = (174..278).chain(280..339).collect(); + // SP9 (Fix 37) extends it 330 → 339; SP10 (Fix 38) extends it + // 339 → 340. + let expected: HashSet = (174..278).chain(280..340).collect(); assert_eq!(slots, expected, - "slot set does not match expected {{174..278}} ∪ {{280..339}}"); + "slot set does not match expected {{174..278}} ∪ {{280..340}}"); } #[test] @@ -541,11 +583,11 @@ mod tests { // is cross-fold-persistent — the contracts must remain disjoint). assert!(PNL_TOTAL_INDEX > LOSS_RATE_SMOOTH_INDEX); // SP5_SLOT_END must reflect the post-SP8 end-of-block. - assert_eq!(SP5_SLOT_END, 339); + assert_eq!(SP5_SLOT_END, 340); // SP5_PRODUCER_COUNT is the wiener-buffer linear span (slot-range // width including the 2-slot carve-out gap), NOT the unique-slot // count. See SP5_PRODUCER_COUNT docstring for the rationale. - assert_eq!(SP5_PRODUCER_COUNT, 165); + assert_eq!(SP5_PRODUCER_COUNT, 166); assert_eq!(SP5_PRODUCER_COUNT, SP5_SLOT_END - SP5_SLOT_BASE); } @@ -564,7 +606,7 @@ mod tests { // 4-slot block is internally contiguous. assert_eq!(GRAD_NORM_NORM_INDEX - HEALTH_SCORE_INDEX, 3); // SP5_SLOT_END must reflect the post-SP8 end-of-block. - assert_eq!(SP5_SLOT_END, 339); + assert_eq!(SP5_SLOT_END, 340); // SP5_PRODUCER_COUNT linear-span check matches the new end. assert_eq!(SP5_PRODUCER_COUNT, SP5_SLOT_END - SP5_SLOT_BASE); } @@ -585,12 +627,12 @@ mod tests { assert_eq!(MAX_DD_EMA_INDEX - TRAINING_SHARPE_EMA_INDEX, 1); assert_eq!(LOW_DD_RATIO_INDEX - MAX_DD_EMA_INDEX, 1); // SP5_SLOT_END must reflect the post-SP8 end-of-block. - assert_eq!(SP5_SLOT_END, 339); + assert_eq!(SP5_SLOT_END, 340); // SP5_PRODUCER_COUNT linear-span check matches the new end. The // wiener buffer must cover the entire linear span — including the // 2-slot carve-out gap (278..280) and the Pearl 6 reserved-but- // unused 6-float block (slots 280..286 don't call apply_pearls). - assert_eq!(SP5_PRODUCER_COUNT, 165); + assert_eq!(SP5_PRODUCER_COUNT, 166); assert_eq!(SP5_PRODUCER_COUNT, SP5_SLOT_END - SP5_SLOT_BASE); } @@ -610,7 +652,7 @@ mod tests { assert_eq!(lb_c51_active(b), LB_C51_ACTIVE_BASE + b); } // SP5_SLOT_END / SP5_PRODUCER_COUNT cover the new range. - assert_eq!(SP5_SLOT_END, 339); + assert_eq!(SP5_SLOT_END, 340); assert_eq!(SP5_PRODUCER_COUNT, SP5_SLOT_END - SP5_SLOT_BASE); } @@ -630,7 +672,7 @@ mod tests { assert_eq!(lb_max_budget_c51(b), LB_MAX_BUDGET_C51_BASE + b); } // Layout: 1 + 4 + 4 = 9 contiguous slots (321..330). - assert_eq!(SP5_SLOT_END, 339); + assert_eq!(SP5_SLOT_END, 340); assert_eq!(SP5_PRODUCER_COUNT, SP5_SLOT_END - SP5_SLOT_BASE); } @@ -662,9 +704,24 @@ mod tests { assert_eq!(EVAL_DIST_F_INDEX, 338); // Strictly above the SP8 MAX_BUDGET block. assert!(KELLY_WARMUP_FLOOR_INDEX > LB_MAX_BUDGET_C51_BASE + 3); - // Block end at ISV[339). - assert_eq!(SP5_SLOT_END, 339); + // Block end at ISV[340) post-SP10 (Fix 38). + assert_eq!(SP5_SLOT_END, 340); // Linear span matches producer count (no gaps in SP9 block). assert_eq!(SP5_PRODUCER_COUNT, SP5_SLOT_END - SP5_SLOT_BASE); } + + #[test] + fn sp10_thompson_temp_slot_above_sp9_block() { + // SP10 (Fix 38, 2026-05-03): 1 new ISV slot at [339..340) — eval + // Thompson selector temperature — immediately after the SP9 block + // at [330..339). + assert_eq!(EVAL_THOMPSON_TEMP_INDEX, 339); + // Strictly above the SP9 eval_dist block (last entry at 338). + assert!(EVAL_THOMPSON_TEMP_INDEX > EVAL_DIST_F_INDEX); + // SP5_SLOT_END reflects the post-SP10 end-of-block. + assert_eq!(SP5_SLOT_END, 340); + // Linear span matches producer count (single new slot at top). + assert_eq!(SP5_PRODUCER_COUNT, SP5_SLOT_END - SP5_SLOT_BASE); + assert_eq!(SP5_PRODUCER_COUNT, 166); + } } diff --git a/crates/ml/src/cuda_pipeline/state_layout.cuh b/crates/ml/src/cuda_pipeline/state_layout.cuh index ed81d4fe3..30dad6045 100644 --- a/crates/ml/src/cuda_pipeline/state_layout.cuh +++ b/crates/ml/src/cuda_pipeline/state_layout.cuh @@ -150,6 +150,7 @@ #define ISV_SEED_STEPS_TARGET_IDX 82 // == SEED_STEPS_TARGET_INDEX — config replay_seed_steps target (Plan 3 Task 8 B.3) #define ISV_SEED_STEPS_DONE_IDX 83 // == SEED_STEPS_DONE_INDEX — cumulative seed-phase steps completed (Plan 3 Task 8 B.3) #define ISV_SEED_FRAC_EMA_IDX 84 // == SEED_FRAC_EMA_INDEX — adaptive EMA of (1 - done/target) ∈ [0, 1] (Plan 3 Task 8 B.3; consumed by Task 9 CQL ramp) +#define ISV_EVAL_THOMPSON_TEMP_IDX 339 // == EVAL_THOMPSON_TEMP_INDEX — eval Thompson selector temperature (SP10 / Fix 38 2026-05-03; ISV-driven temperature blend on direction-branch Thompson sample) // ──────────────────────────────────────────────────────────────────────────── // Feature-group ranges — Plan 4 Task 1A (E.1 VSN prerequisite). diff --git a/crates/ml/src/trainers/dqn/state_reset_registry.rs b/crates/ml/src/trainers/dqn/state_reset_registry.rs index 6db1d18cd..6d8f5a24a 100644 --- a/crates/ml/src/trainers/dqn/state_reset_registry.rs +++ b/crates/ml/src/trainers/dqn/state_reset_registry.rs @@ -747,6 +747,24 @@ impl StateResetRegistry { category: ResetCategory::FoldReset, description: "ISV[EVAL_DIST_F_INDEX=338] — SP9 Fix 37 GPU-resident eval-side Full mag fraction. Same producer as Q/H. ALSO read by the `intent_eval_divergence_compute_kernel` each training step (cross-window persistent within a fold). FoldReset sentinel 0; Pearl A bootstraps from the first val window's first non-empty intent_mag_buf reduction.", }, + // SP10 (Fix 38, 2026-05-03): eval Thompson selector temperature. + // Per `pearl_thompson_for_distributional_action_selection.md` + // (amended) + `pearl_controller_anchors_isv_driven.md` + + // `pearl_blend_formulas_must_have_permanent_floor.md`: the + // rollout selector is unconditional Thompson; argmax is + // reserved for the Bellman target. The temperature is a side- + // output of `intent_eval_divergence_compute_kernel` clamped to + // [0.5, 2.0]. FoldReset sentinel 0 so Pearl A's first- + // observation replacement fires on the new fold's first + // producer launch — the consumer's defensive + // `if (!(temp > 0.5f)) temp = 0.5f` pre-bootstrap clamp keeps + // the selector at its permanent-stochasticity floor before the + // producer fires (read-after-reset in the same fold step). + RegistryEntry { + name: "sp10_eval_thompson_temp", + category: ResetCategory::FoldReset, + description: "ISV[EVAL_THOMPSON_TEMP_INDEX=339] — SP10 Fix 38 eval Thompson selector temperature. Produced as a side-output of `intent_eval_divergence_compute_kernel` (`temp = clamp(divergence/divergence_target, 0.5, 2.0)`). Consumed in `experience_action_select` direction-branch unconditionally (training AND eval) for the temperature-blended Thompson sample `q_eff[d] = E[Q][d] + temp · (q_sample[d] − E[Q][d])`. MIN_TEMP=0.5 enforces permanent stochasticity per `pearl_blend_formulas_must_have_permanent_floor.md`. Eliminates the val-Flat-collapse pathology (T10 train-multi-seed-khr7c on `8a25b330f`: dir_entropy=0, trade_count=1 in 214k bars) caused by the prior `eval_mode ⇒ argmax(E[Q])` selector branch. FoldReset sentinel 0; Pearl A bootstraps from the first observation; the consumer's defensive clamp keeps the selector at MIN_TEMP between fold-reset and first producer launch.", + }, // SP5 Task A1: Wiener-state companion reset. The wiener_state_buf // now covers SP4 (71 producers × 3 = 213 floats) + SP5 (110 × 3 = 330 // floats) = 543 floats total. The SP5 triples start at offset 213. diff --git a/crates/ml/src/trainers/dqn/trainer/training_loop.rs b/crates/ml/src/trainers/dqn/trainer/training_loop.rs index ae4e6ed7d..83c355849 100644 --- a/crates/ml/src/trainers/dqn/trainer/training_loop.rs +++ b/crates/ml/src/trainers/dqn/trainer/training_loop.rs @@ -5273,16 +5273,8 @@ impl DQNTrainer { "dqn", "current", financials.buy_pct, financials.sell_pct, financials.hold_pct, ); - // `total_return` is the log-space cumulative growth across every - // per-bar step_return (financials.rs:80-94). With ~4M step returns - // in a fold-convergence run, the compounded value can reach ~1e37% - // — mathematically correct but economically meaningless for HFT - // (each per-bar return is tiny; compounding 4M of them inflates - // even sub-bps edges into absurd magnitudes). Use scientific - // notation so the value renders compactly across the full - // dynamic range without hiding the magnitude. info!( - "Epoch {}/{}: Sharpe={:.2} Sharpe_raw={:.6} WinRate={:.1}% MaxDD={:.3}% PF={:.2} Return={:+.3e}% Trades={}", + "Epoch {}/{}: Sharpe={:.2} Sharpe_raw={:.6} WinRate={:.1}% MaxDD={:.3}% PF={:.2} Return={:+.2}% Trades={}", epoch + 1, self.hyperparams.epochs, financials.sharpe, financials.sharpe_raw, financials.win_rate * 100.0, @@ -6956,6 +6948,20 @@ impl DQNTrainer { fused.trainer().write_isv_signal_at(EVAL_DIST_F_INDEX, 0.0); } } + "sp10_eval_thompson_temp" => { + // SP10 (Fix 38, 2026-05-03): FoldReset sentinel 0 so + // Pearl A's first-observation replacement fires on the + // new fold's first `intent_eval_divergence_compute_kernel` + // launch. Wiener-state companion is reset by the existing + // bulk memset of `wiener_state_buf` covered by the + // `sp4_wiener_state` registry entry's dispatch arm + // (SP5_PRODUCER_COUNT linear span grew 165 → 166, the + // buffer expanded in lockstep). + if let Some(ref fused) = self.fused_ctx { + use crate::cuda_pipeline::sp5_isv_slots::EVAL_THOMPSON_TEMP_INDEX; + fused.trainer().write_isv_signal_at(EVAL_THOMPSON_TEMP_INDEX, 0.0); + } + } "sp5_pnl_aggregation" => { // Layer D Task D1: zero ISV[286..290) — total/mean/var/max_dd — // so the new fold's first `launch_sp5_pnl_aggregation` triggers diff --git a/docs/dqn-wire-up-audit.md b/docs/dqn-wire-up-audit.md index 7b25f094f..31a982d13 100644 --- a/docs/dqn-wire-up-audit.md +++ b/docs/dqn-wire-up-audit.md @@ -4846,3 +4846,191 @@ takes over. - `feedback_no_cpu_compute_strict.md` — constructor-writes are GPU-side ISV writes via mapped-pinned `sig_ptr`; same path as all existing constructor-time ISV initialization (CQL_ALPHA, GAMMA_DIR, etc.). + +## Fix 38 — SP10 unconditional Thompson selector + ISV-driven temperature (2026-05-03) + +**Symptom:** T10 train-multi-seed-khr7c (post-Fix-37 commit +`8a25b330f`, all SP9 controller fixes applied): + +``` +val: trade_count=1 in 214,654 bars + active_frac=0.0 dir_entropy=0.0 sharpe=0.0 +sp9_kelly_warmup [floor=0.25 (engaged) + divergence=500000+ (eval collapsed) + conf [stat=1.0 bhv=0.0 tmp=0.2]] +``` + +Despite Kelly warmup floor=0.25 (allowing Half-mag exposure), Kelly +controllers all healthy, IQN warnings absent — the model picks ONE +direction (Hold) at every val bar. dir_entropy=0 means deterministic +single-action over the whole 214k window. + +**Root cause:** `experience_action_select` kernel branched on +`eval_mode` and used `argmax(E[Q])` at eval (line 1208 pre-fix). With +Hold's E[Q] ≈ 0 (no position cost) and Short/Long's E[Q] = ε (small edge +minus tx costs), argmax wins Hold deterministically every bar. +Thompson sampling lets positive-Q directions win proportional to their +posterior overlap with Hold. The Fix 33-37 chain addressed +training-side controller dynamics (loss-balance saturation, IQN +target-h-s2, CQL ratio direction, MAX_BUDGET adaptation, Kelly +cold-start exit); SP10 closes the val-side selector that produces the +collapse regardless of controller state. + +The `pearl_thompson_for_distributional_action_selection` was originally +worded as "Thompson at training, argmax at eval"; per +`feedback_trust_code_not_docs.md`, the smoke evidence trumps the +original wording — the pearl is amended to clarify that argmax is +reserved for the Bellman TARGET Q computation only (DDQN target), +NEVER for the rollout SELECTOR. + +**Fix structure (atomic commit per `feedback_no_partial_refactor`):** + +- 1 new ISV slot @ [339..340) (`EVAL_THOMPSON_TEMP_INDEX=339`); + `ISV_TOTAL_DIM` 339 → 340; `SP5_PRODUCER_COUNT` linear span 165 → 166 +- 1 new scratch slot @ [265..266) (`SCRATCH_SP10_THOMPSON_TEMP=265`); + `SP5_SCRATCH_TOTAL` 265 → 266 +- Extend existing `intent_eval_divergence_compute_kernel.cu` (no new + kernel — temperature is a deterministic function of the divergence + already computed; one signal, multiple consumers per + `pearl_engagement_rate_self_correction`): + - 2 new kernel parameters: `divergence_target_isv_index` + (= `KELLY_DIVERGENCE_TARGET_INDEX`) and `scratch_temp_idx` + - 1 new compute branch: + `temp = clamp(divergence / max(div_target, EPS_DIV), 0.5, 2.0)` + - Existing `scratch_idx` parameter renamed to `scratch_div_idx` + (semantic clarity; ABI ordering preserved by inserting new + params after the divergence write) +- Modify `experience_kernels.cu::experience_action_select`: DELETE + `if (eval_mode)` argmax branch entirely; replace direction-selection + block with unconditional temperature-blended Thompson: + - Read `temp = isv_signals_ptr[ISV_EVAL_THOMPSON_TEMP_IDX]` + (defensive clamp to [0.5, 2.0] for cold-start before producer + first observation) + - Per-direction blend: `q_eff = E[Q] + temp · (q_sample − E[Q])`; + argmax over directions on `q_eff` + - `eval_mode` parameter retained in kernel signature (other + branches mag/ord/urg still gate eps-greedy on it; their + argmax-at-eval pathways are different and Pearl 3 of + `pearl_thompson_for_distributional_action_selection` exempts + them) +- New ISV index `#define ISV_EVAL_THOMPSON_TEMP_IDX 339` in + `state_layout.cuh` (mirrors named constant in `sp5_isv_slots.rs`) +- Update launcher `launch_intent_eval_divergence_compute` in + `gpu_dqn_trainer.rs`: 2 new kernel arg slots + a second + `apply_pearls_ad_kernel` chain to smooth the temperature into + ISV[339] +- Update test `test_eval_action_select_eval_argmax_picks_best` → + `test_eval_action_select_thompson_picks_proportionally`: + - Allocate ISV buffer with `EVAL_THOMPSON_TEMP_INDEX = 1.0` (pure + Thompson) and pass to kernel (replaces previous `null_ptr`) + - New assertion: with τ=1.0 and a clear Q gap (E[Q_long]=+0.5 vs + E[Q_flat]=+0.1), best direction wins ≥ 70% of samples (was + ≥ 99% under deterministic argmax) and < 100% (selector is sampling) +- New FoldReset registry entry `sp10_eval_thompson_temp` + dispatch + arm in `reset_named_state` writing sentinel 0; Pearl A's first- + observation replacement fires on the new fold's first + `intent_eval_divergence_compute_kernel` launch +- Pearl amendment in + `pearl_thompson_for_distributional_action_selection.md`: + selector/target distinction clarified; MEMORY.md index entry updated + +**Files touched:** + +- `crates/ml/src/cuda_pipeline/sp5_isv_slots.rs` — new + `EVAL_THOMPSON_TEMP_INDEX=339` constant; `SP5_SLOT_END` 339 → 340; + `SP5_PRODUCER_COUNT` 165 → 166; layout fingerprint update; new + `sp10_thompson_temp_slot_above_sp9_block` test; + `slot_layout_no_overlaps_and_total_correct` extended (164 unique + slots; expected set `{174..278} ∪ {280..340}`) +- `crates/ml/src/cuda_pipeline/gpu_dqn_trainer.rs` — `ISV_TOTAL_DIM` + 339 → 340; `SP5_SCRATCH_TOTAL` 265 → 266; new + `SCRATCH_SP10_THOMPSON_TEMP=265` constant; layout fingerprint update; + `launch_intent_eval_divergence_compute` extended with 2 kernel args + + second `apply_pearls_ad_kernel` chain +- `crates/ml/src/cuda_pipeline/intent_eval_divergence_compute_kernel.cu` + — 2 new parameters (`divergence_target_isv_index`, + `scratch_temp_idx`); existing param renamed to `scratch_div_idx`; + temperature compute branch added (clamp [0.5, 2.0]) +- `crates/ml/src/cuda_pipeline/state_layout.cuh` — new + `ISV_EVAL_THOMPSON_TEMP_IDX 339` define +- `crates/ml/src/cuda_pipeline/experience_kernels.cu` — + `experience_action_select` direction-selection block rewritten: + `if (eval_mode)` argmax DELETED; unconditional temperature-blended + Thompson installed +- `crates/ml/src/cuda_pipeline/mod.rs` — test renamed + `test_eval_action_select_eval_argmax_picks_best` → + `test_eval_action_select_thompson_picks_proportionally`; ISV buffer + setup with `EVAL_THOMPSON_TEMP_INDEX=1.0`; assertions updated +- `crates/ml/src/trainers/dqn/state_reset_registry.rs` — new + `sp10_eval_thompson_temp` FoldReset entry +- `crates/ml/src/trainers/dqn/trainer/training_loop.rs` — new + dispatch arm in `reset_named_state` writing sentinel 0 +- `~/.claude/projects/-home-jgrusewski-Work-foxhunt/memory/pearl_thompson_for_distributional_action_selection.md` + — §4 amended (selector/target distinction); references list extended +- `~/.claude/projects/-home-jgrusewski-Work-foxhunt/memory/MEMORY.md` + — index entry updated +- `docs/dqn-wire-up-audit.md` (this entry) + +**Verification:** + +- `SQLX_OFFLINE=true cargo check -p ml` — clean (only pre-existing 18 + warnings; no new errors or warnings introduced by Fix 38). +- `SQLX_OFFLINE=true cargo test -p ml --lib state_reset` — all 4 + tests pass including the contract test + `every_fold_and_soft_reset_entry_has_dispatch_arm` (now covers the + new `sp10_eval_thompson_temp` entry). +- `SQLX_OFFLINE=true cargo test -p ml --lib sp5_isv_slots` — all 10 + tests pass including the new + `sp10_thompson_temp_slot_above_sp9_block`. + +**Expected smoke-test signature post-fix (5-epoch L40S):** + +- `val_active_frac` > 0.10 with τ ≥ 0.5 (always some action variety) +- `dir_entropy` at eval > 0.3 (multi-direction selection) +- HEALTH_DIAG sp10 line shows τ evolving with divergence; collapsed + states see τ saturate at 2.0 (exploration boost) until divergence + recovers +- T10 50-epoch: `val_trade_count` > 1000 in 214k-bar window (was 1); + `val_active_frac` > 0.20 across folds; sharpe positive net of fold + transitions + +**Per-pearl provenance:** + +- `pearl_thompson_for_distributional_action_selection.md` (amended) — + Thompson is the rollout SELECTOR at all times; argmax ONLY for the + Bellman TARGET Q computation. +- `pearl_controller_anchors_isv_driven.md` — temperature derives from + the SP9 `intent_eval_divergence` canary; not a hardcoded eval-side + constant. +- `pearl_blend_formulas_must_have_permanent_floor.md` — MIN_TEMP=0.5 + enforces permanent stochasticity; eval is never fully deterministic. +- `pearl_engagement_rate_self_correction.md` — extending an existing + producer with a second output (one signal, multiple consumers) is + preferred over allocating a new producer kernel. +- `pearl_first_observation_bootstrap.md` — Pearl A sentinel 0 + + defensive consumer clamp at MIN_TEMP keeps the selector correct from + fold-reset through first producer launch. +- `feedback_no_partial_refactor.md` — atomic commit (slot allocation + + kernel extension + selector rewrite + test rename + dispatch arm + + pearl amendment + audit entry, all in one commit). +- `feedback_isv_for_adaptive_bounds.md` — only Invariant 1 numerical + anchors stay as constants (`MIN_TEMP=0.5`, `MAX_TEMP=2.0`, + `EPS_DIV=1e-6`). +- `feedback_no_cpu_compute_strict.md` — temperature compute lives in + the existing GPU producer kernel; consumer reads ISV via mapped-pinned + `sig_ptr`. +- `feedback_no_atomicadd.md` — extended kernel remains single-block, + single-thread. +- `feedback_no_stubs.md` — temperature is a real signal-driven value + (`divergence/divergence_target` clamped to a permanent floor), not a + placeholder. + +**Forward compatibility:** + +The temperature formula admits more inputs without a kernel signature +change: a future iteration could blend in additional canaries (per-state +uncertainty from Q variance, regime-dependent τ) by reading more ISV +slots inside the existing producer kernel. Other branches (mag/ord/urg) +remain on their current Boltzmann/eps-greedy selectors; if argmax-at-eval +biases are found there, the same SP10 pattern (unconditional sampling + +ISV-driven temperature) applies.