feat(ml-alpha): Phase 3D — three-intervention overtrading fix (A+B+C)
Combined atomic attack on foxhunt's structural overtrading pathology
per omnisearch (2026-06-03) RL HFT literature: foxhunt uniquely uses
Q-distill as the SOLE policy-training mechanism, making it susceptible
to the four-stage Q→π attenuation chain (Goodhart-Skalse 2024) that
blinds π to small persistent fees. Three concurrent fixes attack
different layers:
A. Hold-action logit bias (+log(4) ≈ 1.386 on action 2)
* crates/ml-alpha/src/rl/ppo.rs PolicyHead::new
* crates/ml-alpha/src/rl/multi_head_policy.rs all K heads + build_priors
Counter-balances the structural 4:1 open-vs-hold action prior (4
open variants 0,1,5,6 vs 1 Hold=2). Pre-bias P(open)=36% / P(hold)=9%;
post-bias P(hold)≈29% / P(any open)≈7%. Mid-smoke seed=42 confirms
Hold rises to 71/128 = 55.5% by step 1999.
B. Quadratic-in-trade-size impact-aware cost (Cao et al. 2026,
arXiv:2603.29086 §4)
* crates/ml-alpha/cuda/rl_fused_reward_pipeline.cu Phase 1.5
* 3 new ISV slots 794-796 (α=0.5, β=2.0, enabled=1.0)
cost = α·|Δlots| + β·(Δlots)². 1-lot=2.5; 4-lot flip=34; 8-lot=132
(superlinear). Applied BEFORE shaping so the surfer-scaffold weight
does not amplify or mute. Trail actions (7,8) are no-ops in the
position kernel → cost=0 for them as expected.
C. PPO surrogate gradient restoration with adaptive blend (Cao 2026 §4)
* crates/ml-alpha/cuda/rl_pi_grad_blend.cu (new — element-wise
scale-or-zero operator)
* crates/ml-alpha/src/trainer/integrated.rs Step 7 (π gradient blend)
* 2 new ISV slots 797-798 (weight=0.005, enabled=1.0)
Previously π was trained ONLY by Q-distillation (line 5850 header).
Now: pi_grad = w_ppo·grad_PPO + grad_Q_distill + grad_SAC_entropy.
Restores the direct fee-aware policy-gradient channel that Q-distill
alone cannot transmit. Blend kernel runs BETWEEN surrogate_backward
and rl_q_pi_distill_grad (which uses +=).
Diag emission (E):
* crates/ml-alpha/src/trainer/integrated.rs rewards.{quadratic_cost_alpha,
quadratic_cost_beta, quadratic_cost_enabled, ppo_surrogate_weight,
ppo_surrogate_enabled}
Tests (D):
* multi_head_policy_invariants: updated k1_reduces_to_single_head for
Phase 3D-A bias; new phase_3d_a_hold_bias_propagates_all_heads
invariant verifies Hold dominance in every head at h_t=0. 18/18 pass.
* phase_3d_blend_kernel_invariants (new): 3 GPU-oracle invariants on
rl_pi_grad_blend (disabled-zeros, enabled-scales-linearly, weight-0-
equivalent-to-disabled). 3/3 pass.
* reward_alignment_invariants: 2 new tests (phase_3d_diag_emission +
phase_3d_quadratic_cost_visible_in_rewards) + fix to the existing
surfer_scaffold test (relaxed bootstrap check for Phase 3B-Y pure-pnl
mode default 0.0; absorbs eval drain row). 3/3 pass.
* All 68 ml-alpha lib tests pass.
Verification:
* SQLX_OFFLINE=true cargo build --release --example alpha_rl_train -p
ml-alpha: exit 0
* SQLX_OFFLINE=true cargo check --workspace: exit 0
* ./scripts/determinism-check.sh --quick: DETERMINISTIC (200/200 rows
bit-equal across two same-seed runs)
* FOXHUNT_USE_MULTI_HEAD_POLICY=1 ./scripts/determinism-check.sh --quick:
DETERMINISTIC
* Local Tier 1.5 mid-smoke (seed=42, b=128, 2000 train + 500 eval):
exit 0, completed_clean=true, no NaN, no abort.
Primary kill criterion (total_trades final < 5,000): NOT MET.
Result: 11,767 trades vs 14,691 baseline = 20% reduction. Cao 2026
forecast 96% reduction for pure-PPO/SAC architectures was not
achieved — foxhunt's Q-distill dominance (q_pi_agree_ema = 0.948 in
this run) attenuates the PPO surrogate's fee signal even with the
blend operator. The behavioral signature IS present (Hold dominance
rises from baseline ~36% structural prior to 55.5% at step 1999;
action_entropy = 1.748 within healthy [1.2, 2.04] target).
Regression guard (eval pnl ≥ -$5M): MARGINAL PASS at -$4.96M (-$36k
inside threshold). The Tier 1.5 verdict flags KILL on Pearson +
wr_train + wr_eval + eval_pnl. Pre-cluster, the literature
recommendation is to A/B-ablate each intervention (slots 794-798
individually gated). Mid-smoke architecturally validates that the
three interventions PROPAGATE and do not crash; cluster b=1024 with
longer runs (20k steps) will surface whether the 20% reduction
compounds into a viable policy.
Architectural references:
* pearl_foxhunt_pi_trained_by_q_distillation_not_ppo
* pearl_reward_signal_anti_aligned_with_pnl
* pearl_bootstrap_must_respect_clamp_range
* feedback_no_atomicadd / feedback_no_htod_htoh_only_mapped_pinned
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
This commit is contained in:
@@ -78,6 +78,7 @@ const KERNELS: &[&str] = &[
|
||||
"rl_atom_support_update", // audit 2026-05-24 followup: refreshes atom_supports_d from ISV V_MIN/V_MAX so C51 atom span adapts with reward clamp (Q learning was capped at V_MAX=1.0)
|
||||
"rl_kl_reference_grad",
|
||||
"rl_q_pi_distill_grad", // audit 2026-05-24 vj5f6 followup: KL(softmax(E_Q/τ) || π_new) gradient ADDED to pi_grad_logits — couples Q's improved C51 calibration to action selection (was decoupled per Option B)
|
||||
"rl_pi_grad_blend", // Phase 3D-C (2026-06-03): PPO surrogate × Q-distill blend operator (replaces zero-fill before Q-distill +=); restores direct PG signal to π
|
||||
"action_entropy_per_step", // POST-gate action entropy EMA for SAC α/τ co-tuning
|
||||
"rl_q_distill_lambda_controller",// audit 2026-05-24 rljzl followup: adaptive λ_distill via Schulman bounded step on KL_EMA vs target
|
||||
"rl_unit_state_update", // SP20 P1+P5 audit fix: per-unit trade state machine — detects open/close/reverse transitions, sets up unit slot 0 entry+trail
|
||||
|
||||
@@ -50,6 +50,14 @@
|
||||
// dones + edge_decay_frac_alerted; fades w as agent crosses break-even).
|
||||
#define RL_SURFER_SCAFFOLD_WEIGHT_INDEX 753
|
||||
|
||||
// Phase 3D-B (2026-06-03): quadratic-in-trade-size impact-aware cost
|
||||
// (Cao et al. 2026 arXiv:2603.29086). cost = α·|Δlots| + β·(Δlots)²
|
||||
// applied per env step in Phase 1.5 BETWEEN realized_pnl_delta (Phase 1)
|
||||
// and surfer-scaffold shaping (Phase 5). Disabled when gate slot is ≤ 0.5.
|
||||
#define RL_QUADRATIC_COST_ALPHA_INDEX 794
|
||||
#define RL_QUADRATIC_COST_BETA_INDEX 795
|
||||
#define RL_QUADRATIC_COST_ENABLED_INDEX 796
|
||||
|
||||
#define MAX_UNITS 4
|
||||
|
||||
// PosFlat layout (crates/ml-backtesting/src/lob/mod.rs::PosFlat):
|
||||
@@ -109,6 +117,36 @@ extern "C" __global__ void rl_fused_reward_pipeline(
|
||||
rewards[b] = reward;
|
||||
dones[b] = done;
|
||||
|
||||
// ================================================================
|
||||
// PHASE 1.5 (Phase 3D-B, 2026-06-03): Cao et al. 2026
|
||||
// quadratic-in-trade-size impact-aware cost.
|
||||
//
|
||||
// cost = α·|Δlots| + β·(Δlots)²
|
||||
//
|
||||
// Applied BEFORE Phase 5 shaping so the shaped-vs-raw weight w (slot
|
||||
// 753) does not amplify or mute the cost. Discourages overtrading
|
||||
// even when reward is pnl-aligned. Trail-stop actions (7, 8) are
|
||||
// no-ops in actions_to_market_targets and don't change current_lots,
|
||||
// so Δlots = 0 for those and the cost stays at 0. Hold (2) likewise
|
||||
// produces no position change → no cost. FlatFrom* / HalfFlat* /
|
||||
// open actions all incur position changes proportional to their
|
||||
// size, so they pay an impact penalty scaled by the change magnitude.
|
||||
//
|
||||
// Reward units: pts × lots (matches realized_pnl_delta convention).
|
||||
// Bootstrap α=0.5, β=2.0 yields: 1-lot=2.5, 4-lot flip=34, 8-lot flip=132.
|
||||
// ================================================================
|
||||
if (isv[RL_QUADRATIC_COST_ENABLED_INDEX] > 0.5f) {
|
||||
const float trade_size = fabsf((float)(current_lots - prev_lots));
|
||||
if (trade_size > 0.0f) {
|
||||
const float alpha_cost = isv[RL_QUADRATIC_COST_ALPHA_INDEX];
|
||||
const float beta_cost = isv[RL_QUADRATIC_COST_BETA_INDEX];
|
||||
const float cost = alpha_cost * trade_size
|
||||
+ beta_cost * trade_size * trade_size;
|
||||
reward -= cost;
|
||||
rewards[b] = reward;
|
||||
}
|
||||
}
|
||||
|
||||
// ================================================================
|
||||
// PHASE 2: rl_unit_state_update
|
||||
// ================================================================
|
||||
|
||||
48
crates/ml-alpha/cuda/rl_pi_grad_blend.cu
Normal file
48
crates/ml-alpha/cuda/rl_pi_grad_blend.cu
Normal file
@@ -0,0 +1,48 @@
|
||||
// rl_pi_grad_blend.cu — Phase 3D-C (2026-06-03): PPO surrogate × Q-distill
|
||||
// gradient blend operator.
|
||||
//
|
||||
// Replaces the pre-Q-distill zero-fill of `ss_pi_grad_logits_d` with a
|
||||
// conditional scale/zero step that lets the PPO clipped-surrogate
|
||||
// gradient flow into π alongside the Q-distill term:
|
||||
//
|
||||
// ENABLED (slot RL_PPO_SURROGATE_ENABLED_INDEX > 0.5):
|
||||
// pi_grad[i] *= weight (slot RL_PPO_SURROGATE_WEIGHT_INDEX)
|
||||
// DISABLED:
|
||||
// pi_grad[i] = 0.0 (legacy zero-then-distill path)
|
||||
//
|
||||
// Q-distill then ADDS (`pi_grad_logits[i] += ...`) on top, producing:
|
||||
// pi_grad[i] = weight * grad_PPO[i] + grad_Q_distill[i]
|
||||
//
|
||||
// Why this is needed: per Goodhart-Skalse 2024 (Causes of Misalignment),
|
||||
// the four-stage Q→π attenuation chain (V → Bellman target → softmax(Q/τ)
|
||||
// → KL distill) cannot transmit small persistent fees back to π. Cao et
|
||||
// al. 2026 (arXiv:2603.29086) showed restoring a direct policy-gradient
|
||||
// channel (weight 1e-3 to 1e-2) cuts turnover 96% across TD3 / PPO / SAC
|
||||
// when paired with quadratic impact costs.
|
||||
//
|
||||
// Per `feedback_no_atomicadd`: element-wise, one thread per logit.
|
||||
// Per `feedback_no_nvrtc`: pre-compiled cubin via build.rs.
|
||||
|
||||
#include <stdint.h>
|
||||
|
||||
#define RL_PPO_SURROGATE_WEIGHT_INDEX 797
|
||||
#define RL_PPO_SURROGATE_ENABLED_INDEX 798
|
||||
|
||||
extern "C" __global__ void rl_pi_grad_blend(
|
||||
float* __restrict__ pi_grad_logits, // [B × N_ACTIONS] IN/OUT
|
||||
const float* __restrict__ isv, // ISV bus (RO)
|
||||
int b_size,
|
||||
int n_actions
|
||||
) {
|
||||
const int i = blockIdx.x * blockDim.x + threadIdx.x;
|
||||
const int total = b_size * n_actions;
|
||||
if (i >= total) return;
|
||||
|
||||
const float enabled = isv[RL_PPO_SURROGATE_ENABLED_INDEX];
|
||||
if (enabled > 0.5f) {
|
||||
const float w = isv[RL_PPO_SURROGATE_WEIGHT_INDEX];
|
||||
pi_grad_logits[i] *= w;
|
||||
} else {
|
||||
pi_grad_logits[i] = 0.0f;
|
||||
}
|
||||
}
|
||||
@@ -1976,6 +1976,62 @@ pub const RL_POLICY_GATE_CONTROLLER_BOOTSTRAP_DONE_INDEX: usize = 792;
|
||||
/// natural OFF sentinel and matches a hypothetical clamp-floor of 0.0.
|
||||
pub const RL_POPART_NORMALIZE_ENABLED_INDEX: usize = 793;
|
||||
|
||||
/// Phase 3D-B quadratic-in-trade-size impact-aware cost coefficient α
|
||||
/// (linear term) — applied per env step as `r -= α·|Δlots| + β·(Δlots)²`
|
||||
/// inside `rl_fused_reward_pipeline.cu` Phase 1.5, between the
|
||||
/// `realized_pnl_delta` extraction and the surfer-scaffold shaping.
|
||||
///
|
||||
/// Implements Cao et al. 2026 (arXiv:2603.29086) impact-aware cost
|
||||
/// methodology, which showed TD3 turnover collapse 19% → 1% across PPO,
|
||||
/// SAC, DDPG, A2C when flat fees were replaced with quadratic costs.
|
||||
///
|
||||
/// Bootstrap 0.5 — small linear penalty so 1-lot trades cost 0.5 reward
|
||||
/// units (commensurate with reward = realized_pnl_delta in pts × lots).
|
||||
///
|
||||
/// `pearl_bootstrap_must_respect_clamp_range`: kernel uses the raw slot
|
||||
/// value without further clamping (no controller actively writes here);
|
||||
/// 0.5 is the documented bootstrap and there is no [MIN, MAX] for the
|
||||
/// constant — it stays at 0.5 until an operator decides to tune.
|
||||
pub const RL_QUADRATIC_COST_ALPHA_INDEX: usize = 794;
|
||||
|
||||
/// Phase 3D-B quadratic-in-trade-size impact-aware cost coefficient β
|
||||
/// (quadratic term). Bootstrap 2.0 produces superlinear flip penalties:
|
||||
/// 1-lot = 0.5 + 2 = 2.5 reward units; 4-lot flip = 2 + 32 = 34; 8-lot
|
||||
/// flip = 4 + 128 = 132. Mirrors Cao et al. 2026 quadratic impact curve.
|
||||
pub const RL_QUADRATIC_COST_BETA_INDEX: usize = 795;
|
||||
|
||||
/// Phase 3D-B quadratic-cost gate. `> 0.5f` enables the Phase 1.5
|
||||
/// penalty in `rl_fused_reward_pipeline.cu`; ≤ 0.5 disables it (kernel
|
||||
/// preserves bit-equality with pre-3D when off). Bootstrap 1.0 (ENABLED).
|
||||
///
|
||||
/// Operator-controlled binary gate; clamp range is the kernel's `> 0.5f`
|
||||
/// test so bootstraps 0.0 (off) or 1.0 (on) satisfy
|
||||
/// `pearl_bootstrap_must_respect_clamp_range`.
|
||||
pub const RL_QUADRATIC_COST_ENABLED_INDEX: usize = 796;
|
||||
|
||||
/// Phase 3D-C PPO clipped-surrogate gradient blend weight. Multiplies
|
||||
/// the PPO surrogate gradient before Q-distill adds its own gradient to
|
||||
/// `ss_pi_grad_logits_d`, giving π a direct fee-aware policy-gradient
|
||||
/// signal alongside the Q→π distillation chain (which Goodhart-Skalse
|
||||
/// 2024 showed cannot perceive small persistent fees through the four
|
||||
/// learned-approximator attenuation hops).
|
||||
///
|
||||
/// Bootstrap 0.005 — small but non-zero per Cao et al. 2026 §4
|
||||
/// recommendation of 1e-3 to 1e-2 for blended surrogate weights when
|
||||
/// the primary critic head is well-calibrated.
|
||||
///
|
||||
/// Clamp by interpretation: weight ∈ [0, ∞); 0.005 is positive so satisfies
|
||||
/// `pearl_bootstrap_must_respect_clamp_range` (the natural MIN is 0,
|
||||
/// MAX is undefined; we just don't snap-to-zero unless an operator sets
|
||||
/// `RL_PPO_SURROGATE_ENABLED_INDEX` ≤ 0.5).
|
||||
pub const RL_PPO_SURROGATE_WEIGHT_INDEX: usize = 797;
|
||||
|
||||
/// Phase 3D-C PPO surrogate gradient enable gate. `> 0.5f` runs the
|
||||
/// PPO surrogate backward + scaled accumulation into `ss_pi_grad_logits_d`
|
||||
/// BEFORE Q-distill writes; ≤ 0.5 falls back to the legacy zero-then-distill
|
||||
/// path. Bootstrap 1.0 (ENABLED).
|
||||
pub const RL_PPO_SURROGATE_ENABLED_INDEX: usize = 798;
|
||||
|
||||
/// Last RL-allocated slot index (exclusive).
|
||||
/// Pre-risk-stack: 662. Post-Fix-B: 685. Post-v9 (warmup boundary): 696.
|
||||
/// Post-regime-observer (F1.1): 716. Post-warmed-flag addendum: 717.
|
||||
@@ -1995,4 +2051,6 @@ pub const RL_POPART_NORMALIZE_ENABLED_INDEX: usize = 793;
|
||||
/// Post-Phase 2A-D fix B1 (gate LR multiplier slot 790): 790.
|
||||
/// Post-Phase 2A-D fix B1.3 (gate-LR controller state slots 791-792): 792.
|
||||
/// Post-Phase 3B-Y (PopArt normalize gate slot 793): 793.
|
||||
pub const RL_SLOTS_END: usize = 794;
|
||||
/// Post-Phase 3D-B (quadratic cost α/β/enabled slots 794-796): 796.
|
||||
/// Post-Phase 3D-C (PPO surrogate blend weight/enabled slots 797-798): 798.
|
||||
pub const RL_SLOTS_END: usize = 799;
|
||||
|
||||
@@ -311,6 +311,23 @@ impl MultiHeadPolicy {
|
||||
// Head 2 — HOLD / stable (Phase 0 dominant action)
|
||||
b_heads_init[2 * N_ACTIONS + 2] += 0.5;
|
||||
}
|
||||
// Phase 3D-A (2026-06-03): structural Hold-action elevation across
|
||||
// ALL heads. The 4:1 open-vs-hold action prior (4 opens: 0,1,5,6 vs
|
||||
// 1 hold: 2) biases uniform-init sampling toward trading
|
||||
// (P(open)=36%, P(hold)=9%). Adding +log(4)≈1.386 to every head's
|
||||
// Hold bias counter-balances this at init while preserving the
|
||||
// per-head specialization additions above (Head 0 still gets +0.5
|
||||
// on ShortLarge PLUS +1.386 on Hold; Head 2 gets +0.5 + +1.386 =
|
||||
// +1.886 on Hold). Mirror the change in `build_priors` below so
|
||||
// the aux-KL prior does not fight the new initialization.
|
||||
//
|
||||
// Literature: Cao et al. 2026 (arXiv:2603.29086) action-space-prior
|
||||
// adjustment; one of three concurrent Phase 3D interventions
|
||||
// (B: quadratic impact cost; C: PPO surrogate restoration).
|
||||
let hold_bias: f32 = (4.0_f32).ln();
|
||||
for head in 0..cfg.k {
|
||||
b_heads_init[head * N_ACTIONS + 2] += hold_bias;
|
||||
}
|
||||
let b_heads_d = upload(&stream, &b_heads_init)
|
||||
.context("upload b_heads init")?;
|
||||
|
||||
@@ -822,14 +839,20 @@ fn upload(stream: &Arc<CudaStream>, host: &[f32]) -> Result<CudaSlice<f32>> {
|
||||
// Mirror the per-head init biases from `MultiHeadPolicy::new` so the aux
|
||||
// KL term does not fight the asymmetry-break initialization.
|
||||
//
|
||||
// prior[0] = softmax([+0.5, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0]) (SHORT)
|
||||
// prior[1] = softmax([0, 0, 0, 0, 0, +0.5, +0.5, 0, 0, 0, 0]) (LONG)
|
||||
// prior[2] = softmax([0, 0, +0.5, 0, 0, 0, 0, 0, 0, 0, 0]) (HOLD)
|
||||
// prior[k≥3] = uniform = 1 / N_ACTIONS
|
||||
// Post-Phase 3D-A (2026-06-03): every head now also receives +log(4) on
|
||||
// action 2 (Hold) to counter the 4:1 open-vs-hold action prior. The
|
||||
// priors must mirror that elevation or the aux-KL gradient would push
|
||||
// the policy back toward the (over-opening) uniform distribution.
|
||||
//
|
||||
// prior[0] = softmax([+0.5, 0, +log(4), 0, 0, 0, 0, 0, 0, 0, 0]) (SHORT + HOLD lift)
|
||||
// prior[1] = softmax([0, 0, +log(4), 0, 0, +0.5, +0.5, 0, 0, 0, 0]) (LONG + HOLD lift)
|
||||
// prior[2] = softmax([0, 0, +0.5+log(4), 0, 0, 0, 0, 0, 0, 0, 0]) (HOLD bias + lift)
|
||||
// prior[k≥3] = softmax([0, 0, +log(4), 0, 0, 0, 0, 0, 0, 0, 0]) (HOLD lift only)
|
||||
//
|
||||
// Returns a `K × N_ACTIONS` flat row-major host buffer.
|
||||
fn build_priors(k: usize) -> Vec<f32> {
|
||||
let mut priors = vec![0.0_f32; k * N_ACTIONS];
|
||||
let hold_bias: f32 = (4.0_f32).ln();
|
||||
for head in 0..k {
|
||||
// Logit ramp matching the init bias (zero for unbiased actions).
|
||||
let mut logit = vec![0.0_f32; N_ACTIONS];
|
||||
@@ -840,8 +863,10 @@ fn build_priors(k: usize) -> Vec<f32> {
|
||||
logit[6] = 0.5;
|
||||
}
|
||||
2 => logit[2] = 0.5,
|
||||
_ => {} // uniform — all zeros → softmax = 1/N
|
||||
_ => {} // uniform — all zeros → Hold-lift below is the only bias
|
||||
}
|
||||
// Phase 3D-A: HOLD-action elevation (all heads).
|
||||
logit[2] += hold_bias;
|
||||
// softmax — single-batch row, no overflow risk at this scale.
|
||||
let mx = logit.iter().cloned().fold(f32::NEG_INFINITY, f32::max);
|
||||
let mut s = 0.0_f32;
|
||||
|
||||
@@ -45,7 +45,7 @@ use rand_chacha::ChaCha8Rng;
|
||||
|
||||
use crate::heads::HIDDEN_DIM;
|
||||
use crate::pinned_mem::MappedF32Buffer;
|
||||
use crate::rl::common::N_ACTIONS;
|
||||
use crate::rl::common::{Action, N_ACTIONS};
|
||||
use crate::trainer::raw_launch::{RawArgs, raw_launch};
|
||||
|
||||
const PPO_SURR_CUBIN: &[u8] = include_bytes!(concat!(
|
||||
@@ -190,7 +190,24 @@ impl PolicyHead {
|
||||
let w_host: Vec<f32> = (0..n_out * n_in)
|
||||
.map(|_| rng.gen_range(-scale..scale))
|
||||
.collect();
|
||||
let b_host: Vec<f32> = vec![0.0_f32; n_out];
|
||||
// Phase 3D-A (2026-06-03): Hold-action logit bias.
|
||||
//
|
||||
// The 11-action grid (see `crate::rl::common::Action`) has 4 open
|
||||
// variants (0=ShortLarge, 1=ShortSmall, 5=LongSmall, 6=LongLarge)
|
||||
// vs 1 hold (2=Hold). At uniform random init the structural prior
|
||||
// is P(open)=4/11≈36% and P(hold)=1/11≈9% BEFORE any learning —
|
||||
// a 4:1 bias toward opening. Counter-balance by adding +log(4)≈1.386
|
||||
// to the Hold action bias so the asymmetry-break probabilities
|
||||
// are P(hold)=4/(4+10)≈28.6% and P(any specific open)=1/14≈7.1%,
|
||||
// restoring near-uniform open-vs-hold mass at init.
|
||||
//
|
||||
// Per Cao et al. 2026 + Goodhart-Skalse 2024 literature on
|
||||
// action-space prior bias, this is one of three concurrent fixes
|
||||
// (B: quadratic impact cost; C: PPO surrogate restoration) for
|
||||
// foxhunt's structural overtrading pathology.
|
||||
let hold_bias: f32 = (4.0_f32).ln();
|
||||
let mut b_host: Vec<f32> = vec![0.0_f32; n_out];
|
||||
b_host[Action::Hold as usize] += hold_bias;
|
||||
|
||||
let w_d = upload(&stream, &w_host)?;
|
||||
let b_d = upload(&stream, &b_host)?;
|
||||
|
||||
@@ -319,6 +319,14 @@ const RL_INVENTORY_VARIANCE_UPDATE_CUBIN: &[u8] =
|
||||
include_bytes!(concat!(env!("OUT_DIR"), "/rl_inventory_variance_update.cubin"));
|
||||
const RL_Q_PI_DISTILL_GRAD_CUBIN: &[u8] =
|
||||
include_bytes!(concat!(env!("OUT_DIR"), "/rl_q_pi_distill_grad.cubin"));
|
||||
// Phase 3D-C (2026-06-03): PPO surrogate × Q-distill blend operator.
|
||||
// Scales `ss_pi_grad_logits_d` by ISV[RL_PPO_SURROGATE_WEIGHT_INDEX] when
|
||||
// enabled (slot 798 > 0.5), else zeros it. Run BETWEEN
|
||||
// `policy_head.surrogate_backward_logits` and `rl_q_pi_distill_grad`
|
||||
// (which `+=` adds on top), producing:
|
||||
// pi_grad = w_ppo · grad_PPO + grad_Q_distill
|
||||
const RL_PI_GRAD_BLEND_CUBIN: &[u8] =
|
||||
include_bytes!(concat!(env!("OUT_DIR"), "/rl_pi_grad_blend.cubin"));
|
||||
// λ_distill adaptive controller (rljzl followup 2026-05-24).
|
||||
// Drives λ via Schulman bounded step on KL_EMA toward target.
|
||||
const RL_Q_DISTILL_LAMBDA_CONTROLLER_CUBIN: &[u8] =
|
||||
@@ -856,6 +864,11 @@ pub struct IntegratedTrainer {
|
||||
// Q→π distillation gradient (audit 2026-05-24 vj5f6 followup).
|
||||
_rl_q_pi_distill_grad_module: Arc<CudaModule>,
|
||||
rl_q_pi_distill_grad_fn: CudaFunction,
|
||||
// Phase 3D-C (2026-06-03): PPO surrogate × Q-distill blend operator
|
||||
// (replaces zero-fill before `+=` Q-distill). Scales `ss_pi_grad_logits_d`
|
||||
// by ISV[RL_PPO_SURROGATE_WEIGHT_INDEX] when enabled, else zeros it.
|
||||
_rl_pi_grad_blend_module: Arc<CudaModule>,
|
||||
rl_pi_grad_blend_fn: CudaFunction,
|
||||
_rl_kl_reference_module: Arc<CudaModule>,
|
||||
#[allow(dead_code)]
|
||||
_rl_kl_reference_grad_fn: CudaFunction,
|
||||
@@ -1913,6 +1926,13 @@ impl IntegratedTrainer {
|
||||
let rl_q_pi_distill_grad_fn = rl_q_pi_distill_grad_module
|
||||
.load_function("rl_q_pi_distill_grad")
|
||||
.context("load rl_q_pi_distill_grad")?;
|
||||
// Phase 3D-C (2026-06-03): PPO surrogate blend operator.
|
||||
let rl_pi_grad_blend_module = ctx
|
||||
.load_cubin(RL_PI_GRAD_BLEND_CUBIN.to_vec())
|
||||
.context("load rl_pi_grad_blend cubin")?;
|
||||
let rl_pi_grad_blend_fn = rl_pi_grad_blend_module
|
||||
.load_function("rl_pi_grad_blend")
|
||||
.context("load rl_pi_grad_blend")?;
|
||||
let rl_kl_reference_module = ctx
|
||||
.load_cubin(include_bytes!(concat!(env!("OUT_DIR"), "/rl_kl_reference_grad.cubin")).to_vec())
|
||||
.context("load rl_kl_reference_grad cubin")?;
|
||||
@@ -3154,6 +3174,8 @@ impl IntegratedTrainer {
|
||||
rl_atom_support_update_fn,
|
||||
_rl_q_pi_distill_grad_module: rl_q_pi_distill_grad_module,
|
||||
rl_q_pi_distill_grad_fn,
|
||||
_rl_pi_grad_blend_module: rl_pi_grad_blend_module,
|
||||
rl_pi_grad_blend_fn,
|
||||
_rl_kl_reference_module: rl_kl_reference_module,
|
||||
_rl_kl_reference_grad_fn: rl_kl_reference_grad_fn,
|
||||
_rl_ensemble_action_value_module: rl_ensemble_action_value_module,
|
||||
@@ -3717,7 +3739,7 @@ impl IntegratedTrainer {
|
||||
// (slot, value) pair — pure device write, no HtoD per
|
||||
// `feedback_no_htod_htoh_only_mapped_pinned`.
|
||||
{
|
||||
let isv_constants: [(usize, f32); 245] = [
|
||||
let isv_constants: [(usize, f32); 250] = [
|
||||
// Static seeds for the adaptive reward-clamp controller —
|
||||
// these are the initial values that
|
||||
// `rl_reward_clamp_controller` will replace once it
|
||||
@@ -4157,6 +4179,24 @@ impl IntegratedTrainer {
|
||||
// Targets the 40-50% of `Pearson(reward, Δpnl)` gap attributed
|
||||
// to popart sign-flips by the Phase 3B-followup audit.
|
||||
(crate::rl::isv_slots::RL_POPART_NORMALIZE_ENABLED_INDEX, 0.0_f32),
|
||||
// Phase 3D-B (2026-06-03): quadratic-in-trade-size impact-aware
|
||||
// cost (Cao et al. 2026 arXiv:2603.29086). Penalises overtrading
|
||||
// even when reward is pnl-aligned. Cost = α·|Δlots| + β·(Δlots)²
|
||||
// applied per env step inside `rl_fused_reward_pipeline.cu`
|
||||
// Phase 1.5. Bootstraps target ~$2.5 (1-lot) → ~$132 (8-lot flip)
|
||||
// reduction in reward units (pts × lots).
|
||||
(crate::rl::isv_slots::RL_QUADRATIC_COST_ALPHA_INDEX, 0.5_f32),
|
||||
(crate::rl::isv_slots::RL_QUADRATIC_COST_BETA_INDEX, 2.0_f32),
|
||||
(crate::rl::isv_slots::RL_QUADRATIC_COST_ENABLED_INDEX, 1.0_f32),
|
||||
// Phase 3D-C (2026-06-03): PPO surrogate gradient restoration.
|
||||
// Currently π is trained ONLY by Q-distillation (see header at
|
||||
// line 5850). The four-stage Q→π attenuation chain (Goodhart-
|
||||
// Skalse 2024) blinds π to small persistent fees. Restoring
|
||||
// the PPO clipped-surrogate gradient at weight 0.005 gives π
|
||||
// a direct policy-gradient channel for the quadratic-cost
|
||||
// signal, while Q-distill remains the dominant term.
|
||||
(crate::rl::isv_slots::RL_PPO_SURROGATE_WEIGHT_INDEX, 0.005_f32),
|
||||
(crate::rl::isv_slots::RL_PPO_SURROGATE_ENABLED_INDEX, 1.0_f32),
|
||||
];
|
||||
for (slot, value) in isv_constants.iter() {
|
||||
let slot_i32 = *slot as i32;
|
||||
@@ -5847,21 +5887,70 @@ impl IntegratedTrainer {
|
||||
// BCE/aux (supervised via step_batched) only.
|
||||
let _ = &self.ss_q_grad_h_t_d;
|
||||
|
||||
// ── Step 7: π ← Q distillation (replaces PPO surrogate) ──────
|
||||
// π is trained ONLY by behavioral cloning from Q-Thompson.
|
||||
// No PPO, no advantages, no importance ratios.
|
||||
// grad = λ × (π_θ(a|s) - softmax(E_Q/τ)(a))
|
||||
// This is the cross-entropy gradient that aligns π with Q's
|
||||
// action ranking. qpa is positive by construction.
|
||||
// Zero the grad buffer first (distillation ADDS to it).
|
||||
unsafe {
|
||||
raw_memset_d8_zero(
|
||||
self.ss_pi_grad_logits_d.raw_ptr(),
|
||||
self.ss_pi_grad_logits_d.num_bytes(),
|
||||
self.raw_stream,
|
||||
).map_err(|e| anyhow::anyhow!("zero ss_pi_grad_logits: {:?}", e))?;
|
||||
// ── Step 7: π ← w·PPO_surrogate_grad + Q distillation + SAC entropy ─
|
||||
// Phase 3D-C (2026-06-03): π gradient now blends two sources:
|
||||
//
|
||||
// (a) PPO clipped-surrogate gradient (per-batch, advantage-aware,
|
||||
// fee-sensitive). Written to `ss_pi_grad_logits_d` by
|
||||
// `policy_head.surrogate_backward_logits`, then scaled in place
|
||||
// by `RL_PPO_SURROGATE_WEIGHT_INDEX` via the blend kernel
|
||||
// (or zeroed if the enable gate ≤ 0.5, preserving the legacy
|
||||
// zero-then-distill path for ablation).
|
||||
//
|
||||
// (b) Q→π distillation + SAC entropy (existing). The
|
||||
// `rl_q_pi_distill_grad` kernel uses `+=` so it ADDS on top
|
||||
// of whatever the blend kernel produced. Result:
|
||||
// pi_grad = w_ppo · grad_PPO + grad_Q_distill_and_entropy
|
||||
//
|
||||
// Why both: per Goodhart-Skalse 2024 (Causes of Misalignment), the
|
||||
// four-stage Q→π attenuation chain (V → Bellman target → softmax(Q/τ)
|
||||
// → KL distill) cannot transmit small persistent fees back to π.
|
||||
// Cao et al. 2026 (arXiv:2603.29086) showed restoring a direct
|
||||
// policy-gradient channel (weight 1e-3 to 1e-2) cuts turnover 96%
|
||||
// across TD3/PPO/SAC/DDPG/A2C when paired with quadratic impact
|
||||
// costs. This commit pairs (a)+(b)+Phase 3D-A Hold-bias + Phase 3D-B
|
||||
// quadratic cost (one atomic intervention package).
|
||||
{
|
||||
// (a.1) PPO surrogate backward → ss_pi_grad_logits_d (OVERWRITE).
|
||||
// Reads pi_logits, log_pi_old, actions, advantages from the
|
||||
// same buffers the forward pass used; emits per-logit gradient
|
||||
// including the entropy bonus contribution via the kernel's
|
||||
// `entropy_coef` slot read.
|
||||
self.policy_head.surrogate_backward_logits(
|
||||
&self.pi_logits_d,
|
||||
&self.log_pi_old_d,
|
||||
&self.actions_d,
|
||||
&self.advantages_d,
|
||||
&self.isv_dev_ptr,
|
||||
b_size,
|
||||
&mut self.ss_pi_grad_logits_d,
|
||||
)
|
||||
.context("policy_head.surrogate_backward_logits (Phase 3D-C blend)")?;
|
||||
|
||||
// (a.2) Scale ss_pi_grad_logits_d in place by ppo_surrogate_weight
|
||||
// — OR — zero it if ENABLED ≤ 0.5 (legacy zero-then-distill
|
||||
// path for ablation). One element-wise kernel.
|
||||
let total = (b_size * N_ACTIONS) as i32;
|
||||
let mut args = RawArgs::new();
|
||||
args.push_ptr(self.ss_pi_grad_logits_d.raw_ptr());
|
||||
args.push_ptr(self.isv_dev_ptr);
|
||||
args.push_i32(b_size as i32);
|
||||
args.push_i32(N_ACTIONS as i32);
|
||||
let block: u32 = 128;
|
||||
let grid: u32 = ((total as u32) + block - 1) / block;
|
||||
let mut ptrs = args.build_arg_ptrs();
|
||||
unsafe {
|
||||
raw_launch(
|
||||
self.rl_pi_grad_blend_fn.cu_function(),
|
||||
(grid, 1, 1), (block, 1, 1), 0,
|
||||
self.raw_stream,
|
||||
&mut ptrs[..args.len()],
|
||||
).map_err(|e| anyhow::anyhow!("rl_pi_grad_blend: {:?}", e))?;
|
||||
}
|
||||
}
|
||||
{
|
||||
// (b) Q→π distillation + SAC entropy gradient (ADDS via `+=`).
|
||||
// grad = λ × (π_θ(a|s) − softmax(E_Q/τ)(a)) − SAC_α × ...
|
||||
let b_size_i = b_size as i32;
|
||||
let mut args = RawArgs::new();
|
||||
args.push_ptr(self.q_logits_target_st_d.raw_ptr());
|
||||
@@ -5881,8 +5970,8 @@ impl IntegratedTrainer {
|
||||
}
|
||||
}
|
||||
|
||||
// π gradient is fully determined by Q→π distillation + SAC entropy
|
||||
// (rl_q_pi_distill_grad above). No PPO surrogate or KL penalty.
|
||||
// π gradient = w_ppo · grad_PPO + grad_Q_distill + grad_SAC_entropy
|
||||
// (Phase 3D-C blend; ablation: set RL_PPO_SURROGATE_ENABLED_INDEX ≤ 0.5).
|
||||
|
||||
// Phase 2A-C: gate π backward between legacy and MultiHeadPolicy.
|
||||
// Both paths consume `ss_pi_grad_logits_d` (Q-distill output)
|
||||
@@ -11642,6 +11731,24 @@ impl IntegratedTrainer {
|
||||
isv[RL_NEG_SCALED_REWARD_MAX_EMA_INDEX],
|
||||
"surfer_scaffold_weight":
|
||||
isv[crate::rl::isv_slots::RL_SURFER_SCAFFOLD_WEIGHT_INDEX],
|
||||
// Phase 3D-B (2026-06-03): quadratic-cost slot snapshot.
|
||||
// No per-step batch-aggregated cost is emitted here — the
|
||||
// reward kernel applies the penalty per-batch element
|
||||
// and the change is already visible via `rewards.sum`/`min`
|
||||
// vs. legacy. Slot values are stable per session unless
|
||||
// an operator tunes them mid-run, so emitting them every
|
||||
// step preserves config visibility in the diag.
|
||||
"quadratic_cost_alpha":
|
||||
isv[crate::rl::isv_slots::RL_QUADRATIC_COST_ALPHA_INDEX],
|
||||
"quadratic_cost_beta":
|
||||
isv[crate::rl::isv_slots::RL_QUADRATIC_COST_BETA_INDEX],
|
||||
"quadratic_cost_enabled":
|
||||
isv[crate::rl::isv_slots::RL_QUADRATIC_COST_ENABLED_INDEX],
|
||||
// Phase 3D-C (2026-06-03): PPO surrogate blend slots.
|
||||
"ppo_surrogate_weight":
|
||||
isv[crate::rl::isv_slots::RL_PPO_SURROGATE_WEIGHT_INDEX],
|
||||
"ppo_surrogate_enabled":
|
||||
isv[crate::rl::isv_slots::RL_PPO_SURROGATE_ENABLED_INDEX],
|
||||
},
|
||||
"ppo": {
|
||||
"ratio_clamp_max": isv[RL_PPO_RATIO_CLAMP_MAX_INDEX],
|
||||
|
||||
@@ -222,7 +222,11 @@ fn k1_reduces_to_single_head() -> Result<()> {
|
||||
|
||||
// h_t = 0 so the policy-head matmul produces zero logits;
|
||||
// post-init, the bias is the only source of non-zero pi_logits_k.
|
||||
// At K=1 the head's bias is [+0.5, 0, 0, ..., 0] (Head 0 → ShortLarge).
|
||||
//
|
||||
// Post Phase 3D-A: at K=1 the head's bias is now
|
||||
// [+0.5, 0, +log(4), 0, 0, 0, 0, 0, 0, 0, 0]
|
||||
// (action 0 = ShortLarge → +0.5 per Head 0 specialization;
|
||||
// action 2 = Hold → +log(4) per Phase 3D-A all-heads lift).
|
||||
let h_t = vec![0.0_f32; b_size * HIDDEN_DIM];
|
||||
let regime: Vec<f32> = (0..b_size * REGIME_DIM).map(|i| (i as f32) * 0.01).collect();
|
||||
let h_t_d = upload_f32(&stream, &h_t)?;
|
||||
@@ -234,21 +238,32 @@ fn k1_reduces_to_single_head() -> Result<()> {
|
||||
let pi_logits = read_slice_d_pub(&stream, &pi_logits_out, b_size * N_ACTIONS)?;
|
||||
|
||||
// At K=1: mixture = single head's softmax. With h_t=0 the per-head
|
||||
// pre-softmax logits are exactly b_heads = [+0.5, 0, ..., 0]. So:
|
||||
// p[0] = exp(0.5) / (exp(0.5) + (N_ACTIONS-1))
|
||||
// p[a≠0] = 1 / (exp(0.5) + (N_ACTIONS-1))
|
||||
let denom = 0.5_f32.exp() + (N_ACTIONS as f32 - 1.0);
|
||||
// pre-softmax logits are exactly b_heads above. Let s = exp(0.5),
|
||||
// h = exp(log(4)) = 4, and (N_ACTIONS - 2) = 9 unbiased actions.
|
||||
// denom = s + h + 9
|
||||
// p[0] = s / denom (ShortLarge bias)
|
||||
// p[2] = h / denom (Hold bias)
|
||||
// p[else] = 1 / denom
|
||||
let hold_logit = (4.0_f32).ln();
|
||||
let denom = 0.5_f32.exp() + hold_logit.exp() + (N_ACTIONS as f32 - 2.0);
|
||||
let p0_expected = 0.5_f32.exp() / denom;
|
||||
let phold_expected = hold_logit.exp() / denom;
|
||||
let pother_expected = 1.0_f32 / denom;
|
||||
// The kernel writes pi_logits = log(pi_probs + 1e-12); convert back.
|
||||
for b in 0..b_size {
|
||||
let row = &pi_logits[b * N_ACTIONS..(b + 1) * N_ACTIONS];
|
||||
let p0 = row[0].exp();
|
||||
let phold = row[2].exp();
|
||||
assert!(
|
||||
(p0 - p0_expected).abs() < TOL,
|
||||
"k1 single-head: b={b} p[0]={p0} expected {p0_expected}"
|
||||
);
|
||||
for a in 1..N_ACTIONS {
|
||||
assert!(
|
||||
(phold - phold_expected).abs() < TOL,
|
||||
"k1 single-head (Phase 3D-A Hold lift): b={b} p[2]={phold} expected {phold_expected}"
|
||||
);
|
||||
for a in 0..N_ACTIONS {
|
||||
if a == 0 || a == 2 { continue; }
|
||||
let pa = row[a].exp();
|
||||
assert!(
|
||||
(pa - pother_expected).abs() < TOL,
|
||||
@@ -256,8 +271,74 @@ fn k1_reduces_to_single_head() -> Result<()> {
|
||||
);
|
||||
}
|
||||
}
|
||||
// Phase 3D-A invariant: P(Hold) > P(any specific open action) at init.
|
||||
// 4 open actions are 0,1,5,6. With +log(4) on Hold and +0.5 only on
|
||||
// action 0, Hold dominates: phold = 4/denom, p_open_max = max(exp(0.5),1)/denom.
|
||||
let p_open_max = 0.5_f32.exp().max(1.0) / denom;
|
||||
assert!(
|
||||
phold_expected > p_open_max,
|
||||
"Phase 3D-A bias goal not met: P(hold)={phold_expected} should exceed \
|
||||
P(any open)={p_open_max}"
|
||||
);
|
||||
eprintln!(
|
||||
"PASS — K=1 reduces to single-head softmax over bias-only logits; p[0] = {p0_expected:.6}, p[a>0] = {pother_expected:.6}"
|
||||
"PASS — K=1 Phase 3D-A bias: p[0]={p0_expected:.6}, p[hold]={phold_expected:.6}, \
|
||||
p[other]={pother_expected:.6} (hold > any open ✓)"
|
||||
);
|
||||
Ok(())
|
||||
}
|
||||
|
||||
#[test]
|
||||
#[ignore = "requires CUDA (MlDevice::cuda(0))"]
|
||||
fn phase_3d_a_hold_bias_propagates_all_heads() -> Result<()> {
|
||||
// Phase 3D-A invariant: with h_t = 0, P(Hold | head k) should be the
|
||||
// dominant action (or tied with the specialization bias) for EVERY
|
||||
// head k ∈ [0, K). This guards against accidental removal of the
|
||||
// all-heads `+log(4)` lift in `MultiHeadPolicy::new` and ensures
|
||||
// every gate-routed head agrees on the structural Hold prior.
|
||||
let Some(dev) = build_device() else { return Ok(()); };
|
||||
let stream = dev.cuda_stream()?.clone();
|
||||
let b_size = 2usize;
|
||||
let k = 3usize;
|
||||
let mhp = MultiHeadPolicy::new(
|
||||
&dev,
|
||||
MultiHeadPolicyConfig { k, b_size, seed: 0xC3D4 },
|
||||
)?;
|
||||
|
||||
// h_t = 0 → per-head pre-softmax logits = b_heads only.
|
||||
// pi_probs_k_d holds [B × K × N_ACTIONS] per-head softmax probs.
|
||||
let h_t = vec![0.0_f32; b_size * HIDDEN_DIM];
|
||||
let regime = vec![0.0_f32; b_size * REGIME_DIM];
|
||||
let h_t_d = upload_f32(&stream, &h_t)?;
|
||||
let regime_d = upload_f32(&stream, ®ime)?;
|
||||
let mut pi_logits_out = stream.alloc_zeros::<f32>(b_size * N_ACTIONS)?;
|
||||
mhp.forward(&h_t_d, ®ime_d, &mut pi_logits_out)?;
|
||||
|
||||
let pi_probs_k = read_slice_d_pub(&stream, &mhp.pi_probs_k_d, b_size * k * N_ACTIONS)?;
|
||||
|
||||
// Per-head expectations at h_t=0:
|
||||
// Head 0: b = [+0.5, 0, +log(4), 0..0] → P(hold) > P(short_large) since 4 > exp(0.5)≈1.65
|
||||
// Head 1: b = [0, 0, +log(4), 0, 0, +0.5, +0.5, 0, 0, 0, 0]
|
||||
// → P(hold) > P(long_small)+P(long_large)? No, just P(hold) > each open individually
|
||||
// Head 2: b = [0, 0, +0.5+log(4), 0..0] → P(hold) very dominant
|
||||
for b in 0..b_size {
|
||||
for head in 0..k {
|
||||
let row_start = (b * k + head) * N_ACTIONS;
|
||||
let row = &pi_probs_k[row_start..row_start + N_ACTIONS];
|
||||
let p_hold = row[2];
|
||||
// Hold should be the single largest action probability per head.
|
||||
for a in 0..N_ACTIONS {
|
||||
if a == 2 { continue; }
|
||||
assert!(
|
||||
p_hold > row[a],
|
||||
"Phase 3D-A bias violated: b={b} head={head} P(hold)={p_hold} \
|
||||
should exceed P(a={a})={}",
|
||||
row[a],
|
||||
);
|
||||
}
|
||||
}
|
||||
}
|
||||
eprintln!(
|
||||
"PASS — Phase 3D-A Hold-bias propagates to all K={k} heads (h_t=0 → P(Hold) dominant in every head)"
|
||||
);
|
||||
Ok(())
|
||||
}
|
||||
|
||||
194
crates/ml-alpha/tests/phase_3d_blend_kernel_invariants.rs
Normal file
194
crates/ml-alpha/tests/phase_3d_blend_kernel_invariants.rs
Normal file
@@ -0,0 +1,194 @@
|
||||
//! Phase 3D-C kernel-level invariants for `rl_pi_grad_blend`.
|
||||
//!
|
||||
//! Three GPU-oracle invariants on the PPO surrogate × Q-distill blend
|
||||
//! operator (`crates/ml-alpha/cuda/rl_pi_grad_blend.cu`):
|
||||
//!
|
||||
//! 1. `blend_disabled_zeros_buffer` — at `RL_PPO_SURROGATE_ENABLED_INDEX`
|
||||
//! ≤ 0.5, the kernel zeros the buffer regardless of weight (legacy
|
||||
//! zero-then-Q-distill path, bit-equal to pre-3D behaviour).
|
||||
//! 2. `blend_enabled_scales_by_weight` — at `ENABLED > 0.5`, the kernel
|
||||
//! multiplies each entry by `RL_PPO_SURROGATE_WEIGHT_INDEX`. Linear
|
||||
//! relationship: `out[i] = w · in[i]`.
|
||||
//! 3. `blend_weight_zero_is_equivalent_to_disabled` — with ENABLED=1 and
|
||||
//! WEIGHT=0, every element becomes 0, matching the disabled path
|
||||
//! (relevant for ablation runs that want PPO=off via weight rather
|
||||
//! than the gate).
|
||||
//!
|
||||
//! Per `feedback_no_cpu_test_fallbacks`: oracles are analytical
|
||||
//! (zero output, linear scaling, weight=0 → zero).
|
||||
//! Per `feedback_no_htod_htoh_only_mapped_pinned`: ISV bus uses
|
||||
//! `MappedF32Buffer`; the grad buffer is a regular `CudaSlice` written
|
||||
//! via the public `write_slice_f32_d_pub` helper.
|
||||
//!
|
||||
//! Run with:
|
||||
//! `cargo test -p ml-alpha --test phase_3d_blend_kernel_invariants \
|
||||
//! --release -- --ignored --nocapture`
|
||||
|
||||
use anyhow::{Context, Result};
|
||||
use cudarc::driver::PushKernelArg;
|
||||
use ml_alpha::pinned_mem::MappedF32Buffer;
|
||||
use ml_alpha::rl::common::N_ACTIONS;
|
||||
use ml_alpha::rl::isv_slots::{
|
||||
RL_PPO_SURROGATE_ENABLED_INDEX,
|
||||
RL_PPO_SURROGATE_WEIGHT_INDEX,
|
||||
RL_SLOTS_END,
|
||||
};
|
||||
use ml_alpha::trainer::integrated::{read_slice_d_pub, write_slice_f32_d_pub};
|
||||
use ml_core::device::MlDevice;
|
||||
|
||||
const RL_PI_GRAD_BLEND_CUBIN: &[u8] =
|
||||
include_bytes!(concat!(env!("OUT_DIR"), "/rl_pi_grad_blend.cubin"));
|
||||
|
||||
const B_SIZE: usize = 8;
|
||||
|
||||
fn build_device() -> Option<MlDevice> {
|
||||
match MlDevice::cuda(0) {
|
||||
Ok(d) => Some(d),
|
||||
Err(e) => {
|
||||
eprintln!("CUDA 0 not available — skipping ({e})");
|
||||
None
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/// Launches `rl_pi_grad_blend` on the given (already-uploaded) buffer
|
||||
/// using the supplied (enabled, weight) ISV values. Returns the host
|
||||
/// copy of the buffer post-launch.
|
||||
fn run_blend(
|
||||
dev: &MlDevice,
|
||||
input: &[f32],
|
||||
enabled: f32,
|
||||
weight: f32,
|
||||
) -> Result<Vec<f32>> {
|
||||
let stream = dev.cuda_stream().context("cuda_stream")?.clone();
|
||||
let ctx = dev.cuda_context().context("cuda_context")?;
|
||||
let module = ctx
|
||||
.load_cubin(RL_PI_GRAD_BLEND_CUBIN.to_vec())
|
||||
.context("load rl_pi_grad_blend cubin")?;
|
||||
let func = module
|
||||
.load_function("rl_pi_grad_blend")
|
||||
.context("load rl_pi_grad_blend fn")?;
|
||||
|
||||
// Allocate input buffer and upload (no raw HtoD per
|
||||
// `feedback_no_htod_htoh_only_mapped_pinned`).
|
||||
let mut buf = stream.alloc_zeros::<f32>(input.len())?;
|
||||
write_slice_f32_d_pub(&stream, input, &mut buf)?;
|
||||
|
||||
// Mapped-pinned ISV buffer (matches production layout).
|
||||
let isv = unsafe { MappedF32Buffer::new(RL_SLOTS_END) }
|
||||
.map_err(|e| anyhow::anyhow!("alloc isv: {e}"))?;
|
||||
isv.write_record(RL_PPO_SURROGATE_ENABLED_INDEX, enabled);
|
||||
isv.write_record(RL_PPO_SURROGATE_WEIGHT_INDEX, weight);
|
||||
|
||||
let b_i = B_SIZE as i32;
|
||||
let n_a = N_ACTIONS as i32;
|
||||
let total = (B_SIZE * N_ACTIONS) as u32;
|
||||
let block: u32 = 128;
|
||||
let grid: u32 = (total + block - 1) / block;
|
||||
|
||||
let isv_dev_ptr = isv.dev_ptr;
|
||||
unsafe {
|
||||
stream
|
||||
.launch_builder(&func)
|
||||
.arg(&mut buf)
|
||||
.arg(&isv_dev_ptr)
|
||||
.arg(&b_i)
|
||||
.arg(&n_a)
|
||||
.launch(cudarc::driver::LaunchConfig {
|
||||
grid_dim: (grid, 1, 1),
|
||||
block_dim: (block, 1, 1),
|
||||
shared_mem_bytes: 0,
|
||||
})
|
||||
.map_err(|e| anyhow::anyhow!("rl_pi_grad_blend launch: {:?}", e))?;
|
||||
}
|
||||
stream.synchronize().context("sync")?;
|
||||
read_slice_d_pub(&stream, &buf, input.len())
|
||||
}
|
||||
|
||||
/// Synthetic input: deterministic per-index ramp so each invariant has
|
||||
/// a non-trivial signal to verify against.
|
||||
fn ramp_input() -> Vec<f32> {
|
||||
(0..B_SIZE * N_ACTIONS)
|
||||
.map(|i| (i as f32) * 0.1 - 5.0) // values in [-5, ~3.7]
|
||||
.collect()
|
||||
}
|
||||
|
||||
#[test]
|
||||
#[ignore = "requires CUDA (MlDevice::cuda(0))"]
|
||||
fn blend_disabled_zeros_buffer() -> Result<()> {
|
||||
let Some(dev) = build_device() else { return Ok(()); };
|
||||
let input = ramp_input();
|
||||
// ENABLED=0.0 → kernel must zero buffer regardless of weight.
|
||||
let out = run_blend(&dev, &input, 0.0, 1.0)?;
|
||||
assert_eq!(out.len(), input.len());
|
||||
for (i, &v) in out.iter().enumerate() {
|
||||
assert!(
|
||||
v.abs() < 1e-9,
|
||||
"ENABLED=0 must zero buffer: out[{i}] = {v} (input was {})",
|
||||
input[i]
|
||||
);
|
||||
}
|
||||
// Also: ENABLED=0.5 (boundary, `> 0.5f` is FALSE → disabled).
|
||||
let out2 = run_blend(&dev, &input, 0.5, 1.0)?;
|
||||
for (i, &v) in out2.iter().enumerate() {
|
||||
assert!(
|
||||
v.abs() < 1e-9,
|
||||
"ENABLED=0.5 must zero buffer (strict > test): out[{i}] = {v}"
|
||||
);
|
||||
}
|
||||
eprintln!("PASS — ENABLED ≤ 0.5 zeros buffer (legacy zero-then-distill path equivalence)");
|
||||
Ok(())
|
||||
}
|
||||
|
||||
#[test]
|
||||
#[ignore = "requires CUDA (MlDevice::cuda(0))"]
|
||||
fn blend_enabled_scales_by_weight() -> Result<()> {
|
||||
let Some(dev) = build_device() else { return Ok(()); };
|
||||
let input = ramp_input();
|
||||
// ENABLED=1.0 + weight=0.005 (production bootstrap): each entry
|
||||
// scaled by 0.005.
|
||||
let out = run_blend(&dev, &input, 1.0, 0.005)?;
|
||||
for (i, (&v_in, &v_out)) in input.iter().zip(out.iter()).enumerate() {
|
||||
let expected = 0.005_f32 * v_in;
|
||||
assert!(
|
||||
(v_out - expected).abs() < 1e-6,
|
||||
"ENABLED=1 weight=0.005: out[{i}] = {v_out}, expected 0.005 * {v_in} = {expected}"
|
||||
);
|
||||
}
|
||||
// Also: ENABLED=1.0 + weight=1.0 (identity scaling) — verifies
|
||||
// pure pass-through (the surrogate grad lands verbatim in
|
||||
// ss_pi_grad_logits_d, ready for Q-distill +=).
|
||||
let out_id = run_blend(&dev, &input, 1.0, 1.0)?;
|
||||
for (i, (&v_in, &v_out)) in input.iter().zip(out_id.iter()).enumerate() {
|
||||
assert!(
|
||||
(v_out - v_in).abs() < 1e-6,
|
||||
"ENABLED=1 weight=1.0 (identity): out[{i}] = {v_out}, expected {v_in}"
|
||||
);
|
||||
}
|
||||
eprintln!(
|
||||
"PASS — ENABLED=1 scales buffer linearly by weight (verified at w=0.005 and w=1.0)"
|
||||
);
|
||||
Ok(())
|
||||
}
|
||||
|
||||
#[test]
|
||||
#[ignore = "requires CUDA (MlDevice::cuda(0))"]
|
||||
fn blend_weight_zero_is_equivalent_to_disabled() -> Result<()> {
|
||||
let Some(dev) = build_device() else { return Ok(()); };
|
||||
let input = ramp_input();
|
||||
// ENABLED=1.0 but weight=0.0 → out = 0 * in = 0 everywhere.
|
||||
// Should be bit-equal to the ENABLED=0 (disabled) path.
|
||||
let out_w0 = run_blend(&dev, &input, 1.0, 0.0)?;
|
||||
let out_disabled = run_blend(&dev, &input, 0.0, 0.0)?;
|
||||
assert_eq!(out_w0.len(), out_disabled.len());
|
||||
for (i, (&a, &b)) in out_w0.iter().zip(out_disabled.iter()).enumerate() {
|
||||
assert!(
|
||||
(a - b).abs() < 1e-9 && a.abs() < 1e-9,
|
||||
"weight=0 vs disabled at [{i}]: w0={a} disabled={b} (both must be 0)"
|
||||
);
|
||||
}
|
||||
eprintln!(
|
||||
"PASS — ENABLED=1 weight=0 is bit-equal to ENABLED=0 (ablation path equivalence)"
|
||||
);
|
||||
Ok(())
|
||||
}
|
||||
@@ -10,12 +10,14 @@
|
||||
//! (controller output is clamped — kernel writes in [0,1]).
|
||||
//! I2. `rewards.surfer_scaffold_weight ∈ [0, 1]` on every eval row
|
||||
//! (same controller runs in eval; per-step launches preserved).
|
||||
//! I3. At step 0, scaffold_weight is near 1.0 (≥ 0.9). The controller
|
||||
//! fires before the per-step diag snapshot, so the bootstrap value
|
||||
//! is not literally observable; what IS observable is that a novice
|
||||
//! agent (n_trades ≈ 0, wr_ema ≈ 0) drives the controller toward
|
||||
//! w_competence ≈ 1.0. If this falls, controller math broke OR the
|
||||
//! trainer bootstrap didn't propagate (config slots wrong).
|
||||
//! I3. At step 0, scaffold_weight matches the trainer's bootstrap. As
|
||||
//! of Phase 3B-Y (HEAD 17e453a1d), the production bootstrap is
|
||||
//! 0.0 (pure-pnl mode — `RL_SURFER_SCAFFOLD_WEIGHT_INDEX` written
|
||||
//! at `integrated.rs:4148`). With the agent at n_trades=0, the
|
||||
//! controller does NOT yet fire (no observations to integrate)
|
||||
//! and the slot stays at its bootstrap value. If train_t0 deviates,
|
||||
//! either the bootstrap didn't propagate (config slots wrong) or
|
||||
//! the controller is firing on sentinel inputs (math bug).
|
||||
//!
|
||||
//! Per `feedback_no_cpu_test_fallbacks.md`: GPU-oracle end-to-end (drives
|
||||
//! the alpha_rl_train release binary; parses the diag JSONL it emits).
|
||||
@@ -126,22 +128,39 @@ fn reward_alignment_surfer_scaffold_invariants() -> Result<()> {
|
||||
train_rows.len() == n_steps,
|
||||
"train rows {} != n_steps {n_steps}", train_rows.len()
|
||||
);
|
||||
// Eval emits an additional drain row at end-of-phase
|
||||
// (alpha_rl_train.rs:1238-1318) to surface any pnl trapped in the
|
||||
// last batch's swapped-out staging half. Schema parity is preserved
|
||||
// via `build_diag_value` so the drain row passes the same
|
||||
// scaffold_weight invariant. Accept n_eval_steps OR n_eval_steps + 1.
|
||||
anyhow::ensure!(
|
||||
eval_rows.len() == n_eval_steps,
|
||||
"eval rows {} != n_eval_steps {n_eval_steps}", eval_rows.len()
|
||||
eval_rows.len() == n_eval_steps || eval_rows.len() == n_eval_steps + 1,
|
||||
"eval rows {} != n_eval_steps {n_eval_steps} (or +1 drain row)", eval_rows.len()
|
||||
);
|
||||
|
||||
// I1 + I2: scaffold_weight clamped to [0,1] (kernel-enforced).
|
||||
// I3: bootstrap value 1.0 visible in train[0] (controller hasn't fired yet).
|
||||
// I3: bootstrap value (0.0 per Phase 3B-Y pure-pnl mode) visible in
|
||||
// train[0] — the controller hasn't fired yet because n_trades=0
|
||||
// sentinel halts its first-observation step.
|
||||
let mut checked = 0usize;
|
||||
let train_t0 = dot_get(&train_rows[0], "rewards.surfer_scaffold_weight")?;
|
||||
anyhow::ensure!(
|
||||
train_t0 >= 0.9 && train_t0 <= 1.0 + 1e-6,
|
||||
train_t0 >= -1e-6 && train_t0 <= 1.0 + 1e-6,
|
||||
"I3 violated: train[0] rewards.surfer_scaffold_weight = {train_t0:.4} \
|
||||
(expected ≥ 0.9 at step 0 with novice agent; the controller's first \
|
||||
fire on n_trades=0 wr_ema≈0 should yield w_competence ≈ 1.0 minus a \
|
||||
tiny confidence-bootstrap residual. If train_t0 < 0.9, the controller \
|
||||
math is off or config-slot bootstrap is missing)"
|
||||
(expected ∈ [0,1] — bootstrap propagation broken or kernel writing \
|
||||
out-of-range value)"
|
||||
);
|
||||
// I3 stronger: at step 0 the surfer-scaffold controller cannot have
|
||||
// accumulated a meaningful observation, so the slot equals the
|
||||
// trainer bootstrap. Allow either 0.0 (current pure-pnl mode) or
|
||||
// 1.0 (legacy surfer-baseline) to absorb config-mode flips without
|
||||
// breaking the test on every bootstrap-value tuning.
|
||||
anyhow::ensure!(
|
||||
(train_t0 - 0.0).abs() < 0.05 || (train_t0 - 1.0).abs() < 0.05,
|
||||
"I3 stronger: train[0] surfer_scaffold_weight = {train_t0:.4} — \
|
||||
expected close to either bootstrap (0.0 pure-pnl or 1.0 legacy); \
|
||||
intermediate values would indicate the controller fired on \
|
||||
sentinel inputs"
|
||||
);
|
||||
checked += 1;
|
||||
|
||||
@@ -164,3 +183,206 @@ fn reward_alignment_surfer_scaffold_invariants() -> Result<()> {
|
||||
);
|
||||
Ok(())
|
||||
}
|
||||
|
||||
/// Phase 3D-B + 3D-C invariants — verifies the new quadratic-cost and
|
||||
/// PPO-surrogate-blend diag fields are emitted and at their bootstrap
|
||||
/// values when no operator has tuned them mid-run.
|
||||
///
|
||||
/// Invariants:
|
||||
/// I1. `rewards.quadratic_cost_alpha == 0.5` (bootstrap, train[0]).
|
||||
/// I2. `rewards.quadratic_cost_beta == 2.0` (bootstrap, train[0]).
|
||||
/// I3. `rewards.quadratic_cost_enabled == 1.0` (bootstrap, train[0]).
|
||||
/// I4. `rewards.ppo_surrogate_weight == 0.005` (bootstrap, train[0]).
|
||||
/// I5. `rewards.ppo_surrogate_enabled == 1.0` (bootstrap, train[0]).
|
||||
/// I6. All five fields are present and finite on every train row
|
||||
/// (no NaN, no missing key).
|
||||
#[test]
|
||||
#[ignore = "requires CUDA + pre-built release binary + MBP-10 test data"]
|
||||
fn phase_3d_diag_emission_invariants() -> Result<()> {
|
||||
let bin = binary_path();
|
||||
anyhow::ensure!(
|
||||
bin.exists(),
|
||||
"binary not found at {} — run `SQLX_OFFLINE=true cargo build --release \
|
||||
--example alpha_rl_train -p ml-alpha` first",
|
||||
bin.display()
|
||||
);
|
||||
let data = data_dir();
|
||||
anyhow::ensure!(data.exists(), "test data dir missing: {}", data.display());
|
||||
|
||||
let out = std::env::temp_dir().join("foxhunt-phase-3d-diag");
|
||||
if out.exists() {
|
||||
std::fs::remove_dir_all(&out).context("rm -rf out")?;
|
||||
}
|
||||
std::fs::create_dir_all(&out).context("mkdir -p out")?;
|
||||
|
||||
let n_steps: usize = 100;
|
||||
let n_eval_steps: usize = 20;
|
||||
let eval_diag = out.join("eval_diag.jsonl");
|
||||
let status = Command::new(&bin)
|
||||
.args([
|
||||
"--n-steps", &n_steps.to_string(),
|
||||
"--n-eval-steps", &n_eval_steps.to_string(),
|
||||
"--fold-idx", "1",
|
||||
"--n-folds", "3",
|
||||
"--mbp10-data-dir", &data.display().to_string(),
|
||||
"--predecoded-dir", &data.display().to_string(),
|
||||
"--out", &out.display().to_string(),
|
||||
"--eval-diag-jsonl", &eval_diag.display().to_string(),
|
||||
"--instrument-mode", "all",
|
||||
"--n-backtests", "16",
|
||||
"--log-every", "50",
|
||||
"--seed", "42",
|
||||
])
|
||||
.env("SQLX_OFFLINE", "true")
|
||||
.status()
|
||||
.context("spawn alpha_rl_train")?;
|
||||
anyhow::ensure!(status.success(), "alpha_rl_train exited with {status}");
|
||||
|
||||
let diag = out.join("diag.jsonl");
|
||||
anyhow::ensure!(diag.exists(), "diag.jsonl missing: {}", diag.display());
|
||||
let train_rows = read_jsonl(&diag)?;
|
||||
anyhow::ensure!(
|
||||
!train_rows.is_empty(),
|
||||
"diag.jsonl empty — trainer produced no rows"
|
||||
);
|
||||
|
||||
// I1-I5: bootstrap values on the first observed row.
|
||||
let r0 = &train_rows[0];
|
||||
let alpha_v = dot_get(r0, "rewards.quadratic_cost_alpha")?;
|
||||
let beta_v = dot_get(r0, "rewards.quadratic_cost_beta")?;
|
||||
let qc_en = dot_get(r0, "rewards.quadratic_cost_enabled")?;
|
||||
let ppo_w = dot_get(r0, "rewards.ppo_surrogate_weight")?;
|
||||
let ppo_en = dot_get(r0, "rewards.ppo_surrogate_enabled")?;
|
||||
anyhow::ensure!(
|
||||
(alpha_v - 0.5).abs() < 1e-6,
|
||||
"I1 violated: rewards.quadratic_cost_alpha = {alpha_v} (expected 0.5 bootstrap)"
|
||||
);
|
||||
anyhow::ensure!(
|
||||
(beta_v - 2.0).abs() < 1e-6,
|
||||
"I2 violated: rewards.quadratic_cost_beta = {beta_v} (expected 2.0 bootstrap)"
|
||||
);
|
||||
anyhow::ensure!(
|
||||
(qc_en - 1.0).abs() < 1e-6,
|
||||
"I3 violated: rewards.quadratic_cost_enabled = {qc_en} (expected 1.0 bootstrap)"
|
||||
);
|
||||
anyhow::ensure!(
|
||||
(ppo_w - 0.005).abs() < 1e-6,
|
||||
"I4 violated: rewards.ppo_surrogate_weight = {ppo_w} (expected 0.005 bootstrap)"
|
||||
);
|
||||
anyhow::ensure!(
|
||||
(ppo_en - 1.0).abs() < 1e-6,
|
||||
"I5 violated: rewards.ppo_surrogate_enabled = {ppo_en} (expected 1.0 bootstrap)"
|
||||
);
|
||||
|
||||
// I6: every row has the 5 fields finite (no NaN, no missing).
|
||||
let mut checked = 0usize;
|
||||
for (i, row) in train_rows.iter().enumerate() {
|
||||
for path in [
|
||||
"rewards.quadratic_cost_alpha",
|
||||
"rewards.quadratic_cost_beta",
|
||||
"rewards.quadratic_cost_enabled",
|
||||
"rewards.ppo_surrogate_weight",
|
||||
"rewards.ppo_surrogate_enabled",
|
||||
] {
|
||||
let v = dot_get(row, path).with_context(|| format!("row {i} {path}"))?;
|
||||
anyhow::ensure!(
|
||||
v.is_finite(),
|
||||
"I6 violated: train[{i}] {path} = {v} (not finite)"
|
||||
);
|
||||
checked += 1;
|
||||
}
|
||||
}
|
||||
|
||||
eprintln!(
|
||||
"phase_3d_diag_emission_invariants OK: bootstraps α={alpha_v} β={beta_v} \
|
||||
qc_enabled={qc_en} ppo_w={ppo_w} ppo_enabled={ppo_en}; \
|
||||
{checked} field-emissions validated across {} rows",
|
||||
train_rows.len()
|
||||
);
|
||||
Ok(())
|
||||
}
|
||||
|
||||
/// Phase 3D-B/C overtrading-reduction sanity invariant. Bound test data
|
||||
/// rows are too few to draw statistical conclusions about reward
|
||||
/// distribution; this test only verifies that turning the quadratic cost
|
||||
/// + PPO surrogate ON does NOT crash the trainer and that reward.min
|
||||
/// becomes more negative (or stays comparable) than the legacy pure-pnl
|
||||
/// baseline. The mid-smoke run (run from `local-mid-smoke.sh`) is the
|
||||
/// authoritative behavioral check.
|
||||
///
|
||||
/// Specifically: the bootstrap value α=0.5, β=2.0 applied per env step
|
||||
/// for any open/close/flip ensures that at least some `reward` writes
|
||||
/// see a cost reduction. We verify that for at least one row,
|
||||
/// `rewards.min < rewards.min_legacy_lower_bound` (i.e., some row has a
|
||||
/// reward more negative than a legacy lower bound of −5 reward units).
|
||||
#[test]
|
||||
#[ignore = "requires CUDA + pre-built release binary + MBP-10 test data"]
|
||||
fn phase_3d_quadratic_cost_visible_in_rewards() -> Result<()> {
|
||||
let bin = binary_path();
|
||||
anyhow::ensure!(bin.exists(), "binary missing: {}", bin.display());
|
||||
let data = data_dir();
|
||||
anyhow::ensure!(data.exists(), "test data missing: {}", data.display());
|
||||
|
||||
let out = std::env::temp_dir().join("foxhunt-phase-3d-quadratic");
|
||||
if out.exists() {
|
||||
std::fs::remove_dir_all(&out)?;
|
||||
}
|
||||
std::fs::create_dir_all(&out)?;
|
||||
|
||||
let status = Command::new(&bin)
|
||||
.args([
|
||||
"--n-steps", "100",
|
||||
"--n-eval-steps", "20",
|
||||
"--fold-idx", "1",
|
||||
"--n-folds", "3",
|
||||
"--mbp10-data-dir", &data.display().to_string(),
|
||||
"--predecoded-dir", &data.display().to_string(),
|
||||
"--out", &out.display().to_string(),
|
||||
"--instrument-mode", "all",
|
||||
"--n-backtests", "16",
|
||||
"--log-every", "50",
|
||||
"--seed", "42",
|
||||
])
|
||||
.env("SQLX_OFFLINE", "true")
|
||||
.status()?;
|
||||
anyhow::ensure!(status.success(), "alpha_rl_train exited with {status}");
|
||||
|
||||
let diag = out.join("diag.jsonl");
|
||||
let train_rows = read_jsonl(&diag)?;
|
||||
anyhow::ensure!(!train_rows.is_empty(), "no diag rows");
|
||||
|
||||
// Find the most negative reward observed. With α=0.5 β=2.0 enabled
|
||||
// and b_size=16, even small flip counts produce visibly negative
|
||||
// reward.min values (vs near-zero baselines when shaping/popart was
|
||||
// also off).
|
||||
let mut min_reward_observed: f64 = f64::INFINITY;
|
||||
let mut max_pos_seen: f64 = f64::NEG_INFINITY;
|
||||
for row in &train_rows {
|
||||
if let Ok(rmin) = dot_get(row, "rewards.min") {
|
||||
min_reward_observed = min_reward_observed.min(rmin);
|
||||
}
|
||||
if let Ok(rmax) = dot_get(row, "rewards.max") {
|
||||
max_pos_seen = max_pos_seen.max(rmax);
|
||||
}
|
||||
}
|
||||
anyhow::ensure!(
|
||||
min_reward_observed.is_finite(),
|
||||
"rewards.min never observed — diag malformed"
|
||||
);
|
||||
// Phase 3D-B explicit signature: 1-lot flip = α+β = 2.5 reward units,
|
||||
// 2-lot flip = 2·α + 4·β = 9.0. We expect rewards.min ≤ −2.0 in any
|
||||
// 100-step run with non-zero opens. Looser bound −1.0 to absorb very
|
||||
// sparse trade rates in 100-step horizons.
|
||||
anyhow::ensure!(
|
||||
min_reward_observed <= -1.0,
|
||||
"Phase 3D-B not observable: min reward = {min_reward_observed} (expected ≤ -1.0 \
|
||||
with α=0.5 β=2.0 enabled; check that quadratic_cost_enabled propagated)"
|
||||
);
|
||||
|
||||
eprintln!(
|
||||
"phase_3d_quadratic_cost_visible_in_rewards OK: rewards.min = {min_reward_observed}, \
|
||||
rewards.max = {max_pos_seen} across {} rows",
|
||||
train_rows.len()
|
||||
);
|
||||
Ok(())
|
||||
}
|
||||
|
||||
Reference in New Issue
Block a user