Files
foxhunt/crates/ml-backtesting/tests/threshold_and_cost.rs
jgrusewski fef5939556 feat(ml-backtesting): cold-start stopgap — max-confidence bytecode policy (Q1/Tier1)
The threshold-tuning smoke at 81decf40f produced n_trades=0 despite
74.6% of decisions having max_conv ≥ 0.30 — the linear-weighted-mean
aggregator in decision_policy_default is structurally dilution-bound
at cold-start (per spec §1).

Q1 stopgap: when sim_variants[i].use_cold_start_stopgap = true, the
harness uploads a max-confidence Strategy bytecode program for that
backtest, routing decisions through decision_policy_program with
OP_AGG_MAX_CONFIDENCE. Existing kernel; zero CUDA changes.

Field additions (atomically across BatchedSimConfig + UniformSimParams
+ ResolvedSimVariant + SweepBase.SimVariant) — every UniformSimParams
literal migrated to include use_cold_start_stopgap: false (default).
The sweep YAML's sim_variants entry sets it to true only for the
validation run; production deployability uses Q2's kernel fix instead.

Sweep YAML (config/ml/sweep_smoke.yaml) flipped to use_cold_start_stopgap=true
at threshold=0.0, cost=0.125 — same anchor as the threshold-tuning
smoke that produced n_trades=0, for direct comparison.

This is a VALIDATION step. Cluster smoke at this commit MUST produce
n_trades > 100 + finite metrics. Q2's kernel CBSW immediately follows
and deletes this entire stopgap atomically (field, harness branch,
YAML setting, every literal).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-19 22:00:52 +02:00

80 lines
3.3 KiB
Rust

//! P4 regression: threshold gate + per-fill cost integration.
//!
//! Tight tests focused on the GATE's pre-Kelly skip path and the cost
//! plumbing in apply_fill_to_pos. The full kelly_state_sees_net_return
//! end-to-end test requires a full submit_market → fill → close sequence
//! which is exercised by the production smoke; here we test the kernel
//! contract in isolation.
use anyhow::Result;
use ml_backtesting::sim::{BatchedSimConfig, LobSimCuda, UniformSimParams};
use ml_core::device::MlDevice;
fn cfg_with_threshold(n: usize, threshold: f32, cost: f32) -> BatchedSimConfig {
BatchedSimConfig::from_uniform(n, &UniformSimParams {
target_annual_vol_units: 50.0,
annualisation_factor: 825.0,
max_lots: 5,
latency_ns: 0,
kelly_frac_floor: 0.20,
sharpe_weight_floor: 0.10,
threshold,
cost_per_lot_per_side: cost,
use_cold_start_stopgap: false,
})
}
#[test]
#[ignore = "requires CUDA"]
fn threshold_gate_skips_low_conviction() -> Result<()> {
let dev = match MlDevice::cuda(0) {
Ok(d) => d,
Err(e) => { eprintln!("skipping: cuda device unavailable ({e})"); return Ok(()); }
};
let mut sim = LobSimCuda::new(1, &dev)?;
// p_h = 0.51 → max_conviction = 0.02, well below threshold = 0.10.
sim.broadcast_alpha(&[0.51, 0.51, 0.51, 0.51, 0.51])?;
sim.step_decision_with_latency(0, &cfg_with_threshold(1, 0.10, 0.0))?;
let (side, size) = sim.read_market_target(0)?;
assert_eq!(side, 2, "side should be noop under threshold gate; got side={side}");
assert_eq!(size, 0);
Ok(())
}
#[test]
#[ignore = "requires CUDA"]
fn threshold_gate_allows_high_conviction() -> Result<()> {
let dev = match MlDevice::cuda(0) {
Ok(d) => d,
Err(e) => { eprintln!("skipping: cuda device unavailable ({e})"); return Ok(()); }
};
let mut sim = LobSimCuda::new(1, &dev)?;
// p_h = 0.8 → max_conviction = 0.6, above threshold = 0.10.
// (p=0.7 would clear the gate but ss=0.4 rounds to lots=0; needs p≥0.75
// with kelly_floor=0.20 + max_lots=5 to clear the lots>=1 rounding.)
sim.broadcast_alpha(&[0.8, 0.8, 0.8, 0.8, 0.8])?;
sim.step_decision_with_latency(0, &cfg_with_threshold(1, 0.10, 0.0))?;
let (side, size) = sim.read_market_target(0)?;
assert_eq!(side, 0, "side should be buy with strong alpha; got side={side}");
assert!(size >= 1, "size {size} < 1 — threshold gate may be over-restricting");
Ok(())
}
#[test]
#[ignore = "requires CUDA"]
fn threshold_zero_is_passthrough() -> Result<()> {
// Sanity check: threshold = 0.0 should behave exactly like the
// P1 cold-start path (no gate). max_conviction >= 0 always.
let dev = match MlDevice::cuda(0) {
Ok(d) => d,
Err(e) => { eprintln!("skipping: cuda device unavailable ({e})"); return Ok(()); }
};
let mut sim = LobSimCuda::new(1, &dev)?;
sim.broadcast_alpha(&[0.8, 0.8, 0.8, 0.8, 0.8])?;
sim.step_decision_with_latency(0, &cfg_with_threshold(1, 0.0, 0.0))?;
let (side, size) = sim.read_market_target(0)?;
assert_eq!(side, 0, "threshold=0 with strong alpha should pass through; got side={side}");
assert!(size >= 1, "size {size} < 1 — passthrough behaviour broken");
Ok(())
}