test(dqn): Phase 2 Test 2.B — production vs standalone (KS-fallback)

Plan C Phase 2 T7. #[ignore]-gated GPU test that runs BOTH the
production experience_action_select kernel AND the standalone
direction_thompson_v2_test (added in commit 5de5e546a) on identical
Phase 0 Test 0.D failure-mode inputs and asserts the resulting
direction histograms are statistically indistinguishable.

Path (a) — bit-identical comparison via a shared pre-computed uniform
array — would require editing the standalone v2 kernel's signature to
accept a `float* uniforms` instead of generating its own LCG draws.
Out of scope for the 1-2-hour dispatch. Path (b) — KS-style histogram
comparison — is used.

Setup:
- Single per-direction logits tile (Phase 0 Test 0.D failure mode)
- Production: tile replicated across batch=100 000 samples; Philox
  keyed on (i, timestep=0, ctr) per thread
- Standalone v2: same tile fed to single-tile launcher with n_seeds=100 000;
  LCG keyed on (base_seed=0 + seed_idx) per thread

Assertion:
- KS distance over the {Short, Hold, Long, Flat} histograms ≤ 0.02
  (n=100k empirical-CDF noise floor for matching distributions is
  ~4·sqrt(2/n) ≈ 0.018; algorithm divergences would shift mass by
  O(10%) → KS ≈ O(0.1), easily detected)
- Sanity: both histograms have P(Long+Short) ≥ 0.20 (rules out
  coincidental shared-mode collapse with KS ≈ 0)

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
This commit is contained in:
jgrusewski
2026-04-29 18:17:26 +02:00
parent 9f226f429d
commit 1a80fcab10
2 changed files with 262 additions and 0 deletions

View File

@@ -1204,3 +1204,241 @@ fn test_2a_production_kernel_modes() {
p_active, train_counts,
);
}
// ── Test 2.B helper: standalone v2 kernel launcher ────────────────────────
/// Cached handle for the v2 standalone direction Thompson kernel.
static STANDALONE_V2: OnceLock<CudaFunction> = OnceLock::new();
fn standalone_thompson_v2(stream: &Arc<CudaStream>) -> &'static CudaFunction {
STANDALONE_V2.get_or_init(|| {
let module = stream
.context()
.load_cubin(TEST_KERNEL_CUBIN.to_vec())
.expect("load thompson_test_kernel cubin (v2)");
module
.load_function("direction_thompson_v2_test")
.expect("load direction_thompson_v2_test")
})
}
/// Launch the standalone v2 direction Thompson test kernel and return the
/// per-seed direction picks. Inputs are SINGLE-tile (the v2 wrapper uses
/// `i=0` internally — see `compute_atom_values_test`):
/// - `b_logits` [b0_size * n_atoms]
/// - `support_tile` [b0_size * 3] stride-3 (v_min, v_max, delta_z)
fn launch_standalone_v2(
stream: &Arc<CudaStream>,
b_logits_tile: &[f32],
support_tile: &[f32],
n_atoms: i32,
base_seed: u32,
n_seeds: usize,
) -> Vec<i32> {
assert_eq!(b_logits_tile.len(), (PROD_B0 as usize) * (n_atoms as usize));
assert_eq!(support_tile.len(), (PROD_B0 as usize) * 3);
let kernel = standalone_thompson_v2(stream);
let logits_buf = stream.clone_htod(b_logits_tile).expect("upload v2 logits");
let support_buf = stream.clone_htod(support_tile).expect("upload v2 support");
let buffer = unsafe { MappedI32Buffer::new(n_seeds) }
.expect("alloc MappedI32Buffer for v2 thompson seeds");
let n_seeds_i32 = i32::try_from(n_seeds).expect("n_seeds fits in i32");
let block_dim_x: u32 = 256;
let grid_dim_x: u32 = (n_seeds as u32).div_ceil(block_dim_x);
let null_atom_positions: u64 = 0;
let dev_ptr = buffer.dev_ptr;
unsafe {
stream
.launch_builder(kernel)
.arg(&logits_buf)
.arg(&support_buf)
.arg(&null_atom_positions)
.arg(&PROD_B0)
.arg(&n_atoms)
.arg(&base_seed)
.arg(&n_seeds_i32)
.arg(&dev_ptr)
.launch(LaunchConfig {
grid_dim: (grid_dim_x.max(1), 1, 1),
block_dim: (block_dim_x, 1, 1),
shared_mem_bytes: 0,
})
.expect("launch direction_thompson_v2_test");
}
stream
.synchronize()
.expect("synchronize after standalone v2 launch");
buffer.read_all()
}
/// Two-sample Kolmogorov-Smirnov-style histogram distance for a discrete
/// 4-bin distribution. Returns `max_d |F1(d) - F2(d)|` where F is the
/// empirical CDF.
fn ks_distance_4bin(hist1: &[u32; 4], hist2: &[u32; 4]) -> f32 {
let n1: u32 = hist1.iter().sum();
let n2: u32 = hist2.iter().sum();
let mut cum1 = 0_u32;
let mut cum2 = 0_u32;
let mut max_d = 0.0_f32;
for i in 0..4 {
cum1 += hist1[i];
cum2 += hist2[i];
let d = ((cum1 as f32 / n1 as f32) - (cum2 as f32 / n2 as f32)).abs();
if d > max_d {
max_d = d;
}
}
max_d
}
/// Phase 2 Test 2.B — Production kernel matches Phase 0 standalone wrapper
/// in distribution.
///
/// Bit-equivalent comparison is impossible without modifying the standalone
/// kernel signature: the production direction Thompson uses Philox keyed on
/// `(i, timestep, rng_ctr)`; the standalone v2 uses LCG keyed on
/// `base_seed + seed_idx`. Path (a) (shared uniform array) requires
/// editing `direction_thompson_v2_test` to take a `float* uniforms` instead
/// of generating its own — out of scope for a 1-2-hour dispatch.
///
/// Path (b) — KOLMOGOROV-SMIRNOV-STYLE DISTRIBUTION COMPARISON. Run both
/// kernels on identical inputs over 100 000 samples / seeds. The two output
/// histograms over directions {Short, Hold, Long, Flat} must be statistically
/// indistinguishable: KS statistic ≤ 0.02 (≈ 4 * sqrt(2/100000), well above
/// the noise floor for matching distributions).
///
/// The shared input is the Phase 0 Test 0.D failure mode (Flat/Hold = δ(0),
/// Long/Short = log-Gaussian at v=-0.001, σ=0.05). Both kernels must reach
/// the same active-direction frequency under Thompson. A KS distance >0.02
/// would localise to either: (i) the inverse-CDF math diverges (would
/// shift mass between bins by tens of %); or (ii) the softmax/atom-value
/// helpers diverge (same effect). RNG quality contributes <0.005 noise at
/// n=100k.
#[test]
#[ignore]
fn test_2b_production_matches_standalone() {
let stream = make_test_stream();
let n_atoms: i32 = 21;
let v_min = -0.5_f32;
let v_max = 0.5_f32;
let delta_z = (v_max - v_min) / (n_atoms as f32 - 1.0);
// Build a single per-direction logits tile matching Phase 0 Test 0.D.
let mut tile = vec![0.0_f32; (PROD_B0 as usize) * (n_atoms as usize)];
let center_atom = ((0.0 - v_min) / delta_z).round() as usize;
tile[3 * n_atoms as usize + center_atom] = 10.0; // Flat δ(v=0)
tile[n_atoms as usize + center_atom] = 10.0; // Hold δ(v=0)
for a in 0..n_atoms as usize {
let v = v_min + (a as f32) * delta_z;
let mean = -0.001_f32;
let sigma = 0.05_f32;
let logit = -((v - mean).powi(2)) / (2.0 * sigma * sigma);
tile[2 * n_atoms as usize + a] = logit; // Long
tile[a] = logit; // Short
}
let support_tile = {
let mut s = vec![0.0_f32; (PROD_B0 as usize) * 3];
for d in 0..(PROD_B0 as usize) {
s[d * 3] = v_min;
s[d * 3 + 1] = v_max;
s[d * 3 + 2] = delta_z;
}
s
};
// ── Production kernel ────────────────────────────────────────────────
//
// Replicate the single tile across the production-kernel batch. Each
// thread reads the same tile but draws Philox(i, timestep=0, ctr).
let batch: usize = 100_000;
let mut full_logits = Vec::with_capacity(batch * tile.len());
for _ in 0..batch {
full_logits.extend_from_slice(&tile);
}
let mut full_support = Vec::with_capacity(batch * support_tile.len());
for _ in 0..batch {
full_support.extend_from_slice(&support_tile);
}
let q_values = vec![0.0_f32; batch * PROD_Q_STRIDE];
let mut fixture = ProdActionSelectFixture::new(
Arc::clone(&stream),
batch,
n_atoms,
&q_values,
&full_logits,
&full_support,
);
let prod_actions = fixture.launch(/*eval_mode=*/ false, /*timestep=*/ 0);
let mut prod_hist = [0_u32; 4];
for action in &prod_actions {
let d = decode_dir(*action) as usize;
prod_hist[d] += 1;
}
// ── Standalone v2 kernel ─────────────────────────────────────────────
//
// Standalone takes the SINGLE tile + support tile and runs each thread
// with `seed_idx` keying its LCG. 100 000 seeds matches the production
// batch size for histogram parity.
let n_seeds: usize = 100_000;
let standalone_picks = launch_standalone_v2(
&stream,
&tile,
&support_tile,
n_atoms,
/*base_seed=*/ 0,
n_seeds,
);
let mut stand_hist = [0_u32; 4];
for d in &standalone_picks {
stand_hist[*d as usize] += 1;
}
println!(
"Test 2.B production hist (S/H/L/F): {:?} (n={})",
prod_hist, batch,
);
println!(
"Test 2.B standalone hist (S/H/L/F): {:?} (n={})",
stand_hist, n_seeds,
);
let ks = ks_distance_4bin(&prod_hist, &stand_hist);
println!("Test 2.B KS distance: {:.5}", ks);
// KS threshold 0.02 — at n=100k the empirical CDF noise floor for
// matching distributions is ~4*sqrt(2/n) ≈ 0.018. Allow a small
// additional margin (0.02) for f32 rounding differences between the
// two kernels' independent softmax / inverse-CDF passes.
//
// Algorithm divergences (different inverse-CDF math, different softmax
// normalisation, etc.) would shift mass between bins by O(10%), so KS
// would be O(0.1) — well above 0.02. RNG implementation differences
// alone (Philox vs LCG, both good random sources) contribute < 0.005.
assert!(
ks <= 0.02,
"KS distance prod vs standalone v2 = {:.5} exceeds 0.02 — kernels' \
direction-Thompson distributions diverge.\n prod = {:?}\n stand = {:?}",
ks, prod_hist, stand_hist,
);
// Sanity: both histograms must have non-trivial active-direction mass.
// If both kernels collapsed to the same mode, KS would still be ~0 but
// the test wouldn't actually verify Thompson behaviour. With the Phase
// 0 0.D failure setup we expect P(Long+Short) ≥ 0.20 in both.
let prod_active = (prod_hist[0] + prod_hist[2]) as f32 / batch as f32;
let stand_active = (stand_hist[0] + stand_hist[2]) as f32 / n_seeds as f32;
assert!(
prod_active >= 0.20,
"Production active frac {:.4} < 0.20 — Thompson should explore",
prod_active,
);
assert!(
stand_active >= 0.20,
"Standalone active frac {:.4} < 0.20 — Thompson should explore",
stand_active,
);
}

View File

@@ -2095,5 +2095,29 @@ asserts:
Long+Short picks is ≥ 30% — Thompson sampling must explore
directional alternatives despite their lower E[Q].
### Test 2.B — Production matches Phase 0 standalone (Task T7)
`crates/ml/src/trainers/dqn/distributional_q_tests.rs::test_2b_production_matches_standalone`
launches BOTH the production `experience_action_select` kernel AND the
standalone `direction_thompson_v2_test` kernel (added by commit
`5de5e546a`) on identical Phase 0 Test 0.D failure-mode inputs and
verifies the resulting direction histograms are statistically
indistinguishable.
Bit-equivalence path (a) — drive both kernels from a shared
pre-computed uniform array — was deferred (would require modifying the
standalone v2 kernel's signature). Path (b) — KS-style histogram
comparison — is used: 100 000 samples / seeds per kernel; assert KS
distance ≤ 0.02 (well above the empirical CDF noise floor of
~4·sqrt(2/n) ≈ 0.018 at n=100k for matching distributions). Algorithm
divergences (different inverse-CDF math, different softmax
normalisation) would shift mass between bins by O(10%), giving KS
≈ O(0.1) — easily detected. RNG implementation differences alone
(Philox vs LCG, both good random sources) contribute < 0.005 noise.
A sanity assertion also verifies P(Long+Short) ≥ 0.20 in both
histograms — without this a coincidental shared-mode collapse would
produce KS ≈ 0 without actually exercising Thompson behaviour.
Plan C T11 follow-up A.1 (2026-04-29): reset `prev_epoch_q_mean` and `adaptive_tau` in `DQNTrainer::reset_for_fold` (`crates/ml/src/trainers/dqn/trainer/mod.rs`). Q-drift kill criterion's prev-baseline carried across fold boundaries because `q_value_history` was cleared but its derived `prev_epoch_q_mean` was not — the ratio at fold-N epoch 0 was computed against fold-(N-1) epoch 5's q_mean. Companion `adaptive_tau` modulation state also reset to `hyperparams.tau` baseline. Independent bug fix; applies regardless of Plan C Thompson outcome. Identified by the Goal-1 q-drift research (researcher report 2026-04-29) when reconciling a52d99613's q_mean=5.003 fold-1 spike (no kill — criterion was added later in 1c917e3cb) against Plan C's fold-0 ep2 trigger.