feat(sp21): T2.2 Phase 1 Step B — per-trade tape buffers + readback (atomic)

Replaces Step A's NULL launcher passes with real device buffers.
Both kernel launch sites in gpu_backtest_evaluator.rs (backtest_env_step
and backtest_env_step_batch) now pass real per-trade tape pointers.
The kernel's per-trade emission block fires unconditionally on close
events — single-threaded per-window writes preserve event ordering and
enable race-free counter increment without atomicAdd.

New constants + types:
  - MAX_TRADES_PER_WINDOW = 200_000 (typical eval window bar count;
    per-window memory: 4 SoA buffers × 4 bytes × 200k = 3.2MB)
  - pub struct EvalTrade with 5 fields: bar_index, pnl, holding_bars,
    direction, magnitude. Does NOT include predicted_q / ensemble_var
    — those need entry-time captures (entry_q, entry_var in portfolio
    state) deferred to Phase 1.5.

New struct fields on GpuBacktestEvaluator:
  - per_trade_pnl_buf:          CudaSlice<f32> [n_windows × MAX_TRADES]
  - per_trade_holding_bars_buf: CudaSlice<u32>
  - per_trade_bar_index_buf:    CudaSlice<u32>
  - per_trade_dir_mag_buf:      CudaSlice<u32> (packed dir/mag)
  - per_trade_count_buf:        CudaSlice<u32> [n_windows]

New methods:
  - reset_per_trade_tape (folded into reset_evaluation_state): zeros
    the count buffer at the start of each eval window. SoA value
    buffers don't need zeroing — read up to count[w] only.
  - pub fn read_per_trade_tape(&self) -> Result<Vec<EvalTrade>, MLError>:
    reads count buffer first (cheap), early-returns empty if no trades,
    else reads 4 SoA buffers (~16MB DtoH at PCIe ≈ 1ms) and flattens
    window-major into chronological Vec<EvalTrade>.

Phase 2 follow-up (next commit) — wire read_per_trade_tape to the
enrichment caller in training_loop.rs:1510-1568, replacing
extract_eval_trades_from_metrics (the fake-trade synthesizer).

Phase 1.5 follow-up (if Phase 2 keeps E1+E5) — add entry_q + entry_var
to portfolio_state at trade open, extend per-trade tape with 6th/7th
SoA buffers.

Affected files:
  - crates/ml/src/cuda_pipeline/gpu_backtest_evaluator.rs
    (constants, struct fields, alloc, construction, 2 launcher sites,
    reset_evaluation_state addition, read_per_trade_tape method)

Verification:
  - cargo check -p ml --tests: passes (warnings only)
  - GPU oracle tests: behavior preserved by construction (existing
    WindowMetrics aggregator unaffected — separate kernel)

Plan reference: docs/plans/2026-05-10-sp21-train-eval-coherence-isv-defrost.md
T2.2 multi-phase scope; Phase 1 Step B closure.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
This commit is contained in:
jgrusewski
2026-05-10 21:26:40 +02:00
parent 5d01190e19
commit 7d538d9304
2 changed files with 296 additions and 24 deletions

View File

@@ -96,6 +96,44 @@ const DQN_BACKTEST_CHUNK_SIZE: usize = 512;
/// `KELLY_STATS_SIZE` in `backtest_env_kernel.cu`.
const BACKTEST_KELLY_STATS_SIZE: usize = 4;
/// SP21 T2.2 Phase 1 (2026-05-10) — per-trade tape capacity per window.
/// 200_000 trades is a conservative upper bound: typical eval windows are
/// 200k bars (xmd6b: window_bars=214654) with at most ~1 close per bar
/// (observed trade_count=139251 at trades_per_bar=0.65). Setting capacity
/// equal to the bar count guarantees no trade is dropped under any
/// trading frequency. Memory cost per window: 4 × 4 bytes × 200k = 3.2MB
/// (5 × 4 with the per-window count u32). For typical n_windows=5
/// walk-forward setups: 16MB total — acceptable for eval (fold-boundary
/// alloc, not per-step). The kernel guards against overflow via
/// `if (slot < max_trades_per_window)` so larger-than-expected windows
/// gracefully truncate the tape instead of crashing.
pub const MAX_TRADES_PER_WINDOW: usize = 200_000;
/// SP21 T2.2 Phase 1 (2026-05-10) — host-side per-trade record decoded
/// from the GPU per-trade tape buffers. The kernel writes the SoA tape;
/// `read_per_trade_tape()` decodes packed `dir_mag` and zips the slices
/// into `Vec<EvalTrade>` for the enrichment phase. Field order matches
/// what the enrichment functions in `trainer/enrichment.rs` consume.
#[derive(Debug, Clone, Copy)]
pub struct EvalTrade {
/// Global bar index (0-based step within the window) when the trade
/// CLOSED. Used by E6 (winner indices) for priority-replay tagging.
pub bar_index: u32,
/// Realized per-trade return = `prev_position × (close entry_price)
/// / pre_trade_equity`. Dimensionless (fraction of equity). Used by
/// E1 (Q-correction), E5 (agreement-Sharpe), E6/E7/E8.
pub pnl: f32,
/// Trade duration in bars (capture of `hold_time` BEFORE close). Used
/// by E3 (gamma from holding-bars) and E4 (urgency-error rate).
pub holding_bars: u32,
/// Direction at close: 0=Short, 1=Hold, 2=Long, 3=Flat. Decoded from
/// the kernel-packed `dir_mag` (hi-16 bits).
pub direction: u8,
/// Magnitude at close: 0=Quarter, 1=Half, 2=Full. Decoded from
/// `dir_mag` (lo-16 bits).
pub magnitude: u8,
}
// ── DQN backtest forward config ───────────────────────────────────────────────
/// Configuration for cuBLAS-based DQN forward pass in backtest evaluation.
@@ -406,6 +444,29 @@ pub struct GpuBacktestEvaluator {
done_buf: CudaSlice<i32>, // [n_windows]
actions_buf: CudaSlice<i32>, // [n_windows]
actions_history_buf: CudaSlice<i32>, // [n_windows * max_len]
/// SP21 T2.2 Phase 1 (2026-05-10) — per-trade tape device buffers.
/// Written by `backtest_env_step` / `backtest_env_step_batch` at
/// trade-close events (single thread per window writes its own slot
/// — race-free counter increment without atomicAdd per
/// `feedback_no_atomicadd`). Read host-side via
/// `read_per_trade_tape()` after each eval window completes; the
/// resulting `Vec<EvalTrade>` is the authoritative per-trade signal
/// for the enrichment phase (E1-E8) — replaces the prior
/// fake-trade synthesizer (`extract_eval_trades_from_metrics`)
/// per `feedback_no_stubs`. Reset to zero at `reset_per_trade_tape()`
/// at the start of each eval window.
///
/// Layout: `[n_windows × MAX_TRADES_PER_WINDOW]` for the SoA buffers,
/// `[n_windows]` for the count. SoA chosen over AoS for GPU coalescing
/// (each kernel thread writes one slot per buffer at the same index).
/// `dir_mag` packs direction (hi-16) and magnitude (lo-16) so the
/// per-trade emission stays at 4 buffers + count.
per_trade_pnl_buf: CudaSlice<f32>, // [n_windows * MAX_TRADES_PER_WINDOW]
per_trade_holding_bars_buf: CudaSlice<u32>, // [n_windows * MAX_TRADES_PER_WINDOW]
per_trade_bar_index_buf: CudaSlice<u32>, // [n_windows * MAX_TRADES_PER_WINDOW]
per_trade_dir_mag_buf: CudaSlice<u32>, // [n_windows * MAX_TRADES_PER_WINDOW]
per_trade_count_buf: CudaSlice<u32>, // [n_windows]
/// Intent magnitude history — pure-policy magnitude argmax from each
/// eval step, recorded BEFORE the Hold/Flat dir_idx forcing in the
/// action-select kernel and BEFORE Kelly/margin caps modify exposure
@@ -907,6 +968,28 @@ impl GpuBacktestEvaluator {
.alloc_zeros::<f32>(n_windows * BACKTEST_KELLY_STATS_SIZE)
.map_err(|e| MLError::ModelError(format!("kelly_stats alloc: {e}")))?;
// SP21 T2.2 Phase 1 (2026-05-10) — per-trade tape device buffers.
// Sized at MAX_TRADES_PER_WINDOW per window (3.2MB per window for
// the 4 SoA buffers + count). Zero-init: per_trade_count[w]=0 so
// the first close in each window writes at slot 0. `reset_per_trade_tape`
// re-zeros at the start of each eval window so trade indices don't
// accumulate across folds.
let per_trade_pnl_buf = stream
.alloc_zeros::<f32>(n_windows * MAX_TRADES_PER_WINDOW)
.map_err(|e| MLError::ModelError(format!("per_trade_pnl alloc: {e}")))?;
let per_trade_holding_bars_buf = stream
.alloc_zeros::<u32>(n_windows * MAX_TRADES_PER_WINDOW)
.map_err(|e| MLError::ModelError(format!("per_trade_holding_bars alloc: {e}")))?;
let per_trade_bar_index_buf = stream
.alloc_zeros::<u32>(n_windows * MAX_TRADES_PER_WINDOW)
.map_err(|e| MLError::ModelError(format!("per_trade_bar_index alloc: {e}")))?;
let per_trade_dir_mag_buf = stream
.alloc_zeros::<u32>(n_windows * MAX_TRADES_PER_WINDOW)
.map_err(|e| MLError::ModelError(format!("per_trade_dir_mag alloc: {e}")))?;
let per_trade_count_buf = stream
.alloc_zeros::<u32>(n_windows)
.map_err(|e| MLError::ModelError(format!("per_trade_count alloc: {e}")))?;
// plan_isv_buf: [N, SL_PORTFOLIO_PLAN_DIM=7] — live plan_isv per window.
// Populated each step by backtest_plan_state_isv (val plan_isv parity,
// task #94) and consumed by the next step's backtest_state_gather.
@@ -1123,6 +1206,13 @@ impl GpuBacktestEvaluator {
done_buf,
actions_buf,
actions_history_buf,
// SP21 T2.2 Phase 1 (2026-05-10) — per-trade tape buffers
// (see field-decl docs for layout / lifecycle).
per_trade_pnl_buf,
per_trade_holding_bars_buf,
per_trade_bar_index_buf,
per_trade_dir_mag_buf,
per_trade_count_buf,
intent_mag_buf: None,
picked_action_history_buf: None,
states_buf,
@@ -1881,9 +1971,95 @@ impl GpuBacktestEvaluator {
self.stream.memset_zeros(&mut self.kelly_stats_buf)
.map_err(|e| MLError::ModelError(format!("reset kelly_stats: {e}")))?;
// SP21 T2.2 Phase 1 (2026-05-10) — reset per-trade tape counters
// so the first close in each window writes at slot 0. The SoA
// value buffers (pnl, holding_bars, bar_index, dir_mag) don't
// need zeroing — they're only read up to `count[w]` host-side,
// and stale values past `count[w]` are never visited by the
// kernel writer (which guards on `slot < max_trades_per_window`)
// or the host reader (which slices `[0..count[w]]`).
self.stream.memset_zeros(&mut self.per_trade_count_buf)
.map_err(|e| MLError::ModelError(format!("reset per_trade_count: {e}")))?;
Ok(())
}
/// SP21 T2.2 Phase 1 (2026-05-10) — read the per-trade tape from the
/// device into a host-side `Vec<EvalTrade>`. Called after the eval
/// window completes (post `launch_metrics_and_record_event` +
/// `consume_metrics_after_event`); flattens the per-window SoA
/// buffers into one chronological-by-window concatenated `Vec`.
///
/// The kernel emits trades in close-event order WITHIN each window
/// (single-threaded per-window writes preserve order). The flatten
/// concatenates by window — call sites that need per-window tapes
/// can reconstruct via `count[]` if needed.
///
/// Synchronizes the eval stream before reading — assumes caller has
/// already issued the eval kernels and just needs to wait for them
/// to complete. Returns total trade count + the flat `Vec`.
pub fn read_per_trade_tape(&self) -> Result<Vec<EvalTrade>, MLError> {
// 1. Read per-window counts (cheap — n_windows u32 values).
let mut counts = vec![0u32; self.n_windows];
self.stream
.memcpy_dtoh(&self.per_trade_count_buf, &mut counts)
.map_err(|e| MLError::ModelError(format!("read per_trade_count: {e}")))?;
// 2. Compute total + early-return if no trades. Saves a 16MB+
// DtoH copy when the policy was inactive (rare but possible at
// cold-start before warmup).
let total: usize = counts.iter().map(|&c| c as usize).sum();
if total == 0 {
return Ok(Vec::new());
}
// 3. Read the four SoA buffers in full (n_windows * MAX_TRADES_PER_WINDOW
// each). Slicing window-side would need an extra index-list copy
// per window; reading whole buffers is simpler and stays cheap
// (16MB at PCIe ≈ 1ms — well below per-epoch overhead). The host
// loop below filters down to `count[w]` per window.
let buf_len = self.n_windows * MAX_TRADES_PER_WINDOW;
let mut pnl_host = vec![0.0_f32; buf_len];
let mut holding_host = vec![0u32; buf_len];
let mut bar_idx_host = vec![0u32; buf_len];
let mut dir_mag_host = vec![0u32; buf_len];
self.stream.memcpy_dtoh(&self.per_trade_pnl_buf, &mut pnl_host)
.map_err(|e| MLError::ModelError(format!("read per_trade_pnl: {e}")))?;
self.stream.memcpy_dtoh(&self.per_trade_holding_bars_buf, &mut holding_host)
.map_err(|e| MLError::ModelError(format!("read per_trade_holding_bars: {e}")))?;
self.stream.memcpy_dtoh(&self.per_trade_bar_index_buf, &mut bar_idx_host)
.map_err(|e| MLError::ModelError(format!("read per_trade_bar_index: {e}")))?;
self.stream.memcpy_dtoh(&self.per_trade_dir_mag_buf, &mut dir_mag_host)
.map_err(|e| MLError::ModelError(format!("read per_trade_dir_mag: {e}")))?;
// 4. Flatten window-major: for each window, take the first
// count[w] trades and append.
let mut tape = Vec::with_capacity(total);
for w in 0..self.n_windows {
let count = counts[w] as usize;
// Defensive cap: if a future kernel change skips the
// overflow guard, count COULD exceed MAX_TRADES_PER_WINDOW.
// The min() avoids out-of-bounds slicing.
let count = count.min(MAX_TRADES_PER_WINDOW);
let base = w * MAX_TRADES_PER_WINDOW;
for i in 0..count {
let idx = base + i;
let dm = dir_mag_host[idx];
let direction = ((dm >> 16) & 0xFFFF) as u8;
let magnitude = (dm & 0xFFFF) as u8;
tape.push(EvalTrade {
bar_index: bar_idx_host[idx],
pnl: pnl_host[idx],
holding_bars: holding_host[idx],
direction,
magnitude,
});
}
}
Ok(tape)
}
/// Submit the cuBLAS-based DQN step loop to the stream (chunked).
///
/// Processes `DQN_BACKTEST_CHUNK_SIZE` steps per cuBLAS forward call to
@@ -2244,17 +2420,17 @@ impl GpuBacktestEvaluator {
.arg(ch_conviction) // conviction_ptr: direction-branch Kelly warmup signal
.arg(ch_magnitude_conviction) // magnitude_conviction_ptr: magnitude-branch Kelly warmup signal
// SP21 T2.2 Phase 1 (2026-05-10): per-trade tape outputs.
// ABI extension at the kernel; production buffers wired
// in Phase 1 Step B (follow-up commit). NULL preserves
// bit-identical pre-Phase-1 behavior. Same NULL-tolerance
// pattern as the alpha_per_env / is_win_per_env additive
// contracts in `sp20_aggregate_inputs_kernel`.
.arg(&0u64) // per_trade_pnl_out (NULL)
.arg(&0u64) // per_trade_holding_bars_out (NULL)
.arg(&0u64) // per_trade_bar_index_out (NULL)
.arg(&0u64) // per_trade_dir_mag_out (NULL)
.arg(&0u64) // per_trade_count (NULL → emission no-op)
.arg(&0i32) // max_trades_per_window (unused when count=NULL)
// Real device buffers — same tape semantics as the
// single-step variant; the batched step loop writes
// per-step closes into the same per-window SoA slots
// (counter persists across the batched chunk via the
// shared `per_trade_count_buf`).
.arg(&self.per_trade_pnl_buf)
.arg(&self.per_trade_holding_bars_buf)
.arg(&self.per_trade_bar_index_buf)
.arg(&self.per_trade_dir_mag_buf)
.arg(&self.per_trade_count_buf)
.arg(&(MAX_TRADES_PER_WINDOW as i32))
.launch(LaunchConfig {
grid_dim: (env_blocks.max(1), 1, 1),
block_dim: (256, 1, 1),
@@ -2836,19 +3012,18 @@ impl GpuBacktestEvaluator {
.arg(&0u64) // conviction_ptr (NULL) — fallback 1.0 in-kernel
.arg(&0u64) // magnitude_conviction_ptr (NULL) — fallback 0.0 in-kernel; max(dir, mag) → dir alone
// SP21 T2.2 Phase 1 (2026-05-10): per-trade tape outputs.
// ABI extension at the kernel; production buffers wired in
// Phase 1 Step B (follow-up commit) — passing NULL preserves
// bit-identical behavior in this commit. Per
// `feedback_no_partial_refactor`: the kernel signature
// change + launcher migration land atomically here; the
// host-allocated buffers + readback + enrichment consumer
// wire-up follow incrementally as additive contracts.
.arg(&0u64) // per_trade_pnl_out (NULL)
.arg(&0u64) // per_trade_holding_bars_out (NULL)
.arg(&0u64) // per_trade_bar_index_out (NULL)
.arg(&0u64) // per_trade_dir_mag_out (NULL)
.arg(&0u64) // per_trade_count (NULL → emission no-op)
.arg(&0i32) // max_trades_per_window (unused when count=NULL)
// Real device buffers wired — kernel emits per-trade
// records at every trade close into the SoA buffers, with
// race-free per-window counter increment. Buffers are
// reset to zero at the start of each eval window via
// `reset_per_trade_tape()`. Read host-side via
// `read_per_trade_tape()` after the eval window completes.
.arg(&self.per_trade_pnl_buf)
.arg(&self.per_trade_holding_bars_buf)
.arg(&self.per_trade_bar_index_buf)
.arg(&self.per_trade_dir_mag_buf)
.arg(&self.per_trade_count_buf)
.arg(&(MAX_TRADES_PER_WINDOW as i32))
.launch(env_cfg)
.map_err(|e| {
MLError::ModelError(format!("backtest_env_step launch step {step}: {e}"))

View File

@@ -2,6 +2,103 @@
**Status:** Populated during Plan 1 Task 6 (A.5 orphan audit). Updated on every commit per Invariant 7.
## 2026-05-10 — SP21 T2.2 Phase 1 Step B: per-trade tape buffers + readback (atomic)
**Branch:** `sp20-aux-h-fixed` (off `5d01190e1`).
**Commits:** 1 atomic (this commit). **Step B replaces Step A's NULL
launcher pass with real device buffers, adds host-side readback
infrastructure, and resets the counter at fold start.**
Drops the NULL launcher arguments introduced in Step A. Both kernel
launch sites in `gpu_backtest_evaluator.rs` (`backtest_env_step` at
~:2786 and `backtest_env_step_batch` at ~:2208) now pass real device
pointers. The kernel's per-trade emission block fires unconditionally
on close events — single-threaded per-window writes preserve event
ordering and enable race-free counter increment without atomicAdd.
**New constants + types:**
- `MAX_TRADES_PER_WINDOW = 200_000` — per-window trade-tape capacity.
Set to typical eval window bar count (xmd6b: 214k bars, 139k trades)
so no trade is dropped under any trading frequency. Memory cost
per window: 4 SoA buffers × 4 bytes × 200k = 3.2MB, plus a
count[n_windows] u32. For typical 5-window walk-forward setups:
~16MB total (acceptable for eval — fold-boundary alloc, not per-step).
- `pub struct EvalTrade` (`crates/ml/src/cuda_pipeline/gpu_backtest_evaluator.rs`)
with 5 fields: `bar_index: u32`, `pnl: f32`, `holding_bars: u32`,
`direction: u8`, `magnitude: u8`. **Does NOT include `predicted_q`
or `ensemble_var`** — those require entry-time captures
(`entry_q`, `entry_var` tracked across the trade lifecycle in
portfolio state). Phase 1.5 follow-up will add them.
**New struct fields on `GpuBacktestEvaluator`:**
- `per_trade_pnl_buf: CudaSlice<f32>``[n_windows × MAX_TRADES_PER_WINDOW]`
- `per_trade_holding_bars_buf: CudaSlice<u32>` — same shape
- `per_trade_bar_index_buf: CudaSlice<u32>` — same shape
- `per_trade_dir_mag_buf: CudaSlice<u32>` — same shape (packed)
- `per_trade_count_buf: CudaSlice<u32>``[n_windows]`
**New methods:**
- `reset_per_trade_tape()` (folded into `reset_evaluation_state`):
zeros `per_trade_count_buf` at the start of each eval window. The
SoA value buffers don't need zeroing — read up to `count[w]` only,
stale slots past `count[w]` never visited.
- `pub fn read_per_trade_tape(&self) -> Result<Vec<EvalTrade>, MLError>`:
reads `per_trade_count_buf` (cheap, n_windows u32), early-returns
`Vec::new()` if total = 0, else reads the full 4 SoA buffers
(~16MB DtoH at PCIe ≈ 1ms — well below per-epoch overhead) and
flattens window-major into a chronological `Vec<EvalTrade>` for
the enrichment phase consumer (Phase 2 follow-up).
**Phase 2 follow-up (next commit):**
- Wire `read_per_trade_tape()` to the enrichment caller in
`training_loop.rs:1510-1568`, replacing
`extract_eval_trades_from_metrics` (the fake-trade synthesizer).
- Update `enrichment::EvalTrade` to match the GPU evaluator's
5-field struct, removing the `predicted_q` and `ensemble_var`
fields that have no producer yet.
- E1 (q_correction) and E5 (agreement_threshold) — refactor to
consume signals other than per-trade `predicted_q`/`ensemble_var`
OR DELETE per `feedback_v7_gem_methodology` (measure → wire or
delete; outputs already unused in production code path).
**Phase 1.5 follow-up (if Phase 2 keeps E1/E5):**
- Add `entry_q` and `entry_var` to `portfolio_state`, captured at
trade open. Extend per-trade tape with corresponding 6th and 7th
SoA buffers. Update `EvalTrade` struct + `read_per_trade_tape`.
**Affected files:**
- `crates/ml/src/cuda_pipeline/gpu_backtest_evaluator.rs` — new
`MAX_TRADES_PER_WINDOW` const, `EvalTrade` struct, 5 buffer
fields, alloc in constructor, struct-construction line, replace
NULL with real pointers in both launchers, add reset to
`reset_evaluation_state`, new `read_per_trade_tape` method.
**Verification:**
- `cargo check -p ml --tests` — passes (warnings only).
- GPU oracle tests: behavior preserved by construction. The kernel's
per-trade emission block runs unconditionally on real buffers, but
the existing aggregator/EMA tests don't check per-trade buffers
(they check WindowMetrics aggregates, which are computed in a
separate kernel from `record_kelly_trade_outcome`'s win/loss
accumulators — still correct).
**Open question for Phase 2 design**: with `predicted_q` and
`ensemble_var` deferred to Phase 1.5, E1 and E5 consume signals we
don't yet emit. Three options for Phase 2: (a) delete E1+E5
(unused per `feedback_v7_gem_methodology` 3-step audit — E1 output
discarded, E5 output consumed but no useful computation possible);
(b) refactor E1+E5 to use alternate signals (avg_predicted_q from
WindowMetrics for E1, ensemble_disagreement EMA from ISV for E5);
(c) accept temporary no-op behavior in Phase 2 with explicit
Phase 1.5 follow-up scheduled.
## 2026-05-10 — SP21 T2.2 Phase 1 Step A: per-trade tape kernel ABI extension (atomic)
**Branch:** `sp20-aux-h-fixed` (off `4ab1c132e`).