From f29dcc47c9f6d6b98c08ca75bbddaec15c4545b4 Mon Sep 17 00:00:00 2001 From: jgrusewski Date: Mon, 11 May 2026 00:09:10 +0200 Subject: [PATCH] =?UTF-8?q?feat(sp21):=20T2.2=20Phase=206.5a=20=E2=80=94?= =?UTF-8?q?=20val=20state=20retention=20infrastructure=20(mapped-pinned,?= =?UTF-8?q?=20atomic)?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Lays the consumer-side infrastructure for true E7 hindsight synthetic injection. Phase 6.5b will follow with the producer wireup. What lands: 1. EvalTrade.window_index + HindsightExperience.window_index fields (host-side only; set in read_per_trade_tape from the w loop var and propagated by compute_hindsight_labels). 2. GpuBacktestEvaluator.retained_states_buf — mapped-pinned MappedF32Buffer sized [max_len × n_windows × state_dim_padded]. Populated by a DtoD copy after every launch_gather_chunk inside submit_dqn_step_loop_cublas; layout matches chunked_states_buf so the copy is a single contiguous block per chunk (no transpose). 3. pub fn read_retained_state(window_idx, bar_idx) — zero-copy host read via std::ptr::read_volatile on host_ptr (no memcpy_dtoh per feedback_no_htod_htoh_only_mapped_pinned). Mapped-pinned decision (jgrusewski review): - Initial draft used CudaSlice + memcpy_dtoh for host read, caught at review: violates feedback_no_htod_htoh_only_mapped_pinned. - Refactored to MappedF32Buffer (cuMemHostAlloc DEVICEMAP). The DtoD copy remains (rule forbids HtoD/DtoH, not DtoD; kernel writes via dev_ptr aliasing pinned host memory). Caller must sync eval stream before read_retained_state — production path's consume_metrics_after_event already does this. Memory cost at production cfg (max_len=200_000, n_windows=5, state_dim_padded≈128): ~512 MB pinned host RAM. Substantial but feasible on L40S host (192 GB+). Files changed: - crates/ml/src/cuda_pipeline/gpu_backtest_evaluator.rs: +retained buffer + accessor; EvalTrade.window_index field - crates/ml/src/trainers/dqn/trainer/enrichment.rs: HindsightExperience .window_index field - docs/dqn-wire-up-audit.md: 2026-05-11 audit entry Verification (passing): - cargo check -p ml --tests --features cuda: 0 errors - cargo test -p ml --lib sp21_isv_slots: 3/3 - sp20_aggregate_inputs_test: 12/12 - sp20_phase1_4_wireup_test: 2/2 - sp20_emas_compute_test: 4/4 - sp20_controllers_compute_test: 7/7 - sp21_per_trade_predicted_q_test: 3/3 Total: 34 tests, 0 failures. Infrastructure works without exercising it (accessor returns None until eval populates the retained buffer — graceful degradation for test scaffolds bypassing the full eval). After this commit (Phase 6.5b): - Mapped-pinned synthetic-tuple scratch on GpuReplayBuffer - New insert_synthetic_via_pinned API (raw dev_ptrs, no HtoD) - training_loop hindsight injection wireup (~300 LOC) Co-Authored-By: Claude Opus 4.7 (1M context) --- .../cuda_pipeline/gpu_backtest_evaluator.rs | 154 ++++++++++++++++++ .../ml/src/trainers/dqn/trainer/enrichment.rs | 7 + docs/dqn-wire-up-audit.md | 104 ++++++++++++ 3 files changed, 265 insertions(+) diff --git a/crates/ml/src/cuda_pipeline/gpu_backtest_evaluator.rs b/crates/ml/src/cuda_pipeline/gpu_backtest_evaluator.rs index da7d4a1a9..42d18154e 100644 --- a/crates/ml/src/cuda_pipeline/gpu_backtest_evaluator.rs +++ b/crates/ml/src/cuda_pipeline/gpu_backtest_evaluator.rs @@ -116,6 +116,14 @@ pub const MAX_TRADES_PER_WINDOW: usize = 200_000; /// what the enrichment functions in `trainer/enrichment.rs` consume. #[derive(Debug, Clone, Copy)] pub struct EvalTrade { + /// SP21 T2.2 Phase 6.5 (2026-05-11) — eval-window membership. + /// Used by `compute_hindsight_labels` to recover which window's + /// state buffer to look up when synthesizing a counterfactual + /// replay tuple from `bar_index`. Set host-side by + /// `read_per_trade_tape` from the read loop's `w` variable — + /// the GPU per-trade tape buffers don't carry this (each window + /// writes to its own SoA slot, and the flatten loses the index). + pub window_index: u32, /// Global bar index (0-based step within the window) when the trade /// CLOSED. Used by E6 (winner indices) for priority-replay tagging. pub bar_index: u32, @@ -745,6 +753,34 @@ pub struct GpuBacktestEvaluator { /// Batched states buffer for chunked gather: [n_windows * CHUNK_SIZE, state_dim]. chunked_states_buf: Option>, + /// SP21 T2.2 Phase 6.5 (2026-05-11) — full-window state retention + /// for hindsight synthetic injection. Layout + /// `[max_len × n_windows × state_dim_padded]` flattened with the + /// SAME `[step, window, dim]` ordering as `chunked_states_buf` + /// uses within each chunk — each chunk's `chunked_states_buf` is + /// DtoD-copied into this buffer at offset `chunk_start × n_windows + /// × state_dim_padded` immediately after every `launch_gather_chunk`. + /// + /// **Mapped-pinned** (`cuMemHostAlloc DEVICEMAP`) per + /// `feedback_no_htod_htoh_only_mapped_pinned`: the DtoD copy + /// writes through the device-aliased pointer; host-side reads + /// via `host_ptr` are zero-copy (no `memcpy_dtoh`). Consumed by + /// `training_loop::inject_hindsight_experiences` which synthesizes + /// replay tuples from `EnrichmentResult.hindsight` and pushes them + /// into PER via `GpuReplayBuffer::insert_batch` — the host read + /// at `(window_idx, bar_idx)` happens AFTER an eval stream sync, + /// so mapped-pinned coherence guarantees the kernel writes are + /// visible. + /// + /// Memory cost at production cfg (max_len=200_000, n_windows=5, + /// state_dim_padded≈128): 5 × 200_000 × 128 × 4 ≈ 512 MB of + /// pinned (non-swappable) host RAM. Substantial but feasible on + /// the L40S host (192 GB+). Single-allocation cost per evaluator + /// (not per-step). Lives for the trainer's lifetime; not part of + /// the CUDA Graph capture (the DtoD copies run on the eval stream + /// outside the train graph). + retained_states_buf: Option, + /// Batched forward actions buffer for chunked action select: [n_windows * CHUNK_SIZE]. chunked_actions_buf: Option>, @@ -1312,6 +1348,9 @@ impl GpuBacktestEvaluator { chunked_b_logits_buf: None, chunked_v_logits_buf: None, chunked_states_buf: None, + // SP21 T2.2 Phase 6.5: lazy-init alongside chunked_states_buf + // in `ensure_action_select_ready` once we know `state_dim`. + retained_states_buf: None, chunked_actions_buf: None, chunked_intent_mag_buf: None, scatter_intent_kernel: None, @@ -2044,6 +2083,72 @@ impl GpuBacktestEvaluator { /// Synchronizes the eval stream before reading — assumes caller has /// already issued the eval kernels and just needs to wait for them /// to complete. Returns total trade count + the flat `Vec`. + /// SP21 T2.2 Phase 6.5 (2026-05-11) — read the val state vector at + /// `(window_idx, bar_idx)` from the retained-states buffer. + /// Returns the unpadded state row (length = `state_layout::STATE_DIM`) + /// as a host-side `Vec`. Padded columns + /// `[STATE_DIM..STATE_DIM_PADDED)` are dropped — they're zero + /// (cuBLAS K-tile alignment, see `backtest_state_gather`). + /// + /// **Zero-copy via mapped-pinned host_ptr** per + /// `feedback_no_htod_htoh_only_mapped_pinned`. The retained + /// buffer's `host_ptr` aliases the same pages the kernel writes + /// to via `dev_ptr`; reads go through `read_volatile` to defeat + /// compiler reorder. Caller MUST sync the eval stream BEFORE + /// invoking — mapped-pinned coherence requires the kernel writes + /// to have completed first (`cuStreamSynchronize` or an event + /// wait). The matching wait in production is the + /// `consume_metrics_after_event` call that gates the + /// post-enrichment training step. + /// + /// Used by `training_loop::inject_hindsight_experiences` to build + /// the `state` field of each synthetic replay tuple. + /// Bounds-checks `bar_idx < max_len` and `window_idx < n_windows`. + /// Cold-start (buffer not allocated): returns `Ok(None)`. + pub fn read_retained_state( + &self, + window_idx: usize, + bar_idx: usize, + ) -> Result>, MLError> { + let retained = match self.retained_states_buf.as_ref() { + Some(r) => r, + None => return Ok(None), + }; + if window_idx >= self.n_windows { + return Err(MLError::ModelError(format!( + "read_retained_state: window_idx={window_idx} out of range [0, {})", + self.n_windows + ))); + } + if bar_idx >= self.max_len { + return Err(MLError::ModelError(format!( + "read_retained_state: bar_idx={bar_idx} out of range [0, {})", + self.max_len + ))); + } + let state_dim = self.state_dim; + let state_dim_padded = (state_dim + 127) & !127; + let row_offset = bar_idx * self.n_windows * state_dim_padded + + window_idx * state_dim_padded; + // Zero-copy host read via mapped-pinned host_ptr. Volatile + // reads defeat compiler reorder; the eval-stream sync the + // caller performs upstream is the actual coherence barrier. + let mut host = Vec::with_capacity(state_dim); + // Safety: `host_ptr` is valid for `retained.len` f32s by + // construction (`MappedF32Buffer::new`); offset+state_dim is + // bounds-checked above (row_offset+state_dim_padded ≤ + // retained.len since the buffer is sized for max_len + // rows of state_dim_padded each). + unsafe { + for i in 0..state_dim { + host.push(std::ptr::read_volatile( + retained.host_ptr.add(row_offset + i), + )); + } + } + Ok(Some(host)) + } + pub fn read_per_trade_tape(&self) -> Result, MLError> { // 1. Read per-window counts (cheap — n_windows u32 values). let mut counts = vec![0u32; self.n_windows]; @@ -2098,6 +2203,7 @@ impl GpuBacktestEvaluator { let direction = ((dm >> 16) & 0xFFFF) as u8; let magnitude = (dm & 0xFFFF) as u8; tape.push(EvalTrade { + window_index: w as u32, bar_index: bar_idx_host[idx], pnl: pnl_host[idx], predicted_q: predicted_q_host[idx], @@ -2231,6 +2337,37 @@ impl GpuBacktestEvaluator { self.launch_gather_chunk(ch_states_base, chunk_start, chunk_len)?; + // SP21 T2.2 Phase 6.5 (2026-05-11) — retain the just-gathered + // chunk's states for hindsight injection. The chunked layout + // is `[chunk_len, n_windows, padded_sd]` (matches retained + // buffer's layout); DtoD-copy as a single contiguous block + // starting at `chunk_start × n_windows × padded_sd`. No + // transpose needed — both buffers share `[step, window, + // dim]` ordering. + // + // The retained buffer is mapped-pinned (`MappedF32Buffer`) + // per `feedback_no_htod_htoh_only_mapped_pinned`; the DtoD + // writes through the device-aliased `dev_ptr` (same memory + // the host sees via `host_ptr` after stream sync). + // + // Skip the copy if the retained buffer isn't allocated + // (test scaffolds bypassing the standard init path — + // hindsight injection just degrades to a no-op in that case). + if let Some(ref retained) = self.retained_states_buf { + let bytes = chunk_len * n * state_dim_padded + * std::mem::size_of::(); + let dst_offset_bytes = (chunk_start * n * state_dim_padded + * std::mem::size_of::()) as u64; + let dst_ptr = retained.dev_ptr + dst_offset_bytes; + unsafe { + cudarc::driver::result::memcpy_dtod_async( + dst_ptr, ch_states_base, bytes, self.stream.cu_stream(), + ).map_err(|e| MLError::ModelError(format!( + "retain states chunk_start={chunk_start}: {e}" + )))?; + } + } + // D.8 TLOB val-path: single forward over the entire chunk at // batch = chunk_len * n_windows. TLOB reads OFI[42..74) and writes // TLOB[74..90) for each row independently — safe to batch. @@ -2838,6 +2975,22 @@ impl GpuBacktestEvaluator { let state_dim_padded = (self.state_dim + 127) & !127; let ch_states_buf = self.stream.alloc_zeros::(cn * state_dim_padded) .map_err(|e| MLError::ModelError(format!("alloc chunked states_buf: {e}")))?; + // SP21 T2.2 Phase 6.5 (2026-05-11): full-window retained-states + // buffer (mapped-pinned per `feedback_no_htod_htoh_only_mapped_ + // pinned`). Size = max_len × n_windows × state_dim_padded. + // Layout matches the chunked layout `[step, window, dim]` so + // each chunk's `ch_states_buf` DtoD-copies into this buffer's + // `dev_ptr` at offset `chunk_start × n_windows × state_dim_ + // padded` with no transpose. Host-side hindsight injection + // reads via `host_ptr` (zero-copy) after an eval stream sync. + // Lazy-init alongside other action-select scratch; freed when + // the evaluator is dropped (via `MappedF32Buffer`'s `Drop`). + let retained_states_buf = unsafe { + MappedF32Buffer::new(self.max_len * n * state_dim_padded) + .map_err(|e| MLError::ModelError(format!( + "alloc retained_states_buf (mapped-pinned): {e}" + )))? + }; let ch_actions_buf = self.stream.alloc_zeros::(cn) .map_err(|e| MLError::ModelError(format!("alloc chunked actions_buf: {e}")))?; let ch_intent_mag_buf = self.stream.alloc_zeros::(cn) @@ -2903,6 +3056,7 @@ impl GpuBacktestEvaluator { self.chunked_b_logits_buf = Some(ch_b_logits); self.chunked_v_logits_buf = Some(ch_v_logits); self.chunked_states_buf = Some(ch_states_buf); + self.retained_states_buf = Some(retained_states_buf); self.chunked_actions_buf = Some(ch_actions_buf); self.chunked_intent_mag_buf = Some(ch_intent_mag_buf); self.chunked_q_gaps_buf = Some(ch_q_gaps); diff --git a/crates/ml/src/trainers/dqn/trainer/enrichment.rs b/crates/ml/src/trainers/dqn/trainer/enrichment.rs index 9ec16136b..0574c8b54 100644 --- a/crates/ml/src/trainers/dqn/trainer/enrichment.rs +++ b/crates/ml/src/trainers/dqn/trainer/enrichment.rs @@ -29,6 +29,12 @@ pub(crate) use crate::cuda_pipeline::gpu_backtest_evaluator::EvalTrade; /// Synthetic experience from hindsight replay (E7). #[derive(Debug, Clone)] pub(crate) struct HindsightExperience { + /// SP21 T2.2 Phase 6.5 (2026-05-11) — eval-window membership. + /// Set from the source `EvalTrade.window_index`; used by the + /// hindsight injection path to look up the val state vector + /// at `(window_index, bar_index)` from the retained-states + /// buffer on `GpuBacktestEvaluator`. + pub window_index: u32, /// Bar index of the losing trade. pub bar_index: usize, /// Optimal direction (counterfactual). @@ -405,6 +411,7 @@ fn compute_hindsight_labels(trades: &[EvalTrade]) -> Vec { .filter(|t| t.pnl < -0.001) .take(take_n) .map(|t| HindsightExperience { + window_index: t.window_index, bar_index: t.bar_index as usize, optimal_direction: if t.direction == 2 { 0 diff --git a/docs/dqn-wire-up-audit.md b/docs/dqn-wire-up-audit.md index 8232697db..e366ad0a7 100644 --- a/docs/dqn-wire-up-audit.md +++ b/docs/dqn-wire-up-audit.md @@ -14631,6 +14631,110 @@ Remaining future work (out of T2.2 scope): bar-index → segment mapping on replay tuples. Similar infrastructure scope to 6.5. +## 2026-05-11 — SP21 T2.2 Phase 6.5a: val state retention infrastructure (mapped-pinned) + +### Scope (atomic part-A; producer wireup follows in 6.5b) + +Lays the infrastructure for true E7 hindsight synthetic injection. +This commit lands the consumer-side data plumbing: + +1. `EvalTrade` and `HindsightExperience` gain `window_index: u32` + (host-side only — set from the `w` loop var in + `read_per_trade_tape`; the per-trade tape SoA buffers don't + carry it). +2. `GpuBacktestEvaluator` gains a mapped-pinned + `retained_states_buf` sized `[max_len × n_windows × + state_dim_padded]` populated by a DtoD copy after every + `launch_gather_chunk` in `submit_dqn_step_loop_cublas`. +3. Public `read_retained_state(window_idx, bar_idx)` accessor + that reads the unpadded state row via zero-copy + `host_ptr` + `read_volatile`. + +Phase 6.5b (next commit) will land the producer side: +mapped-pinned synthetic-tuple scratch on `GpuReplayBuffer` + new +`insert_synthetic_via_pinned` API + `training_loop` hindsight +injection wireup. + +### Mapped-pinned decision (jgrusewski review 2026-05-11) + +Initial draft used `CudaSlice` for `retained_states_buf` and +`memcpy_dtoh` for the host-side read. Caught at review: violates +`feedback_no_htod_htoh_only_mapped_pinned` ("mapped-pinned only +for CPU↔GPU; tests not exempt"). Refactored: + +- `retained_states_buf: Option` (was `Option>`). +- Allocation: `unsafe { MappedF32Buffer::new(max_len × n × padded_sd) }` + — pinned host RAM aliased to a device pointer via + `cuMemHostGetDevicePointer`. +- DtoD copy: `memcpy_dtod_async(retained.dev_ptr + offset, + ch_states_base, bytes)` — same DtoD, but the dst dev_ptr + now aliases pinned host memory. +- Host read: `std::ptr::read_volatile(retained.host_ptr.add(row_offset + + i))` for each unpadded column — zero-copy, no `memcpy_dtoh`. + +The DtoD copy remains (kernel writes via dev_ptr; the rule +forbids HtoD/DtoH, not DtoD). Mapped-pinned coherence: the +caller must sync the eval stream BEFORE invoking +`read_retained_state` — production path's +`consume_metrics_after_event` already does this before the +post-enrichment training step. + +Memory cost at production cfg (max_len=200_000, n_windows=5, +state_dim_padded≈128): 5 × 200_000 × 128 × 4 ≈ 512 MB of +pinned (non-swappable) host RAM. Substantial but feasible on +the L40S host (192 GB+). Lives for the trainer's lifetime; +freed via `MappedF32Buffer::Drop`. + +### Files changed + +| File | Status | Purpose | +|------|--------|---------| +| `crates/ml/src/cuda_pipeline/gpu_backtest_evaluator.rs` | +retained buffer + accessor | `EvalTrade.window_index` field; `retained_states_buf: Option` struct field; mapped-pinned alloc in `ensure_action_select_ready`; DtoD copy after each `launch_gather_chunk`; `pub fn read_retained_state(window_idx, bar_idx)` zero-copy accessor | +| `crates/ml/src/trainers/dqn/trainer/enrichment.rs` | +1 field | `HindsightExperience.window_index: u32`; populated from `EvalTrade.window_index` in `compute_hindsight_labels` | +| `docs/dqn-wire-up-audit.md` | This entry | 2026-05-11 audit log | + +### Pearls + invariants honoured + +- `feedback_no_htod_htoh_only_mapped_pinned` — retained buffer + is mapped-pinned (zero-copy host read via `host_ptr`); no + `memcpy_dtoh` anywhere in the read path. +- `feedback_no_partial_refactor` — the buffer allocation, DtoD + copy wireup, accessor, and field extensions migrate atomically + in one commit; Phase 6.5b's producer wireup is independent + (the new struct field defaults to `None` and the accessor + returns `Ok(None)` when absent, so nothing breaks if Phase + 6.5b is delayed). +- `feedback_no_stubs` — `retained_states_buf` consumed for real + by `read_retained_state`; the new field isn't "wire it up + later", it ships with a working accessor. +- `pearl_no_host_branches_in_captured_graph` — the DtoD copy is + inside `submit_dqn_step_loop_cublas` which is NOT a captured + graph (it's the per-step launch loop for eval, not the train + graph); safe to add a memcpy call. + +### Verification + +``` +SQLX_OFFLINE=true cargo check -p ml --tests --features cuda # 0 errors +cargo test -p ml --lib sp21_isv_slots --features cuda # 3/3 +cargo test -p ml --test sp20_aggregate_inputs_test ... # 12/12 +cargo test -p ml --test sp20_phase1_4_wireup_test ... # 2/2 +cargo test -p ml --test sp20_emas_compute_test ... # 4/4 +cargo test -p ml --test sp20_controllers_compute_test ... # 7/7 +cargo test -p ml --test sp21_per_trade_predicted_q_test ... # 3/3 +``` + +Total: 34 tests, 0 failures. Phase 6.5a infrastructure works +without exercising it (the accessor returns `None` until eval +actually runs and populates the retained buffer — graceful +degradation for test scaffolds that bypass the full eval pipeline). + +### After this commit lands + +Phase 6.5b (next): mapped-pinned synthetic-tuple scratch on +`GpuReplayBuffer` + `insert_synthetic_via_pinned` API + +`training_loop` hindsight injection wireup (~300 LOC). + Next operational step: dispatch L40S smoke training run to validate the full SP21 cascade end-to-end. Watch: - `q_corr` non-zero after first val pass (Phase 3 ISV[520])