Files
foxhunt/docs/superpowers/plans/2026-05-26-gpu-hindsight-replay.md
2026-05-26 00:49:42 +02:00

245 lines
9.6 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# GPU-Native Hindsight Experience Replay Implementation Plan
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
**Goal:** Bidirectional HER (backward peak + forward continuation) that injects synthetic transitions into the GPU PER, accelerating wave-exit timing learning 10×.
**Architecture:** Per-batch mid-price ring buffer tracks price during trades. On close, backward kernel computes peak and injects synthetic if peak >> actual. Every step, forward kernel evaluates recently-closed trades and injects if holding would have been better. Both push to GpuReplayBuffer via the same prefix-sum coordination as rl_per_push.
**Tech Stack:** Rust 1.85, cudarc 0.19, CUDA 12.4, pre-compiled cubins
---
## File Map
| File | Responsibility |
|------|---------------|
| `crates/ml-alpha/src/rl/gpu_hindsight.rs` | NEW: GpuHindsight struct + allocations |
| `crates/ml-alpha/src/rl/mod.rs` | Add `pub mod gpu_hindsight;` |
| `crates/ml-alpha/cuda/rl_hindsight_track.cu` | NEW: per-step mid accumulation + peak tracking |
| `crates/ml-alpha/cuda/rl_hindsight_inject.cu` | NEW: backward HER — synthetic push on done |
| `crates/ml-alpha/cuda/rl_hindsight_forward.cu` | NEW: forward HER — evaluate post-exit continuation |
| `crates/ml-alpha/build.rs` | Register 3 new kernels |
| `crates/ml-alpha/src/trainer/integrated.rs` | Wire GpuHindsight + 3 kernel launches into step pipeline |
| `crates/ml-alpha/src/rl/isv_slots.rs` | Add 3 ISV slots (threshold, priority_boost, inject_count) |
| `crates/ml-alpha/examples/alpha_rl_train.rs` | Add hindsight diag to JSONL |
| `crates/ml-alpha/tests/gpu_hindsight_oracle.rs` | NEW: oracle tests |
---
## Task 1: GpuHindsight struct + ISV slots
**Files:**
- Create: `crates/ml-alpha/src/rl/gpu_hindsight.rs`
- Modify: `crates/ml-alpha/src/rl/mod.rs`
- Modify: `crates/ml-alpha/src/rl/isv_slots.rs`
- [ ] **Step 1: Add ISV slots to `isv_slots.rs`**
```rust
pub const RL_HINDSIGHT_THRESHOLD_INDEX: usize = 549;
pub const RL_HINDSIGHT_PRIORITY_BOOST_INDEX: usize = 550;
pub const RL_HINDSIGHT_INJECT_COUNT_INDEX: usize = 551;
pub const RL_HINDSIGHT_FORWARD_LOOKAHEAD_INDEX: usize = 552;
```
- [ ] **Step 2: Create `gpu_hindsight.rs`**
```rust
use std::sync::Arc;
use anyhow::{Context, Result};
use cudarc::driver::{CudaSlice, CudaStream};
use crate::heads::HIDDEN_DIM;
pub const HINDSIGHT_RING_LEN: usize = 512;
pub const CLOSED_RING_SIZE: usize = 256;
pub const CLOSED_RING_FIELDS: usize = 5; // entry_price, exit_step, direction, actual_pnl, batch_idx
pub struct GpuHindsight {
// Backward HER
pub mid_ring_d: CudaSlice<f32>, // [B × HINDSIGHT_RING_LEN]
pub ring_write_idx_d: CudaSlice<u32>, // [B]
pub peak_mid_d: CudaSlice<f32>, // [B]
pub entry_mid_d: CudaSlice<f32>, // [B]
pub position_dir_d: CudaSlice<i32>, // [B] — +1 long, -1 short, 0 flat
// Forward HER
pub closed_ring_d: CudaSlice<f32>, // [CLOSED_RING_SIZE × CLOSED_RING_FIELDS]
pub closed_h_t_d: CudaSlice<f32>, // [CLOSED_RING_SIZE × HIDDEN_DIM]
pub closed_write_head_d: CudaSlice<u32>,// [1]
pub closed_count_d: CudaSlice<u32>, // [1]
pub closed_step_d: CudaSlice<u32>, // [CLOSED_RING_SIZE] — step when trade closed
}
impl GpuHindsight {
pub fn new(stream: &Arc<CudaStream>, b_size: usize) -> Result<Self> {
Ok(Self {
mid_ring_d: stream.alloc_zeros(b_size * HINDSIGHT_RING_LEN).context("her mid_ring")?,
ring_write_idx_d: stream.alloc_zeros(b_size).context("her ring_write_idx")?,
peak_mid_d: stream.alloc_zeros(b_size).context("her peak_mid")?,
entry_mid_d: stream.alloc_zeros(b_size).context("her entry_mid")?,
position_dir_d: stream.alloc_zeros(b_size).context("her position_dir")?,
closed_ring_d: stream.alloc_zeros(CLOSED_RING_SIZE * CLOSED_RING_FIELDS).context("her closed_ring")?,
closed_h_t_d: stream.alloc_zeros(CLOSED_RING_SIZE * HIDDEN_DIM).context("her closed_h_t")?,
closed_write_head_d: stream.alloc_zeros(1).context("her closed_write_head")?,
closed_count_d: stream.alloc_zeros(1).context("her closed_count")?,
closed_step_d: stream.alloc_zeros(CLOSED_RING_SIZE).context("her closed_step")?,
})
}
}
```
- [ ] **Step 3: Add `pub mod gpu_hindsight;` to `mod.rs`**
- [ ] **Step 4: Verify compilation**
- [ ] **Step 5: Commit**
---
## Task 2: Write 3 HER CUDA kernels
**Files:**
- Create: `crates/ml-alpha/cuda/rl_hindsight_track.cu`
- Create: `crates/ml-alpha/cuda/rl_hindsight_inject.cu`
- Create: `crates/ml-alpha/cuda/rl_hindsight_forward.cu`
- Modify: `crates/ml-alpha/build.rs`
- [ ] **Step 1: `rl_hindsight_track.cu`** — per-step mid accumulation
Grid=(ceil(B/32)), Block=(32). Per thread b:
1. Read position_lots from pos_d — if flat, reset ring_write_idx=0, return
2. mid = 0.5 × (bid_px[0] + ask_px[0])
3. If ring_write_idx==0: entry_mid[b] = mid, position_dir[b] = sign(lots)
4. Store mid in mid_ring[b * RING_LEN + min(idx, RING_LEN-1)]
5. Update peak: if dir>0: peak=fmaxf(peak, mid); else peak=fminf(peak, mid)
6. Increment ring_write_idx[b]
- [ ] **Step 2: `rl_hindsight_inject.cu`** — backward HER on trade close
Grid=(1), Block=(B). Shared memory for prefix-sum (same as rl_per_push).
Per thread b:
1. If dones[b] <= 0.5: no injection, skip
2. Compute peak_pnl = |peak_mid[b] - entry_mid[b]| (in price space)
3. actual_pnl = |rewards_raw[b]| / reward_scale (approximate USD)
4. threshold = isv[RL_HINDSIGHT_THRESHOLD_INDEX]
5. If peak_pnl <= actual_pnl × threshold: skip
6. Flag this thread for injection (shared mem prefix-sum coordination)
7. After prefix-sum: inject synthetic transition to replay at assigned write slot:
- h_t = perception.h_t_view()[b] (current state — "at states like this, optimal was peak_pnl")
- h_tp1 = same (self-loop for synthetic)
- action = same as taken
- reward = peak_pnl × reward_scale (back to scaled space)
- priority = isv[RL_HINDSIGHT_PRIORITY_BOOST_INDEX] × max_priority
8. Reset: ring_write_idx[b]=0, peak_mid[b]=entry_mid[b]
9. Thread 0: write inject count to isv[RL_HINDSIGHT_INJECT_COUNT_INDEX]
- [ ] **Step 3: `rl_hindsight_forward.cu`** — forward HER post-exit
Grid=(ceil(CLOSED_RING_SIZE/32)), Block=(32). Per thread i:
1. If closed_ring[i] empty (entry_price==0): return
2. age = current_step - closed_step[i]
3. If age < FORWARD_LOOKAHEAD: return (too early to evaluate)
4. If age > 2 × FORWARD_LOOKAHEAD: expire (clear slot, return)
5. entry_price = closed_ring[i*5 + 0], direction = closed_ring[i*5 + 2], actual_pnl = closed_ring[i*5 + 3]
6. current_mid = 0.5 × (bid_px[0] + ask_px[0])
7. forward_pnl = direction × (current_mid - entry_price)
8. If forward_pnl > actual_pnl × threshold: inject synthetic from closed_h_t[i]
9. Clear slot after evaluation
- [ ] **Step 4: Register in build.rs**
- [ ] **Step 5: Verify nvcc compilation + cargo check**
- [ ] **Step 6: Commit**
---
## Task 3: Wire HER into trainer step pipeline
**Files:**
- Modify: `crates/ml-alpha/src/trainer/integrated.rs`
- [ ] **Step 1: Add GpuHindsight field + kernel handles**
In the struct: `pub hindsight: GpuHindsight` + 3 kernel fn/module pairs.
- [ ] **Step 2: Init in constructor**
```rust
let hindsight = GpuHindsight::new(&stream, b_size)?;
```
Bootstrap ISV slots:
```rust
(RL_HINDSIGHT_THRESHOLD_INDEX, 1.5),
(RL_HINDSIGHT_PRIORITY_BOOST_INDEX, 3.0),
(RL_HINDSIGHT_FORWARD_LOOKAHEAD_INDEX, 50.0),
```
- [ ] **Step 3: Insert into step pipeline**
After the fused reward pipeline (which sets rewards/dones), before rl_per_push:
```rust
// Track mid price for backward HER
launch rl_hindsight_track(pos_d, bid_px, ask_px, &mut hindsight.*)
// After rl_per_push (real transitions pushed):
// Backward HER inject (synthetic transitions for trades where peak >> actual)
launch rl_hindsight_inject(dones, rewards_raw, h_t, isv, &mut gpu_replay.*, &mut hindsight.*)
// Forward HER inject (synthetic for trades that closed LOOKAHEAD steps ago)
launch rl_hindsight_forward(bid_px, ask_px, isv, &mut gpu_replay.*, &mut hindsight.*)
// Tree rebuild covers both real + synthetic pushes
launch rl_per_tree_rebuild
```
Also: on trade close (done=1), push entry to the forward closed_ring for later evaluation.
- [ ] **Step 4: Smoke test**
- [ ] **Step 5: Commit**
---
## Task 4: Diag JSONL integration
**Files:**
- Modify: `crates/ml-alpha/examples/alpha_rl_train.rs`
- [ ] **Step 1: Add hindsight diag block**
Read ISV[RL_HINDSIGHT_INJECT_COUNT_INDEX] and add to JSONL:
```json
"hindsight": {
"inject_count_step": isv[551],
"threshold": isv[549],
"priority_boost": isv[550]
}
```
- [ ] **Step 2: Commit**
---
## Task 5: GPU oracle tests
**Files:**
- Create: `crates/ml-alpha/tests/gpu_hindsight_oracle.rs`
- [ ] **Step 1: Write 3 tests**
1. `hindsight_track_accumulates_peak` — open position, feed rising mid prices, verify peak_mid tracks the max
2. `hindsight_inject_fires_on_done_with_peak` — set peak >> entry, fire done, verify replay got a synthetic transition
3. `hindsight_forward_injects_after_lookahead` — push a closed trade, advance step counter past lookahead, verify injection fires
- [ ] **Step 2: Run on local GPU**
- [ ] **Step 3: Commit**
---
## Kill Criteria
- Track kernel: peak_mid monotonically increases during a long trade with rising prices
- Backward inject: fires only when done=1 AND peak > threshold × actual
- Forward inject: fires only when age >= LOOKAHEAD AND forward_pnl > threshold × actual
- After 10k steps with HER: inject_count > 0 on done-steps (synthetic transitions appearing)
- No destabilization: l_q trajectory same or better than non-HER run