plan: GPU composite reward implementation — 7 tasks

Task 1: CUDA kernel (8-component formula, PORTFOLIO_STRIDE=12)
Task 2: Config pipeline (DQNHyperparameters + TOML profiles)
Task 3: Experience collector (buffer alloc + kernel launch)
Task 4: Hyperopt adapter (31D→38D search space)
Task 5: Remove dead CPU reward code
Task 6: Integration test (smoke + hyperopt + GPU suite)
Task 7: H100 validation

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
This commit is contained in:
jgrusewski
2026-03-22 20:35:57 +01:00
parent 75885bccaf
commit 86863d72e6

View File

@@ -0,0 +1,434 @@
# GPU Composite Reward Implementation Plan
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
**Goal:** Replace the raw PnL reward in the CUDA experience kernel with an 8-component composite reward (DSR, normalized PnL, drawdown penalty, idle penalty, regime scaling, asymmetric loss, position-time decay, transaction costs) — all GPU-native, zero CPU involvement.
**Architecture:** Extend `experience_env_step` kernel with composite reward computation. Portfolio state buffer grows from 3→12 floats per episode. 7 new reward weights flow through TOML config → `DQNHyperparameters``ExperienceCollectorConfig` → kernel args. Hyperopt search space grows from 31D→38D.
**Tech Stack:** CUDA C (experience kernel), Rust (config pipeline, hyperopt adapter), TOML (training profiles)
**Spec:** `docs/superpowers/specs/2026-03-22-gpu-composite-reward-design.md`
---
## File Structure
| File | Responsibility | Action |
|------|---------------|--------|
| `crates/ml/src/cuda_pipeline/experience_kernels.cu` | GPU reward computation | Modify: new kernel params, composite reward formula, stride 3→12 |
| `crates/ml/src/cuda_pipeline/gpu_experience_collector.rs` | Buffer allocation, kernel launch | Modify: PORTFOLIO_STRIDE, allocation, init, launch args |
| `crates/ml/src/trainers/dqn/config.rs` | DQNHyperparameters struct | Modify: add 7 reward fields |
| `crates/ml/src/training_profile.rs` | TOML deserialization | Modify: add RewardSection, apply_to mapping |
| `crates/ml/src/hyperopt/adapters/dqn.rs` | Hyperopt parameter space | Modify: DQNParams 31D→38D, bounds, from/to_continuous |
| `config/training/dqn-production.toml` | Production config | Modify: add [reward] section |
| `config/training/dqn-smoketest.toml` | Smoketest config | Modify: add [reward] section |
| `config/training/dqn-hyperopt.toml` | Hyperopt search bounds | Modify: add reward search bounds + phase_fast |
| `crates/ml-dqn/src/reward.rs` | CPU reward (to be gutted) | Modify: remove RewardNormalizer, CPU DSR |
---
### Task 1: CUDA Kernel — Composite Reward
**Files:**
- Modify: `crates/ml/src/cuda_pipeline/experience_kernels.cu`
- [ ] **Step 1: Add PORTFOLIO_STRIDE constant and update stride references**
At top of file, add:
```cuda
#define PORTFOLIO_STRIDE 12
```
Replace all `* 3` portfolio stride references:
- Line 169: `portfolio_states + (long long)i * 3``* PORTFOLIO_STRIDE`
- Line 420: `portfolio_states + (long long)i * 3``* PORTFOLIO_STRIDE`
Do NOT change `portfolio_sim_kernel` (uses stride 8, separate path).
- [ ] **Step 2: Add reward parameters to `experience_env_step` kernel signature**
Add after existing `float hold_reward` parameter (which will be removed):
```cuda
float w_dsr, float w_pnl, float w_dd, float w_idle,
float dd_threshold, float loss_aversion, float time_decay_rate,
float eta,
const float* __restrict__ features,
int market_dim
```
Remove `float hold_reward` parameter.
- [ ] **Step 3: Replace reward computation (lines 454-465) with composite formula**
Replace the current:
```cuda
float reward = pnl - tx_cost;
if (fabsf(position) < 0.001f) { reward += hold_reward; }
```
With the full 8-component composite. Read portfolio state indices 3-11 at kernel start, write updated values at kernel end. Include:
- DSR with pre-update A/B (Moody & Saffell)
- Normalized PnL with running EMA
- Drawdown penalty with peak_equity guard
- Idle penalty (replaces hold_reward)
- Regime scaling from ADX/CUSUM features
- Asymmetric loss scaling
- Position-time decay
- All division-by-zero guards from spec
- [ ] **Step 4: Update portfolio state initialization in `experience_state_gather`**
Ensure indices 3-11 are preserved across steps (not overwritten by gather kernel). The gather kernel currently reads indices 0-2 for state assembly — it must NOT touch indices 3-11.
- [ ] **Step 5: Compile check**
```bash
SQLX_OFFLINE=true cargo check -p ml 2>&1 | tail -5
```
Expected: compilation errors from Rust side (kernel params changed but Rust not updated yet). CUDA `.cu` file compiles via NVRTC at runtime.
- [ ] **Step 6: Commit**
```bash
git add crates/ml/src/cuda_pipeline/experience_kernels.cu
git commit -m "feat: 8-component composite reward CUDA kernel"
```
---
### Task 2: Rust Config Pipeline — DQNHyperparameters + TOML Profile
**Files:**
- Modify: `crates/ml/src/trainers/dqn/config.rs`
- Modify: `crates/ml/src/training_profile.rs`
- Modify: `config/training/dqn-production.toml`
- Modify: `config/training/dqn-smoketest.toml`
- Modify: `config/training/dqn-hyperopt.toml`
- [ ] **Step 1: Add 7 reward fields to DQNHyperparameters**
In `config.rs`, add to the struct (near existing `hold_penalty` / `transaction_cost_multiplier`):
```rust
pub w_dsr: f64,
pub w_pnl: f64,
pub w_dd: f64,
pub w_idle: f64,
pub dd_threshold: f64,
pub loss_aversion: f64,
pub time_decay_rate: f64,
```
Set defaults in `conservative()`:
```rust
w_dsr: 1.0,
w_pnl: 0.3,
w_dd: 1.0,
w_idle: 0.01,
dd_threshold: 0.02,
loss_aversion: 1.5,
time_decay_rate: 0.0005,
```
Remove `hold_penalty` field (replaced by `w_idle`).
- [ ] **Step 2: Add RewardSection to training_profile.rs**
```rust
#[derive(Debug, Clone, Deserialize, Default)]
pub struct RewardSection {
pub w_dsr: Option<f64>,
pub w_pnl: Option<f64>,
pub w_dd: Option<f64>,
pub w_idle: Option<f64>,
pub dd_threshold: Option<f64>,
pub loss_aversion: Option<f64>,
pub time_decay_rate: Option<f64>,
}
```
Add `pub reward: Option<RewardSection>` to `DqnTrainingProfile`.
Add `apply_to` mapping in the `[reward]` section handler.
- [ ] **Step 3: Add [reward] section to all 3 TOML profiles**
All profiles get the same defaults:
```toml
[reward]
w_dsr = 1.0
w_pnl = 0.3
w_dd = 1.0
w_idle = 0.01
dd_threshold = 0.02
loss_aversion = 1.5
time_decay_rate = 0.0005
```
Remove `hold_reward` from `[experience]` sections.
- [ ] **Step 4: Compile and run profile tests**
```bash
SQLX_OFFLINE=true cargo test -p ml --lib training_profile -- --test-threads=1
```
Expected: all profile tests pass.
- [ ] **Step 5: Commit**
```bash
git add crates/ml/src/trainers/dqn/config.rs crates/ml/src/training_profile.rs config/training/
git commit -m "feat: reward config pipeline — DQNHyperparameters + TOML profiles"
```
---
### Task 3: Experience Collector — Buffer Allocation + Kernel Launch
**Files:**
- Modify: `crates/ml/src/cuda_pipeline/gpu_experience_collector.rs`
- [ ] **Step 1: Add PORTFOLIO_STRIDE constant and update allocation**
```rust
const PORTFOLIO_STRIDE: usize = 12;
```
Replace all `* 3` with `* PORTFOLIO_STRIDE` at lines 575, 579, 581, 1166, 1168.
Update portfolio_init to set indices 7 and 9 to `initial_capital`:
```rust
portfolio_init[off + 7] = initial_capital; // peak_equity
portfolio_init[off + 9] = initial_capital; // prev_equity
```
- [ ] **Step 2: Add reward fields to ExperienceCollectorConfig**
```rust
pub w_dsr: f32,
pub w_pnl: f32,
pub w_dd: f32,
pub w_idle: f32,
pub dd_threshold: f32,
pub loss_aversion: f32,
pub time_decay_rate: f32,
```
Remove `hold_reward` field. Remove `use_dsr` field — DSR is always enabled, no opt-out path.
- [ ] **Step 3: Update kernel launch args**
In the `launch_timestep_loop` method, replace `hold_rw` kernel arg with the 7 reward weights + feature buffer pointer + market_dim.
The features buffer pointer is already available as `self.features_buf` (the market data uploaded for state_gather). Pass it to `experience_env_step` as well.
- [ ] **Step 4: Update all callers that construct ExperienceCollectorConfig**
Search for all places that create `ExperienceCollectorConfig` and add the new fields. The training loop constructs it in `collect_gpu_experiences()`.
- [ ] **Step 5: Compile check**
```bash
SQLX_OFFLINE=true cargo check -p ml 2>&1 | tail -5
```
Expected: clean compilation.
- [ ] **Step 6: Commit**
```bash
git add crates/ml/src/cuda_pipeline/gpu_experience_collector.rs
git commit -m "feat: experience collector — PORTFOLIO_STRIDE=12, reward kernel args"
```
---
### Task 4: Hyperopt Adapter — 31D → 38D Parameter Space
**Files:**
- Modify: `crates/ml/src/hyperopt/adapters/dqn.rs`
- Modify: `config/training/dqn-hyperopt.toml`
- [ ] **Step 1: Add 7 reward fields to DQNParams struct**
Add after existing fields:
```rust
pub w_dsr: f64,
pub w_pnl: f64,
pub w_dd: f64,
pub w_idle: f64,
pub dd_threshold: f64,
pub loss_aversion: f64,
pub time_decay_rate: f64,
```
Remove `hold_penalty_weight` field.
- [ ] **Step 2: Update continuous_bounds() — add 7 dimensions**
Add bounds loaded from TOML `[search_space]`:
```rust
let w_dsr_b = b("w_dsr", (0.1, 2.0));
let w_pnl_b = b("w_pnl", (0.0, 1.0));
let w_dd_b = b("w_dd", (0.0, 5.0));
let w_idle_b = b("w_idle", (0.0, 0.1));
let dd_thresh_b = b("dd_threshold", (0.01, 0.10));
let loss_av_b = b("loss_aversion", (1.0, 3.0));
let time_dec_b = b("time_decay_rate", (0.0001, 0.005));
```
Append to bounds vec (indices 31-37). Remove `hold_penalty_weight` bound.
- [ ] **Step 3: Update from_continuous() and to_continuous()**
`from_continuous`: expect 38 elements, map indices 31-37 to reward fields.
`to_continuous`: output 38 elements, append reward fields.
`param_names`: add 7 names, remove `hold_penalty_weight`.
- [ ] **Step 4: Update [search_space] and [phase_fast] in dqn-hyperopt.toml**
```toml
[search_space]
# ... existing ...
w_dsr = [0.1, 2.0]
w_pnl = [0.0, 1.0]
w_dd = [0.0, 5.0]
w_idle = [0.0, 0.1]
dd_threshold = [0.01, 0.10]
loss_aversion = [1.0, 3.0]
time_decay_rate = [0.0001, 0.005]
[phase_fast]
# ... existing ...
w_dsr = 1.0
w_pnl = 0.3
w_dd = 1.0
w_idle = 0.01
dd_threshold = 0.02
loss_aversion = 1.5
time_decay_rate = 0.0005
```
- [ ] **Step 5: Update SearchSpaceSection in training_profile.rs**
Add 7 `Option<[f64; 2]>` fields to `SearchSpaceSection` and the `bound()` match.
- [ ] **Step 6: Fix all tests that reference 31D or hold_penalty_weight**
Update `test_dqn_params_bounds`, `test_qr_dqn_activation_threshold`, and any other tests that assert dimension count or specific param indices.
- [ ] **Step 7: Run hyperopt tests**
```bash
SQLX_OFFLINE=true cargo test -p ml --lib hyperopt -- --test-threads=1
```
Expected: all pass with 38D search space.
- [ ] **Step 8: Commit**
```bash
git add crates/ml/src/hyperopt/adapters/dqn.rs crates/ml/src/training_profile.rs config/training/dqn-hyperopt.toml
git commit -m "feat: hyperopt 31D→38D — 7 reward weight dimensions"
```
---
### Task 5: Remove Dead CPU Reward Code
**Files:**
- Modify: `crates/ml-dqn/src/reward.rs`
- Modify: `crates/ml/src/trainers/dqn/trainer/training_loop.rs`
- [ ] **Step 1: Remove RewardNormalizer usage from training loop**
The training loop may call `reward_fn.normalize()` or similar. Remove those calls — normalization is now in the GPU kernel.
- [ ] **Step 2: Remove CPU DSR code from reward.rs**
Keep the file but remove the `RewardNormalizer` struct and the CPU DSR computation. Keep any types or traits that are still referenced.
- [ ] **Step 3: Remove `use_dsr` from ExperienceCollectorConfig callers**
Search for all `use_dsr` references and remove them. `w_dsr > 0` implies DSR enabled.
- [ ] **Step 4: Compile check**
```bash
SQLX_OFFLINE=true cargo check --workspace 2>&1 | grep "error\[" | head -5
```
Expected: clean or only unrelated warnings.
- [ ] **Step 5: Commit**
```bash
git add crates/ml-dqn/src/reward.rs crates/ml/src/trainers/dqn/trainer/training_loop.rs
git commit -m "refactor: remove dead CPU reward code (DSR + normalizer now in GPU kernel)"
```
---
### Task 6: Integration Test — Composite Reward on GPU
**Files:**
- Test with existing: `crates/ml/tests/dqn_training_smoke_test.rs`
- [ ] **Step 1: Run smoke test locally**
```bash
SQLX_OFFLINE=true cargo test -p ml --test dqn_training_smoke_test -- --nocapture --test-threads=1
```
Expected: passes with composite reward (TOML profile provides defaults). Q-values should be non-zero. Loss should decrease.
- [ ] **Step 2: Run full hyperopt test suite**
```bash
SQLX_OFFLINE=true cargo test -p ml --lib hyperopt -- --test-threads=1
```
Expected: all pass with 38D space.
- [ ] **Step 3: Run local hyperopt smoketest (2 trials, 2 epochs)**
```bash
SQLX_OFFLINE=true timeout 60 ./target/release/examples/hyperopt_baseline_rl \
--model dqn --phase fast --trials 6 --n-initial 5 --epochs 2 \
--data-dir test_data/futures-baseline --symbol ES.FUT \
--output /tmp/reward_test.json --seed 42
```
Expected: 6 trials complete, Q-values non-zero, Sharpe improves from trial 1.
- [ ] **Step 4: Run all 17 GPU tests**
```bash
SQLX_OFFLINE=true cargo test -p ml -p ml-dqn -p ml-core -- --test-threads=1
```
Expected: all pass.
- [ ] **Step 5: Commit and push**
```bash
git push origin main
```
---
### Task 7: H100 Validation
- [ ] **Step 1: Submit GPU test pipeline**
Verify 17/17 pass on H100 with composite reward.
- [ ] **Step 2: Submit hyperopt smoketest (6 trials × 10 epochs)**
Verify non-zero Q-values, improving Sharpe, all trials complete.
- [ ] **Step 3: Compare metrics**
Compare best Sharpe, win rate, max drawdown with the old PnL-only results (best was Sharpe 0.06, win rate 3%). The composite reward should show improvement.