Task 1: CUDA kernel (8-component formula, PORTFOLIO_STRIDE=12) Task 2: Config pipeline (DQNHyperparameters + TOML profiles) Task 3: Experience collector (buffer alloc + kernel launch) Task 4: Hyperopt adapter (31D→38D search space) Task 5: Remove dead CPU reward code Task 6: Integration test (smoke + hyperopt + GPU suite) Task 7: H100 validation Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
13 KiB
GPU Composite Reward Implementation Plan
For agentic workers: REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (
- [ ]) syntax for tracking.
Goal: Replace the raw PnL reward in the CUDA experience kernel with an 8-component composite reward (DSR, normalized PnL, drawdown penalty, idle penalty, regime scaling, asymmetric loss, position-time decay, transaction costs) — all GPU-native, zero CPU involvement.
Architecture: Extend experience_env_step kernel with composite reward computation. Portfolio state buffer grows from 3→12 floats per episode. 7 new reward weights flow through TOML config → DQNHyperparameters → ExperienceCollectorConfig → kernel args. Hyperopt search space grows from 31D→38D.
Tech Stack: CUDA C (experience kernel), Rust (config pipeline, hyperopt adapter), TOML (training profiles)
Spec: docs/superpowers/specs/2026-03-22-gpu-composite-reward-design.md
File Structure
| File | Responsibility | Action |
|---|---|---|
crates/ml/src/cuda_pipeline/experience_kernels.cu |
GPU reward computation | Modify: new kernel params, composite reward formula, stride 3→12 |
crates/ml/src/cuda_pipeline/gpu_experience_collector.rs |
Buffer allocation, kernel launch | Modify: PORTFOLIO_STRIDE, allocation, init, launch args |
crates/ml/src/trainers/dqn/config.rs |
DQNHyperparameters struct | Modify: add 7 reward fields |
crates/ml/src/training_profile.rs |
TOML deserialization | Modify: add RewardSection, apply_to mapping |
crates/ml/src/hyperopt/adapters/dqn.rs |
Hyperopt parameter space | Modify: DQNParams 31D→38D, bounds, from/to_continuous |
config/training/dqn-production.toml |
Production config | Modify: add [reward] section |
config/training/dqn-smoketest.toml |
Smoketest config | Modify: add [reward] section |
config/training/dqn-hyperopt.toml |
Hyperopt search bounds | Modify: add reward search bounds + phase_fast |
crates/ml-dqn/src/reward.rs |
CPU reward (to be gutted) | Modify: remove RewardNormalizer, CPU DSR |
Task 1: CUDA Kernel — Composite Reward
Files:
-
Modify:
crates/ml/src/cuda_pipeline/experience_kernels.cu -
Step 1: Add PORTFOLIO_STRIDE constant and update stride references
At top of file, add:
#define PORTFOLIO_STRIDE 12
Replace all * 3 portfolio stride references:
- Line 169:
portfolio_states + (long long)i * 3→* PORTFOLIO_STRIDE - Line 420:
portfolio_states + (long long)i * 3→* PORTFOLIO_STRIDE
Do NOT change portfolio_sim_kernel (uses stride 8, separate path).
- Step 2: Add reward parameters to
experience_env_stepkernel signature
Add after existing float hold_reward parameter (which will be removed):
float w_dsr, float w_pnl, float w_dd, float w_idle,
float dd_threshold, float loss_aversion, float time_decay_rate,
float eta,
const float* __restrict__ features,
int market_dim
Remove float hold_reward parameter.
- Step 3: Replace reward computation (lines 454-465) with composite formula
Replace the current:
float reward = pnl - tx_cost;
if (fabsf(position) < 0.001f) { reward += hold_reward; }
With the full 8-component composite. Read portfolio state indices 3-11 at kernel start, write updated values at kernel end. Include:
-
DSR with pre-update A/B (Moody & Saffell)
-
Normalized PnL with running EMA
-
Drawdown penalty with peak_equity guard
-
Idle penalty (replaces hold_reward)
-
Regime scaling from ADX/CUSUM features
-
Asymmetric loss scaling
-
Position-time decay
-
All division-by-zero guards from spec
-
Step 4: Update portfolio state initialization in
experience_state_gather
Ensure indices 3-11 are preserved across steps (not overwritten by gather kernel). The gather kernel currently reads indices 0-2 for state assembly — it must NOT touch indices 3-11.
- Step 5: Compile check
SQLX_OFFLINE=true cargo check -p ml 2>&1 | tail -5
Expected: compilation errors from Rust side (kernel params changed but Rust not updated yet). CUDA .cu file compiles via NVRTC at runtime.
- Step 6: Commit
git add crates/ml/src/cuda_pipeline/experience_kernels.cu
git commit -m "feat: 8-component composite reward CUDA kernel"
Task 2: Rust Config Pipeline — DQNHyperparameters + TOML Profile
Files:
-
Modify:
crates/ml/src/trainers/dqn/config.rs -
Modify:
crates/ml/src/training_profile.rs -
Modify:
config/training/dqn-production.toml -
Modify:
config/training/dqn-smoketest.toml -
Modify:
config/training/dqn-hyperopt.toml -
Step 1: Add 7 reward fields to DQNHyperparameters
In config.rs, add to the struct (near existing hold_penalty / transaction_cost_multiplier):
pub w_dsr: f64,
pub w_pnl: f64,
pub w_dd: f64,
pub w_idle: f64,
pub dd_threshold: f64,
pub loss_aversion: f64,
pub time_decay_rate: f64,
Set defaults in conservative():
w_dsr: 1.0,
w_pnl: 0.3,
w_dd: 1.0,
w_idle: 0.01,
dd_threshold: 0.02,
loss_aversion: 1.5,
time_decay_rate: 0.0005,
Remove hold_penalty field (replaced by w_idle).
- Step 2: Add RewardSection to training_profile.rs
#[derive(Debug, Clone, Deserialize, Default)]
pub struct RewardSection {
pub w_dsr: Option<f64>,
pub w_pnl: Option<f64>,
pub w_dd: Option<f64>,
pub w_idle: Option<f64>,
pub dd_threshold: Option<f64>,
pub loss_aversion: Option<f64>,
pub time_decay_rate: Option<f64>,
}
Add pub reward: Option<RewardSection> to DqnTrainingProfile.
Add apply_to mapping in the [reward] section handler.
- Step 3: Add [reward] section to all 3 TOML profiles
All profiles get the same defaults:
[reward]
w_dsr = 1.0
w_pnl = 0.3
w_dd = 1.0
w_idle = 0.01
dd_threshold = 0.02
loss_aversion = 1.5
time_decay_rate = 0.0005
Remove hold_reward from [experience] sections.
- Step 4: Compile and run profile tests
SQLX_OFFLINE=true cargo test -p ml --lib training_profile -- --test-threads=1
Expected: all profile tests pass.
- Step 5: Commit
git add crates/ml/src/trainers/dqn/config.rs crates/ml/src/training_profile.rs config/training/
git commit -m "feat: reward config pipeline — DQNHyperparameters + TOML profiles"
Task 3: Experience Collector — Buffer Allocation + Kernel Launch
Files:
-
Modify:
crates/ml/src/cuda_pipeline/gpu_experience_collector.rs -
Step 1: Add PORTFOLIO_STRIDE constant and update allocation
const PORTFOLIO_STRIDE: usize = 12;
Replace all * 3 with * PORTFOLIO_STRIDE at lines 575, 579, 581, 1166, 1168.
Update portfolio_init to set indices 7 and 9 to initial_capital:
portfolio_init[off + 7] = initial_capital; // peak_equity
portfolio_init[off + 9] = initial_capital; // prev_equity
- Step 2: Add reward fields to ExperienceCollectorConfig
pub w_dsr: f32,
pub w_pnl: f32,
pub w_dd: f32,
pub w_idle: f32,
pub dd_threshold: f32,
pub loss_aversion: f32,
pub time_decay_rate: f32,
Remove hold_reward field. Remove use_dsr field — DSR is always enabled, no opt-out path.
- Step 3: Update kernel launch args
In the launch_timestep_loop method, replace hold_rw kernel arg with the 7 reward weights + feature buffer pointer + market_dim.
The features buffer pointer is already available as self.features_buf (the market data uploaded for state_gather). Pass it to experience_env_step as well.
- Step 4: Update all callers that construct ExperienceCollectorConfig
Search for all places that create ExperienceCollectorConfig and add the new fields. The training loop constructs it in collect_gpu_experiences().
- Step 5: Compile check
SQLX_OFFLINE=true cargo check -p ml 2>&1 | tail -5
Expected: clean compilation.
- Step 6: Commit
git add crates/ml/src/cuda_pipeline/gpu_experience_collector.rs
git commit -m "feat: experience collector — PORTFOLIO_STRIDE=12, reward kernel args"
Task 4: Hyperopt Adapter — 31D → 38D Parameter Space
Files:
-
Modify:
crates/ml/src/hyperopt/adapters/dqn.rs -
Modify:
config/training/dqn-hyperopt.toml -
Step 1: Add 7 reward fields to DQNParams struct
Add after existing fields:
pub w_dsr: f64,
pub w_pnl: f64,
pub w_dd: f64,
pub w_idle: f64,
pub dd_threshold: f64,
pub loss_aversion: f64,
pub time_decay_rate: f64,
Remove hold_penalty_weight field.
- Step 2: Update continuous_bounds() — add 7 dimensions
Add bounds loaded from TOML [search_space]:
let w_dsr_b = b("w_dsr", (0.1, 2.0));
let w_pnl_b = b("w_pnl", (0.0, 1.0));
let w_dd_b = b("w_dd", (0.0, 5.0));
let w_idle_b = b("w_idle", (0.0, 0.1));
let dd_thresh_b = b("dd_threshold", (0.01, 0.10));
let loss_av_b = b("loss_aversion", (1.0, 3.0));
let time_dec_b = b("time_decay_rate", (0.0001, 0.005));
Append to bounds vec (indices 31-37). Remove hold_penalty_weight bound.
- Step 3: Update from_continuous() and to_continuous()
from_continuous: expect 38 elements, map indices 31-37 to reward fields.
to_continuous: output 38 elements, append reward fields.
param_names: add 7 names, remove hold_penalty_weight.
- Step 4: Update [search_space] and [phase_fast] in dqn-hyperopt.toml
[search_space]
# ... existing ...
w_dsr = [0.1, 2.0]
w_pnl = [0.0, 1.0]
w_dd = [0.0, 5.0]
w_idle = [0.0, 0.1]
dd_threshold = [0.01, 0.10]
loss_aversion = [1.0, 3.0]
time_decay_rate = [0.0001, 0.005]
[phase_fast]
# ... existing ...
w_dsr = 1.0
w_pnl = 0.3
w_dd = 1.0
w_idle = 0.01
dd_threshold = 0.02
loss_aversion = 1.5
time_decay_rate = 0.0005
- Step 5: Update SearchSpaceSection in training_profile.rs
Add 7 Option<[f64; 2]> fields to SearchSpaceSection and the bound() match.
- Step 6: Fix all tests that reference 31D or hold_penalty_weight
Update test_dqn_params_bounds, test_qr_dqn_activation_threshold, and any other tests that assert dimension count or specific param indices.
- Step 7: Run hyperopt tests
SQLX_OFFLINE=true cargo test -p ml --lib hyperopt -- --test-threads=1
Expected: all pass with 38D search space.
- Step 8: Commit
git add crates/ml/src/hyperopt/adapters/dqn.rs crates/ml/src/training_profile.rs config/training/dqn-hyperopt.toml
git commit -m "feat: hyperopt 31D→38D — 7 reward weight dimensions"
Task 5: Remove Dead CPU Reward Code
Files:
-
Modify:
crates/ml-dqn/src/reward.rs -
Modify:
crates/ml/src/trainers/dqn/trainer/training_loop.rs -
Step 1: Remove RewardNormalizer usage from training loop
The training loop may call reward_fn.normalize() or similar. Remove those calls — normalization is now in the GPU kernel.
- Step 2: Remove CPU DSR code from reward.rs
Keep the file but remove the RewardNormalizer struct and the CPU DSR computation. Keep any types or traits that are still referenced.
- Step 3: Remove
use_dsrfrom ExperienceCollectorConfig callers
Search for all use_dsr references and remove them. w_dsr > 0 implies DSR enabled.
- Step 4: Compile check
SQLX_OFFLINE=true cargo check --workspace 2>&1 | grep "error\[" | head -5
Expected: clean or only unrelated warnings.
- Step 5: Commit
git add crates/ml-dqn/src/reward.rs crates/ml/src/trainers/dqn/trainer/training_loop.rs
git commit -m "refactor: remove dead CPU reward code (DSR + normalizer now in GPU kernel)"
Task 6: Integration Test — Composite Reward on GPU
Files:
-
Test with existing:
crates/ml/tests/dqn_training_smoke_test.rs -
Step 1: Run smoke test locally
SQLX_OFFLINE=true cargo test -p ml --test dqn_training_smoke_test -- --nocapture --test-threads=1
Expected: passes with composite reward (TOML profile provides defaults). Q-values should be non-zero. Loss should decrease.
- Step 2: Run full hyperopt test suite
SQLX_OFFLINE=true cargo test -p ml --lib hyperopt -- --test-threads=1
Expected: all pass with 38D space.
- Step 3: Run local hyperopt smoketest (2 trials, 2 epochs)
SQLX_OFFLINE=true timeout 60 ./target/release/examples/hyperopt_baseline_rl \
--model dqn --phase fast --trials 6 --n-initial 5 --epochs 2 \
--data-dir test_data/futures-baseline --symbol ES.FUT \
--output /tmp/reward_test.json --seed 42
Expected: 6 trials complete, Q-values non-zero, Sharpe improves from trial 1.
- Step 4: Run all 17 GPU tests
SQLX_OFFLINE=true cargo test -p ml -p ml-dqn -p ml-core -- --test-threads=1
Expected: all pass.
- Step 5: Commit and push
git push origin main
Task 7: H100 Validation
- Step 1: Submit GPU test pipeline
Verify 17/17 pass on H100 with composite reward.
- Step 2: Submit hyperopt smoketest (6 trials × 10 epochs)
Verify non-zero Q-values, improving Sharpe, all trials complete.
- Step 3: Compare metrics
Compare best Sharpe, win rate, max drawdown with the old PnL-only results (best was Sharpe 0.06, win rate 3%). The composite reward should show improvement.