Files
foxhunt/docs/superpowers/plans/2026-04-16-training-environment-alignment.md
jgrusewski 685746231b plan: Training Environment Alignment — 5 tasks, 3 phases
Phase 1: Position-gated episodes (h100.toml + done flag + soft reset)
Phase 2: Adapt rank norm (skip zeros) + remove hardcoded Q-drift
Phase 3: Aligned Sharpe metrics (un-annualized per-trade)
Phase 4: Build verification + smoke test

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-16 21:49:19 +02:00

356 lines
13 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Training Environment Alignment Implementation Plan
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
**Goal:** Close the 45× gap between training Sharpe (0.19) and validation Sharpe (25) by aligning the training experience collector with the validation backtest — position-gated episodes, adapted reward shaping, and comparable metrics.
**Architecture:** 3 phases, 8 steps. Phase 1 changes episode structure (config + CUDA done flag + soft reset). Phase 2 adapts reward shaping (rank normalization + DSR + Q-drift). Phase 3 aligns metrics (un-annualized per-trade Sharpe). All changes in existing files — no new files created.
**Tech Stack:** Rust 1.85, CUDA 12.4, TOML config
---
## Task 1: Phase 1 — Episode Length + Position-Gated Done + Soft Reset
The most impactful change. Episodes go from 100 fixed bars to 5000 bars with position-gated done flags.
**Files:**
- Modify: `config/gpu/h100.toml`
- Modify: `crates/ml/src/cuda_pipeline/experience_kernels.cu`
- Modify: `crates/ml/src/cuda_pipeline/gpu_experience_collector.rs`
- [ ] **Step 1: Change H100 config**
In `config/gpu/h100.toml`, change:
```toml
[experience]
gpu_timesteps_per_episode = 5000
gpu_n_episodes = 1024
```
Was: `gpu_timesteps_per_episode = 100` (no gpu_n_episodes, defaulted to 4096).
VRAM: `1024 × 5000 × 80 × 4 = 1.6 GB` (fits in H100 80GB).
- [ ] **Step 2: Add position-gated done to experience_env_step**
In `crates/ml/src/cuda_pipeline/experience_kernels.cu`, find line 1758:
```cuda
int done = (next_bar >= total_bars || check_capital_floor(new_portfolio_value, peak_equity)) ? 1 : 0;
```
Replace with:
```cuda
/* Position-gated done: episode ends when model completes a trade cycle
* (was positioned → now flat). V(flat)=0 is a correct terminal anchor.
* Also done at end-of-data or capital floor breach (existing). */
int trade_complete = (fabsf(pre_trade_position) > 0.001f && fabsf(position) < 0.001f);
int capital_breach = check_capital_floor(new_portfolio_value, peak_equity);
int data_end = (next_bar >= total_bars);
int done = (trade_complete || capital_breach || data_end) ? 1 : 0;
```
- [ ] **Step 3: Change portfolio reset at done to soft reset**
In the same kernel, find the episode reset block (around line 1963):
```cuda
if (done) {
```
Inside this block, find where cash and position are reset. Change to SOFT reset — keep equity, clear position state:
The current reset likely does:
```cuda
cash = initial_capital;
position = 0.0f;
```
Change to:
```cuda
/* Soft reset: keep equity (accumulated gains/losses), clear trade state.
* The model starts the next segment with its real P&L, not artificial reset. */
// cash stays at current value (no reset to initial_capital)
// position is already 0 (trade_complete means we went flat)
// Clear trade tracking state:
entry_price = 0.0f;
trade_start_pnl = 0.0f;
// Keep: equity, cash, win/loss counters, DSR state
```
Read the actual reset block carefully before editing — there may be multiple fields being reset. Only reset TRADE-SPECIFIC state, not PORTFOLIO state.
NOTE: For `data_end` and `capital_breach` done, keep the FULL reset (portfolio goes back to initial capital for the next data window). Only `trade_complete` gets the soft reset.
```cuda
if (done) {
if (data_end || capital_breach) {
/* Hard reset: new data window or capital breach — full restart */
cash = initial_capital;
position = 0.0f;
// ... existing full reset code ...
} else {
/* Soft reset (trade_complete): keep equity, clear trade state */
entry_price = 0.0f;
// position already 0 (model chose flat)
}
}
```
- [ ] **Step 4: Update gpu_n_episodes config handling**
In `crates/ml/src/cuda_pipeline/gpu_experience_collector.rs`, check if `gpu_n_episodes` is read from config. If it's hardcoded or uses a different field name, update to use the config value. Search for where `n_episodes` or `alloc_episodes` is determined:
The experience collector's `new()` takes `n_episodes` as a parameter. The caller (training_loop.rs) may auto-scale it. Check that reducing from 4096 to 1024 is handled — the auto-scaler might override the config.
Search `training_loop.rs` for `n_episodes` or `auto_scale` and ensure 1024 is respected.
- [ ] **Step 5: Verify compilation**
```bash
SQLX_OFFLINE=true cargo check -p ml 2>&1 | tail -5
```
- [ ] **Step 6: Commit**
```bash
cd /home/jgrusewski/Work/foxhunt && git add config/gpu/h100.toml crates/ml/src/cuda_pipeline/experience_kernels.cu crates/ml/src/cuda_pipeline/gpu_experience_collector.rs && git commit -m "feat: position-gated episodes + 5000-bar limit — close 45× training/val gap
Episode done flag: timer-based → position-gated (trade complete = done).
V(flat)=0 is correct terminal anchor. Soft reset keeps equity.
H100: 100 bars → 5000 bars, 4096 episodes → 1024 episodes.
VRAM: 1.6GB (fits in 80GB). 12× more experience per epoch.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>"
```
---
## Task 2: Phase 2 — Adapt Rank Normalization for Sparse Rewards
Rank only non-zero rewards (actual trades). Skip the 95% zeros from holding bars.
**Files:**
- Modify: `crates/ml/src/cuda_pipeline/reward_shaping_kernel.cu`
- Modify: `crates/ml/src/cuda_pipeline/gpu_experience_collector.rs`
- [ ] **Step 1: Add zero-filtering to rank normalization kernel**
In `crates/ml/src/cuda_pipeline/reward_shaping_kernel.cu`, the `reward_rank_normalize` kernel processes ALL N rewards. Modify to SKIP near-zero rewards:
After the line that loads the raw reward:
```cuda
float raw_r = (i < N) ? rewards_in[i] : 0.0f;
```
Add zero-filtering:
```cuda
/* Skip near-zero rewards (holding bars with tiny holding cost).
* Only rank actual trade P&L. Zeros get rank 0.5 (neutral). */
int is_trade = (fabsf(raw_r) > 0.001f) ? 1 : 0;
```
Then in the ranking loop, only count against OTHER trades:
```cuda
/* Only rank against other non-zero (trade) rewards */
if (!is_trade) {
rewards_out[i] = 0.0f; /* holding bars: zero reward (no rank) */
return;
}
```
For the counting section, only compare against other trades (where `fabsf(tile_val) > 0.001f`). This requires modifying the tile comparison loop.
Alternatively — simpler approach: after rank normalization, zero out the holding bars:
```cuda
if (!is_trade) {
rewards_out[i] = 0.0f;
}
```
This preserves the existing ranking logic for trades and just masks out the zeros. The model trains on rank-normalized TRADE returns and zero holding bars.
- [ ] **Step 2: Verify and commit**
```bash
SQLX_OFFLINE=true cargo check -p ml 2>&1 | tail -5
```
```bash
cd /home/jgrusewski/Work/foxhunt && git add crates/ml/src/cuda_pipeline/reward_shaping_kernel.cu && git commit -m "feat: rank normalization skips zero (holding) rewards — only ranks trades
With trade-level reward, 95% of bars have reward≈0 (holding cost).
Ranking zeros is meaningless — a $500 winner and $5 winner got similar
ranks. Now: holding bars get reward=0 (no rank), only actual trade
P&L is ranked. Preserves relative ordering among real trades.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>"
```
---
## Task 3: Phase 2 — Remove Hardcoded Q-Drift Penalty
The E1 enrichment (Q-value reality check) handles Q-drift adaptively from eval data. The hardcoded penalty in c51_grad_kernel is redundant and fights the adaptive system.
**Files:**
- Modify: `crates/ml/src/cuda_pipeline/c51_grad_kernel.cu`
- [ ] **Step 1: Remove hardcoded drift penalty**
In `crates/ml/src/cuda_pipeline/c51_grad_kernel.cu`, find lines 136-141:
```cuda
/* Q-mean drift regularization: ADAPTIVE quadratic penalty.
* Small drift → tiny penalty (don't interfere). Large drift → hard correction.
* lambda=0.1 (scaled up from 0.01 for trade-level reward's smaller Q-values). */
float q_mean_ema = q_mean_ema_ptr[0];
float drift_penalty = 0.1f * q_mean_ema * fabsf(q_mean_ema);
d_val_sum += drift_penalty;
```
Replace with:
```cuda
/* Q-mean drift: handled by E1 enrichment (Q-value reality check) which
* computes bias correction from eval data and applies it adaptively.
* Hardcoded penalty removed — it competed with E1 and used a fixed lambda
* that couldn't adapt to changing reward scales. */
(void)q_mean_ema_ptr; /* keep parameter to avoid kernel signature change */
```
NOTE: Keep the `q_mean_ema_ptr` parameter in the kernel signature to avoid changing ALL launch sites. Just don't USE it for the penalty.
- [ ] **Step 2: Verify and commit**
```bash
SQLX_OFFLINE=true cargo check -p ml 2>&1 | tail -5
```
```bash
cd /home/jgrusewski/Work/foxhunt && git add crates/ml/src/cuda_pipeline/c51_grad_kernel.cu && git commit -m "feat: remove hardcoded Q-drift penalty — E1 enrichment handles it adaptively
The quadratic penalty (0.1 * q_mean * |q_mean|) competed with the
E1 enrichment (Q-value reality check) which computes bias correction
from actual eval performance. E1 adapts to changing reward scales;
the hardcoded penalty couldn't. Kernel parameter kept for ABI compat.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>"
```
---
## Task 4: Phase 3 — Aligned Sharpe Metrics
Add un-annualized per-trade Sharpe to both training and validation logging.
**Files:**
- Modify: `crates/ml/src/trainers/dqn/financials.rs`
- Modify: `crates/ml/src/trainers/dqn/trainer/training_loop.rs`
- [ ] **Step 1: Add raw Sharpe to EpochFinancials**
In `crates/ml/src/trainers/dqn/financials.rs`, find the `EpochFinancials` struct and add:
```rust
/// Un-annualized per-trade Sharpe (mean_trade_return / std_trade_return).
/// Directly comparable between training and validation.
pub sharpe_raw: f64,
```
In `compute_epoch_financials`, after the annualized Sharpe computation (around line 101), compute the raw version:
```rust
let sharpe_raw = if std > 1e-10 { mean / std } else { 0.0 };
```
And set `sharpe_raw` in the returned struct.
- [ ] **Step 2: Log raw Sharpe alongside annualized**
In `crates/ml/src/trainers/dqn/trainer/training_loop.rs`, find the epoch Sharpe logging (around line 2198):
```rust
info!(
"Epoch {}/{}: Sharpe={:.2} WinRate={:.1}% ...",
```
Add raw Sharpe:
```rust
info!(
"Epoch {}/{}: Sharpe={:.2} Sharpe_raw={:.4} WinRate={:.1}% ...",
epoch + 1, self.hyperparams.epochs,
financials.sharpe, financials.sharpe_raw, ...
```
Also log the val_Sharpe raw equivalent:
```rust
// After val_Sharpe logging:
let val_bar_annualization = (self.hyperparams.bars_per_day * 252.0).sqrt();
let val_sharpe_raw = val_sharpe / val_bar_annualization;
info!(" val_Sharpe_raw={:.4} (un-annualized per-bar)", val_sharpe_raw);
```
- [ ] **Step 3: Verify and commit**
```bash
SQLX_OFFLINE=true cargo check -p ml 2>&1 | tail -5
```
```bash
cd /home/jgrusewski/Work/foxhunt && git add crates/ml/src/trainers/dqn/financials.rs crates/ml/src/trainers/dqn/trainer/training_loop.rs && git commit -m "feat: aligned Sharpe metrics — un-annualized per-trade for direct comparison
Add sharpe_raw (mean/std, no annualization) to both training and
validation logging. Training uses per-trade. Validation uses per-bar.
Both now report raw alongside annualized for honest comparison.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>"
```
---
## Task 5: Build Verification + Smoke Test
**Files:**
- Modify: `crates/ml/src/trainers/dqn/smoke_tests/generalization.rs`
- [ ] **Step 1: Update smoke test for longer episodes**
The smoke test uses `batch_size=4096` which might conflict with the new `gpu_n_episodes=1024`. Check if the smoke test overrides episode config. If it uses `smoke_params()`, it may not pick up h100.toml. The smoke test runs on RTX 3050 — keep `rtx3050.toml` at 100 timesteps for fast local testing.
Verify the smoke test still passes:
```bash
SQLX_OFFLINE=true FOXHUNT_TEST_DATA=test_data/futures-baseline cargo test -p ml --lib -- test_generalization_components_smoke --include-ignored --nocapture 2>&1 | tail -20
```
- [ ] **Step 2: Run compute-sanitizer**
```bash
FOXHUNT_TEST_DATA=test_data/futures-baseline compute-sanitizer --tool memcheck --print-limit 5 target/debug/deps/ml-* "test_generalization_components_smoke" --test-threads=1 --include-ignored 2>&1 | grep "ERROR SUMMARY"
```
Target: 0 errors.
- [ ] **Step 3: Run full test suite**
```bash
SQLX_OFFLINE=true cargo test -p ml --lib 2>&1 | tail -5
```
Target: 899+ passed, 0 failed.
- [ ] **Step 4: Final commit**
```bash
cd /home/jgrusewski/Work/foxhunt && git add -A && git commit -m "chore: training environment alignment — build verification
Position-gated episodes (5000 bars), adapted rank normalization,
removed hardcoded Q-drift, aligned Sharpe metrics.
Smoke test passes. compute-sanitizer: 0 errors. 899+ tests pass.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>"
```