From 685746231b04cd8bbbc909569739bb9c79025a9b Mon Sep 17 00:00:00 2001 From: jgrusewski Date: Thu, 16 Apr 2026 21:49:19 +0200 Subject: [PATCH] =?UTF-8?q?plan:=20Training=20Environment=20Alignment=20?= =?UTF-8?q?=E2=80=94=205=20tasks,=203=20phases?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Phase 1: Position-gated episodes (h100.toml + done flag + soft reset) Phase 2: Adapt rank norm (skip zeros) + remove hardcoded Q-drift Phase 3: Aligned Sharpe metrics (un-annualized per-trade) Phase 4: Build verification + smoke test Co-Authored-By: Claude Opus 4.6 (1M context) --- ...26-04-16-training-environment-alignment.md | 355 ++++++++++++++++++ 1 file changed, 355 insertions(+) create mode 100644 docs/superpowers/plans/2026-04-16-training-environment-alignment.md diff --git a/docs/superpowers/plans/2026-04-16-training-environment-alignment.md b/docs/superpowers/plans/2026-04-16-training-environment-alignment.md new file mode 100644 index 000000000..ae8c890dd --- /dev/null +++ b/docs/superpowers/plans/2026-04-16-training-environment-alignment.md @@ -0,0 +1,355 @@ +# Training Environment Alignment Implementation Plan + +> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking. + +**Goal:** Close the 45× gap between training Sharpe (0.19) and validation Sharpe (25) by aligning the training experience collector with the validation backtest — position-gated episodes, adapted reward shaping, and comparable metrics. + +**Architecture:** 3 phases, 8 steps. Phase 1 changes episode structure (config + CUDA done flag + soft reset). Phase 2 adapts reward shaping (rank normalization + DSR + Q-drift). Phase 3 aligns metrics (un-annualized per-trade Sharpe). All changes in existing files — no new files created. + +**Tech Stack:** Rust 1.85, CUDA 12.4, TOML config + +--- + +## Task 1: Phase 1 — Episode Length + Position-Gated Done + Soft Reset + +The most impactful change. Episodes go from 100 fixed bars to 5000 bars with position-gated done flags. + +**Files:** +- Modify: `config/gpu/h100.toml` +- Modify: `crates/ml/src/cuda_pipeline/experience_kernels.cu` +- Modify: `crates/ml/src/cuda_pipeline/gpu_experience_collector.rs` + +- [ ] **Step 1: Change H100 config** + +In `config/gpu/h100.toml`, change: + +```toml +[experience] +gpu_timesteps_per_episode = 5000 +gpu_n_episodes = 1024 +``` + +Was: `gpu_timesteps_per_episode = 100` (no gpu_n_episodes, defaulted to 4096). + +VRAM: `1024 × 5000 × 80 × 4 = 1.6 GB` (fits in H100 80GB). + +- [ ] **Step 2: Add position-gated done to experience_env_step** + +In `crates/ml/src/cuda_pipeline/experience_kernels.cu`, find line 1758: + +```cuda +int done = (next_bar >= total_bars || check_capital_floor(new_portfolio_value, peak_equity)) ? 1 : 0; +``` + +Replace with: + +```cuda + /* Position-gated done: episode ends when model completes a trade cycle + * (was positioned → now flat). V(flat)=0 is a correct terminal anchor. + * Also done at end-of-data or capital floor breach (existing). */ + int trade_complete = (fabsf(pre_trade_position) > 0.001f && fabsf(position) < 0.001f); + int capital_breach = check_capital_floor(new_portfolio_value, peak_equity); + int data_end = (next_bar >= total_bars); + int done = (trade_complete || capital_breach || data_end) ? 1 : 0; +``` + +- [ ] **Step 3: Change portfolio reset at done to soft reset** + +In the same kernel, find the episode reset block (around line 1963): + +```cuda + if (done) { +``` + +Inside this block, find where cash and position are reset. Change to SOFT reset — keep equity, clear position state: + +The current reset likely does: +```cuda + cash = initial_capital; + position = 0.0f; +``` + +Change to: +```cuda + /* Soft reset: keep equity (accumulated gains/losses), clear trade state. + * The model starts the next segment with its real P&L, not artificial reset. */ + // cash stays at current value (no reset to initial_capital) + // position is already 0 (trade_complete means we went flat) + // Clear trade tracking state: + entry_price = 0.0f; + trade_start_pnl = 0.0f; + // Keep: equity, cash, win/loss counters, DSR state +``` + +Read the actual reset block carefully before editing — there may be multiple fields being reset. Only reset TRADE-SPECIFIC state, not PORTFOLIO state. + +NOTE: For `data_end` and `capital_breach` done, keep the FULL reset (portfolio goes back to initial capital for the next data window). Only `trade_complete` gets the soft reset. + +```cuda + if (done) { + if (data_end || capital_breach) { + /* Hard reset: new data window or capital breach — full restart */ + cash = initial_capital; + position = 0.0f; + // ... existing full reset code ... + } else { + /* Soft reset (trade_complete): keep equity, clear trade state */ + entry_price = 0.0f; + // position already 0 (model chose flat) + } + } +``` + +- [ ] **Step 4: Update gpu_n_episodes config handling** + +In `crates/ml/src/cuda_pipeline/gpu_experience_collector.rs`, check if `gpu_n_episodes` is read from config. If it's hardcoded or uses a different field name, update to use the config value. Search for where `n_episodes` or `alloc_episodes` is determined: + +The experience collector's `new()` takes `n_episodes` as a parameter. The caller (training_loop.rs) may auto-scale it. Check that reducing from 4096 to 1024 is handled — the auto-scaler might override the config. + +Search `training_loop.rs` for `n_episodes` or `auto_scale` and ensure 1024 is respected. + +- [ ] **Step 5: Verify compilation** + +```bash +SQLX_OFFLINE=true cargo check -p ml 2>&1 | tail -5 +``` + +- [ ] **Step 6: Commit** + +```bash +cd /home/jgrusewski/Work/foxhunt && git add config/gpu/h100.toml crates/ml/src/cuda_pipeline/experience_kernels.cu crates/ml/src/cuda_pipeline/gpu_experience_collector.rs && git commit -m "feat: position-gated episodes + 5000-bar limit — close 45× training/val gap + +Episode done flag: timer-based → position-gated (trade complete = done). +V(flat)=0 is correct terminal anchor. Soft reset keeps equity. +H100: 100 bars → 5000 bars, 4096 episodes → 1024 episodes. +VRAM: 1.6GB (fits in 80GB). 12× more experience per epoch. + +Co-Authored-By: Claude Opus 4.6 (1M context) " +``` + +--- + +## Task 2: Phase 2 — Adapt Rank Normalization for Sparse Rewards + +Rank only non-zero rewards (actual trades). Skip the 95% zeros from holding bars. + +**Files:** +- Modify: `crates/ml/src/cuda_pipeline/reward_shaping_kernel.cu` +- Modify: `crates/ml/src/cuda_pipeline/gpu_experience_collector.rs` + +- [ ] **Step 1: Add zero-filtering to rank normalization kernel** + +In `crates/ml/src/cuda_pipeline/reward_shaping_kernel.cu`, the `reward_rank_normalize` kernel processes ALL N rewards. Modify to SKIP near-zero rewards: + +After the line that loads the raw reward: +```cuda + float raw_r = (i < N) ? rewards_in[i] : 0.0f; +``` + +Add zero-filtering: +```cuda + /* Skip near-zero rewards (holding bars with tiny holding cost). + * Only rank actual trade P&L. Zeros get rank 0.5 (neutral). */ + int is_trade = (fabsf(raw_r) > 0.001f) ? 1 : 0; +``` + +Then in the ranking loop, only count against OTHER trades: +```cuda + /* Only rank against other non-zero (trade) rewards */ + if (!is_trade) { + rewards_out[i] = 0.0f; /* holding bars: zero reward (no rank) */ + return; + } +``` + +For the counting section, only compare against other trades (where `fabsf(tile_val) > 0.001f`). This requires modifying the tile comparison loop. + +Alternatively — simpler approach: after rank normalization, zero out the holding bars: +```cuda + if (!is_trade) { + rewards_out[i] = 0.0f; + } +``` + +This preserves the existing ranking logic for trades and just masks out the zeros. The model trains on rank-normalized TRADE returns and zero holding bars. + +- [ ] **Step 2: Verify and commit** + +```bash +SQLX_OFFLINE=true cargo check -p ml 2>&1 | tail -5 +``` + +```bash +cd /home/jgrusewski/Work/foxhunt && git add crates/ml/src/cuda_pipeline/reward_shaping_kernel.cu && git commit -m "feat: rank normalization skips zero (holding) rewards — only ranks trades + +With trade-level reward, 95% of bars have reward≈0 (holding cost). +Ranking zeros is meaningless — a $500 winner and $5 winner got similar +ranks. Now: holding bars get reward=0 (no rank), only actual trade +P&L is ranked. Preserves relative ordering among real trades. + +Co-Authored-By: Claude Opus 4.6 (1M context) " +``` + +--- + +## Task 3: Phase 2 — Remove Hardcoded Q-Drift Penalty + +The E1 enrichment (Q-value reality check) handles Q-drift adaptively from eval data. The hardcoded penalty in c51_grad_kernel is redundant and fights the adaptive system. + +**Files:** +- Modify: `crates/ml/src/cuda_pipeline/c51_grad_kernel.cu` + +- [ ] **Step 1: Remove hardcoded drift penalty** + +In `crates/ml/src/cuda_pipeline/c51_grad_kernel.cu`, find lines 136-141: + +```cuda + /* Q-mean drift regularization: ADAPTIVE quadratic penalty. + * Small drift → tiny penalty (don't interfere). Large drift → hard correction. + * lambda=0.1 (scaled up from 0.01 for trade-level reward's smaller Q-values). */ + float q_mean_ema = q_mean_ema_ptr[0]; + float drift_penalty = 0.1f * q_mean_ema * fabsf(q_mean_ema); + d_val_sum += drift_penalty; +``` + +Replace with: +```cuda + /* Q-mean drift: handled by E1 enrichment (Q-value reality check) which + * computes bias correction from eval data and applies it adaptively. + * Hardcoded penalty removed — it competed with E1 and used a fixed lambda + * that couldn't adapt to changing reward scales. */ + (void)q_mean_ema_ptr; /* keep parameter to avoid kernel signature change */ +``` + +NOTE: Keep the `q_mean_ema_ptr` parameter in the kernel signature to avoid changing ALL launch sites. Just don't USE it for the penalty. + +- [ ] **Step 2: Verify and commit** + +```bash +SQLX_OFFLINE=true cargo check -p ml 2>&1 | tail -5 +``` + +```bash +cd /home/jgrusewski/Work/foxhunt && git add crates/ml/src/cuda_pipeline/c51_grad_kernel.cu && git commit -m "feat: remove hardcoded Q-drift penalty — E1 enrichment handles it adaptively + +The quadratic penalty (0.1 * q_mean * |q_mean|) competed with the +E1 enrichment (Q-value reality check) which computes bias correction +from actual eval performance. E1 adapts to changing reward scales; +the hardcoded penalty couldn't. Kernel parameter kept for ABI compat. + +Co-Authored-By: Claude Opus 4.6 (1M context) " +``` + +--- + +## Task 4: Phase 3 — Aligned Sharpe Metrics + +Add un-annualized per-trade Sharpe to both training and validation logging. + +**Files:** +- Modify: `crates/ml/src/trainers/dqn/financials.rs` +- Modify: `crates/ml/src/trainers/dqn/trainer/training_loop.rs` + +- [ ] **Step 1: Add raw Sharpe to EpochFinancials** + +In `crates/ml/src/trainers/dqn/financials.rs`, find the `EpochFinancials` struct and add: + +```rust + /// Un-annualized per-trade Sharpe (mean_trade_return / std_trade_return). + /// Directly comparable between training and validation. + pub sharpe_raw: f64, +``` + +In `compute_epoch_financials`, after the annualized Sharpe computation (around line 101), compute the raw version: + +```rust + let sharpe_raw = if std > 1e-10 { mean / std } else { 0.0 }; +``` + +And set `sharpe_raw` in the returned struct. + +- [ ] **Step 2: Log raw Sharpe alongside annualized** + +In `crates/ml/src/trainers/dqn/trainer/training_loop.rs`, find the epoch Sharpe logging (around line 2198): + +```rust +info!( + "Epoch {}/{}: Sharpe={:.2} WinRate={:.1}% ...", +``` + +Add raw Sharpe: +```rust +info!( + "Epoch {}/{}: Sharpe={:.2} Sharpe_raw={:.4} WinRate={:.1}% ...", + epoch + 1, self.hyperparams.epochs, + financials.sharpe, financials.sharpe_raw, ... +``` + +Also log the val_Sharpe raw equivalent: +```rust +// After val_Sharpe logging: +let val_bar_annualization = (self.hyperparams.bars_per_day * 252.0).sqrt(); +let val_sharpe_raw = val_sharpe / val_bar_annualization; +info!(" val_Sharpe_raw={:.4} (un-annualized per-bar)", val_sharpe_raw); +``` + +- [ ] **Step 3: Verify and commit** + +```bash +SQLX_OFFLINE=true cargo check -p ml 2>&1 | tail -5 +``` + +```bash +cd /home/jgrusewski/Work/foxhunt && git add crates/ml/src/trainers/dqn/financials.rs crates/ml/src/trainers/dqn/trainer/training_loop.rs && git commit -m "feat: aligned Sharpe metrics — un-annualized per-trade for direct comparison + +Add sharpe_raw (mean/std, no annualization) to both training and +validation logging. Training uses per-trade. Validation uses per-bar. +Both now report raw alongside annualized for honest comparison. + +Co-Authored-By: Claude Opus 4.6 (1M context) " +``` + +--- + +## Task 5: Build Verification + Smoke Test + +**Files:** +- Modify: `crates/ml/src/trainers/dqn/smoke_tests/generalization.rs` + +- [ ] **Step 1: Update smoke test for longer episodes** + +The smoke test uses `batch_size=4096` which might conflict with the new `gpu_n_episodes=1024`. Check if the smoke test overrides episode config. If it uses `smoke_params()`, it may not pick up h100.toml. The smoke test runs on RTX 3050 — keep `rtx3050.toml` at 100 timesteps for fast local testing. + +Verify the smoke test still passes: + +```bash +SQLX_OFFLINE=true FOXHUNT_TEST_DATA=test_data/futures-baseline cargo test -p ml --lib -- test_generalization_components_smoke --include-ignored --nocapture 2>&1 | tail -20 +``` + +- [ ] **Step 2: Run compute-sanitizer** + +```bash +FOXHUNT_TEST_DATA=test_data/futures-baseline compute-sanitizer --tool memcheck --print-limit 5 target/debug/deps/ml-* "test_generalization_components_smoke" --test-threads=1 --include-ignored 2>&1 | grep "ERROR SUMMARY" +``` + +Target: 0 errors. + +- [ ] **Step 3: Run full test suite** + +```bash +SQLX_OFFLINE=true cargo test -p ml --lib 2>&1 | tail -5 +``` + +Target: 899+ passed, 0 failed. + +- [ ] **Step 4: Final commit** + +```bash +cd /home/jgrusewski/Work/foxhunt && git add -A && git commit -m "chore: training environment alignment — build verification + +Position-gated episodes (5000 bars), adapted rank normalization, +removed hardcoded Q-drift, aligned Sharpe metrics. +Smoke test passes. compute-sanitizer: 0 errors. 899+ tests pass. + +Co-Authored-By: Claude Opus 4.6 (1M context) " +```