Files
foxhunt/docs/superpowers/plans/2026-04-16-training-environment-alignment.md
jgrusewski 685746231b plan: Training Environment Alignment — 5 tasks, 3 phases
Phase 1: Position-gated episodes (h100.toml + done flag + soft reset)
Phase 2: Adapt rank norm (skip zeros) + remove hardcoded Q-drift
Phase 3: Aligned Sharpe metrics (un-annualized per-trade)
Phase 4: Build verification + smoke test

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-16 21:49:19 +02:00

13 KiB
Raw Blame History

Training Environment Alignment Implementation Plan

For agentic workers: REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (- [ ]) syntax for tracking.

Goal: Close the 45× gap between training Sharpe (0.19) and validation Sharpe (25) by aligning the training experience collector with the validation backtest — position-gated episodes, adapted reward shaping, and comparable metrics.

Architecture: 3 phases, 8 steps. Phase 1 changes episode structure (config + CUDA done flag + soft reset). Phase 2 adapts reward shaping (rank normalization + DSR + Q-drift). Phase 3 aligns metrics (un-annualized per-trade Sharpe). All changes in existing files — no new files created.

Tech Stack: Rust 1.85, CUDA 12.4, TOML config


Task 1: Phase 1 — Episode Length + Position-Gated Done + Soft Reset

The most impactful change. Episodes go from 100 fixed bars to 5000 bars with position-gated done flags.

Files:

  • Modify: config/gpu/h100.toml

  • Modify: crates/ml/src/cuda_pipeline/experience_kernels.cu

  • Modify: crates/ml/src/cuda_pipeline/gpu_experience_collector.rs

  • Step 1: Change H100 config

In config/gpu/h100.toml, change:

[experience]
gpu_timesteps_per_episode = 5000
gpu_n_episodes = 1024

Was: gpu_timesteps_per_episode = 100 (no gpu_n_episodes, defaulted to 4096).

VRAM: 1024 × 5000 × 80 × 4 = 1.6 GB (fits in H100 80GB).

  • Step 2: Add position-gated done to experience_env_step

In crates/ml/src/cuda_pipeline/experience_kernels.cu, find line 1758:

int done = (next_bar >= total_bars || check_capital_floor(new_portfolio_value, peak_equity)) ? 1 : 0;

Replace with:

    /* Position-gated done: episode ends when model completes a trade cycle
     * (was positioned → now flat). V(flat)=0 is a correct terminal anchor.
     * Also done at end-of-data or capital floor breach (existing). */
    int trade_complete = (fabsf(pre_trade_position) > 0.001f && fabsf(position) < 0.001f);
    int capital_breach = check_capital_floor(new_portfolio_value, peak_equity);
    int data_end = (next_bar >= total_bars);
    int done = (trade_complete || capital_breach || data_end) ? 1 : 0;
  • Step 3: Change portfolio reset at done to soft reset

In the same kernel, find the episode reset block (around line 1963):

    if (done) {

Inside this block, find where cash and position are reset. Change to SOFT reset — keep equity, clear position state:

The current reset likely does:

    cash = initial_capital;
    position = 0.0f;

Change to:

    /* Soft reset: keep equity (accumulated gains/losses), clear trade state.
     * The model starts the next segment with its real P&L, not artificial reset. */
    // cash stays at current value (no reset to initial_capital)
    // position is already 0 (trade_complete means we went flat)
    // Clear trade tracking state:
    entry_price = 0.0f;
    trade_start_pnl = 0.0f;
    // Keep: equity, cash, win/loss counters, DSR state

Read the actual reset block carefully before editing — there may be multiple fields being reset. Only reset TRADE-SPECIFIC state, not PORTFOLIO state.

NOTE: For data_end and capital_breach done, keep the FULL reset (portfolio goes back to initial capital for the next data window). Only trade_complete gets the soft reset.

    if (done) {
        if (data_end || capital_breach) {
            /* Hard reset: new data window or capital breach — full restart */
            cash = initial_capital;
            position = 0.0f;
            // ... existing full reset code ...
        } else {
            /* Soft reset (trade_complete): keep equity, clear trade state */
            entry_price = 0.0f;
            // position already 0 (model chose flat)
        }
    }
  • Step 4: Update gpu_n_episodes config handling

In crates/ml/src/cuda_pipeline/gpu_experience_collector.rs, check if gpu_n_episodes is read from config. If it's hardcoded or uses a different field name, update to use the config value. Search for where n_episodes or alloc_episodes is determined:

The experience collector's new() takes n_episodes as a parameter. The caller (training_loop.rs) may auto-scale it. Check that reducing from 4096 to 1024 is handled — the auto-scaler might override the config.

Search training_loop.rs for n_episodes or auto_scale and ensure 1024 is respected.

  • Step 5: Verify compilation
SQLX_OFFLINE=true cargo check -p ml 2>&1 | tail -5
  • Step 6: Commit
cd /home/jgrusewski/Work/foxhunt && git add config/gpu/h100.toml crates/ml/src/cuda_pipeline/experience_kernels.cu crates/ml/src/cuda_pipeline/gpu_experience_collector.rs && git commit -m "feat: position-gated episodes + 5000-bar limit — close 45× training/val gap

Episode done flag: timer-based → position-gated (trade complete = done).
V(flat)=0 is correct terminal anchor. Soft reset keeps equity.
H100: 100 bars → 5000 bars, 4096 episodes → 1024 episodes.
VRAM: 1.6GB (fits in 80GB). 12× more experience per epoch.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>"

Task 2: Phase 2 — Adapt Rank Normalization for Sparse Rewards

Rank only non-zero rewards (actual trades). Skip the 95% zeros from holding bars.

Files:

  • Modify: crates/ml/src/cuda_pipeline/reward_shaping_kernel.cu

  • Modify: crates/ml/src/cuda_pipeline/gpu_experience_collector.rs

  • Step 1: Add zero-filtering to rank normalization kernel

In crates/ml/src/cuda_pipeline/reward_shaping_kernel.cu, the reward_rank_normalize kernel processes ALL N rewards. Modify to SKIP near-zero rewards:

After the line that loads the raw reward:

    float raw_r = (i < N) ? rewards_in[i] : 0.0f;

Add zero-filtering:

    /* Skip near-zero rewards (holding bars with tiny holding cost).
     * Only rank actual trade P&L. Zeros get rank 0.5 (neutral). */
    int is_trade = (fabsf(raw_r) > 0.001f) ? 1 : 0;

Then in the ranking loop, only count against OTHER trades:

    /* Only rank against other non-zero (trade) rewards */
    if (!is_trade) {
        rewards_out[i] = 0.0f;  /* holding bars: zero reward (no rank) */
        return;
    }

For the counting section, only compare against other trades (where fabsf(tile_val) > 0.001f). This requires modifying the tile comparison loop.

Alternatively — simpler approach: after rank normalization, zero out the holding bars:

    if (!is_trade) {
        rewards_out[i] = 0.0f;
    }

This preserves the existing ranking logic for trades and just masks out the zeros. The model trains on rank-normalized TRADE returns and zero holding bars.

  • Step 2: Verify and commit
SQLX_OFFLINE=true cargo check -p ml 2>&1 | tail -5
cd /home/jgrusewski/Work/foxhunt && git add crates/ml/src/cuda_pipeline/reward_shaping_kernel.cu && git commit -m "feat: rank normalization skips zero (holding) rewards — only ranks trades

With trade-level reward, 95% of bars have reward≈0 (holding cost).
Ranking zeros is meaningless — a $500 winner and $5 winner got similar
ranks. Now: holding bars get reward=0 (no rank), only actual trade
P&L is ranked. Preserves relative ordering among real trades.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>"

Task 3: Phase 2 — Remove Hardcoded Q-Drift Penalty

The E1 enrichment (Q-value reality check) handles Q-drift adaptively from eval data. The hardcoded penalty in c51_grad_kernel is redundant and fights the adaptive system.

Files:

  • Modify: crates/ml/src/cuda_pipeline/c51_grad_kernel.cu

  • Step 1: Remove hardcoded drift penalty

In crates/ml/src/cuda_pipeline/c51_grad_kernel.cu, find lines 136-141:

    /* Q-mean drift regularization: ADAPTIVE quadratic penalty.
     * Small drift → tiny penalty (don't interfere). Large drift → hard correction.
     * lambda=0.1 (scaled up from 0.01 for trade-level reward's smaller Q-values). */
    float q_mean_ema = q_mean_ema_ptr[0];
    float drift_penalty = 0.1f * q_mean_ema * fabsf(q_mean_ema);
    d_val_sum += drift_penalty;

Replace with:

    /* Q-mean drift: handled by E1 enrichment (Q-value reality check) which
     * computes bias correction from eval data and applies it adaptively.
     * Hardcoded penalty removed — it competed with E1 and used a fixed lambda
     * that couldn't adapt to changing reward scales. */
    (void)q_mean_ema_ptr; /* keep parameter to avoid kernel signature change */

NOTE: Keep the q_mean_ema_ptr parameter in the kernel signature to avoid changing ALL launch sites. Just don't USE it for the penalty.

  • Step 2: Verify and commit
SQLX_OFFLINE=true cargo check -p ml 2>&1 | tail -5
cd /home/jgrusewski/Work/foxhunt && git add crates/ml/src/cuda_pipeline/c51_grad_kernel.cu && git commit -m "feat: remove hardcoded Q-drift penalty — E1 enrichment handles it adaptively

The quadratic penalty (0.1 * q_mean * |q_mean|) competed with the
E1 enrichment (Q-value reality check) which computes bias correction
from actual eval performance. E1 adapts to changing reward scales;
the hardcoded penalty couldn't. Kernel parameter kept for ABI compat.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>"

Task 4: Phase 3 — Aligned Sharpe Metrics

Add un-annualized per-trade Sharpe to both training and validation logging.

Files:

  • Modify: crates/ml/src/trainers/dqn/financials.rs

  • Modify: crates/ml/src/trainers/dqn/trainer/training_loop.rs

  • Step 1: Add raw Sharpe to EpochFinancials

In crates/ml/src/trainers/dqn/financials.rs, find the EpochFinancials struct and add:

    /// Un-annualized per-trade Sharpe (mean_trade_return / std_trade_return).
    /// Directly comparable between training and validation.
    pub sharpe_raw: f64,

In compute_epoch_financials, after the annualized Sharpe computation (around line 101), compute the raw version:

        let sharpe_raw = if std > 1e-10 { mean / std } else { 0.0 };

And set sharpe_raw in the returned struct.

  • Step 2: Log raw Sharpe alongside annualized

In crates/ml/src/trainers/dqn/trainer/training_loop.rs, find the epoch Sharpe logging (around line 2198):

info!(
    "Epoch {}/{}: Sharpe={:.2} WinRate={:.1}% ...",

Add raw Sharpe:

info!(
    "Epoch {}/{}: Sharpe={:.2} Sharpe_raw={:.4} WinRate={:.1}% ...",
    epoch + 1, self.hyperparams.epochs,
    financials.sharpe, financials.sharpe_raw, ...

Also log the val_Sharpe raw equivalent:

// After val_Sharpe logging:
let val_bar_annualization = (self.hyperparams.bars_per_day * 252.0).sqrt();
let val_sharpe_raw = val_sharpe / val_bar_annualization;
info!("  val_Sharpe_raw={:.4} (un-annualized per-bar)", val_sharpe_raw);
  • Step 3: Verify and commit
SQLX_OFFLINE=true cargo check -p ml 2>&1 | tail -5
cd /home/jgrusewski/Work/foxhunt && git add crates/ml/src/trainers/dqn/financials.rs crates/ml/src/trainers/dqn/trainer/training_loop.rs && git commit -m "feat: aligned Sharpe metrics — un-annualized per-trade for direct comparison

Add sharpe_raw (mean/std, no annualization) to both training and
validation logging. Training uses per-trade. Validation uses per-bar.
Both now report raw alongside annualized for honest comparison.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>"

Task 5: Build Verification + Smoke Test

Files:

  • Modify: crates/ml/src/trainers/dqn/smoke_tests/generalization.rs

  • Step 1: Update smoke test for longer episodes

The smoke test uses batch_size=4096 which might conflict with the new gpu_n_episodes=1024. Check if the smoke test overrides episode config. If it uses smoke_params(), it may not pick up h100.toml. The smoke test runs on RTX 3050 — keep rtx3050.toml at 100 timesteps for fast local testing.

Verify the smoke test still passes:

SQLX_OFFLINE=true FOXHUNT_TEST_DATA=test_data/futures-baseline cargo test -p ml --lib -- test_generalization_components_smoke --include-ignored --nocapture 2>&1 | tail -20
  • Step 2: Run compute-sanitizer
FOXHUNT_TEST_DATA=test_data/futures-baseline compute-sanitizer --tool memcheck --print-limit 5 target/debug/deps/ml-* "test_generalization_components_smoke" --test-threads=1 --include-ignored 2>&1 | grep "ERROR SUMMARY"

Target: 0 errors.

  • Step 3: Run full test suite
SQLX_OFFLINE=true cargo test -p ml --lib 2>&1 | tail -5

Target: 899+ passed, 0 failed.

  • Step 4: Final commit
cd /home/jgrusewski/Work/foxhunt && git add -A && git commit -m "chore: training environment alignment — build verification

Position-gated episodes (5000 bars), adapted rank normalization,
removed hardcoded Q-drift, aligned Sharpe metrics.
Smoke test passes. compute-sanitizer: 0 errors. 899+ tests pass.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>"