Files
foxhunt/docs/superpowers/plans/2026-03-24-master-trading-agent.md
jgrusewski be2672596d docs: master plan — maximize all 25 DQN features for profitable ES trading
12 tasks covering:
- Reward v5: trade-aware hybrid (dense + sparse) with patience multiplier
- Dynamic trailing stop as environment physics (regime-adaptive)
- Activate 11 dormant features: IQN CVaR, HER, ensemble, DT, CQL, curiosity,
  count bonus, entropy, Kelly sizing, Q-gap conviction, CVaR action selection
- Three-phase training: behavioral cloning → online RL → DT refinement
- Ensemble consensus action selection with uncertainty-based sizing
- Full hyperopt search space (50+ dimensions)
- Smoke test validation: model must produce winning trades

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-24 22:00:45 +01:00

24 KiB
Raw Blame History

MASTER PLAN: Maximum Profitable DQN Trading Agent

For agentic workers: REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (- [ ]) syntax for tracking.

Goal: Activate all 25 DQN features into a unified, maximally profitable ES futures swing trading agent — nothing left on the shelf.

Architecture: Three-phase training pipeline (Behavioral Cloning → Full-Stack Online RL → Decision Transformer Refinement) with trade-aware hybrid reward, dynamic trade management as environment physics, ensemble consensus action selection, and IQN risk-aware position sizing. Every auxiliary head (IQN, ensemble, attention, curiosity, CQL, DT) injects gradient into the shared trunk via the fused CUDA training pipeline.

Tech Stack: CUDA kernels (nvcc precompiled), cuBLAS SGEMM, Rust async, cudarc 0.19.3, CUDA Graphs with split forward/adam phases.


Design Philosophy

We Are Building a Swing Trader, Not a Scalper

  • Timeframe: 1-minute bars, holding 5-100+ bars (5 min to 2+ hours)
  • Edge: Pattern recognition across 4 timeframes (5m/15m/1h/4h) + regime detection
  • Sizing: Conviction-based via ensemble agreement + IQN risk quantiles + Kelly optimal
  • Exits: Model learns when to exit from the reward signal, with dynamic guardrails
  • Goal: Fewer, higher-quality trades that capture multi-bar moves

The Agent Decides WHAT, the Environment Manages HOW

Agent Controls Environment Controls
Entry timing (when to trade) Transaction costs (real friction)
Exposure level (how much) Capital floor (25% max drawdown)
Order type (limit vs market) CVaR position scaling (tail risk)
Exit timing (when to close) Kelly sizing (optimal fraction)
Dynamic trailing stop (profit lock)

The Reward Teaches Trading, Not Just Returns

reward_t = trade_shaping(t) + exit_bonus(t)

trade_shaping(t):
  - IN TRADE: unrealized_change / equity × regime_quality  (dense, per-bar)
  - FLAT: 0.0  (doing nothing is free)

exit_bonus(t):
  - AT EXIT: trade_return × patience_multiplier  (sparse, strong)
  - patience_multiplier = sqrt(hold_time / regime_expected_hold)
  - Rewards patience: same return in fewer bars = less reward

Feature Activation Map

Tier 1: Foundation (14 features — already ACTIVE)

These run every step. No changes needed.

# Feature Status
1 Branching DQN (9×3×3) ACTIVE
2 C51 Distributional (51 atoms) ACTIVE
3 MSE Warmup + Blended Ramp ACTIVE
4 Multi-Timeframe Features (16-dim) ACTIVE
5 4-Head Attention on h_s2 ACTIVE
6 Expert Demonstrations (MA crossover) ACTIVE
7 N-Step Returns (n=3) ACTIVE
8 PER Beta Annealing ACTIVE
9 Noisy Networks (σ=0.5) ACTIVE
10 Circuit Breaker ACTIVE
11 Gradient Clipping ACTIVE
12 Soft Target Updates (τ=0.005) ACTIVE
13 Spectral Normalization ACTIVE
14 Financial Metrics Pipeline ACTIVE

Tier 2: Activate via Config (5 features — ACTIVE when flag > 0)

Just need hyperopt to search over non-zero values.

# Feature Config Flag Default
15 IQN Dual-Head iqn_lambda 0.25
16 HER Counterfactual her_ratio 0.0 → 0.2
17 Count Bonus (UCB) count_bonus_coefficient 0.0 → 0.1
18 Q-Gap Conviction q_gap_threshold 0.0 → 0.05
19 Regime PER Scaling (hardcoded true) ACTIVE

Tier 3: Wire into Pipeline (4 features — code exists, needs connection)

GPU kernels exist but aren't called from the training loop.

# Feature What's Missing
20 Kelly Sizing Apply kelly_fraction to position in experience kernel
21 CVaR Action Selection Remove hardcoded false, pass IQN quantiles to selector
22 Curiosity Reward Wire GpuCuriosityTrainer into experience collector
23 Decision Transformer Integrate DT pretrain into Phase 1 of training pipeline

Tier 4: Implement GPU Path (3 features — need new CUDA kernels)

Config flags exist but no GPU implementation.

# Feature What's Needed
24 CQL Conservative Loss GPU kernel for conservative Q-penalty on OOD actions
25 Ensemble Value Backward cuBLAS backward through value heads for KL diversity gradient

Tier 5: Reward & Trade Management Redesign (the core)

# Component Description
26 Trade-Aware Hybrid Reward Dense per-bar shaping + sparse trade-completion
27 Dynamic Trade Management Regime-aware trailing stop as environment physics
28 Three-Phase Training Pipeline BC → Online RL → DT Refinement

File Structure

File Action Responsibility
crates/ml/src/cuda_pipeline/experience_kernels.cu Rewrite reward section Trade-aware hybrid reward, dynamic trailing stop, Kelly/CVaR scaling
crates/ml/src/cuda_pipeline/gpu_dqn_trainer.rs Modify CQL loss kernel, ensemble backward, entropy in fused path
crates/ml/src/cuda_pipeline/gpu_experience_collector.rs Modify Wire curiosity reward into experience loop
crates/ml/src/cuda_pipeline/gpu_action_selector.rs Modify CVaR-aware action selection from IQN quantiles
crates/ml/src/trainers/dqn/fused_training.rs Modify DT auxiliary loss step, CQL step, curiosity training step
crates/ml/src/trainers/dqn/config.rs Modify Activate defaults for all Tier 2-3 features
crates/ml/src/trainers/dqn/trainer/training_loop.rs Modify Three-phase pipeline orchestration
crates/ml/src/trainers/dqn/trainer/constructor.rs Modify Remove hardcoded flags, respect config
crates/ml/src/trainers/dqn/trainer/action.rs Modify Ensemble consensus + IQN CVaR + Kelly sizing in action path
crates/ml/src/hyperopt/adapters/dqn.rs Modify Add search ranges for all new config fields
crates/ml/src/trainers/dqn/smoke_tests/training_stability.rs Modify Assert winning trades, positive Sharpe potential

Task 1: Trade-Aware Hybrid Reward (experience_kernels.cu)

This is the most critical task. The reward function determines what the model optimizes for.

Files:

  • Modify: crates/ml/src/cuda_pipeline/experience_kernels.cu — rewrite the reward section

Reward Design

/* ═══════════════════════════════════════════════════════════════════
 * REWARD v5: Trade-Aware Hybrid
 *
 * Two components, two timescales:
 *
 * 1. DENSE (every bar): Trade-level unrealized P&L change
 *    reward_dense = (unrealized_t - unrealized_{t-1}) / equity
 *    This is NOT portfolio return — it's specifically about THIS trade.
 *    When flat: reward_dense = 0.0 (no trade, no signal)
 *
 * 2. SPARSE (at trade exit): Full trade return with patience bonus
 *    reward_sparse = trade_return × sqrt(hold_time / expected_hold)
 *    expected_hold = base_hold × regime_factor(ADX)
 *    This rewards holding: same return over more bars = more reward
 *
 * The dense component keeps gradients flowing every bar.
 * The sparse component teaches the actual goal: profitable trades.
 *
 * Loss aversion: negative returns weighted 1.5x (both components)
 * ═══════════════════════════════════════════════════════════════════ */

Dynamic Trade Management (Environment Physics)

/* Dynamic trailing stop — activates when trade is profitable.
 *
 * trail_distance = base_trail × vol_scale(CUSUM) × trend_scale(ADX)
 *   - High vol (CUSUM > 0.5): wider trail (let noise pass)
 *   - Strong trend (ADX > 30): wider trail (let trend run)
 *   - Ranging (ADX < 20): tight trail (capture quick profits)
 *
 * The trailing stop is the ONLY auto-exit. No static SL/TP.
 * The model learns entry quality + exit timing through the reward.
 * The trailing stop only prevents giving back large gains.
 *
 * Minimum hold: regime-adaptive, learned via hold_time state feature
 *   - The model SEES hold_time in its state (slot +3 in state_gather)
 *   - It learns that exiting before min_hold means re-entering costs 2x tx
 *   - No hard override — just the natural tx cost penalty
 */
  • Step 1: Remove the old reward v2 8-component block (DSR, normalized PnL, drawdown, idle, hold penalty, profit-take, trade completion)
  • Step 2: Implement reward v5 dense component: per-bar unrealized P&L change for the current trade
  • Step 3: Implement reward v5 sparse component: trade-completion return × patience multiplier at exit
  • Step 4: Implement dynamic trailing stop: regime-aware trail distance, activates when unrealized > 0
  • Step 5: Keep capital floor protection (25% drawdown = game over, done=1)
  • Step 6: Keep CVaR/Kelly position scaling (environment physics, unchanged)
  • Step 7: Run SQLX_OFFLINE=true cargo check -p ml — verify clean build
  • Step 8: Commit: "feat: reward v5 — trade-aware hybrid with dynamic trailing stop"

Task 2: Activate Tier 2 Config Defaults

Files:

  • Modify: crates/ml/src/trainers/dqn/config.rs

Set non-zero defaults for features that are already wired but disabled by zero config values.

  • Step 1: Set her_ratio: 0.2 (was 0.0) — 20% of batch gets HER relabeling
  • Step 2: Set count_bonus_coefficient: Some(0.1) (was 0.0) — UCB exploration active
  • Step 3: Set q_gap_threshold: 0.05 (was 0.0) — conviction filter active with warmup ramp
  • Step 4: Set curiosity_weight: 0.05 (was 0.0) — intrinsic motivation for novel states
  • Step 5: Set use_cql: true (was false in practice) — conservative Q-learning
  • Step 6: Set ensemble_count: 3 (was 1) — 3-head ensemble for uncertainty
  • Step 7: Set dt_pretrain_epochs: 3 (was 0) — Decision Transformer pretraining
  • Step 8: Set use_cvar_action_selection: true in hyperparams (was hardcoded false)
  • Step 9: Set enable_kelly_sizing: true, kelly_fractional: 0.5, kelly_max_fraction: 0.25
  • Step 10: Run SQLX_OFFLINE=true cargo check -p ml — verify clean build
  • Step 11: Commit: "feat: activate all dormant features via config defaults"

Task 3: Wire Kelly Sizing into Experience Kernel

Files:

  • Modify: crates/ml/src/cuda_pipeline/experience_kernels.cu

Kelly criterion is already computed in the experience kernel (ps[14:19] track win/loss stats). It already scales target_position at line ~633. Verify it's working and not gated by a dead condition.

  • Step 1: Read the Kelly block in the experience kernel. Verify it reads ps[14:17] (win_count, loss_count, sum_wins, sum_losses) and computes kelly_f
  • Step 2: Verify the Kelly scaling is NOT inside a dead code branch (no if use_kelly guard that's always false)
  • Step 3: Verify the Kelly half-Kelly safety multiplier (0.5×) is applied
  • Step 4: Add Kelly fraction to the state_gather kernel as an observable feature — the model should SEE its own Kelly sizing to learn from it
  • Step 5: Commit: "fix: verify Kelly sizing active in experience kernel + expose as state feature"

Task 4: Wire CVaR Action Selection from IQN

Files:

  • Modify: crates/ml/src/trainers/dqn/trainer/constructor.rs — remove hardcoded false
  • Modify: crates/ml/src/trainers/dqn/trainer/action.rs — use IQN CVaR for position scaling

The IQN head has compute_cvar_q() that returns CVaR (worst-case quantile) Q-values. Use these to scale position sizing: when CVaR is very negative (high downside risk), reduce position.

  • Step 1: In constructor.rs, change use_cvar_action_selection: false to use_cvar_action_selection: hyperparams.use_cvar_action_selection
  • Step 2: In action.rs, after Q-value selection but before action output: if IQN is active, compute CVaR Q-values for the selected action
  • Step 3: Use CVaR to scale the selected exposure: effective_exposure = selected_exposure × cvar_scale where cvar_scale = sigmoid(cvar_q / threshold)
  • Step 4: Pass the CVaR-scaled exposure to the GPU action selector
  • Step 5: Commit: "feat: IQN CVaR-aware position scaling in action selection"

Task 5: Wire Curiosity into GPU Experience Collector

Files:

  • Modify: crates/ml/src/cuda_pipeline/gpu_experience_collector.rs
  • Modify: crates/ml/src/trainers/dqn/fused_training.rs

The GpuCuriosityTrainer exists and its forward model can predict next-state. The prediction error IS the curiosity reward. Wire it into the experience collection loop.

  • Step 1: In gpu_experience_collector.rs, after the env_step kernel writes rewards, launch the curiosity forward model on (state, action) pairs
  • Step 2: Compute prediction error: curiosity_reward = ||predicted_next_state - actual_next_state||²
  • Step 3: Add curiosity_reward × curiosity_scale to the experience rewards buffer via SAXPY kernel
  • Step 4: In fused_training.rs, after the main training step, train the curiosity forward model on the sampled batch (update the prediction model)
  • Step 5: Commit: "feat: GPU curiosity reward in experience collection + forward model training"

Task 6: Implement CQL Conservative Loss (GPU Kernel)

Files:

  • Modify: crates/ml/src/cuda_pipeline/gpu_dqn_trainer.rs
  • Modify: crates/ml/src/trainers/dqn/fused_training.rs

CQL adds a penalty for Q-values on out-of-distribution actions: L_cql = α × (log_sum_exp(Q(s,a)) - Q(s, a_data)). This prevents Q-value overestimation on actions the model hasn't tried.

  • Step 1: Add inline CUDA kernel cql_penalty_kernel in gpu_dqn_trainer.rs:
    • Inputs: Q-values for all actions [B, total_actions], actions taken [B]
    • Computes: logsumexp(Q) - Q(s, a_taken) per sample
    • Writes: per-sample CQL penalty [B]
  • Step 2: Add cql_penalty_grad_kernel — gradient of CQL penalty w.r.t. Q-values
  • Step 3: Compile both kernels in GpuDqnTrainer::new()
  • Step 4: Add launch_cql_penalty() and launch_cql_grad() methods
  • Step 5: In fused_training.rs run_full_step(), add CQL as Step 5g (after ensemble):
    if self.hyperparams.use_cql {
        let cql_loss = self.trainer.launch_cql_penalty(cql_alpha)?;
        self.trainer.launch_cql_grad()?;
        // CQL gradient already in d_value/d_adv logits, flows through backward
    }
    
  • Step 6: Commit: "feat: GPU CQL conservative loss kernel — prevents Q-value overestimation"

Task 7: Implement Ensemble Value Backward

Files:

  • Modify: crates/ml/src/trainers/dqn/fused_training.rs

The ensemble computes KL diversity gradient but doesn't backpropagate it through the value heads to the trunk. This means the trunk doesn't learn to produce diverse representations.

  • Step 1: In run_ensemble_step(), after KL gradient is computed, perform cuBLAS SGEMM backward through each ensemble value head
  • Step 2: Accumulate the trunk gradient from all K heads via SAXPY into grad_buf (scaled by diversity_weight)
  • Step 3: The combined gradient (C51 + IQN + attention + ensemble_diversity) all flow into the single Adam update
  • Step 4: Commit: "feat: ensemble value backward — KL diversity gradient flows to shared trunk"

Task 8: Integrate Decision Transformer into Phase 1

Files:

  • Modify: crates/ml/src/trainers/dqn/trainer/training_loop.rs
  • Modify: crates/ml/src/trainers/dqn/fused_training.rs

Decision Transformer pretraining uses historical trajectories to bootstrap the policy. In Phase 1 (first dt_pretrain_epochs epochs), the DT loss is an auxiliary objective alongside MSE.

  • Step 1: In training_loop.rs, during the first dt_pretrain_epochs epochs:
    • Build trajectory sequences from the experience buffer: (state, action, return-to-go) tuples
    • Upload to GPU via DT's trajectory builder kernel
  • Step 2: In fused_training.rs, add run_dt_pretrain_step():
    • DT forward: embed → causal attention → action prediction
    • DT loss: cross-entropy on predicted vs actual actions
    • DT backward: gradient through DT layers → trunk embedding gradient
    • SAXPY into grad_buf (scaled by dt_lambda)
  • Step 3: After dt_pretrain_epochs, disable DT auxiliary loss (let online RL take over)
  • Step 4: Commit: "feat: Decision Transformer pretraining in Phase 1 — behavioral cloning from trajectories"

Task 9: Ensemble Consensus Action Selection

Files:

  • Modify: crates/ml/src/trainers/dqn/trainer/action.rs

With K=3 ensemble heads, action selection should use the COMMITTEE, not just head 0:

  • Step 1: For each state, forward through all K value heads to get K sets of Q-values
  • Step 2: Compute mean Q (consensus) and variance (uncertainty) across heads
  • Step 3: Action selection: argmax on mean Q (consensus trade)
  • Step 4: Position sizing: scale by 1 / (1 + uncertainty) — high disagreement → smaller position
  • Step 5: If uncertainty > threshold: override to Flat (heads can't agree → sit out)
  • Step 6: Commit: "feat: ensemble consensus action selection — uncertainty-based sizing and sit-out"

Task 10: Three-Phase Training Pipeline

Files:

  • Modify: crates/ml/src/trainers/dqn/trainer/training_loop.rs

Wire the three training phases into the epoch loop:

Phase 1: Behavioral Cloning + MSE Warmup (epochs 0..dt_pretrain_epochs)

  • MSE loss (c51_alpha ≈ 0)
  • Expert demonstrations at high ratio (expert_ratio = 0.5)
  • Decision Transformer auxiliary loss
  • HER relabeling (learn from counterfactuals)

Phase 2: Full-Stack Online RL (epochs dt_pretrain_epochs..end)

  • Blended MSE→C51 ramp (c51_alpha ramps to 1.0)
  • Expert ratio decays to 0
  • All auxiliary heads active: IQN, ensemble, attention, curiosity, CQL
  • HER continues (counterfactual learning)
  • Dynamic trade management (trailing stops, CVaR scaling)

Phase 3: Refinement (last 20% of epochs)

  • Pure C51 loss (c51_alpha = 1.0)

  • No expert demonstrations

  • Ensemble variance-based exploration (replace epsilon-greedy)

  • Shrink-and-perturb for plasticity maintenance

  • Step 1: Add phase detection logic at epoch boundary:

    let phase = if epoch < dt_pretrain_epochs { Phase::BehavioralCloning }
                else if epoch < total_epochs * 80 / 100 { Phase::OnlineRL }
                else { Phase::Refinement };
    
  • Step 2: Configure each auxiliary head's weight based on phase

  • Step 3: Configure exploration strategy based on phase (epsilon → noisy → ensemble)

  • Step 4: Configure expert demo ratio based on phase (high → decay → zero)

  • Step 5: Commit: "feat: three-phase training pipeline — BC → Online RL → Refinement"


Task 11: Hyperopt Search Space for All Features

Files:

  • Modify: crates/ml/src/hyperopt/adapters/dqn.rs

Add all new config fields to the hyperopt search space so PSO can optimize them:

  • Step 1: Add her_ratio to search space: range [0.0, 0.5]
  • Step 2: Add curiosity_weight to search space: range [0.0, 0.2]
  • Step 3: Add cql_alpha to search space: range [0.0, 1.0]
  • Step 4: Add ensemble_count to search space: range [1, 5] (integer)
  • Step 5: Add dt_pretrain_epochs to search space: range [0, 10]
  • Step 6: Add base_trail_distance to search space: range [0.002, 0.02]
  • Step 7: Add patience_multiplier_scale to search space: range [0.5, 3.0]
  • Step 8: Add dense_reward_weight to search space: range [0.01, 0.5]
  • Step 9: Add sparse_reward_weight to search space: range [0.5, 5.0]
  • Step 10: Update from_continuous(), to_continuous(), param_names(), continuous_bounds()
  • Step 11: Update dimension assertions in tests
  • Step 12: Commit: "feat: full hyperopt search space — all 25 features searchable"

Task 12: Smoke Test Validation — Winning Trades

Files:

  • Modify: crates/ml/src/trainers/dqn/smoke_tests/training_stability.rs

The ultimate test: does the model produce winning trades on real data?

  • Step 1: In test_50_epoch_convergence, assert:
    // With 50 epochs of training on real ES data, the model MUST have some winning trades
    let wins = metrics.additional_metrics.get("winning_trades").copied().unwrap_or(0.0);
    let total = metrics.additional_metrics.get("total_trades").copied().unwrap_or(0.0);
    assert!(total > 50.0, "Must execute >50 trades over 50 epochs");
    assert!(wins > 0.0, "Must have at least 1 winning trade — 0 wins is a bug");
    let win_rate = wins / total.max(1.0);
    assert!(win_rate > 0.1, "Win rate must be >10% — random is 50%. Got {:.1}%", win_rate * 100.0);
    
  • Step 2: In test_trading_model_behavior, assert:
    // Even in 3 epochs, the dynamic trailing stop should lock in some winners
    let wins = metrics.additional_metrics.get("winning_trades").copied().unwrap_or(0.0);
    assert!(wins > 0.0, "Even 3 epochs must produce at least 1 winning trade");
    
  • Step 3: Run all smoke tests: FOXHUNT_TEST_DATA=... cargo test -p ml --lib -- smoke_tests --ignored
  • Step 4: Commit: "test: assert winning trades — the model must learn to profit"

Key Design Decisions

Why Trade-Aware Reward Instead of Pure Per-Bar?

Per-bar mark-to-market (reward v4) gives identical signal for a 1-bar scalp and a 50-bar trend-follow. The model has no reason to prefer multi-bar holding. Trade-aware reward solves this by:

  1. Dense component: unrealized P&L change per bar (not portfolio return) — specifically about THIS trade
  2. Sparse component: trade return × patience bonus — rewards holding through noise

Why Dynamic Trailing Stop Instead of Learned Exits?

The model CAN learn exit timing (it sees hold_time, unrealized P&L, ADX in its state). But learning exits requires 10,000+ trades of experience. The trailing stop provides a SAFETY NET while the model develops exit skill:

  • Locks in profits when a trade reverses (prevents giving back winners)
  • Width adapts to regime (tight in ranging, wide in trending)
  • The model learns to ENTER positions that work WITH the trailing stop

As the model improves, the trailing stop triggers less often (the model exits BEFORE the trail). The trail becomes insurance, not the primary exit mechanism.

Why All 25 Features at Once?

Ablation studies can happen on H100. The smoke test validates the PIPELINE works. Hyperopt on H100 finds the optimal CONFIG. We don't know which features the ES futures market rewards — that's hyperopt's job. Our job is to give hyperopt the maximum search space.

Why Three Phases?

Phase 1 (BC + DT) bootstraps from expert knowledge — the model starts better than random. Phase 2 (Online RL) refines from live experience — the model discovers patterns experts can't encode. Phase 3 (Refinement) consolidates — shrink-and-perturb prevents dead neurons, ensemble exploration prevents local optima.


Execution Order

Tasks 1-3 are the foundation (reward + configs + Kelly). Must be done first. Tasks 4-8 activate auxiliary heads (can be parallelized). Task 9 wires ensemble consensus (depends on Task 7). Task 10 orchestrates the pipeline (depends on all above). Task 11 enables hyperopt (depends on all above). Task 12 validates everything (last).

Task 1 (reward) ──┐
Task 2 (configs) ──┤
Task 3 (Kelly) ────┤
                   ├──→ Tasks 4-8 (parallel) ──→ Task 9 ──→ Task 10 ──→ Task 11 ──→ Task 12