Files
foxhunt/WAVE10_REWARD_SYSTEM_REDESIGN_PROPOSAL.md
jgrusewski 00ef9e2866 Wave 15: Complete FactoredAction migration to 45-action system
Major Changes:
- Migrated from 3-action TradingAction to 45-action FactoredAction
- 45 actions: 5 exposure × 3 order types × 3 urgency levels
- Absolute exposure model (target positions -1.0 to +1.0)
- Transaction cost differentiation (Market 0.15%, LimitMaker 0.05%, IoC 0.10%)
- Fixed action diversity threshold (1.11% → 0.5% for 45-action space)

Bug Fixes:
- Bug #15: Incomplete FactoredAction integration (code existed but unused)
- Bug #16: Runtime crash in action diversity checking (hardcoded 3-action match)

Code Changes (13 files, ~464 lines):
- ml/src/dqn/action_space.rs: Core FactoredAction + 4 helper methods
- ml/src/trainers/dqn.rs: Action diversity refactored (3→45 dynamic)
- ml/src/dqn/reward.rs: calculate_reward() signature updated
- ml/src/dqn/portfolio_tracker.rs: execute_action() absolute exposure
- ml/src/dqn/dqn.rs: WorkingDQN action selection migrated
- ml/tests/*.rs: 9 test files updated with FactoredAction assertions

Test Results:
- 1-epoch smoke test: 100% action diversity (45/45 actions, 80.2s)
- 10-epoch production: 87.8% readiness (79/90 scorecard, 14.0 min)
- Loss convergence: 96.9% reduction (119K → 3.6K)
- Action diversity: 100% → 44% (healthy specialization)
- Checkpoint reliability: 12/12 files saved (100%)
- DQN tests: 195/195 passing (100%)
- ML baseline: 1,514/1,515 passing (99.93%)

Production Status:  CERTIFIED (87.8% readiness)
Go/No-Go:  GO FOR 100-EPOCH PRODUCTION TRAINING

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
2025-11-11 23:27:02 +01:00

24 KiB
Raw Blame History

Wave 10 DQN Reward System Redesign - Elite-Tier Proposal

Date: 2025-11-08 Status: 🔴 CRITICAL - Action diversity collapse detected (100% HOLD actions) Target: Elite-tier HFT performance with robust action diversity


1. Executive Summary

Problem: Wave 10 production model (Epoch 100) exhibits complete action diversity collapse on validation data:

  • BUY: 0 actions (0.0%)
  • SELL: 0 actions (0.0%)
  • HOLD: 20,480 actions (100.0%)
  • Q-values: HOLD=234.82, BUY=0.0, SELL=0.0

Root Cause: Current reward system over-penalizes active trading, leading to learned passivity.

Proposed Solution: Multi-component elite-tier reward system combining:

  1. Intrinsic Reward Shaping (AIRS): Adaptive exploration bonuses
  2. Entropy Regularization: Policy diversity maintenance
  3. Multi-Objective Optimization: Balanced Sharpe/activity/drawdown
  4. Curiosity-Driven Exploration: Novelty-based intrinsic rewards
  5. Ensemble Model Fusion: Leverage existing Transformer/LSTM/PPO models

2. Current System Analysis

2.1 Current Reward Function

Location: ml/src/dqn/reward.rs (lines 50-150, estimated)

Current Implementation (inferred from training logs):

fn calculate_reward(
    &self,
    position: Position,
    entry_price: f64,
    exit_price: f64,
    action: Action,
) -> f64 {
    let pnl = match (position, action) {
        (Position::Long, Action::Sell) => exit_price - entry_price,
        (Position::Short, Action::Buy) => entry_price - exit_price,
        _ => 0.0,
    };

    let hold_penalty = if action == Action::Hold {
        -0.01 * self.hold_penalty_weight  // Current: -0.01 * 3.747 = -0.037
    } else {
        0.0
    };

    pnl + hold_penalty
}

Problem Diagnosis:

  1. Binary reward structure: Only rewards closed trades (P&L), ignores unrealized gains
  2. Weak hold penalty: -0.037 insufficient to overcome learned risk aversion
  3. No exploration incentives: No intrinsic rewards for action diversity
  4. No entropy term: Policy collapse not penalized
  5. Single objective: Only optimizes P&L, ignores Sharpe/drawdown/activity

2.2 Q-Value Collapse Analysis

Training Epoch 95 vs Epoch 100:

Metric Epoch 95 Epoch 100 Change
BUY % 45.2% 1.7% -96.2%
SELL % 9.6% 2.1% -78.1%
HOLD % 45.2% 96.2% +112.8%
Validation Loss 20,630 20,643 +0.06%
Avg Q-value ~150 166.5 +11.0%

Hypothesis: Model learned that:

  1. HOLD actions avoid negative rewards (no hold penalty strong enough)
  2. Active trading (BUY/SELL) risks negative P&L
  3. Safe policy (all HOLD) maximizes expected return
  4. Validation loss stabilized → exploitation phase → diversity collapse

3. Elite-Tier Reward System Design

3.1 Multi-Component Reward Function

Mathematical Formulation:

R_total(s, a, s') = α₁·R_extrinsic(s, a, s')
                  + α₂·R_intrinsic(s, a, s')
                  + α₃·R_entropy(π)
                  + α₄·R_curiosity(s, s')
                  + α₅·R_ensemble(s, a)

Component Weights (adaptive):

  • α₁ = 0.40 (Extrinsic: P&L, Sharpe, drawdown)
  • α₂ = 0.25 (Intrinsic: Action diversity, exploration)
  • α₃ = 0.15 (Entropy: Policy stochasticity)
  • α₄ = 0.10 (Curiosity: State novelty)
  • α₅ = 0.10 (Ensemble: Model agreement/disagreement bonus)

3.2 Component Specifications

Component 1: Enhanced Extrinsic Reward

fn calculate_extrinsic_reward(
    &self,
    position: &Position,
    entry_price: f64,
    exit_price: f64,
    action: Action,
    portfolio_value: f64,
    max_drawdown: f64,
) -> f64 {
    // P&L component (40% weight)
    let pnl = self.calculate_pnl(position, entry_price, exit_price, action);
    let pnl_normalized = pnl / portfolio_value;  // Normalize by portfolio size

    // Sharpe ratio component (30% weight) - rolling 100-bar window
    let returns = self.returns_buffer.push(pnl_normalized);
    let sharpe = self.calculate_rolling_sharpe(&returns, window=100);

    // Drawdown penalty (20% weight)
    let dd_penalty = -max_drawdown.abs() * 10.0;  // Heavy penalty for large drawdowns

    // Activity incentive (10% weight) - reward non-HOLD actions
    let activity_bonus = if action != Action::Hold {
        0.05  // Fixed bonus for active trading
    } else {
        -0.10  // Stronger hold penalty (10x current)
    };

    0.40 * pnl_normalized
        + 0.30 * sharpe
        + 0.20 * dd_penalty
        + 0.10 * activity_bonus
}

Key Improvements:

  • Multi-objective: Balances P&L, Sharpe, drawdown, activity
  • Normalized P&L: Relative to portfolio size (scale-invariant)
  • Rolling Sharpe: Rewards consistent returns, not just total P&L
  • 10x stronger hold penalty: -0.10 vs current -0.01

Component 2: Intrinsic Reward (AIRS-Inspired)

struct IntrinsicRewardModule {
    action_counts: HashMap<Action, u64>,  // Track action distribution
    target_buy_ratio: f64,   // Target: 40-50%
    target_sell_ratio: f64,  // Target: 10-15%
    target_hold_ratio: f64,  // Target: 35-50%
}

fn calculate_intrinsic_reward(
    &mut self,
    action: Action,
    episode_step: u64,
) -> f64 {
    // Update action counts
    *self.action_counts.entry(action).or_insert(0) += 1;
    let total_actions = self.action_counts.values().sum::<u64>() as f64;

    // Current action distribution
    let buy_ratio = self.action_counts[&Action::Buy] as f64 / total_actions;
    let sell_ratio = self.action_counts[&Action::Sell] as f64 / total_actions;
    let hold_ratio = self.action_counts[&Action::Hold] as f64 / total_actions;

    // Diversity bonus: Reward actions that move distribution toward target
    let diversity_bonus = match action {
        Action::Buy => {
            if buy_ratio < self.target_buy_ratio {
                (self.target_buy_ratio - buy_ratio) * 2.0  // Stronger for underrepresented
            } else {
                0.0
            }
        },
        Action::Sell => {
            if sell_ratio < self.target_sell_ratio {
                (self.target_sell_ratio - sell_ratio) * 2.0
            } else {
                0.0
            }
        },
        Action::Hold => {
            // Penalize HOLD if overrepresented
            if hold_ratio > self.target_hold_ratio {
                -(hold_ratio - self.target_hold_ratio) * 5.0  // Heavy penalty
            } else {
                0.0
            }
        },
    };

    // Exploration bonus (decays over time)
    let exploration_bonus = (1.0 / (1.0 + episode_step as f64 / 1000.0)) * 0.5;

    diversity_bonus + exploration_bonus
}

Key Features:

  • Adaptive diversity bonuses: Rewards underrepresented actions
  • Heavy HOLD penalty: 5x multiplier when HOLD exceeds 50%
  • Time-decaying exploration: Strong early, weak late
  • Target ratios: BUY 40-50%, SELL 10-15%, HOLD 35-50%

Component 3: Entropy Regularization

fn calculate_entropy_bonus(
    &self,
    q_values: &Tensor,  // [batch_size, num_actions]
) -> f64 {
    // Convert Q-values to action probabilities via softmax
    let action_probs = q_values.softmax(-1, Kind::Float);  // Shape: [batch_size, 3]

    // Calculate Shannon entropy: H(π) = -Σ π(a|s) * log(π(a|s))
    let log_probs = action_probs.log();
    let entropy = -(action_probs * log_probs).sum(Kind::Float);  // Shape: [batch_size]

    // Average entropy across batch
    let avg_entropy = entropy.mean(Kind::Float).double_value(&[]);

    // Entropy bonus: Reward high entropy (stochastic policies)
    // Maximum entropy for 3 actions: log(3) ≈ 1.099
    // Normalize to [0, 1] and scale
    let normalized_entropy = avg_entropy / 1.099;

    // Strong bonus for entropy > 0.7 (diverse policy)
    if normalized_entropy > 0.7 {
        normalized_entropy * 2.0
    } else {
        // Penalty for low entropy (deterministic policy)
        -(0.7 - normalized_entropy) * 3.0
    }
}

Key Features:

  • Softmax Q-values: Converts Q-values to stochastic policy
  • Shannon entropy: Measures policy diversity
  • Normalized bonus: 2x bonus for high entropy, 3x penalty for low
  • Threshold: 0.7 normalized entropy (diverse vs deterministic)

Component 4: Curiosity-Driven Exploration

struct CuriosityModule {
    state_embeddings: Vec<Tensor>,  // Historical state embeddings
    forward_model: ForwardDynamicsModel,  // Predicts s_{t+1} from (s_t, a_t)
}

fn calculate_curiosity_reward(
    &mut self,
    state: &Tensor,
    action: Action,
    next_state: &Tensor,
) -> f64 {
    // Encode states to embeddings (use first 32 features)
    let state_embedding = state.narrow(1, 0, 32);  // Shape: [batch, 32]
    let next_state_embedding = next_state.narrow(1, 0, 32);

    // Forward model prediction
    let predicted_next_state = self.forward_model.predict(state, action);

    // Prediction error = novelty/surprise
    let prediction_error = (predicted_next_state - next_state_embedding)
        .pow_tensor_scalar(2)
        .mean(Kind::Float)
        .double_value(&[]);

    // Novelty bonus: Reward exploration of novel states
    // Clip to prevent excessive rewards for noisy states
    let novelty_bonus = prediction_error.clamp(0.0, 5.0);

    // Update forward model (online learning)
    self.forward_model.train_step(state, action, next_state_embedding);

    novelty_bonus
}

// Simple forward dynamics model (2-layer MLP)
struct ForwardDynamicsModel {
    fc1: nn::Linear,  // 32 + 3 (action one-hot) → 64
    fc2: nn::Linear,  // 64 → 32
}

impl ForwardDynamicsModel {
    fn predict(&self, state: &Tensor, action: Action) -> Tensor {
        // One-hot encode action
        let action_onehot = Tensor::zeros(&[state.size()[0], 3], (Kind::Float, state.device()));
        action_onehot.narrow(1, action as i64, 1).fill_(1.0);

        // Concatenate state + action
        let input = Tensor::cat(&[state.narrow(1, 0, 32), action_onehot], 1);

        // Forward pass
        input.apply(&self.fc1).relu().apply(&self.fc2)
    }

    fn train_step(&mut self, state: &Tensor, action: Action, target: Tensor) {
        // SGD update with MSE loss
        let pred = self.predict(state, action);
        let loss = (pred - target).pow_tensor_scalar(2).mean(Kind::Float);
        loss.backward();
        // Optimizer step (Adam, lr=1e-4)
    }
}

Key Features:

  • Forward dynamics model: Learns to predict next state
  • Prediction error as novelty: High error = novel/surprising state
  • Online learning: Forward model updates during training
  • Clipped rewards: Prevents noise exploitation (max 5.0)

Component 5: Ensemble Model Fusion

struct EnsembleOracle {
    transformer: Arc<TransformerModel>,  // ml/src/transformers/
    lstm: Arc<LSTMModel>,                // ml/src/lstm/
    ppo: Arc<PPOPolicy>,                 // ml/src/ppo/
}

fn calculate_ensemble_reward(
    &self,
    state: &Tensor,
    dqn_action: Action,
) -> f64 {
    // Get predictions from all models
    let transformer_pred = self.transformer.predict(state);  // Returns action probabilities
    let lstm_pred = self.lstm.predict(state);
    let ppo_pred = self.ppo.predict(state);

    // Convert to action selections
    let transformer_action = transformer_pred.argmax(-1, false);
    let lstm_action = lstm_pred.argmax(-1, false);
    let ppo_action = ppo_pred.argmax(-1, false);

    // Agreement bonus: Reward when DQN agrees with ensemble majority
    let votes = vec![
        transformer_action.int64_value(&[0]) as usize,
        lstm_action.int64_value(&[0]) as usize,
        ppo_action.int64_value(&[0]) as usize,
    ];

    let mut vote_counts = HashMap::new();
    for vote in votes {
        *vote_counts.entry(vote).or_insert(0) += 1;
    }

    let majority_action = *vote_counts.iter().max_by_key(|(_, count)| *count).unwrap().0;

    // Agreement bonus
    let agreement_bonus = if dqn_action as usize == majority_action {
        0.5  // Strong bonus for ensemble agreement
    } else {
        // Small bonus for disagreement (exploration value)
        0.1
    };

    // Diversity bonus: Reward when models disagree (indicates uncertainty)
    let num_unique_actions = vote_counts.len();
    let diversity_bonus = match num_unique_actions {
        3 => 0.3,  // All models disagree (high uncertainty)
        2 => 0.1,  // Moderate disagreement
        1 => 0.0,  // Full agreement (low uncertainty)
        _ => 0.0,
    };

    agreement_bonus + diversity_bonus
}

Key Features:

  • Multi-model oracle: Leverages Transformer, LSTM, PPO predictions
  • Majority voting: Identifies consensus action
  • Agreement bonus: Rewards DQN for aligning with ensemble
  • Diversity bonus: Rewards exploration in high-uncertainty states

4. Implementation Plan

4.1 File Structure

ml/src/dqn/
├── reward.rs                  # Current reward implementation
├── reward_elite.rs            # NEW: Elite-tier multi-component reward
├── intrinsic_rewards.rs       # NEW: AIRS-inspired intrinsic rewards
├── curiosity.rs               # NEW: Forward dynamics model
├── ensemble_oracle.rs         # NEW: Multi-model ensemble fusion
└── portfolio_tracker.rs       # Existing: Portfolio state tracking

4.2 Phase 1: Core Reward Redesign (Week 1)

Goal: Implement enhanced extrinsic + intrinsic rewards

Tasks:

  1. Create reward_elite.rs with multi-component reward function
  2. Implement IntrinsicRewardModule with action diversity tracking
  3. Add rolling Sharpe ratio calculation (100-bar window)
  4. Integrate with existing PortfolioTracker
  5. Add unit tests (20+ test cases)

Files Modified:

  • ml/src/dqn/reward_elite.rs (NEW, ~400 lines)
  • ml/src/dqn/intrinsic_rewards.rs (NEW, ~200 lines)
  • ml/src/dqn/mod.rs (add module exports)
  • ml/src/trainers/dqn.rs (integrate new reward function)

Test Coverage:

#[cfg(test)]
mod tests {
    #[test]
    fn test_extrinsic_reward_long_profit() { ... }

    #[test]
    fn test_extrinsic_reward_short_profit() { ... }

    #[test]
    fn test_intrinsic_diversity_bonus() { ... }

    #[test]
    fn test_intrinsic_hold_penalty() { ... }

    #[test]
    fn test_rolling_sharpe_calculation() { ... }

    #[test]
    fn test_adaptive_weight_scaling() { ... }

    // ... 15+ more test cases
}

4.3 Phase 2: Entropy Regularization (Week 2)

Goal: Add policy entropy bonus to prevent collapse

Tasks:

  1. Implement calculate_entropy_bonus() in reward_elite.rs
  2. Modify Q-value selection to use softmax (currently argmax)
  3. Add entropy tracking to training logs
  4. Add entropy visualization to TensorBoard

Files Modified:

  • ml/src/dqn/reward_elite.rs (add entropy module)
  • ml/src/dqn/dqn.rs (modify action selection)
  • ml/src/trainers/dqn.rs (add entropy logging)

Expected Impact:

  • Current: Deterministic policy (entropy ≈ 0)
  • Target: Stochastic policy (entropy > 0.7 × log(3) = 0.77)
  • Action diversity: HOLD < 50%, BUY > 30%, SELL > 10%

4.4 Phase 3: Curiosity-Driven Exploration (Week 3)

Goal: Add forward dynamics model for novelty detection

Tasks:

  1. Create curiosity.rs with ForwardDynamicsModel
  2. Implement online learning updates during training
  3. Add state embedding buffer (32 dimensions)
  4. Integrate with main reward function

Files Modified:

  • ml/src/dqn/curiosity.rs (NEW, ~300 lines)
  • ml/src/dqn/reward_elite.rs (integrate curiosity module)
  • ml/src/trainers/dqn.rs (add forward model checkpointing)

Hyperparameters:

CuriosityConfig {
    embedding_dim: 32,           // State embedding size
    hidden_dim: 64,              // Forward model hidden layer
    learning_rate: 1e-4,         // Forward model optimizer
    max_reward: 5.0,             // Clip curiosity reward
    update_frequency: 1,         // Train every step
}

4.5 Phase 4: Ensemble Model Fusion (Week 4)

Goal: Leverage existing Transformer/LSTM/PPO models

Tasks:

  1. Create ensemble_oracle.rs with multi-model interface
  2. Load pre-trained models (Transformer, LSTM, PPO)
  3. Implement majority voting + disagreement bonus
  4. Add ensemble logging to training

Files Modified:

  • ml/src/dqn/ensemble_oracle.rs (NEW, ~250 lines)
  • ml/src/dqn/reward_elite.rs (integrate ensemble module)
  • ml/examples/train_dqn.rs (add --use-ensemble flag)

Model Loading:

// Load pre-trained models from trained_models/
let transformer = TransformerModel::load("ml/trained_models/tft_best_model.safetensors")?;
let lstm = LSTMModel::load("ml/trained_models/mamba2_best_model.safetensors")?;
let ppo = PPOPolicy::load("ml/trained_models/ppo_best_model.safetensors")?;

let ensemble = EnsembleOracle {
    transformer: Arc::new(transformer),
    lstm: Arc::new(lstm),
    ppo: Arc::new(ppo),
};

Expected Impact:

  • Consensus signals: Higher confidence trades
  • Disagreement signals: Exploration opportunities
  • Multi-strategy fusion: Robustness to regime changes

4.6 Phase 5: Validation & Hyperopt (Week 5)

Goal: Validate new reward system and tune component weights

Tasks:

  1. Retrain DQN with elite reward system (50 epochs)
  2. Run validation backtest on unseen data
  3. Launch hyperopt campaign (30 trials) to tune α₁-α₅ weights
  4. Compare against Wave 10 baseline

Hyperopt Search Space:

HyperoptSpace {
    alpha_extrinsic: (0.30, 0.50),      // α₁
    alpha_intrinsic: (0.15, 0.35),      // α₂
    alpha_entropy: (0.10, 0.25),        // α₃
    alpha_curiosity: (0.05, 0.15),      // α₄
    alpha_ensemble: (0.05, 0.15),       // α₅

    // Constraint: Σ αᵢ = 1.0
}

Success Criteria:

  • Action diversity: BUY > 30%, SELL > 10%, HOLD < 50%
  • Validation Sharpe > 1.5
  • Q-value diversity: σ(Q) > 10.0
  • No collapse over 100 epochs

5. Expected Outcomes

5.1 Performance Metrics

Baseline (Wave 10, Epoch 100):

Action Distribution:
  BUY:   0.0%  (0 actions)
  SELL:  0.0%  (0 actions)
  HOLD: 100.0% (20,480 actions)

Q-Values:
  BUY:   0.0
  SELL:  0.0
  HOLD: 234.82

Sharpe Ratio: N/A (no trades)
Win Rate: N/A
Drawdown: N/A

Target (Elite Reward System):

Action Distribution:
  BUY:  40-50%  (8,192-10,240 actions)
  SELL: 10-15%  (2,048-3,072 actions)
  HOLD: 35-50%  (7,168-10,240 actions)

Q-Values:
  BUY:   180-220
  SELL:  170-200
  HOLD:  160-190
  σ(Q):  > 10.0 (diversity)

Sharpe Ratio: > 2.0
Win Rate: > 55%
Drawdown: < 20%

5.2 Training Dynamics

Expected Changes:

  1. Epoch 0-20 (Exploration): High entropy (>0.8), diverse actions
  2. Epoch 20-50 (Learning): Sharpe improves, entropy stabilizes (0.7-0.8)
  3. Epoch 50-100 (Refinement): Stable action distribution, no collapse
  4. Epoch 100+ (Validation): Maintains diversity on unseen data

Monitoring:

  • Track entropy every epoch (target: > 0.7)
  • Track action distribution every 10 epochs (target: BUY 40-50%)
  • Track Q-value standard deviation (target: > 10.0)
  • Early stopping if entropy < 0.5 for 5 consecutive epochs

5.3 Cost Estimates

Development Time:

  • Phase 1 (Core): 3-4 days
  • Phase 2 (Entropy): 2-3 days
  • Phase 3 (Curiosity): 3-4 days
  • Phase 4 (Ensemble): 2-3 days
  • Phase 5 (Validation): 2-3 days Total: 12-17 days (~3-4 weeks)

GPU Compute:

  • Retraining (50 epochs): ~6 minutes (RTX 3050 Ti)
  • Hyperopt (30 trials): ~3 hours (RTX A4000, $0.75)
  • Validation backtests: ~5 minutes total

Expected ROI:

  • Development cost: ~$2,400-$3,400 (17 days × $20/hr)
  • Performance gain: +2.0 Sharpe vs 0.0 baseline = INFINITE ROI
  • Break-even: First successful trade

6. Risk Analysis

6.1 Technical Risks

Risk Likelihood Impact Mitigation
Reward complexity slows training MEDIUM MEDIUM Start with Phase 1-2 only, add components incrementally
Ensemble overhead (inference latency) LOW MEDIUM Cache model predictions, use only during training
Hyperparameter tuning difficulty HIGH HIGH Use Optuna, 30+ trials, conservative priors
Overfitting to intrinsic rewards MEDIUM HIGH Cap intrinsic component at 25% total reward
Forward model instability MEDIUM MEDIUM Clip gradients, small learning rate (1e-4)

6.2 Fallback Plans

If Phase 1-2 fail to improve diversity:

  • Revert to simple multi-objective (Sharpe + activity + entropy)
  • Increase hold penalty from -0.10 to -0.50
  • Use epsilon-greedy with ε=0.2 during validation

If ensemble overhead too high:

  • Use ensemble only during training, disable for inference
  • Sample ensemble predictions (e.g., every 10 steps)
  • Use lightweight models (LSTM only, skip Transformer)

If hyperopt finds poor parameters:

  • Manual tuning with grid search
  • Use Wave 10 parameters as baseline, modify reward only
  • Consider transfer learning from Wave 9 model

7. Implementation Checklist

Phase 1: Core Reward Redesign

  • Create ml/src/dqn/reward_elite.rs
  • Implement calculate_extrinsic_reward() with multi-objective
  • Implement IntrinsicRewardModule with action diversity
  • Add rolling Sharpe ratio calculation
  • Write 20+ unit tests
  • Integrate with trainers/dqn.rs
  • Run smoke test (5 epochs, verify no crashes)

Phase 2: Entropy Regularization

  • Implement calculate_entropy_bonus()
  • Modify Q-value selection to softmax
  • Add entropy tracking to logs
  • Add TensorBoard entropy visualization
  • Test on 10-epoch training run
  • Verify entropy > 0.7

Phase 3: Curiosity-Driven Exploration

  • Create ml/src/dqn/curiosity.rs
  • Implement ForwardDynamicsModel (2-layer MLP)
  • Add online learning updates
  • Test forward model convergence
  • Integrate with reward function
  • Run 20-epoch validation

Phase 4: Ensemble Model Fusion

  • Create ml/src/dqn/ensemble_oracle.rs
  • Load pre-trained Transformer/LSTM/PPO models
  • Implement majority voting
  • Add disagreement bonus
  • Test inference latency (target: < 500μs)
  • Run 10-epoch training with ensemble

Phase 5: Validation & Hyperopt

  • Retrain DQN with elite reward (50 epochs)
  • Run validation backtest (unseen data)
  • Launch hyperopt campaign (30 trials, tune α₁-α₅)
  • Compare against Wave 10 baseline
  • Document final parameters
  • Update CLAUDE.md with results
  • Commit production model

8. Success Criteria

Definition of Success (ALL must be met):

  1. Action Diversity: BUY > 30%, SELL > 10%, HOLD < 50% on validation data
  2. Q-Value Diversity: Standard deviation σ(Q) > 10.0 (no collapse)
  3. Sharpe Ratio: > 1.5 on validation backtest (2.0 stretch goal)
  4. Training Stability: No entropy collapse over 100 epochs (entropy > 0.7)
  5. Inference Latency: < 500μs per action (ensemble overhead acceptable)
  6. Test Coverage: 100% pass rate (all existing + new tests)

Go/No-Go Decision:

  • 5-6 criteria met: PROCEED TO PRODUCTION
  • ⚠️ 3-4 criteria met: ITERATE (1-2 more cycles)
  • 0-2 criteria met: FALLBACK (Revert to Wave 9 + manual tuning)

9. References

2025 State-of-the-Art Research

  1. Potential-Based Reward Shaping: Ng et al. (1999), revisited with linear shifts (2024-2025 papers)
  2. AIRS (Automatic Intrinsic Reward Shaping): Adaptive intrinsic reward selection
  3. Entropy Regularization: Maximum entropy RL for robust policies
  4. Curiosity-Driven Exploration: ICM (Intrinsic Curiosity Module), Pathak et al. (2017), modern variants (2024-2025)
  5. Multi-Objective RL: Pareto optimization for trading (Sharpe/profit/drawdown balance)
  6. Ensemble Model Fusion: Transformer + LSTM + RL hybrid architectures (2024-2025)

Internal Documentation

  • CLAUDE.md: System architecture, Wave 10 campaign results
  • ml/src/dqn/reward.rs: Current reward implementation (baseline)
  • ml/src/dqn/portfolio_tracker.rs: Portfolio state tracking (218 lines, 9/9 tests)
  • ml/src/hyperopt/adapters/dqn.rs: Hyperopt integration
  • /tmp/ml_training/wave10_production/WAVE10_FINAL_CAMPAIGN_REPORT.md: 476-line analysis

10. Approval and Next Steps

Recommended Action:

  1. IMMEDIATE: Review this proposal with user
  2. Short-term (Week 1): Implement Phase 1-2 (core + entropy)
  3. Medium-term (Week 2-3): Implement Phase 3-4 (curiosity + ensemble)
  4. Long-term (Week 4-5): Hyperopt campaign and production deployment

Required Approvals:

  • Technical design review (user approval)
  • Resource allocation (3-4 weeks dev time)
  • GPU budget ($0.75 for hyperopt)

Contact: @user for questions/feedback


Document Status: READY FOR REVIEW Version: 1.0 Last Updated: 2025-11-08