Major Changes: - Migrated from 3-action TradingAction to 45-action FactoredAction - 45 actions: 5 exposure × 3 order types × 3 urgency levels - Absolute exposure model (target positions -1.0 to +1.0) - Transaction cost differentiation (Market 0.15%, LimitMaker 0.05%, IoC 0.10%) - Fixed action diversity threshold (1.11% → 0.5% for 45-action space) Bug Fixes: - Bug #15: Incomplete FactoredAction integration (code existed but unused) - Bug #16: Runtime crash in action diversity checking (hardcoded 3-action match) Code Changes (13 files, ~464 lines): - ml/src/dqn/action_space.rs: Core FactoredAction + 4 helper methods - ml/src/trainers/dqn.rs: Action diversity refactored (3→45 dynamic) - ml/src/dqn/reward.rs: calculate_reward() signature updated - ml/src/dqn/portfolio_tracker.rs: execute_action() absolute exposure - ml/src/dqn/dqn.rs: WorkingDQN action selection migrated - ml/tests/*.rs: 9 test files updated with FactoredAction assertions Test Results: - 1-epoch smoke test: 100% action diversity (45/45 actions, 80.2s) - 10-epoch production: 87.8% readiness (79/90 scorecard, 14.0 min) - Loss convergence: 96.9% reduction (119K → 3.6K) - Action diversity: 100% → 44% (healthy specialization) - Checkpoint reliability: 12/12 files saved (100%) - DQN tests: 195/195 passing (100%) - ML baseline: 1,514/1,515 passing (99.93%) Production Status: ✅ CERTIFIED (87.8% readiness) Go/No-Go: ✅ GO FOR 100-EPOCH PRODUCTION TRAINING 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com>
766 lines
24 KiB
Markdown
766 lines
24 KiB
Markdown
# Wave 10 DQN Reward System Redesign - Elite-Tier Proposal
|
||
|
||
**Date**: 2025-11-08
|
||
**Status**: 🔴 CRITICAL - Action diversity collapse detected (100% HOLD actions)
|
||
**Target**: Elite-tier HFT performance with robust action diversity
|
||
|
||
---
|
||
|
||
## 1. Executive Summary
|
||
|
||
**Problem**: Wave 10 production model (Epoch 100) exhibits complete action diversity collapse on validation data:
|
||
- **BUY**: 0 actions (0.0%)
|
||
- **SELL**: 0 actions (0.0%)
|
||
- **HOLD**: 20,480 actions (100.0%)
|
||
- **Q-values**: HOLD=234.82, BUY=0.0, SELL=0.0
|
||
|
||
**Root Cause**: Current reward system over-penalizes active trading, leading to learned passivity.
|
||
|
||
**Proposed Solution**: Multi-component elite-tier reward system combining:
|
||
1. **Intrinsic Reward Shaping (AIRS)**: Adaptive exploration bonuses
|
||
2. **Entropy Regularization**: Policy diversity maintenance
|
||
3. **Multi-Objective Optimization**: Balanced Sharpe/activity/drawdown
|
||
4. **Curiosity-Driven Exploration**: Novelty-based intrinsic rewards
|
||
5. **Ensemble Model Fusion**: Leverage existing Transformer/LSTM/PPO models
|
||
|
||
---
|
||
|
||
## 2. Current System Analysis
|
||
|
||
### 2.1 Current Reward Function
|
||
|
||
**Location**: `ml/src/dqn/reward.rs` (lines 50-150, estimated)
|
||
|
||
**Current Implementation** (inferred from training logs):
|
||
```rust
|
||
fn calculate_reward(
|
||
&self,
|
||
position: Position,
|
||
entry_price: f64,
|
||
exit_price: f64,
|
||
action: Action,
|
||
) -> f64 {
|
||
let pnl = match (position, action) {
|
||
(Position::Long, Action::Sell) => exit_price - entry_price,
|
||
(Position::Short, Action::Buy) => entry_price - exit_price,
|
||
_ => 0.0,
|
||
};
|
||
|
||
let hold_penalty = if action == Action::Hold {
|
||
-0.01 * self.hold_penalty_weight // Current: -0.01 * 3.747 = -0.037
|
||
} else {
|
||
0.0
|
||
};
|
||
|
||
pnl + hold_penalty
|
||
}
|
||
```
|
||
|
||
**Problem Diagnosis**:
|
||
1. **Binary reward structure**: Only rewards closed trades (P&L), ignores unrealized gains
|
||
2. **Weak hold penalty**: -0.037 insufficient to overcome learned risk aversion
|
||
3. **No exploration incentives**: No intrinsic rewards for action diversity
|
||
4. **No entropy term**: Policy collapse not penalized
|
||
5. **Single objective**: Only optimizes P&L, ignores Sharpe/drawdown/activity
|
||
|
||
### 2.2 Q-Value Collapse Analysis
|
||
|
||
**Training Epoch 95 vs Epoch 100**:
|
||
|
||
| Metric | Epoch 95 | Epoch 100 | Change |
|
||
|--------|----------|-----------|--------|
|
||
| BUY % | 45.2% | 1.7% | **-96.2%** |
|
||
| SELL % | 9.6% | 2.1% | -78.1% |
|
||
| HOLD % | 45.2% | 96.2% | +112.8% |
|
||
| Validation Loss | 20,630 | 20,643 | +0.06% |
|
||
| Avg Q-value | ~150 | 166.5 | +11.0% |
|
||
|
||
**Hypothesis**: Model learned that:
|
||
1. HOLD actions avoid negative rewards (no hold penalty strong enough)
|
||
2. Active trading (BUY/SELL) risks negative P&L
|
||
3. Safe policy (all HOLD) maximizes expected return
|
||
4. Validation loss stabilized → exploitation phase → diversity collapse
|
||
|
||
---
|
||
|
||
## 3. Elite-Tier Reward System Design
|
||
|
||
### 3.1 Multi-Component Reward Function
|
||
|
||
**Mathematical Formulation**:
|
||
|
||
```
|
||
R_total(s, a, s') = α₁·R_extrinsic(s, a, s')
|
||
+ α₂·R_intrinsic(s, a, s')
|
||
+ α₃·R_entropy(π)
|
||
+ α₄·R_curiosity(s, s')
|
||
+ α₅·R_ensemble(s, a)
|
||
```
|
||
|
||
**Component Weights** (adaptive):
|
||
- α₁ = 0.40 (Extrinsic: P&L, Sharpe, drawdown)
|
||
- α₂ = 0.25 (Intrinsic: Action diversity, exploration)
|
||
- α₃ = 0.15 (Entropy: Policy stochasticity)
|
||
- α₄ = 0.10 (Curiosity: State novelty)
|
||
- α₅ = 0.10 (Ensemble: Model agreement/disagreement bonus)
|
||
|
||
### 3.2 Component Specifications
|
||
|
||
#### Component 1: Enhanced Extrinsic Reward
|
||
```rust
|
||
fn calculate_extrinsic_reward(
|
||
&self,
|
||
position: &Position,
|
||
entry_price: f64,
|
||
exit_price: f64,
|
||
action: Action,
|
||
portfolio_value: f64,
|
||
max_drawdown: f64,
|
||
) -> f64 {
|
||
// P&L component (40% weight)
|
||
let pnl = self.calculate_pnl(position, entry_price, exit_price, action);
|
||
let pnl_normalized = pnl / portfolio_value; // Normalize by portfolio size
|
||
|
||
// Sharpe ratio component (30% weight) - rolling 100-bar window
|
||
let returns = self.returns_buffer.push(pnl_normalized);
|
||
let sharpe = self.calculate_rolling_sharpe(&returns, window=100);
|
||
|
||
// Drawdown penalty (20% weight)
|
||
let dd_penalty = -max_drawdown.abs() * 10.0; // Heavy penalty for large drawdowns
|
||
|
||
// Activity incentive (10% weight) - reward non-HOLD actions
|
||
let activity_bonus = if action != Action::Hold {
|
||
0.05 // Fixed bonus for active trading
|
||
} else {
|
||
-0.10 // Stronger hold penalty (10x current)
|
||
};
|
||
|
||
0.40 * pnl_normalized
|
||
+ 0.30 * sharpe
|
||
+ 0.20 * dd_penalty
|
||
+ 0.10 * activity_bonus
|
||
}
|
||
```
|
||
|
||
**Key Improvements**:
|
||
- **Multi-objective**: Balances P&L, Sharpe, drawdown, activity
|
||
- **Normalized P&L**: Relative to portfolio size (scale-invariant)
|
||
- **Rolling Sharpe**: Rewards consistent returns, not just total P&L
|
||
- **10x stronger hold penalty**: -0.10 vs current -0.01
|
||
|
||
#### Component 2: Intrinsic Reward (AIRS-Inspired)
|
||
```rust
|
||
struct IntrinsicRewardModule {
|
||
action_counts: HashMap<Action, u64>, // Track action distribution
|
||
target_buy_ratio: f64, // Target: 40-50%
|
||
target_sell_ratio: f64, // Target: 10-15%
|
||
target_hold_ratio: f64, // Target: 35-50%
|
||
}
|
||
|
||
fn calculate_intrinsic_reward(
|
||
&mut self,
|
||
action: Action,
|
||
episode_step: u64,
|
||
) -> f64 {
|
||
// Update action counts
|
||
*self.action_counts.entry(action).or_insert(0) += 1;
|
||
let total_actions = self.action_counts.values().sum::<u64>() as f64;
|
||
|
||
// Current action distribution
|
||
let buy_ratio = self.action_counts[&Action::Buy] as f64 / total_actions;
|
||
let sell_ratio = self.action_counts[&Action::Sell] as f64 / total_actions;
|
||
let hold_ratio = self.action_counts[&Action::Hold] as f64 / total_actions;
|
||
|
||
// Diversity bonus: Reward actions that move distribution toward target
|
||
let diversity_bonus = match action {
|
||
Action::Buy => {
|
||
if buy_ratio < self.target_buy_ratio {
|
||
(self.target_buy_ratio - buy_ratio) * 2.0 // Stronger for underrepresented
|
||
} else {
|
||
0.0
|
||
}
|
||
},
|
||
Action::Sell => {
|
||
if sell_ratio < self.target_sell_ratio {
|
||
(self.target_sell_ratio - sell_ratio) * 2.0
|
||
} else {
|
||
0.0
|
||
}
|
||
},
|
||
Action::Hold => {
|
||
// Penalize HOLD if overrepresented
|
||
if hold_ratio > self.target_hold_ratio {
|
||
-(hold_ratio - self.target_hold_ratio) * 5.0 // Heavy penalty
|
||
} else {
|
||
0.0
|
||
}
|
||
},
|
||
};
|
||
|
||
// Exploration bonus (decays over time)
|
||
let exploration_bonus = (1.0 / (1.0 + episode_step as f64 / 1000.0)) * 0.5;
|
||
|
||
diversity_bonus + exploration_bonus
|
||
}
|
||
```
|
||
|
||
**Key Features**:
|
||
- **Adaptive diversity bonuses**: Rewards underrepresented actions
|
||
- **Heavy HOLD penalty**: 5x multiplier when HOLD exceeds 50%
|
||
- **Time-decaying exploration**: Strong early, weak late
|
||
- **Target ratios**: BUY 40-50%, SELL 10-15%, HOLD 35-50%
|
||
|
||
#### Component 3: Entropy Regularization
|
||
```rust
|
||
fn calculate_entropy_bonus(
|
||
&self,
|
||
q_values: &Tensor, // [batch_size, num_actions]
|
||
) -> f64 {
|
||
// Convert Q-values to action probabilities via softmax
|
||
let action_probs = q_values.softmax(-1, Kind::Float); // Shape: [batch_size, 3]
|
||
|
||
// Calculate Shannon entropy: H(π) = -Σ π(a|s) * log(π(a|s))
|
||
let log_probs = action_probs.log();
|
||
let entropy = -(action_probs * log_probs).sum(Kind::Float); // Shape: [batch_size]
|
||
|
||
// Average entropy across batch
|
||
let avg_entropy = entropy.mean(Kind::Float).double_value(&[]);
|
||
|
||
// Entropy bonus: Reward high entropy (stochastic policies)
|
||
// Maximum entropy for 3 actions: log(3) ≈ 1.099
|
||
// Normalize to [0, 1] and scale
|
||
let normalized_entropy = avg_entropy / 1.099;
|
||
|
||
// Strong bonus for entropy > 0.7 (diverse policy)
|
||
if normalized_entropy > 0.7 {
|
||
normalized_entropy * 2.0
|
||
} else {
|
||
// Penalty for low entropy (deterministic policy)
|
||
-(0.7 - normalized_entropy) * 3.0
|
||
}
|
||
}
|
||
```
|
||
|
||
**Key Features**:
|
||
- **Softmax Q-values**: Converts Q-values to stochastic policy
|
||
- **Shannon entropy**: Measures policy diversity
|
||
- **Normalized bonus**: 2x bonus for high entropy, 3x penalty for low
|
||
- **Threshold**: 0.7 normalized entropy (diverse vs deterministic)
|
||
|
||
#### Component 4: Curiosity-Driven Exploration
|
||
```rust
|
||
struct CuriosityModule {
|
||
state_embeddings: Vec<Tensor>, // Historical state embeddings
|
||
forward_model: ForwardDynamicsModel, // Predicts s_{t+1} from (s_t, a_t)
|
||
}
|
||
|
||
fn calculate_curiosity_reward(
|
||
&mut self,
|
||
state: &Tensor,
|
||
action: Action,
|
||
next_state: &Tensor,
|
||
) -> f64 {
|
||
// Encode states to embeddings (use first 32 features)
|
||
let state_embedding = state.narrow(1, 0, 32); // Shape: [batch, 32]
|
||
let next_state_embedding = next_state.narrow(1, 0, 32);
|
||
|
||
// Forward model prediction
|
||
let predicted_next_state = self.forward_model.predict(state, action);
|
||
|
||
// Prediction error = novelty/surprise
|
||
let prediction_error = (predicted_next_state - next_state_embedding)
|
||
.pow_tensor_scalar(2)
|
||
.mean(Kind::Float)
|
||
.double_value(&[]);
|
||
|
||
// Novelty bonus: Reward exploration of novel states
|
||
// Clip to prevent excessive rewards for noisy states
|
||
let novelty_bonus = prediction_error.clamp(0.0, 5.0);
|
||
|
||
// Update forward model (online learning)
|
||
self.forward_model.train_step(state, action, next_state_embedding);
|
||
|
||
novelty_bonus
|
||
}
|
||
|
||
// Simple forward dynamics model (2-layer MLP)
|
||
struct ForwardDynamicsModel {
|
||
fc1: nn::Linear, // 32 + 3 (action one-hot) → 64
|
||
fc2: nn::Linear, // 64 → 32
|
||
}
|
||
|
||
impl ForwardDynamicsModel {
|
||
fn predict(&self, state: &Tensor, action: Action) -> Tensor {
|
||
// One-hot encode action
|
||
let action_onehot = Tensor::zeros(&[state.size()[0], 3], (Kind::Float, state.device()));
|
||
action_onehot.narrow(1, action as i64, 1).fill_(1.0);
|
||
|
||
// Concatenate state + action
|
||
let input = Tensor::cat(&[state.narrow(1, 0, 32), action_onehot], 1);
|
||
|
||
// Forward pass
|
||
input.apply(&self.fc1).relu().apply(&self.fc2)
|
||
}
|
||
|
||
fn train_step(&mut self, state: &Tensor, action: Action, target: Tensor) {
|
||
// SGD update with MSE loss
|
||
let pred = self.predict(state, action);
|
||
let loss = (pred - target).pow_tensor_scalar(2).mean(Kind::Float);
|
||
loss.backward();
|
||
// Optimizer step (Adam, lr=1e-4)
|
||
}
|
||
}
|
||
```
|
||
|
||
**Key Features**:
|
||
- **Forward dynamics model**: Learns to predict next state
|
||
- **Prediction error as novelty**: High error = novel/surprising state
|
||
- **Online learning**: Forward model updates during training
|
||
- **Clipped rewards**: Prevents noise exploitation (max 5.0)
|
||
|
||
#### Component 5: Ensemble Model Fusion
|
||
```rust
|
||
struct EnsembleOracle {
|
||
transformer: Arc<TransformerModel>, // ml/src/transformers/
|
||
lstm: Arc<LSTMModel>, // ml/src/lstm/
|
||
ppo: Arc<PPOPolicy>, // ml/src/ppo/
|
||
}
|
||
|
||
fn calculate_ensemble_reward(
|
||
&self,
|
||
state: &Tensor,
|
||
dqn_action: Action,
|
||
) -> f64 {
|
||
// Get predictions from all models
|
||
let transformer_pred = self.transformer.predict(state); // Returns action probabilities
|
||
let lstm_pred = self.lstm.predict(state);
|
||
let ppo_pred = self.ppo.predict(state);
|
||
|
||
// Convert to action selections
|
||
let transformer_action = transformer_pred.argmax(-1, false);
|
||
let lstm_action = lstm_pred.argmax(-1, false);
|
||
let ppo_action = ppo_pred.argmax(-1, false);
|
||
|
||
// Agreement bonus: Reward when DQN agrees with ensemble majority
|
||
let votes = vec![
|
||
transformer_action.int64_value(&[0]) as usize,
|
||
lstm_action.int64_value(&[0]) as usize,
|
||
ppo_action.int64_value(&[0]) as usize,
|
||
];
|
||
|
||
let mut vote_counts = HashMap::new();
|
||
for vote in votes {
|
||
*vote_counts.entry(vote).or_insert(0) += 1;
|
||
}
|
||
|
||
let majority_action = *vote_counts.iter().max_by_key(|(_, count)| *count).unwrap().0;
|
||
|
||
// Agreement bonus
|
||
let agreement_bonus = if dqn_action as usize == majority_action {
|
||
0.5 // Strong bonus for ensemble agreement
|
||
} else {
|
||
// Small bonus for disagreement (exploration value)
|
||
0.1
|
||
};
|
||
|
||
// Diversity bonus: Reward when models disagree (indicates uncertainty)
|
||
let num_unique_actions = vote_counts.len();
|
||
let diversity_bonus = match num_unique_actions {
|
||
3 => 0.3, // All models disagree (high uncertainty)
|
||
2 => 0.1, // Moderate disagreement
|
||
1 => 0.0, // Full agreement (low uncertainty)
|
||
_ => 0.0,
|
||
};
|
||
|
||
agreement_bonus + diversity_bonus
|
||
}
|
||
```
|
||
|
||
**Key Features**:
|
||
- **Multi-model oracle**: Leverages Transformer, LSTM, PPO predictions
|
||
- **Majority voting**: Identifies consensus action
|
||
- **Agreement bonus**: Rewards DQN for aligning with ensemble
|
||
- **Diversity bonus**: Rewards exploration in high-uncertainty states
|
||
|
||
---
|
||
|
||
## 4. Implementation Plan
|
||
|
||
### 4.1 File Structure
|
||
|
||
```
|
||
ml/src/dqn/
|
||
├── reward.rs # Current reward implementation
|
||
├── reward_elite.rs # NEW: Elite-tier multi-component reward
|
||
├── intrinsic_rewards.rs # NEW: AIRS-inspired intrinsic rewards
|
||
├── curiosity.rs # NEW: Forward dynamics model
|
||
├── ensemble_oracle.rs # NEW: Multi-model ensemble fusion
|
||
└── portfolio_tracker.rs # Existing: Portfolio state tracking
|
||
```
|
||
|
||
### 4.2 Phase 1: Core Reward Redesign (Week 1)
|
||
|
||
**Goal**: Implement enhanced extrinsic + intrinsic rewards
|
||
|
||
**Tasks**:
|
||
1. Create `reward_elite.rs` with multi-component reward function
|
||
2. Implement `IntrinsicRewardModule` with action diversity tracking
|
||
3. Add rolling Sharpe ratio calculation (100-bar window)
|
||
4. Integrate with existing `PortfolioTracker`
|
||
5. Add unit tests (20+ test cases)
|
||
|
||
**Files Modified**:
|
||
- `ml/src/dqn/reward_elite.rs` (NEW, ~400 lines)
|
||
- `ml/src/dqn/intrinsic_rewards.rs` (NEW, ~200 lines)
|
||
- `ml/src/dqn/mod.rs` (add module exports)
|
||
- `ml/src/trainers/dqn.rs` (integrate new reward function)
|
||
|
||
**Test Coverage**:
|
||
```rust
|
||
#[cfg(test)]
|
||
mod tests {
|
||
#[test]
|
||
fn test_extrinsic_reward_long_profit() { ... }
|
||
|
||
#[test]
|
||
fn test_extrinsic_reward_short_profit() { ... }
|
||
|
||
#[test]
|
||
fn test_intrinsic_diversity_bonus() { ... }
|
||
|
||
#[test]
|
||
fn test_intrinsic_hold_penalty() { ... }
|
||
|
||
#[test]
|
||
fn test_rolling_sharpe_calculation() { ... }
|
||
|
||
#[test]
|
||
fn test_adaptive_weight_scaling() { ... }
|
||
|
||
// ... 15+ more test cases
|
||
}
|
||
```
|
||
|
||
### 4.3 Phase 2: Entropy Regularization (Week 2)
|
||
|
||
**Goal**: Add policy entropy bonus to prevent collapse
|
||
|
||
**Tasks**:
|
||
1. Implement `calculate_entropy_bonus()` in `reward_elite.rs`
|
||
2. Modify Q-value selection to use softmax (currently argmax)
|
||
3. Add entropy tracking to training logs
|
||
4. Add entropy visualization to TensorBoard
|
||
|
||
**Files Modified**:
|
||
- `ml/src/dqn/reward_elite.rs` (add entropy module)
|
||
- `ml/src/dqn/dqn.rs` (modify action selection)
|
||
- `ml/src/trainers/dqn.rs` (add entropy logging)
|
||
|
||
**Expected Impact**:
|
||
- Current: Deterministic policy (entropy ≈ 0)
|
||
- Target: Stochastic policy (entropy > 0.7 × log(3) = 0.77)
|
||
- Action diversity: HOLD < 50%, BUY > 30%, SELL > 10%
|
||
|
||
### 4.4 Phase 3: Curiosity-Driven Exploration (Week 3)
|
||
|
||
**Goal**: Add forward dynamics model for novelty detection
|
||
|
||
**Tasks**:
|
||
1. Create `curiosity.rs` with `ForwardDynamicsModel`
|
||
2. Implement online learning updates during training
|
||
3. Add state embedding buffer (32 dimensions)
|
||
4. Integrate with main reward function
|
||
|
||
**Files Modified**:
|
||
- `ml/src/dqn/curiosity.rs` (NEW, ~300 lines)
|
||
- `ml/src/dqn/reward_elite.rs` (integrate curiosity module)
|
||
- `ml/src/trainers/dqn.rs` (add forward model checkpointing)
|
||
|
||
**Hyperparameters**:
|
||
```rust
|
||
CuriosityConfig {
|
||
embedding_dim: 32, // State embedding size
|
||
hidden_dim: 64, // Forward model hidden layer
|
||
learning_rate: 1e-4, // Forward model optimizer
|
||
max_reward: 5.0, // Clip curiosity reward
|
||
update_frequency: 1, // Train every step
|
||
}
|
||
```
|
||
|
||
### 4.5 Phase 4: Ensemble Model Fusion (Week 4)
|
||
|
||
**Goal**: Leverage existing Transformer/LSTM/PPO models
|
||
|
||
**Tasks**:
|
||
1. Create `ensemble_oracle.rs` with multi-model interface
|
||
2. Load pre-trained models (Transformer, LSTM, PPO)
|
||
3. Implement majority voting + disagreement bonus
|
||
4. Add ensemble logging to training
|
||
|
||
**Files Modified**:
|
||
- `ml/src/dqn/ensemble_oracle.rs` (NEW, ~250 lines)
|
||
- `ml/src/dqn/reward_elite.rs` (integrate ensemble module)
|
||
- `ml/examples/train_dqn.rs` (add --use-ensemble flag)
|
||
|
||
**Model Loading**:
|
||
```rust
|
||
// Load pre-trained models from trained_models/
|
||
let transformer = TransformerModel::load("ml/trained_models/tft_best_model.safetensors")?;
|
||
let lstm = LSTMModel::load("ml/trained_models/mamba2_best_model.safetensors")?;
|
||
let ppo = PPOPolicy::load("ml/trained_models/ppo_best_model.safetensors")?;
|
||
|
||
let ensemble = EnsembleOracle {
|
||
transformer: Arc::new(transformer),
|
||
lstm: Arc::new(lstm),
|
||
ppo: Arc::new(ppo),
|
||
};
|
||
```
|
||
|
||
**Expected Impact**:
|
||
- Consensus signals: Higher confidence trades
|
||
- Disagreement signals: Exploration opportunities
|
||
- Multi-strategy fusion: Robustness to regime changes
|
||
|
||
### 4.6 Phase 5: Validation & Hyperopt (Week 5)
|
||
|
||
**Goal**: Validate new reward system and tune component weights
|
||
|
||
**Tasks**:
|
||
1. Retrain DQN with elite reward system (50 epochs)
|
||
2. Run validation backtest on unseen data
|
||
3. Launch hyperopt campaign (30 trials) to tune α₁-α₅ weights
|
||
4. Compare against Wave 10 baseline
|
||
|
||
**Hyperopt Search Space**:
|
||
```rust
|
||
HyperoptSpace {
|
||
alpha_extrinsic: (0.30, 0.50), // α₁
|
||
alpha_intrinsic: (0.15, 0.35), // α₂
|
||
alpha_entropy: (0.10, 0.25), // α₃
|
||
alpha_curiosity: (0.05, 0.15), // α₄
|
||
alpha_ensemble: (0.05, 0.15), // α₅
|
||
|
||
// Constraint: Σ αᵢ = 1.0
|
||
}
|
||
```
|
||
|
||
**Success Criteria**:
|
||
- ✅ Action diversity: BUY > 30%, SELL > 10%, HOLD < 50%
|
||
- ✅ Validation Sharpe > 1.5
|
||
- ✅ Q-value diversity: σ(Q) > 10.0
|
||
- ✅ No collapse over 100 epochs
|
||
|
||
---
|
||
|
||
## 5. Expected Outcomes
|
||
|
||
### 5.1 Performance Metrics
|
||
|
||
**Baseline (Wave 10, Epoch 100)**:
|
||
```
|
||
Action Distribution:
|
||
BUY: 0.0% (0 actions)
|
||
SELL: 0.0% (0 actions)
|
||
HOLD: 100.0% (20,480 actions)
|
||
|
||
Q-Values:
|
||
BUY: 0.0
|
||
SELL: 0.0
|
||
HOLD: 234.82
|
||
|
||
Sharpe Ratio: N/A (no trades)
|
||
Win Rate: N/A
|
||
Drawdown: N/A
|
||
```
|
||
|
||
**Target (Elite Reward System)**:
|
||
```
|
||
Action Distribution:
|
||
BUY: 40-50% (8,192-10,240 actions)
|
||
SELL: 10-15% (2,048-3,072 actions)
|
||
HOLD: 35-50% (7,168-10,240 actions)
|
||
|
||
Q-Values:
|
||
BUY: 180-220
|
||
SELL: 170-200
|
||
HOLD: 160-190
|
||
σ(Q): > 10.0 (diversity)
|
||
|
||
Sharpe Ratio: > 2.0
|
||
Win Rate: > 55%
|
||
Drawdown: < 20%
|
||
```
|
||
|
||
### 5.2 Training Dynamics
|
||
|
||
**Expected Changes**:
|
||
1. **Epoch 0-20** (Exploration): High entropy (>0.8), diverse actions
|
||
2. **Epoch 20-50** (Learning): Sharpe improves, entropy stabilizes (0.7-0.8)
|
||
3. **Epoch 50-100** (Refinement): Stable action distribution, no collapse
|
||
4. **Epoch 100+** (Validation): Maintains diversity on unseen data
|
||
|
||
**Monitoring**:
|
||
- Track entropy every epoch (target: > 0.7)
|
||
- Track action distribution every 10 epochs (target: BUY 40-50%)
|
||
- Track Q-value standard deviation (target: > 10.0)
|
||
- Early stopping if entropy < 0.5 for 5 consecutive epochs
|
||
|
||
### 5.3 Cost Estimates
|
||
|
||
**Development Time**:
|
||
- Phase 1 (Core): 3-4 days
|
||
- Phase 2 (Entropy): 2-3 days
|
||
- Phase 3 (Curiosity): 3-4 days
|
||
- Phase 4 (Ensemble): 2-3 days
|
||
- Phase 5 (Validation): 2-3 days
|
||
**Total**: 12-17 days (~3-4 weeks)
|
||
|
||
**GPU Compute**:
|
||
- Retraining (50 epochs): ~6 minutes (RTX 3050 Ti)
|
||
- Hyperopt (30 trials): ~3 hours (RTX A4000, $0.75)
|
||
- Validation backtests: ~5 minutes total
|
||
|
||
**Expected ROI**:
|
||
- Development cost: ~$2,400-$3,400 (17 days × $20/hr)
|
||
- Performance gain: +2.0 Sharpe vs 0.0 baseline = **INFINITE ROI**
|
||
- Break-even: First successful trade
|
||
|
||
---
|
||
|
||
## 6. Risk Analysis
|
||
|
||
### 6.1 Technical Risks
|
||
|
||
| Risk | Likelihood | Impact | Mitigation |
|
||
|------|-----------|--------|------------|
|
||
| Reward complexity slows training | MEDIUM | MEDIUM | Start with Phase 1-2 only, add components incrementally |
|
||
| Ensemble overhead (inference latency) | LOW | MEDIUM | Cache model predictions, use only during training |
|
||
| Hyperparameter tuning difficulty | HIGH | HIGH | Use Optuna, 30+ trials, conservative priors |
|
||
| Overfitting to intrinsic rewards | MEDIUM | HIGH | Cap intrinsic component at 25% total reward |
|
||
| Forward model instability | MEDIUM | MEDIUM | Clip gradients, small learning rate (1e-4) |
|
||
|
||
### 6.2 Fallback Plans
|
||
|
||
**If Phase 1-2 fail to improve diversity**:
|
||
- Revert to simple multi-objective (Sharpe + activity + entropy)
|
||
- Increase hold penalty from -0.10 to -0.50
|
||
- Use epsilon-greedy with ε=0.2 during validation
|
||
|
||
**If ensemble overhead too high**:
|
||
- Use ensemble only during training, disable for inference
|
||
- Sample ensemble predictions (e.g., every 10 steps)
|
||
- Use lightweight models (LSTM only, skip Transformer)
|
||
|
||
**If hyperopt finds poor parameters**:
|
||
- Manual tuning with grid search
|
||
- Use Wave 10 parameters as baseline, modify reward only
|
||
- Consider transfer learning from Wave 9 model
|
||
|
||
---
|
||
|
||
## 7. Implementation Checklist
|
||
|
||
### Phase 1: Core Reward Redesign
|
||
- [ ] Create `ml/src/dqn/reward_elite.rs`
|
||
- [ ] Implement `calculate_extrinsic_reward()` with multi-objective
|
||
- [ ] Implement `IntrinsicRewardModule` with action diversity
|
||
- [ ] Add rolling Sharpe ratio calculation
|
||
- [ ] Write 20+ unit tests
|
||
- [ ] Integrate with `trainers/dqn.rs`
|
||
- [ ] Run smoke test (5 epochs, verify no crashes)
|
||
|
||
### Phase 2: Entropy Regularization
|
||
- [ ] Implement `calculate_entropy_bonus()`
|
||
- [ ] Modify Q-value selection to softmax
|
||
- [ ] Add entropy tracking to logs
|
||
- [ ] Add TensorBoard entropy visualization
|
||
- [ ] Test on 10-epoch training run
|
||
- [ ] Verify entropy > 0.7
|
||
|
||
### Phase 3: Curiosity-Driven Exploration
|
||
- [ ] Create `ml/src/dqn/curiosity.rs`
|
||
- [ ] Implement `ForwardDynamicsModel` (2-layer MLP)
|
||
- [ ] Add online learning updates
|
||
- [ ] Test forward model convergence
|
||
- [ ] Integrate with reward function
|
||
- [ ] Run 20-epoch validation
|
||
|
||
### Phase 4: Ensemble Model Fusion
|
||
- [ ] Create `ml/src/dqn/ensemble_oracle.rs`
|
||
- [ ] Load pre-trained Transformer/LSTM/PPO models
|
||
- [ ] Implement majority voting
|
||
- [ ] Add disagreement bonus
|
||
- [ ] Test inference latency (target: < 500μs)
|
||
- [ ] Run 10-epoch training with ensemble
|
||
|
||
### Phase 5: Validation & Hyperopt
|
||
- [ ] Retrain DQN with elite reward (50 epochs)
|
||
- [ ] Run validation backtest (unseen data)
|
||
- [ ] Launch hyperopt campaign (30 trials, tune α₁-α₅)
|
||
- [ ] Compare against Wave 10 baseline
|
||
- [ ] Document final parameters
|
||
- [ ] Update CLAUDE.md with results
|
||
- [ ] Commit production model
|
||
|
||
---
|
||
|
||
## 8. Success Criteria
|
||
|
||
**Definition of Success** (ALL must be met):
|
||
|
||
1. ✅ **Action Diversity**: BUY > 30%, SELL > 10%, HOLD < 50% on validation data
|
||
2. ✅ **Q-Value Diversity**: Standard deviation σ(Q) > 10.0 (no collapse)
|
||
3. ✅ **Sharpe Ratio**: > 1.5 on validation backtest (2.0 stretch goal)
|
||
4. ✅ **Training Stability**: No entropy collapse over 100 epochs (entropy > 0.7)
|
||
5. ✅ **Inference Latency**: < 500μs per action (ensemble overhead acceptable)
|
||
6. ✅ **Test Coverage**: 100% pass rate (all existing + new tests)
|
||
|
||
**Go/No-Go Decision**:
|
||
- ✅ 5-6 criteria met: **PROCEED TO PRODUCTION**
|
||
- ⚠️ 3-4 criteria met: **ITERATE (1-2 more cycles)**
|
||
- ❌ 0-2 criteria met: **FALLBACK (Revert to Wave 9 + manual tuning)**
|
||
|
||
---
|
||
|
||
## 9. References
|
||
|
||
### 2025 State-of-the-Art Research
|
||
|
||
1. **Potential-Based Reward Shaping**: Ng et al. (1999), revisited with linear shifts (2024-2025 papers)
|
||
2. **AIRS (Automatic Intrinsic Reward Shaping)**: Adaptive intrinsic reward selection
|
||
3. **Entropy Regularization**: Maximum entropy RL for robust policies
|
||
4. **Curiosity-Driven Exploration**: ICM (Intrinsic Curiosity Module), Pathak et al. (2017), modern variants (2024-2025)
|
||
5. **Multi-Objective RL**: Pareto optimization for trading (Sharpe/profit/drawdown balance)
|
||
6. **Ensemble Model Fusion**: Transformer + LSTM + RL hybrid architectures (2024-2025)
|
||
|
||
### Internal Documentation
|
||
|
||
- **CLAUDE.md**: System architecture, Wave 10 campaign results
|
||
- **ml/src/dqn/reward.rs**: Current reward implementation (baseline)
|
||
- **ml/src/dqn/portfolio_tracker.rs**: Portfolio state tracking (218 lines, 9/9 tests)
|
||
- **ml/src/hyperopt/adapters/dqn.rs**: Hyperopt integration
|
||
- **/tmp/ml_training/wave10_production/WAVE10_FINAL_CAMPAIGN_REPORT.md**: 476-line analysis
|
||
|
||
---
|
||
|
||
## 10. Approval and Next Steps
|
||
|
||
**Recommended Action**:
|
||
1. **IMMEDIATE**: Review this proposal with user
|
||
2. **Short-term** (Week 1): Implement Phase 1-2 (core + entropy)
|
||
3. **Medium-term** (Week 2-3): Implement Phase 3-4 (curiosity + ensemble)
|
||
4. **Long-term** (Week 4-5): Hyperopt campaign and production deployment
|
||
|
||
**Required Approvals**:
|
||
- [ ] Technical design review (user approval)
|
||
- [ ] Resource allocation (3-4 weeks dev time)
|
||
- [ ] GPU budget ($0.75 for hyperopt)
|
||
|
||
**Contact**: @user for questions/feedback
|
||
|
||
---
|
||
|
||
**Document Status**: ✅ READY FOR REVIEW
|
||
**Version**: 1.0
|
||
**Last Updated**: 2025-11-08
|