feat(ml): WAVE 29 DQN Codebase Cleanup & Refactoring Campaign

BREAKING CHANGES:
- Removed orphaned dqn.rs monolithic trainer (4,975 lines)
- Removed orphaned dqn_ensemble.rs module (816 lines)
- Removed orphaned tft.rs and tft_complete_int8_integration_test.rs
- TFT trainer split into modular directory structure

DQN Module Refactoring:
- Split trainers/dqn.rs into modular structure (config.rs, statistics.rs, trainer.rs)
- Fixed hyperopt 39D search space (continuous params only)
- Boolean flags (use_dueling, use_double_dqn, use_per, use_noisy_nets) are now FIXED architectural decisions
- use_distributional defaults to false (Candle BUG #36 - scatter_add gradient issues)

Clean Module Structure:
- ml/src/trainers/dqn/ directory with proper mod.rs exports
- ml/src/trainers/tft/ directory with config.rs, types.rs, model.rs, trainer.rs, tests.rs
- All P0 features validated: TD-error clamping, batch diversity, LR scheduler, priority staleness

Documentation:
- Added comprehensive docs in docs/codebase-cleanup/
- ADR-001 for DQN refactoring decisions
- Rainbow DQN component matrix and quick reference guides

Build Status: Compiles with zero errors

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
This commit is contained in:
jgrusewski
2025-11-27 23:46:13 +01:00
parent 2c1acda2f3
commit 2df1ea92e1
763 changed files with 247870 additions and 1714 deletions

View File

@@ -0,0 +1,451 @@
# Agent 20: Reward Normalization Analysis Report
**Date**: 2025-11-27
**Task**: Review reward normalization for potential overfitting issues
**Status**: ✅ COMPLETE - Analysis Finished
---
## Executive Summary
**VERDICT**: The reward normalization system is **well-designed** with appropriate parameters. However, there are **two potential concerns** related to the effective window size and normalization algorithm that warrant attention.
### Key Findings
1.**EMA Algorithm**: Correctly chosen over Welford's for non-stationary markets
2. ⚠️ **Window Size**: 100-step effective window may be too short for regime detection
3.**Normalization Range**: Proper clipping at ±3.0σ preserves economic signal
4. ⚠️ **Overfitting Risk**: EMA parameters (α=0.01, β=0.01) are fixed, not hyperparameter-tuned
---
## 1. RewardNormalizer Implementation Analysis
### Architecture (Lines 14-165 in `/ml/src/dqn/reward.rs`)
```rust
pub struct RewardNormalizer {
mean: f64, // Running mean via EMA
variance: f64, // Running variance via EMA
alpha: 0.01, // Mean decay rate (100-step window)
beta: 0.01, // Variance decay rate (100-step window)
initialized: bool,
epsilon: 1e-8, // Numerical stability
}
```
### Algorithm: Exponential Moving Average (EMA)
**Update Rule** (Lines 101-114):
```rust
if !self.initialized {
self.mean = value;
self.variance = 1.0;
self.initialized = true;
} else {
// EMA update for mean
self.mean = self.alpha * value + (1.0 - self.alpha) * self.mean;
// EMA update for variance
let diff = value - self.mean;
self.variance = self.beta * diff.powi(2) + (1.0 - self.beta) * self.variance;
}
```
**Normalization** (Lines 121-135):
```rust
let std = self.variance.sqrt();
if std < self.epsilon {
return value;
}
(value - self.mean) / std
```
---
## 2. Parameter Analysis
### 2.1 Window Size: α = β = 0.01 → Effective Window ≈ 100 Steps
**Effective Window Calculation**:
- For decay rate α, effective window = 1/α
- α = 0.01 → **100 steps**
**Decay Profile** (Lines 95-98):
```
After 100 steps: weight = (1-0.01)^100 = 0.366 (63.4% decay)
After 200 steps: weight = (1-0.01)^200 = 0.134 (86.6% decay)
After 500 steps: weight = (1-0.01)^500 = 0.007 (99.3% decay)
```
**Assessment**:
-**Reactive**: Adapts quickly to changing reward distributions
- ⚠️ **Too Short?**: May overfit to recent 100 episodes, missing longer-term trends
- ⚠️ **Regime Detection**: Trading regimes can last 500+ steps (hours/days)
**Recommendation**:
```rust
// Consider longer windows for regime stability
alpha: 0.001, // 1000-step window (more stable)
beta: 0.005, // 200-step window (variance adapts faster)
```
---
## 3. Normalization Range & Clipping
### 3.1 Clipping at ±3.0σ (Line 557)
```rust
let normalized_reward = if let Some(normalizer) = &mut self.normalizer {
let norm = normalizer.normalize(final_reward_f64);
normalizer.update(final_reward_f64); // Update AFTER normalization
norm.clamp(-3.0, 3.0) // Clip to ±3σ
} else {
final_reward_f64
};
```
**Assessment**:
-**Appropriate Range**: ±3σ covers 99.7% of normal distribution
-**Preserves Signal**: Wide enough to not distort economic information
-**Prevents Explosion**: Bounds extreme outliers (e.g., rare large P&L events)
**Expected Q-Values** (Line 556):
```
Q* = r / (1 - γ)
≈ 3.0 / (1 - 0.9626)
≈ 80
```
---
## 4. Overfitting Risk Assessment
### 4.1 Fixed vs. Tunable Parameters
**Current Implementation**:
```rust
alpha: 0.01, // FIXED (hardcoded)
beta: 0.01, // FIXED (hardcoded)
```
**Concern**: These are **not exposed as hyperparameters** in the reward config.
**Evidence from Codebase**:
- `RewardConfig` (Lines 167-213) does NOT include `alpha`/`beta` parameters
- Hyperopt tuning (checked via `ml/src/hyperopt/`) does NOT optimize EMA decay rates
- All trials use the same fixed 100-step window
### 4.2 Potential Overfitting Scenarios
#### Scenario A: Too Short Window (α = 0.01)
```
Market regime: Volatile (500 steps) → Calm (500 steps) → Volatile (500 steps)
EMA behavior: Quickly forgets old volatility, normalizes to recent 100 steps
Result: Rewards appear stable in volatile regime → agent underestimates risk
```
#### Scenario B: Too Long Window (α = 0.001)
```
Market regime: Calm (1000 steps) → Flash crash (10 steps) → Recovery (50 steps)
EMA behavior: Slow to react, uses stale calm-regime statistics
Result: Crash rewards seem extreme → agent avoids trading during recovery
```
**Current Risk**: **MODERATE**
- 100-step window is a reasonable middle ground
- But lack of hyperparameter tuning means we don't know if it's optimal
---
## 5. Normalization Correctness (BUG #41 Fix)
### 5.1 Why EMA Over Welford's Algorithm
**Documentation** (Lines 25-29):
> Welford's algorithm gives equal weight to all historical samples, causing mean drift in non-stationary markets. When the market produces losses, the mean drifts negative and stays there, inverting reward signs. EMA gives exponentially decaying weight to old samples, adapting to regime changes.
**Assessment**:
-**Correct Choice**: EMA is appropriate for non-stationary trading environments
-**Prevents Mean Drift**: Old regimes decay exponentially, not linearly
-**BUG #41 Fix**: Addresses the Welford's algorithm issue mentioned in prior audit reports
---
## 6. Integration with Reward Function
### 6.1 Update Order (CRITICAL FIX, Lines 544-552)
```rust
// CRITICAL: Normalize BEFORE update to avoid zeroing first reward
let norm = normalizer.normalize(final_reward_f64); // Uses PREVIOUS stats
normalizer.update(final_reward_f64); // Updates for NEXT call
```
**Comment** (Lines 545-548):
> Bug: If we update(x) then normalize(x), first call returns 0:
> update(x): mean = x, variance = 1.0
> normalize(x): returns (x - x) / 1.0 = 0
> Solution: normalize(x) using PREVIOUS stats, THEN update with new value
**Assessment**:
-**Order Correct**: Prevents first-reward zero bug
-**Proper Implementation**: Update happens AFTER normalization
---
## 7. Testing Coverage
### 7.1 Existing Tests
**Found in Codebase**:
- `ml/tests/dqn_reward_normalization_test.rs` (11 unit tests)
- `ml/tests/dqn_reward_calculation_test.rs` (12 tests)
**Coverage**:
- ✅ BUY/SELL symmetry
- ✅ Zero P&L normalization
- ✅ Negative P&L handling
- ✅ Large P&L normalization
### 7.2 Missing Tests (from TEST_COVERAGE_GAPS_AUDIT.md)
**RewardNormalizer Update During Training**:
- No test verifies `normalizer.update()` is called after each reward
- **Risk**: Normalizer may use stale epoch-1 statistics for all epochs
**Normalized Reward Range (-1.0 to 1.0)**:
- No test verifies rewards stay within expected bounds
- **Risk**: Rewards may explode if normalization fails
---
## 8. Recommendations
### Priority 1: Add Window Size Hyperparameter
**Current**:
```rust
impl RewardNormalizer {
pub fn new() -> Self {
Self {
alpha: 0.01, // FIXED
beta: 0.01, // FIXED
// ...
}
}
}
```
**Proposed**:
```rust
pub struct RewardConfig {
// ... existing fields ...
/// EMA decay rate for mean (default: 0.01 = 100-step window)
pub reward_normalizer_alpha: Decimal,
/// EMA decay rate for variance (default: 0.01 = 100-step window)
pub reward_normalizer_beta: Decimal,
}
impl Default for RewardConfig {
fn default() -> Self {
Self {
// ... existing defaults ...
reward_normalizer_alpha: Decimal::try_from(0.01).unwrap(),
reward_normalizer_beta: Decimal::try_from(0.01).unwrap(),
}
}
}
```
**Hyperopt Search Space**:
```python
# Add to hyperopt configuration
"reward_normalizer_alpha": hp.loguniform("rn_alpha", -7, -2), # 0.001 to 0.1
"reward_normalizer_beta": hp.loguniform("rn_beta", -7, -2), # 0.001 to 0.1
```
### Priority 2: Test Normalization Update During Training
**Proposed Test** (add to `ml/tests/dqn_reward_normalization_test.rs`):
```rust
#[test]
fn test_normalizer_updates_during_training() {
let config = RewardConfig::default();
let mut reward_fn = RewardFunction::new(config);
// Verify initial state
assert_eq!(reward_fn.normalizer.as_ref().unwrap().count(), 0);
// Run 100 reward calculations
for i in 0..100 {
let action = FactoredAction::new(/* ... */);
let current_state = create_test_state(/* ... */);
let next_state = create_test_state(/* ... */);
reward_fn.calculate_reward(action, &current_state, &next_state, &[])?;
}
// Verify normalizer updated 100 times
assert_eq!(reward_fn.normalizer.as_ref().unwrap().count(), 100);
// Verify mean/variance changed from initialization
let (mean, std) = reward_fn.normalizer.as_ref().unwrap().get_stats();
assert_ne!(mean, 0.0);
assert_ne!(std, 1.0);
}
```
### Priority 3: Add Regime-Aware Window Sizing
**Proposed Enhancement**:
```rust
pub struct RewardNormalizer {
mean: f64,
variance: f64,
alpha: f64,
beta: f64,
// NEW: Adaptive window sizing
regime_detector: Option<RegimeDetector>,
alpha_fast: f64, // 0.01 (100-step) for volatile regimes
alpha_slow: f64, // 0.001 (1000-step) for stable regimes
}
impl RewardNormalizer {
pub fn update(&mut self, value: f64) {
// Detect regime change
let alpha = if let Some(detector) = &self.regime_detector {
if detector.is_volatile() {
self.alpha_fast // React quickly to volatility
} else {
self.alpha_slow // Use longer window for stability
}
} else {
self.alpha // Default fixed window
};
// EMA update with adaptive window
self.mean = alpha * value + (1.0 - alpha) * self.mean;
// ... variance update ...
}
}
```
---
## 9. Comparison with Alternative Approaches
### 9.1 Welford's Algorithm (Original Implementation)
**Pros**:
- ✅ Exact online mean/variance calculation
- ✅ No hyperparameters to tune
**Cons**:
- ❌ Equal weight to all samples → mean drift in non-stationary markets (BUG #41)
- ❌ Cannot adapt to regime changes
**Verdict**: ❌ **Not suitable for trading** (already replaced with EMA)
### 9.2 Batch Normalization (Per-Episode)
**Pros**:
- ✅ No window size parameter
- ✅ Normalizes each episode independently
**Cons**:
- ❌ Cannot normalize online (needs full episode)
- ❌ Breaks temporal correlation across episodes
**Verdict**: ❌ **Not suitable for online RL**
### 9.3 Quantile Normalization
**Pros**:
- ✅ Robust to outliers
- ✅ No assumption of normal distribution
**Cons**:
- ❌ Requires storing large buffer of rewards
- ❌ More complex to implement
**Verdict**: ⚠️ **Worth exploring** for extreme market events
---
## 10. Final Assessment
### Current Parameters
| Parameter | Value | Effective Window | Assessment |
|-----------|-------|------------------|------------|
| `alpha` (mean decay) | 0.01 | 100 steps | ⚠️ May be too reactive |
| `beta` (variance decay) | 0.01 | 100 steps | ⚠️ May be too reactive |
| Clipping range | ±3.0σ | 99.7% coverage | ✅ Appropriate |
| Algorithm | EMA | O(1) memory | ✅ Correct choice |
### Risk Analysis
| Risk Type | Severity | Likelihood | Mitigation |
|-----------|----------|------------|------------|
| **Too short window** | MEDIUM | MEDIUM | Expose as hyperparameter, test 500-1000 step windows |
| **Reward hacking** | LOW | LOW | Clipping at ±3σ prevents extreme values |
| **Gradient issues** | LOW | LOW | Normalization to ~N(0,1) stabilizes gradients |
| **Overfitting to recent data** | MEDIUM | MEDIUM | Use regime-aware adaptive windows |
### Does Normalization Preserve Reward Structure?
**Yes**, with caveats:
1.**Relative Ordering**: Normalization preserves `r1 > r2 > r3` ordering
2.**Economic Signal**: ±3σ clipping is wide enough to not distort information
3. ⚠️ **Magnitude Information**: Raw reward magnitude is lost (by design)
4. ⚠️ **Temporal Correlation**: 100-step window may smooth out short-term patterns
---
## 11. Summary & Action Items
### What's Working Well
1.**EMA Algorithm**: Correct choice for non-stationary markets (BUG #41 fix)
2.**Update Order**: Normalize BEFORE update (prevents first-reward zero bug)
3.**Clipping Range**: ±3σ is appropriate for preserving signal
4.**Numerical Stability**: epsilon=1e-8 prevents division by zero
### What Needs Attention
1. ⚠️ **Window Size**: 100 steps may be too short for regime detection
2. ⚠️ **Fixed Parameters**: α/β not exposed as hyperparameters for tuning
3. ⚠️ **Missing Tests**: No integration test for normalizer update during training
### Recommended Actions
| Action | Priority | Effort | Impact |
|--------|----------|--------|--------|
| Add α/β as hyperparameters | HIGH | LOW | Enables hyperopt tuning |
| Test 500-1000 step windows | HIGH | MEDIUM | Reduces overfitting risk |
| Add normalizer update test | MEDIUM | LOW | Catches regression bugs |
| Implement regime-aware windows | LOW | HIGH | Advanced optimization |
---
## Appendix: Effective Window Size Cheat Sheet
| Alpha (α) | Effective Window | Use Case |
|-----------|------------------|----------|
| 0.05 | 20 steps | Very reactive (intraday trading) |
| 0.01 | **100 steps** | **Current default** (balanced) |
| 0.005 | 200 steps | Moderate stability |
| 0.001 | 1000 steps | Stable regimes (multi-day) |
| 0.0001 | 10000 steps | Very stable (trend-following) |
**Recommendation**: Test α ∈ {0.001, 0.005, 0.01} via hyperopt to find optimal balance.
---
**Agent 20 Analysis Complete**
**Next Steps**: Coordinate with Agent 21 (Hyperparameter Tuner) to add α/β to search space.