# Risk Management & Reward Shaping Integration Analysis **Date**: 2025-11-27 **Location**: `/home/jgrusewski/Work/foxhunt` **Scope**: DQN Risk-Reward Integration --- ## Executive Summary This analysis examines how risk management integrates with reward shaping in the DQN system. The architecture shows **sophisticated multi-layer integration** with proper separation of concerns, but reveals **critical signal leakage risks** and **unit mismatches** that could cause reward overfitting to risk patterns. **Key Findings:** - ✅ **Strong Architecture**: EMA-based normalization, Kelly sizing, circuit breakers all properly implemented - ⚠️ **Signal Leakage Risk**: Reward function directly observes risk penalties → agent may learn to game risk metrics - ⚠️ **Unit Mismatch Fixed**: Hold penalty scaling bug fixed (1/1000 scale factor), but monitoring needed - ✅ **Training Stability**: Percentage-based rewards + normalization prevents gradient explosion - ⚠️ **Kelly Integration**: Not directly in reward function (good!), but requires 10+ trades minimum --- ## 1. Reward Normalization & Training Stability ### Implementation: `/ml/src/dqn/reward.rs` (Lines 14-165) **Architecture:** ```rust pub struct RewardNormalizer { mean: f64, // EMA of rewards variance: f64, // EMA of squared deviations alpha: 0.01, // Mean decay (100-step window) beta: 0.01, // Variance decay (100-step window) epsilon: 1e-8, // Numerical stability } // EMA update (NOT Welford's algorithm): mean = alpha * value + (1 - alpha) * mean variance = beta * (value - mean)^2 + (1 - beta) * variance ``` **Why EMA vs. Welford's Algorithm?** - **Welford's**: Equal weight to all history → mean drift in non-stationary markets - **EMA**: Exponential decay → adapts to regime changes (old samples decay as (1-α)^t) - **Window**: α=0.01 → 100-step effective window (63% decay after 100 steps) **Normalization Process:** 1. **Calculate raw reward** (percentage-based P&L) 2. **Normalize using PREVIOUS stats** (critical: avoids zeroing first reward) 3. **Update EMA** with current reward (for NEXT normalization) 4. **Clip to ±3.0 std** (prevents extreme outliers, preserves signal) **Expected Impact:** - Rewards centered at 0.0, scaled to std=1.0 - Q-values stabilize at: `r/(1-γ) ≈ 3.0/(1-0.9626) ≈ 80` - Gradient variance reduced significantly (BUG #41 fix) **Validation:** ```rust // Line 233-252: Mandatory percentage P&L validation if !use_percentage_pnl { return Err(MLError::ConfigError { reason: "use_percentage_pnl=false causes gradient explosion. \ Q-values explode to ±50,000..." }); } ``` **Status**: ✅ **Robust & Well-Documented** --- ## 2. Risk Penalties in Reward Function ### Implementation: `/ml/src/dqn/reward.rs` (Lines 478-573) **Reward Components (Weighted Sum):** ```rust final_reward = pnl_weight * pnl_reward // P&L (%) - risk_weight * risk_penalty // Position size penalty - cost_weight * cost_penalty // Transaction costs + diversity_bonus // Entropy regularization ``` ### 2.1 P&L Reward (Lines 640-712) **Formula:** ```rust pnl_reward = (next_value - current_value) / current_value // Percentage return ``` **Range**: ±0.02 (±2% typical per step) **Scale**: Already in percentage units → no 100x amplification **Status**: ✅ **Properly scaled for C51 distributional RL** ### 2.2 Risk Penalty (Lines 713-741) **Formula:** ```rust position_size = portfolio_features[1] // Signed: +Long, -Short, 0=Flat risk_penalty = if position_size > 0.8 { (position_size - 0.8) * 5.0 // Escalating penalty for >80% exposure } else { 0.0 } ``` **Issues:** 1. **Hardcoded threshold (0.8)**: Not tied to portfolio volatility or market regime 2. **Linear escalation**: 5x multiplier may be too aggressive or too weak depending on scale 3. **No drawdown integration**: Doesn't use actual drawdown data from `RiskMetrics` **Signal Leakage Risk:** - Reward directly observes position size → agent learns to hover at 79% to avoid penalty - **Recommendation**: Use Kelly-suggested position sizes as reference, not absolute thresholds ### 2.3 Transaction Cost Penalty (Lines 743-835) **Formula (FIXED):** ```rust // BUG FIX: Use action.transaction_cost() instead of spread estimation tx_cost_rate = action.transaction_cost() // Market: 0.15%, Limit: 0.05%, IoC: 0.10% cost_penalty = position_change * tx_cost_rate ``` **Before Fix**: Spread × 0.5 ≈ 0.05% (3x underestimated market orders) **After Fix**: Actual exchange fees (0.05-0.15%) **Status**: ✅ **Fixed in BUG #24** ### 2.4 Hold Reward / Penalty (Lines 866-907) **Formula (FIXED):** ```rust // BUG FIX: Scale hold_penalty_weight to percentage units hold_penalty_pct = hold_penalty_weight / 1000.0 // 0.5-2.0 → 0.0005-0.002 (0.05-0.2%) volatility = |next_log_return| // Absolute log return as price velocity hold_reward = if volatility < movement_threshold { +0.001 // Low volatility: reward holding } else { -hold_penalty_pct // High volatility: penalize holding } ``` **Before Fix**: 0.5-2.0 scalar (333-4000x mismatch with tx costs) **After Fix**: 0.0005-0.002 (5-20 basis points, aligned with 0.05-0.15% tx costs) **Ratio**: `hold_penalty / tx_cost ≈ 0.001 / 0.001 = 1.0` (balanced) **Status**: ✅ **Fixed in Phase 1 scaling adjustment** ### 2.5 Diversity Penalty (Lines 344-386, 517-526) **Formula:** ```rust entropy = -Σ(p_i * log2(p_i)) // Shannon entropy of last 100 actions diversity_bonus = if entropy < 0.5 { -0.1 // Penalty for low diversity (< 50% of max entropy) } else { 0.0 } ``` **Range**: Entropy ∈ [0, 1.585] (log2(3) for 3 actions) **Threshold**: 0.5 (31.5% of max) triggers penalty **Weight**: -0.1 (100x stronger than hold_reward) **Purpose**: Prevent action collapse (e.g., 98% SELL bug) **Status**: ✅ **Effective in-training regularization** --- ## 3. Kelly Sizing Integration ### Implementation: `/risk/src/kelly_sizing.rs` (Lines 1-342) **Kelly Criterion Formula:** ```rust // f* = (bp - q) / b // where: // b = odds ratio (average_win / average_loss) // p = win_rate // q = loss_rate (1 - p) // // Example: // win_rate = 0.6, avg_win = $50, avg_loss = $30 // b = 50/30 = 1.667 // f* = (1.667 * 0.6 - 0.4) / 1.667 = 0.36 (36% of capital) ``` **Adjustments:** 1. **Fractional Kelly**: Multiply by 0.25-0.5 (reduces risk, prevents over-betting) 2. **Confidence Scaling**: Based on sample size + win rate distribution 3. **Hard Caps**: Min 1%, Max 10% (prevents extreme bets) **Integration Point:** - **NOT in reward function** (good! avoids circular dependency) - Used in **position sizing** after action selection - Requires **minimum 10 trades** to calculate (lines 140-149) **Critical Issue:** ```rust // Line 140-149: Insufficient data handling if trades.len() < 10 { return Err(RiskError::DataUnavailable { reason: format!( "Insufficient trade history: {} trades (minimum 10 required)", trades.len() ), }); } ``` **Problem**: Early training (< 10 trades) will FAIL with error, not default sizing **Recommendation**: Use conservative default (e.g., 2% position size) until 10+ trades accumulated **Status**: ✅ **Kelly sizing isolated from rewards, but needs warmup handling** --- ## 4. Circuit Breaker Integration ### 4.1 Simplified Circuit Breaker: `/ml/src/dqn/circuit_breaker.rs` **Purpose**: Lightweight training throttling (no Redis, no broker) **States**: Closed → Open → HalfOpen → Closed **Thresholds:** - **Failure threshold**: 5 consecutive losses → OPEN - **Success threshold**: 3 consecutive wins → CLOSED (from HalfOpen) - **Timeout**: 60 seconds cooldown - **Half-open calls**: 2 test trades allowed **Trigger Logic:** ```rust // Open circuit after 5 consecutive failures if consecutive_failures >= 5 { state = Open; // Block all requests open_timestamp = now(); } // After 60s timeout, transition to HalfOpen if elapsed >= 60s { state = HalfOpen; // Allow 2 test trades } // Close after 3 consecutive successes in HalfOpen if consecutive_successes >= 3 { state = Closed; // Resume normal trading } ``` **Integration with Reward:** - Circuit breaker does **NOT affect reward calculation** - Only blocks action execution (prevents runaway losses) - Proper separation of concerns **Status**: ✅ **Well-isolated from reward function** ### 4.2 Risk Crate Circuit Breaker: `/ml/src/dqn/risk_integration.rs` **Purpose**: Full production-grade circuit breaker with Redis coordination **Components:** 1. **TrainingBrokerService**: Tracks portfolio value + daily P&L 2. **DQNRiskCircuitBreaker**: Wraps `risk::RealCircuitBreaker` 3. **Redis Coordination**: Multi-process safe (not needed for single-process training) **Configuration:** ```rust CircuitBreakerConfig { daily_loss_percentage: 10.0%, // $10K on $100K capital position_limit_percentage: 5.0%, // Max 5% portfolio per position max_consecutive_violations: 5, auto_recovery_enabled: false, // Manual recovery for safety cooldown_period_secs: 300, // 5 minutes } ``` **Reward Integration:** ```rust // Line 222-238: Record reward and check circuit breaker pub async fn record_reward(&self, reward: f64) -> RiskResult<()> { self.broker_service.record_reward(reward).await; // Check circuit breaker ONLY on losses if reward < 0.0 { let is_open = self.check().await?; if is_open { error!("Circuit breaker TRIGGERED"); error!("Daily Loss: ${:.2}", state.current_daily_loss); } } Ok(()) } ``` **Status**: ✅ **Properly isolated, but requires Redis (may not be needed for training)** --- ## 5. Signal Leakage Analysis ### 5.1 What is Signal Leakage? **Definition**: Agent learns to game risk metrics instead of learning profitable trading strategies. **Example:** - Agent observes: `risk_penalty = 0 when position_size < 0.8` - Agent learns: "Stay at 79% position to maximize reward" - **Problem**: Agent avoids legitimate risk management instead of learning optimal position sizing ### 5.2 Current Leakage Points | Component | Leakage Risk | Mechanism | Severity | |-----------|-------------|-----------|----------| | **Position Size Penalty** | HIGH | Hardcoded 0.8 threshold → agent learns to hover at 79% | 🔴 Critical | | **Transaction Costs** | LOW | Agent learns to minimize trading (good!) | 🟢 Benign | | **Hold Penalty** | MEDIUM | Agent may learn to avoid holding during high volatility even when correct | 🟡 Moderate | | **Diversity Penalty** | LOW | Entropy threshold only prevents collapse, doesn't leak strategy | 🟢 Benign | | **Kelly Sizing** | NONE | Not in reward function → no leakage | 🟢 Safe | | **Circuit Breaker** | NONE | Only blocks execution, doesn't affect rewards | 🟢 Safe | ### 5.3 Detected Leakage Example **Hypothesis**: Agent learns to maintain 79% position to avoid risk penalty **Evidence to Check:** 1. Distribution of position sizes in trained agent (should see spike at 79%) 2. Correlation between position size and reward (should be independent after removing P&L) 3. Agent behavior when position_size = 0.81 (should immediately reduce to 0.79) **Recommendation:** ```rust // Instead of hardcoded threshold, use Kelly-suggested position let kelly_position = kelly_sizer.calculate_kelly_fraction(symbol, strategy)?; let risk_penalty = if position_size > kelly_position * 1.2 { (position_size - kelly_position) * adaptive_multiplier } else { 0.0 }; ``` This makes the penalty **dynamic** and tied to actual risk analysis, not arbitrary thresholds. --- ## 6. Integration Point Summary ### 6.1 Reward Function Architecture ``` ┌─────────────────────────────────────────────────────┐ │ DQN Reward Function (reward.rs) │ ├─────────────────────────────────────────────────────┤ │ │ │ Base Reward = pnl_weight * pnl_reward │ │ - risk_weight * risk_penalty │ ← LEAKAGE RISK │ - cost_weight * cost_penalty │ ← Safe (realistic) │ + hold_reward / -hold_penalty │ ← Scaled correctly │ │ │ Final Reward = Base + diversity_bonus │ │ │ │ Normalized Reward = EMA_normalize(Final) │ ← Stability layer │ .clamp(-3.0, 3.0) │ │ │ └─────────────────────────────────────────────────────┘ ↓ ┌─────────────────────────────────────────────────────┐ │ Risk Management (Separate Layer) │ ├─────────────────────────────────────────────────────┤ │ │ │ Kelly Sizing (kelly_sizing.rs) │ ← NOT in rewards │ - Position size optimization │ │ - Requires 10+ trade history │ │ - Fractional Kelly (0.25-0.5) │ │ │ │ Circuit Breakers (circuit_breaker.rs) │ ← Execution layer │ - Simplified: 5 failures → OPEN │ │ - Risk Crate: Daily loss % → OPEN │ │ - Does NOT modify rewards │ │ │ └─────────────────────────────────────────────────────┘ ``` ### 6.2 Data Flow ``` Training Step: 1. Select action (DQN policy network) 2. Calculate raw reward (P&L + penalties) 3. Normalize reward (EMA + clipping) 4. Update Q-network (TD error) 5. [Separate] Kelly sizing for next position 6. [Separate] Circuit breaker check Execution Step: 1. Select action (trained policy) 2. [Pre-check] Circuit breaker allow? 3. [Pre-check] Kelly position size 4. Execute trade 5. [Post-check] Record outcome for Kelly 6. [Post-check] Update circuit breaker ``` **Key Observation**: Kelly and circuit breakers are **post-decision**, not pre-decision factors. They don't directly influence reward calculation during training. --- ## 7. Potential Issues & Recommendations ### Issue 1: Position Size Penalty (Hardcoded Threshold) **Severity**: 🔴 High **Location**: `/ml/src/dqn/reward.rs:718-741` **Problem:** ```rust let risk_penalty = if position_size > 0.8 { (position_size - 0.8) * 5.0 // Hardcoded 0.8 threshold } else { 0.0 } ``` **Signal Leakage**: Agent learns to maintain 79% position regardless of market conditions. **Recommendation**: ```rust // Use Kelly-suggested position as dynamic threshold let kelly_fraction = self.kelly_sizer.calculate_kelly_fraction(symbol, strategy)?; let safe_threshold = kelly_fraction.adjusted_kelly_fraction * 1.2; // 20% buffer let risk_penalty = if position_size > safe_threshold { (position_size - safe_threshold).powi(2) * adaptive_weight // Quadratic penalty } else { 0.0 } ``` ### Issue 2: Drawdown Not Used in Penalties **Severity**: 🟡 Medium **Location**: `/ml/src/dqn/reward.rs:319-329` **Problem**: `RiskMetrics` struct includes `max_drawdown`, but it's **never used** in risk penalty calculation. **Current Usage**: ```rust pub struct RiskMetrics { pub var_95: Decimal, // ❌ Not used pub max_drawdown: Decimal, // ❌ Not used pub sharpe_ratio: Decimal, // ❌ Not used pub volatility: Decimal, // ❌ Not used } ``` **Recommendation**: ```rust // Add drawdown penalty component let drawdown_penalty = if max_drawdown > threshold { (max_drawdown - threshold) * drawdown_weight } else { 0.0 }; final_reward = pnl_reward - risk_penalty - cost_penalty - drawdown_penalty; ``` ### Issue 3: Kelly Sizing Warmup **Severity**: 🟡 Medium **Location**: `/risk/src/kelly_sizing.rs:140-149` **Problem**: Training fails with error when < 10 trades recorded. **Current Behavior**: ```rust if trades.len() < 10 { return Err(RiskError::DataUnavailable { ... }); } ``` **Recommendation**: ```rust if trades.len() < 10 { // Use conservative default during warmup return Ok(KellyResult { position_fraction: 0.02, // 2% default use_kelly: false, confidence: 0.0, sample_size: trades.len(), ... }); } ``` ### Issue 4: EMA Normalization Cold Start **Severity**: 🟢 Low (already handled correctly) **Location**: `/ml/src/dqn/reward.rs:548-552` **Correct Implementation**: ```rust // CRITICAL FIX: Normalize BEFORE update to avoid zeroing first reward let norm = normalizer.normalize(final_reward_f64); // Use PREVIOUS stats normalizer.update(final_reward_f64); // Update for NEXT call ``` **Status**: ✅ Fixed (proper order prevents first-reward zeroing) ### Issue 5: Hold Penalty Unit Mismatch **Severity**: 🟢 Low (fixed, but monitor) **Location**: `/ml/src/dqn/reward.rs:881-887` **Fixed Implementation**: ```rust let hold_penalty_scale = Decimal::try_from(1000.0).unwrap_or(Decimal::ONE); let hold_penalty_pct = self.config.hold_penalty_weight / hold_penalty_scale; // Result: 0.5-2.0 → 0.0005-0.002 (5-20 basis points) ``` **Ratio Check**: - Hold penalty: 0.0005-0.002 (0.05-0.2%) - Transaction costs: 0.0005-0.0015 (0.05-0.15%) - Ratio: 0.33-1.33 (balanced) **Status**: ✅ Fixed in Phase 1, monitor for over/under-trading patterns --- ## 8. Testing & Validation Recommendations ### 8.1 Unit Tests (Add These) **Test 1: Position Size Leakage Detection** ```rust #[test] fn test_no_position_size_gaming() { let mut reward_fn = RewardFunction::new(RewardConfig::default()); // Test position sizes around threshold let positions = vec![0.75, 0.79, 0.80, 0.81, 0.85]; let rewards: Vec<_> = positions.iter().map(|&pos| { let state = create_state_with_position(pos); reward_fn.calculate_reward(action, &state, &next_state, &[]) }).collect(); // Check for discontinuity at 0.8 threshold assert!( (rewards[2] - rewards[1]).abs() < 0.1, "Reward should be smooth across threshold, not create gaming incentive" ); } ``` **Test 2: Kelly Warmup Handling** ```rust #[tokio::test] async fn test_kelly_warmup_does_not_fail() { let sizer = KellySizer::new(KellyConfig::default()); // Should return default sizing, not error let result = sizer.calculate_kelly_fraction(&Symbol::from("AAPL"), "strategy"); assert!(result.is_ok()); let kelly = result.unwrap(); assert_eq!(kelly.position_fraction, 0.02); // Conservative default assert!(!kelly.use_kelly); } ``` **Test 3: Hold Penalty Scaling** ```rust #[test] fn test_hold_penalty_magnitude() { let config = RewardConfig { hold_penalty_weight: Decimal::try_from(1.0).unwrap(), ..Default::default() }; let reward_fn = RewardFunction::new(config); // High volatility state let volatile_state = create_state_with_log_return(0.05); // 5% move let hold_reward = reward_fn.calculate_hold_reward(&state, &volatile_state)?; // Check penalty is in basis points (not dollars) assert!( hold_reward.abs() < Decimal::try_from(0.01).unwrap(), "Hold penalty should be in percentage units (< 1%)" ); } ``` ### 8.2 Integration Tests (Add These) **Test 4: End-to-End Reward + Risk Flow** ```rust #[tokio::test] async fn test_reward_risk_integration() { let mut reward_fn = RewardFunction::new(RewardConfig::default()); let kelly_sizer = KellySizer::new(KellyConfig::default()); let circuit_breaker = CircuitBreaker::new(CircuitBreakerConfig::default()); // Simulate 50 trading steps for step in 0..50 { // 1. Calculate reward let reward = reward_fn.calculate_reward(action, &state, &next_state, &actions)?; // 2. Record for Kelly (post-decision) if step >= 10 { kelly_sizer.add_trade_outcome(TradeOutcome { ... }); } // 3. Check circuit breaker (post-decision) if reward < 0.0 { circuit_breaker.record_failure(); assert!(circuit_breaker.allow_request()); // Should not block training } } // Verify Kelly sizing became available after warmup let kelly_result = kelly_sizer.calculate_kelly_fraction(&symbol, "strategy"); assert!(kelly_result.is_ok()); assert!(kelly_result.unwrap().sample_size >= 10); } ``` ### 8.3 Monitoring Metrics (Add These) **Metric 1: Position Size Distribution** ```rust // Track histogram of position sizes during training // Alert if spike at 0.79 (indicates gaming) let position_histogram = agent.get_position_histogram(); assert!( position_histogram[79] < 0.2, // < 20% of actions at 79% "Agent may be gaming position size threshold" ); ``` **Metric 2: Reward Component Variance** ```rust // Track variance of each reward component let component_stats = reward_fn.get_component_stats(); assert!( component_stats.risk_penalty_variance < component_stats.pnl_variance, "Risk penalty should not dominate reward signal" ); ``` **Metric 3: Kelly vs. Actual Position Divergence** ```rust // Compare Kelly-suggested vs. actual positions let kelly_fraction = kelly_sizer.calculate_kelly_fraction(&symbol, "strategy")?; let actual_position = agent.current_position_size(); let divergence = (actual_position - kelly_fraction).abs(); assert!( divergence < 0.2, "Agent should approximately follow Kelly sizing, divergence: {:.2}%", divergence * 100.0 ); ``` --- ## 9. Conclusion ### Strengths ✅ 1. **EMA-based normalization** prevents mean drift in non-stationary markets 2. **Percentage-based P&L** achieves scale-invariance (BUG #17 fix) 3. **Transaction cost fix** uses actual exchange fees (BUG #24 fix) 4. **Hold penalty scaling fix** aligns units with transaction costs (Phase 1 fix) 5. **Kelly sizing isolated** from reward function (no circular dependency) 6. **Circuit breakers separated** from reward calculation (proper layering) ### Weaknesses ⚠️ 1. **Position size penalty uses hardcoded threshold** → signal leakage risk 2. **Drawdown not used in penalties** despite being in `RiskMetrics` struct 3. **Kelly sizing requires 10+ trades** → early training failures 4. **No validation tests** for leakage detection or component interaction ### Critical Recommendations 🔴 1. **Replace hardcoded 0.8 threshold** with Kelly-based dynamic thresholds 2. **Add drawdown penalty** to reward function (use `RiskMetrics.max_drawdown`) 3. **Implement Kelly warmup** (use 2% default for first 10 trades) 4. **Add leakage detection tests** (position size distribution, reward component variance) 5. **Monitor training metrics** for position gaming and reward component dominance ### Expected Impact - **Training Stability**: Current EMA + percentage P&L is robust ✅ - **Risk Management**: Kelly + circuit breakers properly isolated ✅ - **Reward Overfitting**: Moderate risk due to hardcoded thresholds ⚠️ - **Signal Leakage**: High risk if agent learns to game position size threshold 🔴 --- ## File Locations | Component | File Path | Lines | |-----------|-----------|-------| | Reward Normalization | `/ml/src/dqn/reward.rs` | 14-165 | | Reward Function | `/ml/src/dqn/reward.rs` | 389-941 | | Risk Penalties | `/ml/src/dqn/reward.rs` | 713-907 | | Kelly Sizing | `/risk/src/kelly_sizing.rs` | 1-342 | | Circuit Breaker (Simple) | `/ml/src/dqn/circuit_breaker.rs` | 1-435 | | Risk Integration | `/ml/src/dqn/risk_integration.rs` | 1-355 | | Reward Coordinator | `/ml/src/dqn/reward_coordinator.rs` | 1-569 | --- ## Appendix: Reward Formula Breakdown **Complete Reward Calculation:** ```rust // Step 1: Calculate base components pnl_reward = (next_value - current_value) / current_value // ±0.02 typical risk_penalty = if position_size > 0.8 { (position_size - 0.8) * 5.0 } else { 0.0 } cost_penalty = position_change * tx_cost_rate // 0.05-0.15% hold_reward = if volatility < threshold { +0.001 } else { -hold_penalty_pct } diversity_bonus = if entropy < 0.5 { -0.1 } else { 0.0 } // Step 2: Weighted sum base_reward = pnl_weight * pnl_reward // Default: 1.0 - risk_weight * risk_penalty // Default: 0.1 - cost_weight * cost_penalty // Default: 1.0 + hold_reward // Conditional ±0.001 final_reward = base_reward + diversity_bonus // Step 3: Normalize and clip normalized = (final_reward - mean) / std clipped = normalized.clamp(-3.0, 3.0) // Step 4: Update EMA for next step mean = 0.01 * final_reward + 0.99 * mean variance = 0.01 * (final_reward - mean)^2 + 0.99 * variance ``` **Expected Ranges:** - Raw reward: -0.05 to +0.05 (±5% moves rare) - Normalized reward: -3.0 to +3.0 (clipped) - Q-values: -80 to +80 (r/(1-γ) with γ=0.9626) - TD errors: ±10 typical (after normalization) --- **End of Analysis**