Files
foxhunt/WAVE_5_A3_DEPENDENCY_REPORT.md
jgrusewski 17d94e654c feat(dqn): Wave 10 - Architectural improvements and bug fixes
Wave 10 Summary:
- A1-A4: Architecture upgrades (4x network, LeakyReLU, Xavier init, diagnostics)
- A5-A6: Integration testing and production validation
- A7: Research hyperopt vs manual tuning (manual recommended)
- A8-A12: HOLD penalty tuning and critical bug fixes

Architecture Changes:
- Network expansion: [128,64,32] → [256,128,64] (2.5x parameters)
- LeakyReLU activation (alpha=0.01) to prevent dead neurons
- Xavier/Glorot initialization for better gradient flow
- Real-time diagnostic monitoring (Q-values, dead neurons, gradients)

Critical Bugs Fixed:
- Bug #1: HOLD penalty not wired to reward calculation
- Bug #2: Zero price error in calculate_hold_reward (velocity-based fix)
- Huber loss default enabled (Wave 9)
- Shape mismatch fix (Wave 8)

Test Results:
- Integration tests: 149/152 passing (98%)
- New tests: 40+ tests added across 15 files
- Xavier init: 5/5 tests passing
- HOLD penalty wiring: 4/4 tests passing
- Zero price fix: 4/4 tests passing

Known Issues:
- HOLD bias persists at ~100% despite penalties
- Gradient collapse: 217 instances per training run (norm=0.0)
- Reversed penalty effect: Higher penalties → worse Q-spread
- Root cause: Gradient clipping bottleneck (max_norm=10.0 vs penalty signal)

Phase 1 Trials (all completed without crashes):
- Penalty 0.5: Q-spread 250 pts, HOLD 100%
- Penalty 1.0: Q-spread 251 pts, HOLD 100%
- Penalty 2.0: Q-spread 255 pts, HOLD 100% (+ Q-value explosion)

Next Steps: Architectural investigation via parallel agent debugging

🤖 Generated with Claude Code (https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
2025-11-06 00:38:23 +01:00

23 KiB

Wave 5-A3: Entropy Reward Tests - Dependency Analysis

Agent: Wave 5-A3 Task: Create comprehensive tests for entropy-based reward system Status: ⚠️ BLOCKED - Waiting for Wave 5-A1 and Wave 5-A2 to complete Date: 2025-11-05


Executive Summary

Wave 5-A3 cannot proceed until Wave 5-A1 (entropy regularization in reward function) and Wave 5-A2 (recent actions tracking in trainer) are implemented. This report documents:

  1. Current State: The codebase does NOT have entropy-based rewards implemented
  2. Dependencies: What Wave 5-A1 and 5-A2 must implement
  3. Test Plan: Complete test specification ready for implementation once dependencies are met
  4. Impact Analysis: 50+ existing tests need updating

Dependency Analysis

Wave 5-A1: Entropy Regularization (NOT IMPLEMENTED)

Expected Changes to ml/src/dqn/reward.rs:

// NEW: Entropy penalty configuration
pub struct RewardConfig {
    // ... existing fields ...
    pub entropy_penalty_weight: Decimal,  // NEW: Weight for entropy penalty (e.g., 0.1)
    pub entropy_threshold: Decimal,       // NEW: Minimum acceptable entropy (e.g., 0.5)
}

impl RewardFunction {
    /// Calculate reward for a state transition
    pub fn calculate_reward(
        &mut self,
        action: TradingAction,
        current_state: &TradingState,
        next_state: &TradingState,
        recent_actions: &[TradingAction],  // NEW: 4th parameter
    ) -> Result<Decimal, MLError> {
        // ... existing reward calculation ...

        // NEW: Calculate entropy penalty
        let entropy_penalty = self.calculate_entropy_penalty(recent_actions)?;
        let final_reward = base_reward - entropy_penalty;

        Ok(final_reward)
    }

    // NEW: Entropy calculation function
    fn calculate_entropy_penalty(&self, recent_actions: &[TradingAction]) -> Result<Decimal, MLError> {
        if recent_actions.is_empty() {
            return Ok(Decimal::ZERO);  // No penalty if no history
        }

        // Calculate action distribution
        let total = recent_actions.len() as f64;
        let buy_count = recent_actions.iter().filter(|a| matches!(a, TradingAction::Buy)).count() as f64;
        let sell_count = recent_actions.iter().filter(|a| matches!(a, TradingAction::Sell)).count() as f64;
        let hold_count = recent_actions.iter().filter(|a| matches!(a, TradingAction::Hold)).count() as f64;

        // Calculate entropy: H = -Σ(p_i * log2(p_i))
        let entropy = calculate_shannon_entropy(&[buy_count/total, sell_count/total, hold_count/total]);

        // Apply penalty if entropy below threshold
        if entropy < self.config.entropy_threshold {
            let penalty = (self.config.entropy_threshold - entropy) * self.config.entropy_penalty_weight;
            Ok(penalty)
        } else {
            Ok(Decimal::ZERO)
        }
    }
}

// Helper function
fn calculate_shannon_entropy(probabilities: &[f64]) -> Decimal {
    let entropy = -probabilities.iter()
        .filter(|&&p| p > 0.0)
        .map(|&p| p * p.log2())
        .sum::<f64>();
    Decimal::try_from(entropy).unwrap_or(Decimal::ZERO)
}

Status: NOT FOUND in current codebase or uncommitted changes


Wave 5-A2: Recent Actions Tracking (NOT IMPLEMENTED)

Expected Changes to ml/src/trainers/dqn.rs:

pub struct DQNTrainer {
    // ... existing fields ...
    pub recent_actions: VecDeque<TradingAction>,  // NEW: Sliding window (100 actions)
    pub action_window_size: usize,                 // NEW: Default 100
}

impl DQNTrainer {
    pub fn new(...) -> Result<Self> {
        // ... existing code ...

        Ok(Self {
            // ... existing fields ...
            recent_actions: VecDeque::with_capacity(100),
            action_window_size: 100,
        })
    }

    // Update action tracking in train_step() or similar
    fn track_action(&mut self, action: TradingAction) {
        self.recent_actions.push_back(action);
        if self.recent_actions.len() > self.action_window_size {
            self.recent_actions.pop_front();
        }
    }

    // Update all calculate_reward() calls to pass recent_actions
    // EXAMPLE:
    let reward = reward_fn.calculate_reward(
        action,
        &current_state,
        &next_state,
        &self.recent_actions.iter().cloned().collect::<Vec<_>>()  // NEW
    )?;
}

Status: NOT FOUND in current codebase or uncommitted changes


Current State Verification

Uncommitted Changes: Only Bug Fix Campaign changes (gradient clipping, portfolio tracking)

$ git diff ml/src/dqn/reward.rs | grep -i "entropy\|recent_action"
# NO MATCHES - entropy code not implemented

$ git diff ml/src/trainers/dqn.rs | grep -i "entropy\|recent_action"
# NO MATCHES - recent_actions tracking not implemented

Existing Tests: 50+ tests call calculate_reward() with 3 parameters (will break when signature changes)


Test Plan (Ready for Implementation)

New Test File: ml/tests/dqn_entropy_reward_test.rs

Total Tests: 12 tests covering all entropy scenarios

1. Entropy Calculation Tests (6 tests)

#[test]
fn test_entropy_balanced_actions() {
    // Setup: 33 BUY, 33 SELL, 34 HOLD (100 total)
    let recent_actions = vec![
        vec![TradingAction::Buy; 33],
        vec![TradingAction::Sell; 33],
        vec![TradingAction::Hold; 34],
    ].concat();

    let entropy = calculate_action_entropy(&recent_actions);

    // Expected: ~1.58 bits (near maximum 1.585 for 3 actions)
    assert!(entropy > Decimal::try_from(1.55).unwrap());
    assert!(entropy < Decimal::try_from(1.60).unwrap());

    // Expected penalty: 0.0 (balanced distribution)
    let penalty = calculate_entropy_penalty(&recent_actions, 0.5, 0.1);
    assert_eq!(penalty, Decimal::ZERO);
}

#[test]
fn test_entropy_extreme_bias() {
    // Setup: 99% HOLD, 0.5% BUY, 0.5% SELL
    let recent_actions = vec![
        vec![TradingAction::Hold; 99],
        vec![TradingAction::Buy; 1],
    ].concat();

    let entropy = calculate_action_entropy(&recent_actions);

    // Expected: ~0.08 bits (very low)
    assert!(entropy < Decimal::try_from(0.15).unwrap());

    // Expected penalty: -0.1 (full penalty for entropy << threshold)
    let penalty = calculate_entropy_penalty(&recent_actions, 0.5, 0.1);
    let expected = Decimal::try_from(0.042).unwrap(); // (0.5 - 0.08) * 0.1
    assert!((penalty - expected).abs() < Decimal::try_from(0.01).unwrap());
}

#[test]
fn test_entropy_threshold_boundary() {
    // Test entropy exactly at 0.5 threshold
    // 75% HOLD, 12.5% BUY, 12.5% SELL → entropy ≈ 0.5
    let recent_actions = vec![
        vec![TradingAction::Hold; 75],
        vec![TradingAction::Buy; 12],
        vec![TradingAction::Sell; 13],
    ].concat();

    let entropy = calculate_action_entropy(&recent_actions);

    // Entropy should be very close to threshold
    assert!((entropy - Decimal::try_from(0.5).unwrap()).abs() < Decimal::try_from(0.05).unwrap());

    // Penalty should be minimal (< 0.01)
    let penalty = calculate_entropy_penalty(&recent_actions, 0.5, 0.1);
    assert!(penalty < Decimal::try_from(0.01).unwrap());
}

#[test]
fn test_entropy_with_empty_window() {
    // No recent actions yet
    let recent_actions = vec![];

    let entropy = calculate_action_entropy(&recent_actions);

    // Expected: Maximum entropy (1.0) - no bias assumed
    assert_eq!(entropy, Decimal::ONE);

    // Expected penalty: 0.0
    let penalty = calculate_entropy_penalty(&recent_actions, 0.5, 0.1);
    assert_eq!(penalty, Decimal::ZERO);
}

#[test]
fn test_entropy_with_small_window() {
    // Only 10 actions (below typical 100 window size)
    let recent_actions = vec![
        TradingAction::Buy,
        TradingAction::Buy,
        TradingAction::Hold,
        TradingAction::Sell,
        TradingAction::Buy,
        TradingAction::Hold,
        TradingAction::Buy,
        TradingAction::Sell,
        TradingAction::Hold,
        TradingAction::Buy,
    ];

    let entropy = calculate_action_entropy(&recent_actions);

    // Expected: Moderate entropy (5 BUY, 2 SELL, 3 HOLD)
    // H = -(0.5*log2(0.5) + 0.2*log2(0.2) + 0.3*log2(0.3)) ≈ 1.49
    assert!(entropy > Decimal::try_from(1.40).unwrap());
    assert!(entropy < Decimal::try_from(1.55).unwrap());
}

#[test]
fn test_entropy_sliding_window() {
    // 101 actions total - verify only last 100 are used
    let mut recent_actions = vec![TradingAction::Buy; 50]; // First 50 BUY (will be dropped)
    recent_actions.extend(vec![TradingAction::Hold; 100]); // Next 100 HOLD

    // Apply sliding window (take last 100)
    let windowed = &recent_actions[recent_actions.len() - 100..];

    let entropy = calculate_action_entropy(windowed);

    // Expected: 0.0 (100% HOLD in window)
    assert_eq!(entropy, Decimal::ZERO);

    // Expected penalty: Maximum (0.5 - 0.0) * 0.1 = 0.05
    let penalty = calculate_entropy_penalty(windowed, 0.5, 0.1);
    assert_eq!(penalty, Decimal::try_from(0.05).unwrap());
}

2. Reward Integration Tests (3 tests)

#[test]
fn test_reward_includes_diversity_penalty() {
    let config = RewardConfig {
        entropy_penalty_weight: Decimal::try_from(0.1).unwrap(),
        entropy_threshold: Decimal::try_from(0.5).unwrap(),
        ..Default::default()
    };
    let mut reward_fn = RewardFunction::new(config);

    // Extreme bias: 99% HOLD
    let recent_actions = vec![TradingAction::Hold; 99];

    let current_state = create_test_state(100.0, 10000.0);
    let next_state = create_test_state(105.0, 10500.0); // 5% gain

    let reward = reward_fn.calculate_reward(
        TradingAction::Buy,
        &current_state,
        &next_state,
        &recent_actions,
    ).unwrap();

    // Expected: base_reward (~0.05 for 5% gain) - penalty (~0.042)
    // Reward should be reduced by penalty
    assert!(reward < Decimal::try_from(0.01).unwrap());
}

#[test]
fn test_reward_no_penalty_for_balanced() {
    let config = RewardConfig {
        entropy_penalty_weight: Decimal::try_from(0.1).unwrap(),
        entropy_threshold: Decimal::try_from(0.5).unwrap(),
        ..Default::default()
    };
    let mut reward_fn = RewardFunction::new(config);

    // Balanced: 33% each
    let recent_actions = vec![
        vec![TradingAction::Buy; 33],
        vec![TradingAction::Sell; 33],
        vec![TradingAction::Hold; 34],
    ].concat();

    let current_state = create_test_state(100.0, 10000.0);
    let next_state = create_test_state(105.0, 10500.0); // 5% gain

    let reward = reward_fn.calculate_reward(
        TradingAction::Buy,
        &current_state,
        &next_state,
        &recent_actions,
    ).unwrap();

    // Expected: base_reward (~0.05) with NO penalty
    assert!(reward > Decimal::try_from(0.04).unwrap());
}

#[test]
fn test_penalty_strength_relative_to_pnl() {
    // Verify -0.1 penalty is ~10% of typical P&L reward
    let config = RewardConfig {
        entropy_penalty_weight: Decimal::try_from(0.1).unwrap(),
        entropy_threshold: Decimal::try_from(0.5).unwrap(),
        ..Default::default()
    };

    // Typical P&L rewards range from -1.0 to +1.0 for 10% moves
    let typical_pnl_reward = Decimal::try_from(0.5).unwrap();

    // Max penalty (entropy = 0) = (0.5 - 0) * 0.1 = 0.05
    let max_penalty = Decimal::try_from(0.05).unwrap();

    // Verify penalty is ~10% of typical reward
    let ratio = max_penalty / typical_pnl_reward;
    assert_eq!(ratio, Decimal::try_from(0.1).unwrap()); // 10%

    // Verify penalty is noticeable but not dominating
    assert!(max_penalty < typical_pnl_reward / Decimal::from(2)); // < 50% of reward
}

3. Batch Rewards Tests (3 tests)

#[test]
fn test_batch_rewards_with_entropy() {
    let config = RewardConfig {
        entropy_penalty_weight: Decimal::try_from(0.1).unwrap(),
        entropy_threshold: Decimal::try_from(0.5).unwrap(),
        ..Default::default()
    };
    let mut reward_fn = RewardFunction::new(config);

    // Batch of 5 actions
    let actions = vec![
        TradingAction::Buy,
        TradingAction::Sell,
        TradingAction::Hold,
        TradingAction::Buy,
        TradingAction::Sell,
    ];

    let states = (0..5).map(|i| create_test_state(100.0 + i as f64, 10000.0)).collect::<Vec<_>>();
    let next_states = (0..5).map(|i| create_test_state(101.0 + i as f64, 10100.0)).collect::<Vec<_>>();

    // Growing action history (entropy increases as we add actions)
    let mut recent_actions = vec![];
    let mut rewards = vec![];

    for i in 0..5 {
        recent_actions.push(actions[i]);
        let reward = reward_fn.calculate_reward(
            actions[i],
            &states[i],
            &next_states[i],
            &recent_actions,
        ).unwrap();
        rewards.push(reward);
    }

    // Verify all rewards are finite
    assert!(rewards.iter().all(|r| r.is_finite()));

    // Verify entropy penalty decreases as diversity increases
    // (rewards should increase as entropy increases)
    assert!(rewards.len() == 5);
}

#[test]
fn test_entropy_updates_per_action() {
    // Verify entropy is recalculated after each action
    let config = RewardConfig {
        entropy_penalty_weight: Decimal::try_from(0.1).unwrap(),
        entropy_threshold: Decimal::try_from(0.5).unwrap(),
        ..Default::default()
    };
    let mut reward_fn = RewardFunction::new(config);

    let mut recent_actions = vec![TradingAction::Hold; 50]; // Start with 50% HOLD

    // Calculate entropy before adding new action
    let entropy_before = calculate_action_entropy(&recent_actions);

    // Add diverse action (BUY)
    recent_actions.push(TradingAction::Buy);

    // Calculate entropy after
    let entropy_after = calculate_action_entropy(&recent_actions);

    // Entropy should increase (more diversity)
    assert!(entropy_after > entropy_before);
}

#[test]
fn test_entropy_window_overflow() {
    // Test that sliding window correctly drops old actions
    let config = RewardConfig {
        entropy_penalty_weight: Decimal::try_from(0.1).unwrap(),
        entropy_threshold: Decimal::try_from(0.5).unwrap(),
        ..Default::default()
    };

    // Create 150 actions (exceeds 100 window)
    let mut all_actions = vec![TradingAction::Hold; 100]; // First 100 HOLD
    all_actions.extend(vec![TradingAction::Buy; 50]);     // Next 50 BUY

    // Window should only contain last 100 (50 HOLD + 50 BUY)
    let windowed = &all_actions[all_actions.len() - 100..];

    let entropy = calculate_action_entropy(windowed);

    // Expected: H(0.5, 0.5, 0.0) = 1.0 bit (perfect split BUY/HOLD)
    assert!((entropy - Decimal::ONE).abs() < Decimal::try_from(0.05).unwrap());
}

Impact Analysis: Existing Tests Requiring Updates

Total Files: 50+ test files Total Call Sites: 100+ calculate_reward() calls

Required Updates

BEFORE (current signature):

let reward = reward_fn.calculate_reward(action, &current_state, &next_state)?;

AFTER (new signature):

let recent_actions = vec![]; // Empty for unit tests (no entropy penalty)
let reward = reward_fn.calculate_reward(action, &current_state, &next_state, &recent_actions)?;

Test Files Requiring Updates (50+ files)

ml/tests/dqn_portfolio_tracking_integration_test.rs     (2 call sites)
ml/tests/dqn_reward_function_unit_test.rs                (17 call sites)
ml/tests/dqn_penalty_effectiveness_test.rs               (5 call sites)
ml/tests/dqn_dynamic_hold_reward_test.rs                 (9 call sites)
ml/tests/dqn_reward_normalization_test.rs                (8 call sites)
ml/tests/dqn_diversity_penalty_test.rs                   (6 call sites)
ml/tests/dqn_reward_comprehensive_test.rs                (12 call sites)
... (43 more files)

Estimated Update Effort: 2-3 hours (automated search/replace + manual verification)


Implementation Roadmap

Step 1: Verify Dependencies Complete (WAITING)

# Check Wave 5-A1 (entropy in reward.rs)
grep -n "calculate_entropy_penalty\|entropy_penalty_weight" ml/src/dqn/reward.rs

# Check Wave 5-A2 (recent_actions in trainer)
grep -n "recent_actions: VecDeque\|action_window_size" ml/src/trainers/dqn.rs

Expected Output:

  • Found calculate_entropy_penalty() function
  • Found recent_actions field in DQNTrainer
  • Found calculate_reward() with 4 parameters

Step 2: Create Test File (30 minutes)

# Create test file
touch ml/tests/dqn_entropy_reward_test.rs

# Add to ml/tests/mod.rs if needed

Step 3: Implement 12 Tests (2 hours)

  • 6 entropy calculation tests
  • 3 reward integration tests
  • 3 batch/window tests

Step 4: Update Existing Tests (2-3 hours)

# Find all call sites
rg "calculate_reward\(" ml/tests/dqn*.rs -l | wc -l
# Expected: 50+ files

# Update signature (semi-automated)
rg "calculate_reward\(" ml/tests/dqn*.rs -A 2 -B 2

Step 5: Compile & Run (10 minutes)

# Compile only
cargo test --package ml --test dqn_entropy_reward_test --features cuda --no-run

# Run tests
cargo test --package ml --test dqn_entropy_reward_test --features cuda

# Run all DQN tests
cargo test --package ml --lib --features cuda dqn

Expected Results:

  • New tests: 12/12 passing
  • Existing tests: 147/147 passing (updated signatures)
  • Total DQN tests: 159/159 (147 + 12)

Blockers & Risks

Critical Blockers

  1. Wave 5-A1 NOT Implemented

    • calculate_entropy_penalty() function missing
    • entropy_penalty_weight field missing from RewardConfig
    • recent_actions parameter NOT added to calculate_reward()
  2. Wave 5-A2 NOT Implemented

    • recent_actions: VecDeque field missing from DQNTrainer
    • Action tracking logic missing
    • No sliding window implementation

Risks

  1. Breaking Changes: All existing tests will break when signature changes
  2. Timeline: 50+ test files need updates (2-3 hours manual work)
  3. Coordination: Wave 5-A1 and 5-A2 must coordinate on parameter format

Recommendations

For Wave 5-A1 Agent

  1. Add entropy configuration to RewardConfig:

    pub entropy_penalty_weight: Decimal,  // Suggested: 0.1
    pub entropy_threshold: Decimal,       // Suggested: 0.5
    
  2. Update calculate_reward() signature:

    pub fn calculate_reward(
        &mut self,
        action: TradingAction,
        current_state: &TradingState,
        next_state: &TradingState,
        recent_actions: &[TradingAction],  // NEW
    ) -> Result<Decimal, MLError>
    
  3. Implement calculate_entropy_penalty():

    • Use Shannon entropy formula: H = -Σ(p_i * log2(p_i))
    • Return penalty proportional to entropy deficit below threshold
    • Handle empty recent_actions gracefully (no penalty)
  4. Add unit tests in reward.rs:

    • Test entropy calculation with known distributions
    • Test penalty scaling
    • Test edge cases (empty, single action, etc.)

For Wave 5-A2 Agent

  1. Add fields to DQNTrainer:

    pub recent_actions: VecDeque<TradingAction>,
    pub action_window_size: usize,  // Default: 100
    
  2. Track actions in training loop:

    • Call self.track_action(action) after each step
    • Maintain sliding window (drop oldest when > 100)
  3. Update all calculate_reward() calls:

    • Convert VecDeque to Vec for passing to reward function
    • Pass as 4th parameter
  4. Add integration tests in dqn.rs:

    • Test action window correctly limits to 100
    • Test window correctly slides (FIFO)
    • Test entropy changes across training steps

For Wave 5-A3 Agent (This Agent)

  1. Wait for A1 and A2 to complete ⏸️

  2. Monitor for completion signals:

    • Git commit messages mentioning "Wave 5-A1" or "Wave 5-A2"
    • Presence of calculate_entropy_penalty() in codebase
    • Presence of recent_actions field in trainer
  3. Once unblocked:

    • Implement 12 new tests in dqn_entropy_reward_test.rs
    • Update 50+ existing test files
    • Verify 159/159 tests passing

Success Criteria

Entropy Implementation (Wave 5-A1)

  • calculate_entropy_penalty() function exists
  • Entropy calculation mathematically correct (Shannon entropy)
  • Penalty scales linearly with entropy deficit
  • Empty recent_actions handled gracefully
  • Unit tests in reward.rs passing

Action Tracking (Wave 5-A2)

  • recent_actions field exists in DQNTrainer
  • Sliding window correctly maintains 100 actions
  • Actions tracked after each training step
  • All calculate_reward() calls updated
  • Integration tests passing

Testing (Wave 5-A3)

  • 12 new entropy tests created
  • All 12 tests passing
  • 50+ existing tests updated
  • All existing tests still passing
  • Total test count: 159/159 (147 + 12)

Next Actions

IMMEDIATE:

  1. ⏸️ WAIT for Wave 5-A1 and Wave 5-A2 completion signals
  2. 👀 MONITOR git commits and codebase for changes
  3. 📋 PREPARE test implementation (this report serves as blueprint)

UPON UNBLOCKING:

  1. Verify dependencies implemented
  2. 🧪 Create dqn_entropy_reward_test.rs
  3. 🔧 Update existing test files
  4. Run full test suite
  5. 📊 Report results

Appendix: Helper Functions

/// Calculate Shannon entropy for action distribution
fn calculate_action_entropy(recent_actions: &[TradingAction]) -> Decimal {
    if recent_actions.is_empty() {
        return Decimal::ONE; // Assume maximum entropy (no bias)
    }

    let total = recent_actions.len() as f64;
    let buy_count = recent_actions.iter().filter(|a| matches!(a, TradingAction::Buy)).count() as f64;
    let sell_count = recent_actions.iter().filter(|a| matches!(a, TradingAction::Sell)).count() as f64;
    let hold_count = recent_actions.iter().filter(|a| matches!(a, TradingAction::Hold)).count() as f64;

    let probabilities = [buy_count/total, sell_count/total, hold_count/total];

    let entropy = -probabilities.iter()
        .filter(|&&p| p > 0.0)
        .map(|&p| p * p.log2())
        .sum::<f64>();

    Decimal::try_from(entropy).unwrap_or(Decimal::ZERO)
}

/// Calculate entropy penalty given configuration
fn calculate_entropy_penalty(
    recent_actions: &[TradingAction],
    entropy_threshold: f64,
    penalty_weight: f64,
) -> Decimal {
    let entropy = calculate_action_entropy(recent_actions);
    let threshold = Decimal::try_from(entropy_threshold).unwrap();
    let weight = Decimal::try_from(penalty_weight).unwrap();

    if entropy < threshold {
        (threshold - entropy) * weight
    } else {
        Decimal::ZERO
    }
}

/// Create test state with price and portfolio value
fn create_test_state(price: f64, portfolio_value: f64) -> TradingState {
    TradingState {
        price_features: vec![price, price, price, price], // [close, high, low, open]
        technical_indicators: vec![0.5; 4],
        market_features: vec![0.001, 1000.0, 0.0, 0.0],
        portfolio_features: vec![portfolio_value / 10000.0, 0.0, 0.0], // Normalized
    }
}

End of Report