Files
foxhunt/docs/plans/2026-03-05-hyperopt-degenerate-fix.md
2026-03-05 17:06:53 +01:00

22 KiB

Hyperopt Degenerate Trial Fix — Implementation Plan

For Claude: REQUIRED SUB-SKILL: Use superpowers:executing-plans to implement this plan task-by-task.

Goal: Fix the flat objective plateau (1.25) that makes 7/9 hyperopt trials indistinguishable, giving PSO zero gradient signal in 26D space.

Architecture: Three surgical changes: (1) Add exposure-aware eval to EvaluationEngine so factored actions generate trades from partial position changes, (2) Add graduated trade-count penalty to extract_objective() so different failure modes produce unique objective values, (3) Wire the factored eval into the hyperopt backtest loop (replace legacy 3-action collapse).

Tech Stack: Rust, candle ML framework, EvaluationEngine, DQN hyperopt adapter, PSO optimizer


Task 1: Add process_bar_factored() to EvaluationEngine

Files:

  • Modify: crates/ml/src/evaluation/engine.rs (add new fields + method after line 184)

Step 1: Add new fields to EvaluationEngine struct

Add current_exposure field to track continuous position exposure. The existing current_position + trades + kelly_fraction remain unchanged.

In crates/ml/src/evaluation/engine.rs, add after line 60 (pub kelly_fraction: f64,):

    /// Current position exposure for factored evaluation (-1.0 to +1.0)
    /// None = not using factored eval, Some(f64) = active factored tracking
    pub current_exposure: f64,
    /// Entry price for current exposure-based position
    pub exposure_entry_price: f32,
    /// Bar index where current exposure was first entered
    pub exposure_entry_bar: usize,

Step 2: Initialize new fields in constructors

In new_with_kelly() (line 69), add after kelly_fraction,:

            current_exposure: 0.0,
            exposure_entry_price: 0.0,
            exposure_entry_bar: 0,

Step 3: Add process_bar_factored() method

Add after close_position() (after line 184):

    /// Process a bar using the full 45-action factored space.
    ///
    /// Tracks continuous exposure (-1.0 to +1.0) and generates trades on
    /// any exposure change (including partial: Long100→Long50 = sell 0.5).
    /// Transaction costs use the FactoredAction's order type.
    pub fn process_bar_factored(
        &mut self,
        bar_idx: usize,
        bar: &OHLCVBarF32,
        action: &crate::common::action::FactoredAction,
    ) {
        let target = action.target_exposure(); // -1.0, -0.5, 0.0, +0.5, +1.0
        let delta = target - self.current_exposure;

        // Update legacy action counts for compatibility with metrics
        if target > 0.0 {
            self.action_counts[0] += 1; // buy
        } else if target < 0.0 {
            self.action_counts[2] += 1; // sell
        } else {
            self.action_counts[1] += 1; // hold
        }

        const EPSILON: f64 = 1e-6;
        if delta.abs() < EPSILON {
            return; // No position change
        }

        // Record trade for the exposure change
        let effective_delta = delta.abs() * self.kelly_fraction;
        let fee_rate = action.transaction_cost() as f64; // 0.0015 Market, 0.0005 Limit, 0.001 IoC

        // PnL from the portion being closed (if reducing or reversing)
        let closing_size = if delta.signum() != self.current_exposure.signum() && self.current_exposure.abs() > EPSILON {
            // Closing part (or all) of existing position
            self.current_exposure.abs().min(delta.abs())
        } else if delta.abs() < self.current_exposure.abs() && delta.signum() == -self.current_exposure.signum() {
            delta.abs()
        } else {
            0.0
        };

        if closing_size > EPSILON {
            let price_diff = bar.close - self.exposure_entry_price;
            let direction_sign = if self.current_exposure > 0.0 { 1.0_f32 } else { -1.0_f32 };
            let gross_pnl = price_diff * direction_sign * (closing_size * self.kelly_fraction) as f32;
            let tx_cost = (self.exposure_entry_price.abs() + bar.close.abs()) as f64
                * 0.5 * closing_size * self.kelly_fraction * fee_rate;

            self.trades.push(Trade {
                entry_bar_idx: self.exposure_entry_bar,
                exit_bar_idx: bar_idx,
                entry_price: self.exposure_entry_price,
                exit_price: bar.close,
                direction: if self.current_exposure > 0.0 {
                    "long".to_owned()
                } else {
                    "short".to_owned()
                },
                pnl: gross_pnl - tx_cost as f32,
            });
        }

        // Update exposure state
        if (target.abs()) > EPSILON {
            // Opening or adjusting — reset entry if crossing zero or first entry
            if self.current_exposure.abs() < EPSILON || target.signum() != self.current_exposure.signum() {
                self.exposure_entry_price = bar.close;
                self.exposure_entry_bar = bar_idx;
            }
        }
        self.current_exposure = target;
    }

    /// Close any remaining factored exposure at end of backtest
    pub fn close_factored_position(&mut self, bar_idx: usize, bar: &OHLCVBarF32) {
        const EPSILON: f64 = 1e-6;
        if self.current_exposure.abs() < EPSILON {
            return;
        }
        let price_diff = bar.close - self.exposure_entry_price;
        let direction_sign = if self.current_exposure > 0.0 { 1.0_f32 } else { -1.0_f32 };
        let size = self.current_exposure.abs() * self.kelly_fraction;
        let gross_pnl = price_diff * direction_sign * size as f32;
        let tx_cost = bar.close.abs() as f64 * size * 0.0015; // Market order for forced close

        self.trades.push(Trade {
            entry_bar_idx: self.exposure_entry_bar,
            exit_bar_idx: bar_idx,
            entry_price: self.exposure_entry_price,
            exit_price: bar.close,
            direction: if self.current_exposure > 0.0 {
                "long".to_owned()
            } else {
                "short".to_owned()
            },
            pnl: gross_pnl - tx_cost as f32,
        });
        self.current_exposure = 0.0;
    }

Step 4: Run cargo check -p ml

Run: SQLX_OFFLINE=true cargo check -p ml Expected: compiles with 0 errors

Step 5: Commit

git add crates/ml/src/evaluation/engine.rs
git commit -m "feat(ml): add process_bar_factored() for exposure-aware backtesting

Tracks continuous position exposure (-1.0 to +1.0) instead of binary
Buy/Sell/Hold. Partial position changes (Long100→Long50) now generate
trades, giving hyperopt richer evaluation signal."

Task 2: Add unit tests for process_bar_factored()

Files:

  • Modify: crates/ml/src/evaluation/engine.rs (add #[cfg(test)] mod tests at bottom)

Step 1: Write tests

Add at the end of crates/ml/src/evaluation/engine.rs:

#[cfg(test)]
mod tests {
    use super::*;
    use crate::common::action::{ExposureLevel, FactoredAction, OrderType, Urgency};

    fn bar(close: f32) -> OHLCVBarF32 {
        OHLCVBarF32 { timestamp: 0, open: close, high: close, low: close, close, volume: 0.0 }
    }

    fn market_action(exposure: ExposureLevel) -> FactoredAction {
        FactoredAction::new(exposure, OrderType::Market, Urgency::Normal)
    }

    #[test]
    fn factored_same_exposure_no_trade() {
        let mut engine = EvaluationEngine::new(10000.0);
        let b = bar(100.0);
        engine.process_bar_factored(0, &b, &market_action(ExposureLevel::Long100));
        engine.process_bar_factored(1, &b, &market_action(ExposureLevel::Long100));
        assert_eq!(engine.trades.len(), 0, "Same exposure should generate no trades");
    }

    #[test]
    fn factored_partial_close_generates_trade() {
        let mut engine = EvaluationEngine::new(10000.0);
        engine.process_bar_factored(0, &bar(100.0), &market_action(ExposureLevel::Long100));
        engine.process_bar_factored(1, &bar(110.0), &market_action(ExposureLevel::Long50));
        assert_eq!(engine.trades.len(), 1, "Long100→Long50 should generate 1 trade");
        assert!(engine.trades[0].pnl > 0.0, "Price went up on long = profit");
        assert_eq!(engine.current_exposure, 0.5);
    }

    #[test]
    fn factored_reversal_generates_trade() {
        let mut engine = EvaluationEngine::new(10000.0);
        engine.process_bar_factored(0, &bar(100.0), &market_action(ExposureLevel::Long100));
        engine.process_bar_factored(1, &bar(105.0), &market_action(ExposureLevel::Short100));
        assert!(engine.trades.len() >= 1, "Reversal should generate at least 1 trade");
        assert_eq!(engine.current_exposure, -1.0);
    }

    #[test]
    fn factored_flat_from_long_closes() {
        let mut engine = EvaluationEngine::new(10000.0);
        engine.process_bar_factored(0, &bar(100.0), &market_action(ExposureLevel::Long100));
        engine.process_bar_factored(1, &bar(95.0), &market_action(ExposureLevel::Flat));
        assert_eq!(engine.trades.len(), 1);
        assert!(engine.trades[0].pnl < 0.0, "Price went down on long = loss");
        assert_eq!(engine.current_exposure, 0.0);
    }

    #[test]
    fn factored_close_at_end() {
        let mut engine = EvaluationEngine::new(10000.0);
        engine.process_bar_factored(0, &bar(100.0), &market_action(ExposureLevel::Short50));
        engine.close_factored_position(1, &bar(90.0));
        assert_eq!(engine.trades.len(), 1);
        assert!(engine.trades[0].pnl > 0.0, "Price down on short = profit");
        assert_eq!(engine.current_exposure, 0.0);
    }

    #[test]
    fn factored_all_buy_still_one_trade_at_close() {
        let mut engine = EvaluationEngine::new(10000.0);
        for i in 0..100 {
            engine.process_bar_factored(i, &bar(100.0 + i as f32), &market_action(ExposureLevel::Long100));
        }
        // All same exposure → 0 trades during loop
        assert_eq!(engine.trades.len(), 0);
        // But close at end generates 1
        engine.close_factored_position(100, &bar(200.0));
        assert_eq!(engine.trades.len(), 1);
    }

    #[test]
    fn factored_alternating_generates_many_trades() {
        let mut engine = EvaluationEngine::new(10000.0);
        for i in 0..10 {
            let action = if i % 2 == 0 {
                market_action(ExposureLevel::Long100)
            } else {
                market_action(ExposureLevel::Short100)
            };
            engine.process_bar_factored(i, &bar(100.0), &action);
        }
        // Each reversal closes + opens = 1 trade per transition after first
        assert!(engine.trades.len() >= 9, "Alternating should generate many trades: got {}", engine.trades.len());
    }
}

Step 2: Run tests

Run: SQLX_OFFLINE=true cargo test -p ml --lib evaluation::engine::tests -- --nocapture Expected: all 7 tests pass

Step 3: Commit

git add crates/ml/src/evaluation/engine.rs
git commit -m "test(ml): add unit tests for process_bar_factored()"

Task 3: Add graduated trade insufficiency penalty to extract_objective()

Files:

  • Modify: crates/ml/src/hyperopt/adapters/dqn.rs:3105-3258 (extract_objective() method)

Step 1: Add calculate_trade_insufficiency_penalty() function

Add before extract_objective() (before line 3105):

/// Calculate graduated penalty for insufficient trade count.
///
/// Produces different objective values for different failure modes so PSO
/// gets gradient signal instead of a flat plateau at 1.25.
///
/// Returns penalty in [0.0, 10.0] (additive on objective, higher = worse).
fn calculate_trade_insufficiency_penalty(total_trades: usize, total_bars: usize) -> f64 {
    // Expected minimum trades: ~1 trade per 500 bars (8-hour avg hold at 1min)
    let min_expected = (total_bars / 500).max(10);

    if total_trades == 0 {
        10.0 // Model does nothing at all
    } else if total_trades < 10 {
        5.0 + 5.0 * (1.0 - total_trades as f64 / 10.0) // 5.0-10.0 range
    } else if total_trades < min_expected {
        2.0 * (1.0 - total_trades as f64 / min_expected as f64) // 0.0-2.0 smooth
    } else {
        0.0 // Sufficient trades → rely on Sharpe/Sortino
    }
}

Step 2: Wire penalty into extract_objective()

In extract_objective(), replace lines 3168-3230 (the if let Some(backtest) branch) with:

        let objective_total = if let Some(backtest) = &metrics.backtest_metrics {
            // TRADE INSUFFICIENCY PENALTY: gives PSO gradient across degenerate plateau
            let trade_penalty = calculate_trade_insufficiency_penalty(
                backtest.total_trades,
                backtest.total_trades.max(1) * 500, // Approximate total_bars from trade count context
            );

            // Short-circuit: if too few trades, skip composite score (it's all zeros anyway)
            if backtest.total_trades < 10 {
                info!(
                    "DEGENERATE TRIAL: {} trades → trade_penalty={:.2} (objective={:.2})",
                    backtest.total_trades, trade_penalty, trade_penalty
                );
                trade_penalty
            } else {
                // Component 1: Multi-objective composite score (60% weight)
                let composite_score =
                    0.4 * backtest.sortino_ratio +
                    0.3 * backtest.calmar_ratio +
                    0.2 * backtest.sharpe_ratio +
                    0.1 * backtest.omega_ratio;

                // Tail risk penalty
                let cvar_penalty = if backtest.cvar_95 < -0.05 { 10.0 } else { 0.0 };

                // Component 2: HFT activity score (25% weight)
                let hft_activity = calculate_hft_activity_score_wave10(buy_pct, sell_pct, hold_pct);

                // Combine: base objective + trade insufficiency penalty
                let base_objective =
                    -0.60 * composite_score + cvar_penalty + -0.25 * hft_activity + 0.15 * stability_penalty_raw;

                let objective = base_objective + trade_penalty;

                info!(
                    "OBJECTIVE: {:.4} = base {:.4} + trade_penalty {:.4} | trades={} composite={:.4}",
                    objective, base_objective, trade_penalty, backtest.total_trades, composite_score
                );

                objective
            }

Important: We need total_bars in extract_objective(). Since it's not in DQNMetrics, we approximate from trade context. Alternatively, we can use a constant (224K bars is the standard dataset). Let me check if we can just use a constant:

The training data is always the full validation set (~224K bars for ES.FUT). Rather than passing total_bars through metrics, use the simpler approach: hardcode min_expected = 100 (conservative — even 100 trades on 224K bars means the model is barely active, but at least it's distinguishable from 0).

Revised calculate_trade_insufficiency_penalty():

fn calculate_trade_insufficiency_penalty(total_trades: usize) -> f64 {
    const MIN_VIABLE_TRADES: usize = 100;

    if total_trades == 0 {
        10.0
    } else if total_trades < 10 {
        5.0 + 5.0 * (1.0 - total_trades as f64 / 10.0)
    } else if total_trades < MIN_VIABLE_TRADES {
        2.0 * (1.0 - total_trades as f64 / MIN_VIABLE_TRADES as f64)
    } else {
        0.0
    }
}

Step 3: Run check

Run: SQLX_OFFLINE=true cargo check -p ml Expected: compiles with 0 errors

Step 4: Commit

git add crates/ml/src/hyperopt/adapters/dqn.rs
git commit -m "feat(ml): add graduated trade insufficiency penalty to hyperopt objective

Different failure modes now produce unique objective values:
- 0 trades → 10.0 (model does nothing)
- 1 trade → 9.5 (single-direction collapse)
- 50 trades → 1.0 (insufficient activity)
- 100+ trades → 0.0 (rely on Sharpe/Sortino)

This breaks the flat 1.25 plateau that gave PSO zero gradient signal."

Task 4: Add unit tests for trade insufficiency penalty

Files:

  • Modify: crates/ml/src/hyperopt/adapters/dqn.rs (add tests to existing mod tests)

Step 1: Write tests

Add to the existing mod tests block (after the last #[test] fn):

    #[test]
    fn test_trade_insufficiency_penalty_zero_trades() {
        let penalty = calculate_trade_insufficiency_penalty(0);
        assert!((penalty - 10.0).abs() < 1e-6, "0 trades = max penalty: {}", penalty);
    }

    #[test]
    fn test_trade_insufficiency_penalty_one_trade() {
        let penalty = calculate_trade_insufficiency_penalty(1);
        assert!(penalty > 9.0 && penalty < 10.0, "1 trade near max: {}", penalty);
    }

    #[test]
    fn test_trade_insufficiency_penalty_graduated() {
        let p0 = calculate_trade_insufficiency_penalty(0);
        let p1 = calculate_trade_insufficiency_penalty(1);
        let p5 = calculate_trade_insufficiency_penalty(5);
        let p50 = calculate_trade_insufficiency_penalty(50);
        let p100 = calculate_trade_insufficiency_penalty(100);
        let p500 = calculate_trade_insufficiency_penalty(500);

        // Strictly decreasing (key property for PSO gradient)
        assert!(p0 > p1, "0 > 1: {} > {}", p0, p1);
        assert!(p1 > p5, "1 > 5: {} > {}", p1, p5);
        assert!(p5 > p50, "5 > 50: {} > {}", p5, p50);
        assert!(p50 > p100, "50 > 100: {} > {}", p50, p100);
        assert!((p100 - 0.0).abs() < 1e-6, "100+ = no penalty: {}", p100);
        assert!((p500 - 0.0).abs() < 1e-6, "500 = no penalty: {}", p500);
    }

    #[test]
    fn test_trade_insufficiency_no_flat_plateau() {
        // The whole point: different trade counts → different penalties
        let penalties: Vec<f64> = (0..20).map(|t| calculate_trade_insufficiency_penalty(t)).collect();
        for i in 0..19 {
            assert!(
                penalties[i] > penalties[i + 1] || (penalties[i] - penalties[i + 1]).abs() < 1e-6,
                "penalty[{}]={} should be >= penalty[{}]={}",
                i, penalties[i], i + 1, penalties[i + 1]
            );
        }
        // Verify no two adjacent values are the same (no plateaus in 0-10 range)
        for i in 0..9 {
            assert!(
                (penalties[i] - penalties[i + 1]).abs() > 0.01,
                "PLATEAU at trades={}: penalty[{}]={:.4} == penalty[{}]={:.4}",
                i, i, penalties[i], i + 1, penalties[i + 1]
            );
        }
    }

Step 2: Run tests

Run: SQLX_OFFLINE=true cargo test -p ml --lib hyperopt::adapters::dqn::tests::test_trade_insufficiency -- --nocapture Expected: all 4 tests pass

Step 3: Commit

git add crates/ml/src/hyperopt/adapters/dqn.rs
git commit -m "test(ml): add unit tests for trade insufficiency penalty gradient"

Task 5: Wire factored eval into hyperopt backtest loop

Files:

  • Modify: crates/ml/src/hyperopt/adapters/dqn.rs:2764-2790 (backtest loop)

Step 1: Replace legacy action collapse with factored eval

Replace lines 2772-2788:

                        let factored = crate::dqn::FactoredAction::from_index(action_idx)?;
                        let legacy = factored.to_legacy_action();
                        let action = match legacy {
                            crate::dqn::TradingAction::Buy => Action::Buy,
                            crate::dqn::TradingAction::Sell => Action::Sell,
                            crate::dqn::TradingAction::Hold => Action::Hold,
                        };

                        let bar = OHLCVBarF32 {
                            timestamp: bar_idx as i64,
                            open: close,
                            high: close,
                            low: close,
                            close,
                            volume: 0.0,
                        };
                        engine.process_bar(bar_idx, &bar, action);
                        ohlcv_bars.push(bar);

With:

                        let factored = crate::dqn::FactoredAction::from_index(action_idx)?;

                        let bar = OHLCVBarF32 {
                            timestamp: bar_idx as i64,
                            open: close,
                            high: close,
                            low: close,
                            close,
                            volume: 0.0,
                        };
                        engine.process_bar_factored(bar_idx, &bar, &factored);
                        ohlcv_bars.push(bar);

Step 2: Replace close_position with close_factored_position at backtest end

Replace line 2818 (engine.close_position(ohlcv_bars.len() - 1, last_bar);) with:

                    engine.close_factored_position(ohlcv_bars.len() - 1, last_bar);

Step 3: Run check

Run: SQLX_OFFLINE=true cargo check -p ml Expected: compiles with 0 errors

Step 4: Run full test suite

Run: SQLX_OFFLINE=true cargo test -p ml --lib -- --nocapture 2>&1 | tail -5 Expected: all tests pass (2500+)

Step 5: Run clippy

Run: SQLX_OFFLINE=true cargo clippy -p ml --lib -- -D warnings 2>&1 | tail -5 Expected: 0 warnings

Step 6: Commit

git add crates/ml/src/hyperopt/adapters/dqn.rs
git commit -m "feat(ml): wire factored eval into hyperopt backtest loop

Replace legacy Buy/Sell/Hold collapse with exposure-aware evaluation.
Long100→Long50 now generates a partial-close trade instead of being
a no-op. Combined with graduated trade penalty, PSO now gets gradient
signal across the entire 26D search space."

Task 6: Verify full workspace builds and run integration tests

Files:

  • No files modified (verification only)

Step 1: Full workspace check

Run: SQLX_OFFLINE=true cargo check --workspace Expected: compiles with 0 errors

Step 2: Run ml lib tests

Run: SQLX_OFFLINE=true cargo test -p ml --lib 2>&1 | tail -3 Expected: test result: ok. 2500+ passed; 0 failed

Step 3: Run clippy on workspace

Run: SQLX_OFFLINE=true cargo clippy --workspace -- -D warnings 2>&1 | tail -5 Expected: 0 errors, 0 warnings

Step 4: Commit (no changes, verification only)

No commit needed — this is verification only.