Files
foxhunt/docs/plans/2026-02-20-validation-stack-design.md
jgrusewski c3b5e124f0 chore: update .gitignore and add design plan docs
Ignore ML checkpoints, trained model safetensors, stray ml/ml/ dir,
and .claude/worktrees/. Clean up duplicate hive-mind-prompt entries.
Add 17 design/implementation plan docs from 2026-02-20 to 2026-02-22.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-24 00:30:15 +01:00

264 lines
9.5 KiB
Markdown

# Statistical Validation Stack Design
## Goal
Build an algorithm-agnostic statistical validation harness that answers: **"Is this trading strategy statistically real, or overfit noise?"** using walk-forward validation, Deflated Sharpe Ratio, Probability of Backtest Overfitting, Monte Carlo permutation tests, and per-regime performance breakdown.
## Context
The foxhunt codebase has working DQN (Rainbow + IQN + CQL), PPO, TFT, and Mamba2 models with comprehensive backtesting metrics (Sharpe, Sortino, Calmar, VaR, CVaR, Omega) and hyperopt infrastructure — but no statistical validation proving any of it actually works out-of-sample. Walk-forward validation is stubbed (`walk_forward_split()` returns `vec![]`). DSR, PBO, and permutation tests don't exist.
## Architecture
Layered pipeline with trait-based strategy interface. Five independent modules composed by a harness orchestrator:
```
ml/src/validation/
├── mod.rs # ValidatableStrategy trait, TimeSeriesData, re-exports
├── walk_forward.rs # Walk-forward splitter with embargo periods
├── statistical.rs # DSR + PBO (CSCV) + Monte Carlo permutation tests
├── regime_analysis.rs # Per-regime Sharpe breakdown (uses existing regime module)
└── harness.rs # Orchestrator: WF → evaluate → stats → regime → report
```
### Relationship to Existing Modules
- `ml/src/evaluation/` — Per-trade metrics (Sharpe, Sortino, etc.). Used _within_ each walk-forward fold.
- `ml/src/backtesting/` — Barrier method backtesting. Orthogonal, not replaced.
- `ml/src/regime/` — Regime detection (trending, ranging, volatile, CUSUM, Bayesian changepoint). Reused directly for per-regime analysis.
- `ml/src/hyperopt/` — Trial count feeds into DSR's `num_trials_tested` parameter.
## Component Design
### 1. ValidatableStrategy Trait
```rust
pub trait ValidatableStrategy: Send {
/// Train the model on the given time-series data
fn train(&mut self, data: &TimeSeriesData) -> Result<(), MLError>;
/// Evaluate: generate daily PnL returns on test data (post-training)
fn evaluate(&self, data: &TimeSeriesData) -> Result<Vec<f64>, MLError>;
/// Human-readable name for reporting
fn name(&self) -> &str;
/// Reset model weights for fresh fold training
fn reset(&mut self) -> Result<(), MLError>;
}
```
Algorithm-specific wrappers implement this trait:
- `DqnStrategy` wraps `DQN` — calls `train_step()` and `select_action()`
- `PpoStrategy` wraps PPO trainer
- `TftStrategy` wraps TFT
- Future models implement the same trait
### 2. TimeSeriesData
```rust
pub struct TimeSeriesData {
pub timestamps: Vec<DateTime<Utc>>,
pub features: Vec<Vec<f32>>, // [num_bars, feature_dim]
pub prices: Vec<f64>, // close prices for PnL computation
pub returns: Vec<f64>, // daily log returns (derived)
}
```
Constructable from existing MBP10 loader output. The `features` field matches what `DQNConfig.state_dim` expects (51 features).
### 3. Walk-Forward Splitter
```rust
pub struct WalkForwardConfig {
pub train_bars: usize, // training window size in bars
pub test_bars: usize, // test window size in bars
pub embargo_bars: usize, // gap between train end and test start
pub step_bars: usize, // advance per fold
pub min_train_samples: usize, // minimum bars required for training
}
pub struct Fold {
pub fold_index: usize,
pub train_range: Range<usize>, // index range into TimeSeriesData
pub embargo_range: Range<usize>,
pub test_range: Range<usize>,
}
pub fn walk_forward_split(
num_bars: usize,
config: &WalkForwardConfig,
) -> Vec<Fold>;
```
**Embargo period**: Prevents information leakage from features with lookback windows. If you use 20-bar moving averages, embargo should be >= 20 bars.
### 4. Deflated Sharpe Ratio (DSR)
Reference: Bailey & Lopez de Prado (2014), "The Deflated Sharpe Ratio"
```rust
pub struct DsrResult {
pub observed_sharpe: f64,
pub expected_max_sharpe: f64, // E[max(SR_1, ..., SR_N)] under null
pub deflated_sharpe: f64, // DSR statistic
pub pvalue: f64, // P(SR* > SR_obs | H0)
}
pub fn deflated_sharpe_ratio(
observed_sharpe: f64,
num_trials: usize, // hyperopt trials run
sharpe_variance: f64, // variance of per-fold Sharpe estimates
skewness: f64, // skewness of returns
kurtosis: f64, // excess kurtosis of returns
num_observations: usize, // total test bars
) -> DsrResult;
```
Key formula: `DSR = Phi[(SR_obs - SR*) / sigma_SR]` where:
- `SR* = sqrt(V[SR]) * ((1 - gamma) * Phi_inv(1 - 1/N) + gamma * Phi_inv(1 - 1/(N*e)))` (approximation from the paper)
- `sigma_SR = sqrt((1 - skew*SR + (kurt-1)/4 * SR^2) / (num_obs - 1))`
- `gamma = 0.5772...` (Euler-Mascheroni constant)
### 5. Probability of Backtest Overfitting (PBO)
Reference: Bailey et al. (2017), "Probability of Backtest Overfitting"
Uses Combinatorially Symmetric Cross-Validation (CSCV):
```rust
pub struct PboResult {
pub pbo: f64, // probability of overfitting [0, 1]
pub num_combinations: usize, // C(N, N/2) combinations tested
pub logit_distribution: Vec<f64>, // distribution of logit(rank) values
}
pub fn probability_of_backtest_overfitting(
per_fold_returns: &[Vec<f64>], // returns per fold (N folds)
strategy_sharpes: &[f64], // Sharpe per fold
) -> PboResult;
```
Algorithm:
1. Given N folds, generate C(N, N/2) combinations
2. For each combination: IS half = training folds, OOS half = test folds
3. Compute performance metric (Sharpe) on IS and OOS halves
4. Find IS-best strategy, measure its OOS rank
5. Compute logit(rank): `logit = ln(rank / (N - rank))`
6. PBO = proportion of combinations where logit > 0 (IS-best underperforms OOS median)
### 6. Monte Carlo Permutation Tests
```rust
pub struct PermutationResult {
pub observed_sharpe: f64,
pub null_sharpes: Vec<f64>, // Sharpe from each permutation
pub pvalue: f64, // fraction >= observed
pub num_permutations: usize,
}
pub fn permutation_test(
daily_returns: &[f64],
num_permutations: usize, // default: 10_000
seed: u64, // reproducibility
) -> PermutationResult;
```
Shuffles the daily returns time series (destroying temporal structure and any signal), recomputes Sharpe on shuffled data. The p-value = fraction of permuted Sharpes >= observed Sharpe. This is independent from DSR — DSR corrects for multiple testing, permutation tests verify the signal itself.
### 7. Per-Regime Analysis
```rust
pub struct RegimeMetrics {
pub regime: RegimeType, // from existing regime module
pub sharpe: f64,
pub num_bars: usize,
pub win_rate: f64,
pub avg_return: f64,
}
pub fn per_regime_breakdown(
daily_returns: &[f64],
timestamps: &[DateTime<Utc>],
prices: &[f64],
) -> HashMap<RegimeType, RegimeMetrics>;
```
Uses the existing `RegimeType` enum from `ml/src/dqn/regime_conditional.rs` (Trending, Ranging, Volatile). Labels each bar using the existing regime detection infrastructure, then groups returns by regime and computes per-regime metrics.
### 8. ValidationReport
```rust
pub struct ValidationReport {
pub strategy_name: String,
// Walk-forward
pub per_fold_sharpes: Vec<f64>,
pub aggregate_sharpe: f64,
pub num_folds: usize,
// DSR
pub deflated_sharpe_ratio: f64,
pub dsr_pvalue: f64,
pub num_trials_tested: usize,
// PBO
pub pbo: f64,
pub pbo_num_combinations: usize,
// Monte Carlo
pub monte_carlo_pvalue: f64,
pub num_permutations: usize,
// Regime breakdown
pub per_regime_metrics: HashMap<RegimeType, RegimeMetrics>,
// Overall verdict
pub verdict: ValidationVerdict,
}
pub enum ValidationVerdict {
Pass, // DSR p < 0.05 AND PBO < 0.25 AND MC p < 0.05
Marginal, // At least one test passes at relaxed threshold
Fail, // No statistical evidence of real signal
}
```
### 9. Harness Orchestrator
```rust
pub struct ValidationHarness {
pub wf_config: WalkForwardConfig,
pub num_permutations: usize,
pub num_trials: usize,
}
impl ValidationHarness {
pub fn validate<S: ValidatableStrategy>(
&self,
strategy: &mut S,
data: &TimeSeriesData,
) -> Result<ValidationReport, MLError> {
// 1. Split data via walk-forward
// 2. For each fold: reset strategy, train, evaluate, collect returns
// 3. Compute per-fold Sharpe
// 4. Compute DSR from aggregate stats
// 5. Compute PBO via CSCV on fold returns
// 6. Run Monte Carlo permutation test on concatenated test returns
// 7. Label bars with regime, compute per-regime metrics
// 8. Determine verdict
// 9. Return ValidationReport
}
}
```
## Dependencies
- `chrono` — timestamps (already in workspace)
- `rand` — permutation shuffling and CSCV sampling (already in workspace)
- Standard normal CDF — implement via rational approximation (no external dep needed, ~10 lines of code)
## Success Criteria
1. Walk-forward splitter produces correct non-overlapping folds with embargo
2. DSR correctly penalizes strategies found via many hyperopt trials
3. PBO > 0.5 on random strategies, PBO < 0.25 on synthetically profitable ones
4. Permutation test p < 0.05 for strategies with genuine signal
5. Per-regime breakdown matches known market conditions
6. All existing 1823 tests continue passing
7. Algorithm-agnostic: DQN and PPO both work through ValidatableStrategy trait