Closes the patience-based early-stopping bug that ran xmd6b 30 epochs
past peak val performance (epoch 2: val_Sharpe=90, total_pnl=0.44 →
epoch 30: val_Sharpe=26, total_pnl=0.16) — a 70% loss of alpha to
training-induced overfitting.
Two intertwined bugs, fixed atomically:
T1.1a (wrong-source) at training_loop.rs:7234:
- was: self.early_stopping.should_stop(-log_output.epoch_sharpe, epoch)
reading the TRAINING ROLLOUT Sharpe (Thompson-noisy, in-sample,
oscillates even when the model is frozen)
- now: self.early_stopping.should_stop(log_output.val_loss, min_delta, epoch)
reading the deterministic-backtest val_loss
- The comment 5733 lines earlier (line 1499) explicitly says "Use
val_Sharpe (deterministic backtest), NOT epoch_sharpe" — patience
path was the inconsistency, backtracking already honored it.
T1.1b (hardcoded threshold) in early_stopping.rs:
- was: EarlyStopping::new(patience, min_delta) with min_delta=0.001
constructor-set, struct field, structurally meaningless
against the val_loss noise floor (typical val_sharpe deltas
are O(1-10), so 0.001 essentially never gates)
- now: EarlyStopping::new(patience), should_stop(val_loss, min_delta,
epoch) with min_delta computed per-call from
ISV[VAL_SHARPE_VAR_EMA_INDEX=351] as
sqrt(var_ema).max(0.5)
- Floor 0.5 covers cold-start before var_ema bootstraps from sentinel
per pearl_blend_formulas_must_have_permanent_floor.
- Per feedback_isv_for_adaptive_bounds + feedback_adaptive_not_tuned.
T2.3 (test signature update) absorbed:
- 6 existing unit tests migrated to new should_stop signature.
- 1 NEW test (test_min_delta_can_change_per_call) verifying per-call
threshold change works correctly.
- EarlyStopping::min_delta struct field deleted.
- Atomic per feedback_no_partial_refactor.
Affected files:
- crates/ml/src/trainers/dqn/early_stopping.rs (struct + tests)
- crates/ml/src/trainers/dqn/trainer/constructor.rs (new() arg)
- crates/ml/src/trainers/dqn/trainer/training_loop.rs (call site)
Verification:
- cargo check -p ml --tests: passes
- cargo test -p ml --lib early_stopping: 8/8 pass
Behavioral expectation post-fix: xmd6b-shape runs (val_Sharpe rising
31→90 epochs 0-2, declining 90→26 epochs 3-30) will trigger early-stop
near the peak. With patience=5 and var_ema bootstrapping by epoch 2-3,
the controller detects "no improvement of ≥ 1σ for 5 consecutive
epochs" by ~epoch 7-8 and stops, saving ~22 epochs of overfitting.
Plan reference: docs/plans/2026-05-10-sp21-train-eval-coherence-isv-defrost.md
Tier 1 status: T1.1a ✓, T1.1b ✓, T2.3 ✓ (this commit). T1.2 + T1.4 next.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>