5 new pearls for closing val/OOS gap to <15%: G11: Q-value anchoring to Flat baseline (focus on alpha) G12: Predictive coding auxiliary loss (self-supervised trunk) G13: Sharpe-aware reward shaping (align reward with goal) G14: Confidence-weighted replay (suppress noise samples) G15: Action commitment penalty (anti-churn beyond costs) Total: 15 generalization components, ~550 LOC. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>