docs: generalization gap action plan (domain randomization + CQL + multi-fold)
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
This commit is contained in:
48
docs/superpowers/plans/2026-03-29-generalization-gap.md
Normal file
48
docs/superpowers/plans/2026-03-29-generalization-gap.md
Normal file
@@ -0,0 +1,48 @@
|
||||
# DQN Generalization Gap — Action Plan
|
||||
|
||||
## Problem
|
||||
In-sample Sharpe +0.57, OOS Sharpe -1.30 on H100. Model performs worse than random OOS.
|
||||
100K+ IS trades vs 211 OOS trades — model learned IS-specific trading patterns.
|
||||
|
||||
## Root Causes (from codebase analysis)
|
||||
1. **Deterministic episodes** — same start points every epoch, enables memorization
|
||||
2. **CQL alpha=0.1** — effectively disabled (1.5% of gradient budget)
|
||||
3. **Fixed simulation params** — spread=1bp, tx_cost=1.0, fill_prob=0.85 never vary
|
||||
4. **Single-fold validation** — 80/20 time split, no purging/embargo
|
||||
5. **No data augmentation** — pixel-perfect price sequences
|
||||
|
||||
## Execution Order
|
||||
|
||||
### Phase 1: Quick Wins (1-2 days)
|
||||
|
||||
**1a. Domain Randomization (Rust only, per-epoch)**
|
||||
- `training_loop.rs` lines 900-990: randomize before kernel launch
|
||||
- spread: U[0.5bp, 3.0bp], tx_cost: U[0.5, 2.5], fill probs: randomized
|
||||
- Episode starts: add ±25% stride jitter (line 935)
|
||||
- Variable episode length: U[250, 750] timesteps
|
||||
|
||||
**1b. CQL Alpha + Gradient Budget**
|
||||
- `config.rs` line 1509: `cql_alpha: 0.1 → 1.0`
|
||||
- `fused_training.rs` line 51: `CQL_GRAD_BUDGET: 0.15 → 0.25`
|
||||
- C51 budget: 0.70 → 0.60
|
||||
- Add `cql_alpha` warmup synchronized with c51_warmup
|
||||
|
||||
### Phase 2: Walk-Forward (3-5 days)
|
||||
|
||||
**2a. Enable GPU walk-forward by default**
|
||||
- `config.rs`: `enable_gpu_walk_forward: true`
|
||||
- Configure 6-8 folds with purge gap = n_steps bars
|
||||
- Embargo gap = 100 bars (DSR EMA window)
|
||||
- Model selection: median OOS Sharpe across folds
|
||||
|
||||
### Phase 3: Multi-Instrument (2-3 weeks)
|
||||
|
||||
**3a. ES + NQ dual training**
|
||||
- Shared trunk, instrument-specific advantage heads
|
||||
- Normalized rewards (vol-adjusted, instrument-agnostic)
|
||||
- Data loading for multiple .dbn.zst files
|
||||
|
||||
## Expected Impact
|
||||
- Phase 1: 50-70% generalization gap reduction
|
||||
- Phase 2: 30-50% PBO reduction
|
||||
- Phase 3: 8-12% OOS improvement
|
||||
Reference in New Issue
Block a user