Create PpoStrategy and PpoLstmStrategy implementing ValidatableStrategy to run PPO through walk-forward validation with DSR, PBO, and permutation tests. Both variants validated on real 6E.FUT data (29,937 bars, 15 folds). Key implementation detail: LSTM hidden states are detached from the computation graph after each step to prevent stack overflow from unbounded graph growth across 30k+ sequential forward passes. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
19 KiB
19 KiB