Setting dones[b]=1 in rl_drawdown_stop without actually closing the lobsim position created a lie: Q trained on false trade-closes while the position stayed open. Result: done_count=0 from step 99 onward, 990 trades total in 2000 steps (should be ~80k), training collapsed. The drawdown penalty (per-step min(0, unrealized_r) × rate) is kept — it creates exit gradient without corrupting the lobsim state. Hard stop-loss needs to override the ACTION (to FlatL/FlatS) BEFORE the lobsim step, not set dones after. Tracked for future implementation. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2.1 KiB
2.1 KiB