Proves PPO training pipeline works end-to-end on production-sized state (54 features). Key insight: critic_lr=1e-4 (10x lower than default) prevents value loss divergence and shows 31.4% reduction. Assertions: epochs completed, value loss bounded (<1000), all losses finite, policy loss bounded by clipping, checkpoints saved, explained variance not catastrophic. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
9.2 KiB
9.2 KiB