Exposes gradient_accumulation_steps in PpoHyperparameters so hyperopt and training binaries can configure effective batch scaling. The actual accumulation logic already existed in PPO::update_mlp(). Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>