- Add debug_assert_eq! guards in 4 train_baseline functions to catch bar/feature length misalignment at debug time (#4) - Remove "last sample targets itself" block in hyperopt PPO adapter that created ~0 return sample biasing toward HOLD (#5) - Align hyperopt state_dim 54→51 and num_actions 45→3 to match train_baseline architecture, making tuned hyperparams transferable (#6) - Use greedy_action() in evaluate_baseline PPO eval for deterministic results Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
29 KiB
29 KiB