Two issues causing all 20 hyperopt trials to have identical f64::MAX
objective (TPE optimizer blind):
1. Val-loss plateau early stopping fired at epoch 5-6 of every 8-epoch
trial (plateau_window=5 too aggressive for short runs). Disabled
early_stopping_enabled for hyperopt; gradient-collapse patience
still active as safety net.
2. Penalty metrics used f64::MAX for gradient_norm/q_value_std which
produced ~3.6e+308 objective. Changed to 100.0 so TPE can still
differentiate between early-stopped trials by other metric fields.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>