A: Switch backtest eval from Gumbel softmax to greedy argmax (batch_greedy_actions)
so hyperopt Sharpe reflects the agent's actual learned policy, not noisy sampling.
C: Disable reward normalization (enable_normalization=false). EMA normalizer with
±3.0 clipping was flattening the reward landscape, preventing the agent from
distinguishing large winners from scratch trades.
D: Wire tx_cost_bps (0.1 bps for IBKR ES) through to EvaluationEngine via
new_with_fee_rate(). Previously hardcoded at 15 bps (150x mismatch with actual
commission costs), massively penalizing every trade in backtest.
E: Scale PnL reward by agent's target exposure in calculate_pnl_reward().
Previously, a Short100 action received POSITIVE reward when market went up
(pct_return ignored position direction). Now: reward = pct_return × exposure.
2735 tests pass, 0 clippy warnings.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>