4 root causes of 45-action DQN collapsing to 1-6 actions:
1. Batch epsilon ignoring noisy_epsilon_floor: select_actions_batch()
and select_actions_batch_gpu() used get_epsilon() which returns 0.0
with noisy nets — zero random exploration in the training path.
Added get_effective_epsilon() that respects noisy_epsilon_floor.
2. Entropy coefficient too weak: default 0.05 with bounds (0.01, 0.2)
produced max ~0.19 bonus vs TD loss of 4+. Bumped default to 0.1,
widened bounds to (0.05, 0.5) for effective anti-collapse.
3. count_bonus_coefficient not in search space: was hardcoded at 0.1
in from_continuous(). Promoted to 31st search dimension with bounds
(0.05, 1.0) so PSO/TPE can optimize exploration strength.
4. Diversity penalty too coarse: objective had <10 unique actions
short-circuit but nothing for 10-20. Added graduated penalty that
linearly ramps from 3.0 (10 actions) to 0.0 (20 actions).
Also fixes pre-existing clippy impl_trait_in_params in optimizer.rs.
2720 tests pass, 0 clippy warnings.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>