Sharpe calculation:
- Use per-trade returns (sum_returns/sum_sq_returns) instead of per-bar
step_returns. Per-bar Sharpe collapsed variance → bogus 19.75 with PF=0.03.
- Annualize by sqrt(trades_per_year) not sqrt(bars_per_year).
Batch sizing:
- Cap auto-computed batch_size at 8192 (VRAM ceiling of 2M is OOM limit, not
optimal RL batch).
- Add VRAM floor: batch_size < ceiling/4096 gets scaled UP (128 → 512 on H100).
- Only let hyperopt override batch_size when explicitly non-zero — preserve
profile's batch_size=0 auto-compute sentinel.
Replay buffer:
- Divide per_max_memory_bytes by 3.0 for regime heads (PER budget was 3x too
large, causing OOM cascade 74M → 37M → 18M → 9M → 4.6M).
- HER buffer uses original_buffer_size (pre-autosizer), not inflated 74M.
Q-value clipping:
- Wire hyperparams.q_clip_min/max to DQNConfig (was hardcoded ±500, production
TOML has ±50). Prevents Q-value overestimation ratio of 94.6x at epoch 2.
Training stability:
- Anti-LR warmup: skip first 5 epochs (early Sharpe unreliable from random
policy). Prevents bogus 3x LR boost at epoch 2.
- min_replay_size from profile (1000), not hyperopt batch_size (128).
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>