jgrusewski
86b950df10
fix(ml): DQN hyperopt overhaul — 7 root cause fixes (B1-B3, C1-C4)
B1: Eval mode — disable noisy layer noise, use softmax action selection
(was greedy argmax, causing train/eval policy mismatch)
B3: Per-bar portfolio state sync in eval (was frozen within 1024-bar chunks)
C1: Extrinsic-only replay buffer — curiosity reward no longer stored
(was corrupting Q-values to learn novelty instead of trading P&L)
C2: Single exploration mechanism — noisy nets only. Removed count bonus
from Q-values and epsilon floor (was triple-stacking exploration)
C3: Neutral hold reward (0.0) — removed hold_penalty_weight from 31D→30D
search space (was biasing Q-values toward excessive trading)
C4: Re-enabled early stopping with adaptive plateau_window = epochs/2
2720 tests pass, 0 failures, 0 clippy warnings.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-06 10:23:51 +01:00
..
2026-03-01 22:10:43 +01:00
2026-02-27 01:33:18 +01:00
2026-02-27 01:33:18 +01:00
2026-03-06 10:23:51 +01:00
2026-03-06 00:32:09 +01:00
2026-03-06 08:42:03 +01:00
2026-03-06 00:32:09 +01:00
2026-03-06 08:42:03 +01:00
2026-03-06 00:32:09 +01:00