Add q_gap_threshold to action selection kernel: when greedy Q(best) - Q(flat) < threshold, default to flat. Teaches model to trade only with conviction. 39D search space (was 38D). Default 0.0 (disabled), hyperopt range [0.0, 0.5]. Remove use_branching parameter from experience_action_select — GPU pipeline always uses branching DQN. Flat mode was dead code. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
9.6 KiB
9.6 KiB