Files
foxhunt/crates/ml/src
jgrusewski 5f2e8930d3 feat: trade lifecycle fix — 3-layer hold enforcement (action masking + kernel override + reward scaling)
4 bugs fixed:
1. Hold guard was dead code (!exiting_trade always false) — replaced with unified guard
2. Reversals bypassed hold — new guard covers exits AND reversals
3. Single-bar trades got full reward — reward scaled by hold_time/min_hold_bars
4. No action masking — experience_action_select now forces current exposure during hold

3 layers of defense:
- Layer 1: Action masking in experience_action_select (masks Q-values + blocks random exploration)
- Layer 2: Kernel override in experience_env_step (catches epsilon exploration slipthrough)
- Layer 3: Reward scaling (smooth gradient: 10% floor to 100% at min_hold_bars)

Config: min_hold_bars added to DQNHyperparameters (default 5), TOML, hyperopt [2,20].
Trailing stop respects min_hold_bars (was hardcoded > 2.0f).
Search space: 46D (was 45D).

Also fixes flaky test_batch_hierarchical_vs_flat_softmax_diversity assertion.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-25 23:32:15 +01:00
..