Files
foxhunt/docs
jgrusewski 1c63c32297 spec(dqn): reframe Phase 3 as direction-branch dead-code cleanup
User correctly challenged the "band-aid removal" framing. All fixes
shipped during the val-Flat-collapse investigation addressed real bugs
at their respective layers and should be preserved:

  - Kelly cap warm-branch (0c9d1ee39): post-decision physics layer.
    Thompson-independent. KEEP.
  - Train Return display + Sharpe annualization (non-tau parts of
    7a3d88646): display/metric layer. Thompson-independent. KEEP.
  - Direction Boltzmann tau-floor (tau part of 7a3d88646) +
    adaptive eps_dir floor (d54b49efc): gates inside direction-branch
    action selection. Phase 2 replaces direction-branch action
    selection wholesale (eps-greedy + Boltzmann → Thompson), so these
    direction-only code paths become structurally unreachable.

The latter two are NOT band-aids being removed because Thompson is
better. They are dead code being cleaned up because Thompson replaces
the surrounding mechanism. Magnitude/order/urgency branches keep their
existing eps-greedy + Boltzmann + tau-floor + EPS_FLOOR paths intact.

Reframed Phase 3 deliverable: "direction-branch dead-code cleanup"
with explicit rationale (per feedback_no_legacy_aliases.md and
feedback_no_partial_refactor.md). 0.5 day budget instead of 1.

Also clarified eval action selection: argmax of (E[Q_C51]+E[Q_IQN])/2
is correct. Bellman backup is a Q-learning UPDATE rule, not an
action-selection rule. Once Q is learned, optimal policy is greedy
argmax of learned Q. Online Bellman lookahead at eval would require
a forward model of market dynamics — not available, not standard.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-26 23:41:05 +02:00
..