Revises the C51-bias spec after deeper review surfaced 12 design gaps,
of which 4 were critical:
1. Train-only vs train+eval ambiguity — UCB at eval would conflate
"model recommends Long" with "model is uncertain about Long",
inflating reported edge. CRITICAL for trading where eval drives
real capital decisions.
2. Thompson sampling is more principled than UCB:
- parameter-free (no κ to tune)
- uses distribution directly without scalar reduction
- naturally explore-exploit balanced via distribution shape
3. c51_alpha is the wrong blend weight (it's C51-vs-MSE-warmup, not
C51-vs-IQN). Equal-weight average of C51 and IQN samples is the
structural choice — no tuned blend weight needed.
4. The bias might be CORRECT BEHAVIOUR — model rationally choosing
Flat when no edge has been discovered. Phase 0 must include a
synthetic-edge test (controlled MDP with KNOWN positive Long
expected value) to verify Thompson can discover edge if it exists.
Other gaps fixed:
- Eval at argmax E[Q] (not Boltzmann, not Thompson)
- Pearl wording broadened to cover ensembles + future methods
- Ensemble Q-head added to aggregation contract table
- Explicit caveat: NEVER extend Thompson to magnitude branch (would
worsen existing magnitude saturation)
- Phase 0.F uses CONVERGED checkpoint (≥30 epochs), not 2-epoch run
- L4 long smoke (30 epochs, ~1 hour) added — Thompson edge discovery
needs longer feedback loop than 5 epochs
- Phase 3 explicitly removes eps-floor + tau-floor band-aids
(Thompson replaces direction Boltzmann; band-aids become dead code)
- Conviction stays E[Q]-based, not sample-based (avoid Kelly cap
jitter from stochastic samples)
Architecture (Thompson only, no UCB):
TRAINING: dir_idx = argmax(0.5 × (sample_C51(d) + sample_IQN(d)))
magnitude/order/urgency: existing Boltzmann + ε-greedy
EVAL: dir_idx = argmax(0.5 × (E[Q_C51] + E[Q_IQN]))
magnitude/order/urgency: existing Boltzmann (eval mode)
Direction-branch ε-greedy + Boltzmann are REMOVED — Thompson is the
exploration mechanism. No new GPU buffers; existing C51 atoms + IQN
quantiles passed to action_select.
5-7 days active work across 4 sub-plans; each gets its own
writing-plans cycle.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>