Files
foxhunt/docs/superpowers
jgrusewski 308c54484e spec(dqn): full interaction matrix + outside-the-box Thompson improvements
Per-user direction: every existing mechanism in the DQN system MUST be
explicitly considered for interaction with Thompson, and Thompson itself
MUST be examined for system-specific improvements beyond vanilla.

INTERACTION MATRIX (3 categories, 20 mechanisms):

Category 1 — Compose with Thompson (no change required):
  Counterfactual reward, B.2 novelty bonus, PopArt, Saboteur,
  Curiosity, NoisyNets/VSN, Distillation, CQL, Polyak target EMA,
  HER, PER, Replay warm-start, Multi-fold validation harness.

Category 2 — Trivially adapt to Thompson (one-line changes):
  D7/N7 contrarian sign flip (negate the SAMPLE), cosine epsilon
  schedule (still applies to mag/ord/urg), per-sample epsilon (IQL
  expectile gap), adaptive Boltzmann tau (still applies to mag/ord/urg).

Category 3 — Take precedence over Thompson (hard constraints):
  Plan-based action lock (Thompson sample discarded if plan active),
  per-magnitude Kelly cap, trail stop, capital floor breach.

Critical insights from the audit:

1. NoisyNets is ALREADY a form of training-time Thompson at the
   parameter level. Output-space Thompson stacks on top —
   total exploration = parameter-space ⊗ output-space (multiplicative).

2. Curiosity is ORTHOGONAL to Thompson — Thompson explores actions
   whose Q is uncertain; curiosity explores states whose dynamics
   are uncertain. Both axes desirable; no conflict.

3. Plan lock takes precedence; same as currently with Boltzmann.

OUTSIDE-THE-BOX v2 ENHANCEMENTS (deferred to follow-up specs):

  v2.1 Triple-source Thompson (C51 + IQN + Ensemble) — incorporate
       the existing ensemble Q-head as 3rd uncertainty source.
  v2.2 Persistent Thompson (anti-churn for HFT) — bias sampling
       toward current direction, ISV-driven; reduces tx_cost from
       Long/Short oscillation across bars.
  v2.3 CVaR-aware eval (risk-adjusted deployment) — eval picks
       argmax(E[Q] − λ·CVaR_α[Q]); risk-aware decision making for
       production with real capital.
  v2.4 Information-Directed Sampling (Russo & Van Roy 2014) — picks
       action minimizing regret²/info_gain; more efficient than
       vanilla Thompson when learning saturates.
  v2.5 Hierarchical Thompson on (trade vs no-trade) → (which
       direction) — addresses 50/50 structural advantage of no-trade.
  v2.6 Composition with curiosity-driven exploration — explicit
       coupling beyond reward-side composition.

Each v2 enhancement gets its own spec/plan when prioritised. Vanilla
Thompson is v1; ships first; verified independently.

ALSO FIXED (from earlier self-review):

  - Pearl claim softened: only the C51 Hold/Flat bias is directly
    attributed; other historic bugs had different mechanisms.
  - TFT entry removed from contract table — TFT is a Variable
    Selection Network (feature processor), not a Q-head. Replaced
    with generic "Future Q-head additions" placeholder.
  - Eval direction = argmax E[Q] explicitly flagged as a behavior
    change from current Boltzmann-with-tau (val_dir_dist will be
    more concentrated than current).
  - Phase 0.E budget reduced (1-2 hours, not 1 day) — synthetic
    bandit is ~50 lines of Rust, not full RL training loop.
  - Phase 0 enumerates existing checkpoints before training new one.
  - Architecture diagram parenthesis fixed.
  - Conviction implementation note: compute E[Q] once, reuse for
    conviction AND eval-mode argmax — no redundant computation.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-26 23:51:17 +02:00
..