Files
foxhunt/docs
jgrusewski 25ba051157 spec(dqn): critical review pass — fix all majors, mediums, minors
Self-review identified 8 major + 8 medium + 15 minor issues. All fixed:

MAJORS:
  M1 Test 0.A: clarified — compare argmax(E[Q]), Boltzmann(E[Q]),
     and Thompson sampling distributions; assertion is on relative
     ordering across all three.
  M2 IQN quantile count: replaced hardcoded `5` with N_IQN_QUANTILES
     constant (defined per task #147 fixed-quantile design).
  M3 "Converged checkpoint" definition: ≥60 epochs trained AND
     val_sharpe stabilised (no >10% change over last 10 epochs).
     Cites prior 60-epoch validation runs (task #80, train-7rgqd).
  M4 R3 reframed: replaced "no issue" handwave with explicit
     by-design tradeoff acknowledgment + cost analysis. Wasted
     exploration is the cost of finding out whether edge exists.
  M5 Test 0.D σ_long=0.05 justified: chosen to match expected order
     of magnitude given typical |return| ~ 50bps; Phase 0.F
     validates against real checkpoint.
  M6 rng_ctr post-increment: clarified — matches existing pattern
     at experience_kernels.cu:858 (no behaviour change).
  M7 train_active_frac instrumentation: NEW Phase 2 deliverable —
     existing HEALTH_DIAG only has val_active_frac, but L3 verifies
     training-time active_frac. Spec now explicitly adds this
     ~10-line metrics.rs change as a Phase 2 deliverable.
  M8 eps_dir cleanup code-level detail: explicit reference to
     experience_kernels.cu lines 814-865; remove eps_dir from both
     static EPS_FLOOR clamp AND adaptive boost block; verify
     variable can be removed from kernel signature via grep.

MEDIUMS:
  Med1 Current C51/IQN combination: clarified that compute_expected_q
       blends per training schedule; Phase 2 replaces with explicit
       0.5*E_C51 + 0.5*E_IQN equal weighting; Phase 0.F verifies.
  Med2 Eval mode phrasing: "eval mode already sets eps=0 in existing
       kernel" — no semantic override, factually correct.
  Med3 Magnitude σ claim: clarified — magnitude branch likely has σ
       bias in OPPOSITE direction (Full has larger σ; UCB would
       prefer Full and worsen saturation). Empirical verification
       deferred. Phase 0.F should also report per-magnitude σ.
  Med4 Hierarchical sampling claim corrected: it's not about
       balancing 50/50 (already 50/50). It's about decoupling
       cluster-best decisions; clarified.
  Med5 n_atoms vs N_IQN_QUANTILES: clarified — n_atoms variable per
       config (currently 51); N_IQN_QUANTILES fixed at 5.
  Med6 Conviction code: removed pseudo-code; references existing
       implementation at experience_kernels.cu:1091; provides
       implementation hint for E[Q] reuse.
  Med7 Q-target propagation: clarified — uses full distribution
       (C51 atom projection / IQN quantile regression), not just
       E[Q]. Thompson modifies action selection only.
  Med8 References: added Thompson 1933 (original), Bellemare 2017
       (C51), Dabney 2018 (IQN) for theoretical foundations.

MINORS:
  Min1 Date updated to 2026-04-27.
  Min2-3 Argmax monotonic /2 simplified out — argmax(a+b) =
         argmax((a+b)/2). Code clarity improved.
  Min4 P(argmax picks Long) = 0 deterministic; reframed assertion.
  Min5 Test 0.F structural assertions added: σ_C51[FLAT] < 0.01 ×
       σ_C51[LONG]; same for IQN; E[Q_FLAT] > E[Q_LONG]; argmax
       picks FLAT; Thompson P(LONG)+P(SHORT) ≥ 0.20.
  Min6 -INFINITY → CUDART_INF_F (CUDA convention).
  Min7 dir_idx scope: comment notes it's declared earlier in kernel.
  Min8 action_select args: explicit — three buffers exist on GPU
       but not currently passed; new params, no new buffers.
  Min9 Phase 0 time math: 5 hours tests + 1 hour enumeration + 2
       hours 0.F + (3 hours runtime if checkpoint training needed,
       runs in parallel). Honest budget.
  Min10 "Two-stream" → "5-Layer Gate" header.
  Min11 Plan 5 reference uses full path consistently.
  Min12 Plan B time budget: explicit 1 day if pass; 2-5 days if bug.
  Min13 active_frac: clarified Long+Short combined, not per direction.
  Min14 train_active_frac: now in Phase 2 deliverables (see M7).
  Min15 "20 mechanisms" → 21, with sub-counts in section headers.

Spec now 615 lines, comprehensive coverage of:
  - Pearl + theoretical foundation
  - Problem statement (with bias-might-be-correct caveat)
  - Architecture (Thompson at training, argmax at eval)
  - Train vs eval distinct selectors with behavior-change disclosure
  - Direction-only scope with magnitude σ-bias warning
  - Conviction stays E[Q]-based (no Kelly cap jitter)
  - Interaction matrix: 21 mechanisms in 3 categories
  - 6 v2 enhancements documented + deferred
  - Phase 0/1/2/3 with tests, exit gates, time budgets
  - 5-layer verification + train_active_frac instrumentation
  - 8 risks with mitigations + 5 stop conditions
  - What v1 doesn't touch (referencing interaction matrix)
  - References (Thompson 1933, Bellemare 2017, Dabney 2018, etc.)
  - Aggregation contract (project-wide pearl, enforced)

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-27 00:08:02 +02:00
..