84de278dfee5b1d12f480a2eaabc42e9e6213c19
180 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
40e737a181 |
docs(sp14): spec v3 — K=4 direction + ISV-adaptive β rate limiter
Two further user-flagged corrections from second-critical review:
1. Direction Q-head emits K=4 actions, NOT K=3.
Authoritative source: state_layout.cuh:123-126
#define DIR_SHORT 0 // open/maintain short
#define DIR_HOLD 1 // keep current position (no-op)
#define DIR_LONG 2 // open/maintain long
#define DIR_FLAT 3 // close all to zero
The "default: 3 — Short/Flat/Long" comment at gpu_dqn_trainer.rs:2438
is stale pre-SP13. Production callsites all set branch_0_size: 4.
SP13 added DIR_HOLD as a separate fourth direction action (Hold-pricing).
The config default comment was never updated when DIR_HOLD landed.
This is exactly the feedback_trust_code_not_docs failure mode — a
single stale comment would have silently corrupted the q_disagreement
signal (Hold/Flat being indices 1/3 instead of just Flat=1).
Updates:
- B.1 action-space context: 4 actions with Hold AND Flat both
non-committal (Hold = keep position, Flat = exit to zero)
- B.2.3 q_disagreement mapping: K=4↔K=2 with both Hold and Flat
masked from disagreement signal (no new directional commitment
to evaluate); only Short and Long picks contribute to disagreement
- Edge case handling for all-Hold/all-Flat batches
2. Adaptive β rate limiter (was structural β=0.9).
Per feedback_isv_for_adaptive_bounds, β should be signal-driven
not hardcoded. v3 derives β from variance of α_grad_raw, mirroring
the k_aux/k_q variance-driven steepness pattern in B.2.5.
Formula:
β = clip(β_base + variance_alpha_raw / variance_ref_alpha,
[β_base, β_max])
β_base = 0.5 (light smoothing baseline; ~2-step half-life)
β_max = 0.95 (heavy smoothing; ~20-step half-life)
Stable α_grad_raw → β = β_base (preserves directional intent)
Volatile α_grad_raw → β → β_max (dampens jitter)
Adds 2 ISV slots:
ALPHA_GRAD_RAW_VARIANCE_EMA_INDEX (Welford variance)
BETA_RATE_LIMITER_ADAPTIVE_INDEX (current β value)
ISV slot count: 11 → 13 (net +2 for variance + adaptive β).
LOC estimate: ~1150 → ~1180 (negligible delta; 3 Welford
variances now in alpha_grad_compute_kernel instead of 2).
HEALTH_DIAG pearl_egf_diag emit updated to expose all three
adaptive scalars (α, β, k_aux/k_q) plus all three driving
variances (var_alpha, var_aux, var_q) for full observability.
Verified against current code at HEAD
|
||
|
|
d243a6f080 |
docs(sp14): spec v2 — critical review corrections
8 corrections from code-anchored critical review at HEAD
|
||
|
|
eaf4adcb98 |
docs(sp14): spec — aux→Q wire + Earned Gradient Flow pearl
Designs the SP14 chain on top of SP13 Layer B (HEAD
|
||
|
|
ba83fcd1f5 |
docs(sp13): v3 spec + plan — Hold-pricing replaces Hold-elimination
P0a.T3 v2 implementer's audit revealed `DirectionAction` enum doesn't exist; the codebase uses an 8-variant fused `ExposureLevel` (ShortSmall/Half/Full, Hold, LongSmall/Half/Full, Flat) with cross-crate consumers across 77 files and 32+ test files pinning the 8-variant invariant. Atomic Hold elimination would cascade massively. User insight (2026-05-04): Hold being FREE is the bug, not Hold itself. MFT trading legitimately needs multi-bar holds; we want the model to use them deliberately, not as a CQL-bias lazy default. Holding isn't free in the real world — broker fees, margin interest, opportunity cost. v3 reframes as Hold-pricing: - 4-way action space stays; ExposureLevel::Hold stays; no cross-crate cascade - 3 new ISV slots (380-382): HOLD_COST_INDEX, HOLD_RATE_TARGET_INDEX, HOLD_RATE_OBSERVED_EMA_INDEX - Hold-rate observer: small GPU kernel + Pearls A+D smoothing - Hold-cost controller: 5-line deficit-driven formula (excess > target → cost rises 1×→5× base; observed ≤ target → relax) - Per-bar reward subtraction at action == DIR_HOLD site - 2 GPU oracle tests for the controller P0a.T3 cuts from ~250 LOC + 32-test cascade → ~120 LOC additive. T1+T2 already-staged work unchanged. T4/T5/Layer B/C/D structure preserved. Tension with pearl_event_driven_reward_density_alignment acknowledged in spec — per-bar Hold cost is exposure-NEGATIVE (away from Hold), models real economic carry, ISV-bounded by controller. Inverse of the pearl's failure mode. Faithful reward modeling, not artificial shaping. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
0ca45ef61d |
docs(sp13): v2 spec + plan after critical review
Spec v2 supersedes v1 with five major fixes from the critical review: - Drop alpha-vs-benchmark reward (mathematically tautological — Long trades had alpha = -costs always against always-long benchmark) - Split Phase 0 into 0a (Hold-only, clean test of user hypothesis) + 0b (aux amplification probe, only if 0a partial), avoiding the v1 confounded experiment - Bound direction-skill bonus relative to |alpha| (cap_ratio × |alpha|) to prevent reward gaming on small-alpha correct-direction losers; spec adds 8-quadrant worked-example matrix verifying the no-negative-EV invariant - Replace 5-epoch absolute-threshold gate with 10-epoch trajectory criterion to avoid false-negatives from aux head underconvergence - Aux_w controller adds dual-EMA stagnation detector (decays toward base when no improvement) — prevents permanent destabilization of Q-head in data-limited case Plan v2 mirrors spec changes: - Phase 0a (Hold elimination + dir_acc instrumentation, ~300 LOC, 1 atomic commit) - Phase 0b (aux_w controller replacement, conditional, ~80 LOC) - Layer B (aux head regression -> binary classification, ~120 LOC) - Layer C (skill bonus + luck discount with calibrated bounds, ~150 LOC) - Layer D (30-epoch validation + 3 new pearls) ISV slot allocation [372..380): drops slot 371 (BENCHMARK_PNL_CUMULATIVE), adds 374 (AUX_DIR_ACC_LONG_EMA — stagnation), 379 (SKILL_BONUS_CAP_RATIO). Plan integrates Explore-agent touch-list (50-70 sites) with corrected enum ordering: Short=0, Long=1, Flat=2 (preserves codebase Short-first convention, avoids 30+ stale-comment churn). Adds direction-bias signal handling at experience_kernels.cu:5159 (Hold's 0.5 softening gate vanishes -> [1, 1, 0]). Three new pearls planned for Layer D close-out: - pearl_redefine_success_for_predictive_skill - pearl_skill_bonus_must_be_alpha_bounded (calibration lesson) - pearl_reward_quadrant_audit_required (meta-pearl) Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
ecab584c32 |
spec(sp12): per-trade event-driven reward composition (v3)
Three architectural changes in one unified design to push the model from
HFT noise extraction (62% trade rate, sharpe-gaming) toward MFT alpha-hunting:
1. Asymmetric bounded cap (-10/+5) — restores loss aversion erased by
SP11 symmetric cap. 2:1 ratio matches Kahneman/Tversky prospect theory.
Anchor: pearl_audit_unboundedness_for_implicit_asymmetry.
2. Min-hold soft penalty with temperature curriculum — patience requirement
at exit. Soft factor = deficit/(deficit+T), T anneals 50→5 over 50 epochs.
Forces commitment without paralysis in early training.
3. Zero per-bar shaping (gate micro/opp_cost on events) — eliminates
continuous-reward gradient that pulls toward continuous exposure.
Anchor: pearl_event_driven_reward_density_alignment.
Combined: reward fires only on trade events with prospect-theory loss
aversion + commitment requirement. Pure per-trade event-driven Q-learning
properly aligned with per-trade P&L objective.
~50 LOC across 3-4 files. No new ISV slots in Phase 1 (constants only).
Cost ~€1.30 (€0.30 smoke + €1.00 30-epoch validation).
Empirical motivation: train-multi-seed-pmbwn 50-epoch on commit
|
||
|
|
85069ba75e |
spec(sp11): amend B1b smoke-recovery — z-score normalization for mag-ratio canary
smoke-test-6wd2c on commit |
||
|
|
9395b983cc |
spec(sp11): fix all 13 review issues — straight-up implementation
Critical bugs fixed:
1. Z-score formula: was mean(Δ)/mean(|Δ|), bounded to [-1,+1] — sigmoid
never saturated, controller stuck in [0.27, 0.73]. Now true Z-score:
delta_ema / sqrt(delta_var_ema) with two-pass var. Renamed canary
slot 351 from VAL_SHARPE_STD_EMA to VAL_SHARPE_VAR_EMA.
2. Component weight renormalization added — Σweights = 1 enforced
post-floor so PopArt's reward-magnitude EMA stays uncontaminated.
3. SABOTEUR_MIN bound violation — engagement-floor (0.1) was silently
pulling output below 0.5 stated minimum. Added post-multiplication
clamp to [SABOTEUR_MIN, SABOTEUR_MAX].
4. ratios[] undeclared — now __shared__ float ratios[6] block-loaded
once from ISV.
5. §3.6 slot count off-by-five (15 → 20). All 20 slots get FoldReset
entries; cold-start window behavior specified explicitly (no NaN path).
Architectural fixes:
6. Q-overconfidence concern addressed in scope — controller's
REWARD_CF_WEIGHT_INDEX covers CQL conservatism. No SP11′ deferral;
if T10 fails, extend controller outputs (e.g., target-update τ).
7. Curiosity recompute-at-replay specified (§3.5.1) — replay buffer
stores base reward only; curiosity added per-tuple at replay time
against current visit-count and current ISV. Eliminates stale-signal
replay contamination that would worsen ep1-overfitting.
Smaller fixes:
8. Novelty signal specified — SimHash 42×16 projection → 16-bit code,
1M-slot per-bucket count table, 1/sqrt(1+count). Hash table lives
in state-reset registry (cleared at fold boundary).
9. Saboteur engagement operationally defined — per-bar
|reward_with_saboteur − reward_without_saboteur| > EPS_ENGAGEMENT,
block tree-reduce, EMA'd. EPS_ENGAGEMENT = 0.01 × PnL_EMA.
10. Duplicate sentence in §7 component-starvation paragraph removed.
11. Validation criteria strengthened (§9): aggregation across 9
trajectories specified; primary metric = median peak-epoch ≥ 10
(baseline median peak-epoch = 1); secondary = drop-from-peak ≤ 5%
by ep20; tertiary = ep20 mean ≥ baseline ep1 mean. §9.3 fix-forward
response codified — no rollback.
12. Phases split per project pattern — Layer A additive (3 commits:
slots, canaries, controller), Layer B atomic consumer migration
(1 commit, including replay-time curiosity), Layer C validation +
close-out. No falsification gate; smoke is validation only.
13. Pearls A+D applied to controller outputs (§3.4.1) — chained
apply_pearls_ad_kernel after controller writes scratch, smoothed
values land in ISV. Consumers read smoothed slots.
Constants surviving (Invariant-1 anchors only, all rate-not-regime):
EPS_DIV, WEIGHT_HARD_FLOOR, SABOTEUR_MIN/MAX, CURIOSITY_PERMANENT_FRACTION,
CURIOSITY_BOUND_FRACTION, WEIGHT_FLOOR_FRACTION, ENGAGEMENT_FLOOR.
Spec is now full straight-up implementation; smoke validates Layer A
infrastructure and Layer B consumers, T10 validates SP11 success metric.
~1550 LOC across 3 commits in Layer A + 1 atomic commit in Layer B +
close-out in Layer C.
|
||
|
|
d49fbbe4e2 |
spec(sp11): reframe curiosity as feature + add permanent floor
User correction: curiosity is the *fix* for the ep1-peak overfitting pathology, not a hazard to defend against. Reframed §7 from "Risks" to "Design notes" — curiosity bound is a signal-relative scale, not a defensive cap. Fix the contradiction this exposed in the formula: previous `curiosity_pressure = stagnant_or_worse * curiosity_bound` went to zero when improving, which would cancel the always-on exploration the §7 narrative now relies on. Replace with permanent-floor pattern per pearl_blend_formulas_must_have_permanent_floor: curiosity_floor = 0.2 * curiosity_bound (CURIOSITY_PERMANENT_FRACTION) curiosity_dynamic = stagnant_or_worse * curiosity_bound curiosity_pressure = max(curiosity_dynamic, curiosity_floor) Now curiosity is always ≥ 20% of bound (anti-overfitting baseline) and rises toward the bound when stagnant (stagnation breaker). Updated unit-test guidance to assert pressure > 0 even at z=+10. CURIOSITY_PERMANENT_FRACTION=0.2 added to Invariant-1 fraction list. Saboteur-rising-with-improvement reframed as adversarial-load feature rather than over-stress risk. |
||
|
|
a72a23e6b1 |
spec(sp11): reward as controlled subsystem — fully ISV-driven
Brainstorm spec for SP11. Resolves the policy-stagnation pathology
surfaced in T10 train-multi-seed-xkjkb seed-0 ep0-14: model finds a
stable fixed point at ep1 (peak val sharpe 80.61), then OVERFITS to
it across remaining epochs (decline 80.61 → 70.58). Q-values grow but
val performance declines because reward function has no improvement
pressure.
Architecture (every input ISV-driven):
- Z-score-driven adaptation (no hardcoded "improving" threshold):
improvement_z = val_sharpe_delta_ema / max(val_sharpe_std_ema, EPS)
- 10 ISV outputs: 6 component weights + curiosity_pressure +
saboteur_intensity_mult + adaptive weight_floor + curiosity_bound
- 5 ISV canaries: val_sharpe_delta + val_sharpe_std (Z-score noise
estimate) + 6 per-component grad ratios + saboteur engagement +
PnL magnitude EMA (signal-relative curiosity bound)
- 4 new producer kernels (controller + 3 canary computers)
- Audit + migrate hardcoded cf_weight=0.3 in mse_loss_kernel.cu:318
and c51_loss_kernel.cu:789, plus other shaping multipliers
- NEW reward dimension: curiosity bonus, bounded by PnL magnitude
Per pearl_controller_anchors_isv_driven: every threshold replaced with
sigmoid(z) — no constants encode "what counts as improving". Per
pearl_blend_formulas_must_have_permanent_floor: every weight has
adaptive floor preventing zero-out. Per pearl_engagement_rate_self_
correction: saboteur intensity self-corrects via engagement rate canary.
Per pearl_cold_start_exit_signal_or: improvement signal OR'd from
multiple canaries so single-signal-failure doesn't stall controller.
New pearl authored alongside spec: pearl_reward_as_controlled_subsystem
— meta-principle that every reward path degree of freedom is a unified
controller output. Subsumes controller-anchor pearl at the reward layer.
Scope: 20 ISV slots, 4 producer kernels, audit + migration of
hardcoded reward shaping constants, 1 atomic commit.
~1300-1700 LOC; ~2.5-3 hours subagent work.
Success metric: val_sharpe[ep20] > val_sharpe[ep1] (Fix 33-38 baseline
peaked at ep1; SP11 should shift peak later as model continues
learning).
|
||
|
|
6ae9bc0055 |
spec(sp10): unconditional Thompson selector + ISV-driven temperature
Brainstorm spec for SP10. Resolves the val-Flat-collapse pathology that persisted through Fix 33-37: the eval-time argmax in experience_action_ select picks Hold deterministically every val bar (dir_entropy=0, trade_count=1 in 214k bars) regardless of controller state. Architecture: - Delete `if (eval_mode)` argmax branch in direction-selector kernel - Use temperature-blended Thompson: q_eff = E[Q] + τ × (Thompson - E[Q]) - τ = clamp(intent_eval_divergence / divergence_target, 0.5, 2.0) - τ self-corrects: collapse → high τ; healthy → low τ; permanent 0.5 floor - Reuses SP9's intent_eval_divergence_compute_kernel (extended, not new) Per pearl_controller_anchors_isv_driven: τ is ISV-driven, no hardcoded constants beyond Invariant 1 numerical anchors (clamp range). Per pearl_blend_formulas_must_have_permanent_floor: MIN_TEMP=0.5 ensures eval ALWAYS has stochasticity. The pearl_thompson_for_distributional_action_selection was about Bellman TARGET argmax (selector/target symmetry). It does NOT prohibit Thompson at the rollout selector. SP10 amends the pearl to clarify. Scope: 1 ISV slot, kernel modification (no new kernel — extend SP9's producer), 1 consumer kernel rewrite, test update, pearl amendment, audit doc Fix 38. ~300-500 LOC, single atomic commit per feedback_no_partial_refactor. |
||
|
|
38d2408604 |
spec(sp9): Kelly cold-start warmup floor — fully ISV-driven design
Brainstorm spec for SP9 (Fix 37). Resolves the val-Flat-collapse pathology
surfaced in T10 train-multi-seed-wsnc6 ep1-3: training intent_dist_f=0.47
but eval_dist_f=0.00 because Kelly cap pinned eval mag to Quarter, with
trade_count=1 in 214k bars (chicken-and-egg: cap closed → no trades → no
Kelly samples → cap stays closed).
Architecture per pearl_controller_anchors_isv_driven and
pearl_cold_start_exit_signal_or:
final_kelly_f = max(measured_kelly_f, warmup_floor)
warmup_floor = base_floor × (1 − combined_confidence)
base_floor = clamp(q_var_mag / q_var_mag_ema, MIN, MAX)
combined_confidence = max(statistical, behavioral, temporal)
All thresholds ISV-driven via Pearl D Wiener-α; only Invariant 1 numerical
anchors (clamp ranges, EPS) remain as constants.
Scope: 9 new ISV slots, 6 new GPU producer kernels including a
mandatory eval_dist GPU migration (eliminating DtoH readback at
gpu_backtest_evaluator.rs:1353 per feedback_no_cpu_compute_strict).
~1000-1200 LOC, single atomic commit per feedback_no_partial_refactor.
Pearls authored as part of brainstorm:
- pearl_controller_anchors_isv_driven (regime-encoded constants)
- pearl_cold_start_exit_signal_or (OR'd signals across statistical /
behavioral / temporal axes)
|
||
|
|
8f484f3312 |
spec(sp7): v2 — flatness-gated targets + Wiener α + sentinel-bootstrap
Self-review found 4 violations of feedback_isv_for_adaptive_bounds in v1: 1. TARGET_CQL_OVER_IQN=2.0 hardcoded → flatness-modulated per-branch target_ratio = ANCHOR × (1 − flatness[b]). Realizes the original pearl_q_flatness_gates_conservatism brainstorm idea. 2. TARGET_C51_OVER_IQN=1.0 hardcoded → mirror direction: target_ratio = ANCHOR × flatness[b]. 3. ALPHA_META=0.01 fixed → Wiener-optimal α from per-branch state-tracker EMAs (16 new ISV slots: diff_var/sample_var × 2 heads × 4 branches). 4. Cascaded floor cleanup: compute_adaptive_budgets line 3384 floors reads to 0.02/0.05 always, fighting the controller. Replaced with sentinel-aware bootstrap: cold-start uses BOOTSTRAP value, post-write takes ISV value verbatim. Acceptance criteria rewritten to be ISV-relative where the corresponding signal exists (q-spread vs NOISY_SIGMA, label_scale vs NOISY_SIGMA band). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
0808ebe3e5 |
spec(sp7): loss-balance controller — adapt CQL/C51 budgets to grad-norm ratio
Outcome-driven controller replacing Pearl 2's hardcoded-0 CQL budget and
floor-pinned C51 budget with multiplicative ratio adaptation. Targets
`grad_split_bwd cql/iqn = 2.0` and `c51/iqn = 1.0` per slice (trunk, dir,
mag). Slow EMA (α=0.01) consistent with existing controller pearls; Pearl
A sentinel-bootstrap on fold boundary; Pearl D Wiener-optimal smoothing.
Replaces ghost docstring at fused_training.rs:3409 ("0.10×(1−regime)×health"
formula was never implemented). Pearl 2 keeps owning IQN budget (reference
denominator) and FLATNESS_BASE (consumed by NoisyNet σ pearl); stops
writing CQL/C51/ENS slots so the new controller is the sole driver.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
ca6e0007da |
docs(sp6): brainstorm spec — per-branch consumer infrastructure refactor
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
6e6e0fa114 | docs(sp5): remove duplicate Layer A close-out section (superseded by Layer D) | ||
|
|
af8fe57d1c |
docs(sp5): apply critical review fixes — Layer D, Kelly carve-out, Pearl 4 mitigations, acceptance loosened
Critical self-review surfaced 15 issues; user directed fixes:
- "a no cpu path": KEEP host-EMA close-out (rule compliance) — but split
into separate Layer D atomic commit (was mis-scoped as Layer A)
- Pearl 4 kept: 3 concrete risks documented (constant-β proof break,
β2 memory reset destabilization, ε numerical envelope), structural
envelope bounds added, ALPHA_META halved, ε-only fall-back path defined
- Pearl 6 Kelly cross-fold persistence carve-out: separate slot range
280..286, NOT in SP4 fold-reset registry, Invariant 1 architectural exception
- Commit ordering Pearls 1 → 3 → 2 → 4 → 5 → 6 → 8 → 1-ext (resolves
Pearl 2 circular dependency on 1+3)
- Acceptance criteria: correctness gates (must pass) + performance gates
(loosened to "not catastrophically negative", within 2σ of pre-SP5)
- Pearl 7 timing: explicitly Layer C step 4, post-Layer-B + 3-seed validation
- Pearl 8 enumeration: 4 slots (TRAIL_DIST_PER_DIR per direction)
- Pearl 9 collapsed: 0 slots (Thompson achieved via Pearl 1's atom adaptation)
- Total slot count corrected: 110 (was 120-128 inconsistent)
Layer structure: A (additive, 8 commits) → B (atomic, 11 consumers) → C
(validation + Pearl 7 investigation) → D (host-EMA close-out, separate
atomic commit). Layer D split off from A's "close-out" because PnL
aggregation pipeline migration is its own architectural concern.
User final review pending before invoking writing-plans skill.
|
||
|
|
3361944456 |
docs(sp5): resolve open questions Q1=c, Q2=a, Q3=a
Per user direction:
Q1 (Layer A granularity): per-pearl commits (~9 commits, decision c)
Q2 (Pearl 7 timing): investigate-only in SP5; fix in follow-up
if Bin(2,0.5) persists post-Pearls-1-3 (decision a)
Q3 (Pearl 4 Adam β): include as designed; accept theoretical risk
with Pearls A+D + EPS_CLAMP_FLOOR mitigation;
Layer C smoke monitors for destabilization;
fall-back to ε-only if observed (decision a)
Spec section updated:
- Layer A description: per-pearl commit structure
- Pearl 7 framed as investigation-only with conditional follow-up
- Pearl 4 documents theoretical caveat + mitigation + fallback
User final review pending before invoking writing-plans skill.
|
||
|
|
5f1c1eec51 |
docs(sp5): comprehensive per-branch + per-group adaptation spec — close all hardcoded-multiplier deferrals
SP5 design covers every known adaptive-parameter deferral in the DQN training loop in a single coherent project. After SP5: zero hardcoded multipliers, every adaptive value ISV-driven via Pearls A+D. 9 pearls + 1 sweep close-out + 1 validation milestone: 1-3. Per-branch atom span / loss budget / NoisyNet σ (52 slots) 4. Per-group Adam β1/β2/ε ISV-driven (24 slots) 5. Per-branch IQN τ schedule (20 slots) 6. Kelly cap signal-driven floors (6 slots) 7. dist_q/h/f Bin(2,0.5) audit + action_select fix (0-8 slots) 8. Trail stop signal-driven thresholds (6-8 slots) 9. Thompson direction-branch temperature (4 slots) 1-ext. Per-branch C51 num_atoms (4 slots) Layer A close-out: 5 host-EMA host→GPU migrations Validation: 3-seed × 50-epoch acceptance gate Total: 120-128 new ISV slots, ~5000-7500 LOC, 11-13 producer kernels, ~12 consumer migrations. Layer A (additive infrastructure, ~15 commits) → Layer B (atomic consumer migration, single coordinated commit) → Layer C (validation + cleanup). Mirrors SP4's layer pattern. Triggering data: train-multi-seed-cv2mw 50-epoch L40S baseline (terminated F0 ep10) revealed magnitude head Q-flatness, eval collapse, and frozen action distributions. Plus all SP4 close-out + sweep deferrals folded in per user direction "no deferrals — make a single plan based on ALL findings". Spec at: docs/superpowers/specs/2026-05-01-sp5-magnitude-differentiation-and-eval-collapse-design.md User review pending before invoking writing-plans skill. |
||
|
|
24accea774 |
fix(sp4): migrate fast/slow grad-norm EMA update to GPU per feedback_no_cpu_compute_strict
Layer C close-out C1 redesigned. Original plan claimed
grad_norm_slow_ema_pinned was orphan post-Mech-6 migration; verification
surfaced a SECOND live consumer (fold_warmup_factor_update, commit
|
||
|
|
d8666a232f |
docs(sp4): fix all 11 review findings — spec is now source of truth
Resolved every issue from the critical self-review:
1. Param-group count: 7 → 8 throughout. Slot total: 36 → 40 (8 groups × 3
per-group families = 24 + 16 single = 40). IQL high-tau and low-tau
are 2 distinct param groups (separate buffers + Adam states), not
one. The "(Resolved during implementation)" hand-wave deleted.
2. Reset semantics section rewritten — no longer references Xavier-
derived bootstraps. Sentinel 0 per Pearl A; Wiener-state triples
reset to 0 alongside ISV slots (160 reset entries total).
3-4. Weight decay and L1 lambda hardcoded α/bootstrap removed. Both
now follow universal Pearls A+D contract (sentinel-detect +
Wiener adaptive). For L1, the natural curriculum (λ ramps as
gradient differentiates) emerges from the signal itself, no
"Bootstrap = 0.0" constant needed.
5. ε naming collision resolved: ε_div = 1e-8 (Pearl D division-safety),
ε_clamp_floor = 1.0 (Pearl A consumer cold-start floor). Distinct
names, distinct purposes.
6. Per-signal kernel signature updated: removed stale `ema_alpha` arg,
added `wiener_state` pointer + `wiener_state_offset` + uniform
`alpha_meta` (structural meta-EMA constant).
7. Pearl C engagement counter race fixed. Replaced `clamp_engage_buf
[group] += 1` (race) with proper register-then-tree-reduce pattern
mirroring `dqn_grad_norm_kernel`. Per-thread register counter →
block-shared tree-reduce → single block-leader writes to
`clamp_engage_per_block_buf[engage_buf_offset + blockIdx.x]`.
Host sums across blocks. No atomicAdd, consistent with feedback.
8. Pearl D state buffer allocation spelled out: `wiener_state_buf:
MappedF32Buffer` of size 141 floats (47 slots × 3), per
`feedback_no_htod_htoh_only_mapped_pinned`. Reset-registry entries:
40 + 40×3 (new bound slots + Wiener triples) + 7×3 (retrofit
existing producers' Wiener states) = 181 reset entries.
9. `grad_norm_slow_ema` retirement spelled out in Layer C: removed
entirely (sole consumer Mech 6 migrates to ISV[GRAD_CLIP_BOUND]).
Also documented Q_ABS_REF=16 and H_S2_RMS_EMA=96 transition to
"monitoring-only ISV" — producers stay running, only the orphaned
Mech 1/2/5/6/9/10 consumer reads removed.
10. Histogram bin range fixed: linear-spaced bins from 0 to step_max
(avoids log(0) singularity from earlier log-spaced design).
Linear is also better for p99: top-of-distribution gets ~0.4%
resolution per bin. Degenerate "step_max == 0" branch handles
all-zero signals gracefully (skip ISV update, leave previous bound).
11. Pearl D-subsumes-Pearl-A claim CORRECTED: mathematically wrong.
At t=0, Pearl D's formula yields `x_mean[0] = 0`, not `x[0]`.
Pearl A's first-observation replacement requires an explicit
sentinel-detection branch in the producer. Both pearls are
necessary and complementary; they are not hierarchical.
Slot count math now consistent: 1 target_q + 4 atom_pos +
3×8 per-group + 1 grad_clip + 1 h_s2 + 8 wd_rate + 1 l1_lambda = 40.
Producer count: 15 fused (1 + 4 + 8 + 1 + 1).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
074b4f0865 |
docs(sp4): fold all 4 pearls into spec — A, B, C, D
PEARL A — First-observation bootstrap: eliminates Xavier-derived formulas (2.33 z-score, √2 std, √(2/K_in)) from the bootstrap section. Sentinel ISV[X]=0 at fold reset; producer step 0 detects sentinel, replaces directly with step_observation. Subsequent steps EMA-blend. Consumer cold-start safety via .max(1.0) numerical floor only (Adam-ε category, not magnitude). Truly zero magnitude constants in bootstrap. PEARL B — Fused per-param-group statistics oracle: producer count 36 → 14. Per-group fused kernel reads (params, grads, adam_m, adam_v) once and computes WEIGHT_BOUND, ADAM_M_BOUND, ADAM_V_BOUND, WD_RATE in one multi-pass operation. Trunk's oracle adds Pass E for L1_LAMBDA gradient-direction entropy. 4× memory bandwidth reduction. Cleaner conceptual unit (per-param-group bounds = one oracle). PEARL C — Engagement-rate self-correction: detects post-clamp feedback-loop saturation. For in-kernel clamps, theoretical engagement rate = 1% (top 1% by p99 definition). Producer-side rate-deficit EMA detects sustained deviation; force-bumps bound to step_max when detected. Resolves the "in-kernel feedback loop accepted" limitation from first draft. Per-Adam-kernel block-shared-memory engagement counter (no atomicAdd), block-wide reduce, host-side rate-deficit EMA. PEARL D — Wiener-optimal adaptive α: replaces all hardcoded EMA rates across 14 new producers AND 7 existing pre-SP4 producers. Per-step: α* = diff_var / (diff_var + sample_var + ε_num). Theoretically optimal under Wiener-filter analysis; subsumes Pearl A as t=0 edge case (both vars=0 → α=1 → first-observation replacement). On stationary signals: α→0 (smooth). On non-stationary: α→1 (track). Eliminates the recursion problem (α controlled by signal stats, not another α). ε_num = 1e-8 (Adam-ε numerical category). 3 floats of state per producer. 36+ hardcoded α values eliminated codebase-wide. Limitations section restructured: 6 of 8 first-draft limitations RESOLVED by pearls (in-kernel feedback loop, smoke time-budget, producer plumbing, stale-bound, 8 carved-out items, magnitude bootstrap formulas). Remaining 6 limitations are genuinely irreducible (F0 launch-scheduling variance, Layer B atomic flip risk, novel-pearl field-validation, equilibrium-formula non-stationarity, Pearl C counter-state cost, Pearl D state cost). SP4 now closes 100% of the magnitude/regularization/EMA-rate surface. No hardcoded scalars remain in the entire bound/clamp/EMA chain. Producer count: 14. ISV slot count: 36. Unit tests: 36. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
21b6058e07 |
docs(sp4): close all carve-outs — weight decay + L1 lambda fully signal-driven
Folds two new producer signals into SP4 scope, eliminating the earlier "carved out" exception: WEIGHT DECAY (per param-group, 7 new ISV slots): λ = |w·g| / max(||w||², ε) Derived from equilibrium analysis of d/dt(||w||²) = 2(w·g) - 2λ||w||². The equilibrium gradient-projection-onto-weight-direction divided by weight norm. Same theoretical-derivation category as Adam β values. EMA half-life α=0.005 (~140 steps). Bootstrap 1.0. L1 LAMBDA (trunk only, 1 new ISV slot — NOVEL PEARL): λ = (mean(|g|) / mean(|w|)) × D where D = (log K - H_observed) / log K is gradient-direction entropy deficit and H_observed = -Σ p[i]·log p[i], p[i] = ||g[:,i]|| / Σ ||g[:,j]|| L1 regularization-strength derives from gradient-direction entropy deficit across input features. When gradient is uniform across features (D≈0): network hasn't differentiated, λ=0 (no pruning). When gradient concentrates on few features (D≈1): network has identified what matters, λ ramps up to prune the rest. Self-curriculum — L1 strength tracks the emergence of feature differentiation. This extends pearl_adaptive_moe_lambda (regularization strength = EMA- tracked deficit of regularized quantity) to feature-redundancy domain. Pearl-name candidate (post-validation): pearl_signal_driven_regularisation_strength. Bootstrap λ=0 means cold-start = no L1 pruning; ramp-up only after gradient differentiates. Worst-case behavior is "L1 disabled" — graceful. Total ISV slot count: 28 → 36. Total producer kernels: 28 → 36. Effort estimate: 3000-4000 → 3500-4500 LOC, 1-1.5 → 1.5-2 weeks. Acknowledged limitations updated: removed item #8 (carve-out) since no carve-outs remain. Added items for L1 pearl novelty (untested) and weight decay equilibrium-formula non-stationarity. Both have graceful worst-case behavior and explicit validation criteria (#10 and #11) to detect anomalies. SP4 now closes 100% of the magnitude/regularization surface — no hardcoded scalars remain in the entire chain. AdamW config fields for weight_decay and l1_lambda removed from HyperParams to prevent accidental hardcoding regression. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
506f9443fe |
docs(sp4): fix critical issue J — post-clamp diagnostic dead path
Second-pass self-review found a new critical issue: under "diagnostic = clamp engagement", post-clamp diagnostics (Mech 9 weights, Mech 5 Adam m/v slots 36-43, slots 44-45) NEVER fire — because post-clamp |v| ≤ bound by construction, so producer-side comparison always returns false. Fix: split diagnostic implementation by clamp location. - Buffer-based clamps (Mech 1, 2, 10): diagnostic stays in producer (reads pre-clamp buffer, fires when step_max > bound). - In-kernel clamps (Mech 6, 9): diagnostic lives INSIDE the Adam kernel at the clamp step. Each Adam kernel takes a `diag_slot` arg alongside `weight_clamp_max_abs`; on clamp engagement, writes `nan_flags_buf[diag_slot] = 1`. Idempotent per-thread store (all threads writing 1, race-free per existing convention). No atomicAdd. Without this fix, slots 36-45 would be dead diagnostics under SP4 — detecting nothing, providing no signal. With this fix, every diagnostic slot fires on its corresponding clamp engagement regardless of where the clamp lives. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
cfe57109e3 |
docs(sp4): revise spec — fix 8 critical issues from self-review
Self-review found multiple problems with the first draft. This revision addresses all 8 critical issues: 1. P² parallelization claim was WRONG — P² is sequential. Replaced with dynamic-range histogram (3-pass single-block kernel: max-reduce → log-spaced bin → cumulative-from-top to find p99). 256 bins → ~0.4% quantile precision; numerical-precision derivation in same theoretical-constant category as floating-point precision. 2. Bootstrap values had wrong magnitudes — used Xavier σ instead of p99 of max-element. Recomputed: WEIGHT_BOUND[group] = 2.33 × √(2/K_in[group]); H_S2_BOUND = 2.33 × √2 ≈ 3.3. All bootstraps now from theoretical p99 under Xavier-init, not std. 3. EMA half-life of 700 steps (α=0.001) didn't converge in 5-epoch smoke. Revised α=0.005 (~140 steps) for weight/Adam producers — reaches ~99% convergence within one fold's training. Smoke now validates steady-state behavior, not just bootstrap. 4. Weight decay, L1 lambda, CLIP_MULTIPLIER were listed in scope but undesignable in producer-consumer pattern. CARVED OUT explicitly: weight decay + L1 λ → separate research-spec; CLIP_MULTIPLIER and MIN_CLIP subsumed by SP4's GRAD_CLIP_BOUND slot. 5. F0 ≥ 45 acceptance criterion was uncertain. Revised to F0 ≥ 37.5 (matches the 1e30-effectively-unclamped diagnostic smoke). The ~8-point F0 variance from launch-scheduling-shift is independent of clamp value; SP4 cannot guarantee F0=45 even with ideal design. 6. Per-param-group p99 plumbing concretized: each producer takes (offset, length) launch args; main DQN params buffer sliced into trunk/value/branch using existing param_sizes layout knowledge. 7. Pre/post-clamp feedback loop EXPLICIT: producer runs BEFORE consumer for buffer-based clamps (h_s2, target_q, atom_pos — in captured graph immediately before clamp). Producer runs AFTER for in-kernel clamps (weights, Adam m/v — feedback loop accepted with documented soft-anchor dynamics). No more hand-waving. 8. Unit test strategy: per-producer kernel test with synthetic Gaussian input → known p99 ≈ 2.33 → assert |computed - 2.33|<5%. 28 tests total. Catches algorithm bugs before L40S smoke. Effort estimate revised UP: 3500-4500 LOC (from 2-3000), 1.5-2.5 weeks (from 1-2). 28 producer kernels + 28 unit tests is more plumbing than first draft assumed. Self-review limitations section now explicit about: in-kernel feedback loop accepted, smoke time-budget marginally validates Adam EMAs, F0 may not return to 45, Layer B atomic-flip risk, plumbing density, stale-by-one-step bound, histogram-precision is design choice not tuning, carved-out items remain hardcoded post-SP4. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
e00736ce9b |
docs(sp4): signal-driven magnitude control design spec
Comprehensive design replacing every hardcoded magnitude multiplier in the SP3 mechanism stack (Mechs 1, 2, 5, 6, 9, 10) plus pre-SP3 mechs in the same magnitude-control surface, per feedback_isv_for_adaptive_bounds and feedback_adaptive_not_tuned. Core principle: the BOUND lives in an ISV slot, computed by a producer kernel as p99 EMA of observed signal magnitude. Consumer reads the slot and clamps directly — no multiplier between ISV read and clamp. Cold- start ε from theoretical-init bootstrap (Xavier, etc.) — same theoretical- constant category as Adam β values, not tuning knobs. Architecture: - 28 new ISV slots (7 base bounds × per-param-group split where appropriate) - Per-signal P² (Jain-Chlamtac) quantile producer kernels - Diagnostic = clamp engagement (sticky flag from producer's max-comparison) - Migration in 3 layers: additive infra → atomic consumer flip → smoke Out of scope: theoretical/structural constants (Adam β, Xavier formula, attention 1/√d, hidden_dim, num_atoms). EMA rates stay as documented statistical-design parameters (half-life ≈ observation time-window). Estimated: ~2-3000 LOC across 3 layers, 1-2 weeks, 1 L40S smoke. Awaiting user spec review before writing implementation plan. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
ed727b51c0 |
spec(dqn): SP2 + SP3 combined — F0 regression fix + Q-learning structural stability
Designs the fixes for the two SP1 leftover issues identified at SP1 closure
(commit
|
||
|
|
792812baa1 |
spec(dqn): SP1 numerical stability — revise per critical review
9 substantive issues addressed inline:
1. ISV-driven design elevated from 'if applicable' to MANDATORY for
all dynamic bounds in SP1 fixes. Numerical-stability ε bounds are
the only carve-out (Invariant 1). Hardcoded tuning constants for
dynamic ranges explicitly rejected.
2. F0 Sharpe regression criterion changed from absolute (≥55) to
ratio-based (≥95% of latest baseline; floor 53.08 currently).
Prevents iterative erosion across multiple fix commits.
3. 'F1 trending positive' replaced with concrete monotone-improvement
test: Best Sharpe at last epoch ≥ Best Sharpe at first epoch of
the same fold.
4. Pass criterion distinguishes NaN-CLAMPED-TO-ZERO (failure) from
'Genuine grad collapse' (legitimate observation, permitted) per
the existing infrastructure from commit
|
||
|
|
dab5990287 |
spec(dqn): SP1 numerical stability investigation — F1 NaN root-cause
Brainstorm session 2026-04-29 produced this spec scoping Sub-project 1 of a three-sub-project decomposition addressing the F1 ep2 NaN explosion that persists across 9+ defensive layers in Plan C Phase 2 (commits e445d07a..e9096c7be). Decomposition: - SP1 (this spec): F1 NaN root-cause; γ audit + β always-landing instrumentation + surgical fix + multi-fold validation - SP2 (future): numerical stability framework — codify guards - SP3 (future): Q-learning structural stability — target-Q clip, pessimistic ensemble, atom-range governance Operating principles: - No deferrals (anomalies fixed within SP1, not punted) - Combined fixes — rich commits (per feedback_no_partial_refactor) - Always-landing diagnostic instrumentation (24-slot nan_flags_buf expands to 48; permanent regression sentinel) - F0 Best Sharpe ≥ 55 preserved (no regression on the working path) Pass criterion: all 3 folds train 5 epochs, F0+F1+F2 Sharpe ≥ 0, zero NaN-CLAMPED-TO-ZERO log lines, all 48 flag slots remain zero. 5 design sections approved iteratively: Architecture, Components, Data Flow, Error Handling, Testing. Next: writing-plans produces implementation plan. |
||
|
|
8a535681b7 |
spec(dqn): Plan C Phase 2 ACTIVE — UCB asymmetry regression triggers Thompson resumption
The original PAUSED state was motivated by measurement-artefact hunt exit. The bug-hunt cycle is complete (commits |
||
|
|
fc4addaabf |
spec+plan(moe): correct gate input dim 42 → STATE_DIM=128
T1.6 implementer correctly identified that the gate input dim is
ml_core::state_layout::STATE_DIM=128, not the literal 42 the spec/plan
incorrectly stated. The 42-dim figure was the bar-feature subset; the
actual state vector is 128-dim (42 features + portfolio + MTF + OFI
padded to 128 for cuBLAS alignment).
Updated spec §3 architecture diagram, §4.1 gate subnetwork description
+ parameter count (3,272 → 8,776), and plan header architecture line.
Implementation in commit
|
||
|
|
76e55047ec |
spec+plan(moe): add no-HtoD/HtoH constraint; mapped pinned only
Per feedback_no_htod_htoh_only_mapped_pinned.md (newly recorded): every CPU<->GPU path in this redesign uses mapped pinned memory exclusively. No cudaMemcpy HtoD, no Vec-to-Vec defensive copies, including in test code. CPU is strictly read-only on the production surface. Plan changes: - New Task 2.0 promotes MappedF32Buffer / MappedI32Buffer from distributional_q_tests.rs local definitions to a shared crates/ml/src/cuda_pipeline/mapped_pinned.rs module so all kernel test wrappers (Test 0.F, upcoming MoE tests) share one implementation. Adds write_from_slice helper for direct host_ptr write (no memcpy). - Task 2.1 test wrapper rewritten to allocate mapped pinned buffers + write to host_ptr + read GPU-written output via host_ptr. No more memcpy_stod / memcpy_dtov in test code. Spec: new section 6.4 codifies the mapped-pinned-only constraint and references the shared module + reference implementation. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
8629a9e7c2 |
spec(dqn): MoE replacement for vestigial RegimeConditionalDQN
End-to-end investigation (2026-04-27) confirmed RegimeConditionalDQN is vestigial decoration — 3 heads constructed at training start but only trending_head ever receives gradient updates. GpuDqnTrainer (the actual production GPU trainer) has zero references to RegimeType/regime routing; experience replay inserts go to trending_head.memory only; ranging_head and volatile_head stay at random init for the entire training run. Several support APIs (get_count_bonuses_branched, config, get_state_dim) hardcode-delegate to trending_head, ignoring the regime split entirely. Per `feedback_no_hiding.md` (wire up or delete) and the user's preference to fix not delete: the design wires regime conditioning properly via Mixture-of-Experts replacing the vestigial 3-head architecture. Pearl introduced and saved as `pearl_learned_gate_subsumes_handcoded.md`: when the network already sees the heuristic's inputs, a learned gate strictly subsumes any hand-coded discretization. This is the load-bearing rationale — ADX/CUSUM are already at state indices 40/41, so threshold-based regime classification is a strict information bottleneck the gate can recover and improve on. Design summary: - Architecture: shared GRN trunk -> K=8 small expert MLPs (256->64->256 bottleneck per expert, ~33k params each) -> learned gating network (state[42]->64->8 softmax) -> mixed h_s2 -> existing branching heads + C51 + IQN dual head. Soft full mixture (no top-k hardcoding); gate emerges peaky or flat from data. Anti-collapse load-balancing aux loss with default lambda=0.01 (configurable hyperparameter, not a kernel constant) prevents init-noise-dominated single-expert lock-in without forcing uniform utilization. User-confirmed signal: "collapses don't recover well in this codebase". - 9 new ISV slots (118-126: per-expert utilization EMA + gate entropy EMA), GPU-driven producer per `pearl_cold_path_no_exception_to_gpu_drives.md`. - 3 new small CUDA kernels (moe_mixture_forward/backward, moe_load_balance_loss) + 1 ISV producer; everything else is cuBLAS- reusable. CUDA Graph capture compatible. - Atomic deletion (no fallback): regime_conditional.rs (~700 LOC), RegimeType enum, classify_from_features, RegimeMetrics, RegimeClassConfig, 4 DQNConfig regime threshold fields, per-regime 3-file checkpoint format. DQNAgentType becomes thin wrapper over single DQN. Old checkpoints fail loudly with layout-fingerprint mismatch. - 5-layer testing strategy (unit kernels, smoke, gate-differentiation validation, L40S production validation with explicit kill criteria, architecture-hash backward-incompat). Out of scope (explicit): top-k routing, per-expert action heads, hierarchical MoE, regime-conditional CountBonus/NoisySigma broadcast, expert warm-start from existing trending checkpoint, CVaR action selection on mixed C51 distribution. Precondition: the in-progress use_* flag cleanup + count_bonus [f32; N] refactor lands as its own commit before MoE implementation begins, per `feedback_no_partial_refactor.md`. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
0e8804a770 |
spec(dqn): reframe Thompson rollout PAUSED (not SUPERSEDED)
Retracts the prior SUPERSEDED footer (commit |
||
|
|
42ffd6aadc |
spec(dqn): mark Thompson rollout SUPERSEDED — motivating pathology was a measurement artefact
The val-Flat-collapse / Short-collapse / C51 expected-Q bias hypothesis that motivated the 4-plan distributional-RL Thompson rollout was largely a measurement artefact in the diagnostic infrastructure, not a real policy pathology. Three layered bugs in actions_history_buf init + reader + mag_stats attribution conspired to inflate val_dir_dist Short, inflate active_frac, and pin wr_h/wr_f to zero. After fixing all three (commits |
||
|
|
021bb0ef73 |
spec(dqn): pivot Phase 0+2 tests to GPU-direct (no CPU mirror)
User correctly identified that CPU mirror function tests don't test
the production GPU code path. A bug shared between mirror and kernel
(translated identically wrong) would slip through. Mirror tests + a
single GPU bridge test were a weak compromise.
GPU-direct testing strategy:
- All Phase 0 kernel-correctness tests (0.A, 0.B, 0.C, 0.D, 0.F):
launch tiny test-only kernels with the SAME math the Phase 2
production kernel will use; assert properties of the output.
- Test 0.E (synthetic edge discovery): stays CPU. It tests an
ALGORITHMIC PROPERTY of Thompson exploration (does it discover
edge if edge exists?), not a kernel correctness property.
- All Phase 2 unit tests (2.A-2.D): GPU-direct against the
modified production kernel.
- Phase 0.F (real checkpoint extraction): unchanged — already GPU.
Local development uses RTX 3050 GPU (per memory user_dev_environment.md).
CI runs --ignored flag to skip GPU tests on CPU-only runners.
Time budget: Phase 0 was 1-2 days (CPU mirror); now 2-3 days
(GPU-direct, includes kernel wrapper setup half-day).
Other delta:
- Phase 0 deliverable file renamed: distributional_q.rs ->
distributional_q_tests.rs (no mirror functions, just tests +
kernel wrappers).
- Phase 2 unit tests rephrased to launch production kernel rather
than compare against CPU mirror.
- L1 verification gate runtime: seconds -> minutes (GPU launch
overhead per test).
The user's intuition was right: testing production directly is the
honest approach. Mirror was an optimization that traded correctness
for speed; with local GPU available the optimization isn't needed.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
25ba051157 |
spec(dqn): critical review pass — fix all majors, mediums, minors
Self-review identified 8 major + 8 medium + 15 minor issues. All fixed:
MAJORS:
M1 Test 0.A: clarified — compare argmax(E[Q]), Boltzmann(E[Q]),
and Thompson sampling distributions; assertion is on relative
ordering across all three.
M2 IQN quantile count: replaced hardcoded `5` with N_IQN_QUANTILES
constant (defined per task #147 fixed-quantile design).
M3 "Converged checkpoint" definition: ≥60 epochs trained AND
val_sharpe stabilised (no >10% change over last 10 epochs).
Cites prior 60-epoch validation runs (task #80, train-7rgqd).
M4 R3 reframed: replaced "no issue" handwave with explicit
by-design tradeoff acknowledgment + cost analysis. Wasted
exploration is the cost of finding out whether edge exists.
M5 Test 0.D σ_long=0.05 justified: chosen to match expected order
of magnitude given typical |return| ~ 50bps; Phase 0.F
validates against real checkpoint.
M6 rng_ctr post-increment: clarified — matches existing pattern
at experience_kernels.cu:858 (no behaviour change).
M7 train_active_frac instrumentation: NEW Phase 2 deliverable —
existing HEALTH_DIAG only has val_active_frac, but L3 verifies
training-time active_frac. Spec now explicitly adds this
~10-line metrics.rs change as a Phase 2 deliverable.
M8 eps_dir cleanup code-level detail: explicit reference to
experience_kernels.cu lines 814-865; remove eps_dir from both
static EPS_FLOOR clamp AND adaptive boost block; verify
variable can be removed from kernel signature via grep.
MEDIUMS:
Med1 Current C51/IQN combination: clarified that compute_expected_q
blends per training schedule; Phase 2 replaces with explicit
0.5*E_C51 + 0.5*E_IQN equal weighting; Phase 0.F verifies.
Med2 Eval mode phrasing: "eval mode already sets eps=0 in existing
kernel" — no semantic override, factually correct.
Med3 Magnitude σ claim: clarified — magnitude branch likely has σ
bias in OPPOSITE direction (Full has larger σ; UCB would
prefer Full and worsen saturation). Empirical verification
deferred. Phase 0.F should also report per-magnitude σ.
Med4 Hierarchical sampling claim corrected: it's not about
balancing 50/50 (already 50/50). It's about decoupling
cluster-best decisions; clarified.
Med5 n_atoms vs N_IQN_QUANTILES: clarified — n_atoms variable per
config (currently 51); N_IQN_QUANTILES fixed at 5.
Med6 Conviction code: removed pseudo-code; references existing
implementation at experience_kernels.cu:1091; provides
implementation hint for E[Q] reuse.
Med7 Q-target propagation: clarified — uses full distribution
(C51 atom projection / IQN quantile regression), not just
E[Q]. Thompson modifies action selection only.
Med8 References: added Thompson 1933 (original), Bellemare 2017
(C51), Dabney 2018 (IQN) for theoretical foundations.
MINORS:
Min1 Date updated to 2026-04-27.
Min2-3 Argmax monotonic /2 simplified out — argmax(a+b) =
argmax((a+b)/2). Code clarity improved.
Min4 P(argmax picks Long) = 0 deterministic; reframed assertion.
Min5 Test 0.F structural assertions added: σ_C51[FLAT] < 0.01 ×
σ_C51[LONG]; same for IQN; E[Q_FLAT] > E[Q_LONG]; argmax
picks FLAT; Thompson P(LONG)+P(SHORT) ≥ 0.20.
Min6 -INFINITY → CUDART_INF_F (CUDA convention).
Min7 dir_idx scope: comment notes it's declared earlier in kernel.
Min8 action_select args: explicit — three buffers exist on GPU
but not currently passed; new params, no new buffers.
Min9 Phase 0 time math: 5 hours tests + 1 hour enumeration + 2
hours 0.F + (3 hours runtime if checkpoint training needed,
runs in parallel). Honest budget.
Min10 "Two-stream" → "5-Layer Gate" header.
Min11 Plan 5 reference uses full path consistently.
Min12 Plan B time budget: explicit 1 day if pass; 2-5 days if bug.
Min13 active_frac: clarified Long+Short combined, not per direction.
Min14 train_active_frac: now in Phase 2 deliverables (see M7).
Min15 "20 mechanisms" → 21, with sub-counts in section headers.
Spec now 615 lines, comprehensive coverage of:
- Pearl + theoretical foundation
- Problem statement (with bias-might-be-correct caveat)
- Architecture (Thompson at training, argmax at eval)
- Train vs eval distinct selectors with behavior-change disclosure
- Direction-only scope with magnitude σ-bias warning
- Conviction stays E[Q]-based (no Kelly cap jitter)
- Interaction matrix: 21 mechanisms in 3 categories
- 6 v2 enhancements documented + deferred
- Phase 0/1/2/3 with tests, exit gates, time budgets
- 5-layer verification + train_active_frac instrumentation
- 8 risks with mitigations + 5 stop conditions
- What v1 doesn't touch (referencing interaction matrix)
- References (Thompson 1933, Bellemare 2017, Dabney 2018, etc.)
- Aggregation contract (project-wide pearl, enforced)
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
308c54484e |
spec(dqn): full interaction matrix + outside-the-box Thompson improvements
Per-user direction: every existing mechanism in the DQN system MUST be
explicitly considered for interaction with Thompson, and Thompson itself
MUST be examined for system-specific improvements beyond vanilla.
INTERACTION MATRIX (3 categories, 20 mechanisms):
Category 1 — Compose with Thompson (no change required):
Counterfactual reward, B.2 novelty bonus, PopArt, Saboteur,
Curiosity, NoisyNets/VSN, Distillation, CQL, Polyak target EMA,
HER, PER, Replay warm-start, Multi-fold validation harness.
Category 2 — Trivially adapt to Thompson (one-line changes):
D7/N7 contrarian sign flip (negate the SAMPLE), cosine epsilon
schedule (still applies to mag/ord/urg), per-sample epsilon (IQL
expectile gap), adaptive Boltzmann tau (still applies to mag/ord/urg).
Category 3 — Take precedence over Thompson (hard constraints):
Plan-based action lock (Thompson sample discarded if plan active),
per-magnitude Kelly cap, trail stop, capital floor breach.
Critical insights from the audit:
1. NoisyNets is ALREADY a form of training-time Thompson at the
parameter level. Output-space Thompson stacks on top —
total exploration = parameter-space ⊗ output-space (multiplicative).
2. Curiosity is ORTHOGONAL to Thompson — Thompson explores actions
whose Q is uncertain; curiosity explores states whose dynamics
are uncertain. Both axes desirable; no conflict.
3. Plan lock takes precedence; same as currently with Boltzmann.
OUTSIDE-THE-BOX v2 ENHANCEMENTS (deferred to follow-up specs):
v2.1 Triple-source Thompson (C51 + IQN + Ensemble) — incorporate
the existing ensemble Q-head as 3rd uncertainty source.
v2.2 Persistent Thompson (anti-churn for HFT) — bias sampling
toward current direction, ISV-driven; reduces tx_cost from
Long/Short oscillation across bars.
v2.3 CVaR-aware eval (risk-adjusted deployment) — eval picks
argmax(E[Q] − λ·CVaR_α[Q]); risk-aware decision making for
production with real capital.
v2.4 Information-Directed Sampling (Russo & Van Roy 2014) — picks
action minimizing regret²/info_gain; more efficient than
vanilla Thompson when learning saturates.
v2.5 Hierarchical Thompson on (trade vs no-trade) → (which
direction) — addresses 50/50 structural advantage of no-trade.
v2.6 Composition with curiosity-driven exploration — explicit
coupling beyond reward-side composition.
Each v2 enhancement gets its own spec/plan when prioritised. Vanilla
Thompson is v1; ships first; verified independently.
ALSO FIXED (from earlier self-review):
- Pearl claim softened: only the C51 Hold/Flat bias is directly
attributed; other historic bugs had different mechanisms.
- TFT entry removed from contract table — TFT is a Variable
Selection Network (feature processor), not a Q-head. Replaced
with generic "Future Q-head additions" placeholder.
- Eval direction = argmax E[Q] explicitly flagged as a behavior
change from current Boltzmann-with-tau (val_dir_dist will be
more concentrated than current).
- Phase 0.E budget reduced (1-2 hours, not 1 day) — synthetic
bandit is ~50 lines of Rust, not full RL training loop.
- Phase 0 enumerates existing checkpoints before training new one.
- Architecture diagram parenthesis fixed.
- Conviction implementation note: compute E[Q] once, reuse for
conviction AND eval-mode argmax — no redundant computation.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
1c63c32297 |
spec(dqn): reframe Phase 3 as direction-branch dead-code cleanup
User correctly challenged the "band-aid removal" framing. All fixes shipped during the val-Flat-collapse investigation addressed real bugs at their respective layers and should be preserved: - Kelly cap warm-branch ( |
||
|
|
5c325f8947 |
spec(dqn): pivot to Thompson sampling — distributional RL action selection
Revises the C51-bias spec after deeper review surfaced 12 design gaps,
of which 4 were critical:
1. Train-only vs train+eval ambiguity — UCB at eval would conflate
"model recommends Long" with "model is uncertain about Long",
inflating reported edge. CRITICAL for trading where eval drives
real capital decisions.
2. Thompson sampling is more principled than UCB:
- parameter-free (no κ to tune)
- uses distribution directly without scalar reduction
- naturally explore-exploit balanced via distribution shape
3. c51_alpha is the wrong blend weight (it's C51-vs-MSE-warmup, not
C51-vs-IQN). Equal-weight average of C51 and IQN samples is the
structural choice — no tuned blend weight needed.
4. The bias might be CORRECT BEHAVIOUR — model rationally choosing
Flat when no edge has been discovered. Phase 0 must include a
synthetic-edge test (controlled MDP with KNOWN positive Long
expected value) to verify Thompson can discover edge if it exists.
Other gaps fixed:
- Eval at argmax E[Q] (not Boltzmann, not Thompson)
- Pearl wording broadened to cover ensembles + future methods
- Ensemble Q-head added to aggregation contract table
- Explicit caveat: NEVER extend Thompson to magnitude branch (would
worsen existing magnitude saturation)
- Phase 0.F uses CONVERGED checkpoint (≥30 epochs), not 2-epoch run
- L4 long smoke (30 epochs, ~1 hour) added — Thompson edge discovery
needs longer feedback loop than 5 epochs
- Phase 3 explicitly removes eps-floor + tau-floor band-aids
(Thompson replaces direction Boltzmann; band-aids become dead code)
- Conviction stays E[Q]-based, not sample-based (avoid Kelly cap
jitter from stochastic samples)
Architecture (Thompson only, no UCB):
TRAINING: dir_idx = argmax(0.5 × (sample_C51(d) + sample_IQN(d)))
magnitude/order/urgency: existing Boltzmann + ε-greedy
EVAL: dir_idx = argmax(0.5 × (E[Q_C51] + E[Q_IQN]))
magnitude/order/urgency: existing Boltzmann (eval mode)
Direction-branch ε-greedy + Boltzmann are REMOVED — Thompson is the
exploration mechanism. No new GPU buffers; existing C51 atoms + IQN
quantiles passed to action_select.
5-7 days active work across 4 sub-plans; each gets its own
writing-plans cycle.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
494125c104 |
spec(dqn): distributional RL aggregation — UCB on C51+IQN σ for action selection
Designs the structural fix for the C51 expected-Q Hold/Flat bias exposed
by the Kelly val-Flat-collapse fix. The bias is a manifestation of a
deeper project-wide pearl:
"Distributional RL aggregation discards uncertainty;
action selection must restore it."
Any value head representing Q as a distribution (atoms, quantiles,
ensembles) MUST expose both E[Q] and σ(Q) to action selection. Boltzmann
on E[Q] alone produces structural bias toward low-variance actions
regardless of expected payoff — the C51+IQN Flat-attractor is one
instance of this lost-information pattern.
Fix: extract σ(Q) from BOTH C51 atoms (closed form) and IQN quantiles
(IQR/1.349), blend by loss-time weight, feed Q_eff = E[Q] + κ·σ
(κ=1.0 structural identity) to direction-branch Boltzmann ONLY.
4-phase implementation:
Phase 0 — TDD hypothesis verification (Rust mirror functions + 5 unit
tests including GPU integration on real checkpoint)
Phase 1 — Audit existing reward levers (B.2, CF, PopArt, Q-target)
via 6 unit tests; fix any bugs found
Phase 2 — UCB integration: new compute_q_with_uncertainty kernel,
modified action_select, Rust orchestration, project-wide
aggregation contract in dqn-wire-up-audit.md
Phase 3 — Verification per Plan 5 Task 5 multi-seed × multi-fold
5-layer verification gate; 8 risks with mitigations; 4 stop conditions
that halt execution and force redesign.
Existing band-aid fixes (Kelly cap, eps-floor, tau-floor) stay — they
address symptoms at different layers. UCB adds the missing aggregation
step that was the common root across all the symptoms.
Direction-branch only — magnitude/order/urgency don't have the
Flat-attractor (atom-mass collapse asymmetry).
5-7 days active work across 4 sub-plans; each gets its own
writing-plans cycle.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
f3a8a5ff62 |
spec(dqn-v2): D.8 TLOB pivot — cuBLAS port using existing DQN infra, no ONNX, no pretraining
User pearl (2026-04-24): the DQN already has every primitive TLOB needs (gpu_attention, batched_forward/backward, GpuLinear, cuBLASLt handles, cuda_autograd). Port TLOB's Q/K/V + attention as a composition of existing primitives. Random-init, trainable end-to-end from the DQN's reward signal. Uniform with Mamba2, IQL, atoms/γ/τ/ε — all of which already follow this pattern. Eliminates: - ONNX Runtime dependency (was already dead — stripped from ml-supervised) - Separate supervised pretraining pipeline - "Freeze vs fine-tune" false dichotomy - Pretrained-checkpoint-file-not-found failure mode Prerequisites for the cuBLAS-native design (all satisfied): - gpu_attention.rs exists - GpuLinear trainable layer exists - cuda_autograd over cuBLAS exists - MBP-10 data already in the DQN data pipeline What was "BLOCKED on prerequisites" in the Task 6C audit referred to the OLD pretrained-ONNX design. The cuBLAS-native design has all prerequisites satisfied — Task 6C can proceed under the new scope. Pearl captured in memory: pearl_tlob_no_pretraining.md. Generalises to any future attention/state-space/Neural ODE module: port to cuBLAS, random init, let the DQN teach it. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
cf36091eeb |
spec+plan(dqn-v2): GPU drives, CPU reads (§4.C.6 revision)
User raised that the CPU-controller pattern (where CPU computes adaptive values from ISV and writes back to ISV) violates the architectural principle: GPU kernels should compute adaptive decisions; CPU-side code is pure observation. Replaces the AdaptiveController trait with read-only AdaptiveMonitor: - No update() — CPU doesn't compute adaptive values. - No write_output() — CPU doesn't write adaptive values to ISV. - read() returns the GPU-computed value from ISV. - diagnose() emits HEALTH_DIAG snapshot. - observe() + fire_rate() track how often the GPU-computed value changes. Reclassifies the 9 adaptive mechanisms: - 6 reactive get GPU kernel + CPU monitor: atoms, gamma, kelly_cap, tau (Polyak EMA), epsilon, grad_balancer (last is already GPU-driven). - 3 static get ISV constructor-write, no monitor: cql_alpha, conviction_floor, plan_threshold. New ISV slots for GPU-written adaptive outputs (EPSILON_EFF, TAU_EFF, GAMMA_EFF, KELLY_CAP_EFF) and CPU-born inputs (EPOCH_IDX, TOTAL_EPOCHS), plus slots for the 3 static configs. Rationale: - Unified adaptive machinery, no CPU-side special cases. - Per-sample granularity available (Expected SARSA τ in c51_loss_kernel is already the exemplar — reads ISV q_gap + health per sample). - Zero CPU→GPU config transfer in any path. - ISV is single source of truth for every adaptive value. Plan 1 Tasks 8-17 restructured: - Task 8 creates AdaptiveMonitor trait (not AdaptiveController). - Tasks 9, 10, 11, 13, 14, 17: GPU kernel + monitor pairs. - Tasks 12, 15, 16: static ISV writes only. Follow-up commits will revert Batch A ( |
||
|
|
fb6d0f40f9 |
audit(dqn-v2): A.5 orphan audit complete
Populated docs/dqn-wire-up-audit.md with every pub module and CUDA kernel in the DQN path. Each entry classified Wired / Partial / Orphan / Ghost / OUT-of-DQN-scope with the action plan linking to the plan+task that resolves any non-Wired status. No Orphan left unclassified. Orphans fall into three buckets: 1. Scheduled for wiring by a later Plan (gpu_statistics → Plan 2 D.2; tlob_loader → Plan 2 D.8). 2. OUT-of-DQN-scope because supervised consumers exist (PPO kernels, xLSTM, KAN trainable adapter, flash_attention, benchmarks). 3. Genuinely unused — escalated to user review in the task output, not deleted autonomously (streaming_dbn_loader, unified_data_loader, training/orchestrator, training_pipeline, inference_validator, model_loader_integration, paper_trading/mod.rs, portfolio_transformer, regime_detection/mod.rs). Summary: 109 total modules/kernels, 74 wired, 7 partial, 11 orphan, 0 ghost, 17 OUT-of-DQN-scope. Plan 1 Task 6. Spec §4.A.5, Invariant 2. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
1353e71ae6 |
spec+plan(dqn-v2): layout fingerprint replaces schema version (§4.A.2)
User raised a backward-compat concern: integer "schema version" semantically implies a family of coexisting versions with upgrade paths between them, inviting the forbidden pattern (feedback_no_legacy_aliases.md, feedback_no_partial_refactor.md). Replace with a compile-time structural fingerprint: - ISV[0..2) stores a u64 FNV-1a hash of the slot layout (split across two f32 lanes preserving raw bits). - The hash is computed by a const fn over the slot list; any slot change automatically updates the fingerprint. No human decides "what version is this now". - Checkpoint load is fail-fast only. Error message does NOT mention migration as an option. - Pre-commit hook rejects any `fn migrate_isv|upgrade_isv` to make the no-migration rule structurally enforced (landed as part of Plan 1 Task 5 hook extension). - StateResetRegistry entry renamed ISV_SCHEMA_VERSION → ISV_LAYOUT_FINGERPRINT (Task 2's landed code touched in Task 5's commit per no-partial-refactor). Updated: - spec §4.A.2 (fingerprint design + rationale) - spec §5 landing-order note (fingerprint auto-updates on layout change) - Plan 1 architecture line - Plan 1 Task 2 StateResetRegistry test + entry naming - Plan 1 Task 5 — full rewrite of implementation steps - Plan 1 exit criteria #6 - Plan 2 pre-plan verification (grep fingerprint constants, not version==0) - Plan 2 Task 6D.1 (update fingerprint seed, not bump version) - Plan 2 Task 6D exit criteria #7 - Plan 3 pre-plan verification - docs/isv-slots.md ISV[0..2) row No code changed. Plan 1 Task 5 is not yet implemented — this commit realigns the spec + plans so the implementer subagent works from the corrected design. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
336ee40b9d |
spec(dqn): add Invariant 9 — no deferred work, no stubs, concrete E decisions
User-mandated addition: Invariant 9 — No deferred work, no stubs, no TODO/FIXME Every commit lands complete. No stubs (placeholder return values), no TODO/FIXME/XXX/HACK/TBD markers, no half-finished implementations. If it can't finish in this commit, it doesn't start. Stubs and deferrals train the network on semantic emptiness — invisible in convergence metrics but burn GPU time optimising against partly-fake signal. Authority: feedback_no_stubs.md, feedback_no_todo_fixme.md, feedback_no_quickfixes.md elevated to first-class spec enforcement. Pre-commit hook greps for forbidden markers. Part E decisions made concrete (per Invariant 9): - E.1 TFT VSN: IN (full VSN extension) - E.2 GRN: IN if A.5 audit finds absent (otherwise mark existing as canonical) - E.3 Multi-quantile heads: IN (5/25/50/75/95 quantile decomposition) - E.4 Encoder-decoder: IN (extends D.3 to full trunk/head separation) - E.5 Interpretability: IN (attention-weight ISV diagnostics) - E.6 Auxiliary heads: IN (next-return + regime-classification) - xLSTM: OUT (redundant with Mamba2+TLOB; YAGNI) - KAN: OUT (function approx not a bottleneck) No "evaluate later" items. Every concept either lands or explicitly doesn't, with rationale. §8 decisions tightened: - D.8 TLOB mode: decision criterion = per-step cost benchmark ≤ 5ms - C.6 controller order: fixed in spec (atoms→gamma→Kelly→cql_alpha→ tau→epsilon→conviction_floor→plan_threshold→balancer) - A.3 resource plan: auto-switch L40S→H100 at 12 GPU-hour threshold - A.4 metric band init: [mean − 3σ, mean + 3σ] from last 3 good runs No discretionary "we'll see" remain. All decisions data/order/budget/ statistic-driven. Counter updates: "seven invariants" → "nine" in §3 header and §5 cleanup contract. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
3a62e23056 |
spec(dqn): review fixes — clarify hot-path scope, add Invariant 8 named dims
Self-review pass on the DQN v2 unified design spec (
|
||
|
|
d13b535862 |
spec(dqn): DQN v2 unified integrated policy system design
Design spec consolidating every hard lesson from the val-Flat-collapse investigation, the task #92 gradient-pathology triad, the ISV v-range unification work, the f64→f32 ABI cleanup, and many sessions of fold-1 grad explosions and orphan-feature accumulation. Seven non-negotiable invariants: 1. ISV-driven adaptive bounds (no tuned constants) 2. Wire-It-Up (no orphans) 3. GPU-only hot path (zero memcpy DtoH; pinned-zero-copy only) 4. Pinned-only host memory 5. Diagnostic-first 6. Convergence discipline (primary guardrail, multi-seed × multi-fold) 7. Wire-up + diagnostic audit per-commit Five Parts, landing as one coordinated rollout: A. Foundation — reset registry, ISV contract, multi-seed harness, regression detection, orphan audit, hot-path purity audit. B. Flat-trap escape — ISV-driven reward shaping, trade-attempt bonus, replay warm-start via scripted-policy portfolio, adaptive plan threshold. C. Refinement — quantile atom support, reward attribution, state-dist divergence signal, temporal timing bonus, CQL-seed coupling, adaptive controller unification refactor. D. Temporal architecture — Mamba2 backward completion, per-branch gamma, horizon-decomposed value function, generalised temporal reward, soft fold transitions, plan temporal dynamics, liquid audit, TLOB-DQN integration. E. Supervised→DQN concept audit — TFT VSN, GRN, multi-quantile, xLSTM, KAN, encoder-decoder, interpretability, auxiliary heads. Tiered success criteria: Tier 1 — convergence discipline (must pass first). Tier 2 — behavioural parity (val trades escape Flat-trap). Tier 3 — profitability (val Sharpe > 1.0 on window). ISV slots grow from 37 to 72. Five new audit docs tracked by pre-commit hook. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
9deda5f65b |
feat(dqn): ISV-unified per-branch Q-support range (spec 2026-04-23)
Unifies the Q-support range source across atom grid / warm-start quantile
clamp consumers via the ISV signal bus. One broadcast written at epoch
boundary from per-branch Q-stats EMAs, read by the two consumers that
previously held disagreeing ranges. Target observation: atom utilisation
≥40% (up from 11-15% on train-fpxnw).
Phase 0 — per-branch Q-stats kernel Rust plumbing:
* Load q_stats_per_branch_reduce alongside legacy q_stats_reduce
* Add per_branch_q_stats_pinned (28 f32 = 4 × 7, device-mapped)
* PerBranchQValueStats struct: [QValueStatsResult; 4]
* reduce_current_q_stats_per_branch launches the new kernel with the
four branch (off, size) pairs derived from config.branch_N_size
Phase 1 — ISV v-range plumbing (zero behavioural change at epoch 1):
* ISV_NETWORK_DIM=23 preserved for w_isv_fc1 sizing; ISV_TOTAL_DIM=31
allocates 8 additional slots for per-branch (centre, half-width)
* Slot constants V_CENTER_DIR..V_HALF_URG covering slots 23..30
* eval_q_mean_ema / eval_q_std_ema / eval_ema_initialized promoted
to [f32; 4] / [bool; 4]; scalar setters preserved for trajectory
backtracking (broadcast same value to all branches)
* Bootstrap at construction: centre=0, half=(v_max-v_min)/2 → the
byte-identical [config.v_min, config.v_max] span per branch before
any Q observations arrive
* reset_eval_v_range_state resets the 4 per-branch EMAs AND the 8 ISV
slots to bootstrap values; legacy eval_v_range_pinned[2] still reset
(deferred removal — spec Phase 3)
* update_eval_v_range reworked: signature takes PerBranchQValueStats and
per_branch_q_gaps. Maintains 4 independent adaptive-rate EMAs,
computes (centre, half) per branch with min_half_floor=0.1×(v_max-v_min)
and clamps to config bounds, writes 8 ISV slots. Branch-0 (direction)
centre±half is also mirrored into the legacy eval_v_range_pinned for
consumers that have not yet migrated to the per-branch bus.
Phase 2a/2b — atom grid per-branch v-range:
* adaptive_atom_positions kernel signature changed from
(v_min: float, v_max: float) to (branch_idx: int, isv_signals: float*);
reads centre/half from ISV slots 23+2·b, 24+2·b. Eliminates the f64→f32
ABI trap (spec Phase 2 side-effect) since the only per-branch range
path is now pointer-based.
* recompute_atom_positions passes branch_idx + isv_signals_dev_ptr per
branch; no scalar v_min/v_max arg remains.
Phase 2c — warm-start quantile clamp per-branch from ISV:
* warm_start_atom_positions reads per-branch (centre, half) from pinned
ISV host memory, clamps shared reward-quantile vector into each
branch's adaptive range before tiling into atom_positions_buf.
Bootstrap makes this equivalent to the pre-spec config.v_{min,max}
clamp until the first Q observation lands.
Deviations from spec:
* Phase 2d (per_sample_support_buf → [N, 4, 3]) NOT implemented. The
spec's premise was that per_sample_support is host-tiled from
eval_v_range, but the active path in this codebase has it filled by
iql_compute_per_sample_support (V(s)-centered, per-sample, already
adaptive) — orthogonal to the ISV bus. Migrating that kernel to
per-branch output would require rewriting iql_value_kernel +
iql_support_floor + C51/MSE loss kernel indexing in lockstep, which
the "no unrelated refactoring" constraint disallows. The loss-kernel
Bellman projection today uses V-centered bounds that are themselves
adaptive; the ISV v-range fix still lands the primary win (atom grid
+ warm-start agreement) without touching IQL.
Compile verified: cargo check -p ml + --workspace pass (SQLX_OFFLINE,
CARGO_INCREMENTAL=0, sccache). No TODO/FIXME/XXX introduced.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
baa4151fab |
docs(isv): signal quality audit for DQN ISV bus — L40S oscillation diagnosis
Diagnostic-only audit of all 22 ISV slots using train-mdh86 HEALTH_DIAG (20 epochs, same oscillation pattern as train-rq6n8). Identifies three dead writers (slots 2, 3, 4), one strongly anti-correlated signal (slot 12 learning_health, r=-0.765 vs Sharpe), sample-0 bias in four regime slots (8-11), and EMA-pollution risk on the Q-scale EMAs (16, 21) that feed Kelly conviction + C51 bin weighting. Produces a per-slot table with writer/reader citations and an ordered fix list. Hypothesises the observed +34 → -67 Sharpe swing is driven by health mis-reporting "improving" while policy diverges, q_dir_abs_ref contamination collapsing Kelly, and the zero-writer TD-error EMA disabling micro-reward regime awareness. No code changes — reconnaissance only. |