d243a6f0806268950bd5c2e3fc0c290aaf52d479
329 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
d243a6f080 |
docs(sp14): spec v2 — critical review corrections
8 corrections from code-anchored critical review at HEAD
|
||
|
|
eaf4adcb98 |
docs(sp14): spec — aux→Q wire + Earned Gradient Flow pearl
Designs the SP14 chain on top of SP13 Layer B (HEAD
|
||
|
|
0ad5b6fa42 |
chore(sp13): script bug fix + Layer B implementation notes
While P0b smoke (train-sw4ws on
|
||
|
|
ba83fcd1f5 |
docs(sp13): v3 spec + plan — Hold-pricing replaces Hold-elimination
P0a.T3 v2 implementer's audit revealed `DirectionAction` enum doesn't exist; the codebase uses an 8-variant fused `ExposureLevel` (ShortSmall/Half/Full, Hold, LongSmall/Half/Full, Flat) with cross-crate consumers across 77 files and 32+ test files pinning the 8-variant invariant. Atomic Hold elimination would cascade massively. User insight (2026-05-04): Hold being FREE is the bug, not Hold itself. MFT trading legitimately needs multi-bar holds; we want the model to use them deliberately, not as a CQL-bias lazy default. Holding isn't free in the real world — broker fees, margin interest, opportunity cost. v3 reframes as Hold-pricing: - 4-way action space stays; ExposureLevel::Hold stays; no cross-crate cascade - 3 new ISV slots (380-382): HOLD_COST_INDEX, HOLD_RATE_TARGET_INDEX, HOLD_RATE_OBSERVED_EMA_INDEX - Hold-rate observer: small GPU kernel + Pearls A+D smoothing - Hold-cost controller: 5-line deficit-driven formula (excess > target → cost rises 1×→5× base; observed ≤ target → relax) - Per-bar reward subtraction at action == DIR_HOLD site - 2 GPU oracle tests for the controller P0a.T3 cuts from ~250 LOC + 32-test cascade → ~120 LOC additive. T1+T2 already-staged work unchanged. T4/T5/Layer B/C/D structure preserved. Tension with pearl_event_driven_reward_density_alignment acknowledged in spec — per-bar Hold cost is exposure-NEGATIVE (away from Hold), models real economic carry, ISV-bounded by controller. Inverse of the pearl's failure mode. Faithful reward modeling, not artificial shaping. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
0ca45ef61d |
docs(sp13): v2 spec + plan after critical review
Spec v2 supersedes v1 with five major fixes from the critical review: - Drop alpha-vs-benchmark reward (mathematically tautological — Long trades had alpha = -costs always against always-long benchmark) - Split Phase 0 into 0a (Hold-only, clean test of user hypothesis) + 0b (aux amplification probe, only if 0a partial), avoiding the v1 confounded experiment - Bound direction-skill bonus relative to |alpha| (cap_ratio × |alpha|) to prevent reward gaming on small-alpha correct-direction losers; spec adds 8-quadrant worked-example matrix verifying the no-negative-EV invariant - Replace 5-epoch absolute-threshold gate with 10-epoch trajectory criterion to avoid false-negatives from aux head underconvergence - Aux_w controller adds dual-EMA stagnation detector (decays toward base when no improvement) — prevents permanent destabilization of Q-head in data-limited case Plan v2 mirrors spec changes: - Phase 0a (Hold elimination + dir_acc instrumentation, ~300 LOC, 1 atomic commit) - Phase 0b (aux_w controller replacement, conditional, ~80 LOC) - Layer B (aux head regression -> binary classification, ~120 LOC) - Layer C (skill bonus + luck discount with calibrated bounds, ~150 LOC) - Layer D (30-epoch validation + 3 new pearls) ISV slot allocation [372..380): drops slot 371 (BENCHMARK_PNL_CUMULATIVE), adds 374 (AUX_DIR_ACC_LONG_EMA — stagnation), 379 (SKILL_BONUS_CAP_RATIO). Plan integrates Explore-agent touch-list (50-70 sites) with corrected enum ordering: Short=0, Long=1, Flat=2 (preserves codebase Short-first convention, avoids 30+ stale-comment churn). Adds direction-bias signal handling at experience_kernels.cu:5159 (Hold's 0.5 softening gate vanishes -> [1, 1, 0]). Three new pearls planned for Layer D close-out: - pearl_redefine_success_for_predictive_skill - pearl_skill_bonus_must_be_alpha_bounded (calibration lesson) - pearl_reward_quadrant_audit_required (meta-pearl) Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
ecab584c32 |
spec(sp12): per-trade event-driven reward composition (v3)
Three architectural changes in one unified design to push the model from
HFT noise extraction (62% trade rate, sharpe-gaming) toward MFT alpha-hunting:
1. Asymmetric bounded cap (-10/+5) — restores loss aversion erased by
SP11 symmetric cap. 2:1 ratio matches Kahneman/Tversky prospect theory.
Anchor: pearl_audit_unboundedness_for_implicit_asymmetry.
2. Min-hold soft penalty with temperature curriculum — patience requirement
at exit. Soft factor = deficit/(deficit+T), T anneals 50→5 over 50 epochs.
Forces commitment without paralysis in early training.
3. Zero per-bar shaping (gate micro/opp_cost on events) — eliminates
continuous-reward gradient that pulls toward continuous exposure.
Anchor: pearl_event_driven_reward_density_alignment.
Combined: reward fires only on trade events with prospect-theory loss
aversion + commitment requirement. Pure per-trade event-driven Q-learning
properly aligned with per-trade P&L objective.
~50 LOC across 3-4 files. No new ISV slots in Phase 1 (constants only).
Cost ~€1.30 (€0.30 smoke + €1.00 30-epoch validation).
Empirical motivation: train-multi-seed-pmbwn 50-epoch on commit
|
||
|
|
85069ba75e |
spec(sp11): amend B1b smoke-recovery — z-score normalization for mag-ratio canary
smoke-test-6wd2c on commit |
||
|
|
e0d3abd9d2 |
plan(sp11): patch all 5 violations + 2 gaps from review
feedback_no_htod_htoh_only_mapped_pinned (tests not exempt):
- All test fixtures converted from htod_copy/dtoh_sync_copy to
MappedF32Buffer with host_slice / host_slice_mut access.
- Novelty hash buffer + projection matrix changed from
cudarc::CudaSlice<f32> + alloc_zeros to MappedF32Buffer.
feedback_no_cpu_forwards (CPU is read-only):
- Projection matrix initialization changed from host-side StdRng +
host_slice_mut writes to a one-shot GPU init kernel
(novelty_simhash_proj_init_kernel) using Philox seeded from
config.seed. No host RNG, no host writes.
feedback_no_cpu_compute_strict (saboteur multiplication):
- B1 step 5 reverted from Rust-side `read_isv_slot * scale` to
GPU-side: pass base scale + ISV pointer + slot index to the
saboteur perturbation kernel; multiplication happens on-device.
feedback_trust_code_not_docs (grad-ratio terminology):
- Spec §3.3.1 "per-component grad EMA" was wrong — SP4 grad-balancer
is per-branch (4 slots), not per-reward-component. Renamed:
reward_component_grad_ratio_compute_kernel
→ reward_component_mag_ratio_compute_kernel
REWARD_COMPONENT_GRAD_RATIO_BASE
→ REWARD_COMPONENT_MAG_RATIO_BASE
Source: existing REWARD_POPART_EMA_INDEX..+6 (per-component reward
magnitude EMAs from SP4 reward_component_ema_kernel). Semantic
equivalent for the controller's exploit/diversify blend.
feedback_no_stubs (dead parameter):
- Removed `eps_div_idx_unused` from mag_ratio kernel signature.
Saboteur engagement: missing producer specified
- Spec §3.3.1's two-reward-arrays formulation replaced with single
`saboteur_delta_reward_buf` produced by the saboteur perturbation
kernel itself (single reward computation, diff emitted as side
output). Engagement kernel signature simplified to one input array.
PNL_REWARD_MAGNITUDE_EMA_INDEX (slot 359): producer wired
- mag-ratio kernel mirrors `isv[REWARD_POPART_EMA_INDEX]` to
scratch_out[6]; chained apply_pearls_ad targets slot 359.
Replay sample kernel location specified
- graph_utility_kernels.cu:71 (gather_f32_scalar). New sibling kernel
`gather_replay_reward_with_curiosity` defined; replaces the existing
scalar gather (no legacy alias per feedback_no_legacy_aliases).
novelty_simhash_lookup runs before, novelty_simhash_update after.
A2 controller test placeholders → full GPU oracle assertions
- Three controller tests (z=0 midpoint, weight renorm, saboteur clamp)
have full mapped-pinned fixtures with assertions on weight sum,
individual values, post-clamp bounds.
Plan now passes:
- feedback_no_htod_htoh_only_mapped_pinned (tests + production)
- feedback_no_cpu_forwards (CPU never writes/computes for GPU)
- feedback_no_cpu_compute_strict (all multiplications GPU-side)
- feedback_no_atomicadd (race-tolerated non-atomic, safety documented)
- feedback_no_partial_refactor (Layer B atomic; saboteur kernel sig
change touches all callers in the same commit)
- feedback_no_stubs (no dead parameters)
- feedback_trust_code_not_docs (corrected spec terminology)
- feedback_wire_everything_up (every new field has init + producer)
- feedback_no_legacy_aliases (old gather_f32_scalar replaced, deleted)
1447 lines, +268 from previous version.
|
||
|
|
d8b44e0829 |
plan(sp11): implementation plan for reward-as-controlled-subsystem
Task-by-task plan for the SP11 spec at HEAD
|
||
|
|
9395b983cc |
spec(sp11): fix all 13 review issues — straight-up implementation
Critical bugs fixed:
1. Z-score formula: was mean(Δ)/mean(|Δ|), bounded to [-1,+1] — sigmoid
never saturated, controller stuck in [0.27, 0.73]. Now true Z-score:
delta_ema / sqrt(delta_var_ema) with two-pass var. Renamed canary
slot 351 from VAL_SHARPE_STD_EMA to VAL_SHARPE_VAR_EMA.
2. Component weight renormalization added — Σweights = 1 enforced
post-floor so PopArt's reward-magnitude EMA stays uncontaminated.
3. SABOTEUR_MIN bound violation — engagement-floor (0.1) was silently
pulling output below 0.5 stated minimum. Added post-multiplication
clamp to [SABOTEUR_MIN, SABOTEUR_MAX].
4. ratios[] undeclared — now __shared__ float ratios[6] block-loaded
once from ISV.
5. §3.6 slot count off-by-five (15 → 20). All 20 slots get FoldReset
entries; cold-start window behavior specified explicitly (no NaN path).
Architectural fixes:
6. Q-overconfidence concern addressed in scope — controller's
REWARD_CF_WEIGHT_INDEX covers CQL conservatism. No SP11′ deferral;
if T10 fails, extend controller outputs (e.g., target-update τ).
7. Curiosity recompute-at-replay specified (§3.5.1) — replay buffer
stores base reward only; curiosity added per-tuple at replay time
against current visit-count and current ISV. Eliminates stale-signal
replay contamination that would worsen ep1-overfitting.
Smaller fixes:
8. Novelty signal specified — SimHash 42×16 projection → 16-bit code,
1M-slot per-bucket count table, 1/sqrt(1+count). Hash table lives
in state-reset registry (cleared at fold boundary).
9. Saboteur engagement operationally defined — per-bar
|reward_with_saboteur − reward_without_saboteur| > EPS_ENGAGEMENT,
block tree-reduce, EMA'd. EPS_ENGAGEMENT = 0.01 × PnL_EMA.
10. Duplicate sentence in §7 component-starvation paragraph removed.
11. Validation criteria strengthened (§9): aggregation across 9
trajectories specified; primary metric = median peak-epoch ≥ 10
(baseline median peak-epoch = 1); secondary = drop-from-peak ≤ 5%
by ep20; tertiary = ep20 mean ≥ baseline ep1 mean. §9.3 fix-forward
response codified — no rollback.
12. Phases split per project pattern — Layer A additive (3 commits:
slots, canaries, controller), Layer B atomic consumer migration
(1 commit, including replay-time curiosity), Layer C validation +
close-out. No falsification gate; smoke is validation only.
13. Pearls A+D applied to controller outputs (§3.4.1) — chained
apply_pearls_ad_kernel after controller writes scratch, smoothed
values land in ISV. Consumers read smoothed slots.
Constants surviving (Invariant-1 anchors only, all rate-not-regime):
EPS_DIV, WEIGHT_HARD_FLOOR, SABOTEUR_MIN/MAX, CURIOSITY_PERMANENT_FRACTION,
CURIOSITY_BOUND_FRACTION, WEIGHT_FLOOR_FRACTION, ENGAGEMENT_FLOOR.
Spec is now full straight-up implementation; smoke validates Layer A
infrastructure and Layer B consumers, T10 validates SP11 success metric.
~1550 LOC across 3 commits in Layer A + 1 atomic commit in Layer B +
close-out in Layer C.
|
||
|
|
d49fbbe4e2 |
spec(sp11): reframe curiosity as feature + add permanent floor
User correction: curiosity is the *fix* for the ep1-peak overfitting pathology, not a hazard to defend against. Reframed §7 from "Risks" to "Design notes" — curiosity bound is a signal-relative scale, not a defensive cap. Fix the contradiction this exposed in the formula: previous `curiosity_pressure = stagnant_or_worse * curiosity_bound` went to zero when improving, which would cancel the always-on exploration the §7 narrative now relies on. Replace with permanent-floor pattern per pearl_blend_formulas_must_have_permanent_floor: curiosity_floor = 0.2 * curiosity_bound (CURIOSITY_PERMANENT_FRACTION) curiosity_dynamic = stagnant_or_worse * curiosity_bound curiosity_pressure = max(curiosity_dynamic, curiosity_floor) Now curiosity is always ≥ 20% of bound (anti-overfitting baseline) and rises toward the bound when stagnant (stagnation breaker). Updated unit-test guidance to assert pressure > 0 even at z=+10. CURIOSITY_PERMANENT_FRACTION=0.2 added to Invariant-1 fraction list. Saboteur-rising-with-improvement reframed as adversarial-load feature rather than over-stress risk. |
||
|
|
a72a23e6b1 |
spec(sp11): reward as controlled subsystem — fully ISV-driven
Brainstorm spec for SP11. Resolves the policy-stagnation pathology
surfaced in T10 train-multi-seed-xkjkb seed-0 ep0-14: model finds a
stable fixed point at ep1 (peak val sharpe 80.61), then OVERFITS to
it across remaining epochs (decline 80.61 → 70.58). Q-values grow but
val performance declines because reward function has no improvement
pressure.
Architecture (every input ISV-driven):
- Z-score-driven adaptation (no hardcoded "improving" threshold):
improvement_z = val_sharpe_delta_ema / max(val_sharpe_std_ema, EPS)
- 10 ISV outputs: 6 component weights + curiosity_pressure +
saboteur_intensity_mult + adaptive weight_floor + curiosity_bound
- 5 ISV canaries: val_sharpe_delta + val_sharpe_std (Z-score noise
estimate) + 6 per-component grad ratios + saboteur engagement +
PnL magnitude EMA (signal-relative curiosity bound)
- 4 new producer kernels (controller + 3 canary computers)
- Audit + migrate hardcoded cf_weight=0.3 in mse_loss_kernel.cu:318
and c51_loss_kernel.cu:789, plus other shaping multipliers
- NEW reward dimension: curiosity bonus, bounded by PnL magnitude
Per pearl_controller_anchors_isv_driven: every threshold replaced with
sigmoid(z) — no constants encode "what counts as improving". Per
pearl_blend_formulas_must_have_permanent_floor: every weight has
adaptive floor preventing zero-out. Per pearl_engagement_rate_self_
correction: saboteur intensity self-corrects via engagement rate canary.
Per pearl_cold_start_exit_signal_or: improvement signal OR'd from
multiple canaries so single-signal-failure doesn't stall controller.
New pearl authored alongside spec: pearl_reward_as_controlled_subsystem
— meta-principle that every reward path degree of freedom is a unified
controller output. Subsumes controller-anchor pearl at the reward layer.
Scope: 20 ISV slots, 4 producer kernels, audit + migration of
hardcoded reward shaping constants, 1 atomic commit.
~1300-1700 LOC; ~2.5-3 hours subagent work.
Success metric: val_sharpe[ep20] > val_sharpe[ep1] (Fix 33-38 baseline
peaked at ep1; SP11 should shift peak later as model continues
learning).
|
||
|
|
6ae9bc0055 |
spec(sp10): unconditional Thompson selector + ISV-driven temperature
Brainstorm spec for SP10. Resolves the val-Flat-collapse pathology that persisted through Fix 33-37: the eval-time argmax in experience_action_ select picks Hold deterministically every val bar (dir_entropy=0, trade_count=1 in 214k bars) regardless of controller state. Architecture: - Delete `if (eval_mode)` argmax branch in direction-selector kernel - Use temperature-blended Thompson: q_eff = E[Q] + τ × (Thompson - E[Q]) - τ = clamp(intent_eval_divergence / divergence_target, 0.5, 2.0) - τ self-corrects: collapse → high τ; healthy → low τ; permanent 0.5 floor - Reuses SP9's intent_eval_divergence_compute_kernel (extended, not new) Per pearl_controller_anchors_isv_driven: τ is ISV-driven, no hardcoded constants beyond Invariant 1 numerical anchors (clamp range). Per pearl_blend_formulas_must_have_permanent_floor: MIN_TEMP=0.5 ensures eval ALWAYS has stochasticity. The pearl_thompson_for_distributional_action_selection was about Bellman TARGET argmax (selector/target symmetry). It does NOT prohibit Thompson at the rollout selector. SP10 amends the pearl to clarify. Scope: 1 ISV slot, kernel modification (no new kernel — extend SP9's producer), 1 consumer kernel rewrite, test update, pearl amendment, audit doc Fix 38. ~300-500 LOC, single atomic commit per feedback_no_partial_refactor. |
||
|
|
38d2408604 |
spec(sp9): Kelly cold-start warmup floor — fully ISV-driven design
Brainstorm spec for SP9 (Fix 37). Resolves the val-Flat-collapse pathology
surfaced in T10 train-multi-seed-wsnc6 ep1-3: training intent_dist_f=0.47
but eval_dist_f=0.00 because Kelly cap pinned eval mag to Quarter, with
trade_count=1 in 214k bars (chicken-and-egg: cap closed → no trades → no
Kelly samples → cap stays closed).
Architecture per pearl_controller_anchors_isv_driven and
pearl_cold_start_exit_signal_or:
final_kelly_f = max(measured_kelly_f, warmup_floor)
warmup_floor = base_floor × (1 − combined_confidence)
base_floor = clamp(q_var_mag / q_var_mag_ema, MIN, MAX)
combined_confidence = max(statistical, behavioral, temporal)
All thresholds ISV-driven via Pearl D Wiener-α; only Invariant 1 numerical
anchors (clamp ranges, EPS) remain as constants.
Scope: 9 new ISV slots, 6 new GPU producer kernels including a
mandatory eval_dist GPU migration (eliminating DtoH readback at
gpu_backtest_evaluator.rs:1353 per feedback_no_cpu_compute_strict).
~1000-1200 LOC, single atomic commit per feedback_no_partial_refactor.
Pearls authored as part of brainstorm:
- pearl_controller_anchors_isv_driven (regime-encoded constants)
- pearl_cold_start_exit_signal_or (OR'd signals across statistical /
behavioral / temporal axes)
|
||
|
|
bcbd16f8c5 |
plan(sp7): v2 — fix 5 self-review violations from v1
Critical re-review of v1 caught: 1. Task 6 retained Pearl 2 kernel signature with (void) no-op args "to save 3-site cascade" — that's the partial-refactor anti-pattern feedback_no_partial_refactor explicitly forbids. v2 does the full contract change: signature shrinks 9→6 args, launcher migrates atomically. 2. SCRATCH_PEARL_2_C51/_CQL/_ENS would become orphan constants — feedback_wire_everything_up says delete or wire. v2 deletes them in T6. 3. Audit-doc updates were separated into a single late commit — would fail check_audit_doc_updates pre-commit gate on T1-T6. v2 folds Fix 31 stub into T1 and extends it commit-by-commit. 4. T5 referenced grad_decomp_*_result_dev_ptr field names that don't exist — actual layout is one shared 27-float pinned buffer with per-component byte offsets (iqn=0, cql_sx=24, c51=36). v2 uses pointer arithmetic from grad_decomp_result_dev_ptr. 5. Tasks 1, 2, 4 had "if file shape X then Y, otherwise Z" placeholder branches — writing-plans skill forbids these. v2 has concrete patches based on reading the actual files. Also corrected: T2 fixes the stale "regime_stability allocator" description in state_reset_registry.rs alongside the new entries. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
30f7032d12 |
plan(sp7): implementation plan for loss-balance controller
10 tasks: ISV slots → reset registry → kernel → build manifest → trainer load+launcher → Pearl 2 retreat → atomic launch+consumer+stale-doc → audit doc + memory pearl → L40S 5-epoch smoke verify → L40S 50-epoch full validation. Each task is producer-only or atomic-pair commits per feedback_no_partial_refactor. Verification gates (cargo check, cuda build, smoke acceptance criteria) embedded in each task. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
8f484f3312 |
spec(sp7): v2 — flatness-gated targets + Wiener α + sentinel-bootstrap
Self-review found 4 violations of feedback_isv_for_adaptive_bounds in v1: 1. TARGET_CQL_OVER_IQN=2.0 hardcoded → flatness-modulated per-branch target_ratio = ANCHOR × (1 − flatness[b]). Realizes the original pearl_q_flatness_gates_conservatism brainstorm idea. 2. TARGET_C51_OVER_IQN=1.0 hardcoded → mirror direction: target_ratio = ANCHOR × flatness[b]. 3. ALPHA_META=0.01 fixed → Wiener-optimal α from per-branch state-tracker EMAs (16 new ISV slots: diff_var/sample_var × 2 heads × 4 branches). 4. Cascaded floor cleanup: compute_adaptive_budgets line 3384 floors reads to 0.02/0.05 always, fighting the controller. Replaced with sentinel-aware bootstrap: cold-start uses BOOTSTRAP value, post-write takes ISV value verbatim. Acceptance criteria rewritten to be ISV-relative where the corresponding signal exists (q-spread vs NOISY_SIGMA, label_scale vs NOISY_SIGMA band). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
0808ebe3e5 |
spec(sp7): loss-balance controller — adapt CQL/C51 budgets to grad-norm ratio
Outcome-driven controller replacing Pearl 2's hardcoded-0 CQL budget and
floor-pinned C51 budget with multiplicative ratio adaptation. Targets
`grad_split_bwd cql/iqn = 2.0` and `c51/iqn = 1.0` per slice (trunk, dir,
mag). Slow EMA (α=0.01) consistent with existing controller pearls; Pearl
A sentinel-bootstrap on fold boundary; Pearl D Wiener-optimal smoothing.
Replaces ghost docstring at fused_training.rs:3409 ("0.10×(1−regime)×health"
formula was never implemented). Pearl 2 keeps owning IQN budget (reference
denominator) and FLATNESS_BASE (consumed by NoisyNet σ pearl); stops
writing CQL/C51/ENS slots so the new controller is the sole driver.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
6b01b70c3e |
docs(sp6): implementation plan — 3-sub-project parallel refactor (Pearls 2, 3, 5)
Three independent sub-projects (one per Pearl), each an atomic commit, each running in its own git worktree (sp6-pearl-2/3/5). Pearl 2 expands compute_adaptive_budgets() to per-branch [f32;4] arrays + trunk-mean scalars, dispatches branch-correction SAXPY sub-launches. Pearl 3 replaces noise_sigma: f32 with noise_sigma_per_branch: [f32;4] in ExperienceCollectorConfig and extends add_advantage_noise to per-branch lookup. Pearl 5 adds 4-branch tau buffer triplets to GpuIqnHead, runs 4 IQN forward passes per step (Approach B), each contributing iqn_trunk/4 to prevent 4x gradient accumulation. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
ca6e0007da |
docs(sp6): brainstorm spec — per-branch consumer infrastructure refactor
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
11979d7c08 |
docs(sp5): comprehensive implementation plan — 4 layers, ~21 tasks
SP5 implementation plan covering:
Layer A: 8 per-pearl commits (Tasks A0-A8)
A0: ISV slot constants foundation
A1: Pearl 1 per-branch atom span + Q-stats source
A2: Pearl 3 per-branch NoisyNet σ
A3: Pearl 2 per-branch loss budget (after Pearls 1+3)
A4: Pearl 4 per-group Adam β/β/ε
A5: Pearl 5 per-branch IQN τ schedule
A6: Pearl 6 cross-fold-persistent Kelly
A7: Pearl 8 per-direction trail distance
A8: Pearl 1-ext per-branch num_atoms
Layer B: 1 atomic commit (Task B1) — 11 consumer migrations
Layer C: validation + cleanup (Tasks C1-C6)
C1: Local GPU unit tests
C2: L40S 5-epoch smoke
C3: L40S 3-seed × 50-epoch full validation
C4: Pearl 7 investigation (post-validation)
C5: Audit doc + 8 memory pearls
C6: Layer C commit
Layer D: separate atomic commit (Tasks D1-D4)
D1: PnL aggregation kernel
D2: Health composition kernel
D3: Training metrics EMA kernel
D4: Layer D atomic commit
Total: 21 numbered tasks, ~110 ISV slots, 9-11 producer kernels,
11 consumer migrations.
Plan follows SP4's high-fidelity-for-Task-A1 + differential-pattern-
for-A2-A8 structure. Each task has concrete file paths, code blocks,
verification commands, exact commit messages.
Self-review:
- Spec coverage: all 9 pearls + 4 layers covered
- No TBD/TODO placeholders in step bodies
- Type/slot consistency verified across tasks
Refs: docs/superpowers/specs/2026-05-01-sp5-magnitude-differentiation-and-eval-collapse-design.md (HEAD
|
||
|
|
6e6e0fa114 | docs(sp5): remove duplicate Layer A close-out section (superseded by Layer D) | ||
|
|
af8fe57d1c |
docs(sp5): apply critical review fixes — Layer D, Kelly carve-out, Pearl 4 mitigations, acceptance loosened
Critical self-review surfaced 15 issues; user directed fixes:
- "a no cpu path": KEEP host-EMA close-out (rule compliance) — but split
into separate Layer D atomic commit (was mis-scoped as Layer A)
- Pearl 4 kept: 3 concrete risks documented (constant-β proof break,
β2 memory reset destabilization, ε numerical envelope), structural
envelope bounds added, ALPHA_META halved, ε-only fall-back path defined
- Pearl 6 Kelly cross-fold persistence carve-out: separate slot range
280..286, NOT in SP4 fold-reset registry, Invariant 1 architectural exception
- Commit ordering Pearls 1 → 3 → 2 → 4 → 5 → 6 → 8 → 1-ext (resolves
Pearl 2 circular dependency on 1+3)
- Acceptance criteria: correctness gates (must pass) + performance gates
(loosened to "not catastrophically negative", within 2σ of pre-SP5)
- Pearl 7 timing: explicitly Layer C step 4, post-Layer-B + 3-seed validation
- Pearl 8 enumeration: 4 slots (TRAIL_DIST_PER_DIR per direction)
- Pearl 9 collapsed: 0 slots (Thompson achieved via Pearl 1's atom adaptation)
- Total slot count corrected: 110 (was 120-128 inconsistent)
Layer structure: A (additive, 8 commits) → B (atomic, 11 consumers) → C
(validation + Pearl 7 investigation) → D (host-EMA close-out, separate
atomic commit). Layer D split off from A's "close-out" because PnL
aggregation pipeline migration is its own architectural concern.
User final review pending before invoking writing-plans skill.
|
||
|
|
3361944456 |
docs(sp5): resolve open questions Q1=c, Q2=a, Q3=a
Per user direction:
Q1 (Layer A granularity): per-pearl commits (~9 commits, decision c)
Q2 (Pearl 7 timing): investigate-only in SP5; fix in follow-up
if Bin(2,0.5) persists post-Pearls-1-3 (decision a)
Q3 (Pearl 4 Adam β): include as designed; accept theoretical risk
with Pearls A+D + EPS_CLAMP_FLOOR mitigation;
Layer C smoke monitors for destabilization;
fall-back to ε-only if observed (decision a)
Spec section updated:
- Layer A description: per-pearl commit structure
- Pearl 7 framed as investigation-only with conditional follow-up
- Pearl 4 documents theoretical caveat + mitigation + fallback
User final review pending before invoking writing-plans skill.
|
||
|
|
5f1c1eec51 |
docs(sp5): comprehensive per-branch + per-group adaptation spec — close all hardcoded-multiplier deferrals
SP5 design covers every known adaptive-parameter deferral in the DQN training loop in a single coherent project. After SP5: zero hardcoded multipliers, every adaptive value ISV-driven via Pearls A+D. 9 pearls + 1 sweep close-out + 1 validation milestone: 1-3. Per-branch atom span / loss budget / NoisyNet σ (52 slots) 4. Per-group Adam β1/β2/ε ISV-driven (24 slots) 5. Per-branch IQN τ schedule (20 slots) 6. Kelly cap signal-driven floors (6 slots) 7. dist_q/h/f Bin(2,0.5) audit + action_select fix (0-8 slots) 8. Trail stop signal-driven thresholds (6-8 slots) 9. Thompson direction-branch temperature (4 slots) 1-ext. Per-branch C51 num_atoms (4 slots) Layer A close-out: 5 host-EMA host→GPU migrations Validation: 3-seed × 50-epoch acceptance gate Total: 120-128 new ISV slots, ~5000-7500 LOC, 11-13 producer kernels, ~12 consumer migrations. Layer A (additive infrastructure, ~15 commits) → Layer B (atomic consumer migration, single coordinated commit) → Layer C (validation + cleanup). Mirrors SP4's layer pattern. Triggering data: train-multi-seed-cv2mw 50-epoch L40S baseline (terminated F0 ep10) revealed magnitude head Q-flatness, eval collapse, and frozen action distributions. Plus all SP4 close-out + sweep deferrals folded in per user direction "no deferrals — make a single plan based on ALL findings". Spec at: docs/superpowers/specs/2026-05-01-sp5-magnitude-differentiation-and-eval-collapse-design.md User review pending before invoking writing-plans skill. |
||
|
|
24accea774 |
fix(sp4): migrate fast/slow grad-norm EMA update to GPU per feedback_no_cpu_compute_strict
Layer C close-out C1 redesigned. Original plan claimed
grad_norm_slow_ema_pinned was orphan post-Mech-6 migration; verification
surfaced a SECOND live consumer (fold_warmup_factor_update, commit
|
||
|
|
c06169136c |
docs(sp4): implementation plan for signal-driven magnitude control
Comprehensive 3-layer implementation plan for the SP4 spec at docs/superpowers/specs/2026-04-30-sp4-signal-driven-magnitude-control-design.md. Layer A — additive infrastructure (Tasks A1-A18, no behavior change): - A1: 40 ISV slot index constants in new sp4_isv_slots.rs module - A2: wiener_state_buf (141 floats) + clamp_engage_per_block_buf (2048 ints) - A3: Pearls A+D shared helper (sp4_wiener_ema.rs) + 5 unit tests - A4: linear-histogram p99 device function (sp4_histogram_p99.cuh) + GPU test - A5-A9: 5 producer kernels (target_q + atom_pos × 4 branches + param_group_oracle × 8 groups + grad_norm + h_s2 = 15 fused producers per Pearl B) - A10-A11: wire 4 captured-graph + 10 cold-path launch sites - A12: 181 StateResetRegistry entries (40 ISV + 141 Wiener-state float + 2048 Pearl C int) - A13: retrofit 7 existing pre-SP4 producers with Pearls A+D - A14-A15: Pearl C engagement counter in 5 Adam kernels + host-side rate-deficit EMA + force-bump trigger - A16-A18: 40-test integration validation + L40S regression smoke + audit doc draft Layer B — atomic consumer migration (Task B1, ONE coordinated commit): - Mech 1, 2, 5, 6, 9, 10 consumers flip simultaneously to ISV-driven - Adam kernels' weight_decay arg from ISV[WD_RATE[group]] - L1 lambda from ISV[L1_LAMBDA_TRUNK_INDEX] - HyperParams config fields removed - Per `feedback_no_partial_refactor` Layer C — validation + cleanup (Tasks C1-C5): - C1: retire grad_norm_slow_ema infrastructure (Mech 6 migrated away) - C2: remove dead Mech 5 fused-kernel parameters - C3: L40S smoke validation against spec pass criteria - C4: comprehensive Invariant 7 audit-doc entry - C5: 5 new pearl memory entries + SP4 close-out project entry Estimated effort: 3500-4500 LOC across ~25 commits, 1.5-2 weeks of focused work, 1 L40S smoke validation. Bite-sized TDD tasks with exact file paths, complete code blocks, and exact build/test commands per task. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
d8666a232f |
docs(sp4): fix all 11 review findings — spec is now source of truth
Resolved every issue from the critical self-review:
1. Param-group count: 7 → 8 throughout. Slot total: 36 → 40 (8 groups × 3
per-group families = 24 + 16 single = 40). IQL high-tau and low-tau
are 2 distinct param groups (separate buffers + Adam states), not
one. The "(Resolved during implementation)" hand-wave deleted.
2. Reset semantics section rewritten — no longer references Xavier-
derived bootstraps. Sentinel 0 per Pearl A; Wiener-state triples
reset to 0 alongside ISV slots (160 reset entries total).
3-4. Weight decay and L1 lambda hardcoded α/bootstrap removed. Both
now follow universal Pearls A+D contract (sentinel-detect +
Wiener adaptive). For L1, the natural curriculum (λ ramps as
gradient differentiates) emerges from the signal itself, no
"Bootstrap = 0.0" constant needed.
5. ε naming collision resolved: ε_div = 1e-8 (Pearl D division-safety),
ε_clamp_floor = 1.0 (Pearl A consumer cold-start floor). Distinct
names, distinct purposes.
6. Per-signal kernel signature updated: removed stale `ema_alpha` arg,
added `wiener_state` pointer + `wiener_state_offset` + uniform
`alpha_meta` (structural meta-EMA constant).
7. Pearl C engagement counter race fixed. Replaced `clamp_engage_buf
[group] += 1` (race) with proper register-then-tree-reduce pattern
mirroring `dqn_grad_norm_kernel`. Per-thread register counter →
block-shared tree-reduce → single block-leader writes to
`clamp_engage_per_block_buf[engage_buf_offset + blockIdx.x]`.
Host sums across blocks. No atomicAdd, consistent with feedback.
8. Pearl D state buffer allocation spelled out: `wiener_state_buf:
MappedF32Buffer` of size 141 floats (47 slots × 3), per
`feedback_no_htod_htoh_only_mapped_pinned`. Reset-registry entries:
40 + 40×3 (new bound slots + Wiener triples) + 7×3 (retrofit
existing producers' Wiener states) = 181 reset entries.
9. `grad_norm_slow_ema` retirement spelled out in Layer C: removed
entirely (sole consumer Mech 6 migrates to ISV[GRAD_CLIP_BOUND]).
Also documented Q_ABS_REF=16 and H_S2_RMS_EMA=96 transition to
"monitoring-only ISV" — producers stay running, only the orphaned
Mech 1/2/5/6/9/10 consumer reads removed.
10. Histogram bin range fixed: linear-spaced bins from 0 to step_max
(avoids log(0) singularity from earlier log-spaced design).
Linear is also better for p99: top-of-distribution gets ~0.4%
resolution per bin. Degenerate "step_max == 0" branch handles
all-zero signals gracefully (skip ISV update, leave previous bound).
11. Pearl D-subsumes-Pearl-A claim CORRECTED: mathematically wrong.
At t=0, Pearl D's formula yields `x_mean[0] = 0`, not `x[0]`.
Pearl A's first-observation replacement requires an explicit
sentinel-detection branch in the producer. Both pearls are
necessary and complementary; they are not hierarchical.
Slot count math now consistent: 1 target_q + 4 atom_pos +
3×8 per-group + 1 grad_clip + 1 h_s2 + 8 wd_rate + 1 l1_lambda = 40.
Producer count: 15 fused (1 + 4 + 8 + 1 + 1).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
074b4f0865 |
docs(sp4): fold all 4 pearls into spec — A, B, C, D
PEARL A — First-observation bootstrap: eliminates Xavier-derived formulas (2.33 z-score, √2 std, √(2/K_in)) from the bootstrap section. Sentinel ISV[X]=0 at fold reset; producer step 0 detects sentinel, replaces directly with step_observation. Subsequent steps EMA-blend. Consumer cold-start safety via .max(1.0) numerical floor only (Adam-ε category, not magnitude). Truly zero magnitude constants in bootstrap. PEARL B — Fused per-param-group statistics oracle: producer count 36 → 14. Per-group fused kernel reads (params, grads, adam_m, adam_v) once and computes WEIGHT_BOUND, ADAM_M_BOUND, ADAM_V_BOUND, WD_RATE in one multi-pass operation. Trunk's oracle adds Pass E for L1_LAMBDA gradient-direction entropy. 4× memory bandwidth reduction. Cleaner conceptual unit (per-param-group bounds = one oracle). PEARL C — Engagement-rate self-correction: detects post-clamp feedback-loop saturation. For in-kernel clamps, theoretical engagement rate = 1% (top 1% by p99 definition). Producer-side rate-deficit EMA detects sustained deviation; force-bumps bound to step_max when detected. Resolves the "in-kernel feedback loop accepted" limitation from first draft. Per-Adam-kernel block-shared-memory engagement counter (no atomicAdd), block-wide reduce, host-side rate-deficit EMA. PEARL D — Wiener-optimal adaptive α: replaces all hardcoded EMA rates across 14 new producers AND 7 existing pre-SP4 producers. Per-step: α* = diff_var / (diff_var + sample_var + ε_num). Theoretically optimal under Wiener-filter analysis; subsumes Pearl A as t=0 edge case (both vars=0 → α=1 → first-observation replacement). On stationary signals: α→0 (smooth). On non-stationary: α→1 (track). Eliminates the recursion problem (α controlled by signal stats, not another α). ε_num = 1e-8 (Adam-ε numerical category). 3 floats of state per producer. 36+ hardcoded α values eliminated codebase-wide. Limitations section restructured: 6 of 8 first-draft limitations RESOLVED by pearls (in-kernel feedback loop, smoke time-budget, producer plumbing, stale-bound, 8 carved-out items, magnitude bootstrap formulas). Remaining 6 limitations are genuinely irreducible (F0 launch-scheduling variance, Layer B atomic flip risk, novel-pearl field-validation, equilibrium-formula non-stationarity, Pearl C counter-state cost, Pearl D state cost). SP4 now closes 100% of the magnitude/regularization/EMA-rate surface. No hardcoded scalars remain in the entire bound/clamp/EMA chain. Producer count: 14. ISV slot count: 36. Unit tests: 36. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
21b6058e07 |
docs(sp4): close all carve-outs — weight decay + L1 lambda fully signal-driven
Folds two new producer signals into SP4 scope, eliminating the earlier "carved out" exception: WEIGHT DECAY (per param-group, 7 new ISV slots): λ = |w·g| / max(||w||², ε) Derived from equilibrium analysis of d/dt(||w||²) = 2(w·g) - 2λ||w||². The equilibrium gradient-projection-onto-weight-direction divided by weight norm. Same theoretical-derivation category as Adam β values. EMA half-life α=0.005 (~140 steps). Bootstrap 1.0. L1 LAMBDA (trunk only, 1 new ISV slot — NOVEL PEARL): λ = (mean(|g|) / mean(|w|)) × D where D = (log K - H_observed) / log K is gradient-direction entropy deficit and H_observed = -Σ p[i]·log p[i], p[i] = ||g[:,i]|| / Σ ||g[:,j]|| L1 regularization-strength derives from gradient-direction entropy deficit across input features. When gradient is uniform across features (D≈0): network hasn't differentiated, λ=0 (no pruning). When gradient concentrates on few features (D≈1): network has identified what matters, λ ramps up to prune the rest. Self-curriculum — L1 strength tracks the emergence of feature differentiation. This extends pearl_adaptive_moe_lambda (regularization strength = EMA- tracked deficit of regularized quantity) to feature-redundancy domain. Pearl-name candidate (post-validation): pearl_signal_driven_regularisation_strength. Bootstrap λ=0 means cold-start = no L1 pruning; ramp-up only after gradient differentiates. Worst-case behavior is "L1 disabled" — graceful. Total ISV slot count: 28 → 36. Total producer kernels: 28 → 36. Effort estimate: 3000-4000 → 3500-4500 LOC, 1-1.5 → 1.5-2 weeks. Acknowledged limitations updated: removed item #8 (carve-out) since no carve-outs remain. Added items for L1 pearl novelty (untested) and weight decay equilibrium-formula non-stationarity. Both have graceful worst-case behavior and explicit validation criteria (#10 and #11) to detect anomalies. SP4 now closes 100% of the magnitude/regularization surface — no hardcoded scalars remain in the entire chain. AdamW config fields for weight_decay and l1_lambda removed from HyperParams to prevent accidental hardcoding regression. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
506f9443fe |
docs(sp4): fix critical issue J — post-clamp diagnostic dead path
Second-pass self-review found a new critical issue: under "diagnostic = clamp engagement", post-clamp diagnostics (Mech 9 weights, Mech 5 Adam m/v slots 36-43, slots 44-45) NEVER fire — because post-clamp |v| ≤ bound by construction, so producer-side comparison always returns false. Fix: split diagnostic implementation by clamp location. - Buffer-based clamps (Mech 1, 2, 10): diagnostic stays in producer (reads pre-clamp buffer, fires when step_max > bound). - In-kernel clamps (Mech 6, 9): diagnostic lives INSIDE the Adam kernel at the clamp step. Each Adam kernel takes a `diag_slot` arg alongside `weight_clamp_max_abs`; on clamp engagement, writes `nan_flags_buf[diag_slot] = 1`. Idempotent per-thread store (all threads writing 1, race-free per existing convention). No atomicAdd. Without this fix, slots 36-45 would be dead diagnostics under SP4 — detecting nothing, providing no signal. With this fix, every diagnostic slot fires on its corresponding clamp engagement regardless of where the clamp lives. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
cfe57109e3 |
docs(sp4): revise spec — fix 8 critical issues from self-review
Self-review found multiple problems with the first draft. This revision addresses all 8 critical issues: 1. P² parallelization claim was WRONG — P² is sequential. Replaced with dynamic-range histogram (3-pass single-block kernel: max-reduce → log-spaced bin → cumulative-from-top to find p99). 256 bins → ~0.4% quantile precision; numerical-precision derivation in same theoretical-constant category as floating-point precision. 2. Bootstrap values had wrong magnitudes — used Xavier σ instead of p99 of max-element. Recomputed: WEIGHT_BOUND[group] = 2.33 × √(2/K_in[group]); H_S2_BOUND = 2.33 × √2 ≈ 3.3. All bootstraps now from theoretical p99 under Xavier-init, not std. 3. EMA half-life of 700 steps (α=0.001) didn't converge in 5-epoch smoke. Revised α=0.005 (~140 steps) for weight/Adam producers — reaches ~99% convergence within one fold's training. Smoke now validates steady-state behavior, not just bootstrap. 4. Weight decay, L1 lambda, CLIP_MULTIPLIER were listed in scope but undesignable in producer-consumer pattern. CARVED OUT explicitly: weight decay + L1 λ → separate research-spec; CLIP_MULTIPLIER and MIN_CLIP subsumed by SP4's GRAD_CLIP_BOUND slot. 5. F0 ≥ 45 acceptance criterion was uncertain. Revised to F0 ≥ 37.5 (matches the 1e30-effectively-unclamped diagnostic smoke). The ~8-point F0 variance from launch-scheduling-shift is independent of clamp value; SP4 cannot guarantee F0=45 even with ideal design. 6. Per-param-group p99 plumbing concretized: each producer takes (offset, length) launch args; main DQN params buffer sliced into trunk/value/branch using existing param_sizes layout knowledge. 7. Pre/post-clamp feedback loop EXPLICIT: producer runs BEFORE consumer for buffer-based clamps (h_s2, target_q, atom_pos — in captured graph immediately before clamp). Producer runs AFTER for in-kernel clamps (weights, Adam m/v — feedback loop accepted with documented soft-anchor dynamics). No more hand-waving. 8. Unit test strategy: per-producer kernel test with synthetic Gaussian input → known p99 ≈ 2.33 → assert |computed - 2.33|<5%. 28 tests total. Catches algorithm bugs before L40S smoke. Effort estimate revised UP: 3500-4500 LOC (from 2-3000), 1.5-2.5 weeks (from 1-2). 28 producer kernels + 28 unit tests is more plumbing than first draft assumed. Self-review limitations section now explicit about: in-kernel feedback loop accepted, smoke time-budget marginally validates Adam EMAs, F0 may not return to 45, Layer B atomic-flip risk, plumbing density, stale-by-one-step bound, histogram-precision is design choice not tuning, carved-out items remain hardcoded post-SP4. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
e00736ce9b |
docs(sp4): signal-driven magnitude control design spec
Comprehensive design replacing every hardcoded magnitude multiplier in the SP3 mechanism stack (Mechs 1, 2, 5, 6, 9, 10) plus pre-SP3 mechs in the same magnitude-control surface, per feedback_isv_for_adaptive_bounds and feedback_adaptive_not_tuned. Core principle: the BOUND lives in an ISV slot, computed by a producer kernel as p99 EMA of observed signal magnitude. Consumer reads the slot and clamps directly — no multiplier between ISV read and clamp. Cold- start ε from theoretical-init bootstrap (Xavier, etc.) — same theoretical- constant category as Adam β values, not tuning knobs. Architecture: - 28 new ISV slots (7 base bounds × per-param-group split where appropriate) - Per-signal P² (Jain-Chlamtac) quantile producer kernels - Diagnostic = clamp engagement (sticky flag from producer's max-comparison) - Migration in 3 layers: additive infra → atomic consumer flip → smoke Out of scope: theoretical/structural constants (Adam β, Xavier formula, attention 1/√d, hidden_dim, num_atoms). EMA rates stay as documented statistical-design parameters (half-life ≈ observation time-window). Estimated: ~2-3000 LOC across 3 layers, 1-2 weeks, 1 L40S smoke. Awaiting user spec review before writing implementation plan. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
9b35fe9d45 |
plan(dqn): SP2 + SP3 implementation — fused NaN kernel + 5-mechanism Q-stability
Implementation plan for the approved combined design spec
(commit
|
||
|
|
ed727b51c0 |
spec(dqn): SP2 + SP3 combined — F0 regression fix + Q-learning structural stability
Designs the fixes for the two SP1 leftover issues identified at SP1 closure
(commit
|
||
|
|
05958d3a0c |
plan(dqn): SP1 numerical stability — F1 NaN root-cause investigation
Implementation plan for SP1 (Sub-project 1 of 3) of the numerical stability investigation. Follows the γ + β methodology from the spec at docs/superpowers/specs/2026-04-29-numerical-stability-investigation-design.md. 8 tasks across 4 phases: - Phase A (Task 1): γ read-only audit producing docs/dqn-backward-nan-audit.md - Phase B (Tasks 2-5): always-landing β instrumentation expanding nan_flags_buf 24→48 with 12 new backward-kernel NaN check slots + 12 reserved slots for future coverage - Phase C (Task 6): surgical fix(es), content-driven by audit + smoke topology, ISV-driven for any dynamic bound (mandatory) - Phase D (Task 7): multi-fold L40S smoke validation against 7 pass criteria (F0 ≥ 95% baseline, F1+F2 monotone improvement, zero NaN-CLAMPED-TO-ZERO, all 48 NaN flag slots remain at zero) - Closure (Task 8): audit doc closure, memory entry, SP2/SP3 handoff Operating principles (mandatory per spec): - No deferrals — anomalies discovered during investigation get fixed within SP1, not punted to SP2/SP3 - Combined RELATED fixes ship as rich commits (per feedback_no_partial_refactor) - ISV-driven design for any dynamic bound Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
792812baa1 |
spec(dqn): SP1 numerical stability — revise per critical review
9 substantive issues addressed inline:
1. ISV-driven design elevated from 'if applicable' to MANDATORY for
all dynamic bounds in SP1 fixes. Numerical-stability ε bounds are
the only carve-out (Invariant 1). Hardcoded tuning constants for
dynamic ranges explicitly rejected.
2. F0 Sharpe regression criterion changed from absolute (≥55) to
ratio-based (≥95% of latest baseline; floor 53.08 currently).
Prevents iterative erosion across multiple fix commits.
3. 'F1 trending positive' replaced with concrete monotone-improvement
test: Best Sharpe at last epoch ≥ Best Sharpe at first epoch of
the same fold.
4. Pass criterion distinguishes NaN-CLAMPED-TO-ZERO (failure) from
'Genuine grad collapse' (legitimate observation, permitted) per
the existing infrastructure from commit
|
||
|
|
dab5990287 |
spec(dqn): SP1 numerical stability investigation — F1 NaN root-cause
Brainstorm session 2026-04-29 produced this spec scoping Sub-project 1 of a three-sub-project decomposition addressing the F1 ep2 NaN explosion that persists across 9+ defensive layers in Plan C Phase 2 (commits e445d07a..e9096c7be). Decomposition: - SP1 (this spec): F1 NaN root-cause; γ audit + β always-landing instrumentation + surgical fix + multi-fold validation - SP2 (future): numerical stability framework — codify guards - SP3 (future): Q-learning structural stability — target-Q clip, pessimistic ensemble, atom-range governance Operating principles: - No deferrals (anomalies fixed within SP1, not punted) - Combined fixes — rich commits (per feedback_no_partial_refactor) - Always-landing diagnostic instrumentation (24-slot nan_flags_buf expands to 48; permanent regression sentinel) - F0 Best Sharpe ≥ 55 preserved (no regression on the working path) Pass criterion: all 3 folds train 5 epochs, F0+F1+F2 Sharpe ≥ 0, zero NaN-CLAMPED-TO-ZERO log lines, all 48 flag slots remain zero. 5 design sections approved iteratively: Architecture, Components, Data Flow, Error Handling, Testing. Next: writing-plans produces implementation plan. |
||
|
|
5de5e546a4 |
fix(dqn): Plan C T2 amendment — match production adaptive C51 support
Plan C as authored assumed:
- c51_probs_dir = post-softmax probabilities (didn't exist on collector;
only raw exp_b_logits is materialised)
- atom_values = single global linear support (production uses per-sample
per-direction adaptive [v_min, v_max, delta_z] + optional atom_positions)
- iqn_quantiles_dir = rollout-side IQN inference (doesn't exist; IQN is
training-only)
Amended kernel signature uses the production architecture:
- b_logits_dir [N, b0_size, n_atoms] — raw direction-branch logits
- per_sample_support [N, b0_size, 3] — adaptive per-direction support
- atom_positions [b0_size, n_atoms] — non-linear positions (NULL = linear)
- n_atoms
Direction-branch Thompson is now single-distribution over C51 (with adaptive
support), not joint C51+IQN. Single-distribution still provides the
principled posterior sample that fixes the UCB selector/target asymmetry —
the goal of Plan C is preserved, just sourced from the existing rollout-time
distribution instead of a non-existent IQN inference path.
New device-inline helpers: softmax_c51_inline, compute_atom_values_inline.
Plan + audit doc updated. Phase 0 standalone test kernel gained two new
entry points (direction_thompson_v2_test, argmax_eq_v2_test) matching the
amended production API; original Plan A entry points retained for tests
0.B-0.F. Tasks 3+4 (buffer wiring) unblocked — collector/evaluator already
have per_sample_support_buf and exp_b_logits in production.
|
||
|
|
8a535681b7 |
spec(dqn): Plan C Phase 2 ACTIVE — UCB asymmetry regression triggers Thompson resumption
The original PAUSED state was motivated by measurement-artefact hunt exit. The bug-hunt cycle is complete (commits |
||
|
|
fc4addaabf |
spec+plan(moe): correct gate input dim 42 → STATE_DIM=128
T1.6 implementer correctly identified that the gate input dim is
ml_core::state_layout::STATE_DIM=128, not the literal 42 the spec/plan
incorrectly stated. The 42-dim figure was the bar-feature subset; the
actual state vector is 128-dim (42 features + portfolio + MTF + OFI
padded to 128 for cuBLAS alignment).
Updated spec §3 architecture diagram, §4.1 gate subnetwork description
+ parameter count (3,272 → 8,776), and plan header architecture line.
Implementation in commit
|
||
|
|
76e55047ec |
spec+plan(moe): add no-HtoD/HtoH constraint; mapped pinned only
Per feedback_no_htod_htoh_only_mapped_pinned.md (newly recorded): every CPU<->GPU path in this redesign uses mapped pinned memory exclusively. No cudaMemcpy HtoD, no Vec-to-Vec defensive copies, including in test code. CPU is strictly read-only on the production surface. Plan changes: - New Task 2.0 promotes MappedF32Buffer / MappedI32Buffer from distributional_q_tests.rs local definitions to a shared crates/ml/src/cuda_pipeline/mapped_pinned.rs module so all kernel test wrappers (Test 0.F, upcoming MoE tests) share one implementation. Adds write_from_slice helper for direct host_ptr write (no memcpy). - Task 2.1 test wrapper rewritten to allocate mapped pinned buffers + write to host_ptr + read GPU-written output via host_ptr. No more memcpy_stod / memcpy_dtov in test code. Spec: new section 6.4 codifies the mapped-pinned-only constraint and references the shared module + reference implementation. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
69141d6266 |
plan(dqn): MoE regime redesign — bite-sized implementation plan
Implementation plan for the approved MoE spec at
docs/superpowers/specs/2026-04-27-moe-regime-redesign-design.md.
6 phases:
- Phase 0: use_* flag cleanup + count_bonus refactor (precondition,
partially in stash@{0})
- Phase 1: MoE infrastructure (additive — moe_lambda hyperparameter,
9 ISV slots, MoeGate/MoeExpert skeletons, params_buf layout extension)
- Phase 2: 4 new CUDA kernels (mixture forward/backward,
load_balance_loss, expert_util_ema_update) with TDD unit tests
- Phase 3: wire MoE into the training graph (forward, backward, loss
aggregation, ISV producer launch, HEALTH_DIAG aux_moe line)
- Phase 4: atomic deletion of vestigial RegimeConditionalDQN (per
feedback_no_partial_refactor.md)
- Phase 5: Layer 3/5 tests + L40S 30-epoch validation with explicit
kill criteria
Each task has bite-sized steps (2-5 min each) with exact file paths,
copy-pasteable code, exact commands, and TDD pattern (test fails first,
implement to make pass, commit).
Self-review: spec section coverage verified, no placeholders, type
consistency verified across phases.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
8629a9e7c2 |
spec(dqn): MoE replacement for vestigial RegimeConditionalDQN
End-to-end investigation (2026-04-27) confirmed RegimeConditionalDQN is vestigial decoration — 3 heads constructed at training start but only trending_head ever receives gradient updates. GpuDqnTrainer (the actual production GPU trainer) has zero references to RegimeType/regime routing; experience replay inserts go to trending_head.memory only; ranging_head and volatile_head stay at random init for the entire training run. Several support APIs (get_count_bonuses_branched, config, get_state_dim) hardcode-delegate to trending_head, ignoring the regime split entirely. Per `feedback_no_hiding.md` (wire up or delete) and the user's preference to fix not delete: the design wires regime conditioning properly via Mixture-of-Experts replacing the vestigial 3-head architecture. Pearl introduced and saved as `pearl_learned_gate_subsumes_handcoded.md`: when the network already sees the heuristic's inputs, a learned gate strictly subsumes any hand-coded discretization. This is the load-bearing rationale — ADX/CUSUM are already at state indices 40/41, so threshold-based regime classification is a strict information bottleneck the gate can recover and improve on. Design summary: - Architecture: shared GRN trunk -> K=8 small expert MLPs (256->64->256 bottleneck per expert, ~33k params each) -> learned gating network (state[42]->64->8 softmax) -> mixed h_s2 -> existing branching heads + C51 + IQN dual head. Soft full mixture (no top-k hardcoding); gate emerges peaky or flat from data. Anti-collapse load-balancing aux loss with default lambda=0.01 (configurable hyperparameter, not a kernel constant) prevents init-noise-dominated single-expert lock-in without forcing uniform utilization. User-confirmed signal: "collapses don't recover well in this codebase". - 9 new ISV slots (118-126: per-expert utilization EMA + gate entropy EMA), GPU-driven producer per `pearl_cold_path_no_exception_to_gpu_drives.md`. - 3 new small CUDA kernels (moe_mixture_forward/backward, moe_load_balance_loss) + 1 ISV producer; everything else is cuBLAS- reusable. CUDA Graph capture compatible. - Atomic deletion (no fallback): regime_conditional.rs (~700 LOC), RegimeType enum, classify_from_features, RegimeMetrics, RegimeClassConfig, 4 DQNConfig regime threshold fields, per-regime 3-file checkpoint format. DQNAgentType becomes thin wrapper over single DQN. Old checkpoints fail loudly with layout-fingerprint mismatch. - 5-layer testing strategy (unit kernels, smoke, gate-differentiation validation, L40S production validation with explicit kill criteria, architecture-hash backward-incompat). Out of scope (explicit): top-k routing, per-expert action heads, hierarchical MoE, regime-conditional CountBonus/NoisySigma broadcast, expert warm-start from existing trending checkpoint, CVaR action selection on mixed C51 distribution. Precondition: the in-progress use_* flag cleanup + count_bonus [f32; N] refactor lands as its own commit before MoE implementation begins, per `feedback_no_partial_refactor.md`. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
0e8804a770 |
spec(dqn): reframe Thompson rollout PAUSED (not SUPERSEDED)
Retracts the prior SUPERSEDED footer (commit |
||
|
|
42ffd6aadc |
spec(dqn): mark Thompson rollout SUPERSEDED — motivating pathology was a measurement artefact
The val-Flat-collapse / Short-collapse / C51 expected-Q bias hypothesis that motivated the 4-plan distributional-RL Thompson rollout was largely a measurement artefact in the diagnostic infrastructure, not a real policy pathology. Three layered bugs in actions_history_buf init + reader + mag_stats attribution conspired to inflate val_dir_dist Short, inflate active_frac, and pin wr_h/wr_f to zero. After fixing all three (commits |
||
|
|
3c61ee52ce |
plan(dqn): Thompson sampling rollout — Plans A+B+C+D (4 sequential phases)
Implementation plans for the distributional-RL aggregation spec
(docs/superpowers/specs/2026-04-26-distributional-rl-aggregation-design.md).
Plan A — Phase 0: TDD hypothesis verification (8 tasks)
- Standalone GPU test kernel: sample_c51_inverse_cdf,
sample_iqn_quantile_interp, compute_e_c51, compute_e_iqn,
thompson_direction_test, argmax_eq_test
- 6 unit tests (5 GPU + 1 CPU synthetic edge)
- GPU integration on converged checkpoint (Test 0.F)
Plan B — Phase 1: existing-lever audit (8 tasks)
- 6 unit tests covering B.2/CF/PopArt/Q-target audit fixture
- Exit gate: all 6 PASS = no reward-shaping bug; Phase 2 unblocked
Plan C — Phase 2: Thompson sampling integration (11 tasks)
- Direction branch of experience_action_select rewritten:
eps-greedy + Boltzmann → Thompson (training) + argmax E[Q] (eval)
- C51 + IQN buffers wired to gpu_dqn_trainer + gpu_backtest_evaluator
- train_active_frac HEALTH_DIAG instrumentation
- 4 GPU-direct tests against production kernel
- Aggregation Contract table + memory pearl
Plan D — Phase 3: long verification + dead-code cleanup (8 tasks)
- Test 3.A: regression anchor against original C51 Flat bias
- L4: 1 seed × 6 folds × 30 epoch (~1 hour)
- L5: 5 seed × 6 fold matrix per Plan 5 Task 5 (Tier 1+2+3 PASS)
- Direction-only eps_dir adaptive boost + Boltzmann tau-floor cleanup
(per feedback_no_partial_refactor.md)
- nsys regression check; final architecture/spec footer
Plans gated sequentially: each phase exit gate must pass before next.
|
||
|
|
021bb0ef73 |
spec(dqn): pivot Phase 0+2 tests to GPU-direct (no CPU mirror)
User correctly identified that CPU mirror function tests don't test
the production GPU code path. A bug shared between mirror and kernel
(translated identically wrong) would slip through. Mirror tests + a
single GPU bridge test were a weak compromise.
GPU-direct testing strategy:
- All Phase 0 kernel-correctness tests (0.A, 0.B, 0.C, 0.D, 0.F):
launch tiny test-only kernels with the SAME math the Phase 2
production kernel will use; assert properties of the output.
- Test 0.E (synthetic edge discovery): stays CPU. It tests an
ALGORITHMIC PROPERTY of Thompson exploration (does it discover
edge if edge exists?), not a kernel correctness property.
- All Phase 2 unit tests (2.A-2.D): GPU-direct against the
modified production kernel.
- Phase 0.F (real checkpoint extraction): unchanged — already GPU.
Local development uses RTX 3050 GPU (per memory user_dev_environment.md).
CI runs --ignored flag to skip GPU tests on CPU-only runners.
Time budget: Phase 0 was 1-2 days (CPU mirror); now 2-3 days
(GPU-direct, includes kernel wrapper setup half-day).
Other delta:
- Phase 0 deliverable file renamed: distributional_q.rs ->
distributional_q_tests.rs (no mirror functions, just tests +
kernel wrappers).
- Phase 2 unit tests rephrased to launch production kernel rather
than compare against CPU mirror.
- L1 verification gate runtime: seconds -> minutes (GPU launch
overhead per test).
The user's intuition was right: testing production directly is the
honest approach. Mirror was an optimization that traded correctness
for speed; with local GPU available the optimization isn't needed.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
25ba051157 |
spec(dqn): critical review pass — fix all majors, mediums, minors
Self-review identified 8 major + 8 medium + 15 minor issues. All fixed:
MAJORS:
M1 Test 0.A: clarified — compare argmax(E[Q]), Boltzmann(E[Q]),
and Thompson sampling distributions; assertion is on relative
ordering across all three.
M2 IQN quantile count: replaced hardcoded `5` with N_IQN_QUANTILES
constant (defined per task #147 fixed-quantile design).
M3 "Converged checkpoint" definition: ≥60 epochs trained AND
val_sharpe stabilised (no >10% change over last 10 epochs).
Cites prior 60-epoch validation runs (task #80, train-7rgqd).
M4 R3 reframed: replaced "no issue" handwave with explicit
by-design tradeoff acknowledgment + cost analysis. Wasted
exploration is the cost of finding out whether edge exists.
M5 Test 0.D σ_long=0.05 justified: chosen to match expected order
of magnitude given typical |return| ~ 50bps; Phase 0.F
validates against real checkpoint.
M6 rng_ctr post-increment: clarified — matches existing pattern
at experience_kernels.cu:858 (no behaviour change).
M7 train_active_frac instrumentation: NEW Phase 2 deliverable —
existing HEALTH_DIAG only has val_active_frac, but L3 verifies
training-time active_frac. Spec now explicitly adds this
~10-line metrics.rs change as a Phase 2 deliverable.
M8 eps_dir cleanup code-level detail: explicit reference to
experience_kernels.cu lines 814-865; remove eps_dir from both
static EPS_FLOOR clamp AND adaptive boost block; verify
variable can be removed from kernel signature via grep.
MEDIUMS:
Med1 Current C51/IQN combination: clarified that compute_expected_q
blends per training schedule; Phase 2 replaces with explicit
0.5*E_C51 + 0.5*E_IQN equal weighting; Phase 0.F verifies.
Med2 Eval mode phrasing: "eval mode already sets eps=0 in existing
kernel" — no semantic override, factually correct.
Med3 Magnitude σ claim: clarified — magnitude branch likely has σ
bias in OPPOSITE direction (Full has larger σ; UCB would
prefer Full and worsen saturation). Empirical verification
deferred. Phase 0.F should also report per-magnitude σ.
Med4 Hierarchical sampling claim corrected: it's not about
balancing 50/50 (already 50/50). It's about decoupling
cluster-best decisions; clarified.
Med5 n_atoms vs N_IQN_QUANTILES: clarified — n_atoms variable per
config (currently 51); N_IQN_QUANTILES fixed at 5.
Med6 Conviction code: removed pseudo-code; references existing
implementation at experience_kernels.cu:1091; provides
implementation hint for E[Q] reuse.
Med7 Q-target propagation: clarified — uses full distribution
(C51 atom projection / IQN quantile regression), not just
E[Q]. Thompson modifies action selection only.
Med8 References: added Thompson 1933 (original), Bellemare 2017
(C51), Dabney 2018 (IQN) for theoretical foundations.
MINORS:
Min1 Date updated to 2026-04-27.
Min2-3 Argmax monotonic /2 simplified out — argmax(a+b) =
argmax((a+b)/2). Code clarity improved.
Min4 P(argmax picks Long) = 0 deterministic; reframed assertion.
Min5 Test 0.F structural assertions added: σ_C51[FLAT] < 0.01 ×
σ_C51[LONG]; same for IQN; E[Q_FLAT] > E[Q_LONG]; argmax
picks FLAT; Thompson P(LONG)+P(SHORT) ≥ 0.20.
Min6 -INFINITY → CUDART_INF_F (CUDA convention).
Min7 dir_idx scope: comment notes it's declared earlier in kernel.
Min8 action_select args: explicit — three buffers exist on GPU
but not currently passed; new params, no new buffers.
Min9 Phase 0 time math: 5 hours tests + 1 hour enumeration + 2
hours 0.F + (3 hours runtime if checkpoint training needed,
runs in parallel). Honest budget.
Min10 "Two-stream" → "5-Layer Gate" header.
Min11 Plan 5 reference uses full path consistently.
Min12 Plan B time budget: explicit 1 day if pass; 2-5 days if bug.
Min13 active_frac: clarified Long+Short combined, not per direction.
Min14 train_active_frac: now in Phase 2 deliverables (see M7).
Min15 "20 mechanisms" → 21, with sub-counts in section headers.
Spec now 615 lines, comprehensive coverage of:
- Pearl + theoretical foundation
- Problem statement (with bias-might-be-correct caveat)
- Architecture (Thompson at training, argmax at eval)
- Train vs eval distinct selectors with behavior-change disclosure
- Direction-only scope with magnitude σ-bias warning
- Conviction stays E[Q]-based (no Kelly cap jitter)
- Interaction matrix: 21 mechanisms in 3 categories
- 6 v2 enhancements documented + deferred
- Phase 0/1/2/3 with tests, exit gates, time budgets
- 5-layer verification + train_active_frac instrumentation
- 8 risks with mitigations + 5 stop conditions
- What v1 doesn't touch (referencing interaction matrix)
- References (Thompson 1933, Bellemare 2017, Dabney 2018, etc.)
- Aggregation contract (project-wide pearl, enforced)
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
308c54484e |
spec(dqn): full interaction matrix + outside-the-box Thompson improvements
Per-user direction: every existing mechanism in the DQN system MUST be
explicitly considered for interaction with Thompson, and Thompson itself
MUST be examined for system-specific improvements beyond vanilla.
INTERACTION MATRIX (3 categories, 20 mechanisms):
Category 1 — Compose with Thompson (no change required):
Counterfactual reward, B.2 novelty bonus, PopArt, Saboteur,
Curiosity, NoisyNets/VSN, Distillation, CQL, Polyak target EMA,
HER, PER, Replay warm-start, Multi-fold validation harness.
Category 2 — Trivially adapt to Thompson (one-line changes):
D7/N7 contrarian sign flip (negate the SAMPLE), cosine epsilon
schedule (still applies to mag/ord/urg), per-sample epsilon (IQL
expectile gap), adaptive Boltzmann tau (still applies to mag/ord/urg).
Category 3 — Take precedence over Thompson (hard constraints):
Plan-based action lock (Thompson sample discarded if plan active),
per-magnitude Kelly cap, trail stop, capital floor breach.
Critical insights from the audit:
1. NoisyNets is ALREADY a form of training-time Thompson at the
parameter level. Output-space Thompson stacks on top —
total exploration = parameter-space ⊗ output-space (multiplicative).
2. Curiosity is ORTHOGONAL to Thompson — Thompson explores actions
whose Q is uncertain; curiosity explores states whose dynamics
are uncertain. Both axes desirable; no conflict.
3. Plan lock takes precedence; same as currently with Boltzmann.
OUTSIDE-THE-BOX v2 ENHANCEMENTS (deferred to follow-up specs):
v2.1 Triple-source Thompson (C51 + IQN + Ensemble) — incorporate
the existing ensemble Q-head as 3rd uncertainty source.
v2.2 Persistent Thompson (anti-churn for HFT) — bias sampling
toward current direction, ISV-driven; reduces tx_cost from
Long/Short oscillation across bars.
v2.3 CVaR-aware eval (risk-adjusted deployment) — eval picks
argmax(E[Q] − λ·CVaR_α[Q]); risk-aware decision making for
production with real capital.
v2.4 Information-Directed Sampling (Russo & Van Roy 2014) — picks
action minimizing regret²/info_gain; more efficient than
vanilla Thompson when learning saturates.
v2.5 Hierarchical Thompson on (trade vs no-trade) → (which
direction) — addresses 50/50 structural advantage of no-trade.
v2.6 Composition with curiosity-driven exploration — explicit
coupling beyond reward-side composition.
Each v2 enhancement gets its own spec/plan when prioritised. Vanilla
Thompson is v1; ships first; verified independently.
ALSO FIXED (from earlier self-review):
- Pearl claim softened: only the C51 Hold/Flat bias is directly
attributed; other historic bugs had different mechanisms.
- TFT entry removed from contract table — TFT is a Variable
Selection Network (feature processor), not a Q-head. Replaced
with generic "Future Q-head additions" placeholder.
- Eval direction = argmax E[Q] explicitly flagged as a behavior
change from current Boltzmann-with-tau (val_dir_dist will be
more concentrated than current).
- Phase 0.E budget reduced (1-2 hours, not 1 day) — synthetic
bandit is ~50 lines of Rust, not full RL training loop.
- Phase 0 enumerates existing checkpoints before training new one.
- Architecture diagram parenthesis fixed.
- Conviction implementation note: compute E[Q] once, reuse for
conviction AND eval-mode argmax — no redundant computation.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|