Commit Graph

128 Commits

Author SHA1 Message Date
jgrusewski
3ab87d7672 plan(dqn-v2): Plan 3 self-review fix — remove unimplemented!() from example
Self-review caught that the CqlAlphaSeedCoupledController read_signals
example used unimplemented!("plumbed by caller"), which violates the very
Invariant 9 the plan enforces. Replaced with a concrete implementation
that reads SEED_FRACTION_EMA_INDEX from the ISV bus (a Plan 1 reserved
controller-signal slot).

No other plan had placeholder violations. The remaining TBD/TODO/FIXME
grep hits across Plans 1-5 are either the pre-commit hook's rejection
patterns (correctly quoted) or enforcement scans that DETECT the markers
in audit docs.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-24 10:07:10 +02:00
jgrusewski
74071af908 plan(dqn-v2): Plan 5 — validation substrate + final pass implementation plan
Fifth and final plan decomposing the DQN v2 unified spec
(docs/superpowers/specs/2026-04-24-dqn-v2-unified-design.md §2, §4.A.3,
§4.A.4, §4.A.4.1). No new policy mechanisms — orchestration, automation,
regression detection, and the final sign-off run.

Covers spec sections:
- §4.A.3 Multi-seed × multi-fold Argo harness (N=5, K=6 defaults via
  --multi-seed / --folds flags on argo-train.sh)
- §4.A.4 Regression detection hard-stop (2N consecutive error-band
  metrics → self-SIGTERM with TerminationReason)
- §4.A.4.1 nsys profile harness with --profile flag and
  compare-nsys-profiles.py for 20% per-epoch regression detection
- §2 Tier 1 + Tier 2 + Tier 3 final exit criteria via scripts/validation/
  suite (tiered tier1/tier2/tier3 check scripts + wrapper)
- §2 DQN v2 rollout close-out (spec marked LANDED, retrospective,
  plan-level validation docs archived)

Plan structure: 6 tasks — 4 infrastructure (harness, regression-detect,
nsys, validation-scripts) + final-run task + rollout-closeout task. The
final-run task is the sole exit-gate for the DQN v2 rollout: ALL THREE
tiers must PASS on the first final-validation run (per stop-on-anomaly:
retrying with different seeds hoping for a different outcome is anti-
pattern).

Preserves all 9 invariants. Every mechanism wired, no stubs, no TODO/FIXME.
Retrospective written from actual implementation experience after Plans
1-5 all land.

This is the terminal plan. Plans 1-5 together constitute the complete
DQN v2 rollout. After Plan 5's final-run exit passes, the spec is marked
LANDED and the rollout is declared complete.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-24 10:06:25 +02:00
jgrusewski
691d769bb3 plan(dqn-v2): Plan 4 — supervised→DQN concept adoption implementation plan
Fourth of five sequential plans decomposing the DQN v2 unified spec
(docs/superpowers/specs/2026-04-24-dqn-v2-unified-design.md §4.E).

Covers the six Part E IN decisions:
- §4.E.1 TFT Variable Selection Network across 6 feature groups
  (market/OFI/TLOB/MTF/portfolio/plan_isv) with vsn_feature_selection kernel
- §4.E.2 Gated Residual Network — CONDITIONAL branch (ADOPT vs CANONICALISE)
  based on docs/ml-supervised-to-dqn-concept-audit.md row
- §4.E.3 Multi-quantile IQN heads (5/25/50/75/95) replacing single CVaR output
- §4.E.4 Encoder-Decoder separation — explicit StateEncoder + per-branch
  ValueDecoder with D.3 horizon-decomposed V_short/V_long sub-heads
- §4.E.5 Attention-weight interpretability — 7 new ISV slots [65..72) for
  per-group attention focus EMAs + Mamba2 retention proxy
- §4.E.6 Multi-task auxiliary heads — next-bar return MSE + 5-bar regime CE,
  ISV-coupled aux-weight schedule (sharpe-reactive)

ISV_TOTAL_DIM seals at 72 with this plan's allocations. Part E audit doc
closes out all rows (zero TBD/evaluate remain per Invariant 9).

Plan structure: 8 tasks (6 feature tasks + audit close-out + validation).
TDD-disciplined steps. Task 2 documents the CONDITIONAL branch decision
pathway explicitly (ADOPT vs CANONICALISE Branch A/B structure).

Plan 4 exit gate: Tier 1 convergence RETAINED (no regression vs Plan 3
baseline) + aux heads produce measurable signal. Tier 2 + Tier 3 checked
by Plan 5.

Preserves all 9 invariants. Zero stubs. Zero TODO/FIXME. Every Part E item
either lands fully or is explicit OUT (xLSTM, KAN); no third path.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-24 10:03:13 +02:00
jgrusewski
c70bbd5d15 plan(dqn-v2): Plan 3 — behavioural core implementation plan
Third of five sequential plans decomposing the DQN v2 unified spec
(docs/superpowers/specs/2026-04-24-dqn-v2-unified-design.md).

Covers spec sections:
- §4.C.2 Reward-component attribution (6 ISV slots [52..58), reward_split[..]
  in HEALTH_DIAG) — lands first as diagnostic substrate for the rest
- §4.B.1 ISV-driven continuous Flat opp-cost (scales by isv[Q_DIR_ABS_REF])
- §4.B.2 Novelty-weighted trade-attempt bonus (2 ISV slots [50, 51])
- §4.B.4 Adaptive plan-MLP activation threshold (1 ISV slot [49] +
  PlanMlpThresholdController via the AdaptiveController trait)
- §4.C.4 Temporal timing bonus (trade-exit, bars_early / hold_time)
- §4.D.4 Generalised temporal reward coupling (3 ISV slots [59..62):
  persistence credit, regime-shift penalty, conviction consistency)
- §4.C.3 State-distribution divergence signal (3 ISV slots [58, 63, 64])
- §4.B.3 Replay buffer seeded warm-start (4 scripted policies,
  ≥100K experiences, CQL α decoupled from seed phase)
- §4.C.5 CQL alpha schedule coupled to seed-phase decay via
  CqlAlphaSeedCoupledController

Plan structure: 9 implementation tasks + pre-plan verification + exit-gate
validation task (Task 10). Each task TDD-disciplined with bite-sized
steps. Task 1 (C.2 diagnostic audit) ordered first so subsequent
reward-shape tasks can verify contributions rather than log-archaeology.

Plan 3 exit gate: multi-seed (N=5) × multi-fold (K=6) run passes Tier 1
(convergence) + Tier 2 behavioural subset (val trades/bar ≥ 0.005,
val_active_frac > 0.2). Tier 3 profitability deferred to Plan 5.

Preserves all 9 invariants. Zero stubs. Zero TODO/FIXME. Every new module
wired to its consumer in the same task it lands.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-24 09:59:32 +02:00
jgrusewski
a5101eb2f8 plan(dqn-v2): Plan 2 — temporal core implementation plan
Second of five sequential plans decomposing the DQN v2 unified spec
(docs/superpowers/specs/2026-04-24-dqn-v2-unified-design.md).

Covers spec sections:
- §4.C.1 Quantile-based atom support (8 new ISV slots, q_quantile_reduce kernel)
- §4.D.1 Mamba2 backward pipeline completion (no atomicAdd — per-sample arrays + host reduce)
- §4.D.2 Per-branch gamma via AdaptiveController (4 new ISV slots)
- §4.D.5 Soft fold-boundary transitions (extends StateResetRegistry with SoftReset category)
- §4.D.7 Liquid Time-constant audit (trace fire rate, decide wire-or-delete)
- §4.D.3 + §4.D.6 + §4.D.8 coordinated state-layout migration (horizon-decomposed V +
  plan_isv[6] + TLOB integration with atomic ISV schema version bump 1→2)

Plan structure: 7 tasks + pre-plan verification, all TDD-disciplined with
bite-sized steps (test → fail → implement → pass → commit). Task 6 is
explicitly atomic to preserve Invariant 5 (state-layout consistency).

Plan 2 prerequisites: Plan 1 landed (StateResetRegistry, AdaptiveController
trait, ISV schema version, audit docs). Plan 2 exit: all 9 invariants
preserved; convergence-scaffolding run passes with 16 new ISV slots
populated; ISV schema version == 2.

Preserves all 9 invariants. No stubs. No TODO/FIXME. Every new module
wired to production path in the same task it lands in.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-24 09:53:11 +02:00
jgrusewski
e5f6afca8d plan(dqn-v2): Plan 1 — substrate & refactor implementation plan
First of 5 implementation plans for the DQN v2 unified integrated
policy system spec (d13b53586, 336ee40b9). Plan 1 covers:

  Task 1: Audit doc scaffolding + pre-commit hook (Invariant 7, 9)
  Task 2: StateResetRegistry definition (A.1)
  Task 3: Wire StateResetRegistry into fold-boundary reset (A.1)
  Task 4: Named-dimension refactor — ps[], plan_isv[], plan_params[],
          branch indices, dir/mag sub-indices (Invariant 8)
  Task 5: ISV schema version at ISV[0], fail-fast on checkpoint
          mismatch (A.2)
  Task 6: Orphan audit populated (A.5)
  Task 7: Hot-path purity audit populated; MIGRATE calls fixed (A.6)
  Task 8: AdaptiveController trait + harness (C.6)
  Tasks 9–17: Migrate each adaptive controller to the trait in spec-
            specified order (atoms → gamma → Kelly → cql_alpha → tau
            → epsilon → conviction_floor → plan_threshold → balancer)
  Task 18: Plan 1 validation run (smoke + 3-epoch L40S)

Plan 1 landing criteria: all 9 invariants preserved across every
commit, all smoke tests pass, audit docs fully populated. Plan 2
(temporal core) starts only after Plan 1 passes exit criteria.

Plans 2–5 cover Parts B, C (minus C.6 already in Plan 1), D, E, and
final validation. Planned as separate documents in docs/superpowers/
plans/2026-04-24-dqn-v2-plan-{2..5}-*.md.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-24 09:41:22 +02:00
jgrusewski
d1068de2b8 plan(policy-quality): reframe DELETE tasks per no-functionality-removal rule
User directive: "we don't remove functionality at all, we fix what's
broken". Standing rule captured in
~/.claude/projects/-home-jgrusewski-Work-foxhunt/memory/feedback_no_functionality_removal.md.

Phase 2 plan reframed:
  * Task 2.3 (R5 micro-reward): was DELETE; now FIX — either (a) set a
    non-zero scale that delivers real signal, or (b) document why it's
    intentionally disabled while preserving the code path.
  * Task 2.4 (R6 loss-aversion): unchanged — the relocation to C51
    Bellman target is consolidation, not removal. Invariant preserved.
  * Task 2.6 (E4 entropy-reg): was DELETE-CANDIDATE; now TUNE — if the
    signal doesn't reach magnitude, re-route it rather than remove.
  * Task 2.7 (C4 grad-clip): was DELETE-CANDIDATE; now
    KEEP-AND-DOCUMENT — gradient clipping is a safety feature; run the
    ablation for diagnostic, not for removal.
  * H9 delete-magnitude-branch fallback: rejected permanently. On
    magnitude convergence failure the fallback is per-magnitude reward
    shaping / per-bin advantage weighting / curriculum / state enrichment.

Task 2.5 wiring-bug sweep unchanged (correctness fixes, not removal).

This commit updates only the planning narrative. Task execution hasn't
started on 2.3/2.6/2.7 yet, so no code changes required here. When
those tasks run, they will implement the fix-paths, not the
delete-paths.
2026-04-22 11:18:28 +02:00
jgrusewski
7e78bf4f85 plan(policy-quality): Phase 2 pivot — H4 REJECTED by Task 2.0 data
Task 2.0 instrumentation (commits d60e5375a / 980f3b07f / 41b0c559c)
revealed two silent bugs in Task 0.4's grad_ratio_mag_dir accessor:

  Bug A — readback size mismatch: pinned buffer allocated at
  total_params, but grad_buf length is total_params+cutlass_tile_pad,
  so size check always failed → Err silently coerced to 0.0 by proxy.

  Bug B — readback timing wipe: estimate_avg_q_value_with_early_stopping
  in process_epoch_boundary replays the forward graph, which zeros
  grad_buf. Any subsequent readback sees all zeros.

Both fixed in 41b0c559c; snapshot hoisted to top of process_epoch_boundary.

With those bugs fixed, the measured gradient ratio is NOT 0.0000 — it is
50–400× mag/dir across 60 epochs. Magnitude branch is over-fed, not
starved. Direction gradient is small but non-zero (~2e-2 to 7e0).
Direction policy is observably healthy (Short/Hold/Long/Flat 38/12/42/14%).

Track 1 triage's H4 CONFIRMED verdict was a measurement artefact
produced by Bugs A+B. H4 as originally defined is now REJECTED.

Plan changes:
  - Add new "Task 2.0 findings" section after cross-cutting concerns,
    documenting bug A, bug B, observed data, and revised strategy.
  - Task 2.1 marked DEFERRED (not deleted — kept for reference).
  - Task 2.2 promoted to PRIMARY fix.
  - New fallback: H9 delete-magnitude-branch as replacement for Task 2.1
    if Task 2.2 alone is insufficient.
  - Task 2.0 inventory row updated with LANDED status and commit chain.

Net effect: Phase 2 simplifies. The hardest task (2.1 three-branch
architectural fix) is skipped; primary path is Task 2.2 (~25 LOC).
2026-04-22 11:13:48 +02:00
jgrusewski
5a4d561451 plan(policy-quality): Task 2.0 revised approach — in-graph pinned snapshots
First Task 2.0 dispatch escalated BLOCKED: the four loss-component
backward kernels are captured inside the fused training graph, so
host-side snapshot-between-components isn't possible mid-graph without
a force-ungraphed diagnostic step (~210 LOC + cross-stream sync risk).

Revised approach (chosen after cost analysis):

  cudaMemcpyAsync(device → pinned host) IS captureable in a CUDA graph.
  Even better: DtoD into per-component scratch buffers, then an in-graph
  reduction kernel computes per-component (mag_norm, dir_norm) and writes
  8 floats to a pinned result slot. Only the 8-float result crosses
  PCIe (at epoch boundary), keeping per-step PCIe traffic to zero.

Changes to the plan's Step 2 + Step 3 + Step 4:
  - Step 2: added 4 device-side scratch buffers (one per component,
    ~10 MB each = 40 MB device) + 8-float pinned result slot + new
    reduction kernel grad_decomp_kernel.cu spec'd out.
  - Step 3: clarified that DtoD snapshot + backward + reduction kernel
    are ALL captured in the graph; graph replays them every step;
    no force-ungraphed dance needed.
  - Step 4: added refresh_grad_component_norms() accessor that reads
    the 8-float pinned slot at epoch boundary (zero-copy) and populates
    the host-side cache.

Approach matches Task 0.4 pattern (commit bb42c9963) extended four-fold.
No atomicAdd (reduction uses shared-mem tree), no non-captured replay,
no cross-stream sync risk.

LOC estimate: ~90 (was ~210 for the rejected option).
2026-04-22 09:46:17 +02:00
jgrusewski
2a54edc92d plan(policy-quality): Phase 2 plan refinements (tolerance bands + decision tables)
Three polish items from the plan review:

1. Task 2.4 Step 5 — replaced vague "must not regress" with a numeric
   tolerance-band table: multi_fold best_val_metric ±15% per fold,
   Best Sharpe ≥ 20 floor, and hard F_Half/F_Full ≥ 0.05 + cf_flip ≥ 0.1
   BLOCKERS. Anchors to the baseline metrics doc values captured at
   policy-quality-baseline.

2. Task 2.6 — converted the ad-hoc text decision at Step 1 into the same
   five-row decision table format Task 2.1 uses (DELETE / KEEP / KEEP-as-
   safety-net / INCONCLUSIVE / DELETE-because-inseparable). Step 2
   ablation got its own 4-row outcome table keyed to ent_mag delta,
   multi-fold Sharpe regression, and NaN appearance.

3. Task 2.7 Step 1 — added "Repeat 3× with different seeds" clause and
   explicit total wall-clock note: ~75 min (3 × 25 min) for the ablation
   pass before the Step 2 decision. Aligns with the Cross-cutting concern
   #6 about sample-noise rejection.

No task count change (still 11: 2.0–2.10). Net code-delta estimate
unchanged. Standing-rule compliance unchanged (no stubs / no atomics
/ no quickfixes / no hiding / no feature flags / no push-per-task).
2026-04-22 09:29:16 +02:00
jgrusewski
1ce99efc53 plan(policy-quality): Phase 2 implementation — synthesis of 4 tracks
Consolidates Phase 1 triage findings from Tracks 1-4 into an 11-task
Phase 2 plan at docs/superpowers/plans/2026-04-21-policy-quality-phase2.md.

Task inventory:
  - 2.0  Per-component gradient decomposition (H4 keystone diagnostic)
  - 2.1  H4 fix — magnitude-head gradient starvation (decision tree from 2.0)
  - 2.2  H10 fix — stable argmax tie-break at eval
  - 2.3  DELETE R5 micro-reward
  - 2.4  DELETE R6 loss-aversion; relocate neg-tail to C51 target smoothing
  - 2.5  Wiring-bug sweep (7 bugs: C1 fire, epsilon gen_range, if !true,
         sigma_mean scale, fold-boundary reset, stale docstring, C5 ISV null)
  - 2.6  E4 entropy-reg DELETE-or-KEEP (data-driven, post-2.0)
  - 2.7  C4 adaptive grad-clip ablation + DELETE-or-KEEP
  - 2.8  L40S validation run — all 4 tracks re-measured
  - 2.9  Mandatory-gate verification + phase3-results.md
  - 2.10 Tag policy-quality-phase2-complete (and policy-quality-v1 if soft pass)

Matches Phase 0/1 plan formatting (checkbox steps, concrete file paths,
code snippets, bash commands, per-task commit templates). References
project standing rules (no quickfixes, no stubs, no atomic-adds on hot
paths, no feature flags, no hiding errors) and the pinned-readback
pattern from Task 0.4 (commit bb42c9963).
2026-04-22 09:23:52 +02:00
jgrusewski
944fecd913 docs(policy-quality): main-branch workflow — cut feat/wip branches
User preference: everything lands on main, ~500 LOC is small enough
that feature-branch isolation isn't needed.

Spec changes:
  * §2.3 outcome paths: drop "unmerged branch" vocabulary
  * §3 architecture: main-branch timeline, no feat/, no wip/
  * §4 Phase 0: commits on main, tag policy-quality-baseline at end
  * §5 Phase 1: experiments are worktree edits reverted after, only
    triage docs commit (to main)
  * §6 Phase 2: commits land on main, author chooses granularity,
    each leaves smokes green
  * §7.3 outcome handling: iterate on main, baseline tag for rollback
  * §7.4 closing: tag policy-quality-v1, update status

Plan changes:
  * Task 0.1: verify-starting-state (not branch creation)
  * Task 1.x: temp-edit-then-revert (not wip branches)
  * Task 1.5: triage docs commit directly to main
  * Phase 4: tag + close (not merge + cleanup)
  * Rollback: one git reset --hard command
2026-04-21 21:06:33 +02:00
jgrusewski
1456094a5a plan(policy-quality): implementation plan for 4-track V7 audit
Produced by superpowers:writing-plans from the approved spec at
docs/superpowers/specs/2026-04-21-policy-quality-design.md.

Plan structure:
  * Phase 0 (tasks 0.1-0.17): measurement substrate — TDD-detailed,
    exact code + commands. Produces 6 new smoke tests + ~25 HEALTH_DIAG
    fields across 4 tracks.
  * Phase 1 (tasks 1.0-1.5): parallel investigation tracks — process
    tasks (run instrumentation, record evidence, classify hypotheses).
    No code to pre-write; output is 4 triage documents + a Phase-2
    planning doc.
  * Phase 2 (template only): convergence fix. Concrete tasks produced
    by a Phase-2-specific plan AFTER Phase 1 triage — can't pre-write
    the fixes without knowing which hypotheses confirm.
  * Phase 3 (tasks 3.1-3.4): L40S validation gates, surrogate-noise,
    branch decision (merge or iterate).
  * Phase 4 (tasks 4.1-4.4): merge via --no-ff, tag, cleanup.
  * Rollback procedure: git reset --hard policy-quality-baseline.
2026-04-21 20:56:45 +02:00
jgrusewski
ab7567875a plan: adaptive learning dynamics — 19 tasks across 5 phases
Implementation plan for unified LearningHealth system (all gems/pearls/novels):
- Phase A (A1-A4): ISV extension, LearningHealth module, training loop wiring,
  HEALTH_DIAG logging
- Phase B (B1-B4): G2 uncertainty-gated CQL, G3 health-coupled tau,
  G4 temp-continuous Expected SARSA, G5 adaptive gradient budget
- Phase C (C1-C4): P1 health-weighted PER priorities, P2/P3 subsumed by A3,
  P4 temporal-coupled gamma
- Phase D (D1-D8): N1 temporal self-distillation, N2 Q-gap barrier,
  N3 health-triggered plasticity, N4 CF curriculum, N5 information bottleneck,
  N6 ensemble oracle, N7 contrarian override, N8 meta-Q collapse predictor
- Phase E (E1-E2): collapse-recovery smoke test + L40S deployment verification

Spec: docs/superpowers/specs/2026-04-20-adaptive-learning-dynamics-design.md
Target: WinRate >55% by epoch 20, Q-gap stays above 0.1 (no collapse).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-20 19:12:41 +02:00
jgrusewski
bc8fde9a2b plan: unified state layout implementation — 8 tasks
Task 1: Rust constants (ml-core/state_layout.rs)
Task 2: CUDA header (state_layout.cuh + assemble_state())
Task 3: Refactor experience_state_gather to use shared assembly
Task 4: Delete backtest_gather_kernel, replace with shared function
Task 5: Remove configurable state_dim everywhere (131 call sites)
Task 6: Update OFI_DIAG indices + checkpoint validation
Task 7: Layout match integration test
Task 8: Deploy and verify on H100

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-20 14:36:03 +02:00
jgrusewski
063fd27166 feat: target_dim 4→6 + spec v5 with pearls (bar duration, book CoM, retrospective hold)
target_dim expansion: adds raw_open (OHLCV) and mid_price_open
(MBP-10 midpoint at bar formation) to fxcache targets. FXCACHE_VERSION
2→3 for auto-rebuild. Legacy v2 files handled with close-price fallback.

Spec v5 adds 3 pearls:
- Bar duration encoding in Mamba2 (continuous-time SSM awareness)
- Order book center of mass from all 10 MBP-10 levels (aggression signal)
- Retrospective hold quality bonus (teaches exit timing)

Plus: Hold action (4th direction), DSR Sharpe EMA fix, counterfactual
magnitude/order sign fix, MFT mid-price mark-to-market.

OFI embed MLP now 18→10 (was 16→8). Mamba2 width SH2+10 (was SH2+8).
Attention width D+10 (was D+8).

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-19 23:47:04 +02:00
jgrusewski
f7cf02f363 plan: OFI momentum + dense micro-reward — 9 tasks, full gradient flow
9-task implementation plan covering:
- fxcache target_dim 4→6 (raw_open + mid_price_open from MBP-10)
- OFI data pipeline fix (ofi_gpu wiring, indices 42→66, PORTFOLIO_STRIDE 30→38)
- Dense micro-reward (OFI momentum × price confirm × adaptive cost tolerance)
- OFI embed MLP (16→8 cuBLAS) with full backward gradient flow
- Mamba2 history enrichment (SH2→SH2+8) + d_h_history backward
- Attention enrichment (D→D+8) + d_input_scratch exposure

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-19 23:21:24 +02:00
jgrusewski
d751d2ad4f plan: kernel optimization — 8 tasks, 7 kernels → cuBLAS + batch-parallel
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-19 12:15:04 +02:00
jgrusewski
4f7b1f4804 plan: true single-graph — 7 tasks, zero CPU on hot path
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-19 00:06:05 +02:00
jgrusewski
c054b4e1d8 plan: unified single-graph — 5 tasks, zero ungraphed kernel launches
Task 1: CUfunction audit (map every kernel to its child graph)
Task 2: Expose submit methods (post_aux + maintenance + prev_grad_buf)
Task 3: Add new child graphs (capture + launch 7 children)
Task 4: Verify remaining outside-graph code is kernel-free
Task 5: Deploy and validate (target: adam ~30ms, epoch ~64s)

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-18 21:26:40 +02:00
jgrusewski
f851119e2d plan: Phase 2 multi-stream parallelism — 7 tasks
Task 1: Allocate aux streams + workspaces + events
Task 2: IQN/Attention accept workspace+stream override
Task 3: Fork-join in submit_aux_ops (IQN+Attention parallel)
Task 4: Switch capture mode GLOBAL → RELAXED
Task 5: Forward backward branch parallelism
Task 6: Fix adam child event timing (verify)
Task 7: Deploy + verify child graph breakdown

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-18 16:07:06 +02:00
jgrusewski
45d8f4661c plan: rewrite Task 9 for child graph architecture
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-17 23:58:32 +02:00
jgrusewski
750b228ed9 plan: Trade Plan Head — 5 tasks, hierarchical plan-based trading
Task 1: PORTFOLIO_STRIDE 23→30, plan slots ps[23-29]
Task 2: trade_plan_forward kernel + 4 weight tensors (82→86)
Task 3: Plan activation + enforcement in env_step
Task 4: Direction lock in action_select + counter-plan Q (N14)
Task 5: Smoke test + compute-sanitizer

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-17 00:48:18 +02:00
jgrusewski
ddb7b170bc plan: Tick Microstructure Intelligence — 4 tasks, OFI_DIM 8→20
Task 1: IncrementalMicrostructureCalculator (12 new features, O(1)/tick)
Task 2: fxcache OFI_DIM 8→20 atomic update (15 files, 30+ locations)
Task 3: Extend precompute pipeline with MicrostructureState
Task 4: Regenerate fxcache + smoke test + compute-sanitizer

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-17 00:02:02 +02:00
jgrusewski
6dd2f3e68b plan: Introspective State Vector — 9 tasks, pure GPU implementation
Task 1: ISV scratch scalars (piggyback on existing kernels)
Task 2: isv_signal_update GPU kernel + temporal history
Task 3: ISV encoder MLP + weight tensors 68→78
Task 4: Drift-conditioned gamma (C51 scalar→per-sample buffer)
Task 5: Branch confidence routing (replaces regime_branch_gate)
Task 6: Risk branch wider input SH2→SH2+9
Task 7: Recursive confidence head (predict own TD-error)
Task 8: Target network ISV online-only
Task 9: Smoke test + compute-sanitizer verification (0 errors)

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-16 22:46:54 +02:00
jgrusewski
685746231b plan: Training Environment Alignment — 5 tasks, 3 phases
Phase 1: Position-gated episodes (h100.toml + done flag + soft reset)
Phase 2: Adapt rank norm (skip zeros) + remove hardcoded Q-drift
Phase 3: Aligned Sharpe metrics (un-annualized per-trade)
Phase 4: Build verification + smoke test

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-16 21:49:19 +02:00
jgrusewski
8d413f7759 plan: Self-Improving Training Loop — 6 tasks, 9 enrichments
Enrichment module (enrichment.rs) + training loop wiring:
E1-E3: Q-correction, adaptive epsilon, dynamic gamma
E4-E5: Trade autopsy per-branch LR, ensemble agreement
E6-E8: Winner distillation, hindsight labels, curriculum weights
E9: State confidence (deferred — needs K-means kernel)

~215 lines Rust, no new CUDA kernels. 6 tasks.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-16 17:34:14 +02:00
jgrusewski
5c4f153f26 plan: Learned Risk Management — 6 tasks, 5th branch risk_budget [0,1]
3 CUDA kernels (forward, apply, backward), NUM_WEIGHT_TENSORS 64→68,
per-sample CVaR alpha + commitment lambda, separate Adam, ~33K params.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-16 01:32:16 +02:00
jgrusewski
5b3183504f plan: OOS Generalization Enhancement — 9 tasks, 10 components
3-layer defense: compress (AdamW+L1, gamma anneal, regime dropout),
align (cost curriculum, walk-forward), exploit (epistemic gate, branch
independence, counterfactual, temporal consistency).
~480 lines total. Companion to OOS Performance Enhancement plan.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-15 20:33:03 +02:00
jgrusewski
57a17c6d81 plan: add Tasks 9d-9e — Q-mean drift regularization + new component LR warmup
Task 9d: Q-mean EMA drift penalty (lambda=0.01) in C51 grad kernel.
Prevents +0.47 Q-mean drift observed in train-vpb4w.

Task 9e: 500-step LR warmup for new tensor groups (regime gate,
adaptive atoms, Mamba2, multi-horizon). Prevents gradient shock
on new components introduced to trained network.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-15 19:43:50 +02:00
jgrusewski
31ad6d5dec plan: add Tasks 9a-9c — missing backward/training for atoms, Mamba2, multi-horizon
Task 9a: Adaptive atom position training via atom entropy gradient.
Numerical finite-difference gradient on spacing_raw, SGD at LR=1e-3.

Task 9b: Mamba2 BPTT backward through K=8 scan steps.
d_W_A/B/C accumulated, separate Adam at LR=1e-4.

Task 9c: Multi-horizon C51 loss — 3 n-step passes (1/5/20 bar),
3 C51 loss launches, regime-weighted gradient blend (ADX-based).

Fixes the "untrained component" gap identified in plan review.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-15 19:22:27 +02:00
jgrusewski
f2ddba5189 plan: OOS Performance Enhancement — 9 tasks, 8 components
Task 1: 3 CUDA kernels (regime gate, adaptive atoms, mamba2 scan)
Task 2: Adaptive DSR (training Sharpe EMA → w_dsr scaling)
Task 3: CVaR objective (risk-sensitive Bellman target)
Task 4: Regime branch gating (20 params, ADX/CUSUM → branch importance)
Task 5: Adaptive atom positions (204 params, learned C51 spacing)
Task 6: Ensemble heads (config 1→3, existing infrastructure)
Task 7: Mamba2 temporal scan (12K params, 8-bar context)
Task 8: Multi-horizon prediction (144K params, 1/5/20 bar horizons)
Task 9: Build verification

NUM_WEIGHT_TENSORS: 50→64 across tasks 4,5,8.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-15 18:56:34 +02:00
jgrusewski
f14440be55 plan: add Tasks 10-11 — Q-mean centering + atom entropy recovery (execute first)
Task 10: q_mean_reduce + q_mean_subtract two-phase kernel. Zero params.
Fixes Q-mean drift 0→+0.47 from bootstrapping bias.

Task 11: Adaptive entropy_coeff based on utilization_ema. Zero params.
Fixes atom utilization collapse 100%→20%.

Execution order: 10, 11, then 1-9.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-15 08:40:09 +02:00
jgrusewski
fcf34ac5c8 plan: Supervised Architecture Transfer — 9 tasks, 7 components
Task 1: 7 CUDA kernels (quantile, graph, KAN, diffusion, OFI concat)
Task 2: Quantile Q-select (replaces epsilon-greedy, zero params)
Task 3: Graph message passing (4 directed edges, 60 params)
Task 4: KAN spline gates (replaces sigmoid, 9,216 params)
Task 5: Diffusion Q-refinement (3-step denoise, 2,772 params)
Task 6: TLOB microstructure injection (widen order/urgency FC)
Task 7: xLSTM temporal context (CPU matrix memory, 528 params)
Task 8: RK4 adaptive ODE (replaces Euler in liquid tau)
Task 9: Build verification + smoke tests

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-15 08:29:51 +02:00
jgrusewski
cc232f644d plan: Cross-Pollinated Branch Intelligence — 9 tasks, 64 steps
Task 1: Dead code cleanup (11 items across 6 files)
Task 2: 6 CUDA kernels (VSN, GLU, Q-attention, selectivity)
Task 3: Weight layout 26→42 + separate buffers + liquid_mod pinned
Task 4: VSN + GLU forward integration
Task 5: VSN + GLU backward integration
Task 6: Cross-branch Q attention (624 params, 2-head on 12 Q-values)
Task 7: Liquid tau per-branch dynamics (replaces spread_velocity)
Task 8: Mamba-2 selective replay (learned PER curriculum)
Task 9: Build verification + smoke tests

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-14 23:26:13 +02:00
jgrusewski
f49ae6858e fix: Layer 3 plan — add Q_dir normalization by delta_z
Raw E[Q] ~ 0.1 while h_s2 ~ 1.0 post-ReLU. Without normalization,
Q_dir columns learn 10x slower. Dividing by delta_z scales to ~1-10.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-14 22:30:29 +02:00
jgrusewski
8d2ed799e6 plan: direction-conditioned magnitude head (Layer 3) — 7 tasks
Tasks: CUDA kernel, param widening, forward conditioning, backward split,
zero-init, experience collector, H100 integration test.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-14 22:19:41 +02:00
jgrusewski
e9e4bd54ac plan: Adaptive Training Dynamics Layers 1+2 implementation plan
Task 1: Reward rank normalization kernel (SNR amplification)
Task 2: Q-gap momentum for spread gradient (plateau modulation)
Task 3: H100 integration test (50 epochs, compare to train-w6qfd)

Layers 3+4 (dir→mag conditioning, variance sizing) deferred to
separate plan after L1+L2 validation.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-14 21:42:57 +02:00
jgrusewski
f0cda76229 docs: IQL full integration implementation plan — 9 tasks
6 new CUDA kernels, C51/grad/epsilon kernel mods, dual-tau IQL,
v_range dead code deletion, per-sample everything.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-13 14:56:27 +02:00
jgrusewski
6621177f78 docs: implementation plan for single shared cuBLAS handle
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-13 11:32:23 +02:00
jgrusewski
1f6a8436d3 feat: v9 Differential Sharpe Ratio reward — optimize for consistency, not PnL
Fundamental shift: reward optimizes Sharpe ratio contribution per trade
instead of directional PnL magnitude. Targets Sharpe 5+ via many small
consistent gains (spread capture, execution quality) instead of few
large directional bets.

DSR implementation (Moody & Saffell, 2001):
- Per-episode EMA statistics in portfolio state slots [3]-[5]
- DSR = (B*R - 0.5*A*R²) / (B - A²)^1.5 at trade completion
- 10-trade warmup, clamped [-5, +5]
- w_dsr=5.0 (primary reward signal)

Reward hierarchy restructured:
- DSR: 5.0 (NEW — primary, rewards Sharpe consistency)
- Directional PnL: 2.0 (was 10.0 — demoted to secondary)
- Order credit: 1.0 (was 0.1 — spread capture amplified 10×)
- Urgency credit: 0.5 (was 0.1 — fill quality amplified 5×)
- Inventory penalty: -0.005×|pos|/max_pos per bar (NEW)
- Dense micro-reward: 0.0 (removed — noisy direction signal)

Expected: WinRate increases (many small spread captures), per-trade
variance decreases, Sharpe rises from consistency not prediction.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-12 14:33:56 +02:00
jgrusewski
eb7139c436 feat: adaptive C51 v_range — Q-values 0.008→0.525 (65× improvement)
Fundamental fix: C51 atom spacing (dz=0.6) was larger than the Q-value
range (0.02), making the distributional loss unable to resolve rewards.
The network was blind to its own learning signal.

Adaptive v_range via device buffer (same pattern as tau_buf):
- v_range_buf[2] on GPU, all C51/MSE/CQL/expected_q kernels read via
  pointer (graph captures stable address, reads value at replay time)
- Cold start: v_range=±1.0 → dz=2/51=0.039 → 25× atom resolution
- Monotonic expansion based on Q-stats every 50 training steps
- Never shrinks (avoids invalidating Bellman targets in replay buffer)

Kernel changes (5 files):
- c51_loss_kernel.cu, mse_loss_kernel.cu, mse_grad_kernel.cu,
  cql_grad_kernel.cu, compute_expected_q: scalar v_min/v_max → pointer

Also fixed:
- PopArt warmup 100→1 (start normalizing from step 1)
- PopArt variance floor 0.0001→1e-8 (allow sigma < 0.01)
- Backtest evaluator: allocate own v_range_buf for compute_expected_q
- num_atoms 51→52 in all TOMLs (TF32 4-element alignment)

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-12 07:58:02 +02:00
jgrusewski
ad5c2b7ded refactor: exp collector zero-copy — direct pointer into trainer params_buf
Dead code sweep (167 lines deleted):
- online_params_flat (synced but never read by forward)
- online_params_f32 (never synced, always zeros — ROOT CAUSE of garbage Q-values)
- DuelingWeightSet/BranchingWeightSet/CuriosityWeightSet/RmsNormWeightSet fields
- sync_weights_flat, sync_weights_f32, flatten_online_weights methods
- sync_gpu_weights epoch-boundary DtoD copy

Zero-copy replacement:
- trainer_params_ptr: u64 points directly into trainer's params_buf
- param_sizes computed from real config (bottleneck_dim=16, market_dim=42)
- CublasForward with correct s1_input_dim=54 for bottleneck
- Bottleneck forward activated (bn_hidden + tanh + concat)
- Fused context init moved before Phase 2 (experience collection)

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-11 23:10:25 +02:00
jgrusewski
dd6143936e fix: IQN 4-branch runtime + order counterfactual + per-branch entropy + histogram
IQN kernel rewrite (critical — was corrupting 75% of gradient budget):
- Eliminate compile-time BRANCH_*_SIZE defines (BRANCH_0_SIZE was 5, should be 3)
- All 9 IQN kernels now take runtime b0/b1/b2/b3 params
- 4-branch forward, backward, decode (was 3-branch, missing magnitude)
- branch_actions buffer B*3 → B*4

Action diversity infrastructure:
- 3-cycle counterfactual: dir mirror / mag relabel / order-type cost delta (50× amplified)
- Per-branch entropy: order (d==2) gets 3× boost in C51 + MSE grad kernels
- Position histogram expanded 6→12 bins (all 4 branches get entropy bonus)
- Boltzmann temperature floor raised to 0.5 for order + urgency branches
- Smoke test epochs 3→10 for diversity verification

Note: experience collector still needs bottleneck forward path — currently reads
trainer's bottleneck-layout weights with state_dim K, producing misaligned Q-values.
Next: eliminate weight copy, use direct pointer views into trainer's params_buf.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-11 22:40:47 +02:00
jgrusewski
a662e54e86 plan: VarStore elimination — zero-copy pointer views into flat params_buf (8 tasks)
DuelingWeightSet becomes raw u64 pointers into params_buf, not separate
CudaSlice allocations. Eliminates ALL D2D copies between weight
representations. flatten/unflatten cycle deleted entirely.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-11 13:35:15 +02:00
jgrusewski
370319f7c8 plan: ABI Complete — per-branch epsilon, IQN anti-collapse, hindsight relabeling, vaccine, ensemble, bottleneck (8 tasks)
Unified plan combining Sub-plans B+C + novel findings:
1. Fix c51_alpha ramp (eliminate 3 graph recaptures)
2. Per-branch epsilon (magnitude 2×, Boltzmann for order/urgency)
3. IQN quantile spread as anti-collapse bonus
4. Hindsight magnitude relabeling
5. Gradient vaccine with diverse batches
6. Ensemble variance at inference
7. Per-branch bottleneck width
8. Full verification + deploy

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-11 12:27:56 +02:00
jgrusewski
55be1776b5 plan: ABI Foundation — 7-level exposure + PopArt + tau coupling + PER clear (6 tasks, 27 steps)
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-11 11:28:59 +02:00
jgrusewski
f6991dd8b2 plan: dead code cleanup + GPU per-phase CUDA event profiling (6 tasks, 32 steps) 2026-04-10 22:23:16 +02:00
jgrusewski
445c2197f2 feat: unified Argo train workflow — replace 7 templates with 1
New: train-template.yaml with smart caching:
  - ensure-binary: checks /data/bin/{sha}/ on PVC, compiles only on cache miss
  - ensure-fxcache: runs precompute only if cache doesn't exist
  - gpu-warmup: parallel autoscale during compile
  - hyperopt → train-best → evaluate → upload-results

Deleted 7 redundant templates:
  - compile-and-train-template.yaml
  - train-dqn-template.yaml
  - train-baseline-rl-template.yaml
  - train-supervised-template.yaml
  - training-workflow-template.yaml
  - precompute-features-template.yaml
  - train-ppo-template.yaml

Deleted: scripts/argo-precompute.sh (absorbed into ensure-fxcache step)
Rewritten: scripts/argo-train.sh (single template, --sha for commit pinning)
Updated: kustomization.yaml

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-10 20:36:03 +02:00
jgrusewski
4842a04011 refactor: eliminate ALL bf16 from entire workspace — pure f32/TF32 pipeline
Complete bf16 elimination across all crates (ml, ml-core, ml-dqn, ml-ppo,
ml-supervised). Zero half::bf16, __nv_bfloat16, or CudaSlice<half::bf16>
references remain (verified by grep).

CUDA: All 60+ .cu kernels and .cuh headers converted to native float.
  - Half-precision intrinsics (__hmul, __hadd, __hdiv) → float operators
  - atomicAddBF16 → native atomicAdd(float)
  - bf16 wrapper functions → f32 identity passthroughs

Rust: All CudaSlice<half::bf16> → CudaSlice<f32> across 90+ files.
  - htod_f32_to_bf16/dtoh_bf16_to_f32 → htod_f32/dtoh_f32 (direct, no conversion)
  - Deleted bf16 mirror infrastructure (DuelingWeightSetBf16, alloc_bf16_mirror, etc.)
  - Renamed params_bf16→params_flat, d_value_logits_bf16→d_value_logits, etc.
  - Fixed .to_f32() sed damage on Decimal::to_f32() and rng.f32()

FxCache: Single f32 disk format (was bf16/f64 dual-version).
  - Deleted --bf16 CLI flag from precompute_features
  - PVC cache files need regeneration via precompute_features

TF32 tensor cores activated via cublasLtMatmul CUBLAS_COMPUTE_32F — no
explicit TF32 types needed. Storage is pure f32 everywhere.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-10 18:52:21 +02:00