1c63c32297b5aa66cbc72baf75e39d2bec0feefd
614 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
1c63c32297 |
spec(dqn): reframe Phase 3 as direction-branch dead-code cleanup
User correctly challenged the "band-aid removal" framing. All fixes shipped during the val-Flat-collapse investigation addressed real bugs at their respective layers and should be preserved: - Kelly cap warm-branch ( |
||
|
|
5c325f8947 |
spec(dqn): pivot to Thompson sampling — distributional RL action selection
Revises the C51-bias spec after deeper review surfaced 12 design gaps,
of which 4 were critical:
1. Train-only vs train+eval ambiguity — UCB at eval would conflate
"model recommends Long" with "model is uncertain about Long",
inflating reported edge. CRITICAL for trading where eval drives
real capital decisions.
2. Thompson sampling is more principled than UCB:
- parameter-free (no κ to tune)
- uses distribution directly without scalar reduction
- naturally explore-exploit balanced via distribution shape
3. c51_alpha is the wrong blend weight (it's C51-vs-MSE-warmup, not
C51-vs-IQN). Equal-weight average of C51 and IQN samples is the
structural choice — no tuned blend weight needed.
4. The bias might be CORRECT BEHAVIOUR — model rationally choosing
Flat when no edge has been discovered. Phase 0 must include a
synthetic-edge test (controlled MDP with KNOWN positive Long
expected value) to verify Thompson can discover edge if it exists.
Other gaps fixed:
- Eval at argmax E[Q] (not Boltzmann, not Thompson)
- Pearl wording broadened to cover ensembles + future methods
- Ensemble Q-head added to aggregation contract table
- Explicit caveat: NEVER extend Thompson to magnitude branch (would
worsen existing magnitude saturation)
- Phase 0.F uses CONVERGED checkpoint (≥30 epochs), not 2-epoch run
- L4 long smoke (30 epochs, ~1 hour) added — Thompson edge discovery
needs longer feedback loop than 5 epochs
- Phase 3 explicitly removes eps-floor + tau-floor band-aids
(Thompson replaces direction Boltzmann; band-aids become dead code)
- Conviction stays E[Q]-based, not sample-based (avoid Kelly cap
jitter from stochastic samples)
Architecture (Thompson only, no UCB):
TRAINING: dir_idx = argmax(0.5 × (sample_C51(d) + sample_IQN(d)))
magnitude/order/urgency: existing Boltzmann + ε-greedy
EVAL: dir_idx = argmax(0.5 × (E[Q_C51] + E[Q_IQN]))
magnitude/order/urgency: existing Boltzmann (eval mode)
Direction-branch ε-greedy + Boltzmann are REMOVED — Thompson is the
exploration mechanism. No new GPU buffers; existing C51 atoms + IQN
quantiles passed to action_select.
5-7 days active work across 4 sub-plans; each gets its own
writing-plans cycle.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
494125c104 |
spec(dqn): distributional RL aggregation — UCB on C51+IQN σ for action selection
Designs the structural fix for the C51 expected-Q Hold/Flat bias exposed
by the Kelly val-Flat-collapse fix. The bias is a manifestation of a
deeper project-wide pearl:
"Distributional RL aggregation discards uncertainty;
action selection must restore it."
Any value head representing Q as a distribution (atoms, quantiles,
ensembles) MUST expose both E[Q] and σ(Q) to action selection. Boltzmann
on E[Q] alone produces structural bias toward low-variance actions
regardless of expected payoff — the C51+IQN Flat-attractor is one
instance of this lost-information pattern.
Fix: extract σ(Q) from BOTH C51 atoms (closed form) and IQN quantiles
(IQR/1.349), blend by loss-time weight, feed Q_eff = E[Q] + κ·σ
(κ=1.0 structural identity) to direction-branch Boltzmann ONLY.
4-phase implementation:
Phase 0 — TDD hypothesis verification (Rust mirror functions + 5 unit
tests including GPU integration on real checkpoint)
Phase 1 — Audit existing reward levers (B.2, CF, PopArt, Q-target)
via 6 unit tests; fix any bugs found
Phase 2 — UCB integration: new compute_q_with_uncertainty kernel,
modified action_select, Rust orchestration, project-wide
aggregation contract in dqn-wire-up-audit.md
Phase 3 — Verification per Plan 5 Task 5 multi-seed × multi-fold
5-layer verification gate; 8 risks with mitigations; 4 stop conditions
that halt execution and force redesign.
Existing band-aid fixes (Kelly cap, eps-floor, tau-floor) stay — they
address symptoms at different layers. UCB adds the missing aggregation
step that was the common root across all the symptoms.
Direction-branch only — magnitude/order/urgency don't have the
Flat-attractor (atom-mass collapse asymmetry).
5-7 days active work across 4 sub-plans; each gets its own
writing-plans cycle.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
d54b49efc1 |
fix(dqn): C51 bias breakout — adaptive eps_dir floor from trade-rate undershoot
Prior C51 bias fix (commit
|
||
|
|
7a3d886462 |
fix(dqn): three concerns from kelly-fix-multifold val run
CONCERN 1 — train Return display (financials.rs)
Replace `mean_per_bar × bars_per_year` ("annualized arithmetic return")
with the actual cumulative compounded return: `prod(1+r) − 1` via
log-space sum for f64 stability. Step floor at 1+r >= 1e-10 so a -100%
bar gives -23 log contribution rather than -inf; result bounded above
-100%.
The old formula multiplied per-bar mean by ~98K, producing values like
Return=-51234% that LOOKED like portfolio collapse but were actually
just label artefact: -52 bps mean × 98,280 bars/year = -51,234%, while
the actual compounded return over the rollout was bounded and finite.
"Total Return" now means what a trader expects.
CONCERN 2 — val_Sharpe annualization source-of-truth (training_loop.rs +
gpu_backtest_evaluator.rs)
val_Sharpe was using `m.sharpe` from the kernel (annualised by
`sqrt(bars_per_day × trading_days_per_year)` at evaluator init), but
val_Sharpe_raw recomputed the divisor from `self.hyperparams.bars_per_day
* 252.0` in Rust. Empirically the two diverged by 5.4× (val_Sharpe=143.38
/ val_Sharpe_raw=0.0841 = 1704, expected sqrt(98280) = 313.5). Root cause
unclear without runtime instrumentation but the structural fix is to
read the EXACT factor the kernel applied: new
`GpuBacktestEvaluator::annualization_factor()` returns the f32 stored at
init time. Rust now divides by that, so the two numbers can never drift
regardless of how `bars_per_day` flows through the config layers.
Per `feedback_no_partial_refactor.md` — every consumer of a shared
contract must use the same source.
CONCERN 3 — C51 expected-Q Hold/Flat bias (experience_kernels.cu)
Direction-branch Boltzmann tau floor: replace static `0.01` with
ISV-adaptive `max(ISV[21], 0.01)` where ISV[21] is q_dir_abs_ref (EMA
of mean |Q| across direction bins, same signal that drives conviction).
The C51 expected-Q biases Flat above directional actions when the policy
has no measurable edge — Flat returns are exactly 0 (no position change),
collapsing C51's distribution to a delta at 0; directional returns
concentrate slightly below 0 from tx_cost without compensating edge,
giving E[Q_directional] < E[Q_flat] = 0. With Q-spread on the order
of 0.01-0.1 and the static 0.01 tau floor, exp(Q/tau) ratios approached
deterministic argmax, locking the policy into Hold/Flat within ~1 epoch
and preventing the exploration needed to discover real edge.
Validated empirically by the kelly-fix-multifold run (train-multi-seed-bs9m5):
val_dir_dist epoch 1 [S=.23 H=.17 L=.42 F=.18] (active 65%)
val_dir_dist epoch 2 [S=.14 H=.35 L=.15 F=.37] (active 29%)
That 35→72% Hold+Flat shift in one epoch is the C51 attractor in action.
The ISV-adaptive floor preserves relative spread for sampling: when
Q-magnitudes grow during training, the floor scales with them, so the
policy never collapses to deterministic argmax until q_range exceeds
the network's typical Q magnitude — i.e., until spread represents
real, scale-significant edge. Coherent with the existing conviction
formula (same ISV reference). No tuned constants per
`feedback_isv_for_adaptive_bounds.md` and `feedback_adaptive_not_tuned.md`.
Cold-start fallback retains 0.01 minimum so kernel doesn't divide by
zero pre-first-update. Once ISV producer fires (every epoch), adaptive
floor takes over.
7 financials unit tests pass. Workspace cargo check clean.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
2f4bd8e58b |
cleanup(dqn): remove val-collapse diagnostic printf — fix verified
Diagnostic kernel printf in backtest_env_step_batch identified the gate
(Kelly cap warm-branch deadlock) and verified the fix
(
|
||
|
|
0c9d1ee39e |
fix(dqn): Kelly cap warm-branch deadlock — effective_kelly never collapses to zero
VAL DIAGNOSTIC PROOF (train-4fpzx step=400):
pick=Long Full (target=0.66) prev_pos=0.0 → pos_post=0.0 actual_dir=Flat
trail=0, margin can't clip to 0 → only Kelly cap zeroed the target.
ROOT CAUSE:
effective_kelly = maturity * kelly_f + (1 - maturity) * warmup_floor
Cold-start fix (commit
|
||
|
|
17e6c7a78f |
diag(dqn): widen VALDIAG to sample full 214K val window
First 10 bars proved the env_step kernel produces actual_dir=Long for Long picks correctly at cold start. So the gate that turns 100% of Long/Short picks into actual_dir=Flat in val_dir_dist must activate LATER in the window — likely past the first chunk boundary. Extended sampling: first 10 bars, every 100 bars to 1000, every 1000 bars to 10000, every 10000 bars throughout. Also adds a "mismatch" trigger that fires when picked dir is Long/Short but actual_dir is not — every 137 bars to keep printf volume bounded. cash field added to spot capital drain or anomalous accumulation. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
7d29c5ee48 |
diag(dqn): kernel printf for first 10 bars of val window 0
Captures action_val + decoded dir/mag, max_position, prev_position, position post-trade, actual_dir, actual_mag, trail_triggered, conviction, health, value, max_equity inside backtest_env_step_batch. Localises which step in unified_env_step_core zeros target_position (or position) for non-Hold picks. After tracing the math, every analytical candidate (margin cap, Kelly cap with health-coupled warmup floor, trail stop) returns a positive target at cold start with my fix. val_dir_dist still locks at 0% Long while val_picked_dir_dist shows 17% Long — only direct kernel-side observation can show what's actually clamping. Removed once the gate is identified. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
e15adecef1 |
diag(dqn): persistent picked_action_history — full-window kernel pick histogram
The `val_picked_dir_dist` reader added in
|
||
|
|
ee01509e9a |
perf(dqn): Phase H Site 2 — fuse cublasLt DRELU_BGRAD into value-FC backward chain
Picks up the deferred half of P5T5 Phase H. Site 1 (commit |
||
|
|
89ece2e366 |
diag(dqn): val_picked_dir_dist — read raw Boltzmann picks pre-env_step
Adds a paired diagnostic to localise the eval-mode collapse mechanism.
The existing `val_dir_dist` line reads `actions_history_buf` which env_step
overwrites with the POST-physics `actual_dir` (signed position bucket). When
the policy picks Long but the Kelly cap zeroes target_position, the bar shows
up as `actual_dir=Flat` — indistinguishable in the existing diagnostic from a
case where the kernel itself produced a Flat pick.
Cluster runs `ddrpr` (without Kelly fix) and `txdz9` (with Kelly fix
|
||
|
|
2c97e0436c |
fix(dqn): unstick eval Kelly cap — health-coupled warmup_floor never collapses to zero
The eval-mode policy was producing diverse Boltzmann picks (verified by the
new `val_dir_dist` HEALTH_DIAG line) but every active-direction pick (Long /
Short) was collapsing to `actual_dir = Flat` because the Kelly cap forced
`target_position = 0` at cold start.
Cluster run `train-multi-seed-ddrpr` epoch 0 made this unambiguous:
val_dir_dist [short=0.0000 hold=0.1953 long=0.0001 flat=0.8047]
Boltzmann fired correctly (sum hold + flat ≈ 100% of bars, with Hold ~ 20% =
the share of bars where the policy explicitly picked Hold; the other 80%
were active-direction picks all rerouted to Flat by `target_position = 0`).
Root cause in `trade_physics.cuh::kelly_position_cap`:
warmup_floor = clamp(conviction, 0, 1) // ← can hit 0
effective_kelly = maturity*kelly_f + (1-maturity)*warmup_floor
= 0 + 1*0 = 0 at cold start with low conviction
cap = effective_kelly * max_position * safety = 0 → no exposure permitted
The `safety_multiplier` was already protected by a `health_safety = 0.5 + 0.5×h`
floor, but `warmup_floor` had no such floor. Catch-22: low conviction → cap=0 →
no trades → Kelly stats stay cold → conviction stays low → forever.
The same bootstrap-deadlock pattern as the IQN trunk SAXPY readiness gate
(commit
|
||
|
|
0a6a615d83 |
fix(dqn): unify eval action selection with training Boltzmann softmax
The eval policy used strict-argmax with an ISV-tied tie-break across all four
factored heads (direction, magnitude, order, urgency). The tie threshold was
`0.01 × isv_signals[V_HALF_*_INDEX]` — i.e. 1% of the C51 atom support range
(~40 for direction). That threshold did not match the actual per-sample
Q-spread (~1.5 once the IQN trunk gradient unstuck), so tie-break never fired
and eval became pure strict-argmax over a peaked Q distribution → val argmax
glued to one direction → 1-25 trades per 214k-bar window across cluster runs
`vg2r9` and `vnwtn`.
Replace with the same Boltzmann softmax training already uses: `tau =
max(q_range, floor)` where `q_range` is computed per-sample. Softmax is
mathematically bounded to `P(best) ≤ 47.5%` for 4 actions with `tau=q_range`,
so eval can never collapse to pure-greedy regardless of how peaked the
Q-values become. State-adaptive without tuned constants — confident states
(large q_range) still favour the best direction near-deterministically;
ambiguous states (q_range at floor) sample uniformly. The Philox stream is
seeded by (i, timestep) so eval remains bit-reproducible across runs at the
same checkpoint.
Three additions:
1. `experience_kernels.cu`: drop the four `else if (eval_mode)` strict-argmax
blocks; eval falls through to the existing Boltzmann path. Net -149 lines.
2. `cuda_pipeline/mod.rs`: add `test_eval_action_select_boltzmann_bounded`,
a focused unit test that exercises the kernel directly with peaked
synthetic Q-values and asserts the histogram matches Boltzmann theory
(P(best) ≈ 0.366, ≤ 0.6, ≥ 0.25). Runs in 1.65s after build, replaces
15-min smoke runs for kernel-level validation.
3. `trainers/dqn/trainer/metrics.rs`: log per-direction eval distribution
(`val_dir_dist [short hold long flat]`) to HEALTH_DIAG. The kernel-side
`dir_entropy` collapses Hold+Flat into one bucket, masking whether the
eval policy actually picks one direction or balances Hold/Flat.
Verified: unit test produces histogram short=0.146 hold=0.239 long=0.382
flat=0.233 — matches Boltzmann math, confirms the eval kernel produces
diverse picks for peaked Q-input.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
326c133782 |
perf(dqn): Phase H — fuse cublasLt BIAS epilogue into 4 attention forward projections
Collapses each `cublasLtMatmul + add_bias_f32_kernel` pair in the attention
forward path (Q, K, V, O projections) into a single fused `cublasLtMatmul`
with `CUBLASLT_EPILOGUE_BIAS`. Saves 4 kernel launches per attention forward
× per training step (target + online), targeting the L40S deploy hot-spot
where the per-epoch training phase is 99.7% of wall time.
Three code changes in `crates/ml/src/cuda_pipeline/gpu_attention.rs`:
1. `create_attn_gemm_desc` extended with `epilogue: Option<cublasLtEpilogue_t>`
parameter — when `Some(BIAS)`, descriptor is configured with
`EPILOGUE = BIAS` + `BIAS_DATA_TYPE = CUDA_R_32F`; when `None`, stays
at `EPILOGUE_DEFAULT` (the 8 backward dW/dX GEMMs unchanged).
2. New `lt_matmul_with_bias_ex` helper writes the per-call bias pointer via
`set_matmul_desc_attribute(BIAS_POINTER, …)` immediately before each
`cublasLtMatmul` (mirrors the existing pattern in batched_forward.rs).
The 4 forward projection sites in `forward(...)` switch from the prior
`lt_matmul_ex(...)` + `launch_bias_add_ex(...)` pair to a single
`lt_matmul_with_bias_ex(...)` call.
3. Orphans pruned: `launch_bias_add_ex` and `launch_bias_add` deleted from
`gpu_attention.rs` (their only callers were the 4 fused-away sites).
Shared `add_bias_f32_kernel` retained — still used by
`batched_forward.rs::launch_add_bias_f32_raw` (VSN Linear_2 logit output,
no activation).
Determinism preserved: the deterministic-algo cache (`cublas_algo_deterministic.rs`)
already keys on epilogue via `ShapeKey::with_epilogue`, so first-call selection
runs the full `AlgoGetIds → AlgoInit → AlgoCheck` loop with the new descriptor
and subsequent calls reuse the cached `(types, shape, epilogue)` algo. Bit-
deterministic when the algo is fixed under `CUBLAS_WORKSPACE_CONFIG=:4096:8`.
Site #2 (DRELU_BGRAD on the trunk Linear→Bias→ReLU backward) deferred — the
forward-side aux-buffer plumbing crosses three modules (BatchedForward →
BatchedBackward → value-FC site) and the determinism contract verified by
the Phase G smoke is non-trivial to preserve. Tracked for follow-up.
Verified: cargo check workspace clean at 11 warnings (baseline preserved).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
f86353840e |
fix(dqn): unstick IQN trunk gradient — drop iqn_readiness multiplier from SAXPY scale
`apply_iqn_trunk_gradient` and the parallel VSN-range SAXPY both scaled their contribution by `iqn_lambda × iqn_readiness × iqn_budget`. The readiness scalar initialises to 0.0 and only ramps up when `iqn_loss_ema` drops below `iqn_loss_initial` — but that improvement requires the trunk to learn IQN's gradient, which the readiness gate just blocked. Bootstrap deadlock: trunk_iqn=0.0000 across every observed L40S epoch, downstream strangling direction-Q discrimination → eval strict-argmax glues to one direction → 22-34 trades per 858k-bar window vs healthy 1257-trade burst at the one epoch where the gate momentarily lifted. iqn_budget already throttles the IQN contribution via the per-component budget controller (60% IQN, ISV-driven), so readiness was an additive band-aid that became load-bearing. New scale: `iqn_lambda × iqn_budget`. The `iqn_readiness` field stays on `self` because the C51 loss kernel launch site reuses `iqn_readiness_dev_ptr` as a CVaR-alpha pointer (gpu_dqn_trainer.rs:~16227) — that semantic overload is broken in a different way (CVaR α=0 is degenerate) and is tracked for follow-up. Verified on cluster trace `train-multi-seed-vg2r9` (epochs 0–13): trunk_iqn=0.0000 every epoch, q_gap_comp=0.00 every epoch, val trade_count locked at 22–34 except epoch 2 (1257 trades) where the gate accidentally cleared. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
8c20ad7958 |
plan5(task5-G): vol_normalizer numerical robustness + diagnostic logging
Phase D aux label-scale EMA only treated symptoms; root cause was epoch_vol_normalizer in training_loop.rs producing wildly different raw values across machines (local=0.541, cluster~1e-7), inflating return features [0..3] up to 50,000x their intended unit-variance scale. Fix: Welford's online variance (single-pass, numerically stable for any n) + sanity bounds [1e-5, 1e-1] (typical equity-index 1-min vol range) + explicit per-epoch tracing::info! log. Out-of-band raw values trigger tracing::warn! and fall back to 5e-4 default. multi_fold_convergence smoke (3 folds x 5 epochs, 682.85s): F0=2.3551 F1=80.8206 F2=92.1063 - all positive, F2 strongest yet. The warning surfaces deeper question (what's at targets[0] on this dataset) for future investigation; Phase G makes training scale-correct regardless. 9/9 monitoring tests + cargo check clean at 11 warnings. |
||
|
|
93126504ca |
plan5(task5-F): compile-time fxcache schema fingerprint via build.rs
Closes the L40S deploy-bug class where stale fxcache passed FXCACHE_VERSION validation despite incompatible feature semantics. Root cause of the original failure: extract_ohlcv_features column 0 changed from raw price -> log-return without anyone bumping the manually-maintained FXCACHE_VERSION const, so the L40S PVC's older cache loaded clean and the trainer fed raw prices into the aux head expecting log-returns (aux_next_bar_mse=2.587e7). Fix: * crates/ml/build.rs::emit_feature_schema_hash() runs unconditionally (before the existing CUDA-feature gate so non-CUDA builds also pick up the env var) and FNV-1a-hashes the raw bytes of the three schema-defining sources -- crates/ml/src/features/extraction.rs, crates/ml/src/fxcache.rs, crates/ml-core/src/state_layout.rs -- mixing in each file's relative path + length so renames / reorderings also bump the hash. Stable across rust versions and machines (FNV-1a, not std::hash::DefaultHasher). Emits cargo:rustc-env=FEATURE_SCHEMA_HASH=<decimal_u64> + three cargo:rerun-if-changed= lines. * crates/ml/src/fxcache.rs::FEATURE_SCHEMA_HASH consumes the env var via env! + const u64::from_str_radix(_, 10) (const-stable since rust 1.83; workspace MSRV 1.85). FxCacheHeader grows a feature_schema_hash: u64 field; header size 64->72 bytes; FXCACHE_VERSION bumped 5->6 to flag the wire-format change. validate() strict-checks the hash alongside magic / version / dims; mismatch bails with a descriptive error pointing at "source files defining feature extraction / state layout / fxcache format have changed since this cache was built." The existing precompute_features.rs:218 delete-and-regen-on-Err path handles recovery automatically; the Argo ensure-fxcache step is unchanged. * FXCACHE_VERSION docstring now declares it tracks WIRE-FORMAT changes only -- schema-level changes (feature column semantics, dimensionality) are tracked automatically by FEATURE_SCHEMA_HASH. Removes the manual ritual that broke the L40S deploy. * docs/dqn-wire-up-audit.md entry under Plan 5 Task 5 Phase F. Cost: cosmetic edits (whitespace, comments) to the three schema sources trigger one cache regen on next deploy (~5 min for full L40S dataset, ~40 s for local ES.FUT). Acceptable trade -- false negatives (missed schema drift) are not. Validation: * cargo check workspace clean at 11 warnings (baseline preserved). * Local ES.FUT cache regen confirmed: existing v5 file rejected with "Stale FxCache version: 5 (expected 6). Delete and regenerate.", regenerated v6 cache loads clean on retry (40 s, 175874 bars). * Auto-detection verified: comment-only edit to extraction.rs line 1 changed emitted hash 5046469432341222878 -> 7772630163018944575; revert returned the hash deterministically to 5046469432341222878. * multi_fold_convergence smoke PASSED (1 passed, 0 failed; 689.24 s, ~11.5 min). All 3 folds produced best-checkpoints. Per-fold best Sharpe: F0=-9.7831 (epoch 1), F1=37.9597 (epoch 2), F2=40.4789 (epoch 5). aux next_bar_mse range across all 15 epochs: 6.097e-2 -- 4.722e-1 (O(0.1), not 1e7 as in the L40S regression). No new pip/cargo deps (FNV-1a is ~10 LOC stdlib). No fingerprint change (LAYOUT_FINGERPRINT_CURRENT untouched -- this is fxcache wire-format, not GPU param layout). Files touched: * crates/ml/build.rs * crates/ml/src/fxcache.rs * docs/dqn-wire-up-audit.md Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
43d173a4eb |
plan5(task5-E): reset regression-detection streaks at fold boundary
P5T2 bug found during P5T5 Phase D smoke: walk-forward folds accumulated consecutive_warn/consecutive_error streaks across boundaries because reset_for_fold did not clear MetricBandsRegistry. F2 tripped termination at 6 consecutive even though no fold individually crossed 2N=6. Fix: add MetricBandsRegistry::reset_streaks() that clears both HashMaps; called from DQNTrainer::reset_for_fold. New unit test reset_streaks_clears_consecutive_counters proves per-fold isolation. 9/9 monitoring unit tests pass. cargo check clean at 11 warnings. |
||
|
|
eca26a1feb |
fix(dqn-v2): P4T6/P5T5 — ISV-driven aux next-bar label-scale EMA + normalize-before-MSE
Defends the aux next-bar regression head against underlying-data scale.
Label = `next_states[:, 0]` carries log returns (~1e-3) in local fxcache
but raw price (~5000) in the L40S fxcache. First L40S deploy attempt
(workflow `train-multi-seed-7j8zc`) produced `aux next_bar_mse = 2.587e7`
and grad_norm=126,769 because the unnormalised label dominated the
residual; the trunk learned garbage off the corrupted aux-gradient SAXPY.
Per `feedback_adaptive_not_tuned.md` + `feedback_isv_for_adaptive_bounds.md`:
runtime-adaptive ISV-driven EMA normalisation, NOT a tuned constant.
Changes:
- New ISV slot `AUX_LABEL_SCALE_EMA_INDEX=117`; ISV_TOTAL_DIM 117→118.
FoldReset → 1.0 (multiplicative identity, NOT 0.0). Tail-appended after
fingerprint slots so the layout grows monotonically.
- New GPU producer kernel `aux_label_scale_ema_update` in
aux_heads_loss_ema_kernel.cu: single-block 256-thread shmem reduction
over the just-gathered `aux_nb_label_buf [B]`, EMA-blends `mean(|label|)`
into ISV[117] at α=0.05.
- `aux_next_bar_loss_reduce` + `aux_next_bar_backward` kernel signatures
grow `+isv_dev_ptr+isv_label_scale_index` args; both kernels divide
label by `max(isv[117], 1e-6)` before the residual `(pred - label/scale)`
so loss + gradient stay unit-scale regardless of underlying data
magnitude. Graph-capture-stable: ISV device pointer + slot index pair
are stable; the scalar updates per step via the new producer kernel.
- `aux_heads_forward` Step 2b launches the producer between strided_gather
and the loss reduce (same captured graph, same stream → ordering
enforced).
- `aux_heads_backward` reads ISV[117] via the same device pointer.
- HEALTH_DIAG aux line gains `label_scale={:.3e}` 4th field for
observability.
- state_reset_registry adds `isv_aux_label_scale_ema` FoldReset entry;
reset_named_state dispatch arm writes 1.0 (not 0.0).
- layout_fingerprint shifts `0x26f7b1deb94cb226` → `0x829bc87b42f2feee`
(checkpoint-incompatible, no migrator per spec §4.A.2).
Validation:
- cargo check --workspace clean at 11 warnings (workspace baseline preserved).
- multi_fold_convergence smoke (RTX 3050 Ti, 591s): 1 passed.
- Fold 0: Best Sharpe = -9.7831 (matches seed=42 historical baseline -9.78).
- Fold 1: Best Sharpe = 65.3679 (within seed-noise of historical 65.96).
- Fold 2: terminated by regression-detection (avg_grad_norm escalation),
pre-existing pathology unrelated to aux head — checkpoint saved before
termination.
- HEALTH_DIAG aux line shows `label_scale=3.59e-2` to `4.35e-2` (matches
expected log-return mean-abs magnitude); `next_bar_mse=6.29e-2` to
`4.93e-1` (O(1), the new baseline post-normalisation — was O(1e-4)
pre-fix as numerical artefact of `pred ≈ 0` − tiny unnormalised label).
Aux grad_norm contribution stays bounded; explosion in F2 is downstream
C51/CQL, not aux.
Spec-aligned: aux head still regresses on `next_states[:, 0]` per spec
§4.E.6; the fix is the normalisation, not the source.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
fcf76701f4 |
plan5(task5-B): pivot multi-seed Argo from N×K (seed,fold) to N seed-only fanout
The first L40S deploy attempt (workflow `train-multi-seed-z2llf`, terminated)
failed at startup with `error: unexpected argument '--fold' found` on every
job: `train_baseline_rl` is a multi-fold walk-forward executor that accepts
`--max-folds K`, NOT `--fold N`. The original P5T1 harness assumed the
opposite and fanned out N seeds × K folds = N*K jobs, each invoking the
binary with `--seed N --fold K`.
User chose Path B: pivot to one job per seed (each runs all K folds via the
existing `--max-folds` mechanism). Per-job runtime is K× longer, but fanout
drops from N*K=30 → N=5 (matches L40S pool capacity better) and the binary
contract becomes the one the binary actually has.
4 surface changes:
1. crates/ml/examples/train_baseline_rl.rs — add `--seed N` CLI arg
(default 42 — historic implicit value). Sets `FOXHUNT_SEED` env var at
startup BEFORE any CUDA module spins up. Logs the seed value at the
training start banner.
2. crates/ml/src/cuda_pipeline/mod.rs — add `global_seed()` (reads
`FOXHUNT_SEED`, default 42) + `mix_seed(base)` (SplitMix64 avalanche
so adjacent global seeds produce uncorrelated module seeds). Six call
sites updated to mix the global seed into their previously-hardcoded
constants:
- trainer/action.rs: GpuActionSelector seed (0xDEAD_BEEF_CAFE) + the
epsilon-greedy fallback StdRng (0xAC7_DEF0).
- cuda_pipeline/gpu_iqn_head.rs: IQN Xavier-init RNG (0x1CA_1234).
- cuda_pipeline/gpu_iql_trainer.rs: V(s) Xavier-init RNG (0x1C1_9ABC).
- cuda_pipeline/gpu_her.rs: random-donor RNG (0x4E4_5678).
- cuda_pipeline/gpu_ppo_collector.rs: rng_seeds Vec for PPO
experience-collector init + reset (0xAA0_5EED).
- trainer/training_loop.rs: per-epoch regime_dropout_seed.
3. infra/k8s/argo/train-multi-seed-template.yaml — drop `fold` parameter
from `train-single` template; binary invoked as `--seed "$SEED"
--max-folds {{workflow.parameters.folds}}` so the walk-forward sweep
happens inside the single training process. Drop `FOLD` env var. Update
the nsys-rep upload filename to drop the fold suffix. Update banners /
doc comments to reflect "one-job-per-seed" semantics.
4. scripts/argo-train.sh — matrix generator drops the inner fold loop.
Each emitted task carries only `seed=${s}` and depends on the same
ensure-fxcache + gpu-warmup. The dry-run synthetic marker switches from
`seed=${s} fold=${f}` to `seed=${s} max_folds=${FOLDS}` so test harnesses
count the new shape correctly.
5. scripts/tests/test_multi_seed_harness.sh — assertions updated:
- `--multi-seed 3 --folds 2` produces 3 tasks (was 6).
- Rendered binary command must include `--max-folds
{{workflow.parameters.folds}}` placeholder.
- Rendered template must declare `folds` workflow parameter (so
`argo submit -p folds=K` overrides the default).
- Rendered binary command must NOT contain any per-fold flag — this
catches the failure mode that broke the first L40S deploy.
- Backward-compat: `--multi-seed 1 --folds 1` preserves the existing
single-template path (no DAG matrix tasks emitted).
6. docs/dqn-wire-up-audit.md — adds 1 Wired row documenting the pivot,
the new `--seed`/`mix_seed` plumbing, all 6 RNG call sites, and the
end-to-end seed-variation verification result.
Validation:
cargo check --workspace clean at 11 warnings (workspace baseline preserved).
cargo build --release --example train_baseline_rl succeeds; --help shows
the new --seed flag with documented default 42.
Seed-variation end-to-end test on RTX 3050 Ti (1 fold × 2 epochs each):
--seed 42 → F0 best Sharpe = -9.7831, best_val_metric = 1.957244,
epoch-2 train Sharpe = -16.12, val_Sharpe = +1.11.
--seed 999 → F0 best Sharpe = +92.9341, best_val_metric = 2.161012,
epoch-2 train Sharpe = +92.93, val_Sharpe = -0.25.
Different best Sharpe / best_val_metric / epoch-2 train + val Sharpe
across seeds proves the seed actually propagates through the RNG init
paths and is not just accepted-and-ignored. The seed=42 numbers match
the prompt's "deterministic baseline" expectation (F0 = -9.7831 was
bit-identical pre-pivot because no global-seed plumbing existed).
./scripts/argo-train.sh dqn --multi-seed 5 --folds 6 --dry-run produces
exactly 5 WorkflowTask markers (train-s0..train-s4), each with
`--max-folds {{workflow.parameters.folds}}` in the binary invocation.
All 3 harness tests PASS:
- test_multi_seed_harness.sh: 5 PASS lines, exit 0.
- test_nsys_harness.sh: 4 PASS lines + ALL PASS, exit 0.
- test_tier_checks.sh: PASS overall (good-fixture passes, bad-fixture
surfaces expected check rejections), exit 0.
Backward compat: existing single-job `argo-train.sh` callers (no
`--multi-seed`, no `--folds`) route to the original `train-template.yaml`
unchanged. `--seed 42` is a no-op offset for the SplitMix64 mix at the call
sites — the trajectory shifts only when the user passes `--seed` explicitly,
matching the prompt's "default 42 (historic implicit value)" requirement.
L40S pool: argo-train.sh defaults `--gpu-pool ci-training-h100`; user passes
`--gpu-pool ci-training-l40s` at deploy time. No script default change
(per constraint 5).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
fbee2a00f5 |
plan5(task5-A): wire tier 2/3 val_* metrics into HEALTH_DIAG
Plan 5 Task 4 left every tier-2/tier-3 check failing with "metric missing
from aggregate" because the existing 'Validation backtest:' free-form log
line was not parseable by the aggregate-multi-seed-metrics.py block-keyed
parser. Phase A closes that gap end-to-end (CPU-only, no kernel touch):
* metrics.rs::compute_validation_loss — emit a new
HEALTH_DIAG[<epoch>]: val [sharpe=… sortino=… win_rate=…
max_drawdown=… trade_count=… calmar=…
omega_ratio=… total_pnl=… var_95=… cvar_95=…
trades_per_bar=… active_frac=… dir_entropy=…
sharpe_annualised=… profit_factor=…
window_bars=…]
block immediately after the existing 'Validation backtest:' line. All
16 keys derive from the existing GpuBacktestEvaluator WindowMetrics
reduction (no new GPU work):
- sharpe / sortino / win_rate / max_drawdown / total_trades /
calmar / omega_ratio / total_pnl / var_95 / cvar_95 / buy_count /
sell_count / hold_count come straight from m.*
- window_bars = buy + sell + hold (kernel tallies one direction
per bar)
- trades_per_bar = total_trades / window_bars
- active_frac = (buy + sell) / window_bars (kernel folds Hold AND
Flat into hold_count, so 'active' = bars where the policy chose
Short or Long — meets the Tier-2 'not always Hold' intent)
- dir_entropy = -Σ p ln p over the 3-bucket {short, hold-or-flat,
long} distribution. Documented limitation: max log(3) ≈ 1.099
vs spec's 4-bucket 0.8·log(4) ≈ 1.109 ceiling — tier2 dir_entropy
threshold is unreachable from this 3-bucket distribution; resolution
tracked in audit row.
- sharpe_annualised = m.sharpe alias (kernel already multiplies by
sqrt(bars_per_day · 252) at backtest_metrics_kernel:266)
- profit_factor = m.omega_ratio alias (kernel's omega computes
gain_sum/loss_sum at threshold 0, equivalent to per-step PF;
trade-level PF deferred — needs boundary-aware kernel work)
* mod.rs — adds last_val_metrics: Option<[f32; 14]> on DQNTrainer to
snapshot the WindowMetrics-derived values for downstream consumers
(smoke tests, future telemetry).
* constructor.rs — initialises the new field to None.
* aggregate-multi-seed-metrics.py — switches the block→key joiner from
'__' to '_' so 'val [sharpe=…]' surfaces as the bare 'val_sharpe'
aggregate key the tier check scripts and synthetic test fixtures
already expect. The pre-existing '__' joiner was an oversight in
Plan 5 Task 1B that was never validated against actual aggregator
output (the aggregator emitted 90 'block__key' metrics that nothing
consumed; the synthetic good_tier1.json / bad_tier1.json fixtures
were always shaped as 'val_sharpe', confirming the single-underscore
convention was intended). Renaming the 90 existing keys is safe — no
consumers had locked in on the '__' form.
* docs/dqn-wire-up-audit.md — updates Plan 5 Task 4 row to reference
the now-landed wiring and adds a new row documenting the val [...]
HEALTH_DIAG block pipeline + aggregator joiner change + the deferred
4-bucket dir_dist + trade-level PF caveats.
Validation:
cargo check --workspace clean at 11 warnings.
multi_fold_convergence smoke (629s, 3 folds × 5 epochs on RTX 3050 Ti)
PASSES with 3/3 fold checkpoints. Per-fold best Sharpe: -9.78 / 42.46 /
88.18 (within smoke noise band — no perturbation from the additive
CPU-only HEALTH_DIAG line).
scripts/aggregate-multi-seed-metrics.py against /tmp/p5t5a-smoke.log
produces 3 streams (one per fold), 16 val_* keys all present:
val_sharpe, val_sharpe_annualised, val_sortino, val_win_rate,
val_max_drawdown, val_trade_count, val_calmar, val_omega_ratio,
val_total_pnl, val_var_95, val_cvar_95, val_trades_per_bar,
val_active_frac, val_dir_entropy, val_profit_factor, val_window_bars.
check_tier2.py / check_tier3.py rejection messages are now substantive
(threshold-based) rather than "missing key":
Tier 2: trades_per_bar PASS @ 0.0127; active_frac FAIL @ 0.058 (model
mostly Hold on 5-epoch smoke); dir_entropy FAIL @ 0.18 (within
documented 3-bucket vs 4-bucket caveat).
Tier 3: sharpe_annualised FAIL @ -0.25 (5-epoch smoke not converged);
win_rate skipped (192 trades ≤ 500 noise gate); profit_factor
FAIL @ 0.18 (untrained policy).
Real validation pass requires the L40S 60-epoch run (Phase C).
Deferred (out of T5 Phase A scope):
- val_dir_dist_{short,hold,long,flat} per-direction breakdown — kernel
intentionally collapses Hold+Flat for trade-cycle counting; Tier-2's
log(4) threshold needs either a kernel-level split or a 3-bucket
threshold tweak in check_tier2.py.
- Trade-level profit_factor (sum-winner-PnL / sum-loser-PnL) vs the
per-step omega-equivalent emitted here.
- avg_q_value bare-key aggregation — the metric is logged via separate
Prometheus + tracing paths but not inside any HEALTH_DIAG block; out
of T5 Phase A scope and pre-existing.
|
||
|
|
0d373da490 |
plan5(task4): tiered-exit validation script suite (tier1/2/3 checks)
Creates scripts/validation/ with per-tier exit checks consuming the
aggregate JSON from scripts/aggregate-multi-seed-metrics.py (P5T1B):
check_tier1.py — convergence (std/mean ≤ 0.15 on val_sharpe /
avg_q_value / train_loss; avg_q_value max ≤ 500 fold-1 explosion
guard; placeholders for Q-saturation + hot-path-DtoH per spec).
check_tier2.py — behavioural (val_trades_per_bar ≥ 0.005,
val_active_frac > 0.2, dir argmax entropy > 0.8·log4 with
val_dir_entropy primary + val_dir_dist_* fallback).
check_tier3.py — profitability (val_sharpe_annualised > 1.0 with
val_sharpe per-bar fallback, val_win_rate ≥ 0.52 gated on
>500 trades, val_profit_factor mean ≥ 1.1 AND cross-seed std < 0.3).
check_all_tiers.py — subprocess wrapper, exits 0 only if all pass.
Stdlib-only (statistics / argparse / json / subprocess) — no new deps.
Defensive missing-metric handling: each check FAILs with an explanatory
message when its required aggregate key is absent rather than silently
passing, so missing HEALTH_DIAG metrics are surfaced loudly.
Test harness scripts/validation/tests/test_tier_checks.sh exercises
good + bad fixtures across all four scripts and against the wrapper.
Audit row added to docs/dqn-wire-up-audit.md documenting the suite +
the deferred metrics list (val_trades_per_bar, val_active_frac,
val_dir_entropy/_dist_*, val_sharpe_annualised, val_win_rate,
val_profit_factor, val_trade_count) that HEALTH_DIAG must emit before
tiers 2/3 can ever PASS on real data — tracked for Plan 5 Task 5
pre-flight wire-up.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
2606506cd8 |
plan5(task3): A.4.1 nsys profile harness with regression-comparison script
- argo-train.sh: --profile flag forces multi-seed render path so the nsys wrapper + foxhunt-training-artifacts upload step are visible in --dry-run YAML without cluster contact (test surface). - train-multi-seed-template.yaml: new `profile` parameter (default "false") gates the per-(seed, fold) `nsys profile --capture-range=cudaProfilerApi` wrapper and the `mc cp` upload to foxhunt-training-artifacts/profiles/<sha>/. mc binary fetched on-demand (ci-builder image lacks it). MinIO creds optional — upload warn-skips if absent. - Dockerfile.foxhunt-training-runtime: install nsight-systems-cli unpinned (pinning the stale 2024.4.1.61-1 from earlier plans breaks builds when apt index advances). - minio.yaml: add foxhunt-training-artifacts bucket to minio-init. - compare-nsys-profiles.py: V0 regression detector — compares cuda_gpu_kern_sum total_ns / epoch_count between two profiles; exits 1 on >20% slowdown. NVTX per-epoch ranges deferred to T5. - tests/test_nsys_harness.sh: dry-run grep test — verifies both required strings appear when --profile is set, and that the default (no --profile) path keeps profile=false in the rendered template. - dqn-wire-up-audit.md: Plan 5 Task 3 row added documenting the harness + the baseline-capture deferral to T5. Backward compat: test_multi_seed_harness.sh from P5T1 still PASS. |
||
|
|
6cdfbff8d6 |
plan5(task2): A.4 regression-detection hard-stop on 2N consecutive error-band
Adds the convergence guardrail: every per-epoch HEALTH_DIAG metric is
checked against the bands in config/metric-bands.toml; N consecutive
warn-band epochs emit a tracing::warn; 2N consecutive error-band epochs
return Err(CommonError::RegressionDetected{...}) cleanly from the
training loop, which propagates to the train_baseline_rl subprocess
exit code (no libc::raise — clean Rust error path).
Wire-points:
- New module: crates/ml/src/trainers/dqn/trainer/monitoring.rs
- MetricBands {warn_low, warn_high, error_low, error_high}
- BandSettings {consecutive_epochs_for_warn, consecutive_epochs_for_error}
- MetricBandsRegistry: load_from_toml + update_and_check
- TerminationReason {RegressionWarn, RegressionError}
- NaN treated as out-of-band (consecutive++; never resets streak)
- Unknown metrics return None (silent OK per Invariant 7 audit)
- crates/common/src/error.rs: new CommonError::RegressionDetected variant
carrying {metric, value, band, consecutive}
- crates/ml/src/trainers/dqn/trainer/constructor.rs: load
config/metric-bands.toml at trainer init; warn-only on missing file
(backward compat for environments without the config)
- crates/ml/src/trainers/dqn/trainer/training_loop.rs: harvest per-epoch
metrics (parallel emit alongside HEALTH_DIAG), feed each through
registry.update_and_check; on Some(TerminationReason::RegressionError)
emit final HEALTH_DIAG[N]: TERMINATED_BY_REGRESSION line and return Err
- services/trading_service/src/error.rs: minimal handler for the new
CommonError variant (existing pattern)
Validation:
- 8 unit tests in monitoring::tests pass (band logic, NaN, warn-only
behaviour, error-streak threshold, unknown-metric, invalid TOML)
- regression_detection GPU smoke (3.19s): trainer with intentionally
narrow train_loss error band [0, 1e-9] self-terminates at epoch 5
after 6 consecutive error-band epochs; final HEALTH_DIAG line emits
TERMINATED_BY_REGRESSION with metric/value/consecutive/band fields
- multi_fold_convergence smoke (650s, --release): all 3 folds train
to completion, all 3 checkpoints saved, no false-positive
termination on the populated metric bands. Per-fold best train
Sharpe: F0=-9.7831 (bit-baseline), F1=25.8272, F2=39.2687. F1/F2
on the lower end of observed noise distribution
({74.56, 61.10, 71.53, 25.83} for F1; {88.20, 61.57, 65.96, 39.27}
for F2) but training healthy throughout: aux clauses fire every
epoch, sharpe_ema recovers from F0 collapse (-9.78 → +14.8 by start
of F2), no regression detection trips.
config/metric-bands.toml populated for the metrics emitted by
HEALTH_DIAG today (avg_q_value, train_loss, val_sharpe, train_sharpe,
aux_next_bar_mse, aux_regime_ce, isv_* slot EMAs, sharpe_ema, etc.).
Bands derived from current cleanroom smoke + permissive defaults
where only one sample exists; populate-metric-bands-from-runs.py will
tighten them after Plan 5 Task 5's multi-seed pass produces real
distributions.
Constraints honoured: GPU-only in hot path (band check is CPU-side
post-HEALTH_DIAG, off the captured graph); no atomicAdd; no stubs;
no // ok: band-aids; no tuned constants beyond the toml-loaded bands;
no .unwrap() introduced; cargo check clean at 11 warnings (workspace
baseline preserved, plus ml-dqn pre-existing 1 warning).
Audit doc: new row added documenting monitoring.rs module, the
CommonError variant, the training_loop wire-point, and the design
choice that band-checks run AFTER HEALTH_DIAG emit (not before) so
the diag log already reflects the metric values that triggered any
termination.
Plan 5 T1 (multi-seed harness) landed at c6634254e+47c8b783c; T2
(this) gives the regression hard-stop that the multi-seed final
pass (T5) consumes to bail out early on bad seeds.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
e47d067390 |
plan4(task7): Part E audit close-out — every supervised concept landed or OUT
Updates the supervised → DQN concept audit doc to its terminal state per Plan 4 Task 7. Every Part E row + the cross-Plan-2 D.1/D.8 rows now cite the commit SHA in which they landed; xLSTM/KAN remain OUT-intentional (redundant with Mamba2+TLOB and not a bottleneck respectively); Liquid is AUDITED-LANDED (deleted from DQN per D.7's identity-at-fixed-point finding). Landed SHAs: - E.1 (TFT VSN): |
||
|
|
647f15f9db |
feat(dqn-v2): Plan 4 Task 6 Commit B — aux heads behavioral wiring
Activates the dormant Commit A scaffolding: forward + backward + ISV
producer + aux loss combination + HEALTH_DIAG observability. Two
auxiliary heads (next-bar return MSE + 5-class regime CE) train end-
to-end off the trunk's h_s2 activation, with their gradient flowing
into the trunk's bw_d_h_s2 accumulator (ISV-scaled by aux_weight)
before encoder_backward_chain consumes it.
Wire-points:
- Forward (launch_cublas_forward, after forward_online_raw saves
save_h_s2, BEFORE stochastic depth): regime label builder →
strided_gather (col 0 of next_states → dedicated aux_nb_label_buf,
NOT aliased) → next-bar forward → regime forward → next-bar MSE
reduce → regime CE reduce.
- Backward (launch_cublas_backward_to, AFTER backward_full fills
bw_d_h_s2, BEFORE encoder_backward_chain): per-head backward
emits per-sample partials + per-sample dh_s2; aux_param_grad_reduce
collapses partials → final, then SAXPY into grad_buf[119..127);
both dh_s2 buffers SAXPY into bw_d_h_s2 with alpha=aux_weight.
- ISV producer (training_loop.rs per-step): launch_aux_heads_loss_ema
alongside launch_h_s2_rms_ema / launch_vsn_mask_ema; updates
ISV[113][114] EMAs.
- Aux-weight refresh (training_loop.rs per-step): clamp(0.1 ×
LEARNING_HEALTH × (1 − tanh(0.1 × sharpe_ema)), 0.05, 0.3); pushed
via FusedTrainingCtx::set_aux_weight delegate.
- HEALTH_DIAG aux clause: emits per-epoch aux[next_bar_mse=…
regime_ce=… w=…] from the ISV slots populated by the producer.
Lessons applied from previous WIP attempt (stashed):
- Dedicated aux_nb_label_buf [B] f32 field (NOT aliased over
aux_partial_nb_b2) — partial-refactor invariant honoured cleanly.
- Init log "AuxHeadsForwardOps initialized: K_nb=1 K_rg=5 aux_h=32 …"
for observability sanity check.
- HEALTH_DIAG aux clause verified firing in smoke (5 lines, one per
epoch, with non-zero loss EMAs and ISV-driven aux_weight evolution
0.050 → 0.105 over 5 epochs).
Smoke (multi_fold_convergence --release, 3 folds × 5 epochs on RTX
3050 Ti, 684.37s, all 3 dqn_fold{N}_best.safetensors written):
F0=-9.7831 F1=71.5327 F2=65.9598
Within baseline noise band [F1: 61.10–74.56, F2: 61.57–88.20] from
two clean Commit A baseline runs. F0 is bit-reproducible across all
runs (deterministic cold-start collapse under --release profile —
pre-existing, not from this task).
HEALTH_DIAG observability (5 epochs):
aux [next_bar_mse=9.962e-5 regime_ce=2.099e-5 w=0.050]
aux [next_bar_mse=2.295e-4 regime_ce=3.924e-5 w=0.088]
aux [next_bar_mse=3.261e-4 regime_ce=6.893e-5 w=0.099]
aux [next_bar_mse=4.307e-4 regime_ce=6.819e-5 w=0.102]
aux [next_bar_mse=4.824e-4 regime_ce=6.634e-5 w=0.105]
Aux losses grow as expected (cold-start Xavier predictions diverge
from labels until trunk learns); aux_weight rises with LEARNING_HEALTH
and sharpe_ema; nothing pinned at clamp boundaries — formula working.
Constraints honoured: GPU-only (every reduce + SAXPY + EMA on-device,
zero DtoH in hot path; only HEALTH_DIAG cold-path read on CPU); no
atomicAdd; no stubs; no // ok: band-aids; no buffer aliasing tricks;
no tuned constants beyond the documented 0.1 aux base + [0.05, 0.3]
numerical-stability clamp; partial-refactor invariant honoured
(bw_d_h_s2 accumulator semantics extended, not replaced;
encoder_backward_chain consumes through unchanged pointer).
target_ema_update NOT extended for aux heads (Commit A design
decision preserved): aux heads are online-only supervised heads, NOT
used in Bellman bootstrapping. target_params_buf[119..127) remains at
Xavier init forever.
cargo check clean at 11 warnings (workspace baseline preserved); no
fingerprint change (Commit A already shifted to 0x26f7b1deb94cb226).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
5478e7c821 |
feat(dqn-v2): Plan 4 Task 6 Commit A — aux heads scaffolding (kernels + params [119..127) + 2 ISV slots, additive)
Multi-task auxiliary heads (E.6) — additive scaffolding ONLY. NO
production callers in this commit; Commit B wires forward/backward +
training-loop loss accumulation.
Two `Linear(SH2 → 32) → ELU → Linear(32 → K)` MLPs branch off h_s2:
- Next-bar return regression head (K=1, MSE loss).
- 5-class regime classification head (K=5, cross-entropy loss).
Lands in this commit (additive, no behavioural change):
1. Two CUDA kernel files (kernel count 62 → 64):
- `aux_heads_kernel.cu` — 8 entry points (per-head forward, backward,
loss-reduce; shared regime label-builder + per-tensor batch-reduce).
Single-block-per-sample, shmem-tree reductions, no atomicAdd.
- `aux_heads_loss_ema_kernel.cu` — single-thread single-block ISV
producer mirroring h_s2_rms_ema_kernel.cu's footprint.
2. Rust orchestrator `cuda_pipeline/gpu_aux_heads.rs` —
`AuxHeadsForwardOps` + `AuxHeadsBackwardOps` mirror gpu_grn.rs's
raw-u64-pointer ABI for graph-capture safety. NO callers.
3. Param tensors at [119..127) (NUM_WEIGHT_TENSORS 119 → 127):
aux_nb_{w1,b1,w2,b2}, aux_rg_{w1,b1,w2,b2}. Xavier on weights,
zero on biases. Adam SAXPY iterates 0..total_params (count-driven),
so the new range gets covered automatically; with no producer for
grad_buf[119..127) in this commit, Adam keeps params at Xavier
init (intended dormant state).
4. Two new ISV slots tail-appended after VSN:
AUX_NEXT_BAR_MSE_EMA_INDEX=113, AUX_REGIME_CE_EMA_INDEX=114.
Producer: aux_heads_loss_ema_update. Cold-start 0.0; FoldReset → 0.0.
Diagnostic only; no GPU consumer kernel reads slots [113..115).
5. Fingerprint pair shifted [111,112] → [115,116]; ISV_TOTAL_DIM
113 → 117. layout_fingerprint_seed extended with the 2 new ISV
slot names + 8 new PARAM_AUX_* entries terminating at
PARAM_TOTAL_TENSORS=127. New LAYOUT_FINGERPRINT_CURRENT =
0x26f7b1deb94cb226 (was 0x1b28321bb816f246). Checkpoint break by
intent — fail-fast at constructor load on mismatch.
6. Two new FoldReset entries in state_reset_registry.rs
(isv_aux_next_bar_mse_ema, isv_aux_regime_ce_ema) WITH matching
dispatch arm in training_loop.rs::reset_named_state (the wire
VSN-rc2 missed and chain-final fix made dispatch-mandatory).
7. Polyak EMA target sync NOT extended for aux heads — design
decision: aux heads are online-only supervised heads, not used in
Bellman bootstrapping. target_params_buf[119..127) initialised by
xavier_init_params_buf (same Xavier values as online) and never
updated; verified no consumer reads target_params_buf past
padded_byte_offset(119).
Constraints honoured: GPU-only (every reduction GPU-side, zero DtoH);
no atomicAdd (shmem-tree reductions throughout backward + final
batch-dim collapse); no stubs (every Rust function is real
implementation; "no callers yet" applies only to orchestrator wrappers
waiting on Commit B); partial-refactor invariant honoured
(NUM_WEIGHT_TENSORS, Adam range, xavier init, fingerprint seed,
ISV_TOTAL_DIM, FoldReset entries + dispatch arm all migrate together);
no `// ok:` band-aids; no tuned constants.
cargo check --workspace clean at 11 warnings (workspace baseline
preserved). Smoke deferred to Commit B (no behaviour change in this
commit — aux head params are dormant Xavier weights).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
31e0f219a5 |
feat(dqn-v2): Plan 4 Task 1B-iv chain final — VSN backward + Polyak EMA target sync fix
Combines uncommitted iv-main + iv-ext aux paths + Fix A (Polyak EMA
extension covering VSN [95..119)) + Fix B (missed dispatch arm), and
strips rc2/rc3 band-aids that turned out to be unnecessary once the
upstream cause was fixed.
Root cause (focused HtoD/DtoD/pinned-memory audit):
target_ema_update only ran the EMA kernel over non_isv_params at
byte range covering [0..FIRST_ISV_TENSOR=77), skipping the 24 VSN
tensors at [95..119) added in 1B-ii. Online VSN trained from the
1B-iv backward chain; target VSN stayed at alloc_zeros for the
entire run. Bellman target used softmax(zeros) = uniform 1/6 mask
while online's mask drifted toward [market=0.13, portfolio=0.25].
The systematic Q-estimate divergence collapsed fold-2 magnitude
branch and pinned F0 best Sharpe at 21.14 across every prior
magnitude-scaling / state-isolation attempt.
Fix A: extend target_ema_update with a second dqn_ema_kernel launch
covering params_buf[vsn_param_byte_offset..vsn_param_total) so
target VSN tracks online VSN through the same Polyak EMA the rest
of the network uses. Same kernel signature, on-device, no DtoH.
Fix B: add the missed isv_vsn_aux_grad_scale dispatch arm in
training_loop.rs::reset_named_state that the rc2 work added
without its dispatch counterpart.
Cleanup (Phase B): strip rc2 ISV slot 113 + dqn_scale_f32_isv_scaled
+ aux_bottleneck_vsn_backward_dispatch indirection AND rc3 split-
Adam vsn_m_buf/vsn_v_buf — band-aids for the symptom Fix A actually
addresses. VSN params return to the main Adam state; aux-path
SAXPYs use the original direct-saxpy pattern from 1B-iv-ext.
Fingerprint reverts to the 1B-ii value 0x1b28321bb816f246 (no
checkpoint break beyond what 1B-ii already required).
Smoke (multi_fold_convergence --release-test, 3 folds × 5 epochs on
RTX 3050 Ti, 671.68s, all 3 dqn_fold{N}_best.safetensors written):
F0=93.4114 F1=73.0430 F2=73.0749
geom-mean = 79.31 vs 71.24 floor = +11.3%
F0 +36% over 1B-iii baseline; F2 +18%; F1 -14% (60-epoch L40S
run will validate equilibration; short-horizon asymmetry expected).
Constraints honoured: GPU-only (target EMA on-device), no atomicAdd,
no stubs, no // ok: band-aids, no tuned constants (Polyak EMA tau
shared with existing main-range launch), partial-refactor invariant
(dqn_ema_kernel signature unchanged — both launches use identical
args, different byte offsets/counts).
Lesson: 4 prior remedies (rc, rc2, rc3, rc4) and 1 diagnostic run
(E1) all chased downstream symptoms (gradient scale, Adam variance,
kernel sync). The upstream cause was a 1-line gap in Polyak target
sync that didn't include post-Plan-4 tensors. Same shape as the
2c.3a follow-up bottleneck Linear bug (commit
|
||
|
|
f2609fdcc7 |
feat(dqn-v2): Plan 4 Task 1B-iii — VSN forward orchestrator + 3 wire sites + ISV producer launch
Wire the VSN feature-selection forward at all 3 production passes so the
gated state actually flows through downstream consumers. The 24 VSN param
tensors at indices [95..119) (1B-ii) and the 6 ISV mask-EMA slots at
[105..111) (1B-ii) acquire their first production callers.
VSN orchestrator on `CublasGemmSet::vsn_forward` (in `batched_forward.rs`)
runs the per-group MLP via 6× `(sgemm_f32_ldb Linear_1[g] → fused
add_bias_relu → sgemm Linear_2[g] → bias-add → strided_scatter into col g
of the assembled [B, num_groups] logits buffer)` followed by the
`vsn_softmax_and_gate_forward` kernel (1 block per sample, 256 threads,
24-byte shmem). Per-group state slice is selected via `state_in_ptr +
gb * sizeof(f32)` with `ldb = state_dim_padded` so cuBLAS reads the
col-major `[state_dim_padded, B]` view starting at the group's first
feature column — no slice copy needed. The fused `add_bias_relu` writes
the post-ReLU activation back into `h1_ptr` so the buffer IS the
save-for-backward storage on the online pass; the `strided_scatter`
(reused from `experience_kernels.cu`) uses `src_stride=1, dst_stride=
num_groups` and dst pointer offset by `g * sizeof(f32)` to interleave
the 6 per-group logits into the assembled buffer without atomicAdd.
3 production wire sites all consume the orchestrator:
* `launch_cublas_forward` online-on-states pass — saves `vsn_logits_buf`
/ `vsn_mask_buf` / per-group `vsn_save_h1_g{0..5}_buf` for 1B-iv
backward; bottleneck Linear, `bn_tanh_concat`, and `launch_concat_ofi`
read the gated state via the `states_for_bn` local.
* `launch_cublas_forward` target-on-next_states pass — uses `tg_w_ptrs`
(Polyak EMA of online VSN weights — the existing target-EMA loop
covers tensors [95..119) automatically since `target_params_buf` is
sized to `total_params + cutlass_tile_pad`); throwaway logits/mask
scratches; single `vsn_h1_inference_scratch` shared across the 6
groups (target is inference-only — no save).
* `submit_forward_ops_ddqn` DDQN-online-on-next_states pass — uses
`on_w_ptrs` (DDQN argmax runs the online net on next_states);
throwaway scratches; same `vsn_h1_inference_scratch`.
Per-pass distinct mask + logits scratches prevent target/DDQN from
clobbering the online pass's saved buffers — the ISV producer reads
`vsn_mask_buf` (online) only.
ISV producer launcher `launch_vsn_mask_ema(ema_alpha)` mirrors
`launch_h_s2_rms_ema` (single-block 256-thread shmem reduce, no
atomicAdd, no DtoH); wired in `training_loop.rs` at the per-step ISV
producer block alongside `launch_h_s2_rms_ema` and
`launch_iqn_quantile_ema`. Both 1B-i cubin statics
(`VSN_FEATURE_SELECTION_CUBIN`, `VSN_MASK_EMA_CUBIN`) lose their
`#[allow(dead_code)]` annotations — they now have runtime
`load_cubin(...)` consumers in `CublasGemmSet::new` and
`GpuDqnTrainer::new` respectively. The forward kernel gains a kernel
handle on `CublasGemmSet`; the backward kernel handle is loaded here
(stays `#[allow(dead_code)]` until 1B-iv consumes it).
VSN dW = 0 in this commit — no backward chain extension yet (1B-iv
adds it). Expected behaviour analysis: the cold-start
`softmax(≈small_random_logits) ≈ 1/6 + ε` per group acts as a
near-uniform multiplicative scale on every state row that the
downstream GRN trunk's LayerNorm absorbs at the first cuBLAS Linear
— so the smoke either matches the baseline (if the gate is truly
uniform) or strictly improves (if the random per-group bias acts as
a tiny attention-like signal that the trunk amplifies through
training despite zero VSN dW).
No fingerprint change (no new ISV slot or param tensor — pure
orchestrator + buffer + wire-in). No DtoH in any new path. No
atomicAdd. No stubs (the only `Option<CudaFunction>` is gated by
`debug_assert!` + `expect()`, never silent-skipped). All 3 forward
paths attach VSN at the same point and downstream consumers all read
the gated buffer — no path is left reading raw `states_buf` /
`next_states_buf`.
Smoke: `cargo test … multi_fold_convergence --ignored --release`,
594.94s, 3 folds × 5 epochs on RTX 3050 Ti — `test result: ok`.
Per-fold best train Sharpe 68.83 / 84.84 / 61.95 at epochs 2 / 5 / 4
(vs post-1B-ii baseline 6.53 / 80.11 / 66.72 → geom-mean 39.27);
this commit geom-mean = (68.83 × 84.84 × 61.95)^(1/3) ≈ 71.24,
+81% lift vs baseline — vastly above the spec's ≥27.5 acceptance
threshold. No NaN/Inf, no fingerprint mismatch (unchanged at
`0x1b28321bb816f246`), no panic, no producer-launch warning.
cargo check clean at 11 warnings (workspace baseline preserved);
cargo build --release -p ml clean.
+663 / -31 LOC across:
- crates/ml/src/cuda_pipeline/batched_forward.rs (orchestrator +
kernel handle fields + load + scatter wire-up)
- crates/ml/src/cuda_pipeline/gpu_dqn_trainer.rs (19 new buffer
fields + allocations + launch_vsn_mask_ema + 3 wire sites + cubin
static dead_code retire + scatter wire-up at construction)
- crates/ml/src/trainers/dqn/trainer/training_loop.rs (per-step
producer launch alongside h_s2_rms_ema)
- docs/dqn-wire-up-audit.md (1B-iii entry)
2 Orphan-by-design rows (1B-i kernels) reclassify Wired.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
dde38d800d |
feat(dqn-v2): Plan 4 Task 1B-ii — VSN params (24 tensors) + 6 ISV slots (additive)
Per-group VSN MLP parameter tensors and per-group importance-mask EMA ISV
slots land additively — no production callers in this commit. The
forward orchestrator + cuBLAS Linear_1/Linear_2 wire-in lands in 1B-iii;
backward orchestrator + autograd integration lands in 1B-iv.
Param tensors (24 new, indices [95..119)):
* Per-group quad (`vsn_w1_g{g}`, `vsn_b1_g{g}`, `vsn_w2_g{g}`,
`vsn_b2_g{g}`) for each `FEATURE_GROUP_RANGES` entry
(market / ofi / tlob / mtf / portfolio / plan_isv).
* `VSN_HIDDEN_DIM = 16`. Per-group element count =
16*group_dim_g + 33; sum = 16*121 + 6*33 = 2134 floats.
* Xavier init on weights, zero on biases — initial logits ≈ 0 yield
softmax mask ≈ 1/SL_NUM_FEATURE_GROUPS (uniform cold-start).
* `NUM_WEIGHT_TENSORS` 95 → 119; `compute_param_sizes()` and
`xavier_init_params_buf()` rewritten with a runtime per-group loop
reading `FEATURE_GROUP_RANGES` from `ml_core::state_layout`
(Task 1A's prerequisite).
* Adam state, target params, and grad buffers grow automatically
since their allocations use `compute_total_params(cfg) +
cutlass_tile_pad`.
ISV slots (6 new, indices [105..111)):
* `VSN_MASK_GROUP_0_EMA_INDEX = 105` (market) through
`VSN_MASK_GROUP_5_EMA_INDEX = 110` (plan_isv).
* Cold-start `1.0 / SL_NUM_FEATURE_GROUPS = 1/6` (uniform prior — no
group preference at init).
* FoldReset reapplies the same uniform value at fold boundary; 6 new
`RegistryEntry` rows + a single dispatch arm in
`training_loop.rs::reset_named_state` covering all six names.
* Producer kernel `vsn_mask_ema_update` (1B-i) overwrites these slots
once the launch site is wired in 1B-iii.
Fingerprint:
* Pair shifted 103/104 → 111/112; `ISV_TOTAL_DIM` 105 → 113.
* `layout_fingerprint_seed()` extended with the 6 new ISV slot names
+ 24 new `PARAM_VSN_*` entries terminating at
`PARAM_TOTAL_TENSORS=119`.
* New `LAYOUT_FINGERPRINT_CURRENT = 0x1b28321bb816f246`
(was `0x5789155b683ab59c`). Checkpoint break by intent — fail-fast
at constructor load on mismatch, no migration path per
`feedback_no_legacy_aliases.md`.
Cubin refs:
* `VSN_FEATURE_SELECTION_CUBIN` and `VSN_MASK_EMA_CUBIN` added as
`pub(crate) static` `#[allow(dead_code)]` byte arrays — the
`include_bytes!()` runs at compile time, but no `load_cubin(...)` /
`load_function(...)` call exists yet (deferred to 1B-iii).
Mirrors the 2c.3a pattern (param tensor reshuffle without runtime
callers — that landed cleanly because the runtime sites were panic-
gated; here, the runtime sites simply don't exist yet so no gating
needed).
cargo check clean at 11 warnings (workspace baseline preserved); cargo
build compiles 62/62 cubins clean (unchanged — no new `.cu` files).
Smoke deferred — behaviour byte-identical to the
|
||
|
|
423b36646f |
feat(dqn-v2): Plan 4 Task 1B-i — VSN kernel module landing (additive, no callers)
Two new CUDA kernel files registered in build.rs (kernel count 60 → 62, all compile clean under nvcc sm_80). Module is additive — ZERO production callers; consumers wired in 1B-ii/iii/iv. vsn_feature_selection_kernel.cu (262 LOC, cubin 31,392 B): - vsn_softmax_and_gate_forward: per-sample (B blocks × 256 threads) numerically-stable softmax over 6 feature groups + per-feature gate- multiply with passthrough for the 7-element padding tail. - vsn_softmax_and_gate_backward: full softmax-over-groups Jacobian (d_logit[g] = mask[g] * (d_dot[g] - sum_h(mask[h] * d_dot[h]))) plus per-feature d_state via mask broadcast. Per-block shmem tree reduce for d_dot, no atomicAdd. vsn_mask_ema_kernel.cu (78 LOC, cubin 6,688 B): - vsn_mask_ema_update: single-block 256-thread reduction over the per- sample mask buffer → 6-element batch-mean → per-slot EMA into ISV[first..first+6). Mirrors h_s2_rms_ema_kernel.cu's shmem-reduce pattern exactly. Producer-only. Caller-side cuBLAS GEMMs (Linear_1[g] / ReLU / Linear_2[g]) feed logits to the forward kernel — mirrors gpu_grn.rs's contract pattern. Layout row-major [B, state_dim_padded] for state, [B, num_groups] for the mask. All reductions GPU-side per pearl_cold_path_no_exception_to_gpu_drives.md. No checkpoint break, no fingerprint change, no new param tensors, no new ISV slots in this commit. cargo check clean at 11 warnings (baseline preserved); cargo build compiles 62/62 cubins clean (was 60). Smoke deferred to 1B-iii where the forward orchestrator wire-in actually validates something — multi_fold_convergence byte-identical to the post- bottleneck-fix baseline (geom-mean 39.27 / per-fold 6.53 / 80.11 / 66.72). 2 new Orphan-by-design rows (will retire at 1B-iii). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
f3e3ac3477 |
fix(dqn-v2): bottleneck Linear runtime indices + missing target/DDQN bn paths (2c.3a follow-up)
Three related bugs from the 2c.3a GRN trunk reshuffle that the bottleneck-Linear runtime sites silently inherited: 1. **on_w_ptrs[24/25] -> [33/34]** (4 sites in gpu_dqn_trainer.rs). 2c.3a's +9 shift migrated branch/value/VSN/GLU/KAN consumers but missed the bottleneck Linear's forward GEMM (line ~14098), forward bias-add (~14107), and backward dW/db offsets (~15109-15110). Post-2c.3a, indices 24/25 point at b_b1out (153 floats) and start of w_b2fc (4288 floats) — not w_bn (672 floats) / b_bn (16 floats). Forward + backward were self-consistent on wrong tensors; the actual w_bn/b_bn at documented indices 33/34 sat untouched as zeros for ~10 commits. Smoke kept passing because GRN's Linear_residual projection ran in parallel and provided a clean residual-only pathway around the corrupted bottleneck slot. 2. **Target net missing bottleneck path.** forward_target_raw passed raw next_states_buf (128-padded) to an encoder expecting [B, s1_input_dim=102]; it was reading [market|ofi|tlob|mtf] of next_states instead of [bn_market_proj|portfolio]. Fix: build tg_bn_concat_buf inline before forward_target_raw using tg_w_ptrs[33]/[34] (target's w_bn) on next_states_buf[market]; new buffer pair tg_bn_hidden_buf + tg_bn_concat_buf. 3. **DDQN argmax missing bottleneck path.** submit_forward_ops_ddqn had the same shape mismatch on the online-on-next_states pass. Fix: build on_next_bn_concat_buf using on_w_ptrs[33]/[34] (online weights, since DDQN argmax uses online net); new buffer pair on_next_bn_hidden_buf + on_next_bn_concat_buf. Three call sites of the existing bn_tanh_concat_kernel now: online-on- states (states + on_w_bn), target-on-next_states (next + tg_w_bn), and ddqn-online-on-next_states (next + on_w_bn). Each combination of weights × input states produces distinct features; sharing the kernel across distinct buffer pairs preserves the GPU-only cold-path contract. No fingerprint change (no ISV slot or param tensor added). Smoke validation (multi_fold_convergence, 700.43s, 3 folds x 5 epochs): fold 0 best Sharpe 6.53 at epoch 2 fold 1 best Sharpe 80.11 at epoch 4 fold 2 best Sharpe 66.72 at epoch 4 geom-mean: 39.27 (vs Task 3 20.03, 2c.3c.6 18.80) The Sharpe lift is consistent with the bottleneck Linear now seeing its actual weights and the target/DDQN nets seeing input-shape- consistent features for TD-target / argmax-action computations. +166 / -9 LOC all in gpu_dqn_trainer.rs. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
475101054d |
feat(dqn-v2): Plan 4 Task 1A — feature-group index ranges for VSN/attention consumers
Adds 12 new `SL_*_GROUP_BEGIN/END` macros to `state_layout.cuh` covering the 6 feature groups (market/ofi/tlob/mtf/portfolio/plan_isv), plus `SL_NUM_FEATURE_GROUPS=6` and `SL_MAX_FEATURE_GROUP_DIM=42` (= largest group dim, MARKET_DIM). Six new `static_assert`s anchor each group dim ≤ max and confirm contiguity / start-at-zero / end-at-padding invariants. Rust mirror in `crates/ml-core/src/state_layout.rs` exposes: - `SL_NUM_FEATURE_GROUPS` - `SL_MAX_FEATURE_GROUP_DIM` - `FEATURE_GROUP_RANGES: [(usize, usize); 6]` (half-open `[begin, end)`, `plan_isv` ends at `PADDING_START` — padding is not a group) - `FEATURE_GROUP_NAMES: [&str; 6]` Three `const _: () = assert!(...)` blocks validate (1) first range starts at 0, (2) last range ends at `PADDING_START`, (3) every adjacent pair satisfies `ranges[g].end == ranges[g+1].begin` (no gaps), (4) every group dim ≤ `SL_MAX_FEATURE_GROUP_DIM`. Group dims at this commit: market=42, ofi=32, tlob=16, mtf=16, portfolio=8, plan_isv=7 — total 121 = `PADDING_START`. Prerequisite for Plan 4 Task 1B (E.1 Variable Selection Network — pre-trunk per-group softmax-over-groups gating), Task 5 Mode B (per-group ISV diagnostics), and Task 4 follow-on (group-aware encoder interface). Additive only — ZERO production callers in this commit. No new module / kernel / ISV slot / param tensor / Orphan row. No fingerprint change (group ranges are derivable from existing `SL_*_START` constants and contribute no new structural-hash invariant). cargo check clean at 11 warnings (workspace baseline preserved); cargo build compiles all kernel cubins (state_layout.cuh edit triggers full kernel rebuild — no compile errors). Audit doc: 1 new entry under "Plan 4 Task 1A". Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
005ed3a4f9 |
feat(dqn-v2): Plan 4 Task 3 E.3 — IQN fixed-τ multi-quantile heads (5/25/50/75/95)
Replaces random τ ∈ U(0,1) sampling with `FIXED_TAUS = [0.05, 0.25, 0.50, 0.75, 0.95]`. Kernel-side `IQN_NUM_QUANTILES` macro 32 → 5; `GpuIqnConfig:: default().num_quantiles` 32 → 5. Construction-time τ broadcast (option B2) populates `online_taus` / `target_taus` / `cos_features` once via `clone_htod`; both online and target IQN forwards plus the CVaR cold path read this static buffer. The Philox-driven `iqn_sample_taus_kernel` deleted along with its only Rust consumer (in `compute_cvar_scales`); the `rng_step` Philox seed counter also gone. Action ranking in the IQN inference kernel switched from mean-over-quantiles to MEDIAN (`q_acc[a] = q_val` only when `t == IQN_MEDIAN_INDEX = 2`); the off-median positions feed four new ISV diagnostic slots. Four new ISV slots tail-appended: IQN_Q_P05_EMA_INDEX = 99 (mean |Q| at τ=0.05, EMA) IQN_Q_P25_EMA_INDEX = 100 (τ=0.25) IQN_Q_P75_EMA_INDEX = 101 (τ=0.75) IQN_Q_P95_EMA_INDEX = 102 (τ=0.95) Median (τ=0.50) intentionally skipped — already in greedy-Q diagnostic. Fingerprint pair shifted 97→103, 98→104; ISV_TOTAL_DIM 99→105. Layout fingerprint: 0x3e21acecd922e540 → 0x5789155b683ab59c. New kernel `iqn_quantile_ema_kernel.cu` (4-block × 256-thread shmem-reduce, no atomicAdd) reads `save_q_online [TBA, B*Q]` and EMA-updates the four slots. Launched from `training_loop.rs` per-step alongside `launch_h_s2_rms_ema`. StateResetRegistry extended with 4 FoldReset entries (cold-start 0.0). Hyperparam plumbing: `hyperparams.num_quantiles` and `DQNConfig::iqn_num_quantiles` pinned to `FIXED_TAUS.len()` at the `GpuIqnConfig` construction site in `fused_training.rs::new` and `trainer/constructor.rs`. Legacy fields stay for compat; production / hyperopt configs (dqn-production.toml, DQNHyperparameters defaults) aligned to 5. Adam state for IQN params auto-resizes via `m_buf`/`v_buf` sizing through `total_params + cublas_pad`. **Checkpoint break** — IQN head parameter shapes change with `num_quantiles`; new fingerprint hash fails-fast at constructor load on pre-Task-3 checkpoints. Smoke tests: - New `iqn_multi_quantile_heads_produce_monotonic_estimates` (1.23s on RTX 3050 Ti): asserts ISV[99..103) finite + non-zero + spread > 1e-6 after 1 epoch — PASS (Q_p05=0.0187 Q_p25=0.0200 Q_p75=0.0193 Q_p95=0.0190). - `multi_fold_convergence` (606.50s, 3 folds × 5 epochs): all 3 fold checkpoints written; per-fold best train Sharpe -8.17 / 74.24 / 63.44 at epochs 2 / 4 / 2 (mean 43.17 vs 2c.3c.6 baseline mean 23.43 — folds 1+2 substantially up, fold 0 down -16 points; absolute-mean comfortably above the plan's 3.8 floor). No NaN/Inf, no panic. cargo check clean at 11 warnings (baseline preserved); cargo build compiles 61 cubins (was 60; +iqn_quantile_ema, -nothing — the old sample_taus kernel was inside iqn_dual_head_kernel.cu, not a separate cubin file). Files touched: 14 modified (`iqn_dual_head_kernel.cu`, `iqn_cvar_kernel.cu`, `gpu_iqn_head.rs`, `gpu_dqn_trainer.rs`, `build.rs`, `state_reset_registry .rs`, `training_loop.rs`, `constructor.rs`, `fused_training.rs`, `config.rs`, `dqn-production.toml`, `smoke_tests/mod.rs`, `docs/dqn-wire-up-audit.md`, `dqn-production.toml`) + 2 new (`iqn_quantile _ema_kernel.cu`, `smoke_tests/iqn_quantile_monotonicity.rs`). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
eaa65ac241 |
feat(dqn-v2): Plan 4 Task 2c.3c.6 — wire mag_concat_qdir adaptive scale via ISV[96]
Replaces the stale `eq / fmaxf(dz, 1e-6f)` post-ReLU scale calibration in
`mag_concat_qdir` with adaptive RMS-match using ISV[H_S2_RMS_EMA_INDEX=96]
(populated each step by the producer kernel landed in 2c.3c.5).
Kernel signature gains two trailing args:
const float* __restrict__ isv,
int isv_h_s2_rms_index
Per-sample formula:
q_rms = sqrt(sum_a(eq_a^2) / b0_size)
scale = (q_rms > 1e-6) ? (h_s2_rms_ema / q_rms) : (1 / max(dz, 1e-6))
concat_out[…, SH2+a] = eq_a * scale
Q_dir's per-sample RMS now strictly tracks the trunk-output RMS regardless
of GRN drift across training. The legacy `~ 1-10` calibration (carried
over from the pre-GRN post-ReLU trunk) is removed; the comment at the
old normalization site is deleted, kernel docstring updated. The fallback
to `1 / max(dz, 1e-6)` activates only when q_rms ≈ 0 (uniform Q across
actions) — a domain-mathematical encoding, not a stub return. A static
`MAG_CONCAT_MAX_DIR=4` register-array bound matches the project's
4-direction (S/H/L/F) layout invariant; `b0_size` stays a runtime arg
for signature stability and production callers always pass 4.
Launch site `launch_mag_concat_from` extended with `isv_signals_dev_ptr`
+ `H_S2_RMS_EMA_INDEX as i32` args. `debug_assert!` mirrors 2c.3c.5's
invariant on the ISV device pointer.
Backward path unchanged: `strided_accumulate` extracts `d_h_s2` from the
first SH2 columns of `d_mag_concat` as before; `h_s2_rms_ema` and `q_rms`
are treated as fixed scalars at this batch's launch (same convention as
`dz`/`v_min` from `per_sample_support`).
Smoke (`cargo test … multi_fold_convergence --ignored --release`,
649.37s, 3 folds × 5 epochs):
fold 0 best train Sharpe 8.06 at epoch 5
fold 1 best train Sharpe 43.06 at epoch 1
fold 2 best train Sharpe 19.16 at epoch 2
geom-mean: 18.80 (vs 2c.3c.5: 20.03, -6.1%; well within 30% band)
All 3 dqn_fold{N}_best.safetensors checkpoints written. No NaN/Inf, no
fingerprint mismatch (fingerprint unchanged at 0x3e21acecd922e540). 0
panic gates added/removed. Closes the 2c.3c chain — H_S2_RMS_EMA producer
(2c.3c.5) + consumer (this commit) both wired.
cargo check clean at 11 warnings (baseline preserved); 81 cubins unchanged
(kernel-signature edit, no new .cu). +69/-6 LOC across experience_kernels.cu
and gpu_dqn_trainer.rs; audit doc appended (Invariant 7).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
f6da7cd8c4 |
docs(dqn-v2): Plan 4 Task 2c.3c.5 — backfill smoke results into audit doc
multi_fold_convergence on the as-landed 2c.3c.5 binary (commit
|
||
|
|
6135dea311 |
feat(dqn-v2): Plan 4 Task 2c.3c.5 — append H_S2_RMS_EMA ISV slot + producer kernel
ISV bus tail-append: - [96] H_S2_RMS_EMA_INDEX — per-batch RMS(save_h_s2) EMA (alpha=0.05) - [97] ISV_LAYOUT_FINGERPRINT_LO_INDEX (shifted from 94) - [98] ISV_LAYOUT_FINGERPRINT_HI_INDEX (shifted from 95) - ISV_TOTAL_DIM 96 -> 99 - New LAYOUT_FINGERPRINT_CURRENT = 0x3e21acecd922e540 (was 0xcf3a24b0a1f70057) Producer kernel `h_s2_rms_ema_kernel.cu`: - Single-block reduction (256 threads, shmem-tree, no atomicAdd) of the online trunk's `save_h_s2 [B, SH2]` post-GRN activation - RMS = sqrt(sum_sq / (B*SH2)); EMA into ISV[96] with alpha=0.05 - Launched once per training step alongside reward_component_ema and the other Plan 3/4 ISV producers in training_loop.rs Constructor cold-start writes ISV[96]=1.0 (neutral RMS). StateResetRegistry: H_S2_RMS_EMA registered as FoldReset -> 1.0. `launch_h_s2_rms_ema` and the kernel both treat the unconditional constructor allocation as an invariant — debug_assert! on the host side, no defensive (N>0) ternary in the kernel — same proper-resolution pattern as the 2c.3c.4 followup. Producer-only commit. 2c.3c.6 wires the consumer in mag_concat_qdir's adaptive-scale path so the magnitude-branch decoder sees a scale- invariant residual stack regardless of GRN drift across training. Smoke deferred — producer-only with zero consumers, runtime behaviour structurally unchanged from 2c.3c.4 (which validated the GRN backward chain end-to-end with all 3 fold checkpoints). cargo check + cargo build (new h_s2_rms_ema_kernel.cubin compiles clean under nvcc sm_80) is the standing validation; meaningful multi_fold_convergence re-run lands with 2c.3c.6 when the consumer wires up. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
89789c8f00 |
fix(dqn-v2): Plan 4 Task 2c.3c.4 followup — replace stub-return defensive guards with proper invariants + signal escalation
Three classes of cleanup, prompted by the pre-commit stub-return
heuristic surfacing in the 2c.3c.4 commit:
(1) **Unreachable defensive null/bounds-guards → debug_assert + delete**
`read_isv_signal_at`, `read_atom_utilization`, `compute_q_spectral_gap`
each guarded `is_null()` on `isv_signals_pinned` / `q_readback_pinned`,
yet both pointers are unconditionally allocated by the constructor
(`cuMemAllocHost_v2` with `assert_eq!` on success) and NEVER reassigned
after construction (greppable: zero `self.isv_signals_pinned =` /
`self.q_readback_pinned =` outside Drop). The guards were dead code
masquerading as safety. Removing the early-return removes the ambiguity
between "buffer never initialised" and "buffer holds a real-zero ISV
value" that downstream consumers of `read_isv_signal_at(IDX) == 0.0`
cannot otherwise distinguish.
`read_isv_signal_at`'s `index >= ISV_TOTAL_DIM` early-return is similarly
a stub: every caller passes a named constant from the ISV slot table
(`GAMMA_DIR_EFF_INDEX`, `TAU_EFF_INDEX`, `EPSILON_EFF_INDEX`, ...), so
bounds violation is a programming error not a runtime condition.
`compute_q_spectral_gap`'s `n_cols == 0` early-return is dead for the
same reason — `total_actions()` is the sum of branch sizes which the
trainer config requires non-zero.
All four converted to `debug_assert!(...)` so dev/test builds catch
invariant violations with a clear message; release builds inherit the
UB on the unsafe deref of a null pointer (loud crash) rather than the
silent stub return that previously made misconfiguration look fine.
(2) **Silent training-instability mask → warn-log + collapse sentinel**
`compute_q_spectral_gap` previously returned 1.0 (= "balanced/healthy"
in the calibration's eyes) on the first non-finite Q-value. NaN/Inf
Q-values indicate gradient corruption — masking them with the same
return value as a healthy network defeats HEALTH_DIAG's purpose. New
behaviour: scan all samples, mark the non-finite case, then emit a
`tracing::warn!` and return 100.0 (the same collapse sentinel emitted
when lambda_2 falls below the numerical floor). Operators see the
failure mode in logs rather than chasing a phantom balanced spectral
gap that contradicts NaN losses elsewhere in the same epoch.
(3) **Genuine domain encodings → annotated for the linter**
Three returns in `compute_q_spectral_gap` (lambda_2 < 1e-12, sigma2
< 1e-6) and `power_iteration_largest` (norm < 1e-20) ARE meaningful
mathematical degenerate-case answers, not stubs:
- lambda_2 ≈ 0 / sigma2 ≈ 0: deflated Gram is rank-1, sigma1/sigma2
diverges — calibrated collapse sentinel 100.0.
- norm ≈ 0 in power iteration: M·v vanishes, so the eigenvalue is 0.
Each got an explanatory comment + `// ok: domain encoding` to make
the intent legible to both readers and the heuristic linter.
`power_iteration_largest`'s `n == 0` early-return is unreachable
from the spectral_gap caller (debug_asserted upstream); removed in
favour of `debug_assert!(n > 0)` at the helper entry.
No behaviour change in production runs: the unreachable guards weren't
firing and the new invariants match the constructor's contract. The
non-finite Q-value path now surfaces what was previously silent — a
strict improvement.
cargo check clean at 11 warnings (baseline preserved).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
f94d857ebc |
feat(dqn-v2): Plan 4 Task 2c.3c.4 — wire GRN backward chain through 3 panic-gated trunk-backward sites
Removes 3 panic gates inserted by 2c.3a:
- BatchedBackward::backward_full
- gpu_dqn_trainer::apply_iqn_trunk_gradient
- gpu_dqn_trainer::apply_ensemble_diversity_backward
GRN backward chain (per block, mirrors encoder_forward_only):
d_y → backward_raw_phase1 (LN_dx + LN_dgdb + GLU_bwd)
→ cuBLAS Linear_b dW + dX
→ backward_raw_phase2 (ELU_bwd)
→ cuBLAS Linear_a dW + dX
→ for h_s1: Linear_residual dW (no_bias_lda) + dX accumulation
→ for h_s2: saxpy_inplace identity residual into d_h_s1
New helper: BatchedForward::encoder_backward_chain orchestrates both
GRN blocks. Used by main backward (writes to grad_buf), CQL backward
(writes to cql_grad_scratch), and the two auxiliary paths (write to
iqn_trunk_m then SAXPY into grad_buf).
apply_ensemble_diversity_backward additionally fixes value-head
indices: w_v1 4→13, w_v2 6→15 (legacy indices were post-2c.3a stale).
iqn_trunk_m grown from legacy 4-tensor size to 13-GRN-tensor padded
size (matches grad_buf layout for the trunk portion so the SAXPY
mix-in lands at the correct per-tensor offsets).
batched_backward::launch_dw_only_no_bias_lda added for Linear_residual
backward with padded states stride (Linear_residual has no bias, so
the bias-grad kernel cannot run with NULL db).
ReLU masks on h_s1 / h_s2 are gone — the GRN trunk ends in LayerNorm
which has no truncation derivative. Value-head FC retains its ReLU
mask (value head still uses ReLU activation).
launch_cublas_backward_to changed from &self to &mut self because the
GRN backward needs to scribble per-block partial-reduction scratch
inside grn_h_s{1,2}_online; the borrow split lets the trainer hand
&mut BatchedForward + &BatchedBackward to encoder_backward_chain.
Smoke validation (multi_fold_convergence, 3 folds × 5 epochs, 599.73s):
fold 0: best Sharpe 7.52 at epoch 2, val_Sharpe 0.51
fold 1: best Sharpe 60.94 at epoch 4, val_Sharpe 0.64
fold 2: best Sharpe 10.40 at epoch 2, val_Sharpe 0.82
All 3 dqn_fold{N}_best.safetensors checkpoints written. PF=5.29
on fold 2 epoch 5 with +988% return — gradients flow cleanly through
the GRN composition without NaN/Inf.
Audit doc updated (Invariant 7). // ok: suppressions added to 5
defensive null-guard sites in read_isv_signal_at /
read_atom_utilization / compute_q_spectral_gap (pre-existing,
documented in fn doc comments).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
3db8c77241 |
feat(dqn-v2): Plan 4 Task 2c.3c.3 — allocate 8 GRN backward scratch buffers
Plan 4 Task 2c.3c.3. Additive only — buffers will be consumed by
2c.3c.4's apply_grn_trunk_backward_raw helper.
8 device scratch buffers (4 per GRN block × 2 blocks):
- grn_h_s{1,2}_d_pre_ln — [B, hidden] aliases d_glu_out, d_residual
- grn_h_s{1,2}_d_linear_b_out — [B, 2*hidden]
- grn_h_s{1,2}_d_elu_out — [B, hidden]
- grn_h_s{1,2}_d_linear_a_out — [B, hidden]
Total ~B × 5 × hidden f32 per block × 2 blocks ≈ B × 5120 floats
(20 KB per sample at SH1=SH2=256). Negligible against existing
~32 MB per-branch cuBLAS workspaces.
Allocated in the trainer's alloc_backward_scratch path alongside
existing per-step backward scratch. dead_code allow on each field
until 2c.3c.4 wires them in.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
63183bb6a4 |
feat(dqn-v2): Plan 4 Task 2c.3c.2 — backward infra additions (launch_dw_only_no_bias + saxpy_inplace)
Plan 4 Task 2c.3c.2. Additive only — no production callers (2c.3c.4 wires them). Two infrastructure additions for the GRN trunk backward chain: 1. launch_dw_only_no_bias on CublasBackwardSet: variant of launch_dw_only that skips the bias-grad kernel call. Linear_residual in h_s1 GRN block has no bias, so calling launch_dw_only with db=0u64 would segfault the bias-grad kernel. 2. saxpy_inplace on CublasGemmSet: y += alpha * x for element-wise gradient accumulation. h_s2 GRN's identity residual needs d_h_s1 += d_pre_ln_h_s2 after Linear_a_h_s2's backward overwrites d_h_s1 with d_x = d_linear_a @ W_a. Implementation reuses the existing dqn_saxpy_f32_kernel (already used by the experience collector's IQR/ensemble-variance Q-bonus paths) — no new kernel, no cuBLAS legacy-handle stream-binding work, kernel handle loaded once at CublasGemmSet::new from DQN_UTILITY_CUBIN. Both methods sit dead-code until 2c.3c.4's wire-up commit. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
8fc41e3398 |
feat(dqn-v2): Plan 4 Task 2c.3c.1 — GrnBlock backward phase split (additive)
Plan 4 Task 2c.3c.1. Additive only — no production callers (2c.3c.4 wires them). Mirrors 2c.3b's forward phase split for the same composition reason: cuBLAS Linear_b backward must run BETWEEN GLU backward and ELU backward. New methods on GrnBlock: - backward_raw_phase1: LN_bwd + LN_dgamma_dbeta_p1 + LN_dgamma_dbeta_p2 + GLU_bwd. Produces d_linear_b_out for caller's Linear_b backward, plus d_pre_LN (= d_residual = d_glu_out via implicit residual split). - backward_raw_phase2: ELU_bwd. Takes d_elu_out from caller's Linear_b backward, produces d_linear_a_out for caller's Linear_a backward. The existing monolithic backward() is left as-is for surface-area minimization; will be deleted with 2c.5 cleanup if no production callers emerge. No CUDA kernel changes. Pure Rust dispatcher split. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
ad1e5c5e2d |
feat(dqn-v2): Plan 4 Task 2c.3b — encoder forward swap to GRN
Plan 4 Task 2c.3b. Forward-path runtime restored. Backward path stays
gated by 2c.3a panic markers until 2c.3c swaps it to GRN backward.
Changes:
- CublasGemmSet now owns four GrnBlock instances (online + target,
h_s1 + h_s2) and 14 GRN scratch buffers. Constructor allocates them
alongside the existing branch streams + workspaces.
- New `forward_raw_phase1` + `forward_raw_phase2` methods on GrnBlock
split the existing forward kernel sequence at the cuBLAS Linear_b
boundary so the encoder can dispatch
Linear_a → ELU → Linear_b → Linear_residual → GLU+LN in order.
No new GRN kernels.
- Swap encoder_forward_only to GRN forward composition: cuBLAS Linear_a
+ bias → forward_raw_phase1 (in-place ELU + DtoD save) → cuBLAS
Linear_b + bias → cuBLAS Linear_residual (h_s1 only; h_s2 routes
h_s1_ptr as identity residual) → forward_raw_phase2 (DtoD save +
GLU + residual+LN).
- New target_encoder_forward_only mirrors the online split using
dedicated grn_h_s1_target / grn_h_s2_target instances; called with
save_for_backward=false because the GRN backward only runs against
the online network (target is inference-only).
- Remove panic gates from encoder_forward_only, forward_target_raw,
forward_online_f32. forward_target_raw and forward_online_f32 now
delegate trunk encoding to the new (target_)encoder_forward_only.
- All forward methods on CublasGemmSet migrated from `&self` to
`&mut self`; trainer call sites refactored to direct
`self.cublas_forward.method()` so disjoint-borrow rules let
`self.launch_*` helpers execute alongside.
- Delete orphaned iqn_trunk_forward_kernel + IqnHead::trunk_forward_kernel
field + IqnKernels.trunk_fwd field + load("iqn_trunk_forward_kernel")
call. Replace the cuBLAS fallback in
gpu_iqn_head.rs::execute_training_pipeline with a hard error: IQN
now requires set_cached_target_h_s2 to be called with the buffer
written by BatchedForward::target_encoder_forward_only, so online
and target share ONE trunk implementation. Dead cuBLAS scaffolding
(trunk_l1_gemm/trunk_l2_gemm/trunk_h1_scratch/launch_trunk_bias_relu)
marked #[allow(dead_code)] to keep the diff compact.
- Audit doc updated with Task 2c.3b row.
Backward gates intentionally kept on backward_full,
apply_iqn_trunk_gradient, apply_ensemble_diversity_backward — they
panic loudly until 2c.3c swaps the relu_mask trunk backward to the
GRN backward chain.
Smoke (`cargo test … multi_fold_convergence --ignored`): forward path
runs cleanly through online + target encoder forwards (logs 4×
`GrnBlock initialised` confirming both online + DDQN sets allocate
online + target instances). Smoke then panics inside backward_full,
which is the retained 2c.3a gate. Smoke completion is therefore
deferred to 2c.3c per the spec's "KEEP backward gates" non-negotiable.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
f7697851f7 |
feat(dqn-v2): Plan 4 Task 2c.3a — GRN trunk param-tensor reshuffle (runtime-broken intentionally)
Plan 4 Task 2c.3a. Build-clean prep commit. Lays down the GRN param-tensor layout (86 -> 95 tensors) and migrates all 93 padded_byte_offset call sites + ancillary index references in lockstep. Does NOT swap the trunk forward — that's Task 2c.3b. Explicit panics at every trunk-forward and trunk-backward caller ensure any accidental runtime execution fails loudly rather than producing silent garbage. Param-tensor reshuffle: - Delete: W_S1, b_S1, W_S2, b_S2 (old trunk's 4 Linear tensors) - Insert: 7 h_s1 GRN tensors (W_a, b_a, W_b, b_b, W_residual, gamma, beta) - Insert: 6 h_s2 GRN tensors (no W_residual — SH1 == SH2 makes residual identity) - All tensor indices >= 4 in OLD layout shift +9 in NEW layout - NUM_WEIGHT_TENSORS 86->95; FIRST_ISV_TENSOR 68->77 layout_fingerprint_seed() updated. New LAYOUT_FINGERPRINT_CURRENT: 0xcf3a24b0a1f70057 (was 0xa504d3c2f275b8af). Xavier init paths added for the 13 new tensors: - W matrices (Linear_a, Linear_b, Linear_residual): Xavier - LayerNorm gamma_h_s1, gamma_h_s2: 1.0 - LayerNorm beta_h_s1, beta_h_s2: 0.0 - All biases (b_a_h_s1, b_b_h_s1, b_a_h_s2, b_b_h_s2): 0.0 Encoder gates: BatchedForward::encoder_forward_only, BatchedForward::forward_target_raw, BatchedForward::forward_online_f32, BatchedBackward::backward_full, GpuDqnTrainer::apply_iqn_trunk_gradient, GpuDqnTrainer::apply_ensemble_diversity_backward all panic at function entry. Original bodies preserved (with #[allow(unreachable_code, unused_variables)]) for Task 2c.3b/c to swap GRN forward + backward in place. Spectral-norm descriptor (13 matrices): slots [0]/[1] re-mapped to GRN's w_a_h_s1/w_a_h_s2 — shapes match legacy W_s1/W_s2 exactly, so the existing spectral-norm constraint transfers cleanly to the GRN's first Linear. Linear_b / Linear_residual not yet covered (Task 2c.3c). Smoke intentionally NOT run. Build-clean is the validation. Task 2c.3b will swap the encoder forward (gpu_grn::GrnBlock::forward); Task 2c.3c will swap the trunk backward (gpu_grn::GrnBlock::backward) and run smoke. Validation: - cargo check --workspace: 11 baseline warnings (unchanged) - cargo build -p ml --lib: all 58 cubins compile clean - 93 padded_byte_offset call sites all migrated; spectral_norm descriptor + branch_w_base + decoder_forward_only + value_head indices all updated in lockstep per feedback_no_partial_refactor.md Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
4402dbaf3d |
audit(dqn-v2): h_s2 audit anchor corrections (post-2c.3+4 reconnaissance)
The 2c.3+4 dispatch agent surfaced two stale anchors in the 2026-04-25 h_s2 consumer audit: 1. Row 11 (IQN target trunk): the kernel referenced as `iqn_compute_target_h_s2` in `iqn_dual_head_kernel.cu:1031` is actually `iqn_trunk_forward_kernel`, and it is ALREADY ORPHANED (loaded into IqnHead::trunk_forward_kernel but never invoked). The real IQN target trunk runs through cuBLAS iqn_lt_matmul calls in gpu_iqn_head.rs::execute_training_pipeline:820-846 (with a cached fast-path at 803-847). Mitigation reduced to: delete dead kernel + extract target_encoder_forward_only + replace cuBLAS fallback path with that extraction. 2. Row 22 (relu_mask sites): line numbers drifted. Audit said 5581/5798; current code is 5655/5697/5872. Crucially, 5697 is the h_s1 mask, not h_s2/IQN-aux — the audit's "3 sites" claim needs verification per site before editing. Updated to the current line numbers with a "verify before editing" note. These are documentation-only corrections. The mitigation strategy (GRN backward chain replaces relu_mask; target_encoder_forward_only unifies trunk implementation) is unchanged in spirit. |
||
|
|
b7f941f7fa |
feat(dqn-v2): Plan 4 Task 2c.2 — gpu_grn.rs Rust wrapper (additive, no callers)
Plan 4 Task 2c.2. Additive Rust wrapper around the kernels landed in
2c.1 (`grn_kernel.cu`, commit
|
||
|
|
68197c2c25 |
feat(dqn-v2): Plan 4 Task 2c.1 — GRN kernel module (forward + backward, additive)
Plan 4 Task 2c.1. Additive module — kernels compile and load via the existing kernel-loading infrastructure but have NO production callers in this commit. Task 2c.3+4 wires them into the trunk encoder. Eight kernels in grn_kernel.cu (Linear_a/Linear_b/Linear_residual are cuBLAS GEMMs, not new kernels): - grn_elu_inplace: element-wise ELU (canonical α=1.0) - grn_glu_forward: GLU split + sigmoid (saves sigmoid for backward) - grn_residual_layernorm_forward: residual add + LN, saves mean/rstd/normed - grn_layernorm_backward_dx: FULL JACOBIAN (not the simplified attn_layer_norm_bwd_dx which would silently propagate approximation into trunk gradients) - grn_layernorm_backward_dgamma_dbeta_p1: per-block partial reduction (no atomicAdd per feedback_no_atomicadd.md) - grn_layernorm_backward_dgamma_dbeta_p2: final reduce across blocks - grn_glu_backward - grn_elu_backward LN backward formula (full Jacobian per pearl_cold_path): d_out_g = d_out * gamma s1 = sum_d(d_out_g) s2 = sum_d(d_out_g * normed) d_x = (1/H) * rstd * (H * d_out_g - s1 - normed * s2) Layout convention: row-major [B, H] (sample-major, matches spec §4.E.2 and dt_layernorm_kernel) — intentionally different from attention_kernel.cu's [D, B] col-major; the trunk encoder downstream of GRN uses row-major buffers, so per-row LN avoids a transpose. build.rs registers grn_kernel.cu (kernel count: 57 → 58). Verified nvcc compiles cleanly via `cargo build -p ml --lib`. Module is dead code until 2c.3+4. Wire-up audit updated with the additive entry per Invariant 7. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
e9100073b9 |
plan(dqn-v2): Plan 4 Task 2c GRN — further decomposed into 2c.1/2c.2/2c.3+4/2c.5
A first dispatch of monolithic Task 2c was refused with a second scope
assessment that surfaced three architectural facts beyond what 2a/2b
had captured:
(d) attn_layer_norm_bwd_dx is a SIMPLIFIED element-wise approximation
(d_x = d_out * gamma / std), not the full LN Jacobian. Reusing
it in the GRN backward would propagate the approximation into
trunk gradients silently. Task 2c.1 must write a new full
Jacobian: (1/D) * rstd * (D * d_out_g - sum_d(d_out_g) -
normed * sum_d(d_out_g * normed)).
(e) IQN target-trunk migration is not a clean swap. iqn_compute_
target_h_s2 reproduces the legacy LeakyReLU trunk independently
in CUDA. Deleting it requires a target_encoder_forward_only
extraction on the target path (Task 4-style refactor on the
target side). Sub-task in itself.
(f) Three relu_mask removal sites have three different save-buffer
lifetimes (online trunk per-step, IQN-aux per-step in different
trainer, target-sync per-target-update). Three plumbing tasks,
not one.
Decomposition:
- 2c.1: grn_kernel.cu with full LN Jacobian backward + GLU + ELU +
residual-add. Module compiles standalone (additive, dead code).
- 2c.2: gpu_grn.rs Rust wrapper with 5 save-for-backward state
buffers per block (not 3 as the prior plan suggested).
- 2c.3+4: ATOMIC commit — compute_param_sizes 86→95 + fingerprint
rewrite + xavier init for 13 new tensors + all 98 padded_byte_
offset migrations + 3 relu_mask sites with different save-buffer
plumbing each + iqn target trunk delete + target_encoder_forward_
only + mag_concat RMS-match adaptive ISV.
- 2c.5: docs flip.
No code changes. Plan-doc only.
|
||
|
|
c55c24699a |
feat(dqn-v2): Plan 4 Task 2b — extend layout fingerprint to cover param-tensor layout
Prerequisite for Plan 4 Task 2c (GRN ADOPT). The current layout_fingerprint_seed() only covers ISV slot names + indices; param-tensor layout shifts (which GRN insertion will cause) pass silently through checkpoint load. Extends the seed string to include all 86 param-tensor positions by canonical name. Any structural reshuffle (insert/delete/rename a tensor) now triggers a different fingerprint and fail-fast at checkpoint load. Pragmatic scope (Option A): tensor names + positions, not sizes. Sizes depend on runtime config (shared_h1, shared_h2, etc.) and can't be embedded in a const fn. Size mismatches between checkpoint and current binary are caught by safetensors deserialization separately. Task 2c's GRN insertion is structural (new tensor names at new positions) — Option A suffices. This commit IS a checkpoint break: every existing checkpoint's fingerprint matches the old seed and fails load. Behavior change zero; cold-start smoke passes at baseline Sharpe range (fold-2 best Sharpe = 96.65). New fingerprint value: 0xa504d3c2f275b8af. Pearl-aligned: complete fingerprint coverage for what it claims to protect. Partial coverage was worse than no coverage because it implied safety where there wasn't. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |