9fb980da2bf71e1e847fbc2d199f61771e1f560c
1074 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
9fb980da2b |
fixup(class-a-audit-batch-4b): invert MIN_HOLD_TEMPERATURE dir_acc → temp mapping
Post-commit kernel-formula audit of `compute_min_hold_penalty` at
trade_physics.cuh:567-577 showed the original `temp = TEMP_MIN +
(TEMP_MAX - TEMP_MIN) × skill` mapping was BACKWARDS for the
WR-plateau scenario this 8-commit chain targets.
Formula recap:
soft_factor = deficit / (deficit + T)
HIGH T → soft_factor → 0 → forgiving (no exit penalty)
LOW T → soft_factor → 1 → sharp (max exit penalty)
The deleted schedule's `START=50 → END=5` therefore meant
"permissive early, strict late" (classic explore-then-exploit), not
"permissive when committed, strict when uncertain" as the
substituted mapping assumed.
WR-plateau scenario: dir_acc ≈ 0.46-0.48 (model committed but
wrong) — the model needs a permissive (HIGH T) exit ramp to bail
out of bad committed directions, not maximum exit penalty (LOW T)
locking it into them.
Inverted the mapping using `confusion = 1 - skill`:
- dir_acc ≈ 0.5 (random / plateau) → confusion=1 → temp=50
(permissive, plateau exit ramp)
- dir_acc → 1.0 (saturated skill) → confusion=0 → temp=5
(strict, force informed commitment)
This preserves the deleted schedule's epoch-0 anchor (T=50
cold-start) while reacting to actual realized skill instead of an
epoch-time anchor.
Files touched (atomic per `feedback_no_partial_refactor`):
- min_hold_temperature_update_kernel.cu — formula + header doctring.
- sp14_isv_slots.rs — slot-doc comment expanded with new mapping
+ fixup-rationale paragraph.
- gpu_aux_trunk.rs — `MinHoldTemperatureUpdateOps` docstring.
- gpu_dqn_trainer.rs — cubin docstring + ISV_TOTAL_DIM doc-comment.
- tests/sp14_oracle_tests.rs — Tests 3 + 4 flipped (high dir_acc
now expects temp=TEMP_MIN; low dir_acc now expects TEMP_MAX);
Tests 1, 2, 5 numerically unchanged (sentinel guard fires before
the mapping; midpoint dir_acc=0.75 has skill=confusion=0.5 so
pre/post-fixup numerics match).
- docs/dqn-wire-up-audit.md — Item 4 entry rewritten with the
fixup rationale + mapping table updated.
Verification:
cargo check -p ml --tests --all-targets PASS
cargo test sp14_audit_4b (lib) 4/4 PASS (slot layout
locks unchanged)
GPU oracle re-run deferred to next L40S smoke (kernel was rebuilt
in-place so cubin contents change; the 5 oracles encode the new
mapping and will exercise the new cubin).
Per `pearl_first_observation_bootstrap.md` (sentinel→target
replace; sentinel anchors unchanged) +
`pearl_symmetric_clamp_audit.md` (bilateral clamp on target_temp
preserved) + `feedback_no_partial_refactor.md`.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
0b9ea77dc4 |
fix(class-a-audit-batch-4b): plan_threshold floor adaptive + MIN_HOLD_TEMPERATURE ISV-driven
Per Class A audit-fix Batch 4-B (final 2 of 4 deferred items from
P1-wiring/P1-producer). Completes the 8-commit WR-plateau intervention
chain. Validation deferred to next L40S smoke.
Item 3: plan_threshold adaptive floor (Design Y - inline producer)
- NEW slot PLAN_THRESHOLD_FLOOR_ADAPTIVE_INDEX=459
- Writes slow-EMA shadow of `0.5 * readiness_ema` from inside the
existing update kernel (no new file, no new launch).
- Pearl-A first-observation bootstrap (sentinel 0.1 matches pre-fix
hardcoded value for bit-identical cold-start) + Welford alpha=0.005
slow EMA.
- Bilateral clamp [0.05, 0.50] (probability units) per
pearl_symmetric_clamp_audit.
- Consumer reads isv[459] as the floor in the same launch's final
fmaxf; cold-start sentinel REPLACES with threshold_target so the
pre-fix `fmaxf(0.1, 0.5*ema)` semantic is preserved bit-identical
for any readiness EMA above 0.20.
Item 4: MIN_HOLD_TEMPERATURE -> ISV-driven (driving signal: dir_acc skill)
- NEW slot MIN_HOLD_TEMPERATURE_ADAPTIVE_INDEX=460
- NEW kernel min_hold_temperature_update_kernel.cu (single-thread
cold-path, per-epoch boundary launch).
- Driving signal: dir_acc skill = clamp((short_ema - 0.5)/0.5, 0, 1)
from ISV[AUX_DIR_ACC_SHORT_EMA_INDEX=373]. When committing skillfully
(high dir_acc) -> temp HIGH (permissive). When at random baseline
(~0.5) -> temp LOW (sharp commitment pressure). Substituted for
the audit-spec's `dir_entropy_deficit` because no dir_entropy ISV
slot exists - dir_acc skill is the closest semantically-equivalent
signal that preserves the spec intent.
- Pearl-A bootstrap (sentinel 50.0 matches the deleted
MIN_HOLD_TEMPERATURE_START=50 anchor) + alpha=0.05 mid-cadence EMA.
- Bounds [5, 50] (matches the deleted schedule range).
- Decouples temperature from epoch number - the old schedule pinned
LOW temp (sharp) at end of training, exactly when a WR-plateaued
model needed forgiveness to escape.
- DELETED: state_layout.cuh::MIN_HOLD_TEMPERATURE_{START, END, DECAY}
#defines + training_loop.rs::min_hold_temperature_for_epoch helper
function (kept docstring tombstone explaining the deletion). Both
call sites migrated to the new ISV reader. Per
feedback_no_legacy_aliases + feedback_no_partial_refactor.
ISV_TOTAL_DIM: 459 -> 461.
Cumulative WR-plateau fix series (final commit, #8):
-
|
||
|
|
7e9a8f6ef1 |
fix(class-a-audit-batch-4a): DD saturation floor adaptive + legacy DD path Case A
Per Class A audit-fix Batch 4-A (deferred from P1-wiring/P1-producer due to
audit-doc errors). Fixes 2 of 4 deferred items; Batch 4-B handles plan_threshold
floor + MIN_HOLD_TEMPERATURE in a separate commit.
Item 1: DD saturation floor (the upper end of the DD ramp at trade_physics.cuh:154
in apply_margin_cap, NOT line 548 as the audit doc claimed — that line is a
magnitude action constant; the actual saturation floor lives in apply_margin_cap)
- NEW slot DD_SATURATION_FLOOR_ADAPTIVE_INDEX=458
- Producer dd_saturation_floor_update_kernel.cu — p75(per-env DD_MAX) × 1.5
via Welford `mean + Z_75 × sigma` estimator with `max(p75, mean)` robustness
guard, mirrors P0-A REWARD_POS_CAP producer pattern (Pearl-A bootstrap +
Welford α=0.01)
- Cold-start fallback: 0.25f (DD_SATURATION_FLOOR_DEFAULT in state_layout.cuh)
- Bounds: [0.10, 0.50] (Category-1 dimensional safety)
- Distinct from SP15_DD_THRESHOLD_INDEX=421 (the SP15 quadratic DD-penalty
*trigger* threshold, a *lower* bound; this slot is the *upper* end of the
linear position-size scaling ramp dd_scale = max(0.05, 1.0 − dd_frac/floor))
- Threaded `isv_signals_ptr` into `apply_margin_cap` with NULL-tolerant
cold-start fallback to DD_SATURATION_FLOOR_DEFAULT
- 4 oracle tests (Pearl-A bootstrap, no-DD guard, bounds clamp, Welford EMA)
Item 2: Legacy compute_drawdown_penalty path → Case A (DELETED)
- Decision rationale: SP15's quadratic asymmetric DD penalty
(compute_sp15_final_reward_kernel.cu:154 via sp15_dd_penalty helper) runs
unconditionally as a post-modifier on the SP11-composed reward with
ISV-driven λ_dd (slot 420) and DD threshold (slot 421). Layering the legacy
linear-ramp penalty inside the SP11 composer on top of the SP15 quadratic
creates double-counting of DD shaping — exactly the code-smell the Class A
audit was designed to eliminate. Per `feedback_no_legacy_aliases.md` and
`feedback_no_partial_refactor.md`.
- Atomic deletion across:
- `compute_drawdown_penalty` device function (trade_physics.cuh)
- Single call site at experience_kernels.cu:3822
- `dd_threshold` and `w_dd` kernel arguments
- `w_dd` Rust config field (gpu_experience_collector.rs +
trainers/dqn/config.rs DQNHyperparameters)
- `w_dd` profile section + dispatch (training_profile.rs RewardSection,
OptimizableParameterRanges, FixedRewardParameters, ParamLookup
dispatch, profile→hyperparam mapping, test assertion)
- `w_dd *= rki` risk-intensity multiplier (config.rs)
- `w_dd` TOML keys (dqn-hyperopt.toml × 2, dqn-localdev.toml,
dqn-production.toml, dqn-smoketest.toml)
- Stale doc comments on hyperopt/adapters/dqn.rs + config.rs
risk_intensity field
- `config.dd_threshold` SURVIVES (still consumed by `launch_sp15_dd_state`
as the dd_budget for DD_PCT scaling). Documented in field comment.
ISV_TOTAL_DIM: 458 → 459 (Item 1 adds 1 slot; Item 2 is pure deletion)
Cumulative WR-plateau fix series (this is commit 7):
- Class C bug 1 + P0-B (
|
||
|
|
87d597d5d7 |
fix(class-a-p1-producer): adaptive Bayesian Kelly priors per-fold-end
Replace 4 hardcoded Bayesian priors in Kelly cap calculations
(prior_wins=2.0, prior_losses=2.0, prior_sum_wins=0.01,
prior_sum_losses=0.01) with ISV-driven slow-EMA values fed by a new
producer kernel that aggregates realized PS_KELLY_* fields across envs
from the same portfolio_state buffer kelly_cap_update_kernel reads
from at the same per-epoch boundary.
Slots [454..458):
- KELLY_PRIOR_WINS, KELLY_PRIOR_LOSSES,
KELLY_PRIOR_SUM_WINS, KELLY_PRIOR_SUM_LOSSES
ISV_TOTAL_DIM bumps 454 -> 458.
Producer:
- kelly_bayesian_priors_update_kernel.cu (per-fold-end / per-epoch
boundary, block-tree-reduce in shmem, no atomicAdd). Pearl-A
first-observation bootstrap (sentinels 2.0/2.0/0.01/0.01 match
pre-P1-Producer hardcoded values for bit-identical cold-start) +
alpha=0.005 slow EMA. Bounds counts in [0.5, 100], sums in [0.001,
1.0] per feedback_isv_for_adaptive_bounds. Launched RIGHT BEFORE
launch_kelly_cap_update so that kernel sees the freshly-blended
priors.
Consumers migrated atomically (same commit):
- kelly_cap_update_kernel.cu:39-42 -> ISV[454..458) with cold-start
fallback via kelly_prior_or_default helper (range guard).
- trade_physics.cuh::kelly_position_cap:304-307 -> NULL-tolerant
isv_signals_ptr threaded through apply_kelly_cap (single caller in
unified_env_step_core line 898 already had the bus pointer; mirrors
the existing kelly_f_smooth / kelly_warmup_floor_sp9 patterns).
State reset:
- 4 FoldReset registry entries (sp14_p1_kelly_prior_*) +
4 dispatch arms in training_loop.rs::reset_named_state.
Oracle tests (sp14_oracle_tests.rs, GPU-gated #[ignore]):
- Pearl-A bootstrap (sentinel -> REPLACE with aggregated targets)
- No realized trades -> ISV preserved bit-exactly
- Bounds clamp on extreme aggregates
- Slow EMA blend after bootstrap
CPU tests passing:
- sp14_p1_kelly_prior_slot_layout_locked
- all_sp14_p1_slots_fit_within_isv_total_dim
- every_fold_and_soft_reset_entry_has_dispatch_arm (C.10 lesson)
- layout_fingerprint_bumps_after_sp14_wire
DEFERRED — Item 2 (MIN_HOLD_TEMPERATURE EMA): the audit-spec said
"MIN_HOLD_TEMPERATURE = 0.5f hardcoded somewhere" but the actual code
has MIN_HOLD_TEMPERATURE_{START=50.0f, END=5.0f, DECAY=20.0f} as a
PER-EPOCH ANNEALING SCHEDULE driven by min_hold_temperature_for_epoch
in training_loop.rs:68-73. The kernel takes T as a runtime scalar
specifically to enable Phase 2 ISV-driven lift "without recompiling
cubin" per the SP12 v3 design comment. The audit's claim of "0.5f
hardcoded" does not match reality; the right Phase 2 signal is a
separate spec decision and should not be guessed at per
feedback_no_quickfixes. Reporting back per the prompt's hard rule #8.
Per feedback_isv_for_adaptive_bounds + feedback_no_partial_refactor +
pearl_controller_anchors_isv_driven.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
657972a4b5 |
fix(class-a-p0a-downstream): DD penalty + MIN_HOLD_PENALTY_MAX scale to POS_CAP_ADAPTIVE
Per Class A audit P0-A downstream batch — both constants were tuned for fixed REWARD_POS_CAP=5.0f. P0-A made POS_CAP adaptive via isv[452]; this commit propagates the ratio to keep the Kahneman 2:1 asymmetry coherent across the reward shaping chain. Items: 1. DD penalty -5.0f → -1.0f * isv[REWARD_POS_CAP_ADAPTIVE_INDEX] (helper signature change in trade_physics.cuh::compute_drawdown_penalty adds dd_penalty_scale parameter; sole call site in experience_kernels.cu:3775 resolves the ISV value with the same defensive guard as sp15_apply_sp12_cap) 2. MIN_HOLD_PENALTY_MAX 3.0f → 0.6f * isv[REWARD_POS_CAP_ADAPTIVE_INDEX] (existing 60% ratio from state_layout.cuh comment line 252-253 preserved; resolved at the call site mirroring the effective_min_hold_target precedent for slot 451) Cold-start fallbacks preserved: - DD penalty: REWARD_POS_CAP=5.0f when ISV at sentinel/out-of-range - MIN_HOLD_PENALTY_MAX: kernel-passed 3.0f from gpu_experience_collector.rs:399 (bit-identical pre-P0-A behavior) Defensive guard at both consumer sites: ISV must be in [REWARD_POS_CAP_MIN_BOUND=1.0, REWARD_POS_CAP_MAX_BOUND=50.0] AND not within 1e-6f of SENTINEL_REWARD_POS_CAP=5.0f. Mirrors the existing sp15_apply_sp12_cap and segment-complete cap fallback patterns. Note: the SP15 quadratic DD penalty path (compute_sp15_final_reward_ kernel.cu::sp15_dd_penalty) is already fully ISV-driven via slots 420 (λ_dd) and 421 (dd_threshold) — only the legacy compute_drawdown_ penalty (linear ramp, slot-free) had the hardcoded -5.0f. The audit recommendation suggested ratio = 1.0 for MIN_HOLD_PENALTY_MAX assuming the value was 5.0f; actual is 3.0f and the existing tuning comment locks the ratio at 60% — pure wiring uses 0.6. Cumulative WR-plateau fix series: - Class C bug 1 + P0-B ( |
||
|
|
c4b6d6ef29 |
fix(class-a-p1-wiring): var_floor q_gap-only adaptive (1 of 4 wireable)
Per Class A audit P1 wiring batch — only 1 of 4 items was wireable as spec'd. The other 3 are deferred per the spec's "DEFER not BLOCK" rule because their target ISV slots either don't exist or are unit-incompatible with the consumer site. Items: 1. trade_physics.cuh:548 (DD floor_dd 0.25f) — DEFERRED. Slot 421 (DD_THRESHOLD_INDEX) holds the DD trigger threshold (~0.05), NOT the penalty saturation floor (~0.25). Different parameters; substituting would break the ramp formula `(dd-thr)/(floor-thr)`. Needs a new slot for the saturation floor. 2. gpu_experience_collector.rs:343-344 (dd_threshold + w_dd config) — DEFERRED. No W_DD slot exists in any sp*_isv_slots.rs. The legacy compute_drawdown_penalty path (experience_kernels.cu:3769) uses config scalars; the SP15 reward axis (compute_sp15_final_reward_kernel) already reads slot 421 directly. Wiring needs a new W_DD slot. 3. plan_threshold_update_kernel.cu:44 (plan threshold floor) — DEFERRED. Unit mismatch: ema and plan_threshold are probabilities ∈[0,1]; Q_DIR_ABS_REF (slot 21) is a Q-value magnitude EMA (5–50). Scaling a probability by a Q-magnitude is dimensionally meaningless. 4. experience_kernels.cu:2471 (var_floor formula) — DONE. Drop the hardcoded 0.25f lower bound. Was `fminf(fmaxf(q_gap*0.5f, 0.25f), 1.0f)`; now `fminf(q_gap*0.5f, 1.0f)`. NULL-q_gaps fallback uses ε=1e-6f for numerical safety (vs prior 0.25f). var_scale itself remains (0,1] from `1/(1+sqrt(var_q))` so the multiplicative chain stays bounded. Cumulative WR-plateau fix series: - Class C bug 1 + P0-B ( |
||
|
|
394de7d434 |
feat(class-a-p0a): REWARD_POS/NEG_CAP → ISV-driven adaptive caps from realized return distribution
Per Class A audit ranking, the highest-suspected-impact fix for the months-long WR-stuck-at-46-48% plateau across 11 superprojects. Hardcoded REWARD_POS_CAP=+5.0f / REWARD_NEG_CAP=-10.0f (state_layout.cuh:266-267) was structurally clipping the upper tail of realized alpha. Controller signals (sharpe EMA, var_q, q_gap) all derive from this CAPPED buffer, so the controller cannot select for trades it cannot see. Selectivity gradient evaporates — small wins clip to +5 alongside large wins also clipping to +5. The state_layout comment lines 261-264 explicitly deferred Phase 2 (ISV-driven) "IF Phase 1 validation reveals adaptive need". Phase 1 has been running 11 SPs without budging WR — adaptive need revealed. Architecture: - 2 new ISV slots [452..454): REWARD_POS_CAP_ADAPTIVE, REWARD_NEG_CAP_ADAPTIVE. - Producer kernel reward_cap_update_kernel.cu: block-tree-reduce Welford `mean + Z_99 × sigma` p99 estimator over winning realized returns + conservative `max(p99, max_win)` takeover, × 1.5 safety factor → POS cap. NEG cap = -2 × POS cap (preserves Kahneman 2:1 asymmetry per pearl_audit_unboundedness_for_implicit_asymmetry — asymmetry stays, but moved from hardcoded scalar to producer-time multiplier; single source of truth, no consumer applies the 2× ratio itself). - Pearl-A first-observation bootstrap from sentinel (5.0 / -10.0, matching pre-P0-A hardcoded values for bit-identical cold-start). Welford EMA α=0.01 thereafter (slow blend — reward distribution is the foundation of training and shouldn't move fast). - Bounds: POS in [1, 50], NEG in [-100, -2] (Category-1 dimensional safety per feedback_isv_for_adaptive_bounds, NOT tuning). - 3 consumer sites migrated atomically per feedback_no_partial_refactor: experience_kernels.cu:3112-3114 (segment_complete cap), compute_sp15_final_reward_kernel.cu:163 (Stage 4 helper invocation), sp15_reward_axis_helpers.cuh:211 (sp15_apply_sp12_cap device fn signature change to take isv ptr). - Cold-start fallback: when ISV slot at sentinel OR outside [REWARD_POS_CAP_MIN_BOUND=1, REWARD_POS_CAP_MAX_BOUND=50], consumers fall back to original macros (still defined in state_layout.cuh). - 2 new device-ptr accessors on the experience collector (step_ret_per_sample_dev_ptr, trade_close_per_sample_dev_ptr) — reuses existing per-sample buffers; no new buffer allocated. - Per-epoch boundary launch (cold path) in training_loop.rs alongside launch_aux_horizon_chain. - Reset registry entries + dispatch arms in reset_named_state per the C.10 lesson (missing dispatch causes runtime crash). - Layout fingerprint seed updated: ISV_TOTAL_DIM 452→454 + AUX_PRED_HORIZON_BARS=450 + AVG_WIN_HOLD_TIME_BARS=451 (previously missing from seed) + REWARD_POS_CAP_ADAPTIVE=452 + REWARD_NEG_CAP_ADAPTIVE=453. Per feedback_isv_for_adaptive_bounds: every adaptive bound in ISV. Verification: - cargo check -p ml --tests --all-targets: clean (19 pre-existing warnings, 0 new). - sp14_isv_slots tests: 8/8 pass (4 layout + 4 fits-within). - sp14_oracle_tests with --features cuda --ignored: 8/8 pass (4 existing q_disagreement/dir_concat + 4 new P0-A tests covering Pearl-A bootstrap, no-winning-trades preservation, bounds clamping to [1,50], Welford α=0.01 EMA blend). - sp15_phase1_oracle_tests with --features cuda --ignored: 36/36 pass (no regression from sp15_apply_sp12_cap signature change). Cumulative WR-plateau fix series: - Class C bug 1 ( |
||
|
|
316db416bb |
fix(class-a-p0c): MIN_HOLD_TARGET → ISV[AVG_WIN_HOLD_TIME_BARS_INDEX=451] (adaptive)
Per Class A audit: MIN_HOLD_TARGET=30.0f hardcoded was creating a
deterministic gradient pushing trades toward 30-bar holds regardless of
edge expiry. User's trading frequency is between HFT-MFT and varies by
regime; a 30-bar fixed target kills MFT-frequency alpha when the
optimal hold for current data is shorter (or longer).
The producer slot ISV[AVG_WIN_HOLD_TIME_BARS_INDEX=451] already exists
from SP14 Layer C Phase C.4b (commit
|
||
|
|
8f218cab24 |
fix(q-side): replay buffer intent→realized + Kelly warmup floor wiring (WR-plateau root causes)
Two independent bugs surfaced by Class C (frame-shift) + Class A (hardcoded bounds) audits, both implicated in the months-long WR-stuck-at-46-48% plateau across 11 superprojects. ## Bug 1: Replay buffer Sutton's deadly triad experience_kernels.cu:2275 was writing the original INTENT action to out_actions, but the reward in the replay buffer was computed from the REALIZED position (post-enforcement: Kelly cap, capital floor, trail-stop, broker cap can clamp Long→Flat). Replay buffer stored (s, intent, r_realized, s'). Q(s, Long) was therefore trained against r(s, Flat) whenever env clamped the intent. This explains the train_active_frac=0.40 vs val_active_frac=0.05 gap: train measures intent (40% Long/Short), eval measures realized (5% Long/Short). The 8× gap is env physics draining intent. Fix: after unified_env_step_core resolves actual_dir_core/actual_mag_core, overwrite out_actions[out_off] with the realized action (same encoding as backtest_env_kernel.cu:323-330, which has been doing it correctly all along). Order/urgency preserved from intent. ## Bug 2: Kelly cap update kernel ignored existing ISV warmup floor kelly_cap_update_kernel.cu:53 hardcoded the kelly_f floor at 0.0f. Cold path (per-epoch boundary). Per project_magnitude_eval_collapse_kelly_capped, this collapses kelly_cap to 0 → max position pinned to Quarter for cold start. The val-mag pathology. The warmup floor producer (ISV[KELLY_WARMUP_FLOOR_INDEX=330], SP9 Fix 37) was already populated and consumed by the per-step path at trade_physics.cuh:377-384, but this cold-path kernel never read it. Partial wiring. Fix: replace fmaxf(kelly_f, 0.0f) with fmaxf(kelly_f, isv[330]). One-line change. ## Predicted effect - train_active_frac and val_active_frac should converge (Bug 1 inflated train by counting overridden intents) - Magnitude distribution should escape Quarter-only (Bug 2 was pinning it) - WR ceiling at 46-48% may finally move (Bug 1 broke Bellman consistency; Bug 2 prevented edge realization) Falsification: 5-epoch L40S smoke. If unmoved by ep5, the plateau is deeper still (Class A P0-A REWARD_POS_CAP/NEG_CAP next). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
976f9b9807 |
fix(sp14-c.10): missing reset_named_state dispatch arm for sp14_q_disagreement_variance_ema
Phase C.1's atomic α deletion preserved Q_DISAGREEMENT_VARIANCE_EMA_INDEX=389
(HEALTH_DIAG diagnostic) and its state_reset_registry entry, but the
corresponding dispatch arm in reset_named_state was missing — even though
the inline comment at training_loop.rs:8024 falsely claimed it existed.
train-rqd8r (commit
|
||
|
|
10e647c141 |
test(sp14-c9): synthetic smoke for aux trunk gradient chain + C.8/C.9 audit close-out
C.8 (ISV-driven aux trunk Adam β1/β2/ε/LR/grad-clip) was already complete in C.5a
commit
|
||
|
|
0e61de408f |
feat(sp14-c6): h_s2_aux_rms_ema producer — ISV[449] per-collector-step
Single-block 256-thread CUDA kernel computing RMS(h_s2_aux [B, SH2]) and EMA-blending the step observation into ISV[H_S2_AUX_RMS_EMA_INDEX=449] directly. Pearl-A first-observation bootstrap embedded in kernel body (sentinel 0.0 → replace); fixed α=0.05 EMA blend thereafter. ISV slot 449 is outside the SP4/SP5 wiener buffer linear span so the scratch+apply_pearls_ad_kernel path is not available — self-contained Pearl-A logic mirrors the avg_win_hold_time_update_kernel precedent (slot 451). No atomicAdd; shmem block-tree-reduce only. Launched after aux_trunk_forward in the collector per-step hot path. - h_s2_aux_rms_ema_kernel.cu — new CUDA kernel (81 lines) - build.rs — cubin manifest entry - gpu_dqn_trainer.rs — H_S2_AUX_RMS_EMA_CUBIN static - gpu_aux_trunk.rs — HS2AuxRmsEmaOps struct + launch() - gpu_experience_collector.rs — field + constructor + hot-path launch - aux_trunk_oracle_tests.rs — h_s2_aux_rms_ema_pearl_a_bootstrap test - dqn-wire-up-audit.md — Phase C.6 audit entry cargo check -p ml --tests: clean (only pre-existing warnings) Oracle test: 1 new test added (requires GPU to run) Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
b26b189925 |
feat(sp14-c.5b): atomic contract migration — h_s2 → h_s2_aux + revert zero-fills
Atomic flip of the aux heads' input from Q's GRN trunk output `save_h_s2` to the SEPARATE aux trunk's output `h_s2_aux`. The aux trunk now trains its own w1/w2/w3/b1/b2/b3 from CE loss (next-bar + regime); Q's encoder is structurally protected by `aux_trunk_backward`'s missing `dx_in` output param (encoder boundary stop-grad enforced at the kernel-set level). Reverts the C.0 stop-grad band-aid commits (`872bd7392`, `411a30473`): the zero-fills in `aux_next_bar_backward` + `aux_regime_backward` Step 3 are replaced with the genuine SAXPY-back-to-input gradient (`dh_s2_aux[b,j] = sum_k sh_dh_pre[k] * w1[k,j]`). The leak that motivated stop-grad is now blocked structurally rather than by data zero-fill — aux gradient flows through the aux trunk's own params, never into Q's encoder. Wired in this commit (atomic, ~330 LOC): - 4 kernel signatures renamed `h_s2 → h_s2_aux` / `dh_s2_out → dh_s2_aux_out` (`aux_next_bar_forward`, `aux_regime_forward`, `aux_next_bar_backward`, `aux_regime_backward`); Rust wrappers in `gpu_aux_heads.rs` follow - Trainer fwd: insert `aux_trunk_forward_ops.launch(...)` in `aux_heads_forward` Step 0, populating `h_s2_aux` from `save_h_s1` (encoder layer-1 output, dim=shared_h1=256). Both head fwds redirect input pointer from `save_h_s2` to `h_s2_aux` - Trainer bwd: SAXPY both `aux_dh_s2_*_buf` into `dh_s2_aux_accum` (pre-zeroed each step via graph-safe `cuMemsetD32Async`); then `aux_trunk_backward_ops.launch(...)` propagates through w3/w2/w1 + b3/b2/b1; then `launch_aux_trunk_adam_update` applies global L2-norm clip + per-tensor Adam updates over 6 grad tensors - Collector fwd: insert `exp_aux_trunk_forward_ops.launch(...)` after `forward_online_f32`, reading `exp_h_s1_f32` and writing `exp_h_s2_aux`; redirect `exp_aux_heads_fwd.forward_next_bar` input from `exp_h_s2_f32` to `exp_h_s2_aux` - Pre-capture host-write of ISV-driven LR + grad-clip + step counter into mapped-pinned buffers in `launch_cublas_backward_to` (BEFORE `aux_heads_backward`); same `&mut self` pattern as `step_ofi_embed_adam` Verification: - `cargo check -p ml --tests` clean (1m02s, only pre-existing warnings) - `aux_trunk_oracle_tests` + `sp14_oracle_tests` 12/12 pass: - aux_trunk gradient check: max_rel_err=1.33e-2 (tol=2e-2) — matches C.4 baseline - aux_trunk_backward_does_not_write_dx: kernel source clean of dx_in/dx_in_out - aux_sign_label_lookahead_mask: 60/100 masked, 40/100 valid - 9 other oracle tests pass bit-identically Plan: docs/superpowers/plans/2026-05-07-sp14-layer-c-separate-aux-trunk.md §C.5b Audit: docs/dqn-wire-up-audit.md "SP14 Layer C Phase C.5b" section |
||
|
|
1edd71a2c1 |
feat(sp14-c.5a-fixup): missing scratch buffers + collector ptr scaffolding (dead code)
Phase C.5a-fixup — completes C.5a's additive infrastructure so C.5b can be
a true atomic contract flip with zero scaffolding work mixed in:
- Allocate dh_aux1_pre_scratch [B_max, AUX_TRUNK_H1=256] +
dh_aux2_pre_scratch [B_max, AUX_TRUNK_H2=128] (missed in C.5a — both
required by aux_trunk_backward.launch per gpu_aux_trunk.rs:266-267).
- Collector struct gains 6× u64 aux_trunk_{w1,b1,w2,b2,w3,b3}_ptr fields,
exp_aux_trunk_forward_ops: AuxTrunkForwardOps field (constructed in
ctor on collector's stream), and set_trainer_aux_trunk_param_ptrs
setter — mirrors the existing set_trainer_params_ptr zero-copy pattern.
- Trainer gains aux_trunk_param_ptrs() -> (u64×6) accessor returning
raw_ptr() for all 6 aux trunk parameter tensors.
- training_loop wires the new setter at both existing
set_trainer_params_ptr call sites (initial fused-ctx init + fold-boundary
re-init).
NO contract change: wire sites still call save_h_s2. The setter is called
and the 6 aux trunk param ptrs are populated, but no collector-side launch
reads them yet — C.5b atomically inserts aux_trunk_forward.launch(...)
post-forward_online_f32 and switches the aux head input pointer.
Graph-capture audit: aux_heads_backward IS INSIDE the captured `forward`
child graph (call chain: submit_forward_ops_main → launch_cublas_backward
→ launch_cublas_backward_to → aux_heads_backward; capture begins at
fused_training.rs:2964 / capture_child_graph). The existing function body
is fully device-side (zero host writes) — capture-safe by construction.
C.5b's new aux_trunk Adam launch is also fully device-side and will sit
inside the same captured region. The host writes for aux_trunk_t_pinned
must use the existing GPU-side increment_step_counters kernel chain
(submit_counters_ops, line 22799) — NOT host-side aux_trunk_adam_step
+= 1 inside capture. ISV-driven LR/clip writes happen pre-capture
(cold-path); the captured graph reads via aux_trunk_lr_dev_ptr /
aux_trunk_grad_clip_dev_ptr. This avoids the &self → &mut self ripple on
aux_heads_backward (gap 4 in C.5b implementer's blocker report). Full
wiring strategy + alternative (pre-capture host-write) documented in
docs/dqn-wire-up-audit.md C.5a-fixup section.
Verification:
- cargo check -p ml --tests --all-targets: clean (no new warnings).
- cargo test -p ml --test aux_trunk_oracle_tests --test sp14_oracle_tests
--release -- --ignored --nocapture: 12/12 pass (8 aux_trunk + 4 sp14;
bit-identical to C.5a baseline — pure scaffolding, no regression).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
c90de98594 |
feat(sp14-c.5a): allocate aux trunk fwd/bwd buffers + Adam launcher (dead code)
Phase C.5a — additive infrastructure for the aux trunk wire-up. Allocates saved-fwd buffers (h_s2_aux, h_aux1, h_aux2), gradient buffers (6× aux_trunk_*_grad), accumulator (dh_s2_aux_accum), and dedicated Adam launcher (launch_aux_trunk_adam_update) reading β1/β2/ε/LR/grad-clip from ISV[444..449). No contract change. No call sites for the new launcher yet. C.5b atomically wires these in. Phase C.5 was split (authorized 2026-05-08) after the original implementer flagged ~600-800 LOC scope across 4 files with correctness windows. C.5a is purely additive; C.5b is the genuine ~300 LOC atomic migration. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
3b71d21834 |
feat(sp14-c): aux prediction horizon ISV-driven (multi-bar pivot)
Aux's original label was (p_{t+1} > p_t) — pure HFT-scale microstructure
noise that's unlearnable at our HFT-MFT trading frequency. Migrated to
(p_{t+H} > p_t) where H is read from ISV[AUX_PRED_HORIZON_BARS_INDEX=450].
Adaptive producer drives H from observed avg winning hold time:
- Pearl-A first-observation bootstrap: replace sentinel H=60 directly
on first valid observation
- Steady-state Wiener-α EMA blend, slow (α=0.01) for stable horizon
(no target-variance EMA available, fallback per
pearl_wiener_optimal_adaptive_alpha)
- "No winning trades yet" guard keeps sentinel until first valid observation
Lookahead truncation: labels at t where t+H >= total_bars are masked
(sentinel -1, loss-reduce skips). The existing aux_next_bar_loss_reduce
in aux_heads_kernel.cu already supports the -1 mask convention via the
B_valid count — no new valid_mask parameter needed.
Step 5b finding: Case B — existing per-sample buffers
(hold_at_exit_per_sample, trade_profitable_per_sample) populated by
unified_env_step_core, but no aggregate ISV slot. Added new aggregator
slot AVG_WIN_HOLD_TIME_BARS_INDEX=451 + new producer kernel
avg_win_hold_time_update_kernel.cu (block-tree-reduce, no atomicAdd).
ISV_TOTAL_DIM bumped 450 → 452.
ATOMIC migration per feedback_no_partial_refactor: both label kernels
(aux_sign_label_kernel.cu trajectory + aux_sign_label_per_step_kernel.cu
per-rollout-step) migrated together to the new
(targets, bar_indices, isv, isv_h_idx, out_labels, total, total_bars)
signature. The lookahead host-passed scalar argument is removed; H is
read from ISV inside the kernel (broadcast value, single read per
thread, on-device clamp [1, 240]).
Producer chain (per-epoch boundary): new
GpuDqnTrainer::launch_aux_horizon_chain orchestrates
avg_win_hold_time_update → aux_horizon_update sequentially alongside
launch_kelly_cap_update at the existing epoch-boundary slot in
training_loop.rs.
Trunk math (C.2/C.3/C.4) unchanged — separate aux trunk is label-
agnostic. Validation in C.10 will use H=60 cold-start; the adaptive
producer drives H from real winning-trade observations.
Tests (8 oracle, 5 new + 3 preserved):
- aux_trunk_forward_matches_numpy_reference (C.3) ✓
- aux_trunk_backward_gradient_check (C.4) ✓
- aux_trunk_backward_does_not_write_dx (C.4) ✓
- aux_sign_label_h_bar_horizon (NEW) ✓
- aux_sign_label_lookahead_mask (NEW) ✓
- aux_horizon_pearl_a_bootstrap (NEW) ✓
- aux_horizon_converges_to_steady_target (NEW) ✓
- aux_horizon_holds_sentinel_with_no_winning_trades (NEW) ✓
8/8 pass on RTX 3050 Ti.
Phase C.4b of SP14 Layer C separate-aux-trunk refactor.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
5d584dc751 |
feat(sp14-c): aux trunk backward kernel + gradient check + stop-grad invariant test
Backward propagates dh_s2_aux through w3/w2/w1 with block-tree-reduce
(no atomicAdd per feedback_no_atomicadd). Critical: kernel set does NOT
write dx_in — encoder gradient remains Q-shaped only. Stop-grad
invariant verified via parameter-list structural enforcement (kernels
literally cannot reference an `dx_in_out` pointer they don't accept) +
kernel source inspection that strips comments and asserts no `dx_in`
write pattern.
Three kernels in aux_trunk_backward_kernel.cu:
- aux_trunk_bwd_dh_pre: per-sample, computes dh_aux2_pre [B, H2] +
dh_aux1_pre [B, H1] using ELU' from POST-activation form
(`(y > 0) ? 1 : (1 + y)` mirrors aux_elu_bwd_from_post in
aux_heads_kernel.cu).
- aux_trunk_bwd_dW_reduce: generic outer-product reduce
`dW[k, j] = sum_b A[b, k] * B[b, j]`. One block per output
element, shmem-tree reduce over batch. Used 3× (dW3, dW2, dW1).
- aux_trunk_bwd_db_reduce: generic batch-reduce `db[j] = sum_b
B[b, j]`. One block per output element. Used 3× (db3, db2, db1).
Memory-efficient: no per-sample partials (avoids B×163,072 floats for
production topology). Per-element reduction means O(P) blocks each
doing O(B) work in shmem.
Rust wrapper AuxTrunkBackwardOps in gpu_aux_trunk.rs orchestrates seven
launches in fixed sequence (capture-friendly, no host branches per
pearl_no_host_branches_in_captured_graph). All three CudaFunction
handles pre-loaded once at construction. Field added to GpuDqnTrainer
alongside aux_trunk_forward_ops; constructor mirrors C.3 pattern.
Tests (all pass on RTX 3050 Ti, sub-ULP forward, 1.33e-2 max rel-err
backward gradient at smallest sampled gradient):
- aux_trunk_forward_matches_numpy_reference (C.3 — preserved).
- aux_trunk_backward_gradient_check (NEW): central-difference
numerical gradient at 16 sampled dW3 indices vs analytic from
backward kernel. Loss = 0.5 * ||h_s2_aux||^2 so dh_s2_aux =
h_s2_aux. EPS=1e-3, B=4, ENC=H1=H2=AUX=32 (33 forwards in ~2s).
REL_TOL = 2e-2 (f32 finite-difference noise floor for
small-gradient tail; production topology is dimension-independent
given runtime args).
- aux_trunk_backward_does_not_write_dx (NEW): reads kernel source,
strips C-style comments (so design-discussion text mentioning
`dx_in` doesn't false-positive), asserts no `dx_in` / `dx_in_out`
symbol survives in code. Complements the structural enforcement
(kernel signatures don't accept `dx_in_out` pointer).
Phase C.4 of SP14 Layer C separate-aux-trunk refactor. Module is
additive — wire-up into collector backward chain + Adam updates lands
in Phase C.5 (atomic).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
cb6bca4629 |
feat(sp14-c): aux trunk forward kernel + Rust wrapper + oracle test
3-layer MLP forward (Linear→ELU→Linear→ELU→Linear). Pre-loaded CudaFunction for graph-capture safety per pearl_no_host_branches_in_captured_graph. Oracle test verifies bit-for-bit match against numpy reference within 1e-4 tol. Saves h_aux1 and h_aux2 to global memory for backward. Phase C.3 of SP14 Layer C separate-aux-trunk refactor (plan: docs/superpowers/plans/2026-05-07-sp14-layer-c-separate-aux-trunk.md). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
4926fb7c65 |
feat(sp14-c): allocate aux trunk params (~132K) + Adam m/v state
3-layer MLP: encoder_out_dim → 256 → 128 → AUX_HIDDEN_DIM. Kaiming-He weights (Box-Muller from LCG-uniform), zero biases, separate Adam m/v buffers (12 state tensors). Allocated in trainer constructor — collector borrows via raw_ptr at wire-up time per existing OFI-embed / q-attn ownership pattern (no parallel param mirror needed). Not yet wired to forward/backward — pure allocation per Phase C.2 design. Topology dimensions resolved against actual codebase: - encoder_out_dim = config.shared_h1 (= SH1 = 256 in production) - AUX_HIDDEN_DIM = config.shared_h2 (= SH2 = 256, matches existing aux head's input dim per aux_heads_kernel.cu:118 `h_s2 [B, SH2]`) Total params: 65,536 + 256 + 32,768 + 128 + 32,768 + 256 = 131,712. Audit doc updated per Invariant 7. Phase C.2 of SP14 Layer C separate-aux-trunk refactor (plan: docs/superpowers/plans/2026-05-07-sp14-layer-c-separate-aux-trunk.md). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
4f372a49a8 |
refactor(sp14-c): atomic α machinery deletion + aux trunk ISV slot allocation
Phase C.1 of SP14 Layer C separate-aux-trunk refactor. Single atomic
commit per feedback_no_partial_refactor and feedback_no_legacy_aliases —
no DEPRECATED stage.
Deleted (per C.0 audit
|
||
|
|
a7a162d1ff |
docs(sp14-c-preflight): catalogue α-machinery launch sites pending atomic deletion
Phase C.0 of SP14 Layer C separate-aux-trunk refactor. Pure audit doc entry — no source code change. Establishes "before" state of α machinery (Coupling B) launch sites + ISV slots, ahead of atomic deletion in Phase C.7. Files slated for deletion: alpha_grad_compute_kernel.cu, sp14_scale_wire_col_kernel.cu. ISV slots: VAR_AUX/VAR_Q/VAR_ALPHA/ ALPHA_RAW/ALPHA_SMOOTHED/GATE1_OPEN_STATE/GATE2_OPEN_STATE (7 total). Preserved: q_disagreement_* slots [383..390) + producer (diagnostic-only); dir_concat_qaux_kernel (Coupling A: forward feature wire, survives unchanged). Plan: docs/superpowers/plans/2026-05-07-sp14-layer-c-separate-aux-trunk.md Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
0ec791734d |
fix(sp15-cubin-preload): pre-load 6 SP15 launchers' CudaFunction handles in experience collector
Bug audit Pattern 2 — 6 SP15 launchers in gpu_dqn_trainer.rs were doing load_cubin + load_function PER CALL inside the launcher body, called from gpu_experience_collector.rs's per-rollout-step body (~thousands of times per epoch x 4096 envs x 1000 timesteps). Same architectural bug class as commits |
||
|
|
fa1f299147 |
fix(sp7-cadence): migrate SP5 Pearl 2 budget + SP7 loss-balance controller from process_epoch_boundary to per-step submit_aux_ops
Bug audit finding #2 (post train-d2b2s diagnostic — same Pattern 1 class as SP14 B.11 commit |
||
|
|
411a304731 |
fix(aux-head-regime): stop-gradient on aux_regime_backward dh_s2 — completes today's stop-gradient pair
Mirrors commit |
||
|
|
872bd73927 |
fix(aux-head): stop-gradient on aux's h_s2 input — fixes aux-loss-rises-during-training pathology
Root cause from train-v8ztm 9-epoch HEALTH_DIAG aux next_bar_mse trajectory: - Ep 0: 0.352 (learnable signal — below random baseline ln(2)≈0.693) - Ep 9: 0.717 (above random baseline — aux is now WORSE than random) - aux_dir_acc_long stuck at 0.19 (anti-correlated with truth) Aux head's backward gradient was flowing back to shared trunk activation h_s2 via dh_s2_out write at aux_heads_kernel.cu:599-613. Q-loss gradient on h_s2 dominates (larger magnitude, structurally different objective: cumulative discounted reward vs next-bar direction). h_s2 evolves to support Q's task; aux's CE loss climbs as h_s2 features become anti-aligned with direction prediction. Fix: stop-gradient. Aux reads h_s2 via forward, trains its own w1/b1/w2/b2 from CE loss, but does NOT propagate to h_s2. Q-loss is the sole shaping force on h_s2. Aux must adapt to whatever h_s2 happens to be — if the representation has direction signal, aux's params will extract it; if not, aux can't learn (separate-trunk Option 2 deferred for that case). This was the SEVENTH fix in today's chain (after 6 SP14 EGF cadence/ gate/saturation fixes). The EGF was a scaffold over a broken aux head; fixing aux first is the architectural prerequisite for EGF to route useful signal. Verification: cargo check clean; sp14_oracle_tests 7/7 pass. Validation: aux next_bar_mse should now DECREASE during training in the next L40S smoke (vs the rising-from-0.35-to-0.72 pattern in v8ztm). Deferred follow-up: aux_regime_backward has the same architecture (propagates dh_s2 to trunk). Same fix is a candidate once next_bar result validates the approach. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
c260dca8bd |
fix(sp14-β): remove training-time EGF launches from submit_aux_ops — collector path is canonical
Atomic step 5 of β migration. SP14 EGF producer chain now fires
ONLY from the experience collector (commit
|
||
|
|
c691bd381a |
feat(sp14-β): wire collector-native SP14 producer chain after aux forward
Step 4 of β migration: 3 SP14/SP13-EGF producers fire per-rollout-step
in collector using collector-owned kernel handles + collector stream.
Reads rollout-time q_values (post-expected-Q, pre-IQR/ensemble/noise)
+ exp_aux_nb_softmax. Writes to shared ISV.
Producer order preserved (matches trainer submit_aux_ops chain):
1. SP13 dir-acc reduce → 2 fixed-α EMAs → aux_pred to ISV[375]
2. SP14 q_disagreement_update (reads aux softmax + q_values)
3. SP14 alpha_grad_compute (pure ISV state machine)
Same kernel gate (commit
|
||
|
|
296ba92282 |
feat(sp14-β): wire collector-side aux-head forward + label producer per-rollout-step
Step 3 of β migration: collector now runs aux_next_bar_forward on rollout state every step. Label producer (new thin variant aux_sign_label_per_step_kernel) derives sign(price[t+1] - price[t]) per env using bar = episode_starts[ep] + t. Aux predictions feed the EGF kernel chain (step 4), NOT the Q-head's input (rollout Q-head still sees raw h_s2, dir_qaux_concat_ptr remains 0u64). Placement: AFTER captured forward graph, BEFORE expected_q kernel. Same-stream serial ordering reads exp_h_s2_f32 populated by forward_online_f32 inside the captured graph. Cold-start gated on trainer_params_ptr != 0 to skip the test-scaffold path where the trainer hasn't wired its params yet. Files added: aux_sign_label_per_step_kernel.cu (66 lines). Files modified: build.rs (+8 lines, register cubin), gpu_experience_collector.rs (+106 lines: struct field, cubin static, load in new(), per-step launch block). Compile clean; sp14_oracle_tests 2/2 non-GPU pass. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
88eb7aa241 |
feat(sp14-β): allocate rollout-sized aux buffers + AuxHeadsForwardOps in collector
Step 2 of β migration: 5 new buffers sized to alloc_episodes (vs trainer's batch_size). AuxHeadsForwardOps instance is collector- owned and stream-bound to collector stream via AuxHeadsForwardOps::new(&stream). Param tensors shared with trainer via existing f32_weight_ptrs_from_base path. Buffer sizing: exp_aux_nb_hidden_buf [alloc_episodes × 32], exp_aux_nb_logits/softmax_buf [alloc_episodes × 2], exp_aux_nb_label_buf [alloc_episodes] i32, exp_aux_dir_acc_buf [6] mapped-pinned (post-B1.1a 6-float layout matches trainer). Compile clean; sp14_oracle_tests 2/2 non-GPU pass (7 GPU tests ignored on RTX 3050 Ti host). No aux forward yet — buffers allocated, ready for wire. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
09202aa991 |
feat(sp14-β): pre-load SP14 EGF kernel handles in experience collector (Option B)
Step 1 of β migration — Option B (collector-native): collector loads its own CudaFunction handles for the 3 SP14 producer kernels plus their sub-kernels (4 total: aux_dir_acc_reduce, aux_pred_to_isv_tanh, q_disagreement_update, alpha_grad_compute). Mirrors SP13 hold_rate pattern at gpu_experience_collector.rs:1820. Cleaner than cross-component launcher calls (avoids trainer-stream / collector-stream race; no signature surgery on the existing trainer launchers). Cubin static decls flipped to pub(crate) so the collector can re-load on its own stream. Compile clean; sp14_oracle_tests pass (2/2 non-GPU; GPU-gated 7 ignored on RTX 3050 Ti host). No launches yet — additive infrastructure commit. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
9d0c124cee |
fix(sp14-egf): gate q_disagreement EMA update on total_cnt > 0 — fixes training-time decay-to-zero of rollout signal
Root cause from train-6fcml 5-epoch trajectory (commit
|
||
|
|
5608b866b6 |
fix(producer-cadence): migrate 4 more per-step ISV producers from process_epoch_boundary to per-step hot path
Continuation of SP14 B.11 cadence fix (commit
|
||
|
|
200f05fcef |
fix(sp14-B.11): move EGF producer chain into per-step training loop — fixes per-epoch staleness
Root cause from train-v8ztm 10-ep validation (commit |
||
|
|
1396b62ec6 |
fix(sp15-wave5-followup): pre-load sp15_baseline + cost_net cubins — fixes hyperopt-trial CUDA_ERROR_ILLEGAL_ADDRESS
Root cause: 5 SP15 evaluation launchers (cost_net_sharpe + 4 baseline_*) were doing `load_cubin` + `load_function` PER-CALL inside `GpuBacktestEvaluator`'s eval hot loop. Pattern is fragile across CUDA context lifetimes — works in single-pass train-best context (smoke train-9bcwm verified), fails in hyperopt-trial child stream context (workflow train-xggfc trial 1 failed at "load sp15_baseline_kernels cubin: ILLEGAL_ADDRESS"; after the host-side load corrupted the trial's context, trials 2-20 all cascade-failed at "Fork CUDA stream for trial"). Fix (atomic, matches |
||
|
|
5d63762ab3 |
fix(sp15-wave4.1b-OOB-followup): pre-load bn_tanh_concat_dd_kernel — fixes forward-capture SEGV
Root cause: SP15 Wave 4.1b ( |
||
|
|
bfc3ffa9dc |
diag(sp15-wave5): in-capture CAPTURE_PHASE_* checkpoints
Smoke train-fp7xx printed all 16 Phase-6/7/8 checkpoints clean through PER_PRIORITY_DONE. CAPTURE_DONE missing, exit changed 139→143 (SIGTERM) — capture_training_graph hangs or takes ~40s past PER_PRIORITY_DONE before Argo terminates the pod. Adds 14 CAPTURE_PHASE_* checkpoints across the 12 child captures + parent compose: BEGIN / PER_SAMPLE / COUNTERS / SPECTRAL / FORWARD / DDQN / AUX POST_AUX / ADAM_GRAD / ADAM_UPDATE / MAINTENANCE / IQL_MODULATE PER_PRIORITY / CHILDREN_STORED / PARENT_COMPOSED Diagnostic-only. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
a8d6c33040 |
diag(sp15-wave5): extend STEP0_PHASE_* checkpoints into Phase 6/7/8
Smoke train-xq9hg got past all 13 original checkpoints (BEGIN through NAN_CHECKS_DONE) cleanly. SEGV is downstream in Phase 6/7/8. Adds 16 more checkpoints: TLOB_BWD_ADAM, MAMBA2_BWD, MAMBA2_ADAM, OFI_EMBED_BWD, OFI_EMBED_ADAM, PRUNING, BRANCH_GRAD_BALANCE, GRAD_NORM, ADAM_OPS, Q_MAG_BIN, Q_DIR_BIN, ISV_UPDATE, MAINTENANCE, IQL_MODULATE, PER_PRIORITY, CAPTURE. Diagnostic-only — no behavior change. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
e9d9afbd61 |
diag(sp15-wave5): stderr checkpoints in run_full_step step-0 ungraphed path
L40S smoke train-jfbzr (commit
|
||
|
|
23e9a1f78c |
fix(cuda): compute_expected_q stride 13 vs denoise_target_q_buf size 12 — OOB writes threads 60-63
Pre-existing latent bug surfaced by compute-sanitizer after the SP15 NULL-pointer fix at |
||
|
|
2e37af29d4 |
fix(sp15-p1.3.b-followup-OOB): wire ISV bus + SP15 control pointers BEFORE first collect — fixes "load sp15_dd_state cubin" CUDA_ILLEGAL_ADDRESS
Root cause: Wave 4.1b (s1_input_dim 102→103, |
||
|
|
c16b3b5a80 |
feat(sp15-p1.3.b-followup-B): per-(env,t) dd_trajectory + PER sampler — fixes Wave 4.3 uniform-batch limitation
Phase 1.3.b-followup (
|
||
|
|
5b394f1035 |
feat(sp15-p1.3.b-followup): per-env DD redesign — Path A env-0-canonical → Path B per-env tile + reduction
Closes the Phase 1.3.b deferred per-env redesign per feedback_no_partial_refactor. Path A (env-0-canonical, commit |
||
|
|
483cef454c |
feat(sp15-p3.5.4.c): production caller + OR-gate consumer for plasticity injection — closes 3.5.4 end-to-end
Wave 4.2 (
|
||
|
|
19f6cce510 |
feat(sp15-wave4.3 / 3.5.5.b): PER sampler integration — recovery transitions oversampled
Phase 3.5.5 (
|
||
|
|
ef08611d3f |
feat(sp15-wave4.2 / 3.5.4.b): cuRAND Kaiming-He weight reset for plasticity injection
Phase 3.5.4 (
|
||
|
|
a54f53e4ed |
feat(sp15-wave4.1c): behavioral KL test — dd_pct trunk integration shifts policy distribution
Closes out Wave 4.1 (Phase 1.5.b consumer migration). Wave 4.1a ( |
||
|
|
eb9515e41c |
feat(sp15-wave4.1b): consumer migration — s1_input_dim 102→103, GRN w_s1 reshape, 4 forward + 3 backward sites
Atomic consumer migration that flips production callers of bn_tanh_concat_kernel
over to Wave 4.1a's bn_tanh_concat_dd_kernel and bumps s1_input_dim from 102 to
103 across the entire trunk forward + backward path. Eliminates the documented
Wave 4.1a transient orphan.
Changes:
- s1_input_dim formula bump (bn_dim + portfolio_dim → bn_dim + portfolio_dim + 1)
at compute_param_sizes, trainer ctor's CublasGemmSet, xavier_init_params_buf,
the experience collector's CublasGemmSet, and CublasBackwardSet::new for the
backward gemm cache. cuBLAS gemm caches re-key automatically (fresh HashMap).
- GRN w_a_h_s1[0] / w_residual_h_s1[4] reshape [shared_h1, 102] → [shared_h1, 103]
via compute_param_sizes + xavier_init's fan_dims. Xavier-uniform init covers
the new dd_pct column (bounded [0,1] — Xavier's small-magnitude assumption is
appropriate; differs from SP14's aux_softmax_diff zero-init which was driven
by the bidirectional ±1 range).
- bn_concat_dim() accessor +1 (TLOB backward row stride).
- 5 concat_dim local-var bumps (1 alloc + 3 forward + 2 backward + 1 in
experience collector).
- 4 forward-call migrations to launch_sp15_bn_concat_dd: DDQN argmax pass
(~27158), online forward (~27467), target forward (~27666), experience
collector forward (~3853). Each takes self.isv_signals_dev_ptr; the kernel
reads ISV[DD_PCT_INDEX=406] on-device and broadcasts.
- 3 GRN backward sites (main, ensemble, CQL) flow through encoder_backward_chain
which uses s1_input_dim — bumped automatically. dd_pct column gradient is
silently discarded by vsn_d_gated_state_portfolio_pad_kernel (reads
[bn_dim..bn_dim+portfolio_dim) only) and bn_tanh_backward_kernel (reads
[0..bn_dim) only). Correct: dd_pct sources from ISV bus, no learnable input.
- Legacy bn_tanh_concat_kernel field DELETED from trainer struct alongside its
loader, tuple-element, destructuring, assignment (5 mechanical sites for the
one dead field). Function tuple shrinks 44→43 elements. Kernel symbol stays
in the cubin source for SP15 oracle parity tests.
- mag_concat / OFI concat audit verdict: DECOUPLED from s1_input_dim. They
widen shared_h2, not the trunk INPUT dim.
- test_gpu_backtest_evaluator_state_dim_calculation migrated to assert
STATE_DIM == 128 (was 96, stale per feedback_trust_code_not_docs).
Atomic per feedback_no_partial_refactor: every consumer of s1_input_dim and
bn_concat_buf row-stride migrated together. Eliminates Wave 4.1a transient-
orphan launcher per feedback_wire_everything_up. Legacy field deleted per
feedback_no_legacy_aliases.
Tests: cargo check clean (18 pre-existing unrelated warnings). Wave 4.1a
oracle parity test (bn_tanh_concat_dd_kernel_writes_dd_pct_column) still
passes. ML lib suite went from 945 pass / 14 fail (pre-Wave-4.1b baseline) to
947 pass / 12 fail post-Wave-4.1b — improved by +2 (state_dim_calculation
migration + ensemble checkpoint round-trip flake resolved).
Refs: SP15 Wave 4.1a (
|
||
|
|
a8da1cb9cf |
feat(sp15-wave4.1a): bn_tanh_concat appends dd_pct column from ISV — bottleneck-aware Phase 1.5 consumer migration
The standalone dd_pct_concat_kernel from Phase 1.5 was bottleneck-
incompatible — it operated on raw [B, 128] state, but production trunk
consumes [B, s1_input_dim] = [B, 102] post-bottleneck. Wave 4.1a fixes
this at the kernel level; Wave 4.1b lands the consumer migration
(s1_input_dim 102→103, GRN w_s1 reshape, 3 forward + 3 backward sites).
Spec correction (per feedback_trust_code_not_docs): the spec's
'state_dim 48→49' is stale terminology pre-STATE_DIM 48→112→128
evolution. Production s1_input_dim is bottleneck_dim + (STATE_DIM −
market_dim) = 16 + (128 − 42) = 102. Wave 4.1b will bump this to 103.
NEW bn_tanh_concat_dd_kernel in dqn_utility_kernels.cu:
- Fuses dd_pct append into the same launch as bn_tanh + portfolio
concat (output shape [B, bn_dim + portfolio_dim + 1])
- Reads isv[DD_PCT_INDEX=406] (set by Wave 1.3.b dd_state_kernel
per-step), broadcasts the scalar across batch as the appended
last column
DELETED standalone dd_pct_concat_kernel.cu + launch_sp15_dd_pct_concat
+ cubin manifest entry per feedback_no_legacy_aliases (zero production
callers — only test consumer; bottleneck-on path is canonical).
Test helpers added (used by Wave 4.1c behavioral KL test).
Phase 1.5 oracle test migrated to bn_tanh_concat_dd_kernel contract:
test name bn_tanh_concat_dd_kernel_writes_dd_pct_column passes on
RTX 3050 Ti.
Layout fingerprint already covers Phase 1.5 via the existing
TRUNK_INPUT_DD_PCT=sp15_phase_1_5; marker — pre-SP15 checkpoints
already break.
fxcache schema_hash auto-bumps from file content hashes (per task
P5T5 Phase F mechanism); no manual schema bump needed.
Atomic per feedback_no_partial_refactor for the kernel-signature
contract change. Consumer wiring (s1_input_dim propagation, GRN
reshape, forward/backward call sites) deferred to Wave 4.1b's atomic
commit per the established 3a/3b split precedent — kernel + launcher
land first.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
4320820ae2 |
feat(sp15-wave3b): host-side wire-up — eliminates 5 orphan launchers via GpuBacktestEvaluator constructor signature change
Second half of the Wave 3 val-cost-streams refactor (3a kernel-side
foundation landed at
|
||
|
|
e968f4ded9 |
feat(sp15-wave3a): kernel-side foundation — baseline output buffers + position_history derivation
Wave 3a half of the val-cost-streams refactor (3b host-side wire-up
follows). Atomically migrates the kernel-side contracts; 5 launchers
remain orphan transiently awaiting 3b production callers.
Baseline kernels (1.4):
- 4 baseline_*_kernel signatures gain 'out: float*' parameter writing
per-window [mean, std, raw_sharpe] (matches 1.1.b sharpe_per_bar shape)
- ISV writes to slots 409, 410, 412, 416 removed entirely
(per-window output is correct for WindowMetrics consumption;
ISV-scalar writes were spec scaffolding for a single-fold-aggregate
version that 1.4.b's per-window contract supersedes)
- 4 ISV slot constants removed from sp15_isv_slots.rs
- state_reset_registry: NO entries to remove (verified via grep —
the 4 slots never had registry entries / dispatch arms in the first
place; they were single-fold-aggregate scalars defaulted at every
fold start by the constructor-write that initialises the ISV bus).
Task 4 from the dispatch is a no-op; the
every_fold_and_soft_reset_entry_has_dispatch_arm regression test
continues to pass unchanged.
- 4 oracle tests migrated to output-buffer assertion
- layout_fingerprint_seed string updated (4 retired entries removed,
4 trunk-shared entries retained; layout-break-class change)
New action_decoding_helpers.cuh:
- Extracts factored_action_to_dir_idx + factored_action_to_position
__device__ helpers (the latter is a higher-level position state-
machine helper not previously available)
- Mirrors trade_physics.cuh::decode_direction_4b semantics exactly so
on-policy and counterfactual paths agree on factored-action meaning
- Single source of truth for action→direction→position mapping;
consumers #include the header
New position_history_derivation_kernel.cu (post-loop derivation for
cost_net_sharpe consumer in Wave 3b):
- Reads actions_history_buf, reconstructs per-bar position_history
(-1/0/+1), side_ind (1.0 on position change), rt_ind (1.0 on
transition-to-flat from non-flat) via sequential walk (single
block per window, no atomicAdd per feedback_no_atomicadd)
- New launcher launch_sp15_position_history_derivation in
gpu_dqn_trainer.rs
- New cubin manifest entry in build.rs
- 1 oracle test covering 8-bar Short→Hold→Long→Hold→Flat→Long→Flat→
Short sequence; expected position/side_ind/rt_ind triples match
hand-computed values
Atomic per feedback_no_partial_refactor for the ISV-contract change
(every consumer of slots 409/410/412/416 migrated in this commit; their
consumers were the 4 oracle tests, all migrated). The orphan launcher
transient state for the 5 baselines + derivation kernel is explicitly
the 3a/3b split point — production callers land in 3b's
GpuBacktestEvaluator::new constructor signature change.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|