da5e564ccfad43071f579a824ef122dc7017d837
748 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
da5e564ccf |
feat(sp22-vnext): Phase B4b-2 — replay-buffer scatter + trainer setter
Second half of Phase B4b. Wires the trade-outcome label column through
the replay buffer (struct field, allocation, scatter on insert, gather
on sample, direct-to-trainer pointer + setter) and connects the
trainer's aux_to_label_buf to receive per-batch sampled labels.
Completes the end-to-end producer → ring → trainer i32 path the K=3
sparse CE consumer reads.
Replay buffer changes (crates/ml-dqn/src/gpu_replay_buffer.rs):
- New struct fields: aux_outcome_labels (capacity-sized ring),
sample_aux_outcome_labels (mbs-sized fallback gather), trainer_aux_
outcome_labels_ptr (direct-path destination)
- New aux_outcome_labels_ptr field on GpuBatchPtrs
- insert_batch signature: new aux_outcome arg between aux_conf_in and
bs. Scatters via existing K-generic scatter_insert_i32 (same kernel
the K=2 aux_sign_labels uses).
- sample_proportional direct gather when trainer_aux_outcome_labels_ptr
!= 0; fallback gather otherwise. Mirrors aux_sign_labels direct/
fallback semantic exactly.
- New setter set_trainer_aux_outcome_labels_ptr mirrors
set_trainer_aux_conf_ptr.
Collector emission:
- New aux_outcome_labels field on GpuExperienceBatch
- Populated at end of collect_experiences_gpu via dtod_clone_i32 from
Phase B4b-1's per-(env, t) producer scratch.
Trainer wireup:
- aux_to_label_buf_ptr() accessor on GpuDqnTrainer (mirrors
aux_nb_label_buf_ptr)
- trainer_aux_to_label_buf_ptr() delegating accessor on
FusedTrainingCtx
- New set_trainer_aux_outcome_labels_ptr call in training_loop at the
same site where set_trainer_buffers + set_trainer_aux_conf_ptr fire.
Two call sites updated (lines ~835 + ~2857).
- insert_batch call in training_loop passes &gpu_batch.aux_outcome
_labels as new arg.
Test fixtures updated: 4 smoke test files + 3 unit-test fixtures in
gpu_replay_buffer.rs alloc zero-init aux_outcome i32 arg.
End-to-end chain complete:
trade_outcome_label_kernel (A2)
→ collector per-step launch (B3)
→ collector emission (B4b-1)
→ replay-buffer insert + scatter (B4b-2)
→ PER sample + direct gather (B4b-2)
→ trainer aux_heads_forward.loss_reduce (B4) reads sparse {-1,0,1,2}
→ trainer aux_heads_backward (B4) computes per-sample partials
→ Adam SAXPY (B1+B4) updates W1, b1, W2, b2 at [163..167)
The "degraded predict-Profit-everywhere" cold-start from Phase B4 is
resolved. K=3 head trains on real sparse trade-outcome labels.
Verification:
- cargo check -p ml clean (21 warnings, none new).
- cargo test -p ml --lib → 1016/0 on clean runs; pre-existing
NoisyLinear flake still surfaces ~30-50% of runs (unrelated to
vNext work — see
|
||
|
|
20e1aea27c |
feat(sp22-vnext): Phase B4 — trainer-side replay-batch chain wireup
Wires the trade-outcome head's forward + loss reduce + backward +
per-sample partial reduce + SAXPY into the trainer's aux_heads_forward
and aux_heads_backward methods. Adam SAXPY for the 4 new weight tensors
at [163..167) extends uniformly via the existing aux_param_specs array
iteration.
Changes to aux_heads_forward:
- Steps 7 + 8 appended after K=2 next-bar head's loss reduce
- K=3 forward reads weights at [163..167) (Phase B1) → writes to
aux_to_* save-for-backward tiles (Phase B2)
- K=3 loss_reduce writes aux_to_loss_scalar_buf + aux_to_valid_count_buf
Changes to aux_heads_backward:
- K=3 backward appended after K=5 regime backward, emits per-sample
partials (dW1, db1, dW2, db2) + per-sample dh_s2_aux
- aux_param_specs array extended from 8 → 12 entries. Per-tensor
reduce + SAXPY loop iterates uniformly, scaling each by aux_weight
- K=3 dh_s2_aux SAXPY appended after K=5's, all three heads'
gradients flow into aux trunk's dh_s2_aux_accum (encoder stop-grad
enforced structurally by aux_trunk_backward's missing dx_in output)
Label semantic (cold-start): aux_to_label_buf is alloc_zeros (all 0
= Profit) until Phase B4b lands replay-buffer label scatter. Model
trains on "predict Profit everywhere" — degraded but well-defined
(no NaN). Mirrors K=2 head's known-degraded state between B1.1a
(forward landed) and B1.1b (label producer wired).
Adam SAXPY: existing global SAXPY iterates 0..NUM_WEIGHT_TENSORS
(now 167) — 4 new weight slots get gradient SAXPYs followed by
Adam m/v updates uniformly. Architectural payoff of Phase B1's
NUM_WEIGHT_TENSORS bump.
Test flake mitigation: added bind_to_thread() to
ensemble::adapters::dqn::tests::shared_device() mirroring the
cuda_stream() test-helper pattern from the fix sweep at
|
||
|
|
b98dc2730d |
feat(sp22-vnext): Phase B2 — trade-outcome trainer saved-tensor + partial buffers
Adds 11 buffer fields + 2 orchestrator ops handles (fwd + bwd) to the
trainer struct, mirroring the existing aux_nb_* / aux_partial_nb_*
pattern at K=3 instead of K=2.
Trainer struct additions:
- aux_to_fwd: AuxTradeOutcomeForwardOps (Phase B0 scaffold)
- aux_to_bwd: AuxTradeOutcomeBackwardOps
- aux_to_hidden_buf [B, H=128] saved post-ELU
- aux_to_logits_buf [B, K=3] saved logits
- aux_to_softmax_buf [B, K=3] saved softmax (3 future consumers)
- aux_to_label_buf [B] i32 sparse {-1, 0, 1, 2}
- aux_to_loss_scalar_buf [1] mean CE
- aux_to_valid_count_buf [1] B_valid for backward
- aux_dh_s2_to_buf [B, SH2] SAXPYs into dh_s2_aux_accum
- aux_partial_to_w1 [B, H, SH2] per-sample dW1
- aux_partial_to_b1 [B, H] per-sample db1
- aux_partial_to_w2 [B, K=3, H] per-sample dW2
- aux_partial_to_b2 [B, K=3] per-sample db2
Memory: aux_partial_to_w1 = 256 MB at B=2048 — identical to K=2 head's
partial size (same SH2, same H). Total new aux-to footprint ≈ 260 MB.
The existing aux_param_grad_final_buf scratch is sized to the largest
tensor across all aux heads; trade-outcome head's largest is W1 [H, SH2]
= 32,768 floats — identical to K=2/K=5 W1s. No resize needed.
Cold-start label semantics: alloc_zeros yields label 0 (Profit) for
every sample. Until the producer wires in (B3), the trainer's CE loss
treats every sample as "should have predicted Profit" — degraded but
well-defined (no NaN). Mirrors the K=2 head's known-degraded state
between B1.1a and B1.1b.
No FoldReset registration: these buffers are overwritten every batch
— no stale-state-leak risk across folds (matches the existing aux_nb_*
pattern).
Phase B3 next: collector-side rollout buffers + forward chain wireup
into collect_experiences_gpu (per-env softmax → per-(i, t) fan-out
scatter for trainer's aux_to_softmax_buf population).
Audit: docs/dqn-wire-up-audit.md Phase B2 section.
Cargo check clean (21 warnings, none new on aux_to_* fields).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
205e46c171 |
feat(sp22-vnext): Phase B1 — trade-outcome head weight tensors + Xavier init
Adds the 4 weight tensors (W1, b1, W2, b2) for the SP22 H6 vNext
trade-outcome aux head into the trainer's flat params_buf at indices
[163..167). Adam machinery (m/v moment buffers, SAXPY iteration over
0..NUM_WEIGHT_TENSORS) picks up the new tensors uniformly — no
per-tensor wiring needed.
Changes:
- NUM_WEIGHT_TENSORS bumped 163 → 167. Most of the 54 references are
&[u64; NUM_WEIGHT_TENSORS] array-size generics that resize
uniformly with the constant.
- compute_param_sizes() adds 4 new size entries:
[163] aux_to_w1 [H=128, SH2=256] = 32,768 floats
[164] aux_to_b1 [H=128] = 128 floats
[165] aux_to_w2 [K=3, H=128] = 384 floats
[166] aux_to_b2 [K=3] = 3 floats
Total: 33,283 floats = ~133 KB params, ~266 KB Adam state.
- compute_param_sizes() debug_assert updated 163 → 167.
- Xavier fan_dims added: (H, SH2) for W1, (0, 0) for biases (zero-init),
(K=3, H) for W2. Cold-start: logits ≈ 0 → softmax ≈ uniform 1/3 → no
Profit/Stop/Timeout preference per pearl_first_observation_bootstrap.
SH2 stays at 256 in this commit (mirrors K=2 head exactly). The spec's
Phase B input concat (256 → 262 with plan_params 6-dim) will re-shape
slot [163] to 128 × 262 = 33,536 floats later — small touch-up vs the
full B1 commit.
Verification:
- cargo check -p ml clean (21 warnings, none new).
- 14 pre-existing test failures stashed-verified unrelated (OFI features
missing, ISV slot count drift — independent of NUM_WEIGHT_TENSORS).
Phase B2 next: saved-tensor + per-sample partial buffers (hidden_post,
logits, softmax, valid_count, dW*_partial, dh_s2_aux_out).
Audit: docs/dqn-wire-up-audit.md Phase B1 section.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
108a426c38 |
feat(sp22-vnext): Phase B0 — AuxTradeOutcome ops struct scaffolding
First Rust-side commit of Phase B for the SP22 H6 vNext trade-outcome aux head. Pure orchestrator scaffolding mirroring AuxHeadsForwardOps / AuxHeadsBackwardOps. Zero production callers — additive per the gpu_grn / aux_trunk commit-by-commit ordering convention (scaffold → trainer fields → collector wireup). Adds: - gpu_dqn_trainer.rs: 4 cubin embeds (TRADE_OUTCOME_LABEL_CUBIN, AUX_TRADE_OUTCOME_FORWARD_CUBIN, AUX_TRADE_OUTCOME_LOSS_REDUCE_CUBIN, AUX_TRADE_OUTCOME_BACKWARD_CUBIN). All #[allow(dead_code)] until B1+. - gpu_aux_heads.rs: AUX_OUTCOME_K = 3 constant (parallel to AUX_NEXT_BAR_K). - gpu_aux_heads.rs: AuxTradeOutcomeForwardOps struct holding 3 kernel handles (forward + loss_reduce + label producer). Launch methods: forward(), loss_reduce(), compute_label(). Per-env label kernel uses grid ceil(n_envs/256) — pure per-env map, no reduction. - gpu_aux_heads.rs: AuxTradeOutcomeBackwardOps struct holding 1 kernel handle (backward). Launch method: backward(). Reuses existing K-generic aux_param_grad_reduce from AuxHeadsBackwardOps — no separate reducer needed. Contract shapes mirror AuxHeadsForwardOps/AuxHeadsBackwardOps so subsequent wireup commits plug in with minimal contract drift. Phase B proper (input concat 256→262 with plan_params) becomes a small change touching only the buffer fill + W1 shape once B1-B4 land the wireup. Phase B1 next: trainer struct fields for W1/b1/W2/b2 + Adam state + Xavier init + reset registry entries. Audit: docs/dqn-wire-up-audit.md Phase B0 section. Cargo check clean (21 warnings, none new). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
57cf5c1e08 |
fix(sp22): defensive NaN guards on Step 8/11 dW + Adam path
Bisect cut #1 (train-6zbcn) DEFINITIVELY identified Step 8/11 as the NaN source. With launches disabled, q_var/q_by_action/v_a_means all CLEAN. Forward atom-shift kernels confirmed innocent. Defensive fix (belt-and-suspenders): 1. c51_aux_dw_kernel: clamp NaN/inf total to 0 at dW write site. 2. adam_w_aux_kernel: early-return if any input non-finite. 3. compute_expected_q: guard w_aux[a] read with isfinite (in addition to state_121). Together these prevent any NaN propagation through atom-shift regardless of where the upstream NaN originated. Step 8/11 launches RESTORED in submit_adam_ops. The underlying NaN source is still unidentified but defensively neutralised. Smoke verification next. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
a1945de031 |
bisect(sp22): disable Step 8/11 launches (c51_aux_dw + adam_w_aux)
User-requested static-analysis-first bisect (Option 2): comment out both launch_c51_aux_dw and launch_adam_w_aux in submit_adam_ops to isolate forward vs backward as NaN source. With these disabled: - Forward atom-shift kernels still run (no-op with W=0) - W stays at [0,0,0,0] (no Adam update) - dW kernel doesn't run Hypothesis test: - Clean smoke -> backward is the bug - NaN persists -> forward is the bug Audit doc updated. Cargo check clean. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
5106e3b117 |
feat(sp22): H6 Phase 3 α Phase C1 — collector W ptr setter
Wires the trainer's `w_aux_to_q_dir [4]` device pointer into the GpuExperienceCollector so rollout-time action selection sees the active atom-shift for the direction branch. Changes: - gpu_experience_collector.rs: new field aux_w_to_q_dir_dev_ptr (u64, default 0) + setter set_aux_w_to_q_dir_ptr(). Both rollout launchers (compute_expected_q + quantile_q_select) now pass W ptr + exp_states_f32 + STATE_DIM_PADDED instead of NULL placeholders. NULL-safe via the kernels' existing aux_shift_active gating. - gpu_dqn_trainer.rs: w_aux_to_q_dir field promoted to pub(crate) for cross-module access via raw_ptr(). - trainers/dqn/trainer/training_loop.rs: new wire-up block after SP15 warm-count setter, mirroring the established setter pattern. NULL-safe on test scaffolds where fused_ctx or collector is absent. End-state — rollout activation: - compute_expected_q + quantile_q_select now return shifted E[Q] / quantile-blends per direction action during rollout. - With trainer's W trained each step (Step 8+11 Adam), the rollout policy's direction-action distribution actively reflects the learned aux→policy coupling. state_121's per-env value drives a per-(env, action) bias of magnitude W[a] (≤0.5 initial prior, learned thereafter). Verification: cargo check -p ml --lib clean (0 errors, 21 pre-existing warnings). Trainer + collector now smoke-ready end-to-end for the trainer/rollout side. Eval-side activation pending Phase D. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
5d163e0e1d |
feat(sp22): H6 Phase 3 α — adaptive W via dW backward + Adam (B9 Steps 8+11)
Step 8 — c51_aux_dw_kernel (new):
- Per-action block tree-reduce: grid=(4,1,1), block=(256,1,1). One block
per W index a, tree-reduces dW[a] across batch via warp shuffle + shmem.
Zero atomicAdd per pearl_no_atomicadd.
- Per-sample contributions:
a == a_d: dW[a] += inv_batch × isw × (SP_b/dz) × state_121
a == a*: dW[a] += inv_batch × isw × (-γ(1-done)) × (SP_b/dz) × next_state_121
c51_loss_kernel forward — new scratch outputs:
- aux_target_a_dir_buf[B] (i32): saves best_next_a for d==0 after Step c
sampling.
- aux_proj_logdiff_dir_buf[B] (f32): saves SP_b = Σ_n p_target_n ×
(current_lp[upper_n] - current_lp[lower_n]) after Step d's projection
via re-derivation of lower_n/upper_n (matching Huber compression +
clamp arithmetic of block_bellman_project_f).
Step 11 — adam_w_aux_kernel (new):
- Standard Adam with bias correction, grid=(1,1,1), block=(4,1,1).
- Graph-capture-safe: lr via self.lr_dev_ptr pointer arg; step via
self.ptrs.t_buf pointer arg (matches main Adam pattern). beta/eps
as value args from sp5_isv_slots constants.
- Bias-correction denominator floored at 1e-30 to avoid /0.
Trainer wiring (submit_adam_ops):
- launch_c51_aux_dw + launch_adam_w_aux added right after
launch_adam_update. Both inside the captured adam_child graph.
- New trainer fields: aux_target_a_dir_buf, aux_proj_logdiff_dir_buf,
c51_aux_dw_kernel, adam_w_aux_kernel. Cubin statics SP22_C51_AUX_DW_CUBIN
+ SP22_ADAM_W_AUX_CUBIN added.
NULL-safety:
- aux_shift_active=false in c51_loss_kernel forward → both scratch
buffers stay at alloc_zeros 0 → dW reads 0 → Adam W is a no-op.
- aux_target_a_dir_out / aux_proj_logdiff_dir_out are NULL-tolerant.
Deferred (deliberate scope):
- dL/dstate_121 backward (c51 → aux head): refinement, not correctness;
aux head trains via own supervised CE loss.
- Phase C1 collector W ptr setter.
- Phase D (eval-side aux infrastructure).
Verification:
- cargo build -p ml --lib: 0 errors, 21 pre-existing warnings.
- nvcc full recompile clean (1m05s for sm_89 target).
- All forward atom-shift consumers + adaptive W backward + Adam now wired.
End-state: adaptive W trains from structural prior [-0.5, 0, +0.5, 0]
via projection log-diff gradient. Aux head trains independently via
supervised CE. Together they form learned cross-coupling from aux
direction predictions to dir-branch Q distribution shifts. Smoke can
now measure adaptive W's effect on WR.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
a98f299823 |
feat(sp22): H6 Phase 3 α — atom-shift forward-side complete (B9 Steps 7+9+10)
Steps 7+9+10 land all FORWARD-path consumers of atom positions:
Step 7 — c51_loss_kernel atom-shift threading:
- 5 new args (w_aux, batch_states, next_batch_states, aux_dir_prob_index,
state_dim). NULL-safe — any NULL collapses to 0 shift, bit-identical
to pre-Phase-3-α.
- Per-sample state_121 + next_state_121 hoisted once at top of kernel.
- eq_per_action[a] += W[a]*next_state_121 in dir branch (d==0) before
Expected SARSA target-action sampling. Biases a* toward aux-aligned.
- Bellman projection: effective_reward = reward + γ*Δ_target*(1-done)
- Δ_online substituted for reward in block_bellman_project_f call.
Math derivation: see prior commit
|
||
|
|
195c051e7a |
feat(sp22): H6 Phase 3 α — atom-shift wired through compute_expected_q (B9 Step 5+6)
W structural prior init kernel (Step 5): - aux_w_prior_init_kernel.cu: 4-thread one-time init writing W[a] = [-0.5, 0.0, +0.5, 0.0] (Short / Hold / Long / Flat). Hardcoded values in kernel — no HtoD per feedback_no_htod_htoh_only_mapped_pinned.md. - build.rs registration + gpu_dqn_trainer.rs cubin static + handle field + new() launch after alloc_zeros for w_aux_to_q_dir. compute_expected_q atom-shift (Step 6): - experience_kernels.cu signature grows 4 args (w_aux, batch_states, aux_dir_prob_index, state_dim). NULL-safe — collapses to 0 shift when either pointer is NULL, bit-identical to pre-Phase-3-α. - Per-action inner loop computes aux_atom_shift = w_aux[a] * state_121 once per (b, a) for d==0 only. Inner z-loop applies z_val += aux_atom_shift before all S/TZ/TZ²/TLM accumulators consume it. Launcher updates (3 trainer sites + 1 collector site): - populate_q_out: passes W + current states (online path). - replay_forward_for_q_values: passes W + current states (online replay). - compute_denoise_target_q: passes W + next_states (target on s'; W shared across online/target). - gpu_experience_collector.rs: passes NULL W + NULL states until Phase C1 wires the trainer's W ptr through a setter. Architectural notes: - Action selection (every compute_expected_q call) now uses shifted atom positions for direction branch (d==0). Other branches stay bit-identical (shift = 0). Step 7 (c51_loss_kernel) + Step 8 (c51_grad_kernel) will close the loop on training loss + W gradient in the same atomic commit to avoid the gradient mismatch trap per feedback_no_partial_refactor.md. - Adam wireup for w_aux_to_q_dir lands at Step 11; until then W stays at structural prior values (no Adam step modifies it). Verification: - cargo check -p ml --lib: 0 errors, 21 pre-existing warnings (Phase 3b baseline parity). - Audit doc updated with B9 Step 5+6 checkpoint entry. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
08fd5803c4 |
refactor(sp22): H6 Phase 3 α cleanup + runbook revision for atom-shift design
Same-session cleanup of Phase 3b scalar-bias residue (commit cb80b74ce's
additions that became architecturally redundant under the 2026-05-13
atom-shift revision). The scalar-bias α design is mathematically
ineffective in C51 distributional Q-learning (softmax-shift-invariance
→ W gradient = 0 → never trains). The revised design threads
`aux_atom_shift[b, a] = W[a] * state_121[b]` through compute_expected_q
+ c51_loss_kernel + c51_grad_kernel + mag_concat_qdir; dW + dstate
gradients integrate into c51_grad_kernel's projection backward.
Files
─────
- crates/ml/build.rs:
Removed `aux_to_q_dir_bias_kernel.cu` + `aux_to_q_dir_bias_backward_kernel.cu`
from kernels_with_common. Replaced with comment block documenting
C51 softmax-invariance reason + spec/audit pointers.
- crates/ml/src/cuda_pipeline/gpu_dqn_trainer.rs:
Removed `pub(crate) static SP22_AUX_TO_Q_DIR_BIAS_CUBIN` and
`SP22_AUX_TO_Q_DIR_BIAS_BWD_CUBIN` static declarations.
Removed 3 struct fields (aux_to_q_dir_bias_kernel,
aux_to_q_dir_bias_backward_dw_kernel,
aux_to_q_dir_bias_backward_dstate_kernel) and their new()
loading blocks + struct construction entries.
KEPT: w_aux_to_q_dir + adam_m_w_aux + adam_v_w_aux + dw_aux_buf
(still needed for atom-shift Adam-trained W).
Updated W doc-comment to describe atom-shift threading (was
scalar-bias) and structural-prior initialization `[-0.5, 0.0,
+0.5, 0.0]`.
- crates/ml/src/cuda_pipeline/aux_to_q_dir_bias_kernel.cu +
crates/ml/src/cuda_pipeline/aux_to_q_dir_bias_backward_kernel.cu:
Source files stay on disk as committed dead code (commit
|
||
|
|
cb80b74ce9 |
feat(sp22): H6 Phase 3b — α infrastructure loaded (NOT yet wired)
Phase 3b builds on Phase 3a ( |
||
|
|
464bc5f7a4 |
feat(sp22): H6 Phase 3 — Phase A foundation (WIP checkpoint, additive)
Phase A of the SP22 H6 Phase 3 implementation runbook
(docs/plans/2026-05-13-sp22-h6-phase3-alpha-beta-runbook.md). Purely
additive: new ISV slots + new kernel files + build.rs registration
+ state_reset_registry entries. Nothing in the running training or
eval pipeline consumes this infrastructure yet — the existing
6-component contract is untouched. Phase A is committable as a clean
WIP foundation; Phases B-F (7-component contract migration, α
plumbing, A2 eval-side wiring, verification, smoke) resume in a
future session.
Files
─────
- crates/ml/src/cuda_pipeline/sp22_isv_slots.rs (NEW)
REWARD_AUX_ALIGN_EMA_INDEX = 536 — EMA of 7th reward component
SP22_AUX_ALIGN_SCALE_INDEX = 537 — SP11-controller scale_β
- crates/ml/src/cuda_pipeline/gpu_dqn_trainer.rs
ISV_TOTAL_DIM bumped 536 → 538 (SP22 H6 Phase 3 +2 slots)
Layout fingerprint extended with SLOT_536 + SLOT_537 entries
- crates/ml/src/cuda_pipeline/mod.rs
pub mod sp22_isv_slots
- crates/ml/src/trainers/dqn/state_reset_registry.rs
sp22_reward_aux_align_ema (FoldReset, sentinel 0)
sp22_aux_align_scale (FoldReset, sentinel 0)
- crates/ml/src/cuda_pipeline/aux_to_q_dir_bias_kernel.cu (NEW)
α forward: Q_dir[b, a] += W_aux[a] * state_121[b]
Reads state_121 from encoder INPUT (batch_states), bypassing
the cold encoder weight for slot 121. No atomicAdd, no host
branches, capture-safe.
- crates/ml/src/cuda_pipeline/aux_to_q_dir_bias_backward_kernel.cu (NEW)
Two kernels in one .cu:
- aux_to_q_dir_bias_backward_dw: per-action block tree-reduce
computing dW[a] = Σ_b state_121[b] * dq_dir[b, a]
- aux_to_q_dir_bias_backward_dstate: per-sample sum
dstate_121[b] += Σ_a W[a] * dq_dir[b, a] (option-(i) gradient
routing: both α and encoder paths get signal)
- crates/ml/build.rs
Both new cubins registered in kernels_with_common (compiled
successfully by nvcc; cubins exist at OUT_DIR/aux_to_q_dir_bias_*.cubin)
- docs/dqn-wire-up-audit.md
Phase A entry documenting the checkpoint + remaining-work
breakdown for next session.
Verification
────────────
- cargo check -p ml --features cuda: 0 errors, 21 pre-existing
warnings (Phase 2 baseline parity)
- nvcc compiles both new cubins (3.6 KB fwd, 10.6 KB bwd)
- No runtime impact: no consumer wires the new infrastructure yet
Resumes in
──────────
Future session executes Phases B-F per the runbook:
- B: 11 tasks — 7-component contract migration cascade
- C: 1 task — α collector rollout-side wiring
- D: 7 tasks — A2 eval-side aux trunk + α + state-gather
- E: 7 tasks — verification gates
- F: 5 tasks — audit doc append + atomic commit (incl. this Phase A
+ Phase B-F changes) + push + smoke + verdict
Estimated remaining: ~28-43 hr engineering + ~37 min smoke wall-clock.
Phase B partial work (300-line diff for experience_kernels.cu + reward_component_ema_kernel.cu — 7-stride migration + β producer + preamble update) saved at /tmp/sp22-h6-phase3-b-partial.patch (local-only).
Refs
────
- docs/plans/2026-05-12-sp22-h6-phase3-alpha-beta.md (spec)
- docs/plans/2026-05-13-sp22-h6-phase3-alpha-beta-runbook.md (runbook)
- pearl_no_partial_refactor (Phase A is additive; safe to commit alone)
- pearl_no_atomicadd (backward block tree-reduce, no atomics)
- pearl_no_host_branches_in_captured_graph (kernel capture-safety)
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
7fc9799343 |
feat(sp22): H6 Phase 1 — aux→policy state bridge (atomic)
Wires the aux head's per-env directional probability into policy STATE
slot 121 (AUX_DIR_PROB_INDEX = PADDING_START + 0), preserving the trunk-
separation invariant from `pearl_separate_aux_trunk_when_shared_starves`.
H1 (label horizon) confirmed aux learns 78% dir-acc within-fold at H=200
but the policy was walled off; this commit conducts that signal through
the state input with a one-step lag.
Mechanism (rollout-time, collector-only)
────────────────────────────────────────
- State[env, t] reads `prev_aux_dir_prob[env]` (= p_up from step t-1).
- After aux forward at step t, the new copy kernel writes
`aux_softmax[env, 1]` → `prev_aux_dir_prob[env]` for step t+1.
- Cold-start + FoldReset seed the buffer to 0.5 (neutral; p_up = 50%)
via the pure-GPU `fill_f32` kernel — no HtoD per
`feedback_no_htod_htoh_only_mapped_pinned.md`.
- Launch sits in the same `isv_signals && trainer_params != 0` gate as
the aux forward, so when aux is skipped the cache keeps its previous
(sentinel or last-good) value instead of copying stale `alloc_zeros`.
Three state-gather kernels updated atomically (per
`feedback_no_partial_refactor.md`):
- `experience_state_gather` — training, reads `aux_dir_prob_per_env[i]`
- `backtest_state_gather` — eval (single-step), NULL → 0.5 (A3 fallback)
- `backtest_state_gather_chunk` — eval (chunked), NULL → 0.5 (A3 fallback)
`assemble_state` gained a 7th param `float aux_dir_prob` written to
`out[SL_PADDING_START + 0]`; the remaining 6 padding slots stay zero
for 8-alignment.
Phase 1 scope = training-side + eval A3 NULL fallback. A2 (aux trunk
forward in eval) is deferred per the runbook — gates on whether the
smoke moves WR off the 50.1–50.2% plateau.
New files
─────────
- crates/ml/src/cuda_pipeline/aux_softmax_to_per_env_kernel.cu
Modified
────────
- crates/ml-core/src/state_layout.rs (+AUX_DIR_PROB_INDEX)
- crates/ml/src/cuda_pipeline/state_layout.cuh (+assemble_state param)
- crates/ml/src/cuda_pipeline/experience_kernels.cu
(3 state-gather kernels + NULL-defensive sentinel)
- crates/ml/src/cuda_pipeline/gpu_experience_collector.rs
(per-env buffer + 2 kernel handles + cold-start fill + copy launch
+ FoldReset re-fill + state-gather arg)
- crates/ml/src/cuda_pipeline/gpu_backtest_evaluator.rs
(NULL aux_dir_prob_per_env at all 3 launchers for A3)
- crates/ml/src/cuda_pipeline/gpu_dqn_trainer.rs
(SP22_AUX_SOFTMAX_TO_PER_ENV_CUBIN static)
- crates/ml/src/cuda_pipeline/gpu_action_selector.rs
(EPSILON_GREEDY_CUBIN → pub(crate) so collector reuses fill_f32)
- crates/ml/build.rs (register new kernel)
- docs/dqn-wire-up-audit.md
(## 2026-05-12 — SP22 H6 implementation: Phase 1 entry)
Verification (all three gates clean, no smoke yet)
──────────────────────────────────────────────────
- cargo check -p ml --features cuda: 0 errors
- gpu_backtest_validation: 4/4 expected-passing tests still pass
(the 2 PnL-assertion failures are pre-existing per the runbook)
- compute-sanitizer --tool=memcheck: 0 CUDA errors
Refs
────
- docs/plans/2026-05-12-sp22-h6-aux-policy-state-bridge.md
- docs/plans/2026-05-12-sp22-h6-next-session-prompt.md
- pearl_separate_aux_trunk_when_shared_starves
- pearl_first_observation_bootstrap (sentinel = 0.5 cold-start)
- feedback_no_htod_htoh_only_mapped_pinned (fill_f32 not HtoD)
- feedback_no_partial_refactor (3 state-gather kernels atomic)
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
ad99b79e07 |
feat(sp21): T2.2 Phase 7.5 — E8 curriculum_weights → per-segment PER insert priority boost (atomic)
Closes the "true E8 per-segment PER sampling" deferral from Phase 7.
Phase 7 wired E8's SCALAR concentration to per_update_pa's alpha
boost; Phase 7.5 wires the FULL Vec<f32> of per-segment weights to
per_insert_pa's priority boost.
What lands:
1. 8 new ISV slots [528..536): CURRICULUM_WEIGHT_{0..8}_INDEX.
2. ISV_TOTAL_DIM 528 → 536 (bus extension); fingerprint adds 8 SLOT
entries; CURRICULUM_N_SEGMENTS=8 const + curriculum_weight_index
accessor.
3. per_insert_pa kernel reads isv[528 + seg_id] where seg_id = i % 8
(round-robin segment tag); effective priority × N_SEGMENTS ×
weight[seg_id]. Uniform weights → no-op (× 1.0); cold-start
sentinel → no-op; 0.1× floor against pathological zero-weight
segments preventing sticky exclusion.
4. Producer in training_loop writes 8 ISV slots from
result.curriculum_weights[0..8].
Segment tagging rationale:
- Naïve approach (tag tuples by val-curriculum-segment id) is
infeasible — val and training have separate coordinate systems
(same problem documented in Phase 5+6 audit re E6 winner indices).
- Round-robin via `i % 8` distributes experience-collector's typical
512+ tuple batch evenly across 8 segments. Over time buffer has
equal representation per segment; E8 weights redirect sampling
pressure toward "hard" segments at insert time.
- HEURISTIC mapping (doesn't preserve val-segment semantics) but
consumes the curriculum_weights vector for real PER priority
redistribution — Phase 7.5's stated goal.
Files changed:
- crates/ml/src/cuda_pipeline/sp21_isv_slots.rs: 8 new slot consts
+ N_SEGMENTS + curriculum_weight_index accessor + tests
- crates/ml/src/cuda_pipeline/gpu_dqn_trainer.rs: ISV_TOTAL_DIM bump
+ fingerprint
- crates/ml-dqn/src/per_kernels.cu: per_insert_pa per-segment boost
- crates/ml/src/trainers/dqn/trainer/training_loop.rs: producer wireup
- docs/dqn-wire-up-audit.md: 2026-05-11 audit entry
Verification (passing):
- cargo check -p ml --tests --features cuda: 0 errors
- cargo test -p ml --lib sp21_isv_slots: 4/4 (new curriculum_weight_
index test)
- sp20_aggregate_inputs_test: 12/12
- sp20_phase1_4_wireup_test: 2/2
- sp20_emas_compute_test: 4/4
- sp20_controllers_compute_test: 7/7
- sp21_per_trade_predicted_q_test: 3/3
Total: 35 tests, 0 failures.
SP21 T2.2 cascade — TRULY fully complete (13 atomic commits). Every
enrichment output E1-E8 wires to a real consumer. No remaining
deferrals or hardcoded controller anchors in SP21 T2.2 scope.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
34d19955ff |
feat(sp21): T2.2 Phase 7 — E8 curriculum_concentration → PER alpha boost (atomic)
Wires E8's curriculum_weights distribution into the per_update_pa alpha boost composition via a new scalar signal: `compute_curriculum_concentration(weights) = 1 - entropy/log(n)`. Same pattern as Phases 5+6 (E6 winner_concentration + E7 hindsight_magnitude); all three feed the same boost_delta sum. Plan revision (mirrors Phase 5+6 pattern): - Original plan: "wire E8 (curriculum weights) → segment sampling weights." Literal interpretation requires new curriculum-segment abstraction in PER (segments don't exist — flat ring buffer today). - Resolution: signal-driven from the SHAPE of the weights distribution (entropy concentration), not the CONTENT (per-segment weighted sampling). Feeds existing per_update_pa kernel; no new segment-sampling kernel. - True per-segment PER sampling deferred to Phase 7.5 (mirrors Phase 6.5 deferral for E7's literal injection consumer). Signal semantics (compute_curriculum_concentration): - Input: Vec<f32> of E8's per-segment weights (sum=1). - Output: 1 - entropy/log(n) ∈ [0, 1]. - 0.0 = uniform (segments equally hard) - 1.0 = one segment dominates (concentrated difficulty) - Single-segment trivially → 1.0 - Cold start (empty input) → 0.0 sentinel - Zero-weight segments skipped (p log p → 0) ISV slot allocation: - CURRICULUM_CONCENTRATION_INDEX = 527 - ISV_TOTAL_DIM 527 → 528 (bus extension) - Layout fingerprint adds SLOT_527 entry Kernel change (4 lines, no ABI churn): - per_update_pa reads isv_signals[527] (already-existing arg from Phase 5+6) - curriculum_term = clamp(curric_conc, 0, 1) × 0.1 ∈ [0, 0.1] - boost_delta upper bound 0.4 → 0.5 to accommodate new term - Cold-start short-circuit predicate extended to all 3 signals Producer wireup: - enrichment.rs: new helper compute_curriculum_concentration; EnrichmentResult.curriculum_concentration field; run_enrichments populates it; log line extended - training_loop.rs: post-enrichment block writes ISV[527] Files changed: - crates/ml/src/cuda_pipeline/sp21_isv_slots.rs: +1 slot const - crates/ml/src/cuda_pipeline/gpu_dqn_trainer.rs: ISV_TOTAL_DIM bump + fingerprint - crates/ml-dqn/src/per_kernels.cu: 4-line boost composition extension - crates/ml/src/trainers/dqn/trainer/enrichment.rs: helper + field - crates/ml/src/trainers/dqn/trainer/training_loop.rs: write ISV - docs/dqn-wire-up-audit.md: 2026-05-11 audit entry Verification (passing): - cargo check -p ml --tests --features cuda: 0 errors - cargo test -p ml --lib sp21_isv_slots: 3/3 - sp20_aggregate_inputs_test: 12/12 - sp20_phase1_4_wireup_test: 2/2 - sp20_emas_compute_test: 4/4 - sp20_controllers_compute_test: 7/7 - sp21_per_trade_predicted_q_test: 3/3 Total: 34 tests, 0 failures. After this commit (Phase 8 + 6.5 + 7.5): - Phase 8: signal-drive remaining controller GAINS in enrichment - Phase 6.5: true E7 hindsight synthetic injection (deferred) - Phase 7.5: true E8 per-segment PER sampling (deferred) Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
a6087a0b23 |
feat(sp21): T2.2 Phase 5+6 combined — E6/E7 signal-driven PER alpha scaling (atomic)
Wires E6 (compute_winner_indices) + E7 (compute_hindsight_labels) outputs into the PER priority update kernel via signal-driven ISV slots. Phases 5 and 6 combined into one atomic commit per feedback_no_deferrals_for_complementary_fixes — both attack the same consumer kernel (per_update_pa) with non-overlapping refactor scopes (E6 contributes one ISV-mediated scalar, E7 contributes another, kernel composes both into alpha_boost). Plan revision (briefed and accepted 2026-05-11): - E6's literal Vec<usize> of val bar indices CANNOT directly bump training-PER priorities — val and training have separate coord systems (no bar-index → buffer-slot mapping). - E7's true synthetic injection requires val state retention infrastructure (n_windows × max_len × state_dim scratch buffer) — significant scope expansion deferred to Phase 6.5. - The MEANINGFUL signal in both phases is scalar aggregations of the per-trade tape: E6 winner concentration (top-decile-mean P&L / all-mean P&L) and E7 hindsight magnitude (mean |counterfactual_pnl|). Both compose multiplicatively into per_update_pa's alpha_eff. ISV slot allocation: - WINNER_CONCENTRATION_INDEX = 525 - HINDSIGHT_MAGNITUDE_INDEX = 526 - ISV_TOTAL_DIM 525 → 527 (bus extension) - Layout fingerprint adds 2 SLOT entries - Compile-time test asserts SP21_SLOT_END (527) ≤ ISV_TOTAL_DIM Signal semantics: - winner_concentration: top_decile_mean_pnl / max(all_mean_pnl, EPS). > 1.0 = top decile dominates; ≈ 1.0 = uniform; ≤ 0 → 0.0 sentinel (losing strategy, signal ill-defined). - hindsight_magnitude: mean(|counterfactual_pnl|). Higher = larger missed opportunities. - Cold start: both at 0.0 sentinel; kernel short-circuits to alpha_eff = alpha (no-op). ABI surgery (minimal): - per_update_pa: ONE new arg `const float* __restrict__ isv_signals` at end. NULL-tolerant. Reads slots 525/526 device-side, composes boost_delta (bounded [0, 0.4]), applies alpha_eff = alpha + delta. - update_priorities_gpu launcher: passes existing self.isv_signals_dev_ptr (already on struct, settable via set_isv_signals_ptr — same infra per_insert_pa uses for recovery_oversample). Zero new public API. Producer wireup: - enrichment.rs: 2 new helpers (compute_winner_concentration, compute_hindsight_magnitude); EnrichmentResult gains 2 fields (winner_concentration, hindsight_magnitude); run_enrichments populates both; log line extended. - training_loop.rs: post-enrichment block writes both ISV slots via fused.trainer().write_isv_signal_at. Files changed: - crates/ml/src/cuda_pipeline/sp21_isv_slots.rs: +2 slot consts - crates/ml/src/cuda_pipeline/gpu_dqn_trainer.rs: ISV_TOTAL_DIM bump + fingerprint - crates/ml-dqn/src/per_kernels.cu: per_update_pa ABI + boost logic - crates/ml-dqn/src/gpu_replay_buffer.rs: launcher passes ISV ptr - crates/ml/src/trainers/dqn/trainer/enrichment.rs: 2 helpers + 2 fields + populate - crates/ml/src/trainers/dqn/trainer/training_loop.rs: write 2 slots - docs/dqn-wire-up-audit.md: 2026-05-11 audit entry Verification (passing): - cargo check -p ml --tests --features cuda: 0 errors - cargo test -p ml --lib sp21_isv_slots: 3/3 - sp20_aggregate_inputs_test: 12/12 - sp20_phase1_4_wireup_test: 2/2 - sp20_emas_compute_test: 4/4 - sp20_controllers_compute_test: 7/7 - sp21_per_trade_predicted_q_test: 3/3 Total: 34 tests, 0 failures. Behavioral gate (alpha_eff lift on post-warmup val passes) is the upcoming smoke training run. Deferred follow-up (Phase 6.5): true E7 hindsight synthetic experience injection via val state retention + per_insert_pa extension. Significant infrastructure (n_windows × max_len × state_dim scratch + insert API extension). Documented in audit. After this commit (T2.2 Phases 7-8 + 6.5 deferred): - Phase 7: E8 curriculum weights → segment sampling - Phase 8: signal-drive remaining controller GAINS - Phase 6.5: synthetic hindsight injection (deferred — needs val state buffer infrastructure) Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
47e67011c9 |
feat(sp21): T2.2 Phase 4.5 — re-instate Pearl C engagement tracking for branches (atomic)
Closes the Phase 4 deferral. The 4 DqnBranches sub-launches now write
per-block engagement counts to 4 distinct ranges in
clamp_engage_per_block_buf; pearl_c_post_adam_engagement_check
aggregates all 4 ranges into one per-group rate-deficit EMA —
semantically equivalent to the pre-Phase-4 single-launch tracking.
Design: Option (b) sub-block offsetting, mirroring the existing
Curiosity pattern exactly. SP4_ENGAGE_EXTRA_BRANCHES_SUBLAUNCHES = 3
reserves 3 extra slots beyond Curiosity's tail; sub-launch 0 (Dir)
reuses the canonical DqnBranches slot 2; sub-launches 1/2/3 (Mag,
Order, Urgency) get extra slots 11/12/13. This keeps ParamGroup at
8 entries (no taxonomy growth, no 28 new ISV slot allocations).
SP4_ENGAGE_BUF_LEN: 11 → 14 × MAX_BLOCKS_PER_ADAM (45056 → 57344
i32 slots, +12 KiB on a non-hot-path mapped-pinned buffer).
Branch-to-offset mapping:
| Sub-launch | Offset constant | Slot | Buffer offset |
|------------|------------------------------------|------|---------------|
| Dir (b0) | SP4_ENGAGE_OFFSET_BRANCH_DIR | 2 | 8 192 |
| Mag (b1) | SP4_ENGAGE_OFFSET_BRANCH_MAG | 11 | 45 056 |
| Order (b2) | SP4_ENGAGE_OFFSET_BRANCH_ORDER | 12 | 49 152 |
| Urgency(b3)| SP4_ENGAGE_OFFSET_BRANCH_URGENCY | 13 | 53 248 |
pearl_c_post_adam_engagement_check generalized: a `multi_range` flag
covers both Curiosity (4 sub-launches W1/b1/W2/b2) and DqnBranches
(4 sub-launches Dir/Mag/Order/Urgency). Read AND zero-out paths use
the same flag — no duplication of branching logic.
Files changed:
- crates/ml/src/cuda_pipeline/sp4_isv_slots.rs: SP4_ENGAGE_EXTRA_
BRANCHES_SUBLAUNCHES const; SP4_ENGAGE_BUF_LEN formula extended;
SP4_ENGAGE_OFFSET_BRANCH_{DIR,MAG,ORDER,URGENCY}; layout test
updated to 14× MAX_BLOCKS_PER_ADAM
- crates/ml/src/cuda_pipeline/gpu_dqn_trainer.rs: launch_adam_update
branch sub-launches use new offsets; pearl_c_post_adam_engagement_
check aggregates DqnBranches multi-range
- docs/dqn-wire-up-audit.md: 2026-05-11 audit entry
Verification (passing):
- cargo check -p ml --tests --features cuda: 0 errors
- cargo test -p ml --lib sp4_isv_slots: 2/2 (engage buf layout asserts
14× + branch offsets correct)
- cargo test -p ml --lib sp21_isv_slots: 3/3
- sp20_aggregate_inputs_test: 12/12
- sp20_phase1_4_wireup_test: 2/2
- sp20_emas_compute_test: 4/4
- sp20_controllers_compute_test: 7/7
- sp21_per_trade_predicted_q_test: 3/3
Total: 33 tests, 0 failures. Behavioral gate: HEALTH_DIAG emits
per-branch engagement-rate-deficit EMAs in upcoming smoke run.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
b7c4f84ea0 |
feat(sp21): T2.2 Phase 4 — wire E4 branch_lr_scale → per-branch Adam (atomic)
Wires EnrichmentResult::branch_lr_scale (E4's [f32; 4] clamped [0.5,
2.0] per-branch LR multiplier) into the DQN Adam optimizer. Splits
the existing single DqnBranches Adam sub-launch into 4 per-branch
sub-launches, each consuming its own LR scale from ISV[521..525).
Plan amendment from on-paper design:
- Plan said "via existing per-group Adam infrastructure" but per-group
operates at PARAM_GROUP granularity (8 groups, all 4 action branches
lumped into DqnBranches). E4 needs per-action-BRANCH granularity.
- Resolution: split DqnBranches sub-launch into 4 per-branch sub-launches
using the already-canonical branch byte ranges (4 param tensors per
branch: Dir [17..21), Mag [21..25), Order [25..29), Urgency [29..33)).
Coverage invariant updated to 7-way (was 4-way).
- Pearl C engagement tracking on branches DEFERRED to Phase 4.5: the
shared DqnBranches engagement counter offset would collide on writes
if 4 sub-launches use the same offset. Pearl C is a diagnostic system
(not load-bearing) so its temporary unavailability for branches is
acceptable; Trunk + Value + Trunk-extras Pearl C still active.
Phase 4.5 follow-up will re-instate via ParamGroup expansion or
sub-block offsetting scheme.
ISV slot allocation:
- BRANCH_LR_SCALE_{DIR,MAG,ORDER,URGENCY}_INDEX = 521..525
- ISV_TOTAL_DIM 521 → 525 (bus extension)
- Layout fingerprint adds 4 SLOT entries
- New branch_lr_scale_index(branch_idx) accessor for clean mapping
- Cold-start floor: launcher reads ISV, floors at 1.0 if at sentinel
0.0 (per pearl_first_observation_bootstrap — first emit replaces
directly with no intermediate state)
ABI surgery:
- dqn_adam_update_kernel: ONE new arg float lr_scale at end of
signature; lr = *lr_ptr * lr_scale inside kernel
- 5 callers migrated atomically per feedback_no_partial_refactor:
- launch_adam_update (main DQN): 7 sub-launches with per-branch
lr_scale (Trunk + Value + 4 branch + Trunk-extras)
- 4 post-aux launchers (ofi_embed/aux_trunk/denoise/sel) pass 1.0
- decision_transformer Adam launch passes 1.0
Producer wireup (training_loop.rs post-enrichment block):
```rust
for (branch_idx, &scale) in result.branch_lr_scale.iter().enumerate() {
fused.trainer().write_isv_signal_at(
branch_lr_scale_index(branch_idx),
scale,
);
}
```
Files changed:
- crates/ml/src/cuda_pipeline/sp21_isv_slots.rs: +4 slots + accessor + tests
- crates/ml/src/cuda_pipeline/gpu_dqn_trainer.rs: ISV_TOTAL_DIM bump +
fingerprint + 7-way Adam split with per-branch lr_scale
- crates/ml/src/cuda_pipeline/dqn_utility_kernels.cu: lr_scale arg + apply
- crates/ml/src/cuda_pipeline/decision_transformer.rs: lr_scale=1.0
- crates/ml/src/trainers/dqn/trainer/training_loop.rs: write 4 ISV slots
- docs/dqn-wire-up-audit.md: 2026-05-11 audit entry
Verification (passing):
- cargo check -p ml --tests --features cuda: 0 errors
- cargo test -p ml --lib sp21_isv_slots --features cuda: 3/3 (bus bounds
+ slot uniqueness + branch index mapping)
- sp20_aggregate_inputs_test: 12/12
- sp20_phase1_4_wireup_test: 2/2
- sp20_emas_compute_test: 4/4
- sp20_controllers_compute_test: 7/7
- sp21_per_trade_predicted_q_test: 3/3
Total: 31 tests, 0 failures. Behavioral gate (per-branch LR divergence)
is the upcoming smoke training run.
After this commit (T2.2 Phases 5-7 + 8 + 4.5):
- Phase 5: E6 winner indices → PER priority bumps
- Phase 6: E7 hindsight → replay buffer injection
- Phase 7: E8 curriculum weights → segment sampling
- Phase 8: signal-drive remaining controller GAINS
- Phase 4.5: re-instate Pearl C engagement tracking for branches
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
21f911151a |
feat(sp21): T2.2 Phase 3 — wire E1 q_correction → Bellman target offset (atomic)
Wires `EnrichmentResult::q_correction` (Phase 1.5+2's E1 clamp(-mean(predicted_q − pnl), ±10.0) from the per-trade tape) into the training loss kernels' Bellman target computation. Both mse_loss_batched and c51_loss_kernel::block_bellman_project_f apply q_correction as additive shift on the Bellman target — blended (1-α)·MSE + α·C51 sees identical Q-target shifts (symmetric application per feedback_no_partial_refactor). ISV slot allocation (per pearl_controller_anchors_isv_driven + feedback_isv_for_adaptive_bounds — adaptive Q-target shift IS adaptive bound, lives in ISV bus): - New Q_CORRECTION_INDEX = 520 in cuda_pipeline::sp21_isv_slots - ISV_TOTAL_DIM 520 → 521 (bus extension) - Layout fingerprint adds SLOT_520_Q_CORRECTION - New module registered in cuda_pipeline/mod.rs - Compile-time test verifies SP21_SLOT_END <= ISV_TOTAL_DIM Sign convention: q_correction is the additive correction (E1 already negates the bias). predicted_q > pnl → q_correction < 0 → "lower Q-targets". predicted_q < pnl → q_correction > 0 → "raise Q-targets". Cold-start sentinel 0.0 (no trades yet) is no-op until first val pass writes (per pearl_first_observation_bootstrap). ABI surgery (minimal): - mse_loss_batched: ONE new scalar arg (float q_correction) at end - c51_loss_batched: ZERO new args (block_bellman_project_f already takes isv_signals; reads slot 520 from there) - launch_mse_loss: reads ISV[520] host-side via existing read_isv_signal_at, passes scalar - training_loop.rs (post-enrichment): writes result.q_correction to ISV[520] via existing write_isv_signal_at Files changed: - crates/ml/src/cuda_pipeline/sp21_isv_slots.rs (NEW): slot const + bounds-check tests - crates/ml/src/cuda_pipeline/mod.rs: pub mod sp21_isv_slots - crates/ml/src/cuda_pipeline/gpu_dqn_trainer.rs: ISV_TOTAL_DIM bump + fingerprint update + launch_mse_loss launcher arg - crates/ml/src/cuda_pipeline/mse_loss_kernel.cu: q_correction arg + application to target_q - crates/ml/src/cuda_pipeline/c51_loss_kernel.cu: read isv_signals[520] in block_bellman_project_f, add to t_z - crates/ml/src/trainers/dqn/trainer/training_loop.rs: write_isv_signal_at( Q_CORRECTION_INDEX, result.q_correction) post-enrichment - docs/dqn-wire-up-audit.md: 2026-05-11 audit entry Verification: - cargo check -p ml --tests --features cuda: 0 errors - cargo test -p ml --lib sp21_isv_slots --features cuda: 2/2 (bus bounds + slot uniqueness) - sp20_aggregate_inputs_test: 12/12 - sp20_phase1_4_wireup_test: 2/2 - sp20_emas_compute_test: 4/4 - sp20_controllers_compute_test: 7/7 - sp21_per_trade_predicted_q_test: 3/3 Total: 30 tests, 0 failures. Behavioral gate (q_correction → non- zero ISV[520] → Q-target shift) is the upcoming smoke training run. After this commit (T2.2 Phases 4-8): - Phase 4: E4 per-branch LR scaling via per-group Adam - Phase 5: E6 winner indices → PER priority bumps - Phase 6: E7 hindsight → replay buffer injection - Phase 7: E8 curriculum weights → segment sampling - Phase 8: signal-drive remaining controller GAINS Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
ab4a7db33c |
feat(sp20): c51_loss launcher aux_conf arg + Phase 5 gate tests
Threads `self.aux_conf_at_state_buf` into the `c51_loss_batched` launch
in `GpuDqnTrainer::launch_c51_loss`. Position matches the kernel's
appended trailing arg from the previous commit.
Tests added in `crates/ml-dqn/src/gpu_replay_buffer.rs::tests`:
- `aux_gate_high_confidence_passes_full_target` (CPU pure-math):
gate(aux_conf=0.5, threshold=0.10, temp=0.05) > 0.99 proves
high-confidence reward pass-through.
- `aux_gate_low_confidence_attenuates_reward` (CPU pure-math):
gate(aux_conf=0.02, threshold=0.10, temp=0.05) < 0.20 proves
the uncertain-state neutralizer semantic.
- `aux_gate_temp_floor_keeps_gate_finite` (CPU pure-math):
sweeps {temp, aux_conf, threshold} and asserts finite gate ∈ [0,1]
across the ISV-controllable parameter range — proves the
fmaxf(temp, 1e-3) floor keeps the kernel numerically safe.
- `aux_conf_direct_to_trainer_gather_populates_destination` (GPU
behavioral): wires a fresh CudaSlice<f32> as the trainer
destination, inserts 8 transitions with strictly-positive distinct
aux_conf values, samples 1, asserts the trainer destination
buffer post-sample holds a value from the inserted set (NOT the
alloc_zeros sentinel) — proves the direct-gather wiring actually
populates the trainer buffer with non-trivial data.
All 3 CPU math tests + 1 GPU integration test pass on RTX 3050.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
d3a057af4f |
feat(sp20): allocate trainer aux_conf_at_state_buf + wire PER direct-gather
Phase 5 plumbing — consumer-side wire-up. Allocates `aux_conf_at_state_buf: CudaSlice<f32>` ([batch_size]) on `GpuDqnTrainer`, exposes the raw_ptr via `aux_conf_at_state_buf_ptr()` and the `FusedTrainerCtx` delegating accessor `trainer_aux_conf_at_state_buf_ptr()`, and invokes `GpuReplayBuffer::set_trainer_aux_conf_ptr` at both fused_ctx init sites in training_loop.rs (init + re-init, atomic per `feedback_no_partial_refactor`). The PER `gather_f32_scalar` now writes the SAMPLED bar's per-batch aux_conf directly into the trainer's f32 buffer on every step — same direct-to-trainer pattern as the SP13 B1.1b `aux_nb_label_buf` (i32) wire-up immediately above the new call. The c51_loss_batched reward gate (lands in the next commit) reads this buffer to compute `gate = sigmoid((aux_conf - ISV[AUX_CONF_THRESHOLD]) / ISV[AUX_GATE_TEMP])` and applies `r_used = gate * reward` at the Bellman projection. `alloc_zeros` cold-start: 0.0 sentinel → at threshold ≈ 0.10 and temp ≈ 0.05 the gate is `sigmoid(-2) ≈ 0.12` → reward is mostly suppressed pre-population. This is the "graceful degradation" semantic from the Phase 3 Task 3.4 audit doc spec §4.4. Once PER's direct-gather populates from the producer ring on the first sample step, the per-bar aux_conf values drive the gate as designed. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
235e838422 |
feat(sp20): Phase 1.4 (Path C) — kernel-arg refactor + aggregation kernel + wire-up
Atomic Phase 1.4 commit per `feedback_no_partial_refactor`. Wires the
SP20 fused-producer chain (Stats → Aggregate → EMAs → Controllers) into
GpuExperienceCollector's per-rollout-step path with one new aggregation
kernel + Phase 1.2 EMA kernel-signature refactor + production wire-up
+ tests + audit/spec/plan amendments, all in one commit.
## Path C decision rationale
The Phase 1.2 EMA kernel originally took 9 scalar value-args. The Phase
1.4 wire-up site (per-rollout-step in GpuExperienceCollector) needs to
feed per-env GPU-resident trade signals (trade_close_per_sample,
step_ret_per_sample, hold_at_exit_per_sample, packed factored
actions_out) into the kernel — forbidden via host sync
(`feedback_cpu_is_read_only`) and via memcpy_dtoh/htod
(`feedback_no_htod_htoh_only_mapped_pinned`). Path C refactors the EMA
kernel to take a device struct pointer (`SP20EmaInputs* ema_inputs`) +
adds an aggregation kernel that writes the struct on the GPU. Phase 1.2
had zero production consumers yet, so the sig change + Phase 1.4 wire-up
land atomically.
## What this commit contains
1. **EMA kernel signature refactor** (Phase 1.2 → Path C):
`sp20_emas_compute_kernel.cu` swaps 9 scalar args for a single
`const SP20EmaInputs* __restrict__ ema_inputs` device pointer.
Math semantics bit-identical. Rust `EmaInputs` mirrored as
`#[repr(C)]` byte-for-byte; new `pack_inputs_into_f32_view` helper
for tests + collectors that fill the struct via a mapped-pinned
f32-aliased buffer.
2. **New aggregation kernel** (`sp20_aggregate_inputs_kernel.cu`):
Per-env arrays (sliced to current rollout step) + sp20_stats outputs
(p50/std) + aux_dir_acc_reduce_kernel output [0] → SP20EmaInputs
struct. Block tree-reduce (4 stripes × bdim × 4 bytes shmem); no
atomicAdd. Aggregation rules per spec §4.5:
- is_close = OR over envs
- is_win = (≥ 0.5 of closed envs were wins)
- trade_duration = round(mean(hold_at_exit) over closed envs)
- action_is_hold = strict majority (count*2 > n_envs) HOLD
- alpha = 0.0 (Phase 2 forward ref — reward kernel)
- per_bar_hold_reward = 0.0 (Phase 3.2 forward ref — Hold-cost dual)
- aux_logits_p50/std/dir_acc forwarded from upstream
3. **Production wire-up in GpuExperienceCollector**:
4 kernel handles + 4 mapped-pinned buffers (struct fields, alloc,
init); per-rollout-step launch sequence after env_step (step 5c):
`Stats → Aggregate → EMAs → Controllers`. All on the same stream;
stream-implicit producer→consumer ordering. Gated on
`isv_signals_dev_ptr != 0 && trainer_params_ptr != 0` (matches the
existing SP14-β EGF chain pattern).
4. **Fold-boundary reset**:
3 new RegistryEntry records (sp20_ema_inputs_buf,
sp20_emas_internal_buf, sp20_emas_obs_count_buf) + 13 new dispatch
arms in training_loop.rs (10 SP20 ISV slot resets + 3 buffer resets).
Closes the pre-existing `every_fold_and_soft_reset_entry_has_dispatch_arm`
regression (was failing on fresh check after the SP20 ISV slot
registrations landed in commit
|
||
|
|
f5eed1fa79 | spec(sp20): reserve 10 ISV slots [510..520) for WR-first reward optimization | ||
|
|
37964ae2a4 |
feat(sp19 commit a): ISV slot reservations [507..510) for multi-horizon reward blend
Path (B) Commit A — additive infrastructure for the SP19 producer-side
multi-horizon reward augmentation. Three ISV slots reserve indices for
a future Path (A) refactor that would let a controller drive per-batch
horizon weights; producer (Commit B) hardcodes equal-thirds 1/3 each at
fxcache-write time so the slot reservations are dormant in this branch.
Pure additive — changes NO runtime behaviour. The dispatch arms write
the equal-thirds sentinel at every fold boundary, but no kernel
consumes the slots. This is forward-compatible reservation per
feedback_isv_for_adaptive_bounds, NOT a half-fix; the producer wiring
is complete (next commit), the slot reservations are documented
forward-compatibility for the TARGET_DIM-bumping refactor.
Slot allocation:
- 507 REWARD_HORIZON_WEIGHT_1BAR_INDEX sentinel = 1/3
- 508 REWARD_HORIZON_WEIGHT_5BAR_INDEX sentinel = 1/3
- 509 REWARD_HORIZON_WEIGHT_30BAR_INDEX sentinel = 1/3
Touches:
- crates/ml/src/cuda_pipeline/sp14_isv_slots.rs: 3 slot constants +
sentinel + range markers + sp19_reward_horizon_slot_layout_locked +
all_sp19_slots_fit_within_isv_total_dim lock tests.
- crates/ml/src/cuda_pipeline/gpu_dqn_trainer.rs: ISV_TOTAL_DIM 507 →
510 + layout_fingerprint_seed extension (3 SLOT_* lines + a
SP19_PRODUCER_HARDCODED_HORIZON_BLEND token registering producer-time
consumption).
- crates/ml/src/cuda_pipeline/state_layout.cuh: 3 #define mirrors of
the Rust slot indices + SENTINEL_REWARD_HORIZON_WEIGHT_DEFAULT macro.
- crates/ml/src/trainers/dqn/state_reset_registry.rs: 3 RegistryEntry {
FoldReset } with the "SP19 Path (B) reservation" marker + lock-test
expansion 24 → 27 entries.
- crates/ml/src/trainers/dqn/trainer/training_loop.rs: 3 dispatch arms
in reset_named_state writing SENTINEL_REWARD_HORIZON_WEIGHT_DEFAULT.
- docs/dqn-wire-up-audit.md: SP19 Commit A entry documenting the
reservation rationale + verification steps.
Verification:
SQLX_OFFLINE=true CUDA_COMPUTE_CAP=86 cargo check --workspace clean
cargo test -p ml --lib sp19_reward_horizon_slot_layout_locked passes
cargo test -p ml --lib all_sp19_slots_fit_within_isv_total_dim passes
cargo test -p ml --lib sp18_fold_reset_entries_present passes (27)
cargo test -p ml --lib every_fold_and_soft_reset_entry_has_dispatch_arm passes
bash scripts/audit_sp18_consumers.sh --check exit 0
Atomic-refactor invariant (HARD — feedback_no_partial_refactor): NO
L40S DISPATCH between Commit A and Commit B. The producer-side blend
+ fxcache version bump + behavioural test land in the next commit on
this branch.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
5ea5aa9b8e |
feat(sp18 v2 P3.T1-T5): adaptive HOLD_REWARD_POS/NEG_CAP producer kernel
Lifts the Phase 2 caps from sentinel-driven cold-start (5.0/-10.0
fixed) to p99(|step_ret|) over Long/Short trade closes × 1.5 safety
factor with Wiener-optimal alpha blend. The Phase 2 consumer
(compute_sp18_hold_opportunity_cost) sees producer-driven slots
[483]/[484] from epoch 1 onward; sentinel branch in the consumer
remains as the cold-start fallback path (zero-trade-close epoch ⇒
producer early-returns ⇒ slots stay at sentinel ⇒ consumer falls
through to the +5/-10 macro defaults — bit-identical to pre-Phase-3).
Mirrors SP14 P0-A reward_cap_update_kernel structural template with
three differences: (1) filter (is_close && step_ret != 0) — both
winners AND losers (the consumer needs the magnitude scale of *all*
Long/Short closes); (2) Welford-derived Wiener-α (slots [487..493))
replaces fixed α=0.01, with floor at WELFORD_ALPHA_MIN=0.4 per
pearl_wiener_alpha_floor_for_nonstationary (the policy-realised
distribution is intrinsically non-stationary as the policy adapts);
(3) bounds [0.5, 50.0] (vs. position-side [1.0, 50.0]).
Atomic single-commit per feedback_no_partial_refactor:
- crates/ml/src/cuda_pipeline/hold_reward_cap_update_kernel.cu (NEW)
- crates/ml/build.rs cubin manifest entry
- HoldRewardCapUpdateOps in gpu_aux_trunk.rs (new struct + impl)
- HOLD_REWARD_CAP_UPDATE_CUBIN static + struct field +
launch_hold_reward_cap_update method + constructor instantiation +
field-init in gpu_dqn_trainer.rs (5 sites)
- Per-epoch boundary launch in training_loop.rs right AFTER
launch_reward_cap_update (shared step_ret/trade_close source buffers,
independent ISV slot pairs)
- HEALTH_DIAG[N]: hold_reward_cap [pos={:.4} neg={:.4} fire_rate={:.4}]
- 3 GPU oracle tests (T5 producer-drives-slots, Pearl-A REPLACE,
no-closes preserves-isv) — all pass on local RTX 3050 Ti
- Phase 3 close-out sections in docs/sp18-wireup-audit.md and
docs/dqn-wire-up-audit.md
Pearls applied: feedback_no_atomicadd, pearl_first_observation_bootstrap,
pearl_wiener_optimal_adaptive_alpha, pearl_wiener_alpha_floor_for_nonstationary,
pearl_no_host_branches_in_captured_graph, pearl_symmetric_clamp_audit,
pearl_audit_unboundedness_for_implicit_asymmetry (NEG = -2 × POS at
producer time, single source of truth), feedback_isv_for_adaptive_bounds,
pearl_fused_per_group_statistics_oracle.
Validation: cargo check --workspace clean; 3 GPU oracle tests pass on
local RTX 3050 Ti; scripts/audit_sp18_consumers.sh --check exits 0
(no fingerprint drift in tracked sections).
Plan: docs/superpowers/plans/2026-05-08-sp18-reward-shape-hold-attractor.md
Phase 4-5 (B-leg target-net forward + q_next replacement) follows.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
e877e8aefe |
chore(sp18): archaeology cleanup — strip post-deletion markers + ISV_TOTAL_DIM giant docstring
Pure-comment follow-up to SP18 Phase 1 atomic deletion ( |
||
|
|
3c318953a7 |
feat(sp18 v2 P1.T3-T6): atomic delete SP13/SP16 hold_cost_scale chain
Deletes the entire reactive Hold-cost controller chain from SP13 P0a / SP16
P2 / SP16 T3 atomically — replaced by the SP18 D-leg structural opportunity-
cost reward landing in Phase 2.
Per DD7(c): observation chain preserved (slots 380, 382). Slot 380 becomes
constructor-init-only diagnostic at HOLD_COST_BASE=0.005, never updated.
18 consumer sites deleted (8 from plan checklist + 10 from A1-A10 audit
expansion):
Production code:
- 3 reward subtraction sites (replaced with Phase 2 placeholder comments)
- hold_cost_scale_update_kernel.cu (file deleted)
- HoldCostScaleUpdateOps struct in gpu_aux_trunk.rs (~120 lines)
- launch_hold_cost_scale_update in gpu_dqn_trainer.rs (forwarder)
- Host controller block in training_loop.rs (~60 lines)
- Slot constants [461..468) in sp14_isv_slots.rs
- HOLD_COST_CONTROLLER_GAIN/FLOOR/CEIL constants in sp13_isv_slots.rs
- state_layout.cuh mirror constants for [461..468)
- build.rs cubin manifest entry
- 7 fold-reset registry entries + 7 dispatch arms
Tests:
- sp14_oracle_tests.rs lines 2173-2926 (11 tests + helpers, ~750 lines)
— these include_bytes! the deleted cubin and cannot survive
- lock_sp18_v2_pp4_retired_chain test (deleted, no inverse-contract
replacement — locking a deletion is pointless ceremony per user call)
Layout fingerprint: range [461..468) is RESERVED gap (SP14-C.1 pattern;
ISV_TOTAL_DIM stays at 507 for checkpoint compatibility). Layout fingerprint
seed updated atomically: HOLD_COST_SCALE + HCS_* slots replaced with
RESERVED_GAP_461_TO_468=SP18_P1_RETIRED marker.
INTERIM STATE: 3 reward sites are now missing the per-bar Hold cost. This
is forbidden to L40S-dispatch until Phase 2 lands the
compute_sp18_hold_opportunity_cost device fn that replaces the deleted
subtraction. The audit doc records this hold.
Validation: cargo check --workspace clean; cargo test -p ml --lib passes
(state_reset_registry::every_fold_and_soft_reset_entry_has_dispatch_arm,
sp16_t3_wiener_welford_slot_layout_locked, sp18_combined_slot_layout_locked,
sp18_shrink_perturb_slot_layout_locked, sp18_fold_reset_entries_present —
all OK). Audit fingerprint: per-bar Hold-cost subtraction sites count
13 → 0 (smoking gun for complete deletion).
Plan: docs/superpowers/plans/2026-05-08-sp18-reward-shape-hold-attractor.md
Audit: docs/sp18-wireup-audit.md (full A1-A10 expansion + post-deletion
fingerprint)
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
0b737ec30a |
feat(sp18-v2): ISV-adaptive shrink-and-perturb at fold transition
Replaces the hardcoded `α=0.8 / σ=0.01` constants in `reset_for_fold`'s
`shrink_and_perturb` call with ISV-driven adaptive bounds per
`feedback_isv_for_adaptive_bounds` and `feedback_adaptive_not_tuned`.
The driving signal is the relative weight drift `||current - best||₂ /
||best||₂` already computed each epoch by the SP18 v2 weight-drift
diagnostic kernel (commit
|
||
|
|
27f8e332da |
plan(sp18 v2): Phase 0 Task 0.2 — B-leg V_SHARE trajectory + TD-error magnitude diagnostic
Adds the SP18 v2 Phase 0 B-leg observability scaffold per the plan's
Task 0.2 spec:
- New kernel `td_error_mag_ema_kernel.cu` (single-block × 256 threads):
reads `td_errors_buf [B]` (already populated by C51/MSE loss for PER
priority recomputation, post-train-step), block tree-reduces
`mean(|td_errors[b]|)` (no atomicAdd), and EMA-blends into
`ISV[TD_ERROR_MAG_EMA_INDEX=493]` via Pearl-A first-observation
bootstrap (sentinel 0.0 → REPLACE) + fixed α=`WELFORD_ALPHA_MIN=0.4`
per `pearl_wiener_alpha_floor_for_nonstationary`. The TDB_* Welford
accumulators in slots [498..504) are RESERVED for the Phase 4
q_next_target Wiener-α chain — not used in Phase 0.
- Cubin manifest entry in `crates/ml/build.rs` + `TD_ERROR_MAG_EMA_
CUBIN` re-export in `gpu_dqn_trainer.rs`.
- `sp18_td_error_mag_ema_kernel` field on `GpuDqnTrainer` + cubin load
on the trainer's stream + `launch_sp18_td_error_mag_ema_update()`
cold-path launcher + `read_sp18_td_error_mag_ema()` convenience
wrapper.
- New `sp18_v_share_history: [[f32; 4]; 5]` field on `DQNTrainer` —
fixed-size ring buffer of the last 5 epochs of per-branch V_SHARE
EMA readings (slots [478..482) per SP17 Phase 3.2). Initialised to
`[[NaN; 4]; 5]`; epochs 0–3 emit `nan` as the slope and skip the
ISV write; epoch 4 onward computes `(EMA[now] - EMA[now-4]) / 4`
per branch and writes the dir-branch slope to
`ISV[V_SHARE_TREND_DIAG_INDEX=496]`.
- Two new HEALTH_DIAG lines in `training_loop.rs` at the per-epoch
boundary (right after the SP18 reward_decomp line):
HEALTH_DIAG[N]: v_share_traj [dir_slope=X mag_slope=Y ord_slope=Z urg_slope=W]
HEALTH_DIAG[N]: td_error_pre [magnitude_ema=X]
V_SHARE slope is host-side computation against the ring buffer
(`(now - now_m4) / 4` per branch). TD-error magnitude is post-blend
ISV slot 493 read (producer fires inside `read_sp18_td_error_mag_
ema()`). Pre-fix baseline for the B-DD9 ratio gate
(`avg(|TD-error|) ratio post-fix / pre-fix ∈ [0.5, 5.0]`).
- New GPU oracle tests in `crates/ml/tests/sp18_hold_reward_oracle_
tests.rs`:
* `td_error_mag_ema_pearl_a_bootstrap` — synthetic td_errors with
closed-form mean(|td|)=1.125; pre-populate slot at sentinel;
assert post-launch slot equals the mean (Pearl-A direct-replace).
* `td_error_mag_ema_blend_post_bootstrap` — synthetic td_errors
with mean(|td|)=0.5; pre-populate slot at non-sentinel 1.0;
assert blend equals `(1 - 0.4) × 1.0 + 0.4 × 0.5 = 0.8`.
Pure observability — no production-path consumer in this commit. No
reward changes, no Bellman target changes, no kernel modifications to
the action-selection or training paths. Per `feedback_no_partial_
refactor` the kernel + cubin manifest + buffer + launcher + ring
buffer + HEALTH_DIAG emit + GPU oracle tests all land atomically.
Verification:
SQLX_OFFLINE=true CUDA_COMPUTE_CAP=86 cargo check --workspace clean.
All 4 GPU oracle tests pass on RTX 3050 Ti (2.09s):
reward_decomp_per_action_gpu_oracle, reward_decomp_empty_bin,
td_error_mag_ema_pearl_a_bootstrap, td_error_mag_ema_blend_post_bootstrap.
Existing slot lock + state_reset_registry tests still pass.
Plan: docs/superpowers/plans/2026-05-08-sp18-reward-shape-hold-attractor.md
§ Phase 0 Task 0.2.
Audit: docs/dqn-wire-up-audit.md § "SP18 v2 Phase 0 Task 0.2".
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
15b50ac38f |
plan(sp18 v2): Phase 0 Task 0.1 — D-leg per-action reward decomposition diagnostic
Adds the SP18 v2 Phase 0 D-leg observability scaffold per the plan's
"Phase 0 — diagnostic emit (NO functional change)" task:
- New kernel `reward_decomp_diag_kernel.cu`: block tree-reduce
(4 blocks × 256 threads, one block per direction-axis bin) reading
`reward_components_per_sample [N×6]` + `actions_out [N]` and emitting
5 per-bin stats (mean r_micro / mean r_opp_cost / mean r_popart /
mean |reward| / fire_rate) into a 20-float row-major output. Bin
order: Hold(0)→Long(1)→Short(2)→Flat(3); col order: micro→opp→
popart→abs→fire. Empty-bin guard emits 0.0 (NOT NaN) per the
consumer-side KILL CRITERION arithmetic contract.
- Cubin manifest entry in `crates/ml/build.rs` + `REWARD_DECOMP_DIAG_
CUBIN` re-export in `gpu_dqn_trainer.rs`.
- 20-float `MappedF32Buffer sp18_reward_decomp_diag_buf` field on
`GpuDqnTrainer` + accessor pair (`sp18_reward_decomp_diag_dev_ptr`
for the writer-side launcher; `read_sp18_reward_decomp_diag` for the
HEALTH_DIAG reader). Buffer is constructor-zeroed so cold-start
HEALTH_DIAG emits a deterministic zero block.
- `sp18_reward_decomp_diag_kernel` field on `GpuExperienceCollector` +
cubin load on the collector's stream + `launch_sp18_reward_decomp_
diag(n, b1, b2, b3, out_dev_ptr)` launcher. Wired in
`training_loop.rs` at the per-step boundary, BEFORE
`launch_reward_component_ema_inplace` (which `memset_zeros` the
source buffer after consuming it) per `pearl_canary_input_freshness_
launch_order`.
- New per-epoch HEALTH_DIAG line emit at the existing per-epoch
boundary (after the SP17 dueling line):
HEALTH_DIAG[N]: reward_decomp [hold(micro=X opp=Y popart=Z abs=W
fire=F) long(...) short(...) flat(...)]
Reads the mapped-pinned 20-float diag buffer directly via the
collector→trainer host_ptr — no DtoH copy.
- New `crates/ml/tests/sp18_hold_reward_oracle_tests.rs`:
* `reward_decomp_per_action_cpu_oracle` (CPU oracle pinning the
per-bin reduction math against a 4-sample synthetic batch).
* `reward_decomp_per_action_gpu_oracle` (GPU oracle, ignored unless
`--ignored`; asserts kernel matches CPU oracle bit-for-bit within
1e-6 f32 budget).
* `reward_decomp_empty_bin_emits_zero_not_nan` (empty-bin contract
guard).
Pure observability — no production-path consumer in this commit. No
reward changes, no Bellman target changes, no kernel modifications to
the action-selection or training paths. Per
`feedback_no_partial_refactor` the kernel + cubin manifest + buffer +
launcher + production wire-up + HEALTH_DIAG emit + GPU oracle test all
land atomically.
Verification:
SQLX_OFFLINE=true CUDA_COMPUTE_CAP=86 cargo check --workspace clean.
CPU oracle test passes; GPU oracle + empty-bin guard both pass on
RTX 3050 Ti (2.13s).
Plan: docs/superpowers/plans/2026-05-08-sp18-reward-shape-hold-attractor.md
§ Phase 0 Task 0.1.
Audit: docs/dqn-wire-up-audit.md § "SP18 v2 Phase 0 Task 0.1".
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
16100912c3 |
plan(sp18 v2): allocate 22 ISV slots [483..505) for D-leg + B-leg
Adds combined SP18 D-leg Hold-reward adaptive-cap slots [483..493) + B-leg TD(λ) Q(s') bootstrap diagnostics + PopArt reset flag at [493..505) per spec DD6 (D-leg 10 slots) + B-DD13 (B-leg 12 slots). D-leg [483..493) — 10 slots: 483 HOLD_REWARD_POS_CAP (sentinel 5.0, bounds [0.5, 50]) 484 HOLD_REWARD_NEG_CAP (sentinel -10.0) 485 HOLD_REWARD_DECOMP_DIAG (sentinel 0.0) 486 HOLD_OPP_COST_FIRE_RATE_EMA (sentinel 0.0) 487..493 HRC_* Welford accumulators (mirrors SP16 T3 HCS_* pattern) B-leg [493..505) — 12 slots: 493 TD_ERROR_MAG_EMA (HEALTH_DIAG B-DD9 ratio gate input) 494 Q_NEXT_TARGET_P99 (target-Q bootstrap bound check) 495 Q_NEXT_MINUS_REWARD_P99 (sanity: should be O(1) post-fix) 496 V_SHARE_TREND_DIAG (B-leg synergy probe) 497 POPART_RESET_FLAG (sentinel 1.0 — one-shot, B-DD11) 498..504 TDB_* Welford accumulators (mirrors HRC_* pattern) 504 RESERVED (B-leg follow-up) Pearl-A first-observation bootstrap sentinels match position-side SP14 P0-A REWARD_POS_CAP_ADAPTIVE pattern (POS=5.0, NEG=-10.0). Slot 497 POPART_RESET_FLAG sentinel = 1.0 per B-DD11 — host writes 1.0 once at first SP18 epoch, kernel zeroes after consuming, gating the per-fold PopArt slot 63 EMA reset at the SP18 deployment boundary. ISV_TOTAL_DIM bumped 483 → 505; layout_fingerprint_seed updated with all 22 new slot names; state_layout.cuh C-side mirror in lockstep (continues SP14-P0A/P1/audit-fix-4A/4B mirror precedent — SP16/SP17 slots intentionally not mirrored per existing pattern, only SP18 gets fresh mirror entries). `feedback_no_partial_refactor`: both legs share an ISV section + a single fingerprint bump; SP13 [380..383) and SP16 [461..474) slots remain ALLOCATED but RETIRED in PP.4 (sentinel 0.0, no producer launch — RESERVED-gap pattern from SP14-C.1 preserves checkpoint compatibility). Audit doc updated per Invariant 7 with Pre-Phase PP.2 entry. Spec: docs/superpowers/specs/2026-05-08-sp18-reward-shape-hold-attractor-design.md Plan: docs/superpowers/plans/2026-05-08-sp18-reward-shape-hold-attractor.md Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
b6b17d46bb |
feat(sp17-3.2): V_share + advantage_clip_bound producers + extended emit
Phase 3.2 lands the remaining two SP17 dueling-Q diagnostic producers
atomically with kernel + launcher + Rust wrapper + extended HEALTH_DIAG
emit + GPU oracle tests per `feedback_wire_everything_up`.
V_share[d] = |E[V]| / (|E[V]| + |E[A_centered, picked]|)
where picked = argmax_a Σ_z A_raw[i, a, z] (max-Q semantic — tractable
per-batch without depending on actions_history_buf which is collector-
time state stale relative to the cuBLAS forward at HEALTH_DIAG cadence).
Pearl-A bootstrap (sentinel 0.5) + α=WELFORD_ALPHA_MIN=0.4 + bilateral
[0, 1] clamp per `pearl_symmetric_clamp_audit`. 4 blocks × 256 threads.
advantage_clip_bound = p99(|A_centered|) × ADVANTAGE_CLIP_SAFETY_FACTOR=1.5
via sp4_histogram_p99 (block tree-reduce + per-warp tile binning, NO
atomicAdd per `pearl_fused_per_group_statistics_oracle`). EMA α=0.01
slow per-fold + bilateral clamp [0.1, 100.0] per
`pearl_symmetric_clamp_audit`. Pearl-A bootstrap (sentinel 1.0).
Single block × 256 threads + flat |A_centered| scratch buffer
(mapped-pinned, sized to B × Σ_d b_d × NA).
Observability-only — the actual clipping wire-up is Phase 5 follow-up.
The Phase 1 mean-zero contract (commits eabcf8d52..6f53d676f) makes
A_centered a meaningful signal; this commit observes it.
Extended HEALTH_DIAG line:
HEALTH_DIAG[N]: dueling [v_share=(d=X m=Y o=Z u=W)]
[a_var=(d=A m=B o=C u=D)] [clip=K]
GPU oracle tests on RTX 3050 Ti (all pass, 13/13 SP17 tests):
- v_share_per_branch_matches_closed_form: synthetic V=2.0 + linear A
per branch; closed-form V_share = 2/(2 + |K_d × (n_d-1)/2|);
ε=1e-4. Pearl-A bootstrap REPLACES on first launch.
- advantage_clip_bound_tracks_p99_safety: synthetic A with action-
dominant + per-(i,z) jitter (the jitter is REQUIRED — pathologically
lockstep values undercount in sp4_histogram_p99's non-atomic warp
tile binning per the kernel's documented "1/(256×32) loss for
uniformly distributed signals" qualifier; concentrated values violate
the assumption. Real |A_centered| in production is continuous, so
this is a test-data-only effect.) ε=0.20 (jitter + linear histogram
quantization).
Plan: docs/superpowers/plans/2026-05-08-sp17-dueling-q-network.md
Phase 3.2.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
1e70cd5e59 |
feat(sp17-3.1): A_var_ema per-branch producer + HEALTH_DIAG emit
Phase 3 of SP17 dueling-Q identifiability — first of three diagnostic
producers landed atomically with kernel + launcher + Rust wrapper +
HEALTH_DIAG emit + GPU oracle test per `feedback_wire_everything_up`.
Per branch d ∈ {dir, mag, ord, urg}:
Var_d = (1/(B × n_d × NA)) Σ_{i, a, z} (A[i, a, z] − mean_a A[*, z])²
Block tree-reduce (no atomicAdd, `feedback_no_atomicadd`); 4 blocks ×
256 threads. Pearl-A first-observation bootstrap (sentinel 0.0 →
REPLACE on first launch); steady-state α = WELFORD_ALPHA_MIN=0.4 per
`pearl_wiener_alpha_floor_for_nonstationary` — the structural-control
floor preserves catch-up bandwidth without storing 24 Welford
accumulator slots for a cold-path-cadence diagnostic.
Cold-path emit: single launch per HEALTH_DIAG cadence (epoch boundary)
right after `v_a_means`. New line:
HEALTH_DIAG[N]: dueling [a_var=(d=X m=Y o=Z u=W)]
The line will be extended with V_share + advantage_clip_bound in
Phase 3.2, then finalised in Phase 3.3.
GPU oracle test on RTX 3050 Ti: synthetic A constructed so each branch
d has a closed-form Var(A_centered); kernel readback matches expected
value within ε=1e-4 (f32 rounding budget for ~8×n×51 accumulator
length).
Plan: docs/superpowers/plans/2026-05-08-sp17-dueling-q-network.md
Phase 3.1.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
3107edb8f7 |
feat(sp17): Thompson V-wire-in (Commit C)
User design call DD10 option (b): Thompson direction selector now reads
softmax(V[z] + (A[a, z] - mean_a A[*, z])) instead of pre-SP17
softmax(A[a, z]). Without V wired in, action selection responded to a
DIFFERENT distribution than compute_expected_q's E[Q] (the Bellman-target
ranking) — the very pathology SP17 is fixing at the action layer.
Architectural change closes the contract gap atomically across:
Kernel signature (`experience_action_select`):
- Added `const float* __restrict__ v_logits_dir` after `b_logits_dir`.
NULL is invalid (no fallback) — kernel hard-requires V.
- New per-thread `a_mean_per_atom_dir[THOMPSON_MAX_ATOMS]` reduction
computes mean A across the 4 direction actions per atom z (sample-local
register, no atomicAdd).
- Pass 1 (E[Q] for temperature blend + conviction): builds per-action
`combined_logits_d[z] = V[z] + (A[a,z] - mean_a A[*,z])` and feeds
`softmax_c51_inline` instead of raw `b_logits_dir + d * n_atoms`.
- Pass 2 (Thompson sample): same combined-logits rebuild per direction.
- `softmax_c51_inline` device helper UNCHANGED — keeping centering at
caller maintains a narrower contract; thompson_test_kernel and other
potential callers stay unaffected.
- THOMPSON_MAX_ATOMS=128 ceiling preserved; __trap() on overflow.
QValueProvider trait extension:
- `compute_q_and_b_logits_to` now takes `v_logits_out_ptr: u64` and
DtoD-copies `on_v_logits_buf` per sub-iteration (atomic per
feedback_no_partial_refactor — every consumer migrates in lockstep).
Trainer + evaluator wire-up:
- `gpu_dqn_trainer.rs::on_v_logits_buf_ptr() -> u64` (new pub fn, mirror
of existing `on_b_logits_buf_ptr`).
- `gpu_backtest_evaluator.rs::chunked_v_logits_buf` field allocated
[chunk_n * NA + 32*3] (cuBLAS tail-safety pad).
- Both call sites (collector + evaluator) pass v_logits arg in launch.
GPU oracle test (RTX 3050 Ti, 5/5 PASS):
thompson_direction_select_reads_v_logits — runs production cubin
twice on identical A logits with V=[0,0,0] vs V=[10,0,0]; asserts
q_gap_v_dominant < 50% of q_gap_v_zero. If V is being IGNORED
(regression), both runs produce IDENTICAL q_gaps and the test fails
with a clear message. Uses MappedF32Buffer / MappedI32Buffer per
feedback_no_htod_htoh_only_mapped_pinned.
Note on the plan's "raw argmax = Hold but centered argmax = Long" test:
Mathematical analysis shows softmax-with-constant-shift preserves
action ordering (mean subtraction adds the same per-atom constant to
every action's logits), so the plan's specific assertion isn't
algebraically constructable with simple A/V. The replacement test
(V-dependence of q_gap) is more sensitive — it fails on the actual
regression case (V ignored ⇒ identical q_gaps) the plan was trying
to detect.
Verification:
cargo check --workspace → clean
cargo test sp17_dueling_oracle_tests --features cuda
-- --ignored → 5/5 PASS
⚠ INTERIM STATE: aux-CQL barrier_gradient_direction +
ib_gradient_direction still read raw advantage. Commit D closes them;
Commit E annotates the pre-SP17 c51_loss/c51_grad already-centered sites.
Plan: docs/superpowers/plans/2026-05-08-sp17-dueling-q-network.md
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
cc4746b48d |
plan(sp17): HEALTH_DIAG v_a_means baseline (PRE-CENTERING)
Adds per-epoch readout of the V/A breakdown so we can observe the pre-SP17 baseline distribution before the identifiability projection lands. Also instruments the kill criterion for Phase 0: HEALTH_DIAG[N]: v_a_means [v=X a_dir=Y a_mag=... a_ord=... a_urg=...] If train-multi-seed-b5gmp's Q(Flat) over-attribution is V-driven, V should be elevated (~0.4+) while a_dir is small. If V is balanced and A is doing the work, dueling cannot help and we abort SP17. Block-tree-reduce kernel \`v_a_means_diag_kernel\` (no atomicAdd per \`feedback_no_atomicadd\`); MappedF32Buffer per \`feedback_no_htod_htoh_only_mapped_pinned\`. Plan: docs/superpowers/plans/2026-05-08-sp17-dueling-q-network.md Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
a225926e5f |
plan(sp17): allocate ISV slots [474..483) for dueling-Q diagnostics
9 slots: A_var_ema × 4 branches (per-branch advantage variance EMA) + V_share × 4 branches (V/(V+A) magnitude share) + adaptive advantage_clip_bound. All Pearl-A first-observation bootstrap with fold-reset registry entries + dispatch arms in trainer/training_loop.rs::reset_named_state. Bump ISV_TOTAL_DIM 474 → 483 + layout fingerprint seed. Per `feedback_isv_for_adaptive_bounds`: advantage_clip_bound is producer-tracked from p99(|A_centered|) × 1.5 safety factor, never hardcoded. Bounds [0.1, 100.0] are Category-1 dimensional safety floors. No consumer change in this commit (additive infrastructure). Phase 1 mean-centering producer (atomic across compute_expected_q + c51_loss + c51_grad + mag_concat_qdir) lands in the next 4 commits. Plan: docs/superpowers/plans/2026-05-08-sp17-dueling-q-network.md Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
641aa0dfde |
fix(sp16-t3): Wiener-optimal adaptive α per pearl — hold_cost_scale + min_hold_temperature
Per train-multi-seed-hjzss validation: SP16 T1+T2 chain was structurally
landed but BEHAVIORALLY INERT in 5-epoch smoke (bit-identical to pfh9n
baseline through epoch 3). Root cause: hardcoded `alpha = 0.05f` in both
producer kernels violates feedback_isv_for_adaptive_bounds AND prevents
convergence in short runs (~60 epochs needed from cold start).
Fix per pearl_wiener_optimal_adaptive_alpha:
α = diff_var / (diff_var + sample_var + ε)
Where sample_var = running variance of target signal (Welford accumulator)
and diff_var = running variance of consecutive one-step differences.
Cold-start: target jumps 1.0 → 6.4 → 7.0 → high diff_var → α ≈ 0.6+
→ near-bootstrap responsiveness in epochs 1-3
Steady-state: signal stabilizes → diff_var drops → α decays naturally
→ smoothing emerges without hardcoded constant
Adds 12 new ISV slots (6 per producer):
- HCS_TARGET_MEAN/M2, HCS_DIFF_MEAN/M2, HCS_PREV_TARGET, HCS_SAMPLE_COUNT
- MHT_TARGET_MEAN/M2, MHT_DIFF_MEAN/M2, MHT_PREV_TARGET, MHT_SAMPLE_COUNT
ISV_TOTAL_DIM 462 → 474.
Both kernels migrated atomically. Pearl-A bootstrap preserved (sentinel
on prev_blended triggers REPLACE; cold-start α=1.0 when N<3 samples).
Defensive bounds [WELFORD_ALPHA_MIN=0.01, WELFORD_ALPHA_MAX=0.95] on the
Wiener-derived α to guard against denormal/underflow corner cases.
HEALTH_DIAG[N] emit extended with `alpha=...` and `sample_count=...` for
direct trajectory observation in validation smoke.
Behavioral tests verify:
- α high during signal jumps (>0.3 at epoch 3 post-cold-start)
- α low in steady state (mean tail α<0.4 under converging signal)
- Pearl-A bootstrap fires on first observation (Welford state advances
regardless of REPLACE branch)
- α stays within [WELFORD_ALPHA_MIN, WELFORD_ALPHA_MAX] over 50 epochs
(post-cold-start; cold-start α=1.0 by design)
- No 0.05f hardcoded literal remains in blend math (regression-locked
via host-only string scan)
5 GPU + host tests pass: sp16_phase3_alpha_high_during_signal_jump,
alpha_low_in_steady_state, pearl_a_bootstrap_first_obs,
alpha_naturally_bounded, no_hardcoded_alpha. sp14 + sp15 oracle suites
unchanged (34 GPU tests + 4 host tests).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
1a3bcf97b8 |
feat(sp16-p2): adaptive Hold cost scale via ISV[461]
Per train-multi-seed-pfh9n post-mortem: observed_hold_rate climbed 0.25 → 0.52 across training while cost penalty (~0.006) was 100× smaller than per-bar reward magnitudes (popart=0.97, cf=0.65). Hold action was effectively free, allowing Q(Hold) to dominate via structural low-variance bias. Fix: scale Hold cost adaptively. ISV[HOLD_COST_SCALE_INDEX=461] tracks: scale = clamp(1.0 + 24.0 × max(0, observed - target) / max(target, 0.01), 1.0, 25.0) Effective cost at 100% overrun (observed=2× target): 0.006 × 25 = 0.15, competitive with per-bar reward magnitudes (~0.01-0.1). At/below target: scale = 1.0 (no extra penalty). Pearl-A bootstrap + Welford slow EMA (α=0.05). Mirrors T1's MIN_HOLD_TEMPERATURE pattern (same input signals: ISV[382] observed, ISV[381] target). Producer: hold_cost_scale_update_kernel.cu — single-thread cold-path, per-epoch boundary, AFTER MIN_HOLD_TEMPERATURE in training_loop.rs. Consumer migration (atomic per feedback_no_partial_refactor): 3 sites in experience_kernels.cu — segment_complete branch (line ~3089), per-bar positioned-Hold branch (line ~3553), per-bar flat-Hold branch (line ~3617). Cold-start fallback: scale=1.0 when slot ≤ 0 or out-of-bounds (bit-identical pre-Phase-2 cost magnitude). ISV_TOTAL_DIM: 461 → 462. Behavioral tests (5/5 PASS on RTX 3050): - sp16_phase2_hold_cost_scale_climbs_with_overrun - sp16_phase2_hold_cost_scale_at_target_is_one - sp16_phase2_hold_cost_scale_under_target_is_one - sp16_phase2_hold_cost_scale_bounds_clamp - sp16_phase2_hold_cost_scale_pearl_a_bootstrap Regression: SP14 oracle suite 30/30 PASS, SP15 phase 1 oracle suite 36/36 PASS. Instrumentation: HEALTH_DIAG[N]: hold_cost_scale_diag obs/tgt/norm/scale. Per feedback_isv_for_adaptive_bounds + feedback_no_partial_refactor. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
0426ce8887 |
fix(sp16-p1): MIN_HOLD_TEMPERATURE signal swap to hold-rate overrun + slot 330 non-bug closure
Per train-multi-seed-pfh9n post-mortem follow-up: slot 460 stuck at 50 in Fold 1 was NOT a launch-lifecycle bug. Producer fires per-epoch but kernel had early-return guard on AUX_DIR_ACC_SHORT_EMA (slot 373) at sentinel 0.5. Slot 373 reset on fold boundary; aux dir-acc EMA either didn't fire or settled within ε of 0.5 → kernel kept early-returning. Fix: drop slot 373 dependency entirely. Drive temperature from observed hold-rate vs target overrun: overrun = max(0, observed_hold_rate - target_hold_rate) overrun_norm = clamp(overrun / max(target, 0.01), 0, 1) new_temp = TEMP_MIN + (TEMP_MAX - TEMP_MIN) × overrun_norm blended_temp = Welford EMA α=0.05 with Pearl-A bootstrap When over-holding: temp HIGH → exit ramp permissive (matches design intent). When at/under target: temp LOW → exit penalty strict. Survives fold reset: hold-rate measurement starts fresh with real data immediately, no chained-input-sentinel masking. Slot 330 (KELLY_WARMUP_FLOOR) investigated and confirmed NON-BUG: producer behaves correctly per pearl_kelly_cap_signal_driven_floors cross-fold- persistence. floor=0 post-warmup is correct steady state. Behavioral tests: - sp16_phase1_min_hold_temp_climbs_with_hold_overrun - sp16_phase1_min_hold_temp_strict_when_at_target - sp16_phase1_min_hold_temp_strict_when_under_target - sp16_phase1_min_hold_temp_no_longer_reads_slot_373 Instrumentation: HEALTH_DIAG[N]: min_hold_temp_diag obs/target/overrun/temp. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
fb24614f07 |
feat(sp16-p0): Q-by-action HEALTH_DIAG diagnostic for Hold-bias observation
Per train-multi-seed-pfh9n post-mortem: structural-Q-bias hypothesis (Adam's m/sqrt(v) prefers low-variance Hold over noisy direction Q-targets) needs direct per-action Q observations to verify or refute. WR plateaued at ~0.46 across both folds while Hold% climbed 15% → 52% and Q-mean climbed 0.19 → 0.41 — but per-action Q was unobservable. Adds HEALTH_DIAG emit each epoch: HEALTH_DIAG[N]: q_by_action [hold=X long=Y short=Z flat=W] Reads via host-side averaging of `q_out_buf [B, total_actions]` direction-branch columns [0..b0=4]. No new kernel needed — q_out_buf is already populated row-major by compute_expected_q. Modeled directly on the existing Task 0.3 magnitude-bucket diagnostic (same dtoh-and- average cold path). Action-index ordering canonical, see state_layout.cuh:123-126: DIR_SHORT=0, DIR_HOLD=1, DIR_LONG=2, DIR_FLAT=3. Emit slot order is [hold, long, short, flat] (Hold first because the hypothesis is about Hold's ascent dominating direction Q-magnitudes). New surface: - gpu_dqn_trainer.rs: q_dir_means_cached field + update_q_dir_means_cached method + q_direction_action_means accessor + free static helper compute_q_dir_means_from_host_buf (factored for testability). - fused_training.rs: FusedTrainingCtx wrappers parallel to q_mag pair. - training_loop.rs: emit block adjacent to q_var_per_branch. Behavioral test: sp16_phase0_q_by_action_diagnostic_reads_four_action_means seeds known per-column constants, asserts canonical-dir-idx → emit-slot mapping at 1e-3 tolerance, and includes sentinel-leak guard (non-direction columns at 99.0 fail any wrong index→slot mapping with margin >12). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
9fb980da2b |
fixup(class-a-audit-batch-4b): invert MIN_HOLD_TEMPERATURE dir_acc → temp mapping
Post-commit kernel-formula audit of `compute_min_hold_penalty` at
trade_physics.cuh:567-577 showed the original `temp = TEMP_MIN +
(TEMP_MAX - TEMP_MIN) × skill` mapping was BACKWARDS for the
WR-plateau scenario this 8-commit chain targets.
Formula recap:
soft_factor = deficit / (deficit + T)
HIGH T → soft_factor → 0 → forgiving (no exit penalty)
LOW T → soft_factor → 1 → sharp (max exit penalty)
The deleted schedule's `START=50 → END=5` therefore meant
"permissive early, strict late" (classic explore-then-exploit), not
"permissive when committed, strict when uncertain" as the
substituted mapping assumed.
WR-plateau scenario: dir_acc ≈ 0.46-0.48 (model committed but
wrong) — the model needs a permissive (HIGH T) exit ramp to bail
out of bad committed directions, not maximum exit penalty (LOW T)
locking it into them.
Inverted the mapping using `confusion = 1 - skill`:
- dir_acc ≈ 0.5 (random / plateau) → confusion=1 → temp=50
(permissive, plateau exit ramp)
- dir_acc → 1.0 (saturated skill) → confusion=0 → temp=5
(strict, force informed commitment)
This preserves the deleted schedule's epoch-0 anchor (T=50
cold-start) while reacting to actual realized skill instead of an
epoch-time anchor.
Files touched (atomic per `feedback_no_partial_refactor`):
- min_hold_temperature_update_kernel.cu — formula + header doctring.
- sp14_isv_slots.rs — slot-doc comment expanded with new mapping
+ fixup-rationale paragraph.
- gpu_aux_trunk.rs — `MinHoldTemperatureUpdateOps` docstring.
- gpu_dqn_trainer.rs — cubin docstring + ISV_TOTAL_DIM doc-comment.
- tests/sp14_oracle_tests.rs — Tests 3 + 4 flipped (high dir_acc
now expects temp=TEMP_MIN; low dir_acc now expects TEMP_MAX);
Tests 1, 2, 5 numerically unchanged (sentinel guard fires before
the mapping; midpoint dir_acc=0.75 has skill=confusion=0.5 so
pre/post-fixup numerics match).
- docs/dqn-wire-up-audit.md — Item 4 entry rewritten with the
fixup rationale + mapping table updated.
Verification:
cargo check -p ml --tests --all-targets PASS
cargo test sp14_audit_4b (lib) 4/4 PASS (slot layout
locks unchanged)
GPU oracle re-run deferred to next L40S smoke (kernel was rebuilt
in-place so cubin contents change; the 5 oracles encode the new
mapping and will exercise the new cubin).
Per `pearl_first_observation_bootstrap.md` (sentinel→target
replace; sentinel anchors unchanged) +
`pearl_symmetric_clamp_audit.md` (bilateral clamp on target_temp
preserved) + `feedback_no_partial_refactor.md`.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
0b9ea77dc4 |
fix(class-a-audit-batch-4b): plan_threshold floor adaptive + MIN_HOLD_TEMPERATURE ISV-driven
Per Class A audit-fix Batch 4-B (final 2 of 4 deferred items from
P1-wiring/P1-producer). Completes the 8-commit WR-plateau intervention
chain. Validation deferred to next L40S smoke.
Item 3: plan_threshold adaptive floor (Design Y - inline producer)
- NEW slot PLAN_THRESHOLD_FLOOR_ADAPTIVE_INDEX=459
- Writes slow-EMA shadow of `0.5 * readiness_ema` from inside the
existing update kernel (no new file, no new launch).
- Pearl-A first-observation bootstrap (sentinel 0.1 matches pre-fix
hardcoded value for bit-identical cold-start) + Welford alpha=0.005
slow EMA.
- Bilateral clamp [0.05, 0.50] (probability units) per
pearl_symmetric_clamp_audit.
- Consumer reads isv[459] as the floor in the same launch's final
fmaxf; cold-start sentinel REPLACES with threshold_target so the
pre-fix `fmaxf(0.1, 0.5*ema)` semantic is preserved bit-identical
for any readiness EMA above 0.20.
Item 4: MIN_HOLD_TEMPERATURE -> ISV-driven (driving signal: dir_acc skill)
- NEW slot MIN_HOLD_TEMPERATURE_ADAPTIVE_INDEX=460
- NEW kernel min_hold_temperature_update_kernel.cu (single-thread
cold-path, per-epoch boundary launch).
- Driving signal: dir_acc skill = clamp((short_ema - 0.5)/0.5, 0, 1)
from ISV[AUX_DIR_ACC_SHORT_EMA_INDEX=373]. When committing skillfully
(high dir_acc) -> temp HIGH (permissive). When at random baseline
(~0.5) -> temp LOW (sharp commitment pressure). Substituted for
the audit-spec's `dir_entropy_deficit` because no dir_entropy ISV
slot exists - dir_acc skill is the closest semantically-equivalent
signal that preserves the spec intent.
- Pearl-A bootstrap (sentinel 50.0 matches the deleted
MIN_HOLD_TEMPERATURE_START=50 anchor) + alpha=0.05 mid-cadence EMA.
- Bounds [5, 50] (matches the deleted schedule range).
- Decouples temperature from epoch number - the old schedule pinned
LOW temp (sharp) at end of training, exactly when a WR-plateaued
model needed forgiveness to escape.
- DELETED: state_layout.cuh::MIN_HOLD_TEMPERATURE_{START, END, DECAY}
#defines + training_loop.rs::min_hold_temperature_for_epoch helper
function (kept docstring tombstone explaining the deletion). Both
call sites migrated to the new ISV reader. Per
feedback_no_legacy_aliases + feedback_no_partial_refactor.
ISV_TOTAL_DIM: 459 -> 461.
Cumulative WR-plateau fix series (final commit, #8):
-
|
||
|
|
7e9a8f6ef1 |
fix(class-a-audit-batch-4a): DD saturation floor adaptive + legacy DD path Case A
Per Class A audit-fix Batch 4-A (deferred from P1-wiring/P1-producer due to
audit-doc errors). Fixes 2 of 4 deferred items; Batch 4-B handles plan_threshold
floor + MIN_HOLD_TEMPERATURE in a separate commit.
Item 1: DD saturation floor (the upper end of the DD ramp at trade_physics.cuh:154
in apply_margin_cap, NOT line 548 as the audit doc claimed — that line is a
magnitude action constant; the actual saturation floor lives in apply_margin_cap)
- NEW slot DD_SATURATION_FLOOR_ADAPTIVE_INDEX=458
- Producer dd_saturation_floor_update_kernel.cu — p75(per-env DD_MAX) × 1.5
via Welford `mean + Z_75 × sigma` estimator with `max(p75, mean)` robustness
guard, mirrors P0-A REWARD_POS_CAP producer pattern (Pearl-A bootstrap +
Welford α=0.01)
- Cold-start fallback: 0.25f (DD_SATURATION_FLOOR_DEFAULT in state_layout.cuh)
- Bounds: [0.10, 0.50] (Category-1 dimensional safety)
- Distinct from SP15_DD_THRESHOLD_INDEX=421 (the SP15 quadratic DD-penalty
*trigger* threshold, a *lower* bound; this slot is the *upper* end of the
linear position-size scaling ramp dd_scale = max(0.05, 1.0 − dd_frac/floor))
- Threaded `isv_signals_ptr` into `apply_margin_cap` with NULL-tolerant
cold-start fallback to DD_SATURATION_FLOOR_DEFAULT
- 4 oracle tests (Pearl-A bootstrap, no-DD guard, bounds clamp, Welford EMA)
Item 2: Legacy compute_drawdown_penalty path → Case A (DELETED)
- Decision rationale: SP15's quadratic asymmetric DD penalty
(compute_sp15_final_reward_kernel.cu:154 via sp15_dd_penalty helper) runs
unconditionally as a post-modifier on the SP11-composed reward with
ISV-driven λ_dd (slot 420) and DD threshold (slot 421). Layering the legacy
linear-ramp penalty inside the SP11 composer on top of the SP15 quadratic
creates double-counting of DD shaping — exactly the code-smell the Class A
audit was designed to eliminate. Per `feedback_no_legacy_aliases.md` and
`feedback_no_partial_refactor.md`.
- Atomic deletion across:
- `compute_drawdown_penalty` device function (trade_physics.cuh)
- Single call site at experience_kernels.cu:3822
- `dd_threshold` and `w_dd` kernel arguments
- `w_dd` Rust config field (gpu_experience_collector.rs +
trainers/dqn/config.rs DQNHyperparameters)
- `w_dd` profile section + dispatch (training_profile.rs RewardSection,
OptimizableParameterRanges, FixedRewardParameters, ParamLookup
dispatch, profile→hyperparam mapping, test assertion)
- `w_dd *= rki` risk-intensity multiplier (config.rs)
- `w_dd` TOML keys (dqn-hyperopt.toml × 2, dqn-localdev.toml,
dqn-production.toml, dqn-smoketest.toml)
- Stale doc comments on hyperopt/adapters/dqn.rs + config.rs
risk_intensity field
- `config.dd_threshold` SURVIVES (still consumed by `launch_sp15_dd_state`
as the dd_budget for DD_PCT scaling). Documented in field comment.
ISV_TOTAL_DIM: 458 → 459 (Item 1 adds 1 slot; Item 2 is pure deletion)
Cumulative WR-plateau fix series (this is commit 7):
- Class C bug 1 + P0-B (
|
||
|
|
87d597d5d7 |
fix(class-a-p1-producer): adaptive Bayesian Kelly priors per-fold-end
Replace 4 hardcoded Bayesian priors in Kelly cap calculations
(prior_wins=2.0, prior_losses=2.0, prior_sum_wins=0.01,
prior_sum_losses=0.01) with ISV-driven slow-EMA values fed by a new
producer kernel that aggregates realized PS_KELLY_* fields across envs
from the same portfolio_state buffer kelly_cap_update_kernel reads
from at the same per-epoch boundary.
Slots [454..458):
- KELLY_PRIOR_WINS, KELLY_PRIOR_LOSSES,
KELLY_PRIOR_SUM_WINS, KELLY_PRIOR_SUM_LOSSES
ISV_TOTAL_DIM bumps 454 -> 458.
Producer:
- kelly_bayesian_priors_update_kernel.cu (per-fold-end / per-epoch
boundary, block-tree-reduce in shmem, no atomicAdd). Pearl-A
first-observation bootstrap (sentinels 2.0/2.0/0.01/0.01 match
pre-P1-Producer hardcoded values for bit-identical cold-start) +
alpha=0.005 slow EMA. Bounds counts in [0.5, 100], sums in [0.001,
1.0] per feedback_isv_for_adaptive_bounds. Launched RIGHT BEFORE
launch_kelly_cap_update so that kernel sees the freshly-blended
priors.
Consumers migrated atomically (same commit):
- kelly_cap_update_kernel.cu:39-42 -> ISV[454..458) with cold-start
fallback via kelly_prior_or_default helper (range guard).
- trade_physics.cuh::kelly_position_cap:304-307 -> NULL-tolerant
isv_signals_ptr threaded through apply_kelly_cap (single caller in
unified_env_step_core line 898 already had the bus pointer; mirrors
the existing kelly_f_smooth / kelly_warmup_floor_sp9 patterns).
State reset:
- 4 FoldReset registry entries (sp14_p1_kelly_prior_*) +
4 dispatch arms in training_loop.rs::reset_named_state.
Oracle tests (sp14_oracle_tests.rs, GPU-gated #[ignore]):
- Pearl-A bootstrap (sentinel -> REPLACE with aggregated targets)
- No realized trades -> ISV preserved bit-exactly
- Bounds clamp on extreme aggregates
- Slow EMA blend after bootstrap
CPU tests passing:
- sp14_p1_kelly_prior_slot_layout_locked
- all_sp14_p1_slots_fit_within_isv_total_dim
- every_fold_and_soft_reset_entry_has_dispatch_arm (C.10 lesson)
- layout_fingerprint_bumps_after_sp14_wire
DEFERRED — Item 2 (MIN_HOLD_TEMPERATURE EMA): the audit-spec said
"MIN_HOLD_TEMPERATURE = 0.5f hardcoded somewhere" but the actual code
has MIN_HOLD_TEMPERATURE_{START=50.0f, END=5.0f, DECAY=20.0f} as a
PER-EPOCH ANNEALING SCHEDULE driven by min_hold_temperature_for_epoch
in training_loop.rs:68-73. The kernel takes T as a runtime scalar
specifically to enable Phase 2 ISV-driven lift "without recompiling
cubin" per the SP12 v3 design comment. The audit's claim of "0.5f
hardcoded" does not match reality; the right Phase 2 signal is a
separate spec decision and should not be guessed at per
feedback_no_quickfixes. Reporting back per the prompt's hard rule #8.
Per feedback_isv_for_adaptive_bounds + feedback_no_partial_refactor +
pearl_controller_anchors_isv_driven.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
394de7d434 |
feat(class-a-p0a): REWARD_POS/NEG_CAP → ISV-driven adaptive caps from realized return distribution
Per Class A audit ranking, the highest-suspected-impact fix for the months-long WR-stuck-at-46-48% plateau across 11 superprojects. Hardcoded REWARD_POS_CAP=+5.0f / REWARD_NEG_CAP=-10.0f (state_layout.cuh:266-267) was structurally clipping the upper tail of realized alpha. Controller signals (sharpe EMA, var_q, q_gap) all derive from this CAPPED buffer, so the controller cannot select for trades it cannot see. Selectivity gradient evaporates — small wins clip to +5 alongside large wins also clipping to +5. The state_layout comment lines 261-264 explicitly deferred Phase 2 (ISV-driven) "IF Phase 1 validation reveals adaptive need". Phase 1 has been running 11 SPs without budging WR — adaptive need revealed. Architecture: - 2 new ISV slots [452..454): REWARD_POS_CAP_ADAPTIVE, REWARD_NEG_CAP_ADAPTIVE. - Producer kernel reward_cap_update_kernel.cu: block-tree-reduce Welford `mean + Z_99 × sigma` p99 estimator over winning realized returns + conservative `max(p99, max_win)` takeover, × 1.5 safety factor → POS cap. NEG cap = -2 × POS cap (preserves Kahneman 2:1 asymmetry per pearl_audit_unboundedness_for_implicit_asymmetry — asymmetry stays, but moved from hardcoded scalar to producer-time multiplier; single source of truth, no consumer applies the 2× ratio itself). - Pearl-A first-observation bootstrap from sentinel (5.0 / -10.0, matching pre-P0-A hardcoded values for bit-identical cold-start). Welford EMA α=0.01 thereafter (slow blend — reward distribution is the foundation of training and shouldn't move fast). - Bounds: POS in [1, 50], NEG in [-100, -2] (Category-1 dimensional safety per feedback_isv_for_adaptive_bounds, NOT tuning). - 3 consumer sites migrated atomically per feedback_no_partial_refactor: experience_kernels.cu:3112-3114 (segment_complete cap), compute_sp15_final_reward_kernel.cu:163 (Stage 4 helper invocation), sp15_reward_axis_helpers.cuh:211 (sp15_apply_sp12_cap device fn signature change to take isv ptr). - Cold-start fallback: when ISV slot at sentinel OR outside [REWARD_POS_CAP_MIN_BOUND=1, REWARD_POS_CAP_MAX_BOUND=50], consumers fall back to original macros (still defined in state_layout.cuh). - 2 new device-ptr accessors on the experience collector (step_ret_per_sample_dev_ptr, trade_close_per_sample_dev_ptr) — reuses existing per-sample buffers; no new buffer allocated. - Per-epoch boundary launch (cold path) in training_loop.rs alongside launch_aux_horizon_chain. - Reset registry entries + dispatch arms in reset_named_state per the C.10 lesson (missing dispatch causes runtime crash). - Layout fingerprint seed updated: ISV_TOTAL_DIM 452→454 + AUX_PRED_HORIZON_BARS=450 + AVG_WIN_HOLD_TIME_BARS=451 (previously missing from seed) + REWARD_POS_CAP_ADAPTIVE=452 + REWARD_NEG_CAP_ADAPTIVE=453. Per feedback_isv_for_adaptive_bounds: every adaptive bound in ISV. Verification: - cargo check -p ml --tests --all-targets: clean (19 pre-existing warnings, 0 new). - sp14_isv_slots tests: 8/8 pass (4 layout + 4 fits-within). - sp14_oracle_tests with --features cuda --ignored: 8/8 pass (4 existing q_disagreement/dir_concat + 4 new P0-A tests covering Pearl-A bootstrap, no-winning-trades preservation, bounds clamping to [1,50], Welford α=0.01 EMA blend). - sp15_phase1_oracle_tests with --features cuda --ignored: 36/36 pass (no regression from sp15_apply_sp12_cap signature change). Cumulative WR-plateau fix series: - Class C bug 1 ( |
||
|
|
0e61de408f |
feat(sp14-c6): h_s2_aux_rms_ema producer — ISV[449] per-collector-step
Single-block 256-thread CUDA kernel computing RMS(h_s2_aux [B, SH2]) and EMA-blending the step observation into ISV[H_S2_AUX_RMS_EMA_INDEX=449] directly. Pearl-A first-observation bootstrap embedded in kernel body (sentinel 0.0 → replace); fixed α=0.05 EMA blend thereafter. ISV slot 449 is outside the SP4/SP5 wiener buffer linear span so the scratch+apply_pearls_ad_kernel path is not available — self-contained Pearl-A logic mirrors the avg_win_hold_time_update_kernel precedent (slot 451). No atomicAdd; shmem block-tree-reduce only. Launched after aux_trunk_forward in the collector per-step hot path. - h_s2_aux_rms_ema_kernel.cu — new CUDA kernel (81 lines) - build.rs — cubin manifest entry - gpu_dqn_trainer.rs — H_S2_AUX_RMS_EMA_CUBIN static - gpu_aux_trunk.rs — HS2AuxRmsEmaOps struct + launch() - gpu_experience_collector.rs — field + constructor + hot-path launch - aux_trunk_oracle_tests.rs — h_s2_aux_rms_ema_pearl_a_bootstrap test - dqn-wire-up-audit.md — Phase C.6 audit entry cargo check -p ml --tests: clean (only pre-existing warnings) Oracle test: 1 new test added (requires GPU to run) Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
b26b189925 |
feat(sp14-c.5b): atomic contract migration — h_s2 → h_s2_aux + revert zero-fills
Atomic flip of the aux heads' input from Q's GRN trunk output `save_h_s2` to the SEPARATE aux trunk's output `h_s2_aux`. The aux trunk now trains its own w1/w2/w3/b1/b2/b3 from CE loss (next-bar + regime); Q's encoder is structurally protected by `aux_trunk_backward`'s missing `dx_in` output param (encoder boundary stop-grad enforced at the kernel-set level). Reverts the C.0 stop-grad band-aid commits (`872bd7392`, `411a30473`): the zero-fills in `aux_next_bar_backward` + `aux_regime_backward` Step 3 are replaced with the genuine SAXPY-back-to-input gradient (`dh_s2_aux[b,j] = sum_k sh_dh_pre[k] * w1[k,j]`). The leak that motivated stop-grad is now blocked structurally rather than by data zero-fill — aux gradient flows through the aux trunk's own params, never into Q's encoder. Wired in this commit (atomic, ~330 LOC): - 4 kernel signatures renamed `h_s2 → h_s2_aux` / `dh_s2_out → dh_s2_aux_out` (`aux_next_bar_forward`, `aux_regime_forward`, `aux_next_bar_backward`, `aux_regime_backward`); Rust wrappers in `gpu_aux_heads.rs` follow - Trainer fwd: insert `aux_trunk_forward_ops.launch(...)` in `aux_heads_forward` Step 0, populating `h_s2_aux` from `save_h_s1` (encoder layer-1 output, dim=shared_h1=256). Both head fwds redirect input pointer from `save_h_s2` to `h_s2_aux` - Trainer bwd: SAXPY both `aux_dh_s2_*_buf` into `dh_s2_aux_accum` (pre-zeroed each step via graph-safe `cuMemsetD32Async`); then `aux_trunk_backward_ops.launch(...)` propagates through w3/w2/w1 + b3/b2/b1; then `launch_aux_trunk_adam_update` applies global L2-norm clip + per-tensor Adam updates over 6 grad tensors - Collector fwd: insert `exp_aux_trunk_forward_ops.launch(...)` after `forward_online_f32`, reading `exp_h_s1_f32` and writing `exp_h_s2_aux`; redirect `exp_aux_heads_fwd.forward_next_bar` input from `exp_h_s2_f32` to `exp_h_s2_aux` - Pre-capture host-write of ISV-driven LR + grad-clip + step counter into mapped-pinned buffers in `launch_cublas_backward_to` (BEFORE `aux_heads_backward`); same `&mut self` pattern as `step_ofi_embed_adam` Verification: - `cargo check -p ml --tests` clean (1m02s, only pre-existing warnings) - `aux_trunk_oracle_tests` + `sp14_oracle_tests` 12/12 pass: - aux_trunk gradient check: max_rel_err=1.33e-2 (tol=2e-2) — matches C.4 baseline - aux_trunk_backward_does_not_write_dx: kernel source clean of dx_in/dx_in_out - aux_sign_label_lookahead_mask: 60/100 masked, 40/100 valid - 9 other oracle tests pass bit-identically Plan: docs/superpowers/plans/2026-05-07-sp14-layer-c-separate-aux-trunk.md §C.5b Audit: docs/dqn-wire-up-audit.md "SP14 Layer C Phase C.5b" section |
||
|
|
1edd71a2c1 |
feat(sp14-c.5a-fixup): missing scratch buffers + collector ptr scaffolding (dead code)
Phase C.5a-fixup — completes C.5a's additive infrastructure so C.5b can be
a true atomic contract flip with zero scaffolding work mixed in:
- Allocate dh_aux1_pre_scratch [B_max, AUX_TRUNK_H1=256] +
dh_aux2_pre_scratch [B_max, AUX_TRUNK_H2=128] (missed in C.5a — both
required by aux_trunk_backward.launch per gpu_aux_trunk.rs:266-267).
- Collector struct gains 6× u64 aux_trunk_{w1,b1,w2,b2,w3,b3}_ptr fields,
exp_aux_trunk_forward_ops: AuxTrunkForwardOps field (constructed in
ctor on collector's stream), and set_trainer_aux_trunk_param_ptrs
setter — mirrors the existing set_trainer_params_ptr zero-copy pattern.
- Trainer gains aux_trunk_param_ptrs() -> (u64×6) accessor returning
raw_ptr() for all 6 aux trunk parameter tensors.
- training_loop wires the new setter at both existing
set_trainer_params_ptr call sites (initial fused-ctx init + fold-boundary
re-init).
NO contract change: wire sites still call save_h_s2. The setter is called
and the 6 aux trunk param ptrs are populated, but no collector-side launch
reads them yet — C.5b atomically inserts aux_trunk_forward.launch(...)
post-forward_online_f32 and switches the aux head input pointer.
Graph-capture audit: aux_heads_backward IS INSIDE the captured `forward`
child graph (call chain: submit_forward_ops_main → launch_cublas_backward
→ launch_cublas_backward_to → aux_heads_backward; capture begins at
fused_training.rs:2964 / capture_child_graph). The existing function body
is fully device-side (zero host writes) — capture-safe by construction.
C.5b's new aux_trunk Adam launch is also fully device-side and will sit
inside the same captured region. The host writes for aux_trunk_t_pinned
must use the existing GPU-side increment_step_counters kernel chain
(submit_counters_ops, line 22799) — NOT host-side aux_trunk_adam_step
+= 1 inside capture. ISV-driven LR/clip writes happen pre-capture
(cold-path); the captured graph reads via aux_trunk_lr_dev_ptr /
aux_trunk_grad_clip_dev_ptr. This avoids the &self → &mut self ripple on
aux_heads_backward (gap 4 in C.5b implementer's blocker report). Full
wiring strategy + alternative (pre-capture host-write) documented in
docs/dqn-wire-up-audit.md C.5a-fixup section.
Verification:
- cargo check -p ml --tests --all-targets: clean (no new warnings).
- cargo test -p ml --test aux_trunk_oracle_tests --test sp14_oracle_tests
--release -- --ignored --nocapture: 12/12 pass (8 aux_trunk + 4 sp14;
bit-identical to C.5a baseline — pure scaffolding, no regression).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|