diff --git a/docs/dqn-wire-up-audit.md b/docs/dqn-wire-up-audit.md index 28fcf1366..d1ae14d1d 100644 --- a/docs/dqn-wire-up-audit.md +++ b/docs/dqn-wire-up-audit.md @@ -2,77 +2,80 @@ **Status:** Populated during Plan 1 Task 6 (A.5 orphan audit). Updated on every commit per Invariant 7. -## 2026-05-10 — SP20 Phase 5 (commit 3/4): launcher arg + tests +## 2026-05-10 — SP20 Phase 5: aux→Q reward gate (close-out) -### Launcher +**Branch:** `sp20-phase-5` (off `d1a8ec206`). +**Commits:** 4 atomic — trainer plumbing, kernel gate, launcher+tests, +this consolidation. -`GpuDqnTrainer::launch_c51_loss` (in `gpu_dqn_trainer.rs`) appends one -`.arg(&self.aux_conf_at_state_buf)` after the existing -`isv_signals_dev_ptr` argument. Position matches the kernel's appended -trailing arg from commit 2/4. +Lands the consumer-side closure of the Phase 3 Task 3.4 producer→ring +schema. Once the model is in fused-training mode, PER's +`gather_f32_scalar` writes the SAMPLED bar's `aux_conf` (K=2 +peak-softmax-above-uniform confidence) into the trainer's +`aux_conf_at_state_buf` on every step, and the c51_loss kernel reads +that buffer at the kernel-entry reward-setup site to gate the reward +before Bellman projection. -### Tests added (`crates/ml-dqn/src/gpu_replay_buffer.rs::tests`) +### Design: gate the REWARD, not the target - - **`aux_gate_high_confidence_passes_full_target`** (CPU pure-math). - Verifies `gate(aux_conf=0.5, threshold=0.10, temp=0.05) > 0.99` — - proves the "high confidence ⇒ near-identity reward" structural - invariant. - - **`aux_gate_low_confidence_attenuates_reward`** (CPU pure-math). - Verifies `gate(aux_conf=0.02, threshold=0.10, temp=0.05) < 0.20` — - proves the "uncertain-state neutralizer" semantic from Phase 3 - Task 3.4 audit doc spec §4.4. - - **`aux_gate_temp_floor_keeps_gate_finite`** (CPU pure-math). - Sweeps {temp ∈ [0, 0.5]} × {aux_conf ∈ [0, 0.5]} × {threshold ∈ - [0.01, 0.20]} and asserts `gate.is_finite()` ∧ `gate ∈ [0, 1]` for - all inputs — proves the `fmaxf(temp, 1e-3)` floor keeps the kernel - numerically stable across the full ISV-controllable parameter - range. - - **`aux_conf_direct_to_trainer_gather_populates_destination`** (GPU - behavioral). Wires a fresh `CudaSlice` as the trainer - destination via `set_trainer_aux_conf_ptr(dst.raw_ptr())`, inserts - 8 transitions with strictly-positive distinct aux_conf values, - samples 1, asserts the destination buffer post-sample contains a - value from the inserted set (NOT the alloc_zeros sentinel of 0.0). - Proves the Phase 5 wiring (PER's `gather_f32_scalar` direct-path) - actually populates the trainer-side buffer with non-trivial data. - -### Files modified - -| File | Status | Purpose | -|------|--------|---------| -| `crates/ml/src/cuda_pipeline/gpu_dqn_trainer.rs` | +1 launcher arg | Thread aux_conf_at_state_buf into c51_loss kernel | -| `crates/ml-dqn/src/gpu_replay_buffer.rs` | +4 tests | Gate-formula + direct-gather verification | - -### Verification +Per the user-supplied design rationale (cf. preceding-agent report): ``` -SQLX_OFFLINE=true cargo check --workspace --examples # green -SQLX_OFFLINE=true cargo test -p ml-dqn --lib aux_gate # 3/3 pass (CPU) -SQLX_OFFLINE=true cargo test -p ml-dqn --lib aux_conf # 3/3 pass (CPU + GPU integration) +r_used = gate * reward +where gate = sigmoid((aux_conf - threshold) / temp) ``` -## 2026-05-10 — SP20 Phase 5 (commit 2/4): c51_loss_kernel reward gate +Mathematical interpretation: at low aux confidence (gate→0), +`r_used → 0`, so the Bellman target becomes `gamma * Q(s', a')` — the +Q value at the current state collapses toward `gamma * Q(s', a')`, i.e., +the model gets no reward feedback on uncertain transitions. Effectively +"don't update Q on uncertain transitions." This avoids needing +`q_mean_a` entirely — a cleaner formulation than the original Phase 3 +Task 3.4 audit-doc spec §4.4 sketch (which blended toward +`mean_a Q(s, a)`). -Adds the consumer kernel-side gate. New kernel arg `const float* -__restrict__ aux_conf_at_state` appended to `c51_loss_batched`'s signature -(after the existing `isv_signals` arg). Gate computation runs once per -sample at the kernel-entry reward-setup site (immediately after the -existing `#27 ensemble_disagreement` adjustment, so the gate composes -cleanly with the existing reward-shaping cascade): +### Components -```cuda -if (aux_conf_at_state != NULL && isv_signals != NULL) { - float aux_conf = aux_conf_at_state[sample_id]; - float threshold = isv_signals[518]; /* AUX_CONF_THRESHOLD_INDEX */ - float temp = fmaxf(isv_signals[519], 1e-3f); /* AUX_GATE_TEMP_INDEX */ - float gate = 1.0f / (1.0f + __expf(-(aux_conf - threshold) / temp)); - reward = gate * reward; -} -``` + 1. **Trainer-side buffer** (`crates/ml/src/cuda_pipeline/gpu_dqn_trainer.rs`): + `aux_conf_at_state_buf: CudaSlice` `[batch_size]`, allocated via + `alloc_f32(&stream, aux_b, "aux_conf_at_state_buf")`. Sentinel 0.0 + matches Phase 3 Task 3.4 producer's NULL-fallback. Accessor + `aux_conf_at_state_buf_ptr()` mirrors the SP13 B1.1b + `aux_nb_label_buf_ptr()` direct-gather accessor. + 2. **FusedTrainerCtx delegating accessor** + (`crates/ml/src/trainers/dqn/fused_training.rs`): + `trainer_aux_conf_at_state_buf_ptr()` — same delegation pattern as the + 6 sibling getters (`trainer_states_buf_ptr` etc.). + 3. **PER direct-gather wire-up** + (`crates/ml/src/trainers/dqn/trainer/training_loop.rs`): two call + sites — init (line ~826) and re-init (line ~2645) — both invoke + `agent_w.memory_mut().gpu.set_trainer_aux_conf_ptr( + fused.trainer_aux_conf_at_state_buf_ptr())` immediately after + `set_trainer_buffers(...)`. Atomic per `feedback_no_partial_refactor`. + 4. **Kernel signature + gate** + (`crates/ml/src/cuda_pipeline/c51_loss_kernel.cu`): new trailing + arg `const float* __restrict__ aux_conf_at_state` after + `isv_signals`. Gate computation runs once per sample at the + reward-setup site (after the `#27 ensemble_disagreement` + adjustment, before `block_bellman_project_f` is called per branch): -The gated reward propagates through every branch's `block_bellman_project_f` -call without any per-branch changes — the gate is semantically a per-sample -modifier on the reward signal, which is per-sample by construction. + ```cuda + if (aux_conf_at_state != NULL && isv_signals != NULL) { + float aux_conf = aux_conf_at_state[sample_id]; + float threshold = isv_signals[518]; /* AUX_CONF_THRESHOLD_INDEX */ + float temp = fmaxf(isv_signals[519], 1e-3f); /* AUX_GATE_TEMP_INDEX */ + float gate = 1.0f / (1.0f + __expf(-(aux_conf - threshold) / temp)); + reward = gate * reward; + } + ``` + + The gated reward propagates through every branch's + `block_bellman_project_f` call without per-branch changes — the gate + is semantically a per-sample modifier on the reward, which is + per-sample by construction. + 5. **Launcher** (`GpuDqnTrainer::launch_c51_loss`): one + `.arg(&self.aux_conf_at_state_buf)` appended after the existing + `isv_signals_dev_ptr` argument. ### NULL-tolerance contract @@ -84,109 +87,124 @@ behaviour): - `isv_signals == NULL` ⇒ cannot read threshold/temp, degrade to identity rather than producing garbage. +### Default-state semantics ("graceful degradation") + +`alloc_zeros` cold-start: 0.0 sentinel. The Phase 5 gate computes: + +``` +gate = sigmoid((0 - 0.10) / 0.05) = sigmoid(-2) ≈ 0.12 +``` + +so `r_used ≈ 0.12 × reward`. The model gets minimal reward feedback at +the Bellman target until PER's first gather populates the buffer with +real per-bar aux_conf values. After that, the per-bar values from the +producer drive the gate as designed. + ### IQN: explicitly out of scope `iqn_dual_head_kernel.cu` is the auxiliary distributional loss in this -codebase; C51 is the production loss. The Phase 5 gate lives in C51 only -per the user-supplied design rationale — adding it to IQN is more -complexity for marginal gain. +codebase; C51 is production. The Phase 5 gate lives in C51 only per the +user-supplied design rationale — adding it to IQN is more complexity for +marginal gain. -### Files modified +### `reward_bias` interaction + +The kernel's existing `reward_bias` term (Task 2.Y-ext v5: ISV-driven +direction-branch deficit, computed per-branch inside +`block_bellman_project_f`) is unchanged. The Phase 5 gate fires on the +raw `reward` BEFORE `block_bellman_project_f` is called, so the gated +reward is what enters the per-branch projection: `t_z = (gate * reward ++ reward_bias) + gamma * z * (1 - done)`. The two mechanisms compose +cleanly: + + - `gate * reward` damps the reward signal to zero on uncertain + transitions. + - `reward_bias` lifts the tradable bin's Q-mean toward + `max_pathology_q + lead_scale` when training health is poor, + independent of the gate (it's a per-branch direction-stretch, not + a reward modifier). + +### Files modified (4 commits, 4 files net + 1 audit doc) | File | Status | Purpose | |------|--------|---------| -| `crates/ml/src/cuda_pipeline/c51_loss_kernel.cu` | +1 kernel arg, +gate computation | Reward-gate consumer | - -### Verification - -``` -SQLX_OFFLINE=true cargo check -p ml --lib # green -``` - -GPU verification deferred to Commit 3/4 (launcher + tests). - -## 2026-05-10 — SP20 Phase 5 (commit 1/4): trainer aux_conf_at_state_buf + PER direct-gather wire-up - -Phase 5 consumer-side plumbing — atomic across all struct boundaries per -`feedback_no_partial_refactor`. Lands the trainer-side buffer the c51_loss -reward-gate kernel will consume in commit 2/4. - -### Trainer-side allocation - -`GpuDqnTrainer.aux_conf_at_state_buf: CudaSlice` — `[batch_size]` f32. -Allocated via `alloc_f32(&stream, aux_b, "aux_conf_at_state_buf")` in the -constructor, alongside the existing aux-heads buffers. Sentinel: `0.0` -(alloc_zeros) — matches the K=2 uniform-prior default the Phase 3 Task 3.4 -producer writes when `aux_logits_per_env == NULL`. - -### Trainer-side accessor - -`pub(crate) fn aux_conf_at_state_buf_ptr(&self) -> u64` — mirrors the -`aux_nb_label_buf_ptr` accessor used by SP13 B1.1b's i32 direct-gather. The -wrapping `FusedTrainerCtx::trainer_aux_conf_at_state_buf_ptr` delegates to -the inner trainer accessor (same delegation pattern as the 6 sibling getters). - -### Replay-buffer wire-up (already plumbed in Phase 3 Task 3.4) - -`GpuReplayBuffer::set_trainer_aux_conf_ptr(u64)` was added in Phase 3 Task 3.4 -(comment at `gpu_replay_buffer.rs:1163-1165` explicitly says: "Phase 5 will -call this once after the trainer's `aux_conf_at_state_buf` is allocated"). -This commit calls it. - -### Training loop wire-up (atomic — both fused_ctx init sites) - -`crates/ml/src/trainers/dqn/trainer/training_loop.rs` — two call sites mirror -the existing 6-arg `set_trainer_buffers` block: - - - Line ~826 (init site): `agent_w.memory_mut().gpu.set_trainer_aux_conf_ptr( - fused.trainer_aux_conf_at_state_buf_ptr())` immediately after - `set_trainer_buffers(...)`. - - Line ~2645 (re-init site after fused_ctx tear-down): same call. Both - sites stay in sync per `feedback_no_partial_refactor`. - -### Default-state semantics ("graceful degradation") - -`alloc_zeros` cold-start: 0.0 sentinel. The Phase 5 gate kernel (commit 2/4) -will compute: - -``` -gate = sigmoid((aux_conf - threshold) / temp) - = sigmoid((0 - 0.10) / 0.05) // typical ISV values per Phase 1.3 - = sigmoid(-2) - ≈ 0.12 -``` - -so `r_used = 0.12 × reward` ≈ "reward mostly suppressed". This is the -"uncertain-state neutralizer" semantic from Phase 3 Task 3.4 audit doc -spec §4.4 — pre-population (no aux signal yet), the model gets minimal -reward feedback at the Bellman target → Q collapses to `γ × Q(s', a')`, -i.e., "don't update Q on uncertain transitions." Once PER's -`gather_f32_scalar` populates from the producer ring on the first sample -step, the per-bar aux_conf values drive the gate as designed. - -### Files modified - -| File | Status | Purpose | -|------|--------|---------| -| `crates/ml/src/cuda_pipeline/gpu_dqn_trainer.rs` | +field, +alloc, +accessor | Trainer-side buffer ownership | +| `crates/ml/src/cuda_pipeline/gpu_dqn_trainer.rs` | +field, +alloc, +accessor, +1 launcher arg | Trainer ownership + launch threading | | `crates/ml/src/trainers/dqn/fused_training.rs` | +delegating accessor | FusedTrainerCtx wrapper | | `crates/ml/src/trainers/dqn/trainer/training_loop.rs` | +2 call sites | PER direct-gather wire-up | +| `crates/ml/src/cuda_pipeline/c51_loss_kernel.cu` | +1 kernel arg, +gate computation | Reward-gate consumer | +| `crates/ml-dqn/src/gpu_replay_buffer.rs` | +4 tests | Gate-formula + direct-gather verification | +| `docs/dqn-wire-up-audit.md` | This entry (consolidated) | Audit log | + +### Tests + + - **`aux_gate_high_confidence_passes_full_target`** (CPU pure-math). + Verifies `gate(aux_conf=0.5, threshold=0.10, temp=0.05) > 0.99` — + proves the "high confidence ⇒ near-identity reward" structural + invariant. + - **`aux_gate_low_confidence_attenuates_reward`** (CPU pure-math). + Verifies `gate(aux_conf=0.02, threshold=0.10, temp=0.05) < 0.20` — + proves the "uncertain-state neutralizer" semantic. + - **`aux_gate_temp_floor_keeps_gate_finite`** (CPU pure-math). Sweeps + {temp ∈ [0, 0.5]} × {aux_conf ∈ [0, 0.5]} × {threshold ∈ [0.01, + 0.20]} and asserts `gate.is_finite()` ∧ `gate ∈ [0, 1]` for all + inputs — proves the `fmaxf(temp, 1e-3)` floor keeps the kernel + numerically stable across the full ISV-controllable parameter + range. + - **`aux_conf_direct_to_trainer_gather_populates_destination`** (GPU + behavioral). Wires a fresh `CudaSlice` as the trainer + destination via `set_trainer_aux_conf_ptr(dst.raw_ptr())`, inserts + 8 transitions with strictly-positive distinct aux_conf values, + samples 1, asserts the destination buffer post-sample contains a + value from the inserted set (NOT the alloc_zeros sentinel of 0.0). + Proves the Phase 5 wiring (PER's `gather_f32_scalar` direct-path) + actually populates the trainer-side buffer with non-trivial data. ### Verification ``` SQLX_OFFLINE=true cargo check --workspace --examples # green +SQLX_OFFLINE=true cargo test -p ml-dqn --lib aux_gate # 3/3 pass (CPU) +SQLX_OFFLINE=true cargo test -p ml-dqn --lib aux_conf # 3/3 pass (CPU + GPU integration) ``` -Forward-references (commits 2-4 of this Phase 5 lands in this branch): +Pre-existing failures unrelated to this commit (pre-existing on base +`d1a8ec206`): 16 ml-dqn lib tests fail with "State dimension mismatch: +expected 128, got 4" — these are `branching` / `dueling` / +`distributional_dueling` / `ensemble_network` / `network` forward-pass +tests that are stale w.r.t. the current state-dim layout and were red +before this commit. - - Commit 2/4: `c51_loss_kernel.cu` adds `aux_conf_at_state` arg + gate - computation at the Bellman projection. NULL-tolerant (gate=1.0 - no-op). - - Commit 3/4: launcher in `gpu_dqn_trainer.rs::launch_c51_loss` threads - `self.aux_conf_at_state_buf.raw_ptr()` as the new last arg; unit - tests for the gate formula. - - Commit 4/4: audit-doc consolidation + close-out. +### Plan accuracy errata + +The user-supplied plan said `Add aux_conf_at_state_buf field to GpuBatch +struct in fused_training.rs` (commit 1 task 2). On closer inspection, +`GpuBatch` (in `crates/ml-dqn/src/replay_buffer_type.rs`) does NOT +currently carry the SP13 B1.1b `aux_sign_labels_ptr` either — that +direct-gather buffer lives only on the trainer (`aux_nb_label_buf`) and +is consumed in-kernel by `launch_aux_*` reading +`self.aux_nb_label_buf.raw_ptr()` directly. Mirroring this pattern, +`aux_conf_at_state_buf` is a trainer-only buffer; the launcher reads +`self.aux_conf_at_state_buf` directly without plumbing through +`GpuBatch`. The PER direct-gather populates it via the +`set_trainer_aux_conf_ptr` mechanism. No `GpuBatch` field was added. + +### Confidence + +Medium-high that the gate fires correctly on real data: + + - Math invariants verified (3 CPU tests). + - Direct-gather wiring verified end-to-end (GPU behavioral test — + inserts → samples → reads back the trainer destination buffer + populated by PER's `gather_f32_scalar`). + - Kernel signature change compiles + matches the launcher arg list. + - `feedback_no_partial_refactor`-compliant: every consumer in the + chain (kernel signature, launcher, replay-buffer setter) updated + atomically across the 4 commits. + +The remaining unknown is end-to-end smoke (whether the gated reward +actually changes training dynamics in the expected direction). That +verification is deferred to a future smoke run; per the controller +instruction, no smokes were dispatched in this implementation pass. ## 2026-05-10 — MAX_UPLOAD_BYTES 2GB → 8GB