Per `pearl_q_thompson_actor_makes_pi_dead_weight` follow-up + #35
deferral: the K-loop in step_with_lobsim was running full
step_synthetic K times per env step (Q + π + V + encoder + LR
controller emit + ISV refresh + EMA inputs). At K=4 (default) or K=8
(prior K_MAX) this caused PPO overshoot — KL excursions to 12.44 in
f2ggr, policy drift faster than the env step rate, gradient
overtraining on the same env-step's h_t.
## Fix: extract dqn_replay_step helper
New public method `dqn_replay_step(b_size)`:
1. Forward Q on sampled_h_t + sampled_h_tp1 (Double-DQN argmax)
2. Bellman target via TARGET net at h_tp1 + select + project
3. Q backward (logits → grad_w/b/h_t)
4. Per-batch reduce → grad_w/grad_b
5. Q Adam (uses LR already set by step_synthetic — no re-fire of
the LR controller per K iter)
6. Writes td_per_sample_d for PER priority update by caller
Discards Q's grad_h_t per R7d stop-grad (same as step_synthetic).
What dqn_replay_step does NOT do:
* π forward / surrogate / Adam — runs once per env step in
step_synthetic
* V forward / backward / Adam — same
* Encoder backward / grad combine — same
* LR controller emit + ISV mirror refresh — same
* EMA inputs (entropy, KL, advantage_var, td_kurtosis) — same
## K-loop in step_with_lobsim
for k_iter in 0..k_updates {
let per_indices = sample_and_gather(b_size)?;
if k_iter == 0 {
stats = step_synthetic(snapshots)?; // full update
} else {
dqn_replay_step(b_size)?; // Q-only
}
// PER priority update
}
Result:
* Q gets K Adam updates per env step (K-fold variance reduction)
* π + V + encoder get 1 Adam update per env step (no overshoot)
* LR controllers fire once per env step (no double-counting of
plateau detection)
* At b_size=16 with low advantage_var_ratio (batch averaging
reduces noise), K-loop typically settles at K=1 — the split
becomes a no-op in the steady state. At b_size=1 fallback or
high-noise regimes, the split materially reduces PPO drift.
## Code duplication
dqn_replay_step duplicates ~120 lines of Q-section code from
step_synthetic. Acceptable temporary tech debt — full dedupe would
require restructuring step_synthetic to call dqn_replay_step
internally, which is a larger refactor with regression risk. Marked
TODO for a follow-up commit once the b_size=16 + π-actor + K-split
architecture is empirically validated.
## Verified gates (local sm_86)
G1 isv_bootstrap ✅
G3 controllers ✅
G4 target_update ✅
integrated_smoke ✅
## No smoke yet
alpha-rl-9k9x6 (commit 3737feb66, π-actor + b_size=16) is still in
flight; submitting a new smoke would compete for L40S GPU. This
commit lands on remote; smoke will be submitted after 9k9x6 lands
and we've analyzed whether the b_size=16 + π-actor architecture
worked. If 9k9x6 shows Q learning unblocked, this split is polish.
If it doesn't, the split becomes the next experiment.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>