Derived the projection-arithmetic transformation that collapses per-(b, a)
atom-position shift into a SINGLE per-sample reward adjustment:
effective_reward = reward + gamma_eff * Delta_target * (1-done) - Delta_online
where
Delta_online = W[a_d] * state_121[b] (taken-action shift on online support)
Delta_target = W[a*] * next_state_121[b] (sampled-target shift on target dist)
The Bellman projection (block_bellman_project_f body) needs NO arithmetic
change — only the reward arg passed in. This dramatically reduces Step 7's
kernel surgery surface area.
For Step 8 backward, derived the gradient paths:
dL/dDelta_online = (1/dz) * sum_n p_target_n * (log_p_online[upper_n] - log_p_online[lower_n])
dL/dDelta_target = -gamma*(1-done) * dL/dDelta_online
dL/dW[a_d] += dL/dDelta_online * state_121
dL/dW[a*] += dL/dDelta_target * next_state_121
dL/dstate_121[b] += dL/dDelta_online * W[a_d]
dL/dnext_state_121[b] += dL/dDelta_target * W[a*]
Critical constraint documented: Steps 7+8+11 MUST land atomically. Splitting
Step 7 forward from Step 8 backward creates a silent gradient path through
the aux head (missing dDelta_online/dnetwork_weights). Splitting Step 8 from
Step 11 Adam accumulates dW without updating W — no learning. Per
feedback_no_partial_refactor.md.
The math now makes Step 7+8 a transcription job rather than exploration.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>