Files
foxhunt/docs/plans
jgrusewski 7eae832f24 docs(sp22): H6 Phase 3 runbook — atom-shift math derivation for Steps 7+8
Derived the projection-arithmetic transformation that collapses per-(b, a)
atom-position shift into a SINGLE per-sample reward adjustment:

  effective_reward = reward + gamma_eff * Delta_target * (1-done) - Delta_online

where
  Delta_online = W[a_d] * state_121[b]      (taken-action shift on online support)
  Delta_target = W[a*] * next_state_121[b]  (sampled-target shift on target dist)

The Bellman projection (block_bellman_project_f body) needs NO arithmetic
change — only the reward arg passed in. This dramatically reduces Step 7's
kernel surgery surface area.

For Step 8 backward, derived the gradient paths:
  dL/dDelta_online = (1/dz) * sum_n p_target_n * (log_p_online[upper_n] - log_p_online[lower_n])
  dL/dDelta_target = -gamma*(1-done) * dL/dDelta_online

  dL/dW[a_d]            += dL/dDelta_online * state_121
  dL/dW[a*]             += dL/dDelta_target * next_state_121
  dL/dstate_121[b]      += dL/dDelta_online * W[a_d]
  dL/dnext_state_121[b] += dL/dDelta_target * W[a*]

Critical constraint documented: Steps 7+8+11 MUST land atomically. Splitting
Step 7 forward from Step 8 backward creates a silent gradient path through
the aux head (missing dDelta_online/dnetwork_weights). Splitting Step 8 from
Step 11 Adam accumulates dW without updating W — no learning. Per
feedback_no_partial_refactor.md.

The math now makes Step 7+8 a transcription job rather than exploration.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-13 01:45:40 +02:00
..