Files
foxhunt/crates
jgrusewski 7eccfc53c9 feat: Q-gap floor gradient — perpetual action differentiation pressure
Adds a spread gradient to the C51 advantage logits that pushes the
taken action's distribution toward higher atoms and non-taken toward
lower. Scale = inv_batch * delta_z (adaptive to per-sample atom
resolution, zero hardcoded constants).

This gradient is ORTHOGONAL to the Bellman equation — it depends on
atom position, not target match. Active on ALL samples, providing
perpetual pressure to differentiate Q-values even when the C51
cross-entropy gradient vanishes at convergence. Prevents the Q-gap
plateau where all actions have identical Q-values.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-14 21:07:23 +02:00
..