Files
foxhunt/crates
jgrusewski 5ba9f376c4 fix(magnitude): mean-logit Bellman target — close the last C51 feedback path
The MSE and C51 loss kernels computed Bellman targets using C51's
distributional softmax expected Q for the argmax over next-state actions.
For magnitude (d==1), this structurally favored Small (tight distribution
→ higher softmax expected Q), creating an irrecoverable target feedback
loop even when C51 gradient was zeroed.

Fix: for magnitude branch (d==1), use mean-logit Q (average of V+A
across atoms without softmax weighting) for both argmax selection AND
target Q computation. This is variance-neutral — only the average
advantage level matters, not the distributional shape.

Applied to all 3 kernels:
- mse_loss_kernel.cu: argmax + target_eq
- c51_loss_kernel.cu: argmax
- expected_q_kernel.cu: expected Q for backtest evaluator

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-09 00:28:03 +02:00
..