The forward pass (action selection) must apply the same advantage standardization as the loss kernels. Otherwise the network's raw advantages can encode stale C51 preferences that don't match the standardized training signal. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>