apply_reward_scale.cu's post-scale [-3,+1] clamp was masking ~99.99% of
realized tail magnitudes from the agent's reward signal: local b=16 smoke
showed train pnl_cum_usd −$0.63 (clamp-truncated) vs realized_pnl_cum_usd
−$8,574.63 (raw raw_rewards sum), a 13,600× compression. Per van Hasselt
2016, popart standardization + F4 envelope are designed to handle tail
magnitudes; the clamp fights them.
Changes:
- RL_REWARD_CLAMP_ENABLED_INDEX (slot 724): default 0 (disabled). When 0
the kernel skips the asymmetric clamp; scaled rewards pass through to
rewards[b] unchanged. Legacy behavior restored by setting to 1.
- DiagInputs.realized_pnl_cum_usd: parallel counter computed from
raw_rewards (pre-scale, pre-clamp shaped pnl). Compare against
trading.pnl_cum_usd to surface clamp-truncation gaps.
- Trainer accumulates realized_pnl_cum_usd in both train + eval loops
per closed-trade done-step (same pattern as pnl_cum_usd).
- EXPECTED_LEAVES 651 → 652 for the new diag leaf.
Validation:
- 200+100 b=16 fold-1 smoke clean; train leaves = eval leaves = 652;
realized_pnl_cum_usd diverges from pnl_cum_usd as expected when clamp
disabled (the reward-hacking gap is now observable).
- compute-sanitizer pending (cluster).
B-8 (popart σ_welford disaggregation) and B-9 (C51 Bellman-target
saturation observability) specs at docs/superpowers/specs/2026-06-01-*
build on this slot allocation (725, 726-729 next).
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>