Two issues found in the H100 production run:
1. Primary gradient budget dropped from 1.0 to 0.10 when C51 alpha
ramped to 1.0 (c51_frac=0.10 with IQN at 60%). Branch heads
depend entirely on primary gradient — starved → grad_norm=0.000
at epoch 7. Fix: floor c51_frac at 0.30 so branch heads always
get meaningful gradient.
2. val_loss identical at epochs 1-2 (6.228316) because the backtest
evaluator's CUDA Graph was never invalidated between epochs.
Fix: call invalidate_dqn_graph() before each evaluation.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>