Current Flat opp-cost was conviction-driven but |Q|-scale was implicit.
When training drifts Q magnitudes into ±50, opp-cost becomes invisible
relative to action Q-values → Flat wins argmax. Fix: multiply by
isv_signals_ptr[21] (Q_DIR_ABS_REF_INDEX, EMA of max(|Q_mean|) across
direction bins, populated by update_eval_v_range / q_stats_kernel).
Self-scaling: opp-cost tracks |Q| proportionally throughout training.
No tuned multiplier; relies on existing ISV slot. Floor 1e-3 for cold-
start protection before EMA is warm. rc[4] is now written in the Flat
branch so reward_component_ema kernel picks it up into ISV[67].
No parallel paths — the old formula IS modified in place.
Plan 3 Task 2. Spec §4.B.1.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>