Val-Flat-collapse fix#3 revision (task #94, 2026-04-24). Prior
commit 543e3c11b used `-0.5 × holding_cost_rate × vol_proxy` — the
0.5 is a hardcoded tuned constant violating
`feedback_isv_for_adaptive_bounds.md` and
`feedback_adaptive_not_tuned.md`.
Replace with ISV-driven per-sample conviction, already computed by
the action-select kernel and threaded into env_step via
`conviction_ptr → conviction_core`:
reward_flat = -shaping_scale * holding_cost_rate
* conviction_core * vol_proxy_flat
`conviction_core ∈ [0, 1]` = direction-branch Q-range normalised by
`ISV[21]` (q_dir_abs_ref EMA). Self-adapting properties:
- Cold start / ISV[21] uninitialised → fallback 1.0 → full penalty,
encourages early exploration out of the flat equilibrium.
- Low conviction (uncertain direction) → penalty scales toward 0
→ doing nothing is acceptable when there's no signal (matches
real-world: flat cost is only real when there's opportunity).
- High conviction (strong directional edge) → penalty scales up
→ Flat becomes expensive ONLY where the model itself says
there's an edge. Forces the policy to take action exactly
where it has belief, not blindly.
Continuity: conviction is continuous ∈ [0, 1], no step function. The
temporal Mamba2 layers in the trunk feed the Q-values that drive
conviction, so conviction inherits temporal history — the model can
learn "market has been signalling for N bars, time to try" without
an explicit time-since-last-trade feature.
Training-time exploration (Boltzmann sampling, epsilon floor 2%,
NoisyNets σ, count bonus) is unchanged. The Flat-cost shifts the
learned Q-values so that deterministic val argmax picks trade-
actions where the training-time exploration already found edge.
Hold semantics retained:
- Hold while position≠0 (in-trade stance) → positioned-bar
holding-cost branch (unchanged)
- Hold while position=0 AND Flat (no-op outcomes) → this branch
→ conviction-scaled opportunity cost.
The distinction is structural by portfolio state, not by dir label.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>