Files
foxhunt/crates
jgrusewski 33376525ac fix(dqn): Flat opportunity cost scales with ISV-driven conviction, not tuned constant
Val-Flat-collapse fix #3 revision (task #94, 2026-04-24). Prior
commit 543e3c11b used `-0.5 × holding_cost_rate × vol_proxy` — the
0.5 is a hardcoded tuned constant violating
`feedback_isv_for_adaptive_bounds.md` and
`feedback_adaptive_not_tuned.md`.

Replace with ISV-driven per-sample conviction, already computed by
the action-select kernel and threaded into env_step via
`conviction_ptr → conviction_core`:

    reward_flat = -shaping_scale * holding_cost_rate
                * conviction_core * vol_proxy_flat

`conviction_core ∈ [0, 1]` = direction-branch Q-range normalised by
`ISV[21]` (q_dir_abs_ref EMA). Self-adapting properties:

  - Cold start / ISV[21] uninitialised → fallback 1.0 → full penalty,
    encourages early exploration out of the flat equilibrium.
  - Low conviction (uncertain direction) → penalty scales toward 0
    → doing nothing is acceptable when there's no signal (matches
    real-world: flat cost is only real when there's opportunity).
  - High conviction (strong directional edge) → penalty scales up
    → Flat becomes expensive ONLY where the model itself says
    there's an edge. Forces the policy to take action exactly
    where it has belief, not blindly.

Continuity: conviction is continuous ∈ [0, 1], no step function. The
temporal Mamba2 layers in the trunk feed the Q-values that drive
conviction, so conviction inherits temporal history — the model can
learn "market has been signalling for N bars, time to try" without
an explicit time-since-last-trade feature.

Training-time exploration (Boltzmann sampling, epsilon floor 2%,
NoisyNets σ, count bonus) is unchanged. The Flat-cost shifts the
learned Q-values so that deterministic val argmax picks trade-
actions where the training-time exploration already found edge.

Hold semantics retained:
  - Hold while position≠0 (in-trade stance) → positioned-bar
    holding-cost branch (unchanged)
  - Hold while position=0 AND Flat (no-op outcomes) → this branch
    → conviction-scaled opportunity cost.
The distinction is structural by portfolio state, not by dir label.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-24 00:54:32 +02:00
..