fix(magnitude): var_scale permanent floor — eval Full reachable via conviction
Smoke (with mag_concat off-by-one fixed and fxcache regenerated):
- training MAG_DIST: Q=0.287 H=0.301 F=0.412 (healthy diversity)
- INTENT at eval: Q=0.045 H=0.045 F=0.911 (network learned Full)
- realised at eval: Q=0.592 H=0.408 F=0.000 (Full silenced)
Even though q_f >> q_h >> q_q (network strongly prefers Full), eval
realised Full = 0.000. Position-sizing pipeline composes four
multiplicative shrinks at experience_kernels.cu:1615-1672:
effective_max_pos = max_position
× cvar_scale (line 1621)
× q_gap_conviction (line 1628, clamped [0.25,1])
× kelly_f (line 1649, only if >20 trades)
× var_scale (line 1671, = 1/(1+sqrt(var_q)))
Each term well-bounded individually but composing them silences the
policy at validation when var_q persists high (var_q ≈ 30 → var_scale
≈ 0.15; even with conviction = 1.0, kelly_f = 1.0, the compound 0.15
falls in the Quarter bucket [0, 0.375)). Network INTENT reaches the
target_position kernel correctly; var_scale strips it back out.
Same family as val-Flat-collapse (warm-branch Kelly = 0 from balanced
priors) — fix is the same pearl
(`pearl_blend_formulas_must_have_permanent_floor.md`):
var_scale = max(var_scale, q_gap_conviction)
The q_gap-derived conviction is already an adaptive ISV-coupled signal
(line 1628), already clamped to [0.25, 1.0]. Using it as a permanent
floor on var_scale lets the policy's magnitude intent reach the
realised position when conviction is high, regardless of variance.
No tuned constants — feedback_adaptive_not_tuned + feedback_isv_for_adaptive_bounds
both honoured by reusing the existing adaptive bound rather than
introducing a new threshold.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -1654,7 +1654,19 @@ extern "C" __global__ void experience_env_step(
|
||||
/* Distributional variance position sizing (Kelly from C51 atoms).
|
||||
* High Var[Q] = uncertain outcome → smaller position.
|
||||
* Low Var[Q] = confident → full position.
|
||||
* Scale: 1 / (1 + sqrt(Var[Q_taken])). Bounded (0, 1]. */
|
||||
* Scale: 1 / (1 + sqrt(Var[Q_taken])). Bounded (0, 1].
|
||||
*
|
||||
* Permanent floor: when the policy's own conviction signal (q_gap) is
|
||||
* high, var_scale must not silence the magnitude intent. Same pearl as
|
||||
* the Kelly warmup_floor (`pearl_blend_formulas_must_have_permanent_floor.md`):
|
||||
* cascading multiplicative shrinks (cvar × q_gap_conviction × kelly × var_scale)
|
||||
* compose to near-zero at validation when var_q persists high, even
|
||||
* though every individual term is "well-bounded". The fix is the same:
|
||||
* floor var_scale at the q_gap-derived conviction already applied
|
||||
* upstream — it IS the adaptive bound (per feedback_isv_for_adaptive_bounds.md),
|
||||
* not a tuned constant. With this floor, high conviction overrides
|
||||
* variance shrinkage so Full intent at eval can produce Full realized
|
||||
* positions even before var_q has fully decayed. */
|
||||
if (q_variance != NULL) {
|
||||
/* Taken action's factored index: sum of branch-local indices */
|
||||
int taken_branch_idx = dir_idx; /* direction is branch 0 */
|
||||
@@ -1664,6 +1676,11 @@ extern "C" __global__ void experience_env_step(
|
||||
/* Use max variance across the two exposure branches (direction + magnitude) */
|
||||
float var_q = fmaxf(var_taken, var_mag);
|
||||
float var_scale = 1.0f / (1.0f + sqrtf(var_q));
|
||||
/* Permanent floor — same q_gap-derived conviction clamped at line 1628. */
|
||||
float var_floor = (q_gaps != NULL)
|
||||
? fminf(fmaxf(q_gaps[i] * 0.5f, 0.25f), 1.0f)
|
||||
: 0.25f;
|
||||
var_scale = fmaxf(var_scale, var_floor);
|
||||
/* Task 0.3: persist the per-sample Kelly scale for host-side epoch-mean
|
||||
* reduction. Host filters on `> 0.0f` to distinguish measured samples
|
||||
* from unreached/Var[Q]-disabled default-zero slots. */
|
||||
|
||||
@@ -2,6 +2,20 @@
|
||||
|
||||
**Status:** Populated during Plan 1 Task 6 (A.5 orphan audit). Updated on every commit per Invariant 7.
|
||||
|
||||
var_scale permanent floor for eval magnitude collapse (2026-04-27): the
|
||||
EVAL_DIST=Quarter 0.59 / Half 0.41 / Full 0.00 pathology persists even with
|
||||
intent=0.911 for Full because the position-sizing pipeline composes four
|
||||
multiplicative shrinks (cvar_scale × q_gap_conviction × kelly_f × var_scale)
|
||||
each well-bounded individually but compounding to ≪ 0.75. Kelly already
|
||||
fixed via permanent-floor pearl (`pearl_blend_formulas_must_have_permanent_floor.md`);
|
||||
applied the same pearl to var_scale at `experience_kernels.cu:1666`:
|
||||
`var_scale = max(var_scale, q_gap_conviction)`. The conviction signal IS
|
||||
the adaptive bound (per `feedback_isv_for_adaptive_bounds.md`), already
|
||||
clamped to [0.25, 1.0] at line 1628. With this floor, high conviction
|
||||
overrides Var[Q] shrinkage so the policy's magnitude intent reaches the
|
||||
realised position even before var_q decays. No tuned constants — same
|
||||
adaptive q_gap signal used upstream.
|
||||
|
||||
data_source globalization to mbp10 (2026-04-27): completes the alignment
|
||||
started in the earlier "fxcache data_source alignment" entry below. Smoke
|
||||
(`dqn-smoketest.toml`) and localdev (`dqn-localdev.toml`) profiles
|
||||
|
||||
Reference in New Issue
Block a user