fix(rl): sparse-aware EMA + Q→π distillation breaks defensive trap
Two coupled fixes addressing vj5f6 findings:
(1) WIN_clamp oscillation — sparse-aware EMA
vj5f6 showed WIN_clamp oscillating 1.0 ↔ 67.0 across 40k steps.
Root cause: the Wiener-α blend in rl_reward_clamp_controller
treated pos_max=0 as "no win this step ≡ win magnitude is zero,"
exponentially decaying the EMA toward 0 during dry-spell windows
(no closed winning trades). With α=0.4, ten dry steps decayed EMA
by 0.6^10 ≈ 0.006, collapsing WIN back to MIN_WIN=1.0 floor.
Fix: only update pos_max_ema AND clip_rate_ema AND MARGIN when
pos_max > 0. A dry step is "no signal," not "zero signal." The
EMA retains its last winning-period estimate; the controller
doesn't ratchet on stale data.
(2) Q→π distillation — couples Q's improved calibration to π
vj5f6 showed l_q dropping 100× (2.37 → 0.02) but reward economics
IDENTICAL to 8xwq8 (no C51 V_MAX lift). Per Option B, π drives
action selection but is trained by PPO surrogate using advantage
= returns - V. V regression doesn't benefit from C51 calibration,
so Q's improved knowledge stays trapped in the critic head.
Deep audit revealed a self-reinforcing defensive trap:
Q learned "big positions lose money" → π_target favors small
actions → π picks a3+a4 (tiny long / Hold) → position lots ≈ 0
→ rewards mostly 0 → V learns "everything is 0" → V_pred ≈ 0
→ advantage = returns - V_pred ≈ 0 → PPO gradient ≈ 0 → π
frozen at defensive attractor → loop. Trade count dropped 3×
(rdgzl 25k → 8xwq8/vj5f6 9k closes per 10k steps), win rate
inversely correlated with l_q (50% early → 22% late) because
only forced closes happen (stops = losses).
Fix: new rl_q_pi_distill_grad.cu computes
π_target = softmax(E_Q[s,*] / τ)
∂L/∂logits[a] = λ × (π_new(a) - π_target(a))
and ADDS this gradient to pi_grad_logits AFTER the PPO surrogate
backward. Couples Q's preferences directly into π's update without
going through advantage. λ=0.01 (small, PPO dominant), τ=1.0
(canonical Boltzmann). 3 new ISV slots (λ + τ + KL_ema diag).
Diag exposes c51_v_max/v_min, q_distill_lambda/temperature, and
q_distill_kl_ema so the adaptation + distillation loop is observable.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
This commit is contained in:
@@ -70,6 +70,7 @@ const KERNELS: &[&str] = &[
|
||||
"rl_pi_action_kernel", // audit Option B: π drives action selection via multinomial sampling from softmax(pi_logits); Q becomes pure critic
|
||||
"rl_reward_clamp_controller", // audit 2026-05-24: adaptive [-LOSS, +WIN] clamp from positive-tail EMA; replaces static [-3, +1] that crushed winning-trade signal in rmgm5
|
||||
"rl_atom_support_update", // audit 2026-05-24 followup: refreshes atom_supports_d from ISV V_MIN/V_MAX so C51 atom span adapts with reward clamp (Q learning was capped at V_MAX=1.0)
|
||||
"rl_q_pi_distill_grad", // audit 2026-05-24 vj5f6 followup: KL(softmax(E_Q/τ) || π_new) gradient ADDED to pi_grad_logits — couples Q's improved C51 calibration to action selection (was decoupled per Option B)
|
||||
];
|
||||
|
||||
// Cache bust v31 — five new reduce / derive kernels populate the input
|
||||
|
||||
Reference in New Issue
Block a user