Files
foxhunt/crates/ml-alpha/examples
jgrusewski b15bc71fb6 feat(rl): specialized wr-targeting LOSS aversion controller
The previous adaptive LOSS controller (rl_reward_clamp_controller) drove
LOSS toward 1.0 floor via observed L/W trade ratio, killing Q's loss
aversion and producing wr=0.27 trend-follower. The dd049d9a4 baseline
achieved wr=0.567 with LOSS=3.0 STATIC, but disabling adaptation
caused NaN at step 4 (other components depend on LOSS-driven signals).

Solution: build a SPECIALIZED ISV controller that drives LOSS toward
maintaining target wr, decoupled from observed L/W trade ratio.

New: rl_loss_aversion_controller.cu
- Single-thread Schulman bounded-step controller
- Input: wr_ema (slot 590, EMA of trade outcomes on done events)
- Target: wr_target (slot 591, bootstrap 0.55 — surfer)
- Output: LOSS clamp (slot 453)
- ±10% adjustment per fire, asymmetric dead-zone (loss aversion bias)
- Bounds [1.5, 5.0], sparse-aware (skips if wr_ema = 0)

Modified rl_reward_clamp_controller.cu:
- Added wr_ema update on done events (single source of truth alongside
  win/loss magnitude EMAs)
- LOSS slot 453 ownership transferred to rl_loss_aversion_controller

Local smoke (1000 steps, b=128, seed=16962):
- wr_ema trajectory: 0.32 → 0.42 → 0.53 → 0.49 (climbing toward 0.55)
- LOSS auto-adjusts: 4.71 → 4.38 → 4.29 → 3.52 (tightening as wr rises)
- Action entropy stable 2.0-2.2 (no collapse)
- Reward clamp confirmed active (r_min saturates at -LOSS)
- No NaN, completed_clean: true

Per pearl_loss_clamp_controls_entropy_stability: LOSS≥3.0 correlates
with high wr + stable entropy. Specialized controller maintains this
structural property regardless of observed trade ratio noise.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-28 19:22:10 +02:00
..