Files
foxhunt/crates/ml-backtesting
jgrusewski 78b781a523 fix(rl): root-cause overfit fixes — value leak + environment timeout
Two architectural fixes addressing the "never close" overfit attractor
that caused walk-forward fold 0 to converge to dones=0 by step 2000.

1. Asymmetric clip on unrealized_R in trade_context (value leak fix):
   trade_context[1] = fminf(0.0f, unrealized_R)

   The model previously saw unrealized_R as a state feature. Q learned
   "high unrealized = high V(state)" → Q(close) < V(hold) → never close.
   Now Q only sees the LOSS side: losing positions visible (cut losses),
   winning positions invisible (close decision driven by general policy,
   not state-conditioned on profit). Realized PnL on close still
   teaches "take profits" via Bellman backup.

2. Enable max_hold_ns = 60s (environment constraint):
   LobSimCuda::max_hold_ns_d was alloc_zeros (disabled). The lobsim
   has session-gap force-close (>1 hour gaps) but no per-trade timeout,
   letting the agent hold positions indefinitely within sessions.
   Now alpha_rl_train sets max_hold to 60s via new upload_max_hold_ns
   method backed by fill_u64.cu kernel (device-side, complies with
   feedback_no_htod_htoh_only_mapped_pinned).

These are environment/architectural fixes, NOT reward shaping hacks.
Together they address why the model COULD learn "never close" and
ensure dones signal density for Q-learning.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-28 11:13:39 +02:00
..