Two architectural fixes addressing the "never close" overfit attractor
that caused walk-forward fold 0 to converge to dones=0 by step 2000.
1. Asymmetric clip on unrealized_R in trade_context (value leak fix):
trade_context[1] = fminf(0.0f, unrealized_R)
The model previously saw unrealized_R as a state feature. Q learned
"high unrealized = high V(state)" → Q(close) < V(hold) → never close.
Now Q only sees the LOSS side: losing positions visible (cut losses),
winning positions invisible (close decision driven by general policy,
not state-conditioned on profit). Realized PnL on close still
teaches "take profits" via Bellman backup.
2. Enable max_hold_ns = 60s (environment constraint):
LobSimCuda::max_hold_ns_d was alloc_zeros (disabled). The lobsim
has session-gap force-close (>1 hour gaps) but no per-trade timeout,
letting the agent hold positions indefinitely within sessions.
Now alpha_rl_train sets max_hold to 60s via new upload_max_hold_ns
method backed by fill_u64.cu kernel (device-side, complies with
feedback_no_htod_htoh_only_mapped_pinned).
These are environment/architectural fixes, NOT reward shaping hacks.
Together they address why the model COULD learn "never close" and
ensure dones signal density for Q-learning.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>