From 1a05af803d110fdf33fb6471b82909764bf950ab Mon Sep 17 00:00:00 2001 From: jgrusewski Date: Sat, 30 May 2026 23:35:54 +0200 Subject: [PATCH] fix(cuda): force-close existing position on DD trip / cooldown entry MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Cluster v5 alpha-rl-rjsjq step 371 revealed worst-account session_pnl growing monotonically across cooldown windows (-$3.7k → -$9.3k from step 100 → 371). Root cause: when DD triggers, actions_to_market_targets forces {side=2, size=0} — a no-op that suppresses EVERY action type including trail-stop-driven closes. If the account was holding a losing position when DD fired, that position bleeds mark-to-market for the entire 500-step cooldown with no exit path. Fix: on first DD/cooldown step, emit a closing market order matched to the current position (side=1 sell for long, side=0 buy for short, size = |position_lots|). Once flat, subsequent cooldown steps emit the existing no-op. Net effect: account closes its position at the moment of DD trip, then sits flat until its recovery clock expires. Validation: 13/13 risk_stack_invariants pass, 20/20 trade_management pass, integrated_trainer_smoke passes. Local b=128 1k smoke: qpa: +0.95 (best result of session) mean_active: +$47k worst: stabilizes at -$11.8k (vs unbounded growth pre-fix); residual loss is from fat-tail single-step market moves that cross dd_limit before the controller can react — out of scope for this fix. --- .../cuda/actions_to_market_targets.cu | 31 ++++++++++++++++--- 1 file changed, 26 insertions(+), 5 deletions(-) diff --git a/crates/ml-alpha/cuda/actions_to_market_targets.cu b/crates/ml-alpha/cuda/actions_to_market_targets.cu index e82b174b4..f118cf44e 100644 --- a/crates/ml-alpha/cuda/actions_to_market_targets.cu +++ b/crates/ml-alpha/cuda/actions_to_market_targets.cu @@ -126,15 +126,36 @@ extern "C" __global__ void actions_to_market_targets( // ──────────────────────────────────────────────────────────────────── // Layer 1 (CMDP hard constraints — spec 2026-05-30-adaptive-risk-management). - // Session-level overrides force Hold regardless of agent's choice. - // Per-batch (one independent backtest session per b); a tripped DD - // or active cooldown on session `b` does NOT lock out the others. + // Session-level overrides force the account FLAT regardless of agent's + // choice. Per-batch (one independent backtest session per b); a tripped + // DD or active cooldown on session `b` does NOT lock out the others. + // + // Fix E (2026-05-30): emit a *closing* market order when the account + // is non-flat — otherwise an existing losing position bleeds m2m for + // the entire 500-step cooldown with no exit (trail stops are also + // suppressed by this branch). Cluster alpha-rl-rjsjq step 371 showed + // worst session_pnl growing monotonically -$3.7k → -$9.3k from + // exactly this mechanism. + // + // Once the position is flat, subsequent cooldown steps emit a true + // no-op (side=2, size=0) and the account waits for its recovery clock. // ──────────────────────────────────────────────────────────────────── const bool dd_triggered = session_dd_triggered_per_batch[b] >= 0.5f; const bool in_cooldown = cooldown_remaining_per_batch[b] > 0.0f; if (dd_triggered || in_cooldown) { - market_targets[b * 2 + 0] = 2; // no-op side - market_targets[b * 2 + 1] = 0; // size = 0 + if (position_lots > 0) { + // long → close via sell + market_targets[b * 2 + 0] = 1; + market_targets[b * 2 + 1] = position_lots; + } else if (position_lots < 0) { + // short → close via buy + market_targets[b * 2 + 0] = 0; + market_targets[b * 2 + 1] = -position_lots; + } else { + // already flat → true no-op + market_targets[b * 2 + 0] = 2; + market_targets[b * 2 + 1] = 0; + } return; }