Files
foxhunt/crates/ml-backtesting/cuda/decision_policy.cu
jgrusewski b44a97ff94 diag(crt): empirical measurement battery for signal × controller × cost analysis
Pure-instrumentation commit. Zero behavior change. The Gate CRT.1 smoke
p9cnk (commit 1e656948b) produced 222k trades and 29688% drawdown,
worse than the failed CRT.A. The structural fixes (multi-horizon
conviction, no-trade band, composite exit) all work as designed — but
the trade cadence is fundamentally too fast for the signal/cost ratio.
Before another round of threshold tuning, measure the structural truth.

Adds device-side counters dumped at smoke close (single end-of-run
memcpy_dtoh batch in read_diagnostics(), NOT per-event):

Group A — per-horizon signal persistence:
  - flip_count[h]:        total direction-sign flips
  - sum_run_length[h]:    cumulative events between flips (u64)
  - run_length_hist[h]:   5-bucket log-scale histogram (1-9, 10-99,
                          100-999, 1k-9.9k, 10k+)
  Output: mean_run_len per horizon = signal half-life proxy

Group B — conviction smoothed-EMA distribution:
  - conv_hist[10]:        10-bucket histogram over [0, 1)
  Output: shape of conviction distribution at decision time

Group C — hold-time distribution:
  - hold_hist[6]:         <1s, 1-10s, 10-60s, 1-10m, 10-60m, >1h
  Output: empirical distribution of trade hold times

Group D — outcome by entry-conviction:
  - outcome_n[10]:        trades per conviction bucket
  - outcome_sum_pnl[10]:  cumulative realized PnL per bucket (price-units)
  - outcome_n_wins[10]:   wins per bucket
  Output: validates "high conviction = better trades" hypothesis OR
          surfaces signal miscalibration

Per-backtest memory: ~330 bytes. Negligible vs existing state.

All counters single-writer-per-block (threadIdx.x == 0 convention; no
atomicAdd per feedback_no_atomicadd). All updates kernel-internal (no
host roundtrip per event). The end-of-run dump uses memcpy_dtoh in
read_diagnostics() — called once at harness termination from the
post-stop_ctrl_counters block, never per-event. Per
feedback_no_htod_htoh_only_mapped_pinned, the hot path is untouched.

N_HORIZONS = 5 (heads.rs canonical {30, 100, 300, 1000, 6000}); the
plan text says 4 but the workspace constant is 5 — used the code value.

Output at end of cluster smoke as a `crt_diag` log block:
  crt_diag h30: flips=X mean_run_len=Y events
  crt_diag h30 run_length_hist: 1-9:N 10-99:N 100-999:N ...
  ...
  crt_diag conv_ema_hist: [0.0-0.1]:N [0.1-0.2]:N ...
  crt_diag hold_time_hist: <1s:N 1-10s:N ...
  crt_diag outcome_by_entry_conv[0.0-0.1]: n=N win_rate=X.X% mean_pnl_pu=Y.YYY
  ...

The output drives the next round of CRT.1 threshold design (delta_floor,
composite exit factors, min-hold cooldown) from data instead of guesses.

Per pearl_adaptive_not_tuned and pearl_controller_anchors_isv_driven:
thresholds will become ISV-derived in CRT.2; CRT.1 must first establish
defensible defaults from the empirical signal characteristics.

Verified: stop_controller (22/22 pass), decision_floor_coldstart (3/3
pass), threshold_and_cost (3/3 pass), parallel_sim_correctness (1/1
pass), lob_sim_integrated_fuzz (3/3 pass). Two lob_sim_fixtures
failures (fix_decision_alpha_buy_close, fix_decision_program_h4_only)
were pre-existing at HEAD 1e656948b — not caused by this commit.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-20 21:56:07 +02:00

999 lines
47 KiB
Plaintext
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
// decision_policy.cu — alpha → per-horizon Kelly → aggregate → market target.
//
// Hardcoded v1 policy = Strategy::default_for(max_lots) from C3:
// for each h in 0..N_HORIZONS:
// compute signed sizing from (alpha[h], IsvKellyState[h])
// aggregate via WeightedByRealizedSharpe:
// w[h] = max(0, recent_sharpe[h]) / sum_of_positive_sharpes
// final = sum_h w[h] * signed_size[h]
// final → side+size in market_targets[b]
//
// The bytecode VM from spec §6 is intentionally not implemented in v1 —
// the default Strategy::default_for produces exactly this shape and
// alternative compositions are deferred. Bytecode plumbing remains in
// src/policy/mod.rs ready for v2 expansion.
//
// Single-writer per block (thread 0). See spec §5.
#include "lob_state.cuh"
// Minimum closed-trade count before the variance-derived `cap_units`
// constraint is trusted. With N < this threshold the cap falls back to
// the host-supplied `max_lots` budget — a single-sample variance proxy
// (`ret²`) systematically over-estimates Var[ret] for a directionally
// biased loss/win and locks `cap_lots` to ~zero, which kills further
// trading. Threshold of 10 picked as the standard rule-of-thumb for
// "first moment estimate becomes informative" — same role as the
// pearl_first_observation_bootstrap pattern applied to the variance side.
#define MIN_TRADES_FOR_VAR_CAP 10u
// CRT Phase A0.5 corrective: record a single max-conviction f32 per
// decision-tick into a host-allocated growing device buffer. The same
// max_conv calculation as the per-backtest threshold gate (decision_policy_default
// line ~179) but driven once per decision rather than per backtest, with
// the result staged on-device for a single end-of-run DtoV. Lets the
// harness keep the percentile-tuning side-channel alive without round-
// tripping host memory on the per-event hot path.
//
// Launch: grid=(1,1,1), block=(1,1,1). Single thread.
// convictions_d : [max_decisions] device buffer
// write_idx : host-managed write head (passed as scalar)
// Writes one f32 into convictions_d[write_idx]. Bounds-clamps to
// max_decisions so a long-running smoke that exceeds the allocation
// can't OOB. Out-of-bound decisions just stop being recorded; the
// host-side counter knows the real decision_count.
extern "C" __global__ void record_max_conviction_to_slot(
const float* __restrict__ alpha_probs, // [N_HORIZONS]
float* __restrict__ convictions, // [max_decisions]
int write_idx,
int max_decisions
) {
if (threadIdx.x != 0 || blockIdx.x != 0) return;
if (write_idx < 0 || write_idx >= max_decisions) return;
float max_conv = 0.0f;
#pragma unroll
for (int h = 0; h < N_HORIZONS; ++h) {
float c = fabsf(alpha_probs[h] - 0.5f) * 2.0f;
// Match the harness's prior clamp: clamp(|p-0.5|*2, 0, 1).
if (c < 0.0f) c = 0.0f;
if (c > 1.0f) c = 1.0f;
if (c > max_conv) max_conv = c;
}
convictions[write_idx] = max_conv;
}
// Per spec §3 + §5: shared stop-check helper called from both
// decision_policy_default and decision_policy_program. Pure-device,
// no host branches (CUDA Graph capture safe per
// pearl_no_host_branches_in_captured_graph).
//
// Returns 1 if a stop fired (and market_targets[b] was written to
// the (3, 0) force-flat encoding); 0 otherwise. Caller checks return
// value to decide whether to skip the VM / sizing logic.
//
// Skeleton: gates on flat position only. SL/trail/averaging logic
// lands in plan Tasks 4-6 of the ISV-stop-controller series.
__device__ static int stop_check_isv(
int b,
const Pos* __restrict__ positions,
const unsigned char* __restrict__ isv_kelly_base,
const unsigned int* __restrict__ open_horizon_masks,
const float* __restrict__ atr_mid_ema,
float* __restrict__ trail_hwm,
const Book* __restrict__ books,
const float* __restrict__ cost_per_lot_per_side,
int* __restrict__ market_targets,
// S2.1: mh_kernel_calls counts stop_check_isv entries with non-zero position.
// Kept as an ongoing generic diagnostic (always-non-decreasing means SL/trail
// are being evaluated). The 6 max_hold-specific counters were removed when
// max_hold enforcement moved to resting_orders_step (event-rate).
unsigned int* __restrict__ mh_kernel_calls
) {
const Pos& pos = positions[b];
if (pos.position_lots == 0) return 0;
// S2.1: count every entry with a non-zero position.
mh_kernel_calls[b] += 1u;
// Spec §5 controller math — SL only in this task; trail in Task 5.
const float atr = atr_mid_ema[b];
const unsigned int mask = open_horizon_masks[b];
const IsvKellyState* isv = reinterpret_cast<const IsvKellyState*>(
isv_kelly_base + (size_t)b * N_HORIZONS * sizeof(IsvKellyState)
);
float ema_loss = 0.0f;
float ema_win = 0.0f;
float var_avg = 0.0f;
int n_open = 0;
#pragma unroll
for (int h = 0; h < N_HORIZONS; ++h) {
if (mask & (1u << h)) {
ema_loss += isv[h].pnl_ema_loss;
ema_win += isv[h].pnl_ema_win;
var_avg += isv[h].realised_return_var;
n_open += 1;
}
}
if (n_open > 0) {
ema_loss /= (float)n_open;
ema_win /= (float)n_open;
var_avg /= (float)n_open;
}
// Per pearl_trade_level_vol_for_stop_distance.md (2026-05-20):
// Trade-level noise floor = sqrt(realised_return_var) bootstrapped from
// cost². Cold-start (var=0): trade_vol = cost (structural minimum,
// sub-cost is sub-fee noise). Post-bootstrap (>= MIN_TRADES_FOR_VAR_CAP
// closes): sqrt(var) dominates — controller learns its own scale.
// cost appears once as a bootstrap sentinel; never as a distance multiplier.
const float cost = cost_per_lot_per_side[b];
const float trade_vol = sqrtf(fmaxf(var_avg, cost * cost));
// pearl_blend_formulas_must_have_permanent_floor — max, not blend.
const float sl_distance = fmaxf(fmaxf(ema_loss, atr), trade_vol);
const float trail_distance = fmaxf(fmaxf(ema_win, atr), trade_vol);
const bool is_long = pos.position_lots > 0;
const Book& bk = books[b];
const float close_px = is_long ? bk.bid_px[0] : bk.ask_px[0];
const float entry_px = pos.vwap_entry;
const float unrealized_pl_per_lot =
is_long ? (close_px - entry_px) : (entry_px - close_px);
const bool sl_fired = unrealized_pl_per_lot <= -sl_distance;
const float new_hwm = fmaxf(trail_hwm[b], unrealized_pl_per_lot);
trail_hwm[b] = new_hwm;
const bool trail_fired =
(new_hwm > trail_distance)
&& (unrealized_pl_per_lot <= new_hwm - trail_distance);
if (sl_fired || trail_fired) {
// §7 force-flat encoding.
market_targets[b * 2 + 0] = 3;
market_targets[b * 2 + 1] = 0;
return 1;
}
return 0;
}
// CRT.1 C1.2: Wiener-α adaptive EMA on the multi-horizon ISV-weighted
// conviction scalar (spec §4.4). Caller computes |conviction_signed| from
// the per-horizon weighted-signed sum and passes it as `conviction_abs`.
//
// Updates per-backtest second-order EMAs of sample variance + difference
// variance, derives adaptive α with floor 0.4 per
// pearl_wiener_alpha_floor_for_nonstationary, and blends. First-observation
// bootstrap (prev_ema == 0): replace directly per
// pearl_first_observation_bootstrap.
//
// Writes new EMA to conviction_ema[b], updates the two second-order
// variance EMAs, and returns the smoothed conviction magnitude. Direction
// is recovered by the caller from `conviction_signed` so that smoothing
// is applied only to magnitude — sign flips respond at event rate while
// magnitude jitter is suppressed.
//
// Single-thread-per-backtest invariant: caller must already have
// gated on threadIdx.x == 0.
// CRT.1 C1.4: composite exit_signal safety circuit-breaker per spec §4.3.
// Backup exit path for cases the primary `target → 0 from collapsed
// conviction` mechanism (C1.2) misses — most importantly when conviction
// stays bounded above zero but the trade is deep in unrealized loss.
//
// Three risk dimensions, max-of-three composite:
// degradation_term = (1 conv_ema_now / conviction_at_entry) / 2.0
// (50% conviction decay relative to entry fires)
// disagreement_term = disagree_consecutive_events / 32.0
// (32 events of short/long horizon disagreement)
// pnl_decay_term = max(pnl_adjusted_conv_ema, 0) / 1.0
// (pnl-adjusted conviction flips and grows < 1)
//
// Reads the trajectory fields populated by the decision kernel above.
// Writes force-flat (3, 0) on fire. Returns 1 if fired, 0 otherwise.
//
// Thresholds (2.0 / 32.0 / 1.0) are spec-bootstrap defaults; they migrate
// to ISV-driven values in CRT.2 alongside the rest of the controller anchors
// (pearl_controller_anchors_isv_driven). Single-thread-per-backtest invariant.
__device__ static int composite_exit_check(
int b,
const Pos* __restrict__ positions,
const unsigned char* __restrict__ open_trade_state,
int* __restrict__ market_targets
) {
if (positions[b].position_lots == 0) return 0;
const unsigned char* st_ro = open_trade_state + (size_t)b * 64;
const float conviction_at_entry = *reinterpret_cast<const float*>(st_ro + 20);
const float conv_ema_now = *reinterpret_cast<const float*>(st_ro + 24);
const float pnl_adjusted_conv_ema = *reinterpret_cast<const float*>(st_ro + 44);
const unsigned int disagree_n = *reinterpret_cast<const unsigned int*>(st_ro + 56);
const float degradation_factor = 2.0f;
const float disagreement_factor = 32.0f;
const float pnl_decay_factor = 1.0f;
const float degradation_term = (conviction_at_entry > 1e-6f)
? (1.0f - conv_ema_now / conviction_at_entry) / degradation_factor
: 0.0f;
const float disagreement_term = (float)disagree_n / disagreement_factor;
const float pnl_decay_term = fmaxf(-pnl_adjusted_conv_ema, 0.0f) / pnl_decay_factor;
float exit_signal = degradation_term;
if (disagreement_term > exit_signal) exit_signal = disagreement_term;
if (pnl_decay_term > exit_signal) exit_signal = pnl_decay_term;
if (exit_signal >= 1.0f) {
market_targets[b * 2 + 0] = 3;
market_targets[b * 2 + 1] = 0;
return 1;
}
return 0;
}
// CRT.1 C1.4: update intra-trade trajectory fields in open_trade_state.
// Called by the decision kernel each event AFTER the multi-horizon
// conviction is computed and EMA-smoothed. Writes the four offsets the
// composite_exit_check reads on the next stop_check pass:
// offset 24: conviction_ema (mirror of conviction_ema_per_b[b])
// offset 44: pnl_adjusted_conviction_ema (signed by sign of unrealized)
// offset 48: peak_unrealized_pnl (max tracker)
// offset 56: disagreement_consecutive_events (short vs long horizon)
//
// Skipped when no position is open — the composite check is a guard
// against losing-trade edge cases and trajectory state is meaningless
// pre-entry. Single-thread-per-backtest invariant: caller must already
// have gated on threadIdx.x == 0.
__device__ static void update_open_trade_trajectory(
int b,
const Pos* __restrict__ positions,
const Book* __restrict__ books,
const float* __restrict__ alpha_probs,
float conv_ema_now,
unsigned char* __restrict__ open_trade_state
) {
const Pos& pos = positions[b];
if (pos.position_lots == 0) return;
unsigned char* st = open_trade_state + (size_t)b * 64;
// Mirror the current smoothed conviction into open_trade_state so
// composite_exit_check reads it without needing the conviction_ema_per_b
// pointer plumbed in.
*reinterpret_cast<float*>(st + 24) = conv_ema_now;
// Unrealized P&L = (close_px vwap_entry) × dir × |lots|.
// Long: close at best bid; short: close at best ask. Matches the
// close_px used by stop_check_isv's SL/trail computation.
const Book& bk = books[b];
const bool is_long = pos.position_lots > 0;
const float close_px = is_long ? bk.bid_px[0] : bk.ask_px[0];
const float lots_abs = is_long ? (float)pos.position_lots
: -(float)pos.position_lots;
const float dir_long = is_long ? 1.0f : -1.0f;
const float unrealized = (close_px - pos.vwap_entry) * dir_long * lots_abs;
// pnl-adjusted conviction: sign of unrealized × conv_ema_now. EMA on
// top with the same fixed-α=0.1 used by the conviction second-order
// EMAs in update_conviction_ema. First-observation bootstrap per
// pearl_first_observation_bootstrap (prev == 0 ⇒ replace directly).
const float pnl_sign = (unrealized >= 0.0f) ? 1.0f : -1.0f;
const float prev_pnl_adj = *reinterpret_cast<const float*>(st + 44);
const float new_pnl_adj_inst = pnl_sign * conv_ema_now;
const float new_pnl_adj = (prev_pnl_adj == 0.0f)
? new_pnl_adj_inst
: (0.9f * prev_pnl_adj + 0.1f * new_pnl_adj_inst);
*reinterpret_cast<float*>(st + 44) = new_pnl_adj;
// peak_unrealized_pnl tracker.
const float prev_peak = *reinterpret_cast<const float*>(st + 48);
if (unrealized > prev_peak) {
*reinterpret_cast<float*>(st + 48) = unrealized;
}
// disagreement_consecutive_events — short-horizon (h=0) vs long-horizon
// (h=N_HORIZONS-1) directional sign disagreement. Reset to 0 when they
// align, increment when they oppose. Used by composite_exit_check's
// disagreement_term.
const float p_short = alpha_probs[0];
const float p_long = alpha_probs[N_HORIZONS - 1];
const bool dir_short_up = (p_short > 0.5f);
const bool dir_long_up = (p_long > 0.5f);
const unsigned int prev_disagree = *reinterpret_cast<const unsigned int*>(st + 56);
if (dir_short_up != dir_long_up) {
*reinterpret_cast<unsigned int*>(st + 56) = prev_disagree + 1u;
} else {
*reinterpret_cast<unsigned int*>(st + 56) = 0u;
}
}
__device__ static float update_conviction_ema(
int b,
float conviction_abs,
float* __restrict__ conviction_ema,
float* __restrict__ conviction_diff_var_ema,
float* __restrict__ conviction_sample_var_ema
) {
// Defensive clamp to [0, 1] — conviction is bounded by construction
// (sum of magnitude_h × weight_h × direction_h divided by total |weight_h|,
// with magnitude_h ∈ [0, 1]), but bound it before feeding into a
// variance EMA in case of NaN propagation through near-zero weights.
if (!(conviction_abs > 0.0f)) conviction_abs = 0.0f; // also catches NaN
if (conviction_abs > 1.0f) conviction_abs = 1.0f;
// Update second-order EMAs (diff_var, sample_var) with fixed α=0.1.
// These track the variance of the signal and the variance of changes,
// and are themselves bootstrapped with the first observation.
const float prev_ema = conviction_ema[b];
const float diff = conviction_abs - prev_ema;
const float diff_var_prev = conviction_diff_var_ema[b];
const float new_diff_var = (diff_var_prev == 0.0f)
? (diff * diff) // first-observation bootstrap
: (0.9f * diff_var_prev + 0.1f * diff * diff);
conviction_diff_var_ema[b] = new_diff_var;
const float sample_var_prev = conviction_sample_var_ema[b];
const float new_sample_var = (sample_var_prev == 0.0f)
? (conviction_abs * conviction_abs) // first-observation bootstrap
: (0.9f * sample_var_prev + 0.1f * conviction_abs * conviction_abs);
conviction_sample_var_ema[b] = new_sample_var;
// Wiener α with floor at 0.4 (pearl_wiener_alpha_floor_for_nonstationary).
const float eps_v = 1e-6f;
const float alpha_raw = new_diff_var / (new_diff_var + new_sample_var + eps_v);
const float alpha_active = (alpha_raw > 0.4f) ? alpha_raw : 0.4f;
// First-observation bootstrap: prev_ema==0 → replace directly. Otherwise
// blend with Wiener-α.
const float new_ema = (prev_ema == 0.0f)
? conviction_abs
: (alpha_active * conviction_abs + (1.0f - alpha_active) * prev_ema);
conviction_ema[b] = new_ema;
return new_ema;
}
// CRT.1 C1.2: Compute multi-horizon ISV-weighted conviction (spec §4.4).
// Replaces the scalar `max(|p-0.5|)` approach used in v2 A2 (1d889d2de) +
// A2.1 (fe2498769) which failed Gate 1 catastrophically. Direction comes
// from the per-horizon SIGNED sum, so disagreeing horizons cancel — the
// structural fix the scalar EMA could not provide.
//
// Per-horizon weight = max(net_edge_h, eps_edge) / (var_h + cost²) with
// `eps_edge = cost · 0.01` (spec §4.4: ε = cost / 100). The cost² term
// in the denominator bootstraps variance pre-trade per
// pearl_trade_level_vol_for_stop_distance; net_edge is one-sided positive
// per pearl_audit_unboundedness_for_implicit_asymmetry so losing-edge
// horizons drop out instead of flipping the sign. Normalizing by
// `total_abs_weight` makes the formula scale-invariant per
// pearl_zscore_normalization_for_magnitude_asymmetric_signals.
//
// Returns conviction_signed ∈ [-1, +1]. Caller derives:
// conviction_abs = fabsf(conviction_signed)
// direction = sign(conviction_signed)
__device__ static float compute_multi_horizon_conviction(
const float* __restrict__ alpha_probs,
const IsvKellyState* __restrict__ isv,
float cost
) {
const float eps_edge = cost * 0.01f; // spec §4.4 ε floor on net_edge
float weighted_sum_signed = 0.0f;
float total_abs_weight = 0.0f;
#pragma unroll
for (int h = 0; h < N_HORIZONS; ++h) {
const float p_h = alpha_probs[h];
const float direction_h = (p_h > 0.5f) ? 1.0f : -1.0f;
const float magnitude_h = fabsf(p_h - 0.5f) * 2.0f;
const float net_edge_h = fmaxf(isv[h].pnl_ema_win - isv[h].pnl_ema_loss, eps_edge);
const float weight_h = net_edge_h / (isv[h].realised_return_var + cost * cost);
weighted_sum_signed += magnitude_h * weight_h * direction_h;
total_abs_weight += fabsf(weight_h);
}
return (total_abs_weight > 1e-9f)
? (weighted_sum_signed / total_abs_weight)
: 0.0f;
}
extern "C" __global__ void decision_policy_default(
const float* __restrict__ alpha_probs, // [N_HORIZONS] broadcast
const unsigned char* __restrict__ isv_kelly_base, // [n_backtests * N_HORIZONS * sizeof(IsvKellyState)]
const Pos* __restrict__ positions, // [n_backtests]
const unsigned int* __restrict__ program_lens, // [n_backtests] — non-zero ⇒ skip (program kernel handles it)
int* __restrict__ market_targets, // [n_backtests * 2] — written
unsigned int* __restrict__ open_horizon_masks, // [n_backtests] — written (entry attribution)
// Per-backtest envelope cap (CRT.1 C1.2: target = direction × conv_ema × max_lots).
const int* __restrict__ max_lots_per_b, // [n_backtests]
// Threshold gate on the smoothed multi-horizon conviction (spec §5.2).
const float* __restrict__ threshold_per_b, // [n_backtests]
// CRT.1 C1.2: cost feeds the §4.4 conviction weight denominator
// (cost² bootstrap per pearl_trade_level_vol_for_stop_distance) AND
// stop_check_isv's SL/trail-distance floor (round-trip cost).
const float* __restrict__ cost_per_lot_per_side_per_b, // [n_backtests]
// Stop-controller state (ISV-driven stop controller spec §3 + §5).
const float* __restrict__ atr_mid_ema_per_b, // [n_backtests]
float* __restrict__ trail_hwm_per_b, // [n_backtests] — read + write
const Book* __restrict__ books_per_b, // [n_backtests * BOOK_LEVELS] (via lob_state.cuh)
// S2.1: mh_kernel_calls diagnostic counter (non-zero-position entries).
unsigned int* __restrict__ mh_kernel_calls_per_b, // [n_backtests]
// Wiener-α conviction-EMA state (per-backtest). The three slots stay
// from v2 A2 — the EMA mechanism is reused on the multi-horizon scalar.
float* __restrict__ conviction_ema_per_b, // [n_backtests] — read + write
float* __restrict__ conviction_diff_var_ema_per_b, // [n_backtests] — read + write
float* __restrict__ conviction_sample_var_ema_per_b, // [n_backtests] — read + write
// CRT.1 C1.4: trajectory state + composite exit_signal circuit-breaker.
// pnl_track populated conviction_at_entry on the open transition;
// this kernel updates the intra-trade fields each event so the composite
// exit check has fresh inputs on the next pass.
unsigned char* __restrict__ open_trade_state_per_b, // [n_backtests * 64]
// CRT.diag Group A: per-horizon signal persistence counters.
// Observe-only; zero behavior impact. Single-writer-per-block
// (caller already gated on threadIdx.x == 0). End-of-run dump only.
unsigned int* __restrict__ diag_flip_count, // [n_backtests * N_HORIZONS]
unsigned int* __restrict__ diag_current_run_length, // [n_backtests * N_HORIZONS]
unsigned long long* __restrict__ diag_sum_run_length, // [n_backtests * N_HORIZONS]
unsigned int* __restrict__ diag_run_length_hist, // [n_backtests * N_HORIZONS * 5]
signed char* __restrict__ diag_prev_dir_signed, // [n_backtests * N_HORIZONS]
// CRT.diag Group B: conviction-EMA histogram.
unsigned int* __restrict__ diag_conv_hist, // [n_backtests * 10]
int n_backtests
) {
int b = blockIdx.x;
if (b >= n_backtests || threadIdx.x != 0) return;
if (program_lens[b] != 0u) return; // backtest b uses a custom bytecode program; skip default
// Stop-check (spec §3) — pre-VM. Returns 1 if force-flat was written.
if (stop_check_isv(b, positions, isv_kelly_base, open_horizon_masks,
atr_mid_ema_per_b, trail_hwm_per_b, books_per_b,
cost_per_lot_per_side_per_b,
market_targets,
mh_kernel_calls_per_b)) {
return;
}
// CRT.1 C1.2: multi-horizon ISV-weighted conviction (spec §4.4).
// Replaces v2 A2 (1d889d2de) + A2.1 (fe2498769) scalar approach that
// failed Gate 1 catastrophically on smokes vjmwc + lkrdf (155k trades,
// 9100% drawdown, Sharpe -15.7). Direction now comes from the
// per-horizon SIGNED sum so disagreeing horizons cancel — the
// structural fix the scalar EMA could not provide.
const IsvKellyState* isv = reinterpret_cast<const IsvKellyState*>(
isv_kelly_base + (size_t)b * N_HORIZONS * sizeof(IsvKellyState)
);
const float cost = cost_per_lot_per_side_per_b[b];
const float conviction_signed = compute_multi_horizon_conviction(
alpha_probs, isv, cost);
const float conviction_abs = fabsf(conviction_signed);
const float conviction_dir = (conviction_signed >= 0.0f) ? 1.0f : -1.0f;
// Wiener-α adaptive EMA on the multi-horizon conviction magnitude.
// Runs ahead of the threshold gate so the EMA tracks all events
// (gating off an unsmoothed tick would let the EMA grow stale on long
// gated stretches).
const float conv_ema = update_conviction_ema(
b, conviction_abs,
conviction_ema_per_b,
conviction_diff_var_ema_per_b,
conviction_sample_var_ema_per_b);
// CRT.diag Group A: per-horizon direction-flip + run-length tracking.
// Observe-only. Single-writer-per-block (already gated on thread 0).
#pragma unroll
for (int h = 0; h < N_HORIZONS; ++h) {
const float p = alpha_probs[h];
const signed char new_dir = (p > 0.5f) ? (signed char)1
: ((p < 0.5f) ? (signed char)-1 : (signed char)0);
const signed char prev_dir = diag_prev_dir_signed[b * N_HORIZONS + h];
const unsigned int prev_run = diag_current_run_length[b * N_HORIZONS + h];
if (prev_dir != 0 && new_dir != 0 && prev_dir != new_dir) {
// Direction flipped — record the prev run length.
diag_flip_count[b * N_HORIZONS + h] += 1u;
diag_sum_run_length[b * N_HORIZONS + h] += (unsigned long long)prev_run;
int bk;
if (prev_run < 10u) bk = 0;
else if (prev_run < 100u) bk = 1;
else if (prev_run < 1000u) bk = 2;
else if (prev_run < 10000u) bk = 3;
else bk = 4;
diag_run_length_hist[(b * N_HORIZONS + h) * 5 + bk] += 1u;
diag_current_run_length[b * N_HORIZONS + h] = 1u; // start new run
} else {
diag_current_run_length[b * N_HORIZONS + h] = prev_run + 1u;
}
diag_prev_dir_signed[b * N_HORIZONS + h] = new_dir;
}
// CRT.diag Group B: smoothed-conviction histogram (10 buckets over [0, 1)).
{
float c = conv_ema;
if (c < 0.0f) c = 0.0f;
if (c >= 1.0f) c = 0.999999f;
int bk = (int)(c * 10.0f);
if (bk < 0) bk = 0;
if (bk > 9) bk = 9;
diag_conv_hist[b * 10 + bk] += 1u;
}
// CRT.1 C1.4: refresh intra-trade trajectory state (offsets 24, 44, 48,
// 56) for the composite exit_signal. Skipped internally when flat.
update_open_trade_trajectory(b, positions, books_per_b, alpha_probs,
conv_ema, open_trade_state_per_b);
// CRT.1 C1.4: composite exit_signal safety circuit-breaker (spec §4.3).
// Catches the edge case where conviction stays bounded above zero but
// the trade is deep in unrealized loss — the primary `target → 0 from
// collapsed conviction` path (C1.2) cannot reach.
if (composite_exit_check(b, positions, open_trade_state_per_b, market_targets)) {
return;
}
// Threshold gate on the smoothed multi-horizon conviction. Per spec §5.2:
// gating on the smoothed signal rather than the raw signal makes the
// pre-registered percentile thresholds robust to event-rate jitter.
if (conv_ema < threshold_per_b[b]) {
market_targets[b * 2 + 0] = 2;
market_targets[b * 2 + 1] = 0;
open_horizon_masks[b] = 0u;
return;
}
// Build per-horizon attribution mask: every horizon whose signed
// contribution agrees with the multi-horizon direction gets credit.
// This replaces the v2 sharpe-weight attribution because the v3
// weights come from the §4.4 formula directly.
unsigned int attribution_mask = 0u;
#pragma unroll
for (int h = 0; h < N_HORIZONS; ++h) {
const float dir_h = (alpha_probs[h] > 0.5f) ? 1.0f : -1.0f;
if (dir_h == conviction_dir) attribution_mask |= (1u << h);
}
// Target = direction × smoothed conviction × envelope_max. No
// aggregate rescale step — the v2 `final_size *= (conv_ema/raw_max_conv)`
// bug that amplified weak signals is structurally absent because the
// smoothed conviction enters the size formula multiplicatively, never
// ratio'd against the raw scalar. Spec §4.2.
const int max_lots = max_lots_per_b[b];
const float envelope_max = (float)max_lots;
const float target_lots_f = conviction_dir * conv_ema * envelope_max;
const int target_lots = (int)truncf(
target_lots_f + (target_lots_f >= 0.0f ? 0.5f : -0.5f));
if (target_lots == 0) {
// No new lots ordered. Use force-flat encoding if currently open
// so the harness reduces toward zero; use no-op when already flat.
const int code = (positions[b].position_lots == 0) ? 2 : 3;
market_targets[b * 2 + 0] = code;
market_targets[b * 2 + 1] = 0;
return;
}
// If currently flat, this is an opening trade: record attribution mask
// so pnl_track + isv_kelly_update can credit the right horizons on close.
if (positions[b].position_lots == 0) {
open_horizon_masks[b] = attribution_mask;
}
const int side = (target_lots > 0) ? 0 : 1;
const int abs_sz = (target_lots > 0) ? target_lots : -target_lots;
market_targets[b * 2 + 0] = side;
market_targets[b * 2 + 1] = abs_sz;
}
// Bytecode VM — interprets the flat Program emitted by Strategy::flatten()
// (Rust src/policy/mod.rs) against current alpha probs + per-horizon
// IsvKellyState. Replaces the hardcoded default policy when the host
// has uploaded a non-empty program for backtest b.
//
// Stack layout: parallel float-value + u32-attribution-mask stacks.
// Each EmitPerHorizonSize pushes its (signed_size, mask=(1<<h)). The
// aggregators pop n values + n masks, push the aggregated value with
// OR'd mask. The terminal WriteOrder pops one value+mask, converts to
// {side, abs_size} in market_targets[b] and records the attribution
// mask in open_horizon_masks[b] for entry-time crediting.
//
// Layout MUST stay in sync with src/policy/mod.rs::{OpCode, Instruction}.
#define OP_NOOP 0
#define OP_EVAL_REGIME 1
#define OP_EMIT_PER_HORIZON_SIZE 2
#define OP_AGG_WEIGHTED_SHARPE 3
#define OP_AGG_MEAN 4
#define OP_AGG_MAX_CONFIDENCE 5
#define OP_APPLY_CONFLICT 6
#define OP_WRITE_ORDER 7
#define OP_PUSH_SCALAR 8
#define OP_BRANCH_IF_REGIME 9
#define STACK_CAP 32
struct Instruction {
unsigned char op;
unsigned char arg0;
unsigned short arg1;
unsigned int arg2;
};
extern "C" __global__ void decision_policy_program(
const float* __restrict__ alpha_probs,
const unsigned char* __restrict__ isv_kelly_base,
const Pos* __restrict__ positions,
const Instruction* __restrict__ programs, // [n_backtests * max_instructions]
const unsigned int* __restrict__ program_lens, // [n_backtests]
const unsigned int* __restrict__ regimes, // [n_backtests] — current regime id
int* __restrict__ market_targets, // [n_backtests * 2]
unsigned int* __restrict__ open_horizon_masks,
// P1: per-backtest sim parameter arrays. See decision_policy_default header for contract.
const float* __restrict__ target_annual_vol_units_per_b,
const float* __restrict__ annualisation_factor_per_b,
int max_instructions,
const int* __restrict__ max_lots_per_b,
const float* __restrict__ kelly_frac_floor_per_b,
const float* __restrict__ sharpe_weight_floor_per_b,
// P4: threshold gate (same as decision_policy_default).
const float* __restrict__ threshold_per_b,
const float* __restrict__ cost_per_lot_per_side_per_b, // [n_backtests] — round-trip cost floor for SL/trail
// Stop-controller state (ISV-driven stop controller spec §3 + §5).
const float* __restrict__ atr_mid_ema_per_b, // [n_backtests]
float* __restrict__ trail_hwm_per_b, // [n_backtests] — read + write
const Book* __restrict__ books_per_b, // [n_backtests * BOOK_LEVELS]
// S2.1: mh_kernel_calls diagnostic counter (non-zero-position entries).
unsigned int* __restrict__ mh_kernel_calls_per_b, // [n_backtests]
// CRT-A2: Wiener-α conviction-EMA state (per-backtest).
float* __restrict__ conviction_ema_per_b, // [n_backtests] — read + write
float* __restrict__ conviction_diff_var_ema_per_b, // [n_backtests] — read + write
float* __restrict__ conviction_sample_var_ema_per_b, // [n_backtests] — read + write
// CRT.1 C1.4: trajectory state + composite exit_signal circuit-breaker.
// Mirrors decision_policy_default — see that kernel for the contract.
unsigned char* __restrict__ open_trade_state_per_b, // [n_backtests * 64]
// CRT.diag Group A: per-horizon signal persistence counters (observe-only).
unsigned int* __restrict__ diag_flip_count, // [n_backtests * N_HORIZONS]
unsigned int* __restrict__ diag_current_run_length, // [n_backtests * N_HORIZONS]
unsigned long long* __restrict__ diag_sum_run_length, // [n_backtests * N_HORIZONS]
unsigned int* __restrict__ diag_run_length_hist, // [n_backtests * N_HORIZONS * 5]
signed char* __restrict__ diag_prev_dir_signed, // [n_backtests * N_HORIZONS]
// CRT.diag Group B: conviction-EMA histogram.
unsigned int* __restrict__ diag_conv_hist, // [n_backtests * 10]
int n_backtests
) {
int b = blockIdx.x;
if (b >= n_backtests || threadIdx.x != 0) return;
const unsigned int plen = program_lens[b];
if (plen == 0u) {
// Skip — this backtest uses the hardcoded default kernel separately.
return;
}
// Stop-check (spec §3).
if (stop_check_isv(b, positions, isv_kelly_base, open_horizon_masks,
atr_mid_ema_per_b, trail_hwm_per_b, books_per_b,
cost_per_lot_per_side_per_b,
market_targets,
mh_kernel_calls_per_b)) {
return;
}
// CRT.1 C1.2: multi-horizon ISV-weighted conviction (spec §4.4).
// Mirrors decision_policy_default — see that kernel for the full
// rationale. The VM's per-horizon emit + aggregate opcodes still
// execute, but the terminal OP_WRITE_ORDER ignores the stack's
// signed `final_size` and uses the §4.4 conviction-driven target
// instead. The attribution mask from the stack IS still consumed
// at OP_WRITE_ORDER so the VM continues to credit horizons on entry.
const Instruction* prog = programs + (size_t)b * max_instructions;
const IsvKellyState* isv = reinterpret_cast<const IsvKellyState*>(
isv_kelly_base + (size_t)b * N_HORIZONS * sizeof(IsvKellyState)
);
const unsigned int my_regime = regimes[b];
const float cost = cost_per_lot_per_side_per_b[b];
const float conviction_signed = compute_multi_horizon_conviction(
alpha_probs, isv, cost);
const float conviction_abs = fabsf(conviction_signed);
const float conviction_dir = (conviction_signed >= 0.0f) ? 1.0f : -1.0f;
// Wiener-α adaptive EMA on the multi-horizon conviction magnitude.
// Runs ahead of the threshold gate so the EMA tracks all events.
const float conv_ema = update_conviction_ema(
b, conviction_abs,
conviction_ema_per_b,
conviction_diff_var_ema_per_b,
conviction_sample_var_ema_per_b);
// CRT.diag Group A: per-horizon direction-flip + run-length tracking.
// Mirrors decision_policy_default. Observe-only; single-writer-per-block.
#pragma unroll
for (int h = 0; h < N_HORIZONS; ++h) {
const float p = alpha_probs[h];
const signed char new_dir = (p > 0.5f) ? (signed char)1
: ((p < 0.5f) ? (signed char)-1 : (signed char)0);
const signed char prev_dir = diag_prev_dir_signed[b * N_HORIZONS + h];
const unsigned int prev_run = diag_current_run_length[b * N_HORIZONS + h];
if (prev_dir != 0 && new_dir != 0 && prev_dir != new_dir) {
diag_flip_count[b * N_HORIZONS + h] += 1u;
diag_sum_run_length[b * N_HORIZONS + h] += (unsigned long long)prev_run;
int bk;
if (prev_run < 10u) bk = 0;
else if (prev_run < 100u) bk = 1;
else if (prev_run < 1000u) bk = 2;
else if (prev_run < 10000u) bk = 3;
else bk = 4;
diag_run_length_hist[(b * N_HORIZONS + h) * 5 + bk] += 1u;
diag_current_run_length[b * N_HORIZONS + h] = 1u;
} else {
diag_current_run_length[b * N_HORIZONS + h] = prev_run + 1u;
}
diag_prev_dir_signed[b * N_HORIZONS + h] = new_dir;
}
// CRT.diag Group B: smoothed-conviction histogram.
{
float c = conv_ema;
if (c < 0.0f) c = 0.0f;
if (c >= 1.0f) c = 0.999999f;
int bk = (int)(c * 10.0f);
if (bk < 0) bk = 0;
if (bk > 9) bk = 9;
diag_conv_hist[b * 10 + bk] += 1u;
}
// CRT.1 C1.4: refresh intra-trade trajectory state for composite exit.
update_open_trade_trajectory(b, positions, books_per_b, alpha_probs,
conv_ema, open_trade_state_per_b);
// CRT.1 C1.4: composite exit_signal safety circuit-breaker (spec §4.3).
if (composite_exit_check(b, positions, open_trade_state_per_b, market_targets)) {
return;
}
// Threshold gate on the smoothed multi-horizon conviction (spec §5.2).
if (conv_ema < threshold_per_b[b]) {
market_targets[b * 2 + 0] = 2;
market_targets[b * 2 + 1] = 0;
open_horizon_masks[b] = 0u;
return;
}
// Per-backtest scalar reads for VM opcodes (Kelly/Sharpe aggregations).
const float target_annual_vol_units = target_annual_vol_units_per_b[b];
const float annualisation_factor = annualisation_factor_per_b[b];
const int max_lots = max_lots_per_b[b];
const float kelly_frac_floor = kelly_frac_floor_per_b[b];
const float sharpe_weight_floor = sharpe_weight_floor_per_b[b];
float stack_val[STACK_CAP];
unsigned int stack_mask[STACK_CAP];
int sp = 0;
unsigned int pc = 0;
while (pc < plen && pc < (unsigned int)max_instructions) {
const Instruction ins = prog[pc];
pc += 1;
switch (ins.op) {
case OP_NOOP: break;
case OP_PUSH_SCALAR: {
if (sp >= STACK_CAP) break;
const float v = __int_as_float((int)ins.arg2);
stack_val[sp] = v;
stack_mask[sp] = 0u;
sp += 1;
break;
}
case OP_EVAL_REGIME: {
if (sp >= STACK_CAP) break;
stack_val[sp] = (my_regime == (unsigned int)ins.arg0) ? 1.0f : 0.0f;
stack_mask[sp] = 0u;
sp += 1;
break;
}
case OP_BRANCH_IF_REGIME: {
if (my_regime != (unsigned int)ins.arg0) {
pc = ins.arg2;
}
break;
}
case OP_EMIT_PER_HORIZON_SIZE: {
if (sp >= STACK_CAP) break;
const int h = (int)ins.arg0;
const int leaf_max_lots = (int)ins.arg1;
if (h < 0 || h >= N_HORIZONS) { stack_val[sp]=0.0f; stack_mask[sp]=0u; sp+=1; break; }
const float p_h = alpha_probs[h];
const IsvKellyState& s = isv[h];
const float sig_mag = fabsf(p_h - 0.5f) * 2.0f;
const float dir = (p_h > 0.5f) ? 1.0f : -1.0f;
float kelly_frac;
if (s.pnl_ema_win > 1e-9f) {
float kf = (s.win_rate_ema * s.pnl_ema_win
- (1.0f - s.win_rate_ema) * s.pnl_ema_loss)
/ s.pnl_ema_win;
kf = fmaxf(0.0f, fminf(1.0f, kf));
kelly_frac = fmaxf(kelly_frac_floor, kf);
} else {
kelly_frac = kelly_frac_floor;
}
const float effective_cap = fminf((float)leaf_max_lots, (float)max_lots);
// Same MIN_TRADES_FOR_VAR_CAP gate as decision_policy_default.
float cap_lots;
if (s.n_trades_seen >= MIN_TRADES_FOR_VAR_CAP
&& s.realised_return_var > 1e-9f)
{
const float cap_units = target_annual_vol_units
/ sqrtf(s.realised_return_var * annualisation_factor);
cap_lots = fminf(cap_units, effective_cap);
} else {
cap_lots = effective_cap;
}
const float ss = dir * sig_mag * kelly_frac * cap_lots;
stack_val[sp] = ss;
stack_mask[sp] = 1u << h;
sp += 1;
break;
}
case OP_AGG_MEAN: {
const int n = (int)ins.arg0;
if (n <= 0 || sp < n) break;
float sum = 0.0f;
unsigned int mask_or = 0u;
for (int i = sp - n; i < sp; ++i) {
sum += stack_val[i];
mask_or |= stack_mask[i];
}
sp -= n;
stack_val[sp] = sum / (float)n;
stack_mask[sp] = mask_or;
sp += 1;
break;
}
case OP_AGG_WEIGHTED_SHARPE: {
const int n = (int)ins.arg0;
if (n <= 0 || sp < n) break;
float w_sum = 0.0f, wx_sum = 0.0f;
unsigned int mask_or = 0u;
for (int i = sp - n; i < sp; ++i) {
const unsigned int m = stack_mask[i];
// Look up recent_sharpe for the single horizon this slot
// attributes to. If the mask has multiple bits set (a
// nested aggregator already merged horizons) we fall back
// to uniform weight 1.0 because per-horizon attribution
// collapsed.
float w;
if (m != 0u && (m & (m - 1u)) == 0u) {
// Single-bit mask — recover horizon index.
int h = 0;
unsigned int mm = m;
while ((mm & 1u) == 0u && h < N_HORIZONS) { mm >>= 1; h += 1; }
// Floor lets cold-start aggregate; matches the
// decision_policy_default weight pattern.
w = fmaxf(sharpe_weight_floor, isv[h].recent_sharpe);
} else {
w = 1.0f;
}
w_sum += w;
wx_sum += w * stack_val[i];
mask_or |= m;
}
sp -= n;
stack_val[sp] = (w_sum > 1e-9f) ? (wx_sum / w_sum) : 0.0f;
stack_mask[sp] = mask_or;
sp += 1;
break;
}
case OP_AGG_MAX_CONFIDENCE: {
const int n = (int)ins.arg0;
if (n <= 0 || sp < n) break;
int best = sp - n;
for (int i = sp - n + 1; i < sp; ++i) {
if (fabsf(stack_val[i]) > fabsf(stack_val[best])) best = i;
}
const float bv = stack_val[best];
const unsigned int bm = stack_mask[best];
sp -= n;
stack_val[sp] = bv;
stack_mask[sp] = bm;
sp += 1;
break;
}
case OP_APPLY_CONFLICT: {
// v1: no-op (mean of children already aggregated; conflict
// resolution only matters when multiple cells share state,
// which v1 doesn't). Preserved for v2 Portfolio mode.
break;
}
case OP_WRITE_ORDER: {
if (sp <= 0) break;
sp -= 1;
// Pop the VM stack's aggregated value (unused under v3 §4.4
// sizing) and keep the attribution mask for entry crediting.
const unsigned int attribution = stack_mask[sp];
// CRT.1 C1.2: target = direction × smoothed conviction × envelope_max.
// Mirrors decision_policy_default. The v2 `final_size *= scale`
// amplification bug is structurally absent — smoothed conviction
// multiplies the envelope directly, never ratio'd against a raw
// scalar.
const float envelope_max = (float)max_lots;
const float target_lots_f = conviction_dir * conv_ema * envelope_max;
const int target_lots = (int)truncf(
target_lots_f + (target_lots_f >= 0.0f ? 0.5f : -0.5f));
if (target_lots == 0) {
const int code = (positions[b].position_lots == 0) ? 2 : 3;
market_targets[b * 2 + 0] = code;
market_targets[b * 2 + 1] = 0;
} else {
if (positions[b].position_lots == 0) {
open_horizon_masks[b] = attribution;
}
const int side = (target_lots > 0) ? 0 : 1;
const int abs_sz = (target_lots > 0) ? target_lots : -target_lots;
market_targets[b * 2 + 0] = side;
market_targets[b * 2 + 1] = abs_sz;
}
break;
}
default: break;
}
}
}
// Updates per-horizon IsvKellyState[h] for every horizon flagged in
// `closed_horizon_mask[b]`. Called by the host after pnl_track_step
// detects close transitions. `realised_return[b]` is the closed-trade
// per-lot return in price units (positive = win, negative = loss).
//
// Wiener-α floor 0.4 per pearl_wiener_alpha_floor_for_nonstationary.md.
// Welford variance updates a running estimate around mean 0 (simplified;
// v2 may switch to ema-centered).
//
// recent_sharpe composite re-computed from the freshly updated EMAs.
extern "C" __global__ void isv_kelly_update_on_close(
unsigned char* isv_kelly_base,
const unsigned int* closed_horizon_mask, // [n_backtests]
const float* realised_return, // [n_backtests]
int n_backtests
) {
int b = blockIdx.x;
if (b >= n_backtests || threadIdx.x != 0) return;
const unsigned int mask = closed_horizon_mask[b];
if (mask == 0u) return;
const float ret = realised_return[b];
IsvKellyState* isv = reinterpret_cast<IsvKellyState*>(
isv_kelly_base + (size_t)b * N_HORIZONS * sizeof(IsvKellyState)
);
const float alpha = 0.4f; // Wiener-α floor (non-stationary policy P&L)
#pragma unroll
for (int h = 0; h < N_HORIZONS; ++h) {
if (!(mask & (1u << h))) continue;
IsvKellyState& s = isv[h];
const bool win = ret > 0.0f;
if (s.n_trades_seen == 0u) {
// First-observation bootstrap: replace EMAs directly (no zero-bias warmup).
s.pnl_ema_win = win ? ret : 0.0f;
s.pnl_ema_loss = win ? 0.0f : -ret;
s.win_rate_ema = win ? 1.0f : 0.0f;
s.realised_return_var = ret * ret; // single-sample variance proxy
} else {
if (win) {
s.pnl_ema_win = (1.0f - alpha) * s.pnl_ema_win + alpha * ret;
s.win_rate_ema = (1.0f - alpha) * s.win_rate_ema + alpha * 1.0f;
} else {
s.pnl_ema_loss = (1.0f - alpha) * s.pnl_ema_loss + alpha * (-ret);
s.win_rate_ema = (1.0f - alpha) * s.win_rate_ema + alpha * 0.0f;
}
// Welford-ish running variance around 0.
const float prev_n = (float)s.n_trades_seen;
s.realised_return_var =
(prev_n * s.realised_return_var + ret * ret) / (prev_n + 1.0f);
}
s.n_trades_seen += 1u;
// Composite recent_sharpe = (mean_winning_contribution mean_losing_contribution)
// / sqrt(realised_return_var)
const float numerator = s.pnl_ema_win * s.win_rate_ema
- s.pnl_ema_loss * (1.0f - s.win_rate_ema);
const float denom = sqrtf(fmaxf(s.realised_return_var, 1e-9f));
s.recent_sharpe = numerator / denom;
}
}