Files
foxhunt/crates
jgrusewski 04f973b071 refactor(env): shared Kelly cap + stats update across training and validation
Addresses the "why isn't this a shared module?" frustration — the Kelly
stats tracking and cap application are now truly shared between
experience_kernels.cu (training) and backtest_env_kernel.cu (validation)
via trade_physics.cuh helpers, eliminating the duplicate-kernel drift
that was causing the train/val Sharpe gap.

Shared helpers added to trade_physics.cuh:
- apply_kelly_cap(target, stats, max_position, safety)
- record_kelly_trade_outcome(prev_pos, curr_pos, entry_price, close,
                             equity, &win_count, &loss_count,
                             &sum_wins, &sum_losses)

Architecture: Kelly stats live in a SEPARATE buffer (kelly_stats_buf)
in the validation env, stride 4, so the 8-slot portfolio_buf remains
consumable by backtest_state_gather without needing that kernel's
stride-8 indexing to change. Training stores its stats inline in ps[14..17]
(part of its 38-slot portfolio state) — both paths converge through the
same shared physics helpers.

Results on E1 smoke test (20-epoch):
  Training Sharpe_raw   ≈  +0.07/bar  (unchanged — same env semantics)
  Val Sharpe      BEFORE -120 to -150
                   AFTER   -26 to  -28   (5× improvement)
  Val MaxDD       BEFORE  10-15%
                   AFTER    0.6-0.7%     (20× improvement)
  Val Sharpe_raw  BEFORE  -0.39
                   AFTER  -0.09          (4× improvement)
  Final q_gap     0.1475  (mechanism still protecting against collapse)

The catastrophic val losses were primarily from uncapped leverage in
backtest — the agent could max out position even during collapsing-policy
epochs. Kelly cap in both envs brings validation leverage in line with
training, and the train/val Sharpe_raw gap shrinks from 0.4 to 0.2.

Files:
- trade_physics.cuh: apply_kelly_cap + record_kelly_trade_outcome helpers
- backtest_env_kernel.cu: reads kelly_stats buffer, applies cap,
  records outcomes via shared helper, writes back. Both single-step
  and batched variants updated symmetrically.
- gpu_backtest_evaluator.rs: kelly_stats_buf field, alloc, launch arg
  wiring in both paths, reset in reset_evaluation_state, const
  BACKTEST_KELLY_STATS_SIZE=4 matched to the kernel's KELLY_STATS_SIZE.

Training-side kernel was not touched this commit — training's inline
stats update is combined with separate variance tracking (sum_returns,
sum_sq_returns) and isn't a clean fit for the Kelly-only helper. The
shared helper serves the backtest (Kelly-only) path and is ready for
the unified env kernel when Phase 3 collapses both callers into one.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-21 01:16:02 +02:00
..