Files
foxhunt/config/ml
jgrusewski 4f71ab32ae feat(alpha): random-uniform policy baseline (10K episodes, horizon 600)
Phase E.0 Task 7c. Ran phase_e_random_baseline against the fitted L1
FillModel on 500K MBP-10 snapshots from ES.FUT 2024-Q1. Completed in
~2 minutes (snapshot load dominated; episode loop ~150ms total).

Results:
  mean reward        = -5185.13
  std reward         = 4952.85
  p05                = -13972.31
  p25                =  -7251.85
  p50 (median)       =  -2804.56
  p75                =  -1787.90
  p95                =   -954.75   (best 5% of random episodes still lose)
  kill threshold     =  +4720.57   (= mean + 2σ; E.1 DQN must exceed)
  avg fills/ep       =   139.22    (~1 fill every 4.3 steps)

These numbers feed ISV slots:
  547 (RANDOM_BASELINE_MEAN_INDEX) = -5185.13
  548 (RANDOM_BASELINE_STD_INDEX)  =  4952.85

Interpretation: the broken fitter (β_spread = -40 → near-zero limit fill
probability at typical spreads) causes the random policy to over-rely on
market orders, paying full spread + fee on every flip. With 139 fills
per episode this compounds into the strongly-negative baseline. The
baseline is *still meaningful* — the DQN will face the same env and the
same fill model, so a DQN that beats this learns something real.

Open follow-up for Phase E.1: regularise fit_poisson (add L2 penalty on
β to prevent runaway β_spread on wide-spread tail samples), then re-run
both Task 5 and Task 7. Until then, the current baseline is the
operational reference point.
2026-05-15 13:37:56 +02:00
..