Phase E.0 Task 7c. Ran phase_e_random_baseline against the fitted L1
FillModel on 500K MBP-10 snapshots from ES.FUT 2024-Q1. Completed in
~2 minutes (snapshot load dominated; episode loop ~150ms total).
Results:
mean reward = -5185.13
std reward = 4952.85
p05 = -13972.31
p25 = -7251.85
p50 (median) = -2804.56
p75 = -1787.90
p95 = -954.75 (best 5% of random episodes still lose)
kill threshold = +4720.57 (= mean + 2σ; E.1 DQN must exceed)
avg fills/ep = 139.22 (~1 fill every 4.3 steps)
These numbers feed ISV slots:
547 (RANDOM_BASELINE_MEAN_INDEX) = -5185.13
548 (RANDOM_BASELINE_STD_INDEX) = 4952.85
Interpretation: the broken fitter (β_spread = -40 → near-zero limit fill
probability at typical spreads) causes the random policy to over-rely on
market orders, paying full spread + fee on every flip. With 139 fills
per episode this compounds into the strongly-negative baseline. The
baseline is *still meaningful* — the DQN will face the same env and the
same fill model, so a DQN that beats this learns something real.
Open follow-up for Phase E.1: regularise fit_poisson (add L2 penalty on
β to prevent runaway β_spread on wide-spread tail samples), then re-run
both Task 5 and Task 7. Until then, the current baseline is the
operational reference point.