Rename binary alpha_compose_backtest → alpha_baseline and remove the
boolean flags whose features are now mandatory:
--c51 (always C51 distributional Q)
--temporal (always Mamba2 temporal encoder)
--isv-continual (controller always fires per eval episode)
--regime-scale (vol-EMA regime defense always on)
--pruned-actions (FALSIFIED 2026-05-15 per pearl_action_pruning_falsified)
Every dependent code path was stripped, not just gated:
- Linear-Q kernels (lq_fwd, lq_grad, munch_kernel) and their cubin loads
are gone — C51 is the only Q-network.
- Single-env push_kernel / h_store_kernel loads removed; the backtest
has been batched-parallel-env since T14 and only the _batched
variants are called here. (The smoke binary still uses single-env
variants because one env per episode is its job.)
- Dead transition buffers removed: states_dev, next_states_dev,
actions_dev, rewards_dev, dones_dev, q_current_dev, q_next_dev,
target_dev, single_state_dev, single_q_dev, probs_current_dev,
probs_next_dev, m_dev, single_probs_dev, single-env state_pinned,
action_pinned, window_tensor, h_enriched_buf_dev.
- Dead constants and helpers: PRUNED_ACTIONS, N_WEIGHTS, N_BIASES,
epsilon_greedy, epsilon_greedy_gated.
End-to-end verification on the existing Q1 fxcache (rebuild was OOM
locally; full multi-quarter validation is the next phase):
cost=0.0000 best τ=0.250 Sharpe_ann=+36.83 win=0.984 trades/ep=83.3
cost=0.0625 best τ=0.250 Sharpe_ann=+38.53 win=0.996 trades/ep=83.2
cost=0.1250 best τ=0.250 Sharpe_ann=+38.37 win=0.994 trades/ep=83.3
cost=0.2500 best τ=0.250 Sharpe_ann=+34.24 win=0.990 trades/ep=85.6
cost=0.5000 best τ=0.250 Sharpe_ann=+31.83 win=0.946 trades/ep=84.8
Numbers track the prior T16-flag config (within stochastic noise),
confirming the conditional-stripping was a pure simplification — no
behavioral change, just a smaller, honester binary.
Also updated:
- scripts/alpha_pipeline.sh — A/B conditions collapse to fixed-cost
vs cost-randomized training (the only opt-in left).
- scripts/walk_forward_cv.sh — drop legacy flags, pass --window-k only.
- crates/ml/src/env/loaders.rs — module doc-comment updated.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Adds a sliding-window walk-forward harness for the T10 backtest:
- New load_snapshots_from_fxcache_at(start_offset, ...) loader variant
reads bars [start_offset..start_offset+max_snapshots) from the fxcache.
Alpha-cache lookups use absolute bar indices, so the same
alpha_logits_cache.bin works across folds.
- New --data-start-offset CLI flag on alpha_compose_backtest.
- scripts/walk_forward_cv.sh runs 3 folds (window=700K, train_frac=0.6)
at offsets 0 / 600K / 1.2M, producing /tmp/cv_fold_{A,B,C}.json plus
an aggregated mean±stddev Sharpe table across folds.
Walk-forward result (alpha_logits_cache trained on bars 0..1.57M, so
fold C eval is fully past the stacker cut):
cost fold-A fold-B fold-C mean ± stddev
0.0000 +91.52 -21.44 +46.74 +38.94 ± 56.88
0.0625 +84.94 -27.97 +38.42 +31.79 ± 56.74
0.1250 +79.91 -31.22 +33.51 +27.40 ± 55.82
0.2500 +72.77 -45.41 +15.16 +14.17 ± 59.09
0.5000 +50.52 -59.82 -12.75 -7.35 ± 55.37
Fold B (mid-quarter, bars 600K..1.3M) is a disaster — win rate
collapses to 0-22% across all costs. Folds A and C succeed strongly.
Cross-fold SD ≈ mean, so the policy is regime-dependent and cannot
be reliably deployed without regime detection.
Mean Sharpe at half-tick (+27.40) is still ~7× the stateless
Phase 1d.4 baseline (-4.0), so the temporal encoder adds real value
on average — but the single-window +62 OOS celebrated earlier was
a cherry-picked favorable regime, not a deployment-ready result.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>