Files
foxhunt/docs/superpowers/specs/2026-05-15-phase-e-temporal-encoder-design.md
jgrusewski eb9047fc30 docs(phase-e): E.4 temporal encoder design + E.4.A implementation plan
Design doc (specs/): TFT-style architecture for Phase E execution
policy — sliding window → Mamba2 SSM → GRN trunk → MoE regime gate
→ C51 head → Thompson selector. Two core pillars added per user:
  A) Full L1-L10 LOB depth input via hybrid MBP-10 peek
  B) ISV-continual-learning: controllers fire at training AND
     inference; Q-net weights frozen at inference but effective
     policy adapts via ISV modulation

Plan doc (plans/): 14-task implementation plan for E.4.A foundation
(window buffer + L1-L10 depth + Mamba2 forward+backward + ISV-eval
controllers). Falsification gates: smoke R_mean improvement ≥ 50%,
backtest cost=0 Sharpe ≥ +8 (no regression vs C51-flat +10.41),
half-tick Sharpe ≥ -8 (closes 5pt+ of 10pt gap to Phase 1d.4
baseline -4.0).

TGGN (foxhunt Temporal Graph Gated Network) explicitly deferred to
Phase E.5+: existing CPU graph implementation + GPU adapter at
ml-supervised/src/tgnn/ — marginal benefit for single-instrument ES
futures vs the TFT-Mamba2 stack; revisit for multi-instrument
extension or production HFT inference layer.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-15 20:42:35 +02:00

23 KiB
Raw Blame History

Phase E Temporal Encoder + Reasoning Architecture — Design

Status: Design doc (not yet implementation plan). Awaiting user approval. Date: 2026-05-15 Author: session 9c4e48ad-0f44-4989-8dab-09ec6870540c Predecessors: pearl_c51_thompson_closed_phase_e3_gap.md, pearl_action_pruning_falsified.md, project_phase_e3_close.md Production reference stack: mamba2_temporal_kernel.cu, tlob_kernel.cu, attention_kernel.cu, grn_kernel.cu, vsn_feature_selection_kernel.cu, moe_kernels.cu, aux_heads_kernel.cu, aux_trunk_forward_kernel.cu, pearl_1_atom_kernel.cu

1. Goal

Lift the Phase E execution policy from stateless linear C51 Q to a stateful, regime-aware, auxiliary-supervised Q-network. Close the remaining 10pt Sharpe gap to the Phase 1d.4 baseline at half-tick by adopting the production trainer's temporal-reasoning stack — same components that lifted Phase 1d.2's snapshot-stream AUC from 0.50 → 0.66 at K=6000.

2. Why now

The Phase E.3 close + C51 follow-up establish:

  • Alpha signal is real and gets transmitted (rvr = +1.045σ across linear and C51).
  • Calibration was the dominant bottleneck (C51 closed +26pt of Sharpe gap at cost=0 vs linear Q).
  • Action variance was NOT the bottleneck (pruning falsified, -22pt Sharpe).
  • Fill economics was NOT the bottleneck (real spread ≈ fixed for ES — 76% of bars at 1-tick floor).
  • ⚠️ The remaining 10pt half-tick gap is trade-count economics × value-estimation precision under current stateless architecture.

The production trainer already implements the techniques needed:

  • Temporal memory (Mamba2, TLOB, attention) — pearl_state_amplifies_short_horizon_into_long_horizon proved this works on ES futures.
  • Regime-aware decisions (MoE) — pearl_snapshot_alpha_is_regime_conditional showed spread-Q4 acc=0.747 vs middle quintiles below chance; the alpha lives in specific regimes.
  • Dense auxiliary supervision (aux heads + separate aux trunk) — pearl_separate_aux_trunk_when_shared_starves is the canonical fix for sparse-PnL training.
  • Adaptive atom support (pearl_per_branch_c51_atom_span) — eliminates the hardcoded [-10, +10] choice.

Smart-borrow philosophy: adopt the techniques (kernels, controller signals, ISV slots) without importing the specialization (4-branch action factorization, magnitude bins, direction-specific reward biases). Same approach that worked for the C51 borrow.

2.5. Core architectural pillars (added by user 2026-05-15)

Pillar A: Full L1-L10 LOB depth input

The DBN MBP-10 files contain real L1-L10 bid/ask data (verified dbn_parser.rs:721-728 and 866-873 correctly copy all 10 levels). The current Phase E env synthesizes L2/L3 at ±tick offsets because the fxcache doesn't carry depth — only L1-derived features (spread_bps, l1_imbalance) plus the 81-dim Block-S feature vector.

To use real L4-L10, options:

  1. Hybrid loader: at env construction, peek MBP-10 by timestamp for each fxcache bar — adds depth without breaking alpha_cache alignment.
  2. Fxcache rebuild: extend the fxcache schema to store full depth.
  3. MBP-10 direct mode: skip fxcache entirely; loses alpha_cache.

Pick (1) for the staged rollout. Adds ~120 floats per snapshot (L1-L10 × bid+ask × {px, sz, ct} = 60 fields × 2 sides = up to 120). Phase E.4.A.2 task.

Pillar B: ISV-continual-learning (train AND inference)

Foxhunt's controller pattern already produces and consumes ISV slots GPU-resident via pearl_engagement_rate_self_correction and friends. Phase E currently freezes ISV at training end. The proposal: keep controllers firing during inference / eval / deployment.

What stays frozen at inference: Q-net weights (W, b for the C51 head; Mamba2 SSM parameters; GRN/MoE/VSN weights).

What adapts at inference via ISV controllers:

  • Slot 543 — stacker threshold (engagement-rate controller continues responding to observed trade rate)
  • Slot 545 — observed-rate Wiener-α EMA (always updating)
  • Slot 546 — Kelly attenuation (tightens after drawdowns)
  • Slot atom_headroom (Pearl-1) — adaptive C51 support widens/narrows as observed Q-scale shifts
  • Slot ~126 — MoE gate entropy EMA (drives MoE λ controller)
  • ISV[12] — health composition (if we adopt it) gates ALL downstream adaptations

Effective dynamic weight: the policy's decision rule action = ThompsonSelect(MoE_mix(experts), threshold_gate) is a function of the ISV slots. As ISV evolves at inference, the effective policy adapts WITHOUT touching neural-network parameters. This is "dynamic weight via ISV modulation" — lightweight, safe, composable with frozen Q-net.

Optional aggressive variant — LoRA at inference: add a low-rank adapter W' = W + α·U·V^T where U ∈ R^{H×r}, V ∈ R^{r×H}, r << H. Update U, V via online SGD on observed reward error during inference. Adds literal dynamic weights (not just ISV). Out of scope for Phase E.4 core — possible Phase E.5 extension if ISV-only adaptation isn't enough.

Why this composes: Mamba2 hidden state accumulates within-episode temporal context (fast adaptation). ISV controllers accumulate across-episode regime context (slow adaptation). Together they form a two-timescale online adaptation system on top of a frozen representational core.

Failure modes to guard against:

  • ISV runaway in low-data regimes — bound all controller updates by Wiener-α floor per pearl_wiener_alpha_floor_for_nonstationary.
  • Threshold spiraling to no-trade — engagement-rate controller has a target floor; verify it activates at inference.
  • Adversarial market manipulation against the live ISV — for HFT deployment, controllers should have rate limits on update magnitude.

3. Architecture (target)

┌──────────────────────────────────────────────────────────────────────┐
│ Per-step input: state[10]  +  rolling window buffer[K, 10]           │
│ (current state pushed into a circular buffer of last K snapshots)    │
└────────────────────────────────┬─────────────────────────────────────┘
                                 │
                                 ▼
┌──────────────────────────────────────────────────────────────────────┐
│ Variable Selection Network (VSN)                                     │
│   • Per-state softmax gate over feature groups → vsn_mask[10]        │
│   • Element-wise gated features → x_in[K, 10]                        │
│   • mask EMA → ISV[~120..125]  (interpretability)                    │
└────────────────────────────────┬─────────────────────────────────────┘
                                 │
                                 ▼
┌──────────────────────────────────────────────────────────────────────┐
│ Mamba2 SSM Temporal Encoder                                          │
│   • State-space accumulation over the K-bar window                   │
│   • Output: h_temporal ∈ R^H  (compressed temporal context)          │
│   • SAME kernel as Phase 1d.2 — proven on ES futures                 │
└────────────────────────────────┬─────────────────────────────────────┘
                                 │
                                 ▼
┌──────────────────────────────────────────────────────────────────────┐
│ Gated Residual Network (GRN) Trunk                                   │
│   • h_temporal → h_s2 ∈ R^H'  (shared trunk encoding)                │
│   • Gating + residual; more expressive than linear, cheaper than     │
│     full transformer                                                 │
└──────────────────────────────┬─┴─────────────────────────────────────┐
                               │                                       │
              ┌────────────────┘                                       │
              │  (main path)                                           │
              ▼                                                        ▼
┌──────────────────────────────────┐    ┌──────────────────────────────┐
│ MoE Regime Gate                  │    │ Aux Trunk (SEPARATE)          │
│   • K_e=4 experts, gate by h_s2  │    │   • Linear→ELU→Linear→ELU→    │
│   • Each expert is a C51 head    │    │     Linear MLP                │
│     [9 actions × N_atoms atoms]  │    │   • stop_grad at encoder      │
│   • Output: gate-mixed probs     │    │     boundary                  │
│     p(a, k | s)                  │    │   • Independent Adam          │
│   • gate_entropy_ema → ISV[126]  │    │   • Aux heads:                │
│     drives MoE λ controller      │    │     1) next-bar return MSE    │
│     (existing in production)     │    │     2) 5-class regime CE      │
└────────────────────────┬─────────┘    └──────────────────────────────┘
                         │
                         ▼
┌──────────────────────────────────────────────────────────────────────┐
│ C51 categorical output: probs[B, 9, N_atoms]                         │
│   • Adaptive atom support [v_min, v_max] from Pearl-1 controller     │
│   • ISV-driven instead of hardcoded                                  │
└────────────────────────────────┬─────────────────────────────────────┘
                                 │
                                 ▼
┌──────────────────────────────────────────────────────────────────────┐
│ GPU Thompson selector (existing) + alpha_confidence threshold gate    │
│   • Inverse-CDF over each expert's mixed probs                        │
│   • Mapped-pinned action output (current implementation, unchanged)   │
└──────────────────────────────────────────────────────────────────────┘

Components (justification)

Component Production kernel Phase E role Compat note
Sliding-window state buffer new (CPU or GPU circular buffer) Provides K bars of context per step New buffer mgmt; reset on episode boundary
VSN vsn_feature_selection_kernel.cu Per-state feature-importance gating Use 10 features instead of 6 groups; mask EMA into ISV
Mamba2 SSM mamba2_temporal_kernel.cu Temporal encoder Same kernel; D=10 input, configure hidden width
GRN grn_kernel.cu Gated trunk encoding Direct reuse
MoE gate moe_kernels.cu + moe_lambda_eff_kernel.cu Regime-aware Q-head dispatch K_e=4 (matches spread-quintile regime structure)
C51 head alpha_c51.cu (existing) Per-expert categorical output Already works
Aux trunk aux_trunk_forward/backward_kernel.cu Separate trunk for aux supervision Stop-grad at encoder per pearl_separate_aux_trunk_when_shared_starves
Aux heads aux_heads_kernel.cu + aux_heads_loss_ema_kernel.cu Next-bar MSE + regime CE Adapt label-builders to Phase E targets
Pearl-1 atom span pearl_1_atom_kernel.cu Adaptive [v_min, v_max] for C51 Replaces hardcoded support; ISV-driven
Thompson selector alpha_c51.cu (existing) Action selection Reads gate-mixed probs; no kernel change

TFT correspondence

TFT component Provided by
Variable Selection Network vsn_feature_selection_kernel.cu (direct)
Gated Residual Network grn_kernel.cu (direct)
Static covariate encoder absent — Phase E has no static covariates; OK
LSTM seq-to-seq encoder substituted by Mamba2 SSM (better; subquadratic)
Interpretable multi-head attention optional Phase E.4.E follow-up via attention_kernel.cu
Quantile output substituted by C51 categorical (already working)

The proposed architecture is structurally a TFT, with two strict upgrades vs the original Lim et al. 2021 design:

  1. Mamba2 instead of LSTM (linear-time state accumulation, no vanishing-gradient).
  2. C51 categorical output instead of quantile loss (proven calibration win).

4. State representation changes

Current Phase E state vector (10-dim, single timestep):

[alpha_logit, alpha_confidence, spread_bps, l1_imbalance, ofi_sum_5,
 mid_drift_5, position, step_normalized, log_tau, log_event_rate]

Proposed: same 10 features but K-bar history for all features. Window length K starts at 16 (~16 bars of context); validate empirically against K=32, K=64. Mamba2 handles long-horizon well (Phase 1d.2 used K=6000); the cost is GPU memory not architectural limit.

Per-step input shape: [B, K, 10]. The buffer is maintained as a circular array; reset on episode start zeros it (or seeds with the first K observed states once available).

5. Training changes

  • Forward: per step, push current state to buffer, run window → VSN → Mamba2 → GRN → (MoE → C51 head) + (aux trunk → aux heads).
  • Backward: Q-loss (C51 CE) flows through MoE→GRN→Mamba2→VSN. Aux-loss flows through aux trunk only (stop-grad at encoder boundary, per pearl_separate_aux_trunk_when_shared_starves).
  • Loss balance: per pearl_loss_balance_controller, signal-modulated target × Wiener-α. Phase E currently uses Q-loss only; aux-loss weight starts at 1.0 and is controller-driven.
  • Optimizer: keep plain SGD initially (C51 works with SGD). Switch to Adam in Phase E.4.D if signal plateaus.
  • PER: NOT in scope for E.4 (current on-policy training works); revisit only if learning is sample-starved.

6. Phasing

Each phase produces a runnable smoke + backtest. Each phase ships as a discrete experiment we can A/B against the prior best.

Phase E.4.A — Foundation (1 week)

  • Phase E.4.A.1: Sliding-window state buffer (GPU-resident, circular)
  • Phase E.4.A.2: Hybrid MBP-10 depth loader — extend load_snapshots_from_fxcache to peek MBP-10 by timestamp and populate real L4-L10 bid/ask in SnapshotRow.bid_l[3..10] / ask_l[3..10]. Preserves alpha_cache alignment.
  • Phase E.4.A.3: Extend state vector to include depth features (L1-L10 cumulative size, spread-curve slope, etc.)
  • Phase E.4.A.4: Mamba2 forward + backward (already exists; wire into Phase E)
  • Phase E.4.A.5: GRN trunk (already exists; wire)
  • Phase E.4.A.6: C51 head reads h_s2 instead of state directly
  • Phase E.4.A.7: ISV-continual-learning toggle — controllers fire at eval time too (already in smoke at training; lift to backtest eval).
  • Smoke + backtest vs C51-flat-baseline (target: maintain +10.4 at cost=0, lift half-tick)
  • Falsification: if Sharpe at cost=0 drops, the temporal architecture isn't lifting — pause and diagnose before adding more.

Phase E.4.B — Regime gating (1 week)

  • Add VSN feature selection
  • Add MoE 4-expert gating + λ controller (existing kernel)
  • Validate expert utilization (should specialize across spread quintiles)
  • A/B vs E.4.A

Phase E.4.C — Aux supervision (3-4 days)

  • Add aux trunk + heads (next-bar return MSE, 5-class regime CE)
  • Validate aux losses converge (pearl_separate_aux_trunk_when_shared_starves gate: CE < 0.1 AND dir_acc > 0.95)
  • Validate Q-learning isn't destabilized
  • A/B vs E.4.B

Phase E.4.D — Adaptive atoms + polish (3-4 days)

  • Pearl-1 atom span controller (replaces hardcoded [-10, +10])
  • Add health composition (ISV[12]) feeding controller / atom span
  • Trade-rate target sweep at high-cost regime (closes the trade-count economics)
  • Final 2D sweep, write close-out memo

Phase E.4.E — Optional follow-ups (defer)

  • TFT interpretable multi-head attention
  • IQN quantile head composed with C51
  • Pearl-4 Adam adaptive hyperparams

7. Falsification criteria

For the architecture upgrade to be worth the engineering cost:

Smoke must show (vs C51-flat-baseline R_mean = -1.1):

  • R_mean improvement ≥ 50% (i.e., R_mean ≥ -0.5)
  • rvr maintained (≥ +1.04σ)
  • Action entropy NOT collapsed to Wait-only (>0.5 × ln(9))

Backtest must show (vs C51-flat-baseline Sharpe_ann at half-tick = -13.8):

  • Half-tick Sharpe_ann ≥ -8 (closes 5pt+ of the remaining 10pt gap to Phase 1d.4 baseline at -4.0)
  • Trade rate ≤ 70/ep at best τ (moving toward Phase 1d.4's 20-50/ep regime)
  • Cost=0 Sharpe_ann ≥ +8 (not regressed below the C51-flat result)

If these don't hit by Phase E.4.B end: rollback to C51-flat and reconsider. The temporal architecture is a STRATEGIC bet — if it doesn't deliver, we have a known-good fallback (the C51-flat policy already shipped in this session).

8. Risks

  • State-buffer memory: K=64 × 10 features × float = 2.5KB per env-instance. For batch backtest with 500 episodes the working set is fine. K=600 (full smoke horizon) is 24KB/instance — also fine.
  • Mamba2 init: Phase 1d.2 used random init + trainable end-to-end (pearl_tlob_no_pretraining). Should work for Phase E too.
  • Aux head label leak: pearl_trend_scanning_purged_cv_doesnt_sterilize_forward_features warns that forward-window labels leak into forward-horizon targets. Mitigation: aux supervision uses NEXT BAR (single-step) target, not multi-bar trend-scanning label.
  • MoE expert collapse: if all 4 experts learn the same C51 distribution, MoE adds parameters without benefit. Mitigation: moe_lambda_eff_kernel.cu λ controller penalizes gate entropy collapse; existing kernel does this.
  • Loss balance instability: with Q-loss + 2 aux losses, the encoder receives 3 gradient streams. pearl_loss_balance_controller is the production answer; adopt at Phase E.4.C.

9. Open questions

  • Window length K: start at 16, sweep 16/32/64. Mamba2 doesn't impose a ceiling.
  • MoE expert count K_e: start at 4 (matches spread-quintile structure from Phase 1c). Validate by checking expert utilization per regime.
  • Aux head regime classes: 5 classes for spread quintiles, OR 5 classes for volatility quintiles? Pick spread (alpha-conditional dimension from Phase 1c).
  • Should we keep the threshold gate? The current --train-threshold is a CPU-side gate. Once the Q-net is temporal + regime-aware, the gate may be redundant. Test with --train-threshold 0.0 in E.4.B.

10. Out of scope / future research

Out of scope (Phase E.4 core):

  • Different markets (validated as not the bottleneck for ES futures spread economics)
  • Action-space changes (pruning falsified; full 9 stays)
  • PER / off-policy replay (current on-policy training has no observable sample-starvation)
  • LoRA dynamic weights at inference (aggressive ISV-continual-learning variant — Phase E.5 if needed)

Future research (Phase E.5+ or production-tier):

  • TGGN graph reasoning over L1-L10 depth (ml-supervised/src/tgnn/): foxhunt has an existing CPU-based Temporal Graph Gated Network for HFT (Performance targets <500ns graph build, <1μs GNN inference). The trainable adapter is GPU-native (crates/ml/src/tgnn/trainable_adapter.rs using cuBLAS GpuLinear + GpuAdamW) but the message-passing layer is CPU. Honest assessment for Phase E.4: marginal benefit over the TFT-Mamba2 stack — the 81-dim Block-S features already encode LOB structure and Mamba2 captures temporal context. TGGN's strongest value is in (a) multi-instrument cross-asset graphs, (b) full MBP-10 depth utilization, (c) production HFT sub-1μs latency. With Pillar A (L4-L10 input wired), the depth case becomes stronger — revisit as Phase E.5 candidate.
  • TFT interpretable multi-head attentionattention_kernel.cu + attention_backward_kernel.cu. Adds variable-importance via attention weights. Already deferred in section 6 as Phase E.4.E.
  • IQN dual head composed with C51 — joint sampling (pearl_thompson_for_distributional_action_selection). Trade-off: +2-3× training overhead for marginal exploration gain.
  • Pearl-4 Adam adaptive hyperparams — replace plain SGD.
  • xLSTM / KAN / Liquid alternative encoders — adapters exist in crates/ml/src/hyperopt/adapters/. Test only if Mamba2 plateaus.
  • Multi-instrument extension (ES + NQ + RTY graph for cross-asset signals) — opens up TGGN's strongest use case.

11. Approval gate

This is a design proposal, not yet an implementation plan. Before invoking superpowers:writing-plans to produce a step-by-step implementation:

  • User approved the TFT-Mamba2 architecture ("the idea is great we use that", 2026-05-15)
  • TGGN clarified — foxhunt-specific Temporal Graph Gated Networks (ml-supervised/src/tgnn/). Existing CPU implementation + GPU adapter. Deferred to Phase E.5+ research per Section 10.
  • L4-L10 depth source verified — DBN MBP-10 files contain real L1-L10. Pillar A (Section 2.5) integrates them via hybrid MBP-10 peek alongside fxcache.
  • ISV-continual-learning concept adopted as Pillar B (Section 2.5) — core architectural feature, not optional.
  • User confirms phasing (4 phases × ~3 weeks total)
  • User confirms scope: stop after E.4.D? Or extend with E.4.E?
  • User approves the falsification criteria (smoke + backtest thresholds)

After approval, the writing-plans skill produces task-by-task TDD implementation steps for Phase E.4.A first.