spec(ml-alpha): CfC + PPO greenfield design

Approved design document for the ml-alpha rebuild. Seven sections plus
appendices, all 7 sections + ISV-driven auto-tuning addition approved
via the brainstorming-skill flow.

Key architectural decisions:
- Full closed-loop scope: training + offline inference + live IBKR
- Three components: ml-alpha library, alpha_train binary, AlphaPpoStrategy
  in trading_agent_service (reuses existing data_acquisition,
  trading_service, broker_gateway, and the 1177-LOC IBKR adapter that
  already exists in trading_engine)
- Snapshot-level CfC perception trunk (Closed-form Continuous-time, the
  trainable LTC variant from Hasani 2022) — chosen over Mamba2 for
  native multi-time-scale via per-cell learnable τ; hard validation
  gate requires CfC AUC ≥ Mamba2 baseline before continuing
- 5 multi-horizon heads at {30, 100, 300, 1000, 6000} snapshots forward
- Decision-stride 50 snapshots (~1 min); empirically validated this
  session (per-bar→stride=200 flipped 3-fold mean Sharpe from -4.29 to
  +1.78 at quarter-tick cost)
- Action: target_position categorical ∈ {-10, ..., +10}; max_train=10
  fixed at architecture level, max_live ≤ 10 runtime-cappable
- Reward in price units (position-size invariant): dense per-segment
  PnL − C_trade·|Δcontracts| − C_vol·vol_excess; C_trade=0.034 fixed
  (IBKR $1.70 RT ÷ $50/point), C_vol=0 default
- Full online learning with EWC anchoring + 80/20 offline/live replay
  buffer + kill-switch on σ-band metric deviations (loss, mean reward,
  hit-rate, KL, entropy)
- All hyperparameters that vary during training are ISV-driven
  (controller outputs, not magic numbers) per
  pearl_controller_anchors_isv_driven.md
- Fully GPU-resident on hot path per feedback_cpu_is_read_only.md
- $35k starting capital; conservative max_live ramp 1→2→3 contracts
  with manual config bumps gated on realized track record
- Documented IBKR-specific quirks (PDT, mid-session margin spikes,
  rollover, halts, rate limits) with concrete design responses

Supersedes the abandoned 2026-05-16-alpha-ppo-trainer.md (DQN-port plan)
which inherited the wrong architectural shape.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
This commit is contained in:
jgrusewski
2026-05-16 19:59:51 +02:00
parent 10f4bcc15b
commit d4ea54d780

View File

@@ -0,0 +1,635 @@
# ml-alpha — Greenfield CfC + PPO Trading System Design
**Date**: 2026-05-16
**Status**: Approved (all 7 sections + ISV-driven auto-tuning addition)
**Supersedes**: `2026-05-16-alpha-ppo-trainer.md` (the earlier DQN-port plan, abandoned per user direction to rebuild from first principles)
## TL;DR
A clean greenfield rebuild of `ml-alpha` as a full closed-loop alpha trading system for ES.FUT on IBKR with $35k starting capital. Three components: (1) a Rust model library that owns a CfC perception trunk + 5 multi-horizon alpha heads + a CfC actor-critic policy, (2) an offline training binary, (3) an `AlphaPpoStrategy` integration in the existing `trading_agent_service` for live runtime. Everything GPU-resident on the hot path. Online learning with EWC anchoring + replay buffer (80% offline / 20% live) + kill-switch on σ-band metric deviations. Hyperparameters that vary during training are ISV-driven (signal-driven controllers, no magic numbers). Reuse the existing IBKR adapter, data ingestion, execution engine, and risk manager.
## Why this design exists
The session that led here started with the DQN smoke baseline (`alpha_baseline.rs`) running at per-bar trade frequency, then discovered via stride sweep that **the framing was the bug**: decision cadence at per-bar level was coin-flipping on noise. Switching to `decision_stride=200` with scaled training flipped 3-fold mean Sharpe at quarter-tick cost from **-4.29 (per-bar) to +1.78 (stride=200 + scaled)**, with variance collapsing from ±8.8 to ±1.15.
That validated the alpha signal at IBKR-realistic costs but exposed that `alpha_baseline.rs` was a smoke trainer — not the right vehicle for production. The first attempt to port the production DQN trainer (`gpu_dqn_trainer.rs`, 1.9 MB) was abandoned because (a) DQN's per-bar Bellman target is structurally wrong for minute-horizon alpha, (b) we have no validation that the existing DQN architecture fits this problem, (c) porting DQN means inheriting DQN-shaped assumptions. User direction: rebuild from first principles, no backward-compat, no transitional layers — "consistent in development without endless gates or backward compats."
## Architecture summary (one paragraph)
A snapshot-level CfC perception trunk (Closed-form Continuous-time recurrent network, ~128 cells, each with its own learnable time constant) reads Databento MBP-10 snapshots and emits 5 multi-horizon direction probabilities (P(up at h) for h ∈ {30, 100, 300, 1000, 6000} snapshots forward). A small CfC actor-critic policy reads a 20-dimensional state (5 alpha probs + 8-dim perception projection + position state + spread / vol / time features) and outputs a target position categorical over {-10, ..., 0, ..., +10}. The policy fires every `decision_stride=50` snapshots (~1 minute). Trades happen when `target ≠ previous_target`. Reward per decision step is dense per-segment P&L minus IBKR commission minus optional vol penalty. The model is fully GPU-resident on the hot path. Offline training is two-phase: (A) supervised multi-horizon BCE on the perception, with a hard validation gate requiring CfC AUC ≥ existing Mamba2 baseline before continuing, and (B) PPO on the frozen perception with chronological walk-forward CV over 6 folds spanning 9 quarters of historical data. Live runtime reuses `data_acquisition_service` (Databento streaming already built), `trading_agent_service` (where a new `AlphaPpoStrategy` plugs in), `trading_service::execution_engine` + `risk_manager`, and `crates/trading_engine/src/brokers/interactive_brokers.rs` (the IBKR adapter, 1177 LOC, already exists). Online updates run every 4 hours or 10K live decisions, applying PPO with an EWC term anchored to the offline-trained baseline, drawing 80% from a frozen offline replay buffer + 20% from live trajectories, with a kill-switch that rolls back weights if any of {loss, mean reward, hit-rate, KL divergence, entropy} deviates outside σ-band thresholds.
---
## Section 1 — System scope and integration
### Component boundaries
| | Where it lives | Role | Approx new LOC |
|---|---|---|---|
| `ml-alpha` Rust library | `crates/ml-alpha/` (gut existing files, rebuild clean) | CfC perception + heads + CfC policy + EWC anchor + Adam state + replay buffer types + training kernels | ~3000 |
| `alpha_train` binary | `crates/ml-alpha/examples/alpha_train.rs` | Offline training entry: Phase A (perception pretrain), then Phase B (PPO) per fold | ~300 |
| `AlphaPpoStrategy` | `services/trading_agent_service/src/strategies/alpha_ppo.rs` (NEW) | Live runtime adapter: owns `AlphaModel` + `ReplayBuffer`, called per snapshot, emits target signals | ~250 |
### What gets reused (no new work)
| Component | Existing path |
|---|---|
| Databento streaming | `crates/data/src/providers/databento_streaming.rs` |
| Live MBP-10 ingestion | `services/data_acquisition_service/` |
| Order execution + routing | `services/trading_service/src/core/execution_engine.rs`, `broker_routing.rs`, `order_manager.rs`, `position_manager.rs` |
| Risk manager (limits, kill-switch hook) | `services/trading_service/src/core/risk_manager.rs` |
| Broker gateway | `services/broker_gateway_service/` |
| IBKR adapter (1177 LOC, `ibapi` v10.19+) | `crates/trading_engine/src/brokers/interactive_brokers.rs` |
### Public API of the model library
```rust
// crates/ml-alpha/src/lib.rs (public surface — total ~5 types, ~10 methods)
pub struct AlphaModel { /* CfC weights + heads + policy + EWC anchor + Adam state, all GPU-resident */ }
impl AlphaModel {
pub fn load(weights_path: &Path) -> Result<Self>;
pub fn save(&self, weights_path: &Path) -> Result<()>;
// Hot path — called per snapshot during live. Target latency p99 ≤ 500μs.
pub fn perception_forward(&mut self, snapshot_pinned: &MappedPinnedSnapshot) -> AlphaProbsDevicePtr;
// Decision point — called every decision_stride snapshots.
pub fn policy_action(&self, state_gpu: &AlphaStateGpu) -> (Action, LogProb, Value);
// Slow path — periodic online update; ~30s wall on L40S.
pub fn update(&mut self, batch: &MiniBatch) -> UpdateStatus;
}
pub struct ReplayBuffer { /* offline + live trajectories, GPU-resident */ }
impl ReplayBuffer {
pub fn record_step(&mut self, transition: Transition);
pub fn complete_last_segment_reward(&mut self, reward: f32);
pub fn sample_mixed(&self, n: usize, offline_frac: f32) -> MiniBatch;
pub fn load_offline_from(&mut self, path: &Path) -> Result<()>;
}
pub enum UpdateStatus {
Applied { metrics: TrainingMetrics },
RolledBack { reason: String, deviating_metric: String, sigma: f32 },
}
pub struct AlphaState { /* 20-dim vector, constructed by trading_agent */ }
#[repr(i8)]
pub enum Action { /* target_position ∈ {-10, ..., +10} */ }
```
GPU constraint: weights, hidden states, replay buffer, Adam state, and EWC anchor all live on the GPU. CPU only touches: snapshot bytes from Databento socket (via mapped-pinned), the chosen Action enum (~1 byte), the UpdateStatus enum, and serialized weights at save/load. No CPU shadow buffers. No GPU→CPU roundtrips on the hot path. Mirrors the foxhunt rules: `feedback_cpu_is_read_only.md`, `feedback_no_htod_htoh_only_mapped_pinned.md`, `feedback_no_atomicadd.md`, `feedback_no_nvrtc.md`.
---
## Section 2 — Model architecture
### Snapshot feature vector (input to CfC perception, ~32 dims)
Per MBP-10 snapshot, build a fixed-shape feature vector for the CfC trunk. Built on GPU once when the snapshot lands (via mapped-pinned), reused across all downstream calls.
| Index | Field | Source | Notes |
|---|---|---|---|
| 0 | mid log-return since prev snapshot | (mid_t mid_{t-1}) / mid_{t-1}, then log | Direction signal |
| 1 | L1 spread in ticks | (best_ask best_bid) / tick_size | Liquidity signal |
| 26 | Bid-side per-level size (L1..L5) | Normalized log-size | Top-of-book depth |
| 711 | Ask-side per-level size (L1..L5) | Normalized log-size | Top-of-book depth |
| 1216 | Per-level OFI (L1..L5, signed) | Standard OFI from `ml_features::ofi_calculator` | Order flow imbalance |
| 17 | Trade arrival count since prev snapshot | Integer count, log-scaled | Trade-flow proxy |
| 18 | Trade signed volume since prev snapshot | Sum(buy_size) Sum(sell_size), normalized | Trade-flow direction |
| 19 | Snapshot inter-arrival time (Δt in ms) | From MBP-10 timestamp | Already used by CfC for time-constant integration; included here for the input network too |
| 2031 | Reserved for v2 (Hawkes intensity, microprice deviation, multi-level imbalances, etc.) | Initially zero | Architectural headroom |
L5 depth chosen (not full L10) to keep the input dim modest. L6-L10 are rarely informative for ES.FUT given liquidity is concentrated in L1-L3. Reserved slots 20-31 allow extension without retraining if v2 features prove necessary.
### Perception trunk (CfC, snapshot-level, stateful)
Closed-form Continuous-time recurrent network (Hasani et al. 2022). The practical trainable variant of LTC — replaces ODE solve with a closed-form recurrence, ~50× faster than naive LTC, preserves the "each neuron has its own time-constant" property.
- 128 CfC cells in the perception trunk
- Each cell has its own learnable time constant `τ_i`. Time-constant diversity gives multi-horizon natively: some cells specialize to fast dynamics, others to slow.
- Forward equation per snapshot:
```
h_new_i = h_old_i · exp(-Δt / τ_i)
+ (1 exp(-Δt / τ_i)) · tanh(W_in_i · snap_features + W_rec_i · h_old + b_i)
```
where `Δt` is the actual nanoseconds-since-last-snapshot from MBP-10 timestamps. Irregular timestamps are native — no aggregation or padding to fixed bar intervals.
- Hidden state `h ∈ ^128` lives on device, persists across snapshots. During training: one state per parallel env. During live: single state, persists across calls (re-initialized only at startup or after explicit reset).
### Multi-horizon alpha heads (5 thin linear projections)
For each `h ∈ {30, 100, 300, 1000, 6000}` snapshots forward: a `128 → 1` linear layer + sigmoid produces P(up at h). Trained jointly with the trunk via multi-horizon BCE loss.
### Projection layer (perception → policy state)
A learned `W_proj ∈ ^(8×128)` linear layer projects the perception hidden state to an 8-dim summary that flows into the policy state. Trained jointly with perception (so the projection learns what the policy needs beyond just the explicit alpha probabilities).
### Policy network (CfC actor-critic)
| Component | Architecture |
|---|---|
| Actor | 20-dim state → CfC(64) → linear → 21 logits (target position ∈ {-10, ..., +10}) |
| Critic | 20-dim state → CfC(64) → linear → 1 scalar (V(s)) |
CfC time-constants in the policy let the policy "remember" recent context across decision steps. Useful because the policy fires every `decision_stride=50` snapshots, and the gap between calls is meaningful state.
### EWC anchor
- Stores baseline weights `θ_anchor` (copy of model weights after offline training)
- Stores diagonal Fisher information `F_i = E_val[(∂L_PPO / ∂θ_i)²]` (one scalar per parameter)
- Online updates add a penalty term: `λ_ewc · Σ_i F_i · (θ_current_i θ_anchor_i)²`
- `θ_anchor` is frozen — only updated by explicit operator action
### Parameter count
| Component | Params |
|---|---|
| CfC perception trunk (128 cells, ~20 input dims) | ~20K |
| Multi-horizon heads (5 × 128 → 1) | ~700 |
| Projection (128 → 8) | ~1K |
| CfC actor (20 → 64 → 21) | ~6K |
| CfC critic (20 → 64 → 1) | ~5K |
| EWC anchor + Fisher diagonal | 2× weight count |
| AdamW optimizer state | 2× weight count |
| **Total trainable weights** | **~35K** |
| **Total device memory (weights + state + optimizer)** | **~1 MB** |
Inference is microseconds per snapshot. Training is minutes per fold even on consumer GPU.
---
## Section 3 — Action space, reward, state
### Action space
Single categorical head, target_position ∈ {-10, -9, ..., 0, ..., +9, +10} (21 outputs).
- Architecture is fixed at `max_train = 10` (21 outputs always).
- Live inference takes runtime parameter `max_live ≤ max_train`. Logits outside `[-max_live, +max_live]` are set to `-inf` on-device before softmax; sample or argmax operates on the remaining live range. Zero retraining cost to dial down at deployment.
- Trade-event semantics are derived: `trade_happens = (target_t ≠ target_{t-1})`.
### Reward (dense per-segment, all units in PRICE UNITS, not dollars)
```
r_t = ΔPnL_segment_t ← P&L over the decision_stride window (50 snapshots)
C_trade · |target_t target_{t-1}| ← IBKR commission per contract changed
C_vol · max(0, |segment_return_t| r_threshold) ← vol penalty (default coef 0)
```
All quantities in **price units of the underlying (ES.FUT points)**, not dollars. Conversion: $ amount = price units × $50/point × contracts. Keeping the reward in price units makes it position-size invariant: a 1-point move is the same reward regardless of how many contracts are held; the policy already accounts for `target` magnitude. Dollar P&L is computed only for risk monitoring (daily loss limit, margin) — never in the reward function.
`ΔPnL_segment_t` formal definition:
```
ΔPnL_segment_t = target_t · (mid_close_of_segment mid_open_of_segment)
```
where `mid_open_of_segment` is the mid-price at the snapshot at which the segment started (the previous decision), and `mid_close_of_segment` is the mid at the current decision snapshot.
Constants:
- `C_trade = 0.034` price units per contract (IBKR ES: $0.85/side × 2 = $1.70 RT ÷ $50/point = 0.034)
- `C_vol = 0.0` initially (vol penalty disabled; ISV-driven controller can ramp up if vol becomes a problem — see Section 7)
- `r_threshold = 0.25` (1 tick — first tick treated as expected noise)
### State vector (20 floats, GPU-resident)
| Index | Field | Source | Normalization |
|---|---|---|---|
| 04 | 5 multi-horizon alpha probs (h=30/100/300/1000/6000) | Perception heads (sigmoid) | [0, 1] |
| 512 | 8-dim projection of perception hidden | Learned `W_proj: 128 → 8` | ~unit-variance from layer-norm |
| 13 | Current target position | Strategy state | `current / max_train` ∈ [-1, +1] |
| 14 | Unrealized P&L (in ticks since entry) | Strategy state | clipped to [-20, +20], scaled by /20 |
| 15 | Time-since-entry (decision steps) | Strategy state | clipped to [0, 100], scaled by /100 |
| 16, 17 | Minute-of-day sin, cos | Wall clock | [-1, +1] |
| 18 | Spread quantile (rolling 500-bar) | Live computation | [0, 1] |
| 19 | Recent realized vol (rolling 100-bar log-return std) | Live computation | scaled by historical 95th percentile |
---
## Section 3.5 — Capital envelope and IBKR broker constraints
### Capital math
Initial capital: **$35,000**. ES.FUT contract value ≈ $50/point × ~5500 points ≈ $275k notional per contract.
IBKR margin requirements (subject to broker discretion):
- Day-trading initial margin: ~$500-1,500/contract
- Overnight initial margin: ~$13,200-15,000/contract (exchange minimum ~$11,500 plus IBKR cushion)
- Maintenance margin (intraday): ~$1,000-1,500/contract
Conservative `max_live` ramp plan:
| Phase | `max_live` | Notional exposure | Trigger |
|---|---|---|---|
| Initial deployment | **1 contract** | ~7.8× leverage | Day-1 |
| After 100+ trades with positive realized Sharpe | 2 | ~16× leverage | Manual config bump |
| After 500+ trades, ≥3-month positive track record | 3 | ~23× leverage | Manual config bump |
| Beyond 3 | — | — | Requires more capital or better margin tier |
The model is always trained with `max_train = 10` (network output dim). The live cap grows by manual config bump only — never automatically.
### IBKR quirks (the "tricks" to design around)
| # | Quirk | Required design response |
|---|---|---|
| 1 | Pattern Day Trader rule: ≥4 day-trades in 5 biz days requires ≥$25k equity | Monitor day-trade count; halt strategy if equity drops near $25k threshold to avoid disqualification |
| 2 | Margin can be raised mid-session by IBKR risk dept | Pre-trade `whatif` margin check via IBKR API before submit; reject if margin insufficient |
| 3 | Liquidation if margin breach: IBKR force-closes at market | Maintain ≥40% margin cushion; if margin utilization > 60%, force flat *before* IBKR does |
| 4 | Order rejection without warning | Robust order-status state machine; treat rejection as halt-and-alert; manual ack to resume |
| 5 | Contract rollover: ES quarterly expiry (Mar/Jun/Sep/Dec) | Auto-detect rollover via front-month volume; flat positions in expiring contract 5 days before expiry; open in next month |
| 6 | Market halts (LULDs, circuit breakers) | Subscribe to TWS halt status; force-cancel pending orders; log + alert if positioned |
| 7 | API rate limits (>50 messages/sec) | Bound order submission rate; coalesce target changes within decision_stride |
| 8 | Commission tier (Pro vs Lite, Smart vs direct routing) | Use Smart Routing for ES; pin `$0.85/side` as assumed commission |
| 9 | Symbol drift on rollover (ESM6 → ESU6) | Symbol-lookup layer; never hardcode contract month |
| 10 | Daily loss circuit by IBKR | Match internal threshold; trigger flat earlier than IBKR would |
### Pre-deployment IBKR audit tasks
The existing `crates/trading_engine/src/brokers/interactive_brokers.rs` (1177 LOC, `ibapi` crate) needs an audit before go-live:
1. Verify it handles ES futures specifically (the adapter is built but may be stock/option focused)
2. Confirm `whatif` pre-trade margin check is exposed
3. Confirm rollover detection and symbol-update logic exists
4. Confirm halt-status subscription
5. Audit rejection-code catalog coverage
6. Check PDT counter integration with `risk_manager`
---
## Section 4 — Training pipeline
### Phase A: Perception pretrain (supervised, multi-horizon BCE)
**Data path:**
- Source: predecoded MBP-10 sidecars from the PVC (already built earlier this session). 61M snapshots / 9Q.
- Loader: zero-copy mmap of sidecars → mapped-pinned device staging buffer → CfC reads device memory directly. No CPU shadow.
- Sample shape: sequences of `seq_len = 64` consecutive snapshots, anchored at random valid positions.
- Labels: multi-horizon labels at `{30, 100, 300, 1000, 6000}` snapshots forward, computed once during data pipeline init.
**Training step (per minibatch):**
1. Sample `batch_size = 256` sequences. Upload bar indices to device via mapped-pinned.
2. CfC perception forward over `seq_len = 64`, producing 128-dim hidden + 8-dim projection at every position.
3. Multi-horizon heads emit 5 logits per position → sigmoid.
4. Loss = `(1/5) · Σ_h BCE(σ(logit_h), label_h)` averaged over valid positions.
5. Backward → AdamW step.
**Defaults:**
- Optimizer: AdamW, lr=3e-3, weight_decay=1e-2
- Epochs: 5 over the training window
- Time-constant init: log-uniform over [10ms, 1000s]
### Validation gate: CfC must meet Mamba2
After Phase A on the first 2Q (Q1-Q2 2024):
- Measure val AUC at each of the 5 horizons on Q2 (held out from training)
- Reference: Mamba2 stacker baseline AUC ≈ 0.66-0.71 at K=6000 from prior session runs
- **Gate**: CfC AUC ≥ Mamba2 AUC at every horizon (within 0.01 tolerance). If not met, pause and investigate; do not proceed to Phase B.
If the gate fails, the revert path is clean: the perception layer is the only component that depends on CfC. A revert to Mamba2 trunk (existing code, validated) leaves the policy, env, reward, online update, and runtime integration untouched.
### Phase B: PPO policy training on frozen perception
**Data path:**
- Perception is frozen; its forward outputs (5 alpha probs + 8-dim projection) are computed once per snapshot and cached in a device tensor.
- Iterated in temporal order — we're running an env, not sampling.
**Rollout collection (on-device):**
- `n_par = 64` parallel envs. Each env maintains (current_position, current_pnl, time_since_entry, last_trade_pos) on device.
- Step loop: every snapshot, perception forward (frozen). Every `decision_stride = 50` snapshots, build AlphaState, policy forward, sample action, compute reward over the 50-snapshot segment, store transition.
- After `rollout_length = 2048` decisions per env: stop collecting, run PPO updates.
**PPO update:**
1. Compute GAE: γ=0.99, λ=0.95 (both fixed)
2. K=4 PPO epochs, minibatches of 256 decisions
3. Per minibatch: PPO clip loss + value MSE + entropy bonus + EWC term; AdamW step
4. KL early-stop within K epochs if `KL(current || rollout) > 0.02`
### Walk-forward CV (chronological, 6 folds, expanding window)
| Fold | Train | Val (select within-fold hyperparams) | Test (held-out) |
|---|---|---|---|
| 1 | Q1-Q4 2024 | Q1 2025 | Q2 2025 |
| 2 | Q1 2024 - Q1 2025 | Q2 2025 | Q3 2025 |
| 3 | Q1 2024 - Q2 2025 | Q3 2025 | Q4 2025 |
| 4 | Q1 2024 - Q3 2025 | Q4 2025 | Q1 2026 |
| 5 | Q1 2024 - Q4 2025 | Q1 2026 | (no test — future) |
| 6 | All 9Q | (final calibration) | (live) |
Fold 4's test (Q1 2026) is true forward OOS. Reported test Sharpe is the mean of folds 1-4 test segments, with std as the uncertainty band.
---
## Section 5 — Live runtime integration
### Data flow at runtime
```
data_acquisition_service trading_agent_service trading_service
Databento stream ──→ snapshot event ──→ AlphaPpoStrategy: ──→ emit target signal
1. mapped-pinned snapshot ──→ risk_manager (limits, kill-switch, PDT, margin)
2. model.perception_forward (GPU) ──→ execution_engine
3. every 50 snaps: model.policy ──→ broker_routing → IBKR adapter
4. build Action ↓
5. buffer.record_step broker_gateway_service
IBKR TWS / Gateway
```
### AlphaPpoStrategy lifecycle (sketch)
```rust
pub struct AlphaPpoStrategy {
model: AlphaModel, // CfC + policy + EWC anchor + Adam, all GPU
buffer: ReplayBuffer, // offline + live trajectories, GPU
snap_pinned: MappedPinnedSnapshot, // CPU→GPU staging for one snapshot
decision_counter: usize,
last_target: i8,
last_state_gpu: AlphaStateGpu,
config: AlphaConfig, // max_live, decision_stride, update_cadence, ...
metrics: AlphaMetrics,
update_clock: UpdateClock,
}
impl Strategy for AlphaPpoStrategy {
fn on_snapshot(&mut self, snap: &Mbp10Snapshot) -> Option<Signal> {
self.snap_pinned.write_volatile(snap);
self.model.perception_forward(&self.snap_pinned);
self.decision_counter += 1;
if self.decision_counter % self.config.decision_stride != 0 { return None; }
self.build_state_gpu(snap);
let (action, log_prob, value) = self.model.policy_action(&self.last_state_gpu);
let target = action.clamp_to(self.config.max_live);
let signal = (target != self.last_target).then(|| Signal::SetTargetPosition { contracts: target });
self.buffer.record_step(Transition { state: self.last_state_gpu.clone(), action: target, log_prob, value });
self.last_target = target;
if self.update_clock.tick() {
let batch = self.buffer.sample_mixed(8192, /* offline_frac */ 0.8);
match self.model.update(&batch) {
UpdateStatus::Applied { metrics } => self.metrics.update_applied(metrics),
UpdateStatus::RolledBack { reason, .. } => {
self.metrics.update_rolled_back(&reason);
self.config.online_updates_enabled = false; // auto-disable; manual ack required
tracing::error!(reason, "online update rolled back; auto-updates DISABLED");
}
}
}
signal
}
fn on_fill(&mut self, fill: &Fill) {
self.update_pnl_and_position(fill);
self.buffer.complete_last_segment_reward(self.pnl_delta(fill));
}
}
```
### Persistence and restart semantics
| Element | Persistence | Restart behavior |
|---|---|---|
| Model weights | Atomic write to `/data/ml-alpha/weights-<ts>.bin` on every successful `update()` | Load latest at startup |
| Perception hidden state | Not persisted (v1) | Cold-start: zeros; emit flat signal for first 60s while CfC saturates |
| Replay buffer | Live trajectories written to QuestDB; offline trajectories loaded from PVC | Rebuild: load offline + replay live from QuestDB since last weight snapshot |
| Position state | Owned by `trading_service::position_manager` | Reconcile with broker on startup |
| EWC anchor + Fisher matrix | Saved alongside weights | Loaded with weights |
### Observability
- Per snapshot: latency p50/p99 (Prometheus histogram)
- Per decision: action sampled, alpha probs at each horizon (QuestDB row)
- Per fill: realized P&L, position change, commission (QuestDB row)
- Per update: training metrics, kill-switch margin (Prometheus)
- Alerts: kill-switch fired, daily loss limit hit, margin breach, IBKR rejection
---
## Section 6 — Online learning mechanics
### EWC anchor — what gets frozen, what gets learned
After Phase B finishes on the training window:
1. Take trained weights `θ_anchor` and freeze them as reference.
2. Compute diagonal Fisher information `F_i = E_val[(∂L_PPO / ∂θ_i)²]` per parameter.
3. Persist `(θ_anchor, F)` to disk alongside live weights.
During online updates:
```
loss = loss_PPO + λ_ewc · Σ_i F_i · (θ_current_i θ_anchor_i)²
```
`λ_ewc` is the conservatism dial — initial value 1e3 but driven by a controller during operation (see Section 7).
### Replay buffer composition
| Component | Source | Fraction | Lifecycle |
|---|---|---|---|
| Offline trajectories | Saved during Phase B (last ~3 training quarters, ~50K decisions) | 80% (controller-tunable) | Loaded once at startup; GPU-resident; never updated |
| Live trajectories | Collected by `AlphaPpoStrategy::on_snapshot` | 20% (controller-tunable) | Ring buffer of ~10K most recent live decisions; mirror written to QuestDB |
Each minibatch: 8192 transitions = 6553 offline + 1638 live (at default 80/20).
### Update cadence
Whichever fires first:
- Time-based: every `update_interval_secs = 14_400` (4 hours)
- Volume-based: every `update_interval_transitions = 10_000` live decisions
The update lasts ~30s on L40S; `policy_action` calls block momentarily during atomic weights swap (~10ms typical).
### Atomic weights swap with kill-switch
```rust
fn update(&mut self, batch: &MiniBatch) -> UpdateStatus {
let snapshot_path = self.write_weights_snapshot()?; // versioned snapshot first
let candidate_weights = self.weights.clone_gpu(); // GPU-side copy
let metrics = self.run_ppo_epochs_on(&candidate_weights, batch);
match self.kill_switch.check(&metrics) {
KillSwitchVerdict::Ok => {
std::mem::swap(&mut self.weights, &mut candidate_weights);
self.kill_switch.record(metrics.clone());
UpdateStatus::Applied { metrics }
}
KillSwitchVerdict::Reject { reason, deviating_metric, sigma } => {
// candidate_weights dropped — live weights untouched
UpdateStatus::RolledBack { reason, deviating_metric, sigma }
}
}
}
```
### Kill-switch criteria
Rolling baseline tracked from last 20 successful updates.
| Metric | Reject condition | Why |
|---|---|---|
| `loss_ppo` | > baseline_mean + 3σ | Training divergence |
| `mean_reward_per_decision` | < baseline_mean 2σ | Strategy degrading |
| `hit_rate` | absolute < 0.45 | Below empirical floor |
| `kl_divergence(candidate \|\| anchor)` | > 0.1 | Policy drifting; EWC has failed |
| `entropy` | < 0.1 nats | Policy collapse (deterministic) |
If any fires: weights stay at pre-update state, `online_updates_enabled = false` in config, operator alert via Prometheus + log. Manual ack required to re-enable.
### Snapshot retention
Every successful update creates `weights-<timestamp>-<idx>.bin` (~150 KB). Retain rolling last 100. Rollback CLI: `alpha-control rollback --to <timestamp>`.
The anchor `(θ_anchor, F)` is FROZEN automatically — only updated by explicit operator command after a multi-month track record.
---
## Section 7 — Testing strategy, ISV-driven hyperparameters, risks
### Testing tiers (in order of cost / fidelity)
| Tier | What | When | Pass criterion |
|---|---|---|---|
| Unit tests | CfC cell, EWC loss, GAE, PPO clip, reward formula, action masking, replay buffer | Continuously | All pass |
| Bit-equiv | CfC forward CPU vs GPU on small inputs | Per commit | < 1e-5 relative error |
| Smoke training | Phase A perception on 1Q subset | Before each merge | val AUC > 0.55 at K=6000 |
| **CfC vs Mamba2 gate** | Phase A perception on 2Q | Before Phase B | CfC AUC ≥ Mamba2 AUC at every horizon |
| Walk-forward Sharpe | 6-fold CV with PPO on 9Q | Before paper trading | Mean test Sharpe > 0 at IBKR commission; std < mean |
| **Paper trading** | Live MBP-10 + simulated execution (no broker orders); 5 trading days | Before live capital | Live P&L tracks backtest ±15%; no kill-switch fires; latency p99 < 500μs |
| Live with $1k cap | Real money, `max_live=1`, manual kill armed, 5-day window | Before scaling | Realized Sharpe > 0; zero IBKR rejections; rollover clean if it occurs |
### ISV-driven hyperparameters (per `pearl_controller_anchors_isv_driven.md`)
Most "hyperparameters" are controller outputs, not config constants. The ml-alpha trainer gets its own ISV bus (~32 slots).
#### Controller-driven (signal-driven adaptation)
| Hyperparam | Controller signal | Update rule |
|---|---|---|
| `entropy_coef` | Entropy deficit = (target_entropy current_entropy), target = 70% of uniform | `entropy_coef ← entropy_coef · exp(η · deficit)` |
| `lr_policy` | Loss-slope EMA + grad-norm stability | Wiener-α blend toward target_loss_slope = -ε per epoch; floor at 1e-5 |
| `λ_ewc` | Weight drift magnitude `‖θ θ_anchor‖²` vs target | Raise λ if drift > 1.2× target; lower if < 0.5× target |
| `clip_eps` (PPO) | KL divergence trajectory; target KL = 0.01 | Raise clip if KL too low (under-update); lower if too high (over-update) |
| `value_coef` | Ratio of value loss to policy loss; target = 1:5 | Adjust value_coef to match target ratio |
| Replay offline/live mix | Reward gap = `live_reward_mean offline_reward_mean` | Increase live fraction if gap > Nσ (regime shift signal); decrease if gap < 0 |
| `C_vol` (vol penalty) | Variance of returns vs target | Raise coef if return std > target × 1.5; keep at 0 otherwise |
All controllers use the standard foxhunt pattern: Wiener-α optimal blending + permanent floor (per `pearl_wiener_alpha_floor_for_nonstationary.md` and `pearl_blend_formulas_must_have_permanent_floor.md`).
#### Fixed parameters (genuinely architectural)
| Parameter | Value | Why constant |
|---|---|---|
| `decision_stride` | 50 snapshots | Architecture choice; changing requires retraining |
| `seq_len` (perception) | 64 snapshots | Architecture choice |
| `n_horizons` | 5 (30/100/300/1000/6000) | Architecture choice |
| `max_train` | 10 contracts | Architecture choice (network output dim) |
| `n_par` (parallel envs) | 64 | Hardware fit |
| `batch_size` (perception) | 256 sequences | Hardware fit |
| `batch_size` (PPO update) | 8192 | Hardware fit |
| GAE γ | 0.99 | Standard; controller would risk subtle bugs |
| GAE λ | 0.95 | Standard |
| K (PPO epochs) | 4 | KL early-stop within K catches over-update |
| Kill-switch σ thresholds | 3σ loss, 2σ reward | Safety floor — operator-controlled, not signal-driven |
#### ISV slot allocation (32 slots)
```
ml-alpha ISV bus
─────────────────────────────────
0..5 Perception output: 5 alpha probs at h={30,100,300,1000,6000}
5..13 Perception hidden projection (8 dims)
13..18 Training diagnostics:
13: loss_ppo EMA
14: kl_divergence (vs anchor)
15: entropy_current
16: grad_norm EMA
17: weight_drift |θ θ_anchor|²
18..24 Controller outputs (effective hyperparams):
18: lr_policy_eff
19: lr_perception_eff
20: entropy_coef_eff
21: clip_eps_eff
22: value_coef_eff
23: λ_ewc_eff
24..28 Replay buffer state:
24: mean_live_reward EMA
25: mean_offline_reward EMA
26: live_offline_mix_eff (0..1)
27: buffer_fill_fraction
28..32 Operational / live:
28: current_pos_eff
29: realized_pnl_session_eff
30: kill_switch_margin (σ-distance from rolling baseline)
31: rolling_baseline_idx
```
### Known risks
| Risk | Severity | Mitigation |
|---|---|---|
| CfC fails to match Mamba2 | High | Validation gate stops Phase B; perception layer is replaceable in isolation; revert to Mamba2 if needed |
| PPO entropy collapse | Medium | Kill-switch fires on entropy < 0.1; entropy_coef controller raises automatically |
| Replay buffer staleness | Medium | 80% offline data is from fixed historical window; if regime shifts dramatically the offline anchor is obsolete. Manual remix or new offline trajectory generation needed periodically |
| IBKR margin spike on FOMC | Medium | Pre-trade `whatif` rejects oversized orders; daily loss limit catches systemic damage |
| Snapshot CfC compute slower than 500μs | Medium | Profile early; if perception forward > 100μs per snapshot may need TensorRT export for live (not training) |
| Decision-stride / horizon mismatch | Low | Decision_stride=50 << longest horizon 6000; plenty of decision points. Verified this session |
| EWC λ mis-tuning | Low | Now controller-driven, not hardcoded |
| Contract rollover bugs | Low (but high impact) | Existing trading_engine handles symbol changes; audit before go-live |
### What this design does NOT include
- HFT / sub-millisecond latency (target ≤500μs per snapshot, not nanoseconds)
- Cross-asset universe expansion (ES.FUT only)
- Hedging logic (outright ES only)
- Options strategies
- Hyperparameter sweep over architecture (the architecture itself — number of CfC cells, horizons, decision_stride — is fixed)
- Auto-rollback of `θ_anchor` (anchor moves only by explicit operator action)
- Time-of-day filter (overnight trading allowed; operator-controlled risk)
---
## Implementation order
1. **IBKR adapter audit** (1-2 days) — verify ES futures, `whatif`, rollover, halt, PDT counter
2. **ml-alpha library scaffold + ISV slot allocation** (1 day)
3. **CfC perception forward + multi-horizon heads + Phase A trainer** (3-4 days)
4. **Phase A run + CfC-vs-Mamba2 validation gate on 2Q** (1 day cluster) — GATE
5. **CfC policy (actor-critic) + reward + state construction** (2-3 days)
6. **PPO loss + GAE + replay buffer + EWC term + Phase B trainer** (3-4 days)
7. **ISV controllers (entropy, lr, λ_ewc, clip_eps, value_coef, replay mix, C_vol)** (2-3 days)
8. **Online update + kill-switch + atomic swap** (2 days)
9. **AlphaPpoStrategy integration in trading_agent_service** (2-3 days)
10. **Walk-forward CV (6 folds, full 9Q) on cluster** (1-2 days, mostly cluster wall)
11. **Paper trading on live MBP-10 + simulated execution (5 day window)**
12. **Live deploy with `max_live=1` (5 day window)**
13. **Ramp `max_live` per the table in Section 3.5**
Estimated end-to-end timeline to live: **~3-4 weeks of focused work**, including paper trading.
## Appendix A — Memory entries this design embodies
- `feedback_cpu_is_read_only.md` — no CPU compute, GPU is sole arithmetic engine
- `feedback_no_htod_htoh_only_mapped_pinned.md` — snapshot bytes via mapped-pinned, no HtoD
- `feedback_no_atomicadd.md` — block tree-reduce only (relevant for the multi-horizon BCE + EWC term kernels)
- `feedback_no_nvrtc.md` — pre-compiled cubins for any CUDA kernels
- `pearl_controller_anchors_isv_driven.md` — all tunables ISV-controlled
- `feedback_isv_for_adaptive_bounds.md` — bounds live in ISV
- `feedback_adaptive_not_tuned.md` — fixes are signal-driven, not tuned constants
- `pearl_wiener_alpha_floor_for_nonstationary.md` — floor of 0.4 on Wiener-α in non-stationary loops
- `pearl_blend_formulas_must_have_permanent_floor.md` — max-with-floor, not blend
- `pearl_first_observation_bootstrap.md` — EMAs bootstrap from sentinel = 0; first observation replaces directly
- `pearl_bar_abstraction_can_be_alpha_ceiling.md` — bar aggregation can cap alpha; snapshot-level chosen for this reason
- `pearl_state_amplifies_short_horizon_into_long_horizon.md` — stateful encoder (Mamba2/CfC) gives multi-minute alpha from snapshot input
- `pearl_phase1d4_backtest_cost_edge_frontier.md` — frictionless Sharpe ~4.4 baseline; break-even quarter-tick; this session pushed to ~1-tick at stride=200
## Appendix B — Today's session findings that motivated this design
- Per-bar trading = coin-flip on noise. Decision-stride 200 with scaled training flipped Fold A Sharpe at quarter-tick cost from -3.7 to +2.52; 3-fold mean from -4.29 to +0.78; with 16× scaled training to +1.78. The signal exists at minute-horizon; the framing of "decide every bar" was the bug.
- Variance across folds collapsed from ±8.8 (per-bar) to ±0.80 (stride=200) — removing noise-mining made the alpha consistent.
- IBKR commission at $1.70 RT translates to 0.034 price units — well below the new break-even of ~1 tick. Economic deployment is plausible.
- `alpha_baseline.rs` was a smoke trainer; the production DQN trainer (`gpu_dqn_trainer.rs`) is engineered for a different problem regime (per-bar Bellman, branched discrete actions); porting it would inherit the wrong shape. Greenfield rebuild with PPO + multi-horizon perception + CfC is the right move.