diff --git a/docs/plans/2026-03-01-deferred-stubs-design.md b/docs/plans/2026-03-01-deferred-stubs-design.md new file mode 100644 index 000000000..bca396373 --- /dev/null +++ b/docs/plans/2026-03-01-deferred-stubs-design.md @@ -0,0 +1,136 @@ +# Deferred Stubs Implementation — Design + +**Date**: 2026-03-01 +**Branch**: `feature/deferred-stubs` +**Scope**: 3 deferred stubs from 2026-02-24 audit + +## Context + +The stub audit (2026-02-24) identified 18 findings. 8 were fixed with real implementations, +3 with honest errors, 3 acceptable, and 4 deferred. Of those 4: + +- WS streams (stream_model_metrics, stream_signal_strength) — **already fixed** (verified 2026-03-01, both publish real data via tokio channels) +- Feature importance — **correctly stubbed** (no feasible path without SHAP/external lib) + +That leaves **3 real deferred items** to implement: + +| Stub | Location | Current Behavior | +|------|----------|------------------| +| Retrain Model | `trading_service/enhanced_ml.rs:1018` | `Status::unavailable` | +| Portfolio Positions | `risk-data/src/var.rs:641` | Returns empty `Vec` | +| ML Confidence | `trading_agent_service/autonomous_scaling.rs:665` | Liquidity heuristic `score * 0.9 + 0.1` | + +## Area 1: Retrain Model — Fine-Tune via gRPC + +### Problem + +`retrain_model()` returns `Status::unavailable("not yet wired to ml_training_service")`. +The ml_training_service has `StartTraining` RPC but no fine-tune mode. + +### Design + +1. **Proto change**: Add `TrainingMode` enum (`FULL`, `FINE_TUNE`) and fields to + `StartTrainingRequest` in `services/ml_training_service/proto/ml_training.proto`: + - `TrainingMode mode` — selects full vs fine-tune + - `string resume_checkpoint_path` — path to last checkpoint + - `uint32 max_epochs` — fine-tune uses fewer epochs (default 10) + +2. **trading_service client**: Add `MlTrainingServiceClient` to `EnhancedMLServiceImpl` + via `OnceCell` lazy initialization. Endpoint configured via env var + `ML_TRAINING_SERVICE_URL` (default `http://ml-training-service:50052`). + +3. **retrain_model() handler**: + - Look up `RuntimeModelInfo` for the requested model_name + - Map model_name → `CommonModelType` enum + - Build `StartTrainingRequest` with mode=FINE_TUNE, checkpoint path, max_epochs=10 + - Call `client.start_training(request).await` + - Return `RetrainModelResponse { job_id, status: "started" }` + +4. **ml_training_service handler**: In `start_training()`: + - If mode=FINE_TUNE: load checkpoint, train on recent 1-month data window, 10 epochs + - If mode=FULL: existing behavior (walk-forward, full training) + +### Risk + +Low-Medium. Proto change is additive (backward compatible). The fine-tune code path +in ml_training_service builds on existing training infrastructure. + +## Area 2: Portfolio Positions — Broker + DB Fallback + +### Problem + +`get_portfolio_positions()` in `risk-data/src/var.rs` returns an empty Vec, +causing VaR calculations to fail with "No positions found for portfolio". + +### Design + +1. **PositionProvider trait** in `crates/risk-data/src/position_provider.rs`: + ```rust + #[async_trait] + pub trait PositionProvider: Send + Sync { + async fn get_positions(&self, portfolio_id: &str) + -> RiskDataResult>; + } + ``` + +2. **DatabasePositionProvider**: Queries `portfolio_positions` table via SQLX. + Schema: `(id, portfolio_id, symbol, quantity, avg_entry_price, current_price, + side, updated_at)`. Uses existing SQLX offline mode. + +3. **BrokerPositionProvider**: Calls broker_gateway_service gRPC `GetPositions` RPC. + Falls back gracefully if broker is unreachable. + +4. **Fallback chain** in `get_portfolio_positions()`: + - Try broker (live positions, most accurate) + - Try DB (cached positions from fill events) + - Return error if both fail + +5. **Position sync**: Trading engine's fill handler writes/updates positions in DB + on each fill event. This keeps the DB cache fresh even when broker is offline. + +6. **SQL migration**: Add `portfolio_positions` table via SQLX migration. + +### Risk + +Medium. Requires DB schema change and new SQLX queries. The broker_gateway is +Phase 2 (FIX routing MVP), so broker path may not work immediately — DB path +is the primary functional path. + +## Area 3: ML Confidence — Ensemble RPC + +### Problem + +`compute_tier_allocations()` in `autonomous_scaling.rs` uses a liquidity-based +heuristic for ML confidence instead of calling the ensemble service. + +### Design + +1. **Add optional gRPC client**: `Option` on the autonomous + scaling service struct. Initialized from env var `ML_SERVICE_URL`. + +2. **Call GetEnsembleVote**: Per instrument in `compute_tier_allocations()`: + - Build `EnsembleRequest { model_name: "all", symbol: Some(instrument.symbol) }` + - Call `client.get_ensemble_vote(request).await` + - Extract confidence from response + +3. **Fallback**: On RPC failure, use existing heuristic `liquidity_score * 0.9 + 0.1` + +4. **Caching**: Cache ensemble responses for 5 seconds per symbol to avoid + excessive RPC calls during allocation computation. + +### Risk + +Low. The fallback is already tested and functional. This is purely additive. + +## Execution Order + +1. **Retrain Model** (smallest scope, unblocks model lifecycle) +2. **ML Confidence** (low complexity, improves trading accuracy) +3. **Portfolio Positions** (largest scope, requires DB migration) + +## Verification + +After each area: +- `SQLX_OFFLINE=true cargo check --workspace` +- `SQLX_OFFLINE=true cargo clippy -p --lib -- -D warnings` +- `SQLX_OFFLINE=true cargo test -p --lib`