Design for 3 remaining deferred stubs from 2026-02-24 audit: - Retrain Model: fine-tune via gRPC with TrainingMode proto field - Portfolio Positions: PositionProvider trait with broker + DB fallback - ML Confidence: ensemble RPC with liquidity heuristic fallback Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
137 lines
5.2 KiB
Markdown
137 lines
5.2 KiB
Markdown
# Deferred Stubs Implementation — Design
|
|
|
|
**Date**: 2026-03-01
|
|
**Branch**: `feature/deferred-stubs`
|
|
**Scope**: 3 deferred stubs from 2026-02-24 audit
|
|
|
|
## Context
|
|
|
|
The stub audit (2026-02-24) identified 18 findings. 8 were fixed with real implementations,
|
|
3 with honest errors, 3 acceptable, and 4 deferred. Of those 4:
|
|
|
|
- WS streams (stream_model_metrics, stream_signal_strength) — **already fixed** (verified 2026-03-01, both publish real data via tokio channels)
|
|
- Feature importance — **correctly stubbed** (no feasible path without SHAP/external lib)
|
|
|
|
That leaves **3 real deferred items** to implement:
|
|
|
|
| Stub | Location | Current Behavior |
|
|
|------|----------|------------------|
|
|
| Retrain Model | `trading_service/enhanced_ml.rs:1018` | `Status::unavailable` |
|
|
| Portfolio Positions | `risk-data/src/var.rs:641` | Returns empty `Vec<PortfolioPosition>` |
|
|
| ML Confidence | `trading_agent_service/autonomous_scaling.rs:665` | Liquidity heuristic `score * 0.9 + 0.1` |
|
|
|
|
## Area 1: Retrain Model — Fine-Tune via gRPC
|
|
|
|
### Problem
|
|
|
|
`retrain_model()` returns `Status::unavailable("not yet wired to ml_training_service")`.
|
|
The ml_training_service has `StartTraining` RPC but no fine-tune mode.
|
|
|
|
### Design
|
|
|
|
1. **Proto change**: Add `TrainingMode` enum (`FULL`, `FINE_TUNE`) and fields to
|
|
`StartTrainingRequest` in `services/ml_training_service/proto/ml_training.proto`:
|
|
- `TrainingMode mode` — selects full vs fine-tune
|
|
- `string resume_checkpoint_path` — path to last checkpoint
|
|
- `uint32 max_epochs` — fine-tune uses fewer epochs (default 10)
|
|
|
|
2. **trading_service client**: Add `MlTrainingServiceClient` to `EnhancedMLServiceImpl`
|
|
via `OnceCell` lazy initialization. Endpoint configured via env var
|
|
`ML_TRAINING_SERVICE_URL` (default `http://ml-training-service:50052`).
|
|
|
|
3. **retrain_model() handler**:
|
|
- Look up `RuntimeModelInfo` for the requested model_name
|
|
- Map model_name → `CommonModelType` enum
|
|
- Build `StartTrainingRequest` with mode=FINE_TUNE, checkpoint path, max_epochs=10
|
|
- Call `client.start_training(request).await`
|
|
- Return `RetrainModelResponse { job_id, status: "started" }`
|
|
|
|
4. **ml_training_service handler**: In `start_training()`:
|
|
- If mode=FINE_TUNE: load checkpoint, train on recent 1-month data window, 10 epochs
|
|
- If mode=FULL: existing behavior (walk-forward, full training)
|
|
|
|
### Risk
|
|
|
|
Low-Medium. Proto change is additive (backward compatible). The fine-tune code path
|
|
in ml_training_service builds on existing training infrastructure.
|
|
|
|
## Area 2: Portfolio Positions — Broker + DB Fallback
|
|
|
|
### Problem
|
|
|
|
`get_portfolio_positions()` in `risk-data/src/var.rs` returns an empty Vec,
|
|
causing VaR calculations to fail with "No positions found for portfolio".
|
|
|
|
### Design
|
|
|
|
1. **PositionProvider trait** in `crates/risk-data/src/position_provider.rs`:
|
|
```rust
|
|
#[async_trait]
|
|
pub trait PositionProvider: Send + Sync {
|
|
async fn get_positions(&self, portfolio_id: &str)
|
|
-> RiskDataResult<Vec<PortfolioPosition>>;
|
|
}
|
|
```
|
|
|
|
2. **DatabasePositionProvider**: Queries `portfolio_positions` table via SQLX.
|
|
Schema: `(id, portfolio_id, symbol, quantity, avg_entry_price, current_price,
|
|
side, updated_at)`. Uses existing SQLX offline mode.
|
|
|
|
3. **BrokerPositionProvider**: Calls broker_gateway_service gRPC `GetPositions` RPC.
|
|
Falls back gracefully if broker is unreachable.
|
|
|
|
4. **Fallback chain** in `get_portfolio_positions()`:
|
|
- Try broker (live positions, most accurate)
|
|
- Try DB (cached positions from fill events)
|
|
- Return error if both fail
|
|
|
|
5. **Position sync**: Trading engine's fill handler writes/updates positions in DB
|
|
on each fill event. This keeps the DB cache fresh even when broker is offline.
|
|
|
|
6. **SQL migration**: Add `portfolio_positions` table via SQLX migration.
|
|
|
|
### Risk
|
|
|
|
Medium. Requires DB schema change and new SQLX queries. The broker_gateway is
|
|
Phase 2 (FIX routing MVP), so broker path may not work immediately — DB path
|
|
is the primary functional path.
|
|
|
|
## Area 3: ML Confidence — Ensemble RPC
|
|
|
|
### Problem
|
|
|
|
`compute_tier_allocations()` in `autonomous_scaling.rs` uses a liquidity-based
|
|
heuristic for ML confidence instead of calling the ensemble service.
|
|
|
|
### Design
|
|
|
|
1. **Add optional gRPC client**: `Option<MlServiceClient>` on the autonomous
|
|
scaling service struct. Initialized from env var `ML_SERVICE_URL`.
|
|
|
|
2. **Call GetEnsembleVote**: Per instrument in `compute_tier_allocations()`:
|
|
- Build `EnsembleRequest { model_name: "all", symbol: Some(instrument.symbol) }`
|
|
- Call `client.get_ensemble_vote(request).await`
|
|
- Extract confidence from response
|
|
|
|
3. **Fallback**: On RPC failure, use existing heuristic `liquidity_score * 0.9 + 0.1`
|
|
|
|
4. **Caching**: Cache ensemble responses for 5 seconds per symbol to avoid
|
|
excessive RPC calls during allocation computation.
|
|
|
|
### Risk
|
|
|
|
Low. The fallback is already tested and functional. This is purely additive.
|
|
|
|
## Execution Order
|
|
|
|
1. **Retrain Model** (smallest scope, unblocks model lifecycle)
|
|
2. **ML Confidence** (low complexity, improves trading accuracy)
|
|
3. **Portfolio Positions** (largest scope, requires DB migration)
|
|
|
|
## Verification
|
|
|
|
After each area:
|
|
- `SQLX_OFFLINE=true cargo check --workspace`
|
|
- `SQLX_OFFLINE=true cargo clippy -p <affected_crates> --lib -- -D warnings`
|
|
- `SQLX_OFFLINE=true cargo test -p <affected_crates> --lib`
|