- Remove all HTTPS/TLS from MinIO (plain HTTP for internal cluster traffic) - Fix sccache 0% cache hit rate (rustls rejected self-signed MinIO cert) - Remove hardcoded URLs from k8s_dispatcher.rs (S3_ENDPOINT, TRAINING_RUNTIME_IMAGE, CALLBACK_ENDPOINT now required env vars) - Update GitLab registry S3 credentials to HTTP endpoint - Fix PVC manifest (20Gi → 100Gi to match cluster) - Fix nodeSelector: infra/foxhunt → platform (match actual node pool) - Fix rclone trailing backslash causing chmod to be parsed as rclone args - Remove minio-ca-cert ConfigMap references from all manifests - Update trading-service GPU overlay to l40s pool 20 files changed, -118 lines net Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
ml_training_service
Model training orchestration and lifecycle management for DQN, PPO, TFT, Mamba2, TLOB, and Liquid models with progress tracking and artifact storage.
Key Types
MlTrainingServiceImpl-- main gRPC serviceJobTracker-- training job state machineCheckpointManager-- model artifact persistence
Features
minimal(default) -- minimal ML feature set for financial modelsgpu-- SIMD GPU acceleration (requires CUDA)mock-data-- mock training data (testing, bypasses database)
Configuration
GRPC_PORT-- gRPC listen portDATABASE_URL-- PostgreSQL for job metadata and training history- Prometheus metrics on port 9094
Testing
SQLX_OFFLINE=true cargo test -p ml_training_service --lib