Files
foxhunt/services/ml_training_service
jgrusewski 6d731b860d fix(infra): use NodePort localhost:30500 for all container image refs
containerd on Kapsule nodes uses host DNS which can't resolve
.svc.cluster.local names. Switch all image references from
gitlab-registry.foxhunt.svc.cluster.local:5000 to localhost:30500
(NodePort on the GitLab registry). Also make training runtime image
configurable via TRAINING_RUNTIME_IMAGE env var.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-05 14:31:54 +01:00
..

ml_training_service

Model training orchestration and lifecycle management for DQN, PPO, TFT, Mamba2, TLOB, and Liquid models with progress tracking and artifact storage.

Key Types

  • MlTrainingServiceImpl -- main gRPC service
  • JobTracker -- training job state machine
  • CheckpointManager -- model artifact persistence

Features

  • minimal (default) -- minimal ML feature set for financial models
  • gpu -- SIMD GPU acceleration (requires CUDA)
  • mock-data -- mock training data (testing, bypasses database)

Configuration

  • GRPC_PORT -- gRPC listen port
  • DATABASE_URL -- PostgreSQL for job metadata and training history
  • Prometheus metrics on port 9094

Testing

SQLX_OFFLINE=true cargo test -p ml_training_service --lib