- Helm values (controller + server on platform node, MinIO artifact repo) - WorkflowTemplate: parameterized 5-step DAG (fetch→hyperopt→train→eval→upload) - Nginx proxy for argo.fxhnt.ai → Argo Server :2746 - DNS A record for argo.fxhnt.ai - MinIO bucket foxhunt-training-results for Argo artifacts - Kustomization for kubectl apply -k Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
4.8 KiB
4.8 KiB
Argo Workflows Training Orchestration Design
Goal
Replace GitLab CI training jobs with Argo Workflows for ML training orchestration only. GitLab CI continues to handle build/test/deploy. Argo provides DAG scheduling, parameterized workflows, and native MinIO artifact integration.
Architecture
- Argo Workflows controller installed in
foxhuntnamespace via Helm - Argo Events (EventSource + Sensor) for GitLab webhook → workflow submission
- Single parameterized WorkflowTemplate covering the full training pipeline per model
- MinIO as native Argo artifact repository (already deployed at
minio.foxhunt.svc.cluster.local:9000)
Scope
In scope (Argo): hyperopt, train, evaluate — all GPU training jobs Out of scope (stays in GitLab CI): prepare, test, compile, deploy, web dashboard build
WorkflowTemplate Parameters
| Parameter | Default | Description |
|---|---|---|
model |
(required) | Model name: dqn, ppo, tft, mamba2, tggn, tlob, liquid, kan, xlstm, diffusion |
gpu-pool |
ci-training |
Node pool: ci-training (L40S) or ci-training-h100 (H100) |
hyperopt-trials |
5 |
Number of hyperopt PSO trials |
hyperopt-epochs |
8 |
Epochs per hyperopt trial |
train-epochs |
50 |
Epochs for final training with best hyperparams |
symbol |
ES.FUT |
Trading symbol for data |
data-dir |
/data/cache/futures-baseline |
Path to training data |
Workflow Steps (per model)
fetch-binary → hyperopt → train-best → evaluate → upload-results
- fetch-binary: rclone copy from
s3://foxhunt-binaries/training/to workspace - hyperopt: Run
hyperopt_baseline_rlorhyperopt_baseline_supervisedwith trial/epoch params - train-best: Run
train_baseline_rlortrain_baseline_supervisedwith best hyperparams from step 2 - evaluate: Run
evaluate_baselineagainst test data window - upload-results: Upload checkpoints + eval results to MinIO
s3://foxhunt-training-results/
Binary Mapping
| Models | Binary (hyperopt) | Binary (train) |
|---|---|---|
| dqn, ppo | hyperopt_baseline_rl |
train_baseline_rl |
| tft, mamba2, tggn, tlob, liquid, kan, xlstm, diffusion | hyperopt_baseline_supervised |
train_baseline_supervised |
All models use evaluate_baseline for evaluation.
Pod Templates
Two pod templates matching existing runtime images:
-
gpu-training:
foxhunt-training-runtime:latest(CUDA 12.4 + cuDNN + rclone + nvrtc)- GPU node selector:
k8s.scaleway.com/pool-name: {{workflow.parameters.gpu-pool}} - Resources: requests 1 GPU, 32Gi memory, 8 CPU
- tolerations:
dedicated=gpu:NoSchedule
- GPU node selector:
-
cpu-utility:
foxhunt-runtime:latest(debian bookworm + rclone)- For fetch-binary and upload-results steps
- Resources: requests 2 CPU, 4Gi memory
Argo Events Integration
- EventSource: GitLab webhook receiver on
argo-events.fxhnt.ai(or internal service) - Sensor: Filters push events, maps changed paths to model triggers:
crates/ml/src/dqn/**→ trigger DQN workflowcrates/ml/src/trainers/ppo/**→ trigger PPO workflowcrates/ml/src/models/tft/**→ trigger TFT workflow- etc.
- Manual trigger:
argo submitCLI or Argo UI atargo.fxhnt.ai
DNS & Networking
argo.fxhnt.ai→ Tailscale proxy → Argo Server (port 2746)- Argo Server with
--auth-mode=server(no SSO needed initially) - Add nginx proxy block in
infra/k8s/gitlab/tailscale-proxy.yaml - Add DNS A record in
infra/modules/dns/main.tf
MinIO Artifact Repository
Argo native artifact config pointing to existing MinIO:
artifactRepositoryRef:
configMap: artifact-repositories
key: default-v1
ConfigMap:
s3:
endpoint: minio.foxhunt.svc.cluster.local:9000
insecure: true
bucket: foxhunt-training-results
accessKeySecret:
name: minio-credentials
key: access-key
secretKeySecret:
name: minio-credentials
key: secret-key
Migration Plan
Phase 1: Install & Validate
- Install Argo Workflows controller via Helm (minimal config)
- Create artifact repository ConfigMap
- Create WorkflowTemplate
- Manually submit test workflow for DQN
- Verify it runs on GPU, produces checkpoints
Phase 2: Wire Events & Parallel Run
- Install Argo Events (EventSource + Sensor)
- Configure GitLab webhook
- Run both GitLab CI training AND Argo Workflows in parallel
- Compare results, validate reliability
Phase 3: Cutover
- Remove training jobs from
.gitlab-ci.yml - Argo becomes sole training orchestrator
- GitLab CI only: prepare → test → compile → deploy
Constraints
- Keep GitLab CI working until Argo fully validated
- Not everything runs at once — models can be triggered independently
- Single POP2-32C-128G node for CPU work, GPU nodes scale 0→1
- MinIO credentials already in k8s secret
minio-credentials