Files
foxhunt/docs/plans/2026-03-04-argo-workflows-training-design.md
jgrusewski a2468846e1 feat(argo): add Argo Workflows infrastructure for training orchestration
- Helm values (controller + server on platform node, MinIO artifact repo)
- WorkflowTemplate: parameterized 5-step DAG (fetch→hyperopt→train→eval→upload)
- Nginx proxy for argo.fxhnt.ai → Argo Server :2746
- DNS A record for argo.fxhnt.ai
- MinIO bucket foxhunt-training-results for Argo artifacts
- Kustomization for kubectl apply -k

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-04 22:58:40 +01:00

4.8 KiB

Argo Workflows Training Orchestration Design

Goal

Replace GitLab CI training jobs with Argo Workflows for ML training orchestration only. GitLab CI continues to handle build/test/deploy. Argo provides DAG scheduling, parameterized workflows, and native MinIO artifact integration.

Architecture

  • Argo Workflows controller installed in foxhunt namespace via Helm
  • Argo Events (EventSource + Sensor) for GitLab webhook → workflow submission
  • Single parameterized WorkflowTemplate covering the full training pipeline per model
  • MinIO as native Argo artifact repository (already deployed at minio.foxhunt.svc.cluster.local:9000)

Scope

In scope (Argo): hyperopt, train, evaluate — all GPU training jobs Out of scope (stays in GitLab CI): prepare, test, compile, deploy, web dashboard build

WorkflowTemplate Parameters

Parameter Default Description
model (required) Model name: dqn, ppo, tft, mamba2, tggn, tlob, liquid, kan, xlstm, diffusion
gpu-pool ci-training Node pool: ci-training (L40S) or ci-training-h100 (H100)
hyperopt-trials 5 Number of hyperopt PSO trials
hyperopt-epochs 8 Epochs per hyperopt trial
train-epochs 50 Epochs for final training with best hyperparams
symbol ES.FUT Trading symbol for data
data-dir /data/cache/futures-baseline Path to training data

Workflow Steps (per model)

fetch-binary → hyperopt → train-best → evaluate → upload-results
  1. fetch-binary: rclone copy from s3://foxhunt-binaries/training/ to workspace
  2. hyperopt: Run hyperopt_baseline_rl or hyperopt_baseline_supervised with trial/epoch params
  3. train-best: Run train_baseline_rl or train_baseline_supervised with best hyperparams from step 2
  4. evaluate: Run evaluate_baseline against test data window
  5. upload-results: Upload checkpoints + eval results to MinIO s3://foxhunt-training-results/

Binary Mapping

Models Binary (hyperopt) Binary (train)
dqn, ppo hyperopt_baseline_rl train_baseline_rl
tft, mamba2, tggn, tlob, liquid, kan, xlstm, diffusion hyperopt_baseline_supervised train_baseline_supervised

All models use evaluate_baseline for evaluation.

Pod Templates

Two pod templates matching existing runtime images:

  • gpu-training: foxhunt-training-runtime:latest (CUDA 12.4 + cuDNN + rclone + nvrtc)

    • GPU node selector: k8s.scaleway.com/pool-name: {{workflow.parameters.gpu-pool}}
    • Resources: requests 1 GPU, 32Gi memory, 8 CPU
    • tolerations: dedicated=gpu:NoSchedule
  • cpu-utility: foxhunt-runtime:latest (debian bookworm + rclone)

    • For fetch-binary and upload-results steps
    • Resources: requests 2 CPU, 4Gi memory

Argo Events Integration

  • EventSource: GitLab webhook receiver on argo-events.fxhnt.ai (or internal service)
  • Sensor: Filters push events, maps changed paths to model triggers:
    • crates/ml/src/dqn/** → trigger DQN workflow
    • crates/ml/src/trainers/ppo/** → trigger PPO workflow
    • crates/ml/src/models/tft/** → trigger TFT workflow
    • etc.
  • Manual trigger: argo submit CLI or Argo UI at argo.fxhnt.ai

DNS & Networking

  • argo.fxhnt.ai → Tailscale proxy → Argo Server (port 2746)
  • Argo Server with --auth-mode=server (no SSO needed initially)
  • Add nginx proxy block in infra/k8s/gitlab/tailscale-proxy.yaml
  • Add DNS A record in infra/modules/dns/main.tf

MinIO Artifact Repository

Argo native artifact config pointing to existing MinIO:

artifactRepositoryRef:
  configMap: artifact-repositories
  key: default-v1

ConfigMap:

s3:
  endpoint: minio.foxhunt.svc.cluster.local:9000
  insecure: true
  bucket: foxhunt-training-results
  accessKeySecret:
    name: minio-credentials
    key: access-key
  secretKeySecret:
    name: minio-credentials
    key: secret-key

Migration Plan

Phase 1: Install & Validate

  • Install Argo Workflows controller via Helm (minimal config)
  • Create artifact repository ConfigMap
  • Create WorkflowTemplate
  • Manually submit test workflow for DQN
  • Verify it runs on GPU, produces checkpoints

Phase 2: Wire Events & Parallel Run

  • Install Argo Events (EventSource + Sensor)
  • Configure GitLab webhook
  • Run both GitLab CI training AND Argo Workflows in parallel
  • Compare results, validate reliability

Phase 3: Cutover

  • Remove training jobs from .gitlab-ci.yml
  • Argo becomes sole training orchestrator
  • GitLab CI only: prepare → test → compile → deploy

Constraints

  • Keep GitLab CI working until Argo fully validated
  • Not everything runs at once — models can be triggered independently
  • Single POP2-32C-128G node for CPU work, GPU nodes scale 0→1
  • MinIO credentials already in k8s secret minio-credentials