Files
foxhunt/docs/plans/2026-02-28-s3-binary-share-implementation.md
jgrusewski 359f847026 docs: add S3 binary share implementation plan (15 tasks)
Detailed step-by-step plan covering: S3 bucket creation, CI builder
updates, 3 base image Dockerfiles, CI pipeline rewiring, 8 cache PVCs,
all service deployment YAML updates, training job template, and cleanup.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-28 17:50:37 +01:00

41 KiB
Raw Blame History

S3 Binary Share Implementation Plan

For Claude: REQUIRED SUB-SKILL: Use superpowers:executing-plans to implement this plan task-by-task.

Goal: Replace per-service Docker images with generic base images + S3-based binary distribution to eliminate the Kaniko build stage and speed up pod startup.

Architecture: CI compiles binaries and uploads them to s3://foxhunt-binaries/latest/. Pods run pre-pulled generic base images with an initContainer that fetches the binary from S3 into an emptyDir. A per-service PVC cache provides fallback if S3 is unavailable. The build stage (9 Kaniko jobs) is entirely eliminated.

Tech Stack: Scaleway Object Storage (S3), rclone, Kubernetes initContainers, Terragrunt/Terraform, GitLab CI

Design doc: docs/plans/2026-02-28-s3-binary-share-design.md


Task 1: Create S3 bucket via Terraform

Files:

  • Modify: infra/modules/object-storage/main.tf (append new resource)
  • Modify: infra/modules/object-storage/outputs.tf (append new output)

Step 1: Add the foxhunt-binaries bucket resource

In infra/modules/object-storage/main.tf, append after the sccache resource (line ~62):

resource "scaleway_object_bucket" "binaries" {
  name   = "${var.bucket_name_prefix}-binaries"
  region = var.region

  lifecycle_rule {
    enabled = true
    prefix  = "archive/"

    expiration {
      days = 30
    }
  }

  versioning {
    enabled = false
  }
}

Step 2: Add the output

In infra/modules/object-storage/outputs.tf, append:

output "binaries_bucket_name" {
  description = "Name of the binaries distribution bucket"
  value       = scaleway_object_bucket.binaries.name
}

Step 3: Validate locally

Run: cd infra/live/production/object-storage && terragrunt validate Expected: Success (no errors)

Step 4: Commit

git add infra/modules/object-storage/main.tf infra/modules/object-storage/outputs.tf
git commit -m "infra: add foxhunt-binaries S3 bucket for binary distribution"

Note: The actual terragrunt apply runs in CI (infra-apply job) or manually. The bucket name will be foxhunt-binaries (prefix foxhunt + -binaries).


Task 2: Add rclone to CI builder images

Files:

  • Modify: infra/docker/Dockerfile.ci-builder (add rclone install after sccache)
  • Modify: infra/docker/Dockerfile.ci-builder-cpu (add rclone install after sccache)

Step 1: Add rclone to Dockerfile.ci-builder

In infra/docker/Dockerfile.ci-builder, after the sccache install block (line ~44, after chmod +x /usr/local/bin/sccache), add:

# rclone for S3 binary uploads (compile → S3)
RUN curl -fsSL https://downloads.rclone.org/v1.69.1/rclone-v1.69.1-linux-amd64.zip -o /tmp/rclone.zip \
    && unzip -j /tmp/rclone.zip '*/rclone' -d /usr/local/bin/ \
    && rm /tmp/rclone.zip \
    && chmod +x /usr/local/bin/rclone

Also add rclone --version to the verify line (line ~60):

RUN rustc --version && cargo --version && git --version && protoc --version && sccache --version && make --version && nvcc --version && rclone --version

Step 2: Add rclone to Dockerfile.ci-builder-cpu

Same change in infra/docker/Dockerfile.ci-builder-cpu, after the sccache install (line ~43), add the same rclone block. Update verify line (line ~47):

RUN rustc --version && cargo --version && git --version && protoc --version && sccache --version && make --version && mold --version && rclone --version

Step 3: Commit

git add infra/docker/Dockerfile.ci-builder infra/docker/Dockerfile.ci-builder-cpu
git commit -m "infra: add rclone to CI builder images for S3 binary uploads"

Note: CI builder images are rebuilt via the build-ci-builder / build-ci-builder-cpu Kaniko jobs in the prepare stage. Trigger manually from the pipeline UI after merging.


Task 3: Create 3 generic base image Dockerfiles

Files:

  • Create: infra/docker/Dockerfile.foxhunt-runtime
  • Create: infra/docker/Dockerfile.foxhunt-training-runtime
  • Create: infra/docker/Dockerfile.foxhunt-node-builder

Step 1: Create foxhunt-runtime

Write infra/docker/Dockerfile.foxhunt-runtime:

# Generic runtime base image for all Foxhunt service pods
# Contains: rclone (S3 binary fetch), grpc_health_probe (K8s health checks)
# Binaries are fetched by initContainer at pod startup, not baked into this image.
# Rebuild: only when system deps change (libssl, rclone, grpc_health_probe)

FROM debian:bookworm-slim

RUN apt-get update && apt-get install -y --no-install-recommends \
    ca-certificates \
    libssl3 \
    curl \
    unzip \
    && curl -fsSL https://downloads.rclone.org/v1.69.1/rclone-v1.69.1-linux-amd64.zip -o /tmp/rclone.zip \
    && unzip -j /tmp/rclone.zip '*/rclone' -d /usr/local/bin/ \
    && rm /tmp/rclone.zip \
    && apt-get purge -y unzip \
    && rm -rf /var/lib/apt/lists/*

# grpc_health_probe for Kubernetes health checks
RUN curl -fsSL https://github.com/grpc-ecosystem/grpc-health-probe/releases/download/v0.4.25/grpc_health_probe-linux-amd64 \
    -o /usr/local/bin/grpc_health_probe \
    && chmod +x /usr/local/bin/grpc_health_probe

RUN groupadd -g 1000 foxhunt \
    && useradd -u 1000 -g foxhunt -m -s /bin/false foxhunt

USER foxhunt

Step 2: Create foxhunt-training-runtime

Write infra/docker/Dockerfile.foxhunt-training-runtime:

# Generic GPU runtime base image for Foxhunt training pods
# Contains: CUDA 12.4 + cuDNN + NVRTC + rclone
# Binaries are fetched by initContainer at pod startup, not baked into this image.
# Rebuild: only on CUDA version bump

FROM nvidia/cuda:12.4.1-cudnn-runtime-ubuntu22.04

ENV DEBIAN_FRONTEND=noninteractive

RUN apt-get update && apt-get install -y --no-install-recommends \
    ca-certificates \
    libssl3 \
    curl \
    unzip \
    cuda-nvrtc-12-4 \
    && curl -fsSL https://downloads.rclone.org/v1.69.1/rclone-v1.69.1-linux-amd64.zip -o /tmp/rclone.zip \
    && unzip -j /tmp/rclone.zip '*/rclone' -d /usr/local/bin/ \
    && rm /tmp/rclone.zip \
    && apt-get purge -y unzip \
    && rm -rf /var/lib/apt/lists/*

RUN groupadd -g 1000 foxhunt \
    && useradd -u 1000 -g foxhunt -m -s /bin/false foxhunt

ENV NVIDIA_VISIBLE_DEVICES=all
ENV NVIDIA_DRIVER_CAPABILITIES=compute,utility

USER foxhunt

Step 3: Create foxhunt-node-builder

Write infra/docker/Dockerfile.foxhunt-node-builder:

# Node.js CI image for web-dashboard Vite builds + S3 upload
# Contains: Node 22 + rclone
# Rebuild: only on Node major version bump

FROM node:22-slim

RUN apt-get update && apt-get install -y --no-install-recommends \
    curl \
    unzip \
    && curl -fsSL https://downloads.rclone.org/v1.69.1/rclone-v1.69.1-linux-amd64.zip -o /tmp/rclone.zip \
    && unzip -j /tmp/rclone.zip '*/rclone' -d /usr/local/bin/ \
    && rm /tmp/rclone.zip \
    && apt-get purge -y unzip \
    && rm -rf /var/lib/apt/lists/*

Step 4: Commit

git add infra/docker/Dockerfile.foxhunt-runtime \
        infra/docker/Dockerfile.foxhunt-training-runtime \
        infra/docker/Dockerfile.foxhunt-node-builder
git commit -m "infra: add 3 generic base images for S3 binary share"

Task 4: Add base image build jobs to CI prepare stage

Files:

  • Modify: .gitlab-ci.yml (add 3 Kaniko jobs in the prepare stage, after existing builder jobs)

Step 1: Add build-foxhunt-runtime job

In .gitlab-ci.yml, after the build-infra-runner job (around line 188), add:

# --------------------------------------------------------------------------
# Stage 0d: Build generic runtime base images → push to Scaleway CR
# --------------------------------------------------------------------------
build-foxhunt-runtime:
  stage: prepare
  image:
    name: gcr.io/kaniko-project/executor:debug
    entrypoint: [""]
  tags:
    - kapsule
    - docker
  rules:
    - if: $CI_PIPELINE_SOURCE == "push"
      changes:
        - infra/docker/Dockerfile.foxhunt-runtime
      when: on_success
    - when: manual
      allow_failure: true
  before_script:
    - mkdir -p /kaniko/.docker
    - >-
      echo "{\"auths\":{
      \"rg.fr-par.scw.cloud\":{\"username\":\"nologin\",\"password\":\"${SCW_SECRET_KEY}\"},
      \"https://index.docker.io/v1/\":{\"username\":\"${DOCKERHUB_USERNAME}\",\"password\":\"${DOCKERHUB_TOKEN}\"}
      }}" > /kaniko/.docker/config.json
  script:
    - /kaniko/executor
      --context "${CI_PROJECT_DIR}"
      --dockerfile "${CI_PROJECT_DIR}/infra/docker/Dockerfile.foxhunt-runtime"
      --cache=true --cache-repo="${REGISTRY}/cache"
      --destination "${REGISTRY}/foxhunt-runtime:latest"

build-foxhunt-training-runtime:
  stage: prepare
  image:
    name: gcr.io/kaniko-project/executor:debug
    entrypoint: [""]
  tags:
    - kapsule
    - docker
  rules:
    - if: $CI_PIPELINE_SOURCE == "push"
      changes:
        - infra/docker/Dockerfile.foxhunt-training-runtime
      when: on_success
    - when: manual
      allow_failure: true
  before_script:
    - mkdir -p /kaniko/.docker
    - >-
      echo "{\"auths\":{
      \"rg.fr-par.scw.cloud\":{\"username\":\"nologin\",\"password\":\"${SCW_SECRET_KEY}\"},
      \"https://index.docker.io/v1/\":{\"username\":\"${DOCKERHUB_USERNAME}\",\"password\":\"${DOCKERHUB_TOKEN}\"}
      }}" > /kaniko/.docker/config.json
  script:
    - /kaniko/executor
      --context "${CI_PROJECT_DIR}"
      --dockerfile "${CI_PROJECT_DIR}/infra/docker/Dockerfile.foxhunt-training-runtime"
      --cache=true --cache-repo="${REGISTRY}/cache"
      --destination "${REGISTRY}/foxhunt-training-runtime:latest"

build-foxhunt-node-builder:
  stage: prepare
  image:
    name: gcr.io/kaniko-project/executor:debug
    entrypoint: [""]
  tags:
    - kapsule
    - docker
  rules:
    - if: $CI_PIPELINE_SOURCE == "push"
      changes:
        - infra/docker/Dockerfile.foxhunt-node-builder
      when: on_success
    - when: manual
      allow_failure: true
  before_script:
    - mkdir -p /kaniko/.docker
    - >-
      echo "{\"auths\":{
      \"rg.fr-par.scw.cloud\":{\"username\":\"nologin\",\"password\":\"${SCW_SECRET_KEY}\"},
      \"https://index.docker.io/v1/\":{\"username\":\"${DOCKERHUB_USERNAME}\",\"password\":\"${DOCKERHUB_TOKEN}\"}
      }}" > /kaniko/.docker/config.json
  script:
    - /kaniko/executor
      --context "${CI_PROJECT_DIR}"
      --dockerfile "${CI_PROJECT_DIR}/infra/docker/Dockerfile.foxhunt-node-builder"
      --cache=true --cache-repo="${REGISTRY}/cache"
      --destination "${REGISTRY}/foxhunt-node-builder:latest"

Step 2: Also add these image names as variables at the top of .gitlab-ci.yml

After the INFRA_RUNNER_IMAGE variable (around line 59), add:

  # Generic runtime images for S3 binary share (service pods + training pods + node builder)
  RUNTIME_IMAGE: rg.fr-par.scw.cloud/foxhunt-ci/foxhunt-runtime:latest
  TRAINING_RUNTIME_IMAGE: rg.fr-par.scw.cloud/foxhunt-ci/foxhunt-training-runtime:latest
  NODE_BUILDER_IMAGE: rg.fr-par.scw.cloud/foxhunt-ci/foxhunt-node-builder:latest

Step 3: Commit

git add .gitlab-ci.yml
git commit -m "ci: add build jobs for 3 generic base images in prepare stage"

Task 5: Modify compile-services to upload to S3

Files:

  • Modify: .gitlab-ci.ymlcompile-services job

Step 1: Replace the artifacts: block with S3 upload

In the compile-services job, replace the artifacts: block (the final 3 lines, around line 326-328) and append upload steps to the script: section.

The entire compile-services job should become:

compile-services:
  extends: .rust-base-cpu
  stage: compile
  needs: []
  rules:
    - if: $CI_COMMIT_BRANCH == "main" && $CI_PIPELINE_SOURCE == "push"
      changes:
        - Cargo.toml
        - Cargo.lock
        - crates/**
        - bin/**
        - services/**
        - infra/docker/Dockerfile.*
    - if: $CI_PIPELINE_SOURCE == "web" || $CI_PIPELINE_SOURCE == "api"
  script:
    - mkdir -p "$SCCACHE_DIR" && sccache --zero-stats || true
    - PROFILE=${DEV_RELEASE:+dev-release}; PROFILE=${PROFILE:-release}
    - echo "Building services with --profile $PROFILE (CPU-only, no CUDA)"
    - TARGET_DIR="target/$PROFILE"
    - cargo build --profile $PROFILE
        -p trading_service
        -p api_gateway
        -p broker_gateway_service
        -p ml_training_service
        -p backtesting_service
        -p trading_agent_service
        -p data_acquisition_service
        -p web-gateway
    - sccache --show-stats || true
    - mkdir -p build-out/services
    - |
      for bin in trading_service api_gateway broker_gateway_service ml_training_service backtesting_service trading_agent_service data_acquisition_service web-gateway; do
        cp $TARGET_DIR/$bin build-out/services/
        strip build-out/services/$bin
      done
    - ls -lh build-out/services/
    # Upload stripped binaries to S3
    - export RCLONE_S3_PROVIDER=Scaleway RCLONE_S3_ENDPOINT=s3.fr-par.scw.cloud RCLONE_S3_REGION=fr-par
    - export RCLONE_S3_ACCESS_KEY_ID=$SCW_ACCESS_KEY RCLONE_S3_SECRET_ACCESS_KEY=$SCW_SECRET_KEY
    - |
      for bin in trading_service api_gateway broker_gateway_service ml_training_service backtesting_service trading_agent_service data_acquisition_service web-gateway; do
        rclone copyto build-out/services/$bin :s3:foxhunt-binaries/latest/services/$bin
      done
    - rclone copy build-out/services/ :s3:foxhunt-binaries/archive/${CI_COMMIT_SHA}/services/
    - echo "Service binaries uploaded to S3"

Note: No more artifacts: block. The :s3: prefix is rclone's backend syntax when using env-var config (no named remote).

Step 2: Commit

git add .gitlab-ci.yml
git commit -m "ci: compile-services uploads to S3 instead of GitLab artifacts"

Task 6: Modify compile-training to upload to S3

Files:

  • Modify: .gitlab-ci.ymlcompile-training job

Step 1: Replace artifacts with S3 upload

The compile-training job currently ends with artifacts: (around line 383-384). Replace the entire job to match this:

compile-training:
  extends: .rust-base-cpu
  stage: compile
  image: ${CI_BUILDER_IMAGE}
  needs: [compile-services]
  variables:
    CUDA_COMPUTE_CAP: "89"
  rules:
    - if: $CI_COMMIT_BRANCH == "main" && $CI_PIPELINE_SOURCE == "push"
      changes:
        - Cargo.toml
        - Cargo.lock
        - crates/ml/**
        - crates/trading_engine/**
        - crates/common/**
        - crates/risk/**
        - crates/data/**
        - crates/config/**
        - crates/storage/**
        - services/training_uploader/**
    - if: $CI_PIPELINE_SOURCE == "web" || $CI_PIPELINE_SOURCE == "api"
  script:
    - echo "CUDA_COMPUTE_CAP=$CUDA_COMPUTE_CAP"
    - export SCCACHE_DIR="/mnt/sccache/sm_${CUDA_COMPUTE_CAP}"
    - mkdir -p "$SCCACHE_DIR" && sccache --zero-stats || true
    - PROFILE=${DEV_RELEASE:+dev-release}; PROFILE=${PROFILE:-release}
    - echo "Building training binaries with --profile $PROFILE (CUDA stubs)"
    - TARGET_DIR="target/$PROFILE"
    - cargo build --profile $PROFILE -p ml --features ml/cuda
        --example train_baseline_rl
        --example train_baseline_supervised
        --example evaluate_baseline
        --example evaluate_supervised
        --example hyperopt_baseline_rl
        --example hyperopt_baseline_supervised
    - cargo build --profile $PROFILE -p training_uploader
    - sccache --show-stats || true
    - mkdir -p build-out/training
    - |
      for bin in train_baseline_rl train_baseline_supervised evaluate_baseline evaluate_supervised hyperopt_baseline_rl hyperopt_baseline_supervised; do
        cp $TARGET_DIR/examples/$bin build-out/training/
        strip build-out/training/$bin
      done
      cp $TARGET_DIR/training_uploader build-out/training/
      strip build-out/training/training_uploader
    - ls -lh build-out/training/
    # Upload stripped binaries to S3
    - export RCLONE_S3_PROVIDER=Scaleway RCLONE_S3_ENDPOINT=s3.fr-par.scw.cloud RCLONE_S3_REGION=fr-par
    - export RCLONE_S3_ACCESS_KEY_ID=$SCW_ACCESS_KEY RCLONE_S3_SECRET_ACCESS_KEY=$SCW_SECRET_KEY
    - |
      for bin in train_baseline_rl train_baseline_supervised evaluate_baseline evaluate_supervised hyperopt_baseline_rl hyperopt_baseline_supervised training_uploader; do
        rclone copyto build-out/training/$bin :s3:foxhunt-binaries/latest/training/$bin
      done
    - rclone copy build-out/training/ :s3:foxhunt-binaries/archive/${CI_COMMIT_SHA}/training/
    - echo "Training binaries uploaded to S3"

Step 2: Commit

git add .gitlab-ci.yml
git commit -m "ci: compile-training uploads to S3 instead of GitLab artifacts"

Task 7: Add build-web-dashboard, write-manifest, and deploy CI jobs

Files:

  • Modify: .gitlab-ci.yml — add 3 new jobs

Step 1: Add build-web-dashboard job

After the compile-training job, add:

# --------------------------------------------------------------------------
# Stage 3c: Build web-dashboard static assets (Vite)
# --------------------------------------------------------------------------
build-web-dashboard:
  stage: compile
  image: ${NODE_BUILDER_IMAGE}
  needs: []
  tags:
    - kapsule
    - docker
  variables:
    KUBERNETES_CPU_REQUEST: "2000m"
    KUBERNETES_CPU_LIMIT: "4000m"
    KUBERNETES_MEMORY_REQUEST: "2Gi"
    KUBERNETES_MEMORY_LIMIT: "4Gi"
  rules:
    - if: $CI_COMMIT_BRANCH == "main" && $CI_PIPELINE_SOURCE == "push"
      changes:
        - web-dashboard/**
    - if: $CI_PIPELINE_SOURCE == "web" || $CI_PIPELINE_SOURCE == "api"
  script:
    - cd web-dashboard
    - npm ci
    - npm run build
    - export RCLONE_S3_PROVIDER=Scaleway RCLONE_S3_ENDPOINT=s3.fr-par.scw.cloud RCLONE_S3_REGION=fr-par
    - export RCLONE_S3_ACCESS_KEY_ID=$SCW_ACCESS_KEY RCLONE_S3_SECRET_ACCESS_KEY=$SCW_SECRET_KEY
    - rclone sync dist/ :s3:foxhunt-binaries/latest/web-dashboard/dist/
    - echo "Web dashboard assets uploaded to S3"

# --------------------------------------------------------------------------
# Stage 3d: Write manifest after all compiles complete
# --------------------------------------------------------------------------
write-manifest:
  stage: compile
  image: ${RUNTIME_IMAGE}
  needs:
    - job: compile-services
    - job: compile-training
    - job: build-web-dashboard
      optional: true
  tags:
    - kapsule
    - docker
  variables:
    KUBERNETES_CPU_REQUEST: "100m"
    KUBERNETES_MEMORY_REQUEST: "64Mi"
  rules:
    - if: $CI_COMMIT_BRANCH == "main" && $CI_PIPELINE_SOURCE == "push"
    - if: $CI_PIPELINE_SOURCE == "web" || $CI_PIPELINE_SOURCE == "api"
  script:
    - export RCLONE_S3_PROVIDER=Scaleway RCLONE_S3_ENDPOINT=s3.fr-par.scw.cloud RCLONE_S3_REGION=fr-par
    - export RCLONE_S3_ACCESS_KEY_ID=$SCW_ACCESS_KEY RCLONE_S3_SECRET_ACCESS_KEY=$SCW_SECRET_KEY
    - echo "{\"sha\":\"${CI_COMMIT_SHA}\",\"ts\":\"$(date -Iseconds)\"}" | rclone rcat :s3:foxhunt-binaries/latest/MANIFEST.json
    - echo "Manifest written for ${CI_COMMIT_SHA}"

Step 2: Add deploy job (at the end of the file, in the deploy stage)

After the training jobs section, add:

# --------------------------------------------------------------------------
# Stage 5: Deploy — rolling restart of all service deployments
# --------------------------------------------------------------------------
deploy-services:
  stage: deploy
  image:
    name: bitnami/kubectl:latest
    entrypoint: [""]
  tags:
    - kapsule
    - docker
  needs:
    - job: write-manifest
  variables:
    KUBERNETES_CPU_REQUEST: "100m"
    KUBERNETES_MEMORY_REQUEST: "64Mi"
  rules:
    - if: $CI_COMMIT_BRANCH == "main" && $CI_PIPELINE_SOURCE == "push"
      when: manual
      allow_failure: true
    - if: $CI_PIPELINE_SOURCE == "web" || $CI_PIPELINE_SOURCE == "api"
      when: manual
      allow_failure: true
  script:
    - echo "Deploying ${CI_COMMIT_SHA}..."
    - |
      for svc in trading-service api-gateway broker-gateway \
                 ml-training-service backtesting-service \
                 trading-agent-service data-acquisition-service web-gateway; do
        kubectl -n foxhunt rollout restart deployment/$svc
        echo "Restarted $svc"
      done
    - |
      for svc in trading-service api-gateway broker-gateway \
                 ml-training-service backtesting-service \
                 trading-agent-service data-acquisition-service web-gateway; do
        kubectl -n foxhunt rollout status deployment/$svc --timeout=120s
      done
    - echo "All services deployed"

Step 3: Commit

git add .gitlab-ci.yml
git commit -m "ci: add build-web-dashboard, write-manifest, and deploy-services jobs"

Task 8: Delete all Kaniko build jobs from CI

Files:

  • Modify: .gitlab-ci.yml — remove the entire build stage section

Step 1: Delete the following blocks

Remove these sections from .gitlab-ci.yml:

  1. .kaniko-service-base template (the stage: build template starting with "Kaniko base for service images")
  2. .kaniko-training-base template
  3. build-trading-service job
  4. build-api-gateway job
  5. build-broker-gateway job
  6. build-ml-training job
  7. build-backtesting job
  8. build-trading-agent job
  9. build-data-acquisition job
  10. build-web-gateway job
  11. build-training job

These are roughly lines 390-549 in the current .gitlab-ci.yml.

Also remove build from the stages: list (line 31). The stages should become:

stages:
  - prepare
  - test
  - compile
  - train
  - deploy

Step 2: Update training job needs:

All training jobs currently have needs: [build-training] (with optional: true). Change them to depend on compile-training instead:

Find all instances of:

  needs:
    - job: build-training
      optional: true

Replace with:

  needs:
    - job: compile-training
      optional: true

This applies to: train-validate-rl, train-validate-tft, train-validate-mamba2, train-validate-liquid, train-validate-tggn, train-validate-tlob, train-validate-kan, train-validate-xlstm, train-validate-diffusion, hyperopt-ppo, hyperopt-dqn, hyperopt-tft, hyperopt-mamba2, hyperopt-liquid, hyperopt-tggn.

Step 3: Update training job image:

All training jobs currently use image: ${REGISTRY}/training:${CI_COMMIT_SHA}. Change to the generic runtime:

Find: image: ${REGISTRY}/training:${CI_COMMIT_SHA} Replace with: image: ${TRAINING_RUNTIME_IMAGE}

This applies to .train-rl-base and .train-validate-base templates.

Step 4: Update training job scripts — fetch binaries first

In .train-rl-base and .train-validate-base, add binary fetch to before_script::

Replace the existing before_script: in .train-rl-base:

  before_script:
    - export LD_LIBRARY_PATH=$(echo "$LD_LIBRARY_PATH" | tr ':' '\n' | grep -v stubs | tr '\n' ':' | sed 's/:$//')
    - nvidia-smi
    - mkdir -p ${CI_PROJECT_DIR}/output

With:

  before_script:
    - export LD_LIBRARY_PATH=$(echo "$LD_LIBRARY_PATH" | tr ':' '\n' | grep -v stubs | tr '\n' ':' | sed 's/:$//')
    - nvidia-smi
    - mkdir -p ${CI_PROJECT_DIR}/output
    # Fetch training binaries from S3
    - export RCLONE_S3_PROVIDER=Scaleway RCLONE_S3_ENDPOINT=s3.fr-par.scw.cloud RCLONE_S3_REGION=fr-par
    - export RCLONE_S3_ACCESS_KEY_ID=$SCW_ACCESS_KEY RCLONE_S3_SECRET_ACCESS_KEY=$SCW_SECRET_KEY
    - rclone sync :s3:foxhunt-binaries/latest/training/ /usr/local/bin/ --include "*.{rl,supervised,baseline,uploader}*" 2>/dev/null || rclone sync :s3:foxhunt-binaries/latest/training/ /usr/local/bin/
    - chmod +x /usr/local/bin/train_baseline_rl /usr/local/bin/train_baseline_supervised /usr/local/bin/evaluate_baseline /usr/local/bin/evaluate_supervised /usr/local/bin/hyperopt_baseline_rl /usr/local/bin/hyperopt_baseline_supervised /usr/local/bin/training_uploader

Same change for .train-validate-base.

Step 5: Commit

git add .gitlab-ci.yml
git commit -m "ci: delete all Kaniko build jobs, switch training to S3 binary fetch"

Task 9: Create binary cache PVCs

Files:

  • Create: infra/k8s/storage/binary-cache-pvcs.yaml

Step 1: Write the PVC definitions

Write infra/k8s/storage/binary-cache-pvcs.yaml:

# Per-service binary cache PVCs — fallback when S3 is unavailable
# Each service gets its own 500Mi PVC (ReadWriteOnce, scw-bssd)
# Total: 8 × 500Mi = 4Gi, ~€0.08/month
---
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
  name: binary-cache-trading-service
  namespace: foxhunt
  labels:
    app.kubernetes.io/name: binary-cache
    app.kubernetes.io/part-of: foxhunt
spec:
  accessModes: [ReadWriteOnce]
  storageClassName: scw-bssd
  resources:
    requests:
      storage: 500Mi
---
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
  name: binary-cache-api-gateway
  namespace: foxhunt
  labels:
    app.kubernetes.io/name: binary-cache
    app.kubernetes.io/part-of: foxhunt
spec:
  accessModes: [ReadWriteOnce]
  storageClassName: scw-bssd
  resources:
    requests:
      storage: 500Mi
---
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
  name: binary-cache-broker-gateway
  namespace: foxhunt
  labels:
    app.kubernetes.io/name: binary-cache
    app.kubernetes.io/part-of: foxhunt
spec:
  accessModes: [ReadWriteOnce]
  storageClassName: scw-bssd
  resources:
    requests:
      storage: 500Mi
---
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
  name: binary-cache-ml-training-service
  namespace: foxhunt
  labels:
    app.kubernetes.io/name: binary-cache
    app.kubernetes.io/part-of: foxhunt
spec:
  accessModes: [ReadWriteOnce]
  storageClassName: scw-bssd
  resources:
    requests:
      storage: 500Mi
---
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
  name: binary-cache-backtesting-service
  namespace: foxhunt
  labels:
    app.kubernetes.io/name: binary-cache
    app.kubernetes.io/part-of: foxhunt
spec:
  accessModes: [ReadWriteOnce]
  storageClassName: scw-bssd
  resources:
    requests:
      storage: 500Mi
---
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
  name: binary-cache-trading-agent-service
  namespace: foxhunt
  labels:
    app.kubernetes.io/name: binary-cache
    app.kubernetes.io/part-of: foxhunt
spec:
  accessModes: [ReadWriteOnce]
  storageClassName: scw-bssd
  resources:
    requests:
      storage: 500Mi
---
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
  name: binary-cache-data-acquisition-service
  namespace: foxhunt
  labels:
    app.kubernetes.io/name: binary-cache
    app.kubernetes.io/part-of: foxhunt
spec:
  accessModes: [ReadWriteOnce]
  storageClassName: scw-bssd
  resources:
    requests:
      storage: 500Mi
---
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
  name: binary-cache-web-gateway
  namespace: foxhunt
  labels:
    app.kubernetes.io/name: binary-cache
    app.kubernetes.io/part-of: foxhunt
spec:
  accessModes: [ReadWriteOnce]
  storageClassName: scw-bssd
  resources:
    requests:
      storage: 1Gi

Note: web-gateway gets 1Gi because it also caches the dashboard static assets (~5MB).

Step 2: Commit

git add infra/k8s/storage/binary-cache-pvcs.yaml
git commit -m "infra: add 8 binary cache PVCs for S3 fallback during trading"

Apply: kubectl apply -f infra/k8s/storage/binary-cache-pvcs.yaml


Task 10: Update all 8 service deployment YAMLs

This is the largest task. Each deployment gets:

  1. Image changed from per-service GitLab registry → generic foxhunt-runtime from SCW registry
  2. imagePullSecrets changed from gitlab-registryscw-registry
  3. initContainers added for S3 binary fetch with PVC cache fallback
  4. command added to main container to run from /binaries/
  5. volumes updated with binaries emptyDir + binary-cache PVC
  6. Deployment strategy set to maxSurge: 1, maxUnavailable: 0

Files:

  • Modify: infra/k8s/services/trading-service.yaml
  • Modify: infra/k8s/services/api-gateway.yaml
  • Modify: infra/k8s/services/broker-gateway.yaml
  • Modify: infra/k8s/services/ml-training-service.yaml
  • Modify: infra/k8s/services/backtesting-service.yaml
  • Modify: infra/k8s/services/trading-agent-service.yaml
  • Modify: infra/k8s/services/data-acquisition-service.yaml
  • Modify: infra/k8s/services/web-gateway.yaml

The pattern for each service (example: trading-service):

For each deployment YAML, make these changes:

  1. Add deployment strategy after spec.replicas:
  strategy:
    type: RollingUpdate
    rollingUpdate:
      maxSurge: 1
      maxUnavailable: 0
  1. Change imagePullSecrets from gitlab-registry to scw-registry:
      imagePullSecrets:
        - name: scw-registry
  1. Add initContainers before containers:
      initContainers:
        - name: fetch-binary
          image: rg.fr-par.scw.cloud/foxhunt-ci/foxhunt-runtime:latest
          command: ["/bin/sh", "-c"]
          args:
            - |
              set -e
              SERVICE=trading_service
              if rclone copyto :s3:foxhunt-binaries/latest/services/$SERVICE /binaries/$SERVICE 2>/dev/null; then
                chmod +x /binaries/$SERVICE
                cp /binaries/$SERVICE /cache/$SERVICE
                echo "Fetched $SERVICE from S3 ($(stat -c%s /binaries/$SERVICE) bytes)"
              elif [ -f /cache/$SERVICE ]; then
                cp /cache/$SERVICE /binaries/$SERVICE
                chmod +x /binaries/$SERVICE
                echo "WARNING: S3 unavailable, using cached $SERVICE"
              else
                echo "FATAL: No binary available for $SERVICE (S3 down + empty cache)"
                exit 1
              fi
          env:
            - name: RCLONE_S3_PROVIDER
              value: Scaleway
            - name: RCLONE_S3_ENDPOINT
              value: s3.fr-par.scw.cloud
            - name: RCLONE_S3_REGION
              value: fr-par
            - name: RCLONE_S3_ACCESS_KEY_ID
              valueFrom:
                secretKeyRef:
                  name: s3-credentials
                  key: access-key
            - name: RCLONE_S3_SECRET_ACCESS_KEY
              valueFrom:
                secretKeyRef:
                  name: s3-credentials
                  key: secret-key
          volumeMounts:
            - name: binaries
              mountPath: /binaries
            - name: binary-cache
              mountPath: /cache
          resources:
            requests:
              cpu: 100m
              memory: 64Mi
            limits:
              cpu: 500m
              memory: 128Mi
  1. Change container image from git.fxhnt.ai:5050/foxhunt/foxhunt/<service>:latest to:
          image: rg.fr-par.scw.cloud/foxhunt-ci/foxhunt-runtime:latest
  1. Add command to the main container:
          command: ["/binaries/trading_service"]
  1. Add volume mounts to the main container:
          volumeMounts:
            - name: binaries
              mountPath: /binaries
              readOnly: true

(Keep existing volumeMounts like logs, tls-certs)

  1. Add volumes to the pod spec:
      volumes:
        - name: binaries
          emptyDir:
            sizeLimit: 200Mi
        - name: binary-cache
          persistentVolumeClaim:
            claimName: binary-cache-trading-service

(Keep existing volumes like logs, tls-certs)

Per-service binary names and PVC names:

Deployment Binary name PVC name Special
trading-service trading_service binary-cache-trading-service
api-gateway api_gateway binary-cache-api-gateway
broker-gateway broker_gateway_service binary-cache-broker-gateway
ml-training-service ml_training_service binary-cache-ml-training-service Has command: ["./ml_training_service", "serve"] → change to command: ["/binaries/ml_training_service", "serve"]
backtesting-service backtesting_service binary-cache-backtesting-service
trading-agent-service trading_agent_service binary-cache-trading-agent-service
data-acquisition-service data_acquisition_service binary-cache-data-acquisition-service Uses scw-registry already
web-gateway web-gateway binary-cache-web-gateway See below

web-gateway special case:

The initContainer fetches both the binary AND dashboard static assets:

      initContainers:
        - name: fetch-binary
          image: rg.fr-par.scw.cloud/foxhunt-ci/foxhunt-runtime:latest
          command: ["/bin/sh", "-c"]
          args:
            - |
              set -e
              SERVICE=web-gateway
              if rclone copyto :s3:foxhunt-binaries/latest/services/$SERVICE /binaries/$SERVICE 2>/dev/null; then
                chmod +x /binaries/$SERVICE
                cp /binaries/$SERVICE /cache/$SERVICE
                rclone sync :s3:foxhunt-binaries/latest/web-dashboard/dist/ /binaries/static/
                cp -r /binaries/static /cache/static 2>/dev/null || true
                echo "Fetched $SERVICE + static assets from S3"
              elif [ -f /cache/$SERVICE ]; then
                cp /cache/$SERVICE /binaries/$SERVICE
                chmod +x /binaries/$SERVICE
                cp -r /cache/static /binaries/static 2>/dev/null || true
                echo "WARNING: S3 unavailable, using cached $SERVICE"
              else
                echo "FATAL: No binary available for $SERVICE"
                exit 1
              fi
          env:
            # ... same RCLONE_S3_* env vars ...
          volumeMounts:
            - name: binaries
              mountPath: /binaries
            - name: binary-cache
              mountPath: /cache

The web-gateway container gets:

          command: ["/binaries/web-gateway"]

And the STATIC_DIR environment variable (if the binary supports it, otherwise it defaults to ./static/ which would be /binaries/static/ since the workdir would need adjustment — check the binary).

Important: The web-gateway binary currently serves static files from ./static/ (relative to WORKDIR /app). With the new layout, static files are at /binaries/static/. Either: (a) Set env STATIC_DIR=/binaries/static/ if the binary supports it, or (b) Symlink: add ln -sf /binaries/static /app/static to the initContainer. Check the Rust source (services/web-gateway/src/main.rs or similar) to determine the static path config.

Step 1-8: Update each YAML file following the pattern above.

Step 9: Commit

git add infra/k8s/services/trading-service.yaml \
        infra/k8s/services/api-gateway.yaml \
        infra/k8s/services/broker-gateway.yaml \
        infra/k8s/services/ml-training-service.yaml \
        infra/k8s/services/backtesting-service.yaml \
        infra/k8s/services/trading-agent-service.yaml \
        infra/k8s/services/data-acquisition-service.yaml \
        infra/k8s/services/web-gateway.yaml
git commit -m "infra: update all service deployments with S3 binary fetch initContainer"

Task 11: Update GPU overlay deployments

Files:

  • Modify: infra/k8s/services/trading-service-gpu.yaml
  • Modify: infra/k8s/services/ml-training-service-gpu.yaml

Apply the same changes as Task 10 to both GPU overlay files. The only differences from the base manifests are nodeSelector (gpu-inference), tolerations, and resource requests (GPU).

Step 1: Commit

git add infra/k8s/services/trading-service-gpu.yaml \
        infra/k8s/services/ml-training-service-gpu.yaml
git commit -m "infra: update GPU overlay deployments with S3 binary fetch"

Task 12: Update training job template

Files:

  • Modify: infra/k8s/training/job-template.yaml

Step 1: Update the job template

Change the image from rg.fr-par.scw.cloud/foxhunt-ci/training:latest to rg.fr-par.scw.cloud/foxhunt-ci/foxhunt-training-runtime:latest.

Add an initContainer to fetch training binaries. The training job doesn't need a cache PVC (batch jobs can retry).

Replace the initContainers: section (currently the uploader sidecar) to also include binary fetch, and update the main container to use fetched binaries:

      initContainers:
        # Fetch training binaries from S3
        - name: fetch-binaries
          image: rg.fr-par.scw.cloud/foxhunt-ci/foxhunt-training-runtime:latest
          command: ["/bin/sh", "-c"]
          args:
            - |
              set -e
              export RCLONE_S3_PROVIDER=Scaleway RCLONE_S3_ENDPOINT=s3.fr-par.scw.cloud RCLONE_S3_REGION=fr-par
              rclone sync :s3:foxhunt-binaries/latest/training/ /binaries/
              chmod +x /binaries/*
              ls -lh /binaries/
              echo "Fetched training binaries from S3"
          env:
            - name: RCLONE_S3_ACCESS_KEY_ID
              valueFrom:
                secretKeyRef:
                  name: s3-credentials
                  key: access-key
            - name: RCLONE_S3_SECRET_ACCESS_KEY
              valueFrom:
                secretKeyRef:
                  name: s3-credentials
                  key: secret-key
          volumeMounts:
            - name: binaries
              mountPath: /binaries
          resources:
            requests:
              cpu: 100m
              memory: 64Mi
            limits:
              cpu: 500m
              memory: 256Mi
        # Uploader sidecar (native sidecar, restartPolicy: Always)
        - name: uploader
          image: rg.fr-par.scw.cloud/foxhunt-ci/foxhunt-training-runtime:latest
          restartPolicy: Always
          command: ["/binaries/training_uploader"]
          # ... rest of uploader config unchanged, but add binaries mount ...
          volumeMounts:
            - name: output
              mountPath: /output
              readOnly: true
            - name: binaries
              mountPath: /binaries
              readOnly: true

Update main container:

      containers:
        - name: training
          image: rg.fr-par.scw.cloud/foxhunt-ci/foxhunt-training-runtime:latest
          command: ["/binaries/$(TRAINING_BINARY)"]
          volumeMounts:
            - name: training-data
              mountPath: /data
              readOnly: true
            - name: output
              mountPath: /output
            - name: binaries
              mountPath: /binaries
              readOnly: true

Add the binaries volume:

      volumes:
        - name: training-data
          persistentVolumeClaim:
            claimName: training-data-pvc
        - name: output
          emptyDir:
            sizeLimit: 2Gi
        - name: binaries
          emptyDir:
            sizeLimit: 500Mi

Step 2: Commit

git add infra/k8s/training/job-template.yaml
git commit -m "infra: update training job template with S3 binary fetch"

Task 13: Update image-prepuller DaemonSet

Files:

  • Modify: infra/k8s/training/image-prepuller.yaml

Step 1: Change the init container image

Replace:

        - name: pull-training
          image: rg.fr-par.scw.cloud/foxhunt-ci/training:latest
          command: ["echo", "training image pulled"]

With:

        - name: pull-training-runtime
          image: rg.fr-par.scw.cloud/foxhunt-ci/foxhunt-training-runtime:latest
          command: ["echo", "foxhunt-training-runtime image pulled"]

Step 2: Commit

git add infra/k8s/training/image-prepuller.yaml
git commit -m "infra: update image-prepuller to cache foxhunt-training-runtime"

Task 14: Delete old Dockerfiles

Files:

  • Delete: infra/docker/Dockerfile.runtime
  • Delete: infra/docker/Dockerfile.training
  • Delete: infra/docker/Dockerfile.web-gateway-runtime
  • Delete: infra/docker/Dockerfile.web-gateway
  • Delete: infra/docker/Dockerfile.service

Step 1: Remove the files

git rm infra/docker/Dockerfile.runtime \
       infra/docker/Dockerfile.training \
       infra/docker/Dockerfile.web-gateway-runtime \
       infra/docker/Dockerfile.web-gateway \
       infra/docker/Dockerfile.service

Step 2: Commit

git commit -m "infra: delete 5 obsolete Dockerfiles replaced by S3 binary share"

Task 15: Verify and seed — first manual pipeline run

This is a manual verification task, not automated.

Step 1: Build base images first

Trigger the 3 new base image builds from GitLab pipeline UI:

  • build-foxhunt-runtime (manual trigger)
  • build-foxhunt-training-runtime (manual trigger)
  • build-foxhunt-node-builder (manual trigger)

Also trigger CI builder rebuilds (they now include rclone):

  • build-ci-builder (manual trigger)
  • build-ci-builder-cpu (manual trigger)

Step 2: Apply the S3 bucket via Terragrunt

cd infra/live/production/object-storage
terragrunt apply

Step 3: Apply PVCs

kubectl apply -f infra/k8s/storage/binary-cache-pvcs.yaml

Step 4: Run a full pipeline

Push to main — the compile jobs should upload to S3, write-manifest should succeed, and deploy-services should trigger rolling restarts.

Step 5: Verify pods

kubectl -n foxhunt get pods
kubectl -n foxhunt logs deployment/trading-service -c fetch-binary

Step 6: Apply service manifests

kubectl apply -f infra/k8s/services/
kubectl apply -f infra/k8s/training/image-prepuller.yaml

Execution Order Summary

Task Dependencies Can Parallelize
1. S3 bucket TF None Yes
2. rclone in CI builders None Yes (with 1)
3. Base image Dockerfiles None Yes (with 1, 2)
4. Base image CI build jobs 3 No
5. compile-services S3 upload 2 Yes (with 6, 7)
6. compile-training S3 upload 2 Yes (with 5, 7)
7. New CI jobs (dashboard, manifest, deploy) 3 Yes (with 5, 6)
8. Delete Kaniko build jobs 5, 6, 7 No
9. Binary cache PVCs None Yes (with anything)
10. Service deployment YAMLs 3, 9 Yes (with 11, 12)
11. GPU overlay YAMLs 10 Yes (with 12)
12. Training job template 3 Yes (with 10, 11)
13. Image-prepuller update 3 Yes
14. Delete old Dockerfiles 8, 10 No
15. Manual verify + seed All above No