Detailed step-by-step plan covering: S3 bucket creation, CI builder updates, 3 base image Dockerfiles, CI pipeline rewiring, 8 cache PVCs, all service deployment YAML updates, training job template, and cleanup. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
41 KiB
S3 Binary Share Implementation Plan
For Claude: REQUIRED SUB-SKILL: Use superpowers:executing-plans to implement this plan task-by-task.
Goal: Replace per-service Docker images with generic base images + S3-based binary distribution to eliminate the Kaniko build stage and speed up pod startup.
Architecture: CI compiles binaries and uploads them to s3://foxhunt-binaries/latest/. Pods run pre-pulled generic base images with an initContainer that fetches the binary from S3 into an emptyDir. A per-service PVC cache provides fallback if S3 is unavailable. The build stage (9 Kaniko jobs) is entirely eliminated.
Tech Stack: Scaleway Object Storage (S3), rclone, Kubernetes initContainers, Terragrunt/Terraform, GitLab CI
Design doc: docs/plans/2026-02-28-s3-binary-share-design.md
Task 1: Create S3 bucket via Terraform
Files:
- Modify:
infra/modules/object-storage/main.tf(append new resource) - Modify:
infra/modules/object-storage/outputs.tf(append new output)
Step 1: Add the foxhunt-binaries bucket resource
In infra/modules/object-storage/main.tf, append after the sccache resource (line ~62):
resource "scaleway_object_bucket" "binaries" {
name = "${var.bucket_name_prefix}-binaries"
region = var.region
lifecycle_rule {
enabled = true
prefix = "archive/"
expiration {
days = 30
}
}
versioning {
enabled = false
}
}
Step 2: Add the output
In infra/modules/object-storage/outputs.tf, append:
output "binaries_bucket_name" {
description = "Name of the binaries distribution bucket"
value = scaleway_object_bucket.binaries.name
}
Step 3: Validate locally
Run: cd infra/live/production/object-storage && terragrunt validate
Expected: Success (no errors)
Step 4: Commit
git add infra/modules/object-storage/main.tf infra/modules/object-storage/outputs.tf
git commit -m "infra: add foxhunt-binaries S3 bucket for binary distribution"
Note: The actual
terragrunt applyruns in CI (infra-apply job) or manually. The bucket name will befoxhunt-binaries(prefixfoxhunt+-binaries).
Task 2: Add rclone to CI builder images
Files:
- Modify:
infra/docker/Dockerfile.ci-builder(add rclone install after sccache) - Modify:
infra/docker/Dockerfile.ci-builder-cpu(add rclone install after sccache)
Step 1: Add rclone to Dockerfile.ci-builder
In infra/docker/Dockerfile.ci-builder, after the sccache install block (line ~44, after chmod +x /usr/local/bin/sccache), add:
# rclone for S3 binary uploads (compile → S3)
RUN curl -fsSL https://downloads.rclone.org/v1.69.1/rclone-v1.69.1-linux-amd64.zip -o /tmp/rclone.zip \
&& unzip -j /tmp/rclone.zip '*/rclone' -d /usr/local/bin/ \
&& rm /tmp/rclone.zip \
&& chmod +x /usr/local/bin/rclone
Also add rclone --version to the verify line (line ~60):
RUN rustc --version && cargo --version && git --version && protoc --version && sccache --version && make --version && nvcc --version && rclone --version
Step 2: Add rclone to Dockerfile.ci-builder-cpu
Same change in infra/docker/Dockerfile.ci-builder-cpu, after the sccache install (line ~43), add the same rclone block. Update verify line (line ~47):
RUN rustc --version && cargo --version && git --version && protoc --version && sccache --version && make --version && mold --version && rclone --version
Step 3: Commit
git add infra/docker/Dockerfile.ci-builder infra/docker/Dockerfile.ci-builder-cpu
git commit -m "infra: add rclone to CI builder images for S3 binary uploads"
Note: CI builder images are rebuilt via the
build-ci-builder/build-ci-builder-cpuKaniko jobs in the prepare stage. Trigger manually from the pipeline UI after merging.
Task 3: Create 3 generic base image Dockerfiles
Files:
- Create:
infra/docker/Dockerfile.foxhunt-runtime - Create:
infra/docker/Dockerfile.foxhunt-training-runtime - Create:
infra/docker/Dockerfile.foxhunt-node-builder
Step 1: Create foxhunt-runtime
Write infra/docker/Dockerfile.foxhunt-runtime:
# Generic runtime base image for all Foxhunt service pods
# Contains: rclone (S3 binary fetch), grpc_health_probe (K8s health checks)
# Binaries are fetched by initContainer at pod startup, not baked into this image.
# Rebuild: only when system deps change (libssl, rclone, grpc_health_probe)
FROM debian:bookworm-slim
RUN apt-get update && apt-get install -y --no-install-recommends \
ca-certificates \
libssl3 \
curl \
unzip \
&& curl -fsSL https://downloads.rclone.org/v1.69.1/rclone-v1.69.1-linux-amd64.zip -o /tmp/rclone.zip \
&& unzip -j /tmp/rclone.zip '*/rclone' -d /usr/local/bin/ \
&& rm /tmp/rclone.zip \
&& apt-get purge -y unzip \
&& rm -rf /var/lib/apt/lists/*
# grpc_health_probe for Kubernetes health checks
RUN curl -fsSL https://github.com/grpc-ecosystem/grpc-health-probe/releases/download/v0.4.25/grpc_health_probe-linux-amd64 \
-o /usr/local/bin/grpc_health_probe \
&& chmod +x /usr/local/bin/grpc_health_probe
RUN groupadd -g 1000 foxhunt \
&& useradd -u 1000 -g foxhunt -m -s /bin/false foxhunt
USER foxhunt
Step 2: Create foxhunt-training-runtime
Write infra/docker/Dockerfile.foxhunt-training-runtime:
# Generic GPU runtime base image for Foxhunt training pods
# Contains: CUDA 12.4 + cuDNN + NVRTC + rclone
# Binaries are fetched by initContainer at pod startup, not baked into this image.
# Rebuild: only on CUDA version bump
FROM nvidia/cuda:12.4.1-cudnn-runtime-ubuntu22.04
ENV DEBIAN_FRONTEND=noninteractive
RUN apt-get update && apt-get install -y --no-install-recommends \
ca-certificates \
libssl3 \
curl \
unzip \
cuda-nvrtc-12-4 \
&& curl -fsSL https://downloads.rclone.org/v1.69.1/rclone-v1.69.1-linux-amd64.zip -o /tmp/rclone.zip \
&& unzip -j /tmp/rclone.zip '*/rclone' -d /usr/local/bin/ \
&& rm /tmp/rclone.zip \
&& apt-get purge -y unzip \
&& rm -rf /var/lib/apt/lists/*
RUN groupadd -g 1000 foxhunt \
&& useradd -u 1000 -g foxhunt -m -s /bin/false foxhunt
ENV NVIDIA_VISIBLE_DEVICES=all
ENV NVIDIA_DRIVER_CAPABILITIES=compute,utility
USER foxhunt
Step 3: Create foxhunt-node-builder
Write infra/docker/Dockerfile.foxhunt-node-builder:
# Node.js CI image for web-dashboard Vite builds + S3 upload
# Contains: Node 22 + rclone
# Rebuild: only on Node major version bump
FROM node:22-slim
RUN apt-get update && apt-get install -y --no-install-recommends \
curl \
unzip \
&& curl -fsSL https://downloads.rclone.org/v1.69.1/rclone-v1.69.1-linux-amd64.zip -o /tmp/rclone.zip \
&& unzip -j /tmp/rclone.zip '*/rclone' -d /usr/local/bin/ \
&& rm /tmp/rclone.zip \
&& apt-get purge -y unzip \
&& rm -rf /var/lib/apt/lists/*
Step 4: Commit
git add infra/docker/Dockerfile.foxhunt-runtime \
infra/docker/Dockerfile.foxhunt-training-runtime \
infra/docker/Dockerfile.foxhunt-node-builder
git commit -m "infra: add 3 generic base images for S3 binary share"
Task 4: Add base image build jobs to CI prepare stage
Files:
- Modify:
.gitlab-ci.yml(add 3 Kaniko jobs in thepreparestage, after existing builder jobs)
Step 1: Add build-foxhunt-runtime job
In .gitlab-ci.yml, after the build-infra-runner job (around line 188), add:
# --------------------------------------------------------------------------
# Stage 0d: Build generic runtime base images → push to Scaleway CR
# --------------------------------------------------------------------------
build-foxhunt-runtime:
stage: prepare
image:
name: gcr.io/kaniko-project/executor:debug
entrypoint: [""]
tags:
- kapsule
- docker
rules:
- if: $CI_PIPELINE_SOURCE == "push"
changes:
- infra/docker/Dockerfile.foxhunt-runtime
when: on_success
- when: manual
allow_failure: true
before_script:
- mkdir -p /kaniko/.docker
- >-
echo "{\"auths\":{
\"rg.fr-par.scw.cloud\":{\"username\":\"nologin\",\"password\":\"${SCW_SECRET_KEY}\"},
\"https://index.docker.io/v1/\":{\"username\":\"${DOCKERHUB_USERNAME}\",\"password\":\"${DOCKERHUB_TOKEN}\"}
}}" > /kaniko/.docker/config.json
script:
- /kaniko/executor
--context "${CI_PROJECT_DIR}"
--dockerfile "${CI_PROJECT_DIR}/infra/docker/Dockerfile.foxhunt-runtime"
--cache=true --cache-repo="${REGISTRY}/cache"
--destination "${REGISTRY}/foxhunt-runtime:latest"
build-foxhunt-training-runtime:
stage: prepare
image:
name: gcr.io/kaniko-project/executor:debug
entrypoint: [""]
tags:
- kapsule
- docker
rules:
- if: $CI_PIPELINE_SOURCE == "push"
changes:
- infra/docker/Dockerfile.foxhunt-training-runtime
when: on_success
- when: manual
allow_failure: true
before_script:
- mkdir -p /kaniko/.docker
- >-
echo "{\"auths\":{
\"rg.fr-par.scw.cloud\":{\"username\":\"nologin\",\"password\":\"${SCW_SECRET_KEY}\"},
\"https://index.docker.io/v1/\":{\"username\":\"${DOCKERHUB_USERNAME}\",\"password\":\"${DOCKERHUB_TOKEN}\"}
}}" > /kaniko/.docker/config.json
script:
- /kaniko/executor
--context "${CI_PROJECT_DIR}"
--dockerfile "${CI_PROJECT_DIR}/infra/docker/Dockerfile.foxhunt-training-runtime"
--cache=true --cache-repo="${REGISTRY}/cache"
--destination "${REGISTRY}/foxhunt-training-runtime:latest"
build-foxhunt-node-builder:
stage: prepare
image:
name: gcr.io/kaniko-project/executor:debug
entrypoint: [""]
tags:
- kapsule
- docker
rules:
- if: $CI_PIPELINE_SOURCE == "push"
changes:
- infra/docker/Dockerfile.foxhunt-node-builder
when: on_success
- when: manual
allow_failure: true
before_script:
- mkdir -p /kaniko/.docker
- >-
echo "{\"auths\":{
\"rg.fr-par.scw.cloud\":{\"username\":\"nologin\",\"password\":\"${SCW_SECRET_KEY}\"},
\"https://index.docker.io/v1/\":{\"username\":\"${DOCKERHUB_USERNAME}\",\"password\":\"${DOCKERHUB_TOKEN}\"}
}}" > /kaniko/.docker/config.json
script:
- /kaniko/executor
--context "${CI_PROJECT_DIR}"
--dockerfile "${CI_PROJECT_DIR}/infra/docker/Dockerfile.foxhunt-node-builder"
--cache=true --cache-repo="${REGISTRY}/cache"
--destination "${REGISTRY}/foxhunt-node-builder:latest"
Step 2: Also add these image names as variables at the top of .gitlab-ci.yml
After the INFRA_RUNNER_IMAGE variable (around line 59), add:
# Generic runtime images for S3 binary share (service pods + training pods + node builder)
RUNTIME_IMAGE: rg.fr-par.scw.cloud/foxhunt-ci/foxhunt-runtime:latest
TRAINING_RUNTIME_IMAGE: rg.fr-par.scw.cloud/foxhunt-ci/foxhunt-training-runtime:latest
NODE_BUILDER_IMAGE: rg.fr-par.scw.cloud/foxhunt-ci/foxhunt-node-builder:latest
Step 3: Commit
git add .gitlab-ci.yml
git commit -m "ci: add build jobs for 3 generic base images in prepare stage"
Task 5: Modify compile-services to upload to S3
Files:
- Modify:
.gitlab-ci.yml—compile-servicesjob
Step 1: Replace the artifacts: block with S3 upload
In the compile-services job, replace the artifacts: block (the final 3 lines, around line 326-328) and append upload steps to the script: section.
The entire compile-services job should become:
compile-services:
extends: .rust-base-cpu
stage: compile
needs: []
rules:
- if: $CI_COMMIT_BRANCH == "main" && $CI_PIPELINE_SOURCE == "push"
changes:
- Cargo.toml
- Cargo.lock
- crates/**
- bin/**
- services/**
- infra/docker/Dockerfile.*
- if: $CI_PIPELINE_SOURCE == "web" || $CI_PIPELINE_SOURCE == "api"
script:
- mkdir -p "$SCCACHE_DIR" && sccache --zero-stats || true
- PROFILE=${DEV_RELEASE:+dev-release}; PROFILE=${PROFILE:-release}
- echo "Building services with --profile $PROFILE (CPU-only, no CUDA)"
- TARGET_DIR="target/$PROFILE"
- cargo build --profile $PROFILE
-p trading_service
-p api_gateway
-p broker_gateway_service
-p ml_training_service
-p backtesting_service
-p trading_agent_service
-p data_acquisition_service
-p web-gateway
- sccache --show-stats || true
- mkdir -p build-out/services
- |
for bin in trading_service api_gateway broker_gateway_service ml_training_service backtesting_service trading_agent_service data_acquisition_service web-gateway; do
cp $TARGET_DIR/$bin build-out/services/
strip build-out/services/$bin
done
- ls -lh build-out/services/
# Upload stripped binaries to S3
- export RCLONE_S3_PROVIDER=Scaleway RCLONE_S3_ENDPOINT=s3.fr-par.scw.cloud RCLONE_S3_REGION=fr-par
- export RCLONE_S3_ACCESS_KEY_ID=$SCW_ACCESS_KEY RCLONE_S3_SECRET_ACCESS_KEY=$SCW_SECRET_KEY
- |
for bin in trading_service api_gateway broker_gateway_service ml_training_service backtesting_service trading_agent_service data_acquisition_service web-gateway; do
rclone copyto build-out/services/$bin :s3:foxhunt-binaries/latest/services/$bin
done
- rclone copy build-out/services/ :s3:foxhunt-binaries/archive/${CI_COMMIT_SHA}/services/
- echo "Service binaries uploaded to S3"
Note: No more artifacts: block. The :s3: prefix is rclone's backend syntax when using env-var config (no named remote).
Step 2: Commit
git add .gitlab-ci.yml
git commit -m "ci: compile-services uploads to S3 instead of GitLab artifacts"
Task 6: Modify compile-training to upload to S3
Files:
- Modify:
.gitlab-ci.yml—compile-trainingjob
Step 1: Replace artifacts with S3 upload
The compile-training job currently ends with artifacts: (around line 383-384). Replace the entire job to match this:
compile-training:
extends: .rust-base-cpu
stage: compile
image: ${CI_BUILDER_IMAGE}
needs: [compile-services]
variables:
CUDA_COMPUTE_CAP: "89"
rules:
- if: $CI_COMMIT_BRANCH == "main" && $CI_PIPELINE_SOURCE == "push"
changes:
- Cargo.toml
- Cargo.lock
- crates/ml/**
- crates/trading_engine/**
- crates/common/**
- crates/risk/**
- crates/data/**
- crates/config/**
- crates/storage/**
- services/training_uploader/**
- if: $CI_PIPELINE_SOURCE == "web" || $CI_PIPELINE_SOURCE == "api"
script:
- echo "CUDA_COMPUTE_CAP=$CUDA_COMPUTE_CAP"
- export SCCACHE_DIR="/mnt/sccache/sm_${CUDA_COMPUTE_CAP}"
- mkdir -p "$SCCACHE_DIR" && sccache --zero-stats || true
- PROFILE=${DEV_RELEASE:+dev-release}; PROFILE=${PROFILE:-release}
- echo "Building training binaries with --profile $PROFILE (CUDA stubs)"
- TARGET_DIR="target/$PROFILE"
- cargo build --profile $PROFILE -p ml --features ml/cuda
--example train_baseline_rl
--example train_baseline_supervised
--example evaluate_baseline
--example evaluate_supervised
--example hyperopt_baseline_rl
--example hyperopt_baseline_supervised
- cargo build --profile $PROFILE -p training_uploader
- sccache --show-stats || true
- mkdir -p build-out/training
- |
for bin in train_baseline_rl train_baseline_supervised evaluate_baseline evaluate_supervised hyperopt_baseline_rl hyperopt_baseline_supervised; do
cp $TARGET_DIR/examples/$bin build-out/training/
strip build-out/training/$bin
done
cp $TARGET_DIR/training_uploader build-out/training/
strip build-out/training/training_uploader
- ls -lh build-out/training/
# Upload stripped binaries to S3
- export RCLONE_S3_PROVIDER=Scaleway RCLONE_S3_ENDPOINT=s3.fr-par.scw.cloud RCLONE_S3_REGION=fr-par
- export RCLONE_S3_ACCESS_KEY_ID=$SCW_ACCESS_KEY RCLONE_S3_SECRET_ACCESS_KEY=$SCW_SECRET_KEY
- |
for bin in train_baseline_rl train_baseline_supervised evaluate_baseline evaluate_supervised hyperopt_baseline_rl hyperopt_baseline_supervised training_uploader; do
rclone copyto build-out/training/$bin :s3:foxhunt-binaries/latest/training/$bin
done
- rclone copy build-out/training/ :s3:foxhunt-binaries/archive/${CI_COMMIT_SHA}/training/
- echo "Training binaries uploaded to S3"
Step 2: Commit
git add .gitlab-ci.yml
git commit -m "ci: compile-training uploads to S3 instead of GitLab artifacts"
Task 7: Add build-web-dashboard, write-manifest, and deploy CI jobs
Files:
- Modify:
.gitlab-ci.yml— add 3 new jobs
Step 1: Add build-web-dashboard job
After the compile-training job, add:
# --------------------------------------------------------------------------
# Stage 3c: Build web-dashboard static assets (Vite)
# --------------------------------------------------------------------------
build-web-dashboard:
stage: compile
image: ${NODE_BUILDER_IMAGE}
needs: []
tags:
- kapsule
- docker
variables:
KUBERNETES_CPU_REQUEST: "2000m"
KUBERNETES_CPU_LIMIT: "4000m"
KUBERNETES_MEMORY_REQUEST: "2Gi"
KUBERNETES_MEMORY_LIMIT: "4Gi"
rules:
- if: $CI_COMMIT_BRANCH == "main" && $CI_PIPELINE_SOURCE == "push"
changes:
- web-dashboard/**
- if: $CI_PIPELINE_SOURCE == "web" || $CI_PIPELINE_SOURCE == "api"
script:
- cd web-dashboard
- npm ci
- npm run build
- export RCLONE_S3_PROVIDER=Scaleway RCLONE_S3_ENDPOINT=s3.fr-par.scw.cloud RCLONE_S3_REGION=fr-par
- export RCLONE_S3_ACCESS_KEY_ID=$SCW_ACCESS_KEY RCLONE_S3_SECRET_ACCESS_KEY=$SCW_SECRET_KEY
- rclone sync dist/ :s3:foxhunt-binaries/latest/web-dashboard/dist/
- echo "Web dashboard assets uploaded to S3"
# --------------------------------------------------------------------------
# Stage 3d: Write manifest after all compiles complete
# --------------------------------------------------------------------------
write-manifest:
stage: compile
image: ${RUNTIME_IMAGE}
needs:
- job: compile-services
- job: compile-training
- job: build-web-dashboard
optional: true
tags:
- kapsule
- docker
variables:
KUBERNETES_CPU_REQUEST: "100m"
KUBERNETES_MEMORY_REQUEST: "64Mi"
rules:
- if: $CI_COMMIT_BRANCH == "main" && $CI_PIPELINE_SOURCE == "push"
- if: $CI_PIPELINE_SOURCE == "web" || $CI_PIPELINE_SOURCE == "api"
script:
- export RCLONE_S3_PROVIDER=Scaleway RCLONE_S3_ENDPOINT=s3.fr-par.scw.cloud RCLONE_S3_REGION=fr-par
- export RCLONE_S3_ACCESS_KEY_ID=$SCW_ACCESS_KEY RCLONE_S3_SECRET_ACCESS_KEY=$SCW_SECRET_KEY
- echo "{\"sha\":\"${CI_COMMIT_SHA}\",\"ts\":\"$(date -Iseconds)\"}" | rclone rcat :s3:foxhunt-binaries/latest/MANIFEST.json
- echo "Manifest written for ${CI_COMMIT_SHA}"
Step 2: Add deploy job (at the end of the file, in the deploy stage)
After the training jobs section, add:
# --------------------------------------------------------------------------
# Stage 5: Deploy — rolling restart of all service deployments
# --------------------------------------------------------------------------
deploy-services:
stage: deploy
image:
name: bitnami/kubectl:latest
entrypoint: [""]
tags:
- kapsule
- docker
needs:
- job: write-manifest
variables:
KUBERNETES_CPU_REQUEST: "100m"
KUBERNETES_MEMORY_REQUEST: "64Mi"
rules:
- if: $CI_COMMIT_BRANCH == "main" && $CI_PIPELINE_SOURCE == "push"
when: manual
allow_failure: true
- if: $CI_PIPELINE_SOURCE == "web" || $CI_PIPELINE_SOURCE == "api"
when: manual
allow_failure: true
script:
- echo "Deploying ${CI_COMMIT_SHA}..."
- |
for svc in trading-service api-gateway broker-gateway \
ml-training-service backtesting-service \
trading-agent-service data-acquisition-service web-gateway; do
kubectl -n foxhunt rollout restart deployment/$svc
echo "Restarted $svc"
done
- |
for svc in trading-service api-gateway broker-gateway \
ml-training-service backtesting-service \
trading-agent-service data-acquisition-service web-gateway; do
kubectl -n foxhunt rollout status deployment/$svc --timeout=120s
done
- echo "All services deployed"
Step 3: Commit
git add .gitlab-ci.yml
git commit -m "ci: add build-web-dashboard, write-manifest, and deploy-services jobs"
Task 8: Delete all Kaniko build jobs from CI
Files:
- Modify:
.gitlab-ci.yml— remove the entire build stage section
Step 1: Delete the following blocks
Remove these sections from .gitlab-ci.yml:
.kaniko-service-basetemplate (thestage: buildtemplate starting with "Kaniko base for service images").kaniko-training-basetemplatebuild-trading-servicejobbuild-api-gatewayjobbuild-broker-gatewayjobbuild-ml-trainingjobbuild-backtestingjobbuild-trading-agentjobbuild-data-acquisitionjobbuild-web-gatewayjobbuild-trainingjob
These are roughly lines 390-549 in the current .gitlab-ci.yml.
Also remove build from the stages: list (line 31). The stages should become:
stages:
- prepare
- test
- compile
- train
- deploy
Step 2: Update training job needs:
All training jobs currently have needs: [build-training] (with optional: true). Change them to depend on compile-training instead:
Find all instances of:
needs:
- job: build-training
optional: true
Replace with:
needs:
- job: compile-training
optional: true
This applies to: train-validate-rl, train-validate-tft, train-validate-mamba2, train-validate-liquid, train-validate-tggn, train-validate-tlob, train-validate-kan, train-validate-xlstm, train-validate-diffusion, hyperopt-ppo, hyperopt-dqn, hyperopt-tft, hyperopt-mamba2, hyperopt-liquid, hyperopt-tggn.
Step 3: Update training job image:
All training jobs currently use image: ${REGISTRY}/training:${CI_COMMIT_SHA}. Change to the generic runtime:
Find: image: ${REGISTRY}/training:${CI_COMMIT_SHA}
Replace with: image: ${TRAINING_RUNTIME_IMAGE}
This applies to .train-rl-base and .train-validate-base templates.
Step 4: Update training job scripts — fetch binaries first
In .train-rl-base and .train-validate-base, add binary fetch to before_script::
Replace the existing before_script: in .train-rl-base:
before_script:
- export LD_LIBRARY_PATH=$(echo "$LD_LIBRARY_PATH" | tr ':' '\n' | grep -v stubs | tr '\n' ':' | sed 's/:$//')
- nvidia-smi
- mkdir -p ${CI_PROJECT_DIR}/output
With:
before_script:
- export LD_LIBRARY_PATH=$(echo "$LD_LIBRARY_PATH" | tr ':' '\n' | grep -v stubs | tr '\n' ':' | sed 's/:$//')
- nvidia-smi
- mkdir -p ${CI_PROJECT_DIR}/output
# Fetch training binaries from S3
- export RCLONE_S3_PROVIDER=Scaleway RCLONE_S3_ENDPOINT=s3.fr-par.scw.cloud RCLONE_S3_REGION=fr-par
- export RCLONE_S3_ACCESS_KEY_ID=$SCW_ACCESS_KEY RCLONE_S3_SECRET_ACCESS_KEY=$SCW_SECRET_KEY
- rclone sync :s3:foxhunt-binaries/latest/training/ /usr/local/bin/ --include "*.{rl,supervised,baseline,uploader}*" 2>/dev/null || rclone sync :s3:foxhunt-binaries/latest/training/ /usr/local/bin/
- chmod +x /usr/local/bin/train_baseline_rl /usr/local/bin/train_baseline_supervised /usr/local/bin/evaluate_baseline /usr/local/bin/evaluate_supervised /usr/local/bin/hyperopt_baseline_rl /usr/local/bin/hyperopt_baseline_supervised /usr/local/bin/training_uploader
Same change for .train-validate-base.
Step 5: Commit
git add .gitlab-ci.yml
git commit -m "ci: delete all Kaniko build jobs, switch training to S3 binary fetch"
Task 9: Create binary cache PVCs
Files:
- Create:
infra/k8s/storage/binary-cache-pvcs.yaml
Step 1: Write the PVC definitions
Write infra/k8s/storage/binary-cache-pvcs.yaml:
# Per-service binary cache PVCs — fallback when S3 is unavailable
# Each service gets its own 500Mi PVC (ReadWriteOnce, scw-bssd)
# Total: 8 × 500Mi = 4Gi, ~€0.08/month
---
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: binary-cache-trading-service
namespace: foxhunt
labels:
app.kubernetes.io/name: binary-cache
app.kubernetes.io/part-of: foxhunt
spec:
accessModes: [ReadWriteOnce]
storageClassName: scw-bssd
resources:
requests:
storage: 500Mi
---
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: binary-cache-api-gateway
namespace: foxhunt
labels:
app.kubernetes.io/name: binary-cache
app.kubernetes.io/part-of: foxhunt
spec:
accessModes: [ReadWriteOnce]
storageClassName: scw-bssd
resources:
requests:
storage: 500Mi
---
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: binary-cache-broker-gateway
namespace: foxhunt
labels:
app.kubernetes.io/name: binary-cache
app.kubernetes.io/part-of: foxhunt
spec:
accessModes: [ReadWriteOnce]
storageClassName: scw-bssd
resources:
requests:
storage: 500Mi
---
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: binary-cache-ml-training-service
namespace: foxhunt
labels:
app.kubernetes.io/name: binary-cache
app.kubernetes.io/part-of: foxhunt
spec:
accessModes: [ReadWriteOnce]
storageClassName: scw-bssd
resources:
requests:
storage: 500Mi
---
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: binary-cache-backtesting-service
namespace: foxhunt
labels:
app.kubernetes.io/name: binary-cache
app.kubernetes.io/part-of: foxhunt
spec:
accessModes: [ReadWriteOnce]
storageClassName: scw-bssd
resources:
requests:
storage: 500Mi
---
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: binary-cache-trading-agent-service
namespace: foxhunt
labels:
app.kubernetes.io/name: binary-cache
app.kubernetes.io/part-of: foxhunt
spec:
accessModes: [ReadWriteOnce]
storageClassName: scw-bssd
resources:
requests:
storage: 500Mi
---
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: binary-cache-data-acquisition-service
namespace: foxhunt
labels:
app.kubernetes.io/name: binary-cache
app.kubernetes.io/part-of: foxhunt
spec:
accessModes: [ReadWriteOnce]
storageClassName: scw-bssd
resources:
requests:
storage: 500Mi
---
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: binary-cache-web-gateway
namespace: foxhunt
labels:
app.kubernetes.io/name: binary-cache
app.kubernetes.io/part-of: foxhunt
spec:
accessModes: [ReadWriteOnce]
storageClassName: scw-bssd
resources:
requests:
storage: 1Gi
Note: web-gateway gets 1Gi because it also caches the dashboard static assets (~5MB).
Step 2: Commit
git add infra/k8s/storage/binary-cache-pvcs.yaml
git commit -m "infra: add 8 binary cache PVCs for S3 fallback during trading"
Apply:
kubectl apply -f infra/k8s/storage/binary-cache-pvcs.yaml
Task 10: Update all 8 service deployment YAMLs
This is the largest task. Each deployment gets:
- Image changed from per-service GitLab registry → generic
foxhunt-runtimefrom SCW registry imagePullSecretschanged fromgitlab-registry→scw-registryinitContainersadded for S3 binary fetch with PVC cache fallbackcommandadded to main container to run from/binaries/volumesupdated withbinariesemptyDir +binary-cachePVC- Deployment strategy set to
maxSurge: 1, maxUnavailable: 0
Files:
- Modify:
infra/k8s/services/trading-service.yaml - Modify:
infra/k8s/services/api-gateway.yaml - Modify:
infra/k8s/services/broker-gateway.yaml - Modify:
infra/k8s/services/ml-training-service.yaml - Modify:
infra/k8s/services/backtesting-service.yaml - Modify:
infra/k8s/services/trading-agent-service.yaml - Modify:
infra/k8s/services/data-acquisition-service.yaml - Modify:
infra/k8s/services/web-gateway.yaml
The pattern for each service (example: trading-service):
For each deployment YAML, make these changes:
- Add deployment strategy after
spec.replicas:
strategy:
type: RollingUpdate
rollingUpdate:
maxSurge: 1
maxUnavailable: 0
- Change
imagePullSecretsfromgitlab-registrytoscw-registry:
imagePullSecrets:
- name: scw-registry
- Add
initContainersbeforecontainers:
initContainers:
- name: fetch-binary
image: rg.fr-par.scw.cloud/foxhunt-ci/foxhunt-runtime:latest
command: ["/bin/sh", "-c"]
args:
- |
set -e
SERVICE=trading_service
if rclone copyto :s3:foxhunt-binaries/latest/services/$SERVICE /binaries/$SERVICE 2>/dev/null; then
chmod +x /binaries/$SERVICE
cp /binaries/$SERVICE /cache/$SERVICE
echo "Fetched $SERVICE from S3 ($(stat -c%s /binaries/$SERVICE) bytes)"
elif [ -f /cache/$SERVICE ]; then
cp /cache/$SERVICE /binaries/$SERVICE
chmod +x /binaries/$SERVICE
echo "WARNING: S3 unavailable, using cached $SERVICE"
else
echo "FATAL: No binary available for $SERVICE (S3 down + empty cache)"
exit 1
fi
env:
- name: RCLONE_S3_PROVIDER
value: Scaleway
- name: RCLONE_S3_ENDPOINT
value: s3.fr-par.scw.cloud
- name: RCLONE_S3_REGION
value: fr-par
- name: RCLONE_S3_ACCESS_KEY_ID
valueFrom:
secretKeyRef:
name: s3-credentials
key: access-key
- name: RCLONE_S3_SECRET_ACCESS_KEY
valueFrom:
secretKeyRef:
name: s3-credentials
key: secret-key
volumeMounts:
- name: binaries
mountPath: /binaries
- name: binary-cache
mountPath: /cache
resources:
requests:
cpu: 100m
memory: 64Mi
limits:
cpu: 500m
memory: 128Mi
- Change container image from
git.fxhnt.ai:5050/foxhunt/foxhunt/<service>:latestto:
image: rg.fr-par.scw.cloud/foxhunt-ci/foxhunt-runtime:latest
- Add
commandto the main container:
command: ["/binaries/trading_service"]
- Add volume mounts to the main container:
volumeMounts:
- name: binaries
mountPath: /binaries
readOnly: true
(Keep existing volumeMounts like logs, tls-certs)
- Add volumes to the pod spec:
volumes:
- name: binaries
emptyDir:
sizeLimit: 200Mi
- name: binary-cache
persistentVolumeClaim:
claimName: binary-cache-trading-service
(Keep existing volumes like logs, tls-certs)
Per-service binary names and PVC names:
| Deployment | Binary name | PVC name | Special |
|---|---|---|---|
| trading-service | trading_service |
binary-cache-trading-service |
— |
| api-gateway | api_gateway |
binary-cache-api-gateway |
— |
| broker-gateway | broker_gateway_service |
binary-cache-broker-gateway |
— |
| ml-training-service | ml_training_service |
binary-cache-ml-training-service |
Has command: ["./ml_training_service", "serve"] → change to command: ["/binaries/ml_training_service", "serve"] |
| backtesting-service | backtesting_service |
binary-cache-backtesting-service |
— |
| trading-agent-service | trading_agent_service |
binary-cache-trading-agent-service |
— |
| data-acquisition-service | data_acquisition_service |
binary-cache-data-acquisition-service |
Uses scw-registry already |
| web-gateway | web-gateway |
binary-cache-web-gateway |
See below |
web-gateway special case:
The initContainer fetches both the binary AND dashboard static assets:
initContainers:
- name: fetch-binary
image: rg.fr-par.scw.cloud/foxhunt-ci/foxhunt-runtime:latest
command: ["/bin/sh", "-c"]
args:
- |
set -e
SERVICE=web-gateway
if rclone copyto :s3:foxhunt-binaries/latest/services/$SERVICE /binaries/$SERVICE 2>/dev/null; then
chmod +x /binaries/$SERVICE
cp /binaries/$SERVICE /cache/$SERVICE
rclone sync :s3:foxhunt-binaries/latest/web-dashboard/dist/ /binaries/static/
cp -r /binaries/static /cache/static 2>/dev/null || true
echo "Fetched $SERVICE + static assets from S3"
elif [ -f /cache/$SERVICE ]; then
cp /cache/$SERVICE /binaries/$SERVICE
chmod +x /binaries/$SERVICE
cp -r /cache/static /binaries/static 2>/dev/null || true
echo "WARNING: S3 unavailable, using cached $SERVICE"
else
echo "FATAL: No binary available for $SERVICE"
exit 1
fi
env:
# ... same RCLONE_S3_* env vars ...
volumeMounts:
- name: binaries
mountPath: /binaries
- name: binary-cache
mountPath: /cache
The web-gateway container gets:
command: ["/binaries/web-gateway"]
And the STATIC_DIR environment variable (if the binary supports it, otherwise it defaults to ./static/ which would be /binaries/static/ since the workdir would need adjustment — check the binary).
Important: The web-gateway binary currently serves static files from
./static/(relative to WORKDIR/app). With the new layout, static files are at/binaries/static/. Either: (a) Set envSTATIC_DIR=/binaries/static/if the binary supports it, or (b) Symlink: addln -sf /binaries/static /app/staticto the initContainer. Check the Rust source (services/web-gateway/src/main.rsor similar) to determine the static path config.
Step 1-8: Update each YAML file following the pattern above.
Step 9: Commit
git add infra/k8s/services/trading-service.yaml \
infra/k8s/services/api-gateway.yaml \
infra/k8s/services/broker-gateway.yaml \
infra/k8s/services/ml-training-service.yaml \
infra/k8s/services/backtesting-service.yaml \
infra/k8s/services/trading-agent-service.yaml \
infra/k8s/services/data-acquisition-service.yaml \
infra/k8s/services/web-gateway.yaml
git commit -m "infra: update all service deployments with S3 binary fetch initContainer"
Task 11: Update GPU overlay deployments
Files:
- Modify:
infra/k8s/services/trading-service-gpu.yaml - Modify:
infra/k8s/services/ml-training-service-gpu.yaml
Apply the same changes as Task 10 to both GPU overlay files. The only differences from the base manifests are nodeSelector (gpu-inference), tolerations, and resource requests (GPU).
Step 1: Commit
git add infra/k8s/services/trading-service-gpu.yaml \
infra/k8s/services/ml-training-service-gpu.yaml
git commit -m "infra: update GPU overlay deployments with S3 binary fetch"
Task 12: Update training job template
Files:
- Modify:
infra/k8s/training/job-template.yaml
Step 1: Update the job template
Change the image from rg.fr-par.scw.cloud/foxhunt-ci/training:latest to rg.fr-par.scw.cloud/foxhunt-ci/foxhunt-training-runtime:latest.
Add an initContainer to fetch training binaries. The training job doesn't need a cache PVC (batch jobs can retry).
Replace the initContainers: section (currently the uploader sidecar) to also include binary fetch, and update the main container to use fetched binaries:
initContainers:
# Fetch training binaries from S3
- name: fetch-binaries
image: rg.fr-par.scw.cloud/foxhunt-ci/foxhunt-training-runtime:latest
command: ["/bin/sh", "-c"]
args:
- |
set -e
export RCLONE_S3_PROVIDER=Scaleway RCLONE_S3_ENDPOINT=s3.fr-par.scw.cloud RCLONE_S3_REGION=fr-par
rclone sync :s3:foxhunt-binaries/latest/training/ /binaries/
chmod +x /binaries/*
ls -lh /binaries/
echo "Fetched training binaries from S3"
env:
- name: RCLONE_S3_ACCESS_KEY_ID
valueFrom:
secretKeyRef:
name: s3-credentials
key: access-key
- name: RCLONE_S3_SECRET_ACCESS_KEY
valueFrom:
secretKeyRef:
name: s3-credentials
key: secret-key
volumeMounts:
- name: binaries
mountPath: /binaries
resources:
requests:
cpu: 100m
memory: 64Mi
limits:
cpu: 500m
memory: 256Mi
# Uploader sidecar (native sidecar, restartPolicy: Always)
- name: uploader
image: rg.fr-par.scw.cloud/foxhunt-ci/foxhunt-training-runtime:latest
restartPolicy: Always
command: ["/binaries/training_uploader"]
# ... rest of uploader config unchanged, but add binaries mount ...
volumeMounts:
- name: output
mountPath: /output
readOnly: true
- name: binaries
mountPath: /binaries
readOnly: true
Update main container:
containers:
- name: training
image: rg.fr-par.scw.cloud/foxhunt-ci/foxhunt-training-runtime:latest
command: ["/binaries/$(TRAINING_BINARY)"]
volumeMounts:
- name: training-data
mountPath: /data
readOnly: true
- name: output
mountPath: /output
- name: binaries
mountPath: /binaries
readOnly: true
Add the binaries volume:
volumes:
- name: training-data
persistentVolumeClaim:
claimName: training-data-pvc
- name: output
emptyDir:
sizeLimit: 2Gi
- name: binaries
emptyDir:
sizeLimit: 500Mi
Step 2: Commit
git add infra/k8s/training/job-template.yaml
git commit -m "infra: update training job template with S3 binary fetch"
Task 13: Update image-prepuller DaemonSet
Files:
- Modify:
infra/k8s/training/image-prepuller.yaml
Step 1: Change the init container image
Replace:
- name: pull-training
image: rg.fr-par.scw.cloud/foxhunt-ci/training:latest
command: ["echo", "training image pulled"]
With:
- name: pull-training-runtime
image: rg.fr-par.scw.cloud/foxhunt-ci/foxhunt-training-runtime:latest
command: ["echo", "foxhunt-training-runtime image pulled"]
Step 2: Commit
git add infra/k8s/training/image-prepuller.yaml
git commit -m "infra: update image-prepuller to cache foxhunt-training-runtime"
Task 14: Delete old Dockerfiles
Files:
- Delete:
infra/docker/Dockerfile.runtime - Delete:
infra/docker/Dockerfile.training - Delete:
infra/docker/Dockerfile.web-gateway-runtime - Delete:
infra/docker/Dockerfile.web-gateway - Delete:
infra/docker/Dockerfile.service
Step 1: Remove the files
git rm infra/docker/Dockerfile.runtime \
infra/docker/Dockerfile.training \
infra/docker/Dockerfile.web-gateway-runtime \
infra/docker/Dockerfile.web-gateway \
infra/docker/Dockerfile.service
Step 2: Commit
git commit -m "infra: delete 5 obsolete Dockerfiles replaced by S3 binary share"
Task 15: Verify and seed — first manual pipeline run
This is a manual verification task, not automated.
Step 1: Build base images first
Trigger the 3 new base image builds from GitLab pipeline UI:
build-foxhunt-runtime(manual trigger)build-foxhunt-training-runtime(manual trigger)build-foxhunt-node-builder(manual trigger)
Also trigger CI builder rebuilds (they now include rclone):
build-ci-builder(manual trigger)build-ci-builder-cpu(manual trigger)
Step 2: Apply the S3 bucket via Terragrunt
cd infra/live/production/object-storage
terragrunt apply
Step 3: Apply PVCs
kubectl apply -f infra/k8s/storage/binary-cache-pvcs.yaml
Step 4: Run a full pipeline
Push to main — the compile jobs should upload to S3, write-manifest should succeed, and deploy-services should trigger rolling restarts.
Step 5: Verify pods
kubectl -n foxhunt get pods
kubectl -n foxhunt logs deployment/trading-service -c fetch-binary
Step 6: Apply service manifests
kubectl apply -f infra/k8s/services/
kubectl apply -f infra/k8s/training/image-prepuller.yaml
Execution Order Summary
| Task | Dependencies | Can Parallelize |
|---|---|---|
| 1. S3 bucket TF | None | Yes |
| 2. rclone in CI builders | None | Yes (with 1) |
| 3. Base image Dockerfiles | None | Yes (with 1, 2) |
| 4. Base image CI build jobs | 3 | No |
| 5. compile-services S3 upload | 2 | Yes (with 6, 7) |
| 6. compile-training S3 upload | 2 | Yes (with 5, 7) |
| 7. New CI jobs (dashboard, manifest, deploy) | 3 | Yes (with 5, 6) |
| 8. Delete Kaniko build jobs | 5, 6, 7 | No |
| 9. Binary cache PVCs | None | Yes (with anything) |
| 10. Service deployment YAMLs | 3, 9 | Yes (with 11, 12) |
| 11. GPU overlay YAMLs | 10 | Yes (with 12) |
| 12. Training job template | 3 | Yes (with 10, 11) |
| 13. Image-prepuller update | 3 | Yes |
| 14. Delete old Dockerfiles | 8, 10 | No |
| 15. Manual verify + seed | All above | No |