Replace per-service Docker images with generic base images + S3-based binary distribution. CI compiles and uploads stripped binaries to S3. Pods fetch binaries via initContainer with PVC cache fallback for trading resilience. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
15 KiB
S3 Binary Share — Eliminate Docker Builds, Faster Pod Startup
Date: 2026-02-28 Status: Approved
Problem
Every code change triggers a long pipeline: compile → Kaniko packages 9 images → push to registry → pods pull images. The training image alone is ~3GB (CUDA runtime). Even with an image-prepuller DaemonSet, rebuilding and pushing images adds 3-7 minutes per pipeline. During active trading, a pod crash + image pull could cause unacceptable downtime.
Solution
Replace per-service Docker images with generic base images + S3-based binary distribution. CI compiles and uploads stripped binaries to S3. Pods run generic base images (pre-pulled, rarely change) and fetch binaries from S3 via initContainer on startup. A PVC cache per service provides fallback if S3 is unavailable during trading.
Architecture
S3 Bucket Structure
s3://foxhunt-binaries/
├── latest/
│ ├── services/
│ │ ├── trading_service
│ │ ├── api_gateway
│ │ ├── broker_gateway_service
│ │ ├── ml_training_service
│ │ ├── backtesting_service
│ │ ├── trading_agent_service
│ │ ├── data_acquisition_service
│ │ └── web-gateway
│ ├── training/
│ │ ├── train_baseline_rl
│ │ ├── train_baseline_supervised
│ │ ├── evaluate_baseline
│ │ ├── evaluate_supervised
│ │ ├── hyperopt_baseline_rl
│ │ ├── hyperopt_baseline_supervised
│ │ └── training_uploader
│ ├── web-dashboard/
│ │ └── dist/ # Vite build output
│ └── MANIFEST.json # {sha, timestamp}
└── archive/<commit-sha>/ # optional audit trail
- Compile always overwrites
latest/ - Optional: copy to
archive/<sha>/for debugging - Pods always fetch from
latest/ - Lifecycle rule: delete archive entries older than 30 days
- Bucket region: fr-par (same as cluster)
Generic Base Images (3 total)
All pushed to rg.fr-par.scw.cloud/foxhunt-ci/ (Scaleway Container Registry).
1. foxhunt-runtime (services)
FROM debian:bookworm-slim
RUN apt-get update && apt-get install -y --no-install-recommends \
ca-certificates libssl3 curl unzip \
&& curl -fsSL https://downloads.rclone.org/v1.69.1/rclone-v1.69.1-linux-amd64.zip -o /tmp/rclone.zip \
&& unzip -j /tmp/rclone.zip '*/rclone' -d /usr/local/bin/ \
&& rm /tmp/rclone.zip \
&& apt-get purge -y unzip && rm -rf /var/lib/apt/lists/*
RUN curl -fsSL https://github.com/grpc-ecosystem/grpc-health-probe/releases/download/v0.4.25/grpc_health_probe-linux-amd64 \
-o /usr/local/bin/grpc_health_probe && chmod +x /usr/local/bin/grpc_health_probe
RUN groupadd -g 1000 foxhunt && useradd -u 1000 -g foxhunt -m -s /bin/false foxhunt
USER foxhunt
Size: ~100MB. Rebuild: only when libssl/rclone/grpc_health_probe updates.
2. foxhunt-training-runtime (GPU training)
FROM nvidia/cuda:12.4.1-cudnn-runtime-ubuntu22.04
ENV DEBIAN_FRONTEND=noninteractive
RUN apt-get update && apt-get install -y --no-install-recommends \
ca-certificates libssl3 curl unzip cuda-nvrtc-12-4 \
&& curl -fsSL https://downloads.rclone.org/v1.69.1/rclone-v1.69.1-linux-amd64.zip -o /tmp/rclone.zip \
&& unzip -j /tmp/rclone.zip '*/rclone' -d /usr/local/bin/ \
&& rm /tmp/rclone.zip \
&& apt-get purge -y unzip && rm -rf /var/lib/apt/lists/*
RUN groupadd -g 1000 foxhunt && useradd -u 1000 -g foxhunt -m -s /bin/false foxhunt
ENV NVIDIA_VISIBLE_DEVICES=all
ENV NVIDIA_DRIVER_CAPABILITIES=compute,utility
USER foxhunt
Size: ~3GB (CUDA dominates). Rebuild: only on CUDA version bump. Image-prepuller DaemonSet still caches this on GPU nodes.
3. foxhunt-node-builder (Vite build CI job)
FROM node:22-slim
RUN apt-get update && apt-get install -y --no-install-recommends \
curl unzip \
&& curl -fsSL https://downloads.rclone.org/v1.69.1/rclone-v1.69.1-linux-amd64.zip -o /tmp/rclone.zip \
&& unzip -j /tmp/rclone.zip '*/rclone' -d /usr/local/bin/ \
&& rm /tmp/rclone.zip \
&& apt-get purge -y unzip && rm -rf /var/lib/apt/lists/*
Size: ~200MB. Rebuild: only on Node major version bump.
CI Pipeline Changes
New pipeline flow
prepare → test → compile → deploy
Build stage is entirely eliminated.
compile-services (modified)
Same cargo build, but uploads to S3 instead of creating GitLab artifacts:
compile-services:
extends: .rust-base-cpu
stage: compile
needs: []
script:
- mkdir -p "$SCCACHE_DIR" && sccache --zero-stats || true
- PROFILE=${DEV_RELEASE:+dev-release}; PROFILE=${PROFILE:-release}
- TARGET_DIR="target/$PROFILE"
- cargo build --profile $PROFILE
-p trading_service -p api_gateway -p broker_gateway_service
-p ml_training_service -p backtesting_service
-p trading_agent_service -p data_acquisition_service -p web-gateway
- sccache --show-stats || true
- mkdir -p build-out/services
- |
for bin in trading_service api_gateway broker_gateway_service \
ml_training_service backtesting_service \
trading_agent_service data_acquisition_service web-gateway; do
cp $TARGET_DIR/$bin build-out/services/
strip build-out/services/$bin
done
- ls -lh build-out/services/
# Upload to S3
- |
for bin in trading_service api_gateway broker_gateway_service \
ml_training_service backtesting_service \
trading_agent_service data_acquisition_service web-gateway; do
rclone copyto build-out/services/$bin s3:foxhunt-binaries/latest/services/$bin
done
# Archive by SHA
- rclone sync s3:foxhunt-binaries/latest/services/ s3:foxhunt-binaries/archive/${CI_COMMIT_SHA}/services/
compile-training (modified)
Same pattern — uploads to s3:foxhunt-binaries/latest/training/.
build-web-dashboard (new)
build-web-dashboard:
stage: compile
image: rg.fr-par.scw.cloud/foxhunt-ci/foxhunt-node-builder:latest
needs: []
tags: [kapsule, docker]
variables:
KUBERNETES_CPU_REQUEST: "2000m"
KUBERNETES_MEMORY_REQUEST: "2Gi"
rules:
- if: $CI_COMMIT_BRANCH == "main" && $CI_PIPELINE_SOURCE == "push"
changes: [web-dashboard/**]
- if: $CI_PIPELINE_SOURCE == "web" || $CI_PIPELINE_SOURCE == "api"
script:
- cd web-dashboard
- npm ci
- npm run build
- rclone sync dist/ s3:foxhunt-binaries/latest/web-dashboard/dist/
write-manifest (new, runs after all compiles)
write-manifest:
stage: compile
image: rg.fr-par.scw.cloud/foxhunt-ci/foxhunt-runtime:latest
needs: [compile-services, compile-training, build-web-dashboard]
tags: [kapsule, docker]
script:
- echo '{"sha":"'${CI_COMMIT_SHA}'","ts":"'$(date -Iseconds)'"}' \
| rclone rcat s3:foxhunt-binaries/latest/MANIFEST.json
deploy (new)
deploy:
stage: deploy
image: bitnami/kubectl:latest
needs: [write-manifest]
tags: [kapsule, docker]
rules:
- if: $CI_COMMIT_BRANCH == "main" && $CI_PIPELINE_SOURCE == "push"
- if: $CI_PIPELINE_SOURCE == "web" || $CI_PIPELINE_SOURCE == "api"
script:
- rclone cat s3:foxhunt-binaries/latest/MANIFEST.json
- |
for svc in trading-service api-gateway broker-gateway \
ml-training-service backtesting-service \
trading-agent-service data-acquisition-service web-gateway; do
kubectl -n foxhunt rollout restart deployment/$svc
done
- |
for svc in trading-service api-gateway broker-gateway \
ml-training-service backtesting-service \
trading-agent-service data-acquisition-service web-gateway; do
kubectl -n foxhunt rollout status deployment/$svc --timeout=120s
done
Deleted CI jobs
.kaniko-service-basetemplate.kaniko-training-basetemplatebuild-trading-service,build-api-gateway,build-broker-gatewaybuild-ml-training,build-backtesting,build-trading-agentbuild-data-acquisition,build-web-gateway,build-trainingartifacts:sections in compile jobs
Deployment Pattern
Every service deployment uses this initContainer pattern:
initContainers:
- name: fetch-binary
image: rg.fr-par.scw.cloud/foxhunt-ci/foxhunt-runtime:latest
command: ["/bin/sh", "-c"]
args:
- |
set -e
SERVICE=trading_service
if rclone copyto s3:foxhunt-binaries/latest/services/$SERVICE /binaries/$SERVICE 2>/dev/null; then
chmod +x /binaries/$SERVICE
cp /binaries/$SERVICE /cache/$SERVICE
echo "Fetched from S3 ($(stat -c%s /binaries/$SERVICE) bytes)"
elif [ -f /cache/$SERVICE ]; then
cp /cache/$SERVICE /binaries/$SERVICE
chmod +x /binaries/$SERVICE
echo "WARNING: S3 unavailable, using cached binary"
else
echo "FATAL: No binary available (S3 down + empty cache)"
exit 1
fi
env:
- name: RCLONE_S3_PROVIDER
value: Scaleway
- name: RCLONE_S3_ENDPOINT
value: s3.fr-par.scw.cloud
- name: RCLONE_S3_REGION
value: fr-par
- name: RCLONE_S3_ACCESS_KEY_ID
valueFrom:
secretKeyRef:
name: s3-credentials
key: access-key
- name: RCLONE_S3_SECRET_ACCESS_KEY
valueFrom:
secretKeyRef:
name: s3-credentials
key: secret-key
volumeMounts:
- name: binaries
mountPath: /binaries
- name: binary-cache
mountPath: /cache
resources:
requests:
cpu: 100m
memory: 64Mi
limits:
cpu: 500m
memory: 128Mi
containers:
- name: trading-service
image: rg.fr-par.scw.cloud/foxhunt-ci/foxhunt-runtime:latest
command: ["/binaries/trading_service"]
volumeMounts:
- name: binaries
mountPath: /binaries
readOnly: true
volumes:
- name: binaries
emptyDir:
sizeLimit: 200Mi
- name: binary-cache
persistentVolumeClaim:
claimName: binary-cache-trading-service
web-gateway special case
initContainer fetches both binary and dashboard assets:
args:
- |
set -e
SERVICE=web-gateway
if rclone copyto s3:foxhunt-binaries/latest/services/$SERVICE /binaries/$SERVICE 2>/dev/null; then
chmod +x /binaries/$SERVICE
cp /binaries/$SERVICE /cache/$SERVICE
rclone sync s3:foxhunt-binaries/latest/web-dashboard/dist/ /binaries/static/
cp -r /binaries/static /cache/static
echo "Fetched from S3"
elif [ -f /cache/$SERVICE ]; then
cp /cache/$SERVICE /binaries/$SERVICE
chmod +x /binaries/$SERVICE
cp -r /cache/static /binaries/static 2>/dev/null || true
echo "WARNING: S3 unavailable, using cached binary"
else
echo "FATAL: No binary available"
exit 1
fi
Training job template
Uses foxhunt-training-runtime, no cache PVC (batch jobs can retry):
initContainers:
- name: fetch-binaries
image: rg.fr-par.scw.cloud/foxhunt-ci/foxhunt-training-runtime:latest
command: ["/bin/sh", "-c"]
args:
- |
set -e
rclone sync s3:foxhunt-binaries/latest/training/ /binaries/
chmod +x /binaries/*
echo "Fetched training binaries"
volumeMounts:
- name: binaries
mountPath: /binaries
containers:
- name: training
image: rg.fr-par.scw.cloud/foxhunt-ci/foxhunt-training-runtime:latest
command: ["/binaries/$(TRAINING_BINARY)"]
volumeMounts:
- name: binaries
mountPath: /binaries
readOnly: true
Cache PVCs
8 PVCs, one per service deployment (500Mi each, scw-bssd):
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: binary-cache-trading-service
namespace: foxhunt
spec:
accessModes: [ReadWriteOnce]
resources:
requests:
storage: 500Mi
storageClassName: scw-bssd
Total: 8 × 500Mi = 4Gi. Cost: ~€0.08/month.
Resilience During Trading
- Running pods unaffected: Binaries live in emptyDir for the pod's lifetime. S3 outage has zero impact on running services.
- Pod crash + S3 down: initContainer falls back to PVC cache (last known good binary).
- Pod crash + S3 down + empty cache: Only on first-ever deploy. Pod stays in Init:Error; deployment's
maxUnavailable: 0prevents the old pod from terminating. - Deployment strategy:
maxSurge: 1, maxUnavailable: 0ensures old pod keeps running until new pod is ready.
What Gets Deleted
| Item | Notes |
|---|---|
Dockerfile.runtime |
Replaced by foxhunt-runtime |
Dockerfile.training |
Replaced by foxhunt-training-runtime |
Dockerfile.web-gateway-runtime |
Replaced by foxhunt-runtime + initContainer |
Dockerfile.web-gateway |
Full multi-stage builder, no longer needed |
Dockerfile.service |
Full multi-stage builder, no longer needed |
| 9 Kaniko CI jobs | All build-* jobs in .gitlab-ci.yml |
| 10 per-service Docker images | In Scaleway CR and GitLab registry |
What's Kept
| Item | Notes |
|---|---|
Dockerfile.ci-builder |
Compile infrastructure, unchanged |
Dockerfile.ci-builder-cpu |
Compile infrastructure, unchanged |
Dockerfile.infra-runner |
IaC pipeline, unchanged |
| Image-prepuller DaemonSet | Updated to pull foxhunt-training-runtime |
Timing Comparison
| Step | Before | After |
|---|---|---|
| Compile | 8-22 min | 8-22 min (unchanged) |
| Image build | 2-5 min (9 Kaniko) | 0 (eliminated) |
| Registry push | 1-2 min | ~10s (S3 upload) |
| Pod startup | 30s-3min (image pull) | 2-5s (initContainer S3 fetch) |
| Total pipeline | 15-30 min | 8-23 min |
| Pod restart | 30s-3min | 2-5s |
New Infrastructure
| Resource | Spec | Cost |
|---|---|---|
S3 bucket foxhunt-binaries |
fr-par, 30-day lifecycle | ~€0.01/month |
| 8 cache PVCs (500Mi each) | scw-bssd | ~€0.08/month |
| 3 base image Dockerfiles | In infra/docker/ |
0 |
rclone Configuration
rclone in CI and pods is configured via environment variables (no config file needed):
RCLONE_S3_PROVIDER=Scaleway
RCLONE_S3_ENDPOINT=s3.fr-par.scw.cloud
RCLONE_S3_REGION=fr-par
RCLONE_S3_ACCESS_KEY_ID=<from secret>
RCLONE_S3_SECRET_ACCESS_KEY=<from secret>
The s3-credentials K8s secret already exists in the foxhunt namespace (used by training_uploader). CI uses SCW_ACCESS_KEY/SCW_SECRET_KEY variables mapped to AWS_ACCESS_KEY_ID/AWS_SECRET_ACCESS_KEY.
For rclone in CI, the S3 remote is referenced as s3:foxhunt-binaries/... using RCLONE_S3_* env vars.
rclone in ci-builder images
Add rclone to both Dockerfile.ci-builder and Dockerfile.ci-builder-cpu for the upload step:
# Add to both ci-builder Dockerfiles
RUN curl -fsSL https://downloads.rclone.org/v1.69.1/rclone-v1.69.1-linux-amd64.zip -o /tmp/rclone.zip \
&& unzip -j /tmp/rclone.zip '*/rclone' -d /usr/local/bin/ \
&& rm /tmp/rclone.zip