Compare commits
12 Commits
surfer-uni
...
chore/gite
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
79f2ad94a4 | ||
|
|
c3aec12f9a | ||
|
|
b21d9b671d | ||
|
|
08cee0de0d | ||
|
|
80f3fbecef | ||
|
|
042fcd76e5 | ||
|
|
af3815ab6a | ||
|
|
3a18a348ac | ||
|
|
2eb8cd8333 | ||
|
|
2c08b49b4f | ||
|
|
b1a60ca134 | ||
|
|
f4f54c352b |
472
docs/superpowers/plans/2026-06-21-gitea-replace-gitlab.md
Normal file
472
docs/superpowers/plans/2026-06-21-gitea-replace-gitlab.md
Normal file
@@ -0,0 +1,472 @@
|
||||
# Gitea Replaces GitLab + Cockpit Registry to Scaleway — Implementation Plan (Phase 2B)
|
||||
|
||||
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:executing-plans (INLINE, with checkpoints).
|
||||
> **Do NOT run subagent-driven** — this touches live git access, the container registry, and the cockpit
|
||||
> deploy; each cutover step needs a human-confirmed gate. Steps use checkbox (`- [ ]`) syntax.
|
||||
|
||||
**Goal:** Stand up Gitea (chart + existing Postgres) as the git host behind the unchanged `git.fxhnt.ai`,
|
||||
migrate both repos, and move the cockpit image to Scaleway Container Registry — so GitLab can be removed in 2C.
|
||||
|
||||
**Architecture:** Gitea comes up internal-only; repos migrate + validate over a port-forward; the cockpit
|
||||
image moves to Scaleway registry and is build+rollout tested; THEN `git.fxhnt.ai` (HTTPS + SSH `:2222`)
|
||||
cuts over to Gitea in one switch; finally a Gitea webhook → Argo Events auto-deploys the cockpit. GitLab
|
||||
stays running (unaddressed) as the rollback net until 2C.
|
||||
|
||||
**Tech Stack:** Gitea Helm chart (`gitea-charts/gitea`), existing in-cluster PostgreSQL, nginx+socat
|
||||
Tailscale proxy, Scaleway Container Registry, Argo Events, kaniko.
|
||||
|
||||
**Spec:** `docs/superpowers/specs/2026-06-21-gitea-replace-gitlab-design.md`
|
||||
|
||||
**TWO REPOS:** infra changes → `foxhunt` (`~/Work/foxhunt`); cockpit build/deploy → `fxhnt`
|
||||
(`~/Work/fxhnt`). Each gets its own branch + commits.
|
||||
|
||||
---
|
||||
|
||||
## Shared env
|
||||
```bash
|
||||
export SCW_ACCESS_KEY=$(kubectl get secret scaleway-credentials -n foxhunt -o jsonpath='{.data.access-key}'|base64 -d)
|
||||
export SCW_SECRET_KEY=$(kubectl get secret scaleway-credentials -n foxhunt -o jsonpath='{.data.secret-key}'|base64 -d)
|
||||
export SCW_DEFAULT_PROJECT_ID=$(kubectl get secret scaleway-credentials -n foxhunt -o jsonpath='{.data.project-id}'|base64 -d)
|
||||
export SCW_DEFAULT_REGION=fr-par SCW_DEFAULT_ZONE=fr-par-2
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Task 0: Branches
|
||||
- [ ] **Step 1**
|
||||
```bash
|
||||
cd /home/jgrusewski/Work/foxhunt && git checkout -b chore/gitea-replace-gitlab
|
||||
cd /home/jgrusewski/Work/fxhnt && git checkout -b chore/gitea-cockpit-registry
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Task 1: Postgres `gitea` DB + secrets
|
||||
|
||||
**Files:** none committed (cluster secrets). Record secret names only.
|
||||
|
||||
- [ ] **Step 1: Read existing postgres superuser creds**
|
||||
```bash
|
||||
PGU=$(kubectl get deploy postgres -n foxhunt -o jsonpath='{.spec.template.spec.containers[0].env[?(@.name=="POSTGRES_USER")].value}')
|
||||
PGPOD=$(kubectl get pod -n foxhunt -l app=postgres -o jsonpath='{.items[0].metadata.name}' 2>/dev/null || kubectl get pod -n foxhunt | grep '^postgres-' | grep -v backup | awk '{print $1}' | head -1)
|
||||
echo "postgres user=$PGU pod=$PGPOD"
|
||||
```
|
||||
Expected: a username + a running pod name. (If `POSTGRES_USER` is via secretRef not value, read it from the referenced secret instead.)
|
||||
|
||||
- [ ] **Step 2: Generate a Gitea DB password + create role/db**
|
||||
```bash
|
||||
GITEA_DB_PW=$(openssl rand -base64 24 | tr -d '/+=' | head -c 24)
|
||||
kubectl exec -n foxhunt "$PGPOD" -- psql -U "$PGU" -v ON_ERROR_STOP=1 -c \
|
||||
"CREATE ROLE gitea LOGIN PASSWORD '$GITEA_DB_PW';" -c \
|
||||
"CREATE DATABASE gitea OWNER gitea ENCODING 'UTF8';" 2>&1 | grep -iE "CREATE|already exists|Error"
|
||||
echo "$GITEA_DB_PW" # capture for the next step
|
||||
```
|
||||
Expected: `CREATE ROLE` + `CREATE DATABASE` (or `already exists` if re-run).
|
||||
|
||||
- [ ] **Step 3: Create the k8s secrets Gitea needs** (DB creds + admin)
|
||||
```bash
|
||||
GITEA_ADMIN_PW=$(openssl rand -base64 24 | tr -d '/+=' | head -c 24)
|
||||
kubectl create secret generic gitea-db -n foxhunt \
|
||||
--from-literal=password="$GITEA_DB_PW" --dry-run=client -o yaml | kubectl apply -f -
|
||||
kubectl create secret generic gitea-admin -n foxhunt \
|
||||
--from-literal=username=gitadmin --from-literal=password="$GITEA_ADMIN_PW" \
|
||||
--from-literal=email=jeroen@bizworx.nl --dry-run=client -o yaml | kubectl apply -f -
|
||||
echo "ADMIN PW (save in your password manager): $GITEA_ADMIN_PW"
|
||||
```
|
||||
Expected: both secrets `created`/`configured`. **Save the admin password** — needed for Gitea login.
|
||||
|
||||
---
|
||||
|
||||
## Task 2: Gitea Helm values + install (internal-only)
|
||||
|
||||
**Files:** Create `infra/k8s/gitea/values.yaml` (foxhunt repo).
|
||||
|
||||
- [ ] **Step 1: Add the Gitea Helm repo**
|
||||
```bash
|
||||
helm repo add gitea-charts https://dl.gitea.com/charts/ && helm repo update gitea-charts
|
||||
helm search repo gitea-charts/gitea --versions | head -3
|
||||
```
|
||||
Expected: lists chart versions (use the latest stable in Step 3).
|
||||
|
||||
- [ ] **Step 2: Write `infra/k8s/gitea/values.yaml`**
|
||||
```yaml
|
||||
# Gitea — lightweight git host replacing GitLab (Phase 2B). External Postgres (existing in-cluster
|
||||
# `postgres`), no bundled DB/redis/memcached, Actions off (Argo does CI). Internal-only until cutover.
|
||||
replicaCount: 1
|
||||
image:
|
||||
rootless: true
|
||||
|
||||
# Disable all bundled subcharts — reuse the existing in-cluster postgres
|
||||
postgresql:
|
||||
enabled: false
|
||||
postgresql-ha:
|
||||
enabled: false
|
||||
redis-cluster:
|
||||
enabled: false
|
||||
redis:
|
||||
enabled: false
|
||||
|
||||
persistence:
|
||||
enabled: true
|
||||
size: 5Gi
|
||||
storageClass: sbs-default-retain
|
||||
|
||||
resources:
|
||||
requests: { cpu: 100m, memory: 128Mi }
|
||||
limits: { cpu: "1", memory: 512Mi }
|
||||
|
||||
service:
|
||||
http: { type: ClusterIP, port: 3000 }
|
||||
ssh: { type: ClusterIP, port: 22 }
|
||||
|
||||
actions:
|
||||
enabled: false
|
||||
|
||||
gitea:
|
||||
admin:
|
||||
existingSecret: gitea-admin
|
||||
config:
|
||||
server:
|
||||
ROOT_URL: https://git.fxhnt.ai/
|
||||
DOMAIN: git.fxhnt.ai
|
||||
SSH_DOMAIN: git.fxhnt.ai
|
||||
SSH_PORT: "22" # clean git@git.fxhnt.ai clone URLs. Port 22 is free (nothing host-SSHes
|
||||
# the git node); :2222 retired, local remotes updated in Task 5 Step 2b.
|
||||
DISABLE_SSH: "false"
|
||||
database:
|
||||
DB_TYPE: postgres
|
||||
HOST: postgres.foxhunt.svc.cluster.local:5432
|
||||
NAME: gitea
|
||||
USER: gitea
|
||||
service:
|
||||
DISABLE_REGISTRATION: "true"
|
||||
cache:
|
||||
ADAPTER: memory
|
||||
additionalConfigFromEnvs:
|
||||
- name: GITEA__database__PASSWD
|
||||
valueFrom:
|
||||
secretKeyRef: { name: gitea-db, key: password }
|
||||
```
|
||||
|
||||
- [ ] **Step 3: Install (internal-only — no public hostname yet)**
|
||||
```bash
|
||||
cd /home/jgrusewski/Work/foxhunt
|
||||
helm install gitea gitea-charts/gitea -n foxhunt -f infra/k8s/gitea/values.yaml --wait --timeout 5m 2>&1 | tail -5
|
||||
```
|
||||
Expected: `STATUS: deployed`. (If `--wait` times out, continue; verify pod in Step 4.)
|
||||
|
||||
- [ ] **Step 4: Verify Gitea is healthy**
|
||||
```bash
|
||||
kubectl get pods -n foxhunt | grep gitea
|
||||
kubectl run gitea-probe --rm -i --restart=Never -n foxhunt --image=curlimages/curl -- \
|
||||
curl -s http://gitea-http.foxhunt.svc.cluster.local:3000/api/healthz
|
||||
```
|
||||
Expected: a JSON health blob with `"status": "pass"`. Pod `Running`.
|
||||
|
||||
- [ ] **Step 5: Commit values (foxhunt repo)**
|
||||
```bash
|
||||
cd /home/jgrusewski/Work/foxhunt && git add infra/k8s/gitea/values.yaml
|
||||
git commit -m "feat(infra): Gitea Helm values (external postgres, internal-only) — Phase 2B"
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Task 3: Migrate both repos via port-forward + validate
|
||||
|
||||
**Files:** none committed (git data operations).
|
||||
|
||||
- [ ] **Step 1: Port-forward Gitea + create the repos via API**
|
||||
```bash
|
||||
kubectl port-forward -n foxhunt svc/gitea-http 3000:3000 >/tmp/gitea-pf.log 2>&1 & echo $! >/tmp/gitea-pf.pid
|
||||
sleep 3
|
||||
GA=gitadmin; GP=$(kubectl get secret gitea-admin -n foxhunt -o jsonpath='{.data.password}'|base64 -d)
|
||||
for r in fxhnt foxhunt; do
|
||||
curl -s -u "$GA:$GP" -X POST http://localhost:3000/api/v1/user/repos \
|
||||
-H 'Content-Type: application/json' -d "{\"name\":\"$r\",\"private\":true}" -o /dev/null -w "$r: %{http_code}\n"
|
||||
done
|
||||
```
|
||||
Expected: `fxhnt: 201` and `foxhunt: 201`.
|
||||
|
||||
- [ ] **Step 2: Mirror-push both repos from local clones**
|
||||
```bash
|
||||
GA=gitadmin; GP=$(kubectl get secret gitea-admin -n foxhunt -o jsonpath='{.data.password}'|base64 -d)
|
||||
( cd /home/jgrusewski/Work/fxhnt && git push --mirror "http://$GA:$GP@localhost:3000/$GA/fxhnt.git" 2>&1 | tail -2 )
|
||||
( cd /home/jgrusewski/Work/foxhunt && git push --mirror "http://$GA:$GP@localhost:3000/$GA/foxhunt.git" 2>&1 | tail -2 )
|
||||
```
|
||||
Expected: both report refs pushed (no error). foxhunt is 876 MB — may take a minute.
|
||||
|
||||
- [ ] **Step 3: VALIDATE — clone back + diff top SHAs vs GitLab**
|
||||
```bash
|
||||
GA=gitadmin; GP=$(kubectl get secret gitea-admin -n foxhunt -o jsonpath='{.data.password}'|base64 -d)
|
||||
for r in fxhnt foxhunt; do
|
||||
gitea_sha=$(git ls-remote "http://$GA:$GP@localhost:3000/$GA/$r.git" HEAD | awk '{print $1}')
|
||||
gitlab_sha=$(cd /home/jgrusewski/Work/$r && git ls-remote origin HEAD | awk '{print $1}')
|
||||
[ "$gitea_sha" = "$gitlab_sha" ] && echo "$r: MATCH ($gitea_sha)" || echo "$r: MISMATCH gitea=$gitea_sha gitlab=$gitlab_sha"
|
||||
done
|
||||
```
|
||||
Expected: `fxhnt: MATCH` and `foxhunt: MATCH`. **STOP if any MISMATCH.**
|
||||
|
||||
- [ ] **Step 4: Mark foxhunt archived (read-only)**
|
||||
```bash
|
||||
GA=gitadmin; GP=$(kubectl get secret gitea-admin -n foxhunt -o jsonpath='{.data.password}'|base64 -d)
|
||||
curl -s -u "$GA:$GP" -X PATCH "http://localhost:3000/api/v1/repos/$GA/foxhunt" \
|
||||
-H 'Content-Type: application/json' -d '{"archived":true}' -o /dev/null -w "archive foxhunt: %{http_code}\n"
|
||||
kill $(cat /tmp/gitea-pf.pid) 2>/dev/null
|
||||
```
|
||||
Expected: `archive foxhunt: 200`.
|
||||
|
||||
---
|
||||
|
||||
## Task 4: Move cockpit image to Scaleway Container Registry
|
||||
|
||||
**Files:** Modify `~/Work/fxhnt/infra/argo/cockpit-build-deploy.yaml`,
|
||||
`~/Work/fxhnt/infra/k8s/orchestration/dagster.yaml` (+ dashboard manifest if separate).
|
||||
|
||||
- [ ] **Step 1: Create the Scaleway registry pull/push secret**
|
||||
```bash
|
||||
kubectl create secret docker-registry scw-registry -n foxhunt \
|
||||
--docker-server=rg.fr-par.scw.cloud \
|
||||
--docker-username=nologin --docker-password="$SCW_SECRET_KEY" \
|
||||
--dry-run=client -o yaml | kubectl apply -f -
|
||||
```
|
||||
Expected: secret `created`/`configured`. (Scaleway registry auth: username `nologin`, password = a
|
||||
Scaleway secret key.)
|
||||
|
||||
- [ ] **Step 2: Repoint kaniko build → Scaleway registry**
|
||||
|
||||
In `~/Work/fxhnt/infra/argo/cockpit-build-deploy.yaml`, change the kaniko args:
|
||||
```
|
||||
--destination=rg.fr-par.scw.cloud/bizworx/fxhnt-cockpit:latest \
|
||||
--cache-repo=rg.fr-par.scw.cloud/bizworx/cache \
|
||||
```
|
||||
Remove the `--insecure-registry=…gitlab-registry…:5000` and `--skip-tls-verify-registry=…` lines
|
||||
(Scaleway is TLS). Mount the `scw-registry` dockerconfig for kaniko (replace the `gitlab-registry`
|
||||
dockerconfig volume): set the kaniko `DOCKER_CONFIG`/volume to the `scw-registry` secret at
|
||||
`/kaniko/.docker/config.json`.
|
||||
|
||||
- [ ] **Step 3: Repoint the deploy image refs**
|
||||
|
||||
In `~/Work/fxhnt/infra/k8s/orchestration/dagster.yaml` (and the dashboard manifest), change every
|
||||
`image: gitlab-registry.foxhunt.svc.cluster.local:5000/root/foxhunt/fxhnt-cockpit:latest` →
|
||||
`image: rg.fr-par.scw.cloud/bizworx/fxhnt-cockpit:latest`, and add under each pod spec:
|
||||
```yaml
|
||||
imagePullSecrets:
|
||||
- name: scw-registry
|
||||
```
|
||||
|
||||
- [ ] **Step 4: Run one cockpit build → Scaleway + verify the image exists**
|
||||
```bash
|
||||
cd /home/jgrusewski/Work/fxhnt && ./scripts/argo-deploy-cockpit.sh 2>&1 | tail -8
|
||||
# verify the pushed image
|
||||
scw registry image list namespace-id=6561a90a-f4ba-4f44-b43f-be6e27cbdfca 2>/dev/null | grep -i fxhnt-cockpit || echo "(verify via Scaleway console)"
|
||||
```
|
||||
Expected: workflow `Succeeded`; `fxhnt-cockpit` image listed in the `bizworx` namespace.
|
||||
|
||||
- [ ] **Step 5: Verify a rollout pulls from Scaleway**
|
||||
```bash
|
||||
kubectl rollout restart deploy/dagster deploy/fxhnt-dashboard -n foxhunt 2>/dev/null
|
||||
kubectl rollout status deploy/dagster -n foxhunt --timeout=180s
|
||||
kubectl get pods -n foxhunt -o jsonpath='{range .items[*]}{.spec.containers[*].image}{"\n"}{end}' | grep -i "rg.fr-par.scw.cloud/bizworx/fxhnt-cockpit" | head
|
||||
```
|
||||
Expected: rollout succeeds; pods now reference the `rg.fr-par.scw.cloud/bizworx/...` image (no `ImagePullBackOff`).
|
||||
|
||||
- [ ] **Step 6: Commit (fxhnt repo)**
|
||||
```bash
|
||||
cd /home/jgrusewski/Work/fxhnt
|
||||
git add infra/argo/cockpit-build-deploy.yaml infra/k8s/orchestration/dagster.yaml
|
||||
git commit -m "feat(infra): cockpit image -> Scaleway Container Registry (Phase 2B)"
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Task 5: Cut `git.fxhnt.ai` + SSH `:2222` over to Gitea
|
||||
|
||||
**Files:** Modify `infra/k8s/gitlab/tailscale-proxy.yaml` (foxhunt repo). **CHECKPOINT before applying.**
|
||||
|
||||
- [ ] **Step 1: Repoint nginx `git.fxhnt.ai:443` → Gitea**
|
||||
|
||||
In `infra/k8s/gitlab/tailscale-proxy.yaml`, in the `server_name git.fxhnt.ai;` (`listen 443`) block,
|
||||
change:
|
||||
```
|
||||
proxy_pass http://gitlab-webservice-default.foxhunt.svc.cluster.local:8181;
|
||||
```
|
||||
to:
|
||||
```
|
||||
proxy_pass http://gitea-http.foxhunt.svc.cluster.local:3000;
|
||||
```
|
||||
Delete the entire `listen 5050 ssl; server_name git.fxhnt.ai;` registry block (Scaleway registry is
|
||||
external — not proxied).
|
||||
|
||||
- [ ] **Step 2: Repoint the socat SSH proxy to Gitea SSH on `:22` (drop `:2222`)**
|
||||
|
||||
Nothing host-SSHes the git node, so use the standard port `22` and retire `:2222` entirely (clean
|
||||
`git@git.fxhnt.ai` URLs). In the `ssh-proxy` container args, change the listen port to `22` and the target
|
||||
to Gitea:
|
||||
```
|
||||
- "TCP-LISTEN:22,fork,reuseaddr,sndbuf=1048576,rcvbuf=1048576"
|
||||
- "TCP:gitea-ssh.foxhunt.svc.cluster.local:22,sndbuf=1048576,rcvbuf=1048576"
|
||||
```
|
||||
Update the container `ports:` entry from `containerPort: 2222` → `containerPort: 22`. Ensure the Tailscale
|
||||
node/proxy advertises port `22` (update the tailnet sidecar's exposed ports / `serve` config: `22`,
|
||||
replacing `2222`).
|
||||
|
||||
- [ ] **Step 2b: Update local git remotes from `:2222` → default (port 22)**
|
||||
```bash
|
||||
cd /home/jgrusewski/Work/foxhunt && git remote set-url origin ssh://git@git.fxhnt.ai/gitadmin/foxhunt.git
|
||||
cd /home/jgrusewski/Work/fxhnt && git remote set-url origin ssh://git@git.fxhnt.ai/gitadmin/fxhnt.git
|
||||
cd /home/jgrusewski/Work/foxhunt && git remote -v | grep fetch
|
||||
```
|
||||
Expected: remotes now show `ssh://git@git.fxhnt.ai/...` (no `:2222`).
|
||||
|
||||
- [ ] **Step 3: Apply + restart the proxy**
|
||||
```bash
|
||||
cd /home/jgrusewski/Work/foxhunt
|
||||
kubectl apply -f infra/k8s/gitlab/tailscale-proxy.yaml
|
||||
kubectl rollout restart deploy/tailscale-gitlab-proxy -n foxhunt
|
||||
kubectl rollout status deploy/tailscale-gitlab-proxy -n foxhunt --timeout=120s
|
||||
```
|
||||
Expected: rollout complete.
|
||||
|
||||
- [ ] **Step 4: GATE — verify `git.fxhnt.ai` serves Gitea**
|
||||
```bash
|
||||
curl -sk https://git.fxhnt.ai/api/healthz | head -c 200; echo
|
||||
echo "--- SSH on :22 (git@git.fxhnt.ai) ---"; git ls-remote ssh://git@git.fxhnt.ai/gitadmin/fxhnt.git HEAD 2>&1 | head -1
|
||||
```
|
||||
Expected: Gitea healthz JSON; the `:22` ssh ls-remote returns the HEAD sha (proves Gitea web + SSH on the
|
||||
canonical name). **STOP + rollback (revert this file, re-apply) if either fails.**
|
||||
|
||||
- [ ] **Step 5: Commit (foxhunt repo)**
|
||||
```bash
|
||||
cd /home/jgrusewski/Work/foxhunt && git add infra/k8s/gitlab/tailscale-proxy.yaml
|
||||
git commit -m "feat(infra): cut git.fxhnt.ai + SSH :2222 over to Gitea; drop :5050 registry (Phase 2B)"
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Task 6: Wire Gitea push webhook → Argo Events auto-deploy
|
||||
|
||||
**Files:** Create `infra/k8s/argo/events/gitea-push-eventsource.yaml` +
|
||||
`infra/k8s/argo/events/gitea-deploy-sensor.yaml` (foxhunt repo).
|
||||
|
||||
- [ ] **Step 1: Create the webhook eventsource**
|
||||
|
||||
Create `infra/k8s/argo/events/gitea-push-eventsource.yaml`:
|
||||
```yaml
|
||||
apiVersion: argoproj.io/v1alpha1
|
||||
kind: EventSource
|
||||
metadata:
|
||||
name: gitea-push
|
||||
namespace: foxhunt
|
||||
spec:
|
||||
service:
|
||||
ports:
|
||||
- port: 12000
|
||||
targetPort: 12000
|
||||
webhook:
|
||||
fxhnt-push:
|
||||
port: "12000"
|
||||
endpoint: /push
|
||||
method: POST
|
||||
```
|
||||
|
||||
- [ ] **Step 2: Create the sensor that submits the cockpit workflow**
|
||||
|
||||
Create `infra/k8s/argo/events/gitea-deploy-sensor.yaml`:
|
||||
```yaml
|
||||
apiVersion: argoproj.io/v1alpha1
|
||||
kind: Sensor
|
||||
metadata:
|
||||
name: gitea-deploy
|
||||
namespace: foxhunt
|
||||
spec:
|
||||
dependencies:
|
||||
- name: push
|
||||
eventSourceName: gitea-push
|
||||
eventName: fxhnt-push
|
||||
filters:
|
||||
data:
|
||||
- path: body.ref
|
||||
type: string
|
||||
value: ["refs/heads/main", "refs/heads/master"]
|
||||
triggers:
|
||||
- template:
|
||||
name: submit-cockpit
|
||||
argoWorkflow:
|
||||
operation: submit
|
||||
source:
|
||||
resource:
|
||||
apiVersion: argoproj.io/v1alpha1
|
||||
kind: Workflow
|
||||
metadata:
|
||||
generateName: fxhnt-cockpit-
|
||||
namespace: foxhunt
|
||||
spec:
|
||||
workflowTemplateRef:
|
||||
name: fxhnt-cockpit
|
||||
```
|
||||
|
||||
- [ ] **Step 3: Apply + register the webhook in Gitea**
|
||||
```bash
|
||||
cd /home/jgrusewski/Work/foxhunt
|
||||
kubectl apply -f infra/k8s/argo/events/gitea-push-eventsource.yaml -f infra/k8s/argo/events/gitea-deploy-sensor.yaml
|
||||
kubectl get pods -n foxhunt | grep -iE "gitea-push|gitea-deploy"
|
||||
# register webhook (in-cluster URL) on the fxhnt repo
|
||||
GA=gitadmin; GP=$(kubectl get secret gitea-admin -n foxhunt -o jsonpath='{.data.password}'|base64 -d)
|
||||
kubectl port-forward -n foxhunt svc/gitea-http 3000:3000 >/tmp/gitea-pf.log 2>&1 & echo $! >/tmp/gitea-pf.pid; sleep 3
|
||||
curl -s -u "$GA:$GP" -X POST "http://localhost:3000/api/v1/repos/$GA/fxhnt/hooks" -H 'Content-Type: application/json' -d '{
|
||||
"type":"gitea","active":true,"events":["push"],
|
||||
"config":{"url":"http://gitea-push-eventsource-svc.foxhunt.svc.cluster.local:12000/push","content_type":"json"}
|
||||
}' -o /dev/null -w "hook: %{http_code}\n"
|
||||
kill $(cat /tmp/gitea-pf.pid) 2>/dev/null
|
||||
```
|
||||
Expected: eventsource + sensor pods `Running`; `hook: 201`.
|
||||
|
||||
- [ ] **Step 4: GATE — test push triggers a deploy**
|
||||
```bash
|
||||
cd /home/jgrusewski/Work/fxhnt
|
||||
git commit --allow-empty -m "test: trigger Gitea->Argo cockpit deploy"
|
||||
git push origin main # origin already = git.fxhnt.ai (now Gitea)
|
||||
sleep 10
|
||||
kubectl get wf -n foxhunt --sort-by=.metadata.creationTimestamp 2>/dev/null | grep fxhnt-cockpit | tail -2
|
||||
```
|
||||
Expected: a new `fxhnt-cockpit-*` workflow appears (triggered by the push). **If none fires**, check the
|
||||
eventsource pod logs; fallback is the manual `./scripts/argo-deploy-cockpit.sh` (still works).
|
||||
|
||||
- [ ] **Step 5: Commit (foxhunt repo)**
|
||||
```bash
|
||||
cd /home/jgrusewski/Work/foxhunt && git add infra/k8s/argo/events/gitea-push-eventsource.yaml infra/k8s/argo/events/gitea-deploy-sensor.yaml
|
||||
git commit -m "feat(infra): Gitea push webhook -> Argo Events cockpit auto-deploy (Phase 2B)"
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Task 7: Final verification
|
||||
|
||||
- [ ] **Step 1: Footprint + health summary**
|
||||
```bash
|
||||
echo "=== Gitea (new) ==="; kubectl top pods -n foxhunt 2>/dev/null | grep -i gitea
|
||||
echo "=== cockpit healthy on Scaleway image ==="; curl -sk -o /dev/null -w "dashboard.fxhnt.ai: %{http_code}\n" https://dashboard.fxhnt.ai
|
||||
echo "=== git.fxhnt.ai serves Gitea ==="; curl -sk https://git.fxhnt.ai/api/healthz | head -c 80; echo
|
||||
echo "=== GitLab still running (fallback for 2C) ==="; kubectl get pods -n foxhunt | grep -c gitlab | xargs echo "gitlab pods:"
|
||||
```
|
||||
Expected: Gitea ≤~300 MiB; dashboard 200; git.fxhnt.ai healthz pass; GitLab pods still present.
|
||||
|
||||
---
|
||||
|
||||
## Rollback
|
||||
- **Cutover (Task 5):** `git checkout infra/k8s/gitlab/tailscale-proxy.yaml` + `kubectl apply` + rollout
|
||||
restart → `git.fxhnt.ai`/SSH back to GitLab.
|
||||
- **Registry (Task 4):** revert the fxhnt manifests + rollout restart → pods pull the GitLab image again
|
||||
(still present until 2C).
|
||||
- **Gitea:** `helm uninstall gitea -n foxhunt`; drop the `gitea` DB. GitLab was never touched.
|
||||
|
||||
---
|
||||
|
||||
## Acceptance criteria (from spec)
|
||||
- Gitea running on existing postgres, healthz pass, ≤~300 MiB. ✅ Task 2, Task 7
|
||||
- Both repos in Gitea, SHAs match GitLab, foxhunt archived. ✅ Task 3
|
||||
- Cockpit builds to + deploys from Scaleway registry; dashboard 200. ✅ Task 4, Task 7
|
||||
- `git.fxhnt.ai` (HTTPS + SSH :2222) serves Gitea; remotes unchanged. ✅ Task 5
|
||||
- Push to fxhnt main auto-triggers a cockpit deploy. ✅ Task 6
|
||||
- GitLab still running as fallback. ✅ Task 7
|
||||
```
|
||||
264
docs/superpowers/plans/2026-06-21-tfstate-to-scaleway.md
Normal file
264
docs/superpowers/plans/2026-06-21-tfstate-to-scaleway.md
Normal file
@@ -0,0 +1,264 @@
|
||||
# Migrate Terraform State to Scaleway Object Storage — Implementation Plan (Phase 2A)
|
||||
|
||||
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:executing-plans (INLINE, with checkpoints).
|
||||
> **Do NOT run subagent-driven** — state migration is irreversible-adjacent and each module needs a
|
||||
> human-confirmed `plan = No changes` gate. Steps use checkbox (`- [ ]`) syntax.
|
||||
|
||||
**Goal:** Move the 4 production modules' Terraform state from the self-hosted GitLab http backend to
|
||||
external Scaleway Object Storage (bucket `foxhunt-tfstate`), with native S3 lockfile locking.
|
||||
|
||||
**Architecture:** Back up all 4 states from GitLab first; flip the single `remote_state` block in
|
||||
`infra/live/production/root.hcl` from `http`→`s3`; migrate each module in-place with `terragrunt init
|
||||
-migrate-state`; gate each on `terragrunt plan = No changes`. GitLab state is left intact as a fallback
|
||||
(deleted later in 2C).
|
||||
|
||||
**Tech Stack:** terragrunt v0.77.20 + OpenTofu v1.11.5, Scaleway Object Storage (S3-compatible), AWS s3
|
||||
backend.
|
||||
|
||||
**Spec:** `docs/superpowers/specs/2026-06-21-tfstate-to-scaleway-design.md`
|
||||
|
||||
---
|
||||
|
||||
## Shared env (export in EVERY shell that runs terragrunt during this plan)
|
||||
The s3 backend reads **`AWS_*`** env vars; the Scaleway provider reads **`SCW_*`**; reading the OLD
|
||||
GitLab state during `-migrate-state` needs **`TF_HTTP_*`**. Export all three sets:
|
||||
```bash
|
||||
cd /home/jgrusewski/Work/foxhunt
|
||||
export TG_TF_PATH=tofu TERRAGRUNT_TFPATH=tofu
|
||||
# Scaleway provider creds
|
||||
export SCW_ACCESS_KEY=$(kubectl get secret scaleway-credentials -n foxhunt -o jsonpath='{.data.access-key}'|base64 -d)
|
||||
export SCW_SECRET_KEY=$(kubectl get secret scaleway-credentials -n foxhunt -o jsonpath='{.data.secret-key}'|base64 -d)
|
||||
export SCW_DEFAULT_PROJECT_ID=$(kubectl get secret scaleway-credentials -n foxhunt -o jsonpath='{.data.project-id}'|base64 -d)
|
||||
export SCW_DEFAULT_REGION=fr-par SCW_DEFAULT_ZONE=fr-par-2
|
||||
# s3 backend creds (SAME Scaleway keys, AWS-style names)
|
||||
export AWS_ACCESS_KEY_ID="$SCW_ACCESS_KEY"
|
||||
export AWS_SECRET_ACCESS_KEY="$SCW_SECRET_KEY"
|
||||
export AWS_REGION=fr-par
|
||||
# OLD GitLab http backend creds (needed to READ existing state during migrate)
|
||||
export TF_HTTP_USERNAME=root
|
||||
export TF_HTTP_PASSWORD=$(kubectl get secret gitlab-pat -n foxhunt -o jsonpath='{.data.token}'|base64 -d)
|
||||
echo "env: SCW=${SCW_ACCESS_KEY:0:6}… AWS=${AWS_ACCESS_KEY_ID:0:6}… PAT=${TF_HTTP_PASSWORD:0:4}…"
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Task 0: Branch
|
||||
- [ ] **Step 1**
|
||||
```bash
|
||||
cd /home/jgrusewski/Work/foxhunt && git checkout -b chore/tfstate-to-scaleway
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Task 1: Verify prerequisites + back up all 4 GitLab states
|
||||
|
||||
**Files:** none (read-only + /tmp backups)
|
||||
|
||||
- [ ] **Step 1: Confirm the target bucket exists + is empty-ish**
|
||||
```bash
|
||||
scw object bucket list 2>/dev/null | grep foxhunt-tfstate || echo "MISSING — create: scw object bucket create name=foxhunt-tfstate region=fr-par"
|
||||
```
|
||||
Expected: a line containing `foxhunt-tfstate`. If MISSING, run the create shown, then re-check.
|
||||
|
||||
- [ ] **Step 2: Confirm tooling versions**
|
||||
```bash
|
||||
tofu version | head -1; terragrunt --version | head -1
|
||||
```
|
||||
Expected: OpenTofu `v1.11.x` (≥1.10 for `use_lockfile`), terragrunt `v0.77.x`.
|
||||
|
||||
- [ ] **Step 3: Back up each module's CURRENT (GitLab) state** — run with shared env exported
|
||||
```bash
|
||||
for m in kapsule public-gateway dns block-storage; do
|
||||
( cd infra/live/production/$m && terragrunt state pull > /tmp/tfstate-backup-$m.tfstate 2>/dev/null )
|
||||
echo "$m: $(python3 -c "import json;d=json.load(open('/tmp/tfstate-backup-$m.tfstate'));print('serial',d.get('serial'),'resources',len(d.get('resources',[])))" 2>/dev/null || echo 'EMPTY/ERROR')"
|
||||
done
|
||||
```
|
||||
Expected: each module prints a non-zero `resources` count (e.g. kapsule ~5+, dns/public-gateway/block-storage ≥1). **STOP if any backup is EMPTY/ERROR** — that module's GitLab state isn't readable; do not migrate it blind.
|
||||
|
||||
- [ ] **Step 4: Record current state serials for post-migration comparison**
|
||||
```bash
|
||||
grep -H '"serial"' /tmp/tfstate-backup-*.tfstate 2>/dev/null || for m in kapsule public-gateway dns block-storage; do echo "$m $(python3 -c "import json;print(json.load(open('/tmp/tfstate-backup-$m.tfstate'))['serial'])")"; done
|
||||
```
|
||||
Expected: a serial number per module. Keep this output — the migrated state should have the same resource set.
|
||||
|
||||
---
|
||||
|
||||
## Task 2: Switch the backend in root.hcl (http → s3)
|
||||
|
||||
**Files:** Modify `infra/live/production/root.hcl`
|
||||
|
||||
- [ ] **Step 1: Replace the `remote_state` block**
|
||||
|
||||
In `infra/live/production/root.hcl`, replace the entire existing `remote_state { ... }` block (the `backend = "http"`
|
||||
GitLab block) with:
|
||||
```hcl
|
||||
# Remote state in Scaleway Object Storage (bucket foxhunt-tfstate, fr-par) — migrated off GitLab 2026-06-21.
|
||||
# s3 backend reads AWS_* env vars (set to the Scaleway access/secret keys); SCW_* drives the provider.
|
||||
remote_state {
|
||||
backend = "s3"
|
||||
generate = {
|
||||
path = "backend.tf"
|
||||
if_exists = "overwrite"
|
||||
}
|
||||
config = {
|
||||
bucket = "foxhunt-tfstate"
|
||||
key = "${path_relative_to_include()}/terraform.tfstate"
|
||||
region = "fr-par"
|
||||
|
||||
endpoints = {
|
||||
s3 = "https://s3.fr-par.scw.cloud"
|
||||
}
|
||||
|
||||
# Scaleway S3-compat: skip AWS-specific preflight calls
|
||||
skip_credentials_validation = true
|
||||
skip_region_validation = true
|
||||
skip_requesting_account_id = true
|
||||
skip_metadata_api_check = true
|
||||
|
||||
# OpenTofu-native lock (no DynamoDB)
|
||||
use_lockfile = true
|
||||
}
|
||||
}
|
||||
```
|
||||
Leave the `generate "provider"` block and `inputs`/`locals` untouched.
|
||||
|
||||
- [ ] **Step 2: Sanity-check the file parses**
|
||||
```bash
|
||||
cd /home/jgrusewski/Work/foxhunt
|
||||
grep -A2 'backend = "s3"' infra/live/production/root.hcl && grep -c 'backend = "http"' infra/live/production/root.hcl
|
||||
```
|
||||
Expected: shows the `s3` backend line; the `backend = "http"` count is `0`.
|
||||
|
||||
---
|
||||
|
||||
## Task 3: Migrate `kapsule` (canary) + lock smoke test
|
||||
|
||||
**Files:** none (state operation)
|
||||
|
||||
- [ ] **Step 1: Migrate state (with shared env exported)**
|
||||
```bash
|
||||
cd /home/jgrusewski/Work/foxhunt/infra/live/production/kapsule
|
||||
terragrunt init -migrate-state -force-copy -no-color 2>&1 | sed 's/.*tofu: //' | grep -iE "Successfully configured|migrat|Error|s3" | head
|
||||
```
|
||||
Expected: a line like `Terraform has been successfully migrated to the "s3" backend!` (or `Successfully
|
||||
configured the backend "s3"`). **If it errors on checksum** (Scaleway quirk), add `skip_s3_checksum = true`
|
||||
to the `config` block in `infra/live/production/root.hcl` and re-run this step.
|
||||
|
||||
- [ ] **Step 2: GATE — plan must show No changes**
|
||||
```bash
|
||||
terragrunt plan -input=false -no-color 2>&1 | sed 's/.*tofu: //' | grep -iE "No changes|^Plan:|Error" | head
|
||||
```
|
||||
Expected: `No changes. Your infrastructure matches the configuration.`
|
||||
**STOP if it shows any add/change/destroy** — the migrated state doesn't match reality. Revert (Task 7
|
||||
rollback) and investigate before continuing.
|
||||
|
||||
- [ ] **Step 3: Confirm the state object landed in the bucket**
|
||||
```bash
|
||||
scw object bucket list 2>/dev/null >/dev/null # ensure scw configured
|
||||
aws --endpoint-url https://s3.fr-par.scw.cloud s3 ls s3://foxhunt-tfstate/kapsule/ 2>/dev/null || \
|
||||
echo "(aws cli n/a — verify via: mc ls on a scw alias, or scw object)"
|
||||
```
|
||||
Expected: `terraform.tfstate` listed under `kapsule/`. (If `aws` CLI is absent, this is best-effort;
|
||||
the `plan = No changes` in Step 2 already proves the s3 backend is the live source.)
|
||||
|
||||
- [ ] **Step 4: Lock smoke test** — confirm `use_lockfile` works on Scaleway
|
||||
```bash
|
||||
cd /home/jgrusewski/Work/foxhunt/infra/live/production/kapsule
|
||||
# A refresh acquires + releases the lock; success + no leftover .tflock proves locking works
|
||||
terragrunt apply -refresh-only -auto-approve -input=false -no-color 2>&1 | sed 's/.*tofu: //' | grep -iE "Apply complete|Error|lock" | head
|
||||
aws --endpoint-url https://s3.fr-par.scw.cloud s3 ls s3://foxhunt-tfstate/kapsule/ 2>/dev/null | grep -i tflock && echo "WARN: orphan lockfile" || echo "OK: no orphan lockfile"
|
||||
```
|
||||
Expected: `Apply complete!` and `OK: no orphan lockfile`. **If locking errors**, fall back to
|
||||
`use_lockfile = false` (single-maintainer is safe) — document the change in the commit.
|
||||
|
||||
---
|
||||
|
||||
## Task 4: Migrate the remaining 3 modules (`public-gateway`, `dns`, `block-storage`)
|
||||
|
||||
**Files:** none (state operations). Same procedure as Task 3 Steps 1-2, per module, each gated.
|
||||
|
||||
- [ ] **Step 1: Migrate + gate `public-gateway`**
|
||||
```bash
|
||||
cd /home/jgrusewski/Work/foxhunt/infra/live/production/public-gateway
|
||||
terragrunt init -migrate-state -force-copy -no-color 2>&1 | sed 's/.*tofu: //' | grep -iE "migrat|successfully|Error" | head
|
||||
terragrunt plan -input=false -no-color 2>&1 | sed 's/.*tofu: //' | grep -iE "No changes|^Plan:|Error" | head
|
||||
```
|
||||
Expected: migration success line, then `No changes`. **STOP on any drift.**
|
||||
|
||||
- [ ] **Step 2: Migrate + gate `dns`**
|
||||
```bash
|
||||
cd /home/jgrusewski/Work/foxhunt/infra/live/production/dns
|
||||
terragrunt init -migrate-state -force-copy -no-color 2>&1 | sed 's/.*tofu: //' | grep -iE "migrat|successfully|Error" | head
|
||||
terragrunt plan -input=false -no-color 2>&1 | sed 's/.*tofu: //' | grep -iE "No changes|^Plan:|Error" | head
|
||||
```
|
||||
Expected: migration success line, then `No changes`. **STOP on any drift.**
|
||||
|
||||
- [ ] **Step 3: Migrate + gate `block-storage`**
|
||||
```bash
|
||||
cd /home/jgrusewski/Work/foxhunt/infra/live/production/block-storage
|
||||
terragrunt init -migrate-state -force-copy -no-color 2>&1 | sed 's/.*tofu: //' | grep -iE "migrat|successfully|Error" | head
|
||||
terragrunt plan -input=false -no-color 2>&1 | sed 's/.*tofu: //' | grep -iE "No changes|^Plan:|Error" | head
|
||||
```
|
||||
Expected: migration success line, then `No changes`. **STOP on any drift.**
|
||||
|
||||
---
|
||||
|
||||
## Task 5: Final verification + commit
|
||||
|
||||
**Files:** Modify (commit) `infra/live/production/root.hcl`; also commit generated `backend.tf` only if not gitignored.
|
||||
|
||||
- [ ] **Step 1: Re-plan all 4 modules from a clean shell** (proves the s3 backend is authoritative)
|
||||
```bash
|
||||
cd /home/jgrusewski/Work/foxhunt
|
||||
for m in kapsule public-gateway dns block-storage; do
|
||||
echo "=== $m ==="
|
||||
( cd infra/live/production/$m && terragrunt plan -input=false -no-color 2>&1 | sed 's/.*tofu: //' | grep -iE "No changes|^Plan:|Error" | head -1 )
|
||||
done
|
||||
```
|
||||
Expected: every module prints `No changes.`
|
||||
|
||||
- [ ] **Step 2: Confirm all 4 state objects exist in the bucket**
|
||||
```bash
|
||||
for m in kapsule public-gateway dns block-storage; do
|
||||
aws --endpoint-url https://s3.fr-par.scw.cloud s3 ls s3://foxhunt-tfstate/$m/ 2>/dev/null | grep -q terraform.tfstate && echo "$m: state present" || echo "$m: NOT FOUND (verify manually)"
|
||||
done
|
||||
```
|
||||
Expected: `state present` for all 4 (or manual confirmation if `aws` CLI absent — `plan = No changes` is the real proof).
|
||||
|
||||
- [ ] **Step 3: Verify GitLab state still intact (fallback preserved)** — should NOT be deleted in 2A
|
||||
```bash
|
||||
echo "GitLab TF state is intentionally left in place as a 2A fallback; deleted in Phase 2C."
|
||||
```
|
||||
|
||||
- [ ] **Step 4: Commit**
|
||||
```bash
|
||||
cd /home/jgrusewski/Work/foxhunt
|
||||
git add infra/live/production/root.hcl
|
||||
# backend.tf is terragrunt-generated; add only if tracked (usually gitignored)
|
||||
git status --porcelain infra/live/production/*/backend.tf 2>/dev/null
|
||||
git commit -m "chore(infra): migrate Terraform state GitLab -> Scaleway Object Storage (Phase 2A)
|
||||
|
||||
All 4 modules (kapsule/public-gateway/dns/block-storage) migrated to s3 backend
|
||||
(bucket foxhunt-tfstate, fr-par, use_lockfile). plan=No changes verified per module.
|
||||
GitLab http state left intact as fallback (removed in 2C).
|
||||
|
||||
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>"
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Rollback (if any module's post-migration plan shows drift)
|
||||
1. Restore the `http` backend block in `infra/live/production/root.hcl` (git checkout the file).
|
||||
2. Re-attach GitLab state: in each affected module, `terragrunt init -migrate-state -force-copy` (copies
|
||||
back s3→http) OR `terragrunt init -reconfigure` (the GitLab state was never deleted).
|
||||
3. If state is corrupted, restore from `/tmp/tfstate-backup-<module>.tfstate` via `terragrunt state push`.
|
||||
4. The GitLab copy is untouched throughout 2A, so reverting is always possible.
|
||||
|
||||
---
|
||||
|
||||
## Acceptance criteria (from spec)
|
||||
- `infra/live/production/root.hcl` uses the Scaleway s3 backend; committed. ✅ Task 2 + Task 5.4
|
||||
- All 4 modules: `terragrunt plan = No changes` on the new backend. ✅ Task 3.2, Task 4, Task 5.1
|
||||
- State objects at `foxhunt-tfstate/<module>/terraform.tfstate`. ✅ Task 3.3, Task 5.2
|
||||
- Lock smoke test passes. ✅ Task 3.4
|
||||
- GitLab http state intact (fallback). ✅ Task 5.3
|
||||
114
docs/superpowers/specs/2026-06-21-gitea-replace-gitlab-design.md
Normal file
114
docs/superpowers/specs/2026-06-21-gitea-replace-gitlab-design.md
Normal file
@@ -0,0 +1,114 @@
|
||||
# Stand up Gitea + migrate repos + move cockpit registry — Design (Phase 2B)
|
||||
|
||||
**Date:** 2026-06-21
|
||||
**Status:** Design (approved)
|
||||
**Repo:** foxhunt (IaC current home; relocation to fxhnt is Phase 2D)
|
||||
**Part of:** Phase 2 (GitLab → Gitea consolidation, [[project_phase2_gitlab_to_gitea]]). 2A (TF state →
|
||||
Scaleway) DONE. This is **2B**. 2C (GitLab removal) and 2D (IaC relocation) follow.
|
||||
|
||||
## Motivation
|
||||
|
||||
Self-hosted GitLab uses ~6.5 GiB RAM across 17 pods (webservice 1.96 GiB + 3× sidekiq ~3 GiB + gitaly +
|
||||
registry + prometheus + …). The active work is the Python `fxhnt` fund; GitLab is Rust-era overkill.
|
||||
Replace it with **Gitea** (~150–250 MiB, chart + existing Postgres) for git hosting, and move the
|
||||
**cockpit container image** to **Scaleway Container Registry** so GitLab can be removed in 2C without
|
||||
breaking the live cockpit deploy.
|
||||
|
||||
## Decisions (from brainstorming)
|
||||
|
||||
- **Gitea install:** `gitea/gitea` Helm chart. Subcharts `postgresql`/`redis-cluster`/`redis`/`memcached`
|
||||
**disabled**. Built-in **Actions disabled** (Argo does CI/CD). Cache = memory (single replica).
|
||||
- **Database:** reuse the **existing in-cluster `postgres`** (svc `postgres:5432`, already has a backup
|
||||
cronjob). Create a dedicated `gitea` database + role.
|
||||
- **Repos:** migrate **both** — `fxhnt` (active) and `foxhunt` (876 MB, flagged **archived/read-only**) —
|
||||
via `git push --mirror` from existing local clones.
|
||||
- **Canonical hostname stays `git.fxhnt.ai`** (no permanent `gitea.fxhnt.ai`, no remote-URL churn).
|
||||
Gitea comes up internal-only; validate over a port-forward; then cut `git.fxhnt.ai` (HTTPS) + SSH over
|
||||
to Gitea in one switch. **SSH on `:22`** (nothing host-SSHes the git node, so the standard port is free);
|
||||
clone URLs become the clean `git@git.fxhnt.ai` and `:2222` is retired (local remotes updated once).
|
||||
GitLab keeps running (unaddressed) as the 2C-removal fallback.
|
||||
- **Container registry:** cockpit image → **Scaleway Container Registry** (`rg.fr-par.scw.cloud/bizworx/`,
|
||||
namespace exists). Retire the in-cluster GitLab registry. External/managed, no registry pod.
|
||||
- **Argo integration:** Gitea push webhook → Argo Events webhook eventsource → sensor → submit
|
||||
`wftmpl/fxhnt-cockpit`. Replaces the old `gitlab-push-eventsource`.
|
||||
|
||||
## Components
|
||||
|
||||
1. **`infra/k8s/gitea/` (new):** Helm values `values.yaml` + a thin install (release `gitea`, ns
|
||||
`foxhunt`). Storage: 1 PVC `sbs-default-retain` ~5 Gi (foxhunt 876 MB + fxhnt + headroom). Resources:
|
||||
requests 128 Mi/100m, limits 512 Mi. Admin user + secrets from a k8s secret (NOT committed).
|
||||
Gitea config: `server.SSH_PORT`/`SSH_LISTEN_PORT`, `server.ROOT_URL=https://git.fxhnt.ai/`,
|
||||
`server.DISABLE_SSH=false`, DB type `postgres` → `postgres:5432/gitea`.
|
||||
2. **Postgres prep:** one-time `psql` on the existing `postgres` pod — `CREATE ROLE gitea …; CREATE
|
||||
DATABASE gitea OWNER gitea;`. Credentials in a k8s secret consumed by the Gitea release.
|
||||
3. **Exposure:**
|
||||
- HTTPS: extend the nginx `*.fxhnt.ai` proxy (`infra/k8s/gitlab/tailscale-proxy.yaml`) — at cutover,
|
||||
repoint the `git.fxhnt.ai` server block `proxy_pass` → `gitea-http.foxhunt.svc:3000`.
|
||||
- SSH: expose Gitea SSH on `:22` via the Tailscale proxy (socat) at cutover; retire `:2222`; update the
|
||||
two local remotes to `ssh://git@git.fxhnt.ai/...`.
|
||||
- Registry `:5050` server block: removed (Scaleway registry is external, not proxied here).
|
||||
4. **Registry migration:**
|
||||
- `fxhnt/infra/argo/cockpit-build-deploy.yaml`: kaniko `--destination` + `--cache-repo` →
|
||||
`rg.fr-par.scw.cloud/bizworx/fxhnt-cockpit:latest` (+ cache repo); drop the `--insecure-registry`/
|
||||
`--skip-tls-verify` GitLab flags (Scaleway is TLS); auth via a Scaleway registry dockerconfig secret.
|
||||
- `fxhnt/infra/k8s/orchestration/dagster.yaml` + dashboard: image ref →
|
||||
`rg.fr-par.scw.cloud/bizworx/fxhnt-cockpit:latest`; `imagePullSecrets` → Scaleway registry secret.
|
||||
- Create the Scaleway registry pull/push secret (dockerconfigjson) in `foxhunt` ns.
|
||||
5. **Argo Events wiring:** `infra/k8s/argo/events/` — a Gitea webhook eventsource + sensor that submits
|
||||
`fxhnt-cockpit` on push to `fxhnt` default branch (reuses the surviving eventbus + workflow-trigger
|
||||
machinery). Webhook secret in a k8s secret; configured in Gitea repo settings.
|
||||
6. **Deploy script:** `fxhnt/scripts/argo-deploy-cockpit.sh` + the `fxhnt-cockpit` wftmpl clone URL →
|
||||
Gitea (`git.fxhnt.ai`). After cutover the remote name is unchanged, so this is a clone-source confirm.
|
||||
|
||||
## Data flow
|
||||
|
||||
Dev pushes to `git.fxhnt.ai` (Gitea) → webhook → Argo Events sensor → `fxhnt-cockpit` workflow → kaniko
|
||||
builds + pushes to Scaleway registry → `kubectl apply` + rollout → dagster/dashboard pull the new image
|
||||
from Scaleway. Terraform unaffected (state external since 2A).
|
||||
|
||||
## Migration & cutover sequence (each step verified)
|
||||
|
||||
1. Create `gitea` DB on existing postgres + secrets.
|
||||
2. `helm install gitea` (internal svc only). Verify `/api/healthz` 200, admin login via port-forward.
|
||||
3. Port-forward Gitea; `git push --mirror` `fxhnt` and `foxhunt` to it; mark `foxhunt` archived.
|
||||
4. **Validate (pre-cutover):** clone both repos back from the port-forward, diff top commit SHAs vs
|
||||
GitLab — must match.
|
||||
5. **Registry migration:** create Scaleway registry secret; update kaniko build (+cache) + deploy image
|
||||
refs; run one cockpit build → confirm image in Scaleway registry + a test rollout pulls it OK.
|
||||
6. **Cutover:** repoint nginx `git.fxhnt.ai` → Gitea + SSH `:2222` → Gitea; remove the `:5050` block.
|
||||
Verify `git.fxhnt.ai` serves Gitea (web + clone + SSH push).
|
||||
7. **Wire Argo Events:** add Gitea webhook → sensor; test push to `fxhnt` main fires a cockpit deploy.
|
||||
|
||||
## Safety / reversibility
|
||||
|
||||
- GitLab stays fully running throughout 2B (just unaddressed after cutover) — rollback = repoint nginx/SSH
|
||||
back to GitLab. Removal is **2C** only.
|
||||
- The Scaleway registry change is additive until the deploy refs flip; the GitLab registry image stays as
|
||||
fallback until 2C.
|
||||
- All repo data validated (SHA diff) before the public name moves.
|
||||
- Secrets (DB creds, admin, webhook, registry dockerconfig) are k8s secrets — **never committed**.
|
||||
|
||||
## Risks
|
||||
|
||||
- **Registry auth:** Scaleway registry needs a valid API-key dockerconfig for both kaniko push and pod
|
||||
pull. Mitigation: test the build+rollout (step 5) before cutover.
|
||||
- **Gitea↔Postgres coupling:** Gitea shares the existing `postgres`. If that instance is later removed,
|
||||
Gitea breaks. Acceptable (postgres is a kept platform service with backups); noted for 2C/2D.
|
||||
- **SSH port collision at cutover:** GitLab + Gitea can't both own `:2222` on the Tailscale proxy. Cutover
|
||||
flips it atomically; brief push unavailability during the switch (single maintainer — fine).
|
||||
- **Webhook reachability:** Gitea (in-cluster) must reach the Argo Events webhook svc in-cluster — same
|
||||
ns, no Tailscale hop needed.
|
||||
|
||||
## Out of scope
|
||||
|
||||
GitLab helm uninstall + GitLab buckets/state cleanup (**2C**); `git mv` IaC foxhunt → fxhnt (**2D**).
|
||||
|
||||
## Acceptance criteria
|
||||
|
||||
- Gitea running (1 pod, existing postgres), `/api/healthz` 200, ~≤300 MiB.
|
||||
- Both repos in Gitea; clone SHAs match GitLab; `foxhunt` archived.
|
||||
- Cockpit image builds to + deploys from Scaleway Container Registry; `dashboard.fxhnt.ai` 200 on the
|
||||
Scaleway-sourced image.
|
||||
- `git.fxhnt.ai` (HTTPS + SSH `:2222`) serves Gitea; git remotes unchanged.
|
||||
- Push to `fxhnt` main auto-triggers a cockpit deploy via Argo Events.
|
||||
- GitLab still running as fallback (removed in 2C).
|
||||
@@ -0,0 +1,94 @@
|
||||
# Migrate Terraform State to Scaleway Object Storage — Design (Phase 2A)
|
||||
|
||||
**Date:** 2026-06-21
|
||||
**Status:** Design (approved)
|
||||
**Repo:** foxhunt (IaC current home; relocation to fxhnt is Phase 2D)
|
||||
**Part of:** Phase 2 (GitLab → Gitea consolidation). Sub-projects: **2A state migration (this)**, 2B Gitea
|
||||
stand-up + repo migration + Argo wiring, 2C GitLab decommission, 2D IaC relocation into fxhnt.
|
||||
|
||||
## Motivation
|
||||
|
||||
Terraform state for the whole production cluster currently lives in the **self-hosted GitLab** http
|
||||
backend (project ID=1), whose object data sits in **in-cluster MinIO**. That is a chicken-and-egg
|
||||
fragility: the state describing the cluster lives *inside* the cluster it manages. GitLab is also slated
|
||||
for decommission (Phase 2C, replaced by lightweight Gitea). Moving state to **external Scaleway Object
|
||||
Storage** decouples it from both GitLab and the cluster — a prerequisite for removing GitLab and a
|
||||
durability win on its own (state survives cluster loss).
|
||||
|
||||
## Decision (keystone)
|
||||
|
||||
**Target backend: Scaleway Object Storage (native S3), bucket `foxhunt-tfstate`** (already exists,
|
||||
fr-par, created ~3 months ago but never wired). Reuse it (no rename — avoids a needless create; the
|
||||
`foxhunt-` prefix is cosmetic and harmless). Locking via **OpenTofu native S3 lockfile** (`use_lockfile`,
|
||||
OpenTofu v1.11.5 present) — no DynamoDB.
|
||||
|
||||
## Scope — what changes
|
||||
|
||||
1. **`infra/live/production/root.hcl`** — replace the single `remote_state` block: `backend = "http"` (GitLab) →
|
||||
`backend = "s3"` (Scaleway). New config:
|
||||
```hcl
|
||||
remote_state {
|
||||
backend = "s3"
|
||||
generate = { path = "backend.tf", if_exists = "overwrite" }
|
||||
config = {
|
||||
bucket = "foxhunt-tfstate"
|
||||
key = "${path_relative_to_include()}/terraform.tfstate"
|
||||
region = "fr-par"
|
||||
endpoints = { s3 = "https://s3.fr-par.scw.cloud" }
|
||||
skip_credentials_validation = true
|
||||
skip_region_validation = true
|
||||
skip_requesting_account_id = true
|
||||
skip_metadata_api_check = true
|
||||
use_lockfile = true
|
||||
}
|
||||
}
|
||||
```
|
||||
Credentials: Scaleway access/secret keys exported as `AWS_ACCESS_KEY_ID` / `AWS_SECRET_ACCESS_KEY`
|
||||
(the s3 backend reads AWS-style env vars) from the existing `scaleway-credentials` k8s secret.
|
||||
2. **State migration, per module, sequential** (`kapsule`, `public-gateway`, `dns`, `block-storage`):
|
||||
- Back up first: `terragrunt state pull > /tmp/tfstate-backup-<module>.tfstate` (from GitLab, BEFORE
|
||||
changing the backend for that module).
|
||||
- `terragrunt init -migrate-state -force-copy` → copies GitLab state → Scaleway S3.
|
||||
- `terragrunt plan` → **must report `No changes`** before proceeding to the next module.
|
||||
3. **Lock smoke test** (once, on one module): confirm `use_lockfile` works against Scaleway — a lock
|
||||
object is created during apply/plan-with-lock and released after (verify no orphan `.tflock` left).
|
||||
|
||||
## Migration order & gates
|
||||
|
||||
`kapsule` → `public-gateway` → `dns` → `block-storage`. After EACH: backup taken, `-migrate-state`
|
||||
succeeded, `plan = No changes` confirmed. **STOP and revert if any module's post-migration plan shows
|
||||
drift** (revert = restore `http` backend in root.hcl + `terragrunt init` to re-attach GitLab state; the
|
||||
GitLab copy is never deleted in 2A).
|
||||
|
||||
## Safety / reversibility
|
||||
|
||||
- **Old GitLab state is NOT deleted in 2A** — it stays as a live fallback until GitLab is removed in 2C.
|
||||
So the entire migration is reversible by flipping `root.hcl` back to the http backend.
|
||||
- `/tmp` state backups taken before each module (belt-and-suspenders; never committed — contains
|
||||
resource IDs, not secrets, but treat as sensitive).
|
||||
- One module at a time with a `plan = No changes` gate prevents a bad backend config from cascading.
|
||||
- `endpoints.s3` (not deprecated top-level `endpoint`) + the four `skip_*` flags are required for the
|
||||
AWS s3 backend to talk to Scaleway's S3-compatible API without AWS-specific preflight calls.
|
||||
|
||||
## Risks
|
||||
|
||||
- **Scaleway S3 quirks:** some provider/backend versions need `skip_s3_checksum = true` for Scaleway. If
|
||||
`-migrate-state` errors on checksum, add it. (Documented as a known fallback in the plan.)
|
||||
- **Lock semantics:** Scaleway Object Storage must honor conditional-write for `use_lockfile`. The lock
|
||||
smoke test catches this; fallback is to run with locking disabled (single maintainer) if unsupported.
|
||||
- **Credential confusion:** the s3 backend reads `AWS_*` env vars, NOT `SCW_*`. Plan must export both
|
||||
(SCW_* for the provider, AWS_* for the backend) — a common foot-gun.
|
||||
|
||||
## Out of scope (later sub-projects)
|
||||
|
||||
Gitea stand-up + repo migration + Argo wiring (2B); GitLab helm removal + GitLab state/bucket cleanup
|
||||
(2C); `git mv` of IaC from foxhunt → fxhnt + path updates (2D). Files stay in foxhunt for 2A; only the
|
||||
backend changes.
|
||||
|
||||
## Acceptance criteria
|
||||
|
||||
- `infra/live/production/root.hcl` uses the Scaleway s3 backend; committed.
|
||||
- All 4 modules migrated: each shows `terragrunt plan = No changes` against the new backend.
|
||||
- State objects present in `foxhunt-tfstate` at `<module>/terraform.tfstate` (verified via `scw`/`mc`).
|
||||
- Lock smoke test passes (lock acquired + released, no orphan lockfile).
|
||||
- GitLab http state still intact (untouched fallback) — not yet deleted.
|
||||
@@ -1,161 +0,0 @@
|
||||
# Deps Cache Image — Cold-start Compile Acceleration
|
||||
|
||||
This README documents the **`ci-builder-cpu-with-deps:nightly`** image
|
||||
and the Argo plumbing that uses it. Created as part of the CI cold-start
|
||||
optimization sweep (April 2026).
|
||||
|
||||
## What it is
|
||||
|
||||
A Docker image that carries a pre-built `target/` directory at
|
||||
`/cargo-target-prebuilt/` containing the workspace's third-party
|
||||
dependency rlibs. It's pulled by the `compile-services` step of
|
||||
`compile-and-deploy-template.yaml` as an **initContainer** that rsyncs
|
||||
its payload into the `cargo-target-cpu` PVC.
|
||||
|
||||
## Why it exists
|
||||
|
||||
On a fresh node (autoscaler scaled up, fresh PVC, or after a PVC purge),
|
||||
`cargo build --release` would otherwise spend 3-5 minutes compiling
|
||||
~1500 third-party crates before touching workspace member code. With
|
||||
the deps cache image, the cold-start compile path is:
|
||||
|
||||
1. **initContainer pulls image** from in-cluster registry (~30-60s)
|
||||
2. **rsync prebuilt rlibs into PVC** (~30-60s, ~6-8 GB transfer over local FS)
|
||||
3. **cargo build delta-compiles only workspace member code** (~1-2 min)
|
||||
|
||||
Net cold-start saving: **2-4 min per compile pod**.
|
||||
|
||||
## Files
|
||||
|
||||
| File | Role |
|
||||
|------|------|
|
||||
| `infra/docker/Dockerfile.ci-deps-cache` | Builds the deps-cache image (multi-stage). Stage 1 atop ci-builder-cpu, runs `cargo build --release --workspace`. Stage 2 thin Ubuntu + rsync layer carrying just `/cargo-target-prebuilt/`. |
|
||||
| `infra/k8s/argo/refresh-deps-cache-template.yaml` | WorkflowTemplate that drives a Kaniko build of the Dockerfile + push to GitLab registry. Includes a CronWorkflow that runs nightly at 03:00 UTC (suspended by default). |
|
||||
| `infra/k8s/argo/compile-and-deploy-template.yaml` | The `compile-services` step has a `seed-deps-cache` initContainer that rsyncs `/cargo-target-prebuilt/` into the PVC, gated by a stamp file. |
|
||||
|
||||
## How to refresh manually
|
||||
|
||||
```bash
|
||||
# Trigger a one-shot rebuild against current main HEAD.
|
||||
argo submit --from workflowtemplate/refresh-deps-cache -n foxhunt
|
||||
```
|
||||
|
||||
To target a specific commit or bump the cache version (forces re-seed
|
||||
on every compile pod):
|
||||
|
||||
```bash
|
||||
argo submit --from workflowtemplate/refresh-deps-cache -n foxhunt \
|
||||
-p commit-sha=$(git rev-parse main) \
|
||||
-p deps-version=2 \
|
||||
-p image-tag=nightly
|
||||
```
|
||||
|
||||
## Enabling the nightly cron
|
||||
|
||||
The CronWorkflow ships **suspended** so you can validate manually first.
|
||||
To enable:
|
||||
|
||||
```bash
|
||||
kubectl -n foxhunt patch cronworkflow refresh-deps-cache-nightly \
|
||||
--type=merge -p '{"spec":{"suspend":false}}'
|
||||
```
|
||||
|
||||
To disable (return to manual-only):
|
||||
|
||||
```bash
|
||||
kubectl -n foxhunt patch cronworkflow refresh-deps-cache-nightly \
|
||||
--type=merge -p '{"spec":{"suspend":true}}'
|
||||
```
|
||||
|
||||
## Bumping `DEPS_VERSION`
|
||||
|
||||
The compile-services initContainer skips re-seeding when
|
||||
`/cargo-target/cpu_deps_v${DEPS_VERSION}.stamp` exists on the PVC.
|
||||
Bump `DEPS_VERSION` to force every running compile pod to re-seed
|
||||
its PVC from a fresh image. Bump in two places (must match):
|
||||
|
||||
1. `refresh-deps-cache-template.yaml` workflow `deps-version` parameter
|
||||
default value
|
||||
2. `compile-and-deploy-template.yaml` initContainer `DEPS_VERSION` env var
|
||||
|
||||
Bump when:
|
||||
- Cargo.lock has churned significantly (most prebuilt rlibs
|
||||
fingerprint-mismatch cargo's freshness check)
|
||||
- The Rust toolchain version changes (e.g. 1.89 -> 1.90)
|
||||
- A workspace member is renamed or split (rare)
|
||||
|
||||
You do **not** need to bump for routine commits — cargo's own
|
||||
fingerprint check will discard stale rlibs and rebuild as needed.
|
||||
|
||||
## Debugging
|
||||
|
||||
**Symptom: compile-services is still slow on a fresh PVC.**
|
||||
|
||||
Check the seed-deps-cache initContainer logs first:
|
||||
|
||||
```bash
|
||||
argo logs -n foxhunt <workflow-name> -c seed-deps-cache
|
||||
```
|
||||
|
||||
Look for:
|
||||
- `WARN: /cargo-target-prebuilt missing in image` — the image was built
|
||||
but the snapshot dir is empty. The inner `cargo build --workspace`
|
||||
failed in the deps-cache image build. Check the
|
||||
`refresh-deps-cache` workflow logs.
|
||||
- `PVC already seeded for deps v1 (stamp present), skipping.` —
|
||||
expected steady-state on warm pods.
|
||||
- `=== Seeding cargo-target-cpu PVC from prebuilt deps cache ===` +
|
||||
`=== Seed complete ===` — first run on this PVC, working as designed.
|
||||
|
||||
**Symptom: cargo recompiles many third-party deps anyway despite the
|
||||
seed completing.**
|
||||
|
||||
Cargo's fingerprint check is sensitive to: rustc version, the entire
|
||||
target.rustflags list, `RUSTFLAGS` env var, source code mtimes (on
|
||||
build.rs), feature flags. The commit-SHA used to build the deps cache
|
||||
image must match the live compile workflow's flag set. The most common
|
||||
cause of full-rebuild is a `RUSTFLAGS` mismatch between the deps-cache
|
||||
image build and the compile pod (e.g. mold vs wild linker swap).
|
||||
|
||||
To diagnose, look at the cargo build output: lines like
|
||||
`Compiling tokio v1.x.x` for crates that were obviously in the cache
|
||||
indicate fingerprint mismatch.
|
||||
|
||||
**Symptom: PVC fills up faster than before.**
|
||||
|
||||
Each seed adds ~6-8 GB to the PVC. The compile-services step has a
|
||||
prune-at-25GB guard that wipes `release/` and `debug/` if usage
|
||||
exceeds 25GB — when that fires, the next run re-seeds (stamp file is
|
||||
in PVC root, survives the prune). To raise the threshold, edit the
|
||||
`TARGET_SIZE_MB -gt 25000` check in compile-and-deploy-template.yaml.
|
||||
|
||||
## Assumptions that break the cache
|
||||
|
||||
- **rustc toolchain version mismatch** between deps-cache image and
|
||||
compile pod — cargo will full-rebuild. Resolution: rebuild the
|
||||
deps-cache image (refresh-deps-cache).
|
||||
- **Different `RUSTFLAGS` / `target.<triple>.rustflags`** between
|
||||
image build and live compile (e.g. switching CPU target features)
|
||||
— full rebuild. Resolution: rebuild deps-cache image after the
|
||||
rustflags change is committed.
|
||||
- **Cargo.lock SHA divergence** > a few weeks old — most rlibs no
|
||||
longer match. Resolution: rebuild deps-cache image (refresh nightly,
|
||||
or trigger manually).
|
||||
- **Linker swap (e.g. mold -> wild)** — link-output is cached at the
|
||||
fingerprint level; changing linker doesn't trigger rebuild but may
|
||||
produce different final binaries. Not a cache invalidation, just a
|
||||
link-time difference.
|
||||
|
||||
## When to delete this entirely
|
||||
|
||||
If the underlying compile time drops to <30s in cold-start (e.g. via a
|
||||
much faster registry or sccache covering 100% of deps), the
|
||||
deps-cache complexity is no longer worth maintaining. To remove:
|
||||
|
||||
1. Remove the `seed-deps-cache` initContainer block from
|
||||
`compile-and-deploy-template.yaml` (the compile script tolerates
|
||||
an unseeded PVC — cargo just compiles deps from scratch).
|
||||
2. Delete `infra/docker/Dockerfile.ci-deps-cache` and
|
||||
`infra/k8s/argo/refresh-deps-cache-template.yaml`.
|
||||
3. Delete the published image:
|
||||
`crane delete gitlab-registry.foxhunt.svc.cluster.local:5000/root/foxhunt/ci-builder-cpu-with-deps:nightly`.
|
||||
@@ -1,495 +0,0 @@
|
||||
# Alpha walk-forward CV workflow.
|
||||
#
|
||||
# Trains the Mamba2+stacker alpha logits cache on the full 9Q fxcache,
|
||||
# then runs `alpha_baseline` 9 times across disjoint 1.9M-bar windows
|
||||
# (one per quarter). Each fold trains its own execution-policy DQN
|
||||
# from scratch on the front of the window and backtests on the back
|
||||
# (purged train/eval split via `--train-frac`).
|
||||
#
|
||||
# Sequential by design: 9 folds chain in the DAG so a single L40S node
|
||||
# stays warm across the run. With --decision-stride 200 and --horizon
|
||||
# 1200, each fold's DQN training is ~5-10 min on L40S; total wall ~60-90
|
||||
# min for stacker + 9 folds.
|
||||
#
|
||||
# DAG:
|
||||
# ensure-binary ──> ensure-fxcache ──> stacker-train ──> fold-0 ──> ... ──> fold-8
|
||||
#
|
||||
# Usage:
|
||||
# argo submit -n foxhunt --from=wftmpl/alpha-cv \
|
||||
# -p commit-sha=HEAD -p git-branch=main -p decision-stride=200
|
||||
apiVersion: argoproj.io/v1alpha1
|
||||
kind: WorkflowTemplate
|
||||
metadata:
|
||||
name: alpha-cv
|
||||
namespace: foxhunt
|
||||
labels:
|
||||
app.kubernetes.io/name: alpha-cv
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
app.kubernetes.io/component: train
|
||||
spec:
|
||||
entrypoint: pipeline
|
||||
serviceAccountName: argo-workflow
|
||||
archiveLogs: true
|
||||
podMetadata:
|
||||
labels:
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
app.kubernetes.io/component: train
|
||||
securityContext:
|
||||
fsGroup: 0
|
||||
ttlStrategy:
|
||||
secondsAfterCompletion: 3600
|
||||
activeDeadlineSeconds: 28800
|
||||
|
||||
arguments:
|
||||
parameters:
|
||||
- name: commit-sha
|
||||
value: HEAD
|
||||
- name: git-branch
|
||||
value: main
|
||||
- name: cuda-compute-cap
|
||||
value: "89"
|
||||
- name: gpu-pool
|
||||
value: ci-training-l40s
|
||||
- name: symbol
|
||||
value: ES.FUT
|
||||
# Match the fxcache build params (cluster's 9Q fxcache was built
|
||||
# with these on 2026-05-16; see commit e7ce4395e).
|
||||
- name: imbalance-bar-threshold
|
||||
value: "20.0"
|
||||
- name: imbalance-bar-ewma-alpha
|
||||
value: "0.1"
|
||||
- name: volume-bar-size
|
||||
value: "100"
|
||||
- name: data-source
|
||||
value: "mbp10"
|
||||
# Stacker training params (mirror local 2Q smoke).
|
||||
- name: stacker-horizon
|
||||
value: "6000"
|
||||
- name: stacker-seq-len
|
||||
value: "32"
|
||||
- name: stacker-hidden-dim
|
||||
value: "64"
|
||||
- name: stacker-state-dim
|
||||
value: "16"
|
||||
- name: stacker-epochs
|
||||
value: "5"
|
||||
- name: stacker-batch-size
|
||||
value: "128"
|
||||
- name: stacker-lr
|
||||
value: "3e-3"
|
||||
- name: stacker-train-frac
|
||||
value: "0.8"
|
||||
# CV / alpha_baseline params — defaults match the validated 2Q
|
||||
# config that produced +1.78 Sharpe at quarter-tick cost.
|
||||
- name: decision-stride
|
||||
value: "200"
|
||||
- name: fold-window
|
||||
value: "1900000"
|
||||
- name: train-frac
|
||||
value: "0.6"
|
||||
- name: window-k
|
||||
value: "16"
|
||||
- name: horizon
|
||||
value: "1200"
|
||||
- name: n-train-par
|
||||
value: "25"
|
||||
- name: n-train-episodes
|
||||
value: "8000"
|
||||
|
||||
volumes:
|
||||
- name: git-ssh-key
|
||||
secret:
|
||||
secretName: argo-git-ssh-key
|
||||
defaultMode: 256
|
||||
- name: training-data
|
||||
persistentVolumeClaim:
|
||||
claimName: training-data-pvc
|
||||
- name: cargo-target-cuda
|
||||
persistentVolumeClaim:
|
||||
claimName: cargo-target-cuda
|
||||
- name: feature-cache
|
||||
persistentVolumeClaim:
|
||||
claimName: feature-cache-pvc
|
||||
|
||||
templates:
|
||||
# ── DAG ──
|
||||
- name: pipeline
|
||||
dag:
|
||||
tasks:
|
||||
- name: ensure-binary
|
||||
template: ensure-binary
|
||||
- name: stacker-train
|
||||
template: stacker-train
|
||||
dependencies: [ensure-binary]
|
||||
arguments:
|
||||
parameters:
|
||||
- name: sha
|
||||
value: "{{tasks.ensure-binary.outputs.parameters.sha}}"
|
||||
- name: fold-0
|
||||
template: alpha-cv-fold
|
||||
dependencies: [stacker-train]
|
||||
arguments:
|
||||
parameters:
|
||||
- name: sha
|
||||
value: "{{tasks.ensure-binary.outputs.parameters.sha}}"
|
||||
- name: fold-idx
|
||||
value: "0"
|
||||
- name: offset
|
||||
value: "0"
|
||||
- name: fold-1
|
||||
template: alpha-cv-fold
|
||||
dependencies: [fold-0]
|
||||
arguments:
|
||||
parameters:
|
||||
- name: sha
|
||||
value: "{{tasks.ensure-binary.outputs.parameters.sha}}"
|
||||
- name: fold-idx
|
||||
value: "1"
|
||||
- name: offset
|
||||
value: "1900000"
|
||||
- name: fold-2
|
||||
template: alpha-cv-fold
|
||||
dependencies: [fold-1]
|
||||
arguments:
|
||||
parameters:
|
||||
- name: sha
|
||||
value: "{{tasks.ensure-binary.outputs.parameters.sha}}"
|
||||
- name: fold-idx
|
||||
value: "2"
|
||||
- name: offset
|
||||
value: "3800000"
|
||||
- name: fold-3
|
||||
template: alpha-cv-fold
|
||||
dependencies: [fold-2]
|
||||
arguments:
|
||||
parameters:
|
||||
- name: sha
|
||||
value: "{{tasks.ensure-binary.outputs.parameters.sha}}"
|
||||
- name: fold-idx
|
||||
value: "3"
|
||||
- name: offset
|
||||
value: "5700000"
|
||||
- name: fold-4
|
||||
template: alpha-cv-fold
|
||||
dependencies: [fold-3]
|
||||
arguments:
|
||||
parameters:
|
||||
- name: sha
|
||||
value: "{{tasks.ensure-binary.outputs.parameters.sha}}"
|
||||
- name: fold-idx
|
||||
value: "4"
|
||||
- name: offset
|
||||
value: "7600000"
|
||||
- name: fold-5
|
||||
template: alpha-cv-fold
|
||||
dependencies: [fold-4]
|
||||
arguments:
|
||||
parameters:
|
||||
- name: sha
|
||||
value: "{{tasks.ensure-binary.outputs.parameters.sha}}"
|
||||
- name: fold-idx
|
||||
value: "5"
|
||||
- name: offset
|
||||
value: "9500000"
|
||||
- name: fold-6
|
||||
template: alpha-cv-fold
|
||||
dependencies: [fold-5]
|
||||
arguments:
|
||||
parameters:
|
||||
- name: sha
|
||||
value: "{{tasks.ensure-binary.outputs.parameters.sha}}"
|
||||
- name: fold-idx
|
||||
value: "6"
|
||||
- name: offset
|
||||
value: "11400000"
|
||||
- name: fold-7
|
||||
template: alpha-cv-fold
|
||||
dependencies: [fold-6]
|
||||
arguments:
|
||||
parameters:
|
||||
- name: sha
|
||||
value: "{{tasks.ensure-binary.outputs.parameters.sha}}"
|
||||
- name: fold-idx
|
||||
value: "7"
|
||||
- name: offset
|
||||
value: "13300000"
|
||||
- name: fold-8
|
||||
template: alpha-cv-fold
|
||||
dependencies: [fold-7]
|
||||
arguments:
|
||||
parameters:
|
||||
- name: sha
|
||||
value: "{{tasks.ensure-binary.outputs.parameters.sha}}"
|
||||
- name: fold-idx
|
||||
value: "8"
|
||||
- name: offset
|
||||
value: "15200000"
|
||||
|
||||
# ── ensure-binary: compile alpha_train_stacker + alpha_baseline ──
|
||||
- name: ensure-binary
|
||||
outputs:
|
||||
parameters:
|
||||
- name: sha
|
||||
valueFrom:
|
||||
path: /tmp/sha
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: ci-compile-cpu
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
tolerations:
|
||||
- key: node.cilium.io/agent-not-ready
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
container:
|
||||
image: gitlab-registry.foxhunt.svc.cluster.local:5000/root/foxhunt/ci-builder:latest
|
||||
imagePullPolicy: Always
|
||||
command: ["/bin/bash", "-c"]
|
||||
env:
|
||||
- name: SQLX_OFFLINE
|
||||
value: "true"
|
||||
- name: CARGO_TERM_COLOR
|
||||
value: always
|
||||
- name: CARGO_TARGET_DIR
|
||||
value: /cargo-target
|
||||
- name: CARGO_HOME
|
||||
value: /cargo-target/cargo-home
|
||||
- name: RUSTC_WRAPPER
|
||||
value: sccache
|
||||
- name: SCCACHE_DIR
|
||||
value: /cargo-target/sccache
|
||||
- name: SCCACHE_CACHE_SIZE
|
||||
value: "40G"
|
||||
- name: CARGO_INCREMENTAL
|
||||
value: "0"
|
||||
- name: GITLAB_PAT
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: gitlab-pat
|
||||
key: token
|
||||
resources:
|
||||
requests:
|
||||
cpu: "14"
|
||||
memory: 32Gi
|
||||
limits:
|
||||
cpu: "30"
|
||||
memory: 64Gi
|
||||
volumeMounts:
|
||||
- name: git-ssh-key
|
||||
mountPath: /etc/git-ssh
|
||||
readOnly: true
|
||||
- name: cargo-target-cuda
|
||||
mountPath: /cargo-target
|
||||
- name: training-data
|
||||
mountPath: /data
|
||||
args:
|
||||
- |
|
||||
set -e
|
||||
SHA="{{workflow.parameters.commit-sha}}"
|
||||
BRANCH="{{workflow.parameters.git-branch}}"
|
||||
mkdir -p ~/.ssh
|
||||
cp /etc/git-ssh/ssh-privatekey ~/.ssh/id_ed25519
|
||||
chmod 600 ~/.ssh/id_ed25519
|
||||
printf 'Host *\n StrictHostKeyChecking no\n UserKnownHostsFile /dev/null\n' > ~/.ssh/config
|
||||
chmod 600 ~/.ssh/config
|
||||
REPO="ssh://git@gitlab-gitlab-shell.foxhunt.svc.cluster.local:2222/root/foxhunt.git"
|
||||
BUILD="/cargo-target/src"
|
||||
git config --global --add safe.directory "$BUILD"
|
||||
|
||||
if [ "$SHA" = "HEAD" ]; then
|
||||
if [ -d "$BUILD/.git" ]; then
|
||||
cd "$BUILD"; git fetch origin; SHA=$(git rev-parse "origin/$BRANCH"); cd /
|
||||
else
|
||||
git clone --filter=blob:none "$REPO" "$BUILD"; cd "$BUILD"
|
||||
git checkout "origin/$BRANCH"; SHA=$(git rev-parse HEAD); cd /
|
||||
fi
|
||||
fi
|
||||
SHORT_SHA=$(echo "$SHA" | cut -c1-9)
|
||||
echo "Resolved SHA: $SHA (short: $SHORT_SHA)"
|
||||
|
||||
BIN_DIR="/data/bin/$SHORT_SHA"
|
||||
BINARIES="alpha_train_stacker alpha_baseline"
|
||||
ALL_CACHED=true
|
||||
for bin in $BINARIES; do
|
||||
[ ! -x "$BIN_DIR/$bin" ] && ALL_CACHED=false && break
|
||||
done
|
||||
|
||||
if [ "$ALL_CACHED" = "true" ]; then
|
||||
echo "=== Cache HIT ==="
|
||||
ls -lh "$BIN_DIR/"
|
||||
echo "$SHORT_SHA" > /tmp/sha
|
||||
exit 0
|
||||
fi
|
||||
echo "=== Cache MISS: compiling alpha binaries for $SHORT_SHA ==="
|
||||
|
||||
if [ -d "$BUILD/.git" ]; then
|
||||
cd "$BUILD"; git fetch origin
|
||||
CURRENT=$(git rev-parse HEAD 2>/dev/null || echo "none")
|
||||
if [ "$CURRENT" != "$SHA" ]; then
|
||||
git checkout --force "$SHA"; git clean -fd
|
||||
fi
|
||||
else
|
||||
git clone --filter=blob:none "$REPO" "$BUILD"; cd "$BUILD"; git checkout "$SHA"
|
||||
fi
|
||||
|
||||
export PATH="${CARGO_HOME}/bin:${PATH}"
|
||||
export CUDA_COMPUTE_CAP={{workflow.parameters.cuda-compute-cap}}
|
||||
|
||||
# alpha_train_stacker lives in ml-alpha, alpha_baseline in ml.
|
||||
echo "Building alpha_train_stacker (ml-alpha)..."
|
||||
cargo build --release -p ml-alpha --example alpha_train_stacker
|
||||
echo "Building alpha_baseline (ml)..."
|
||||
cargo build --release -p ml --features ml/cuda --example alpha_baseline
|
||||
|
||||
mkdir -p "$BIN_DIR"
|
||||
cp "$CARGO_TARGET_DIR/release/examples/alpha_train_stacker" "$BIN_DIR/"
|
||||
cp "$CARGO_TARGET_DIR/release/examples/alpha_baseline" "$BIN_DIR/"
|
||||
strip "$BIN_DIR/"*
|
||||
# alpha_baseline reads config/ml/alpha_fill_coeffs.json as a
|
||||
# required --fill-coeffs input. The source tree's copy travels
|
||||
# with the binaries so the GPU fold pods don't need to clone.
|
||||
cp "$BUILD/config/ml/alpha_fill_coeffs.json" "$BIN_DIR/" || true
|
||||
ls -lh "$BIN_DIR/"
|
||||
|
||||
cd /data/bin
|
||||
ls -1t | tail -n +6 | while read -r old; do
|
||||
echo "Pruning old cache: $old"; rm -rf "$old"
|
||||
done
|
||||
echo "$SHORT_SHA" > /tmp/sha
|
||||
|
||||
# ── stacker-train: train Mamba2+stacker on the 9Q fxcache ──
|
||||
- name: stacker-train
|
||||
inputs:
|
||||
parameters:
|
||||
- name: sha
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: "{{workflow.parameters.gpu-pool}}"
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
tolerations:
|
||||
- key: nvidia.com/gpu
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
- key: node.cilium.io/agent-not-ready
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
container:
|
||||
image: gitlab-registry.foxhunt.svc.cluster.local:5000/root/foxhunt/training-runtime:latest
|
||||
imagePullPolicy: Always
|
||||
command: ["/bin/bash", "-c"]
|
||||
env:
|
||||
- name: RUST_LOG
|
||||
value: info
|
||||
- name: SQLX_OFFLINE
|
||||
value: "true"
|
||||
resources:
|
||||
requests:
|
||||
cpu: "8"
|
||||
memory: 32Gi
|
||||
nvidia.com/gpu: "1"
|
||||
limits:
|
||||
cpu: "16"
|
||||
memory: 80Gi
|
||||
nvidia.com/gpu: "1"
|
||||
volumeMounts:
|
||||
- name: training-data
|
||||
mountPath: /data
|
||||
- name: feature-cache
|
||||
mountPath: /feature-cache
|
||||
args:
|
||||
- |
|
||||
set -e
|
||||
SHA="{{inputs.parameters.sha}}"
|
||||
BIN="/data/bin/$SHA/alpha_train_stacker"
|
||||
FXCACHE=$(ls -t /feature-cache/*.fxcache | head -1)
|
||||
ALPHA_OUT="/feature-cache/alpha_logits_cache_9q.bin"
|
||||
echo "Using fxcache: $FXCACHE"
|
||||
echo "Output: $ALPHA_OUT"
|
||||
|
||||
"$BIN" \
|
||||
--fxcache-path "$FXCACHE" \
|
||||
--horizon {{workflow.parameters.stacker-horizon}} \
|
||||
--seq-len {{workflow.parameters.stacker-seq-len}} \
|
||||
--hidden-dim {{workflow.parameters.stacker-hidden-dim}} \
|
||||
--state-dim {{workflow.parameters.stacker-state-dim}} \
|
||||
--epochs {{workflow.parameters.stacker-epochs}} \
|
||||
--batch-size {{workflow.parameters.stacker-batch-size}} \
|
||||
--lr {{workflow.parameters.stacker-lr}} \
|
||||
--train-frac {{workflow.parameters.stacker-train-frac}} \
|
||||
--alpha-cache-out "$ALPHA_OUT"
|
||||
ls -lh "$ALPHA_OUT"
|
||||
|
||||
# ── alpha-cv-fold: run alpha_baseline on one window ──
|
||||
- name: alpha-cv-fold
|
||||
inputs:
|
||||
parameters:
|
||||
- name: sha
|
||||
- name: fold-idx
|
||||
- name: offset
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: "{{workflow.parameters.gpu-pool}}"
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
tolerations:
|
||||
- key: nvidia.com/gpu
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
- key: node.cilium.io/agent-not-ready
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
container:
|
||||
image: gitlab-registry.foxhunt.svc.cluster.local:5000/root/foxhunt/training-runtime:latest
|
||||
imagePullPolicy: Always
|
||||
command: ["/bin/bash", "-c"]
|
||||
env:
|
||||
- name: RUST_LOG
|
||||
value: info
|
||||
- name: SQLX_OFFLINE
|
||||
value: "true"
|
||||
resources:
|
||||
requests:
|
||||
cpu: "8"
|
||||
memory: 32Gi
|
||||
nvidia.com/gpu: "1"
|
||||
limits:
|
||||
cpu: "16"
|
||||
memory: 80Gi
|
||||
nvidia.com/gpu: "1"
|
||||
volumeMounts:
|
||||
- name: training-data
|
||||
mountPath: /data
|
||||
- name: feature-cache
|
||||
mountPath: /feature-cache
|
||||
args:
|
||||
- |
|
||||
set -e
|
||||
SHA="{{inputs.parameters.sha}}"
|
||||
FOLD={{inputs.parameters.fold-idx}}
|
||||
OFFSET={{inputs.parameters.offset}}
|
||||
BIN="/data/bin/$SHA/alpha_baseline"
|
||||
FXCACHE=$(ls -t /feature-cache/*.fxcache | head -1)
|
||||
ALPHA="/feature-cache/alpha_logits_cache_9q.bin"
|
||||
# alpha_fill_coeffs.json travels with the source tree; the
|
||||
# binary needs it next to itself or via explicit --fill-coeffs.
|
||||
# The ci-builder image bakes it under /opt/foxhunt/config/ml.
|
||||
FILL="/data/bin/$SHA/alpha_fill_coeffs.json"
|
||||
if [ ! -f "$FILL" ]; then
|
||||
# Fall back to copying from the cloned source on the
|
||||
# cargo-target PVC mounted in ensure-binary; we don't have
|
||||
# that PVC here, so embed via fxcache PVC pre-stage.
|
||||
FILL="/feature-cache/alpha_fill_coeffs.json"
|
||||
fi
|
||||
OUT="/feature-cache/cv_results_9fold/fold_${FOLD}.json"
|
||||
mkdir -p /feature-cache/cv_results_9fold
|
||||
echo "Fold $FOLD: offset=$OFFSET out=$OUT"
|
||||
|
||||
"$BIN" \
|
||||
--fxcache-path "$FXCACHE" \
|
||||
--alpha-cache "$ALPHA" \
|
||||
--fill-coeffs "$FILL" \
|
||||
--data-start-offset $OFFSET \
|
||||
--max-snapshots {{workflow.parameters.fold-window}} \
|
||||
--train-frac {{workflow.parameters.train-frac}} \
|
||||
--window-k {{workflow.parameters.window-k}} \
|
||||
--decision-stride {{workflow.parameters.decision-stride}} \
|
||||
--horizon {{workflow.parameters.horizon}} \
|
||||
--n-train-par {{workflow.parameters.n-train-par}} \
|
||||
--n-train-episodes {{workflow.parameters.n-train-episodes}} \
|
||||
--out-path "$OUT"
|
||||
ls -lh "$OUT"
|
||||
@@ -1,484 +0,0 @@
|
||||
# Alpha perception trainer workflow.
|
||||
#
|
||||
# Stacked Mamba2 -> CfC -> heads supervised pretrain on MBP-10 from the
|
||||
# training-data PVC. Emits `alpha_train_summary.json` to MinIO-backed
|
||||
# feature-cache PVC with per-horizon val AUC for monitoring. Runs on
|
||||
# a single L40S; ~30-90 min wall depending on
|
||||
# n-train-seqs.
|
||||
#
|
||||
# DAG:
|
||||
# ensure-binary ──> train
|
||||
#
|
||||
# Usage:
|
||||
# argo submit -n foxhunt --from=wftmpl/alpha-perception \
|
||||
# -p commit-sha=HEAD -p git-branch=ml-alpha-phase-a
|
||||
apiVersion: argoproj.io/v1alpha1
|
||||
kind: WorkflowTemplate
|
||||
metadata:
|
||||
name: alpha-perception
|
||||
namespace: foxhunt
|
||||
labels:
|
||||
app.kubernetes.io/name: alpha-perception
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
app.kubernetes.io/component: train
|
||||
spec:
|
||||
entrypoint: pipeline
|
||||
serviceAccountName: argo-workflow
|
||||
archiveLogs: true
|
||||
podMetadata:
|
||||
labels:
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
app.kubernetes.io/component: train
|
||||
securityContext:
|
||||
fsGroup: 0
|
||||
ttlStrategy:
|
||||
secondsAfterCompletion: 3600
|
||||
activeDeadlineSeconds: 14400
|
||||
|
||||
arguments:
|
||||
parameters:
|
||||
- name: commit-sha
|
||||
value: HEAD
|
||||
- name: git-branch
|
||||
value: ml-alpha-phase-a
|
||||
- name: cuda-compute-cap
|
||||
value: "89"
|
||||
- name: gpu-pool
|
||||
value: ci-training-l40s
|
||||
# Trainer hyperparameters (mirror the validated smoke config).
|
||||
- name: epochs
|
||||
value: "5"
|
||||
- name: multi-resolution
|
||||
value: "1:32"
|
||||
- name: mamba2-state-dim
|
||||
value: "16"
|
||||
- name: lr-cfc
|
||||
value: "3e-3"
|
||||
- name: lr-mamba2
|
||||
value: "1e-3"
|
||||
- name: n-train-seqs
|
||||
value: "8000"
|
||||
- name: n-val-seqs
|
||||
value: "1000"
|
||||
- name: seed
|
||||
value: "16962"
|
||||
- name: batch-size
|
||||
value: "1"
|
||||
- name: auto-horizon-weights
|
||||
value: "false"
|
||||
- name: instrument-mode
|
||||
value: "all"
|
||||
- name: early-stop-metric
|
||||
value: "mean_auc"
|
||||
- name: early-stop-patience
|
||||
value: "5"
|
||||
- name: cv-fold
|
||||
value: "0"
|
||||
- name: cv-n-folds
|
||||
value: "1"
|
||||
- name: cv-train-window
|
||||
value: "0"
|
||||
- name: smoothness-base-lambda
|
||||
value: "0.0"
|
||||
- name: kernel-step-trace-enable
|
||||
value: "false"
|
||||
- name: kernel-step-trace-path
|
||||
value: ""
|
||||
|
||||
volumes:
|
||||
- name: git-ssh-key
|
||||
secret:
|
||||
secretName: argo-git-ssh-key
|
||||
defaultMode: 256
|
||||
- name: training-data
|
||||
persistentVolumeClaim:
|
||||
claimName: training-data-pvc
|
||||
- name: cargo-target-cuda
|
||||
persistentVolumeClaim:
|
||||
claimName: cargo-target-cuda
|
||||
- name: feature-cache
|
||||
persistentVolumeClaim:
|
||||
claimName: feature-cache-pvc
|
||||
|
||||
templates:
|
||||
# ── DAG ──
|
||||
#
|
||||
# check-cache (tiny alpine pod on platform pool, ~3 sec) inspects
|
||||
# the training-data PVC for /data/bin/$SHA/alpha_train. If present
|
||||
# (cache hit), ensure-binary is SKIPPED via `when:` — the
|
||||
# ~4.8GB ci-builder image pull + sccache compile cycle are avoided
|
||||
# entirely. train still runs, sourcing the binary path from
|
||||
# check-cache's SHA output.
|
||||
#
|
||||
# warmup-gpu fires unconditionally in parallel — triggers the L40S
|
||||
# autoscaler so the node is warm by the time train needs it,
|
||||
# whether or not we compile.
|
||||
- name: pipeline
|
||||
dag:
|
||||
tasks:
|
||||
- name: check-cache
|
||||
template: check-cache
|
||||
- name: ensure-binary
|
||||
template: ensure-binary
|
||||
dependencies: [check-cache]
|
||||
when: "{{tasks.check-cache.outputs.parameters.cache}} == miss"
|
||||
arguments:
|
||||
parameters:
|
||||
- name: sha
|
||||
value: "{{tasks.check-cache.outputs.parameters.sha}}"
|
||||
- name: warmup-gpu
|
||||
template: warmup-gpu
|
||||
- name: train
|
||||
template: train
|
||||
dependencies: [check-cache, ensure-binary]
|
||||
arguments:
|
||||
parameters:
|
||||
- name: sha
|
||||
value: "{{tasks.check-cache.outputs.parameters.sha}}"
|
||||
|
||||
# ── check-cache: probe the training-data PVC for an existing binary ──
|
||||
#
|
||||
# Runs on the platform pool (always up, no autoscaler delay). Tiny
|
||||
# alpine pod, ~3 sec end-to-end. Emits two output parameters:
|
||||
# - sha = the short SHA used for binary cache keying
|
||||
# - cache = "hit" if /data/bin/$SHA/alpha_train exists else "miss"
|
||||
#
|
||||
# Resolves `commit-sha=HEAD` by reading /tmp/head-sha from a
|
||||
# ConfigMap... actually no, simpler: the submission script
|
||||
# (scripts/argo-alpha-perception.sh) pre-resolves HEAD via local
|
||||
# git so commit-sha is always a real SHA in the workflow. This
|
||||
# pod just trims to short SHA and stats the binary.
|
||||
- name: check-cache
|
||||
outputs:
|
||||
parameters:
|
||||
- name: sha
|
||||
valueFrom:
|
||||
path: /tmp/sha
|
||||
- name: cache
|
||||
valueFrom:
|
||||
path: /tmp/cache
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: platform
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
tolerations:
|
||||
- key: node.cilium.io/agent-not-ready
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
container:
|
||||
image: alpine:3.20
|
||||
command: ["/bin/sh", "-c"]
|
||||
resources:
|
||||
requests:
|
||||
cpu: "50m"
|
||||
memory: 32Mi
|
||||
limits:
|
||||
cpu: "100m"
|
||||
memory: 64Mi
|
||||
volumeMounts:
|
||||
- name: training-data
|
||||
mountPath: /data
|
||||
readOnly: true
|
||||
args:
|
||||
- |
|
||||
set -e
|
||||
SHA="{{workflow.parameters.commit-sha}}"
|
||||
if [ "$SHA" = "HEAD" ]; then
|
||||
echo "ERROR: commit-sha cannot be HEAD inside the workflow."
|
||||
echo " Resolve via the submission script (scripts/argo-alpha-perception.sh)."
|
||||
exit 1
|
||||
fi
|
||||
SHORT_SHA=$(echo "$SHA" | cut -c1-9)
|
||||
echo "$SHORT_SHA" > /tmp/sha
|
||||
|
||||
BIN="/data/bin/$SHORT_SHA/alpha_train"
|
||||
if [ -x "$BIN" ]; then
|
||||
SIZE=$(stat -c %s "$BIN")
|
||||
echo "Cache HIT: $BIN ($SIZE bytes)"
|
||||
echo "hit" > /tmp/cache
|
||||
else
|
||||
echo "Cache MISS: $BIN not present, ensure-binary will compile"
|
||||
ls -lh "/data/bin/" 2>/dev/null | head -10 || echo " (no /data/bin/ directory yet)"
|
||||
echo "miss" > /tmp/cache
|
||||
fi
|
||||
|
||||
# ── ensure-binary: compile `alpha_train` example via sccache ──
|
||||
# Only runs when check-cache returned cache=miss. Outputs the SHA
|
||||
# for consistency, but the DAG sources SHA from check-cache directly.
|
||||
- name: ensure-binary
|
||||
inputs:
|
||||
parameters:
|
||||
- name: sha
|
||||
outputs:
|
||||
parameters:
|
||||
- name: sha
|
||||
valueFrom:
|
||||
path: /tmp/sha
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: ci-compile-cpu
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
tolerations:
|
||||
- key: node.cilium.io/agent-not-ready
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
container:
|
||||
image: gitlab-registry.foxhunt.svc.cluster.local:5000/root/foxhunt/ci-builder:latest
|
||||
imagePullPolicy: Always
|
||||
command: ["/bin/bash", "-c"]
|
||||
env:
|
||||
- name: SQLX_OFFLINE
|
||||
value: "true"
|
||||
- name: CARGO_TERM_COLOR
|
||||
value: always
|
||||
- name: CARGO_TARGET_DIR
|
||||
value: /cargo-target
|
||||
- name: CARGO_HOME
|
||||
value: /cargo-target/cargo-home
|
||||
- name: RUSTC_WRAPPER
|
||||
value: sccache
|
||||
- name: SCCACHE_DIR
|
||||
value: /cargo-target/sccache
|
||||
- name: SCCACHE_CACHE_SIZE
|
||||
value: "40G"
|
||||
- name: CARGO_INCREMENTAL
|
||||
value: "0"
|
||||
- name: CUDA_COMPUTE_CAP
|
||||
value: "{{workflow.parameters.cuda-compute-cap}}"
|
||||
resources:
|
||||
requests:
|
||||
cpu: "14"
|
||||
memory: 32Gi
|
||||
limits:
|
||||
cpu: "30"
|
||||
memory: 64Gi
|
||||
volumeMounts:
|
||||
- name: git-ssh-key
|
||||
mountPath: /etc/git-ssh
|
||||
readOnly: true
|
||||
- name: cargo-target-cuda
|
||||
mountPath: /cargo-target
|
||||
- name: training-data
|
||||
mountPath: /data
|
||||
args:
|
||||
- |
|
||||
set -e
|
||||
SHA="{{workflow.parameters.commit-sha}}"
|
||||
BRANCH="{{workflow.parameters.git-branch}}"
|
||||
mkdir -p ~/.ssh
|
||||
cp /etc/git-ssh/ssh-privatekey ~/.ssh/id_ed25519
|
||||
chmod 600 ~/.ssh/id_ed25519
|
||||
printf 'Host *\n StrictHostKeyChecking no\n UserKnownHostsFile /dev/null\n' > ~/.ssh/config
|
||||
chmod 600 ~/.ssh/config
|
||||
REPO="ssh://git@gitlab-gitlab-shell.foxhunt.svc.cluster.local:2222/root/foxhunt.git"
|
||||
BUILD="/cargo-target/src"
|
||||
git config --global --add safe.directory "$BUILD"
|
||||
|
||||
if [ "$SHA" = "HEAD" ]; then
|
||||
if [ -d "$BUILD/.git" ]; then
|
||||
cd "$BUILD"; git fetch origin; SHA=$(git rev-parse "origin/$BRANCH"); cd /
|
||||
else
|
||||
git clone --filter=blob:none "$REPO" "$BUILD"; cd "$BUILD"
|
||||
git checkout "origin/$BRANCH"; SHA=$(git rev-parse HEAD); cd /
|
||||
fi
|
||||
fi
|
||||
SHORT_SHA=$(echo "$SHA" | cut -c1-9)
|
||||
echo "Resolved SHA: $SHA (short: $SHORT_SHA)"
|
||||
|
||||
BIN_DIR="/data/bin/$SHORT_SHA"
|
||||
if [ -x "$BIN_DIR/alpha_train" ]; then
|
||||
echo "=== Cache HIT: $BIN_DIR/alpha_train ==="
|
||||
ls -lh "$BIN_DIR/"
|
||||
echo "$SHORT_SHA" > /tmp/sha
|
||||
exit 0
|
||||
fi
|
||||
echo "=== Cache MISS: compiling alpha_train for $SHORT_SHA ==="
|
||||
|
||||
if [ -d "$BUILD/.git" ]; then
|
||||
cd "$BUILD"; git fetch origin
|
||||
CURRENT=$(git rev-parse HEAD 2>/dev/null || echo "none")
|
||||
if [ "$CURRENT" != "$SHA" ]; then
|
||||
git checkout --force "$SHA"; git clean -fd
|
||||
fi
|
||||
else
|
||||
git clone --filter=blob:none "$REPO" "$BUILD"; cd "$BUILD"; git checkout "$SHA"
|
||||
fi
|
||||
|
||||
export PATH="${CARGO_HOME}/bin:${PATH}"
|
||||
|
||||
FEATURES_FLAG=""
|
||||
if [ "{{workflow.parameters.kernel-step-trace-enable}}" = "true" ]; then
|
||||
FEATURES_FLAG="--features kernel-step-trace"
|
||||
echo "Building with --features kernel-step-trace"
|
||||
# Bust cache: feature-on binary differs from feature-off at same SHA.
|
||||
if [ -x "$BIN_DIR/alpha_train" ]; then
|
||||
echo "Removing cached binary (feature change requires rebuild)"
|
||||
rm -f "$BIN_DIR/alpha_train"
|
||||
fi
|
||||
fi
|
||||
|
||||
echo "Building alpha_train (ml-alpha)..."
|
||||
cargo build --release $FEATURES_FLAG -p ml-alpha --example alpha_train
|
||||
|
||||
mkdir -p "$BIN_DIR"
|
||||
cp "$CARGO_TARGET_DIR/release/examples/alpha_train" "$BIN_DIR/"
|
||||
strip "$BIN_DIR/alpha_train"
|
||||
ls -lh "$BIN_DIR/"
|
||||
|
||||
# Prune old cache (keep last 5)
|
||||
cd /data/bin
|
||||
ls -1t | tail -n +6 | while read -r old; do
|
||||
echo "Pruning old cache: $old"; rm -rf "$old"
|
||||
done
|
||||
echo "$SHORT_SHA" > /tmp/sha
|
||||
|
||||
# ── warmup-gpu: trigger L40S autoscaler scale-up in parallel ──
|
||||
#
|
||||
# Schedules a tiny CPU-only pod on the gpu-pool's nodeSelector.
|
||||
# Kubernetes sees an unschedulable pod (no node currently in the
|
||||
# pool), autoscaler scales the pool 0 → 1. The pod exits the
|
||||
# instant it lands; the node enters scaledown-grace.
|
||||
#
|
||||
# Cluster autoscaler config (see infra/modules/kapsule/main.tf):
|
||||
# scale_down_delay_after_add = "10m" -- 10m before any scaledown consideration
|
||||
# scale_down_unneeded_time = "10m" -- 10m empty before action
|
||||
# Effective grace window is up to ~20m, which always spans
|
||||
# ensure-binary compile (sccache-hit ~10s through cold ~15m).
|
||||
# train lands on a hot node with zero provisioning latency.
|
||||
#
|
||||
# No GPU resource request — that would compete with `train` for
|
||||
# the single GPU and serialise the steps. The nodeSelector +
|
||||
# nvidia.com/gpu toleration are enough to force placement on the
|
||||
# right pool.
|
||||
- name: warmup-gpu
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: "{{workflow.parameters.gpu-pool}}"
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
tolerations:
|
||||
- key: nvidia.com/gpu
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
- key: node.cilium.io/agent-not-ready
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
container:
|
||||
image: alpine:3.20
|
||||
command: ["/bin/sh", "-c"]
|
||||
resources:
|
||||
requests:
|
||||
cpu: "50m"
|
||||
memory: 32Mi
|
||||
limits:
|
||||
cpu: "100m"
|
||||
memory: 64Mi
|
||||
args:
|
||||
- |
|
||||
echo "GPU warmup pod scheduled on $(hostname) — autoscaler scale-up triggered, node now in scaledown-grace window."
|
||||
|
||||
# ── train: run alpha_train on L40S ──
|
||||
- name: train
|
||||
inputs:
|
||||
parameters:
|
||||
- name: sha
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: "{{workflow.parameters.gpu-pool}}"
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
tolerations:
|
||||
- key: nvidia.com/gpu
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
- key: node.cilium.io/agent-not-ready
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
container:
|
||||
image: gitlab-registry.foxhunt.svc.cluster.local:5000/root/foxhunt/foxhunt-training-runtime:latest
|
||||
imagePullPolicy: Always
|
||||
command: ["/bin/bash", "-c"]
|
||||
env:
|
||||
- name: RUST_LOG
|
||||
value: info
|
||||
- name: SQLX_OFFLINE
|
||||
value: "true"
|
||||
- name: CUDA_COMPUTE_CAP
|
||||
value: "{{workflow.parameters.cuda-compute-cap}}"
|
||||
resources:
|
||||
# L40S-1-48G has 8 vCPU (7800m allocatable, after kubelet
|
||||
# overhead) and ~91Gi memory. Request <7800m so the pod
|
||||
# fits alongside the system daemonsets (cilium, csi-node,
|
||||
# nvidia drivers/dcgm/feature-discovery) on the same node.
|
||||
requests:
|
||||
cpu: "6"
|
||||
memory: 16Gi
|
||||
nvidia.com/gpu: "1"
|
||||
limits:
|
||||
cpu: "7"
|
||||
memory: 64Gi
|
||||
nvidia.com/gpu: "1"
|
||||
volumeMounts:
|
||||
- name: training-data
|
||||
mountPath: /data
|
||||
- name: feature-cache
|
||||
mountPath: /feature-cache
|
||||
args:
|
||||
- |
|
||||
set -e
|
||||
SHA="{{inputs.parameters.sha}}"
|
||||
BIN="/data/bin/$SHA/alpha_train"
|
||||
MBP10_DIR="/data/futures-baseline-mbp10/ES.FUT"
|
||||
PREDECODED_DIR="/feature-cache/predecoded"
|
||||
OUT_DIR="/feature-cache/alpha-perception-runs/$SHA"
|
||||
mkdir -p "$OUT_DIR" "$PREDECODED_DIR"
|
||||
|
||||
if [ ! -x "$BIN" ]; then
|
||||
echo "ERROR: binary not found at $BIN"
|
||||
ls -lh /data/bin/ || true
|
||||
exit 1
|
||||
fi
|
||||
if [ ! -d "$MBP10_DIR" ]; then
|
||||
echo "ERROR: MBP-10 data directory not found at $MBP10_DIR"
|
||||
ls -lh /data/ || true
|
||||
exit 1
|
||||
fi
|
||||
echo "Running alpha_train (stacked Mamba2 -> CfC -> heads)"
|
||||
echo " binary: $BIN"
|
||||
echo " mbp10: $MBP10_DIR ($(find "$MBP10_DIR" -name '*.dbn.zst' | wc -l) files)"
|
||||
echo " predecoded: $PREDECODED_DIR"
|
||||
echo " out: $OUT_DIR"
|
||||
|
||||
EXTRA_FLAGS=""
|
||||
if [ "{{workflow.parameters.auto-horizon-weights}}" = "true" ]; then
|
||||
EXTRA_FLAGS="$EXTRA_FLAGS --auto-horizon-weights"
|
||||
fi
|
||||
if [ -n "{{workflow.parameters.instrument-mode}}" ]; then
|
||||
EXTRA_FLAGS="$EXTRA_FLAGS --instrument-mode {{workflow.parameters.instrument-mode}}"
|
||||
fi
|
||||
|
||||
TRACE_FLAG=""
|
||||
if [ -n "{{workflow.parameters.kernel-step-trace-path}}" ]; then
|
||||
TRACE_FLAG="--kernel-step-trace {{workflow.parameters.kernel-step-trace-path}}"
|
||||
fi
|
||||
|
||||
"$BIN" \
|
||||
--mbp10-data-dir "$MBP10_DIR" \
|
||||
--predecoded-dir "$PREDECODED_DIR" \
|
||||
--out "$OUT_DIR" \
|
||||
--epochs {{workflow.parameters.epochs}} \
|
||||
--multi-resolution "{{workflow.parameters.multi-resolution}}" \
|
||||
--mamba2-state-dim {{workflow.parameters.mamba2-state-dim}} \
|
||||
--lr-cfc {{workflow.parameters.lr-cfc}} \
|
||||
--lr-mamba2 {{workflow.parameters.lr-mamba2}} \
|
||||
--n-train-seqs {{workflow.parameters.n-train-seqs}} \
|
||||
--n-val-seqs {{workflow.parameters.n-val-seqs}} \
|
||||
--seed {{workflow.parameters.seed}} \
|
||||
--batch-size {{workflow.parameters.batch-size}} \
|
||||
--early-stop-metric {{workflow.parameters.early-stop-metric}} \
|
||||
--early-stop-patience {{workflow.parameters.early-stop-patience}} \
|
||||
--cv-fold {{workflow.parameters.cv-fold}} \
|
||||
--cv-n-folds {{workflow.parameters.cv-n-folds}} \
|
||||
--cv-train-window {{workflow.parameters.cv-train-window}} \
|
||||
--smoothness-base-lambda {{workflow.parameters.smoothness-base-lambda}} \
|
||||
${TRACE_FLAG} \
|
||||
$EXTRA_FLAGS
|
||||
|
||||
echo "=== Training complete ==="
|
||||
ls -lh "$OUT_DIR/"
|
||||
echo "--- alpha_train_summary.json ---"
|
||||
cat "$OUT_DIR/alpha_train_summary.json"
|
||||
@@ -1,207 +0,0 @@
|
||||
# alpha-rl: single-pod compile + train on GPU.
|
||||
#
|
||||
# Compiles alpha_rl_train incrementally on the GPU node (~3s warm,
|
||||
# ~90s cold) using the cargo-target-cuda PVC, then runs training
|
||||
# immediately. No separate compile node, no binary transfer, no
|
||||
# fxcache step. Predecoded MBP-10 sidecars live on feature-cache PVC.
|
||||
#
|
||||
# Usage:
|
||||
# argo submit -n foxhunt --from=wftmpl/alpha-rl \
|
||||
# -p git-branch=ml-alpha-phase-a
|
||||
# argo submit -n foxhunt --from=wftmpl/alpha-rl \
|
||||
# -p git-branch=ml-alpha-phase-a -p n-steps=50000 -p n-backtests=128
|
||||
apiVersion: argoproj.io/v1alpha1
|
||||
kind: WorkflowTemplate
|
||||
metadata:
|
||||
name: alpha-rl
|
||||
namespace: foxhunt
|
||||
labels:
|
||||
app.kubernetes.io/name: alpha-rl
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
app.kubernetes.io/component: train
|
||||
spec:
|
||||
entrypoint: compile-and-train
|
||||
serviceAccountName: argo-workflow
|
||||
archiveLogs: true
|
||||
podMetadata:
|
||||
labels:
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
app.kubernetes.io/component: train
|
||||
securityContext:
|
||||
fsGroup: 0
|
||||
ttlStrategy:
|
||||
secondsAfterCompletion: 3600
|
||||
activeDeadlineSeconds: 14400
|
||||
|
||||
arguments:
|
||||
parameters:
|
||||
- name: git-branch
|
||||
value: ml-alpha-phase-a
|
||||
- name: gpu-pool
|
||||
value: ci-training-l40s
|
||||
- name: n-steps
|
||||
value: "50000"
|
||||
- name: n-backtests
|
||||
value: "128"
|
||||
- name: seed
|
||||
value: "16962"
|
||||
- name: seq-len
|
||||
value: "32"
|
||||
- name: per-capacity
|
||||
value: "32768"
|
||||
- name: instrument-mode
|
||||
value: "all"
|
||||
- name: log-every
|
||||
value: "5000"
|
||||
- name: fold-idx
|
||||
value: "0"
|
||||
- name: n-folds
|
||||
value: "1"
|
||||
- name: n-eval-steps
|
||||
value: "0"
|
||||
- name: nsys-profile
|
||||
value: "false"
|
||||
- name: use-multi-head-policy
|
||||
value: "0"
|
||||
- name: band-enabled
|
||||
value: "0"
|
||||
|
||||
volumes:
|
||||
- name: git-ssh-key
|
||||
secret:
|
||||
secretName: argo-git-ssh-key
|
||||
defaultMode: 256
|
||||
- name: training-data
|
||||
persistentVolumeClaim:
|
||||
claimName: training-data-pvc
|
||||
- name: cargo-target-cuda
|
||||
persistentVolumeClaim:
|
||||
claimName: cargo-target-cuda
|
||||
- name: feature-cache
|
||||
persistentVolumeClaim:
|
||||
claimName: feature-cache-pvc
|
||||
|
||||
templates:
|
||||
- name: compile-and-train
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: "{{workflow.parameters.gpu-pool}}"
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
tolerations:
|
||||
- key: nvidia.com/gpu
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
- key: node.cilium.io/agent-not-ready
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
container:
|
||||
image: gitlab-registry.foxhunt.svc.cluster.local:5000/root/foxhunt/ci-builder:latest
|
||||
imagePullPolicy: Always
|
||||
command: ["/bin/bash", "-c"]
|
||||
env:
|
||||
- name: SQLX_OFFLINE
|
||||
value: "true"
|
||||
- name: CARGO_TARGET_DIR
|
||||
value: /cargo-target
|
||||
- name: CARGO_HOME
|
||||
value: /cargo-target/cargo-home
|
||||
- name: CARGO_INCREMENTAL
|
||||
value: "1"
|
||||
- name: FOXHUNT_USE_MULTI_HEAD_POLICY
|
||||
value: "{{workflow.parameters.use-multi-head-policy}}"
|
||||
- name: FOXHUNT_BAND_ENABLED
|
||||
value: "{{workflow.parameters.band-enabled}}"
|
||||
resources:
|
||||
requests:
|
||||
cpu: "4"
|
||||
memory: 16Gi
|
||||
nvidia.com/gpu: "1"
|
||||
limits:
|
||||
cpu: "7"
|
||||
memory: 40Gi
|
||||
nvidia.com/gpu: "1"
|
||||
volumeMounts:
|
||||
- name: git-ssh-key
|
||||
mountPath: /etc/git-ssh
|
||||
readOnly: true
|
||||
- name: training-data
|
||||
mountPath: /data
|
||||
readOnly: true
|
||||
- name: feature-cache
|
||||
mountPath: /feature-cache
|
||||
- name: cargo-target-cuda
|
||||
mountPath: /cargo-target
|
||||
args:
|
||||
- |
|
||||
set -e
|
||||
export PATH="${CARGO_HOME}/bin:${PATH}"
|
||||
export LD_LIBRARY_PATH=$(echo "$LD_LIBRARY_PATH" | tr ':' '\n' | grep -v stubs | tr '\n' ':' | sed 's/:$//')
|
||||
|
||||
# SSH
|
||||
mkdir -p ~/.ssh
|
||||
cp /etc/git-ssh/ssh-privatekey ~/.ssh/id_ed25519
|
||||
chmod 600 ~/.ssh/id_ed25519
|
||||
printf 'Host *\n StrictHostKeyChecking no\n UserKnownHostsFile /dev/null\n' > ~/.ssh/config
|
||||
chmod 600 ~/.ssh/config
|
||||
|
||||
nvidia-smi
|
||||
|
||||
# Git
|
||||
BRANCH="{{workflow.parameters.git-branch}}"
|
||||
BUILD="/cargo-target/src"
|
||||
REPO="ssh://git@gitlab-gitlab-shell.foxhunt.svc.cluster.local:2222/root/foxhunt.git"
|
||||
git config --global --add safe.directory "$BUILD"
|
||||
|
||||
if [ -d "$BUILD/.git" ]; then
|
||||
cd "$BUILD"
|
||||
git fetch origin
|
||||
git checkout --force "origin/$BRANCH"
|
||||
git clean -fd
|
||||
else
|
||||
git clone --filter=blob:none "$REPO" "$BUILD"
|
||||
cd "$BUILD"
|
||||
git checkout "origin/$BRANCH"
|
||||
fi
|
||||
SHA=$(git rev-parse --short=9 HEAD)
|
||||
echo "=== Branch: $BRANCH SHA: $SHA ==="
|
||||
|
||||
# Auto-detect GPU compute capability for all crates' build.rs
|
||||
export CUDA_COMPUTE_CAP=$(nvidia-smi --query-gpu=compute_cap --format=csv,noheader | head -1 | tr -d '.')
|
||||
echo "=== GPU arch: sm_${CUDA_COMPUTE_CAP} ==="
|
||||
|
||||
# Clear stale build artifacts when arch changes
|
||||
rm -rf ${CARGO_TARGET_DIR}/build/ml-alpha-* ${CARGO_TARGET_DIR}/build/ml-core-*
|
||||
|
||||
# Compile (~90s cold, ~3s warm after first run)
|
||||
echo "=== Compile ==="
|
||||
time cargo build --release -p ml-alpha --example alpha_rl_train 2>&1
|
||||
|
||||
# Train
|
||||
OUT="/feature-cache/alpha-rl-runs/$SHA/fold{{workflow.parameters.fold-idx}}"
|
||||
mkdir -p "$OUT"
|
||||
echo "=== Train on $(nvidia-smi --query-gpu=name --format=csv,noheader) ==="
|
||||
TRAIN_BIN="${CARGO_TARGET_DIR}/release/examples/alpha_rl_train"
|
||||
if [ "{{workflow.parameters.nsys-profile}}" = "true" ]; then
|
||||
TRAIN_CMD="nsys profile -o $OUT/nsys_trace --stats=true --force-overwrite=true $TRAIN_BIN"
|
||||
else
|
||||
TRAIN_CMD="stdbuf -oL $TRAIN_BIN"
|
||||
fi
|
||||
$TRAIN_CMD \
|
||||
--mbp10-data-dir /data/futures-baseline-mbp10/ES.FUT \
|
||||
--predecoded-dir /feature-cache/predecoded \
|
||||
--out "$OUT" \
|
||||
--n-steps {{workflow.parameters.n-steps}} \
|
||||
--seq-len {{workflow.parameters.seq-len}} \
|
||||
--n-backtests {{workflow.parameters.n-backtests}} \
|
||||
--per-capacity {{workflow.parameters.per-capacity}} \
|
||||
--seed {{workflow.parameters.seed}} \
|
||||
--instrument-mode "{{workflow.parameters.instrument-mode}}" \
|
||||
--fold-idx {{workflow.parameters.fold-idx}} \
|
||||
--n-folds {{workflow.parameters.n-folds}} \
|
||||
--n-eval-steps {{workflow.parameters.n-eval-steps}} \
|
||||
--log-every {{workflow.parameters.log-every}}
|
||||
|
||||
echo "=== Complete: $OUT ==="
|
||||
ls -lh "$OUT/"
|
||||
if [ -f "$OUT/alpha_rl_train_summary.json" ]; then
|
||||
cat "$OUT/alpha_rl_train_summary.json"
|
||||
fi
|
||||
@@ -1,121 +0,0 @@
|
||||
apiVersion: argoproj.io/v1alpha1
|
||||
kind: WorkflowTemplate
|
||||
metadata:
|
||||
name: build-ci-image
|
||||
namespace: foxhunt
|
||||
labels:
|
||||
app.kubernetes.io/name: build-ci-image
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
spec:
|
||||
activeDeadlineSeconds: 3600
|
||||
serviceAccountName: argo-workflow
|
||||
entrypoint: build
|
||||
podMetadata:
|
||||
labels:
|
||||
app.kubernetes.io/component: ci-pipeline
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
ttlStrategy:
|
||||
secondsAfterCompletion: 3600
|
||||
arguments:
|
||||
parameters:
|
||||
- name: commit-sha
|
||||
value: HEAD
|
||||
- name: dockerfile
|
||||
value: Dockerfile.ci-builder
|
||||
- name: image-name
|
||||
value: ci-builder
|
||||
|
||||
templates:
|
||||
# Single-pod build: init container clones repo, kaniko builds image.
|
||||
# Uses emptyDir — no PVC needed, works when called via templateRef.
|
||||
- name: build
|
||||
inputs:
|
||||
parameters:
|
||||
- name: commit-sha
|
||||
value: "{{workflow.parameters.commit-sha}}"
|
||||
- name: dockerfile
|
||||
value: "{{workflow.parameters.dockerfile}}"
|
||||
- name: image-name
|
||||
value: "{{workflow.parameters.image-name}}"
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: ci-compile-cpu
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
tolerations:
|
||||
- key: node.cilium.io/agent-not-ready
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
volumes:
|
||||
- name: workspace
|
||||
emptyDir: {}
|
||||
- name: git-ssh-key
|
||||
secret:
|
||||
secretName: argo-git-ssh-key
|
||||
defaultMode: 256
|
||||
- name: registry-auth
|
||||
secret:
|
||||
secretName: gitlab-registry
|
||||
items:
|
||||
- key: .dockerconfigjson
|
||||
path: config.json
|
||||
initContainers:
|
||||
- name: git-clone
|
||||
image: alpine/git:latest
|
||||
command: ["/bin/sh", "-c"]
|
||||
resources:
|
||||
requests:
|
||||
cpu: 200m
|
||||
memory: 256Mi
|
||||
limits:
|
||||
cpu: "1"
|
||||
memory: 512Mi
|
||||
volumeMounts:
|
||||
- name: git-ssh-key
|
||||
mountPath: /etc/git-ssh
|
||||
readOnly: true
|
||||
- name: workspace
|
||||
mountPath: /workspace
|
||||
args:
|
||||
- |
|
||||
set -ex
|
||||
mkdir -p /root/.ssh
|
||||
cp /etc/git-ssh/ssh-privatekey /root/.ssh/id_ed25519
|
||||
chmod 600 /root/.ssh/id_ed25519
|
||||
printf 'Host *\n StrictHostKeyChecking no\n UserKnownHostsFile /dev/null\n' > /root/.ssh/config
|
||||
chmod 600 /root/.ssh/config
|
||||
|
||||
SHA="{{inputs.parameters.commit-sha}}"
|
||||
REPO="ssh://git@gitlab-gitlab-shell.foxhunt.svc.cluster.local:2222/root/foxhunt.git"
|
||||
git clone --no-checkout --filter=blob:none "$REPO" /workspace/src
|
||||
cd /workspace/src
|
||||
git checkout "$SHA"
|
||||
echo "Checked out $(git rev-parse --short HEAD)"
|
||||
container:
|
||||
image: gcr.io/kaniko-project/executor:debug
|
||||
command: ["/busybox/sh", "-c"]
|
||||
env:
|
||||
- name: DOCKER_CONFIG
|
||||
value: /kaniko/.docker
|
||||
resources:
|
||||
requests:
|
||||
cpu: "4"
|
||||
memory: 8Gi
|
||||
limits:
|
||||
cpu: "8"
|
||||
memory: 16Gi
|
||||
volumeMounts:
|
||||
- name: registry-auth
|
||||
mountPath: /kaniko/.docker
|
||||
readOnly: true
|
||||
- name: workspace
|
||||
mountPath: /workspace
|
||||
args:
|
||||
- |
|
||||
/kaniko/executor \
|
||||
--context=/workspace/src \
|
||||
--dockerfile=/workspace/src/infra/docker/{{inputs.parameters.dockerfile}} \
|
||||
--destination=gitlab-registry.foxhunt.svc.cluster.local:5000/root/foxhunt/{{inputs.parameters.image-name}}:latest \
|
||||
--insecure-registry=gitlab-registry.foxhunt.svc.cluster.local:5000 \
|
||||
--skip-tls-verify-registry=gitlab-registry.foxhunt.svc.cluster.local:5000 \
|
||||
--cache=true \
|
||||
--cache-repo=gitlab-registry.foxhunt.svc.cluster.local:5000/root/foxhunt/cache \
|
||||
--snapshot-mode=redo
|
||||
@@ -1,16 +0,0 @@
|
||||
# infra/k8s/argo/cargo-target-cuda-test-pvc.yaml
|
||||
apiVersion: v1
|
||||
kind: PersistentVolumeClaim
|
||||
metadata:
|
||||
name: cargo-target-cuda-test
|
||||
namespace: foxhunt
|
||||
labels:
|
||||
app.kubernetes.io/name: cargo-target
|
||||
app.kubernetes.io/component: ci-cache
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
spec:
|
||||
accessModes: [ReadWriteOnce]
|
||||
storageClassName: scw-bssd-retain
|
||||
resources:
|
||||
requests:
|
||||
storage: 30Gi
|
||||
@@ -1,43 +0,0 @@
|
||||
# Persistent cargo target directories for incremental compilation.
|
||||
#
|
||||
# Why: sccache can't cache workspace rlib crates (109 non-cacheable per build).
|
||||
# Persisting target/ lets cargo's incremental compilation skip unchanged crates,
|
||||
# reducing typical CI builds from ~20 min (full rebuild) to ~2-3 min.
|
||||
#
|
||||
# Two separate PVCs because compile-services (no cuda feature) and
|
||||
# compile-training (cuda feature) produce incompatible artifacts.
|
||||
#
|
||||
# NOTE: Only one workflow should use each PVC at a time. ci-pipeline and
|
||||
# compile-and-train must not run their compile steps concurrently.
|
||||
---
|
||||
apiVersion: v1
|
||||
kind: PersistentVolumeClaim
|
||||
metadata:
|
||||
name: cargo-target-cpu
|
||||
namespace: foxhunt
|
||||
labels:
|
||||
app.kubernetes.io/name: cargo-target
|
||||
app.kubernetes.io/component: ci-cache
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
spec:
|
||||
accessModes: [ReadWriteOnce]
|
||||
storageClassName: scw-bssd-retain
|
||||
resources:
|
||||
requests:
|
||||
storage: 30Gi
|
||||
---
|
||||
apiVersion: v1
|
||||
kind: PersistentVolumeClaim
|
||||
metadata:
|
||||
name: cargo-target-cuda
|
||||
namespace: foxhunt
|
||||
labels:
|
||||
app.kubernetes.io/name: cargo-target
|
||||
app.kubernetes.io/component: ci-cache
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
spec:
|
||||
accessModes: [ReadWriteOnce]
|
||||
storageClassName: scw-bssd-retain
|
||||
resources:
|
||||
requests:
|
||||
storage: 30Gi
|
||||
@@ -1,685 +0,0 @@
|
||||
apiVersion: argoproj.io/v1alpha1
|
||||
kind: WorkflowTemplate
|
||||
metadata:
|
||||
name: ci-pipeline
|
||||
namespace: foxhunt
|
||||
labels:
|
||||
app.kubernetes.io/name: ci-pipeline
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
spec:
|
||||
activeDeadlineSeconds: 7200
|
||||
serviceAccountName: argo-workflow
|
||||
entrypoint: pipeline
|
||||
onExit: notify-result
|
||||
podMetadata:
|
||||
labels:
|
||||
app.kubernetes.io/component: ci-pipeline
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
archiveLogs: true
|
||||
podGC:
|
||||
strategy: OnPodCompletion
|
||||
ttlStrategy:
|
||||
secondsAfterCompletion: 3600
|
||||
volumes:
|
||||
- name: git-ssh-key
|
||||
secret:
|
||||
secretName: argo-git-ssh-key
|
||||
defaultMode: 256
|
||||
- name: registry-auth
|
||||
secret:
|
||||
secretName: gitlab-registry
|
||||
optional: true
|
||||
items:
|
||||
- key: .dockerconfigjson
|
||||
path: config.json
|
||||
- name: cargo-target-cpu
|
||||
persistentVolumeClaim:
|
||||
claimName: cargo-target-cpu
|
||||
arguments:
|
||||
parameters:
|
||||
- name: commit-sha
|
||||
value: HEAD
|
||||
- name: commits-json
|
||||
value: "[]"
|
||||
templates:
|
||||
# ── DAG: orchestrate pipeline ──
|
||||
# Policy (user directive): push-triggered runs are limited to infrastructure
|
||||
# sync work — docker image rebuilds + argo-template apply + terragrunt apply.
|
||||
# These run automatically when the relevant paths change:
|
||||
# - rebuild-ci-builder* / rebuild-runtime / rebuild-training-runtime: gated on
|
||||
# `detect-changes.docker-images == true` (infra/docker/)
|
||||
# - apply-argo-templates: gated on `detect-changes.needs-argo-templates` (infra/k8s/argo/)
|
||||
# - terragrunt-apply: gated on `detect-changes.needs-infra` (infra/live/, infra/modules/)
|
||||
# All other tasks (test-gate, build-web-dashboard, gpu-test) are `when: "false"` —
|
||||
# they run via manual script invocation (`argo submit` / `./scripts/argo-*.sh`)
|
||||
# or the manual `/compile-deploy` webhook. Training + compile/deploy workflows
|
||||
# are separate templates, always manual.
|
||||
- name: pipeline
|
||||
dag:
|
||||
tasks:
|
||||
- name: detect-changes
|
||||
template: detect-changes
|
||||
|
||||
- name: build-web-dashboard
|
||||
depends: "detect-changes"
|
||||
template: build-web-dashboard
|
||||
# Policy: push-triggered runs are docker image rebuilds only.
|
||||
# Dashboard builds run via manual `argo submit` (scripts).
|
||||
when: "false"
|
||||
|
||||
- name: rebuild-ci-builder
|
||||
depends: "detect-changes"
|
||||
templateRef:
|
||||
name: build-ci-image
|
||||
template: build
|
||||
arguments:
|
||||
parameters:
|
||||
- name: commit-sha
|
||||
value: "{{workflow.parameters.commit-sha}}"
|
||||
- name: dockerfile
|
||||
value: Dockerfile.ci-builder
|
||||
- name: image-name
|
||||
value: ci-builder
|
||||
when: "{{tasks.detect-changes.outputs.parameters.docker-images}} == true"
|
||||
|
||||
- name: rebuild-ci-builder-cpu
|
||||
depends: "detect-changes"
|
||||
templateRef:
|
||||
name: build-ci-image
|
||||
template: build
|
||||
arguments:
|
||||
parameters:
|
||||
- name: commit-sha
|
||||
value: "{{workflow.parameters.commit-sha}}"
|
||||
- name: dockerfile
|
||||
value: Dockerfile.ci-builder-cpu
|
||||
- name: image-name
|
||||
value: ci-builder-cpu
|
||||
when: "{{tasks.detect-changes.outputs.parameters.docker-images}} == true"
|
||||
|
||||
- name: rebuild-runtime
|
||||
depends: "detect-changes"
|
||||
templateRef:
|
||||
name: build-ci-image
|
||||
template: build
|
||||
arguments:
|
||||
parameters:
|
||||
- name: commit-sha
|
||||
value: "{{workflow.parameters.commit-sha}}"
|
||||
- name: dockerfile
|
||||
value: Dockerfile.foxhunt-runtime
|
||||
- name: image-name
|
||||
value: foxhunt-runtime
|
||||
when: "{{tasks.detect-changes.outputs.parameters.docker-images}} == true"
|
||||
|
||||
- name: rebuild-training-runtime
|
||||
depends: "detect-changes"
|
||||
templateRef:
|
||||
name: build-ci-image
|
||||
template: build
|
||||
arguments:
|
||||
parameters:
|
||||
- name: commit-sha
|
||||
value: "{{workflow.parameters.commit-sha}}"
|
||||
- name: dockerfile
|
||||
value: Dockerfile.foxhunt-training-runtime
|
||||
- name: image-name
|
||||
value: foxhunt-training-runtime
|
||||
when: "{{tasks.detect-changes.outputs.parameters.docker-images}} == true"
|
||||
|
||||
- name: test-gate
|
||||
depends: "detect-changes && (rebuild-ci-builder-cpu.Succeeded || rebuild-ci-builder-cpu.Skipped)"
|
||||
template: test-gate
|
||||
arguments:
|
||||
parameters:
|
||||
- name: commit-sha
|
||||
value: "{{workflow.parameters.commit-sha}}"
|
||||
# Policy: push-triggered runs are docker image rebuilds only.
|
||||
# Cargo test + clippy quality gate runs via manual `argo submit`
|
||||
# (scripts/argo-test.sh or CI-driven `argo submit --from=wftmpl/ci-pipeline`).
|
||||
when: "false"
|
||||
|
||||
- name: apply-argo-templates
|
||||
depends: "detect-changes.Succeeded && (rebuild-ci-builder-cpu.Succeeded || rebuild-ci-builder-cpu.Skipped)"
|
||||
template: apply-argo-templates
|
||||
arguments:
|
||||
parameters:
|
||||
- name: commit-sha
|
||||
value: "{{workflow.parameters.commit-sha}}"
|
||||
when: "{{tasks.detect-changes.outputs.parameters.needs-argo-templates}} == true"
|
||||
|
||||
- name: terragrunt-apply
|
||||
depends: "detect-changes.Succeeded && (rebuild-ci-builder-cpu.Succeeded || rebuild-ci-builder-cpu.Skipped)"
|
||||
template: terragrunt-apply
|
||||
arguments:
|
||||
parameters:
|
||||
- name: commit-sha
|
||||
value: "{{workflow.parameters.commit-sha}}"
|
||||
when: "{{tasks.detect-changes.outputs.parameters.needs-infra}} == true"
|
||||
|
||||
# GPU tests: manual only (argo submit --from=wftmpl/gpu-test-pipeline).
|
||||
# Auto-trigger disabled — H100 nodes are expensive, run on demand.
|
||||
- name: gpu-test
|
||||
depends: "detect-changes"
|
||||
template: submit-gpu-test
|
||||
when: "false"
|
||||
|
||||
# ── detect-changes: classify changed files to gate downstream steps ──
|
||||
- name: detect-changes
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: platform
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
container:
|
||||
image: gitlab-registry.foxhunt.svc.cluster.local:5000/root/foxhunt/foxhunt-runtime:latest
|
||||
command: ["/bin/sh", "-c"]
|
||||
env:
|
||||
- name: COMMITS_JSON
|
||||
value: "{{workflow.parameters.commits-json}}"
|
||||
resources:
|
||||
requests:
|
||||
cpu: 100m
|
||||
memory: 128Mi
|
||||
limits:
|
||||
cpu: 500m
|
||||
memory: 256Mi
|
||||
args:
|
||||
- |
|
||||
set -e
|
||||
cat <<'SCRIPT' > /tmp/detect.sh
|
||||
#!/bin/sh
|
||||
set -e
|
||||
|
||||
# Extract changed file paths from webhook JSON using grep+sed (no jq dependency).
|
||||
CHANGED_FILES=$(echo "$COMMITS_JSON" \
|
||||
| grep -oE '"(added|modified|removed)":\[[^]]*\]' \
|
||||
| sed 's/"added"://;s/"modified"://;s/"removed"://' \
|
||||
| tr ',' '\n' | tr -d '[]"' | sed '/^$/d' | sort -u \
|
||||
|| echo "")
|
||||
|
||||
if [ -z "$CHANGED_FILES" ]; then
|
||||
echo "No changed files detected — triggering full rebuild"
|
||||
CHANGED_FILES="Cargo.toml"
|
||||
fi
|
||||
|
||||
echo "=== Changed files ==="
|
||||
echo "$CHANGED_FILES"
|
||||
echo "===================="
|
||||
|
||||
# --- Helper: check if any changed file matches a set of path prefixes ---
|
||||
check_paths() {
|
||||
local patterns="$1"
|
||||
for file in $CHANGED_FILES; do
|
||||
for pattern in $patterns; do
|
||||
case "$file" in
|
||||
${pattern}*) echo "true"; return ;;
|
||||
esac
|
||||
done
|
||||
done
|
||||
echo "false"
|
||||
}
|
||||
|
||||
# --- Boolean gates for downstream DAG tasks ---
|
||||
NEEDS_CODE=$(check_paths "Cargo.toml Cargo.lock crates/ services/ bin/fxt/")
|
||||
NEEDS_DASHBOARD=$(check_paths "web-dashboard/")
|
||||
DOCKER_IMAGES=$(check_paths "infra/docker/")
|
||||
NEEDS_INFRA=$(check_paths "infra/live/ infra/modules/")
|
||||
NEEDS_ARGO_TEMPLATES=$(check_paths "infra/k8s/argo/")
|
||||
ML_CHANGED=$(check_paths "crates/ml crates/ml-dqn crates/ml-core crates/ml-regime crates/ml-features crates/ml-supervised crates/ml-ppo crates/ml-ensemble crates/ml-hyperopt crates/ml-labeling config/training config/gpu")
|
||||
|
||||
echo "=== Build decisions ==="
|
||||
echo "needs-code: $NEEDS_CODE"
|
||||
echo "needs-dashboard: $NEEDS_DASHBOARD"
|
||||
echo "docker-images: $DOCKER_IMAGES"
|
||||
echo "needs-infra: $NEEDS_INFRA"
|
||||
echo "needs-argo: $NEEDS_ARGO_TEMPLATES"
|
||||
echo "ml-changed: $ML_CHANGED"
|
||||
echo "======================"
|
||||
|
||||
mkdir -p /tmp/outputs
|
||||
echo -n "$NEEDS_CODE" > /tmp/outputs/needs-code
|
||||
echo -n "$NEEDS_DASHBOARD" > /tmp/outputs/needs-dashboard
|
||||
echo -n "$DOCKER_IMAGES" > /tmp/outputs/docker-images
|
||||
echo -n "$NEEDS_INFRA" > /tmp/outputs/needs-infra
|
||||
echo -n "$NEEDS_ARGO_TEMPLATES" > /tmp/outputs/needs-argo-templates
|
||||
echo -n "$ML_CHANGED" > /tmp/outputs/ml-changed
|
||||
SCRIPT
|
||||
chmod +x /tmp/detect.sh
|
||||
/tmp/detect.sh
|
||||
outputs:
|
||||
parameters:
|
||||
- name: needs-code
|
||||
valueFrom:
|
||||
path: /tmp/outputs/needs-code
|
||||
- name: needs-dashboard
|
||||
valueFrom:
|
||||
path: /tmp/outputs/needs-dashboard
|
||||
- name: docker-images
|
||||
valueFrom:
|
||||
path: /tmp/outputs/docker-images
|
||||
- name: needs-infra
|
||||
valueFrom:
|
||||
path: /tmp/outputs/needs-infra
|
||||
- name: needs-argo-templates
|
||||
valueFrom:
|
||||
path: /tmp/outputs/needs-argo-templates
|
||||
- name: ml-changed
|
||||
valueFrom:
|
||||
path: /tmp/outputs/ml-changed
|
||||
|
||||
# ── test-gate: clippy + cargo test quality gate ──
|
||||
# Uses ci-builder (with CUDA toolkit) because cudarc 0.19 requires nvcc at build time.
|
||||
- name: test-gate
|
||||
inputs:
|
||||
parameters:
|
||||
- name: commit-sha
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: ci-compile-cpu
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
tolerations:
|
||||
- key: k8s.scaleway.com/pool-name
|
||||
operator: Equal
|
||||
value: ci-compile-cpu
|
||||
effect: NoSchedule
|
||||
sidecars:
|
||||
- name: redis
|
||||
image: redis:7-alpine
|
||||
command: ["redis-server", "--save", "", "--appendonly", "no"]
|
||||
ports:
|
||||
- containerPort: 6379
|
||||
resources:
|
||||
requests:
|
||||
cpu: 100m
|
||||
memory: 64Mi
|
||||
limits:
|
||||
cpu: 500m
|
||||
memory: 128Mi
|
||||
readinessProbe:
|
||||
tcpSocket:
|
||||
port: 6379
|
||||
initialDelaySeconds: 1
|
||||
periodSeconds: 1
|
||||
container:
|
||||
image: gitlab-registry.foxhunt.svc.cluster.local:5000/root/foxhunt/ci-builder:latest
|
||||
imagePullPolicy: Always
|
||||
command: ["/bin/sh", "-c"]
|
||||
env:
|
||||
- name: SQLX_OFFLINE
|
||||
value: "true"
|
||||
- name: CARGO_TERM_COLOR
|
||||
value: always
|
||||
- name: CARGO_TARGET_DIR
|
||||
value: /cargo-target
|
||||
- name: CARGO_HOME
|
||||
value: /cargo-target/cargo-home
|
||||
- name: REDIS_URL
|
||||
value: "redis://localhost:6379"
|
||||
resources:
|
||||
requests:
|
||||
cpu: "14"
|
||||
memory: 16Gi
|
||||
limits:
|
||||
cpu: "30"
|
||||
memory: 32Gi
|
||||
volumeMounts:
|
||||
- name: git-ssh-key
|
||||
mountPath: /etc/git-ssh
|
||||
readOnly: true
|
||||
- name: cargo-target-cpu
|
||||
mountPath: /cargo-target
|
||||
args:
|
||||
- |
|
||||
set -e
|
||||
SHA="{{inputs.parameters.commit-sha}}"
|
||||
|
||||
# SSH setup
|
||||
mkdir -p ~/.ssh
|
||||
cp /etc/git-ssh/ssh-privatekey ~/.ssh/id_ed25519
|
||||
chmod 600 ~/.ssh/id_ed25519
|
||||
printf 'Host *\n StrictHostKeyChecking no\n UserKnownHostsFile /dev/null\n' > ~/.ssh/config
|
||||
chmod 600 ~/.ssh/config
|
||||
|
||||
REPO="ssh://git@gitlab-gitlab-shell.foxhunt.svc.cluster.local:2222/root/foxhunt.git"
|
||||
WORKSPACE="/cargo-target/src"
|
||||
|
||||
git config --global --add safe.directory "$WORKSPACE"
|
||||
|
||||
if [ -d "$WORKSPACE/.git" ]; then
|
||||
cd "$WORKSPACE"
|
||||
CURRENT=$(git rev-parse HEAD 2>/dev/null || echo "none")
|
||||
if [ "$CURRENT" = "$SHA" ]; then
|
||||
echo "=== Already at $SHA ==="
|
||||
else
|
||||
git fetch origin
|
||||
git checkout --force "$SHA"
|
||||
git clean -fd
|
||||
fi
|
||||
else
|
||||
git clone --filter=blob:none "$REPO" "$WORKSPACE"
|
||||
cd "$WORKSPACE"
|
||||
git checkout "$SHA"
|
||||
fi
|
||||
|
||||
export PATH="${CARGO_HOME}/bin:${PATH}"
|
||||
|
||||
echo "=== Waiting for Redis sidecar ==="
|
||||
for i in 1 2 3 4 5; do
|
||||
if printf 'PING\r\n' | nc -w1 127.0.0.1 6379 2>/dev/null | grep -q PONG; then
|
||||
echo "Redis ready"
|
||||
break
|
||||
fi
|
||||
sleep 1
|
||||
done
|
||||
|
||||
echo "=== Running clippy (lib targets) ==="
|
||||
cargo clippy --workspace --lib -- -D warnings 2>&1 | tee /cargo-target/clippy.log
|
||||
|
||||
echo "=== Running tests (lib only, no integration) ==="
|
||||
set +e
|
||||
cargo test --workspace --lib 2>&1 | tee /cargo-target/test-output.log
|
||||
TEST_EXIT=$?
|
||||
set -e
|
||||
if [ "$TEST_EXIT" -ne 0 ]; then
|
||||
echo "=== TEST FAILURES ==="
|
||||
grep -A5 'FAILED\|panicked\|test result: FAILED\|error\[' /cargo-target/test-output.log || true
|
||||
exit "$TEST_EXIT"
|
||||
fi
|
||||
|
||||
echo "=== Test gate passed ==="
|
||||
|
||||
# ── build-web-dashboard: npm build + upload to MinIO ──
|
||||
- name: build-web-dashboard
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: platform
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
container:
|
||||
image: node:22-alpine
|
||||
command: ["/bin/sh", "-c"]
|
||||
env:
|
||||
- name: MINIO_ACCESS_KEY
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: minio-credentials
|
||||
key: access-key
|
||||
- name: MINIO_SECRET_KEY
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: minio-credentials
|
||||
key: secret-key
|
||||
resources:
|
||||
requests:
|
||||
cpu: "1"
|
||||
memory: 1Gi
|
||||
limits:
|
||||
cpu: "2"
|
||||
memory: 2Gi
|
||||
volumeMounts:
|
||||
- name: git-ssh-key
|
||||
mountPath: /etc/git-ssh
|
||||
readOnly: true
|
||||
args:
|
||||
- |
|
||||
set -e
|
||||
SHA="{{workflow.parameters.commit-sha}}"
|
||||
|
||||
# Git clone
|
||||
apk add --no-cache git openssh rclone
|
||||
mkdir -p ~/.ssh
|
||||
cp /etc/git-ssh/ssh-privatekey ~/.ssh/id_ed25519
|
||||
chmod 600 ~/.ssh/id_ed25519
|
||||
printf 'Host *\n StrictHostKeyChecking no\n UserKnownHostsFile /dev/null\n' > ~/.ssh/config
|
||||
chmod 600 ~/.ssh/config
|
||||
|
||||
REPO="ssh://git@gitlab-gitlab-shell.foxhunt.svc.cluster.local:2222/root/foxhunt.git"
|
||||
WORKSPACE="/tmp/workspace"
|
||||
git clone --no-checkout --filter=blob:none "$REPO" "$WORKSPACE"
|
||||
cd "$WORKSPACE"
|
||||
git checkout "$SHA"
|
||||
|
||||
cd "$WORKSPACE/web-dashboard"
|
||||
echo "=== Building web dashboard ==="
|
||||
npm ci
|
||||
npm run build
|
||||
|
||||
echo "=== Uploading to MinIO ==="
|
||||
rclone copy dist/ :s3:foxhunt-binaries/web-dashboard/ \
|
||||
--s3-provider=Minio \
|
||||
--s3-endpoint=http://minio.foxhunt.svc.cluster.local:9000 \
|
||||
--s3-access-key-id="${MINIO_ACCESS_KEY}" \
|
||||
--s3-secret-access-key="${MINIO_SECRET_KEY}" \
|
||||
--s3-no-check-bucket
|
||||
|
||||
echo "=== Web dashboard build + upload done ==="
|
||||
|
||||
# ── terragrunt-apply: apply infra changes on main push ──
|
||||
- name: terragrunt-apply
|
||||
inputs:
|
||||
parameters:
|
||||
- name: commit-sha
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: platform
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
serviceAccountName: argo-workflow
|
||||
container:
|
||||
image: gitlab-registry.foxhunt.svc.cluster.local:5000/root/foxhunt/ci-builder-cpu:latest
|
||||
command: ["/bin/sh", "-c"]
|
||||
env:
|
||||
- name: GITLAB_TOKEN
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: gitlab-pat
|
||||
key: token
|
||||
- name: TF_HTTP_USERNAME
|
||||
value: root
|
||||
- name: TF_HTTP_PASSWORD
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: gitlab-pat
|
||||
key: token
|
||||
- name: SCW_ACCESS_KEY
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: scaleway-credentials
|
||||
key: access-key
|
||||
- name: SCW_SECRET_KEY
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: scaleway-credentials
|
||||
key: secret-key
|
||||
- name: SCW_DEFAULT_PROJECT_ID
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: scaleway-credentials
|
||||
key: project-id
|
||||
- name: GITLAB_TF_STATE_URL
|
||||
value: "http://gitlab-webservice-default.foxhunt.svc.cluster.local:8181"
|
||||
resources:
|
||||
requests:
|
||||
cpu: 200m
|
||||
memory: 256Mi
|
||||
limits:
|
||||
cpu: "1"
|
||||
memory: 512Mi
|
||||
volumeMounts:
|
||||
- name: git-ssh-key
|
||||
mountPath: /etc/git-ssh
|
||||
readOnly: true
|
||||
args:
|
||||
- |
|
||||
set -e
|
||||
|
||||
# Clone repo
|
||||
mkdir -p /root/.ssh
|
||||
cp /etc/git-ssh/ssh-privatekey /root/.ssh/id_ed25519
|
||||
chmod 600 /root/.ssh/id_ed25519
|
||||
printf 'Host *\n StrictHostKeyChecking no\n UserKnownHostsFile /dev/null\n' > /root/.ssh/config
|
||||
chmod 600 /root/.ssh/config
|
||||
|
||||
SHA="{{inputs.parameters.commit-sha}}"
|
||||
REPO="ssh://git@gitlab-gitlab-shell.foxhunt.svc.cluster.local:2222/root/foxhunt.git"
|
||||
git clone --no-checkout --filter=blob:none "$REPO" /workspace/src
|
||||
cd /workspace/src
|
||||
git checkout "$SHA"
|
||||
echo "Checked out $(git rev-parse --short HEAD)"
|
||||
|
||||
tofu --version
|
||||
terragrunt --version
|
||||
|
||||
# Apply each terragrunt module
|
||||
for module in block-storage dns public-gateway kapsule; do
|
||||
MODULE_DIR="infra/live/production/${module}"
|
||||
if [ -d "$MODULE_DIR" ]; then
|
||||
echo "=== Terragrunt plan: ${module} ==="
|
||||
cd "/workspace/src/${MODULE_DIR}"
|
||||
terragrunt init --non-interactive -reconfigure
|
||||
OUTPUT=$(terragrunt plan --non-interactive -detailed-exitcode 2>&1) || EXITCODE=$?
|
||||
EXITCODE=${EXITCODE:-0}
|
||||
|
||||
if [ "$EXITCODE" -eq 2 ]; then
|
||||
echo "=== Terragrunt plan output: ${module} ==="
|
||||
echo "$OUTPUT"
|
||||
echo "=== Terragrunt apply: ${module} ==="
|
||||
terragrunt apply --non-interactive -auto-approve
|
||||
elif [ "$EXITCODE" -eq 0 ]; then
|
||||
echo "=== No changes for ${module} ==="
|
||||
else
|
||||
echo "=== ERROR planning ${module} ==="
|
||||
echo "$OUTPUT"
|
||||
exit 1
|
||||
fi
|
||||
cd /workspace/src
|
||||
fi
|
||||
done
|
||||
|
||||
echo "=== Terragrunt apply complete ==="
|
||||
|
||||
# ── apply-argo-templates: self-apply WorkflowTemplates, EventSources, Sensors ──
|
||||
- name: apply-argo-templates
|
||||
inputs:
|
||||
parameters:
|
||||
- name: commit-sha
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: platform
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
serviceAccountName: argo-workflow
|
||||
container:
|
||||
image: gitlab-registry.foxhunt.svc.cluster.local:5000/root/foxhunt/ci-builder-cpu:latest
|
||||
command: ["/bin/sh", "-c"]
|
||||
resources:
|
||||
requests:
|
||||
cpu: 100m
|
||||
memory: 128Mi
|
||||
limits:
|
||||
cpu: 500m
|
||||
memory: 256Mi
|
||||
volumeMounts:
|
||||
- name: git-ssh-key
|
||||
mountPath: /etc/git-ssh
|
||||
readOnly: true
|
||||
args:
|
||||
- |
|
||||
set -e
|
||||
|
||||
# Clone repo at commit
|
||||
mkdir -p /root/.ssh
|
||||
cp /etc/git-ssh/ssh-privatekey /root/.ssh/id_ed25519
|
||||
chmod 600 /root/.ssh/id_ed25519
|
||||
printf 'Host *\n StrictHostKeyChecking no\n UserKnownHostsFile /dev/null\n' > /root/.ssh/config
|
||||
chmod 600 /root/.ssh/config
|
||||
|
||||
SHA="{{inputs.parameters.commit-sha}}"
|
||||
REPO="ssh://git@gitlab-gitlab-shell.foxhunt.svc.cluster.local:2222/root/foxhunt.git"
|
||||
git clone --no-checkout --filter=blob:none "$REPO" /workspace/src
|
||||
cd /workspace/src
|
||||
git checkout "$SHA"
|
||||
echo "Checked out $(git rev-parse --short HEAD)"
|
||||
|
||||
# Apply all Argo WorkflowTemplates
|
||||
for f in infra/k8s/argo/*-template.yaml; do
|
||||
if [ -f "$f" ]; then
|
||||
echo "=== Applying $(basename $f) ==="
|
||||
kubectl -n foxhunt apply -f "$f"
|
||||
fi
|
||||
done
|
||||
|
||||
# Apply EventSources and Sensors
|
||||
for f in infra/k8s/argo/events/*.yaml; do
|
||||
if [ -f "$f" ]; then
|
||||
echo "=== Applying $(basename $f) ==="
|
||||
kubectl -n foxhunt apply -f "$f"
|
||||
fi
|
||||
done
|
||||
|
||||
# Apply RBAC and NetworkPolicies
|
||||
kubectl -n foxhunt apply -f infra/k8s/argo/ci-deploy-rbac.yaml
|
||||
kubectl -n foxhunt apply -f infra/k8s/argo/argo-workflow-netpol.yaml
|
||||
|
||||
echo "=== Argo templates applied ==="
|
||||
|
||||
# ── submit-gpu-test: launch GPU test workflow as a child workflow ──
|
||||
- name: submit-gpu-test
|
||||
# Argo executor (main + wait) needs 256Mi to track large child workflow status JSON.
|
||||
# Default 64Mi causes OOMKilled when child workflow has many parameters.
|
||||
podSpecPatch: '{"containers":[{"name":"main","resources":{"requests":{"memory":"128Mi"},"limits":{"memory":"256Mi"}}},{"name":"wait","resources":{"requests":{"memory":"128Mi"},"limits":{"memory":"256Mi"}}}]}'
|
||||
resource:
|
||||
action: create
|
||||
setOwnerReference: true
|
||||
successCondition: status.phase == Succeeded
|
||||
failureCondition: status.phase in (Failed, Error)
|
||||
manifest: |
|
||||
apiVersion: argoproj.io/v1alpha1
|
||||
kind: Workflow
|
||||
metadata:
|
||||
generateName: gpu-test-on-push-
|
||||
namespace: foxhunt
|
||||
spec:
|
||||
workflowTemplateRef:
|
||||
name: gpu-test-pipeline
|
||||
arguments:
|
||||
parameters:
|
||||
- name: commit-ref
|
||||
value: "{{workflow.parameters.commit-sha}}"
|
||||
- name: models
|
||||
value: "dqn,ppo,tft"
|
||||
|
||||
# ── notify-result: post workflow outcome to Mattermost ──
|
||||
- name: notify-result
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: platform
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
container:
|
||||
image: gitlab-registry.foxhunt.svc.cluster.local:5000/root/foxhunt/foxhunt-runtime:latest
|
||||
command: ["/bin/sh", "-c"]
|
||||
env:
|
||||
- name: WEBHOOK_URL
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: notification-webhook
|
||||
key: webhook-url
|
||||
optional: true
|
||||
resources:
|
||||
requests:
|
||||
cpu: 50m
|
||||
memory: 32Mi
|
||||
limits:
|
||||
cpu: 100m
|
||||
memory: 64Mi
|
||||
args:
|
||||
- |
|
||||
STATUS="{{workflow.status}}"
|
||||
NAME="{{workflow.name}}"
|
||||
|
||||
if [ -z "$WEBHOOK_URL" ] || echo "$WEBHOOK_URL" | grep -q "PLACEHOLDER"; then
|
||||
echo "No webhook configured, skipping notification"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
if [ "$STATUS" = "Succeeded" ]; then
|
||||
EMOJI=":white_check_mark:"
|
||||
else
|
||||
EMOJI=":x:"
|
||||
fi
|
||||
|
||||
PAYLOAD="{\"username\":\"Argo CI\",\"text\":\"${EMOJI} **${NAME}** — ${STATUS} ({{workflow.duration}}s)\"}"
|
||||
|
||||
curl -sf -X POST -H 'Content-Type: application/json' \
|
||||
-d "$PAYLOAD" "$WEBHOOK_URL" || echo "WARN: webhook post failed"
|
||||
@@ -1,614 +0,0 @@
|
||||
# Compile-and-Deploy: manual workflow for building and deploying service binaries.
|
||||
#
|
||||
# DAG:
|
||||
# create-tag ──> compile-services ──> upload-release ──> deploy-services
|
||||
#
|
||||
# Image builds handled by ci-pipeline, not here.
|
||||
# Training compilation lives in compile-and-train-template.yaml.
|
||||
---
|
||||
apiVersion: argoproj.io/v1alpha1
|
||||
kind: WorkflowTemplate
|
||||
metadata:
|
||||
name: compile-and-deploy
|
||||
namespace: foxhunt
|
||||
labels:
|
||||
app.kubernetes.io/name: compile-and-deploy
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
spec:
|
||||
entrypoint: pipeline
|
||||
onExit: notify-result
|
||||
serviceAccountName: argo-workflow
|
||||
podMetadata:
|
||||
labels:
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
app.kubernetes.io/component: compile-and-deploy
|
||||
archiveLogs: true
|
||||
podGC:
|
||||
strategy: OnPodCompletion
|
||||
ttlStrategy:
|
||||
secondsAfterCompletion: 3600
|
||||
activeDeadlineSeconds: 3600 # 1 hour (compile + deploy)
|
||||
|
||||
arguments:
|
||||
parameters:
|
||||
- name: commit-sha
|
||||
value: HEAD
|
||||
- name: service-packages
|
||||
value: "api trading-service backtesting-service trading-agent-service broker-gateway data-acquisition-service"
|
||||
|
||||
volumes:
|
||||
- name: git-ssh-key
|
||||
secret:
|
||||
secretName: argo-git-ssh-key
|
||||
defaultMode: 256
|
||||
- name: cargo-target-cpu
|
||||
persistentVolumeClaim:
|
||||
claimName: cargo-target-cpu
|
||||
- name: sccache
|
||||
persistentVolumeClaim:
|
||||
claimName: sccache-cpu
|
||||
|
||||
templates:
|
||||
- name: pipeline
|
||||
dag:
|
||||
tasks:
|
||||
- name: create-tag
|
||||
template: create-tag
|
||||
- name: compile-services
|
||||
depends: "create-tag"
|
||||
template: compile-services
|
||||
arguments:
|
||||
parameters:
|
||||
- name: tag
|
||||
value: "{{tasks.create-tag.outputs.parameters.tag}}"
|
||||
- name: service-packages
|
||||
value: "{{workflow.parameters.service-packages}}"
|
||||
- name: upload-release
|
||||
depends: "compile-services"
|
||||
template: upload-release
|
||||
arguments:
|
||||
parameters:
|
||||
- name: tag
|
||||
value: "{{tasks.create-tag.outputs.parameters.tag}}"
|
||||
- name: deploy-services
|
||||
depends: "upload-release"
|
||||
template: deploy-services
|
||||
arguments:
|
||||
parameters:
|
||||
- name: tag
|
||||
value: "{{tasks.create-tag.outputs.parameters.tag}}"
|
||||
- name: deploy-list
|
||||
value: "{{workflow.parameters.service-packages}}"
|
||||
|
||||
# ── create-tag: CalVer auto-tag on code changes ──
|
||||
- name: create-tag
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: platform
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
container:
|
||||
image: curlimages/curl:8.12.1
|
||||
command: ["/bin/sh", "-c"]
|
||||
env:
|
||||
- name: GITLAB_PAT
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: gitlab-pat
|
||||
key: token
|
||||
resources:
|
||||
requests:
|
||||
cpu: 100m
|
||||
memory: 64Mi
|
||||
limits:
|
||||
cpu: 200m
|
||||
memory: 128Mi
|
||||
args:
|
||||
- |
|
||||
set -e
|
||||
GITLAB="http://gitlab-webservice-default.foxhunt.svc.cluster.local:8181"
|
||||
PROJECT_ID=1
|
||||
SHA="{{workflow.parameters.commit-sha}}"
|
||||
|
||||
# Compute CalVer prefix: vYYYY.MM
|
||||
PREFIX="v$(date +%Y.%m)"
|
||||
|
||||
# Query existing tags for this month
|
||||
TAGS=$(curl -sf \
|
||||
-H "PRIVATE-TOKEN: ${GITLAB_PAT}" \
|
||||
"${GITLAB}/api/v4/projects/${PROJECT_ID}/repository/tags?search=${PREFIX}" \
|
||||
|| echo "[]")
|
||||
|
||||
# Parse highest N from vYYYY.MM.N tags
|
||||
LAST_N=$(echo "$TAGS" | grep -oP "\"name\":\"${PREFIX}\.\K[0-9]+" | sort -n | tail -1)
|
||||
if [ -z "$LAST_N" ]; then
|
||||
NEXT_N=1
|
||||
else
|
||||
NEXT_N=$((LAST_N + 1))
|
||||
fi
|
||||
|
||||
TAG="${PREFIX}.${NEXT_N}"
|
||||
echo "Creating tag: ${TAG} at ${SHA}"
|
||||
|
||||
# Create the tag (tolerate failure if tag already exists)
|
||||
HTTP_CODE=$(curl -s -o /tmp/tag_response -w "%{http_code}" -X POST \
|
||||
-H "PRIVATE-TOKEN: ${GITLAB_PAT}" \
|
||||
-d "tag_name=${TAG}" \
|
||||
-d "ref=${SHA}" \
|
||||
"${GITLAB}/api/v4/projects/${PROJECT_ID}/repository/tags")
|
||||
|
||||
RESULT=$(cat /tmp/tag_response 2>/dev/null || echo "{}")
|
||||
|
||||
if [ "$HTTP_CODE" = "201" ]; then
|
||||
echo "Tag ${TAG} created successfully"
|
||||
elif [ "$HTTP_CODE" = "400" ] && echo "$RESULT" | grep -q "already exists"; then
|
||||
echo "Tag ${TAG} already exists — reusing"
|
||||
else
|
||||
echo "WARN: Tag creation returned HTTP ${HTTP_CODE}: ${RESULT}"
|
||||
echo "Falling back to dev tag"
|
||||
TAG="dev-$(echo $SHA | cut -c1-8)"
|
||||
fi
|
||||
|
||||
mkdir -p /tmp/outputs
|
||||
echo -n "$TAG" > /tmp/outputs/tag
|
||||
outputs:
|
||||
parameters:
|
||||
- name: tag
|
||||
valueFrom:
|
||||
path: /tmp/outputs/tag
|
||||
|
||||
# ── compile-services: selective per-binary build, incremental on local RWO PVC ──
|
||||
- name: compile-services
|
||||
metadata:
|
||||
labels:
|
||||
app.kubernetes.io/component: compile
|
||||
inputs:
|
||||
parameters:
|
||||
- name: tag
|
||||
- name: service-packages
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: ci-compile-cpu
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
tolerations:
|
||||
- key: node.cilium.io/agent-not-ready
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
# ── seed-deps-cache: rsync prebuilt third-party rlibs into the PVC ──
|
||||
# On a true cold PVC (e.g. fresh node, first run, or after a purge),
|
||||
# this saves 2-4 min that would otherwise be spent compiling the
|
||||
# ~1500 third-party crates in the workspace dep graph. The image is
|
||||
# rebuilt nightly by refresh-deps-cache (see
|
||||
# refresh-deps-cache-template.yaml). On a warm PVC, the stamp file
|
||||
# short-circuits the whole step in ~50ms.
|
||||
#
|
||||
# Why rsync --ignore-existing instead of cp/overwrite: a warmer PVC
|
||||
# may already have NEWER artifacts from a prior compile; we don't
|
||||
# want to clobber them with potentially-stale prebuilt rlibs. Cargo
|
||||
# will discard rlibs whose fingerprint mismatches anyway.
|
||||
#
|
||||
# If the deps-cache image isn't published yet (first deploy of this
|
||||
# template), set imagePullPolicy: IfNotPresent + the initContainer
|
||||
# is non-blocking-on-failure (`|| true`) so the compile-services
|
||||
# step still runs without the cache prelude.
|
||||
initContainers:
|
||||
- name: seed-deps-cache
|
||||
image: gitlab-registry.foxhunt.svc.cluster.local:5000/root/foxhunt/ci-builder-cpu-with-deps:nightly
|
||||
imagePullPolicy: IfNotPresent
|
||||
env:
|
||||
- name: DEPS_VERSION
|
||||
value: "1"
|
||||
resources:
|
||||
requests:
|
||||
cpu: 500m
|
||||
memory: 512Mi
|
||||
limits:
|
||||
cpu: "2"
|
||||
memory: 1Gi
|
||||
volumeMounts:
|
||||
- name: cargo-target-cpu
|
||||
mountPath: /cargo-target
|
||||
command: ["/bin/sh", "-c"]
|
||||
args:
|
||||
- |
|
||||
set -e
|
||||
STAMP="/cargo-target/cpu_deps_v${DEPS_VERSION}.stamp"
|
||||
if [ -f "$STAMP" ]; then
|
||||
echo "PVC already seeded for deps v${DEPS_VERSION} (stamp present), skipping."
|
||||
exit 0
|
||||
fi
|
||||
if [ ! -d /cargo-target-prebuilt ]; then
|
||||
echo "WARN: /cargo-target-prebuilt missing in image — image not yet published?"
|
||||
echo " Compile will run without prebuilt deps cache (slower cold-start)."
|
||||
exit 0
|
||||
fi
|
||||
echo "=== Seeding cargo-target-cpu PVC from prebuilt deps cache (v${DEPS_VERSION}) ==="
|
||||
du -sh /cargo-target-prebuilt 2>/dev/null || true
|
||||
# --ignore-existing: never clobber newer artifacts already on PVC.
|
||||
# -a: preserve mtimes, perms, links — cargo's freshness check needs accurate mtimes.
|
||||
rsync -a --ignore-existing /cargo-target-prebuilt/ /cargo-target/
|
||||
touch "$STAMP"
|
||||
echo "=== Seed complete ==="
|
||||
du -sh /cargo-target/release 2>/dev/null || true
|
||||
container:
|
||||
image: gitlab-registry.foxhunt.svc.cluster.local:5000/root/foxhunt/ci-builder-cpu:latest
|
||||
imagePullPolicy: Always
|
||||
command: ["/bin/sh", "-c"]
|
||||
env:
|
||||
- name: SQLX_OFFLINE
|
||||
value: "true"
|
||||
- name: CARGO_TERM_COLOR
|
||||
value: always
|
||||
- name: CARGO_TARGET_DIR
|
||||
value: /cargo-target
|
||||
- name: CARGO_HOME
|
||||
value: /cargo-target/cargo-home
|
||||
# sccache: rustc wrapper that content-hashes compile inputs and reuses
|
||||
# object output across pods/commits via the sccache-cpu PVC. Survives
|
||||
# cargo's incremental-cache invalidation from git checkout mtime touches.
|
||||
- name: RUSTC_WRAPPER
|
||||
value: sccache
|
||||
- name: SCCACHE_DIR
|
||||
value: /sccache-cache
|
||||
- name: SCCACHE_CACHE_SIZE
|
||||
value: 20G
|
||||
# Required: workspace `.cargo/config.toml` sets [build] incremental = true,
|
||||
# which forces incremental compilation even on --release. sccache cannot
|
||||
# cache incremental rustc output (documented limitation). Override here
|
||||
# so sccache actually catches. Local dev keeps incremental via config.toml.
|
||||
- name: CARGO_INCREMENTAL
|
||||
value: "0"
|
||||
# Match the cgroup cpu limit below (limits.cpu: "30").
|
||||
# Without this, cargo asks num_cpus::get_physical() — the host's
|
||||
# whole CPU count — and over-subscribes vs the cgroup throttle,
|
||||
# producing scheduling waste under load.
|
||||
- name: CARGO_BUILD_JOBS
|
||||
value: "30"
|
||||
- name: GITLAB_PAT
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: gitlab-pat
|
||||
key: token
|
||||
resources:
|
||||
requests:
|
||||
cpu: "14"
|
||||
memory: 16Gi
|
||||
limits:
|
||||
cpu: "30"
|
||||
memory: 32Gi
|
||||
volumeMounts:
|
||||
- name: git-ssh-key
|
||||
mountPath: /etc/git-ssh
|
||||
readOnly: true
|
||||
- name: cargo-target-cpu
|
||||
mountPath: /cargo-target
|
||||
- name: sccache
|
||||
mountPath: /sccache-cache
|
||||
args:
|
||||
- |
|
||||
set -e
|
||||
SHA="{{workflow.parameters.commit-sha}}"
|
||||
SERVICE_PKGS="{{inputs.parameters.service-packages}}"
|
||||
|
||||
# SSH setup
|
||||
mkdir -p ~/.ssh
|
||||
cp /etc/git-ssh/ssh-privatekey ~/.ssh/id_ed25519
|
||||
chmod 600 ~/.ssh/id_ed25519
|
||||
printf 'Host *\n StrictHostKeyChecking no\n UserKnownHostsFile /dev/null\n' > ~/.ssh/config
|
||||
chmod 600 ~/.ssh/config
|
||||
|
||||
REPO="ssh://git@gitlab-gitlab-shell.foxhunt.svc.cluster.local:2222/root/foxhunt.git"
|
||||
WORKSPACE="/cargo-target/src"
|
||||
|
||||
# PVC may be owned by a different UID from a previous run
|
||||
git config --global --add safe.directory "$WORKSPACE"
|
||||
|
||||
# Persistent checkout on PVC — only changed files get new mtimes,
|
||||
# so cargo skips recompiling unchanged workspace crates.
|
||||
if [ -d "$WORKSPACE/.git" ]; then
|
||||
cd "$WORKSPACE"
|
||||
CURRENT=$(git rev-parse HEAD 2>/dev/null || echo "none")
|
||||
if [ "$CURRENT" = "$SHA" ]; then
|
||||
echo "=== Already at $SHA, skipping checkout ==="
|
||||
else
|
||||
echo "=== Updating checkout: $(echo $CURRENT | cut -c1-8) -> $(echo $SHA | cut -c1-8) ==="
|
||||
git fetch origin
|
||||
git checkout --force "$SHA"
|
||||
git clean -fd
|
||||
fi
|
||||
else
|
||||
echo "=== Initial clone (first run) ==="
|
||||
git clone --filter=blob:none "$REPO" "$WORKSPACE"
|
||||
cd "$WORKSPACE"
|
||||
git checkout "$SHA"
|
||||
fi
|
||||
echo "Checked out $(git rev-parse --short HEAD)"
|
||||
|
||||
# Ensure cargo home registry is on PVC
|
||||
export PATH="${CARGO_HOME}/bin:${PATH}"
|
||||
|
||||
export FOXHUNT_BUILD_VERSION="{{inputs.parameters.tag}}"
|
||||
|
||||
# Prune build artifacts if PVC exceeds 25GB (prevents unbounded growth)
|
||||
TARGET_SIZE_MB=$(du -sm "$CARGO_TARGET_DIR" 2>/dev/null | cut -f1 || echo 0)
|
||||
echo "PVC usage: ${TARGET_SIZE_MB}MB"
|
||||
if [ "$TARGET_SIZE_MB" -gt 25000 ]; then
|
||||
echo "PVC exceeds 25GB limit, pruning build artifacts..."
|
||||
rm -rf "$CARGO_TARGET_DIR/release" "$CARGO_TARGET_DIR/debug"
|
||||
fi
|
||||
|
||||
# Guard: empty package list would build entire workspace
|
||||
if [ -z "$SERVICE_PKGS" ]; then
|
||||
echo "ERROR: service-packages is empty, refusing to build entire workspace"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
# Build only the affected service packages (incremental via persistent target dir)
|
||||
CARGO_ARGS=""
|
||||
for pkg in $SERVICE_PKGS; do
|
||||
CARGO_ARGS="$CARGO_ARGS -p $pkg"
|
||||
done
|
||||
|
||||
# CI-only linker swap: prefer wild over mold. wild is a Rust-native
|
||||
# linker, typically 10-30% faster than mold on release-LTO links;
|
||||
# we have ~5 service binaries each doing release-LTO, so this is
|
||||
# a meaningful 30-90 sec wall-time saving. Both linkers ship in
|
||||
# Dockerfile.ci-builder-cpu — to revert, drop this sed and rebuild.
|
||||
# Sed is idempotent (no-op if wild already substituted from a prior
|
||||
# run on the same PVC checkout) and surgical (single line in
|
||||
# .cargo/config.toml). If wild is missing for any reason, the build
|
||||
# falls back to mold once we revert this hunk.
|
||||
if command -v wild >/dev/null 2>&1; then
|
||||
echo "=== Swapping linker mold -> wild for this CI run ==="
|
||||
sed -i 's|-fuse-ld=mold|-fuse-ld=wild|' .cargo/config.toml
|
||||
grep -n 'fuse-ld' .cargo/config.toml || true
|
||||
else
|
||||
echo "WARN: wild not on PATH, sticking with mold"
|
||||
fi
|
||||
|
||||
echo "=== Building service binaries: $SERVICE_PKGS (incremental) ==="
|
||||
# --locked: skip Cargo.lock resolver work (~5-15s saved) and fail fast
|
||||
# if the lock has drifted (catches accidental Cargo.toml edits without
|
||||
# a re-resolve commit).
|
||||
cargo build --locked --release $CARGO_ARGS
|
||||
|
||||
# Collect built binaries
|
||||
mkdir -p "$WORKSPACE/bin/services"
|
||||
for pkg in $SERVICE_PKGS; do
|
||||
bin_name=$(echo "$pkg" | tr '-' '_')
|
||||
cp "$CARGO_TARGET_DIR/release/$pkg" "$WORKSPACE/bin/services/" 2>/dev/null \
|
||||
|| cp "$CARGO_TARGET_DIR/release/$bin_name" "$WORKSPACE/bin/services/" 2>/dev/null \
|
||||
|| { echo "Binary not found for $pkg"; ls "$CARGO_TARGET_DIR/release/"; exit 1; }
|
||||
done
|
||||
strip "$WORKSPACE/bin/services/"*
|
||||
|
||||
echo "=== Service binaries ==="
|
||||
ls -lh "$WORKSPACE/bin/services/"
|
||||
|
||||
echo "=== Uploading service binaries to GitLab packages ==="
|
||||
GITLAB="http://gitlab-webservice-default.foxhunt.svc.cluster.local:8181"
|
||||
TAG="${FOXHUNT_BUILD_VERSION}"
|
||||
for bin in "$WORKSPACE/bin/services/"*; do
|
||||
BIN_NAME=$(basename "$bin")
|
||||
echo "Uploading ${BIN_NAME} (${TAG})..."
|
||||
curl -f --upload-file "$bin" \
|
||||
-H "PRIVATE-TOKEN: ${GITLAB_PAT}" \
|
||||
"${GITLAB}/api/v4/projects/1/packages/generic/foxhunt-services/${TAG}/${BIN_NAME}"
|
||||
done
|
||||
|
||||
# Update 'latest' per-file (preserves unbuilt binaries from prior runs)
|
||||
echo "=== Updating 'latest' package ==="
|
||||
LATEST_PKG=$(curl -sf -H "PRIVATE-TOKEN: ${GITLAB_PAT}" \
|
||||
"${GITLAB}/api/v4/projects/1/packages?package_name=foxhunt-services&package_version=latest" \
|
||||
| grep -oP '"id":\K[0-9]+' | head -1)
|
||||
for bin in "$WORKSPACE/bin/services/"*; do
|
||||
BIN_NAME=$(basename "$bin")
|
||||
# Delete existing file by name before uploading replacement
|
||||
if [ -n "$LATEST_PKG" ]; then
|
||||
FILE_ID=$(curl -sf -H "PRIVATE-TOKEN: ${GITLAB_PAT}" \
|
||||
"${GITLAB}/api/v4/projects/1/packages/${LATEST_PKG}/package_files" \
|
||||
| grep -oP "\"id\":([0-9]+),\"package_id\":${LATEST_PKG}[^}]*\"file_name\":\"${BIN_NAME}\"" \
|
||||
| grep -oP '"id":\K[0-9]+' | head -1)
|
||||
if [ -n "$FILE_ID" ]; then
|
||||
curl -sf -X DELETE -H "PRIVATE-TOKEN: ${GITLAB_PAT}" \
|
||||
"${GITLAB}/api/v4/projects/1/packages/${LATEST_PKG}/package_files/${FILE_ID}" || true
|
||||
fi
|
||||
fi
|
||||
curl -f --upload-file "$bin" \
|
||||
-H "PRIVATE-TOKEN: ${GITLAB_PAT}" \
|
||||
"${GITLAB}/api/v4/projects/1/packages/generic/foxhunt-services/latest/${BIN_NAME}" || true
|
||||
done
|
||||
|
||||
echo "=== sccache stats ==="
|
||||
sccache --show-stats || true
|
||||
|
||||
echo "=== Service compile + upload done ($SERVICE_PKGS) ==="
|
||||
|
||||
# ── upload-release: create GitLab Release with package links ──
|
||||
- name: upload-release
|
||||
inputs:
|
||||
parameters:
|
||||
- name: tag
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: platform
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
container:
|
||||
image: curlimages/curl:8.12.1
|
||||
command: ["/bin/sh", "-c"]
|
||||
env:
|
||||
- name: GITLAB_PAT
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: gitlab-pat
|
||||
key: token
|
||||
resources:
|
||||
requests:
|
||||
cpu: 100m
|
||||
memory: 64Mi
|
||||
limits:
|
||||
cpu: 200m
|
||||
memory: 128Mi
|
||||
args:
|
||||
- |
|
||||
set -e
|
||||
GITLAB="http://gitlab-webservice-default.foxhunt.svc.cluster.local:8181"
|
||||
PROJECT_ID=1
|
||||
TAG="{{inputs.parameters.tag}}"
|
||||
|
||||
echo "=== Creating GitLab Release ${TAG} ==="
|
||||
|
||||
# Get commits since previous tag for release notes
|
||||
PREV_TAG=$(curl -sf \
|
||||
-H "PRIVATE-TOKEN: ${GITLAB_PAT}" \
|
||||
"${GITLAB}/api/v4/projects/${PROJECT_ID}/repository/tags?per_page=2" \
|
||||
| grep -oP '"name":"\K[^"]+' | sed -n '2p')
|
||||
|
||||
if [ -n "$PREV_TAG" ]; then
|
||||
COMMITS=$(curl -sf \
|
||||
-H "PRIVATE-TOKEN: ${GITLAB_PAT}" \
|
||||
"${GITLAB}/api/v4/projects/${PROJECT_ID}/repository/compare?from=${PREV_TAG}&to=${TAG}" \
|
||||
| grep -oP '"title":"\K[^"]+' | head -20 \
|
||||
| sed 's/^/- /' || echo "- Release ${TAG}")
|
||||
DESCRIPTION="## Changes since ${PREV_TAG}\n\n${COMMITS}"
|
||||
else
|
||||
DESCRIPTION="## Initial release\n\nFirst CalVer release."
|
||||
fi
|
||||
|
||||
# Create the release
|
||||
RESPONSE=$(curl -sf -X POST \
|
||||
-H "PRIVATE-TOKEN: ${GITLAB_PAT}" \
|
||||
-H "Content-Type: application/json" \
|
||||
-d "{
|
||||
\"tag_name\": \"${TAG}\",
|
||||
\"name\": \"${TAG}\",
|
||||
\"description\": \"$(printf '%s' "$DESCRIPTION" | sed 's/"/\\"/g')\"
|
||||
}" \
|
||||
"${GITLAB}/api/v4/projects/${PROJECT_ID}/releases") || {
|
||||
echo "WARN: Release creation failed (may already exist)"
|
||||
}
|
||||
|
||||
echo "Release ${TAG} created"
|
||||
echo "$RESPONSE" | head -5
|
||||
|
||||
# Add package links as release assets
|
||||
for pkg_name in foxhunt-services foxhunt-training; do
|
||||
PKG_CHECK=$(curl -sf \
|
||||
-H "PRIVATE-TOKEN: ${GITLAB_PAT}" \
|
||||
"${GITLAB}/api/v4/projects/${PROJECT_ID}/packages?package_name=${pkg_name}&package_version=${TAG}" \
|
||||
|| echo "[]")
|
||||
|
||||
if echo "$PKG_CHECK" | grep -q "$TAG"; then
|
||||
curl -sf -X POST \
|
||||
-H "PRIVATE-TOKEN: ${GITLAB_PAT}" \
|
||||
-H "Content-Type: application/json" \
|
||||
-d "{
|
||||
\"name\": \"${pkg_name}\",
|
||||
\"url\": \"${GITLAB}/-/packages?type=generic&search=${pkg_name}&version=${TAG}\",
|
||||
\"link_type\": \"package\"
|
||||
}" \
|
||||
"${GITLAB}/api/v4/projects/${PROJECT_ID}/releases/${TAG}/assets/links" || true
|
||||
echo "Linked ${pkg_name} package to release"
|
||||
fi
|
||||
done
|
||||
|
||||
echo "=== Release ${TAG} complete ==="
|
||||
|
||||
# ── deploy-services: selective rolling restart for affected deployments ──
|
||||
- name: deploy-services
|
||||
inputs:
|
||||
parameters:
|
||||
- name: tag
|
||||
- name: deploy-list
|
||||
serviceAccountName: argo-workflow
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: platform
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
container:
|
||||
image: curlimages/curl:8.12.1
|
||||
command: ["/bin/sh", "-c"]
|
||||
resources:
|
||||
requests:
|
||||
cpu: 100m
|
||||
memory: 64Mi
|
||||
limits:
|
||||
cpu: 200m
|
||||
memory: 128Mi
|
||||
args:
|
||||
- |
|
||||
set -e
|
||||
if ! command -v kubectl >/dev/null 2>&1; then
|
||||
echo "=== Installing kubectl ==="
|
||||
curl -sLo /tmp/kubectl "https://dl.k8s.io/release/v1.31.4/bin/linux/amd64/kubectl"
|
||||
chmod +x /tmp/kubectl
|
||||
export PATH="/tmp:$PATH"
|
||||
fi
|
||||
|
||||
TAG="{{inputs.parameters.tag}}"
|
||||
DEPLOY_LIST="{{inputs.parameters.deploy-list}}"
|
||||
|
||||
if [ -z "$DEPLOY_LIST" ]; then
|
||||
echo "=== No services to deploy ==="
|
||||
exit 0
|
||||
fi
|
||||
|
||||
echo "=== Deploying release ${TAG}: $DEPLOY_LIST ==="
|
||||
|
||||
for svc in $DEPLOY_LIST; do
|
||||
echo "Patching $svc with FOXHUNT_RELEASE=${TAG}..."
|
||||
kubectl -n foxhunt patch deployment "$svc" -p "{
|
||||
\"spec\":{\"template\":{
|
||||
\"metadata\":{\"annotations\":{\"foxhunt.io/release\":\"${TAG}\"}},
|
||||
\"spec\":{\"initContainers\":[{
|
||||
\"name\":\"fetch-binary\",
|
||||
\"env\":[{\"name\":\"FOXHUNT_RELEASE\",\"value\":\"${TAG}\"}]
|
||||
}]}
|
||||
}}
|
||||
}" || echo "WARN: $svc patch failed (may not exist)"
|
||||
done
|
||||
|
||||
echo "=== Waiting for rollouts ==="
|
||||
for svc in $DEPLOY_LIST; do
|
||||
kubectl -n foxhunt rollout status deployment "$svc" --timeout=120s || echo "WARN: $svc rollout timeout"
|
||||
done
|
||||
|
||||
echo "=== Deploy ${TAG} complete ($DEPLOY_LIST) ==="
|
||||
|
||||
# ── notify-result: post workflow outcome to Mattermost ──
|
||||
- name: notify-result
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: platform
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
container:
|
||||
image: curlimages/curl:8.12.1
|
||||
command: ["/bin/sh", "-c"]
|
||||
env:
|
||||
- name: WEBHOOK_URL
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: notification-webhook
|
||||
key: webhook-url
|
||||
optional: true
|
||||
resources:
|
||||
requests:
|
||||
cpu: 50m
|
||||
memory: 32Mi
|
||||
limits:
|
||||
cpu: 100m
|
||||
memory: 64Mi
|
||||
args:
|
||||
- |
|
||||
STATUS="{{workflow.status}}"
|
||||
NAME="{{workflow.name}}"
|
||||
|
||||
if [ -z "$WEBHOOK_URL" ] || echo "$WEBHOOK_URL" | grep -q "PLACEHOLDER"; then
|
||||
echo "No webhook configured, skipping notification"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
if [ "$STATUS" = "Succeeded" ]; then
|
||||
EMOJI=":white_check_mark:"
|
||||
else
|
||||
EMOJI=":x:"
|
||||
fi
|
||||
|
||||
PAYLOAD="{\"username\":\"Argo CI\",\"text\":\"${EMOJI} **${NAME}** — ${STATUS} ({{workflow.duration}}s)\"}"
|
||||
|
||||
curl -sf -X POST -H 'Content-Type: application/json' \
|
||||
-d "$PAYLOAD" "$WEBHOOK_URL" || echo "WARN: webhook post failed"
|
||||
@@ -1,12 +0,0 @@
|
||||
apiVersion: v1
|
||||
kind: ConfigMap
|
||||
metadata:
|
||||
name: auto-compile-config
|
||||
namespace: foxhunt
|
||||
labels:
|
||||
app.kubernetes.io/name: auto-compile-config
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
data:
|
||||
# Set to "true" to enable auto-compile on push to main.
|
||||
# kubectl -n foxhunt patch configmap auto-compile-config -p '{"data":{"enabled":"true"}}'
|
||||
enabled: "false"
|
||||
@@ -1,70 +0,0 @@
|
||||
apiVersion: argoproj.io/v1alpha1
|
||||
kind: Sensor
|
||||
metadata:
|
||||
name: ci-pipeline-trigger
|
||||
namespace: foxhunt
|
||||
labels:
|
||||
app.kubernetes.io/name: ci-pipeline-trigger
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
spec:
|
||||
template:
|
||||
serviceAccountName: argo-workflow
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: platform
|
||||
eventBusName: default
|
||||
dependencies:
|
||||
- name: gitlab-push-dep
|
||||
eventSourceName: gitlab-push
|
||||
eventName: gitlab-push
|
||||
filters:
|
||||
data:
|
||||
- path: body.ref
|
||||
type: string
|
||||
value:
|
||||
- "refs/heads/main"
|
||||
# Policy (user directive): push-triggered runs are infra-sync only.
|
||||
# The ci-pipeline DAG runs three conditional task groups on push:
|
||||
# - docker image rebuilds (gated on infra/docker/ changes)
|
||||
# - apply-argo-templates (gated on infra/k8s/argo/ changes)
|
||||
# - terragrunt-apply (gated on infra/live/ or infra/modules/ changes)
|
||||
# Non-infra tasks (test-gate, build-web-dashboard, gpu-test) have `when: "false"`
|
||||
# in the WorkflowTemplate. Training + compile/deploy are separate templates,
|
||||
# always manual.
|
||||
#
|
||||
# All non-push workflows (cargo tests, training, compile/deploy, GPU tests)
|
||||
# run via scripts/manual `argo submit`.
|
||||
|
||||
triggers:
|
||||
- template:
|
||||
name: ci-pipeline
|
||||
argoWorkflow:
|
||||
operation: submit
|
||||
source:
|
||||
resource:
|
||||
apiVersion: argoproj.io/v1alpha1
|
||||
kind: Workflow
|
||||
metadata:
|
||||
generateName: ci-pipeline-
|
||||
namespace: foxhunt
|
||||
spec:
|
||||
serviceAccountName: argo-workflow
|
||||
podGC:
|
||||
strategy: OnPodCompletion
|
||||
ttlStrategy:
|
||||
secondsAfterCompletion: 3600
|
||||
workflowTemplateRef:
|
||||
name: ci-pipeline
|
||||
arguments:
|
||||
parameters:
|
||||
- name: commit-sha
|
||||
- name: commits-json
|
||||
parameters:
|
||||
- src:
|
||||
dependencyName: gitlab-push-dep
|
||||
dataKey: body.checkout_sha
|
||||
dest: spec.arguments.parameters.0.value
|
||||
- src:
|
||||
dependencyName: gitlab-push-dep
|
||||
dataKey: body.commits
|
||||
dataTemplate: "{{ toJson .Input }}"
|
||||
dest: spec.arguments.parameters.1.value
|
||||
@@ -1,15 +0,0 @@
|
||||
apiVersion: argoproj.io/v1alpha1
|
||||
kind: EventSource
|
||||
metadata:
|
||||
name: gitlab-push
|
||||
namespace: foxhunt
|
||||
labels:
|
||||
app.kubernetes.io/name: gitlab-push-eventsource
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
spec:
|
||||
eventBusName: default
|
||||
webhook:
|
||||
gitlab-push:
|
||||
port: "12000"
|
||||
endpoint: /push
|
||||
method: POST
|
||||
@@ -1,93 +0,0 @@
|
||||
apiVersion: networking.k8s.io/v1
|
||||
kind: NetworkPolicy
|
||||
metadata:
|
||||
name: argo-gpu-test-workflow
|
||||
namespace: foxhunt
|
||||
labels:
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
spec:
|
||||
podSelector:
|
||||
matchLabels:
|
||||
app.kubernetes.io/component: gpu-test
|
||||
policyTypes:
|
||||
- Egress
|
||||
egress:
|
||||
# DNS
|
||||
- ports:
|
||||
- port: 53
|
||||
protocol: UDP
|
||||
- port: 53
|
||||
protocol: TCP
|
||||
# Kubernetes API + internal services (service CIDR)
|
||||
- ports:
|
||||
- port: 443
|
||||
protocol: TCP
|
||||
- port: 6443
|
||||
protocol: TCP
|
||||
to:
|
||||
- ipBlock:
|
||||
cidr: 10.32.0.0/16
|
||||
- ipBlock:
|
||||
cidr: 172.16.0.4/32
|
||||
# HTTPS egress (crates.io, etc.)
|
||||
- ports:
|
||||
- port: 443
|
||||
protocol: TCP
|
||||
# Git SSH (GitLab)
|
||||
- ports:
|
||||
- port: 2222
|
||||
protocol: TCP
|
||||
to:
|
||||
- ipBlock:
|
||||
cidr: 100.90.76.85/32
|
||||
- podSelector:
|
||||
matchLabels:
|
||||
app: gitlab-shell
|
||||
# GitLab webservice (git-http)
|
||||
- ports:
|
||||
- port: 8181
|
||||
protocol: TCP
|
||||
to:
|
||||
- podSelector:
|
||||
matchLabels:
|
||||
app: webservice
|
||||
# MinIO (artifact/log storage)
|
||||
- ports:
|
||||
- port: 9000
|
||||
protocol: TCP
|
||||
to:
|
||||
- podSelector:
|
||||
matchLabels:
|
||||
app.kubernetes.io/name: minio
|
||||
# Pushgateway (Prometheus metrics)
|
||||
- ports:
|
||||
- port: 9091
|
||||
protocol: TCP
|
||||
to:
|
||||
- podSelector:
|
||||
matchLabels:
|
||||
app.kubernetes.io/name: pushgateway
|
||||
# OTLP (Tempo)
|
||||
- ports:
|
||||
- port: 4317
|
||||
protocol: TCP
|
||||
to:
|
||||
- podSelector:
|
||||
matchLabels:
|
||||
app.kubernetes.io/name: tempo
|
||||
# Mattermost (notifications)
|
||||
- ports:
|
||||
- port: 8065
|
||||
protocol: TCP
|
||||
to:
|
||||
- podSelector:
|
||||
matchLabels:
|
||||
app.kubernetes.io/name: mattermost
|
||||
# GitLab container registry
|
||||
- ports:
|
||||
- port: 5000
|
||||
protocol: TCP
|
||||
to:
|
||||
- podSelector:
|
||||
matchLabels:
|
||||
app: registry
|
||||
@@ -1,25 +0,0 @@
|
||||
# infra/k8s/argo/gpu-test-nightly-cron.yaml
|
||||
apiVersion: argoproj.io/v1alpha1
|
||||
kind: CronWorkflow
|
||||
metadata:
|
||||
name: gpu-test-nightly
|
||||
namespace: foxhunt
|
||||
labels:
|
||||
app.kubernetes.io/name: gpu-test-nightly
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
spec:
|
||||
schedule: "0 2 * * *"
|
||||
timezone: "UTC"
|
||||
suspend: true # Enable when ready: kubectl patch cronworkflow gpu-test-nightly -n foxhunt -p '{"spec":{"suspend":false}}'
|
||||
concurrencyPolicy: Replace
|
||||
successfulJobsHistoryLimit: 3
|
||||
failedJobsHistoryLimit: 5
|
||||
workflowSpec:
|
||||
workflowTemplateRef:
|
||||
name: gpu-test-pipeline
|
||||
arguments:
|
||||
parameters:
|
||||
- name: models
|
||||
value: "dqn,ppo,tft,mamba2,tggn,tlob,liquid,kan,xlstm,diffusion"
|
||||
- name: commit-ref
|
||||
value: "main"
|
||||
@@ -1,529 +0,0 @@
|
||||
# infra/k8s/argo/gpu-test-pipeline-template.yaml
|
||||
apiVersion: argoproj.io/v1alpha1
|
||||
kind: WorkflowTemplate
|
||||
metadata:
|
||||
name: gpu-test-pipeline
|
||||
namespace: foxhunt
|
||||
labels:
|
||||
app.kubernetes.io/name: gpu-test-pipeline
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
spec:
|
||||
entrypoint: pipeline
|
||||
onExit: notify-result
|
||||
serviceAccountName: argo-workflow
|
||||
podMetadata:
|
||||
labels:
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
app.kubernetes.io/component: gpu-test
|
||||
securityContext:
|
||||
fsGroup: 0
|
||||
archiveLogs: true
|
||||
podGC:
|
||||
strategy: OnPodCompletion
|
||||
ttlStrategy:
|
||||
secondsAfterCompletion: 3600
|
||||
activeDeadlineSeconds: 7200
|
||||
arguments:
|
||||
parameters:
|
||||
- name: commit-ref
|
||||
value: HEAD
|
||||
- name: models
|
||||
value: "dqn,ppo,tft"
|
||||
- name: test-scope
|
||||
value: all
|
||||
- name: gpu-pool
|
||||
value: ci-training-h100
|
||||
- name: cuda-compute-cap
|
||||
value: "90"
|
||||
- name: notify
|
||||
value: "true"
|
||||
volumes:
|
||||
- name: git-ssh-key
|
||||
secret:
|
||||
secretName: argo-git-ssh-key
|
||||
defaultMode: 256
|
||||
- name: cargo-target
|
||||
persistentVolumeClaim:
|
||||
claimName: cargo-target-cuda-test
|
||||
- name: test-data
|
||||
persistentVolumeClaim:
|
||||
claimName: test-data-pvc
|
||||
readOnly: true
|
||||
|
||||
templates:
|
||||
# ── pipeline: DAG entrypoint ──
|
||||
# compile-and-test runs on GPU node (RWO PVC can't be shared cross-node).
|
||||
# gpu-warmup ensures H100 is scaled up before compile starts.
|
||||
- name: pipeline
|
||||
dag:
|
||||
tasks:
|
||||
- name: gpu-warmup
|
||||
template: gpu-warmup
|
||||
- name: compile-and-test
|
||||
template: compile-and-test
|
||||
dependencies: [gpu-warmup]
|
||||
- name: perf-benchmark
|
||||
template: perf-benchmark
|
||||
dependencies: [compile-and-test]
|
||||
|
||||
# ── gpu-warmup: trigger H100 autoscale ──
|
||||
# Requests GPU to force autoscaler to provision node, then releases it.
|
||||
- name: gpu-warmup
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: "{{workflow.parameters.gpu-pool}}"
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
tolerations:
|
||||
- key: nvidia.com/gpu
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
- key: node.cilium.io/agent-not-ready
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
container:
|
||||
image: alpine:3.21
|
||||
command: ["/bin/sh", "-c"]
|
||||
args:
|
||||
- |
|
||||
echo "GPU warmup: triggering node autoscale..."
|
||||
nvidia-smi --query-gpu=name,memory.total --format=csv,noheader || true
|
||||
echo "GPU node ready, releasing for compile-and-test"
|
||||
resources:
|
||||
requests:
|
||||
nvidia.com/gpu: "1"
|
||||
cpu: 100m
|
||||
memory: 64Mi
|
||||
limits:
|
||||
nvidia.com/gpu: "1"
|
||||
cpu: 200m
|
||||
memory: 128Mi
|
||||
|
||||
# ── compile-and-test: compile + run GPU tests in single H100 pod ──
|
||||
- name: compile-and-test
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: "{{workflow.parameters.gpu-pool}}"
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
tolerations:
|
||||
- key: nvidia.com/gpu
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
- key: node.cilium.io/agent-not-ready
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
outputs:
|
||||
parameters:
|
||||
- name: results
|
||||
valueFrom:
|
||||
path: /tmp/outputs/results
|
||||
default: "unknown:FAIL"
|
||||
- name: failures
|
||||
valueFrom:
|
||||
path: /tmp/outputs/failures
|
||||
default: "1"
|
||||
container:
|
||||
image: gitlab-registry.foxhunt.svc.cluster.local:5000/root/foxhunt/ci-builder:latest
|
||||
command: ["/bin/bash", "-c"]
|
||||
env:
|
||||
- name: SQLX_OFFLINE
|
||||
value: "true"
|
||||
- name: CARGO_TERM_COLOR
|
||||
value: always
|
||||
- name: CARGO_TARGET_DIR
|
||||
value: /cargo-target
|
||||
- name: CARGO_HOME
|
||||
value: /cargo-target/cargo-home
|
||||
- name: CUDA_VISIBLE_DEVICES
|
||||
value: "0"
|
||||
- name: CUBLAS_WORKSPACE_CONFIG
|
||||
value: ":4096:8"
|
||||
- name: TEST_DATA_DIR
|
||||
value: /data/test-data
|
||||
- name: RUST_LOG
|
||||
value: info
|
||||
- name: LD_LIBRARY_PATH
|
||||
value: /usr/local/nvidia/lib64:/usr/local/cuda/lib64:/usr/local/cuda/extras/CUPTI/lib64
|
||||
- name: CARGO_PROFILE_TEST_OPT_LEVEL
|
||||
value: "2"
|
||||
volumeMounts:
|
||||
- name: git-ssh-key
|
||||
mountPath: /etc/git-ssh
|
||||
readOnly: true
|
||||
- name: cargo-target
|
||||
mountPath: /cargo-target
|
||||
- name: test-data
|
||||
mountPath: /data/test-data
|
||||
readOnly: true
|
||||
resources:
|
||||
requests:
|
||||
nvidia.com/gpu: "1"
|
||||
cpu: "4"
|
||||
memory: 32Gi
|
||||
limits:
|
||||
nvidia.com/gpu: "1"
|
||||
cpu: "8"
|
||||
memory: 64Gi
|
||||
args:
|
||||
- |
|
||||
set -e
|
||||
|
||||
REF="{{workflow.parameters.commit-ref}}"
|
||||
MODELS="{{workflow.parameters.models}}"
|
||||
SCOPE="{{workflow.parameters.test-scope}}"
|
||||
|
||||
# --- SSH setup (same as compile-and-train) ---
|
||||
mkdir -p ~/.ssh
|
||||
cp /etc/git-ssh/ssh-privatekey ~/.ssh/id_ed25519
|
||||
chmod 600 ~/.ssh/id_ed25519
|
||||
printf 'Host *\n StrictHostKeyChecking no\n UserKnownHostsFile /dev/null\n' > ~/.ssh/config
|
||||
chmod 600 ~/.ssh/config
|
||||
|
||||
REPO="ssh://git@gitlab-gitlab-shell.foxhunt.svc.cluster.local:2222/root/foxhunt.git"
|
||||
BUILD="/cargo-target/src"
|
||||
git config --global --add safe.directory "$BUILD"
|
||||
|
||||
# --- Persistent checkout on PVC ---
|
||||
if [ -d "$BUILD/.git" ]; then
|
||||
cd "$BUILD"
|
||||
echo "=== Fetching latest refs ==="
|
||||
git fetch origin
|
||||
# Resolve REF after fetch so we always get the latest commit.
|
||||
# Try origin/$REF (branch), then $REF directly (tag or SHA).
|
||||
if git rev-parse --verify "origin/$REF" >/dev/null 2>&1; then
|
||||
TARGET=$(git rev-parse "origin/$REF")
|
||||
else
|
||||
TARGET=$(git rev-parse --verify "$REF" 2>/dev/null || echo "$REF")
|
||||
fi
|
||||
CURRENT=$(git rev-parse HEAD 2>/dev/null || echo "none")
|
||||
if [ "$CURRENT" = "$TARGET" ]; then
|
||||
echo "=== Already at $REF ($TARGET) ==="
|
||||
else
|
||||
echo "=== Updating checkout to $REF ($TARGET) ==="
|
||||
git checkout --force --detach "$TARGET"
|
||||
git clean -fd
|
||||
fi
|
||||
else
|
||||
echo "=== Initial clone ==="
|
||||
git clone --filter=blob:none "$REPO" "$BUILD"
|
||||
cd "$BUILD"
|
||||
git fetch origin
|
||||
if git rev-parse --verify "origin/$REF" >/dev/null 2>&1; then
|
||||
git checkout --force --detach "origin/$REF"
|
||||
else
|
||||
git checkout "$REF"
|
||||
fi
|
||||
fi
|
||||
echo "Checked out $(git rev-parse --short HEAD)"
|
||||
|
||||
export PATH="${CARGO_HOME}/bin:${PATH}"
|
||||
export CUDA_COMPUTE_CAP={{workflow.parameters.cuda-compute-cap}}
|
||||
|
||||
# --- PTX cache invalidation ---
|
||||
# Purge stale cached PTX if any CUDA kernel source changed since last run.
|
||||
bash scripts/ptx-cache-invalidate.sh "${CARGO_TARGET_DIR}/.ptx_cache"
|
||||
|
||||
# --- PVC size guard ---
|
||||
TARGET_SIZE_MB=$(du -sm "$CARGO_TARGET_DIR" 2>/dev/null | cut -f1 || echo 0)
|
||||
echo "PVC usage: ${TARGET_SIZE_MB}MB"
|
||||
if [ "$TARGET_SIZE_MB" -gt 25000 ]; then
|
||||
echo "PVC exceeds 25GB, pruning..."
|
||||
rm -rf "$CARGO_TARGET_DIR/release" "$CARGO_TARGET_DIR/debug"
|
||||
fi
|
||||
|
||||
# --- Compile ---
|
||||
echo "=== Compiling test binaries (--features cuda) ==="
|
||||
cargo test -p ml -p ml-dqn -p ml-core --features cuda --no-run 2>&1 | tee /cargo-target/gpu-compile.log
|
||||
echo "=== Compilation done ==="
|
||||
|
||||
# --- Expand "all" ---
|
||||
if [ "$MODELS" = "all" ]; then
|
||||
MODELS="dqn,ppo,tft,mamba2,tggn,tlob,liquid,kan,xlstm,diffusion"
|
||||
fi
|
||||
|
||||
# --- Test runner ---
|
||||
RESULTS=""
|
||||
FAILURES=0
|
||||
|
||||
# Reset CUDA context between test binaries to prevent cuBLAS
|
||||
# CUBLAS_STATUS_NOT_INITIALIZED cascades. Each test binary creates
|
||||
# and destroys hundreds of cuBLAS handles; without a reset, the
|
||||
# driver fails to re-init for the next binary.
|
||||
gpu_context_drain() {
|
||||
nvidia-smi -rgc >/dev/null 2>&1 || true
|
||||
sleep 1
|
||||
}
|
||||
|
||||
run_tests() {
|
||||
local NAME="$1"; shift
|
||||
echo ""
|
||||
echo "========================================"
|
||||
echo " TEST: $NAME"
|
||||
echo "========================================"
|
||||
# --nocapture is a test-binary flag, must come after --
|
||||
# If args already contain --, append after it; otherwise add -- first
|
||||
local HAS_SEP=false
|
||||
for arg in "$@"; do
|
||||
[ "$arg" = "--" ] && HAS_SEP=true && break
|
||||
done
|
||||
set +e
|
||||
if $HAS_SEP; then
|
||||
"$@" --nocapture 2>&1
|
||||
else
|
||||
"$@" -- --nocapture 2>&1
|
||||
fi
|
||||
EXIT=$?
|
||||
set -e
|
||||
if [ $EXIT -eq 0 ]; then
|
||||
RESULTS="${RESULTS}${NAME}:PASS\n"
|
||||
else
|
||||
RESULTS="${RESULTS}${NAME}:FAIL\n"
|
||||
FAILURES=$((FAILURES + 1))
|
||||
fi
|
||||
# Drain CUDA context after each GPU test suite
|
||||
gpu_context_drain
|
||||
}
|
||||
|
||||
# --- Core tests (always run) ---
|
||||
# --test-threads=1 for all GPU lib tests: prevents concurrent cuBLAS
|
||||
# handle creation that causes CUBLAS_STATUS_NOT_INITIALIZED cascades.
|
||||
run_tests "core-lib" cargo test -p ml-core --features cuda --lib -- --test-threads=1
|
||||
run_tests "bayesian" cargo test -p ml --features cuda --test bayesian_changepoint_test -- --test-threads=1
|
||||
|
||||
# --- Per-model tests ---
|
||||
IFS=',' read -ra MODEL_LIST <<< "$MODELS"
|
||||
for MODEL in "${MODEL_LIST[@]}"; do
|
||||
case "$MODEL" in
|
||||
dqn)
|
||||
if [ "$SCOPE" = "lib" ] || [ "$SCOPE" = "all" ]; then
|
||||
# --test-threads=1: GPU lib tests must run serially — each test
|
||||
# creates a cuBLAS handle via Device::new_cuda(0). Under parallel
|
||||
# execution, concurrent cuBLAS init races cause
|
||||
# CUBLAS_STATUS_NOT_INITIALIZED failures (51 test cascade).
|
||||
run_tests "dqn-lib" cargo test -p ml-dqn --features cuda --lib -- --test-threads=1
|
||||
run_tests "dqn-ml-lib" cargo test -p ml --features cuda --lib -- --test-threads=1 dqn
|
||||
fi
|
||||
if [ "$SCOPE" = "integration" ] || [ "$SCOPE" = "all" ]; then
|
||||
# --test-threads=1 for ALL GPU integration tests: parallel execution
|
||||
# corrupts the CUDA primary context (cuDevicePrimaryCtxRetain fails
|
||||
# when multiple threads race on context init/teardown).
|
||||
run_tests "dqn-smoke" cargo test -p ml --features cuda --test smoke_test_real_data -- --test-threads=1
|
||||
# Run each pipeline test in its own cargo test process.
|
||||
# CUDA Graph capture corrupts the async memory pool, making
|
||||
# cuMemAllocAsync fail with CUDA_ERROR_INVALID_VALUE in
|
||||
# subsequent DQNTrainer instances within the same process.
|
||||
run_tests "dqn-pipeline-train" cargo test -p ml --features cuda --test dqn_training_pipeline_test test_dqn_trains_on_es_fut -- --test-threads=1 --exact
|
||||
run_tests "dqn-pipeline-loss" cargo test -p ml --features cuda --test dqn_training_pipeline_test test_dqn_loss_decreases -- --test-threads=1 --exact
|
||||
run_tests "dqn-pipeline-ckpt" cargo test -p ml --features cuda --test dqn_training_pipeline_test test_dqn_checkpoint_save_load -- --test-threads=1 --exact
|
||||
run_tests "dqn-pipeline-qval" cargo test -p ml --features cuda --test dqn_training_pipeline_test test_dqn_q_value_predictions -- --test-threads=1 --exact
|
||||
run_tests "dqn-pipeline-eps" cargo test -p ml --features cuda --test dqn_training_pipeline_test test_dqn_epsilon_greedy -- --test-threads=1 --exact
|
||||
run_tests "dqn-smoke-train" cargo test -p ml --features cuda --test dqn_training_smoke_test -- --test-threads=1
|
||||
run_tests "dqn-early-stop" cargo test -p ml --features cuda --test dqn_early_stopping_termination_test -- --test-threads=1
|
||||
run_tests "dqn-collapse" cargo test -p ml --features cuda --test dqn_action_collapse_fix_test -- --test-threads=1
|
||||
fi
|
||||
;;
|
||||
ppo)
|
||||
if [ "$SCOPE" = "lib" ] || [ "$SCOPE" = "all" ]; then
|
||||
run_tests "ppo-lib" cargo test -p ml --features cuda --lib -- --test-threads=1 ppo
|
||||
fi
|
||||
if [ "$SCOPE" = "integration" ] || [ "$SCOPE" = "all" ]; then
|
||||
run_tests "ppo-barrier" cargo test -p ml --features cuda --test barrier_optimization_test -- --test-threads=1
|
||||
fi
|
||||
;;
|
||||
*)
|
||||
# Supervised models (TFT, Mamba2, TGGN, TLOB, Liquid, KAN, xLSTM, Diffusion)
|
||||
# Map model names to Rust module names where they differ
|
||||
LIB_FILTER="$MODEL"
|
||||
[ "$MODEL" = "tggn" ] && LIB_FILTER="tgnn"
|
||||
if [ "$SCOPE" = "lib" ] || [ "$SCOPE" = "all" ]; then
|
||||
run_tests "${MODEL}-lib" cargo test -p ml --features cuda --lib -- --test-threads=1 "$LIB_FILTER"
|
||||
fi
|
||||
if [ "$SCOPE" = "integration" ] || [ "$SCOPE" = "all" ]; then
|
||||
run_tests "${MODEL}-gpu" cargo test -p ml --features cuda --test supervised_gpu_smoke_test -- --test-threads=1 "test_${MODEL}_gpu_smoke"
|
||||
# Also run model-specific integration tests if they exist
|
||||
if cargo test -p ml --features cuda --test "${MODEL}_integration" --no-run 2>/dev/null; then
|
||||
run_tests "${MODEL}-integ" cargo test -p ml --features cuda --test "${MODEL}_integration" -- --test-threads=1
|
||||
fi
|
||||
fi
|
||||
;;
|
||||
esac
|
||||
done
|
||||
|
||||
# --- Summary ---
|
||||
echo ""
|
||||
echo "========================================"
|
||||
echo " GPU TEST RESULTS"
|
||||
echo "========================================"
|
||||
printf "$RESULTS" | while IFS=: read -r name status; do
|
||||
[ -z "$name" ] && continue
|
||||
printf " %-25s %s\n" "$name" "$status"
|
||||
done
|
||||
echo "========================================"
|
||||
echo " Total failures: $FAILURES"
|
||||
echo "========================================"
|
||||
|
||||
# Write results for notification step
|
||||
mkdir -p /tmp/outputs
|
||||
printf "$RESULTS" > /tmp/outputs/results
|
||||
echo "$FAILURES" > /tmp/outputs/failures
|
||||
|
||||
[ "$FAILURES" -gt 0 ] && exit 1 || exit 0
|
||||
|
||||
# ── perf-benchmark: DQN epoch/s on 3Q data (performance regression guard) ──
|
||||
# Runs after tests pass. Trains DQN on 3Q of ES.FUT data and reports epoch time.
|
||||
# Uses the same binary compiled by compile-and-test (shared cargo-target PVC).
|
||||
# Fails the pipeline if epoch time exceeds 500ms (regression threshold for H100).
|
||||
- name: perf-benchmark
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: "{{workflow.parameters.gpu-pool}}"
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
tolerations:
|
||||
- key: nvidia.com/gpu
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
- key: node.cilium.io/agent-not-ready
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
outputs:
|
||||
parameters:
|
||||
- name: epoch-ms
|
||||
valueFrom:
|
||||
path: /tmp/outputs/epoch-ms
|
||||
default: "9999"
|
||||
container:
|
||||
image: gitlab-registry.foxhunt.svc.cluster.local:5000/root/foxhunt/ci-builder:latest
|
||||
command: ["/bin/bash", "-c"]
|
||||
env:
|
||||
- name: SQLX_OFFLINE
|
||||
value: "true"
|
||||
- name: CARGO_TARGET_DIR
|
||||
value: /cargo-target
|
||||
- name: CARGO_HOME
|
||||
value: /cargo-target/cargo-home
|
||||
- name: CUDA_VISIBLE_DEVICES
|
||||
value: "0"
|
||||
- name: LD_LIBRARY_PATH
|
||||
value: /usr/local/nvidia/lib64:/usr/local/cuda/lib64:/usr/local/cuda/extras/CUPTI/lib64
|
||||
volumeMounts:
|
||||
- name: cargo-target
|
||||
mountPath: /cargo-target
|
||||
- name: test-data
|
||||
mountPath: /data/test-data
|
||||
readOnly: true
|
||||
resources:
|
||||
requests:
|
||||
nvidia.com/gpu: "1"
|
||||
cpu: "4"
|
||||
memory: 16Gi
|
||||
limits:
|
||||
nvidia.com/gpu: "1"
|
||||
cpu: "8"
|
||||
memory: 32Gi
|
||||
args:
|
||||
- |
|
||||
set -e
|
||||
cd /cargo-target/src
|
||||
export PATH="${CARGO_HOME}/bin:${PATH}"
|
||||
export CUDA_COMPUTE_CAP={{workflow.parameters.cuda-compute-cap}}
|
||||
|
||||
echo "========================================"
|
||||
echo " PERF BENCHMARK: DQN epoch/s (3Q ES.FUT)"
|
||||
echo "========================================"
|
||||
|
||||
# Run 5 epochs × 100 steps on 3Q data (--train-months 3 fits in 1 quarter)
|
||||
# Use --step-months 3 to avoid walk-forward window splits
|
||||
OUTPUT=$(cargo run --release --example train_baseline_rl -p ml -- \
|
||||
--model dqn \
|
||||
--data-dir /data/test-data/ohlcv \
|
||||
--mbp10-data-dir /data/test-data/mbp10 \
|
||||
--trades-data-dir /data/test-data/trades \
|
||||
--symbol ES.FUT \
|
||||
--epochs 5 \
|
||||
--train-months 3 --val-months 1 --test-months 1 --step-months 3 \
|
||||
2>&1)
|
||||
|
||||
# Extract epoch times (skip epoch 1 which includes init)
|
||||
EPOCH_TIMES=$(echo "$OUTPUT" | grep "phase breakdown" | grep -v "Epoch 1/" | \
|
||||
sed 's/.*total=\([0-9]*\)ms.*/\1/' | head -4)
|
||||
|
||||
if [ -z "$EPOCH_TIMES" ]; then
|
||||
echo "ERROR: No phase breakdown output found"
|
||||
echo "$OUTPUT" | tail -20
|
||||
mkdir -p /tmp/outputs
|
||||
echo "9999" > /tmp/outputs/epoch-ms
|
||||
exit 1
|
||||
fi
|
||||
|
||||
# Compute average epoch time (epochs 2-5)
|
||||
SUM=0
|
||||
COUNT=0
|
||||
for T in $EPOCH_TIMES; do
|
||||
SUM=$((SUM + T))
|
||||
COUNT=$((COUNT + 1))
|
||||
done
|
||||
AVG=$((SUM / COUNT))
|
||||
|
||||
echo ""
|
||||
echo " Epoch times (ms, excl. epoch 1): $EPOCH_TIMES"
|
||||
echo " Average: ${AVG}ms"
|
||||
echo ""
|
||||
|
||||
mkdir -p /tmp/outputs
|
||||
echo "$AVG" > /tmp/outputs/epoch-ms
|
||||
|
||||
# Regression guard: fail if avg epoch > 500ms on H100
|
||||
THRESHOLD=500
|
||||
if [ "$AVG" -gt "$THRESHOLD" ]; then
|
||||
echo "PERF REGRESSION: ${AVG}ms > ${THRESHOLD}ms threshold"
|
||||
echo "========================================"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
echo " PASS: ${AVG}ms <= ${THRESHOLD}ms threshold"
|
||||
echo "========================================"
|
||||
|
||||
# ── notify-result: post test outcome to Mattermost (onExit) ──
|
||||
- name: notify-result
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: platform
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
container:
|
||||
image: curlimages/curl:8.12.1
|
||||
command: ["/bin/sh", "-c"]
|
||||
env:
|
||||
- name: WEBHOOK_URL
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: notification-webhook
|
||||
key: webhook-url
|
||||
optional: true
|
||||
resources:
|
||||
requests:
|
||||
cpu: 50m
|
||||
memory: 32Mi
|
||||
limits:
|
||||
cpu: 100m
|
||||
memory: 64Mi
|
||||
args:
|
||||
- |
|
||||
NOTIFY="{{workflow.parameters.notify}}"
|
||||
if [ "$NOTIFY" != "true" ]; then
|
||||
echo "Notifications disabled, skipping"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
STATUS="{{workflow.status}}"
|
||||
NAME="{{workflow.name}}"
|
||||
|
||||
if [ -z "$WEBHOOK_URL" ] || echo "$WEBHOOK_URL" | grep -q "PLACEHOLDER"; then
|
||||
echo "No webhook configured, skipping notification"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
if [ "$STATUS" = "Succeeded" ]; then
|
||||
EMOJI=":white_check_mark:"
|
||||
else
|
||||
EMOJI=":x:"
|
||||
fi
|
||||
|
||||
PAYLOAD="{\"username\":\"Argo CI\",\"text\":\"${EMOJI} **GPU Tests** ${NAME} — ${STATUS} ({{workflow.duration}}s)\"}"
|
||||
|
||||
curl -sf -X POST -H 'Content-Type: application/json' \
|
||||
-d "$PAYLOAD" "$WEBHOOK_URL" || echo "WARN: webhook post failed"
|
||||
@@ -1,15 +1,8 @@
|
||||
apiVersion: kustomize.config.k8s.io/v1beta1
|
||||
kind: Kustomization
|
||||
# Rust train/CI/cache templates removed 2026-06-21 (decommission-rust-infra).
|
||||
# Remaining: platform RBAC + network policy used by the active fxhnt-cockpit deploys.
|
||||
resources:
|
||||
- sccache-pvcs.yaml
|
||||
- ci-pipeline-template.yaml
|
||||
- compile-and-deploy-template.yaml
|
||||
- train-template.yaml
|
||||
- build-ci-image-template.yaml
|
||||
- sanitizer-test-template.yaml
|
||||
- nsys-test-template.yaml
|
||||
- smoke-test-template.yaml
|
||||
- refresh-deps-cache-template.yaml
|
||||
- argo-workflow-netpol.yaml
|
||||
- ci-deploy-rbac.yaml
|
||||
- archive-rbac.yaml
|
||||
|
||||
@@ -1,459 +0,0 @@
|
||||
# Real-LOB backtest sweep workflow.
|
||||
#
|
||||
# Fans out a parameter grid across N parallel GPU pods, each running
|
||||
# `fxt-backtest run` against the same MBP-10 data + checkpoint with one
|
||||
# cell's worth of overrides. A final CPU pod runs `fxt-backtest aggregate`
|
||||
# to produce aggregate.parquet + pareto_frontier.json at the sweep root.
|
||||
#
|
||||
# The `# __SWEEP_CELLS__` marker on the dag.tasks line is replaced by
|
||||
# scripts/argo-lob-sweep.sh with N generated WorkflowTask stanzas before
|
||||
# submission. Same convention as train-multi-seed-template.yaml's
|
||||
# # __MATRIX_TASKS__ marker.
|
||||
#
|
||||
# DAG:
|
||||
# ensure-binary ──> [N parallel run-cell-<i> tasks on ci-training-l40s,
|
||||
# each with its own outputDir on the shared PVC]
|
||||
# │
|
||||
# └──> aggregate (CPU node, runs fxt-backtest aggregate
|
||||
# against the sweep root)
|
||||
#
|
||||
# Per `feedback_default_to_l40s_pool.md` (2026-05-09): defaults to
|
||||
# ci-training-l40s + sm_89. Override via --gpu-pool ci-training-h100
|
||||
# for sm_90 / 80 GB.
|
||||
#
|
||||
# Usage:
|
||||
# ./scripts/argo-lob-sweep.sh --grid config/ml/sweep_decision_stride_example.yaml
|
||||
# ./scripts/argo-lob-sweep.sh --grid <path> --sha abc1234 --gpu-pool ci-training-h100
|
||||
# ./scripts/argo-lob-sweep.sh --grid <path> --watch
|
||||
# ./scripts/argo-lob-sweep.sh --grid <path> --dry-run > /tmp/wf.yaml
|
||||
|
||||
apiVersion: argoproj.io/v1alpha1
|
||||
kind: WorkflowTemplate
|
||||
metadata:
|
||||
name: lob-backtest-sweep
|
||||
namespace: foxhunt
|
||||
labels:
|
||||
app.kubernetes.io/name: lob-backtest-sweep
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
app.kubernetes.io/component: train # reuses argo-train-workflow NetworkPolicy egress (port 2222 to gitlab-shell)
|
||||
spec:
|
||||
entrypoint: sweep-matrix
|
||||
onExit: notify-result
|
||||
serviceAccountName: argo-workflow
|
||||
archiveLogs: true
|
||||
podMetadata:
|
||||
labels:
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
app.kubernetes.io/component: train
|
||||
securityContext:
|
||||
fsGroup: 0
|
||||
ttlStrategy:
|
||||
secondsAfterCompletion: 3600
|
||||
# LOB sweep cells are much faster than training (~minutes vs hours);
|
||||
# 2h walltime is generous even for N=128 cells.
|
||||
activeDeadlineSeconds: 7200
|
||||
|
||||
arguments:
|
||||
parameters:
|
||||
- name: commit-sha
|
||||
value: HEAD
|
||||
- name: git-branch
|
||||
value: main
|
||||
- name: cuda-compute-cap
|
||||
value: "89"
|
||||
- name: gpu-pool
|
||||
value: ci-training-l40s
|
||||
- name: data-root
|
||||
value: /mnt/training-data/futures-baseline/ES.FUT
|
||||
- name: predecoded-dir
|
||||
value: /mnt/training-data/futures-baseline/ES.FUT
|
||||
- name: checkpoint
|
||||
value: "" # empty = trunk runs with --seed (random init);
|
||||
# set to a path on the data PVC for trained weights
|
||||
- name: sweep-root
|
||||
value: /mnt/training-data/sweeps/lob-backtest
|
||||
- name: sweep-tag
|
||||
value: default # subdirectory: <sweep-root>/<sweep-tag>/
|
||||
# Operational knobs. argo-lob-sweep.sh sets these per-grid.
|
||||
- name: n-parallel
|
||||
value: "1"
|
||||
- name: latency-ns
|
||||
value: "100000000"
|
||||
- name: target-annual-vol-units
|
||||
value: "50.0"
|
||||
- name: annualisation-factor
|
||||
value: "825.0"
|
||||
- name: max-lots
|
||||
value: "5"
|
||||
- name: max-events
|
||||
value: "0"
|
||||
# Per-step kernel-state JSONL trace path. Empty (default) = disabled:
|
||||
# ensure-binary compiles fxt-backtest WITHOUT --features kernel-step-trace
|
||||
# and run-cell omits the --kernel-step-trace flag. When set to a
|
||||
# non-empty path, ensure-binary rebuilds with the feature enabled
|
||||
# and run-cell passes --kernel-step-trace <path> through to the
|
||||
# binary. Trace is written to that absolute path on the pod (must
|
||||
# resolve to a mounted PVC — usually /feature-cache or /mnt/training-data).
|
||||
- name: kernel-step-trace
|
||||
value: ""
|
||||
|
||||
volumes:
|
||||
- name: training-data
|
||||
persistentVolumeClaim:
|
||||
claimName: training-data-pvc
|
||||
- name: git-ssh-key
|
||||
secret:
|
||||
secretName: argo-git-ssh-key
|
||||
defaultMode: 0400
|
||||
- name: cargo-target-cuda
|
||||
persistentVolumeClaim:
|
||||
claimName: cargo-target-cuda
|
||||
- name: feature-cache
|
||||
persistentVolumeClaim:
|
||||
claimName: feature-cache-pvc
|
||||
|
||||
templates:
|
||||
# ── DAG entrypoint ──────────────────────────────────────────────
|
||||
- name: sweep-matrix
|
||||
dag:
|
||||
tasks:
|
||||
- name: ensure-binary
|
||||
template: ensure-binary
|
||||
# __SWEEP_CELLS__ — replaced by argo-lob-sweep.sh
|
||||
- name: aggregate
|
||||
template: aggregate
|
||||
dependencies: [ensure-binary] # plus all run-cell-* via shell-injected list
|
||||
arguments:
|
||||
parameters:
|
||||
- name: sha
|
||||
value: "{{tasks.ensure-binary.outputs.parameters.sha}}"
|
||||
|
||||
# ── ensure-binary: cache-or-compile fxt-backtest by commit SHA ─
|
||||
# Same shape as train-multi-seed-template.yaml's ensure-binary but
|
||||
# for the fxt-backtest binary only.
|
||||
- name: ensure-binary
|
||||
outputs:
|
||||
parameters:
|
||||
- name: sha
|
||||
valueFrom:
|
||||
path: /tmp/sha
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: ci-compile-cpu
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
tolerations:
|
||||
- key: node.cilium.io/agent-not-ready
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
container:
|
||||
image: gitlab-registry.foxhunt.svc.cluster.local:5000/root/foxhunt/ci-builder:latest
|
||||
imagePullPolicy: Always
|
||||
command: ["/bin/bash", "-c"]
|
||||
env:
|
||||
- name: SQLX_OFFLINE
|
||||
value: "true"
|
||||
- name: CARGO_TARGET_DIR
|
||||
value: /cargo-target
|
||||
- name: CARGO_HOME
|
||||
value: /cargo-target/cargo-home
|
||||
- name: RUSTC_WRAPPER
|
||||
value: sccache
|
||||
- name: SCCACHE_DIR
|
||||
value: /cargo-target/sccache
|
||||
- name: SCCACHE_CACHE_SIZE
|
||||
value: "40G"
|
||||
- name: CARGO_INCREMENTAL
|
||||
value: "0"
|
||||
- name: FOXHUNT_CUDA_ARCH
|
||||
value: "sm_{{workflow.parameters.cuda-compute-cap}}"
|
||||
- name: GITLAB_PAT
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: gitlab-pat
|
||||
key: token
|
||||
resources:
|
||||
requests:
|
||||
cpu: "14"
|
||||
memory: 32Gi
|
||||
limits:
|
||||
cpu: "30"
|
||||
memory: 64Gi
|
||||
volumeMounts:
|
||||
- name: git-ssh-key
|
||||
mountPath: /etc/git-ssh
|
||||
readOnly: true
|
||||
- name: cargo-target-cuda
|
||||
mountPath: /cargo-target
|
||||
- name: training-data
|
||||
mountPath: /mnt/training-data
|
||||
args:
|
||||
- |
|
||||
set -e
|
||||
SHA="{{workflow.parameters.commit-sha}}"
|
||||
BRANCH="{{workflow.parameters.git-branch}}"
|
||||
KERNEL_TRACE="{{workflow.parameters.kernel-step-trace}}"
|
||||
|
||||
mkdir -p ~/.ssh
|
||||
cp /etc/git-ssh/ssh-privatekey ~/.ssh/id_ed25519
|
||||
chmod 600 ~/.ssh/id_ed25519
|
||||
printf 'Host *\n StrictHostKeyChecking no\n UserKnownHostsFile /dev/null\n' > ~/.ssh/config
|
||||
chmod 600 ~/.ssh/config
|
||||
|
||||
REPO="ssh://git@gitlab-gitlab-shell.foxhunt.svc.cluster.local:2222/root/foxhunt.git"
|
||||
BUILD="/cargo-target/src"
|
||||
git config --global --add safe.directory "$BUILD"
|
||||
|
||||
if [ "$SHA" = "HEAD" ]; then
|
||||
if [ -d "$BUILD/.git" ]; then
|
||||
cd "$BUILD"; git fetch origin
|
||||
SHA=$(git rev-parse "origin/$BRANCH"); cd /
|
||||
else
|
||||
git clone --filter=blob:none "$REPO" "$BUILD"
|
||||
cd "$BUILD"; git checkout "origin/$BRANCH"
|
||||
SHA=$(git rev-parse HEAD); cd /
|
||||
fi
|
||||
fi
|
||||
SHORT_SHA=$(echo "$SHA" | cut -c1-9)
|
||||
echo "Resolved SHA: $SHA (short: $SHORT_SHA)"
|
||||
|
||||
# Feature variant: when kernel-step-trace is enabled, build
|
||||
# with the Cargo feature and cache under a distinct subdir
|
||||
# so the default + diagnostic binaries don't clobber each other.
|
||||
FEATURE_FLAGS=""
|
||||
VARIANT="default"
|
||||
if [ -n "$KERNEL_TRACE" ]; then
|
||||
FEATURE_FLAGS="--features kernel-step-trace"
|
||||
VARIANT="kstrace"
|
||||
fi
|
||||
BIN_DIR="/mnt/training-data/bin/$SHORT_SHA-$VARIANT"
|
||||
mkdir -p "$BIN_DIR"
|
||||
if [ -x "$BIN_DIR/fxt-backtest" ]; then
|
||||
echo "=== Cache HIT: fxt-backtest ($VARIANT) present in $BIN_DIR ==="
|
||||
ls -lh "$BIN_DIR/fxt-backtest"
|
||||
echo "$SHORT_SHA-$VARIANT" > /tmp/sha
|
||||
exit 0
|
||||
fi
|
||||
|
||||
echo "=== Cache MISS: compiling fxt-backtest ($VARIANT) for $SHORT_SHA ==="
|
||||
if [ -d "$BUILD/.git" ]; then
|
||||
# Same pattern as alpha-perception-template.yaml: prior
|
||||
# cargo build mutates Cargo.lock (and sometimes other
|
||||
# generated files); --force checkout overwrites them,
|
||||
# `git clean -fd` drops untracked. Without this, the
|
||||
# second SHA's checkout fails on dirty working tree.
|
||||
cd "$BUILD"; git fetch origin
|
||||
git checkout --force "$SHA"; git clean -fd
|
||||
else
|
||||
git clone --filter=blob:none "$REPO" "$BUILD"
|
||||
cd "$BUILD"; git checkout "$SHA"
|
||||
fi
|
||||
|
||||
cargo build -p fxt-backtest --release $FEATURE_FLAGS
|
||||
cp "$CARGO_TARGET_DIR/release/fxt-backtest" "$BIN_DIR/fxt-backtest"
|
||||
echo "$SHORT_SHA-$VARIANT" > /tmp/sha
|
||||
ls -lh "$BIN_DIR/fxt-backtest"
|
||||
|
||||
# ── run-cell: one sweep cell against a GPU pod ─────────────────
|
||||
# Inputs: cell name + every Run arg that varies across the grid.
|
||||
# The shell-rendered DAG tasks supply these per-cell.
|
||||
- name: run-cell
|
||||
inputs:
|
||||
parameters:
|
||||
- name: sha
|
||||
- name: cell-name
|
||||
- name: decision-stride
|
||||
value: "4"
|
||||
- name: latency-ns
|
||||
value: "{{workflow.parameters.latency-ns}}"
|
||||
- name: target-annual-vol-units
|
||||
value: "{{workflow.parameters.target-annual-vol-units}}"
|
||||
- name: annualisation-factor
|
||||
value: "{{workflow.parameters.annualisation-factor}}"
|
||||
- name: max-lots
|
||||
value: "{{workflow.parameters.max-lots}}"
|
||||
- name: max-events
|
||||
value: "{{workflow.parameters.max-events}}"
|
||||
- name: seed
|
||||
value: "0xC0FFEE"
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: "{{workflow.parameters.gpu-pool}}"
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
tolerations:
|
||||
- key: nvidia.com/gpu
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
container:
|
||||
image: gitlab-registry.foxhunt.svc.cluster.local:5000/root/foxhunt/foxhunt-training-runtime:latest
|
||||
imagePullPolicy: Always
|
||||
command: ["/bin/bash", "-c"]
|
||||
resources:
|
||||
requests:
|
||||
cpu: "4"
|
||||
memory: 16Gi
|
||||
nvidia.com/gpu: "1"
|
||||
limits:
|
||||
cpu: "8"
|
||||
memory: 32Gi
|
||||
nvidia.com/gpu: "1"
|
||||
volumeMounts:
|
||||
- name: training-data
|
||||
mountPath: /mnt/training-data
|
||||
- name: feature-cache
|
||||
mountPath: /feature-cache
|
||||
args:
|
||||
- |
|
||||
set -e
|
||||
SHA="{{inputs.parameters.sha}}"
|
||||
CELL="{{inputs.parameters.cell-name}}"
|
||||
BIN="/mnt/training-data/bin/$SHA/fxt-backtest"
|
||||
CKPT="{{workflow.parameters.checkpoint}}"
|
||||
KERNEL_TRACE="{{workflow.parameters.kernel-step-trace}}"
|
||||
OUT="{{workflow.parameters.sweep-root}}/{{workflow.parameters.sweep-tag}}/$CELL"
|
||||
mkdir -p "$OUT"
|
||||
echo "=== sweep cell $CELL on $(hostname) → $OUT ==="
|
||||
|
||||
CKPT_FLAG=""
|
||||
if [ -n "$CKPT" ]; then
|
||||
CKPT_FLAG="--checkpoint $CKPT"
|
||||
fi
|
||||
KERNEL_TRACE_FLAG=""
|
||||
if [ -n "$KERNEL_TRACE" ]; then
|
||||
# Per-cell trace path: append cell name so concurrent run-cell
|
||||
# pods don't clobber each other's JSONL files. Operators get
|
||||
# one trace per cell; aggregate by reading <root>/<sweep>/<cell>/
|
||||
# at analysis time.
|
||||
KERNEL_TRACE_FLAG="--kernel-step-trace $OUT/kernel_step_trace.jsonl"
|
||||
fi
|
||||
|
||||
"$BIN" run \
|
||||
--data "{{workflow.parameters.data-root}}" \
|
||||
--predecoded-dir "{{workflow.parameters.predecoded-dir}}" \
|
||||
--n-parallel "{{workflow.parameters.n-parallel}}" \
|
||||
--decision-stride "{{inputs.parameters.decision-stride}}" \
|
||||
--latency-ns "{{inputs.parameters.latency-ns}}" \
|
||||
--target-annual-vol-units "{{inputs.parameters.target-annual-vol-units}}" \
|
||||
--annualisation-factor "{{inputs.parameters.annualisation-factor}}" \
|
||||
--max-lots "{{inputs.parameters.max-lots}}" \
|
||||
--max-events "{{inputs.parameters.max-events}}" \
|
||||
--seed "{{inputs.parameters.seed}}" \
|
||||
$CKPT_FLAG \
|
||||
$KERNEL_TRACE_FLAG \
|
||||
--out "$OUT"
|
||||
|
||||
echo "=== cell $CELL done ==="
|
||||
ls -lh "$OUT"
|
||||
|
||||
# ── run-sweep: P6 batched flow. Invokes `fxt-backtest sweep` ───
|
||||
# against the FULL grid YAML inside a single GPU pod. When the YAML
|
||||
# carries sim_variants, each cell expands into n_parallel=variants.len()
|
||||
# backtests sharing one forward pass. argo-lob-sweep.sh emits ONE
|
||||
# run-sweep task (instead of N fan-out run-cell tasks) when --batched
|
||||
# is set or sim_variants is detected in the YAML.
|
||||
- name: run-sweep
|
||||
inputs:
|
||||
parameters:
|
||||
- name: sha
|
||||
- name: grid-yaml-b64 # base64-encoded grid YAML, written to a file in-pod
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: "{{workflow.parameters.gpu-pool}}"
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
tolerations:
|
||||
- key: nvidia.com/gpu
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
container:
|
||||
image: gitlab-registry.foxhunt.svc.cluster.local:5000/root/foxhunt/foxhunt-training-runtime:latest
|
||||
imagePullPolicy: Always
|
||||
command: ["/bin/bash", "-c"]
|
||||
resources:
|
||||
requests:
|
||||
cpu: "4"
|
||||
memory: 16Gi
|
||||
nvidia.com/gpu: "1"
|
||||
limits:
|
||||
cpu: "8"
|
||||
memory: 32Gi
|
||||
nvidia.com/gpu: "1"
|
||||
volumeMounts:
|
||||
- name: training-data
|
||||
mountPath: /mnt/training-data
|
||||
- name: feature-cache
|
||||
mountPath: /feature-cache
|
||||
args:
|
||||
- |
|
||||
set -e
|
||||
SHA="{{inputs.parameters.sha}}"
|
||||
BIN="/mnt/training-data/bin/$SHA/fxt-backtest"
|
||||
OUT="{{workflow.parameters.sweep-root}}/{{workflow.parameters.sweep-tag}}"
|
||||
mkdir -p "$OUT"
|
||||
# Write the grid YAML to a known path. base64 in/out keeps Argo
|
||||
# parameter encoding stable across YAML special characters
|
||||
# (curly braces from data_template + sim_variants).
|
||||
GRID_YAML=/tmp/sweep-grid.yaml
|
||||
echo "{{inputs.parameters.grid-yaml-b64}}" | base64 -d > "$GRID_YAML"
|
||||
echo "=== running fxt-backtest sweep on $(hostname) → $OUT ==="
|
||||
echo "=== grid yaml: ==="
|
||||
head -20 "$GRID_YAML"
|
||||
echo "=== ... ==="
|
||||
|
||||
"$BIN" sweep --grid "$GRID_YAML" --out "$OUT"
|
||||
|
||||
echo "=== sweep complete ==="
|
||||
ls -lh "$OUT"
|
||||
|
||||
# ── aggregate: runs fxt-backtest aggregate on the GPU pool. ───
|
||||
# The binary is dynamically linked against libcuda.so.1, which the
|
||||
# ci-compile-cpu pool's host doesn't expose; the aggregate workload
|
||||
# itself is CPU-only (~seconds) but the binary needs the driver libs.
|
||||
# Reusing the GPU pool is cheap (single fast pod) and avoids needing
|
||||
# a separate CPU-only binary.
|
||||
- name: aggregate
|
||||
inputs:
|
||||
parameters:
|
||||
- name: sha
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: "{{workflow.parameters.gpu-pool}}"
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
tolerations:
|
||||
- key: nvidia.com/gpu
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
container:
|
||||
image: gitlab-registry.foxhunt.svc.cluster.local:5000/root/foxhunt/foxhunt-training-runtime:latest
|
||||
imagePullPolicy: Always
|
||||
command: ["/bin/bash", "-c"]
|
||||
resources:
|
||||
requests:
|
||||
cpu: "2"
|
||||
memory: 4Gi
|
||||
# Request a GPU only to make the Scaleway L40S device plugin
|
||||
# mount libcuda.so.1 into the container — aggregate logic is
|
||||
# CPU-only (~seconds). Without the request, the binary's
|
||||
# dynamic loader fails before main(). Cheapest fix; the GPU
|
||||
# is held for ~5s of CPU work, then released.
|
||||
nvidia.com/gpu: "1"
|
||||
limits:
|
||||
cpu: "4"
|
||||
memory: 8Gi
|
||||
nvidia.com/gpu: "1"
|
||||
volumeMounts:
|
||||
- name: training-data
|
||||
mountPath: /mnt/training-data
|
||||
args:
|
||||
- |
|
||||
set -e
|
||||
SHA="{{inputs.parameters.sha}}"
|
||||
BIN="/mnt/training-data/bin/$SHA/fxt-backtest"
|
||||
SWEEP_DIR="{{workflow.parameters.sweep-root}}/{{workflow.parameters.sweep-tag}}"
|
||||
echo "=== aggregate sweep at $SWEEP_DIR ==="
|
||||
"$BIN" aggregate "$SWEEP_DIR"
|
||||
ls -lh "$SWEEP_DIR"
|
||||
|
||||
# ── notify-result: exit hook (placeholder; real impl emits to Slack/MinIO) ─
|
||||
- name: notify-result
|
||||
container:
|
||||
image: alpine:3.20
|
||||
command: ["sh", "-c"]
|
||||
args:
|
||||
- |
|
||||
echo "lob-backtest-sweep workflow {{workflow.name}} finished with status {{workflow.status}}"
|
||||
echo "results under {{workflow.parameters.sweep-root}}/{{workflow.parameters.sweep-tag}}"
|
||||
@@ -1,318 +0,0 @@
|
||||
# infra/k8s/argo/nsys-test-template.yaml
|
||||
#
|
||||
# One-shot Nsight Systems profiling run on L40S. Wraps a smoke test under
|
||||
# `nsys profile` and uploads the .nsys-rep to MinIO at
|
||||
# foxhunt-training-results/profiles/smoke/<short-sha>/.
|
||||
#
|
||||
# Mirrors the nsys integration in train-multi-seed-template (Plan 5 Task 3,
|
||||
# A.4.1) but for smoke tests. Use this for rapid kernel-timing iteration
|
||||
# (per-step bottleneck hunts, SOL %, occupancy diffs) without the full
|
||||
# multi-seed Argo cycle.
|
||||
#
|
||||
# Why L40S not local: laptop RTX 3050 Ti has 4 GB VRAM; nsys on L40S (48 GB)
|
||||
# avoids any contention and runs at near-native speed (typical ~5-10%
|
||||
# slowdown vs uninstrumented).
|
||||
#
|
||||
# Usage:
|
||||
# argo submit --watch -n foxhunt nsys-test-template.yaml \
|
||||
# -p commit-ref=main \
|
||||
# -p test-name=performance::test_real_data_single_epoch
|
||||
#
|
||||
# Or via the wrapper script: scripts/argo-nsys.sh
|
||||
|
||||
apiVersion: argoproj.io/v1alpha1
|
||||
kind: WorkflowTemplate
|
||||
metadata:
|
||||
name: nsys-test
|
||||
namespace: foxhunt
|
||||
labels:
|
||||
app.kubernetes.io/name: nsys-test
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
spec:
|
||||
entrypoint: nsys-run
|
||||
serviceAccountName: argo-workflow
|
||||
podMetadata:
|
||||
labels:
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
app.kubernetes.io/component: nsys-test
|
||||
archiveLogs: true
|
||||
podGC:
|
||||
strategy: OnPodCompletion
|
||||
ttlStrategy:
|
||||
secondsAfterCompletion: 7200
|
||||
activeDeadlineSeconds: 3600 # nsys overhead is small — 1h cap is generous
|
||||
arguments:
|
||||
parameters:
|
||||
- name: commit-ref
|
||||
value: HEAD
|
||||
- name: gpu-pool
|
||||
value: ci-training-l40s
|
||||
- name: cuda-compute-cap
|
||||
value: "89"
|
||||
- name: test-name
|
||||
value: "performance::test_real_data_single_epoch"
|
||||
# nsys capture options. Defaults capture the full kernel timeline.
|
||||
# GPU metrics counters (--gpu-metrics-devices) require elevated CUDA
|
||||
# performance-counter privileges (NVGPUCTRPERM) which the container
|
||||
# lacks; CUDA/NVTX/osrt traces alone still give per-kernel timing and
|
||||
# call-graph context — sufficient for the bottleneck-hunt use case.
|
||||
# Override via --extra-args if running on a node with relaxed perms.
|
||||
- name: nsys-extra-args
|
||||
value: "--trace=cuda,nvtx,osrt"
|
||||
|
||||
volumes:
|
||||
- name: git-ssh-key
|
||||
secret:
|
||||
secretName: argo-git-ssh-key
|
||||
defaultMode: 256
|
||||
- name: cargo-target
|
||||
persistentVolumeClaim:
|
||||
claimName: cargo-target-cuda-test
|
||||
# Use the same training-data-pvc as production training (full 27 months
|
||||
# of OHLCV + MBP-10 + trades) instead of the smaller test-data-pvc.
|
||||
- name: training-data
|
||||
persistentVolumeClaim:
|
||||
claimName: training-data-pvc
|
||||
readOnly: true
|
||||
|
||||
templates:
|
||||
- name: nsys-run
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: "{{workflow.parameters.gpu-pool}}"
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
tolerations:
|
||||
- key: nvidia.com/gpu
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
- key: node.cilium.io/agent-not-ready
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
container:
|
||||
image: gitlab-registry.foxhunt.svc.cluster.local:5000/root/foxhunt/ci-builder:latest
|
||||
command: ["/bin/bash", "-c"]
|
||||
env:
|
||||
- name: SQLX_OFFLINE
|
||||
value: "true"
|
||||
- name: CARGO_TERM_COLOR
|
||||
value: always
|
||||
- name: CARGO_TARGET_DIR
|
||||
value: /cargo-target
|
||||
- name: CARGO_HOME
|
||||
value: /cargo-target/cargo-home
|
||||
- name: CUDA_VISIBLE_DEVICES
|
||||
value: "0"
|
||||
- name: CUBLAS_WORKSPACE_CONFIG
|
||||
value: ":4096:8"
|
||||
- name: FOXHUNT_TEST_DATA
|
||||
value: /data/futures-baseline
|
||||
- name: TEST_DATA_DIR
|
||||
value: /data/futures-baseline
|
||||
# MBP-10 + trades are mandatory per feedback_mbp10_mandatory.md
|
||||
# (OFI features are part of the production state vector).
|
||||
- name: FOXHUNT_MBP10_DATA
|
||||
value: /data/futures-baseline-mbp10
|
||||
- name: FOXHUNT_TRADES_DATA
|
||||
value: /data/futures-baseline-trades
|
||||
- name: RUST_LOG
|
||||
value: info
|
||||
- name: LD_LIBRARY_PATH
|
||||
value: /usr/local/nvidia/lib64:/usr/local/cuda/lib64:/usr/local/cuda/extras/CUPTI/lib64
|
||||
# MinIO upload — same secret + optional pattern as train-multi-seed.
|
||||
# Both refs are `optional: true` so env mount succeeds on clusters
|
||||
# that do not pre-create `minio-credentials` (upload then warn-fails).
|
||||
- name: MINIO_ACCESS_KEY
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: minio-credentials
|
||||
key: access-key
|
||||
optional: true
|
||||
- name: MINIO_SECRET_KEY
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: minio-credentials
|
||||
key: secret-key
|
||||
optional: true
|
||||
volumeMounts:
|
||||
- name: git-ssh-key
|
||||
mountPath: /etc/git-ssh
|
||||
readOnly: true
|
||||
- name: cargo-target
|
||||
mountPath: /cargo-target
|
||||
- name: training-data
|
||||
mountPath: /data
|
||||
readOnly: true
|
||||
resources:
|
||||
requests:
|
||||
nvidia.com/gpu: "1"
|
||||
cpu: "4"
|
||||
memory: 32Gi
|
||||
limits:
|
||||
nvidia.com/gpu: "1"
|
||||
cpu: "8"
|
||||
memory: 64Gi
|
||||
args:
|
||||
- |
|
||||
set -euo pipefail
|
||||
|
||||
REF="{{workflow.parameters.commit-ref}}"
|
||||
TEST_NAME="{{workflow.parameters.test-name}}"
|
||||
EXTRA_ARGS="{{workflow.parameters.nsys-extra-args}}"
|
||||
|
||||
echo "==================================="
|
||||
echo " nsys profile L40S run"
|
||||
echo "==================================="
|
||||
echo " Ref: $REF"
|
||||
echo " Test: $TEST_NAME"
|
||||
echo " Extra: $EXTRA_ARGS"
|
||||
echo " Pool: {{workflow.parameters.gpu-pool}}"
|
||||
echo "==================================="
|
||||
|
||||
# --- SSH setup ---
|
||||
mkdir -p ~/.ssh
|
||||
cp /etc/git-ssh/ssh-privatekey ~/.ssh/id_ed25519
|
||||
chmod 600 ~/.ssh/id_ed25519
|
||||
printf 'Host *\n StrictHostKeyChecking no\n UserKnownHostsFile /dev/null\n' > ~/.ssh/config
|
||||
chmod 600 ~/.ssh/config
|
||||
|
||||
REPO="ssh://git@gitlab-gitlab-shell.foxhunt.svc.cluster.local:2222/root/foxhunt.git"
|
||||
BUILD="/cargo-target/src"
|
||||
git config --global --add safe.directory "$BUILD"
|
||||
|
||||
# --- Persistent checkout on PVC ---
|
||||
if [ -d "$BUILD/.git" ]; then
|
||||
cd "$BUILD"
|
||||
git fetch origin
|
||||
if git rev-parse --verify "origin/$REF" >/dev/null 2>&1; then
|
||||
TARGET=$(git rev-parse "origin/$REF")
|
||||
else
|
||||
TARGET=$(git rev-parse --verify "$REF" 2>/dev/null || echo "$REF")
|
||||
fi
|
||||
CURRENT=$(git rev-parse HEAD 2>/dev/null || echo "none")
|
||||
if [ "$CURRENT" != "$TARGET" ]; then
|
||||
echo "=== Updating checkout to $REF ($TARGET) ==="
|
||||
git checkout --force --detach "$TARGET"
|
||||
git clean -fd
|
||||
fi
|
||||
else
|
||||
git clone --filter=blob:none "$REPO" "$BUILD"
|
||||
cd "$BUILD"
|
||||
git fetch origin
|
||||
if git rev-parse --verify "origin/$REF" >/dev/null 2>&1; then
|
||||
git checkout --force --detach "origin/$REF"
|
||||
else
|
||||
git checkout "$REF"
|
||||
fi
|
||||
fi
|
||||
SHA=$(git rev-parse --short HEAD)
|
||||
echo "Checked out $SHA"
|
||||
|
||||
export PATH="${CARGO_HOME}/bin:/usr/local/cuda/bin:${PATH}"
|
||||
export CUDA_COMPUTE_CAP={{workflow.parameters.cuda-compute-cap}}
|
||||
|
||||
bash scripts/ptx-cache-invalidate.sh "${CARGO_TARGET_DIR}/.ptx_cache"
|
||||
|
||||
# --- Verify nsys ---
|
||||
NSYS_BIN=$(which nsys 2>/dev/null || echo /usr/local/cuda/bin/nsys)
|
||||
if [ ! -x "$NSYS_BIN" ]; then
|
||||
echo "ERROR: nsys not found at $NSYS_BIN"
|
||||
exit 2
|
||||
fi
|
||||
echo "=== nsys version ==="
|
||||
"$NSYS_BIN" --version | head -3
|
||||
|
||||
# --- Compile test binary + train_baseline_rl ---
|
||||
echo "=== Compiling (--release --no-run + train_baseline_rl) ==="
|
||||
cargo test -p ml --release --lib --no-run 2>&1 | tee /cargo-target/nsys-compile.log
|
||||
cargo build -p ml --release --example train_baseline_rl 2>&1 | tee -a /cargo-target/nsys-compile.log
|
||||
|
||||
# Honour CARGO_TARGET_DIR — outputs go to ${CARGO_TARGET_DIR}/release/deps,
|
||||
# NOT the repo-relative target/. (CI workflow sets CARGO_TARGET_DIR=/cargo-target.)
|
||||
# Subshell + || true wraps the head -1 SIGPIPE so pipefail doesn't fire.
|
||||
TARGET_DIR="${CARGO_TARGET_DIR:-target}"
|
||||
TEST_BIN=$( (ls -t "$TARGET_DIR/release/deps/ml-"* 2>/dev/null \
|
||||
| grep -vE '\.(d|rcgu|rmeta|rlib|so|json)$' \
|
||||
| head -1) || true )
|
||||
if [ -z "$TEST_BIN" ] || [ ! -x "$TEST_BIN" ]; then
|
||||
echo "ERROR: could not locate compiled test binary in $TARGET_DIR/release/deps/"
|
||||
ls -la "$TARGET_DIR/release/deps/" 2>&1 | head -20
|
||||
exit 3
|
||||
fi
|
||||
echo "=== Test binary: $TEST_BIN ==="
|
||||
|
||||
# grep -q + pipefail = SIGPIPE on early exit. Capture once, grep
|
||||
# without -q so the consumer reads all input.
|
||||
TEST_LIST=$("$TEST_BIN" --list 2>&1 || true)
|
||||
if ! echo "$TEST_LIST" | grep -F "$TEST_NAME" > /dev/null; then
|
||||
echo "ERROR: test '$TEST_NAME' not found"
|
||||
echo "$TEST_LIST" | grep smoke_tests | (head -30 || true)
|
||||
exit 4
|
||||
fi
|
||||
|
||||
# --- Run under nsys profile ---
|
||||
mkdir -p /tmp/nsys-output
|
||||
POD_NAME="${HOSTNAME:-pod}"
|
||||
NSYS_OUT="/tmp/nsys-output/profile-${POD_NAME}.nsys-rep"
|
||||
|
||||
echo ""
|
||||
echo "==================================="
|
||||
echo " Launching nsys profile..."
|
||||
echo "==================================="
|
||||
START_TS=$(date +%s)
|
||||
set +e
|
||||
"$NSYS_BIN" profile \
|
||||
--output="$NSYS_OUT" \
|
||||
--force-overwrite=true \
|
||||
--stats=true \
|
||||
${EXTRA_ARGS} \
|
||||
-- "$TEST_BIN" "$TEST_NAME" --ignored --nocapture
|
||||
EXIT_CODE=$?
|
||||
set -e
|
||||
END_TS=$(date +%s)
|
||||
ELAPSED=$((END_TS - START_TS))
|
||||
|
||||
echo ""
|
||||
echo "==================================="
|
||||
echo " nsys run complete"
|
||||
echo "==================================="
|
||||
echo " Exit code: $EXIT_CODE"
|
||||
echo " Elapsed: ${ELAPSED}s"
|
||||
ls -la "$NSYS_OUT" 2>/dev/null && \
|
||||
echo " Profile: $(du -h "$NSYS_OUT" | cut -f1)"
|
||||
echo ""
|
||||
|
||||
# --- Upload to MinIO (same pattern as train-multi-seed) ---
|
||||
if [ -f "$NSYS_OUT" ]; then
|
||||
echo "=== Uploading nsys profile to MinIO ==="
|
||||
MC_BIN=$(which mc 2>/dev/null || echo "")
|
||||
if [ -z "$MC_BIN" ]; then
|
||||
MC_BIN=/tmp/mc
|
||||
curl -fsSL https://dl.min.io/client/mc/release/linux-amd64/mc -o "$MC_BIN" || {
|
||||
echo "WARN: failed to download mc — skipping upload"
|
||||
echo "Profile available locally at: $NSYS_OUT (PVC: cargo-target)"
|
||||
exit 0
|
||||
}
|
||||
chmod +x "$MC_BIN"
|
||||
fi
|
||||
"$MC_BIN" alias set foxhunt http://minio.foxhunt.svc.cluster.local:9000 \
|
||||
"${MINIO_ACCESS_KEY:-}" "${MINIO_SECRET_KEY:-}" 2>/dev/null || true
|
||||
# Smoke profiles go to a separate prefix to keep the
|
||||
# production training profile bucket clean.
|
||||
UPLOAD_PATH="foxhunt/foxhunt-training-results/profiles/smoke/${SHA}/profile-${POD_NAME}.nsys-rep"
|
||||
if "$MC_BIN" cp "$NSYS_OUT" "$UPLOAD_PATH"; then
|
||||
echo "=== nsys profile uploaded to ${UPLOAD_PATH} ==="
|
||||
echo ""
|
||||
echo "Download with:"
|
||||
echo " mc cp ${UPLOAD_PATH} ./profile.nsys-rep"
|
||||
echo "Open with: nsys-ui profile.nsys-rep"
|
||||
else
|
||||
echo "WARN: nsys upload failed — profile remains on PVC at $NSYS_OUT"
|
||||
fi
|
||||
fi
|
||||
|
||||
if [ "$EXIT_CODE" -ne 0 ]; then
|
||||
echo ""
|
||||
echo "FAIL: test exited with code $EXIT_CODE"
|
||||
exit "$EXIT_CODE"
|
||||
fi
|
||||
echo ""
|
||||
echo "PASS: nsys profile captured, test passed."
|
||||
@@ -1,181 +0,0 @@
|
||||
# refresh-deps-cache: rebuild ci-builder-cpu-with-deps:nightly image
|
||||
#
|
||||
# Triggered:
|
||||
# - nightly via CronWorkflow (see CronWorkflow at the bottom of this file)
|
||||
# - manually via: argo submit --from workflowtemplate/refresh-deps-cache -n foxhunt
|
||||
#
|
||||
# What it does: clones the repo at HEAD of main, then runs Kaniko to build
|
||||
# infra/docker/Dockerfile.ci-deps-cache and pushes to
|
||||
# gitlab-registry.foxhunt.svc.cluster.local:5000/root/foxhunt/ci-builder-cpu-with-deps:nightly.
|
||||
#
|
||||
# The compile-and-deploy workflow's seed-deps-cache initContainer pulls
|
||||
# this image and rsyncs its /cargo-target-prebuilt/ into the cargo-target-cpu PVC.
|
||||
---
|
||||
apiVersion: argoproj.io/v1alpha1
|
||||
kind: WorkflowTemplate
|
||||
metadata:
|
||||
name: refresh-deps-cache
|
||||
namespace: foxhunt
|
||||
labels:
|
||||
app.kubernetes.io/name: refresh-deps-cache
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
spec:
|
||||
activeDeadlineSeconds: 7200 # 2 hours; full workspace cargo build is slow
|
||||
serviceAccountName: argo-workflow
|
||||
entrypoint: build
|
||||
podMetadata:
|
||||
labels:
|
||||
# Reuse ci-pipeline label so the pod inherits argo-ci-pipeline netpol egress
|
||||
# clone, gitlab-registry:5000 for image push, crates.io HTTPS for cargo
|
||||
# NetworkPolicy clone — same egress targets, no functional difference.
|
||||
app.kubernetes.io/component: ci-pipeline
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
ttlStrategy:
|
||||
secondsAfterCompletion: 7200
|
||||
arguments:
|
||||
parameters:
|
||||
- name: commit-sha
|
||||
value: HEAD
|
||||
- name: deps-version
|
||||
value: "1"
|
||||
- name: image-tag
|
||||
value: nightly
|
||||
|
||||
templates:
|
||||
- name: build
|
||||
inputs:
|
||||
parameters:
|
||||
- name: commit-sha
|
||||
value: "{{workflow.parameters.commit-sha}}"
|
||||
- name: deps-version
|
||||
value: "{{workflow.parameters.deps-version}}"
|
||||
- name: image-tag
|
||||
value: "{{workflow.parameters.image-tag}}"
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: ci-compile-cpu
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
tolerations:
|
||||
- key: node.cilium.io/agent-not-ready
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
volumes:
|
||||
- name: workspace
|
||||
emptyDir:
|
||||
sizeLimit: 30Gi # full workspace + target/release ~10GB
|
||||
- name: git-ssh-key
|
||||
secret:
|
||||
secretName: argo-git-ssh-key
|
||||
defaultMode: 256
|
||||
- name: registry-auth
|
||||
secret:
|
||||
secretName: gitlab-registry
|
||||
items:
|
||||
- key: .dockerconfigjson
|
||||
path: config.json
|
||||
initContainers:
|
||||
- name: git-clone
|
||||
image: alpine/git:latest
|
||||
command: ["/bin/sh", "-c"]
|
||||
resources:
|
||||
requests:
|
||||
cpu: 200m
|
||||
memory: 256Mi
|
||||
limits:
|
||||
cpu: "1"
|
||||
memory: 512Mi
|
||||
volumeMounts:
|
||||
- name: git-ssh-key
|
||||
mountPath: /etc/git-ssh
|
||||
readOnly: true
|
||||
- name: workspace
|
||||
mountPath: /workspace
|
||||
args:
|
||||
- |
|
||||
set -ex
|
||||
mkdir -p /root/.ssh
|
||||
cp /etc/git-ssh/ssh-privatekey /root/.ssh/id_ed25519
|
||||
chmod 600 /root/.ssh/id_ed25519
|
||||
printf 'Host *\n StrictHostKeyChecking no\n UserKnownHostsFile /dev/null\n' > /root/.ssh/config
|
||||
chmod 600 /root/.ssh/config
|
||||
|
||||
SHA="{{inputs.parameters.commit-sha}}"
|
||||
REPO="ssh://git@gitlab-gitlab-shell.foxhunt.svc.cluster.local:2222/root/foxhunt.git"
|
||||
git clone --no-checkout --filter=blob:none "$REPO" /workspace/src
|
||||
cd /workspace/src
|
||||
git checkout "$SHA"
|
||||
echo "Checked out $(git rev-parse --short HEAD)"
|
||||
container:
|
||||
image: gcr.io/kaniko-project/executor:debug
|
||||
command: ["/busybox/sh", "-c"]
|
||||
env:
|
||||
- name: DOCKER_CONFIG
|
||||
value: /kaniko/.docker
|
||||
# Larger than build-ci-image because the inner cargo build is heavy.
|
||||
resources:
|
||||
requests:
|
||||
cpu: "8"
|
||||
memory: 16Gi
|
||||
limits:
|
||||
cpu: "16"
|
||||
memory: 32Gi
|
||||
volumeMounts:
|
||||
- name: registry-auth
|
||||
mountPath: /kaniko/.docker
|
||||
readOnly: true
|
||||
- name: workspace
|
||||
mountPath: /workspace
|
||||
args:
|
||||
- |
|
||||
/kaniko/executor \
|
||||
--context=/workspace/src \
|
||||
--dockerfile=/workspace/src/infra/docker/Dockerfile.ci-deps-cache \
|
||||
--destination=gitlab-registry.foxhunt.svc.cluster.local:5000/root/foxhunt/ci-builder-cpu-with-deps:{{inputs.parameters.image-tag}} \
|
||||
--build-arg=DEPS_VERSION={{inputs.parameters.deps-version}} \
|
||||
--build-arg=BASE_TAG=latest \
|
||||
--insecure-registry=gitlab-registry.foxhunt.svc.cluster.local:5000 \
|
||||
--skip-tls-verify-registry=gitlab-registry.foxhunt.svc.cluster.local:5000 \
|
||||
--cache=true \
|
||||
--cache-repo=gitlab-registry.foxhunt.svc.cluster.local:5000/root/foxhunt/cache \
|
||||
--snapshot-mode=redo
|
||||
|
||||
---
|
||||
# Nightly cron — rebuilds the deps cache image at 03:00 UTC.
|
||||
# Suspended by default; enable with:
|
||||
# kubectl -n foxhunt patch cronworkflow refresh-deps-cache-nightly \
|
||||
# -p '{"spec":{"suspend":false}}' --type=merge
|
||||
apiVersion: argoproj.io/v1alpha1
|
||||
kind: CronWorkflow
|
||||
metadata:
|
||||
name: refresh-deps-cache-nightly
|
||||
namespace: foxhunt
|
||||
labels:
|
||||
app.kubernetes.io/name: refresh-deps-cache-nightly
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
spec:
|
||||
schedules:
|
||||
- "0 3 * * *"
|
||||
timezone: UTC
|
||||
concurrencyPolicy: Forbid
|
||||
successfulJobsHistoryLimit: 2
|
||||
failedJobsHistoryLimit: 3
|
||||
suspend: true
|
||||
workflowSpec:
|
||||
entrypoint: trigger
|
||||
serviceAccountName: argo-workflow
|
||||
ttlStrategy:
|
||||
secondsAfterCompletion: 7200
|
||||
templates:
|
||||
- name: trigger
|
||||
steps:
|
||||
- - name: refresh
|
||||
templateRef:
|
||||
name: refresh-deps-cache
|
||||
template: build
|
||||
arguments:
|
||||
parameters:
|
||||
- name: commit-sha
|
||||
value: HEAD
|
||||
- name: deps-version
|
||||
value: "1"
|
||||
- name: image-tag
|
||||
value: nightly
|
||||
@@ -1,311 +0,0 @@
|
||||
# infra/k8s/argo/sanitizer-test-template.yaml
|
||||
#
|
||||
# One-shot compute-sanitizer run on L40S. Wraps a smoke test under
|
||||
# `compute-sanitizer --tool memcheck` to validate no CUDA memory errors
|
||||
# (out-of-bounds, leaks, sync violations, race conditions).
|
||||
#
|
||||
# Why L40S not local: the laptop RTX 3050 Ti (4 GB VRAM) cannot fit
|
||||
# compute-sanitizer's instrumentation metadata alongside the production
|
||||
# training workload — the sanitizer falls back to "didn't track the
|
||||
# launch" with 60k+ internal-allocation errors. L40S (48 GB) has ample
|
||||
# headroom for memcheck's 2-3× shadow-memory overhead.
|
||||
#
|
||||
# Usage:
|
||||
# argo submit --watch -n foxhunt sanitizer-test-template.yaml \
|
||||
# -p commit-ref=main \
|
||||
# -p test-name=iqn_quantile_monotonicity::iqn_multi_quantile_heads_produce_monotonic_estimates \
|
||||
# -p sanitizer-tool=memcheck
|
||||
#
|
||||
# Or via the wrapper script: scripts/argo-sanitizer.sh
|
||||
|
||||
apiVersion: argoproj.io/v1alpha1
|
||||
kind: WorkflowTemplate
|
||||
metadata:
|
||||
name: sanitizer-test
|
||||
namespace: foxhunt
|
||||
labels:
|
||||
app.kubernetes.io/name: sanitizer-test
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
spec:
|
||||
entrypoint: sanitizer-run
|
||||
serviceAccountName: argo-workflow
|
||||
podMetadata:
|
||||
labels:
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
app.kubernetes.io/component: sanitizer-test
|
||||
archiveLogs: true
|
||||
podGC:
|
||||
strategy: OnPodCompletion
|
||||
ttlStrategy:
|
||||
secondsAfterCompletion: 7200
|
||||
activeDeadlineSeconds: 7200 # 2h cap — sanitizer is slow but bounded
|
||||
arguments:
|
||||
parameters:
|
||||
- name: commit-ref
|
||||
value: HEAD
|
||||
- name: gpu-pool
|
||||
value: ci-training-l40s
|
||||
- name: cuda-compute-cap
|
||||
value: "89" # L40S Ada Lovelace
|
||||
- name: test-name
|
||||
value: "iqn_quantile_monotonicity::iqn_multi_quantile_heads_produce_monotonic_estimates"
|
||||
- name: sanitizer-tool
|
||||
value: memcheck # memcheck | racecheck | synccheck | initcheck
|
||||
- name: sanitizer-extra-args
|
||||
value: ""
|
||||
|
||||
volumes:
|
||||
- name: git-ssh-key
|
||||
secret:
|
||||
secretName: argo-git-ssh-key
|
||||
defaultMode: 256
|
||||
- name: cargo-target
|
||||
persistentVolumeClaim:
|
||||
claimName: cargo-target-cuda-test
|
||||
# Use the same training-data-pvc as production training (full 27 months
|
||||
# of OHLCV + MBP-10 + trades) instead of the smaller test-data-pvc which
|
||||
# only has 3-4 months of MBP-10 — walk-forward needs ≥10 months.
|
||||
- name: training-data
|
||||
persistentVolumeClaim:
|
||||
claimName: training-data-pvc
|
||||
readOnly: true
|
||||
|
||||
templates:
|
||||
- name: sanitizer-run
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: "{{workflow.parameters.gpu-pool}}"
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
tolerations:
|
||||
- key: nvidia.com/gpu
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
- key: node.cilium.io/agent-not-ready
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
container:
|
||||
image: gitlab-registry.foxhunt.svc.cluster.local:5000/root/foxhunt/ci-builder:latest
|
||||
command: ["/bin/bash", "-c"]
|
||||
env:
|
||||
- name: SQLX_OFFLINE
|
||||
value: "true"
|
||||
- name: CARGO_TERM_COLOR
|
||||
value: always
|
||||
- name: CARGO_TARGET_DIR
|
||||
value: /cargo-target
|
||||
- name: CARGO_HOME
|
||||
value: /cargo-target/cargo-home
|
||||
- name: CUDA_VISIBLE_DEVICES
|
||||
value: "0"
|
||||
- name: CUBLAS_WORKSPACE_CONFIG
|
||||
value: ":4096:8"
|
||||
- name: FOXHUNT_TEST_DATA
|
||||
value: /data/futures-baseline
|
||||
- name: TEST_DATA_DIR
|
||||
value: /data/futures-baseline
|
||||
# MBP-10 + trades are mandatory per feedback_mbp10_mandatory.md
|
||||
# (OFI features are part of the production state vector).
|
||||
- name: FOXHUNT_MBP10_DATA
|
||||
value: /data/futures-baseline-mbp10
|
||||
- name: FOXHUNT_TRADES_DATA
|
||||
value: /data/futures-baseline-trades
|
||||
- name: RUST_LOG
|
||||
value: info
|
||||
- name: LD_LIBRARY_PATH
|
||||
value: /usr/local/nvidia/lib64:/usr/local/cuda/lib64:/usr/local/cuda/extras/CUPTI/lib64
|
||||
volumeMounts:
|
||||
- name: git-ssh-key
|
||||
mountPath: /etc/git-ssh
|
||||
readOnly: true
|
||||
- name: cargo-target
|
||||
mountPath: /cargo-target
|
||||
- name: training-data
|
||||
mountPath: /data
|
||||
readOnly: true
|
||||
resources:
|
||||
requests:
|
||||
nvidia.com/gpu: "1"
|
||||
cpu: "4"
|
||||
memory: 32Gi
|
||||
limits:
|
||||
nvidia.com/gpu: "1"
|
||||
cpu: "8"
|
||||
memory: 64Gi
|
||||
args:
|
||||
- |
|
||||
set -euo pipefail
|
||||
|
||||
REF="{{workflow.parameters.commit-ref}}"
|
||||
TEST_NAME="{{workflow.parameters.test-name}}"
|
||||
SANITIZER_TOOL="{{workflow.parameters.sanitizer-tool}}"
|
||||
EXTRA_ARGS="{{workflow.parameters.sanitizer-extra-args}}"
|
||||
|
||||
echo "==================================="
|
||||
echo " compute-sanitizer L40S run"
|
||||
echo "==================================="
|
||||
echo " Ref: $REF"
|
||||
echo " Test: $TEST_NAME"
|
||||
echo " Tool: $SANITIZER_TOOL"
|
||||
echo " Pool: {{workflow.parameters.gpu-pool}}"
|
||||
echo " Compute: {{workflow.parameters.cuda-compute-cap}}"
|
||||
echo "==================================="
|
||||
|
||||
# --- SSH setup (same as gpu-test-pipeline) ---
|
||||
mkdir -p ~/.ssh
|
||||
cp /etc/git-ssh/ssh-privatekey ~/.ssh/id_ed25519
|
||||
chmod 600 ~/.ssh/id_ed25519
|
||||
printf 'Host *\n StrictHostKeyChecking no\n UserKnownHostsFile /dev/null\n' > ~/.ssh/config
|
||||
chmod 600 ~/.ssh/config
|
||||
|
||||
REPO="ssh://git@gitlab-gitlab-shell.foxhunt.svc.cluster.local:2222/root/foxhunt.git"
|
||||
BUILD="/cargo-target/src"
|
||||
git config --global --add safe.directory "$BUILD"
|
||||
|
||||
# --- Persistent checkout on PVC ---
|
||||
if [ -d "$BUILD/.git" ]; then
|
||||
cd "$BUILD"
|
||||
git fetch origin
|
||||
if git rev-parse --verify "origin/$REF" >/dev/null 2>&1; then
|
||||
TARGET=$(git rev-parse "origin/$REF")
|
||||
else
|
||||
TARGET=$(git rev-parse --verify "$REF" 2>/dev/null || echo "$REF")
|
||||
fi
|
||||
CURRENT=$(git rev-parse HEAD 2>/dev/null || echo "none")
|
||||
if [ "$CURRENT" != "$TARGET" ]; then
|
||||
echo "=== Updating checkout to $REF ($TARGET) ==="
|
||||
git checkout --force --detach "$TARGET"
|
||||
git clean -fd
|
||||
fi
|
||||
else
|
||||
echo "=== Initial clone ==="
|
||||
git clone --filter=blob:none "$REPO" "$BUILD"
|
||||
cd "$BUILD"
|
||||
git fetch origin
|
||||
if git rev-parse --verify "origin/$REF" >/dev/null 2>&1; then
|
||||
git checkout --force --detach "origin/$REF"
|
||||
else
|
||||
git checkout "$REF"
|
||||
fi
|
||||
fi
|
||||
echo "Checked out $(git rev-parse --short HEAD)"
|
||||
|
||||
export PATH="${CARGO_HOME}/bin:/usr/local/cuda/bin:${PATH}"
|
||||
export CUDA_COMPUTE_CAP={{workflow.parameters.cuda-compute-cap}}
|
||||
|
||||
# --- PTX cache invalidation ---
|
||||
bash scripts/ptx-cache-invalidate.sh "${CARGO_TARGET_DIR}/.ptx_cache"
|
||||
|
||||
# --- Verify compute-sanitizer is available ---
|
||||
if ! command -v compute-sanitizer >/dev/null 2>&1; then
|
||||
echo "ERROR: compute-sanitizer not found in PATH. CUDA toolkit incomplete."
|
||||
exit 2
|
||||
fi
|
||||
echo "=== compute-sanitizer version ==="
|
||||
compute-sanitizer --version | head -3
|
||||
|
||||
# --- Compile test binary + train_baseline_rl example ---
|
||||
# train_baseline_rl is needed for multi_fold_convergence test which
|
||||
# spawns it as a subprocess via `cargo run --example`. Pre-building
|
||||
# avoids cargo doing it inside the sanitizer-instrumented run.
|
||||
echo "=== Compiling test binary + train_baseline_rl (--release) ==="
|
||||
cargo test -p ml --release --lib --no-run 2>&1 | tee /cargo-target/sanitizer-compile.log
|
||||
cargo build -p ml --release --example train_baseline_rl 2>&1 | tee -a /cargo-target/sanitizer-compile.log
|
||||
|
||||
# Locate the freshest ml-* test binary
|
||||
# Honour CARGO_TARGET_DIR — outputs go to ${CARGO_TARGET_DIR}/release/deps,
|
||||
# NOT the repo-relative target/. (CI workflow sets CARGO_TARGET_DIR=/cargo-target.)
|
||||
# Subshell + || true wraps the head -1 SIGPIPE so pipefail doesn't fire.
|
||||
TARGET_DIR="${CARGO_TARGET_DIR:-target}"
|
||||
TEST_BIN=$( (ls -t "$TARGET_DIR/release/deps/ml-"* 2>/dev/null \
|
||||
| grep -vE '\.(d|rcgu|rmeta|rlib|so|json)$' \
|
||||
| head -1) || true )
|
||||
if [ -z "$TEST_BIN" ] || [ ! -x "$TEST_BIN" ]; then
|
||||
echo "ERROR: could not locate compiled test binary in $TARGET_DIR/release/deps/"
|
||||
ls -la "$TARGET_DIR/release/deps/" 2>&1 | head -20
|
||||
exit 3
|
||||
fi
|
||||
echo "=== Test binary: $TEST_BIN ==="
|
||||
|
||||
# --- Verify the test exists in the binary ---
|
||||
# grep -q + pipefail = SIGPIPE on early exit. Capture once, grep
|
||||
# without -q so the consumer reads all input. Same below for nsys.
|
||||
TEST_LIST=$("$TEST_BIN" --list 2>&1 || true)
|
||||
if ! echo "$TEST_LIST" | grep -F "$TEST_NAME" > /dev/null; then
|
||||
echo "ERROR: test '$TEST_NAME' not found in binary."
|
||||
echo "Available smoke tests:"
|
||||
echo "$TEST_LIST" | grep smoke_tests | (head -30 || true)
|
||||
exit 4
|
||||
fi
|
||||
|
||||
# --- Run under compute-sanitizer ---
|
||||
mkdir -p /tmp/sanitizer-output
|
||||
LOG_FILE=/tmp/sanitizer-output/${SANITIZER_TOOL}.log
|
||||
|
||||
echo ""
|
||||
echo "==================================="
|
||||
echo " Launching compute-sanitizer..."
|
||||
echo "==================================="
|
||||
START_TS=$(date +%s)
|
||||
# --target-processes all so multi_fold_convergence's spawned
|
||||
# `train_baseline_rl` subprocess is also instrumented (the test
|
||||
# itself does no GPU work — the child binary does).
|
||||
set +e
|
||||
compute-sanitizer \
|
||||
--tool "$SANITIZER_TOOL" \
|
||||
--target-processes all \
|
||||
--launch-timeout 1200 \
|
||||
--error-exitcode 99 \
|
||||
--print-limit 200 \
|
||||
--log-file "$LOG_FILE" \
|
||||
${EXTRA_ARGS} \
|
||||
-- "$TEST_BIN" "$TEST_NAME" --ignored --nocapture
|
||||
EXIT_CODE=$?
|
||||
set -e
|
||||
END_TS=$(date +%s)
|
||||
ELAPSED=$((END_TS - START_TS))
|
||||
|
||||
echo ""
|
||||
echo "==================================="
|
||||
echo " Sanitizer run complete"
|
||||
echo "==================================="
|
||||
echo " Exit code: $EXIT_CODE"
|
||||
echo " Elapsed: ${ELAPSED}s"
|
||||
echo " Log: $LOG_FILE ($(wc -l <"$LOG_FILE") lines)"
|
||||
echo ""
|
||||
|
||||
# --- Triage report ---
|
||||
echo "=== Sanitizer error summary ==="
|
||||
grep -E "ERROR SUMMARY|Internal Sanitizer Error" "$LOG_FILE" | tail -5 || echo " (no summary line)"
|
||||
echo ""
|
||||
|
||||
# Categorise errors. compute-sanitizer prints "ERROR SUMMARY: N errors" at the
|
||||
# end. "Internal Sanitizer Error" is sanitizer-internal (allocation failure,
|
||||
# tracking gap) — NOT a real bug in the application code.
|
||||
INTERNAL_ERRORS=$(grep -c "Internal Sanitizer Error" "$LOG_FILE" || true)
|
||||
REAL_ERROR_LINE=$(grep "ERROR SUMMARY:" "$LOG_FILE" | tail -1 || echo "")
|
||||
REAL_ERRORS=$(echo "$REAL_ERROR_LINE" | grep -oE "[0-9]+ errors" | head -1 | grep -oE "[0-9]+" || echo "0")
|
||||
|
||||
echo "=== Triage ==="
|
||||
echo " Internal Sanitizer Errors (instrumentation gaps): $INTERNAL_ERRORS"
|
||||
echo " Reported errors total: $REAL_ERRORS"
|
||||
echo " Real (non-internal) errors: $((REAL_ERRORS - INTERNAL_ERRORS))"
|
||||
echo ""
|
||||
|
||||
# Show first 50 real (non-internal) errors with context
|
||||
echo "=== First non-internal error excerpts ==="
|
||||
grep -vE "Internal Sanitizer Error|^=========\s*$" "$LOG_FILE" \
|
||||
| grep -E "^=========" \
|
||||
| head -50 || echo " (none)"
|
||||
|
||||
# Final pass/fail
|
||||
if [ "$EXIT_CODE" -eq 99 ]; then
|
||||
echo ""
|
||||
echo "FAIL: compute-sanitizer detected real memory errors (--error-exitcode 99 fired)."
|
||||
exit 1
|
||||
elif [ "$EXIT_CODE" -ne 0 ]; then
|
||||
echo ""
|
||||
echo "FAIL: test process exited with code $EXIT_CODE (test failure or sanitizer abort)."
|
||||
exit "$EXIT_CODE"
|
||||
else
|
||||
echo ""
|
||||
echo "PASS: zero real memory errors detected, test passed."
|
||||
fi
|
||||
@@ -1,32 +0,0 @@
|
||||
---
|
||||
apiVersion: v1
|
||||
kind: PersistentVolumeClaim
|
||||
metadata:
|
||||
name: sccache-cpu
|
||||
namespace: foxhunt
|
||||
labels:
|
||||
app.kubernetes.io/name: sccache
|
||||
app.kubernetes.io/component: ci-cache
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
spec:
|
||||
accessModes: [ReadWriteOnce]
|
||||
storageClassName: scw-bssd-retain
|
||||
resources:
|
||||
requests:
|
||||
storage: 20Gi
|
||||
---
|
||||
apiVersion: v1
|
||||
kind: PersistentVolumeClaim
|
||||
metadata:
|
||||
name: sccache-cuda
|
||||
namespace: foxhunt
|
||||
labels:
|
||||
app.kubernetes.io/name: sccache
|
||||
app.kubernetes.io/component: ci-cache
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
spec:
|
||||
accessModes: [ReadWriteOnce]
|
||||
storageClassName: scw-bssd-retain
|
||||
resources:
|
||||
requests:
|
||||
storage: 20Gi
|
||||
@@ -1,308 +0,0 @@
|
||||
# infra/k8s/argo/smoke-test-template.yaml
|
||||
#
|
||||
# Plain smoke-test runner on L40S — no compute-sanitizer, no nsys profile,
|
||||
# no MinIO artefact upload. Mirrors the compile + checkout + GPU + data-mount
|
||||
# layout of nsys-test-template.yaml but executes the test binary directly.
|
||||
#
|
||||
# Use this for fast multi-fold / single-epoch validation runs against the
|
||||
# full training-data PVC (27 months) when local laptop data is too short
|
||||
# (laptop only has 1 quarter of MBP-10/trades).
|
||||
#
|
||||
# Why a separate template: the nsys/sanitizer wrappers add overhead and
|
||||
# extract artefacts not needed for a straight "did the test pass on the
|
||||
# real dataset" gate. Keeping the smoke variant lean keeps the iteration
|
||||
# cycle short (no .nsys-rep upload, no sanitizer instrumentation slowdown).
|
||||
#
|
||||
# Usage:
|
||||
# argo submit --watch -n foxhunt --from=wftmpl/smoke-test \
|
||||
# -p commit-ref=main \
|
||||
# -p test-name=multi_fold_convergence::test_multi_fold_convergence
|
||||
#
|
||||
# Or via the wrapper script: scripts/argo-smoke.sh
|
||||
|
||||
apiVersion: argoproj.io/v1alpha1
|
||||
kind: WorkflowTemplate
|
||||
metadata:
|
||||
name: smoke-test
|
||||
namespace: foxhunt
|
||||
labels:
|
||||
app.kubernetes.io/name: smoke-test
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
spec:
|
||||
entrypoint: smoke-run
|
||||
serviceAccountName: argo-workflow
|
||||
podMetadata:
|
||||
labels:
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
app.kubernetes.io/component: smoke-test
|
||||
archiveLogs: true
|
||||
podGC:
|
||||
strategy: OnPodCompletion
|
||||
ttlStrategy:
|
||||
secondsAfterCompletion: 7200
|
||||
activeDeadlineSeconds: 3600
|
||||
arguments:
|
||||
parameters:
|
||||
- name: commit-ref
|
||||
value: HEAD
|
||||
- name: gpu-pool
|
||||
value: ci-training-l40s
|
||||
- name: cuda-compute-cap
|
||||
value: "89"
|
||||
- name: test-name
|
||||
value: "multi_fold_convergence::test_multi_fold_convergence"
|
||||
# Wipe `/cargo-target/release` for the ml + ml-dqn crates before
|
||||
# compile so the smoke binary is built from a known-clean state. The
|
||||
# PVC's persistent target dir accumulates rmeta/object artefacts
|
||||
# across probes; file deletions (e.g., `regime_conditional.rs` in
|
||||
# ff00af68a) can leave dangling references that perturb downstream
|
||||
# codegen between bisect runs. Defaults to "false" — only set when
|
||||
# bisecting suspected build-cache contamination.
|
||||
- name: clean-cache
|
||||
value: "false"
|
||||
|
||||
volumes:
|
||||
- name: git-ssh-key
|
||||
secret:
|
||||
secretName: argo-git-ssh-key
|
||||
defaultMode: 256
|
||||
- name: cargo-target
|
||||
persistentVolumeClaim:
|
||||
claimName: cargo-target-cuda-test
|
||||
# Same training-data-pvc as production training (full 27 months of
|
||||
# OHLCV + MBP-10 + trades) — laptop's 1-quarter MBP-10 truncates the
|
||||
# fxcache below the smoke's 10-month minimum.
|
||||
#
|
||||
# RW (not readOnly) so the smoke can populate `/data/bin/$SHORT_SHA/`
|
||||
# with the train binaries it just built, letting the next train run's
|
||||
# `ensure-binary` cache check (train-multi-seed-template.yaml:235-246)
|
||||
# hit and skip the 6-min compile. Read-side (fxcache + raw market data)
|
||||
# is unchanged.
|
||||
- name: training-data
|
||||
persistentVolumeClaim:
|
||||
claimName: training-data-pvc
|
||||
|
||||
templates:
|
||||
- name: smoke-run
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: "{{workflow.parameters.gpu-pool}}"
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
tolerations:
|
||||
- key: nvidia.com/gpu
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
- key: node.cilium.io/agent-not-ready
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
container:
|
||||
image: gitlab-registry.foxhunt.svc.cluster.local:5000/root/foxhunt/ci-builder:latest
|
||||
command: ["/bin/bash", "-c"]
|
||||
env:
|
||||
- name: SQLX_OFFLINE
|
||||
value: "true"
|
||||
- name: CARGO_TERM_COLOR
|
||||
value: always
|
||||
- name: CARGO_TARGET_DIR
|
||||
value: /cargo-target
|
||||
- name: CARGO_HOME
|
||||
value: /cargo-target/cargo-home
|
||||
- name: CUDA_VISIBLE_DEVICES
|
||||
value: "0"
|
||||
- name: CUBLAS_WORKSPACE_CONFIG
|
||||
value: ":4096:8"
|
||||
- name: FOXHUNT_TEST_DATA
|
||||
value: /data/futures-baseline
|
||||
- name: TEST_DATA_DIR
|
||||
value: /data/futures-baseline
|
||||
# MBP-10 + trades are mandatory per feedback_mbp10_mandatory.md
|
||||
# (OFI features are part of the production state vector).
|
||||
- name: FOXHUNT_MBP10_DATA
|
||||
value: /data/futures-baseline-mbp10
|
||||
- name: FOXHUNT_TRADES_DATA
|
||||
value: /data/futures-baseline-trades
|
||||
- name: RUST_LOG
|
||||
value: info
|
||||
- name: LD_LIBRARY_PATH
|
||||
value: /usr/local/nvidia/lib64:/usr/local/cuda/lib64
|
||||
volumeMounts:
|
||||
- name: git-ssh-key
|
||||
mountPath: /etc/git-ssh
|
||||
readOnly: true
|
||||
- name: cargo-target
|
||||
mountPath: /cargo-target
|
||||
- name: training-data
|
||||
mountPath: /data
|
||||
# RW so the post-PASS cache-population step can write to
|
||||
# /data/bin/$SHORT_SHA/. Read-side paths (fxcache, MBP-10,
|
||||
# trades) are unaffected.
|
||||
resources:
|
||||
requests:
|
||||
nvidia.com/gpu: "1"
|
||||
cpu: "4"
|
||||
memory: 32Gi
|
||||
limits:
|
||||
nvidia.com/gpu: "1"
|
||||
cpu: "8"
|
||||
memory: 64Gi
|
||||
args:
|
||||
- |
|
||||
set -euo pipefail
|
||||
|
||||
REF="{{workflow.parameters.commit-ref}}"
|
||||
TEST_NAME="{{workflow.parameters.test-name}}"
|
||||
|
||||
echo "==================================="
|
||||
echo " Plain smoke L40S run"
|
||||
echo "==================================="
|
||||
echo " Ref: $REF"
|
||||
echo " Test: $TEST_NAME"
|
||||
echo " Pool: {{workflow.parameters.gpu-pool}}"
|
||||
echo "==================================="
|
||||
|
||||
# --- SSH setup ---
|
||||
mkdir -p ~/.ssh
|
||||
cp /etc/git-ssh/ssh-privatekey ~/.ssh/id_ed25519
|
||||
chmod 600 ~/.ssh/id_ed25519
|
||||
printf 'Host *\n StrictHostKeyChecking no\n UserKnownHostsFile /dev/null\n' > ~/.ssh/config
|
||||
chmod 600 ~/.ssh/config
|
||||
|
||||
REPO="ssh://git@gitlab-gitlab-shell.foxhunt.svc.cluster.local:2222/root/foxhunt.git"
|
||||
BUILD="/cargo-target/src"
|
||||
git config --global --add safe.directory "$BUILD"
|
||||
|
||||
# --- Persistent checkout on PVC ---
|
||||
if [ -d "$BUILD/.git" ]; then
|
||||
cd "$BUILD"
|
||||
git fetch origin
|
||||
if git rev-parse --verify "origin/$REF" >/dev/null 2>&1; then
|
||||
TARGET=$(git rev-parse "origin/$REF")
|
||||
else
|
||||
TARGET=$(git rev-parse --verify "$REF" 2>/dev/null || echo "$REF")
|
||||
fi
|
||||
CURRENT=$(git rev-parse HEAD 2>/dev/null || echo "none")
|
||||
if [ "$CURRENT" != "$TARGET" ]; then
|
||||
echo "=== Updating checkout to $REF ($TARGET) ==="
|
||||
git checkout --force --detach "$TARGET"
|
||||
git clean -fd
|
||||
fi
|
||||
else
|
||||
git clone --filter=blob:none "$REPO" "$BUILD"
|
||||
cd "$BUILD"
|
||||
git fetch origin
|
||||
if git rev-parse --verify "origin/$REF" >/dev/null 2>&1; then
|
||||
git checkout --force --detach "origin/$REF"
|
||||
else
|
||||
git checkout "$REF"
|
||||
fi
|
||||
fi
|
||||
SHA=$(git rev-parse --short HEAD)
|
||||
echo "Checked out $SHA"
|
||||
|
||||
export PATH="${CARGO_HOME}/bin:/usr/local/cuda/bin:${PATH}"
|
||||
export CUDA_COMPUTE_CAP={{workflow.parameters.cuda-compute-cap}}
|
||||
|
||||
bash scripts/ptx-cache-invalidate.sh "${CARGO_TARGET_DIR}/.ptx_cache"
|
||||
|
||||
# --- Optional clean: ml + ml-dqn before compile ---
|
||||
CLEAN_CACHE="{{workflow.parameters.clean-cache}}"
|
||||
if [ "$CLEAN_CACHE" = "true" ]; then
|
||||
echo "=== clean-cache=true — wiping ml + ml-dqn build artefacts ==="
|
||||
cargo clean -p ml --release 2>&1 | tee /cargo-target/smoke-clean.log
|
||||
cargo clean -p ml-dqn --release 2>&1 | tee -a /cargo-target/smoke-clean.log
|
||||
echo "=== Cleaned. Forced fresh compile of ml/ml-dqn (deps stay cached) ==="
|
||||
fi
|
||||
|
||||
# --- Compile test binary + the 3 train binaries ---
|
||||
#
|
||||
# The 3 examples mirror exactly what `ensure-binary` produces
|
||||
# (see train-multi-seed-template.yaml line ~268). Compiling
|
||||
# them here lets the post-PASS step populate the
|
||||
# `/data/bin/$SHORT_SHA/` cache so the next train run at the
|
||||
# same SHA skips its own ensure-binary compile.
|
||||
echo "=== Compiling (--release --no-run + train binaries) ==="
|
||||
cargo test -p ml --release --lib --no-run 2>&1 | tee /cargo-target/smoke-compile.log
|
||||
cargo build -p ml --release \
|
||||
--example train_baseline_rl \
|
||||
--example evaluate_baseline \
|
||||
--example precompute_features \
|
||||
2>&1 | tee -a /cargo-target/smoke-compile.log
|
||||
|
||||
# Honour CARGO_TARGET_DIR — outputs go to ${CARGO_TARGET_DIR}/release/deps,
|
||||
# NOT the repo-relative target/. Subshell + || true wraps the head -1
|
||||
# SIGPIPE so pipefail doesn't fire.
|
||||
TARGET_DIR="${CARGO_TARGET_DIR:-target}"
|
||||
TEST_BIN=$( (ls -t "$TARGET_DIR/release/deps/ml-"* 2>/dev/null \
|
||||
| grep -vE '\.(d|rcgu|rmeta|rlib|so|json)$' \
|
||||
| head -1) || true )
|
||||
if [ -z "$TEST_BIN" ] || [ ! -x "$TEST_BIN" ]; then
|
||||
echo "ERROR: could not locate compiled test binary in $TARGET_DIR/release/deps/"
|
||||
ls -la "$TARGET_DIR/release/deps/" 2>&1 | head -20
|
||||
exit 3
|
||||
fi
|
||||
echo "=== Test binary: $TEST_BIN ==="
|
||||
|
||||
# grep -q + pipefail = SIGPIPE on early exit. Capture once, grep
|
||||
# without -q so the consumer reads all input.
|
||||
TEST_LIST=$("$TEST_BIN" --list 2>&1 || true)
|
||||
if ! echo "$TEST_LIST" | grep -F "$TEST_NAME" > /dev/null; then
|
||||
echo "ERROR: test '$TEST_NAME' not found"
|
||||
echo "$TEST_LIST" | grep smoke_tests | (head -30 || true)
|
||||
exit 4
|
||||
fi
|
||||
|
||||
# --- Run the test ---
|
||||
echo ""
|
||||
echo "==================================="
|
||||
echo " Launching test..."
|
||||
echo "==================================="
|
||||
START_TS=$(date +%s)
|
||||
set +e
|
||||
"$TEST_BIN" "$TEST_NAME" --ignored --nocapture
|
||||
EXIT_CODE=$?
|
||||
set -e
|
||||
END_TS=$(date +%s)
|
||||
ELAPSED=$((END_TS - START_TS))
|
||||
|
||||
echo ""
|
||||
echo "==================================="
|
||||
echo " Smoke run complete"
|
||||
echo "==================================="
|
||||
echo " Exit code: $EXIT_CODE"
|
||||
echo " Elapsed: ${ELAPSED}s"
|
||||
echo " Commit: $SHA"
|
||||
echo " Test: $TEST_NAME"
|
||||
echo "==================================="
|
||||
|
||||
if [ "$EXIT_CODE" -ne 0 ]; then
|
||||
echo ""
|
||||
echo "FAIL: test exited with code $EXIT_CODE"
|
||||
exit "$EXIT_CODE"
|
||||
fi
|
||||
echo ""
|
||||
echo "PASS: smoke completed."
|
||||
|
||||
# --- Populate /data/bin/$SHORT_SHA/ for ensure-binary cache ---
|
||||
# On PASS only — broken binaries must not be cached. Mirrors
|
||||
# train-multi-seed-template.yaml:270-274 exactly: same
|
||||
# destination layout, same strip step. Idempotent: writes are
|
||||
# to a SHA-keyed dir so concurrent smokes at different SHAs
|
||||
# don't collide; same-SHA writes are deterministic compile
|
||||
# output and overwrite-safe.
|
||||
FULL_SHA=$(git rev-parse HEAD)
|
||||
SHORT_SHA=$(echo "$FULL_SHA" | cut -c1-9)
|
||||
BIN_DIR="/data/bin/$SHORT_SHA"
|
||||
BINARIES="train_baseline_rl evaluate_baseline precompute_features"
|
||||
echo ""
|
||||
echo "=== Populating ensure-binary cache at $BIN_DIR ==="
|
||||
mkdir -p "$BIN_DIR"
|
||||
for bin in $BINARIES; do
|
||||
src="${CARGO_TARGET_DIR}/release/examples/$bin"
|
||||
if [ ! -x "$src" ]; then
|
||||
echo "WARN: $src not found (cache pop skipped for $bin)"
|
||||
continue
|
||||
fi
|
||||
cp "$src" "$BIN_DIR/"
|
||||
strip "$BIN_DIR/$bin" 2>/dev/null || true
|
||||
done
|
||||
ls -lh "$BIN_DIR/"
|
||||
echo "=== Cache populated; next train@$SHORT_SHA hits ensure-binary cache ==="
|
||||
@@ -1,605 +0,0 @@
|
||||
# Multi-seed training workflow — Plan 5 Task 5 Phase B (one-job-per-seed).
|
||||
#
|
||||
# Renders an Argo DAG that fans out N seeds into N parallel `train-single` task
|
||||
# instances. Each task receives `seed` via inputs.parameters and runs a
|
||||
# walk-forward training that internally sweeps all K folds via the binary's
|
||||
# `--max-folds K` arg. Per-job runtime is K× longer than the original (seed,
|
||||
# fold) matrix but fanout drops from N*K to N — a better fit for the L40S pool
|
||||
# (5-GPU capacity vs 30 jobs queueing) and a simpler binary contract
|
||||
# (`train_baseline_rl` is a multi-fold walk-forward executor; it does NOT
|
||||
# accept `--fold K`).
|
||||
#
|
||||
# The `# __MATRIX_TASKS__` marker on the dag.tasks line is replaced by
|
||||
# scripts/argo-train.sh with N generated WorkflowTask stanzas before
|
||||
# submission. This avoids hand-writing the matrix and keeps the template
|
||||
# human-readable.
|
||||
#
|
||||
# Usage:
|
||||
# ./scripts/argo-train.sh dqn --multi-seed 5 --folds 6 --tag plan5-final
|
||||
#
|
||||
# DAG (per task):
|
||||
# ensure-binary ──┐
|
||||
# gpu-warmup ─────┼──> ensure-fxcache ──> [N parallel train-single tasks,
|
||||
# one per seed, each runs all K folds]
|
||||
# │
|
||||
# └──> aggregate (manual via
|
||||
# scripts/gather-multi-seed-metrics.sh)
|
||||
---
|
||||
apiVersion: argoproj.io/v1alpha1
|
||||
kind: WorkflowTemplate
|
||||
metadata:
|
||||
name: train-multi-seed
|
||||
namespace: foxhunt
|
||||
labels:
|
||||
app.kubernetes.io/name: train-multi-seed
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
app.kubernetes.io/component: train
|
||||
spec:
|
||||
entrypoint: multi-seed-matrix
|
||||
onExit: notify-result
|
||||
serviceAccountName: argo-workflow
|
||||
archiveLogs: true
|
||||
podMetadata:
|
||||
labels:
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
app.kubernetes.io/component: train
|
||||
securityContext:
|
||||
fsGroup: 0
|
||||
ttlStrategy:
|
||||
secondsAfterCompletion: 3600
|
||||
# Multi-seed runs: N jobs in parallel (one per seed), each running all K folds
|
||||
# in walk-forward sequence. Allow 12h walltime — 6-fold runs are ~6× longer
|
||||
# than the original per-(seed,fold) jobs but easily fit in 12h on L40S.
|
||||
activeDeadlineSeconds: 43200
|
||||
|
||||
arguments:
|
||||
parameters:
|
||||
- name: commit-sha
|
||||
value: HEAD
|
||||
- name: git-branch
|
||||
value: main
|
||||
# Default sm_89 / ci-training-l40s as of 2026-05-14 per
|
||||
# `feedback_default_to_l40s_pool.md`. argo-train.sh's
|
||||
# CUDA_COMPUTE_CAP derivation passes "89" by default to match;
|
||||
# explicit --gpu-pool ci-training-h100 overrides both back to
|
||||
# sm_90 + H100.
|
||||
- name: cuda-compute-cap
|
||||
value: "89"
|
||||
- name: model
|
||||
value: dqn
|
||||
- name: gpu-pool
|
||||
value: ci-training-l40s
|
||||
- name: hyperopt-trials
|
||||
value: "0" # Multi-seed runs typically skip hyperopt (re-use baseline params)
|
||||
- name: hyperopt-epochs
|
||||
value: "8"
|
||||
- name: train-epochs
|
||||
value: "50"
|
||||
- name: symbol
|
||||
value: ES.FUT
|
||||
- name: initial-capital
|
||||
value: "35000"
|
||||
- name: tx-cost-bps
|
||||
value: "0.1"
|
||||
- name: tick-size
|
||||
value: "0.25"
|
||||
- name: spread-ticks
|
||||
value: "1.0"
|
||||
- name: sanitizer
|
||||
value: "none"
|
||||
- name: multi-seed
|
||||
value: "5"
|
||||
- name: folds
|
||||
value: "6"
|
||||
# Plan 5 Task 3 (A.4.1): nsys profile harness toggle. When "true", each
|
||||
# train-single job wraps the training binary under `nsys profile` and
|
||||
# uploads the resulting .nsys-rep to MinIO bucket
|
||||
# foxhunt-training-artifacts/profiles/<short-sha>/. Default off.
|
||||
- name: profile
|
||||
value: "false"
|
||||
# Bar formation params — passed to BOTH precompute_features (when
|
||||
# building fxcache) AND train_baseline_rl (when looking up fxcache).
|
||||
# MUST match between the two for fxcache HIT (cache key includes them
|
||||
# as of 2026-05-09 architectural fix). Defaults match dqn-production.toml.
|
||||
#
|
||||
# imbalance-bar-threshold default: 20.0 (set 2026-05-10). At threshold=0.5
|
||||
# the sampler produced 209M bars from 209M trade ticks (1:1 ratio,
|
||||
# essentially per-tick) on workflow f5wnd, near-OOM in feature extraction
|
||||
# (54Gi/56Gi limit). To match the volume-bar density baseline (~5.74M
|
||||
# bars at 100 contracts), threshold ~20 is the right scale. Override via
|
||||
# argo-train.sh --imbalance-bar-threshold for resolution-sweep experiments.
|
||||
- name: imbalance-bar-threshold
|
||||
value: "20.0"
|
||||
- name: imbalance-bar-ewma-alpha
|
||||
value: "0.1"
|
||||
# Volume bar size (contracts/bar). Used when data-source != "mbp10".
|
||||
# Default 100 matches DEFAULT_VOLUME_BAR_SIZE. Cache key includes this.
|
||||
- name: volume-bar-size
|
||||
value: "100"
|
||||
# Data source mode: "mbp10" (imbalance bars from MBP-10) or "ohlcv"
|
||||
# (volume bars from trades). Threaded into both precompute_features and
|
||||
# the trainer so cache keys align.
|
||||
- name: data-source
|
||||
value: "mbp10"
|
||||
|
||||
volumes:
|
||||
- name: git-ssh-key
|
||||
secret:
|
||||
secretName: argo-git-ssh-key
|
||||
defaultMode: 256
|
||||
- name: training-data
|
||||
persistentVolumeClaim:
|
||||
claimName: training-data-pvc
|
||||
- name: cargo-target-cuda
|
||||
persistentVolumeClaim:
|
||||
claimName: cargo-target-cuda
|
||||
- name: feature-cache
|
||||
persistentVolumeClaim:
|
||||
claimName: feature-cache-pvc
|
||||
|
||||
volumeClaimTemplates:
|
||||
- metadata:
|
||||
name: workspace
|
||||
spec:
|
||||
accessModes: ["ReadWriteOnce"]
|
||||
storageClassName: scw-bssd
|
||||
resources:
|
||||
requests:
|
||||
storage: 5Gi
|
||||
|
||||
templates:
|
||||
# ── DAG: fan out to N train-single tasks (one per seed) ──
|
||||
# The `# __MATRIX_TASKS__` marker is replaced by argo-train.sh with the
|
||||
# generated per-seed WorkflowTask stanzas. Each task sweeps all K folds
|
||||
# via the binary's `--max-folds {{workflow.parameters.folds}}` argument.
|
||||
# The marker MUST stay on its own line for the awk substitution to work.
|
||||
- name: multi-seed-matrix
|
||||
dag:
|
||||
tasks:
|
||||
- name: ensure-binary
|
||||
template: ensure-binary
|
||||
- name: gpu-warmup
|
||||
template: gpu-warmup
|
||||
- name: ensure-fxcache
|
||||
template: ensure-fxcache
|
||||
dependencies: [ensure-binary]
|
||||
arguments:
|
||||
parameters:
|
||||
- name: sha
|
||||
value: "{{tasks.ensure-binary.outputs.parameters.sha}}"
|
||||
# __MATRIX_TASKS__
|
||||
|
||||
# ── ensure-binary: cache-or-compile training binaries by commit SHA ──
|
||||
# (Identical to train-template.yaml. Kept inline rather than templated
|
||||
# via wftmpl-cross-ref to avoid Argo's reluctance to chase template
|
||||
# references across WorkflowTemplates at submission time.)
|
||||
- name: ensure-binary
|
||||
outputs:
|
||||
parameters:
|
||||
- name: sha
|
||||
valueFrom:
|
||||
path: /tmp/sha
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: ci-compile-cpu
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
tolerations:
|
||||
- key: node.cilium.io/agent-not-ready
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
container:
|
||||
image: gitlab-registry.foxhunt.svc.cluster.local:5000/root/foxhunt/ci-builder:latest
|
||||
imagePullPolicy: Always
|
||||
command: ["/bin/bash", "-c"]
|
||||
env:
|
||||
- name: SQLX_OFFLINE
|
||||
value: "true"
|
||||
- name: CARGO_TERM_COLOR
|
||||
value: always
|
||||
- name: CARGO_TARGET_DIR
|
||||
value: /cargo-target
|
||||
- name: CARGO_HOME
|
||||
value: /cargo-target/cargo-home
|
||||
- name: CARGO_INCREMENTAL
|
||||
value: "1"
|
||||
- name: GITLAB_PAT
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: gitlab-pat
|
||||
key: token
|
||||
resources:
|
||||
requests:
|
||||
cpu: "14"
|
||||
memory: 32Gi
|
||||
limits:
|
||||
cpu: "30"
|
||||
memory: 64Gi
|
||||
volumeMounts:
|
||||
- name: git-ssh-key
|
||||
mountPath: /etc/git-ssh
|
||||
readOnly: true
|
||||
- name: cargo-target-cuda
|
||||
mountPath: /cargo-target
|
||||
- name: training-data
|
||||
mountPath: /data
|
||||
args:
|
||||
- |
|
||||
set -e
|
||||
SHA="{{workflow.parameters.commit-sha}}"
|
||||
BRANCH="{{workflow.parameters.git-branch}}"
|
||||
MODEL="{{workflow.parameters.model}}"
|
||||
|
||||
mkdir -p ~/.ssh
|
||||
cp /etc/git-ssh/ssh-privatekey ~/.ssh/id_ed25519
|
||||
chmod 600 ~/.ssh/id_ed25519
|
||||
printf 'Host *\n StrictHostKeyChecking no\n UserKnownHostsFile /dev/null\n' > ~/.ssh/config
|
||||
chmod 600 ~/.ssh/config
|
||||
|
||||
REPO="ssh://git@gitlab-gitlab-shell.foxhunt.svc.cluster.local:2222/root/foxhunt.git"
|
||||
BUILD="/cargo-target/src"
|
||||
git config --global --add safe.directory "$BUILD"
|
||||
|
||||
if [ "$SHA" = "HEAD" ]; then
|
||||
if [ -d "$BUILD/.git" ]; then
|
||||
cd "$BUILD"; git fetch origin
|
||||
SHA=$(git rev-parse "origin/$BRANCH"); cd /
|
||||
else
|
||||
git clone --filter=blob:none "$REPO" "$BUILD"
|
||||
cd "$BUILD"; git checkout "origin/$BRANCH"
|
||||
SHA=$(git rev-parse HEAD); cd /
|
||||
fi
|
||||
fi
|
||||
SHORT_SHA=$(echo "$SHA" | cut -c1-9)
|
||||
echo "Resolved SHA: $SHA (short: $SHORT_SHA)"
|
||||
|
||||
case "$MODEL" in
|
||||
alpha-rl) BINARIES="alpha_rl_train" ;;
|
||||
dqn|ppo) BINARIES="train_baseline_rl evaluate_baseline precompute_features" ;;
|
||||
*) BINARIES="train_baseline_supervised evaluate_supervised precompute_features" ;;
|
||||
esac
|
||||
|
||||
BIN_DIR="/data/bin/$SHORT_SHA"
|
||||
ALL_CACHED=true
|
||||
for bin in $BINARIES; do
|
||||
if [ ! -x "$BIN_DIR/$bin" ]; then ALL_CACHED=false; break; fi
|
||||
done
|
||||
|
||||
if [ "$ALL_CACHED" = "true" ]; then
|
||||
echo "=== Cache HIT: all binaries present in $BIN_DIR ==="
|
||||
ls -lh "$BIN_DIR/"
|
||||
echo "$SHORT_SHA" > /tmp/sha
|
||||
exit 0
|
||||
fi
|
||||
|
||||
echo "=== Cache MISS: compiling binaries for $SHORT_SHA ==="
|
||||
if [ -d "$BUILD/.git" ]; then
|
||||
cd "$BUILD"; git fetch origin
|
||||
CURRENT=$(git rev-parse HEAD 2>/dev/null || echo "none")
|
||||
if [ "$CURRENT" != "$SHA" ]; then
|
||||
git checkout --force "$SHA"; git clean -fd
|
||||
fi
|
||||
else
|
||||
git clone --filter=blob:none "$REPO" "$BUILD"
|
||||
cd "$BUILD"; git checkout "$SHA"
|
||||
fi
|
||||
|
||||
export PATH="${CARGO_HOME}/bin:${PATH}"
|
||||
export CUDA_COMPUTE_CAP={{workflow.parameters.cuda-compute-cap}}
|
||||
|
||||
if [ "$MODEL" = "alpha-rl" ]; then
|
||||
cargo build --release -p ml-alpha --example alpha_rl_train
|
||||
else
|
||||
ML_EXAMPLE_ARGS=""
|
||||
for ex in $BINARIES; do
|
||||
ML_EXAMPLE_ARGS="$ML_EXAMPLE_ARGS --example $ex"
|
||||
done
|
||||
cargo build --release -p ml --features ml/cuda $ML_EXAMPLE_ARGS
|
||||
fi
|
||||
|
||||
mkdir -p "$BIN_DIR"
|
||||
for bin in $BINARIES; do
|
||||
cp "$CARGO_TARGET_DIR/release/examples/$bin" "$BIN_DIR/"
|
||||
done
|
||||
strip "$BIN_DIR/"*
|
||||
echo "$SHORT_SHA" > /tmp/sha
|
||||
|
||||
# ── gpu-warmup: trigger GPU node autoscale during compile ──
|
||||
- name: gpu-warmup
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: "{{workflow.parameters.gpu-pool}}"
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
tolerations:
|
||||
- key: nvidia.com/gpu
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
- key: node.cilium.io/agent-not-ready
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
container:
|
||||
image: busybox:1.37
|
||||
command: ["/bin/sh", "-c"]
|
||||
args:
|
||||
- |
|
||||
echo "GPU warmup: triggering node autoscale..."
|
||||
resources:
|
||||
requests:
|
||||
nvidia.com/gpu: "1"
|
||||
cpu: 100m
|
||||
memory: 64Mi
|
||||
limits:
|
||||
nvidia.com/gpu: "1"
|
||||
cpu: 200m
|
||||
memory: 128Mi
|
||||
|
||||
# ── ensure-fxcache: precompute feature cache (shared across all seeds/folds) ──
|
||||
- name: ensure-fxcache
|
||||
inputs:
|
||||
parameters:
|
||||
- name: sha
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: ci-compile-cpu
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
tolerations:
|
||||
- key: node.cilium.io/agent-not-ready
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
container:
|
||||
image: ubuntu:24.04
|
||||
command: ["/bin/bash", "-c"]
|
||||
env:
|
||||
- name: RUST_LOG
|
||||
value: info
|
||||
resources:
|
||||
requests:
|
||||
cpu: "4"
|
||||
memory: 32Gi
|
||||
limits:
|
||||
cpu: "28"
|
||||
memory: 96Gi
|
||||
volumeMounts:
|
||||
- name: training-data
|
||||
mountPath: /data
|
||||
readOnly: false
|
||||
- name: feature-cache
|
||||
mountPath: /feature-cache
|
||||
args:
|
||||
- |
|
||||
set -e
|
||||
SHA="{{inputs.parameters.sha}}"
|
||||
BINARY="/data/bin/$SHA/precompute_features"
|
||||
if [ ! -x "$BINARY" ]; then
|
||||
echo "ERROR: precompute_features not found at $BINARY"; exit 1
|
||||
fi
|
||||
export RAYON_NUM_THREADS=20
|
||||
if $BINARY \
|
||||
--data-dir /data/futures-baseline \
|
||||
--mbp10-data-dir /data/futures-baseline-mbp10 \
|
||||
--trades-data-dir /data/futures-baseline-trades \
|
||||
--output-dir /feature-cache \
|
||||
--symbol {{workflow.parameters.symbol}} \
|
||||
--data-source {{workflow.parameters.data-source}} \
|
||||
--imbalance-bar-threshold {{workflow.parameters.imbalance-bar-threshold}} \
|
||||
--imbalance-bar-ewma-alpha {{workflow.parameters.imbalance-bar-ewma-alpha}} \
|
||||
--volume-bar-size {{workflow.parameters.volume-bar-size}} \
|
||||
--yes; then
|
||||
echo "=== Feature cache ready ==="
|
||||
else
|
||||
echo "=== Cache stale or missing — regenerating ==="
|
||||
# NOTE: don't blanket-rm /feature-cache/*.fxcache anymore. With
|
||||
# bar-params now in the cache key, multiple valid caches can
|
||||
# coexist (different threshold experiments). The early-exit
|
||||
# check above already validates the specific key — fall through
|
||||
# to regen only the one we need.
|
||||
$BINARY \
|
||||
--data-dir /data/futures-baseline \
|
||||
--mbp10-data-dir /data/futures-baseline-mbp10 \
|
||||
--trades-data-dir /data/futures-baseline-trades \
|
||||
--output-dir /feature-cache \
|
||||
--symbol {{workflow.parameters.symbol}} \
|
||||
--data-source {{workflow.parameters.data-source}} \
|
||||
--imbalance-bar-threshold {{workflow.parameters.imbalance-bar-threshold}} \
|
||||
--imbalance-bar-ewma-alpha {{workflow.parameters.imbalance-bar-ewma-alpha}} \
|
||||
--volume-bar-size {{workflow.parameters.volume-bar-size}} \
|
||||
--yes
|
||||
fi
|
||||
|
||||
# ── train-single: one seed, all folds (training instance) ──
|
||||
# Invoked once per matrix entry. Reads SEED from inputs.parameters and
|
||||
# forwards it via `--seed`; the binary's `--max-folds K` arg drives the
|
||||
# walk-forward sweep over all K folds inside this single process.
|
||||
# Plan 5 Task 5 Phase B pivot: was per-(seed, fold) on the failed deploy
|
||||
# because train_baseline_rl does not accept `--fold N` (it is a multi-fold
|
||||
# executor, not a single-fold one). One-job-per-seed matches the binary's
|
||||
# actual contract and the L40S pool's capacity.
|
||||
- name: train-single
|
||||
inputs:
|
||||
parameters:
|
||||
- name: seed
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: "{{workflow.parameters.gpu-pool}}"
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
tolerations:
|
||||
- key: nvidia.com/gpu
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
- key: node.cilium.io/agent-not-ready
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
container:
|
||||
image: gitlab-registry.foxhunt.svc.cluster.local:5000/root/foxhunt/ci-builder:latest
|
||||
imagePullPolicy: IfNotPresent
|
||||
command: ["/bin/sh", "-c"]
|
||||
env:
|
||||
- name: RUST_LOG
|
||||
value: info
|
||||
- name: SQLX_OFFLINE
|
||||
value: "true"
|
||||
- name: CUBLAS_WORKSPACE_CONFIG
|
||||
value: ":4096:8"
|
||||
- name: FOXHUNT_FEATURE_CACHE_DIR
|
||||
value: /feature-cache
|
||||
- name: SEED
|
||||
value: "{{inputs.parameters.seed}}"
|
||||
# Plan 5 Task 3 (A.4.1): MinIO creds for the optional `mc cp` of
|
||||
# the .nsys-rep artefact at the end of train-single. Both refs are
|
||||
# `optional: true` so the env mount succeeds on clusters that do
|
||||
# not pre-create `minio-credentials` (the upload then warn-fails
|
||||
# gracefully — training itself doesn't depend on these).
|
||||
- name: MINIO_ACCESS_KEY
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: minio-credentials
|
||||
key: access-key
|
||||
optional: true
|
||||
- name: MINIO_SECRET_KEY
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: minio-credentials
|
||||
key: secret-key
|
||||
optional: true
|
||||
resources:
|
||||
requests:
|
||||
nvidia.com/gpu: "1"
|
||||
cpu: "4"
|
||||
memory: 32Gi
|
||||
limits:
|
||||
nvidia.com/gpu: "1"
|
||||
cpu: "8"
|
||||
memory: 64Gi
|
||||
volumeMounts:
|
||||
- name: workspace
|
||||
mountPath: /workspace
|
||||
- name: training-data
|
||||
mountPath: /data
|
||||
readOnly: true
|
||||
- name: feature-cache
|
||||
mountPath: /feature-cache
|
||||
readOnly: true
|
||||
args:
|
||||
- |
|
||||
set -e
|
||||
# Resolve SHA via the binary cache directory layout the
|
||||
# ensure-binary task established. multi-seed-matrix passes the
|
||||
# SHA implicitly via the workflow-scoped /data/bin tree.
|
||||
SHA=$(ls -1t /data/bin | head -1)
|
||||
export PATH="/data/bin/$SHA:$PATH"
|
||||
export LD_LIBRARY_PATH=$(echo "$LD_LIBRARY_PATH" | tr ':' '\n' | grep -v stubs | tr '\n' ':' | sed 's/:$//')
|
||||
nvidia-smi
|
||||
|
||||
MODEL="{{workflow.parameters.model}}"
|
||||
case "$MODEL" in
|
||||
alpha-rl) BINARY=alpha_rl_train ;;
|
||||
dqn|ppo) BINARY=train_baseline_rl ;;
|
||||
*) BINARY=train_baseline_supervised ;;
|
||||
esac
|
||||
|
||||
mkdir -p /workspace/output
|
||||
|
||||
# Plan 5 Task 3 (A.4.1): optional nsys profile wrapper.
|
||||
# When workflow.parameters.profile=="true", wrap the training
|
||||
# binary under `nsys profile --capture-range=cudaProfilerApi`.
|
||||
# The .nsys-rep is uploaded to MinIO bucket
|
||||
# foxhunt-training-artifacts/profiles/<short-sha>/ in the
|
||||
# post-training upload block below.
|
||||
PROFILE="{{workflow.parameters.profile}}"
|
||||
POD_NAME="${HOSTNAME:-pod}"
|
||||
NSYS_OUT="/workspace/output/profile-${POD_NAME}.nsys-rep"
|
||||
NSYS_PREFIX=""
|
||||
if [ "$PROFILE" = "true" ]; then
|
||||
NSYS_BIN=$(which nsys 2>/dev/null || echo "/usr/local/cuda/bin/nsys")
|
||||
if [ -x "$NSYS_BIN" ]; then
|
||||
NSYS_PREFIX="$NSYS_BIN profile --capture-range=cudaProfilerApi --output=${NSYS_OUT} --force-overwrite=true"
|
||||
echo "=== nsys profile enabled: ${NSYS_OUT} ==="
|
||||
else
|
||||
echo "WARN: --profile requested but nsys not found at $NSYS_BIN — running without"
|
||||
fi
|
||||
fi
|
||||
|
||||
echo "=== Training: $MODEL seed=$SEED folds={{workflow.parameters.folds}} ==="
|
||||
if [ "$MODEL" = "alpha-rl" ]; then
|
||||
stdbuf -oL $NSYS_PREFIX ${BINARY} \
|
||||
--mbp10-data-dir /data/futures-baseline-mbp10/{{workflow.parameters.symbol}} \
|
||||
--predecoded-dir /feature-cache/predecoded \
|
||||
--out /workspace/output \
|
||||
--n-steps 50000 \
|
||||
--n-backtests 16 \
|
||||
--seed "$SEED" \
|
||||
--instrument-mode all
|
||||
else
|
||||
stdbuf -oL $NSYS_PREFIX ${BINARY} \
|
||||
--model "$MODEL" \
|
||||
--symbol {{workflow.parameters.symbol}} \
|
||||
--tx-cost-bps {{workflow.parameters.tx-cost-bps}} \
|
||||
--tick-size {{workflow.parameters.tick-size}} \
|
||||
--spread-ticks {{workflow.parameters.spread-ticks}} \
|
||||
--initial-capital {{workflow.parameters.initial-capital}} \
|
||||
--data-dir /data/futures-baseline \
|
||||
--mbp10-data-dir /data/futures-baseline-mbp10 \
|
||||
--trades-data-dir /data/futures-baseline-trades \
|
||||
--output-dir /workspace/output \
|
||||
--epochs {{workflow.parameters.train-epochs}} \
|
||||
--seed "$SEED" \
|
||||
--max-folds {{workflow.parameters.folds}} \
|
||||
--imbalance-bar-threshold {{workflow.parameters.imbalance-bar-threshold}} \
|
||||
--imbalance-bar-ewma-alpha {{workflow.parameters.imbalance-bar-ewma-alpha}} \
|
||||
--volume-bar-size {{workflow.parameters.volume-bar-size}}
|
||||
fi
|
||||
|
||||
echo "=== Training complete: seed=$SEED folds={{workflow.parameters.folds}} ==="
|
||||
|
||||
# Plan 5 Task 3 (A.4.1): upload .nsys-rep to MinIO if profile run.
|
||||
# mc is fetched on-demand (~25 MB single static binary) — the
|
||||
# ci-builder training image does not bundle it. Bucket:
|
||||
# foxhunt-training-artifacts/profiles/<short-sha>/.
|
||||
if [ "$PROFILE" = "true" ] && [ -f "$NSYS_OUT" ]; then
|
||||
echo "=== Uploading nsys profile to MinIO ==="
|
||||
MC_BIN=$(which mc 2>/dev/null || echo "")
|
||||
if [ -z "$MC_BIN" ]; then
|
||||
MC_BIN=/tmp/mc
|
||||
curl -fsSL https://dl.min.io/client/mc/release/linux-amd64/mc -o "$MC_BIN" || {
|
||||
echo "WARN: failed to download mc — skipping upload"
|
||||
exit 0
|
||||
}
|
||||
chmod +x "$MC_BIN"
|
||||
fi
|
||||
"$MC_BIN" alias set foxhunt http://minio.foxhunt.svc.cluster.local:9000 \
|
||||
"$MINIO_ACCESS_KEY" "$MINIO_SECRET_KEY" 2>/dev/null || true
|
||||
"$MC_BIN" cp "$NSYS_OUT" \
|
||||
"foxhunt/foxhunt-training-artifacts/profiles/$SHA/profile-seed${SEED}-${POD_NAME}.nsys-rep" || \
|
||||
echo "WARN: nsys upload failed for seed=$SEED"
|
||||
echo "=== nsys profile upload complete ==="
|
||||
fi
|
||||
|
||||
# ── notify-result: post workflow outcome to Mattermost (onExit) ──
|
||||
- name: notify-result
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: platform
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
container:
|
||||
image: curlimages/curl:8.12.1
|
||||
command: ["/bin/sh", "-c"]
|
||||
env:
|
||||
- name: WEBHOOK_URL
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: notification-webhook
|
||||
key: webhook-url
|
||||
optional: true
|
||||
resources:
|
||||
requests:
|
||||
cpu: 50m
|
||||
memory: 32Mi
|
||||
limits:
|
||||
cpu: 100m
|
||||
memory: 64Mi
|
||||
args:
|
||||
- |
|
||||
STATUS="{{workflow.status}}"
|
||||
NAME="{{workflow.name}}"
|
||||
if [ -z "$WEBHOOK_URL" ] || echo "$WEBHOOK_URL" | grep -q "PLACEHOLDER"; then
|
||||
echo "No webhook configured, skipping notification"; exit 0
|
||||
fi
|
||||
EMOJI=":x:"
|
||||
[ "$STATUS" = "Succeeded" ] && EMOJI=":white_check_mark:"
|
||||
PAYLOAD="{\"username\":\"Argo CI\",\"text\":\"${EMOJI} **${NAME}** (multi-seed) — ${STATUS}\"}"
|
||||
curl -sf -X POST -H 'Content-Type: application/json' \
|
||||
-d "$PAYLOAD" "$WEBHOOK_URL" || echo "WARN: webhook post failed"
|
||||
@@ -1,838 +0,0 @@
|
||||
# Unified training workflow — replaces all model-specific training templates.
|
||||
#
|
||||
# Key innovation: binary caching by commit SHA on PVC.
|
||||
# Check /data/bin/$SHA first; compile only on cache miss.
|
||||
#
|
||||
# DAG:
|
||||
# ensure-binary ──┐
|
||||
# gpu-warmup ─────┼──> ensure-fxcache ──> hyperopt ──> train-best ──> evaluate ──> upload-results
|
||||
#
|
||||
# When hyperopt-trials=0 (the `scripts/argo-precompute.sh` path),
|
||||
# gpu-warmup / hyperopt / train-best / evaluate all skip and the
|
||||
# workflow runs ensure-binary → ensure-fxcache only — no GPU node
|
||||
# gets provisioned for nothing.
|
||||
#
|
||||
# Usage:
|
||||
# argo submit -n foxhunt --from=wftmpl/train
|
||||
# argo submit -n foxhunt --from=wftmpl/train -p model=ppo -p train-epochs=100
|
||||
# argo submit -n foxhunt --from=wftmpl/train -p hyperopt-trials=0 # fxcache-only path
|
||||
apiVersion: argoproj.io/v1alpha1
|
||||
kind: WorkflowTemplate
|
||||
metadata:
|
||||
name: train
|
||||
namespace: foxhunt
|
||||
labels:
|
||||
app.kubernetes.io/name: train
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
app.kubernetes.io/component: train
|
||||
spec:
|
||||
entrypoint: pipeline
|
||||
onExit: notify-result
|
||||
serviceAccountName: argo-workflow
|
||||
archiveLogs: true
|
||||
podMetadata:
|
||||
labels:
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
app.kubernetes.io/component: train
|
||||
securityContext:
|
||||
fsGroup: 0
|
||||
ttlStrategy:
|
||||
secondsAfterCompletion: 3600
|
||||
activeDeadlineSeconds: 28800 # 8 hours
|
||||
|
||||
arguments:
|
||||
parameters:
|
||||
- name: commit-sha
|
||||
value: HEAD
|
||||
- name: git-branch
|
||||
value: main
|
||||
- name: cuda-compute-cap
|
||||
value: "90"
|
||||
- name: model
|
||||
value: dqn
|
||||
# gpu-pool default: ci-training-l40s (set 2026-05-14 per
|
||||
# `feedback_default_to_l40s_pool.md` — SP-chain training has been
|
||||
# standardising on L40S since 2026-05-09; the prior ci-training-h100
|
||||
# default required every SP-run invocation to pass an explicit
|
||||
# `--gpu-pool ci-training-l40s` override. H100 remains opt-in via
|
||||
# `--gpu-pool ci-training-h100` for runs that genuinely need
|
||||
# 80 GB VRAM or sm_90 features). Compute-cap derivation in
|
||||
# `argo-train.sh` matches this default to sm_89 (Ada Lovelace).
|
||||
- name: gpu-pool
|
||||
value: ci-training-l40s
|
||||
- name: hyperopt-trials
|
||||
value: "20"
|
||||
- name: hyperopt-epochs
|
||||
value: "8"
|
||||
- name: train-epochs
|
||||
value: "50"
|
||||
- name: symbol
|
||||
value: ES.FUT
|
||||
- name: initial-capital
|
||||
value: "35000"
|
||||
- name: tx-cost-bps
|
||||
value: "0.1"
|
||||
- name: tick-size
|
||||
value: "0.25"
|
||||
- name: spread-ticks
|
||||
value: "1.0"
|
||||
- name: sanitizer
|
||||
value: "none" # "none", "memcheck", "racecheck", "synccheck"
|
||||
# Bar formation params — passed to BOTH precompute_features (when
|
||||
# building fxcache) AND train_baseline_rl (when looking up fxcache).
|
||||
# MUST match between the two for fxcache HIT (cache key includes them
|
||||
# as of 2026-05-09 architectural fix).
|
||||
#
|
||||
# imbalance-bar-threshold default: 20.0 (set 2026-05-10). At threshold=0.5
|
||||
# the sampler produced 209M bars from 209M trade ticks (1:1 ratio,
|
||||
# essentially per-tick) on workflow f5wnd, near-OOM in feature extraction.
|
||||
# threshold=20 produces ~5-6M bars matching the volume-bar density baseline.
|
||||
- name: imbalance-bar-threshold
|
||||
value: "20.0"
|
||||
- name: imbalance-bar-ewma-alpha
|
||||
value: "0.1"
|
||||
- name: volume-bar-size
|
||||
value: "100"
|
||||
- name: data-source
|
||||
value: "mbp10"
|
||||
|
||||
volumes:
|
||||
- name: git-ssh-key
|
||||
secret:
|
||||
secretName: argo-git-ssh-key
|
||||
defaultMode: 256
|
||||
- name: training-data
|
||||
persistentVolumeClaim:
|
||||
claimName: training-data-pvc
|
||||
- name: cargo-target-cuda
|
||||
persistentVolumeClaim:
|
||||
claimName: cargo-target-cuda
|
||||
- name: feature-cache
|
||||
persistentVolumeClaim:
|
||||
claimName: feature-cache-pvc
|
||||
|
||||
volumeClaimTemplates:
|
||||
- metadata:
|
||||
name: workspace
|
||||
spec:
|
||||
accessModes: ["ReadWriteOnce"]
|
||||
storageClassName: scw-bssd
|
||||
resources:
|
||||
requests:
|
||||
storage: 5Gi
|
||||
|
||||
templates:
|
||||
# ── DAG: orchestrate all steps ──
|
||||
- name: pipeline
|
||||
dag:
|
||||
tasks:
|
||||
- name: ensure-binary
|
||||
template: ensure-binary
|
||||
- name: gpu-warmup
|
||||
template: gpu-warmup
|
||||
# gpu-warmup pre-provisions an L40S node so hyperopt /
|
||||
# train-best don't pay autoscaler latency at start. When
|
||||
# hyperopt-trials==0 (precompute-only path used by
|
||||
# `scripts/argo-precompute.sh`), both downstream consumers
|
||||
# are skipped and the GPU node would sit idle until the
|
||||
# workflow ends. Skip the warmup in that case to avoid
|
||||
# provisioning expensive L40S capacity for nothing.
|
||||
when: "{{workflow.parameters.hyperopt-trials}} != 0"
|
||||
- name: ensure-fxcache
|
||||
template: ensure-fxcache
|
||||
dependencies: [ensure-binary]
|
||||
arguments:
|
||||
parameters:
|
||||
- name: sha
|
||||
value: "{{tasks.ensure-binary.outputs.parameters.sha}}"
|
||||
- name: hyperopt
|
||||
template: hyperopt
|
||||
dependencies: [ensure-binary, gpu-warmup, ensure-fxcache]
|
||||
when: "{{workflow.parameters.hyperopt-trials}} != 0"
|
||||
arguments:
|
||||
parameters:
|
||||
- name: sha
|
||||
value: "{{tasks.ensure-binary.outputs.parameters.sha}}"
|
||||
- name: train-best
|
||||
template: train-best
|
||||
dependencies: [ensure-binary, gpu-warmup, ensure-fxcache, hyperopt]
|
||||
arguments:
|
||||
parameters:
|
||||
- name: sha
|
||||
value: "{{tasks.ensure-binary.outputs.parameters.sha}}"
|
||||
- name: evaluate
|
||||
template: evaluate
|
||||
dependencies: [train-best]
|
||||
arguments:
|
||||
parameters:
|
||||
- name: sha
|
||||
value: "{{tasks.ensure-binary.outputs.parameters.sha}}"
|
||||
- name: upload-results
|
||||
template: upload-results
|
||||
dependencies: [evaluate]
|
||||
|
||||
# ── ensure-binary: cache-or-compile training binaries by commit SHA ──
|
||||
- name: ensure-binary
|
||||
outputs:
|
||||
parameters:
|
||||
- name: sha
|
||||
valueFrom:
|
||||
path: /tmp/sha
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: ci-compile-cpu
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
tolerations:
|
||||
- key: node.cilium.io/agent-not-ready
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
container:
|
||||
image: gitlab-registry.foxhunt.svc.cluster.local:5000/root/foxhunt/ci-builder:latest
|
||||
imagePullPolicy: Always
|
||||
command: ["/bin/bash", "-c"]
|
||||
env:
|
||||
- name: SQLX_OFFLINE
|
||||
value: "true"
|
||||
- name: CARGO_TERM_COLOR
|
||||
value: always
|
||||
- name: CARGO_TARGET_DIR
|
||||
value: /cargo-target
|
||||
- name: CARGO_HOME
|
||||
value: /cargo-target/cargo-home
|
||||
# sccache: per-crate compile-output cache. Disk-local on the same
|
||||
# cargo-target-cuda PVC so no network cost on hit; survives across
|
||||
# pods and commits (cargo's incremental cache is invalidated by
|
||||
# git checkout mtime touches — sccache is content-hash keyed so it
|
||||
# isn't). Binary assumed present in ci-builder image.
|
||||
#
|
||||
# CARGO_INCREMENTAL=0 is required to make sccache effective:
|
||||
# the project's Cargo.toml / .cargo/config.toml set
|
||||
# `incremental = true`, which emits per-query save-analysis
|
||||
# artifacts that sccache does not cache. Turning incremental off
|
||||
# lets rustc emit pure object output that sccache hashes
|
||||
# uniformly — same source + same args → cache hit.
|
||||
- name: RUSTC_WRAPPER
|
||||
value: sccache
|
||||
- name: SCCACHE_DIR
|
||||
value: /cargo-target/sccache
|
||||
- name: SCCACHE_CACHE_SIZE
|
||||
value: "40G"
|
||||
- name: CARGO_INCREMENTAL
|
||||
value: "0"
|
||||
- name: GITLAB_PAT
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: gitlab-pat
|
||||
key: token
|
||||
resources:
|
||||
requests:
|
||||
cpu: "14"
|
||||
memory: 32Gi
|
||||
limits:
|
||||
cpu: "30"
|
||||
memory: 64Gi
|
||||
volumeMounts:
|
||||
- name: git-ssh-key
|
||||
mountPath: /etc/git-ssh
|
||||
readOnly: true
|
||||
- name: cargo-target-cuda
|
||||
mountPath: /cargo-target
|
||||
- name: training-data
|
||||
mountPath: /data
|
||||
args:
|
||||
- |
|
||||
set -e
|
||||
SHA="{{workflow.parameters.commit-sha}}"
|
||||
BRANCH="{{workflow.parameters.git-branch}}"
|
||||
MODEL="{{workflow.parameters.model}}"
|
||||
|
||||
# ── SSH setup ──
|
||||
mkdir -p ~/.ssh
|
||||
cp /etc/git-ssh/ssh-privatekey ~/.ssh/id_ed25519
|
||||
chmod 600 ~/.ssh/id_ed25519
|
||||
printf 'Host *\n StrictHostKeyChecking no\n UserKnownHostsFile /dev/null\n' > ~/.ssh/config
|
||||
chmod 600 ~/.ssh/config
|
||||
|
||||
REPO="ssh://git@gitlab-gitlab-shell.foxhunt.svc.cluster.local:2222/root/foxhunt.git"
|
||||
BUILD="/cargo-target/src"
|
||||
|
||||
git config --global --add safe.directory "$BUILD"
|
||||
|
||||
# ── Resolve SHA ──
|
||||
if [ "$SHA" = "HEAD" ]; then
|
||||
if [ -d "$BUILD/.git" ]; then
|
||||
cd "$BUILD"
|
||||
git fetch origin
|
||||
SHA=$(git rev-parse "origin/$BRANCH")
|
||||
cd /
|
||||
else
|
||||
git clone --filter=blob:none "$REPO" "$BUILD"
|
||||
cd "$BUILD"
|
||||
git checkout "origin/$BRANCH"
|
||||
SHA=$(git rev-parse HEAD)
|
||||
cd /
|
||||
fi
|
||||
fi
|
||||
SHORT_SHA=$(echo "$SHA" | cut -c1-9)
|
||||
echo "Resolved SHA: $SHA (short: $SHORT_SHA)"
|
||||
|
||||
# ── Derive needed binaries from model ──
|
||||
case "$MODEL" in
|
||||
dqn|ppo)
|
||||
BINARIES="hyperopt_baseline_rl train_baseline_rl evaluate_baseline precompute_features"
|
||||
;;
|
||||
*)
|
||||
BINARIES="hyperopt_baseline_supervised train_baseline_supervised evaluate_supervised precompute_features"
|
||||
;;
|
||||
esac
|
||||
|
||||
# ── Cache check ──
|
||||
BIN_DIR="/data/bin/$SHORT_SHA"
|
||||
ALL_CACHED=true
|
||||
for bin in $BINARIES; do
|
||||
if [ ! -x "$BIN_DIR/$bin" ]; then
|
||||
ALL_CACHED=false
|
||||
break
|
||||
fi
|
||||
done
|
||||
|
||||
if [ "$ALL_CACHED" = "true" ]; then
|
||||
echo "=== Cache HIT: all binaries present in $BIN_DIR ==="
|
||||
ls -lh "$BIN_DIR/"
|
||||
echo "$SHORT_SHA" > /tmp/sha
|
||||
exit 0
|
||||
fi
|
||||
|
||||
echo "=== Cache MISS: compiling binaries for $SHORT_SHA ==="
|
||||
|
||||
# ── Clone / checkout ──
|
||||
if [ -d "$BUILD/.git" ]; then
|
||||
cd "$BUILD"
|
||||
git fetch origin
|
||||
CURRENT=$(git rev-parse HEAD 2>/dev/null || echo "none")
|
||||
if [ "$CURRENT" != "$SHA" ]; then
|
||||
echo "Updating checkout: $(echo $CURRENT | cut -c1-8) -> $SHORT_SHA"
|
||||
git checkout --force "$SHA"
|
||||
git clean -fd
|
||||
fi
|
||||
else
|
||||
git clone --filter=blob:none "$REPO" "$BUILD"
|
||||
cd "$BUILD"
|
||||
git checkout "$SHA"
|
||||
fi
|
||||
|
||||
# ── Build ──
|
||||
export PATH="${CARGO_HOME}/bin:${PATH}"
|
||||
export CUDA_COMPUTE_CAP={{workflow.parameters.cuda-compute-cap}}
|
||||
|
||||
ML_EXAMPLE_ARGS=""
|
||||
for ex in $BINARIES; do
|
||||
ML_EXAMPLE_ARGS="$ML_EXAMPLE_ARGS --example $ex"
|
||||
done
|
||||
|
||||
echo "Building: $BINARIES"
|
||||
echo " CUDA_COMPUTE_CAP=$CUDA_COMPUTE_CAP"
|
||||
cargo build --release -p ml --features ml/cuda $ML_EXAMPLE_ARGS
|
||||
|
||||
# ── Install to cache dir ──
|
||||
mkdir -p "$BIN_DIR"
|
||||
for bin in $BINARIES; do
|
||||
cp "$CARGO_TARGET_DIR/release/examples/$bin" "$BIN_DIR/"
|
||||
done
|
||||
strip "$BIN_DIR/"*
|
||||
|
||||
echo "=== Cached binaries ==="
|
||||
ls -lh "$BIN_DIR/"
|
||||
|
||||
# ── Prune old SHAs: keep last 5 ──
|
||||
cd /data/bin
|
||||
ls -1t | tail -n +6 | while read -r old; do
|
||||
echo "Pruning old cache: $old"
|
||||
rm -rf "$old"
|
||||
done
|
||||
|
||||
echo "$SHORT_SHA" > /tmp/sha
|
||||
|
||||
# ── ensure-fxcache: precompute feature cache if needed ──
|
||||
#
|
||||
# Pinned to the high-memory ci-compile-cpu-hm pool (POP2-HM-32C-256G,
|
||||
# 32 vCPU + 256GB RAM, min_size=0). The 9-quarter precompute_features
|
||||
# peaks ~50-60GB during the post-OFI alpha_trades conversion which
|
||||
# OOM-killed the standard 64GB ci-compile-cpu pool twice on
|
||||
# 2026-05-16 (workflows train-wq8b8 + train-2l6p4 both died at
|
||||
# exitCode 137 right after "OFI computed"). The HM pool autoscales
|
||||
# to zero when idle so this pin only costs anything during an
|
||||
# actual fxcache rebuild.
|
||||
- name: ensure-fxcache
|
||||
inputs:
|
||||
parameters:
|
||||
- name: sha
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: ci-compile-cpu-hm
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
tolerations:
|
||||
- key: node.cilium.io/agent-not-ready
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
container:
|
||||
image: ubuntu:24.04
|
||||
command: ["/bin/bash", "-c"]
|
||||
env:
|
||||
- name: RUST_LOG
|
||||
value: info
|
||||
resources:
|
||||
requests:
|
||||
cpu: "4"
|
||||
memory: 16Gi
|
||||
limits:
|
||||
cpu: "28"
|
||||
memory: 200Gi
|
||||
volumeMounts:
|
||||
- name: training-data
|
||||
mountPath: /data
|
||||
readOnly: true
|
||||
- name: feature-cache
|
||||
mountPath: /feature-cache
|
||||
args:
|
||||
- |
|
||||
set -e
|
||||
SHA="{{inputs.parameters.sha}}"
|
||||
BINARY="/data/bin/$SHA/precompute_features"
|
||||
|
||||
if [ ! -x "$BINARY" ]; then
|
||||
echo "ERROR: precompute_features not found at $BINARY"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
# fxcache has built-in version validation (FXCACHE_VERSION in header).
|
||||
# The precompute binary checks existing cache — if version matches,
|
||||
# it skips regeneration. If version mismatches, it fails and we
|
||||
# delete + regenerate. No unconditional rm — cache is reused when valid.
|
||||
echo "=== Running precompute_features (SHA: $SHA) ==="
|
||||
if $BINARY \
|
||||
--data-dir /data/futures-baseline \
|
||||
--mbp10-data-dir /data/futures-baseline-mbp10 \
|
||||
--trades-data-dir /data/futures-baseline-trades \
|
||||
--output-dir /feature-cache \
|
||||
--symbol {{workflow.parameters.symbol}} \
|
||||
--data-source {{workflow.parameters.data-source}} \
|
||||
--imbalance-bar-threshold {{workflow.parameters.imbalance-bar-threshold}} \
|
||||
--imbalance-bar-ewma-alpha {{workflow.parameters.imbalance-bar-ewma-alpha}} \
|
||||
--volume-bar-size {{workflow.parameters.volume-bar-size}} \
|
||||
--yes; then
|
||||
echo "=== Feature cache ready ==="
|
||||
else
|
||||
echo "=== Cache stale or missing — regenerating ==="
|
||||
# NOTE: don't blanket-rm /feature-cache/*.fxcache anymore. With
|
||||
# bar-params now in the cache key, multiple valid caches can
|
||||
# coexist (different threshold experiments).
|
||||
$BINARY \
|
||||
--data-dir /data/futures-baseline \
|
||||
--mbp10-data-dir /data/futures-baseline-mbp10 \
|
||||
--trades-data-dir /data/futures-baseline-trades \
|
||||
--output-dir /feature-cache \
|
||||
--symbol {{workflow.parameters.symbol}} \
|
||||
--data-source {{workflow.parameters.data-source}} \
|
||||
--imbalance-bar-threshold {{workflow.parameters.imbalance-bar-threshold}} \
|
||||
--imbalance-bar-ewma-alpha {{workflow.parameters.imbalance-bar-ewma-alpha}} \
|
||||
--volume-bar-size {{workflow.parameters.volume-bar-size}} \
|
||||
--yes
|
||||
echo "=== Feature cache regenerated ==="
|
||||
fi
|
||||
|
||||
# ── gpu-warmup: trigger GPU node autoscale during compile ──
|
||||
- name: gpu-warmup
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: "{{workflow.parameters.gpu-pool}}"
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
tolerations:
|
||||
- key: nvidia.com/gpu
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
- key: node.cilium.io/agent-not-ready
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
container:
|
||||
image: busybox:1.37
|
||||
command: ["/bin/sh", "-c"]
|
||||
args:
|
||||
- |
|
||||
echo "GPU warmup: triggering node autoscale..."
|
||||
echo "GPU node scheduled, exiting to free resources"
|
||||
resources:
|
||||
requests:
|
||||
nvidia.com/gpu: "1"
|
||||
cpu: 100m
|
||||
memory: 64Mi
|
||||
limits:
|
||||
nvidia.com/gpu: "1"
|
||||
cpu: 200m
|
||||
memory: 128Mi
|
||||
|
||||
# ── hyperopt: PSO/TPE hyperparameter optimization on GPU ──
|
||||
- name: hyperopt
|
||||
inputs:
|
||||
parameters:
|
||||
- name: sha
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: "{{workflow.parameters.gpu-pool}}"
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
tolerations:
|
||||
- key: nvidia.com/gpu
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
- key: node.cilium.io/agent-not-ready
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
container:
|
||||
image: gitlab-registry.foxhunt.svc.cluster.local:5000/root/foxhunt/ci-builder:latest
|
||||
imagePullPolicy: IfNotPresent
|
||||
command: ["/bin/sh", "-c"]
|
||||
env:
|
||||
- name: RUST_LOG
|
||||
value: info
|
||||
- name: SQLX_OFFLINE
|
||||
value: "true"
|
||||
- name: CUBLAS_WORKSPACE_CONFIG
|
||||
value: ":4096:8"
|
||||
- name: FOXHUNT_FEATURE_CACHE_DIR
|
||||
value: /feature-cache
|
||||
resources:
|
||||
requests:
|
||||
nvidia.com/gpu: "1"
|
||||
cpu: "4"
|
||||
memory: 32Gi
|
||||
limits:
|
||||
nvidia.com/gpu: "1"
|
||||
cpu: "8"
|
||||
memory: 64Gi
|
||||
volumeMounts:
|
||||
- name: workspace
|
||||
mountPath: /workspace
|
||||
- name: training-data
|
||||
mountPath: /data
|
||||
readOnly: true
|
||||
- name: feature-cache
|
||||
mountPath: /feature-cache
|
||||
readOnly: true
|
||||
args:
|
||||
- |
|
||||
set -e
|
||||
SHA="{{inputs.parameters.sha}}"
|
||||
export PATH="/data/bin/$SHA:$PATH"
|
||||
export LD_LIBRARY_PATH=$(echo "$LD_LIBRARY_PATH" | tr ':' '\n' | grep -v stubs | tr '\n' ':' | sed 's/:$//')
|
||||
nvidia-smi
|
||||
|
||||
mkdir -p /workspace/output/hyperopt
|
||||
|
||||
MODEL="{{workflow.parameters.model}}"
|
||||
case "$MODEL" in
|
||||
dqn|ppo) BINARY=hyperopt_baseline_rl ;;
|
||||
*) BINARY=hyperopt_baseline_supervised ;;
|
||||
esac
|
||||
|
||||
echo "=== Running hyperopt: $MODEL ($BINARY) ==="
|
||||
echo " Trials: {{workflow.parameters.hyperopt-trials}}, Epochs: {{workflow.parameters.hyperopt-epochs}}"
|
||||
|
||||
${BINARY} \
|
||||
--model "$MODEL" \
|
||||
--phase fast \
|
||||
--trials {{workflow.parameters.hyperopt-trials}} \
|
||||
--epochs {{workflow.parameters.hyperopt-epochs}} \
|
||||
--symbol {{workflow.parameters.symbol}} \
|
||||
--tx-cost-bps {{workflow.parameters.tx-cost-bps}} \
|
||||
--tick-size {{workflow.parameters.tick-size}} \
|
||||
--spread-ticks {{workflow.parameters.spread-ticks}} \
|
||||
--initial-capital {{workflow.parameters.initial-capital}} \
|
||||
--data-dir /data/futures-baseline \
|
||||
--mbp10-data-dir /data/futures-baseline-mbp10 \
|
||||
--trades-data-dir /data/futures-baseline-trades \
|
||||
--base-dir /workspace/output/hyperopt \
|
||||
--output /workspace/output/${MODEL}_hyperopt_results.json
|
||||
|
||||
echo "=== Hyperopt complete ==="
|
||||
cat /workspace/output/${MODEL}_hyperopt_results.json 2>/dev/null || echo "No results file"
|
||||
|
||||
# ── train-best: full training with best hyperparams ──
|
||||
- name: train-best
|
||||
inputs:
|
||||
parameters:
|
||||
- name: sha
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: "{{workflow.parameters.gpu-pool}}"
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
tolerations:
|
||||
- key: nvidia.com/gpu
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
- key: node.cilium.io/agent-not-ready
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
container:
|
||||
image: gitlab-registry.foxhunt.svc.cluster.local:5000/root/foxhunt/ci-builder:latest
|
||||
imagePullPolicy: IfNotPresent
|
||||
command: ["/bin/sh", "-c"]
|
||||
env:
|
||||
- name: RUST_LOG
|
||||
value: info
|
||||
- name: SQLX_OFFLINE
|
||||
value: "true"
|
||||
- name: CUBLAS_WORKSPACE_CONFIG
|
||||
value: ":4096:8"
|
||||
- name: FOXHUNT_FEATURE_CACHE_DIR
|
||||
value: /feature-cache
|
||||
resources:
|
||||
requests:
|
||||
nvidia.com/gpu: "1"
|
||||
cpu: "4"
|
||||
memory: 32Gi
|
||||
limits:
|
||||
nvidia.com/gpu: "1"
|
||||
cpu: "8"
|
||||
memory: 64Gi
|
||||
volumeMounts:
|
||||
- name: workspace
|
||||
mountPath: /workspace
|
||||
- name: training-data
|
||||
mountPath: /data
|
||||
readOnly: true
|
||||
- name: feature-cache
|
||||
mountPath: /feature-cache
|
||||
readOnly: true
|
||||
args:
|
||||
- |
|
||||
set -e
|
||||
SHA="{{inputs.parameters.sha}}"
|
||||
export PATH="/data/bin/$SHA:$PATH"
|
||||
export LD_LIBRARY_PATH=$(echo "$LD_LIBRARY_PATH" | tr ':' '\n' | grep -v stubs | tr '\n' ':' | sed 's/:$//')
|
||||
nvidia-smi
|
||||
|
||||
MODEL="{{workflow.parameters.model}}"
|
||||
case "$MODEL" in
|
||||
dqn|ppo) BINARY=train_baseline_rl ;;
|
||||
*) BINARY=train_baseline_supervised ;;
|
||||
esac
|
||||
|
||||
HYPEROPT_FLAG=""
|
||||
if [ -f "/workspace/output/${MODEL}_hyperopt_results.json" ]; then
|
||||
HYPEROPT_FLAG="--hyperopt-params /workspace/output/${MODEL}_hyperopt_results.json"
|
||||
echo " Using hyperopt results: ${MODEL}_hyperopt_results.json"
|
||||
else
|
||||
echo " No hyperopt results — training with default hyperparams"
|
||||
fi
|
||||
|
||||
echo "=== Training: $MODEL ({{workflow.parameters.train-epochs}} epochs) ==="
|
||||
|
||||
# compute-sanitizer / nsys: optional GPU debugging/profiling
|
||||
SANITIZER="{{workflow.parameters.sanitizer}}"
|
||||
SANITIZER_PREFIX=""
|
||||
if [ "$SANITIZER" = "nsys" ]; then
|
||||
NSYS_BIN=$(which nsys 2>/dev/null || echo "/usr/local/cuda/bin/nsys")
|
||||
if [ -x "$NSYS_BIN" ]; then
|
||||
NSYS_OUT="/feature-cache/nsys_$(date +%Y%m%d_%H%M%S)"
|
||||
mkdir -p /feature-cache
|
||||
SANITIZER_PREFIX="$NSYS_BIN profile -o $NSYS_OUT --cuda-graph-trace=node --stats=true --show-output=true -f true --duration=60"
|
||||
echo " nsys profiling enabled — output: ${NSYS_OUT}.nsys-rep (persistent PVC)"
|
||||
echo " --cuda-graph-trace=node: per-kernel timing inside parent graph"
|
||||
echo " --cudabacktrace=kernel: kernel source attribution"
|
||||
echo " --capture-range=cudaProfilerApi: use cudaProfilerStart/Stop to limit capture"
|
||||
echo " WARNING: ~2x slower — use with 1-2 epochs only"
|
||||
else
|
||||
echo " WARNING: nsys not found at $NSYS_BIN — running without"
|
||||
fi
|
||||
elif [ "$SANITIZER" != "none" ] && [ -n "$SANITIZER" ]; then
|
||||
SANITIZER_BIN=$(which compute-sanitizer 2>/dev/null || echo "/usr/local/cuda/bin/compute-sanitizer")
|
||||
if [ -x "$SANITIZER_BIN" ]; then
|
||||
SANITIZER_PREFIX="$SANITIZER_BIN --tool $SANITIZER --print-limit 20 --error-exitcode 1"
|
||||
echo " compute-sanitizer enabled: --tool $SANITIZER"
|
||||
echo " WARNING: 10-100x slower — use with 1-2 epochs only"
|
||||
else
|
||||
echo " WARNING: compute-sanitizer not found at $SANITIZER_BIN — running without"
|
||||
fi
|
||||
fi
|
||||
|
||||
stdbuf -oL $SANITIZER_PREFIX ${BINARY} \
|
||||
--model "$MODEL" \
|
||||
--symbol {{workflow.parameters.symbol}} \
|
||||
--tx-cost-bps {{workflow.parameters.tx-cost-bps}} \
|
||||
--tick-size {{workflow.parameters.tick-size}} \
|
||||
--spread-ticks {{workflow.parameters.spread-ticks}} \
|
||||
--initial-capital {{workflow.parameters.initial-capital}} \
|
||||
--data-dir /data/futures-baseline \
|
||||
--mbp10-data-dir /data/futures-baseline-mbp10 \
|
||||
--trades-data-dir /data/futures-baseline-trades \
|
||||
--output-dir /workspace/output \
|
||||
--epochs {{workflow.parameters.train-epochs}} \
|
||||
--imbalance-bar-threshold {{workflow.parameters.imbalance-bar-threshold}} \
|
||||
--imbalance-bar-ewma-alpha {{workflow.parameters.imbalance-bar-ewma-alpha}} \
|
||||
--volume-bar-size {{workflow.parameters.volume-bar-size}} \
|
||||
$HYPEROPT_FLAG
|
||||
|
||||
echo "=== Training complete ==="
|
||||
|
||||
# ── evaluate: run evaluation on trained model ──
|
||||
- name: evaluate
|
||||
inputs:
|
||||
parameters:
|
||||
- name: sha
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: "{{workflow.parameters.gpu-pool}}"
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
tolerations:
|
||||
- key: nvidia.com/gpu
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
- key: node.cilium.io/agent-not-ready
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
container:
|
||||
image: gitlab-registry.foxhunt.svc.cluster.local:5000/root/foxhunt/ci-builder:latest
|
||||
imagePullPolicy: IfNotPresent
|
||||
command: ["/bin/sh", "-c"]
|
||||
env:
|
||||
- name: RUST_LOG
|
||||
value: info
|
||||
- name: SQLX_OFFLINE
|
||||
value: "true"
|
||||
resources:
|
||||
requests:
|
||||
nvidia.com/gpu: "1"
|
||||
cpu: "2"
|
||||
memory: 16Gi
|
||||
limits:
|
||||
nvidia.com/gpu: "1"
|
||||
cpu: "4"
|
||||
memory: 32Gi
|
||||
volumeMounts:
|
||||
- name: workspace
|
||||
mountPath: /workspace
|
||||
- name: training-data
|
||||
mountPath: /data
|
||||
readOnly: true
|
||||
- name: feature-cache
|
||||
mountPath: /feature-cache
|
||||
readOnly: true
|
||||
args:
|
||||
- |
|
||||
set -e
|
||||
SHA="{{inputs.parameters.sha}}"
|
||||
export PATH="/data/bin/$SHA:$PATH"
|
||||
export LD_LIBRARY_PATH=$(echo "$LD_LIBRARY_PATH" | tr ':' '\n' | grep -v stubs | tr '\n' ':' | sed 's/:$//')
|
||||
|
||||
MODEL="{{workflow.parameters.model}}"
|
||||
case "$MODEL" in
|
||||
dqn|ppo) BINARY=evaluate_baseline ;;
|
||||
*) BINARY=evaluate_supervised ;;
|
||||
esac
|
||||
|
||||
echo "=== Evaluating: $MODEL ==="
|
||||
# Args must match evaluate_baseline's Args struct — see crates/ml/examples/evaluate_baseline.rs
|
||||
# `--models-dir` (not --checkpoint-dir), `--output` is a FILE path (not --output-dir).
|
||||
mkdir -p /workspace/output/eval
|
||||
${BINARY} \
|
||||
--model "$MODEL" \
|
||||
--symbol {{workflow.parameters.symbol}} \
|
||||
--data-dir /data/futures-baseline \
|
||||
--models-dir /workspace/output \
|
||||
--output /workspace/output/eval/evaluation_report.json \
|
||||
--initial-capital {{workflow.parameters.initial-capital}} \
|
||||
--tx-cost-bps {{workflow.parameters.tx-cost-bps}} \
|
||||
--tick-size {{workflow.parameters.tick-size}} \
|
||||
--spread-ticks {{workflow.parameters.spread-ticks}} || {
|
||||
echo "WARN: Evaluation failed, continuing"
|
||||
}
|
||||
|
||||
echo "=== Evaluation complete ==="
|
||||
|
||||
# ── upload-results: push artifacts to GitLab packages ──
|
||||
- name: upload-results
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: platform
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
container:
|
||||
image: curlimages/curl:8.12.1
|
||||
command: ["/bin/sh", "-c"]
|
||||
env:
|
||||
- name: GITLAB_PAT
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: gitlab-pat
|
||||
key: token
|
||||
resources:
|
||||
requests:
|
||||
cpu: 100m
|
||||
memory: 128Mi
|
||||
limits:
|
||||
cpu: 500m
|
||||
memory: 256Mi
|
||||
volumeMounts:
|
||||
- name: workspace
|
||||
mountPath: /workspace
|
||||
args:
|
||||
- |
|
||||
set -e
|
||||
MODEL="{{workflow.parameters.model}}"
|
||||
SYMBOL="{{workflow.parameters.symbol}}"
|
||||
TIMESTAMP=$(date +%Y%m%d-%H%M%S)
|
||||
GITLAB_API="http://gitlab-webservice-default.foxhunt.svc.cluster.local:8181/api/v4"
|
||||
PKG_NAME="foxhunt-training-results"
|
||||
PKG_VERSION="${MODEL}-${SYMBOL}-${TIMESTAMP}"
|
||||
|
||||
echo "Uploading training artifacts: ${PKG_NAME}/${PKG_VERSION}"
|
||||
|
||||
UPLOADED=0
|
||||
find /workspace/output -type f | while read -r file; do
|
||||
REL_PATH="${file#/workspace/output/}"
|
||||
SAFE_NAME=$(echo "$REL_PATH" | tr '/' '--')
|
||||
echo " Uploading ${REL_PATH} as ${SAFE_NAME}..."
|
||||
curl -f --upload-file "$file" \
|
||||
-H "PRIVATE-TOKEN: ${GITLAB_PAT}" \
|
||||
"${GITLAB_API}/projects/1/packages/generic/${PKG_NAME}/${PKG_VERSION}/${SAFE_NAME}" && \
|
||||
UPLOADED=$((UPLOADED + 1)) || \
|
||||
echo " WARN: Failed to upload ${REL_PATH}"
|
||||
done
|
||||
|
||||
echo "=== Upload complete (${UPLOADED} files) ==="
|
||||
echo "Package: ${PKG_NAME}/${PKG_VERSION}"
|
||||
|
||||
# ── notify-result: post workflow outcome to Mattermost (onExit) ──
|
||||
- name: notify-result
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: platform
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
container:
|
||||
image: curlimages/curl:8.12.1
|
||||
command: ["/bin/sh", "-c"]
|
||||
env:
|
||||
- name: WEBHOOK_URL
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: notification-webhook
|
||||
key: webhook-url
|
||||
optional: true
|
||||
resources:
|
||||
requests:
|
||||
cpu: 50m
|
||||
memory: 32Mi
|
||||
limits:
|
||||
cpu: 100m
|
||||
memory: 64Mi
|
||||
args:
|
||||
- |
|
||||
STATUS="{{workflow.status}}"
|
||||
NAME="{{workflow.name}}"
|
||||
|
||||
if [ -z "$WEBHOOK_URL" ] || echo "$WEBHOOK_URL" | grep -q "PLACEHOLDER"; then
|
||||
echo "No webhook configured, skipping notification"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
if [ "$STATUS" = "Succeeded" ]; then
|
||||
EMOJI=":white_check_mark:"
|
||||
else
|
||||
EMOJI=":x:"
|
||||
fi
|
||||
|
||||
PAYLOAD="{\"username\":\"Argo CI\",\"text\":\"${EMOJI} **${NAME}** — ${STATUS} ({{workflow.duration}}s)\"}"
|
||||
|
||||
curl -sf -X POST -H 'Content-Type: application/json' \
|
||||
-d "$PAYLOAD" "$WEBHOOK_URL" || echo "WARN: webhook post failed"
|
||||
40
infra/k8s/gitea/clusterip-services.yaml
Normal file
40
infra/k8s/gitea/clusterip-services.yaml
Normal file
@@ -0,0 +1,40 @@
|
||||
# The gitea chart's gitea-http/gitea-ssh services are HEADLESS (clusterIP: None), so they resolve to the
|
||||
# pod IP (100.64.x.x). That range OVERLAPS the Tailscale CGNAT range (100.64.0.0/10), so the tailscale
|
||||
# sidecar in the tailscale-gitlab-proxy pod swallows traffic to the pod IP. These normal ClusterIP services
|
||||
# give Gitea a service IP in the service CIDR (10.32.x.x, outside the tailscale range) — the proxy nginx +
|
||||
# socat target these instead. (GitLab worked because its webservice svc was a normal ClusterIP.)
|
||||
apiVersion: v1
|
||||
kind: Service
|
||||
metadata:
|
||||
name: gitea-web
|
||||
namespace: foxhunt
|
||||
labels:
|
||||
app.kubernetes.io/name: gitea
|
||||
spec:
|
||||
type: ClusterIP
|
||||
selector:
|
||||
app.kubernetes.io/name: gitea
|
||||
app.kubernetes.io/instance: gitea
|
||||
ports:
|
||||
- name: http
|
||||
port: 3000
|
||||
targetPort: 3000
|
||||
protocol: TCP
|
||||
---
|
||||
apiVersion: v1
|
||||
kind: Service
|
||||
metadata:
|
||||
name: gitea-sshd
|
||||
namespace: foxhunt
|
||||
labels:
|
||||
app.kubernetes.io/name: gitea
|
||||
spec:
|
||||
type: ClusterIP
|
||||
selector:
|
||||
app.kubernetes.io/name: gitea
|
||||
app.kubernetes.io/instance: gitea
|
||||
ports:
|
||||
- name: ssh
|
||||
port: 22
|
||||
targetPort: 22
|
||||
protocol: TCP
|
||||
61
infra/k8s/gitea/networkpolicy.yaml
Normal file
61
infra/k8s/gitea/networkpolicy.yaml
Normal file
@@ -0,0 +1,61 @@
|
||||
# Gitea NetworkPolicy (foxhunt ns has default-deny-all). Egress: postgres + DNS + in-cluster webhook
|
||||
# target + general HTTPS. Ingress: from platform/argo pods + the tailscale proxy → :3000 (web) and :22 (ssh).
|
||||
apiVersion: networking.k8s.io/v1
|
||||
kind: NetworkPolicy
|
||||
metadata:
|
||||
name: gitea
|
||||
namespace: foxhunt
|
||||
spec:
|
||||
podSelector:
|
||||
matchLabels:
|
||||
app.kubernetes.io/name: gitea
|
||||
policyTypes:
|
||||
- Ingress
|
||||
- Egress
|
||||
ingress:
|
||||
# web (3000) + ssh (22) from platform + argo-workflow pods (cockpit build clones over HTTP)
|
||||
- from:
|
||||
- podSelector:
|
||||
matchExpressions:
|
||||
- key: app.kubernetes.io/part-of
|
||||
operator: In
|
||||
values: ["foxhunt", "argo-workflows"]
|
||||
ports:
|
||||
- { port: 3000, protocol: TCP }
|
||||
- { port: 22, protocol: TCP }
|
||||
# web + ssh from the tailscale proxy (nginx + socat) that fronts git.fxhnt.ai
|
||||
- from:
|
||||
- podSelector:
|
||||
matchLabels:
|
||||
app.kubernetes.io/name: tailscale-gitlab-proxy
|
||||
ports:
|
||||
- { port: 3000, protocol: TCP }
|
||||
- { port: 22, protocol: TCP }
|
||||
egress:
|
||||
# PostgreSQL
|
||||
- to:
|
||||
- podSelector:
|
||||
matchLabels:
|
||||
app.kubernetes.io/name: postgres
|
||||
ports:
|
||||
- { port: 5432, protocol: TCP }
|
||||
# DNS
|
||||
- ports:
|
||||
- { port: 53, protocol: UDP }
|
||||
- { port: 53, protocol: TCP }
|
||||
# Argo Events webhook target (gitea push hook → eventsource :12000, in-cluster)
|
||||
- to:
|
||||
- podSelector:
|
||||
matchExpressions:
|
||||
- key: app.kubernetes.io/part-of
|
||||
operator: In
|
||||
values: ["foxhunt", "argo-workflows"]
|
||||
ports:
|
||||
- { port: 12000, protocol: TCP }
|
||||
# general HTTPS (avatars, external clones) — external only, exclude cluster CIDRs
|
||||
- to:
|
||||
- ipBlock:
|
||||
cidr: 0.0.0.0/0
|
||||
except: ["10.32.0.0/16", "172.16.0.0/16"]
|
||||
ports:
|
||||
- { port: 443, protocol: TCP }
|
||||
64
infra/k8s/gitea/values.yaml
Normal file
64
infra/k8s/gitea/values.yaml
Normal file
@@ -0,0 +1,64 @@
|
||||
# Gitea — lightweight git host replacing GitLab (Phase 2B). External Postgres (existing in-cluster
|
||||
# `postgres`), no bundled DB/redis/memcached, Actions off (Argo does CI). Internal-only until cutover.
|
||||
replicaCount: 1
|
||||
image:
|
||||
rootless: true
|
||||
|
||||
# Disable all bundled subcharts — reuse the existing in-cluster postgres, no redis/memcached.
|
||||
postgresql:
|
||||
enabled: false
|
||||
postgresql-ha:
|
||||
enabled: false
|
||||
redis-cluster:
|
||||
enabled: false
|
||||
redis:
|
||||
enabled: false
|
||||
# chart 12.x replaced redis with valkey — disable both, gitea uses embedded queue/cache/session
|
||||
valkey-cluster:
|
||||
enabled: false
|
||||
valkey:
|
||||
enabled: false
|
||||
|
||||
persistence:
|
||||
enabled: true
|
||||
size: 5Gi
|
||||
storageClass: sbs-default-retain
|
||||
|
||||
resources:
|
||||
requests: { cpu: 100m, memory: 128Mi }
|
||||
limits: { cpu: "1", memory: 512Mi }
|
||||
|
||||
service:
|
||||
http: { type: ClusterIP, port: 3000 }
|
||||
ssh: { type: ClusterIP, port: 22 }
|
||||
|
||||
gitea:
|
||||
admin:
|
||||
existingSecret: gitea-admin
|
||||
config:
|
||||
actions:
|
||||
ENABLED: "false" # Argo does CI/CD; Gitea Actions off (subchart removed in chart 12.x)
|
||||
server:
|
||||
ROOT_URL: https://git.fxhnt.ai/
|
||||
DOMAIN: git.fxhnt.ai
|
||||
SSH_DOMAIN: git.fxhnt.ai
|
||||
SSH_PORT: "22" # clean git@git.fxhnt.ai clone URLs; :2222 retired (Task 5).
|
||||
DISABLE_SSH: "false"
|
||||
database:
|
||||
DB_TYPE: postgres
|
||||
HOST: postgres.foxhunt.svc.cluster.local:5432
|
||||
NAME: gitea
|
||||
USER: gitea
|
||||
service:
|
||||
DISABLE_REGISTRATION: "true"
|
||||
# No redis/memcached — use embedded adapters
|
||||
cache:
|
||||
ADAPTER: memory
|
||||
session:
|
||||
PROVIDER: db
|
||||
queue:
|
||||
TYPE: level
|
||||
additionalConfigFromEnvs:
|
||||
- name: GITEA__database__PASSWD
|
||||
valueFrom:
|
||||
secretKeyRef: { name: gitea-db, key: password }
|
||||
@@ -1,73 +0,0 @@
|
||||
# GitLab Runner — H100 GPU training workloads (hyperopt, walk-forward)
|
||||
# Runs on ci-training-h100 pool (H100 PCIe 1x80GB)
|
||||
# SXM pools have zero Scaleway quota — use PCIe until quota is granted.
|
||||
|
||||
gitlabUrl: http://gitlab-webservice-default.foxhunt.svc.cluster.local:8181
|
||||
# runnerToken set via --set at install time
|
||||
|
||||
replicas: 1
|
||||
|
||||
# Reuse the existing gitlab-runner SA (has pods/secrets/configmaps RBAC)
|
||||
rbac:
|
||||
create: false
|
||||
serviceAccount:
|
||||
create: false
|
||||
name: gitlab-runner
|
||||
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: infra
|
||||
|
||||
runners:
|
||||
cloneUrl: http://gitlab-webservice-default.foxhunt.svc.cluster.local:8181
|
||||
config: |
|
||||
[[runners]]
|
||||
clone_url = "http://gitlab-webservice-default.foxhunt.svc.cluster.local:8181"
|
||||
tag_list = ["kapsule", "h100"]
|
||||
[runners.kubernetes]
|
||||
namespace = "foxhunt"
|
||||
service_account = "gitlab-runner"
|
||||
image = "rust:1.89-slim"
|
||||
privileged = false
|
||||
node_selector_overwrite_allowed = ".*"
|
||||
cpu_request_overwrite_max_allowed = "24000m"
|
||||
cpu_limit_overwrite_max_allowed = "24000m"
|
||||
memory_request_overwrite_max_allowed = "200Gi"
|
||||
memory_limit_overwrite_max_allowed = "200Gi"
|
||||
poll_timeout = 600
|
||||
runtime_class_name = "nvidia"
|
||||
pod_annotations_overwrite_allowed = ".*"
|
||||
# Default resources for H100 training
|
||||
cpu_request = "2000m"
|
||||
cpu_limit = "3800m"
|
||||
memory_request = "4Gi"
|
||||
memory_limit = "8Gi"
|
||||
helper_cpu_request = "100m"
|
||||
helper_cpu_limit = "500m"
|
||||
helper_memory_request = "128Mi"
|
||||
helper_memory_limit = "512Mi"
|
||||
image_pull_secrets = ["gitlab-registry"]
|
||||
[runners.kubernetes.node_selector]
|
||||
"k8s.scaleway.com/pool-name" = "ci-training-h100"
|
||||
[runners.kubernetes.node_tolerations]
|
||||
"nvidia.com/gpu" = "NoSchedule"
|
||||
"node.cilium.io/agent-not-ready" = "NoSchedule"
|
||||
[runners.kubernetes.pod_labels]
|
||||
"app.kubernetes.io/part-of" = "foxhunt-ci"
|
||||
# Request GPU via K8s scheduler so only one training pod runs per GPU
|
||||
[[runners.kubernetes.pod_spec]]
|
||||
name = "build"
|
||||
patch_type = "strategic"
|
||||
patch = '{"containers":[{"name":"build","resources":{"requests":{"nvidia.com/gpu":"1"},"limits":{"nvidia.com/gpu":"1"}}}]}'
|
||||
# H100 PCIe PVCs (separate from L40S to avoid RWO conflicts)
|
||||
[[runners.kubernetes.volumes.pvc]]
|
||||
name = "training-data-h100-pvc"
|
||||
mount_path = "/mnt/training-data"
|
||||
read_only = true
|
||||
[[runners.kubernetes.volumes.pvc]]
|
||||
name = "sccache-h100-pvc"
|
||||
mount_path = "/mnt/sccache"
|
||||
read_only = false
|
||||
|
||||
tags: "kapsule,h100"
|
||||
|
||||
concurrent: 2
|
||||
@@ -1,69 +0,0 @@
|
||||
# GitLab Runner — 2×H100 GPU training workloads (multi-GPU hyperopt)
|
||||
# Runs on ci-training-h100x2 pool (2× H100 PCIe: 160GB total VRAM)
|
||||
# Multi-GPU auto-detected by MultiGpuConfig::detect() in ml crate.
|
||||
|
||||
gitlabUrl: http://gitlab-webservice-default.foxhunt.svc.cluster.local:8181
|
||||
|
||||
replicas: 1
|
||||
|
||||
rbac:
|
||||
create: false
|
||||
serviceAccount:
|
||||
create: false
|
||||
name: gitlab-runner
|
||||
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: infra
|
||||
|
||||
runners:
|
||||
cloneUrl: http://gitlab-webservice-default.foxhunt.svc.cluster.local:8181
|
||||
config: |
|
||||
[[runners]]
|
||||
clone_url = "http://gitlab-webservice-default.foxhunt.svc.cluster.local:8181"
|
||||
tag_list = ["kapsule", "h100x2"]
|
||||
[runners.kubernetes]
|
||||
namespace = "foxhunt"
|
||||
service_account = "gitlab-runner"
|
||||
image = "rust:1.89-slim"
|
||||
privileged = false
|
||||
node_selector_overwrite_allowed = ".*"
|
||||
cpu_request_overwrite_max_allowed = "16000m"
|
||||
cpu_limit_overwrite_max_allowed = "16000m"
|
||||
memory_request_overwrite_max_allowed = "192Gi"
|
||||
memory_limit_overwrite_max_allowed = "192Gi"
|
||||
poll_timeout = 600
|
||||
runtime_class_name = "nvidia"
|
||||
pod_annotations_overwrite_allowed = ".*"
|
||||
cpu_request = "4000m"
|
||||
cpu_limit = "8000m"
|
||||
memory_request = "8Gi"
|
||||
memory_limit = "16Gi"
|
||||
helper_cpu_request = "100m"
|
||||
helper_cpu_limit = "500m"
|
||||
helper_memory_request = "128Mi"
|
||||
helper_memory_limit = "512Mi"
|
||||
image_pull_secrets = ["gitlab-registry"]
|
||||
[runners.kubernetes.node_selector]
|
||||
"k8s.scaleway.com/pool-name" = "ci-training-h100x2"
|
||||
[runners.kubernetes.node_tolerations]
|
||||
"nvidia.com/gpu" = "NoSchedule"
|
||||
"node.cilium.io/agent-not-ready" = "NoSchedule"
|
||||
[runners.kubernetes.pod_labels]
|
||||
"app.kubernetes.io/part-of" = "foxhunt-ci"
|
||||
# Request both GPUs via K8s scheduler for exclusive access
|
||||
[[runners.kubernetes.pod_spec]]
|
||||
name = "build"
|
||||
patch_type = "strategic"
|
||||
patch = '{"containers":[{"name":"build","resources":{"requests":{"nvidia.com/gpu":"2"},"limits":{"nvidia.com/gpu":"2"}}}]}'
|
||||
[[runners.kubernetes.volumes.pvc]]
|
||||
name = "training-data-h100x2-pvc"
|
||||
mount_path = "/mnt/training-data"
|
||||
read_only = true
|
||||
[[runners.kubernetes.volumes.pvc]]
|
||||
name = "sccache-h100x2-pvc"
|
||||
mount_path = "/mnt/sccache"
|
||||
read_only = false
|
||||
|
||||
tags: "kapsule,h100x2"
|
||||
|
||||
concurrent: 2
|
||||
@@ -1,76 +0,0 @@
|
||||
# GitLab Runner — RL training workloads (DQN, PPO)
|
||||
# Runs on ci-training pool (L40S: 48GB VRAM)
|
||||
# Mounts separate PVCs to avoid RWO conflicts with main runner
|
||||
|
||||
gitlabUrl: http://gitlab-webservice-default.foxhunt.svc.cluster.local:8181
|
||||
# runnerToken set via --set at install time
|
||||
|
||||
replicas: 1
|
||||
|
||||
# Reuse the existing gitlab-runner SA (has pods/secrets/configmaps RBAC)
|
||||
rbac:
|
||||
create: false
|
||||
serviceAccount:
|
||||
create: false
|
||||
name: gitlab-runner
|
||||
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: infra
|
||||
|
||||
runners:
|
||||
cloneUrl: http://gitlab-webservice-default.foxhunt.svc.cluster.local:8181
|
||||
config: |
|
||||
[[runners]]
|
||||
clone_url = "http://gitlab-webservice-default.foxhunt.svc.cluster.local:8181"
|
||||
tag_list = ["kapsule", "l40s"]
|
||||
[runners.kubernetes]
|
||||
namespace = "foxhunt"
|
||||
service_account = "gitlab-runner"
|
||||
image = "rust:1.89-slim"
|
||||
privileged = false
|
||||
node_selector_overwrite_allowed = ".*"
|
||||
# Hyperopt RL jobs need 6000m CPU + 20Gi mem (PSO parallel trials)
|
||||
cpu_request_overwrite_max_allowed = "8000m"
|
||||
cpu_limit_overwrite_max_allowed = "8000m"
|
||||
memory_request_overwrite_max_allowed = "48Gi"
|
||||
memory_limit_overwrite_max_allowed = "48Gi"
|
||||
poll_timeout = 600
|
||||
# All RL runner jobs target GPU nodes → set nvidia runtime globally
|
||||
runtime_class_name = "nvidia"
|
||||
# Allow CI jobs to set pod annotations (training jobs expose Prometheus metrics)
|
||||
pod_annotations_overwrite_allowed = ".*"
|
||||
# Default resources for RL training on L40S
|
||||
cpu_request = "2000m"
|
||||
cpu_limit = "3800m"
|
||||
memory_request = "4Gi"
|
||||
memory_limit = "8Gi"
|
||||
helper_cpu_request = "100m"
|
||||
helper_cpu_limit = "500m"
|
||||
helper_memory_request = "128Mi"
|
||||
helper_memory_limit = "512Mi"
|
||||
image_pull_secrets = ["gitlab-registry"]
|
||||
[runners.kubernetes.node_selector]
|
||||
"k8s.scaleway.com/pool-name" = "ci-training"
|
||||
[runners.kubernetes.node_tolerations]
|
||||
"nvidia.com/gpu" = "NoSchedule"
|
||||
"node.cilium.io/agent-not-ready" = "NoSchedule"
|
||||
[runners.kubernetes.pod_labels]
|
||||
"app.kubernetes.io/part-of" = "foxhunt-ci"
|
||||
# Request GPU via K8s scheduler so only one training pod runs per GPU
|
||||
[[runners.kubernetes.pod_spec]]
|
||||
name = "build"
|
||||
patch_type = "strategic"
|
||||
patch = '{"containers":[{"name":"build","resources":{"requests":{"nvidia.com/gpu":"1"},"limits":{"nvidia.com/gpu":"1"}}}]}'
|
||||
# Training PVCs — separate from main runner to avoid RWO Multi-Attach errors
|
||||
[[runners.kubernetes.volumes.pvc]]
|
||||
name = "training-data-l4-pvc"
|
||||
mount_path = "/mnt/training-data"
|
||||
read_only = true
|
||||
[[runners.kubernetes.volumes.pvc]]
|
||||
name = "sccache-l4-pvc"
|
||||
mount_path = "/mnt/sccache"
|
||||
read_only = false
|
||||
|
||||
tags: "kapsule,l40s"
|
||||
|
||||
concurrent: 2
|
||||
@@ -1,86 +0,0 @@
|
||||
# GitLab Runner — Kubernetes executor
|
||||
# Runner manager pod lives on gitlab node pool
|
||||
# Default: build pods spawn on ci-compile-cpu pool (POP2-32C-128G)
|
||||
# CI jobs override node_selector for their target pool
|
||||
|
||||
gitlabUrl: http://gitlab-webservice-default.foxhunt.svc.cluster.local:8181
|
||||
# runnerToken set via --set at install time
|
||||
|
||||
replicas: 1
|
||||
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: platform
|
||||
|
||||
runners:
|
||||
# Override clone URL to internal service (pods can't reach Tailscale IPs)
|
||||
cloneUrl: http://gitlab-webservice-default.foxhunt.svc.cluster.local:8181
|
||||
config: |
|
||||
[[runners]]
|
||||
clone_url = "http://gitlab-webservice-default.foxhunt.svc.cluster.local:8181"
|
||||
tag_list = ["kapsule", "rust", "docker", "gpu"]
|
||||
[runners.kubernetes]
|
||||
namespace = "foxhunt"
|
||||
image = "rust:1.89-slim"
|
||||
privileged = false
|
||||
# Allow CI jobs to override node selector via KUBERNETES_NODE_SELECTOR_* vars
|
||||
# Format: KUBERNETES_NODE_SELECTOR_<LABEL>: "key=value"
|
||||
node_selector_overwrite_allowed = ".*"
|
||||
# Allow CI jobs to override resources via KUBERNETES_*_{REQUEST,LIMIT} vars
|
||||
# Values are maximums (not regexes) — training can request up to 80Gi
|
||||
cpu_request_overwrite_max_allowed = "24000m"
|
||||
cpu_limit_overwrite_max_allowed = "24000m"
|
||||
memory_request_overwrite_max_allowed = "80Gi"
|
||||
memory_limit_overwrite_max_allowed = "80Gi"
|
||||
# Scale-to-zero pools need ~3-5 min to provision; default 180s times out
|
||||
poll_timeout = 600
|
||||
# Allow CI jobs to override runtimeClassName via KUBERNETES_RUNTIME_CLASS_NAME
|
||||
# GPU jobs set "nvidia"; compile/kaniko jobs leave unset (default runc)
|
||||
runtime_class_name_overwrite_allowed = ".*"
|
||||
# Allow CI jobs to set pod annotations (training jobs expose Prometheus metrics)
|
||||
pod_annotations_overwrite_allowed = ".*"
|
||||
# Default resource limits for ci-compile-cpu (POP2-32C-128G)
|
||||
cpu_request = "3500m"
|
||||
cpu_limit = "7800m"
|
||||
memory_request = "12Gi"
|
||||
memory_limit = "28Gi"
|
||||
helper_cpu_request = "100m"
|
||||
helper_cpu_limit = "500m"
|
||||
helper_memory_request = "128Mi"
|
||||
helper_memory_limit = "512Mi"
|
||||
image_pull_secrets = ["gitlab-registry"]
|
||||
# Sub-tables must come AFTER all scalar values (TOML rule)
|
||||
[runners.kubernetes.node_selector]
|
||||
"k8s.scaleway.com/pool-name" = "ci-compile-cpu"
|
||||
[runners.kubernetes.node_tolerations]
|
||||
"nvidia.com/gpu" = "NoSchedule"
|
||||
# Cilium CNI agent takes ~30s to initialize on fresh scale-from-zero nodes
|
||||
"node.cilium.io/agent-not-ready" = "NoSchedule"
|
||||
[runners.kubernetes.pod_labels]
|
||||
"app.kubernetes.io/part-of" = "foxhunt-ci"
|
||||
# Mount training data PVC (Databento futures .dbn.zst files)
|
||||
# Path must NOT be under /data — Redis image's WORKDIR is /data and entrypoint chowns it
|
||||
[[runners.kubernetes.volumes.pvc]]
|
||||
name = "training-data-pvc"
|
||||
mount_path = "/mnt/training-data"
|
||||
read_only = true
|
||||
|
||||
# Runner tags for job matching
|
||||
tags: "kapsule,rust,docker,gpu"
|
||||
|
||||
# Concurrency — 10 allows all 9 Kaniko builds + test to run in parallel
|
||||
concurrent: 10
|
||||
|
||||
# RBAC for runner to spawn pods
|
||||
rbac:
|
||||
create: true
|
||||
rules:
|
||||
- apiGroups: [""]
|
||||
resources: ["pods", "pods/exec", "pods/log", "secrets", "configmaps"]
|
||||
verbs: ["get", "list", "watch", "create", "delete", "update", "patch"]
|
||||
- apiGroups: [""]
|
||||
resources: ["pods/attach"]
|
||||
verbs: ["create", "get"]
|
||||
|
||||
serviceAccount:
|
||||
create: true
|
||||
name: gitlab-runner
|
||||
@@ -75,10 +75,10 @@ spec:
|
||||
- name: ssh-proxy
|
||||
image: alpine/socat:latest
|
||||
args:
|
||||
- "TCP-LISTEN:2222,fork,reuseaddr,sndbuf=1048576,rcvbuf=1048576"
|
||||
- "TCP:gitlab-gitlab-shell.foxhunt.svc.cluster.local:2222,sndbuf=1048576,rcvbuf=1048576"
|
||||
- "TCP-LISTEN:22,fork,reuseaddr,sndbuf=1048576,rcvbuf=1048576"
|
||||
- "TCP:gitea-sshd.foxhunt.svc.cluster.local:22,sndbuf=1048576,rcvbuf=1048576"
|
||||
ports:
|
||||
- containerPort: 2222
|
||||
- containerPort: 22
|
||||
resources:
|
||||
requests:
|
||||
cpu: 25m
|
||||
@@ -152,7 +152,7 @@ data:
|
||||
return 301 https://$host$request_uri;
|
||||
}
|
||||
|
||||
# GitLab — git.fxhnt.ai
|
||||
# Gitea — git.fxhnt.ai (replaced GitLab, Phase 2B cutover)
|
||||
server {
|
||||
listen 443 ssl;
|
||||
server_name git.fxhnt.ai;
|
||||
@@ -164,7 +164,7 @@ data:
|
||||
client_max_body_size 0;
|
||||
|
||||
location / {
|
||||
proxy_pass http://gitlab-webservice-default.foxhunt.svc.cluster.local:8181;
|
||||
proxy_pass http://gitea-web.foxhunt.svc.cluster.local:3000;
|
||||
proxy_set_header Host $http_host;
|
||||
proxy_set_header X-Real-IP $remote_addr;
|
||||
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
|
||||
|
||||
@@ -1,169 +0,0 @@
|
||||
# GPU-enabled overlay for ml-training-service
|
||||
# Apply manually: kubectl apply -f infra/k8s/gpu-overlays/ml-training-service-gpu.yaml
|
||||
# Revert to CPU: kubectl apply -f infra/k8s/services/ml-training-service.yaml
|
||||
#
|
||||
# Binary fetched from MinIO at pod startup — works on any node pool.
|
||||
apiVersion: apps/v1
|
||||
kind: Deployment
|
||||
metadata:
|
||||
name: ml-training-service
|
||||
namespace: foxhunt
|
||||
labels:
|
||||
app.kubernetes.io/name: ml-training-service
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
spec:
|
||||
replicas: 1
|
||||
strategy:
|
||||
type: RollingUpdate
|
||||
rollingUpdate:
|
||||
maxSurge: 0
|
||||
maxUnavailable: 1
|
||||
selector:
|
||||
matchLabels:
|
||||
app.kubernetes.io/name: ml-training-service
|
||||
template:
|
||||
metadata:
|
||||
annotations:
|
||||
prometheus.io/scrape: "true"
|
||||
prometheus.io/port: "9094"
|
||||
prometheus.io/path: "/metrics"
|
||||
labels:
|
||||
app.kubernetes.io/name: ml-training-service
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
spec:
|
||||
serviceAccountName: ml-training-service
|
||||
securityContext:
|
||||
runAsNonRoot: true
|
||||
runAsUser: 1000
|
||||
runAsGroup: 1000
|
||||
fsGroup: 1000
|
||||
seccompProfile:
|
||||
type: RuntimeDefault
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: gpu-inference
|
||||
tolerations:
|
||||
- key: nvidia.com/gpu
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
imagePullSecrets:
|
||||
- name: gitlab-registry
|
||||
initContainers:
|
||||
- name: fetch-binary
|
||||
image: gitlab-registry.foxhunt.svc.cluster.local:5000/root/foxhunt/foxhunt-runtime:latest
|
||||
command: ["/bin/sh", "-c"]
|
||||
args:
|
||||
- |
|
||||
set -e
|
||||
rclone copyto \
|
||||
":s3:foxhunt-binaries/services/ml-training-service" \
|
||||
"/binaries/ml-training-service" \
|
||||
--s3-provider=Minio \
|
||||
--s3-endpoint=http://minio.foxhunt.svc.cluster.local:9000 \
|
||||
--s3-access-key-id="${MINIO_ACCESS_KEY}" \
|
||||
--s3-secret-access-key="${MINIO_SECRET_KEY}" \
|
||||
--s3-no-check-bucket
|
||||
chmod +x /binaries/ml-training-service
|
||||
echo "Fetched ml-training-service ($(stat -c%s /binaries/ml-training-service) bytes)"
|
||||
env:
|
||||
- name: MINIO_ACCESS_KEY
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: minio-credentials
|
||||
key: access-key
|
||||
- name: MINIO_SECRET_KEY
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: minio-credentials
|
||||
key: secret-key
|
||||
volumeMounts:
|
||||
- name: binaries
|
||||
mountPath: /binaries
|
||||
resources:
|
||||
requests:
|
||||
cpu: 100m
|
||||
memory: 64Mi
|
||||
limits:
|
||||
cpu: 500m
|
||||
memory: 128Mi
|
||||
containers:
|
||||
- name: ml-training-service
|
||||
image: gitlab-registry.foxhunt.svc.cluster.local:5000/root/foxhunt/foxhunt-runtime:latest
|
||||
securityContext:
|
||||
allowPrivilegeEscalation: false
|
||||
readOnlyRootFilesystem: true
|
||||
capabilities:
|
||||
drop: ["ALL"]
|
||||
command: ["/binaries/ml-training-service", "serve"]
|
||||
ports:
|
||||
- containerPort: 50053
|
||||
name: grpc
|
||||
- containerPort: 9094
|
||||
name: metrics
|
||||
env:
|
||||
- name: DATABASE_PASSWORD
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: db-credentials
|
||||
key: password
|
||||
- name: DATABASE_URL
|
||||
value: "postgresql://foxhunt:$(DATABASE_PASSWORD)@postgres:5432/foxhunt"
|
||||
- name: REDIS_URL
|
||||
value: "redis://redis:6379"
|
||||
- name: JWT_SECRET
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: jwt-secret
|
||||
key: secret
|
||||
- name: JWT_ISSUER
|
||||
value: foxhunt-api
|
||||
- name: JWT_AUDIENCE
|
||||
value: foxhunt-services
|
||||
- name: S3_ENDPOINT
|
||||
value: "http://minio.foxhunt.svc.cluster.local:9000"
|
||||
- name: S3_BUCKET
|
||||
value: foxhunt-models
|
||||
- name: ENABLE_GPU
|
||||
value: "true"
|
||||
- name: RUST_LOG
|
||||
value: info
|
||||
- name: OTEL_EXPORTER_OTLP_ENDPOINT
|
||||
value: "http://tempo.foxhunt.svc.cluster.local:4317"
|
||||
volumeMounts:
|
||||
- name: binaries
|
||||
mountPath: /binaries
|
||||
readOnly: true
|
||||
- name: tls-certs
|
||||
mountPath: /app/certs/ml-training-service
|
||||
readOnly: true
|
||||
- name: tmp
|
||||
mountPath: /tmp
|
||||
readinessProbe:
|
||||
tcpSocket:
|
||||
port: 50053
|
||||
initialDelaySeconds: 15
|
||||
periodSeconds: 10
|
||||
livenessProbe:
|
||||
tcpSocket:
|
||||
port: 50053
|
||||
initialDelaySeconds: 30
|
||||
periodSeconds: 15
|
||||
failureThreshold: 5
|
||||
resources:
|
||||
requests:
|
||||
nvidia.com/gpu: "1"
|
||||
cpu: "1"
|
||||
memory: 2Gi
|
||||
limits:
|
||||
nvidia.com/gpu: "1"
|
||||
cpu: "4"
|
||||
memory: 8Gi
|
||||
volumes:
|
||||
- name: binaries
|
||||
emptyDir:
|
||||
sizeLimit: 200Mi
|
||||
- name: tls-certs
|
||||
secret:
|
||||
secretName: ml-training-tls
|
||||
- name: tmp
|
||||
emptyDir:
|
||||
sizeLimit: 50Mi
|
||||
@@ -1,156 +0,0 @@
|
||||
# GPU-enabled overlay for trading-service
|
||||
# Apply manually: kubectl apply -f infra/k8s/gpu-overlays/trading-service-gpu.yaml
|
||||
# Revert to CPU: kubectl apply -f infra/k8s/services/trading-service.yaml
|
||||
#
|
||||
# Binary fetched from MinIO at pod startup — works on any node pool.
|
||||
apiVersion: apps/v1
|
||||
kind: Deployment
|
||||
metadata:
|
||||
name: trading-service
|
||||
namespace: foxhunt
|
||||
labels:
|
||||
app.kubernetes.io/name: trading-service
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
spec:
|
||||
replicas: 1
|
||||
strategy:
|
||||
type: RollingUpdate
|
||||
rollingUpdate:
|
||||
maxSurge: 0
|
||||
maxUnavailable: 1
|
||||
selector:
|
||||
matchLabels:
|
||||
app.kubernetes.io/name: trading-service
|
||||
template:
|
||||
metadata:
|
||||
annotations:
|
||||
prometheus.io/scrape: "true"
|
||||
prometheus.io/port: "9092"
|
||||
prometheus.io/path: "/metrics"
|
||||
labels:
|
||||
app.kubernetes.io/name: trading-service
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
spec:
|
||||
securityContext:
|
||||
runAsNonRoot: true
|
||||
runAsUser: 1000
|
||||
runAsGroup: 1000
|
||||
fsGroup: 1000
|
||||
seccompProfile:
|
||||
type: RuntimeDefault
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: l40s
|
||||
tolerations:
|
||||
- key: nvidia.com/gpu
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
imagePullSecrets:
|
||||
- name: gitlab-registry
|
||||
initContainers:
|
||||
- name: fetch-binary
|
||||
image: gitlab-registry.foxhunt.svc.cluster.local:5000/root/foxhunt/foxhunt-runtime:latest
|
||||
command: ["/bin/sh", "-c"]
|
||||
args:
|
||||
- |
|
||||
set -e
|
||||
rclone copyto \
|
||||
":s3:foxhunt-binaries/services/trading-service" \
|
||||
"/binaries/trading-service" \
|
||||
--s3-provider=Minio \
|
||||
--s3-endpoint=http://minio.foxhunt.svc.cluster.local:9000 \
|
||||
--s3-access-key-id="${MINIO_ACCESS_KEY}" \
|
||||
--s3-secret-access-key="${MINIO_SECRET_KEY}" \
|
||||
--s3-no-check-bucket
|
||||
chmod +x /binaries/trading-service
|
||||
echo "Fetched trading-service ($(stat -c%s /binaries/trading-service) bytes)"
|
||||
env:
|
||||
- name: MINIO_ACCESS_KEY
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: minio-credentials
|
||||
key: access-key
|
||||
- name: MINIO_SECRET_KEY
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: minio-credentials
|
||||
key: secret-key
|
||||
volumeMounts:
|
||||
- name: binaries
|
||||
mountPath: /binaries
|
||||
resources:
|
||||
requests:
|
||||
cpu: 100m
|
||||
memory: 64Mi
|
||||
limits:
|
||||
cpu: 500m
|
||||
memory: 128Mi
|
||||
containers:
|
||||
- name: trading-service
|
||||
image: gitlab-registry.foxhunt.svc.cluster.local:5000/root/foxhunt/foxhunt-runtime:latest
|
||||
securityContext:
|
||||
allowPrivilegeEscalation: false
|
||||
readOnlyRootFilesystem: true
|
||||
capabilities:
|
||||
drop: ["ALL"]
|
||||
command: ["/binaries/trading-service"]
|
||||
ports:
|
||||
- containerPort: 50051
|
||||
name: grpc
|
||||
- containerPort: 9092
|
||||
name: metrics
|
||||
env:
|
||||
- name: DATABASE_PASSWORD
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: db-credentials
|
||||
key: password
|
||||
- name: DATABASE_URL
|
||||
value: "postgresql://foxhunt:$(DATABASE_PASSWORD)@postgres:5432/foxhunt"
|
||||
- name: REDIS_URL
|
||||
value: "redis://redis:6379"
|
||||
- name: JWT_SECRET
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: jwt-secret
|
||||
key: secret
|
||||
- name: JWT_ISSUER
|
||||
value: foxhunt-api
|
||||
- name: JWT_AUDIENCE
|
||||
value: foxhunt-services
|
||||
- name: QUESTDB_ILP_HOST
|
||||
value: "questdb:9009"
|
||||
- name: GRPC_PORT
|
||||
value: "50051"
|
||||
- name: ENABLE_GPU_INFERENCE
|
||||
value: "true"
|
||||
- name: RUST_LOG
|
||||
value: info
|
||||
volumeMounts:
|
||||
- name: binaries
|
||||
mountPath: /binaries
|
||||
readOnly: true
|
||||
- name: tmp
|
||||
mountPath: /tmp
|
||||
readinessProbe:
|
||||
exec:
|
||||
command:
|
||||
- grpc_health_probe
|
||||
- -addr=localhost:50051
|
||||
initialDelaySeconds: 10
|
||||
periodSeconds: 10
|
||||
resources:
|
||||
requests:
|
||||
nvidia.com/gpu: "1"
|
||||
cpu: "1"
|
||||
memory: 2Gi
|
||||
limits:
|
||||
nvidia.com/gpu: "1"
|
||||
cpu: "4"
|
||||
memory: 8Gi
|
||||
volumes:
|
||||
- name: binaries
|
||||
emptyDir:
|
||||
sizeLimit: 200Mi
|
||||
- name: tmp
|
||||
emptyDir:
|
||||
sizeLimit: 50Mi
|
||||
@@ -1,4 +1,4 @@
|
||||
# Postgres: accepts connections from foxhunt app pods, GitLab, and Grafana
|
||||
# Postgres: accepts connections from foxhunt app pods, GitLab, Grafana, and Gitea
|
||||
apiVersion: networking.k8s.io/v1
|
||||
kind: NetworkPolicy
|
||||
metadata:
|
||||
@@ -23,6 +23,9 @@ spec:
|
||||
- podSelector:
|
||||
matchLabels:
|
||||
app.kubernetes.io/name: grafana
|
||||
- podSelector: # Gitea (Phase 2B) reuses the in-cluster postgres
|
||||
matchLabels:
|
||||
app.kubernetes.io/name: gitea
|
||||
ports:
|
||||
- protocol: TCP
|
||||
port: 5432
|
||||
|
||||
@@ -1,60 +0,0 @@
|
||||
# DaemonSet image pre-puller — keeps training images cached on GPU nodes
|
||||
# Runs on ci-training (L40S) pool so training jobs skip the pull.
|
||||
# Init containers pull :latest tags, then the main container sleeps forever.
|
||||
---
|
||||
apiVersion: apps/v1
|
||||
kind: DaemonSet
|
||||
metadata:
|
||||
name: image-prepuller
|
||||
namespace: foxhunt
|
||||
labels:
|
||||
app: image-prepuller
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
spec:
|
||||
selector:
|
||||
matchLabels:
|
||||
app: image-prepuller
|
||||
template:
|
||||
metadata:
|
||||
labels:
|
||||
app: image-prepuller
|
||||
spec:
|
||||
affinity:
|
||||
nodeAffinity:
|
||||
requiredDuringSchedulingIgnoredDuringExecution:
|
||||
nodeSelectorTerms:
|
||||
- matchExpressions:
|
||||
- key: k8s.scaleway.com/pool-name
|
||||
operator: In
|
||||
values:
|
||||
- ci-training
|
||||
tolerations:
|
||||
- key: nvidia.com/gpu
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
- key: node.cilium.io/agent-not-ready
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
imagePullSecrets:
|
||||
- name: gitlab-registry
|
||||
initContainers:
|
||||
- name: pull-training-runtime
|
||||
image: gitlab-registry.foxhunt.svc.cluster.local:5000/root/foxhunt/foxhunt-training-runtime:latest
|
||||
command: ["echo", "foxhunt-training-runtime image pulled"]
|
||||
resources:
|
||||
requests:
|
||||
cpu: 10m
|
||||
memory: 16Mi
|
||||
limits:
|
||||
cpu: 10m
|
||||
memory: 16Mi
|
||||
containers:
|
||||
- name: pause
|
||||
image: registry.k8s.io/pause:3.10
|
||||
resources:
|
||||
requests:
|
||||
cpu: 10m
|
||||
memory: 16Mi
|
||||
limits:
|
||||
cpu: 10m
|
||||
memory: 16Mi
|
||||
@@ -1,130 +0,0 @@
|
||||
apiVersion: batch/v1
|
||||
kind: Job
|
||||
metadata:
|
||||
generateName: training-
|
||||
namespace: foxhunt
|
||||
labels:
|
||||
app.kubernetes.io/name: training
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
foxhunt/job-type: training
|
||||
spec:
|
||||
backoffLimit: 1
|
||||
activeDeadlineSeconds: 21600 # 6 hours — hyperopt runs 20 trials × 8 epochs
|
||||
ttlSecondsAfterFinished: 600
|
||||
template:
|
||||
metadata:
|
||||
annotations:
|
||||
prometheus.io/scrape: "true"
|
||||
prometheus.io/port: "9094"
|
||||
prometheus.io/path: "/metrics"
|
||||
labels:
|
||||
app.kubernetes.io/name: training
|
||||
app.kubernetes.io/component: training-workflow
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
foxhunt/job-type: training
|
||||
spec:
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: ci-training-h100
|
||||
tolerations:
|
||||
- key: nvidia.com/gpu
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
# Cilium CNI agent takes ~30s to initialize on fresh scale-from-zero nodes
|
||||
- key: node.cilium.io/agent-not-ready
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
imagePullSecrets:
|
||||
- name: gitlab-registry
|
||||
restartPolicy: Never
|
||||
initContainers:
|
||||
# 1. Fetch training binary from GitLab Generic Package Registry
|
||||
- name: fetch-binary
|
||||
image: gitlab-registry.foxhunt.svc.cluster.local:5000/root/foxhunt/foxhunt-training-runtime:latest
|
||||
command: ["/bin/sh", "-c"]
|
||||
args:
|
||||
- |
|
||||
set -e
|
||||
BINARY="$(TRAINING_BINARY)"
|
||||
curl -fSL -o "/binaries/${BINARY}" \
|
||||
--header "DEPLOY-TOKEN: ${GITLAB_DEPLOY_TOKEN}" \
|
||||
"${GITLAB_API}/projects/1/packages/generic/foxhunt-training/${FOXHUNT_RELEASE}/${BINARY}"
|
||||
chmod +x "/binaries/${BINARY}"
|
||||
echo "Fetched ${BINARY} ${FOXHUNT_RELEASE} ($(stat -c%s /binaries/${BINARY}) bytes)"
|
||||
env:
|
||||
- name: TRAINING_BINARY
|
||||
value: train_baseline_supervised
|
||||
- name: GITLAB_DEPLOY_TOKEN
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: gitlab-deploy-token
|
||||
key: token
|
||||
- name: GITLAB_API
|
||||
value: "http://gitlab-webservice-default.foxhunt.svc.cluster.local:8181/api/v4"
|
||||
- name: FOXHUNT_RELEASE
|
||||
value: "latest"
|
||||
volumeMounts:
|
||||
- name: binaries
|
||||
mountPath: /binaries
|
||||
resources:
|
||||
requests:
|
||||
cpu: 100m
|
||||
memory: 64Mi
|
||||
limits:
|
||||
cpu: 500m
|
||||
memory: 128Mi
|
||||
containers:
|
||||
- name: training
|
||||
image: gitlab-registry.foxhunt.svc.cluster.local:5000/root/foxhunt/foxhunt-training-runtime:latest
|
||||
# Available binaries (copied by initContainer):
|
||||
# train_baseline_rl (for dqn, ppo)
|
||||
# train_baseline_supervised (for tft, mamba2, tggn, tlob, liquid, kan, xlstm, diffusion)
|
||||
# evaluate_baseline
|
||||
# hyperopt_baseline_rl (for dqn, ppo)
|
||||
# hyperopt_baseline_supervised (for tft, mamba2)
|
||||
command: ["/binaries/$(TRAINING_BINARY)"]
|
||||
args:
|
||||
- "--symbol=ES.FUT"
|
||||
- "--data-dir=/data/futures-baseline"
|
||||
- "--mbp10-data-dir=/data/futures-baseline-mbp10"
|
||||
- "--trades-data-dir=/data/futures-baseline-trades"
|
||||
- "--training-profile=$(TRAINING_PROFILE)"
|
||||
- "--output=/output"
|
||||
env:
|
||||
- name: TRAINING_BINARY
|
||||
value: train_baseline_supervised
|
||||
- name: TRAINING_PROFILE
|
||||
value: "dqn-production"
|
||||
- name: RUST_LOG
|
||||
value: info
|
||||
- name: SQLX_OFFLINE
|
||||
value: "true"
|
||||
- name: OTEL_EXPORTER_OTLP_ENDPOINT
|
||||
value: "http://tempo.foxhunt.svc.cluster.local:4317"
|
||||
volumeMounts:
|
||||
- name: training-data
|
||||
mountPath: /data
|
||||
readOnly: true
|
||||
- name: output
|
||||
mountPath: /output
|
||||
- name: binaries
|
||||
mountPath: /binaries
|
||||
readOnly: true
|
||||
resources:
|
||||
requests:
|
||||
nvidia.com/gpu: "1"
|
||||
cpu: "4"
|
||||
memory: 16Gi
|
||||
limits:
|
||||
nvidia.com/gpu: "1"
|
||||
cpu: "8"
|
||||
memory: 32Gi
|
||||
volumes:
|
||||
- name: training-data
|
||||
persistentVolumeClaim:
|
||||
claimName: training-data-pvc
|
||||
- name: output
|
||||
emptyDir:
|
||||
sizeLimit: 2Gi
|
||||
- name: binaries
|
||||
emptyDir:
|
||||
sizeLimit: 500Mi
|
||||
@@ -1,110 +0,0 @@
|
||||
# infra/k8s/training/populate-test-data-job.yaml
|
||||
apiVersion: batch/v1
|
||||
kind: Job
|
||||
metadata:
|
||||
generateName: populate-test-data-
|
||||
namespace: foxhunt
|
||||
labels:
|
||||
app.kubernetes.io/name: populate-test-data
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
spec:
|
||||
backoffLimit: 1
|
||||
ttlSecondsAfterFinished: 300
|
||||
template:
|
||||
spec:
|
||||
restartPolicy: Never
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: ci-training-h100
|
||||
tolerations:
|
||||
- key: nvidia.com/gpu
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
- key: node.cilium.io/agent-not-ready
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
containers:
|
||||
- name: populate
|
||||
image: busybox:1.37
|
||||
command: ["/bin/sh", "-c"]
|
||||
args:
|
||||
- |
|
||||
set -e
|
||||
echo "=== Populating test data PVC ==="
|
||||
|
||||
# All symbols used in CI tests (4 symbols from futures-baseline)
|
||||
SYMBOLS="ES.FUT 6E.FUT ZN.FUT NQ.FUT"
|
||||
|
||||
for SYM in $SYMBOLS; do
|
||||
echo "--- $SYM ---"
|
||||
|
||||
# OHLCV (used by DQN/PPO training tests and walk-forward validation)
|
||||
# Copy 3 quarters per symbol: enough for --train-months 3 --val-months 1 --test-months 1
|
||||
# (5 months window), perf benchmarks, and walk-forward tests. ~6MB/symbol compressed.
|
||||
mkdir -p /test-data/ohlcv/$SYM
|
||||
if ls /training-data/futures-baseline/$SYM/*.dbn.zst >/dev/null 2>&1; then
|
||||
COUNT=0
|
||||
for F in $(ls /training-data/futures-baseline/$SYM/*.dbn.zst 2>/dev/null | head -3); do
|
||||
cp "$F" /test-data/ohlcv/$SYM/
|
||||
echo " OHLCV: copied $(basename $F)"
|
||||
COUNT=$((COUNT + 1))
|
||||
done
|
||||
echo " OHLCV: $COUNT quarters copied"
|
||||
else
|
||||
echo " OHLCV: no source data for $SYM"
|
||||
fi
|
||||
|
||||
# MBP-10 (used by OFI feature loading)
|
||||
mkdir -p /test-data/mbp10/$SYM
|
||||
if ls /training-data/futures-baseline-mbp10/$SYM/*.dbn.zst >/dev/null 2>&1; then
|
||||
COUNT=0
|
||||
for F in $(ls /training-data/futures-baseline-mbp10/$SYM/*.dbn.zst 2>/dev/null | head -3); do
|
||||
cp "$F" /test-data/mbp10/$SYM/
|
||||
echo " MBP-10: copied $(basename $F)"
|
||||
COUNT=$((COUNT + 1))
|
||||
done
|
||||
else
|
||||
echo " MBP-10: no source data for $SYM"
|
||||
fi
|
||||
|
||||
# Trades
|
||||
mkdir -p /test-data/trades/$SYM
|
||||
if ls /training-data/futures-baseline-trades/$SYM/*.dbn.zst >/dev/null 2>&1; then
|
||||
COUNT=0
|
||||
for F in $(ls /training-data/futures-baseline-trades/$SYM/*.dbn.zst 2>/dev/null | head -3); do
|
||||
cp "$F" /test-data/trades/$SYM/
|
||||
echo " Trades: copied $(basename $F)"
|
||||
COUNT=$((COUNT + 1))
|
||||
done
|
||||
else
|
||||
echo " Trades: no source data for $SYM"
|
||||
fi
|
||||
done
|
||||
|
||||
echo ""
|
||||
echo "=== Test data contents ==="
|
||||
find /test-data -type f -exec ls -lh {} \;
|
||||
echo "Total: $(du -sh /test-data | cut -f1)"
|
||||
echo "=== Done ==="
|
||||
resources:
|
||||
requests:
|
||||
nvidia.com/gpu: "1"
|
||||
cpu: 100m
|
||||
memory: 128Mi
|
||||
limits:
|
||||
nvidia.com/gpu: "1"
|
||||
cpu: 500m
|
||||
memory: 256Mi
|
||||
volumeMounts:
|
||||
- name: training-data
|
||||
mountPath: /training-data
|
||||
readOnly: true
|
||||
- name: test-data
|
||||
mountPath: /test-data
|
||||
volumes:
|
||||
- name: training-data
|
||||
persistentVolumeClaim:
|
||||
claimName: training-data-pvc
|
||||
readOnly: true
|
||||
- name: test-data
|
||||
persistentVolumeClaim:
|
||||
claimName: test-data-pvc
|
||||
@@ -1,24 +0,0 @@
|
||||
# MinIO credentials for rclone output sync in training jobs.
|
||||
# No longer needed as a standalone secret — training jobs now use minio-credentials
|
||||
# (deployed via infra/k8s/minio/minio.yaml).
|
||||
#
|
||||
# If you need to recreate manually:
|
||||
# kubectl -n foxhunt create secret generic minio-credentials \
|
||||
# --from-literal=access-key=<MINIO_ACCESS_KEY> \
|
||||
# --from-literal=secret-key=<MINIO_SECRET_KEY> \
|
||||
# --from-literal=root-user=<MINIO_ACCESS_KEY> \
|
||||
# --from-literal=root-password=<MINIO_SECRET_KEY>
|
||||
apiVersion: v1
|
||||
kind: Secret
|
||||
metadata:
|
||||
name: minio-credentials
|
||||
namespace: foxhunt
|
||||
labels:
|
||||
app.kubernetes.io/name: minio-credentials
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
type: Opaque
|
||||
stringData:
|
||||
access-key: REPLACE_ME
|
||||
secret-key: REPLACE_ME
|
||||
root-user: REPLACE_ME
|
||||
root-password: REPLACE_ME
|
||||
@@ -1,15 +0,0 @@
|
||||
apiVersion: v1
|
||||
kind: PersistentVolumeClaim
|
||||
metadata:
|
||||
name: training-output-pvc
|
||||
namespace: foxhunt
|
||||
labels:
|
||||
app.kubernetes.io/name: training-output
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
spec:
|
||||
accessModes:
|
||||
- ReadWriteOnce
|
||||
resources:
|
||||
requests:
|
||||
storage: 50Gi
|
||||
storageClassName: scw-bssd
|
||||
@@ -8,6 +8,7 @@ terraform {
|
||||
|
||||
inputs = {
|
||||
dns_zone = "fxhnt.ai"
|
||||
# Tailscale IP of foxhunt-gitlab proxy pod
|
||||
git_ip = "100.90.76.85"
|
||||
# Tailscale IP of the foxhunt-gitlab proxy pod (live node foxhunt-gitlab; the old 100.90.76.85
|
||||
# was a stale prior IP — its duplicate A records across *.fxhnt.ai were removed 2026-06-21).
|
||||
git_ip = "100.95.225.27"
|
||||
}
|
||||
|
||||
@@ -27,17 +27,17 @@ inputs = {
|
||||
# the 64GB ci-compile-cpu pool (9-quarter precompute_features peaks
|
||||
# ~50-60GB during the post-OFI alpha_trades conversion).
|
||||
# min_size=0 + size=0 initial so this pool costs nothing when idle.
|
||||
enable_ci_compile_cpu_hm_pool = true
|
||||
enable_ci_compile_cpu_hm_pool = false # decommissioned 2026-06-21 — Rust precompute_features retired
|
||||
ci_compile_cpu_hm_type = "POP2-HM-32C-256G"
|
||||
ci_compile_cpu_hm_max_size = 1
|
||||
|
||||
# L40S training pool (48GB VRAM, CUDA CC 89)
|
||||
enable_ci_training_l40s_pool = true
|
||||
enable_ci_training_l40s_pool = false # decommissioned 2026-06-21 — Rust ML training retired
|
||||
ci_training_l40s_type = "L40S-1-48G"
|
||||
ci_training_l40s_max_size = 1
|
||||
|
||||
# H100 training pool (80GB VRAM, CUDA CC 90)
|
||||
enable_ci_training_h100_pool = true
|
||||
enable_ci_training_h100_pool = false # decommissioned 2026-06-21 — Rust RL training retired
|
||||
ci_training_h100_type = "H100-1-80G"
|
||||
ci_training_h100_max_size = 1
|
||||
|
||||
|
||||
@@ -7,21 +7,31 @@ locals {
|
||||
project_id = get_env("SCW_DEFAULT_PROJECT_ID")
|
||||
}
|
||||
|
||||
# Remote state in GitLab Terraform state backend (project root/foxhunt, ID=1)
|
||||
# Token stored in k8s secret gitlab-terraform-state-token
|
||||
# Set TF_HTTP_USERNAME and TF_HTTP_PASSWORD env vars before running terragrunt.
|
||||
# Remote state in Scaleway Object Storage (bucket foxhunt-tfstate, fr-par) — migrated off GitLab 2026-06-21.
|
||||
# The s3 backend reads AWS_* env vars (set them to the Scaleway access/secret keys); SCW_* drives the provider.
|
||||
remote_state {
|
||||
backend = "http"
|
||||
backend = "s3"
|
||||
generate = {
|
||||
path = "backend.tf"
|
||||
if_exists = "overwrite"
|
||||
}
|
||||
config = {
|
||||
address = "${get_env("GITLAB_TF_STATE_URL", "https://tf.fxhnt.ai")}/api/v4/projects/1/terraform/state/${path_relative_to_include()}"
|
||||
lock_address = "${get_env("GITLAB_TF_STATE_URL", "https://tf.fxhnt.ai")}/api/v4/projects/1/terraform/state/${path_relative_to_include()}/lock"
|
||||
unlock_address = "${get_env("GITLAB_TF_STATE_URL", "https://tf.fxhnt.ai")}/api/v4/projects/1/terraform/state/${path_relative_to_include()}/lock"
|
||||
lock_method = "POST"
|
||||
unlock_method = "DELETE"
|
||||
bucket = "foxhunt-tfstate"
|
||||
key = "${path_relative_to_include()}/terraform.tfstate"
|
||||
region = "fr-par"
|
||||
|
||||
endpoints = {
|
||||
s3 = "https://s3.fr-par.scw.cloud"
|
||||
}
|
||||
|
||||
# Scaleway S3-compat: skip AWS-specific preflight calls
|
||||
skip_credentials_validation = true
|
||||
skip_region_validation = true
|
||||
skip_requesting_account_id = true
|
||||
skip_metadata_api_check = true
|
||||
|
||||
# OpenTofu-native lock (no DynamoDB)
|
||||
use_lockfile = true
|
||||
}
|
||||
}
|
||||
|
||||
|
||||
Reference in New Issue
Block a user