Files
foxhunt/docs/plans/2026-03-02-node-pool-split-implementation.md
jgrusewski 58f8478320 infra(tf): add platform pool, rename services → foxhunt in Kapsule module
Terraform changes for 3-pool node split:
- New `platform` pool resource (DEV1-L, databases + monitoring)
- Renamed `services` → `foxhunt` pool (DEV1-L, max_size 1)
- Updated variables, outputs, and live terragrunt inputs
- Updated implementation plan with Terraform-based workflow

Apply in 2 phases:
  Phase 1: terragrunt apply -target=scaleway_k8s_pool.platform
  Phase 2: terragrunt apply (after databases migrated to platform)

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-02 09:59:21 +01:00

28 KiB
Raw Blame History

3-Pool Node Split Implementation Plan

For Claude: REQUIRED SUB-SKILL: Use superpowers:executing-plans to implement this plan task-by-task.

Goal: Split the cluster into 3 isolated node pools (gitlab / platform / foxhunt) by creating a new platform pool, migrating infrastructure + monitoring there, and repurposing one services node as foxhunt.

Architecture: Create platform DEV1-L pool in Scaleway. Move stateless monitoring first (low risk), then stateful databases (requires PVC migration with backups), then rename services → foxhunt for app pods and drop the extra node.

Tech Stack: Scaleway Kapsule, Terragrunt/Terraform, kubectl, pg_dump/pg_restore, sed


Task 1: Create platform Node Pool via Terraform (Phase 1 — additive only)

Context: The cluster is managed by Terragrunt at infra/live/production/kapsule/. We need to create the platform pool first (additive) WITHOUT renaming servicesfoxhunt yet. The rename would destroy the existing services pool before databases are migrated. We do this in two Terraform applies.

Files already modified:

  • infra/modules/kapsule/main.tf — added platform pool resource, renamed servicesfoxhunt
  • infra/modules/kapsule/variables.tf — added platform_type, enable_platform_pool, renamed services_*foxhunt_*
  • infra/modules/kapsule/outputs.tf — added platform_pool_id, renamed services_pool_idfoxhunt_pool_id
  • infra/live/production/kapsule/terragrunt.hcl — updated inputs

Step 1: Run terragrunt plan to preview changes

cd infra/live/production/kapsule
terragrunt plan

Expected output should show:

  • scaleway_k8s_pool.platform[0]create (new platform pool)
  • scaleway_k8s_pool.servicesdestroy (old name being removed)
  • scaleway_k8s_pool.foxhuntcreate (new name for app services)

IMPORTANT: Do NOT apply yet — the destroy of services would evict all pods before databases are migrated.

Step 2: Use -target to apply ONLY the platform pool first

cd infra/live/production/kapsule
terragrunt apply -target=scaleway_k8s_pool.platform

This creates the platform pool without touching the existing services pool.

Step 3: Wait for platform node to become Ready

kubectl get nodes -l k8s.scaleway.com/pool-name=platform -w

Expected: Node shows STATUS=Ready within 2-3 minutes.

Step 4: Verify node resources

kubectl get node -l k8s.scaleway.com/pool-name=platform \
  -o custom-columns="NAME:.metadata.name,CPU:.status.allocatable.cpu,MEM:.status.allocatable.memory"

Expected: CPU ~3800m, MEM ~6450608Ki (DEV1-L allocatable).

Step 5: Commit Terraform changes

git add infra/modules/kapsule/main.tf infra/modules/kapsule/variables.tf \
        infra/modules/kapsule/outputs.tf infra/live/production/kapsule/terragrunt.hcl
git commit -m "infra(tf): add platform pool, rename services → foxhunt in Kapsule module"

Task 2: Move Stateless Monitoring to platform (Loki, Tempo, kube-state-metrics, Pushgateway)

Context: These pods are currently on the gitlab node (loki, tempo, kube-state-metrics) or services node (pushgateway). They have no persistent data that can't be recreated. Moving them is just a nodeSelector change.

Files:

  • Modify: infra/k8s/monitoring/loki.yaml:64
  • Modify: infra/k8s/monitoring/tempo.yaml:60
  • Modify: infra/k8s/monitoring/kube-state-metrics.yaml:118
  • Modify: infra/k8s/monitoring/pushgateway.yaml:20

Step 1: Change nodeSelectors in manifests

In each file, change k8s.scaleway.com/pool-name value to platform:

# loki.yaml line 64: gitlab → platform
sed -i 's/pool-name: gitlab/pool-name: platform/' infra/k8s/monitoring/loki.yaml

# tempo.yaml line 60: gitlab → platform
sed -i 's/pool-name: gitlab/pool-name: platform/' infra/k8s/monitoring/tempo.yaml

# kube-state-metrics.yaml line 118: gitlab → platform
sed -i 's/pool-name: gitlab/pool-name: platform/' infra/k8s/monitoring/kube-state-metrics.yaml

# pushgateway.yaml line 20: services → platform
sed -i 's/pool-name: services/pool-name: platform/' infra/k8s/monitoring/pushgateway.yaml

Step 2: Apply and restart

kubectl apply -f infra/k8s/monitoring/loki.yaml -n foxhunt
kubectl apply -f infra/k8s/monitoring/tempo.yaml -n foxhunt
kubectl apply -f infra/k8s/monitoring/kube-state-metrics.yaml -n foxhunt
kubectl apply -f infra/k8s/monitoring/pushgateway.yaml -n foxhunt

Step 3: Verify pods rescheduled to platform node

kubectl get pods -n foxhunt -l 'app.kubernetes.io/name in (loki,tempo,kube-state-metrics,pushgateway)' \
  -o custom-columns="POD:.metadata.name,NODE:.spec.nodeName,STATUS:.status.phase"

Expected: All 4 pods show a node name containing platform and STATUS=Running.

Step 4: Verify Loki is ingesting

GRAFANA_POD=$(kubectl get pod -n foxhunt -l app.kubernetes.io/name=grafana -o jsonpath='{.items[0].metadata.name}')
kubectl exec -n foxhunt "$GRAFANA_POD" -c grafana -- \
  curl -s "http://loki.foxhunt.svc.cluster.local:3100/loki/api/v1/labels" | head -5

Expected: JSON response with "status": "success" and label names.

Step 5: Commit

git add infra/k8s/monitoring/loki.yaml infra/k8s/monitoring/tempo.yaml \
        infra/k8s/monitoring/kube-state-metrics.yaml infra/k8s/monitoring/pushgateway.yaml
git commit -m "infra: move monitoring stack (loki, tempo, kube-state-metrics, pushgateway) to platform pool"

Task 3: Move Grafana to platform

Context: Grafana is deployed via Helm with values in grafana-values.yaml. It has a 2Gi PVC for dashboards/state, but all dashboards are provisioned from ConfigMaps so PVC data loss is acceptable. The simplest approach is to change the nodeSelector, delete the old PVC, and let a new one get created on the platform node.

Files:

  • Modify: infra/k8s/gitlab/grafana-values.yaml:7

Step 1: Change nodeSelector in Helm values

sed -i 's/pool-name: gitlab/pool-name: platform/' infra/k8s/gitlab/grafana-values.yaml

Verify the file now reads:

nodeSelector:
  k8s.scaleway.com/pool-name: platform

Step 2: Upgrade Grafana via Helm

helm upgrade grafana grafana/grafana \
  -n foxhunt \
  -f infra/k8s/gitlab/grafana-values.yaml \
  --reuse-values

Step 3: If pod is Pending due to PVC bound to old node

The grafana PVC is ReadWriteOnce and may be bound to the old gitlab node. If the pod is stuck Pending:

# Delete the old PVC — dashboards are all provisioned, no data loss
kubectl delete pvc grafana -n foxhunt
# Restart the deployment to trigger new PVC creation
kubectl rollout restart deployment grafana -n foxhunt

Step 4: Wait for Grafana to be ready

kubectl rollout status deployment grafana -n foxhunt --timeout=120s

Expected: deployment "grafana" successfully rolled out

Step 5: Verify Grafana is accessible and dashboards load

GRAFANA_POD=$(kubectl get pod -n foxhunt -l app.kubernetes.io/name=grafana -o jsonpath='{.items[0].metadata.name}')
ADMIN_PASS=$(kubectl get secret grafana -n foxhunt -o jsonpath='{.data.admin-password}' | base64 -d)
kubectl exec -n foxhunt "$GRAFANA_POD" -c grafana -- \
  curl -s -u "admin:${ADMIN_PASS}" http://localhost:3000/api/search?type=dash-db | python3 -c "import sys,json; print(f'{len(json.load(sys.stdin))} dashboards loaded')"

Expected: 13 dashboards loaded

Step 6: Commit

git add infra/k8s/gitlab/grafana-values.yaml
git commit -m "infra: move Grafana to platform pool"

Task 4: Move Tailscale Proxies to platform

Context: Two tailscale deployments: tailscale-gitlab-proxy (in gitlab/) handles SSH to GitLab, and tailscale-subnet-router (in tailscale/) provides network access. Both are stateless and can simply be moved.

Files:

  • Modify: infra/k8s/gitlab/tailscale-proxy.yaml:54
  • Modify: infra/k8s/tailscale/deployment.yaml:20

Step 1: Change nodeSelectors

sed -i 's/pool-name: gitlab/pool-name: platform/' infra/k8s/gitlab/tailscale-proxy.yaml
sed -i 's/pool-name: services/pool-name: platform/' infra/k8s/tailscale/deployment.yaml

Step 2: Apply both

kubectl apply -f infra/k8s/gitlab/tailscale-proxy.yaml -n foxhunt
kubectl apply -f infra/k8s/tailscale/deployment.yaml -n tailscale

Step 3: Verify pods moved

kubectl get pod -n foxhunt -l app.kubernetes.io/name=tailscale-gitlab-proxy \
  -o custom-columns="POD:.metadata.name,NODE:.spec.nodeName"
kubectl get pod -n tailscale -l app.kubernetes.io/name=tailscale-subnet-router \
  -o custom-columns="POD:.metadata.name,NODE:.spec.nodeName"

Expected: Both show node containing platform.

Step 4: Verify GitLab SSH still works through tailscale proxy

ssh -T git@100.90.76.85 -p 2222 2>&1 | head -3

Expected: Welcome to GitLab or similar greeting (not connection refused).

Step 5: Commit

git add infra/k8s/gitlab/tailscale-proxy.yaml infra/k8s/tailscale/deployment.yaml
git commit -m "infra: move tailscale proxies to platform pool"

Task 5: Migrate Redis to platform (Stateless Move)

Context: Redis is used for caching — no persistent data that can't be regenerated. Simply change nodeSelector and restart. No PVC migration needed.

Files:

  • Modify: infra/k8s/databases/redis.yaml:21

Step 1: Change nodeSelector

sed -i 's/pool-name: services/pool-name: platform/' infra/k8s/databases/redis.yaml

Step 2: Apply and wait

kubectl apply -f infra/k8s/databases/redis.yaml -n foxhunt
kubectl rollout status deployment redis -n foxhunt --timeout=60s

Expected: deployment "redis" successfully rolled out

Step 3: Verify Redis is responding

kubectl exec -n foxhunt deployment/redis -- redis-cli ping

Expected: PONG

Step 4: Commit

git add infra/k8s/databases/redis.yaml
git commit -m "infra: move Redis to platform pool"

Task 6: Migrate Postgres to platform (PVC Migration)

Context: Postgres has a 10Gi ReadWriteOnce PVC bound to the current services node. We need to back up the data, create a new PVC on the platform node, restore, and switch over. This causes a few minutes of downtime for postgres.

Files:

  • Modify: infra/k8s/databases/postgres.yaml:38
  • Modify: infra/k8s/gitlab/postgres-init.yaml:15

Step 1: Backup postgres data

# Get current postgres pod
PG_POD=$(kubectl get pod -n foxhunt -l app.kubernetes.io/name=postgres -o jsonpath='{.items[0].metadata.name}')

# Run pg_dumpall to capture everything
kubectl exec -n foxhunt "$PG_POD" -- pg_dumpall -U foxhunt > /tmp/foxhunt-pg-backup.sql
echo "Backup size: $(wc -c < /tmp/foxhunt-pg-backup.sql) bytes"

Expected: Non-zero file size (should be several MB with GitLab + foxhunt data).

Step 2: Scale postgres to 0

kubectl scale deployment postgres -n foxhunt --replicas=0
kubectl get pods -n foxhunt -l app.kubernetes.io/name=postgres

Expected: No pods running.

Step 3: Delete old PVC and create a new one

# Delete the PVC bound to the old node
kubectl delete pvc postgres-pvc -n foxhunt

# Apply the manifest again — it includes the PVC definition and will create a fresh one
# The new PVC will bind to the platform node when the pod schedules there

Step 4: Change nodeSelector in postgres manifest

sed -i 's/pool-name: services/pool-name: platform/' infra/k8s/databases/postgres.yaml

Step 5: Apply manifest and scale up

kubectl apply -f infra/k8s/databases/postgres.yaml -n foxhunt
# The deployment should automatically create 1 replica with the new nodeSelector
kubectl rollout status deployment postgres -n foxhunt --timeout=120s

Expected: Pod starts on the platform node with a fresh (empty) database.

Step 6: Restore data

PG_POD=$(kubectl get pod -n foxhunt -l app.kubernetes.io/name=postgres -o jsonpath='{.items[0].metadata.name}')

# Copy backup into pod and restore
kubectl cp /tmp/foxhunt-pg-backup.sql foxhunt/"$PG_POD":/tmp/backup.sql
kubectl exec -n foxhunt "$PG_POD" -- psql -U foxhunt -f /tmp/backup.sql postgres

Step 7: Verify postgres is healthy

kubectl exec -n foxhunt "$PG_POD" -- psql -U foxhunt -c "SELECT datname FROM pg_database;"

Expected: Lists databases including foxhunt, gitlabhq_production, etc.

Step 8: Update postgres-init nodeSelector too

sed -i 's/pool-name: services/pool-name: platform/' infra/k8s/gitlab/postgres-init.yaml

Step 9: Commit

git add infra/k8s/databases/postgres.yaml infra/k8s/gitlab/postgres-init.yaml
git commit -m "infra: migrate Postgres to platform pool with PVC migration"

Task 7: Migrate QuestDB to platform (PVC Migration)

Context: QuestDB has a 10Gi ReadWriteOnce PVC. It stores ML model metrics (predictions, trade outcomes). Data is nice-to-have but not critical — QuestDB will recreate tables on startup. We can do a simple file copy if desired, or just start fresh.

Files:

  • Modify: infra/k8s/databases/questdb.yaml:38

Step 1: Scale QuestDB to 0

kubectl scale deployment questdb -n foxhunt --replicas=0

Step 2: Delete old PVC

kubectl delete pvc questdb-pvc -n foxhunt

Step 3: Change nodeSelector

sed -i 's/pool-name: services/pool-name: platform/' infra/k8s/databases/questdb.yaml

Step 4: Apply and scale up

kubectl apply -f infra/k8s/databases/questdb.yaml -n foxhunt
kubectl rollout status deployment questdb -n foxhunt --timeout=120s

Expected: QuestDB starts fresh on the platform node.

Step 5: Verify QuestDB is responding

QUESTDB_POD=$(kubectl get pod -n foxhunt -l app.kubernetes.io/name=questdb -o jsonpath='{.items[0].metadata.name}')
kubectl exec -n foxhunt "$QUESTDB_POD" -- curl -s "http://localhost:9000/exec?query=SELECT+1" 2>/dev/null | head -3

Expected: JSON response with result.

Step 6: Commit

git add infra/k8s/databases/questdb.yaml
git commit -m "infra: migrate QuestDB to platform pool (fresh start)"

Task 8: Migrate MinIO to platform (PVC Migration)

Context: MinIO has a 20Gi PVC containing compiled binaries and trained models. We need to preserve this data. Strategy: use a temp pod to mount both old and new PVCs and copy data via rsync.

Files:

  • Modify: infra/k8s/minio/minio.yaml:57,152

Step 1: Record what's in MinIO

# Using minio mc client from the init job approach
MINIO_POD=$(kubectl get pod -n foxhunt -l app.kubernetes.io/name=minio -o jsonpath='{.items[0].metadata.name}')
kubectl exec -n foxhunt "$MINIO_POD" -- ls -la /data/

Note the contents for verification after migration.

Step 2: Scale MinIO to 0

kubectl scale deployment minio -n foxhunt --replicas=0

Step 3: Create a migration pod to copy data

cat <<'EOF' | kubectl apply -n foxhunt -f -
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
  name: minio-data-new
  namespace: foxhunt
  labels:
    app.kubernetes.io/name: minio
    app.kubernetes.io/part-of: foxhunt
spec:
  accessModes:
    - ReadWriteOnce
  storageClassName: scw-bssd
  resources:
    requests:
      storage: 20Gi
---
apiVersion: v1
kind: Pod
metadata:
  name: minio-migrate
  namespace: foxhunt
spec:
  nodeSelector:
    k8s.scaleway.com/pool-name: platform
  restartPolicy: Never
  containers:
    - name: migrate
      image: debian:bookworm-slim
      command: ["sleep", "3600"]
      volumeMounts:
        - name: new-data
          mountPath: /new
  volumes:
    - name: new-data
      persistentVolumeClaim:
        claimName: minio-data-new
EOF

Wait for the migration pod to be Running:

kubectl wait --for=condition=Ready pod/minio-migrate -n foxhunt --timeout=120s

Step 4: Copy data from old PVC to new PVC

Since the old PVC is ReadWriteOnce and bound to a services node, we can't mount it directly on the platform node. Instead, we'll use a second pod on the services node to tar the data and pipe it:

# Create a reader pod on the old node
cat <<'EOF' | kubectl apply -n foxhunt -f -
apiVersion: v1
kind: Pod
metadata:
  name: minio-reader
  namespace: foxhunt
spec:
  nodeSelector:
    k8s.scaleway.com/pool-name: services
  restartPolicy: Never
  containers:
    - name: reader
      image: debian:bookworm-slim
      command: ["sleep", "3600"]
      volumeMounts:
        - name: old-data
          mountPath: /old
          readOnly: true
  volumes:
    - name: old-data
      persistentVolumeClaim:
        claimName: minio-data
EOF

kubectl wait --for=condition=Ready pod/minio-reader -n foxhunt --timeout=120s

# Pipe tar from reader to migrate pod
kubectl exec -n foxhunt minio-reader -- tar cf - -C /old . | \
  kubectl exec -i -n foxhunt minio-migrate -- tar xf - -C /new

# Verify data arrived
kubectl exec -n foxhunt minio-migrate -- ls -la /new/

Expected: Same directory listing as Step 1.

Step 5: Clean up migration pods

kubectl delete pod minio-reader minio-migrate -n foxhunt

Step 6: Swap PVCs

# Delete old PVC
kubectl delete pvc minio-data -n foxhunt

# Rename new PVC (can't rename in k8s — need to update the manifest to use the new name,
# or recreate with the original name)
# Simplest: update minio.yaml to reference minio-data-new

Actually, the simplest approach: edit the minio manifest to use minio-data-new as the PVC name:

sed -i 's/claimName: minio-data/claimName: minio-data-new/' infra/k8s/minio/minio.yaml

Step 7: Change nodeSelectors

sed -i 's/pool-name: services/pool-name: platform/' infra/k8s/minio/minio.yaml

Step 8: Apply and start MinIO

kubectl apply -f infra/k8s/minio/minio.yaml -n foxhunt
kubectl rollout status deployment minio -n foxhunt --timeout=120s

Step 9: Verify MinIO has all data

MINIO_POD=$(kubectl get pod -n foxhunt -l app.kubernetes.io/name=minio -o jsonpath='{.items[0].metadata.name}')
kubectl exec -n foxhunt "$MINIO_POD" -- ls -la /data/

Expected: Same contents as Step 1.

Step 10: Commit

git add infra/k8s/minio/minio.yaml
git commit -m "infra: migrate MinIO to platform pool with data migration"

Task 9: Rename Services → Foxhunt Pool via Terraform (Phase 2 — destructive)

Context: All databases and monitoring are now on the platform node. The services pool only has foxhunt app pods + binary-cache PVCs (disposable). Now we run terragrunt apply to complete the rename: Terraform will destroy the old services pool (2 nodes) and create a new foxhunt pool (1 node). App pods will be evicted temporarily.

Files:

  • Modify: infra/k8s/services/api-gateway.yaml:39
  • Modify: infra/k8s/services/trading-service.yaml:39
  • Modify: infra/k8s/services/broker-gateway.yaml:39
  • Modify: infra/k8s/services/ib-gateway.yaml:26
  • Modify: infra/k8s/services/ml-training-service.yaml:82
  • Modify: infra/k8s/services/trading-agent-service.yaml:39
  • Modify: infra/k8s/services/backtesting-service.yaml:39
  • Modify: infra/k8s/services/data-acquisition-service.yaml:39
  • Modify: infra/k8s/services/web-gateway.yaml:39
  • Modify: infra/k8s/training/idle-reaper.yaml:58

Step 1: Update all foxhunt service manifests to target foxhunt pool

# All services currently say pool-name: services → change to pool-name: foxhunt
for f in infra/k8s/services/api-gateway.yaml \
         infra/k8s/services/trading-service.yaml \
         infra/k8s/services/broker-gateway.yaml \
         infra/k8s/services/ib-gateway.yaml \
         infra/k8s/services/ml-training-service.yaml \
         infra/k8s/services/trading-agent-service.yaml \
         infra/k8s/services/backtesting-service.yaml \
         infra/k8s/services/data-acquisition-service.yaml \
         infra/k8s/services/web-gateway.yaml \
         infra/k8s/training/idle-reaper.yaml; do
  sed -i 's/pool-name: services/pool-name: foxhunt/' "$f"
done

Step 2: Verify changes are correct

grep -r "pool-name: services" infra/k8s/services/ infra/k8s/training/idle-reaper.yaml

Expected: No output (all changed to foxhunt).

grep -r "pool-name: foxhunt" infra/k8s/services/ infra/k8s/training/idle-reaper.yaml | wc -l

Expected: 10 (9 services + 1 idle-reaper).

Step 3: Delete all binary-cache PVCs (they'll be recreated on the new node)

kubectl delete pvc -n foxhunt -l app.kubernetes.io/part-of=foxhunt --field-selector metadata.name!=postgres-pvc,metadata.name!=questdb-pvc,metadata.name!=minio-data-new 2>/dev/null || true

# More precisely — delete just the binary-cache PVCs
for SVC in api-gateway backtesting-service broker-gateway data-acquisition-service \
           ml-training-service trading-agent-service trading-service web-gateway; do
  kubectl delete pvc "binary-cache-${SVC}" -n foxhunt 2>/dev/null || true
done

Step 4: Run Terraform Phase 2 — destroy services pool, create foxhunt pool

cd infra/live/production/kapsule
terragrunt apply

This will:

  • Destroy scaleway_k8s_pool.services (both services nodes get deleted)
  • Create scaleway_k8s_pool.foxhunt (1× DEV1-L)

All pods on the old services nodes get evicted. Wait for the new foxhunt node to become Ready.

Step 5: Wait for new foxhunt node

kubectl get nodes -l k8s.scaleway.com/pool-name=foxhunt -w

Expected: 1 node with STATUS=Ready within 2-3 minutes.

Step 6: Apply updated K8s manifests

kubectl apply -f infra/k8s/services/ -n foxhunt
kubectl apply -f infra/k8s/training/idle-reaper.yaml -n foxhunt

Step 7: Wait for all foxhunt services to be Running

kubectl get pods -n foxhunt -l app.kubernetes.io/part-of=foxhunt -w

Expected: All 9 service pods + ib-gateway Running on the foxhunt node. Binary caches will be populated from S3 via initContainers.

Step 8: Commit

git add infra/k8s/services/ infra/k8s/training/idle-reaper.yaml
git commit -m "infra: rename services → foxhunt pool in all K8s manifests"

Task 10: Move idle-reaper to foxhunt pool

Context: The idle-reaper CronJob currently targets services. Since it manages GPU node scaling, it logically belongs on the foxhunt/services node. Since we kept the pool name as services, no actual change is needed. However, if we want to be explicit:

Files:

  • Verify: infra/k8s/training/idle-reaper.yaml:58 — already says services, no change needed

Step 1: Verify idle-reaper runs on the services (foxhunt) node

kubectl get cronjob gpu-idle-reaper -n foxhunt -o jsonpath='{.spec.jobTemplate.spec.template.spec.nodeSelector}'

Expected: {"k8s.scaleway.com/pool-name":"services"} — already correct.

No commit needed.


Task 11: Full Cluster Verification

Context: Verify the entire cluster is healthy after migration. Check every pool has the right pods, all services are communicating, dashboards work, CI works.

Step 1: Verify node pool distribution

echo "=== Pods by Node Pool ==="
for POOL in gitlab platform services; do
  echo ""
  echo "--- $POOL ---"
  for NODE in $(kubectl get nodes -l k8s.scaleway.com/pool-name="$POOL" -o jsonpath='{.items[*].metadata.name}'); do
    kubectl get pods -A --field-selector spec.nodeName="$NODE" --no-headers 2>/dev/null | \
      awk '{print $1"/"$2}' | sort
  done
done

Expected:

  • gitlab: Only gitlab-* pods + daemonsets (node-exporter, promtail)
  • platform: postgres, redis, minio, questdb, grafana, loki, tempo, kube-state-metrics, tailscale-*, pushgateway + daemonsets
  • services: Only foxhunt app pods (api-gateway, trading-service, etc.) + ib-gateway + daemonsets

Step 2: Verify Grafana dashboards

GRAFANA_POD=$(kubectl get pod -n foxhunt -l app.kubernetes.io/name=grafana -o jsonpath='{.items[0].metadata.name}')
ADMIN_PASS=$(kubectl get secret grafana -n foxhunt -o jsonpath='{.data.admin-password}' | base64 -d)
kubectl exec -n foxhunt "$GRAFANA_POD" -c grafana -- \
  curl -s -u "admin:${ADMIN_PASS}" http://localhost:3000/api/search?type=dash-db | \
  python3 -c "import sys,json; ds=json.load(sys.stdin); print(f'{len(ds)} dashboards'); [print(f'  {d[\"title\"]}') for d in ds]"

Expected: 13 dashboards listed.

Step 3: Verify Loki ingestion

kubectl exec -n foxhunt "$GRAFANA_POD" -c grafana -- \
  curl -s "http://loki.foxhunt.svc.cluster.local:3100/loki/api/v1/labels" | python3 -c "import sys,json; d=json.load(sys.stdin); print(f'Loki status={d[\"status\"]}, {len(d.get(\"data\",[]))} labels')"

Expected: Loki status=success, N labels

Step 4: Verify GitLab CI pipeline

Push a trivial commit or trigger a pipeline manually:

git push  # The commit from earlier tasks triggers CI

Check pipeline status in GitLab UI or:

curl -s --header "PRIVATE-TOKEN: $(cat ~/.gitlab-token)" \
  "http://100.90.76.85:2222/api/v4/projects/1/pipelines?per_page=1" | python3 -c "import sys,json; p=json.load(sys.stdin)[0]; print(f'Pipeline #{p[\"id\"]}: {p[\"status\"]}')"

Expected: Pipeline running or passed.

Step 5: Verify network connectivity between pools

# Test foxhunt services can reach platform databases
kubectl exec -n foxhunt deployment/trading-service -- curl -s http://redis.foxhunt.svc.cluster.local:6379/ping 2>/dev/null || \
  kubectl exec -n foxhunt deployment/trading-service -- nc -zv redis.foxhunt.svc.cluster.local 6379 2>&1

Expected: Connection successful (TCP to Redis on platform node).

Step 6: Check resource utilization on new layout

kubectl top nodes

Expected: 3 nodes visible (gitlab, platform, services) with reasonable utilization.

Step 7: Final commit

git add -A
git status  # Verify only expected changes
git commit -m "infra: complete 3-pool node split (gitlab/platform/foxhunt)

- Created platform pool (DEV1-L) for databases + monitoring
- Migrated Postgres (pg_dump/restore), MinIO (tar copy), QuestDB (fresh)
- Moved Loki, Tempo, Grafana, kube-state-metrics, pushgateway to platform
- Scaled services pool from 2 to 1 node (now foxhunt app services only)
- Same EUR 61/mo total, proper blast radius isolation"

Summary of Changes by File

File Change
infra/k8s/monitoring/loki.yaml gitlabplatform
infra/k8s/monitoring/tempo.yaml gitlabplatform
infra/k8s/monitoring/kube-state-metrics.yaml gitlabplatform
infra/k8s/monitoring/pushgateway.yaml servicesplatform
infra/k8s/gitlab/grafana-values.yaml gitlabplatform
infra/k8s/gitlab/tailscale-proxy.yaml gitlabplatform
infra/k8s/tailscale/deployment.yaml servicesplatform
infra/k8s/databases/redis.yaml servicesplatform
infra/k8s/databases/postgres.yaml servicesplatform
infra/k8s/databases/questdb.yaml servicesplatform
infra/k8s/minio/minio.yaml servicesplatform, PVC name update
infra/k8s/gitlab/postgres-init.yaml servicesplatform

Unchanged (stay on gitlab): values.yaml, runner-values.yaml, runner-rl-values.yaml, pat-rotation.yaml Changed servicesfoxhunt: All 9 foxhunt service YAMLs, idle-reaper.yaml Terraform: infra/modules/kapsule/{main,variables,outputs}.tf, infra/live/production/kapsule/terragrunt.hcl

Rollback Procedure

Before Phase 2 (services pool still exists):

  1. Revert platform nodeSelectors back to gitlab or services
  2. kubectl apply each file
  3. For stateful services (postgres, minio): restore from backups
  4. terragrunt destroy -target=scaleway_k8s_pool.platform to remove platform pool
  5. Revert Terraform changes in git

After Phase 2 (services pool destroyed, foxhunt pool created):

  1. Revert Terraform to recreate services pool: git revert the TF commits, terragrunt apply
  2. Revert all K8s manifests: pool-name: foxhuntpool-name: services
  3. Restore databases from backups
  4. Destroy platform pool