Terraform changes for 3-pool node split: - New `platform` pool resource (DEV1-L, databases + monitoring) - Renamed `services` → `foxhunt` pool (DEV1-L, max_size 1) - Updated variables, outputs, and live terragrunt inputs - Updated implementation plan with Terraform-based workflow Apply in 2 phases: Phase 1: terragrunt apply -target=scaleway_k8s_pool.platform Phase 2: terragrunt apply (after databases migrated to platform) Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
28 KiB
3-Pool Node Split Implementation Plan
For Claude: REQUIRED SUB-SKILL: Use superpowers:executing-plans to implement this plan task-by-task.
Goal: Split the cluster into 3 isolated node pools (gitlab / platform / foxhunt) by creating a new platform pool, migrating infrastructure + monitoring there, and repurposing one services node as foxhunt.
Architecture: Create platform DEV1-L pool in Scaleway. Move stateless monitoring first (low risk), then stateful databases (requires PVC migration with backups), then rename services → foxhunt for app pods and drop the extra node.
Tech Stack: Scaleway Kapsule, Terragrunt/Terraform, kubectl, pg_dump/pg_restore, sed
Task 1: Create platform Node Pool via Terraform (Phase 1 — additive only)
Context: The cluster is managed by Terragrunt at infra/live/production/kapsule/. We need to create the platform pool first (additive) WITHOUT renaming services → foxhunt yet. The rename would destroy the existing services pool before databases are migrated. We do this in two Terraform applies.
Files already modified:
infra/modules/kapsule/main.tf— addedplatformpool resource, renamedservices→foxhuntinfra/modules/kapsule/variables.tf— addedplatform_type,enable_platform_pool, renamedservices_*→foxhunt_*infra/modules/kapsule/outputs.tf— addedplatform_pool_id, renamedservices_pool_id→foxhunt_pool_idinfra/live/production/kapsule/terragrunt.hcl— updated inputs
Step 1: Run terragrunt plan to preview changes
cd infra/live/production/kapsule
terragrunt plan
Expected output should show:
scaleway_k8s_pool.platform[0]— create (new platform pool)scaleway_k8s_pool.services— destroy (old name being removed)scaleway_k8s_pool.foxhunt— create (new name for app services)
IMPORTANT: Do NOT apply yet — the destroy of services would evict all pods before databases are migrated.
Step 2: Use -target to apply ONLY the platform pool first
cd infra/live/production/kapsule
terragrunt apply -target=scaleway_k8s_pool.platform
This creates the platform pool without touching the existing services pool.
Step 3: Wait for platform node to become Ready
kubectl get nodes -l k8s.scaleway.com/pool-name=platform -w
Expected: Node shows STATUS=Ready within 2-3 minutes.
Step 4: Verify node resources
kubectl get node -l k8s.scaleway.com/pool-name=platform \
-o custom-columns="NAME:.metadata.name,CPU:.status.allocatable.cpu,MEM:.status.allocatable.memory"
Expected: CPU ~3800m, MEM ~6450608Ki (DEV1-L allocatable).
Step 5: Commit Terraform changes
git add infra/modules/kapsule/main.tf infra/modules/kapsule/variables.tf \
infra/modules/kapsule/outputs.tf infra/live/production/kapsule/terragrunt.hcl
git commit -m "infra(tf): add platform pool, rename services → foxhunt in Kapsule module"
Task 2: Move Stateless Monitoring to platform (Loki, Tempo, kube-state-metrics, Pushgateway)
Context: These pods are currently on the gitlab node (loki, tempo, kube-state-metrics) or services node (pushgateway). They have no persistent data that can't be recreated. Moving them is just a nodeSelector change.
Files:
- Modify:
infra/k8s/monitoring/loki.yaml:64 - Modify:
infra/k8s/monitoring/tempo.yaml:60 - Modify:
infra/k8s/monitoring/kube-state-metrics.yaml:118 - Modify:
infra/k8s/monitoring/pushgateway.yaml:20
Step 1: Change nodeSelectors in manifests
In each file, change k8s.scaleway.com/pool-name value to platform:
# loki.yaml line 64: gitlab → platform
sed -i 's/pool-name: gitlab/pool-name: platform/' infra/k8s/monitoring/loki.yaml
# tempo.yaml line 60: gitlab → platform
sed -i 's/pool-name: gitlab/pool-name: platform/' infra/k8s/monitoring/tempo.yaml
# kube-state-metrics.yaml line 118: gitlab → platform
sed -i 's/pool-name: gitlab/pool-name: platform/' infra/k8s/monitoring/kube-state-metrics.yaml
# pushgateway.yaml line 20: services → platform
sed -i 's/pool-name: services/pool-name: platform/' infra/k8s/monitoring/pushgateway.yaml
Step 2: Apply and restart
kubectl apply -f infra/k8s/monitoring/loki.yaml -n foxhunt
kubectl apply -f infra/k8s/monitoring/tempo.yaml -n foxhunt
kubectl apply -f infra/k8s/monitoring/kube-state-metrics.yaml -n foxhunt
kubectl apply -f infra/k8s/monitoring/pushgateway.yaml -n foxhunt
Step 3: Verify pods rescheduled to platform node
kubectl get pods -n foxhunt -l 'app.kubernetes.io/name in (loki,tempo,kube-state-metrics,pushgateway)' \
-o custom-columns="POD:.metadata.name,NODE:.spec.nodeName,STATUS:.status.phase"
Expected: All 4 pods show a node name containing platform and STATUS=Running.
Step 4: Verify Loki is ingesting
GRAFANA_POD=$(kubectl get pod -n foxhunt -l app.kubernetes.io/name=grafana -o jsonpath='{.items[0].metadata.name}')
kubectl exec -n foxhunt "$GRAFANA_POD" -c grafana -- \
curl -s "http://loki.foxhunt.svc.cluster.local:3100/loki/api/v1/labels" | head -5
Expected: JSON response with "status": "success" and label names.
Step 5: Commit
git add infra/k8s/monitoring/loki.yaml infra/k8s/monitoring/tempo.yaml \
infra/k8s/monitoring/kube-state-metrics.yaml infra/k8s/monitoring/pushgateway.yaml
git commit -m "infra: move monitoring stack (loki, tempo, kube-state-metrics, pushgateway) to platform pool"
Task 3: Move Grafana to platform
Context: Grafana is deployed via Helm with values in grafana-values.yaml. It has a 2Gi PVC for dashboards/state, but all dashboards are provisioned from ConfigMaps so PVC data loss is acceptable. The simplest approach is to change the nodeSelector, delete the old PVC, and let a new one get created on the platform node.
Files:
- Modify:
infra/k8s/gitlab/grafana-values.yaml:7
Step 1: Change nodeSelector in Helm values
sed -i 's/pool-name: gitlab/pool-name: platform/' infra/k8s/gitlab/grafana-values.yaml
Verify the file now reads:
nodeSelector:
k8s.scaleway.com/pool-name: platform
Step 2: Upgrade Grafana via Helm
helm upgrade grafana grafana/grafana \
-n foxhunt \
-f infra/k8s/gitlab/grafana-values.yaml \
--reuse-values
Step 3: If pod is Pending due to PVC bound to old node
The grafana PVC is ReadWriteOnce and may be bound to the old gitlab node. If the pod is stuck Pending:
# Delete the old PVC — dashboards are all provisioned, no data loss
kubectl delete pvc grafana -n foxhunt
# Restart the deployment to trigger new PVC creation
kubectl rollout restart deployment grafana -n foxhunt
Step 4: Wait for Grafana to be ready
kubectl rollout status deployment grafana -n foxhunt --timeout=120s
Expected: deployment "grafana" successfully rolled out
Step 5: Verify Grafana is accessible and dashboards load
GRAFANA_POD=$(kubectl get pod -n foxhunt -l app.kubernetes.io/name=grafana -o jsonpath='{.items[0].metadata.name}')
ADMIN_PASS=$(kubectl get secret grafana -n foxhunt -o jsonpath='{.data.admin-password}' | base64 -d)
kubectl exec -n foxhunt "$GRAFANA_POD" -c grafana -- \
curl -s -u "admin:${ADMIN_PASS}" http://localhost:3000/api/search?type=dash-db | python3 -c "import sys,json; print(f'{len(json.load(sys.stdin))} dashboards loaded')"
Expected: 13 dashboards loaded
Step 6: Commit
git add infra/k8s/gitlab/grafana-values.yaml
git commit -m "infra: move Grafana to platform pool"
Task 4: Move Tailscale Proxies to platform
Context: Two tailscale deployments: tailscale-gitlab-proxy (in gitlab/) handles SSH to GitLab, and tailscale-subnet-router (in tailscale/) provides network access. Both are stateless and can simply be moved.
Files:
- Modify:
infra/k8s/gitlab/tailscale-proxy.yaml:54 - Modify:
infra/k8s/tailscale/deployment.yaml:20
Step 1: Change nodeSelectors
sed -i 's/pool-name: gitlab/pool-name: platform/' infra/k8s/gitlab/tailscale-proxy.yaml
sed -i 's/pool-name: services/pool-name: platform/' infra/k8s/tailscale/deployment.yaml
Step 2: Apply both
kubectl apply -f infra/k8s/gitlab/tailscale-proxy.yaml -n foxhunt
kubectl apply -f infra/k8s/tailscale/deployment.yaml -n tailscale
Step 3: Verify pods moved
kubectl get pod -n foxhunt -l app.kubernetes.io/name=tailscale-gitlab-proxy \
-o custom-columns="POD:.metadata.name,NODE:.spec.nodeName"
kubectl get pod -n tailscale -l app.kubernetes.io/name=tailscale-subnet-router \
-o custom-columns="POD:.metadata.name,NODE:.spec.nodeName"
Expected: Both show node containing platform.
Step 4: Verify GitLab SSH still works through tailscale proxy
ssh -T git@100.90.76.85 -p 2222 2>&1 | head -3
Expected: Welcome to GitLab or similar greeting (not connection refused).
Step 5: Commit
git add infra/k8s/gitlab/tailscale-proxy.yaml infra/k8s/tailscale/deployment.yaml
git commit -m "infra: move tailscale proxies to platform pool"
Task 5: Migrate Redis to platform (Stateless Move)
Context: Redis is used for caching — no persistent data that can't be regenerated. Simply change nodeSelector and restart. No PVC migration needed.
Files:
- Modify:
infra/k8s/databases/redis.yaml:21
Step 1: Change nodeSelector
sed -i 's/pool-name: services/pool-name: platform/' infra/k8s/databases/redis.yaml
Step 2: Apply and wait
kubectl apply -f infra/k8s/databases/redis.yaml -n foxhunt
kubectl rollout status deployment redis -n foxhunt --timeout=60s
Expected: deployment "redis" successfully rolled out
Step 3: Verify Redis is responding
kubectl exec -n foxhunt deployment/redis -- redis-cli ping
Expected: PONG
Step 4: Commit
git add infra/k8s/databases/redis.yaml
git commit -m "infra: move Redis to platform pool"
Task 6: Migrate Postgres to platform (PVC Migration)
Context: Postgres has a 10Gi ReadWriteOnce PVC bound to the current services node. We need to back up the data, create a new PVC on the platform node, restore, and switch over. This causes a few minutes of downtime for postgres.
Files:
- Modify:
infra/k8s/databases/postgres.yaml:38 - Modify:
infra/k8s/gitlab/postgres-init.yaml:15
Step 1: Backup postgres data
# Get current postgres pod
PG_POD=$(kubectl get pod -n foxhunt -l app.kubernetes.io/name=postgres -o jsonpath='{.items[0].metadata.name}')
# Run pg_dumpall to capture everything
kubectl exec -n foxhunt "$PG_POD" -- pg_dumpall -U foxhunt > /tmp/foxhunt-pg-backup.sql
echo "Backup size: $(wc -c < /tmp/foxhunt-pg-backup.sql) bytes"
Expected: Non-zero file size (should be several MB with GitLab + foxhunt data).
Step 2: Scale postgres to 0
kubectl scale deployment postgres -n foxhunt --replicas=0
kubectl get pods -n foxhunt -l app.kubernetes.io/name=postgres
Expected: No pods running.
Step 3: Delete old PVC and create a new one
# Delete the PVC bound to the old node
kubectl delete pvc postgres-pvc -n foxhunt
# Apply the manifest again — it includes the PVC definition and will create a fresh one
# The new PVC will bind to the platform node when the pod schedules there
Step 4: Change nodeSelector in postgres manifest
sed -i 's/pool-name: services/pool-name: platform/' infra/k8s/databases/postgres.yaml
Step 5: Apply manifest and scale up
kubectl apply -f infra/k8s/databases/postgres.yaml -n foxhunt
# The deployment should automatically create 1 replica with the new nodeSelector
kubectl rollout status deployment postgres -n foxhunt --timeout=120s
Expected: Pod starts on the platform node with a fresh (empty) database.
Step 6: Restore data
PG_POD=$(kubectl get pod -n foxhunt -l app.kubernetes.io/name=postgres -o jsonpath='{.items[0].metadata.name}')
# Copy backup into pod and restore
kubectl cp /tmp/foxhunt-pg-backup.sql foxhunt/"$PG_POD":/tmp/backup.sql
kubectl exec -n foxhunt "$PG_POD" -- psql -U foxhunt -f /tmp/backup.sql postgres
Step 7: Verify postgres is healthy
kubectl exec -n foxhunt "$PG_POD" -- psql -U foxhunt -c "SELECT datname FROM pg_database;"
Expected: Lists databases including foxhunt, gitlabhq_production, etc.
Step 8: Update postgres-init nodeSelector too
sed -i 's/pool-name: services/pool-name: platform/' infra/k8s/gitlab/postgres-init.yaml
Step 9: Commit
git add infra/k8s/databases/postgres.yaml infra/k8s/gitlab/postgres-init.yaml
git commit -m "infra: migrate Postgres to platform pool with PVC migration"
Task 7: Migrate QuestDB to platform (PVC Migration)
Context: QuestDB has a 10Gi ReadWriteOnce PVC. It stores ML model metrics (predictions, trade outcomes). Data is nice-to-have but not critical — QuestDB will recreate tables on startup. We can do a simple file copy if desired, or just start fresh.
Files:
- Modify:
infra/k8s/databases/questdb.yaml:38
Step 1: Scale QuestDB to 0
kubectl scale deployment questdb -n foxhunt --replicas=0
Step 2: Delete old PVC
kubectl delete pvc questdb-pvc -n foxhunt
Step 3: Change nodeSelector
sed -i 's/pool-name: services/pool-name: platform/' infra/k8s/databases/questdb.yaml
Step 4: Apply and scale up
kubectl apply -f infra/k8s/databases/questdb.yaml -n foxhunt
kubectl rollout status deployment questdb -n foxhunt --timeout=120s
Expected: QuestDB starts fresh on the platform node.
Step 5: Verify QuestDB is responding
QUESTDB_POD=$(kubectl get pod -n foxhunt -l app.kubernetes.io/name=questdb -o jsonpath='{.items[0].metadata.name}')
kubectl exec -n foxhunt "$QUESTDB_POD" -- curl -s "http://localhost:9000/exec?query=SELECT+1" 2>/dev/null | head -3
Expected: JSON response with result.
Step 6: Commit
git add infra/k8s/databases/questdb.yaml
git commit -m "infra: migrate QuestDB to platform pool (fresh start)"
Task 8: Migrate MinIO to platform (PVC Migration)
Context: MinIO has a 20Gi PVC containing compiled binaries and trained models. We need to preserve this data. Strategy: use a temp pod to mount both old and new PVCs and copy data via rsync.
Files:
- Modify:
infra/k8s/minio/minio.yaml:57,152
Step 1: Record what's in MinIO
# Using minio mc client from the init job approach
MINIO_POD=$(kubectl get pod -n foxhunt -l app.kubernetes.io/name=minio -o jsonpath='{.items[0].metadata.name}')
kubectl exec -n foxhunt "$MINIO_POD" -- ls -la /data/
Note the contents for verification after migration.
Step 2: Scale MinIO to 0
kubectl scale deployment minio -n foxhunt --replicas=0
Step 3: Create a migration pod to copy data
cat <<'EOF' | kubectl apply -n foxhunt -f -
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: minio-data-new
namespace: foxhunt
labels:
app.kubernetes.io/name: minio
app.kubernetes.io/part-of: foxhunt
spec:
accessModes:
- ReadWriteOnce
storageClassName: scw-bssd
resources:
requests:
storage: 20Gi
---
apiVersion: v1
kind: Pod
metadata:
name: minio-migrate
namespace: foxhunt
spec:
nodeSelector:
k8s.scaleway.com/pool-name: platform
restartPolicy: Never
containers:
- name: migrate
image: debian:bookworm-slim
command: ["sleep", "3600"]
volumeMounts:
- name: new-data
mountPath: /new
volumes:
- name: new-data
persistentVolumeClaim:
claimName: minio-data-new
EOF
Wait for the migration pod to be Running:
kubectl wait --for=condition=Ready pod/minio-migrate -n foxhunt --timeout=120s
Step 4: Copy data from old PVC to new PVC
Since the old PVC is ReadWriteOnce and bound to a services node, we can't mount it directly on the platform node. Instead, we'll use a second pod on the services node to tar the data and pipe it:
# Create a reader pod on the old node
cat <<'EOF' | kubectl apply -n foxhunt -f -
apiVersion: v1
kind: Pod
metadata:
name: minio-reader
namespace: foxhunt
spec:
nodeSelector:
k8s.scaleway.com/pool-name: services
restartPolicy: Never
containers:
- name: reader
image: debian:bookworm-slim
command: ["sleep", "3600"]
volumeMounts:
- name: old-data
mountPath: /old
readOnly: true
volumes:
- name: old-data
persistentVolumeClaim:
claimName: minio-data
EOF
kubectl wait --for=condition=Ready pod/minio-reader -n foxhunt --timeout=120s
# Pipe tar from reader to migrate pod
kubectl exec -n foxhunt minio-reader -- tar cf - -C /old . | \
kubectl exec -i -n foxhunt minio-migrate -- tar xf - -C /new
# Verify data arrived
kubectl exec -n foxhunt minio-migrate -- ls -la /new/
Expected: Same directory listing as Step 1.
Step 5: Clean up migration pods
kubectl delete pod minio-reader minio-migrate -n foxhunt
Step 6: Swap PVCs
# Delete old PVC
kubectl delete pvc minio-data -n foxhunt
# Rename new PVC (can't rename in k8s — need to update the manifest to use the new name,
# or recreate with the original name)
# Simplest: update minio.yaml to reference minio-data-new
Actually, the simplest approach: edit the minio manifest to use minio-data-new as the PVC name:
sed -i 's/claimName: minio-data/claimName: minio-data-new/' infra/k8s/minio/minio.yaml
Step 7: Change nodeSelectors
sed -i 's/pool-name: services/pool-name: platform/' infra/k8s/minio/minio.yaml
Step 8: Apply and start MinIO
kubectl apply -f infra/k8s/minio/minio.yaml -n foxhunt
kubectl rollout status deployment minio -n foxhunt --timeout=120s
Step 9: Verify MinIO has all data
MINIO_POD=$(kubectl get pod -n foxhunt -l app.kubernetes.io/name=minio -o jsonpath='{.items[0].metadata.name}')
kubectl exec -n foxhunt "$MINIO_POD" -- ls -la /data/
Expected: Same contents as Step 1.
Step 10: Commit
git add infra/k8s/minio/minio.yaml
git commit -m "infra: migrate MinIO to platform pool with data migration"
Task 9: Rename Services → Foxhunt Pool via Terraform (Phase 2 — destructive)
Context: All databases and monitoring are now on the platform node. The services pool only has foxhunt app pods + binary-cache PVCs (disposable). Now we run terragrunt apply to complete the rename: Terraform will destroy the old services pool (2 nodes) and create a new foxhunt pool (1 node). App pods will be evicted temporarily.
Files:
- Modify:
infra/k8s/services/api-gateway.yaml:39 - Modify:
infra/k8s/services/trading-service.yaml:39 - Modify:
infra/k8s/services/broker-gateway.yaml:39 - Modify:
infra/k8s/services/ib-gateway.yaml:26 - Modify:
infra/k8s/services/ml-training-service.yaml:82 - Modify:
infra/k8s/services/trading-agent-service.yaml:39 - Modify:
infra/k8s/services/backtesting-service.yaml:39 - Modify:
infra/k8s/services/data-acquisition-service.yaml:39 - Modify:
infra/k8s/services/web-gateway.yaml:39 - Modify:
infra/k8s/training/idle-reaper.yaml:58
Step 1: Update all foxhunt service manifests to target foxhunt pool
# All services currently say pool-name: services → change to pool-name: foxhunt
for f in infra/k8s/services/api-gateway.yaml \
infra/k8s/services/trading-service.yaml \
infra/k8s/services/broker-gateway.yaml \
infra/k8s/services/ib-gateway.yaml \
infra/k8s/services/ml-training-service.yaml \
infra/k8s/services/trading-agent-service.yaml \
infra/k8s/services/backtesting-service.yaml \
infra/k8s/services/data-acquisition-service.yaml \
infra/k8s/services/web-gateway.yaml \
infra/k8s/training/idle-reaper.yaml; do
sed -i 's/pool-name: services/pool-name: foxhunt/' "$f"
done
Step 2: Verify changes are correct
grep -r "pool-name: services" infra/k8s/services/ infra/k8s/training/idle-reaper.yaml
Expected: No output (all changed to foxhunt).
grep -r "pool-name: foxhunt" infra/k8s/services/ infra/k8s/training/idle-reaper.yaml | wc -l
Expected: 10 (9 services + 1 idle-reaper).
Step 3: Delete all binary-cache PVCs (they'll be recreated on the new node)
kubectl delete pvc -n foxhunt -l app.kubernetes.io/part-of=foxhunt --field-selector metadata.name!=postgres-pvc,metadata.name!=questdb-pvc,metadata.name!=minio-data-new 2>/dev/null || true
# More precisely — delete just the binary-cache PVCs
for SVC in api-gateway backtesting-service broker-gateway data-acquisition-service \
ml-training-service trading-agent-service trading-service web-gateway; do
kubectl delete pvc "binary-cache-${SVC}" -n foxhunt 2>/dev/null || true
done
Step 4: Run Terraform Phase 2 — destroy services pool, create foxhunt pool
cd infra/live/production/kapsule
terragrunt apply
This will:
- Destroy
scaleway_k8s_pool.services(both services nodes get deleted) - Create
scaleway_k8s_pool.foxhunt(1× DEV1-L)
All pods on the old services nodes get evicted. Wait for the new foxhunt node to become Ready.
Step 5: Wait for new foxhunt node
kubectl get nodes -l k8s.scaleway.com/pool-name=foxhunt -w
Expected: 1 node with STATUS=Ready within 2-3 minutes.
Step 6: Apply updated K8s manifests
kubectl apply -f infra/k8s/services/ -n foxhunt
kubectl apply -f infra/k8s/training/idle-reaper.yaml -n foxhunt
Step 7: Wait for all foxhunt services to be Running
kubectl get pods -n foxhunt -l app.kubernetes.io/part-of=foxhunt -w
Expected: All 9 service pods + ib-gateway Running on the foxhunt node. Binary caches will be populated from S3 via initContainers.
Step 8: Commit
git add infra/k8s/services/ infra/k8s/training/idle-reaper.yaml
git commit -m "infra: rename services → foxhunt pool in all K8s manifests"
Task 10: Move idle-reaper to foxhunt pool
Context: The idle-reaper CronJob currently targets services. Since it manages GPU node scaling, it logically belongs on the foxhunt/services node. Since we kept the pool name as services, no actual change is needed. However, if we want to be explicit:
Files:
- Verify:
infra/k8s/training/idle-reaper.yaml:58— already saysservices, no change needed
Step 1: Verify idle-reaper runs on the services (foxhunt) node
kubectl get cronjob gpu-idle-reaper -n foxhunt -o jsonpath='{.spec.jobTemplate.spec.template.spec.nodeSelector}'
Expected: {"k8s.scaleway.com/pool-name":"services"} — already correct.
No commit needed.
Task 11: Full Cluster Verification
Context: Verify the entire cluster is healthy after migration. Check every pool has the right pods, all services are communicating, dashboards work, CI works.
Step 1: Verify node pool distribution
echo "=== Pods by Node Pool ==="
for POOL in gitlab platform services; do
echo ""
echo "--- $POOL ---"
for NODE in $(kubectl get nodes -l k8s.scaleway.com/pool-name="$POOL" -o jsonpath='{.items[*].metadata.name}'); do
kubectl get pods -A --field-selector spec.nodeName="$NODE" --no-headers 2>/dev/null | \
awk '{print $1"/"$2}' | sort
done
done
Expected:
gitlab: Only gitlab-* pods + daemonsets (node-exporter, promtail)platform: postgres, redis, minio, questdb, grafana, loki, tempo, kube-state-metrics, tailscale-*, pushgateway + daemonsetsservices: Only foxhunt app pods (api-gateway, trading-service, etc.) + ib-gateway + daemonsets
Step 2: Verify Grafana dashboards
GRAFANA_POD=$(kubectl get pod -n foxhunt -l app.kubernetes.io/name=grafana -o jsonpath='{.items[0].metadata.name}')
ADMIN_PASS=$(kubectl get secret grafana -n foxhunt -o jsonpath='{.data.admin-password}' | base64 -d)
kubectl exec -n foxhunt "$GRAFANA_POD" -c grafana -- \
curl -s -u "admin:${ADMIN_PASS}" http://localhost:3000/api/search?type=dash-db | \
python3 -c "import sys,json; ds=json.load(sys.stdin); print(f'{len(ds)} dashboards'); [print(f' {d[\"title\"]}') for d in ds]"
Expected: 13 dashboards listed.
Step 3: Verify Loki ingestion
kubectl exec -n foxhunt "$GRAFANA_POD" -c grafana -- \
curl -s "http://loki.foxhunt.svc.cluster.local:3100/loki/api/v1/labels" | python3 -c "import sys,json; d=json.load(sys.stdin); print(f'Loki status={d[\"status\"]}, {len(d.get(\"data\",[]))} labels')"
Expected: Loki status=success, N labels
Step 4: Verify GitLab CI pipeline
Push a trivial commit or trigger a pipeline manually:
git push # The commit from earlier tasks triggers CI
Check pipeline status in GitLab UI or:
curl -s --header "PRIVATE-TOKEN: $(cat ~/.gitlab-token)" \
"http://100.90.76.85:2222/api/v4/projects/1/pipelines?per_page=1" | python3 -c "import sys,json; p=json.load(sys.stdin)[0]; print(f'Pipeline #{p[\"id\"]}: {p[\"status\"]}')"
Expected: Pipeline running or passed.
Step 5: Verify network connectivity between pools
# Test foxhunt services can reach platform databases
kubectl exec -n foxhunt deployment/trading-service -- curl -s http://redis.foxhunt.svc.cluster.local:6379/ping 2>/dev/null || \
kubectl exec -n foxhunt deployment/trading-service -- nc -zv redis.foxhunt.svc.cluster.local 6379 2>&1
Expected: Connection successful (TCP to Redis on platform node).
Step 6: Check resource utilization on new layout
kubectl top nodes
Expected: 3 nodes visible (gitlab, platform, services) with reasonable utilization.
Step 7: Final commit
git add -A
git status # Verify only expected changes
git commit -m "infra: complete 3-pool node split (gitlab/platform/foxhunt)
- Created platform pool (DEV1-L) for databases + monitoring
- Migrated Postgres (pg_dump/restore), MinIO (tar copy), QuestDB (fresh)
- Moved Loki, Tempo, Grafana, kube-state-metrics, pushgateway to platform
- Scaled services pool from 2 to 1 node (now foxhunt app services only)
- Same EUR 61/mo total, proper blast radius isolation"
Summary of Changes by File
| File | Change |
|---|---|
infra/k8s/monitoring/loki.yaml |
gitlab → platform |
infra/k8s/monitoring/tempo.yaml |
gitlab → platform |
infra/k8s/monitoring/kube-state-metrics.yaml |
gitlab → platform |
infra/k8s/monitoring/pushgateway.yaml |
services → platform |
infra/k8s/gitlab/grafana-values.yaml |
gitlab → platform |
infra/k8s/gitlab/tailscale-proxy.yaml |
gitlab → platform |
infra/k8s/tailscale/deployment.yaml |
services → platform |
infra/k8s/databases/redis.yaml |
services → platform |
infra/k8s/databases/postgres.yaml |
services → platform |
infra/k8s/databases/questdb.yaml |
services → platform |
infra/k8s/minio/minio.yaml |
services → platform, PVC name update |
infra/k8s/gitlab/postgres-init.yaml |
services → platform |
Unchanged (stay on gitlab): values.yaml, runner-values.yaml, runner-rl-values.yaml, pat-rotation.yaml
Changed services → foxhunt: All 9 foxhunt service YAMLs, idle-reaper.yaml
Terraform: infra/modules/kapsule/{main,variables,outputs}.tf, infra/live/production/kapsule/terragrunt.hcl
Rollback Procedure
Before Phase 2 (services pool still exists):
- Revert
platformnodeSelectors back togitlaborservices kubectl applyeach file- For stateful services (postgres, minio): restore from backups
terragrunt destroy -target=scaleway_k8s_pool.platformto remove platform pool- Revert Terraform changes in git
After Phase 2 (services pool destroyed, foxhunt pool created):
- Revert Terraform to recreate
servicespool:git revertthe TF commits,terragrunt apply - Revert all K8s manifests:
pool-name: foxhunt→pool-name: services - Restore databases from backups
- Destroy platform pool