docs: 3-pool node split design (gitlab/platform/foxhunt)

Isolate platform infra (postgres, redis, minio, monitoring) from foxhunt
app services on separate DEV1-L nodes. Same EUR 61/mo total cost, better
blast radius containment and burst headroom for trading services.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
This commit is contained in:
jgrusewski
2026-03-02 09:49:46 +01:00
parent d87c1cf21f
commit 4da3a16bfe

View File

@@ -0,0 +1,140 @@
# 3-Pool Node Split Design
## Problem
All foxhunt microservices, infrastructure databases, and monitoring share two DEV1-L nodes
in a single `services` pool. At idle, CPU utilization is 4-5% and memory 22-33%. There is
no isolation between platform infrastructure (postgres, redis) and application services
(trading-service, ib-gateway). A misbehaving foxhunt service could starve the database.
Meanwhile, monitoring (grafana, loki, tempo) sits on the gitlab node, consuming ~1 GiB RAM
that GitLab could use for Gitaly caching.
## Decision
Split into 3 purpose-specific node pools. Same total cost (EUR 61/mo), better isolation.
## Layout
| Pool | Instance | vCPU | RAM | Cost/mo | Workload |
|------|----------|------|-----|---------|----------|
| `gitlab` | GP1-XS | 4 | 16G | EUR 27 | GitLab CE (webservice, sidekiq ×2, gitaly, shell ×2, registry, toolbox, prometheus, runners ×2, pat-rotation) |
| `platform` | DEV1-L | 4 | 8G | EUR 17 | Postgres, Redis, MinIO, QuestDB, Grafana, Loki, Tempo, kube-state-metrics, tailscale-proxy |
| `foxhunt` | DEV1-L | 4 | 8G | EUR 17 | api-gateway, trading-service, broker-gateway, ib-gateway, ml-training-service, trading-agent-service, backtesting-service, data-acquisition-service, web-gateway |
CI pools (ci-compile-cpu, ci-training) are unchanged — autoscale 0→1.
## Resource Budget
### Platform node (DEV1-L: 3800m CPU / 6.4 GiB allocatable)
| Pod | CPU req | MEM req | MEM actual |
|-----|---------|---------|------------|
| postgres | 200m | 512Mi | ~800Mi |
| questdb | 200m | 512Mi | ~300Mi |
| redis | 100m | 256Mi | ~100Mi |
| minio | 100m | 256Mi | ~200Mi |
| grafana | — | — | ~190Mi |
| loki | 100m | 128Mi | ~140Mi |
| tempo | 100m | 256Mi | ~630Mi |
| kube-state-metrics | 25m | 64Mi | ~17Mi |
| tailscale-proxy | 25m | 32Mi | ~33Mi |
| **Total** | **~850m** | **~2 GiB** | **~2.4 GiB** |
Headroom: 3 GiB CPU, 4 GiB memory — comfortable.
### Foxhunt node (DEV1-L: 3800m CPU / 6.4 GiB allocatable)
| Pod | CPU req | MEM req | MEM actual |
|-----|---------|---------|------------|
| ib-gateway | 500m | 1Gi | ~500Mi |
| api-gateway | 200m | 256Mi | ~200Mi |
| trading-service | 200m | 256Mi | ~300Mi |
| broker-gateway | 200m | 256Mi | ~150Mi |
| backtesting-service | 200m | 256Mi | ~200Mi |
| trading-agent-service | 200m | 256Mi | ~200Mi |
| ml-training-service | 200m | 256Mi | ~200Mi |
| data-acquisition-service | 100m | 256Mi | ~150Mi |
| web-gateway | 100m | 128Mi | ~100Mi |
| **Total** | **~1.9G** | **~2.9 GiB** | **~2 GiB** |
Headroom: 1.9 GiB CPU, 3.5 GiB memory — room for trading burst.
## Migration Strategy
Brief downtime acceptable (no live trading). Order matters — databases first, then dependents.
### Phase 1: Create `platform` node pool
- Scaleway console or Terraform: DEV1-L, min=1, max=1, autoscale=off
- Wait for node to join cluster and become Ready
### Phase 2: Move stateless monitoring to `platform`
- Change nodeSelector from `gitlab` to `platform` for: grafana, loki, tempo, kube-state-metrics, tailscale-proxy
- `kubectl rollout restart` each — they reschedule to the platform node
- Verify Grafana accessible, Loki ingesting, Tempo receiving traces
### Phase 3: Move stateful infrastructure to `platform`
For each of postgres, questdb, minio (in this order):
1. Scale deployment to 0
2. Take backup (pg_dump for postgres, mc mirror for minio, questdb backup)
3. Delete old PVC (bound to old services node)
4. Create new PVC (will bind to platform node)
5. Restore data to new PVC via a temporary pod
6. Update deployment nodeSelector to `platform`
7. Scale back to 1
8. Verify service healthy
Redis: no persistent data that matters — just change nodeSelector and restart.
### Phase 4: Move foxhunt services
- Change nodeSelector from `services` to `foxhunt` for all 9 foxhunt deployments
- Delete binary-cache PVCs (disposable — services re-fetch from S3 on next init)
- `kubectl rollout restart` all services
- Verify all pods Running, health checks passing
### Phase 5: Rename/reconfigure services pool
- Option A: Rename existing `services` pool to `foxhunt` in Scaleway, scale to 1 node
- Option B: Create new `foxhunt` pool, drain old `services` pool, delete it
- Delete extra services node (was 2 nodes, now need 1)
### Phase 6: Verify
- All pods Running on correct nodes
- Grafana dashboards loading
- Loki queries working
- Network policies still enforced
- GitLab CI pipeline runs successfully
## Network Policies
No changes needed. All network policies use pod label selectors (`app.kubernetes.io/name`,
`app.kubernetes.io/part-of`), not node selectors. Cross-node pod communication works
transparently via Kapsule CNI (Cilium).
## PVC Handling
| PVC | Action |
|-----|--------|
| postgres-pvc (10Gi) | Migrate: pg_dump → new PVC → pg_restore |
| questdb-pvc (10Gi) | Migrate: copy data dir via temp pod |
| minio-data (20Gi) | Migrate: mc mirror to new PVC |
| grafana (2Gi) | Migrate: copy data dir (dashboards are provisioned, so loss is acceptable) |
| binary-cache-* (1Gi each, ×8) | Delete and recreate — populated from S3 on pod init |
| sccache/training PVCs | Unchanged — bound to CI/training pools |
## Risks
| Risk | Likelihood | Impact | Mitigation |
|------|-----------|--------|------------|
| PVC data loss | Low | High | pg_dump + mc mirror before any PVC deletion |
| Pod stuck Pending | Low | Medium | Verify node Ready before migration |
| Memory pressure on platform | Low | Medium | Tempo is the heaviest at 630Mi — can add limits |
| Cross-node latency | Negligible | Low | All nodes in same datacenter, <1ms |
| Scaleway pool creation fails | Low | None | Can abort, everything stays on old nodes |
## Rollback
If anything goes wrong mid-migration:
- Stateful services: restore from backup to original PVCs
- Stateless services: revert nodeSelector to `services`
- Delete the new `platform` pool
- Everything returns to pre-migration state