infra(argo): warmup-gpu exits immediately — relies on Scaleway 15min scaledown grace

Removed the 30s sleep. The warmup pod's purpose is to make the L40S
pool autoscaler scale 0 → 1; once the pod lands on the new node
(Scheduled → Running → Succeeded), the node enters Scaleway Kapsule's
scaledown-grace window (~15 min). That single window covers the
entire range of ensure-binary durations (sccache-hit ~10s through
cold ~15 min), so train always lands on a hot node without holding
the warmup pod open.

Trimmed cpu request 100m → 50m and mem 64Mi → 32Mi: the pod runs
~one shell command then exits; tiny resource footprint = faster
scheduling and no kubelet noise.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
This commit is contained in:
jgrusewski
2026-05-16 23:31:40 +02:00
parent 260b07c492
commit 29cf092551

View File

@@ -218,10 +218,11 @@ spec:
#
# Schedules a tiny CPU-only pod on the gpu-pool's nodeSelector.
# Kubernetes sees an unschedulable pod (no node currently in the
# pool), autoscaler scales the pool from 0 → 1. The pod sleeps
# briefly, exits; the node enters "empty" state with the
# autoscaler's scaledown grace window (~10 min on Scaleway), so
# train can land on it without provisioning latency.
# pool), autoscaler scales the pool 0 → 1. The pod exits the
# instant it lands; the node enters scaledown-grace state for
# ~15 min (Scaleway Kapsule default), which spans the entire
# ensure-binary compile window so train lands on a hot node with
# zero provisioning latency.
#
# No GPU resource request — that would compete with `train` for
# the single GPU and serialise the steps. The nodeSelector +
@@ -243,17 +244,14 @@ spec:
command: ["/bin/sh", "-c"]
resources:
requests:
cpu: "50m"
memory: 32Mi
limits:
cpu: "100m"
memory: 64Mi
limits:
cpu: "200m"
memory: 128Mi
args:
- |
echo "GPU warmup pod scheduled on $(hostname) — autoscaler scale-up triggered."
echo "Sleeping 30s to keep the node alive past provisioning; train will land here."
sleep 30
echo "Warmup complete; node remains in autoscaler scaledown grace window."
echo "GPU warmup pod scheduled on $(hostname) — autoscaler scale-up triggered, node now in scaledown-grace window."
# ── train: run alpha_train on L40S ──
- name: train