Prometheus was configured to scrape pods via prometheus.io/scrape annotations, but all services and training jobs used gitlab.com/ prefix from legacy GitLab-managed Prometheus — causing Prometheus to never discover any foxhunt pods. This resulted in stale/missing metrics on the Grafana training dashboard. 12 files updated across services/, gpu-overlays/, training/, and monitoring/ (node-exporter, dcgm-exporter cleanup). Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
4.8 KiB
4.8 KiB