Replace standalone Prometheus/node-exporter/kube-state-metrics with kube-prometheus-stack Helm chart. Add ServiceMonitor CRDs, PrometheusRule CRDs (HFT + broker alerts), comprehensive network policies, and wire AlertManager to Mattermost via incoming webhook. - Delete old standalone prometheus.yaml, node-exporter.yaml, kube-state-metrics.yaml - Add prometheus-stack-values.yaml (Helm values with Scaleway Kapsule tuning) - Add service-monitors.yaml (foxhunt-services + dcgm-exporter ServiceMonitors) - Add prometheus-rules-hft.yaml and prometheus-rules-broker.yaml (PrometheusRule CRDs) - Rewrite prometheus network policy for operator-managed pods (DNS, kube-state-metrics, operator, alertmanager sidecar, monitoring namespace for DCGM) - Add alertmanager network policy (Mattermost egress) - Upgrade Grafana: AlertManager datasource, unified alerting, fix nodeSelector - Fix Tempo: tolerations for GPU nodes, resource tuning Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
12 KiB
12 KiB