Add QuestDB ILP sink for training metrics, update Prometheus scrape configs, and fix network policies for monitoring stack connectivity. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- Loki (3.4.2): log aggregation on gitlab node, 7-day retention, TSDB storage - Tempo (2.7.1): OTLP trace receiver on gitlab node, 7-day retention - Promtail (3.4.2): DaemonSet log shipper with K8s pod discovery + RBAC - Grafana datasources: Prometheus + Loki + Tempo with trace-to-log correlation - Services: OTEL_EXPORTER_OTLP_ENDPOINT=http://tempo:4317 on ml-training, trading, backtesting services - CI deploy job: applies monitoring manifests alongside service deployments Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>