# API Gateway Metrics - Deployment Guide ## Quick Start ### 1. Run the Monitoring Stack ```bash cd monitoring docker-compose up -d ``` This starts: - **Prometheus** on `http://localhost:9099` (metrics collection) - **Grafana** on `http://localhost:3000` (visualization) - **AlertManager** on `http://localhost:9093` (alert routing) - **PostgreSQL Exporter** on `http://localhost:9187` - **Redis Exporter** on `http://localhost:9121` - **Node Exporter** on `http://localhost:9100` ### 2. Access Grafana Dashboard 1. Open `http://localhost:3000` 2. Login: `admin` / `foxhunt2025` 3. Navigate to **Dashboards** → **API Gateway - Authentication & Performance** ### 3. Run Example to Generate Metrics ```bash cd services/api_gateway cargo run --example metrics_example ``` This will: - Generate 100 successful auth events - Record 4 auth failures - Simulate 50 trading service requests - Simulate 30 backtesting requests - Update health status for all backends - Expose metrics at `http://localhost:9090/metrics` ### 4. View Metrics in Prometheus Open `http://localhost:9099/graph` and try these queries: **Authentication Success Rate:** ```promql 100 * rate(api_gateway_auth_requests_success[1m]) / rate(api_gateway_auth_requests_total[1m]) ``` **Auth Latency p99:** ```promql histogram_quantile(0.99, rate(api_gateway_auth_total_duration_microseconds_bucket[1m])) ``` **Backend Request Rate by Service:** ```promql rate(api_gateway_backend_requests_total[1m]) ``` **Cache Hit Rate:** ```promql 100 * rate(api_gateway_jwt_cache_hits[1m]) / (rate(api_gateway_jwt_cache_hits[1m]) + rate(api_gateway_jwt_cache_misses[1m])) ``` ## Production Integration ### API Gateway Service ```rust use api_gateway::metrics::{GatewayMetrics, metrics_router}; use axum::Server; #[tokio::main] async fn main() -> Result<()> { // Initialize metrics let metrics = GatewayMetrics::new()?; // Start Prometheus exporter on separate port let metrics_addr = "0.0.0.0:9090".parse()?; let metrics_router = metrics_router(metrics.registry()); tokio::spawn(async move { Server::bind(&metrics_addr) .serve(metrics_router.into_make_service()) .await .expect("Metrics server failed"); }); // Create auth interceptor with metrics let auth_interceptor = AuthInterceptor::new( jwt_service, revocation_service, authz_service, rate_limiter, audit_logger, ); // Instrument auth interceptor to record metrics let instrumented_auth = InstrumentedAuthInterceptor::new( auth_interceptor, metrics.auth.clone(), ); // Create gRPC server with metrics let server = Server::builder() .layer(instrumented_auth) .add_service(trading_proxy) .add_service(backtesting_proxy) .add_service(ml_training_proxy) .serve(addr) .await?; Ok(()) } ``` ### Docker Deployment Add metrics port to `docker-compose.yml`: ```yaml services: api-gateway: image: foxhunt/api-gateway:latest ports: - "50051:50051" # gRPC - "9090:9090" # Prometheus metrics environment: - GATEWAY_BIND_ADDR=0.0.0.0:50051 - METRICS_BIND_ADDR=0.0.0.0:9090 - JWT_SECRET_FILE=/run/secrets/jwt_secret - REDIS_URL=redis://redis:6379 networks: - foxhunt-monitoring ``` Update Prometheus to scrape API Gateway: ```yaml # monitoring/prometheus/prometheus.yml scrape_configs: - job_name: 'api_gateway' static_configs: - targets: ['api-gateway:9090'] ``` ## Metrics Overview ### Key Metrics to Monitor | Metric | Target | Alert Threshold | Description | |--------|--------|-----------------|-------------| | `api_gateway_auth_total_duration_microseconds` (p99) | <10μs | >10μs | Total auth latency SLA | | `api_gateway_auth_requests_success` / `total` | >99% | <95% | Auth success rate | | `api_gateway_circuit_breaker_state` | 0 (closed) | 2 (open) | Backend health | | `api_gateway_health_status` | 1 (healthy) | 0 (unhealthy) | Service availability | | `api_gateway_notify_listener_connected` | 1 | 0 | Config hot-reload | ### Performance Targets **Authentication (HFT Requirements):** - JWT extraction: <0.5μs - JWT validation: <1μs - Revocation check: <500ns - RBAC check: <100ns - Rate limit check: <50ns - **Total auth: <10μs (p99)** **Backend Proxy:** - Trading Service: <50ms (p99) - Backtesting Service: <500ms (p99) - ML Training Service: <5000ms (p99) **Configuration:** - Config reload: <50ms (p95) - Config fetch: <25ms (p95) **Cache Performance:** - JWT cache hit rate: >99% - RBAC cache hit rate: >95% - Config cache hit rate: >90% ## Alerting ### Critical Alerts 1. **AuthLatencySLAViolation**: p99 auth latency >10μs for 1m 2. **CircuitBreakerOpen**: Backend circuit breaker opened 3. **BackendServiceUnhealthy**: Health checks failing for 2m 4. **NotifyListenerDisconnected**: Hot-reload capability lost 5. **RedisConnectionFailure**: JWT revocation unavailable ### Warning Alerts 1. **HighAuthFailureRate**: Auth failure rate >10% for 2m 2. **HighBackendLatency**: Backend p99 >100ms for 3m 3. **ConnectionPoolExhaustion**: Pool utilization >90% for 5m 4. **LowCacheHitRate**: RBAC cache hit rate <90% for 5m 5. **ExcessiveRateLimiting**: Rate limit rejections >10/s for 5m ## Grafana Dashboard Panels ### Row 1: Authentication Overview - **Auth Requests**: Total/Success/Failure rate - **Auth Success Rate**: Percentage gauge with thresholds - **Auth SLA Compliance**: <10μs target compliance ### Row 2: Layer Latency Breakdown - **JWT Extraction p99**: Sub-microsecond target - **JWT Validation p99**: <1μs target - **Revocation Check p99**: <500ns target - **RBAC Check p99**: <100ns target - **Rate Limit Check p99**: <50ns target - **Total Auth p99**: <10μs SLA ### Row 3: Error Analysis - **Errors by Type**: Missing JWT, expired, revoked, etc. - **Cache Performance**: JWT and RBAC cache hit rates ### Row 4: Backend Services - **Request Latency**: p99 by service - **Circuit Breaker States**: Open/closed status - **Health Status**: Table view of service health - **Connection Pool**: Utilization percentage ### Row 5: Configuration - **Reload Events**: By type (auth, routing, rate_limit, backend) - **Hot-Reload Latency**: p95 reload time - **NOTIFY Listener**: Connection status indicator ### Row 6: Rate Limiting - **Top 10 Rate-Limited Users**: Highest rejection rates - **Active Entries**: Total users being tracked ## Troubleshooting ### Metrics Not Appearing in Prometheus 1. Check API Gateway metrics endpoint: ```bash curl http://localhost:9090/metrics | grep api_gateway ``` 2. Verify Prometheus target: ```bash curl http://localhost:9099/api/v1/targets | jq '.data.activeTargets[] | select(.job=="api_gateway")' ``` 3. Check Prometheus logs: ```bash docker logs foxhunt-prometheus ``` ### Grafana Dashboard Not Loading 1. Verify Prometheus data source: - Grafana → Configuration → Data Sources - URL should be `http://prometheus:9090` 2. Import dashboard manually: - Grafana → Dashboards → Import - Upload `monitoring/grafana/api_gateway_dashboard.json` ### High Metrics Cardinality If Prometheus memory usage is high: ```bash # Check metric cardinality curl http://localhost:9099/api/v1/label/__name__/values | jq '. | length' # Identify high-cardinality metrics curl 'http://localhost:9099/api/v1/query?query=count({__name__=~".+"}) by (__name__)' | jq '.data.result | sort_by(.value[1] | tonumber) | reverse | .[0:10]' ``` **Solutions:** - Reduce `user_id` label cardinality in per-user metrics - Use recording rules for aggregated metrics - Decrease retention period in Prometheus config ### Alerts Not Firing 1. Verify alert rules syntax: ```bash docker exec foxhunt-prometheus promtool check rules /etc/prometheus/alerts/api_gateway_alerts.yml ``` 2. Check active alerts: ```bash curl http://localhost:9099/api/v1/alerts | jq '.data.alerts' ``` 3. Verify AlertManager: ```bash curl http://localhost:9093/api/v2/alerts | jq ``` ## Performance Impact ### Metrics Overhead - Counter increment: **~50ns** - Histogram observation: **~200ns** - Label lookup: **~10ns** (cached) **Total overhead per authenticated request: <500ns (<0.005% of 10μs SLA)** ### Memory Usage - Prometheus TSDB: ~50MB per million data points - Grafana: ~200MB base + 10MB per dashboard - API Gateway metrics: ~5MB per 100k requests ### Network Bandwidth - Metrics scrape: ~50KB per scrape - Scrape interval: 5s - Bandwidth: ~10KB/s per target ## Production Checklist - [ ] Prometheus deployed with persistent storage - [ ] Grafana deployed with API Gateway dashboard - [ ] AlertManager configured with notification channels - [ ] PostgreSQL exporter connected to config database - [ ] Redis exporter connected to revocation cache - [ ] API Gateway metrics endpoint exposed on :9090 - [ ] Firewall rules allow Prometheus scraping - [ ] HTTPS/TLS configured for Grafana - [ ] Alert routing tested (Slack/PagerDuty) - [ ] Metric retention policy configured (30 days default) - [ ] Backup configured for Prometheus/Grafana data - [ ] Monitoring stack monitored (meta-monitoring) ## Next Steps 1. **Configure Alerting**: - Update Slack webhook in `alertmanager.yml` - Configure PagerDuty for critical alerts - Test alert routing 2. **Customize Dashboard**: - Add panels for business metrics - Configure alerts for your SLAs - Set up user-specific views 3. **Integrate with CI/CD**: - Add metrics checks to deployment pipeline - Canary deployments based on error rates - Automated rollback on SLA violations 4. **Advanced Features**: - Recording rules for complex queries - Federation for multi-cluster setup - Long-term storage (Thanos/Cortex) ## Support For issues or questions: - Check logs: `docker logs foxhunt-prometheus` - Review Grafana docs: https://grafana.com/docs/ - Prometheus docs: https://prometheus.io/docs/ - Internal docs: `services/api_gateway/src/metrics/README.md`