# API Gateway Metrics Architecture ## System Overview ``` ┌─────────────────────────────────────────────────────────────────┐ │ API Gateway Service │ │ │ │ ┌───────────────────────────────────────────────────────────┐ │ │ │ 6-Layer Authentication Pipeline │ │ │ │ │ │ │ │ 1. JWT Extraction ──→ jwt_extraction_duration_us │ │ │ │ 2. Revocation Check ──→ revocation_check_duration_us │ │ │ │ 3. JWT Validation ──→ jwt_validation_duration_us │ │ │ │ 4. RBAC Permission ──→ rbac_check_duration_us │ │ │ │ 5. Rate Limiting ──→ rate_limit_check_duration_us │ │ │ │ 6. Audit Logging ──→ auth_total_duration_us │ │ │ │ │ │ │ │ Total Target: <10μs (p99) │ │ │ └───────────────────────────────────────────────────────────┘ │ │ │ │ │ ▼ │ │ ┌───────────────────────────────────────────────────────────┐ │ │ │ Backend Service Proxies │ │ │ │ │ │ │ │ Trading Service ──→ backend_request_duration_ms │ │ │ │ Backtesting Service ──→ circuit_breaker_state │ │ │ │ ML Training Service ──→ health_status │ │ │ │ │ │ │ └───────────────────────────────────────────────────────────┘ │ │ │ │ │ ▼ │ │ ┌───────────────────────────────────────────────────────────┐ │ │ │ Configuration Hot-Reload (PostgreSQL) │ │ │ │ │ │ │ │ NOTIFY Listener ──→ notify_events_total │ │ │ │ Config Cache ──→ config_cache_hits │ │ │ │ Reload Latency ──→ config_reload_duration_ms │ │ │ │ │ │ │ └───────────────────────────────────────────────────────────┘ │ │ │ │ │ ▼ │ │ ┌───────────────────────────────────────────────────────────┐ │ │ │ Prometheus Metrics Exporter (:9090) │ │ │ │ │ │ │ │ HTTP Endpoint: /metrics │ │ │ │ Format: Prometheus Text │ │ │ │ Total Metrics: 80+ │ │ │ │ │ │ │ └───────────────────────────────────────────────────────────┘ │ └─────────────────────────────────────────────────────────────────┘ │ │ Scrape every 5s ▼ ┌─────────────────────────────────────────────────────────────────┐ │ Prometheus Server (:9099) │ │ │ │ • Time-series database (30 day retention) │ │ • PromQL query engine │ │ • Alert rule evaluation │ │ │ │ Alert Rules: │ │ ├─ AuthLatencySLAViolation (>10μs) │ │ ├─ CircuitBreakerOpen │ │ ├─ BackendServiceUnhealthy │ │ ├─ NotifyListenerDisconnected │ │ └─ HighAuthFailureRate (>10%) │ └─────────────────────────────────────────────────────────────────┘ │ ┌──────────────┴──────────────┐ │ │ ▼ ▼ ┌──────────────────────────┐ ┌──────────────────────────┐ │ Grafana (:3000) │ │ AlertManager (:9093) │ │ │ │ │ │ Dashboards: │ │ Alert Routing: │ │ • Auth Overview │ │ • Slack (#foxhunt-*) │ │ • Layer Latencies │ │ • PagerDuty (critical) │ │ • Error Analysis │ │ • Email (warnings) │ │ • Backend Services │ │ │ │ • Config Hot-Reload │ │ Inhibition Rules: │ │ • Rate Limiting │ │ • Circuit breaker open │ │ │ │ • Backend unhealthy │ │ Panels: 19 total │ │ │ └──────────────────────────┘ └──────────────────────────┘ ``` ## Metrics Flow ### Authentication Path ``` Request arrives │ ├─ Start timer (Instant::now()) │ ├─ JWT Extraction │ └─ metrics.auth.jwt_extraction_duration_us.observe(elapsed_us) │ ├─ Revocation Check (Redis) │ ├─ metrics.auth.revocation_check_duration_us.observe(elapsed_us) │ └─ if revoked: metrics.auth.auth_errors_revoked_jwt.inc() │ ├─ JWT Validation │ ├─ metrics.auth.jwt_validation_duration_us.observe(elapsed_us) │ └─ if invalid: metrics.auth.auth_errors_signature_failed.inc() │ ├─ RBAC Permission Check │ ├─ metrics.auth.rbac_check_duration_us.observe(elapsed_us) │ └─ if denied: metrics.auth.auth_errors_permission_denied.inc() │ ├─ Rate Limiting │ ├─ metrics.auth.rate_limit_check_duration_us.observe(elapsed_us) │ └─ if limited: metrics.auth.record_rate_limit(user_id) │ └─ Total Auth ├─ metrics.auth.auth_total_duration_us.observe(total_elapsed_us) ├─ if total_elapsed_us < 10.0: metrics.auth.auth_sla_met.inc() └─ else: metrics.auth.auth_sla_exceeded.inc() ``` ### Backend Proxy Path ``` Request to Backend Service │ ├─ Start timer │ ├─ Circuit Breaker Check │ └─ if open: metrics.proxy.circuit_breaker_state{service="trading"} = 2 │ ├─ Execute Request │ ├─ On Success: │ │ ├─ metrics.proxy.backend_requests_total{service="trading"}.inc() │ │ ├─ metrics.proxy.backend_requests_success{service="trading"}.inc() │ │ └─ metrics.proxy.backend_request_duration_ms{service, method}.observe(elapsed_ms) │ │ │ └─ On Failure: │ ├─ metrics.proxy.backend_requests_failure{service, error_type}.inc() │ ├─ if consecutive_failures > threshold: │ │ └─ metrics.proxy.record_circuit_breaker_trip(service) │ └─ metrics.proxy.update_health_status(service, false) │ └─ Health Check (periodic) ├─ metrics.proxy.health_check_duration_ms.observe(elapsed_ms) └─ metrics.proxy.health_status{service} = 0 or 1 ``` ### Configuration Hot-Reload Path ``` PostgreSQL NOTIFY Event │ ├─ metrics.config.notify_events_total.inc() │ ├─ Start timer │ ├─ Fetch Configuration │ ├─ metrics.config.config_fetch_duration_ms.observe(elapsed_ms) │ └─ Check cache: │ ├─ if hit: metrics.config.config_cache_hits.inc() │ └─ if miss: metrics.config.config_cache_misses.inc() │ ├─ Validate Configuration │ ├─ if valid: metrics.config.config_validation_success.inc() │ └─ if invalid: metrics.config.config_validation_failure.inc() │ ├─ Reload Configuration │ ├─ metrics.config.record_config_reload(type, elapsed_ms) │ └─ Update type-specific counter: │ ├─ "auth" → metrics.config.config_updates_auth.inc() │ ├─ "routing" → metrics.config.config_updates_routing.inc() │ └─ "backend" → metrics.config.config_updates_backend.inc() │ └─ Listener Health ├─ metrics.config.notify_listener_connected = 1 └─ if reconnect: metrics.config.notify_listener_reconnections.inc() ``` ## Grafana Dashboard Layout ``` ┌─────────────────────────────────────────────────────────────────────┐ │ API Gateway - Authentication & Performance Dashboard │ ├─────────────────────────────────────────────────────────────────────┤ │ │ │ [Authentication Overview] │ │ ┌──────────────────────┐ ┌──────────┐ ┌────────────────┐ │ │ │ Auth Requests/s │ │ Success │ │ SLA Compliance │ │ │ │ Total: 1000 │ │ Rate: 99%│ │ <10μs: 98% │ │ │ │ Success: 990 │ │ │ │ │ │ │ │ Failure: 10 │ │ GREEN │ │ GREEN │ │ │ └──────────────────────┘ └──────────┘ └────────────────┘ │ │ │ │ [Authentication Layer Latencies (μs)] │ │ ┌─────────────────────────────────────────────────────────────────┐ │ │ │ │ │ │ │ 10μs ┤ ╭─ Total p99 │ │ │ │ ├───────────────────────────────────────────┤ │ │ │ │ 5μs ├ ╭─ RBAC │ │ │ │ │ ├─────────────────────────────────┤ │ │ │ │ │ 1μs ├ ╭─ JWT Validation │ │ │ │ │ │ ├───────────┤ │ │ │ │ │ │ 0.5μs ├ JWT Extraction │ │ │ │ │ │ └─────────────────────────────────┴─────────┴─────────────│ │ │ └─────────────────────────────────────────────────────────────────┘ │ │ │ │ [Error Analysis] [Cache Performance] │ │ ┌──────────────────────┐ ┌──────────────────────┐ │ │ │ Errors by Type │ │ JWT Cache: 99.5% │ │ │ │ - Expired: 5/s │ │ RBAC Cache: 98.2% │ │ │ │ - Revoked: 3/s │ │ Config Cache: 95.1% │ │ │ │ - Invalid: 2/s │ │ │ │ │ └──────────────────────┘ └──────────────────────┘ │ │ │ │ [Backend Services] │ │ ┌──────────────────────────────────────────────────────────────────┐│ │ │ Backend Latency (p99 ms) ││ │ │ Trading: 15ms ████░░░░░░░░ ││ │ │ Backtesting: 250ms ████████████████████░░░ ││ │ │ ML Training: 5000ms████████████████████████████████████████████ ││ │ └──────────────────────────────────────────────────────────────────┘│ │ │ │ ┌─────────────────┐ ┌──────────────────────────────────────────┐ │ │ │Circuit Breaker │ │ Health Status │ │ │ │ │ │ Service Status Last Check │ │ │ │Trading: CLOSED │ │ Trading ✅ UP 2s ago │ │ │ │Backtest: CLOSED │ │ Backtesting ✅ UP 1s ago │ │ │ │ML Train: CLOSED │ │ ML Training ✅ UP 3s ago │ │ │ └─────────────────┘ └──────────────────────────────────────────┘ │ │ │ │ [Configuration & Hot-Reload] │ │ ┌────────────────────┐ ┌────────────────────┐ ┌──────────────┐ │ │ │ Config Reloads │ │ Reload Latency p95 │ │ NOTIFY │ │ │ │ Auth: 5 │ │ 25ms │ │ Listener │ │ │ │ Routing: 3 │ │ │ │ │ │ │ │ Backend: 2 │ │ │ │ CONNECTED │ │ │ └────────────────────┘ └────────────────────┘ └──────────────┘ │ │ │ │ [Rate Limiting] │ │ ┌─────────────────────────────────────────────────────────────────┐ │ │ │ Top 10 Rate-Limited Users │ │ │ │ user_123: 50/s ████████████████████ │ │ │ │ user_456: 30/s ████████████ │ │ │ │ user_789: 20/s ████████ │ │ │ └─────────────────────────────────────────────────────────────────┘ │ └─────────────────────────────────────────────────────────────────────┘ ``` ## Alert Flow ``` Metric Exceeds Threshold │ ▼ Prometheus Alert Rule Triggered │ ├─ Evaluates condition (e.g., auth_total_duration_us p99 > 10) └─ Fires alert if condition true for duration │ ▼ AlertManager Receives Alert │ ├─ Apply inhibition rules │ └─ Suppress redundant alerts │ ├─ Route based on labels │ ├─ severity: critical → critical-alerts receiver │ ├─ severity: warning → warning-alerts receiver │ └─ component: auth → auth-alerts receiver │ └─ Send notifications ├─ Slack: #foxhunt-critical ├─ PagerDuty: On-call engineer └─ Email: team@foxhunt.com ``` ## Performance Impact ``` Request Processing Time Breakdown: ┌─────────────────────────────────────────────┐ │ Total Request Processing: 10,500 ns │ ├─────────────────────────────────────────────┤ │ │ │ Authentication: 10,000 ns (95.2%) │ │ ├─ JWT Extraction: 500 ns │ │ ├─ Revocation Check: 300 ns │ │ ├─ JWT Validation: 800 ns │ │ ├─ RBAC Check: 100 ns │ │ ├─ Rate Limit: 50 ns │ │ └─ Audit Logging: 8,250 ns │ │ │ │ Metrics Recording: 500 ns (4.8%) │ │ ├─ Counter increments: 200 ns │ │ ├─ Histogram observations: 250 ns │ │ └─ Label lookups: 50 ns │ │ │ └─────────────────────────────────────────────┘ Metrics Overhead: 500 ns / 10,500 ns = 4.8% SLA Impact: 0.5 μs / 10 μs = 5% of SLA budget ``` ## Data Retention ``` Prometheus Storage: ┌──────────────────────────────────────────┐ │ Metric Type │ Series │ Size/Day │ ├──────────────┼────────┼────────────────┤ │ Counters │ 45 │ 5 MB │ │ Histograms │ 12 │ 15 MB │ │ Gauges │ 23 │ 3 MB │ ├──────────────┼────────┼────────────────┤ │ Total │ 80 │ 23 MB/day │ │ │ │ 30-day retention: ~700 MB │ └──────────────────────────────────────────┘ ``` ## Quick Reference ### Key Metrics to Monitor | Priority | Metric | Target | Alert | |----------|--------|--------|-------| | P0 | `auth_total_duration_microseconds` (p99) | <10μs | >10μs | | P0 | `health_status{service}` | 1 | 0 | | P0 | `circuit_breaker_state{service}` | 0 | 2 | | P1 | `backend_request_duration_milliseconds` (p99) | <50ms | >100ms | | P1 | `auth_requests_success / total` | >99% | <95% | | P2 | `jwt_cache_hits / (hits + misses)` | >99% | <95% | | P2 | `notify_listener_connected` | 1 | 0 | ### PromQL Queries **Auth Success Rate:** ```promql 100 * rate(api_gateway_auth_requests_success[5m]) / rate(api_gateway_auth_requests_total[5m]) ``` **Auth Latency p99:** ```promql histogram_quantile(0.99, rate(api_gateway_auth_total_duration_microseconds_bucket[1m])) ``` **Cache Hit Rate:** ```promql 100 * rate(api_gateway_jwt_cache_hits[1m]) / (rate(api_gateway_jwt_cache_hits[1m]) + rate(api_gateway_jwt_cache_misses[1m])) ``` **Backend Availability:** ```promql avg_over_time(api_gateway_health_status{service="trading"}[5m]) ``` --- *Architecture document for Wave 71 Agent 9 - Metrics Implementation*