# Foxhunt Alerting Architecture **Agent H5** | Production Monitoring System --- ## πŸ—οΈ Alert Flow Architecture ``` β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚ METRICS COLLECTION β”‚ β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€ β”‚ API Gateway (9091) β”‚ Trading Service (9092) β”‚ PostgreSQL β”‚ β”‚ Backtesting (9093) β”‚ ML Training (9094) β”‚ Prometheus β”‚ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β”‚ β–Ό β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚ PROMETHEUS (9090) β”‚ β”‚ Alert Evaluation β”‚ β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€ β”‚ β€’ 8 Alert Groups (32 total alerts) β”‚ β”‚ β€’ Evaluation every 15-60s β”‚ β”‚ β€’ Time-series queries with thresholds β”‚ β”‚ β€’ State: pending β†’ firing β†’ resolved β”‚ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β”‚ β–Ό β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚ ALERTMANAGER (9093) β”‚ β”‚ Routing & Deduplication β”‚ β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€ β”‚ β€’ Group by: alertname, severity, component β”‚ β”‚ β€’ Inhibition rules (suppress redundant alerts) β”‚ β”‚ β€’ Route by severity and component β”‚ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β”‚ β”‚ β”‚ β”‚ β–Ό β–Ό β–Ό β–Ό β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β” β”‚ Slack β”‚ β”‚ Email β”‚ β”‚Webhookβ”‚ β”‚PagerDutyβ”‚ β”‚ 8 channelsβ”‚ β”‚Criticalβ”‚ β”‚Custom β”‚ β”‚ Future β”‚ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”€β”€β”˜ ``` --- ## 🎯 Alert Categories ### 1. Critical Alerts (Immediate Response) **Response Time**: <5 minutes **Channels**: Slack + Email + Webhook ``` Latency β”œβ”€β”€ CriticalP99LatencyAPIGateway (>100ms for 1m) β”œβ”€β”€ CriticalP99LatencyTradingService (>100ms for 1m) └── CriticalOrderProcessingLatency (>100ms for 30s) Availability β”œβ”€β”€ CriticalServiceDown (unreachable for 30s) └── DegradedSystemHealth (<75% services up) Memory β”œβ”€β”€ CriticalMemoryGrowth (>10%/hour for 5m) β”œβ”€β”€ CriticalMemoryUsageAbsolute (>8GB) └── CriticalSystemMemoryPressure (>90% system) Trading/Risk β”œβ”€β”€ CriticalPositionLimitBreach (immediate) β”œβ”€β”€ HighDrawdown (>5%, immediate) β”œβ”€β”€ CriticalMarketDataStale (>5s old) └── RiskCheckFailures (>5 in 5m) Database β”œβ”€β”€ CriticalPostgreSQLDown (30s) └── PostgreSQLConnectionPoolExhaustion (>90%) ``` ### 2. Warning Alerts (Review within hours) **Response Time**: <4 hours **Channels**: Slack only ``` Errors β”œβ”€β”€ HighErrorRateAPIGateway (>1% for 3m) β”œβ”€β”€ HighErrorRateTradingService (>1% for 3m) └── HighOrderRejectionRate (>1% for 3m) Resources β”œβ”€β”€ HighCPUUsage (>80% for 5m) β”œβ”€β”€ DiskSpaceLow (<15% for 5m) └── DiskSpaceCritical (<10% for 2m) ML Health β”œβ”€β”€ HighMLPredictionLatency (>50ms P99) └── MLPredictionErrors (>1% for 3m) Database └── SlowDatabaseQueries (>100ms avg) ``` --- ## πŸ”” Notification Channels ### Slack Channels (8 specialized) ``` #foxhunt-critical-latency β†’ P99 latency violations #foxhunt-critical-outages β†’ Service down alerts #foxhunt-critical-memory β†’ Memory leak detection #foxhunt-critical-risk β†’ Risk management alerts #foxhunt-critical-trading β†’ Trading system alerts #foxhunt-critical-database β†’ Database failures #foxhunt-warnings-errors β†’ Error rate warnings #foxhunt-warnings-resources β†’ CPU/disk warnings #foxhunt-warnings-ml β†’ ML model warnings ``` ### Email Recipients ``` oncall@foxhunt.local β†’ All critical service down risk-team@foxhunt.local β†’ Critical risk alerts monitoring@foxhunt.local β†’ Warning aggregates ``` ### Webhooks ``` http://localhost:5001/webhook β†’ Default http://localhost:5001/critical-latency β†’ Latency alerts http://localhost:5001/critical-service-down β†’ Outages http://localhost:5001/critical-memory β†’ Memory alerts http://localhost:5001/critical-risk β†’ Risk alerts http://localhost:5001/critical-trading β†’ Trading alerts http://localhost:5001/critical-database β†’ Database alerts ``` --- ## πŸ›‘οΈ Alert Inhibition Rules ### Suppression Logic ``` If Service Down ↓ Suppress: All alerts from that service Why: Root cause is service unavailability If System Health Degraded ↓ Suppress: Individual service alerts Why: System-wide issue, not component-specific If Critical Memory Alert ↓ Suppress: Warning memory alerts Why: Critical takes precedence If Database Down ↓ Suppress: Slow queries, connection pool alerts Why: Root cause is database unavailability If Market Data Stale ↓ Suppress: Risk check failures (may be related) Why: Stale data causes risk failures If Position Limit Breached ↓ Suppress: Order rejection alerts Why: Orders rejected due to position limits ``` --- ## πŸ“Š Alert Timing Matrix | Alert Group | Evaluation | Group Wait | Group Interval | Repeat | |-------------|-----------|------------|----------------|--------| | Latency Critical | 15s | 0s | 1m | 15m | | Service Down | 15s | 0s | 30s | 5m | | Memory Critical | 30s | 0s | 2m | 10m | | Trading/Risk | 15s | 0s | 30s-1m | 5-10m | | Error Rates | 15s | 30s | 5m | 2h | | Resources | 30s | 1m | 5m | 4h | | ML Health | 30s | 1m | 10m | 4h | | Aggregate | 1m | 10s | 5m | 4h | --- ## πŸ” Alert States ``` β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚ Normal β”‚ No threshold breach β””β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”˜ β”‚ Threshold breached β–Ό β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚ Pending β”‚ Waiting for "for" duration β””β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”˜ β”‚ Duration met β–Ό β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚ Firing β”‚ Alert active, notifications sent β””β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”˜ β”‚ Issue resolved β–Ό β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚ Resolved β”‚ Recovery notification sent β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ ``` --- ## πŸŽ›οΈ Configuration Files ### 1. Alert Rules **File**: `config/prometheus/rules/production-alerts.yml` ```yaml groups: - name: production-latency-critical interval: 15s rules: - alert: CriticalP99LatencyAPIGateway expr: histogram_quantile(0.99, rate(...)) > 0.1 for: 1m labels: severity: critical component: latency ``` ### 2. AlertManager Routing **File**: `config/prometheus/alertmanager-production.yml` ```yaml route: receiver: 'default-webhook' group_by: ['alertname', 'severity', 'component'] routes: - match: severity: critical component: latency receiver: 'critical-latency' group_wait: 0s ``` ### 3. Prometheus Config **File**: `config/prometheus/prometheus.yml` ```yaml rule_files: - "rules/*.yml" alerting: alertmanagers: - static_configs: - targets: ['alertmanager:9093'] ``` --- ## πŸš€ Quick Commands ### Check Alert Status ```bash # View all firing alerts curl -s http://localhost:9090/api/v1/alerts | \ jq '.data.alerts[] | select(.state == "firing")' # Count alerts by state curl -s http://localhost:9090/api/v1/alerts | \ jq '.data.alerts | group_by(.state) | map({state: .[0].state, count: length})' # View specific alert curl -s http://localhost:9090/api/v1/alerts | \ jq '.data.alerts[] | select(.labels.alertname == "CriticalServiceDown")' ``` ### Reload Configuration ```bash # Reload Prometheus curl -X POST http://localhost:9090/-/reload # Reload AlertManager curl -X POST http://localhost:9093/-/reload ``` ### Test Alerts ```bash # Run test suite ./scripts/test_alerting.sh # Validate configuration ./scripts/validate_h5_alerting.sh ``` --- ## πŸ“ˆ Metrics Dashboard ### Key Prometheus Queries **P99 Latency by Service** ```promql histogram_quantile(0.99, rate(grpc_server_handling_seconds_bucket[1m])) ``` **Error Rate by Service** ```promql sum(rate(grpc_server_handled_total{grpc_code!="OK"}[5m])) by (job) / sum(rate(grpc_server_handled_total[5m])) by (job) ``` **Memory Growth Rate** ```promql ((process_resident_memory_bytes - (process_resident_memory_bytes offset 1h)) / (process_resident_memory_bytes offset 1h)) * 100 ``` **Service Availability** ```promql up{job=~"api_gateway|trading_service|backtesting_service|ml_training_service"} ``` --- ## 🎯 Alert Priorities ### P0 (Critical - Immediate) - Service Down - P99 Latency >100ms - Memory Growth >10%/hour - Position Limit Breach - Drawdown >5% - Market Data Stale ### P1 (Warning - Hours) - Error Rate >1% - CPU >80% - Disk <15% - Slow Queries - ML Prediction Errors ### P2 (Info - Days) - System health degradation - Alert storms - Configuration changes --- ## πŸ“ž Escalation Path ``` Alert Fires ↓ Slack Notification (#critical-*) ↓ If no ACK in 5 minutes ↓ Email to oncall@foxhunt.local ↓ If no ACK in 10 minutes ↓ PagerDuty escalation (future) ↓ If no ACK in 15 minutes ↓ SMS to on-call engineer (future) ``` --- **Last Updated**: 2025-10-18 **Agent**: H5 **Status**: Production Ready **Validation**: 96.2% (26/27 checks passed)