Agent 106: ML health endpoint (HTTP/8095) Agent 107: Redis test fix (serial_test isolation) Agent 108: CLAUDE.md draft update (95-97% → 100%) Agent 109: Prometheus/Grafana setup (31 alerts, 6 dashboards) Agent 110: Deployment docs (9 files + 4 scripts) Agent 111: Security audit prep (0 critical vulnerabilities) Service Health: 4/4 healthy (100%) Tests: 99%+ pass rate Production: ~98% readiness Next: Wave 2 (E2E, load, perf, security validation)
17 KiB
Health Check Verification Procedures
Overview
Foxhunt HFT system implements comprehensive health checks across all services to ensure system reliability and enable automated monitoring. This guide covers health check endpoints, verification procedures, and automated monitoring scripts.
Service Health Endpoints
HTTP Health Endpoints
Each service exposes an HTTP health endpoint for easy monitoring:
API Gateway
curl http://localhost:8080/health
# Expected response:
{
"status": "healthy",
"service": "api_gateway",
"version": "1.0.0",
"uptime_seconds": 3600,
"dependencies": {
"postgres": "healthy",
"redis": "healthy",
"trading_service": "healthy"
}
}
Trading Service
curl http://localhost:8081/health
# Expected response:
{
"status": "healthy",
"service": "trading_service",
"version": "1.0.0",
"uptime_seconds": 3600,
"dependencies": {
"postgres": "healthy",
"redis": "healthy"
}
}
Backtesting Service
curl http://localhost:8083/health
# Expected response:
{
"status": "healthy",
"service": "backtesting_service",
"version": "1.0.0",
"uptime_seconds": 3600,
"models_loaded": 3,
"dependencies": {
"postgres": "healthy"
}
}
ML Training Service
curl http://localhost:8095/health
# Expected response:
{
"status": "healthy",
"service": "ml_training_service",
"version": "1.0.0",
"uptime_seconds": 3600,
"gpu_available": true,
"models_loaded": 3,
"dependencies": {
"postgres": "healthy",
"storage": "healthy"
}
}
gRPC Health Checks
For gRPC services, use grpc_health_probe:
Install grpc_health_probe
# Download latest release
wget https://github.com/grpc-ecosystem/grpc-health-probe/releases/download/v0.4.19/grpc_health_probe-linux-amd64
chmod +x grpc_health_probe-linux-amd64
sudo mv grpc_health_probe-linux-amd64 /usr/local/bin/grpc_health_probe
API Gateway (gRPC)
grpc_health_probe -addr=localhost:50051
# With TLS:
grpc_health_probe -addr=localhost:50051 \
-tls \
-tls-ca-cert=/tmp/foxhunt/certs/ca.crt
# Expected output:
# status: SERVING
Trading Service (gRPC)
grpc_health_probe -addr=localhost:50052
# Expected output:
# status: SERVING
Backtesting Service (gRPC)
grpc_health_probe -addr=localhost:50053
# Expected output:
# status: SERVING
ML Training Service (gRPC)
grpc_health_probe -addr=localhost:50054
# Expected output:
# status: SERVING
Infrastructure Health Checks
PostgreSQL
# Using pg_isready
docker exec foxhunt-postgres pg_isready -U foxhunt
# Expected output:
# /var/run/postgresql:5432 - accepting connections
# Check database connectivity
docker exec foxhunt-postgres psql -U foxhunt -d foxhunt -c 'SELECT 1;'
# Check replication lag (if using replication)
docker exec foxhunt-postgres psql -U foxhunt -d foxhunt -c \
'SELECT pg_last_wal_receive_lsn() - pg_last_wal_replay_lsn() AS replication_lag;'
Redis
# Ping check
docker exec foxhunt-redis redis-cli ping
# Expected output:
# PONG
# Check memory usage
docker exec foxhunt-redis redis-cli INFO memory | grep used_memory_human
# Check keyspace
docker exec foxhunt-redis redis-cli INFO keyspace
HashiCorp Vault
# Health endpoint
curl http://localhost:8200/v1/sys/health
# Expected response (initialized and unsealed):
{
"initialized": true,
"sealed": false,
"standby": false,
"performance_standby": false,
"replication_performance_mode": "disabled",
"replication_dr_mode": "disabled",
"server_time_utc": 1704672000,
"version": "1.15.0"
}
# Check seal status
curl http://localhost:8200/v1/sys/seal-status
InfluxDB
# Health endpoint
curl http://localhost:8086/health
# Expected response:
{
"name": "influxdb",
"message": "ready for queries and writes",
"status": "pass",
"checks": [],
"version": "2.7.0",
"commit": "abc123"
}
# Check database exists
curl -G http://localhost:8086/api/v2/buckets \
-H "Authorization: Token foxhunt-dev-token"
Prometheus Health
# Prometheus health
curl http://localhost:9090/-/healthy
# Expected output:
# Prometheus is Healthy.
# Check all service targets
curl -s http://localhost:9090/api/v1/targets | jq '.data.activeTargets[] | {job: .labels.job, health: .health}'
# Expected output:
# {"job": "api_gateway", "health": "up"}
# {"job": "trading_service", "health": "up"}
# {"job": "backtesting_service", "health": "up"}
# {"job": "ml_training_service", "health": "up"}
# {"job": "postgres", "health": "up"}
# {"job": "redis", "health": "up"}
Grafana Health
# Grafana health
curl http://localhost:3000/api/health
# Expected response:
{
"commit": "abc123",
"database": "ok",
"version": "9.5.0"
}
# Check datasources
curl -u admin:foxhunt123 http://localhost:3000/api/datasources | jq '.[] | {name: .name, type: .type}'
Automated Health Validation Script
Create comprehensive health check script at /home/jgrusewski/Work/foxhunt/scripts/health_check.sh:
#!/bin/bash
# Color codes for output
RED='\033[0;31m'
GREEN='\033[0;32m'
YELLOW='\033[1;33m'
NC='\033[0m' # No Color
# Counters
TOTAL_CHECKS=0
PASSED_CHECKS=0
FAILED_CHECKS=0
# Function to check health
check_health() {
local name=$1
local command=$2
TOTAL_CHECKS=$((TOTAL_CHECKS + 1))
if eval "$command" > /dev/null 2>&1; then
echo -e " ${GREEN}✓${NC} $name: ${GREEN}OK${NC}"
PASSED_CHECKS=$((PASSED_CHECKS + 1))
return 0
else
echo -e " ${RED}✗${NC} $name: ${RED}FAIL${NC}"
FAILED_CHECKS=$((FAILED_CHECKS + 1))
return 1
fi
}
echo "=== Foxhunt Health Check ==="
echo "Timestamp: $(date '+%Y-%m-%d %H:%M:%S')"
# 1. Docker Container Status
echo -e "\n${YELLOW}[1/7] Docker Container Status:${NC}"
check_health "Docker Compose" "docker-compose ps | grep -E '(healthy|Up)'"
docker-compose ps
# 2. Infrastructure Layer
echo -e "\n${YELLOW}[2/7] Infrastructure Health:${NC}"
check_health "PostgreSQL" "docker exec foxhunt-postgres pg_isready -U foxhunt"
check_health "Redis" "docker exec foxhunt-redis redis-cli ping | grep -q PONG"
check_health "Vault" "curl -sf http://localhost:8200/v1/sys/health"
check_health "InfluxDB" "curl -sf http://localhost:8086/health"
# 3. HTTP Health Endpoints
echo -e "\n${YELLOW}[3/7] HTTP Health Endpoints:${NC}"
check_health "API Gateway (HTTP)" "curl -sf http://localhost:8080/health"
check_health "Trading Service (HTTP)" "curl -sf http://localhost:8081/health"
check_health "Backtesting Service (HTTP)" "curl -sf http://localhost:8083/health"
check_health "ML Training Service (HTTP)" "curl -sf http://localhost:8095/health"
# 4. gRPC Services
echo -e "\n${YELLOW}[4/7] gRPC Services:${NC}"
check_health "API Gateway (gRPC)" "grpc_health_probe -addr=localhost:50051"
check_health "Trading Service (gRPC)" "grpc_health_probe -addr=localhost:50052"
check_health "Backtesting Service (gRPC)" "grpc_health_probe -addr=localhost:50053"
check_health "ML Training Service (gRPC)" "grpc_health_probe -addr=localhost:50054"
# 5. Database Connectivity
echo -e "\n${YELLOW}[5/7] Database Connectivity:${NC}"
check_health "PostgreSQL Query" "docker exec foxhunt-postgres psql -U foxhunt -d foxhunt -c 'SELECT 1;'"
check_health "Redis Keyspace" "docker exec foxhunt-redis redis-cli INFO keyspace"
check_health "InfluxDB Buckets" "curl -sf -H 'Authorization: Token foxhunt-dev-token' http://localhost:8086/api/v2/buckets"
# 6. Monitoring Stack
echo -e "\n${YELLOW}[6/7] Monitoring Stack:${NC}"
check_health "Prometheus" "curl -sf http://localhost:9090/-/healthy"
check_health "Grafana" "curl -sf http://localhost:3000/api/health"
check_health "Prometheus Targets" "curl -sf http://localhost:9090/api/v1/targets | jq -e '.data.activeTargets | length > 0'"
# 7. Service Dependencies
echo -e "\n${YELLOW}[7/7] Service Dependencies:${NC}"
# API Gateway dependencies
API_GW_HEALTH=$(curl -sf http://localhost:8080/health | jq -r '.dependencies.postgres')
if [ "$API_GW_HEALTH" = "healthy" ]; then
echo -e " ${GREEN}✓${NC} API Gateway → PostgreSQL: ${GREEN}OK${NC}"
PASSED_CHECKS=$((PASSED_CHECKS + 1))
else
echo -e " ${RED}✗${NC} API Gateway → PostgreSQL: ${RED}FAIL${NC}"
FAILED_CHECKS=$((FAILED_CHECKS + 1))
fi
TOTAL_CHECKS=$((TOTAL_CHECKS + 1))
API_GW_REDIS=$(curl -sf http://localhost:8080/health | jq -r '.dependencies.redis')
if [ "$API_GW_REDIS" = "healthy" ]; then
echo -e " ${GREEN}✓${NC} API Gateway → Redis: ${GREEN}OK${NC}"
PASSED_CHECKS=$((PASSED_CHECKS + 1))
else
echo -e " ${RED}✗${NC} API Gateway → Redis: ${RED}FAIL${NC}"
FAILED_CHECKS=$((FAILED_CHECKS + 1))
fi
TOTAL_CHECKS=$((TOTAL_CHECKS + 1))
# Summary
echo -e "\n${YELLOW}=== Health Check Summary ===${NC}"
echo "Total Checks: $TOTAL_CHECKS"
echo -e "Passed: ${GREEN}$PASSED_CHECKS${NC}"
echo -e "Failed: ${RED}$FAILED_CHECKS${NC}"
if [ $FAILED_CHECKS -eq 0 ]; then
echo -e "\n${GREEN}All health checks passed!${NC}"
exit 0
else
echo -e "\n${RED}Some health checks failed!${NC}"
exit 1
fi
Continuous Health Monitoring
Cron-based Monitoring
Setup periodic health checks:
# Make script executable
chmod +x /home/jgrusewski/Work/foxhunt/scripts/health_check.sh
# Add to crontab (run every 5 minutes)
crontab -e
# Add line:
*/5 * * * * /home/jgrusewski/Work/foxhunt/scripts/health_check.sh >> /var/log/foxhunt_health.log 2>&1
Prometheus Alerting Rules
Create /home/jgrusewski/Work/foxhunt/prometheus/alerts.yml:
groups:
- name: foxhunt_health
interval: 30s
rules:
# Service health alerts
- alert: ServiceDown
expr: up{job=~"api_gateway|trading_service|backtesting_service|ml_training_service"} == 0
for: 1m
labels:
severity: critical
annotations:
summary: "Service {{ $labels.job }} is down"
description: "Service {{ $labels.job }} has been down for more than 1 minute"
# Database health alerts
- alert: PostgreSQLDown
expr: up{job="postgres"} == 0
for: 30s
labels:
severity: critical
annotations:
summary: "PostgreSQL is down"
description: "PostgreSQL database is unreachable"
- alert: RedisDown
expr: up{job="redis"} == 0
for: 30s
labels:
severity: critical
annotations:
summary: "Redis is down"
description: "Redis cache is unreachable"
# High latency alerts
- alert: HighResponseTime
expr: http_request_duration_seconds{quantile="0.99"} > 1
for: 5m
labels:
severity: warning
annotations:
summary: "High response time on {{ $labels.job }}"
description: "P99 latency is {{ $value }}s for {{ $labels.job }}"
# Circuit breaker alerts
- alert: CircuitBreakerOpen
expr: circuit_breaker_state{state="open"} == 1
for: 2m
labels:
severity: warning
annotations:
summary: "Circuit breaker open for {{ $labels.service }}"
description: "Circuit breaker for {{ $labels.service }} has been open for 2 minutes"
# Resource alerts
- alert: HighMemoryUsage
expr: (container_memory_usage_bytes / container_spec_memory_limit_bytes) > 0.9
for: 5m
labels:
severity: warning
annotations:
summary: "High memory usage on {{ $labels.container_name }}"
description: "Memory usage is {{ $value | humanizePercentage }} on {{ $labels.container_name }}"
- alert: HighCPUUsage
expr: (rate(container_cpu_usage_seconds_total[5m]) / container_spec_cpu_quota) > 0.9
for: 5m
labels:
severity: warning
annotations:
summary: "High CPU usage on {{ $labels.container_name }}"
description: "CPU usage is {{ $value | humanizePercentage }} on {{ $labels.container_name }}"
Grafana Dashboard for Health Monitoring
Create Grafana dashboard JSON at /home/jgrusewski/Work/foxhunt/grafana/dashboards/health_overview.json:
{
"dashboard": {
"title": "Foxhunt Health Overview",
"panels": [
{
"id": 1,
"title": "Service Status",
"type": "stat",
"targets": [
{
"expr": "up{job=~\"api_gateway|trading_service|backtesting_service|ml_training_service\"}"
}
],
"fieldConfig": {
"defaults": {
"mappings": [
{"value": 0, "text": "Down", "color": "red"},
{"value": 1, "text": "Up", "color": "green"}
]
}
}
},
{
"id": 2,
"title": "Infrastructure Status",
"type": "stat",
"targets": [
{
"expr": "up{job=~\"postgres|redis|vault\"}"
}
]
},
{
"id": 3,
"title": "Circuit Breaker State",
"type": "graph",
"targets": [
{
"expr": "circuit_breaker_state"
}
]
},
{
"id": 4,
"title": "Health Check Success Rate",
"type": "graph",
"targets": [
{
"expr": "rate(health_check_success_total[5m]) / rate(health_check_total[5m])"
}
]
}
]
}
}
Advanced Health Checks
Deep Health Checks
Beyond basic liveness checks, implement deep health checks:
# Database query performance
curl http://localhost:8081/health/deep
# Expected response includes query timing:
{
"status": "healthy",
"checks": {
"database_query": {
"status": "healthy",
"latency_ms": 15,
"query": "SELECT 1"
},
"redis_ping": {
"status": "healthy",
"latency_ms": 2
},
"model_inference": {
"status": "healthy",
"latency_ms": 45,
"model": "MAMBA-2"
}
}
}
Dependency Chain Validation
Validate entire request chain:
# Create end-to-end health check
curl -X POST http://localhost:8080/health/e2e \
-H "Content-Type: application/json" \
-d '{
"check_trading_flow": true,
"check_ml_inference": true,
"check_backtesting": false
}'
# Response includes timing for each step:
{
"status": "healthy",
"total_latency_ms": 120,
"steps": [
{"step": "auth", "latency_ms": 10, "status": "ok"},
{"step": "rate_limit", "latency_ms": 5, "status": "ok"},
{"step": "trading_service", "latency_ms": 50, "status": "ok"},
{"step": "risk_check", "latency_ms": 30, "status": "ok"},
{"step": "ml_inference", "latency_ms": 25, "status": "ok"}
]
}
Troubleshooting Failed Health Checks
Common Issues and Solutions
1. HTTP Health Endpoint Returns 503
Cause: Service dependencies are unhealthy
Diagnosis:
# Check service logs
docker logs foxhunt-api-gateway --tail 50 | grep -i health
# Check dependency status in health response
curl http://localhost:8080/health | jq '.dependencies'
Fix:
# Identify failed dependency
FAILED_DEP=$(curl -s http://localhost:8080/health | jq -r '.dependencies | to_entries[] | select(.value != "healthy") | .key')
# Restart failed dependency
docker-compose restart $FAILED_DEP
# Verify recovery
sleep 10
curl http://localhost:8080/health
2. gRPC Health Check Times Out
Cause: Service not listening or TLS configuration issue
Diagnosis:
# Check if port is listening
netstat -tulpn | grep 50051
# Verify service is running
docker-compose ps api_gateway
# Check gRPC logs
docker logs foxhunt-api-gateway | grep -i grpc
Fix:
# Verify TLS certificates if using secure connection
grpc_health_probe -addr=localhost:50051 -tls -tls-ca-cert=/tmp/foxhunt/certs/ca.crt
# Or test without TLS
grpc_health_probe -addr=localhost:50051
3. Database Health Check Fails
Cause: Connection pool exhausted or database locked
Diagnosis:
# Check PostgreSQL connections
docker exec foxhunt-postgres psql -U foxhunt -d foxhunt -c \
'SELECT count(*) FROM pg_stat_activity;'
# Check for locks
docker exec foxhunt-postgres psql -U foxhunt -d foxhunt -c \
'SELECT * FROM pg_locks WHERE NOT granted;'
Fix:
# Terminate idle connections
docker exec foxhunt-postgres psql -U foxhunt -d foxhunt -c \
'SELECT pg_terminate_backend(pid) FROM pg_stat_activity WHERE state = '\''idle'\'' AND state_change < NOW() - INTERVAL '\''5 minutes'\'';'
# Restart service to reset connection pool
docker-compose restart api_gateway
Production Health Check Checklist
- All services have HTTP health endpoints
- All gRPC services implement health check protocol
- Prometheus is scraping all service metrics
- Grafana dashboards show health status
- Alerting rules configured for critical services
- Automated health check script runs periodically
- On-call team receives alerts via PagerDuty/Slack
- Runbooks exist for common health check failures
- Deep health checks validate full request chains
- Circuit breakers prevent cascade failures