Files
foxhunt/docs/deployment/DOCKER_TROUBLESHOOTING.md
jgrusewski 39c1028502 🚀 Wave 126 Wave 1 Complete: 6 agents deployed - 4/4 services healthy
Agent 106: ML health endpoint (HTTP/8095)
Agent 107: Redis test fix (serial_test isolation)
Agent 108: CLAUDE.md draft update (95-97% → 100%)
Agent 109: Prometheus/Grafana setup (31 alerts, 6 dashboards)
Agent 110: Deployment docs (9 files + 4 scripts)
Agent 111: Security audit prep (0 critical vulnerabilities)

Service Health: 4/4 healthy (100%)
Tests: 99%+ pass rate
Production: ~98% readiness

Next: Wave 2 (E2E, load, perf, security validation)
2025-10-08 00:11:38 +02:00

13 KiB

Docker Compose Troubleshooting

Service Health Checks

Check Service Status

View the status of all Docker Compose services:

docker-compose ps

Expected output for healthy system:

NAME                        IMAGE                           STATUS
foxhunt-api-gateway         foxhunt/api_gateway:latest      Up (healthy)
foxhunt-trading-service     foxhunt/trading_service:latest  Up (healthy)
foxhunt-backtesting         foxhunt/backtesting:latest      Up (healthy)
foxhunt-ml-training         foxhunt/ml_training:latest      Up (healthy)
foxhunt-postgres            timescale/timescaledb:latest    Up (healthy)
foxhunt-redis               redis:7-alpine                  Up (healthy)
foxhunt-vault               vault:1.15                      Up (healthy)
foxhunt-prometheus          prom/prometheus:latest          Up (healthy)
foxhunt-grafana             grafana/grafana:latest          Up (healthy)

Common Issues

Issue: Service shows "unhealthy"

Symptoms: Docker Compose shows service status as "unhealthy" or "starting"

Diagnosis:

# Check health endpoint
curl http://localhost:8080/health  # Adjust port for service

# Check container logs for errors
docker logs foxhunt-api-gateway --tail 100

# Check health check definition
docker inspect foxhunt-api-gateway | jq '.[].State.Health'

# View health check history
docker inspect foxhunt-api-gateway | jq '.[].State.Health.Log'

Fix:

# Restart unhealthy service
docker-compose restart api_gateway

# If health check fails repeatedly, check service dependencies
docker-compose logs postgres redis  # Check if dependencies are healthy

# Verify health check configuration in docker-compose.yml
docker-compose config | grep -A 10 healthcheck

Issue: Service restarting continuously

Symptoms: Service exits immediately after starting, high restart count

Diagnosis:

# View restart count
docker-compose ps api_gateway

# View full logs (including previous crashes)
docker logs foxhunt-api-gateway

# Check for port conflicts
netstat -tulpn | grep 50051  # Check if port already in use

# View recent container events
docker events --since 10m --filter 'container=foxhunt-api-gateway'

Common Causes & Fixes:

  1. Port Conflict:
# Find process using port
sudo lsof -i :50051

# Kill conflicting process or change port in docker-compose.yml
  1. Database Connection Failed:
# Verify PostgreSQL is accessible
docker exec foxhunt-postgres pg_isready -U foxhunt

# Check database credentials in .env file
grep DATABASE_URL .env

# Run migrations if database exists but schema is missing
cargo sqlx migrate run
  1. Missing Environment Variables:
# Check environment variables passed to container
docker inspect foxhunt-api-gateway | jq '.[].Config.Env'

# Verify .env file exists and is sourced
ls -la .env
docker-compose config | grep -A 5 environment

Issue: "Connection refused" errors

Symptoms: Services cannot connect to each other or external clients cannot reach services

Diagnosis:

# Verify service is listening on correct port
docker exec foxhunt-api-gateway netstat -tulpn

# Check service mesh connectivity
docker network inspect foxhunt_foxhunt-network

# Test internal connectivity between services
docker exec foxhunt-api-gateway ping foxhunt-postgres
docker exec foxhunt-api-gateway curl http://foxhunt-trading-service:50051/health

Fix:

# Ensure all services are on the same network
docker-compose config | grep -A 5 networks

# Restart networking stack
docker-compose down
docker network prune -f
docker-compose up -d

# Verify DNS resolution
docker exec foxhunt-api-gateway nslookup foxhunt-postgres

Issue: "Cannot allocate memory" or OOM errors

Symptoms: Services crash with out-of-memory errors

Diagnosis:

# Check Docker memory usage
docker stats --no-stream

# View container memory limits
docker inspect foxhunt-api-gateway | jq '.[].HostConfig.Memory'

# Check system memory
free -h

Fix:

# Increase container memory limit in docker-compose.yml
services:
  api_gateway:
    deploy:
      resources:
        limits:
          memory: 2G
        reservations:
          memory: 1G

# Or increase Docker daemon memory limit
# Edit /etc/docker/daemon.json:
{
  "default-memory": "2G"
}

# Restart Docker
sudo systemctl restart docker

Rebuild and Redeploy

Clean Rebuild

Perform a complete rebuild from scratch:

# 1. Stop all services
docker-compose down

# 2. Remove volumes (CAUTION: deletes all data)
docker-compose down -v

# 3. Remove old images
docker-compose down --rmi all

# 4. Clean build cache
docker builder prune -af

# 5. Rebuild images from scratch
docker-compose build --no-cache --parallel

# 6. Start services
docker-compose up -d

# 7. Watch logs for startup issues
docker-compose logs -f

# 8. Verify health
docker-compose ps

Individual Service Rebuild

Rebuild and restart a single service:

# Rebuild service image
docker-compose build --no-cache api_gateway

# Restart service with new image
docker-compose up -d api_gateway

# Verify service started correctly
docker-compose logs -f api_gateway

Incremental Updates

For faster development iterations:

# Build without cache for changed code
docker-compose build api_gateway

# Restart with new build
docker-compose up -d api_gateway

# Hot reload without full rebuild (if supported)
docker-compose restart api_gateway

Database Issues

Issue: "Database connection failed"

Diagnosis:

# Verify PostgreSQL is healthy
docker-compose ps postgres

# Check if database exists
docker exec foxhunt-postgres psql -U foxhunt -d foxhunt -c '\l'

# Check database logs
docker logs foxhunt-postgres --tail 50

# Test connection from service container
docker exec foxhunt-api-gateway pg_isready -h postgres -U foxhunt

Fix:

# Restart PostgreSQL
docker-compose restart postgres

# Wait for PostgreSQL to be ready
sleep 10

# Run migrations
cargo sqlx migrate run

# Verify tables created
docker exec foxhunt-postgres psql -U foxhunt -d foxhunt -c '\dt'

Issue: "Migration failed"

Diagnosis:

# Check migration status
cargo sqlx migrate info

# View migration history in database
docker exec foxhunt-postgres psql -U foxhunt -d foxhunt -c \
  'SELECT * FROM _sqlx_migrations ORDER BY version;'

# Check for partial migrations
docker logs foxhunt-postgres | grep -i migration

Fix:

# Revert last migration if corrupt
cargo sqlx migrate revert

# Re-run migrations
cargo sqlx migrate run

# Or manually reset migration state (CAUTION)
docker exec foxhunt-postgres psql -U foxhunt -d foxhunt -c \
  'DELETE FROM _sqlx_migrations WHERE version = <failed_version>;'

Issue: "Redis connection failed"

Diagnosis:

# Verify Redis is running
docker-compose ps redis

# Test Redis connection
docker exec foxhunt-redis redis-cli ping

# Check Redis logs
docker logs foxhunt-redis --tail 50

# Test from service container
docker exec foxhunt-api-gateway redis-cli -h redis ping

Fix:

# Restart Redis
docker-compose restart redis

# Clear Redis data if corrupt
docker exec foxhunt-redis redis-cli FLUSHALL

# Verify connectivity
docker exec foxhunt-redis redis-cli INFO server

Network Issues

Issue: Services can't communicate

Diagnosis:

# List Docker networks
docker network ls

# Inspect Foxhunt network
docker network inspect foxhunt_foxhunt-network

# Verify all services on same network
docker-compose config | grep -A 5 networks

# Check DNS resolution
docker exec foxhunt-api-gateway nslookup foxhunt-postgres
docker exec foxhunt-api-gateway nslookup foxhunt-redis

Fix:

# Recreate network
docker network rm foxhunt_foxhunt-network
docker-compose up -d

# Or force network recreation
docker-compose down
docker-compose up -d --force-recreate

Issue: External access not working

Diagnosis:

# Check port bindings
docker-compose ps
docker port foxhunt-api-gateway

# Verify firewall rules
sudo iptables -L -n | grep 50051

# Test from host machine
curl http://localhost:50051/health
grpc_health_probe -addr=localhost:50051

Fix:

# Ensure ports are exposed in docker-compose.yml
services:
  api_gateway:
    ports:
      - "50051:50050"  # host:container

# Restart with port forwarding
docker-compose down
docker-compose up -d

# Check if SELinux is blocking (on RHEL/CentOS)
sudo setenforce 0  # Temporarily disable to test

Performance Issues

Issue: Slow response times

Diagnosis:

# Check container resource usage
docker stats --no-stream

# Check for CPU throttling
docker inspect foxhunt-api-gateway | jq '.[].HostConfig.CpuQuota'

# Monitor real-time stats
docker stats

# Check for disk I/O bottlenecks
iostat -x 1 10

Fix:

# Increase CPU/memory limits
services:
  api_gateway:
    deploy:
      resources:
        limits:
          cpus: '2.0'
          memory: 4G

# Use host network for lower latency (development only)
services:
  api_gateway:
    network_mode: "host"

# Optimize Docker storage driver
# Edit /etc/docker/daemon.json:
{
  "storage-driver": "overlay2"
}

Issue: High disk usage

Diagnosis:

# Check Docker disk usage
docker system df -v

# Find large images
docker images --format "{{.Repository}}:{{.Tag}} {{.Size}}"

# Check volume usage
docker volume ls
du -sh /var/lib/docker/volumes/*

Fix:

# Clean up unused resources
docker system prune -af --volumes

# Remove old images
docker image prune -af --filter "until=24h"

# Compact database (if WAL files growing)
docker exec foxhunt-postgres psql -U foxhunt -d foxhunt -c 'VACUUM FULL;'

Logging and Debugging

View Service Logs

# View logs for specific service
docker-compose logs api_gateway

# Follow logs in real-time
docker-compose logs -f api_gateway

# View last N lines
docker-compose logs --tail=100 api_gateway

# View logs since timestamp
docker-compose logs --since=2025-01-07T10:00:00 api_gateway

# View all service logs
docker-compose logs -f

Enable Debug Logging

# Set debug log level in .env
echo "RUST_LOG=debug" >> .env

# Restart service to pick up changes
docker-compose restart api_gateway

# Verify debug logs appear
docker-compose logs api_gateway | grep DEBUG

Export Logs for Analysis

# Export logs to file
docker-compose logs api_gateway > api_gateway_logs.txt

# Export with timestamps
docker-compose logs -t api_gateway > api_gateway_logs.txt

# Export all service logs
docker-compose logs > all_services_logs.txt

# Compress for sharing
tar -czf logs_$(date +%Y%m%d_%H%M%S).tar.gz *_logs.txt

Emergency Procedures

Complete System Reset

# WARNING: This deletes ALL data

# 1. Stop and remove all containers
docker-compose down -v --remove-orphans

# 2. Remove all Foxhunt images
docker images | grep foxhunt | awk '{print $3}' | xargs docker rmi -f

# 3. Remove network
docker network rm foxhunt_foxhunt-network

# 4. Clean Docker system
docker system prune -af --volumes

# 5. Rebuild from scratch
docker-compose build --no-cache
docker-compose up -d

# 6. Run migrations
cargo sqlx migrate run

# 7. Verify system
docker-compose ps

Backup Before Troubleshooting

# Backup database
docker exec foxhunt-postgres pg_dump -U foxhunt foxhunt > backup_$(date +%Y%m%d).sql

# Backup Redis
docker exec foxhunt-redis redis-cli --rdb /tmp/dump.rdb
docker cp foxhunt-redis:/tmp/dump.rdb ./redis_backup_$(date +%Y%m%d).rdb

# Backup configuration
cp .env .env.backup
cp docker-compose.yml docker-compose.yml.backup

Restore from Backup

# Restore database
docker exec -i foxhunt-postgres psql -U foxhunt -d foxhunt < backup_20250107.sql

# Restore Redis
docker cp redis_backup_20250107.rdb foxhunt-redis:/tmp/dump.rdb
docker exec foxhunt-redis redis-cli --rdb /tmp/dump.rdb RESTORE

# Restart services
docker-compose restart

Health Check Automation

Automated Health Monitoring Script

Create scripts/docker_health_monitor.sh:

#!/bin/bash

SERVICES=("api_gateway" "trading_service" "backtesting_service" "ml_training_service" "postgres" "redis")
ALERT_EMAIL="ops@foxhunt.trading"

for service in "${SERVICES[@]}"; do
  status=$(docker inspect -f '{{.State.Health.Status}}' foxhunt-${service} 2>/dev/null)

  if [ "$status" != "healthy" ]; then
    echo "ALERT: Service ${service} is ${status}"

    # Send alert (configure your alerting system)
    echo "Service ${service} unhealthy at $(date)" | mail -s "Foxhunt Alert" $ALERT_EMAIL

    # Attempt automatic recovery
    docker-compose restart ${service}
    sleep 30

    # Check if recovered
    status=$(docker inspect -f '{{.State.Health.Status}}' foxhunt-${service})
    if [ "$status" = "healthy" ]; then
      echo "Service ${service} recovered"
    else
      echo "Service ${service} failed to recover - manual intervention required"
    fi
  fi
done

Setup Cron Job

# Add to crontab (run every 5 minutes)
crontab -e

# Add line:
*/5 * * * * /home/foxhunt/scripts/docker_health_monitor.sh >> /var/log/foxhunt_health.log 2>&1

References