Files
foxhunt/docs/deployment/rollback-procedures.md
jgrusewski 7458f1be01 feat(wave12): E2E validation complete - 225-feature pipeline ready
 Validation Results:
- PPO training: 24.2s (1 epoch, 950 samples, dim=225)
- Feature extraction: 105μs/bar (9.5x faster than target)
- Model checkpoint: 293KB (147KB actor + 146KB critic)
- GPU memory: 145MB used (96.4% headroom)
- Zero dimension mismatches

📊 Success Criteria (5/5):
 Feature dimension = 225 (Wave C 201 + Wave D 24)
 Model state_dim = 225
 Training completed without errors
 Checkpoint saved successfully
 No dimension mismatch errors

📁 Training Data Ready:
- ES.FUT: 2.9MB, 180 days
- NQ.FUT: 4.4MB, 180 days
- 6E.FUT: 2.8MB, 180 days
- ZN.FUT: 65KB, 90 days (clean)

🚀 Next: Full production model retraining (4 models, ~10min GPU time)

🤖 Generated with Claude Code (https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
2025-10-22 22:48:04 +02:00

19 KiB

Rollback Procedures

Last Updated: 2025-10-22 Target Environment: All (Development, Staging, Production) RTO (Recovery Time Objective): 15 minutes RPO (Recovery Point Objective): 1 hour


Table of Contents

  1. Overview
  2. Rollback Decision Matrix
  3. Level 1: Configuration Rollback
  4. Level 2: Application Rollback
  5. Level 3: Database Rollback
  6. Level 4: Full System Rollback
  7. Post-Rollback Validation
  8. Incident Documentation

Overview

This document defines standardized rollback procedures for the Foxhunt HFT trading system. Rollbacks are categorized into 4 levels based on severity and scope.

Rollback Levels

Level Scope Downtime Complexity Use Case
L1 Configuration only 0-2 min Low Feature flags, env vars, ConfigMaps
L2 Application code 5-10 min Medium Code bugs, performance issues
L3 Database schema 10-20 min High Schema changes, migration failures
L4 Full system 15-30 min Critical Catastrophic failures

Rollback Decision Matrix

graph TD
    START[Incident Detected]
    START --> ASSESS{Assess Severity}

    ASSESS -->|Config Issue| L1[Level 1:<br/>Config Rollback]
    ASSESS -->|Code Bug| L2[Level 2:<br/>App Rollback]
    ASSESS -->|DB Issue| L3[Level 3:<br/>DB Rollback]
    ASSESS -->|System Failure| L4[Level 4:<br/>Full Rollback]

    L1 --> VALIDATE1{Health OK?}
    L2 --> VALIDATE2{Health OK?}
    L3 --> VALIDATE3{Health OK?}
    L4 --> VALIDATE4{Health OK?}

    VALIDATE1 -->|Yes| DONE[Incident Resolved]
    VALIDATE1 -->|No| ESCALATE1[Escalate to L2]

    VALIDATE2 -->|Yes| DONE
    VALIDATE2 -->|No| ESCALATE2[Escalate to L3]

    VALIDATE3 -->|Yes| DONE
    VALIDATE3 -->|No| ESCALATE3[Escalate to L4]

    VALIDATE4 -->|Yes| DONE
    VALIDATE4 -->|No| POSTMORTEM[Postmortem Required]

Level 1: Configuration Rollback

Duration: 0-2 minutes Risk: Low Downtime: None (rolling restart)

Use Cases

  • Feature flag changes causing errors
  • Environment variable misconfigurations
  • ConfigMap/Secret updates breaking functionality
  • Rate limit adjustments causing service degradation

Prerequisites

# Ensure kubectl access
kubectl cluster-info

# Verify current ConfigMap version
kubectl get configmap foxhunt-config -n foxhunt -o yaml | grep 'resourceVersion'

Procedure

Step 1: Identify Configuration Version to Rollback

# List ConfigMap revision history
kubectl rollout history configmap foxhunt-config -n foxhunt

# Show specific revision
kubectl rollout history configmap foxhunt-config -n foxhunt --revision=2

# Example output:
# REVISION  CHANGE-CAUSE
# 1         Initial deployment
# 2         Enable feature X
# 3         Increase rate limit (CURRENT - PROBLEMATIC)

Step 2: Rollback ConfigMap

# Option 1: Rollback to previous revision
kubectl rollout undo configmap foxhunt-config -n foxhunt

# Option 2: Rollback to specific revision
kubectl rollout undo configmap foxhunt-config -n foxhunt --to-revision=2

# Verify rollback
kubectl get configmap foxhunt-config -n foxhunt -o yaml

Step 3: Trigger Rolling Restart

# Restart affected deployments (Kubernetes)
kubectl rollout restart deployment/foxhunt-api-gateway -n foxhunt
kubectl rollout restart deployment/foxhunt-trading-service -n foxhunt

# Wait for rollout to complete
kubectl rollout status deployment/foxhunt-api-gateway -n foxhunt

# Docker Compose
docker-compose restart api_gateway trading_service

Step 4: Validate Service Health

# Check pod status
kubectl get pods -n foxhunt -l app.kubernetes.io/component=api-gateway

# Verify health endpoints
for i in {1..10}; do
    curl -f http://localhost:8080/health && echo " OK" || echo " FAILED"
    sleep 2
done

Helm Rollback (if deployed via Helm)

# List Helm revisions
helm history foxhunt -n foxhunt

# REVISION  UPDATED                   STATUS      CHART           APP VERSION  DESCRIPTION
# 1         Mon Oct 21 10:00:00 2025  superseded  foxhunt-0.1.0   0.1.0        Install complete
# 2         Tue Oct 22 08:00:00 2025  deployed    foxhunt-0.1.0   0.1.0        Upgrade complete

# Rollback to previous revision
helm rollback foxhunt -n foxhunt

# Rollback to specific revision
helm rollback foxhunt 1 -n foxhunt

# Wait for rollback to complete
helm status foxhunt -n foxhunt

Level 2: Application Rollback

Duration: 5-10 minutes Risk: Medium Downtime: 0-5 minutes (depending on strategy)

Use Cases

  • Code bugs introduced in recent deployment
  • Performance regressions
  • Memory leaks or resource exhaustion
  • Breaking changes in API

Prerequisites

# Identify current deployment version
kubectl describe deployment foxhunt-api-gateway -n foxhunt | grep Image:
# Image: ghcr.io/your-org/foxhunt/api-gateway:v0.1.5

# List available versions (from container registry)
docker images | grep foxhunt/api-gateway
# ghcr.io/your-org/foxhunt/api-gateway   v0.1.5   abc123   2 hours ago   1.2GB
# ghcr.io/your-org/foxhunt/api-gateway   v0.1.4   def456   1 day ago     1.2GB
# ghcr.io/your-org/foxhunt/api-gateway   v0.1.3   ghi789   2 days ago    1.2GB

Procedure

Step 1: Stop Incoming Traffic (Optional - Production Only)

# Scale down API Gateway to prevent new connections
kubectl scale deployment foxhunt-api-gateway -n foxhunt --replicas=0

# Wait for existing connections to drain (30-60 seconds)
sleep 60

# Verify no traffic
kubectl logs -n foxhunt -l app.kubernetes.io/component=api-gateway --tail=10

Step 2: Rollback Deployment

Kubernetes Deployment Rollback:

# Rollback to previous version
kubectl rollout undo deployment/foxhunt-api-gateway -n foxhunt

# Rollback to specific revision
kubectl rollout undo deployment/foxhunt-api-gateway -n foxhunt --to-revision=3

# Check rollout status
kubectl rollout status deployment/foxhunt-api-gateway -n foxhunt
# Waiting for deployment "foxhunt-api-gateway" rollout to finish: 1 out of 3 new replicas have been updated...
# deployment "foxhunt-api-gateway" successfully rolled out

# Verify new image version
kubectl get deployment foxhunt-api-gateway -n foxhunt -o jsonpath='{.spec.template.spec.containers[0].image}'
# ghcr.io/your-org/foxhunt/api-gateway:v0.1.4

Docker Compose Rollback:

# Stop services
docker-compose down

# Checkout previous git version
git log --oneline -10  # Identify commit to rollback to
git checkout <previous-commit-hash>

# Rebuild images
docker-compose build --parallel

# Start services
docker-compose up -d

# Verify versions
docker ps --format "table {{.Names}}\t{{.Image}}\t{{.Status}}"

Step 3: Rollback All Affected Services

# Rollback all application services (parallel execution)
for service in api-gateway trading-service backtesting-service ml-training-service trading-agent; do
    echo "Rolling back foxhunt-$service..."
    kubectl rollout undo deployment/foxhunt-$service -n foxhunt &
done

# Wait for all rollbacks to complete
wait

# Verify all services are healthy
kubectl get pods -n foxhunt -o wide

Step 4: Restore Traffic

# Scale API Gateway back up
kubectl scale deployment foxhunt-api-gateway -n foxhunt --replicas=3

# Monitor pod startup
watch kubectl get pods -n foxhunt -l app.kubernetes.io/component=api-gateway

Level 3: Database Rollback

Duration: 10-20 minutes Risk: High (potential data loss) Downtime: 5-15 minutes

Use Cases

  • Failed database migration
  • Schema corruption
  • Data integrity issues
  • Incompatible schema changes

Prerequisites

# CRITICAL: Verify backup exists BEFORE rollback
kubectl exec -n foxhunt foxhunt-postgresql-0 -- \
    psql -U foxhunt -d foxhunt -c "SELECT version, description, success FROM _sqlx_migrations ORDER BY version DESC LIMIT 5;"

# Verify backup file exists
ls -lh /backup/foxhunt_backup_*.sql
# -rw-r--r-- 1 root root 2.5G Oct 22 08:00 /backup/foxhunt_backup_2025-10-22_08-00-00.sql

Procedure

Step 1: Stop All Application Services

# Kubernetes: Scale down all services (prevents writes during rollback)
for service in api-gateway trading-service backtesting-service ml-training-service trading-agent; do
    kubectl scale deployment/foxhunt-$service -n foxhunt --replicas=0
done

# Docker Compose: Stop application services
docker-compose stop api_gateway trading_service backtesting_service ml_training_service trading_agent

# Verify no active connections
kubectl exec -n foxhunt foxhunt-postgresql-0 -- \
    psql -U foxhunt -d foxhunt -c "SELECT count(*) FROM pg_stat_activity WHERE datname='foxhunt';"
# count
# -----
#     1  (Only postgres itself)

Step 2: Create Emergency Backup (Before Rollback)

# Kubernetes: Backup current state
kubectl exec -n foxhunt foxhunt-postgresql-0 -- \
    pg_dump -U foxhunt -d foxhunt > /backup/foxhunt_emergency_backup_$(date +%Y%m%d_%H%M%S).sql

# Docker Compose: Backup current state
docker-compose exec postgres pg_dump -U foxhunt -d foxhunt > /backup/foxhunt_emergency_backup_$(date +%Y%m%d_%H%M%S).sql

# Verify backup file size
ls -lh /backup/foxhunt_emergency_backup_*.sql

Step 3: Restore from Backup

Option A: Full Database Restore

# Kubernetes: Drop and recreate database
kubectl exec -n foxhunt foxhunt-postgresql-0 -- \
    psql -U foxhunt -d postgres -c "DROP DATABASE foxhunt;"

kubectl exec -n foxhunt foxhunt-postgresql-0 -- \
    psql -U foxhunt -d postgres -c "CREATE DATABASE foxhunt;"

# Restore from backup
kubectl exec -i -n foxhunt foxhunt-postgresql-0 -- \
    psql -U foxhunt -d foxhunt < /backup/foxhunt_backup_2025-10-22_08-00-00.sql

# Verify restoration
kubectl exec -n foxhunt foxhunt-postgresql-0 -- \
    psql -U foxhunt -d foxhunt -c "\dt"

Option B: Migration Rollback (preferred for migration failures)

# Identify failed migration
kubectl exec -n foxhunt foxhunt-postgresql-0 -- \
    psql -U foxhunt -d foxhunt -c "SELECT version, description, success FROM _sqlx_migrations WHERE success=false;"

# Revert specific migration using sqlx
kubectl exec -n foxhunt foxhunt-api-gateway-0 -- sqlx migrate revert

# Or manually rollback migration
kubectl exec -n foxhunt foxhunt-postgresql-0 -- \
    psql -U foxhunt -d foxhunt -f /app/migrations/down/<migration-version>.sql

Step 4: Verify Database Integrity

# Check table counts
kubectl exec -n foxhunt foxhunt-postgresql-0 -- \
    psql -U foxhunt -d foxhunt -c "SELECT schemaname, tablename, n_live_tup FROM pg_stat_user_tables ORDER BY n_live_tup DESC;"

# Verify critical tables exist
kubectl exec -n foxhunt foxhunt-postgresql-0 -- \
    psql -U foxhunt -d foxhunt -c "SELECT tablename FROM pg_tables WHERE schemaname='public';" | grep -E 'orders|positions|regime_states'

# Check for data corruption
kubectl exec -n foxhunt foxhunt-postgresql-0 -- \
    psql -U foxhunt -d foxhunt -c "SELECT count(*) FROM orders WHERE created_at IS NULL;"
# count
# -----
#     0  (Good - no NULL timestamps)

Step 5: Restart Application Services

# Kubernetes: Scale services back up
for service in api-gateway trading-service backtesting-service ml-training-service trading-agent; do
    kubectl scale deployment/foxhunt-$service -n foxhunt --replicas=3
done

# Docker Compose: Start services
docker-compose up -d

# Monitor startup
kubectl get pods -n foxhunt -w

Level 4: Full System Rollback

Duration: 15-30 minutes Risk: Critical Downtime: 15-30 minutes

Use Cases

  • Catastrophic system failure
  • Multiple cascading failures
  • Data corruption across services
  • Infrastructure compromise

Prerequisites

# Verify ALL backups are available
ls -lh /backup/ | grep $(date +%Y-%m-%d)
# foxhunt_backup_2025-10-22_08-00-00.sql
# redis_dump_2025-10-22_08-00-00.rdb
# vault_backup_2025-10-22_08-00-00.tar.gz

# Verify cluster health
kubectl get nodes
# All nodes should be Ready

# Identify last known good version
git log --oneline -20
git tag -l | tail -10

Procedure

Step 1: Declare System Emergency

# Notify team via Slack/PagerDuty
curl -X POST https://hooks.slack.com/services/YOUR/WEBHOOK/URL \
  -H 'Content-Type: application/json' \
  -d '{"text":"🚨 EMERGENCY: Full system rollback initiated for Foxhunt trading system. ETA: 30 minutes."}'

# Set maintenance mode (if available)
kubectl create configmap foxhunt-maintenance -n foxhunt \
  --from-literal=enabled=true \
  --from-literal=message="System maintenance in progress. Trading suspended."

Step 2: Stop All Services

# Kubernetes: Delete all application deployments
kubectl delete deployment -n foxhunt --all

# Docker Compose: Stop all services
docker-compose down

# Verify all stopped
kubectl get pods -n foxhunt
# No resources found in foxhunt namespace.

Step 3: Restore Infrastructure Services

PostgreSQL:

# Restore PostgreSQL data
kubectl exec -n foxhunt foxhunt-postgresql-0 -- \
    psql -U foxhunt -d postgres -c "DROP DATABASE IF EXISTS foxhunt;"

kubectl exec -n foxhunt foxhunt-postgresql-0 -- \
    psql -U foxhunt -d postgres -c "CREATE DATABASE foxhunt;"

kubectl exec -i -n foxhunt foxhunt-postgresql-0 -- \
    psql -U foxhunt -d foxhunt < /backup/foxhunt_backup_2025-10-22_08-00-00.sql

Redis:

# Restore Redis data
kubectl cp /backup/redis_dump_2025-10-22_08-00-00.rdb \
    foxhunt/foxhunt-redis-master-0:/data/dump.rdb

kubectl exec -n foxhunt foxhunt-redis-master-0 -- redis-cli SHUTDOWN SAVE
kubectl delete pod foxhunt-redis-master-0 -n foxhunt

# Wait for Redis to restart
kubectl wait --for=condition=ready pod -l app.kubernetes.io/name=redis -n foxhunt --timeout=120s

Vault:

# Restore Vault data (if self-hosted)
kubectl exec -n foxhunt foxhunt-vault-0 -- vault operator unseal <KEY_1>
kubectl exec -n foxhunt foxhunt-vault-0 -- vault operator unseal <KEY_2>
kubectl exec -n foxhunt foxhunt-vault-0 -- vault operator unseal <KEY_3>

# Restore secrets
tar -xzf /backup/vault_backup_2025-10-22_08-00-00.tar.gz -C /tmp/vault_restore
kubectl cp /tmp/vault_restore foxhunt/foxhunt-vault-0:/vault/restore

kubectl exec -n foxhunt foxhunt-vault-0 -- vault kv put secret/foxhunt/database @/vault/restore/database.json
kubectl exec -n foxhunt foxhunt-vault-0 -- vault kv put secret/foxhunt/redis @/vault/restore/redis.json
kubectl exec -n foxhunt foxhunt-vault-0 -- vault kv put secret/foxhunt/jwt @/vault/restore/jwt.json

Step 4: Checkout Last Known Good Version

# Identify last stable version
git tag -l | grep -E '^v[0-9]+\.[0-9]+\.[0-9]+$' | sort -V | tail -5
# v0.1.1
# v0.1.2
# v0.1.3
# v0.1.4 (last known good)
# v0.1.5 (current - problematic)

# Checkout last stable version
git checkout v0.1.4

# Or use specific commit
git checkout <commit-hash>

# Verify version
git describe --tags
# v0.1.4

Step 5: Rebuild and Redeploy

Kubernetes (Helm):

# Update Helm chart to use previous version
helm upgrade foxhunt ./helm/foxhunt \
  --namespace foxhunt \
  --set global.imageTag=v0.1.4 \
  --force \
  --wait \
  --timeout 15m

# Monitor rollout
kubectl rollout status deployment -n foxhunt --watch

Docker Compose:

# Rebuild images from checked-out version
docker-compose build --parallel --no-cache

# Start all services
docker-compose up -d

# Monitor startup
docker-compose logs -f

Step 6: Comprehensive Health Validation

# Run full validation script
bash /home/jgrusewski/Work/foxhunt/scripts/validate-deployment.sh

# Expected output:
# ✅ PostgreSQL: Healthy
# ✅ Redis: Healthy
# ✅ Vault: Healthy (unsealed)
# ✅ API Gateway: Healthy (3/3 replicas ready)
# ✅ Trading Service: Healthy (5/5 replicas ready)
# ✅ Backtesting Service: Healthy (2/2 replicas ready)
# ✅ ML Training Service: Healthy (1/1 replicas ready)
# ✅ Trading Agent: Healthy (3/3 replicas ready)
# ✅ Database migrations: Up to date (version 044)
# ✅ End-to-end test: PASSED

Step 7: Disable Maintenance Mode

# Remove maintenance ConfigMap
kubectl delete configmap foxhunt-maintenance -n foxhunt

# Notify team
curl -X POST https://hooks.slack.com/services/YOUR/WEBHOOK/URL \
  -H 'Content-Type: application/json' \
  -d '{"text":"✅ Full system rollback complete. Foxhunt trading system operational on version v0.1.4."}'

Post-Rollback Validation

1. Service Health Checks

# Check all service health endpoints
services=("api_gateway:8080" "trading_service:8081" "backtesting_service:8082" "ml_training_service:8095" "trading_agent:8083")

for svc in "${services[@]}"; do
    IFS=':' read -r name port <<< "$svc"
    echo -n "Checking $name... "
    if curl -sf http://localhost:$port/health > /dev/null; then
        echo "✅ Healthy"
    else
        echo "❌ FAILED"
    fi
done

2. Database Integrity Validation

# Run integrity checks
kubectl exec -n foxhunt foxhunt-postgresql-0 -- \
    psql -U foxhunt -d foxhunt -c "
    SELECT
        'orders' as table_name,
        count(*) as row_count,
        count(DISTINCT symbol) as unique_symbols
    FROM orders
    UNION ALL
    SELECT
        'positions' as table_name,
        count(*) as row_count,
        count(DISTINCT symbol) as unique_symbols
    FROM positions;
"

3. End-to-End Smoke Test

# Test TLI authentication
tli auth login --username admin --password $ADMIN_PASSWORD

# Submit test order (paper trading)
tli trade order submit \
    --symbol ES.FUT \
    --action BUY \
    --quantity 1 \
    --order-type MARKET \
    --paper-trade

# Verify order appears in database
kubectl exec -n foxhunt foxhunt-postgresql-0 -- \
    psql -U foxhunt -d foxhunt -c "SELECT * FROM orders ORDER BY created_at DESC LIMIT 1;"

4. Performance Baseline Validation

# Measure API Gateway latency
for i in {1..100}; do
    curl -w "%{time_total}\n" -o /dev/null -s http://localhost:8080/health
done | awk '{sum+=$1; sumsq+=($1)^2} END {
    avg=sum/NR;
    stddev=sqrt(sumsq/NR - avg^2);
    print "Avg: " avg*1000 " ms";
    print "P95: " (avg + 1.645*stddev)*1000 " ms";
    print "P99: " (avg + 2.326*stddev)*1000 " ms"
}'

# Expected:
# Avg: 0.5 ms
# P95: 2.0 ms
# P99: 5.0 ms

Incident Documentation

After every rollback, complete an incident report:

Incident Report Template

# Incident Report: [INCIDENT-ID]

**Date**: 2025-10-22
**Time**: 10:00 UTC
**Severity**: P1 (Critical)
**Duration**: 25 minutes
**Rollback Level**: Level 3 (Database Rollback)

## Summary
Brief description of what happened.

## Impact
- Users affected: 150 active traders
- Trading halted: Yes (25 minutes)
- Data loss: None
- Financial impact: $0 (paper trading only)

## Timeline
- 09:55 UTC: Deployment v0.1.5 initiated
- 10:00 UTC: Database migration 046 applied
- 10:02 UTC: Trading Service unable to connect to database
- 10:05 UTC: Decision to rollback
- 10:10 UTC: Database restored from backup
- 10:15 UTC: Services restarted on v0.1.4
- 10:25 UTC: Full system operational

## Root Cause
Migration 046 had incompatible schema changes.

## Resolution
Rolled back to database backup from 08:00 UTC, restored v0.1.4.

## Action Items
1. [ ] Add database migration testing to CI/CD pipeline
2. [ ] Implement migration dry-run procedure
3. [ ] Update deployment checklist

## Lessons Learned
- Backup frequency (1 hour RPO) was adequate
- Database rollback procedure worked as documented
- Need automated schema compatibility testing

End of Rollback Procedures