# Kubernetes Deployment Guide
**Last Updated**: 2025-10-22
**Target Environment**: Production, Multi-Region
**Prerequisites**: Kubernetes 1.27+, Helm 3.12+, kubectl
**Deployment Time**: 30-45 minutes
---
## Table of Contents
1. [Overview](#overview)
2. [Prerequisites](#prerequisites)
3. [Cluster Setup](#cluster-setup)
4. [Helm Chart Deployment](#helm-chart-deployment)
5. [Service Configuration](#service-configuration)
6. [Scaling Configuration](#scaling-configuration)
7. [Health Checks and Readiness](#health-checks-and-readiness)
8. [Monitoring and Observability](#monitoring-and-observability)
9. [Troubleshooting](#troubleshooting)
---
## Overview
This guide covers deploying Foxhunt to Kubernetes using Helm charts. The deployment supports:
- **High Availability**: 3+ replicas per service
- **Auto-scaling**: HPA based on CPU, memory, and custom metrics
- **Service Mesh**: Istio for traffic management
- **Secrets Management**: External Secrets Operator + Vault
- **Observability**: Prometheus, Grafana, Jaeger
- **GPU Workloads**: ML training service with GPU node pools
### Architecture Diagram
```mermaid
graph TB
subgraph "Ingress Layer"
INGRESS[Ingress Controller
NGINX/Istio]
end
subgraph "Application Layer"
direction LR
GATEWAY[API Gateway
3 replicas]
TRADING[Trading Service
5 replicas]
BACKTEST[Backtesting
2 replicas]
ML[ML Training
1 GPU replica]
AGENT[Trading Agent
3 replicas]
end
subgraph "Data Layer"
POSTGRES[(PostgreSQL
StatefulSet)]
REDIS[(Redis
StatefulSet)]
end
subgraph "Platform Services"
VAULT[Vault]
PROMETHEUS[Prometheus]
GRAFANA[Grafana]
end
INGRESS --> GATEWAY
GATEWAY --> TRADING
GATEWAY --> BACKTEST
GATEWAY --> ML
GATEWAY --> AGENT
TRADING --> POSTGRES
TRADING --> REDIS
BACKTEST --> POSTGRES
ML --> POSTGRES
AGENT --> POSTGRES
AGENT --> REDIS
TRADING --> VAULT
BACKTEST --> VAULT
ML --> VAULT
AGENT --> VAULT
```
---
## Prerequisites
### 1. Cluster Requirements
**Minimum** (Development):
- **Nodes**: 3 worker nodes
- **vCPUs**: 8 cores per node (24 total)
- **RAM**: 16 GB per node (48 GB total)
- **Storage**: 100 GB SSD per node
- **Kubernetes**: 1.27+
**Recommended** (Production):
- **Nodes**: 6 worker nodes (3 app + 2 data + 1 GPU)
- **vCPUs**: 16 cores per node
- **RAM**: 32 GB per node
- **Storage**: 500 GB NVMe SSD per node
- **Kubernetes**: 1.28+
- **GPU Nodes**: 1+ nodes with NVIDIA T4/A100 GPUs
### 2. Software Requirements
```bash
# Install kubectl
curl -LO "https://dl.k8s.io/release/v1.28.0/bin/linux/amd64/kubectl"
sudo install -o root -g root -m 0755 kubectl /usr/local/bin/kubectl
kubectl version --client
# Install Helm
curl https://raw.githubusercontent.com/helm/helm/main/scripts/get-helm-3 | bash
helm version
# Install Istio CLI (optional)
curl -L https://istio.io/downloadIstio | sh -
sudo mv istio-*/bin/istioctl /usr/local/bin/
istioctl version
```
### 3. Cluster Access Validation
```bash
# Verify cluster access
kubectl cluster-info
# Kubernetes control plane is running at https://...
# Check nodes
kubectl get nodes
# NAME STATUS ROLES AGE VERSION
# node-1 Ready master 30d v1.28.0
# node-2 Ready worker 30d v1.28.0
# node-3 Ready worker 30d v1.28.0
# Check available resources
kubectl top nodes
# NAME CPU(cores) CPU% MEMORY(bytes) MEMORY%
# node-2 500m 6% 4096Mi 25%
# node-3 500m 6% 4096Mi 25%
```
---
## Cluster Setup
### 1. Create Namespace
```bash
# Create foxhunt namespace
kubectl create namespace foxhunt
# Label namespace for monitoring
kubectl label namespace foxhunt monitoring=enabled
# Verify namespace
kubectl get namespace foxhunt
# NAME STATUS AGE
# foxhunt Active 1m
```
### 2. Configure Node Pools
**Application Node Pool** (general workloads):
```bash
# Label application nodes
kubectl label nodes node-2 node-3 node-4 workload-type=application
# Verify labels
kubectl get nodes -l workload-type=application
```
**Data Node Pool** (stateful workloads):
```bash
# Label data nodes (with local SSD)
kubectl label nodes node-5 node-6 workload-type=data
# Add taints to prevent non-data workloads
kubectl taint nodes node-5 node-6 workload-type=data:NoSchedule
```
**GPU Node Pool** (ML training):
```bash
# Label GPU nodes
kubectl label nodes gpu-node-1 workload-type=gpu accelerator=nvidia-t4
# Add taints for GPU workloads only
kubectl taint nodes gpu-node-1 nvidia.com/gpu=present:NoSchedule
# Verify GPU availability
kubectl describe node gpu-node-1 | grep nvidia.com/gpu
# nvidia.com/gpu: 1
```
### 3. Install NVIDIA GPU Operator (if using GPUs)
```bash
# Add NVIDIA Helm repository
helm repo add nvidia https://nvidia.github.io/gpu-operator
helm repo update
# Install GPU Operator
helm install gpu-operator nvidia/gpu-operator \
--namespace gpu-operator-resources \
--create-namespace \
--set driver.enabled=true
# Verify GPU operator
kubectl get pods -n gpu-operator-resources
# NAME READY STATUS RESTARTS AGE
# gpu-feature-discovery-xxxxx 1/1 Running 0 2m
# gpu-operator-xxxxx 1/1 Running 0 2m
# nvidia-container-toolkit-daemonset-xxxxx 1/1 Running 0 2m
# nvidia-dcgm-exporter-xxxxx 1/1 Running 0 2m
# nvidia-device-plugin-daemonset-xxxxx 1/1 Running 0 2m
# nvidia-driver-daemonset-xxxxx 1/1 Running 0 2m
```
### 4. Install Istio Service Mesh (optional)
```bash
# Install Istio with minimal profile
istioctl install --set profile=default -y
# Enable sidecar injection for foxhunt namespace
kubectl label namespace foxhunt istio-injection=enabled
# Verify Istio installation
kubectl get pods -n istio-system
# NAME READY STATUS RESTARTS AGE
# istio-ingressgateway-xxxxx 1/1 Running 0 2m
# istiod-xxxxx 1/1 Running 0 2m
```
### 5. Install External Secrets Operator
```bash
# Add External Secrets Helm repository
helm repo add external-secrets https://charts.external-secrets.io
helm repo update
# Install External Secrets Operator
helm install external-secrets \
external-secrets/external-secrets \
--namespace external-secrets-system \
--create-namespace
# Verify installation
kubectl get pods -n external-secrets-system
# NAME READY STATUS RESTARTS AGE
# external-secrets-xxxxx 1/1 Running 0 1m
# external-secrets-cert-controller-xxxxx 1/1 Running 0 1m
# external-secrets-webhook-xxxxx 1/1 Running 0 1m
```
---
## Helm Chart Deployment
### 1. Create Helm Chart Structure
```bash
# Create Helm chart directory
mkdir -p helm/foxhunt
cd helm/foxhunt
# Create chart structure
cat > Chart.yaml < values.yaml < templates/api-gateway-deployment.yaml <<'EOF'
apiVersion: apps/v1
kind: Deployment
metadata:
name: {{ include "foxhunt.fullname" . }}-api-gateway
labels:
{{- include "foxhunt.labels" . | nindent 4 }}
app.kubernetes.io/component: api-gateway
spec:
{{- if not .Values.apiGateway.autoscaling.enabled }}
replicas: {{ .Values.apiGateway.replicaCount }}
{{- end }}
selector:
matchLabels:
{{- include "foxhunt.selectorLabels" . | nindent 6 }}
app.kubernetes.io/component: api-gateway
template:
metadata:
annotations:
prometheus.io/scrape: "true"
prometheus.io/port: "9091"
prometheus.io/path: "/metrics"
labels:
{{- include "foxhunt.selectorLabels" . | nindent 8 }}
app.kubernetes.io/component: api-gateway
spec:
serviceAccountName: {{ include "foxhunt.serviceAccountName" . }}
securityContext:
runAsNonRoot: true
runAsUser: 1000
fsGroup: 1000
containers:
- name: api-gateway
image: "{{ .Values.global.imageRegistry }}/{{ .Values.apiGateway.image.repository }}:{{ .Values.apiGateway.image.tag }}"
imagePullPolicy: {{ .Values.global.imagePullPolicy }}
ports:
- name: grpc
containerPort: 50051
protocol: TCP
- name: http
containerPort: 8080
protocol: TCP
- name: metrics
containerPort: 9091
protocol: TCP
env:
- name: RUST_LOG
value: "info"
- name: ENVIRONMENT
value: {{ .Values.global.environment }}
- name: DATABASE_URL
valueFrom:
secretKeyRef:
name: foxhunt-secrets
key: database-url
- name: REDIS_URL
valueFrom:
secretKeyRef:
name: foxhunt-secrets
key: redis-url
- name: VAULT_ADDR
value: {{ .Values.vault.address }}
- name: VAULT_TOKEN
valueFrom:
secretKeyRef:
name: foxhunt-secrets
key: vault-token
livenessProbe:
exec:
command:
- grpc_health_probe
- -addr=:50051
initialDelaySeconds: 30
periodSeconds: 10
timeoutSeconds: 5
failureThreshold: 3
readinessProbe:
exec:
command:
- grpc_health_probe
- -addr=:50051
initialDelaySeconds: 10
periodSeconds: 5
timeoutSeconds: 5
failureThreshold: 2
resources:
{{- toYaml .Values.apiGateway.resources | nindent 10 }}
volumeMounts:
- name: config
mountPath: /app/config
readOnly: true
volumes:
- name: config
configMap:
name: {{ include "foxhunt.fullname" . }}-config
EOF
```
**HorizontalPodAutoscaler**:
```bash
cat > templates/api-gateway-hpa.yaml <<'EOF'
{{- if .Values.apiGateway.autoscaling.enabled }}
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: {{ include "foxhunt.fullname" . }}-api-gateway
labels:
{{- include "foxhunt.labels" . | nindent 4 }}
app.kubernetes.io/component: api-gateway
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: {{ include "foxhunt.fullname" . }}-api-gateway
minReplicas: {{ .Values.apiGateway.autoscaling.minReplicas }}
maxReplicas: {{ .Values.apiGateway.autoscaling.maxReplicas }}
metrics:
- type: Resource
resource:
name: cpu
target:
type: Utilization
averageUtilization: {{ .Values.apiGateway.autoscaling.targetCPUUtilizationPercentage }}
- type: Resource
resource:
name: memory
target:
type: Utilization
averageUtilization: 80
- type: Pods
pods:
metric:
name: grpc_requests_per_second
target:
type: AverageValue
averageValue: "100"
behavior:
scaleDown:
stabilizationWindowSeconds: 300
policies:
- type: Percent
value: 50
periodSeconds: 60
- type: Pods
value: 2
periodSeconds: 60
selectPolicy: Min
scaleUp:
stabilizationWindowSeconds: 0
policies:
- type: Percent
value: 100
periodSeconds: 15
- type: Pods
value: 4
periodSeconds: 15
selectPolicy: Max
{{- end }}
EOF
```
**Service**:
```bash
cat > templates/api-gateway-service.yaml <<'EOF'
apiVersion: v1
kind: Service
metadata:
name: {{ include "foxhunt.fullname" . }}-api-gateway
labels:
{{- include "foxhunt.labels" . | nindent 4 }}
app.kubernetes.io/component: api-gateway
annotations:
prometheus.io/scrape: "true"
prometheus.io/port: "9091"
spec:
type: {{ .Values.apiGateway.service.type }}
ports:
- port: {{ .Values.apiGateway.service.grpcPort }}
targetPort: grpc
protocol: TCP
name: grpc
- port: {{ .Values.apiGateway.service.httpPort }}
targetPort: http
protocol: TCP
name: http
- port: {{ .Values.apiGateway.service.metricsPort }}
targetPort: metrics
protocol: TCP
name: metrics
selector:
{{- include "foxhunt.selectorLabels" . | nindent 4 }}
app.kubernetes.io/component: api-gateway
EOF
```
**PodDisruptionBudget**:
```bash
cat > templates/api-gateway-pdb.yaml <<'EOF'
apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
name: {{ include "foxhunt.fullname" . }}-api-gateway
labels:
{{- include "foxhunt.labels" . | nindent 4 }}
app.kubernetes.io/component: api-gateway
spec:
minAvailable: 2
selector:
matchLabels:
{{- include "foxhunt.selectorLabels" . | nindent 6 }}
app.kubernetes.io/component: api-gateway
EOF
```
### 3. Create External Secret for Vault Integration
```bash
cat > templates/external-secret.yaml <<'EOF'
apiVersion: external-secrets.io/v1beta1
kind: ExternalSecret
metadata:
name: {{ include "foxhunt.fullname" . }}-secrets
labels:
{{- include "foxhunt.labels" . | nindent 4 }}
spec:
refreshInterval: 1h
secretStoreRef:
name: vault-backend
kind: ClusterSecretStore
target:
name: foxhunt-secrets
creationPolicy: Owner
data:
- secretKey: database-url
remoteRef:
key: secret/foxhunt/database
property: url
- secretKey: redis-url
remoteRef:
key: secret/foxhunt/redis
property: url
- secretKey: vault-token
remoteRef:
key: secret/foxhunt/vault
property: token
- secretKey: jwt-secret
remoteRef:
key: secret/foxhunt/jwt
property: secret
EOF
```
### 4. Deploy Helm Chart
```bash
# Add dependencies
helm dependency update
# Dry-run deployment
helm install foxhunt . \
--namespace foxhunt \
--dry-run \
--debug
# Install chart
helm install foxhunt . \
--namespace foxhunt \
--create-namespace \
--values values.yaml \
--set postgresql.auth.password=$(openssl rand -base64 32) \
--set redis.auth.password=$(openssl rand -base64 32) \
--timeout 10m \
--wait
# Verify deployment
helm list -n foxhunt
# NAME NAMESPACE REVISION UPDATED STATUS CHART APP VERSION
# foxhunt foxhunt 1 2025-10-22 10:00:00.000000000 +0000 UTC deployed foxhunt-0.1.0 0.1.0
# Check pod status
kubectl get pods -n foxhunt
# NAME READY STATUS RESTARTS AGE
# foxhunt-api-gateway-xxxxx 1/1 Running 0 2m
# foxhunt-trading-service-xxxxx 1/1 Running 0 2m
# foxhunt-backtesting-service-xxxxx 1/1 Running 0 2m
# foxhunt-ml-training-service-xxxxx 1/1 Running 0 2m
# foxhunt-trading-agent-xxxxx 1/1 Running 0 2m
# foxhunt-postgresql-0 1/1 Running 0 2m
# foxhunt-redis-master-0 1/1 Running 0 2m
```
---
## Service Configuration
### 1. ConfigMap for Application Configuration
```bash
cat > templates/configmap.yaml <<'EOF'
apiVersion: v1
kind: ConfigMap
metadata:
name: {{ include "foxhunt.fullname" . }}-config
labels:
{{- include "foxhunt.labels" . | nindent 4 }}
data:
config.toml: |
[server]
environment = "{{ .Values.global.environment }}"
[logging]
level = "info"
format = "json"
[metrics]
enabled = true
prometheus_port = 9091
[health]
enabled = true
http_port = 8080
[trading]
max_position_size = 1000
risk_limit_pct = 2.0
[ml]
model_path = "/app/ml/trained_models"
feature_count = 225
gpu_enabled = true
EOF
kubectl apply -f templates/configmap.yaml -n foxhunt
```
### 2. ServiceAccount with RBAC
```bash
cat > templates/serviceaccount.yaml <<'EOF'
apiVersion: v1
kind: ServiceAccount
metadata:
name: {{ include "foxhunt.serviceAccountName" . }}
labels:
{{- include "foxhunt.labels" . | nindent 4 }}
---
apiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata:
name: {{ include "foxhunt.fullname" . }}
labels:
{{- include "foxhunt.labels" . | nindent 4 }}
rules:
- apiGroups: [""]
resources: ["pods", "pods/log"]
verbs: ["get", "list", "watch"]
- apiGroups: [""]
resources: ["configmaps"]
verbs: ["get", "list"]
- apiGroups: [""]
resources: ["secrets"]
verbs: ["get"]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata:
name: {{ include "foxhunt.fullname" . }}
labels:
{{- include "foxhunt.labels" . | nindent 4 }}
roleRef:
apiGroup: rbac.authorization.k8s.io
kind: Role
name: {{ include "foxhunt.fullname" . }}
subjects:
- kind: ServiceAccount
name: {{ include "foxhunt.serviceAccountName" . }}
namespace: {{ .Release.Namespace }}
EOF
kubectl apply -f templates/serviceaccount.yaml -n foxhunt
```
---
## Scaling Configuration
### 1. Vertical Pod Autoscaler (VPA)
```bash
# Install VPA (if not already installed)
git clone https://github.com/kubernetes/autoscaler.git
cd autoscaler/vertical-pod-autoscaler
./hack/vpa-up.sh
# Create VPA for Trading Service
cat > vpa-trading-service.yaml <<'EOF'
apiVersion: autoscaling.k8s.io/v1
kind: VerticalPodAutoscaler
metadata:
name: foxhunt-trading-service-vpa
namespace: foxhunt
spec:
targetRef:
apiVersion: apps/v1
kind: Deployment
name: foxhunt-trading-service
updatePolicy:
updateMode: "Auto" # Or "Recreate" for stateful services
resourcePolicy:
containerPolicies:
- containerName: trading-service
minAllowed:
cpu: 500m
memory: 512Mi
maxAllowed:
cpu: 8000m
memory: 8Gi
controlledResources:
- cpu
- memory
EOF
kubectl apply -f vpa-trading-service.yaml
```
### 2. Cluster Autoscaler Configuration
```bash
# Configure Cluster Autoscaler (AWS EKS example)
cat > cluster-autoscaler.yaml <<'EOF'
apiVersion: apps/v1
kind: Deployment
metadata:
name: cluster-autoscaler
namespace: kube-system
spec:
replicas: 1
selector:
matchLabels:
app: cluster-autoscaler
template:
metadata:
labels:
app: cluster-autoscaler
spec:
serviceAccountName: cluster-autoscaler
containers:
- image: k8s.gcr.io/autoscaling/cluster-autoscaler:v1.28.0
name: cluster-autoscaler
command:
- ./cluster-autoscaler
- --v=4
- --stderrthreshold=info
- --cloud-provider=aws
- --skip-nodes-with-local-storage=false
- --expander=least-waste
- --node-group-auto-discovery=asg:tag=k8s.io/cluster-autoscaler/enabled,k8s.io/cluster-autoscaler/foxhunt
- --balance-similar-node-groups
- --skip-nodes-with-system-pods=false
resources:
limits:
cpu: 100m
memory: 300Mi
requests:
cpu: 100m
memory: 300Mi
EOF
kubectl apply -f cluster-autoscaler.yaml
```
### 3. PodPriority for Critical Services
```bash
cat > pod-priority-classes.yaml <<'EOF'
apiVersion: scheduling.k8s.io/v1
kind: PriorityClass
metadata:
name: foxhunt-critical
value: 1000000
globalDefault: false
description: "Critical services: API Gateway, Trading Service"
---
apiVersion: scheduling.k8s.io/v1
kind: PriorityClass
metadata:
name: foxhunt-high
value: 100000
globalDefault: false
description: "High priority: Trading Agent, ML Training"
---
apiVersion: scheduling.k8s.io/v1
kind: PriorityClass
metadata:
name: foxhunt-normal
value: 10000
globalDefault: true
description: "Normal priority: Backtesting"
EOF
kubectl apply -f pod-priority-classes.yaml
# Update deployments to use priority classes
# Add to Deployment spec.template.spec:
# priorityClassName: foxhunt-critical
```
---
## Health Checks and Readiness
### 1. Liveness and Readiness Probes
**gRPC Health Probe** (already configured in deployment):
```yaml
livenessProbe:
exec:
command:
- grpc_health_probe
- -addr=:50051
initialDelaySeconds: 30
periodSeconds: 10
timeoutSeconds: 5
failureThreshold: 3
readinessProbe:
exec:
command:
- grpc_health_probe
- -addr=:50051
initialDelaySeconds: 10
periodSeconds: 5
timeoutSeconds: 5
failureThreshold: 2
```
### 2. Startup Probe for Slow-Starting Services
```yaml
# Add to ML Training Service deployment
startupProbe:
exec:
command:
- grpc_health_probe
- -addr=:50054
initialDelaySeconds: 0
periodSeconds: 10
timeoutSeconds: 5
failureThreshold: 30 # 5 minutes max startup time
```
---
## Monitoring and Observability
### 1. ServiceMonitor for Prometheus
```bash
cat > templates/servicemonitor.yaml <<'EOF'
apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata:
name: {{ include "foxhunt.fullname" . }}
labels:
{{- include "foxhunt.labels" . | nindent 4 }}
spec:
selector:
matchLabels:
{{- include "foxhunt.selectorLabels" . | nindent 6 }}
endpoints:
- port: metrics
interval: 15s
path: /metrics
namespaceSelector:
matchNames:
- {{ .Release.Namespace }}
EOF
kubectl apply -f templates/servicemonitor.yaml -n foxhunt
```
### 2. Grafana Dashboards
See [grafana-setup.md](../monitoring/grafana-setup.md) for dashboard configuration.
---
## Troubleshooting
### 1. Pod Stuck in Pending State
```bash
# Describe pod to see events
kubectl describe pod -n foxhunt
# Common issues:
# - Insufficient resources: Check node capacity
kubectl top nodes
# - Unschedulable due to taints: Check tolerations
kubectl get pod -n foxhunt -o yaml | grep -A 10 tolerations
# - PVC not bound: Check PersistentVolumeClaims
kubectl get pvc -n foxhunt
```
### 2. CrashLoopBackOff
```bash
# Check logs
kubectl logs -n foxhunt --previous
# Common issues:
# - Database connection failed: Verify DATABASE_URL secret
kubectl get secret foxhunt-secrets -n foxhunt -o yaml
# - Missing GPU: Verify GPU node availability
kubectl get nodes -l accelerator=nvidia-t4
```
### 3. High Memory Usage
```bash
# Check memory usage
kubectl top pods -n foxhunt --sort-by=memory
# Adjust memory limits
helm upgrade foxhunt . \
--namespace foxhunt \
--set tradingService.resources.limits.memory=4Gi \
--reuse-values
```
---
**End of Kubernetes Deployment Guide**