Move 17 library crates into crates/, CLI binary into bin/fxt, consolidate 10 test crates into testing/, split config crate from deployment config files. Root directory reduced from 38+ to ~17 directories. All Cargo.toml paths and build.rs proto refs updated. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
API Gateway Load Testing Framework
Comprehensive load testing infrastructure for validating API Gateway performance under high concurrency.
Overview
This framework provides four test scenarios with detailed metrics collection, time-series analysis, and HTML report generation:
- Normal Load: 1K concurrent clients for 60 seconds
- Spike Load: 0→10K clients in 10s, sustain 60s
- Sustained Load: 100 clients for 24 hours (endurance test)
- Stress Test: Incrementally increase load until failure
Architecture
┌─────────────────────┐
│ Test Orchestrator │
└──────────┬──────────┘
│
├─────────────────┬─────────────────┬─────────────────┐
│ │ │ │
┌──────▼──────┐ ┌──────▼──────┐ ┌──────▼──────┐ ┌──────▼──────┐
│ Normal │ │ Spike │ │ Sustained │ │ Stress │
│ Load │ │ Load │ │ Load │ │ Test │
└──────┬──────┘ └──────┬──────┘ └──────┬──────┘ └──────┬──────┘
│ │ │ │
└─────────────────┴─────────────────┴─────────────────┘
│
┌───────────▼───────────┐
│ Virtual Client Pool │
│ (Authenticated HTTP │
│ + Mixed Workload) │
└───────────┬───────────┘
│
┌───────────▼───────────┐
│ Metrics Collector │
│ - HDR Histogram │
│ - Time Series Data │
│ - Per-Service Stats │
└───────────┬───────────┘
│
┌───────────▼───────────┐
│ Report Generator │
│ - HTML + SVG Charts │
│ - Capacity Analysis │
└───────────────────────┘
Installation
cd /home/jgrusewski/Work/foxhunt/services/api_gateway/load_tests
cargo build --release
Usage
Run Individual Scenarios
# Normal load test (1K clients, 60s)
cargo run --release -- normal \
--gateway-url http://localhost:50050 \
--num-clients 1000 \
--duration-secs 60
# Spike load test (0→10K in 10s, sustain 60s)
cargo run --release -- spike \
--gateway-url http://localhost:50050 \
--target-clients 10000 \
--ramp-up-secs 10 \
--sustain-secs 60
# Sustained load test (100 clients, 24h)
cargo run --release -- sustained \
--gateway-url http://localhost:50050 \
--num-clients 100 \
--duration-secs 86400
# Stress test (incrementally increase until failure)
cargo run --release -- stress \
--gateway-url http://localhost:50050 \
--initial-clients 100 \
--increment 100 \
--increment-interval-secs 60 \
--max-p99-latency-ms 50.0 \
--max-error-rate-pct 5.0
Run All Scenarios
cargo run --release -- all --gateway-url http://localhost:50050
Metrics Collected
Latency Statistics
- Min/Max/Mean: Full latency range
- Percentiles: P50, P90, P95, P99, P99.9
- Standard Deviation: Latency consistency
Request Breakdown
- Total Requests: Aggregate count
- Successful: 2xx responses
- Failed: 4xx/5xx errors
- Timeout: Connection/request timeouts
- Rate Limited: 429 responses
- Circuit Breaker: 503 responses
Time Series Data (1-second intervals)
- Requests per second (RPS)
- P99 latency
- Error rate percentage
- Active client count
Per-Service Statistics
- Trading Service: Order submission, position queries
- Backtesting Service: Backtest execution
- ML Training Service: Model training requests
Workload Distribution
Mixed workload simulates realistic usage:
- 60% - Order submissions
- 30% - Position queries
- 8% - Backtesting requests
- 2% - ML training requests
Each client has random think time (1-50ms) between requests to simulate human behavior.
Report Generation
HTML reports are automatically generated with:
- Summary Cards: Total requests, RPS, error rate, P99 latency
- Latency Table: All percentiles with statistics
- Request Breakdown: Success/failure categorization
- Performance Charts (SVG):
- Requests per second over time
- P99 latency over time
- Error rate over time
- Capacity Recommendations: Based on observed performance
Success Criteria
Normal Load
- Target: <10ms P99 latency, 0% errors
- Pass: Error rate < 1%, P99 < 10ms
- Fail: Error rate ≥ 5%, P99 ≥ 50ms
Spike Load
- Target: Graceful handling, circuit breakers activate
- Pass: Error rate < 10%, circuit breakers respond correctly
- Fail: System crashes, uncontrolled cascading failures
Sustained Load
- Target: No memory leaks, stable latency
- Pass: Latency drift < 5%, error rate stddev < 2%
- Fail: Latency increases > 10%, memory exhaustion
Stress Test
- Target: Identify capacity limits
- Pass: Breaking point identified with clear bottleneck
- Fail: Undefined behavior, data corruption
Example Report Output
Load Test Report: Normal Load Test
Test Period: 2025-10-03 12:00:00 to 2025-10-03 12:01:00
Duration: 60 seconds (0.02 hours)
Summary:
┌──────────────────┬──────────┐
│ Total Requests │ 120,000 │
│ Requests/Second │ 2,000 │
│ Error Rate │ 0.12% │
│ P99 Latency │ 8.5ms │
└──────────────────┴──────────┘
Latency Statistics:
┌──────────┬──────────┐
│ P50 │ 3.2ms │
│ P90 │ 5.1ms │
│ P95 │ 6.8ms │
│ P99 │ 8.5ms │
│ P99.9 │ 12.3ms │
└──────────┴──────────┘
Capacity Recommendation:
✓ System handled 1,000 clients with 0.12% error rate
✓ P99 latency well within 10ms target
✓ Safe for production at this load level
Load Generator Resources
Requirements
- CPU: 4+ cores for 1K clients, 8+ cores for 10K clients
- Memory: 2GB for 1K clients, 8GB for 10K clients
- Network: Low-latency connection to API Gateway
Monitoring Load Generator
The framework monitors its own resource usage to ensure the load generator doesn't become a bottleneck. If you see warnings about load generator CPU/memory, consider:
- Running on a larger machine
- Distributing load across multiple generators
- Reducing concurrent client count
Advanced Configuration
Custom JWT Token
Set authentication parameters in the code:
let auth_client = AuthenticatedClient::new(
gateway_url,
"your-jwt-secret",
"user-id",
"username"
).await?;
Custom Test Duration
All scenarios support custom durations:
# Extended normal load test (5 minutes)
cargo run --release -- normal --duration-secs 300
# Long-running stress test
cargo run --release -- stress --increment-interval-secs 300
Troubleshooting
Connection Refused
Error: Connection refused (os error 111)
Solution: Ensure API Gateway is running at http://localhost:50050
Too Many Open Files
Error: Too many open files (os error 24)
Solution: Increase system file descriptor limit:
ulimit -n 10000
Memory Exhaustion
Error: Cannot allocate memory
Solution: Reduce --num-clients or run on a larger machine
High Latency from Generator
Warning: Load generator CPU > 80%, results may be unreliable
Solution: Use a more powerful machine or reduce client count
Integration with CI/CD
GitHub Actions Example
name: Load Testing
on:
schedule:
- cron: '0 0 * * 0' # Weekly
jobs:
load-test:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v3
- name: Start API Gateway
run: |
docker-compose up -d api_gateway
sleep 10
- name: Run Load Tests
run: |
cd services/api_gateway/load_tests
cargo run --release -- normal
- name: Upload Reports
uses: actions/upload-artifact@v3
with:
name: load-test-reports
path: |
normal_load_report.html
*.svg
Performance Baseline
Expected results for reference hardware (AWS c5.4xlarge):
| Scenario | Clients | RPS | P99 Latency | Error Rate |
|---|---|---|---|---|
| Normal Load | 1,000 | 2,000 | 8ms | <0.1% |
| Spike Load | 10,000 | 8,000 | 25ms | <2% |
| Sustained Load | 100 | 200 | 5ms | <0.01% |
| Stress Test | 5,000 | 5,000 | 45ms | Breaking |
License
Part of the Foxhunt HFT Trading System - MIT OR Apache-2.0