- Updated Cargo.lock with latest compatible versions - ML crate: Added async-stream 0.3 for stream processing - Trading engine: Updated audit trail dependencies - Storage crate: Dependency cleanup and optimization - API gateway load tests: Added benchmarking dependencies - All dependency updates tested with clean compilation
API Gateway Load Testing Framework
Comprehensive load testing infrastructure for validating API Gateway performance under high concurrency.
Overview
This framework provides four test scenarios with detailed metrics collection, time-series analysis, and HTML report generation:
- Normal Load: 1K concurrent clients for 60 seconds
- Spike Load: 0→10K clients in 10s, sustain 60s
- Sustained Load: 100 clients for 24 hours (endurance test)
- Stress Test: Incrementally increase load until failure
Architecture
┌─────────────────────┐
│ Test Orchestrator │
└──────────┬──────────┘
│
├─────────────────┬─────────────────┬─────────────────┐
│ │ │ │
┌──────▼──────┐ ┌──────▼──────┐ ┌──────▼──────┐ ┌──────▼──────┐
│ Normal │ │ Spike │ │ Sustained │ │ Stress │
│ Load │ │ Load │ │ Load │ │ Test │
└──────┬──────┘ └──────┬──────┘ └──────┬──────┘ └──────┬──────┘
│ │ │ │
└─────────────────┴─────────────────┴─────────────────┘
│
┌───────────▼───────────┐
│ Virtual Client Pool │
│ (Authenticated HTTP │
│ + Mixed Workload) │
└───────────┬───────────┘
│
┌───────────▼───────────┐
│ Metrics Collector │
│ - HDR Histogram │
│ - Time Series Data │
│ - Per-Service Stats │
└───────────┬───────────┘
│
┌───────────▼───────────┐
│ Report Generator │
│ - HTML + SVG Charts │
│ - Capacity Analysis │
└───────────────────────┘
Installation
cd /home/jgrusewski/Work/foxhunt/services/api_gateway/load_tests
cargo build --release
Usage
Run Individual Scenarios
# Normal load test (1K clients, 60s)
cargo run --release -- normal \
--gateway-url http://localhost:50050 \
--num-clients 1000 \
--duration-secs 60
# Spike load test (0→10K in 10s, sustain 60s)
cargo run --release -- spike \
--gateway-url http://localhost:50050 \
--target-clients 10000 \
--ramp-up-secs 10 \
--sustain-secs 60
# Sustained load test (100 clients, 24h)
cargo run --release -- sustained \
--gateway-url http://localhost:50050 \
--num-clients 100 \
--duration-secs 86400
# Stress test (incrementally increase until failure)
cargo run --release -- stress \
--gateway-url http://localhost:50050 \
--initial-clients 100 \
--increment 100 \
--increment-interval-secs 60 \
--max-p99-latency-ms 50.0 \
--max-error-rate-pct 5.0
Run All Scenarios
cargo run --release -- all --gateway-url http://localhost:50050
Metrics Collected
Latency Statistics
- Min/Max/Mean: Full latency range
- Percentiles: P50, P90, P95, P99, P99.9
- Standard Deviation: Latency consistency
Request Breakdown
- Total Requests: Aggregate count
- Successful: 2xx responses
- Failed: 4xx/5xx errors
- Timeout: Connection/request timeouts
- Rate Limited: 429 responses
- Circuit Breaker: 503 responses
Time Series Data (1-second intervals)
- Requests per second (RPS)
- P99 latency
- Error rate percentage
- Active client count
Per-Service Statistics
- Trading Service: Order submission, position queries
- Backtesting Service: Backtest execution
- ML Training Service: Model training requests
Workload Distribution
Mixed workload simulates realistic usage:
- 60% - Order submissions
- 30% - Position queries
- 8% - Backtesting requests
- 2% - ML training requests
Each client has random think time (1-50ms) between requests to simulate human behavior.
Report Generation
HTML reports are automatically generated with:
- Summary Cards: Total requests, RPS, error rate, P99 latency
- Latency Table: All percentiles with statistics
- Request Breakdown: Success/failure categorization
- Performance Charts (SVG):
- Requests per second over time
- P99 latency over time
- Error rate over time
- Capacity Recommendations: Based on observed performance
Success Criteria
Normal Load
- Target: <10ms P99 latency, 0% errors
- Pass: Error rate < 1%, P99 < 10ms
- Fail: Error rate ≥ 5%, P99 ≥ 50ms
Spike Load
- Target: Graceful handling, circuit breakers activate
- Pass: Error rate < 10%, circuit breakers respond correctly
- Fail: System crashes, uncontrolled cascading failures
Sustained Load
- Target: No memory leaks, stable latency
- Pass: Latency drift < 5%, error rate stddev < 2%
- Fail: Latency increases > 10%, memory exhaustion
Stress Test
- Target: Identify capacity limits
- Pass: Breaking point identified with clear bottleneck
- Fail: Undefined behavior, data corruption
Example Report Output
Load Test Report: Normal Load Test
Test Period: 2025-10-03 12:00:00 to 2025-10-03 12:01:00
Duration: 60 seconds (0.02 hours)
Summary:
┌──────────────────┬──────────┐
│ Total Requests │ 120,000 │
│ Requests/Second │ 2,000 │
│ Error Rate │ 0.12% │
│ P99 Latency │ 8.5ms │
└──────────────────┴──────────┘
Latency Statistics:
┌──────────┬──────────┐
│ P50 │ 3.2ms │
│ P90 │ 5.1ms │
│ P95 │ 6.8ms │
│ P99 │ 8.5ms │
│ P99.9 │ 12.3ms │
└──────────┴──────────┘
Capacity Recommendation:
✓ System handled 1,000 clients with 0.12% error rate
✓ P99 latency well within 10ms target
✓ Safe for production at this load level
Load Generator Resources
Requirements
- CPU: 4+ cores for 1K clients, 8+ cores for 10K clients
- Memory: 2GB for 1K clients, 8GB for 10K clients
- Network: Low-latency connection to API Gateway
Monitoring Load Generator
The framework monitors its own resource usage to ensure the load generator doesn't become a bottleneck. If you see warnings about load generator CPU/memory, consider:
- Running on a larger machine
- Distributing load across multiple generators
- Reducing concurrent client count
Advanced Configuration
Custom JWT Token
Set authentication parameters in the code:
let auth_client = AuthenticatedClient::new(
gateway_url,
"your-jwt-secret",
"user-id",
"username"
).await?;
Custom Test Duration
All scenarios support custom durations:
# Extended normal load test (5 minutes)
cargo run --release -- normal --duration-secs 300
# Long-running stress test
cargo run --release -- stress --increment-interval-secs 300
Troubleshooting
Connection Refused
Error: Connection refused (os error 111)
Solution: Ensure API Gateway is running at http://localhost:50050
Too Many Open Files
Error: Too many open files (os error 24)
Solution: Increase system file descriptor limit:
ulimit -n 10000
Memory Exhaustion
Error: Cannot allocate memory
Solution: Reduce --num-clients or run on a larger machine
High Latency from Generator
Warning: Load generator CPU > 80%, results may be unreliable
Solution: Use a more powerful machine or reduce client count
Integration with CI/CD
GitHub Actions Example
name: Load Testing
on:
schedule:
- cron: '0 0 * * 0' # Weekly
jobs:
load-test:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v3
- name: Start API Gateway
run: |
docker-compose up -d api_gateway
sleep 10
- name: Run Load Tests
run: |
cd services/api_gateway/load_tests
cargo run --release -- normal
- name: Upload Reports
uses: actions/upload-artifact@v3
with:
name: load-test-reports
path: |
normal_load_report.html
*.svg
Performance Baseline
Expected results for reference hardware (AWS c5.4xlarge):
| Scenario | Clients | RPS | P99 Latency | Error Rate |
|---|---|---|---|---|
| Normal Load | 1,000 | 2,000 | 8ms | <0.1% |
| Spike Load | 10,000 | 8,000 | 25ms | <2% |
| Sustained Load | 100 | 200 | 5ms | <0.01% |
| Stress Test | 5,000 | 5,000 | 45ms | Breaking |
License
Part of the Foxhunt HFT Trading System - MIT OR Apache-2.0