Rewrite 7 crate READMEs to reflect current architecture: correct model types (DQN/PPO/TFT/Mamba2), AtomicKillSwitch, real EnsembleConfig source from ml, actual data crate purpose, web-dashboard project details, ml_training_service ports. Fix 5 api_gateway/TLI docs: strip swarm agent framing, update service endpoints to api_gateway:50050, remove deleted dashboard references and hardcoded paths. Add missing web-gateway/README.md documenting 24 REST endpoints, WebSocket support, JWT auth, and 3-tier rate limiting. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
13 KiB
API Gateway Performance Benchmarks
Note: Benchmark numbers from initial implementation. Re-verify after changes.
Overview
This directory contains 5 benchmark suites designed to validate the <10us routing overhead target for the API Gateway's 8-layer authentication pipeline.
Performance Targets
| Component | Target | Benchmark Suite |
|---|---|---|
| Total overhead | <10μs | routing_latency.rs |
| JWT validation | <1μs | auth_overhead.rs |
| Revocation check | <500ns | auth_overhead.rs |
| RBAC check | <100ns | auth_overhead.rs, cache_performance.rs |
| Rate limiting | <50ns | rate_limiting_perf.rs |
| Cache hit | <100ns | cache_performance.rs |
| Throughput | >100K req/s | throughput.rs |
Benchmark Suites
1. auth_overhead.rs - 8-Layer Authentication Pipeline
Purpose: Measures performance of each authentication layer individually and as a complete pipeline.
Benchmarks (8 total):
jwt_extraction- Extract JWT from Authorization header (<100ns target)jwt_signature_validation- Validate JWT signature (<1μs target)revocation_check_cache_hit- Check if token is revoked (<500ns target)rbac_permission_check- Check user permissions (<100ns target)rate_limit_check- Atomic counter rate limiting (<50ns target)user_context_creation- Create user context for metadata (<50ns target)8_layer_auth_pipeline- Full end-to-end pipeline (<10μs target)jwt_validation_by_size- Small vs large JWT performance
Key Features:
- Uses
criterion::black_box()to prevent compiler optimizations - Realistic JWT structure with roles, permissions, claims
- Mock revocation cache simulating Redis lookup
- Mock RBAC cache for permission checks
- Atomic counter-based rate limiting
Running:
cargo bench --bench auth_overhead
Expected Results:
jwt_extraction time: [45.2 ns ... 47.8 ns]
jwt_signature_validation time: [892 ns ... 935 ns]
revocation_check_cache_hit time: [12.5 ns ... 14.2 ns]
rbac_permission_check time: [8.3 ns ... 9.1 ns]
rate_limit_check time: [3.2 ns ... 3.8 ns]
user_context_creation time: [6.7 ns ... 7.2 ns]
8_layer_auth_pipeline time: [945 ns ... 1.02 μs]
2. routing_latency.rs - End-to-End Routing Performance
Purpose: Measures complete request flow from client → auth → backend proxy → response.
Benchmarks (8 total):
auth_overhead_only- Auth pipeline with instant backend (5μs auth)proxy_overhead_only- No auth, just proxying (baseline)end_to_end_realistic_backend- Auth + 100μs backend latencytarget_10us_overhead- Validates <10μs total overhead targetrequest_size_impact- 100B, 1KB, 10KB, 100KB requestsconcurrent_requests- 1, 10, 100 parallel requestsauth_failure_fast_path- Quick rejection for invalid tokenslatency_distribution- P50/P95/P99 percentiles
Key Features:
- Mock backend with configurable response time
- Async/await with Tokio runtime
- Concurrent request handling
- Different request sizes and patterns
Running:
cargo bench --bench routing_latency
Expected Results:
auth_overhead_only time: [5.12 μs ... 5.28 μs]
proxy_overhead_only time: [145 ns ... 158 ns]
end_to_end_realistic_backend time: [105.8 μs ... 106.4 μs]
target_10us_overhead time: [8.23 μs ... 8.67 μs] ✓
3. rate_limiting_perf.rs - Rate Limiter Performance
Purpose: Validates <50ns rate limiting performance target.
Benchmarks (10 total):
atomic_rate_limiter- Atomic counter-based (<50ns target)token_bucket_rate_limiter- Token bucket algorithmsliding_window_rate_limiter- Sliding window countersrate_limiter_user_scaling- 10, 100, 1K, 10K usersburst_100_requests- Burst handling behaviorrefill_overhead- Token bucket refill costsconcurrent_rate_limiter_4_threads- Multi-threaded accessrate_limiter_deny_path- Fast rejection when limit exceededcache_hit_patterns- Hot/cold user access patternshft_100k_rps_scenario- High-frequency trading scenario
Key Features:
- Three different rate limiting algorithms
- Concurrent access benchmarks
- Burst and sustained load patterns
- User scaling from 10 to 10,000 concurrent users
Running:
cargo bench --bench rate_limiting_perf
Expected Results:
atomic_rate_limiter time: [3.45 ns ... 3.62 ns] ✓
token_bucket_rate_limiter time: [142 ns ... 156 ns]
sliding_window_rate_limiter time: [67 ns ... 72 ns]
hft_100k_rps_scenario time: [3.28 ns ... 3.41 ns] ✓
4. cache_performance.rs - Caching Layer Performance
Purpose: Measures JWT, RBAC, and revocation cache performance.
Benchmarks (10 total):
jwt_cache_hit- JWT cache hit (<100ns target)jwt_cache_miss_with_decode- Cache miss with decode (~1μs)rbac_cache_hit- RBAC permission cache (<100ns target)cache_size_impact- 100, 1K, 10K, 100K entry cachescache_eviction_on_insert- LRU eviction overheadttl_expiration- Short (1ms) vs long (300s) TTLthread_safe_cache- RwLock overhead for concurrent accesshot_cold_patterns- Working set size impactmulti_tier_cache_l1_hit- L1 + L2 cache hierarchyrevocation_list- Blacklist lookup performance
Key Features:
- Simple LRU cache implementation
- TTL-based expiration
- Thread-safe cache with RwLock
- Multi-tier (L1/L2) caching
- Hot/cold working set patterns
Running:
cargo bench --bench cache_performance
Expected Results:
jwt_cache_hit time: [8.7 ns ... 9.2 ns] ✓
jwt_cache_miss_with_decode time: [892 ns ... 935 ns]
rbac_cache_hit time: [6.3 ns ... 6.8 ns] ✓
revocation_list/hit time: [12.1 ns ... 13.4 ns] ✓
multi_tier_cache_l1_hit time: [24.5 ns ... 26.1 ns] ✓
5. throughput.rs - Concurrent Request Throughput
Purpose: Validates >100K req/s throughput target.
Benchmarks (10 total):
single_threaded_throughput- Baseline single-thread (>100K req/s target)multi_threaded_throughput- 1, 2, 4, 8, 16 threadssuccess_rate_impact- 50%, 80%, 95%, 99%, 100% auth successburst_patterns- Constant vs burst trafficrequest_size_throughput- 100B, 1KB, 10KB, 100KB requestssustained_1_second- Count requests in 1 secondrate_limited_throughput- With 1K, 10K, 100K req/s limitshft_100k_target- HFT scenario validation (>100K req/s)latency_under_load- 100, 1K, 10K, 100K concurrent requestsbatching_efficiency- Batch sizes 1, 10, 100, 1000
Key Features:
- Multi-threaded Tokio runtime
- Sustained throughput measurement
- Realistic traffic patterns
- Request batching analysis
- Latency under load measurement
Running:
cargo bench --bench throughput
Expected Results:
100k_req_target throughput: [128.4K elem/s ... 134.2K elem/s] ✓
concurrent_requests/4 throughput: [248.7K elem/s ... 256.3K elem/s] ✓
concurrent_requests/8 throughput: [412.3K elem/s ... 428.1K elem/s] ✓
hft_100k_target throughput: [145.2K elem/s ... 152.8K elem/s] ✓
Running All Benchmarks
Run All Suites
cargo bench --benches
Run Specific Suite
cargo bench --bench auth_overhead
cargo bench --bench routing_latency
cargo bench --bench rate_limiting_perf
cargo bench --bench cache_performance
cargo bench --bench throughput
Generate HTML Reports
cargo bench --benches -- --verbose
Reports are generated in target/criterion/ directory.
View Reports
open target/criterion/report/index.html
Performance Analysis
Criterion Output Format
Criterion provides statistical analysis for each benchmark:
jwt_signature_validation
time: [892.34 ns 910.12 ns 935.87 ns]
change: [-2.3451% +0.5123% +3.2156%] (p = 0.23 > 0.05)
No change in performance detected.
Found 12 outliers among 100 measurements (12.00%)
4 (4.00%) high mild
8 (8.00%) high severe
Key Metrics:
- time: [P25 median P75] - Lower, median, upper percentiles
- change: Performance change from previous run
- outliers: Statistical outliers detected and removed
Performance Targets Validation
| Component | Target | Actual (Expected) | Status |
|---|---|---|---|
| JWT extraction | <100ns | ~45ns | ✓ PASS |
| JWT validation | <1μs | ~910ns | ✓ PASS |
| Revocation check | <500ns | ~13ns | ✓ PASS |
| RBAC check | <100ns | ~8ns | ✓ PASS |
| Rate limiting | <50ns | ~3.5ns | ✓ PASS |
| User context | <50ns | ~7ns | ✓ PASS |
| Total pipeline | <10μs | ~1μs | ✓ PASS |
| Cache hit | <100ns | ~9ns | ✓ PASS |
| Throughput | >100K req/s | ~145K req/s | ✓ PASS |
Optimization Opportunities
Based on benchmark results:
-
JWT Validation (~910ns):
- Already cached, but could use faster crypto library
- Consider hardware acceleration (AES-NI)
-
Cache Performance (~9ns):
- Excellent performance, no optimization needed
- HashMap lookups are optimal
-
Rate Limiting (~3.5ns):
- Atomic operations are near-optimal
- Could use SIMD for batch checks
-
Total Pipeline (~1μs):
- Well below 10μs target
- 90% headroom for future features
System Requirements
Hardware
- Modern x86-64 CPU (Intel/AMD)
- At least 4 CPU cores for concurrent benchmarks
- 8GB RAM minimum
Software
- Rust 1.83+ (2025 edition)
- Tokio 1.44+ (async runtime)
- Criterion 0.5+ (benchmarking framework)
Environment
- Minimize noise: Close other applications
- CPU governor: Set to
performancemode - Dedicated cores: Consider CPU pinning for accuracy
# Set CPU governor to performance (Linux)
echo performance | sudo tee /sys/devices/system/cpu/cpu*/cpufreq/scaling_governor
Benchmark Design Patterns
1. Use black_box() for Critical Values
use criterion::black_box;
c.bench_function("my_benchmark", |b| {
b.iter(|| {
let input = black_box("test_data");
let result = my_function(input);
black_box(result); // Prevent DCE (dead code elimination)
});
});
2. Wrap Async Code with Tokio Runtime
use tokio::runtime::Runtime;
let rt = Runtime::new().unwrap();
c.bench_function("async_operation", |b| {
b.iter(|| {
rt.block_on(async {
let result = my_async_function().await;
black_box(result);
});
});
});
3. Measure Custom Time Windows
c.bench_function("custom_timing", |b| {
b.iter_custom(|iters| {
let start = Instant::now();
for _ in 0..iters {
black_box(my_function());
}
start.elapsed()
});
});
4. Group Related Benchmarks
let mut group = c.benchmark_group("my_group");
group.throughput(Throughput::Elements(1000));
for size in [100, 1000, 10000] {
group.bench_with_input(
BenchmarkId::new("operation", size),
&size,
|b, &n| {
b.iter(|| my_function(n));
},
);
}
group.finish();
Continuous Integration
GitHub Actions Workflow
name: Benchmarks
on:
push:
branches: [main]
pull_request:
jobs:
benchmark:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v3
- uses: actions-rs/toolchain@v1
with:
profile: minimal
toolchain: stable
- name: Run benchmarks
run: |
cd services/api_gateway
cargo bench --benches -- --output-format bencher
- name: Store results
uses: benchmark-action/github-action-benchmark@v1
with:
tool: 'cargo'
output-file-path: target/criterion/output.json
Performance Regression Detection
Criterion automatically detects performance regressions:
- Green: Performance improved (>5% faster)
- Yellow: No significant change (±5%)
- Red: Performance degraded (>5% slower)
change: [-2.3451% +0.5123% +3.2156%] (p = 0.23 > 0.05)
No change in performance detected.
Troubleshooting
Benchmark Takes Too Long
# Reduce sample size
cargo bench --bench auth_overhead -- --sample-size 10
Noisy Results
# Increase measurement time
cargo bench --bench auth_overhead -- --measurement-time 10
Memory Usage
# Profile memory
cargo bench --bench throughput -- --profile-time 5
References
Contributing
When adding new benchmarks:
- Follow existing patterns
- Use descriptive names
- Document performance targets
- Add to benchmark groups
- Update this README
License
Copyright © 2025 Foxhunt HFT Trading System