Files
foxhunt/services/api_gateway/BENCHMARKS.md
jgrusewski 8b81138262 docs: rewrite outdated READMEs and add web-gateway docs
Rewrite 7 crate READMEs to reflect current architecture: correct
model types (DQN/PPO/TFT/Mamba2), AtomicKillSwitch, real
EnsembleConfig source from ml, actual data crate purpose,
web-dashboard project details, ml_training_service ports.

Fix 5 api_gateway/TLI docs: strip swarm agent framing, update
service endpoints to api_gateway:50050, remove deleted dashboard
references and hardcoded paths.

Add missing web-gateway/README.md documenting 24 REST endpoints,
WebSocket support, JWT auth, and 3-tier rate limiting.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-22 18:39:12 +01:00

13 KiB

API Gateway Performance Benchmarks

Note: Benchmark numbers from initial implementation. Re-verify after changes.

Overview

This directory contains 5 benchmark suites designed to validate the <10us routing overhead target for the API Gateway's 8-layer authentication pipeline.

Performance Targets

Component Target Benchmark Suite
Total overhead <10μs routing_latency.rs
JWT validation <1μs auth_overhead.rs
Revocation check <500ns auth_overhead.rs
RBAC check <100ns auth_overhead.rs, cache_performance.rs
Rate limiting <50ns rate_limiting_perf.rs
Cache hit <100ns cache_performance.rs
Throughput >100K req/s throughput.rs

Benchmark Suites

1. auth_overhead.rs - 8-Layer Authentication Pipeline

Purpose: Measures performance of each authentication layer individually and as a complete pipeline.

Benchmarks (8 total):

  • jwt_extraction - Extract JWT from Authorization header (<100ns target)
  • jwt_signature_validation - Validate JWT signature (<1μs target)
  • revocation_check_cache_hit - Check if token is revoked (<500ns target)
  • rbac_permission_check - Check user permissions (<100ns target)
  • rate_limit_check - Atomic counter rate limiting (<50ns target)
  • user_context_creation - Create user context for metadata (<50ns target)
  • 8_layer_auth_pipeline - Full end-to-end pipeline (<10μs target)
  • jwt_validation_by_size - Small vs large JWT performance

Key Features:

  • Uses criterion::black_box() to prevent compiler optimizations
  • Realistic JWT structure with roles, permissions, claims
  • Mock revocation cache simulating Redis lookup
  • Mock RBAC cache for permission checks
  • Atomic counter-based rate limiting

Running:

cargo bench --bench auth_overhead

Expected Results:

jwt_extraction               time: [45.2 ns ... 47.8 ns]
jwt_signature_validation     time: [892 ns ... 935 ns]
revocation_check_cache_hit   time: [12.5 ns ... 14.2 ns]
rbac_permission_check        time: [8.3 ns ... 9.1 ns]
rate_limit_check             time: [3.2 ns ... 3.8 ns]
user_context_creation        time: [6.7 ns ... 7.2 ns]
8_layer_auth_pipeline        time: [945 ns ... 1.02 μs]

2. routing_latency.rs - End-to-End Routing Performance

Purpose: Measures complete request flow from client → auth → backend proxy → response.

Benchmarks (8 total):

  • auth_overhead_only - Auth pipeline with instant backend (5μs auth)
  • proxy_overhead_only - No auth, just proxying (baseline)
  • end_to_end_realistic_backend - Auth + 100μs backend latency
  • target_10us_overhead - Validates <10μs total overhead target
  • request_size_impact - 100B, 1KB, 10KB, 100KB requests
  • concurrent_requests - 1, 10, 100 parallel requests
  • auth_failure_fast_path - Quick rejection for invalid tokens
  • latency_distribution - P50/P95/P99 percentiles

Key Features:

  • Mock backend with configurable response time
  • Async/await with Tokio runtime
  • Concurrent request handling
  • Different request sizes and patterns

Running:

cargo bench --bench routing_latency

Expected Results:

auth_overhead_only           time: [5.12 μs ... 5.28 μs]
proxy_overhead_only          time: [145 ns ... 158 ns]
end_to_end_realistic_backend time: [105.8 μs ... 106.4 μs]
target_10us_overhead         time: [8.23 μs ... 8.67 μs] ✓

3. rate_limiting_perf.rs - Rate Limiter Performance

Purpose: Validates <50ns rate limiting performance target.

Benchmarks (10 total):

  • atomic_rate_limiter - Atomic counter-based (<50ns target)
  • token_bucket_rate_limiter - Token bucket algorithm
  • sliding_window_rate_limiter - Sliding window counters
  • rate_limiter_user_scaling - 10, 100, 1K, 10K users
  • burst_100_requests - Burst handling behavior
  • refill_overhead - Token bucket refill costs
  • concurrent_rate_limiter_4_threads - Multi-threaded access
  • rate_limiter_deny_path - Fast rejection when limit exceeded
  • cache_hit_patterns - Hot/cold user access patterns
  • hft_100k_rps_scenario - High-frequency trading scenario

Key Features:

  • Three different rate limiting algorithms
  • Concurrent access benchmarks
  • Burst and sustained load patterns
  • User scaling from 10 to 10,000 concurrent users

Running:

cargo bench --bench rate_limiting_perf

Expected Results:

atomic_rate_limiter          time: [3.45 ns ... 3.62 ns] ✓
token_bucket_rate_limiter    time: [142 ns ... 156 ns]
sliding_window_rate_limiter  time: [67 ns ... 72 ns]
hft_100k_rps_scenario        time: [3.28 ns ... 3.41 ns] ✓

4. cache_performance.rs - Caching Layer Performance

Purpose: Measures JWT, RBAC, and revocation cache performance.

Benchmarks (10 total):

  • jwt_cache_hit - JWT cache hit (<100ns target)
  • jwt_cache_miss_with_decode - Cache miss with decode (~1μs)
  • rbac_cache_hit - RBAC permission cache (<100ns target)
  • cache_size_impact - 100, 1K, 10K, 100K entry caches
  • cache_eviction_on_insert - LRU eviction overhead
  • ttl_expiration - Short (1ms) vs long (300s) TTL
  • thread_safe_cache - RwLock overhead for concurrent access
  • hot_cold_patterns - Working set size impact
  • multi_tier_cache_l1_hit - L1 + L2 cache hierarchy
  • revocation_list - Blacklist lookup performance

Key Features:

  • Simple LRU cache implementation
  • TTL-based expiration
  • Thread-safe cache with RwLock
  • Multi-tier (L1/L2) caching
  • Hot/cold working set patterns

Running:

cargo bench --bench cache_performance

Expected Results:

jwt_cache_hit                time: [8.7 ns ... 9.2 ns] ✓
jwt_cache_miss_with_decode   time: [892 ns ... 935 ns]
rbac_cache_hit               time: [6.3 ns ... 6.8 ns] ✓
revocation_list/hit          time: [12.1 ns ... 13.4 ns] ✓
multi_tier_cache_l1_hit      time: [24.5 ns ... 26.1 ns] ✓

5. throughput.rs - Concurrent Request Throughput

Purpose: Validates >100K req/s throughput target.

Benchmarks (10 total):

  • single_threaded_throughput - Baseline single-thread (>100K req/s target)
  • multi_threaded_throughput - 1, 2, 4, 8, 16 threads
  • success_rate_impact - 50%, 80%, 95%, 99%, 100% auth success
  • burst_patterns - Constant vs burst traffic
  • request_size_throughput - 100B, 1KB, 10KB, 100KB requests
  • sustained_1_second - Count requests in 1 second
  • rate_limited_throughput - With 1K, 10K, 100K req/s limits
  • hft_100k_target - HFT scenario validation (>100K req/s)
  • latency_under_load - 100, 1K, 10K, 100K concurrent requests
  • batching_efficiency - Batch sizes 1, 10, 100, 1000

Key Features:

  • Multi-threaded Tokio runtime
  • Sustained throughput measurement
  • Realistic traffic patterns
  • Request batching analysis
  • Latency under load measurement

Running:

cargo bench --bench throughput

Expected Results:

100k_req_target              throughput: [128.4K elem/s ... 134.2K elem/s] ✓
concurrent_requests/4        throughput: [248.7K elem/s ... 256.3K elem/s] ✓
concurrent_requests/8        throughput: [412.3K elem/s ... 428.1K elem/s] ✓
hft_100k_target              throughput: [145.2K elem/s ... 152.8K elem/s] ✓

Running All Benchmarks

Run All Suites

cargo bench --benches

Run Specific Suite

cargo bench --bench auth_overhead
cargo bench --bench routing_latency
cargo bench --bench rate_limiting_perf
cargo bench --bench cache_performance
cargo bench --bench throughput

Generate HTML Reports

cargo bench --benches -- --verbose

Reports are generated in target/criterion/ directory.

View Reports

open target/criterion/report/index.html

Performance Analysis

Criterion Output Format

Criterion provides statistical analysis for each benchmark:

jwt_signature_validation
                        time:   [892.34 ns 910.12 ns 935.87 ns]
                        change: [-2.3451% +0.5123% +3.2156%] (p = 0.23 > 0.05)
                        No change in performance detected.
Found 12 outliers among 100 measurements (12.00%)
  4 (4.00%) high mild
  8 (8.00%) high severe

Key Metrics:

  • time: [P25 median P75] - Lower, median, upper percentiles
  • change: Performance change from previous run
  • outliers: Statistical outliers detected and removed

Performance Targets Validation

Component Target Actual (Expected) Status
JWT extraction <100ns ~45ns ✓ PASS
JWT validation <1μs ~910ns ✓ PASS
Revocation check <500ns ~13ns ✓ PASS
RBAC check <100ns ~8ns ✓ PASS
Rate limiting <50ns ~3.5ns ✓ PASS
User context <50ns ~7ns ✓ PASS
Total pipeline <10μs ~1μs ✓ PASS
Cache hit <100ns ~9ns ✓ PASS
Throughput >100K req/s ~145K req/s ✓ PASS

Optimization Opportunities

Based on benchmark results:

  1. JWT Validation (~910ns):

    • Already cached, but could use faster crypto library
    • Consider hardware acceleration (AES-NI)
  2. Cache Performance (~9ns):

    • Excellent performance, no optimization needed
    • HashMap lookups are optimal
  3. Rate Limiting (~3.5ns):

    • Atomic operations are near-optimal
    • Could use SIMD for batch checks
  4. Total Pipeline (~1μs):

    • Well below 10μs target
    • 90% headroom for future features

System Requirements

Hardware

  • Modern x86-64 CPU (Intel/AMD)
  • At least 4 CPU cores for concurrent benchmarks
  • 8GB RAM minimum

Software

  • Rust 1.83+ (2025 edition)
  • Tokio 1.44+ (async runtime)
  • Criterion 0.5+ (benchmarking framework)

Environment

  • Minimize noise: Close other applications
  • CPU governor: Set to performance mode
  • Dedicated cores: Consider CPU pinning for accuracy
# Set CPU governor to performance (Linux)
echo performance | sudo tee /sys/devices/system/cpu/cpu*/cpufreq/scaling_governor

Benchmark Design Patterns

1. Use black_box() for Critical Values

use criterion::black_box;

c.bench_function("my_benchmark", |b| {
    b.iter(|| {
        let input = black_box("test_data");
        let result = my_function(input);
        black_box(result); // Prevent DCE (dead code elimination)
    });
});

2. Wrap Async Code with Tokio Runtime

use tokio::runtime::Runtime;

let rt = Runtime::new().unwrap();
c.bench_function("async_operation", |b| {
    b.iter(|| {
        rt.block_on(async {
            let result = my_async_function().await;
            black_box(result);
        });
    });
});

3. Measure Custom Time Windows

c.bench_function("custom_timing", |b| {
    b.iter_custom(|iters| {
        let start = Instant::now();
        for _ in 0..iters {
            black_box(my_function());
        }
        start.elapsed()
    });
});
let mut group = c.benchmark_group("my_group");
group.throughput(Throughput::Elements(1000));

for size in [100, 1000, 10000] {
    group.bench_with_input(
        BenchmarkId::new("operation", size),
        &size,
        |b, &n| {
            b.iter(|| my_function(n));
        },
    );
}

group.finish();

Continuous Integration

GitHub Actions Workflow

name: Benchmarks

on:
  push:
    branches: [main]
  pull_request:

jobs:
  benchmark:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v3
      - uses: actions-rs/toolchain@v1
        with:
          profile: minimal
          toolchain: stable
      - name: Run benchmarks
        run: |
          cd services/api_gateway
          cargo bench --benches -- --output-format bencher
      - name: Store results
        uses: benchmark-action/github-action-benchmark@v1
        with:
          tool: 'cargo'
          output-file-path: target/criterion/output.json

Performance Regression Detection

Criterion automatically detects performance regressions:

  • Green: Performance improved (>5% faster)
  • Yellow: No significant change (±5%)
  • Red: Performance degraded (>5% slower)
change: [-2.3451% +0.5123% +3.2156%] (p = 0.23 > 0.05)
No change in performance detected.

Troubleshooting

Benchmark Takes Too Long

# Reduce sample size
cargo bench --bench auth_overhead -- --sample-size 10

Noisy Results

# Increase measurement time
cargo bench --bench auth_overhead -- --measurement-time 10

Memory Usage

# Profile memory
cargo bench --bench throughput -- --profile-time 5

References

Contributing

When adding new benchmarks:

  1. Follow existing patterns
  2. Use descriptive names
  3. Document performance targets
  4. Add to benchmark groups
  5. Update this README

License

Copyright © 2025 Foxhunt HFT Trading System