Files
foxhunt/services/load_tests
jgrusewski 82197efb59 🚀 Wave 127 Wave 2: Execution Validation (6 agents)
**Mission**: Validate frameworks created in Wave 126

**Agent 120b: Prometheus Exporters Fix** ⚠️ Code Complete
- Fixed all 4 services (wrong Prometheus registries)
- API Gateway: Now uses GatewayMetrics registry
- Trading Service: Uses TradingMetricsServer
- Backtesting/ML: Created simple_metrics modules
- Built successfully (1m 51s)
- BLOCKER: Docker rebuild needed for deployment

**Agent 122: E2E Test Execution**  BLOCKED
- Fixed Tonic 0.12 → 0.14 migration (all proto enums)
- 54 E2E tests compile successfully
- BLOCKER: JWT auth not implemented in test framework
- Impact: 0/54 tests can execute

**Agent 123: Load Test Execution**  BLOCKED
- Framework validated (7,960-9,354 req/sec client-side)
- HDR histogram metrics working
- BLOCKER: SQL schema mismatch (price vs limit_price)
- Impact: 100% failure rate (477K attempted, 0 successful)

**Agent 124: Benchmark Execution**  PARTIAL
- Authentication: 4.4μs  (<10μs target)
- Order matching: 1-6μs P99  (<50μs target)
- Component latencies validated
- Gap: E2E, risk, ML benchmarks not executed

**Agent 125: PPO Test Fix**  COMPLETE
- Test already passing (575/575 ML tests)
- 100% pass rate in ML crate
- No fix needed (transient failure)

**Agent 126: Security Hardening**  COMPLETE
- RSA 4096-bit certificates generated and deployed
- All services restarted successfully
- H1 security gap closed

**Wave 2 Results**:
- Achievements: Component latency validated, security hardened, GPU working
- Critical Blockers: 3 identified (E2E auth, load test SQL, Prometheus deployment)
- Production Readiness: 91-92% (unchanged - blockers prevent further validation)

**Files Modified** (21):
- services/integration_tests/* (6 files - E2E test compilation fixes)
- services/*/src/main.rs (3 files - Prometheus exporters)
- services/backtesting_service/src/simple_metrics.rs (new)
- services/ml_training_service/src/simple_metrics.rs (new)
- certs/production/* (RSA 4096-bit certificates)
- services/load_tests/tests/* (relocated)

**Critical Blockers Identified**:
1. E2E: JWT Interceptor missing (2-4h fix)
2. Load: SQL schema mismatch (1-2h fix)
3. Prometheus: Docker rebuild needed (30m)

**Validation Report**: /tmp/wave2_gate_validation.md

**Next**: Deploy 3 blocker-fix agents, then Wave 3

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
2025-10-08 09:41:43 +02:00
..

Load Tests - Trading Service Throughput Validation

Overview

Comprehensive load testing suite for validating trading service throughput and performance under various load scenarios.

Test Scenarios

1. Sustained Load (10,000 orders/sec for 60s)

  • Target: 10,000 orders/second sustained throughput
  • Duration: 60 seconds
  • Concurrent Clients: 100
  • Validates: System stability under sustained load

2. Peak Burst (50,000 orders/sec for 10s)

  • Target: 50,000 orders/second peak burst
  • Duration: 10 seconds
  • Concurrent Clients: 500
  • Validates: System behavior under peak load spikes

3. Market Data Streaming (1M updates)

  • Target: 1,000,000 concurrent market data updates
  • Streams: 1,000 concurrent streams
  • Duration: 30 seconds
  • Validates: Streaming infrastructure capacity

4. Connection Pool Saturation (1,000 clients)

  • Target: 1,000 concurrent clients
  • Requests per Client: 100
  • Validates: Connection pool management and resource limits

Usage

Run All Tests

cargo run -p load_tests --release -- --scenario all

Run Individual Scenarios

# Sustained load
cargo run -p load_tests --release -- --scenario sustained

# Peak burst
cargo run -p load_tests --release -- --scenario burst

# Streaming
cargo run -p load_tests --release -- --scenario streaming

# Connection pool
cargo run -p load_tests --release -- --scenario pool

Custom Configuration

cargo run -p load_tests --release -- \
  --scenario sustained \
  --url http://trading-service:50052 \
  --output /path/to/report.md \
  --verbose

Metrics Collected

Throughput Metrics

  • Requests per second (sustained and peak)
  • Total requests processed
  • Success/failure rates

Latency Distribution

  • P50 (median) latency
  • P95 latency
  • P99 latency
  • Maximum latency

Resource Usage

  • Memory consumption (average)
  • Connection pool utilization
  • Stream management overhead

Output Report

Test results are saved as Markdown reports containing:

  • Executive summary
  • Detailed metrics breakdown
  • Latency distribution charts
  • Resource usage analysis
  • Performance recommendations

Default output: /tmp/WAVE_120_AGENT_5_LOAD_TESTING.md

Prerequisites

  1. Trading Service Running:

    docker-compose up -d trading_service
    # OR
    cargo run -p trading_service
    
  2. Database Available:

    docker-compose up -d postgres redis
    
  3. Sufficient System Resources:

    • 8GB+ RAM recommended
    • Multi-core CPU for parallel clients
    • Network bandwidth for 50k+ req/sec

Architecture

Components

  • Scenarios: Test scenario implementations

    • sustained_load.rs: 10k orders/sec for 60s
    • burst_load.rs: 50k orders/sec for 10s
    • streaming_load.rs: 1M market data updates
    • pool_saturation.rs: 1000 concurrent clients
    • comprehensive.rs: All scenarios sequentially
  • Clients: gRPC client implementations

    • trading_client.rs: Trading service client wrapper
  • Metrics: Performance measurement

    • metrics.rs: HDR histogram-based metrics collection
    • monitor.rs: System resource monitoring

Load Generation Pattern

// Concurrent client pattern
for client_id in 0..NUM_CLIENTS {
    tokio::spawn(async move {
        let client = TradingClient::connect(url).await?;

        // Submit orders with rate limiting
        while duration_remaining {
            client.submit_order(...).await?;
            tokio::time::sleep(rate_limit).await;
        }
    });
}

Performance Targets

Sustained Load

  • Throughput: ≥9,000 req/sec
  • Error Rate: <1%
  • P95 Latency: <10ms

Peak Burst

  • Throughput: ≥40,000 req/sec
  • Error Rate: <5%
  • P99 Latency: <50ms

Streaming

  • Updates: ≥900k received
  • Concurrent Streams: 1000
  • Stream Stability: <1% failures

Connection Pool

  • Concurrent Connections: 1000
  • Error Rate: <5%
  • P99 Latency: <100ms

Troubleshooting

Connection Refused

# Verify trading service is running
grpc_health_probe -addr=localhost:50052

High Error Rates

  • Check system resource limits (ulimit, file descriptors)
  • Verify database connection pool size
  • Review trading service logs for errors

Memory Issues

  • Reduce concurrent clients
  • Enable connection pooling
  • Check for memory leaks in trading service

Integration with CI/CD

# .github/workflows/load-test.yml
- name: Run Load Tests
  run: |
    docker-compose up -d
    cargo run -p load_tests --release -- --scenario all

- name: Upload Report
  uses: actions/upload-artifact@v3
  with:
    name: load-test-report
    path: /tmp/WAVE_120_AGENT_5_LOAD_TESTING.md

Wave 120 Objectives

Agent 5 Tasks:

  • Create load_tests package
  • Implement 4 throughput scenarios
  • Measure latency, throughput, error rates
  • Monitor memory usage
  • Run tests against live service
  • Generate performance report

Expected Outcomes:

  • Validate 10k orders/sec sustained capacity
  • Confirm 50k orders/sec peak burst handling
  • Verify 1M concurrent stream updates
  • Validate 1000+ concurrent client support