Files
foxhunt/docs/testing/WAVE4_EXECUTION_GUIDE.md
jgrusewski 7458f1be01 feat(wave12): E2E validation complete - 225-feature pipeline ready
 Validation Results:
- PPO training: 24.2s (1 epoch, 950 samples, dim=225)
- Feature extraction: 105μs/bar (9.5x faster than target)
- Model checkpoint: 293KB (147KB actor + 146KB critic)
- GPU memory: 145MB used (96.4% headroom)
- Zero dimension mismatches

📊 Success Criteria (5/5):
 Feature dimension = 225 (Wave C 201 + Wave D 24)
 Model state_dim = 225
 Training completed without errors
 Checkpoint saved successfully
 No dimension mismatch errors

📁 Training Data Ready:
- ES.FUT: 2.9MB, 180 days
- NQ.FUT: 4.4MB, 180 days
- 6E.FUT: 2.8MB, 180 days
- ZN.FUT: 65KB, 90 days (clean)

🚀 Next: Full production model retraining (4 models, ~10min GPU time)

🤖 Generated with Claude Code (https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
2025-10-22 22:48:04 +02:00

15 KiB

Wave 4 Test Execution Guide

Version: 1.0 Date: 2025-10-22 Status: Production Ready Author: Agent W4-5 (Test Documentation Specialist)


Quick Start

# Run all tests (fastest → slowest)
cargo test --workspace              # All tests (~5-10 min)
cargo test --lib                    # Unit tests only (~30 sec)
cargo test --test "*"               # Integration tests only (~2-3 min)
cargo test -p tli e2e::             # TLI E2E tests (~1 min)

# Run specific test
cargo test test_order_matching_multiple_fills

# Run tests for a single crate
cargo test -p trading_engine
cargo test -p ml
cargo test -p api_gateway

# Run with output
cargo test -- --nocapture           # Show println! output
cargo test -- --show-output         # Show output even for passing tests

# Run tests matching a pattern
cargo test integration::            # All integration tests
cargo test kelly                    # All tests with "kelly" in the name
cargo test regime                   # All Wave D regime tests

1. Unit Tests

Run All Unit Tests

cargo test --workspace --lib

Expected: ~7,000 tests, ~30 seconds

Run Unit Tests for Specific Module

# ML models
cargo test -p ml --lib test_mamba2
cargo test -p ml --lib test_dqn
cargo test -p ml --lib test_ppo
cargo test -p ml --lib test_tft

# Trading engine
cargo test -p trading_engine --lib test_order_matching
cargo test -p trading_engine --lib test_position_manager

# Risk management
cargo test -p risk --lib test_circuit_breaker
cargo test -p risk --lib test_var_calculator

# Adaptive strategy (Wave D)
cargo test -p adaptive-strategy --lib test_kelly_criterion
cargo test -p adaptive-strategy --lib test_regime_detection
cargo test -p adaptive-strategy --lib test_dynamic_stops

Run Doc Tests

cargo test --doc                    # All documentation examples
cargo test -p common --doc          # common crate docs only

2. Integration Tests

Prerequisites

# Start Docker services
docker-compose up -d

# Verify services are healthy
docker-compose ps
# Expected: postgres, redis, vault all "healthy"

# Apply database migrations
cargo sqlx migrate run

Run All Integration Tests

cargo test --workspace --test "*"

Expected: ~1,260 tests, ~2-3 minutes

Service-Specific Integration Tests

# API Gateway
cargo test -p api_gateway --test integration_auth
cargo test -p api_gateway --test integration_routing
cargo test -p api_gateway --test integration_rate_limiting

# Trading Service
cargo test -p trading_service --test integration_order_submission
cargo test -p trading_service --test integration_position_tracking

# Backtesting Service
cargo test -p backtesting_service --test integration_dbn_loading
cargo test -p backtesting_service --test integration_wave_d_backtest

# ML Training Service
cargo test -p ml_training_service --test integration_parquet_training
cargo test -p ml_training_service --test integration_job_spawner
cargo test -p ml_training_service --test integration_grpc_streaming

# Trading Agent Service
cargo test -p trading_agent_service --test integration_kelly_regime
cargo test -p trading_agent_service --test integration_dynamic_stop_loss
cargo test -p trading_agent_service --test test_wave_d_end_to_end

Database Integration Tests

# Run with serial execution (database writes)
cargo test --test "*" -- --test-threads=1

# Or use the serial attribute (automatic)
cargo test integration_database

gRPC Integration Tests

cargo test integration_grpc
cargo test test_grpc_streaming
cargo test test_tonic_health

3. End-to-End (E2E) Tests

TLI E2E Tests

# All TLI E2E tests
cargo test -p tli e2e::

# Specific workflows
cargo test -p tli e2e::test_auth_login_flow
cargo test -p tli e2e::test_trade_submission_flow
cargo test -p tli e2e::test_ml_training_flow
cargo test -p tli e2e::test_regime_detection_flow

Cross-Service E2E Tests

# Run from tests/e2e directory
cargo test -p foxhunt_e2e
cargo test -p foxhunt_e2e e2e_trade_lifecycle
cargo test -p foxhunt_e2e e2e_ml_training_pipeline

Prerequisites:

  • All services running (cargo run -p api_gateway &, etc.)
  • Database seeded with test data
  • Redis available

4. Stress & Load Tests

Run Stress Tests

# Enable stress test feature
cargo test --release --features stress-test

# Specific stress tests
cargo test -p stress_tests --release chaos_testing
cargo test -p stress_tests --release concurrent_order_stress
cargo test -p stress_tests --release memory_leak_detection

Load Testing

# API Gateway load tests
cargo test -p api_gateway_load_tests --release throughput_tests

# Trading Service load tests
cargo test -p trading_service_load_tests --release order_submission_load

# Integration load tests
cargo test -p integration_load_tests --release

Note: Stress tests can take 30 seconds to 5 minutes. Run with --release for accurate performance measurement.


5. Performance Benchmarks

Run All Benchmarks

cargo bench --workspace

Expected: ~46 benchmarks, ~10-15 minutes

Run Specific Benchmarks

# API Gateway benchmarks
cargo bench -p api_gateway auth_overhead
cargo bench -p api_gateway proxy_latency
cargo bench -p api_gateway rate_limiting_perf

# Trading Engine benchmarks
cargo bench -p trading_engine order_matching_latency
cargo bench -p trading_engine e2e_latency

# ML benchmarks
cargo bench -p ml inference_bench
cargo bench -p ml wave_d_features_bench
cargo bench -p ml tft_int8_inference_bench

# Backtesting benchmarks
cargo bench -p backtesting_service dbn_loading_benchmark

Benchmark Output

Benchmarks generate HTML reports in target/criterion/:

# Open benchmark report in browser
firefox target/criterion/report/index.html

# Or use criterion's CLI
cargo criterion --message-format=json > bench_results.json

6. Coverage Reports

Generate Coverage Report

# HTML report (recommended)
cargo llvm-cov --workspace --html --output-dir coverage_report

# Open in browser
firefox coverage_report/index.html

# JSON report (for CI/CD)
cargo llvm-cov --workspace --json --output-path coverage.json

# Summary to stdout
cargo llvm-cov --workspace --summary-only

Coverage for Specific Crate

cargo llvm-cov -p trading_engine --html --output-dir coverage_trading_engine
cargo llvm-cov -p ml --html --output-dir coverage_ml

Coverage with Test Filtering

# Coverage for integration tests only
cargo llvm-cov --workspace --html --test "*" --output-dir coverage_integration

7. Test Filtering & Selection

By Test Name Pattern

cargo test kelly                    # All tests with "kelly" in the name
cargo test regime                   # All Wave D regime tests
cargo test parquet                  # All Parquet-related tests
cargo test grpc                     # All gRPC tests

By Crate

cargo test -p ml                    # All tests in ml crate
cargo test -p api_gateway           # All tests in api_gateway crate
cargo test -p common                # All tests in common crate

By Test Type

cargo test --lib                    # Unit tests only
cargo test --test "*"               # Integration tests only
cargo test --doc                    # Documentation tests only
cargo test --bins                   # Binary tests only
cargo test --benches                # Benchmark tests (without running benchmarks)

Exclude Ignored Tests

cargo test                          # Runs non-ignored tests
cargo test -- --ignored             # Runs only ignored tests
cargo test -- --include-ignored     # Runs all tests (including ignored)

8. Test Parallelization

Default Parallelization

# Auto-detect CPU cores (default)
cargo test

# Equivalent to:
cargo test -- --test-threads={num_cpus}

Serial Execution

# Run all tests sequentially
cargo test -- --test-threads=1

When to use:

  • Database integration tests (avoid conflicts)
  • Tests with global state
  • Tests marked with #[serial]

Custom Thread Count

cargo test -- --test-threads=4     # Use 4 threads
cargo test -- --test-threads=8     # Use 8 threads

9. Test Output Control

Show Output

# Show output only for failing tests (default)
cargo test

# Show output for all tests
cargo test -- --show-output

# Show println! output (deprecated, use --show-output)
cargo test -- --nocapture

Quiet Mode

cargo test --quiet                  # Minimal output
cargo test -q                       # Shorthand

JSON Output (for CI/CD)

cargo test -- --format=json > test_results.json
cargo test -- --format=json --report-time > test_results_with_time.json

10. Test Debugging

Run Single Test with Backtrace

RUST_BACKTRACE=1 cargo test test_order_matching_multiple_fills
RUST_BACKTRACE=full cargo test test_order_matching_multiple_fills

Run Test with Logs

RUST_LOG=debug cargo test test_grpc_streaming
RUST_LOG=trading_engine=trace cargo test test_order_matching

Run Test with GDB/LLDB

# Build test binary
cargo test --no-run test_order_matching_multiple_fills

# Find test binary path
ls -la target/debug/deps/trading_engine-*

# Run with debugger
gdb target/debug/deps/trading_engine-abc123 --args test_order_matching_multiple_fills
lldb target/debug/deps/trading_engine-abc123 -- test_order_matching_multiple_fills

11. CI/CD Integration

GitHub Actions Example

name: Test Suite

on: [push, pull_request]

jobs:
  unit-tests:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v3
      - uses: dtolnay/rust-toolchain@stable
      - uses: Swatinem/rust-cache@v2
      - run: cargo test --workspace --lib

  integration-tests:
    runs-on: ubuntu-latest
    services:
      postgres:
        image: timescaledb/timescaledb:latest
        env:
          POSTGRES_PASSWORD: foxhunt_dev_password
          POSTGRES_USER: foxhunt
          POSTGRES_DB: foxhunt
        options: >-
          --health-cmd pg_isready
          --health-interval 10s
          --health-timeout 5s
          --health-retries 5
      redis:
        image: redis:7-alpine
        options: >-
          --health-cmd "redis-cli ping"
          --health-interval 10s
          --health-timeout 5s
          --health-retries 5

    steps:
      - uses: actions/checkout@v3
      - uses: dtolnay/rust-toolchain@stable
      - uses: Swatinem/rust-cache@v2
      - run: cargo sqlx migrate run
      - run: cargo test --workspace --test "*"

  coverage:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v3
      - uses: dtolnay/rust-toolchain@stable
      - uses: taiki-e/install-action@cargo-llvm-cov
      - run: cargo llvm-cov --workspace --lcov --output-path lcov.info
      - uses: codecov/codecov-action@v3
        with:
          files: lcov.info

12. Common Test Execution Patterns

Quick Local Development Loop

# Fast feedback cycle (30 seconds)
cargo test --lib test_my_new_feature

# When ready for integration
cargo test integration::test_my_new_feature_e2e

Pre-Commit Check

# Run before committing (1-2 minutes)
cargo test --lib && cargo clippy --workspace && cargo fmt -- --check

Pre-PR Check

# Full test suite (5-10 minutes)
docker-compose up -d
cargo sqlx migrate run
cargo test --workspace
cargo test -p tli e2e::
cargo clippy --workspace -- -D warnings
cargo fmt -- --check

Nightly CI

# Comprehensive test suite (30-60 minutes)
cargo test --workspace --release
cargo test --release --features stress-test
cargo bench --workspace
cargo llvm-cov --workspace --html --output-dir coverage_report

13. Troubleshooting

Tests Fail Due to Missing Docker Services

Error: Connection refused (os error 111)

Fix:

docker-compose up -d
docker-compose ps  # Verify all services are healthy

Tests Fail Due to Database Migration

Error: relation "orders" does not exist

Fix:

cargo sqlx migrate run
# Or reset database
docker-compose down -v
docker-compose up -d
cargo sqlx migrate run

Tests Hang Indefinitely

Cause: Deadlock or missing timeout

Fix:

# Add timeout to test
#[tokio::test]
#[timeout(Duration::from_secs(30))]
async fn test_grpc_streaming() { ... }

# Or run with RUST_TEST_TIMEOUT
RUST_TEST_TIMEOUT=30 cargo test

Tests Fail Intermittently (Flaky Tests)

Diagnosis:

# Run test 100 times to reproduce
for i in {1..100}; do
    cargo test test_flaky_test || break
done

Fix:

  • Add #[serial] if test has shared state
  • Add retry logic with tokio-retry
  • Use proper synchronization primitives (not sleeps)

Out of Memory During Tests

Error: SIGKILL or OOM

Fix:

# Run tests sequentially
cargo test -- --test-threads=1

# Or reduce parallel tests
cargo test -- --test-threads=2

14. Performance Optimization Tips

Faster Test Builds

# Use cargo-nextest (parallel test runner)
cargo install cargo-nextest
cargo nextest run --workspace          # ~2x faster than cargo test

# Use sccache (compilation cache)
cargo install sccache
export RUSTC_WRAPPER=sccache

Faster Integration Tests

# Use in-memory SQLite for database tests
#[sqlx::test]
async fn test_database_insert() {
    let pool = SqlitePool::connect(":memory:").await.unwrap();
    // ...
}

Faster Docker Startup

# Pre-pull images
docker-compose pull

# Use docker-compose profiles for selective services
docker-compose --profile testing up -d

15. Test Data Management

Download Test Data

# Download real market data (gitignored)
./scripts/download_test_data.sh

# Expected files:
# test_data/ES.FUT.dbn         (~500MB)
# test_data/NQ.FUT.dbn         (~450MB)
# test_data/ES_FUT_180d.parquet (~1.2GB)
# test_data/NQ_FUT_90d.parquet  (~600MB)

Generate Synthetic Test Data

use common::testing::generate_mock_market_data;

#[test]
fn test_with_synthetic_data() {
    let market_data = generate_mock_market_data(
        symbol: "ES.FUT",
        bars: 1000,
        start_price: 4500.0,
    );
    // Use market_data in test
}

Cleanup Test Data

# Remove generated test artifacts
cargo clean
rm -rf test_data/*.db
rm -rf test_data/*.cache

16. References

  • Test Strategy: WAVE4_TEST_STRATEGY.md
  • Coverage Report: WAVE4_COVERAGE_REPORT.md
  • Debugging Guide: WAVE4_DEBUGGING_GUIDE.md
  • CI/CD Integration: WAVE4_CICD_INTEGRATION.md
  • Test Data Management: WAVE4_TEST_DATA.md

Last Updated: 2025-10-22 Next Review: 2026-01-22 Owner: Engineering Team / QA Lead