# Performance Regression Detection - Quick Start **Goal**: Track performance and automatically fail CI on >10% regression ## Quick Commands ### 1. Record Baseline (First Time) ```bash # DQN baseline cargo run --release -p ml --example quick_performance_benchmark -- \ --output ml/benchmark_results/dqn_baseline.json \ --git-commit $(git rev-parse HEAD) \ --model DQN # PPO baseline cargo run --release -p ml --example quick_performance_benchmark -- \ --output ml/benchmark_results/ppo_baseline.json \ --git-commit $(git rev-parse HEAD) \ --model PPO # MAMBA-2 baseline cargo run --release -p ml --example quick_performance_benchmark -- \ --output ml/benchmark_results/mamba2_baseline.json \ --git-commit $(git rev-parse HEAD) \ --model MAMBA-2 # TFT baseline cargo run --release -p ml --example quick_performance_benchmark -- \ --output ml/benchmark_results/tft_baseline.json \ --git-commit $(git rev-parse HEAD) \ --model TFT ``` ### 2. Check for Regression (CI) ```bash # Run current benchmark cargo run --release -p ml --example quick_performance_benchmark -- \ --output ml/benchmark_results/current.json \ --git-commit $(git rev-parse HEAD) \ --model DQN # Check against baseline cargo run --release -p ml --example check_performance_regression -- \ --baseline ml/benchmark_results/dqn_baseline.json \ --current ml/benchmark_results/current.json \ --output regression_report.md # Exit code 0 = Pass # Exit code 1 = Fail (regression detected) ``` ### 3. Run Tests ```bash cargo test -p ml --test performance_regression_tests ``` ### 4. Import Grafana Dashboard ```bash # Copy dashboard JSON cp ml/grafana/performance_tracking_dashboard.json /var/lib/grafana/dashboards/ # Or import via UI: # Grafana → Dashboards → Import → Upload JSON # File: ml/grafana/performance_tracking_dashboard.json ``` ## Metrics Tracked | Metric | Target | Model-Specific | |--------|--------|----------------| | DBN Load Time | <10ms | No | | Feature Extraction | - | No | | Training Step | - | Yes (100ms-500ms) | | Inference Latency | <50μs | Yes (40μs-55μs) | | Throughput | - | Yes | | Memory Usage | - | Yes (150MB-2GB) | ## Regression Threshold **10%** = Any metric that degrades by >10% fails the build Examples: - ✅ PASS: 0.70ms → 0.75ms (7.1% slower) - ❌ FAIL: 0.70ms → 0.81ms (15.7% slower) ## CI Integration Add to PR workflow: ```yaml # .github/workflows/ci.yml - name: Performance Check run: | # Record current metrics cargo run --release -p ml --example quick_performance_benchmark -- \ --output current.json \ --git-commit ${{ github.sha }} \ --model DQN # Check regression cargo run --release -p ml --example check_performance_regression -- \ --baseline baseline.json \ --current current.json \ --output report.md # Exit code 1 fails the build ``` ## Grafana Setup 1. **Import Dashboard**: - Go to http://localhost:3000 - Dashboards → Import - Upload `ml/grafana/performance_tracking_dashboard.json` 2. **Configure Prometheus**: ```yaml # prometheus.yml scrape_configs: - job_name: 'ml-performance' static_configs: - targets: ['localhost:9094'] ``` 3. **View Metrics**: - DBN Load Time - Inference Latency (by model) - Training Step Time - Memory Usage - Regression Count ## Troubleshooting ### Test Failures ```bash # Run specific test cargo test -p ml test_detect_regression_above_threshold -- --nocapture # Verbose logging RUST_LOG=debug cargo test -p ml --test performance_regression_tests ``` ### Baseline Missing ```bash # Create baseline if it doesn't exist cargo run --release -p ml --example quick_performance_benchmark -- \ --output ml/benchmark_results/dqn_baseline.json \ --git-commit $(git rev-parse HEAD) \ --model DQN ``` ### False Positives If you see false regressions: 1. Check for system load during benchmark 2. Verify consistent hardware (CPU/GPU) 3. Run multiple times and average 4. Adjust threshold if needed: ```bash cargo run --release -p ml --example check_performance_regression -- \ --baseline baseline.json \ --current current.json \ --output report.md \ --threshold 15.0 # Use 15% instead of 10% ``` ## Model-Specific Baselines Each model has different performance characteristics: ``` DQN: 150MB memory, 100ms training, 45μs inference PPO: 200MB memory, 150ms training, 50μs inference MAMBA-2: 400MB memory, 200ms training, 40μs inference TFT: 2GB memory, 500ms training, 55μs inference ``` Always use model-specific baselines: ```bash --baseline ml/benchmark_results/dqn_baseline.json # For DQN --baseline ml/benchmark_results/ppo_baseline.json # For PPO ``` ## File Locations ``` ml/ ├── src/benchmark/performance_tracker.rs # Core implementation ├── tests/performance_regression_tests.rs # Tests (12/12 passing) ├── examples/ │ ├── quick_performance_benchmark.rs # Record metrics │ └── check_performance_regression.rs # Detect regressions ├── benchmark_results/ │ ├── dqn_baseline.json # DQN baseline │ ├── ppo_baseline.json # PPO baseline │ ├── mamba2_baseline.json # MAMBA-2 baseline │ └── tft_baseline.json # TFT baseline └── grafana/performance_tracking_dashboard.json # Grafana dashboard ``` ## Support - **Documentation**: `ml/PERFORMANCE_TRACKING.md` - **Tests**: `cargo test -p ml --test performance_regression_tests` - **Examples**: `ml/examples/quick_performance_benchmark.rs` - **CI Workflow**: `.github/workflows/performance.yml`