Files
foxhunt/monitoring/alertmanager/alertmanager.yml
jgrusewski 35feadf55e 🚀 Wave 160 Phase 6: CUDA Mandatory + TDD Testing + TFT Complete (21 Agents)
## Major Achievements

### 1. CUDA Made Default & Mandatory (Agent 143)
- CUDA now default feature in ml/Cargo.toml
- All training requires GPU (no silent CPU fallback)
- Added get_training_device() helper with fail-fast errors
- Removed --use-gpu flags (GPU mandatory)
- **Impact**: No more wasting time on accidental CPU training

### 2. TFT Training COMPLETE (Agent 144)
-  Training completed successfully in 7.6 minutes
-  Early stopping at epoch 100/200 (best val loss: 0.097318)
-  11 checkpoints saved to ml/trained_models/production/tft/
-  GPU Performance: 99% utilization, 367MB VRAM, 4.4s/epoch
-  10x speedup vs CPU (4.4s vs 43-55s per epoch)
- **Status**: PRODUCTION READY

### 3. TFT CUDA Tensor Contiguity Fix (Agent 142)
- Fixed "matmul not supported for non-contiguous tensors" error
- Added .contiguous() call after narrow() operation in QuantileLayer
- Enabled CUDA-accelerated TFT training
- **Files**: ml/src/tft/quantile_outputs.rs

### 4. MAMBA-2 CUDA Layer Normalization (Agent 145)
- Created CudaLayerNorm wrapper for missing CUDA kernel
- Implemented manual layer norm: γ * (x - μ) / sqrt(σ² + ε) + β
- MAMBA-2 now runs on CUDA (no more "no cuda implementation" error)
- **Files**: ml/src/mamba/mod.rs

### 5. TDD E2E Test Suite (Agent 146) 
- Created comprehensive MAMBA-2 test suite (297 lines)
- 7 tests: shapes, batches, CUDA, gradients, configs
- **16x faster debugging**: 5s per iteration vs 80s
- Already caught dtype mismatch bug (F32 vs F64)
- **Files**: ml/tests/e2e_mamba2_training.rs

## Agent Summary (Agents 126-146)

### Code Fixes (Parallel - Agents 137-141)
- **Agent 137**: MAMBA-2 batch dimension fix (streaming + batch loaders)
- **Agent 138**: Liquid NN API fix (mutable loader, iterator fix)
- **Agent 139**: PPO CheckpointMetadata fix (signature fields)
- **Agent 140**: Paper trading executor (498 lines, 100ms polling)
- **Agent 141**: Real model loading (RealDQNModel, RealPPOModel)

### Infrastructure (Agents 143-146)
- **Agent 143**: CUDA mandatory (Cargo.toml, device helpers)
- **Agent 144**: TFT verification (completion monitoring)
- **Agent 145**: MAMBA-2 CUDA layer norm wrapper
- **Agent 146**: TDD E2E test suite (16x faster debugging)

## Files Modified

### Core ML Infrastructure
- ml/Cargo.toml: Added default = ["minimal-inference", "cuda"]
- ml/src/lib.rs: Added get_training_device() helper (+109 lines)
- ml/src/tft/quantile_outputs.rs: Fixed tensor contiguity
- ml/src/mamba/mod.rs: Added CudaLayerNorm wrapper (+41 lines)

### Training Scripts
- ml/examples/train_tft_dbn.rs: Removed --use-gpu flag
- ml/examples/train_ppo.rs: Removed --use-gpu flag
- ml/examples/train_mamba2_dbn.rs: Forced CUDA-only mode
- ml/examples/train_liquid_dbn.rs: Fixed API usage

### Data Loaders
- ml/src/data_loaders/dbn_sequence_loader.rs: Fixed batch dimensions
- ml/src/data_loaders/streaming_dbn_loader.rs: Fixed batch dimensions

### Trading Service
- services/trading_service/src/paper_trading_executor.rs: New executor (+498 lines)
- services/trading_service/src/services/enhanced_ml.rs: Real model loading
- services/trading_service/src/ensemble_coordinator.rs: Integration

### Tests
- ml/tests/e2e_mamba2_training.rs: New TDD test suite (+297 lines)

### Trainers
- ml/src/trainers/tft.rs: Fixed CheckpointMetadata signature fields

## Performance Metrics

### TFT Training
- Duration: 7.6 minutes (100 epochs with early stopping)
- GPU Utilization: 99%
- GPU Memory: 367MB / 4GB (9%)
- Epoch Time: 4.4 seconds (vs 43-55s on CPU)
- Speedup: 10x vs CPU
- Status:  PRODUCTION READY

### TDD Testing
- Test Execution: 5-10 seconds per test
- Debugging Iteration: 5 seconds (vs 80 seconds before)
- Speedup: 16x faster debugging
- First Bug Found: <1 minute (dtype mismatch)

## Documentation
- 21 comprehensive agent reports
- TDD quick start guide
- CUDA troubleshooting guide
- Training verification procedures

## Next Steps
1. Fix MAMBA-2 dtype mismatch (F32→F64) - 2 minutes
2. Run MAMBA-2 tests until passing - 5-10 minutes
3. Launch full MAMBA-2 training - 200 epochs
4. Launch Liquid NN training

## System Status
- TFT:  COMPLETE (production ready)
- MAMBA-2: 🧪 IN TESTING (TDD suite ready)
- CUDA:  DEFAULT (mandatory for training)
- Tests:  16x faster debugging

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
2025-10-14 23:13:34 +02:00

258 lines
8.1 KiB
YAML
Raw Blame History

This file contains invisible Unicode characters
This file contains invisible Unicode characters that are indistinguishable to humans but may be processed differently by a computer. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# AlertManager Configuration for Foxhunt
#
# Routes alerts to appropriate notification channels
global:
resolve_timeout: 5m
slack_api_url: 'https://hooks.slack.com/services/YOUR/SLACK/WEBHOOK'
pagerduty_url: 'https://events.pagerduty.com/v2/enqueue'
# Alert routing tree
route:
receiver: 'default'
group_by: ['alertname', 'cluster', 'service']
group_wait: 10s
group_interval: 5m
repeat_interval: 4h
# Route alerts based on severity
routes:
# Critical ensemble alerts - IMMEDIATE PagerDuty + Slack
- match:
severity: critical
component: ensemble
receiver: 'ensemble-critical'
group_wait: 0s
repeat_interval: 30m
continue: false
# Warning ensemble alerts - Slack only
- match:
severity: warning
component: ensemble
receiver: 'ensemble-warnings'
group_wait: 15s
repeat_interval: 2h
continue: false
# Info ensemble alerts (A/B test results) - Slack only
- match:
severity: info
component: ensemble
receiver: 'ensemble-info'
group_wait: 5m
repeat_interval: 24h
continue: false
# Critical alerts (non-ensemble) - PagerDuty + Slack
- match:
severity: critical
receiver: 'critical-alerts'
group_wait: 0s
repeat_interval: 1h
# Warning alerts - less urgent
- match:
severity: warning
receiver: 'warning-alerts'
group_wait: 30s
repeat_interval: 4h
# Auth-specific alerts
- match:
component: auth
receiver: 'auth-alerts'
group_by: ['alertname']
# Backend proxy alerts
- match:
component: proxy
receiver: 'backend-alerts'
# Configuration alerts
- match:
component: config
receiver: 'config-alerts'
# Alert receivers (notification channels)
receivers:
# Default receiver (logs only)
- name: 'default'
webhook_configs:
- url: 'http://localhost:9090/api/v1/alerts'
# ========== ENSEMBLE ALERTS ==========
# Ensemble critical alerts - PagerDuty + Slack
- name: 'ensemble-critical'
slack_configs:
- channel: '#foxhunt-ensemble-critical'
title: '🚨 ENSEMBLE CRITICAL: {{ .GroupLabels.alertname }}'
text: |
*Alert:* {{ .GroupLabels.alertname }}
*Type:* {{ .CommonLabels.alert_type }}
*Symbol:* {{ .CommonLabels.symbol }}
{{ range .Alerts }}
*Summary:* {{ .Annotations.summary }}
*Description:* {{ .Annotations.description }}
*Impact:* {{ .Annotations.impact }}
*Action Required:*
{{ .Annotations.action }}
*Runbook:* {{ .Annotations.runbook_url }}
{{ end }}
send_resolved: true
color: '{{ if eq .Status "firing" }}danger{{ else }}good{{ end }}'
pagerduty_configs:
- routing_key: 'YOUR_PAGERDUTY_ENSEMBLE_INTEGRATION_KEY'
severity: 'critical'
description: '{{ .GroupLabels.alertname }}: {{ .CommonAnnotations.summary }}'
details:
alert_type: '{{ .CommonLabels.alert_type }}'
symbol: '{{ .CommonLabels.symbol }}'
impact: '{{ .CommonAnnotations.impact }}'
action: '{{ .CommonAnnotations.action }}'
runbook_url: '{{ .CommonAnnotations.runbook_url }}'
client: 'Foxhunt Ensemble Monitoring'
client_url: 'http://localhost:3000/d/ensemble-ml-production'
# Ensemble warning alerts - Slack only
- name: 'ensemble-warnings'
slack_configs:
- channel: '#foxhunt-ensemble-warnings'
title: '⚠️ ENSEMBLE WARNING: {{ .GroupLabels.alertname }}'
text: |
*Alert:* {{ .GroupLabels.alertname }}
*Type:* {{ .CommonLabels.alert_type }}
*Symbol:* {{ .CommonLabels.symbol }}
{{ range .Alerts }}
*Summary:* {{ .Annotations.summary }}
*Description:* {{ .Annotations.description }}
{{ if .Annotations.action }}*Action:* {{ .Annotations.action }}{{ end }}
{{ end }}
send_resolved: true
color: 'warning'
# Ensemble info alerts (A/B tests, model updates)
- name: 'ensemble-info'
slack_configs:
- channel: '#foxhunt-ensemble-info'
title: ' ENSEMBLE INFO: {{ .GroupLabels.alertname }}'
text: |
*Alert:* {{ .GroupLabels.alertname }}
{{ range .Alerts }}
*Summary:* {{ .Annotations.summary }}
*Description:* {{ .Annotations.description }}
{{ end }}
send_resolved: true
color: 'good'
# ========== GENERAL ALERTS ==========
# Critical alerts (non-ensemble) - PagerDuty + Slack
- name: 'critical-alerts'
slack_configs:
- channel: '#foxhunt-critical'
title: '🚨 CRITICAL: {{ .GroupLabels.alertname }}'
text: '{{ range .Alerts }}{{ .Annotations.summary }}\n{{ .Annotations.description }}{{ end }}'
send_resolved: true
# PagerDuty for on-call rotation
pagerduty_configs:
- routing_key: 'YOUR_PAGERDUTY_SERVICE_KEY'
severity: 'critical'
description: '{{ .GroupLabels.alertname }}'
# Warning alerts - Slack only
- name: 'warning-alerts'
slack_configs:
- channel: '#foxhunt-warnings'
title: '⚠️ Warning: {{ .GroupLabels.alertname }}'
text: '{{ range .Alerts }}{{ .Annotations.summary }}{{ end }}'
send_resolved: true
# Auth-specific alerts
- name: 'auth-alerts'
slack_configs:
- channel: '#foxhunt-auth'
title: '🔐 Auth Alert: {{ .GroupLabels.alertname }}'
text: '{{ range .Alerts }}{{ .Annotations.description }}{{ end }}'
# Backend alerts
- name: 'backend-alerts'
slack_configs:
- channel: '#foxhunt-backend'
title: '🔌 Backend Alert: {{ .GroupLabels.alertname }}'
text: '{{ range .Alerts }}{{ .Annotations.description }}{{ end }}'
# Config alerts
- name: 'config-alerts'
slack_configs:
- channel: '#foxhunt-config'
title: '⚙️ Config Alert: {{ .GroupLabels.alertname }}'
text: '{{ range .Alerts }}{{ .Annotations.description }}{{ end }}'
# Inhibition rules (suppress redundant alerts)
inhibit_rules:
# ========== ENSEMBLE INHIBITION RULES ==========
# If cascade failure detected, suppress individual model failures
- source_match:
alertname: 'EnsembleCascadeFailureDetected'
target_match:
alertname: 'EnsembleModelFailureDetected'
equal: ['cluster']
# If Sharpe ratio drop critical, suppress warning
- source_match:
alertname: 'EnsembleSharpeRatioDropCritical'
target_match:
alertname: 'EnsembleSharpeRatioDropWarning'
equal: ['symbol']
# If high memory, suppress latency alerts (latency caused by swapping)
- source_match:
alertname: 'EnsembleServiceMemoryHigh'
target_match:
alertname: 'EnsembleAggregationLatencyP99High'
equal: ['instance']
# If checkpoint rollback rate high, suppress individual swap failures
- source_match:
alertname: 'EnsembleCheckpointRollbackRateHigh'
target_match:
alertname: 'EnsembleCheckpointSwapFailed'
equal: ['cluster']
# If high disagreement critical, suppress warning
- source_match:
alertname: 'EnsembleHighDisagreementCritical'
target_match:
alertname: 'EnsembleHighDisagreementWarning'
equal: ['symbol']
# If low confidence + high disagreement, suppress both individual alerts
- source_match:
alertname: 'EnsembleLowConfidenceHighDisagreement'
target_match_re:
alertname: 'EnsembleHighDisagreementCritical|EnsembleHighDisagreementWarning'
equal: ['symbol']
# ========== GENERAL INHIBITION RULES ==========
# If circuit breaker is open, suppress high latency alerts
- source_match:
alertname: 'CircuitBreakerOpen'
target_match:
alertname: 'HighBackendLatency'
equal: ['service']
# If backend is unhealthy, suppress other backend alerts
- source_match:
alertname: 'BackendServiceUnhealthy'
target_match_re:
alertname: 'HighBackendLatency|CircuitBreakerOpen'
equal: ['service']
# If NOTIFY listener is down, suppress config alerts
- source_match:
alertname: 'NotifyListenerDisconnected'
target_match_re:
alertname: 'HighConfigReloadLatency|ConfigValidationFailures'