368 lines
9.6 KiB
Markdown
368 lines
9.6 KiB
Markdown
# Foxhunt RunPod Workflow Module
|
|
|
|
Production-ready Python module for deploying and monitoring ML training on RunPod GPU infrastructure.
|
|
|
|
## Features
|
|
|
|
- **Pod Deployment**: REST API with datacenter filtering and automatic GPU selection
|
|
- **S3 Log Monitoring**: Byte-range polling (NO SSH required)
|
|
- **Auto-Termination**: Detect training completion and terminate pods automatically
|
|
- **Rich Output**: Progress tracking with beautiful terminal formatting
|
|
- **Retry Logic**: Exponential backoff with configurable retries
|
|
- **Type Safety**: Full type hints and pydantic validation
|
|
|
|
## Installation
|
|
|
|
```bash
|
|
cd ml/python/foxhunt_runpod
|
|
pip install -r requirements.txt
|
|
```
|
|
|
|
## Configuration
|
|
|
|
Create `.env.runpod` in project root:
|
|
|
|
```env
|
|
# RunPod API
|
|
RUNPOD_API_KEY=rpa_...
|
|
|
|
# S3 Credentials
|
|
RUNPOD_S3_ACCESS_KEY=user_...
|
|
RUNPOD_S3_SECRET=rps_...
|
|
RUNPOD_S3_ENDPOINT=https://s3api-eur-is-1.runpod.io
|
|
RUNPOD_S3_REGION=eur-is-1
|
|
|
|
# Network Volume
|
|
RUNPOD_VOLUME_ID=se3zdnb5o4
|
|
RUNPOD_VOLUME_MOUNT=/runpod-volume
|
|
|
|
# Optional
|
|
RUNPOD_CONTAINER_REGISTRY_AUTH_ID=cmh3...
|
|
```
|
|
|
|
## Quick Start
|
|
|
|
```python
|
|
from foxhunt_runpod import RunPodClient, PodMonitor, S3Client
|
|
|
|
# 1. Deploy pod (auto-selects cheapest GPU)
|
|
client = RunPodClient()
|
|
pod = client.deploy_pod(
|
|
gpu_type="RTX A4000", # Optional: preferred GPU
|
|
command="--parquet-file /runpod-volume/test_data/ES_FUT_180d.parquet --epochs 50"
|
|
)
|
|
|
|
# 2. Monitor training (S3 log tailing)
|
|
monitor = PodMonitor(pod['id'])
|
|
monitor.wait_until_running(timeout=300) # Wait for pod to start
|
|
monitor.stream_s3_logs(follow=True) # Stream logs until completion
|
|
|
|
# 3. Auto-terminate on completion
|
|
monitor.auto_terminate()
|
|
```
|
|
|
|
## API Reference
|
|
|
|
### RunPodClient
|
|
|
|
```python
|
|
client = RunPodClient()
|
|
|
|
# Query available GPUs
|
|
gpus = client.get_available_gpus(min_vram_gb=16, cloud_type="SECURE")
|
|
|
|
# Deploy pod
|
|
pod = client.deploy_pod(
|
|
gpu_type="RTX A4000",
|
|
image="jgrusewski/foxhunt:latest",
|
|
command="--epochs 50",
|
|
container_disk_gb=50,
|
|
env_vars={"KEY": "value"},
|
|
dry_run=False
|
|
)
|
|
|
|
# Get pod status
|
|
status = client.get_pod_status(pod_id)
|
|
|
|
# Terminate pod
|
|
client.terminate_pod(pod_id)
|
|
|
|
# List all pods
|
|
pods = client.list_pods()
|
|
client.display_pods_table(pods)
|
|
```
|
|
|
|
### PodMonitor
|
|
|
|
```python
|
|
monitor = PodMonitor(pod_id)
|
|
|
|
# Wait for pod to start
|
|
monitor.wait_until_running(timeout=300)
|
|
|
|
# Stream logs from S3
|
|
monitor.stream_s3_logs(
|
|
follow=True, # Continue streaming
|
|
poll_interval=5, # Check every 5s
|
|
max_lines=1000 # Limit output
|
|
)
|
|
|
|
# Check completion
|
|
completed, error = monitor.check_completion()
|
|
|
|
# Auto-terminate on completion
|
|
monitor.auto_terminate(wait_for_completion=True)
|
|
|
|
# Display pod info
|
|
monitor.display_pod_info()
|
|
```
|
|
|
|
### S3Client
|
|
|
|
```python
|
|
s3 = S3Client()
|
|
|
|
# Upload binary with progress
|
|
s3.upload_binary(
|
|
local_file=Path("target/release/examples/train_tft_parquet"),
|
|
s3_key="binaries/train_tft_parquet",
|
|
force=False
|
|
)
|
|
|
|
# Tail log file (byte-range request)
|
|
content, new_position = s3.tail_log_file(
|
|
s3_key="logs/training.log",
|
|
start_byte=0,
|
|
max_bytes=1024*1024 # 1MB chunks
|
|
)
|
|
|
|
# Download results
|
|
files = s3.download_results(
|
|
s3_prefix="models/tft/",
|
|
local_dir=Path("./models")
|
|
)
|
|
|
|
# List inventory
|
|
binaries = s3.list_binaries()
|
|
models = s3.list_models()
|
|
```
|
|
|
|
## S3 Log Monitoring Architecture
|
|
|
|
The module monitors training progress via S3 log tailing (NO SSH required):
|
|
|
|
1. **Log Writing**: Training container writes to `/runpod-volume/logs/training.log`
|
|
2. **S3 Sync**: RunPod network volume auto-syncs to S3 bucket
|
|
3. **Byte-Range Polling**: Monitor uses HTTP byte-range requests to tail new content
|
|
4. **Pattern Detection**: Checks for completion/error patterns in logs
|
|
|
|
### Log File Location
|
|
|
|
```
|
|
Pod Container: /runpod-volume/logs/training.log
|
|
↓ (auto-synced)
|
|
S3 Bucket: s3://se3zdnb5o4/logs/training.log
|
|
↓ (byte-range GET)
|
|
PodMonitor: Streams new content every 5s
|
|
```
|
|
|
|
### Completion Detection
|
|
|
|
Automatically detects training completion by matching patterns:
|
|
|
|
**Success Patterns** (configurable):
|
|
- "Training complete"
|
|
- "Model saved to"
|
|
- "✓ Training finished"
|
|
- "SUCCESS:"
|
|
|
|
**Error Patterns**:
|
|
- "CUDA out of memory"
|
|
- "RuntimeError:"
|
|
- "ERROR:"
|
|
- "panic!"
|
|
|
|
## Error Handling
|
|
|
|
All errors inherit from `RunPodError`:
|
|
|
|
```python
|
|
from foxhunt_runpod.errors import (
|
|
RunPodError, # Base exception
|
|
ConfigurationError, # Invalid config
|
|
PodDeploymentError, # Deployment failed
|
|
PodNotFoundError, # Pod doesn't exist
|
|
PodTimeoutError, # Operation timed out
|
|
S3Error, # S3 operation failed
|
|
APIError, # RunPod API error
|
|
NetworkError, # Network failure
|
|
)
|
|
|
|
try:
|
|
pod = client.deploy_pod()
|
|
except PodDeploymentError as e:
|
|
print(f"GPU: {e.gpu_type}, Datacenter: {e.datacenter}")
|
|
except NetworkError as e:
|
|
print(f"Retries: {e.retry_count}")
|
|
```
|
|
|
|
## Configuration Options
|
|
|
|
All settings can be overridden via environment variables or config object:
|
|
|
|
```python
|
|
from foxhunt_runpod.config import RunPodConfig
|
|
|
|
config = RunPodConfig(
|
|
# API settings
|
|
api_timeout=60, # Request timeout (seconds)
|
|
max_retries=3, # Retry attempts
|
|
retry_backoff=2.0, # Exponential backoff multiplier
|
|
|
|
# Monitoring settings
|
|
log_poll_interval=5, # S3 log polling (seconds)
|
|
pod_status_poll_interval=10, # Pod status polling (seconds)
|
|
|
|
# Deployment settings
|
|
datacenters=["EUR-IS-1"], # Preferred datacenters
|
|
cloud_type="SECURE", # SECURE or COMMUNITY
|
|
container_disk_gb=50, # Container disk size
|
|
min_vram_gb=16, # Minimum GPU VRAM
|
|
)
|
|
|
|
client = RunPodClient(config)
|
|
```
|
|
|
|
## Example Scripts
|
|
|
|
### Deploy and Monitor Training
|
|
|
|
```python
|
|
#!/usr/bin/env python3
|
|
"""Deploy TFT training on RunPod with monitoring."""
|
|
|
|
from foxhunt_runpod import RunPodClient, PodMonitor
|
|
|
|
def main():
|
|
# Deploy pod
|
|
client = RunPodClient()
|
|
pod = client.deploy_pod(
|
|
gpu_type="RTX A4000",
|
|
command="--parquet-file /runpod-volume/test_data/ES_FUT_180d.parquet --epochs 50"
|
|
)
|
|
|
|
pod_id = pod['id']
|
|
print(f"Deployed pod: {pod_id}")
|
|
|
|
# Monitor training
|
|
monitor = PodMonitor(pod_id)
|
|
|
|
# Wait for pod to start (5 min timeout)
|
|
monitor.wait_until_running(timeout=300)
|
|
|
|
# Stream logs until completion
|
|
monitor.stream_s3_logs(follow=True)
|
|
|
|
# Auto-terminate
|
|
monitor.auto_terminate()
|
|
|
|
if __name__ == "__main__":
|
|
main()
|
|
```
|
|
|
|
### List and Terminate Pods
|
|
|
|
```python
|
|
#!/usr/bin/env python3
|
|
"""List all pods and terminate old ones."""
|
|
|
|
from foxhunt_runpod import RunPodClient
|
|
from datetime import datetime, timedelta
|
|
|
|
client = RunPodClient()
|
|
|
|
# List all pods
|
|
pods = client.list_pods()
|
|
client.display_pods_table(pods)
|
|
|
|
# Terminate pods older than 2 hours
|
|
cutoff = datetime.now() - timedelta(hours=2)
|
|
|
|
for pod in pods:
|
|
created_at = datetime.fromisoformat(pod.get('createdAt', ''))
|
|
if created_at < cutoff:
|
|
print(f"Terminating old pod: {pod['id']}")
|
|
client.terminate_pod(pod['id'])
|
|
```
|
|
|
|
## Architecture
|
|
|
|
```
|
|
┌─────────────────────────────────────────────────────────────┐
|
|
│ Foxhunt RunPod Module │
|
|
├─────────────────────────────────────────────────────────────┤
|
|
│ │
|
|
│ RunPodClient PodMonitor S3Client │
|
|
│ ├─ get_available_gpus ├─ wait_until_running ├─ upload_binary │
|
|
│ ├─ deploy_pod ├─ stream_s3_logs ├─ tail_log_file │
|
|
│ ├─ get_pod_status ├─ check_completion ├─ download_results │
|
|
│ ├─ terminate_pod └─ auto_terminate └─ list_binaries │
|
|
│ └─ list_pods │
|
|
│ │
|
|
│ RunPodConfig Errors │
|
|
│ ├─ pydantic validation ├─ RunPodError │
|
|
│ └─ .env.runpod loader ├─ PodDeploymentError │
|
|
│ ├─ S3Error │
|
|
│ └─ NetworkError │
|
|
└─────────────────────────────────────────────────────────────┘
|
|
│ │ │
|
|
▼ ▼ ▼
|
|
RunPod REST API RunPod GraphQL RunPod S3 API
|
|
(pod management) (GPU queries) (log tailing)
|
|
```
|
|
|
|
## Testing
|
|
|
|
```bash
|
|
# Deploy test pod (dry run)
|
|
python -c "
|
|
from foxhunt_runpod import RunPodClient
|
|
client = RunPodClient()
|
|
client.deploy_pod(dry_run=True)
|
|
"
|
|
|
|
# Query available GPUs
|
|
python -c "
|
|
from foxhunt_runpod import RunPodClient
|
|
client = RunPodClient()
|
|
gpus = client.get_available_gpus()
|
|
print(f'Found {len(gpus)} GPUs')
|
|
"
|
|
|
|
# Test S3 connection
|
|
python -c "
|
|
from foxhunt_runpod import S3Client
|
|
s3 = S3Client()
|
|
binaries = s3.list_binaries()
|
|
print(f'Binaries: {len(binaries)}')
|
|
"
|
|
```
|
|
|
|
## Production Checklist
|
|
|
|
- [x] Type hints for all functions
|
|
- [x] Pydantic validation for config
|
|
- [x] Retry logic with exponential backoff
|
|
- [x] Comprehensive error handling
|
|
- [x] Rich terminal output
|
|
- [x] S3-based monitoring (NO SSH)
|
|
- [x] Automatic pod termination
|
|
- [x] Configuration via env files
|
|
- [x] Docstrings for all public methods
|
|
- [x] GPU selection by price
|
|
- [x] Datacenter filtering
|
|
|
|
## License
|
|
|
|
Proprietary - Foxhunt HFT Trading System
|