Files
foxhunt/runpod/README.md

368 lines
9.6 KiB
Markdown

# Foxhunt RunPod Workflow Module
Production-ready Python module for deploying and monitoring ML training on RunPod GPU infrastructure.
## Features
- **Pod Deployment**: REST API with datacenter filtering and automatic GPU selection
- **S3 Log Monitoring**: Byte-range polling (NO SSH required)
- **Auto-Termination**: Detect training completion and terminate pods automatically
- **Rich Output**: Progress tracking with beautiful terminal formatting
- **Retry Logic**: Exponential backoff with configurable retries
- **Type Safety**: Full type hints and pydantic validation
## Installation
```bash
cd ml/python/foxhunt_runpod
pip install -r requirements.txt
```
## Configuration
Create `.env.runpod` in project root:
```env
# RunPod API
RUNPOD_API_KEY=rpa_...
# S3 Credentials
RUNPOD_S3_ACCESS_KEY=user_...
RUNPOD_S3_SECRET=rps_...
RUNPOD_S3_ENDPOINT=https://s3api-eur-is-1.runpod.io
RUNPOD_S3_REGION=eur-is-1
# Network Volume
RUNPOD_VOLUME_ID=se3zdnb5o4
RUNPOD_VOLUME_MOUNT=/runpod-volume
# Optional
RUNPOD_CONTAINER_REGISTRY_AUTH_ID=cmh3...
```
## Quick Start
```python
from foxhunt_runpod import RunPodClient, PodMonitor, S3Client
# 1. Deploy pod (auto-selects cheapest GPU)
client = RunPodClient()
pod = client.deploy_pod(
gpu_type="RTX A4000", # Optional: preferred GPU
command="--parquet-file /runpod-volume/test_data/ES_FUT_180d.parquet --epochs 50"
)
# 2. Monitor training (S3 log tailing)
monitor = PodMonitor(pod['id'])
monitor.wait_until_running(timeout=300) # Wait for pod to start
monitor.stream_s3_logs(follow=True) # Stream logs until completion
# 3. Auto-terminate on completion
monitor.auto_terminate()
```
## API Reference
### RunPodClient
```python
client = RunPodClient()
# Query available GPUs
gpus = client.get_available_gpus(min_vram_gb=16, cloud_type="SECURE")
# Deploy pod
pod = client.deploy_pod(
gpu_type="RTX A4000",
image="jgrusewski/foxhunt:latest",
command="--epochs 50",
container_disk_gb=50,
env_vars={"KEY": "value"},
dry_run=False
)
# Get pod status
status = client.get_pod_status(pod_id)
# Terminate pod
client.terminate_pod(pod_id)
# List all pods
pods = client.list_pods()
client.display_pods_table(pods)
```
### PodMonitor
```python
monitor = PodMonitor(pod_id)
# Wait for pod to start
monitor.wait_until_running(timeout=300)
# Stream logs from S3
monitor.stream_s3_logs(
follow=True, # Continue streaming
poll_interval=5, # Check every 5s
max_lines=1000 # Limit output
)
# Check completion
completed, error = monitor.check_completion()
# Auto-terminate on completion
monitor.auto_terminate(wait_for_completion=True)
# Display pod info
monitor.display_pod_info()
```
### S3Client
```python
s3 = S3Client()
# Upload binary with progress
s3.upload_binary(
local_file=Path("target/release/examples/train_tft_parquet"),
s3_key="binaries/train_tft_parquet",
force=False
)
# Tail log file (byte-range request)
content, new_position = s3.tail_log_file(
s3_key="logs/training.log",
start_byte=0,
max_bytes=1024*1024 # 1MB chunks
)
# Download results
files = s3.download_results(
s3_prefix="models/tft/",
local_dir=Path("./models")
)
# List inventory
binaries = s3.list_binaries()
models = s3.list_models()
```
## S3 Log Monitoring Architecture
The module monitors training progress via S3 log tailing (NO SSH required):
1. **Log Writing**: Training container writes to `/runpod-volume/logs/training.log`
2. **S3 Sync**: RunPod network volume auto-syncs to S3 bucket
3. **Byte-Range Polling**: Monitor uses HTTP byte-range requests to tail new content
4. **Pattern Detection**: Checks for completion/error patterns in logs
### Log File Location
```
Pod Container: /runpod-volume/logs/training.log
↓ (auto-synced)
S3 Bucket: s3://se3zdnb5o4/logs/training.log
↓ (byte-range GET)
PodMonitor: Streams new content every 5s
```
### Completion Detection
Automatically detects training completion by matching patterns:
**Success Patterns** (configurable):
- "Training complete"
- "Model saved to"
- "✓ Training finished"
- "SUCCESS:"
**Error Patterns**:
- "CUDA out of memory"
- "RuntimeError:"
- "ERROR:"
- "panic!"
## Error Handling
All errors inherit from `RunPodError`:
```python
from foxhunt_runpod.errors import (
RunPodError, # Base exception
ConfigurationError, # Invalid config
PodDeploymentError, # Deployment failed
PodNotFoundError, # Pod doesn't exist
PodTimeoutError, # Operation timed out
S3Error, # S3 operation failed
APIError, # RunPod API error
NetworkError, # Network failure
)
try:
pod = client.deploy_pod()
except PodDeploymentError as e:
print(f"GPU: {e.gpu_type}, Datacenter: {e.datacenter}")
except NetworkError as e:
print(f"Retries: {e.retry_count}")
```
## Configuration Options
All settings can be overridden via environment variables or config object:
```python
from foxhunt_runpod.config import RunPodConfig
config = RunPodConfig(
# API settings
api_timeout=60, # Request timeout (seconds)
max_retries=3, # Retry attempts
retry_backoff=2.0, # Exponential backoff multiplier
# Monitoring settings
log_poll_interval=5, # S3 log polling (seconds)
pod_status_poll_interval=10, # Pod status polling (seconds)
# Deployment settings
datacenters=["EUR-IS-1"], # Preferred datacenters
cloud_type="SECURE", # SECURE or COMMUNITY
container_disk_gb=50, # Container disk size
min_vram_gb=16, # Minimum GPU VRAM
)
client = RunPodClient(config)
```
## Example Scripts
### Deploy and Monitor Training
```python
#!/usr/bin/env python3
"""Deploy TFT training on RunPod with monitoring."""
from foxhunt_runpod import RunPodClient, PodMonitor
def main():
# Deploy pod
client = RunPodClient()
pod = client.deploy_pod(
gpu_type="RTX A4000",
command="--parquet-file /runpod-volume/test_data/ES_FUT_180d.parquet --epochs 50"
)
pod_id = pod['id']
print(f"Deployed pod: {pod_id}")
# Monitor training
monitor = PodMonitor(pod_id)
# Wait for pod to start (5 min timeout)
monitor.wait_until_running(timeout=300)
# Stream logs until completion
monitor.stream_s3_logs(follow=True)
# Auto-terminate
monitor.auto_terminate()
if __name__ == "__main__":
main()
```
### List and Terminate Pods
```python
#!/usr/bin/env python3
"""List all pods and terminate old ones."""
from foxhunt_runpod import RunPodClient
from datetime import datetime, timedelta
client = RunPodClient()
# List all pods
pods = client.list_pods()
client.display_pods_table(pods)
# Terminate pods older than 2 hours
cutoff = datetime.now() - timedelta(hours=2)
for pod in pods:
created_at = datetime.fromisoformat(pod.get('createdAt', ''))
if created_at < cutoff:
print(f"Terminating old pod: {pod['id']}")
client.terminate_pod(pod['id'])
```
## Architecture
```
┌─────────────────────────────────────────────────────────────┐
│ Foxhunt RunPod Module │
├─────────────────────────────────────────────────────────────┤
│ │
│ RunPodClient PodMonitor S3Client │
│ ├─ get_available_gpus ├─ wait_until_running ├─ upload_binary │
│ ├─ deploy_pod ├─ stream_s3_logs ├─ tail_log_file │
│ ├─ get_pod_status ├─ check_completion ├─ download_results │
│ ├─ terminate_pod └─ auto_terminate └─ list_binaries │
│ └─ list_pods │
│ │
│ RunPodConfig Errors │
│ ├─ pydantic validation ├─ RunPodError │
│ └─ .env.runpod loader ├─ PodDeploymentError │
│ ├─ S3Error │
│ └─ NetworkError │
└─────────────────────────────────────────────────────────────┘
│ │ │
▼ ▼ ▼
RunPod REST API RunPod GraphQL RunPod S3 API
(pod management) (GPU queries) (log tailing)
```
## Testing
```bash
# Deploy test pod (dry run)
python -c "
from foxhunt_runpod import RunPodClient
client = RunPodClient()
client.deploy_pod(dry_run=True)
"
# Query available GPUs
python -c "
from foxhunt_runpod import RunPodClient
client = RunPodClient()
gpus = client.get_available_gpus()
print(f'Found {len(gpus)} GPUs')
"
# Test S3 connection
python -c "
from foxhunt_runpod import S3Client
s3 = S3Client()
binaries = s3.list_binaries()
print(f'Binaries: {len(binaries)}')
"
```
## Production Checklist
- [x] Type hints for all functions
- [x] Pydantic validation for config
- [x] Retry logic with exponential backoff
- [x] Comprehensive error handling
- [x] Rich terminal output
- [x] S3-based monitoring (NO SSH)
- [x] Automatic pod termination
- [x] Configuration via env files
- [x] Docstrings for all public methods
- [x] GPU selection by price
- [x] Datacenter filtering
## License
Proprietary - Foxhunt HFT Trading System