Files
foxhunt/docs/guides/RUNPOD_WORKFLOW_GUIDE.md
jgrusewski e393a8af89 chore(cleanup): Cleanup Wave 3 - Archive reports, organize docs, fix security issues
## Summary
Third major cleanup wave after investigating 287 remaining root files.
Archived historical reports, organized documentation, removed regeneratable
artifacts, and fixed critical security issue.

## Files Cleaned (119 total)
- Archived: 78 files (7 WAVE reports + 71 summaries) → docs/archive/
- Archived: 7 build logs → docs/archive/build_logs/
- Organized: 10 markdown files → docs/guides/ + docs/checklists/
- Deleted: 17 test/coverage artifacts (regeneratable)
- Deleted: 7 empty/obsolete files (docker override, clippy baselines)
- Deleted: 3 large files (119MB - .venv, ppo_hyperopt_output.txt, backup)

## Space Recovered
- Total: ~120.7 MB
- Large files: 119.25 MB (.venv, ppo_hyperopt_output.txt)
- Archives: 1.04 MB (summaries + build logs)
- Test artifacts: 980 KB

## Security Fix (CRITICAL)
- Fixed: certs/security.env removed from git tracking (contained JWT secrets)
- Updated: .gitignore to prevent future tracking of sensitive cert files
- Removed: 4 files from git history (security.env, production.env.template, *.serial)

## Documentation Organization
- Created: docs/archive/ (wave_reports/, summaries/, build_logs/)
- Created: docs/guides/ (7 detailed implementation guides)
- Created: docs/checklists/ (3 operational checklists)
- Retained: 30 essential .md files in root (quick refs, CLAUDE.md)

## Investigation Reports Created
- MARKDOWN_ORGANIZATION_REPORT.md
- TXT_FILES_INVENTORY_AND_ARCHIVAL_PLAN.md
- ROOT_CONFIG_FILES_ANALYSIS_REPORT.md
- DOCKER_ROOT_FILES_ANALYSIS.md
- DATABASE_INITIALIZATION_AND_SETUP_ANALYSIS.md
- (6 additional investigation/index files)

## Cleanup Wave Progress
- Wave 1: 899 files deleted (1,071,884 lines)
- Wave 2: 543 files archived/deleted (~34GB)
- Wave 3: 119 files archived/deleted/organized (~121MB)
- Total: 1,561 files cleaned, ~35.1GB space recovered

## Result
Root directory: 287 files → ~180 files (excluding investigation reports)
Clean, organized, production-ready structure maintained.

Related: Second cleanup wave (previous commit)
2025-10-30 01:46:39 +01:00

744 lines
20 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# RunPod Workflow Guide
**Last Updated**: 2025-10-30
**Status**: ✅ PRODUCTION READY
**Target Audience**: Developers doing fast iteration on ML training
---
## 🚀 Quick Start (3 Commands)
```bash
# 1. Compile binary
cargo build --release --example hyperopt_mamba2_demo --features cuda
# 2. Upload to S3
python3 scripts/upload_binary.py --binary-name hyperopt_mamba2_demo
# 3. Deploy & train
python3 scripts/runpod_deploy.py \
--command "/runpod-volume/binaries/hyperopt_mamba2_demo_cuda_$(date +%Y%m%d_%H%M%S) --epochs 50"
```
**Time**: 3-5 minutes total (compile 2m, upload 10s, deploy 30s)
**Cost**: $0.10-0.30/hour depending on GPU
---
## 📦 Module Overview
### **RunPodClient** - Pod Management
High-level API for deploying and managing GPU pods.
**Key Methods**:
- `get_available_gpus(min_vram=16)` - Query GPU types with pricing
- `deploy_pod(gpu_id, image, command)` - Deploy pod with EUR-IS datacenter filtering
- `terminate_pod(pod_id)` - Stop running pod
- `get_pod_status(pod_id)` - Check pod runtime status
- `list_pods()` - List all active pods
**Features**:
- Automatic datacenter filtering (EUR-IS-1 only for volume mounting)
- Retry logic with exponential backoff
- Graceful fallback if preferred GPU unavailable
- Docker registry authentication
### **PodMonitor** - Status & Logs
Monitor pod lifecycle and stream training logs.
**Key Methods**:
- `wait_until_running(timeout=300)` - Block until pod starts
- `stream_s3_logs(follow=True)` - Tail logs from S3 (NO SSH)
- `check_completion()` - Detect training completion
- `auto_terminate()` - Stop pod when training finishes
- `display_pod_info()` - Show pod details
**Features**:
- S3-based log tailing (byte-range requests)
- Real-time completion detection (regex patterns)
- Automatic termination on completion
- Rich progress display
### **S3Client** - File Operations
Manage binaries, models, and logs on RunPod S3.
**Key Methods**:
- `upload_binary(local_path, s3_key, force=False)` - Upload with progress
- `needs_upload(local_file, s3_key)` - MD5 checksum comparison
- `download_model(s3_key, local_path)` - Pull trained models
- `list_binaries()` - Inventory S3 binaries
- `object_exists(s3_key)` - Check file existence
**Features**:
- MD5 checksum validation (skip unchanged uploads)
- Progress bars with transfer speed
- Configurable retry logic
- Timestamped naming for versioning
### **S3LogMonitor** - Real-time Streaming
Lightweight log monitoring without pod management.
**Key Methods**:
- `stream_logs(pod_id, interval=10)` - Stream training logs
- `check_completion(log_content)` - Detect "Training complete"
**Use Case**: Monitor existing pods without full PodMonitor overhead
---
## 📝 Script Reference
### **scripts/upload_binary.py**
Upload compiled binaries to S3 with versioning.
**Usage**:
```bash
# Auto-find binary in target/release/examples/
python3 scripts/upload_binary.py --binary-name hyperopt_mamba2_demo
# Custom binary path
python3 scripts/upload_binary.py --binary-path ./custom_binary --force
# Skip timestamp (overwrite existing)
python3 scripts/upload_binary.py --binary-name hyperopt_tft_demo --no-timestamp --force
```
**Features**:
- Auto-finds binaries in `target/release/examples/`
- Validates executable permissions
- Checks file size (warns if < 100KB)
- MD5 checksum (skips upload if unchanged)
- Timestamped naming (e.g., `hyperopt_mamba2_demo_cuda_20251030_120000`)
- Rich progress bars with transfer speed
**Output**:
```
✅ Upload complete!
S3 URI: s3://se3zdnb5o4/binaries/hyperopt_mamba2_demo_cuda_20251030_120000
Next steps:
1. Binary available at: /runpod-volume/binaries/hyperopt_mamba2_demo_cuda_20251030_120000
2. Use in deployment: --command '/runpod-volume/binaries/...'
```
### **scripts/runpod_deploy.py**
Deploy GPU pods with automatic GPU selection and monitoring.
**Usage**:
```bash
# Auto-select cheapest GPU
python3 scripts/runpod_deploy.py
# Preferred GPU (falls back if unavailable)
python3 scripts/runpod_deploy.py --gpu-type "RTX 4090"
# Custom training command
python3 scripts/runpod_deploy.py \
--command "/runpod-volume/binaries/train_tft_parquet --epochs 100"
# Enable S3 log monitoring
python3 scripts/runpod_deploy.py --monitor --auto-stop --timeout 2h
# Dry run (no charges)
python3 scripts/runpod_deploy.py --dry-run
```
**Arguments**:
- `--gpu-type` - Preferred GPU (e.g., "RTX 4090", "A100")
- `--image` - Docker image (default: `jgrusewski/foxhunt:latest`)
- `--command` - Training command arguments
- `--container-disk` - Disk size in GB (default: 50)
- `--monitor` - Enable S3 log monitoring
- `--auto-stop` - Auto-terminate on completion
- `--timeout` - Max monitoring time (e.g., "30m", "2h")
- `--dry-run` - Show plan without deploying
**Features**:
- Queries 24+ GPU types with pricing
- Sorts by price (cheapest first)
- Tries GPUs sequentially until one deploys
- EUR-IS-1 datacenter filtering (volume mount requirement)
- Optional real-time log streaming
- Automatic termination on training completion
**Output**:
```
✅ POD DEPLOYED SUCCESSFULLY
Pod ID: abc123xyz
GPU: NVIDIA RTX 4090
Cost: $0.340/hr
Datacenter: EUR-IS-1
📝 NEXT STEPS:
1. Wait 2-3 minutes for pod to initialize
2. Access Jupyter at: https://abc123xyz-8888.proxy.runpod.net
3. SSH access: ssh root@abc123xyz.ssh.runpod.io
```
### **scripts/monitor_hyperopt.sh**
Monitor hyperparameter optimization trials (S3-based).
**Usage**:
```bash
# Monitor latest hyperopt run
./scripts/monitor_hyperopt.sh
# Specify study name
./scripts/monitor_hyperopt.sh mamba2_study_20251030
# Watch mode (auto-refresh every 30s)
watch -n 30 ./scripts/monitor_hyperopt.sh
```
**Output**:
```
=== Hyperopt Monitoring ===
Study: mamba2_study_20251030
Trials: 15/100 completed
Best Loss: 0.0234 (trial 12)
Time Elapsed: 42m 15s
ETA: 3h 12m
```
---
## 🎯 Complete Workflow Examples
### **Example 1: Fast Iteration (Compile → Upload → Deploy → Monitor)**
**Scenario**: Testing code changes on GPU
```bash
# 1. Make code changes
vim ml/src/mamba/mod.rs
# 2. Compile with CUDA
cargo build --release --example hyperopt_mamba2_demo --features cuda
# Time: ~2 minutes
# 3. Upload binary to S3
python3 scripts/upload_binary.py --binary-name hyperopt_mamba2_demo
# Time: ~10 seconds (21MB binary)
# Output: s3://se3zdnb5o4/binaries/hyperopt_mamba2_demo_cuda_20251030_143000
# 4. Deploy pod with new binary
python3 scripts/runpod_deploy.py \
--gpu-type "RTX A4000" \
--command "/runpod-volume/binaries/hyperopt_mamba2_demo_cuda_20251030_143000 --epochs 50" \
--monitor \
--auto-stop \
--timeout 30m
# Time: ~30 seconds (pod startup)
# 5. Logs stream automatically, pod terminates on completion
# Cost: $0.10 (30 min @ $0.17/hr)
```
**Total Time**: 3-5 minutes human time, 30 minutes GPU time
### **Example 2: Production Deployment (100 Trials, 20 Epochs)**
**Scenario**: Full hyperparameter search for production model
```bash
# 1. Compile optimized binary
RUSTFLAGS="-C target-cpu=native" cargo build --release \
--example hyperopt_mamba2_demo --features cuda
# 2. Upload to S3
python3 scripts/upload_binary.py --binary-name hyperopt_mamba2_demo
# Output: .../hyperopt_mamba2_demo_cuda_20251030_150000
# 3. Deploy on high-end GPU
python3 scripts/runpod_deploy.py \
--gpu-type "RTX 4090" \
--command "/runpod-volume/binaries/hyperopt_mamba2_demo_cuda_20251030_150000 --n-trials 100 --epochs 20 --dataset full" \
--monitor \
--auto-stop \
--timeout 4h
# 4. Monitor progress (separate terminal if needed)
./scripts/monitor_hyperopt.sh mamba2_study_20251030
# 5. Download best model after completion
aws s3 cp \
s3://se3zdnb5o4/models/mamba2_best_20251030.safetensors \
./models/ \
--profile runpod \
--endpoint-url https://s3api-eur-is-1.runpod.io
# Cost: ~$1.20 (4 hours @ $0.34/hr)
```
**Total Time**: 4 hours GPU time, auto-terminates
### **Example 3: Parallel Training (Multiple Models)**
**Scenario**: Train DQN, PPO, TFT in parallel
```bash
# 1. Compile all binaries
cargo build --release --examples --features cuda
# Builds: train_dqn, train_ppo, train_tft_parquet
# 2. Upload binaries in parallel
python3 scripts/upload_binary.py --binary-name train_dqn &
python3 scripts/upload_binary.py --binary-name train_ppo &
python3 scripts/upload_binary.py --binary-name train_tft_parquet &
wait
# 3. Deploy 3 pods simultaneously
python3 scripts/runpod_deploy.py \
--command "/runpod-volume/binaries/train_dqn_cuda_20251030_160000 --epochs 100" &
python3 scripts/runpod_deploy.py \
--command "/runpod-volume/binaries/train_ppo_cuda_20251030_160000 --epochs 100" &
python3 scripts/runpod_deploy.py \
--command "/runpod-volume/binaries/train_tft_parquet_cuda_20251030_160000 --epochs 50" &
wait
# 4. Monitor all pods
python3 -c "
from foxhunt_runpod import RunPodClient
client = RunPodClient()
pods = client.list_pods()
client.display_pods_table(pods)
"
# Cost: ~$0.50 (1 hour @ $0.16/hr × 3 pods)
```
**Total Time**: 1 hour parallel execution
### **Example 4: Log Monitoring and Debugging**
**Scenario**: Monitor existing pod, debug training issues
```bash
# 1. List active pods
python3 -c "
from foxhunt_runpod import RunPodClient
client = RunPodClient()
pods = client.list_pods()
for pod in pods:
print(f\"{pod['id']}: {pod['desiredStatus']} ({pod['gpu']['count']}x {pod['machine']['gpuType']['displayName']})\")
"
# 2. Attach to pod logs
python3 -c "
from foxhunt_runpod import PodMonitor
monitor = PodMonitor('abc123xyz')
monitor.stream_s3_logs(follow=True)
"
# 3. Check pod status
python3 -c "
from foxhunt_runpod import RunPodClient
client = RunPodClient()
status = client.get_pod_status('abc123xyz')
print(f\"Runtime: {status['runtime']['status']}\")
print(f\"GPU Usage: {status['runtime'].get('gpuMetrics', 'N/A')}\")
"
# 4. Download logs for offline analysis
aws s3 cp \
s3://se3zdnb5o4/logs/abc123xyz/training.log \
./debug_logs/ \
--profile runpod \
--endpoint-url https://s3api-eur-is-1.runpod.io
# 5. Terminate pod when done debugging
python3 -c "
from foxhunt_runpod import RunPodClient
client = RunPodClient()
client.terminate_pod('abc123xyz')
"
```
---
## 🔧 Module API Reference
### **RunPodClient API**
```python
from foxhunt_runpod import RunPodClient
# Initialize
client = RunPodClient(
api_key="your_key", # Or uses RUNPOD_API_KEY env var
volume_id="se3zdnb5o4", # Network volume ID
registry_auth_id="optional" # Docker registry auth
)
# Query GPUs
gpus = client.get_available_gpus(min_vram=16)
# Returns: [{'id': 'gpu_id', 'name': 'RTX 4090', 'vram': 24, 'price': 0.34, ...}]
# Deploy pod
pod = client.deploy_pod(
gpu_id="NVIDIA RTX 4090", # From gpus list
image="jgrusewski/foxhunt:latest",
command="--epochs 50", # Optional, overrides Dockerfile CMD
container_disk=50, # GB
ports=["8888/http", "22/tcp"],
env={"DEBUG": "1"}, # Environment variables
dry_run=False # True = show plan only
)
# Returns: {'id': 'abc123xyz', 'costPerHr': 0.34, ...}
# Check status
status = client.get_pod_status("abc123xyz")
# Returns: {'runtime': {'status': 'RUNNING', 'gpuMetrics': {...}}, ...}
# List pods
pods = client.list_pods()
# Returns: [{'id': 'abc123', 'desiredStatus': 'RUNNING', ...}, ...]
# Terminate pod
success = client.terminate_pod("abc123xyz")
# Returns: True if terminated
```
### **PodMonitor API**
```python
from foxhunt_runpod import PodMonitor
# Initialize
monitor = PodMonitor(
pod_id="abc123xyz",
config=None, # Uses global config if None
client=None, # Uses default RunPodClient if None
s3_client=None # Uses default S3Client if None
)
# Wait for pod to start
monitor.wait_until_running(
timeout=300, # Seconds
poll_interval=5 # Seconds between checks
)
# Stream logs from S3
monitor.stream_s3_logs(
follow=True, # Continue until completion
poll_interval=10, # Seconds between S3 polls
max_lines=None # Limit output (None = unlimited)
)
# Check if training complete (manual)
complete = monitor.check_completion()
# Returns: True if "Training complete" pattern found
# Auto-terminate when training finishes
monitor.auto_terminate()
# Blocks until completion, then stops pod
# Display pod info
monitor.display_pod_info()
# Prints: GPU type, cost, datacenter, status
```
### **S3Client API**
```python
from foxhunt_runpod import S3Client
# Initialize
s3 = S3Client(config=None) # Uses global config if None
# Check if upload needed (MD5)
needs_upload, reason = s3.needs_upload(
local_file=Path("./binary"),
s3_key="binaries/my_binary"
)
# Returns: (True, "File not found on S3") or (False, "Checksums match")
# Upload binary
success = s3.upload_binary(
local_path=Path("./binary"),
s3_key="binaries/my_binary",
force=False, # True = skip checksum
progress_callback=lambda bytes_transferred: print(bytes_transferred)
)
# Check if object exists
exists = s3.object_exists("binaries/my_binary")
# Returns: True or False
# Get object size
size_bytes = s3.get_object_size("binaries/my_binary")
# List binaries
binaries = s3.list_binaries()
# Returns: [{'name': 'my_binary', 'size': 21000000, 'modified': datetime}, ...]
# Download model
s3.download_model(
s3_key="models/mamba2_best.safetensors",
local_path=Path("./models/mamba2.safetensors"),
progress_callback=lambda bytes_transferred: print(bytes_transferred)
)
```
### **S3LogMonitor API**
```python
from foxhunt_runpod.s3_monitor import S3LogMonitor
# Initialize
monitor = S3LogMonitor(
bucket_name="se3zdnb5o4",
aws_access_key="your_key",
aws_secret_key="your_secret",
endpoint_url="https://s3api-eur-is-1.runpod.io"
)
# Stream logs
monitor.stream_logs(
pod_id="abc123xyz",
interval=10, # Polling interval (seconds)
timeout=7200, # Max time (seconds, None = infinite)
completion_callback=monitor.check_completion # Detect completion
)
# Completion check (used as callback)
is_complete = monitor.check_completion(log_content)
# Returns: True if "Training complete" found
```
---
## 🔍 Troubleshooting
### **Issue: Binary upload fails with "File not executable"**
**Cause**: Binary lacks execute permissions
**Fix**:
```bash
chmod +x target/release/examples/my_binary
python3 scripts/upload_binary.py --binary-name my_binary
```
### **Issue: Pod deployment fails with "No GPUs available in EUR-IS"**
**Cause**: Requested GPU not available in EUR-IS-1 datacenter
**Fix**:
```bash
# 1. Check available GPUs
python3 -c "
from foxhunt_runpod import RunPodClient
client = RunPodClient()
gpus = client.get_available_gpus(min_vram=16)
for gpu in gpus:
print(f\"{gpu['name']}: ${gpu['price']:.3f}/hr (avail: {gpu['global_available']})\")
"
# 2. Try without --gpu-type (auto-selects cheapest)
python3 scripts/runpod_deploy.py
# 3. Or specify different GPU
python3 scripts/runpod_deploy.py --gpu-type "RTX A4000"
```
### **Issue: Log monitoring shows "Log file not found"**
**Cause**: Training hasn't started writing logs yet
**Fix**: Wait 30-60 seconds for pod initialization, then retry
### **Issue: Upload says "up-to-date" but binary changed**
**Cause**: MD5 checksum unchanged (same binary content)
**Fix**:
```bash
# Force upload (ignores checksum)
python3 scripts/upload_binary.py --binary-name my_binary --force
```
### **Issue: Deployment succeeds but training doesn't start**
**Cause**: Docker command incorrect or binary path wrong
**Debug**:
```bash
# 1. Check pod logs
python3 -c "
from foxhunt_runpod import PodMonitor
monitor = PodMonitor('abc123xyz')
monitor.stream_s3_logs(follow=False)
"
# 2. SSH into pod (if needed)
ssh root@abc123xyz.ssh.runpod.io
# 3. Verify binary exists
ls -lh /runpod-volume/binaries/
# 4. Test binary manually
/runpod-volume/binaries/my_binary --help
```
### **Issue: Pod costs more than expected**
**Cause**: Forgot to terminate pod
**Fix**:
```bash
# 1. List active pods
python3 -c "
from foxhunt_runpod import RunPodClient
client = RunPodClient()
pods = client.list_pods()
for pod in pods:
cost = pod.get('costPerHr', 0)
print(f\"{pod['id']}: ${cost:.3f}/hr\")
"
# 2. Terminate all (be careful!)
python3 -c "
from foxhunt_runpod import RunPodClient
client = RunPodClient()
pods = client.list_pods()
for pod in pods:
print(f\"Terminating {pod['id']}...\")
client.terminate_pod(pod['id'])
"
```
### **Issue: S3 upload slow or times out**
**Cause**: Network issues or large binary
**Fix**:
```bash
# 1. Check binary size
ls -lh target/release/examples/my_binary
# Should be 14-21MB. If >50MB, binary may include debug symbols.
# 2. Strip debug symbols
strip target/release/examples/my_binary
# 3. Verify network (ping S3 endpoint)
ping s3api-eur-is-1.runpod.io
# 4. Try with increased timeout (not yet implemented, file issue)
```
---
## 💡 Best Practices
### **Fast Iteration Tips**
1. **Use `--dry-run` First**: Verify deployment plan before charges
```bash
python3 scripts/runpod_deploy.py --dry-run
```
2. **Timestamped Binaries**: Keep versions, easy rollback
```bash
# Default behavior (timestamped)
python3 scripts/upload_binary.py --binary-name my_binary
# Output: my_binary_cuda_20251030_143000
```
3. **Skip Unchanged Uploads**: MD5 checksum automatically skips
```bash
# First run: uploads
python3 scripts/upload_binary.py --binary-name my_binary
# Second run (no changes): skips
python3 scripts/upload_binary.py --binary-name my_binary
# Output: "✓ Binary up-to-date: Checksums match"
```
4. **Use Cheapest GPU First**: Test on RTX A4000/A5000 ($0.16-0.17/hr)
```bash
python3 scripts/runpod_deploy.py # Auto-selects cheapest
```
5. **Enable Auto-Termination**: Never forget to stop pods
```bash
python3 scripts/runpod_deploy.py --monitor --auto-stop --timeout 2h
```
### **Cost Optimization**
| Training Scenario | Recommended GPU | Cost/hr | Duration | Total Cost |
|-------------------|-----------------|---------|----------|------------|
| Quick test (1 epoch) | RTX A5000 | $0.16 | 1 min | $0.003 |
| Dev iteration (10 epochs) | RTX A4000 | $0.17 | 5 min | $0.014 |
| Full training (100 epochs) | RTX 4090 | $0.34 | 30 min | $0.17 |
| Hyperopt (100 trials) | RTX 4090 | $0.34 | 4 hours | $1.36 |
| Production (1000 epochs) | A100 PCIe | $1.19 | 2 hours | $2.38 |
**Tip**: Use `--timeout` to prevent runaway costs
### **Workflow Shortcuts**
**Alias for common commands** (add to `.bashrc`):
```bash
# Fast deploy
alias runpod-deploy='python3 scripts/runpod_deploy.py'
# Fast upload
alias runpod-upload='python3 scripts/upload_binary.py --binary-name'
# List pods
alias runpod-list='python3 -c "from foxhunt_runpod import RunPodClient; c = RunPodClient(); c.display_pods_table(c.list_pods())"'
# Kill all pods (dangerous!)
alias runpod-kill-all='python3 -c "from foxhunt_runpod import RunPodClient; c = RunPodClient(); [c.terminate_pod(p[\"id\"]) for p in c.list_pods()]"'
```
**One-liner workflows**:
```bash
# Compile + upload + deploy
cargo build --release --example my_binary --features cuda && \
python3 scripts/upload_binary.py --binary-name my_binary && \
python3 scripts/runpod_deploy.py --command "/runpod-volume/binaries/my_binary_cuda_$(date +%Y%m%d_%H%M%S)"
# Monitor existing pod
watch -n 10 'aws s3 cp s3://se3zdnb5o4/logs/abc123xyz/training.log - --profile runpod --endpoint-url https://s3api-eur-is-1.runpod.io | tail -20'
```
### **Security Best Practices**
1. **Never commit `.env.runpod`**: Contains secrets
2. **Rotate API keys quarterly**: RunPod dashboard → Settings → API Keys
3. **Use Docker registry auth**: Private images only
4. **Limit S3 bucket permissions**: Read/write only to `/binaries` and `/models`
---
## 📚 Related Documentation
- **RUNPOD_DEPLOY_QUICK_START.md** - Deployment quick reference
- **RUNPOD_PYTHON_QUICK_REF.md** - Module implementation guide
- **ML_TRAINING_PARQUET_GUIDE.md** - Complete training workflows
- **RUNPOD_VOLUME_MOUNT_ARCHITECTURE.md** - S3 volume architecture
- **CLAUDE.md** - System architecture and status
---
## 🎯 Next Steps
1. **Test workflow**: Run Example 1 (Fast Iteration)
2. **Customize**: Add project-specific binaries
3. **Automate**: Create CI/CD pipeline for nightly training
4. **Monitor**: Set up Grafana dashboards for S3 logs
5. **Optimize**: Profile GPU utilization, reduce training time
---
**Questions?** Check troubleshooting section or review existing documentation.