- Fixed PSO budget calculation bug in ml/src/hyperopt/optimizer.rs - Root cause: Division by n_particles in sequential execution - Now correctly calculates max_iters = remaining_trials (no division) - Result: 50 trials complete instead of 23 (100% vs 46%) - Added comprehensive DQN hyperopt results analysis - 39/50 trials analyzed across 2 RunPod deployments - Best hyperparameters identified: LR 4.89e-5 (ultra-low) - Created DQN_HYPEROPT_RESULTS_SUMMARY.md with expert validation - GitLab CI/CD pipeline operational (48 lines fixed) - Fixed YAML syntax errors (unquoted colons) - All 7 jobs validated and working - Warning cleanup complete (136 → 0 warnings) - Removed 143 lines dead code - Fixed visibility, unused imports, Debug traits - Archived Wave D reports to docs/archive/ - 8 early stopping reports moved - Root directory cleaned up 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com>
9.0 KiB
foxhunt-deploy
Rust CLI tool for managing Foxhunt ML training deployments on RunPod GPU infrastructure.
Features
- 🐳 Docker Management: Build and push Docker images to registry
- ☁️ RunPod Deployment: Deploy training jobs to GPU pods with one command
- 📊 S3 Log Monitoring: Stream real-time logs from RunPod S3 buckets
- 🎨 Colorized Output: Easy-to-read log levels (INFO, WARN, ERROR)
- 🔍 Filtering: Regex-based log filtering
- ⚙️ Configuration: Flexible TOML-based configuration
Installation
cd foxhunt-deploy
cargo build --release
# Copy to PATH
sudo cp target/release/foxhunt-deploy /usr/local/bin/
Quick Start
1. Initialize Configuration
foxhunt-deploy init
This creates ~/.foxhunt/config.toml with default settings.
2. Configure Credentials
Edit ~/.foxhunt/config.toml:
[runpod]
api_key = "YOUR_RUNPOD_API_KEY"
default_gpu_type = "RTX A4000"
default_datacenter = "EUR-IS-1"
[docker]
registry = "jgrusewski"
image_name = "foxhunt"
tag = "latest"
[s3]
endpoint = "https://s3api-eur-is-1.runpod.io"
bucket = "se3zdnb5o4"
region = "us-east-1"
poll_interval_secs = 5
Set S3 credentials:
export AWS_ACCESS_KEY_ID="your_runpod_access_key"
export AWS_SECRET_ACCESS_KEY="your_runpod_secret_key"
3. Deploy a Training Job
foxhunt-deploy deploy \
--command "train_ppo --epochs 100" \
--gpu "RTX A4000" \
--datacenter "EUR-IS-1"
4. Monitor Logs
# List available log files
foxhunt-deploy monitor <pod_id> --list
# Show last 50 lines
foxhunt-deploy monitor <pod_id> --tail 50
# Follow logs in real-time
foxhunt-deploy monitor <pod_id> --follow
# Filter for specific patterns
foxhunt-deploy monitor <pod_id> --follow --filter "epoch|loss"
Commands
init
Initialize configuration file.
foxhunt-deploy init
Creates ~/.foxhunt/config.toml with default settings.
build
Build Docker image.
foxhunt-deploy build [OPTIONS]
Options:
--tag <TAG> Docker image tag [default: latest]
--push Push to registry after building
--dockerfile <PATH> Path to Dockerfile [default: Dockerfile.foxhunt-build]
Example:
# Build and push
foxhunt-deploy build --tag v1.0.0 --push
deploy
Deploy training job to RunPod GPU.
foxhunt-deploy deploy [OPTIONS] --command <COMMAND>
Required:
--command <COMMAND> Training command to run
Options:
--name <NAME> Pod name (auto-generated if not provided)
--gpu <GPU> GPU type [default: RTX A4000]
--datacenter <DC> Datacenter [default: EUR-IS-1]
--tag <TAG> Docker image tag [default: latest]
--env <KEY=VALUE> Environment variables (can be repeated)
--volume-size <GB> Network volume size in GB [default: 100]
Examples:
# Deploy PPO training
foxhunt-deploy deploy \
--command "train_ppo --epochs 100" \
--gpu "RTX A4000"
# Deploy with custom env vars
foxhunt-deploy deploy \
--command "train_dqn --data /data/ES_FUT.parquet" \
--env "RUST_LOG=debug" \
--env "CUDA_VISIBLE_DEVICES=0"
# Deploy to specific datacenter
foxhunt-deploy deploy \
--command "train_tft --epochs 50" \
--datacenter "US-TX-1" \
--gpu "RTX 4090"
monitor
Monitor pod logs from S3.
foxhunt-deploy monitor [OPTIONS] <POD_ID>
Arguments:
<POD_ID> Pod ID to monitor
Options:
-f, --follow Follow logs in real-time
-t, --tail <N> Number of recent lines to show
-l, --list List available log files
--filter <PATTERN> Filter logs by regex pattern
Examples:
# List available log files
foxhunt-deploy monitor dy2bn5ninzaxma --list
# Show last 20 lines
foxhunt-deploy monitor dy2bn5ninzaxma --tail 20
# Follow logs in real-time
foxhunt-deploy monitor dy2bn5ninzaxma --follow
# Follow with initial tail (show last 10 lines, then follow)
foxhunt-deploy monitor dy2bn5ninzaxma --follow --tail 10
# Filter for training metrics
foxhunt-deploy monitor dy2bn5ninzaxma --follow --filter "epoch.*loss"
# Filter for errors only
foxhunt-deploy monitor dy2bn5ninzaxma --follow --filter "ERROR|WARN"
run
Execute commands in a running pod.
foxhunt-deploy run <POD_ID> <COMMAND>
Arguments:
<POD_ID> Pod ID
<COMMAND> Command to execute
Example:
foxhunt-deploy run dy2bn5ninzaxma "nvidia-smi"
Configuration
Config File
Location: ~/.foxhunt/config.toml
[runpod]
api_key = "YOUR_API_KEY"
default_gpu_type = "RTX A4000"
default_datacenter = "EUR-IS-1"
[docker]
registry = "jgrusewski"
image_name = "foxhunt"
tag = "latest"
[s3]
endpoint = "https://s3api-eur-is-1.runpod.io"
bucket = "se3zdnb5o4"
region = "us-east-1"
poll_interval_secs = 5
[defaults]
container_disk_gb = 50
volume_path = "/runpod-volume"
volume_size_gb = 100
support_public_ip = false
Environment Variables
Override config values with environment variables:
# RunPod
export RUNPOD_API_KEY="your_key"
# S3 (required for monitor command)
export AWS_ACCESS_KEY_ID="your_access_key"
export AWS_SECRET_ACCESS_KEY="your_secret_key"
# Optional overrides
export S3_ENDPOINT="https://s3api-eur-is-1.runpod.io"
export S3_BUCKET="se3zdnb5o4"
export S3_POLL_INTERVAL="10"
GPU Types
Common GPU options:
RTX A4000- 16GB VRAM, $0.25/hr (recommended)RTX 4090- 24GB VRAM, $0.59/hrA40- 48GB VRAM, $0.79/hrA100 80GB- 80GB VRAM, $1.89/hr
Datacenters
Available datacenters:
EUR-IS-1- Europe (Iceland)US-TX-1- US (Texas)US-OR-1- US (Oregon)
S3 Log Monitoring
How It Works
-
Log Discovery: Automatically finds log files in S3 bucket
- Priority:
training.log>stdout>stderr - Path:
s3://{bucket}/ml_training/{pod_id}/...
- Priority:
-
Real-Time Streaming: Polls S3 every 5 seconds (configurable)
- Tracks last read position
- Only downloads new content (byte-range requests)
-
Colorization: Applies colors based on log level
- 🟢 INFO (green)
- 🟡 WARN (yellow)
- 🔴 ERROR (red)
- 🔵 DEBUG (cyan)
- ⚫ TRACE (dimmed)
-
Completion Detection: Exits when training completes
- Detects keywords: "training complete", "saved final model", etc.
S3 Path Structure
s3://se3zdnb5o4/ml_training/
└── {pod_id}/
├── training_runs/
│ └── {model}/
│ └── run_{timestamp}/
│ ├── logs/
│ │ └── training.log (priority 1)
│ └── checkpoints/
├── stdout (priority 2)
└── stderr (priority 3)
Performance
- Poll Interval: 5 seconds (configurable)
- Latency: ~5-10s lag
- Network: Byte-range downloads (efficient)
- Memory: Streams chunks, doesn't load entire file
Examples
See examples/monitor_demo.sh for interactive demo:
./examples/monitor_demo.sh <pod_id>
Development
Build
cargo build --release
Run Tests
cargo test
Run with Verbose Logging
foxhunt-deploy --verbose monitor <pod_id> --follow
Troubleshooting
Monitor: No logs appearing
Cause: Missing AWS credentials or incorrect pod ID
Solution:
- Check credentials:
echo $AWS_ACCESS_KEY_ID - List log files:
foxhunt-deploy monitor <pod_id> --list - Verify pod ID in RunPod dashboard
Monitor: Permission denied
Cause: Invalid AWS credentials
Solution:
- Verify RunPod API key has S3 access
- Check bucket name in config matches RunPod
- Ensure credentials match RunPod account
Monitor: Slow updates
Cause: High poll interval or network latency
Solution:
- Reduce poll interval in config:
poll_interval_secs = 2 - Check network:
ping s3api-eur-is-1.runpod.io
Deploy: GPU not available
Cause: Requested GPU type unavailable in datacenter
Solution:
- Try different datacenter:
--datacenter "US-TX-1" - Try different GPU:
--gpu "RTX 4090" - Check RunPod dashboard for availability
Architecture
foxhunt-deploy
├── src/
│ ├── cli/ # Command-line interface
│ │ ├── build.rs # Docker build command
│ │ ├── deploy.rs # RunPod deployment command
│ │ ├── monitor.rs # S3 log monitoring command
│ │ └── run.rs # Pod execution command
│ ├── config/ # Configuration management
│ ├── docker/ # Docker operations
│ ├── runpod/ # RunPod API client
│ ├── s3/ # S3 log monitoring
│ │ ├── mod.rs # S3LogClient
│ │ ├── monitor.rs # LogMonitor
│ │ └── parser.rs # Log parsing/colorization
│ ├── utils/ # Utilities (terminal output)
│ └── error.rs # Error handling
└── examples/
└── monitor_demo.sh # Interactive demo
License
Apache-2.0
Credits
Built for the Foxhunt HFT trading system.