Files
foxhunt/MONITOR_LOGS_QUICK_REF.md

8.9 KiB

Monitor Logs Quick Reference

Script: scripts/monitor_logs.py Purpose: Real-time S3 log monitoring for RunPod ML training runs


Requirements

  1. Virtual Environment: MUST activate .venv first

    source .venv/bin/activate
    
  2. Dependencies: Already installed in .venv

    • rich (colored output)
    • boto3 (S3 access)
    • pydantic-settings (config)
  3. Credentials: .env.runpod with S3 credentials


Usage Examples

1. List Recent Runs

# Default: Show 20 most recent runs
python3 scripts/monitor_logs.py

# Custom limit
python3 scripts/monitor_logs.py --limit 10

Output: Formatted table with:

  • Run ID
  • Model type (MAMBA2, DQN, PPO, TFT)
  • Last modified time (relative)
  • Log file size
  • Hyperopt indicator (✓ if trials.json exists)

2. Monitor Specific Run (One-time Snapshot)

# View current log contents (no streaming)
python3 scripts/monitor_logs.py --run-id run_20251029_145151_hyperopt

3. Monitor with Continuous Streaming

# Stream logs in real-time until completion
python3 scripts/monitor_logs.py --run-id run_20251029_145151_hyperopt --follow

Features:

  • Tail new log entries every 5 seconds (configurable)
  • Color-coded output:
    • Red: Errors (CUDA OOM, RuntimeError, etc.)
    • Yellow: Warnings
    • Green: Success messages, completion
  • Auto-detects training completion
  • Shows hyperopt trial count every 30 seconds

4. Monitor with Timeout

# Stream for maximum 30 minutes
python3 scripts/monitor_logs.py --run-id run_20251029_145151_hyperopt --follow --timeout 30m

# Stream for 2 hours
python3 scripts/monitor_logs.py --run-id run_20251029_145151_hyperopt --follow --timeout 2h

# Stream for 45 seconds
python3 scripts/monitor_logs.py --run-id run_20251029_145151_hyperopt --follow --timeout 45s

Timeout Formats:

  • 30s = 30 seconds
  • 15m = 15 minutes
  • 2h = 2 hours

5. Custom Poll Interval

# Check for new logs every 10 seconds (default: 5s)
python3 scripts/monitor_logs.py --run-id run_20251029_145151_hyperopt --follow --interval 10

6. Monitor by Pod ID (Alternative)

# Monitor using pod ID instead of run ID
python3 scripts/monitor_logs.py --pod-id w4srx0tgm5hfgu --follow

S3 Structure

s3://se3zdnb5o4/
├── ml_training/
│   └── training_runs/
│       ├── mamba2/
│       │   └── run_20251029_145151_hyperopt/
│       │       ├── logs/
│       │       │   └── training.log       # Tailed by script
│       │       ├── hyperopt/
│       │       │   └── trials.json        # Checked every 30s
│       │       └── models/
│       │           └── best_model.safetensors
│       ├── dqn/
│       ├── ppo/
│       └── tft/

Output Examples

Listing Runs

                            Recent Training Runs
╭──────────────────────────────┬────────┬───────────────┬──────────┬────────╮
│ Run ID                       │ Model  │ Last Modified │ Log Size │ Trials │
├──────────────────────────────┼────────┼───────────────┼──────────┼────────┤
│ run_20251029_223100_hyperopt │ MAMBA2 │ 39m ago       │    1.0KB │   ✓    │
│ run_20251029_222832_hyperopt │ MAMBA2 │ 51m ago       │     471B │   -    │
│ run_20251029_191027_hyperopt │ MAMBA2 │ 4h ago        │     468B │   -    │
╰──────────────────────────────┴────────┴───────────────┴──────────┴────────╯

Streaming Logs

╭──────────────────────────────── Log Monitor ─────────────────────────────────╮
│ Monitoring Run: run_20251029_145151_hyperopt                                 │
│ Model: MAMBA2                                                                │
│ Log: ml_training/training_runs/mamba2/run_20251029_145151_hyperopt/logs/... │
│ Hyperopt: Yes                                                                │
╰──────────────────────────────────────────────────────────────────────────────╯

Streaming logs... (Press Ctrl+C to stop)
──────────────────────────────────────────────────────────────────────
[2025-10-29 15:08:07] === Starting MAMBA-2 Trial ===
Params: Mamba2Params { ... }
[2025-10-29 15:10:41] Training completed in 154.23s: val_loss=0.089249
📊 Hyperopt trials: 3
──────────────────────────────────────────────────────────────────────
✅ Training completed!

Error Patterns (Auto-Detected)

Script highlights these patterns in RED:

  • CUDA out of memory
  • RuntimeError:
  • AssertionError:
  • FAILED:
  • ERROR:
  • panic!

Script highlights these patterns in GREEN:

  • Training complete
  • Model saved to
  • ✓ Training finished
  • SUCCESS:
  • Hyperparameter optimization complete

Comparison: monitor_logs.py vs monitor_hyperopt.sh

Feature monitor_logs.py monitor_hyperopt.sh
List runs Formatted table Manual S3 listing
Real-time streaming Byte-range tailing Periodic polling
Color output Rich colors ⚠️ Basic ANSI
Auto-completion detection Pattern matching Manual check
Timeout support Configurable None
Hyperopt trials Auto-detected ⚠️ Manual check
Pod monitoring Via PodMonitor SSH only

Recommendation: Use monitor_logs.py for all log monitoring tasks.


Troubleshooting

Error: "Not running in a virtual environment"

# Solution: Activate .venv
source .venv/bin/activate

Error: "Configuration file not found"

# Solution: Ensure .env.runpod exists in project root
ls -la .env.runpod

# If missing, copy from template or restore from backup

Error: "Run not found"

# Solution: List available runs first
python3 scripts/monitor_logs.py

# Verify run ID exactly matches (case-sensitive)

No new logs appearing

  • Check if training is actually running (pod status)
  • Verify S3 sync is working (check S3 directly)
  • Increase poll interval if network is slow (--interval 10)

Integration with Deployment

Standard Workflow

# 1. Deploy training pod
python3 scripts/runpod_deploy.py --gpu-type "RTX A4000" --monitor

# 2. Monitor logs separately (in another terminal)
source .venv/bin/activate
python3 scripts/monitor_logs.py --follow --timeout 2h

# 3. List all runs to find completed ones
python3 scripts/monitor_logs.py --limit 50

Automated Monitoring

# Deploy with auto-monitoring (blocks until complete)
python3 scripts/runpod_deploy.py --monitor --auto-stop --timeout 2h

Architecture

Log Tailing Method

  • NO SSH required: Pure S3 byte-range requests
  • Efficient: Only fetches new bytes since last check
  • Real-time: 5-second polling (configurable)
  • Robust: Handles missing files, network errors

Classes Used

  • PodMonitor: Pod status and log streaming (from foxhunt_runpod)
  • S3Client: S3 operations with retry logic
  • RunPodConfig: Configuration from .env.runpod

Performance

  • S3 API calls: ~12/minute (5s interval)
  • Bandwidth: ~1KB-1MB per check (depends on log size)
  • Latency: <500ms per S3 request
  • Cost: Negligible (S3 GET requests: $0.0004/1000)

Future Enhancements

  1. Web UI: Real-time dashboard with WebSockets
  2. Multi-run monitoring: Monitor multiple runs simultaneously
  3. Alert notifications: Slack/email on completion/errors
  4. Log analytics: Parse metrics, plot training curves
  5. Historical replay: View logs from completed runs with timestamps

  • scripts/runpod_deploy.py: Deploy training pods with optional monitoring
  • scripts/upload_binary.py: Upload training binaries to S3
  • scripts/monitor_hyperopt.sh: Legacy monitoring script (deprecated)

Last Updated: 2025-10-30 Status: Production Ready Tests: Manual validation complete