- Fixed PSO budget calculation bug in ml/src/hyperopt/optimizer.rs - Root cause: Division by n_particles in sequential execution - Now correctly calculates max_iters = remaining_trials (no division) - Result: 50 trials complete instead of 23 (100% vs 46%) - Added comprehensive DQN hyperopt results analysis - 39/50 trials analyzed across 2 RunPod deployments - Best hyperparameters identified: LR 4.89e-5 (ultra-low) - Created DQN_HYPEROPT_RESULTS_SUMMARY.md with expert validation - GitLab CI/CD pipeline operational (48 lines fixed) - Fixed YAML syntax errors (unquoted colons) - All 7 jobs validated and working - Warning cleanup complete (136 → 0 warnings) - Removed 143 lines dead code - Fixed visibility, unused imports, Debug traits - Archived Wave D reports to docs/archive/ - 8 early stopping reports moved - Root directory cleaned up 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com>
314 lines
7.6 KiB
Markdown
314 lines
7.6 KiB
Markdown
# DQN Runpod Monitoring Quick Reference
|
|
|
|
**Purpose**: Quick commands for monitoring DQN training on Runpod
|
|
|
|
---
|
|
|
|
## Real-Time Monitoring
|
|
|
|
### View Training Logs (Automatic with --monitor)
|
|
The deployment script automatically streams logs when using `--monitor` flag.
|
|
|
|
**Manual Log Monitoring**:
|
|
```bash
|
|
# If you need to reconnect to logs
|
|
export PYTHONPATH=/home/jgrusewski/Work/foxhunt:$PYTHONPATH
|
|
source .venv/bin/activate
|
|
|
|
python3 scripts/monitor_logs.py <pod_id>
|
|
```
|
|
|
|
---
|
|
|
|
## Key Metrics to Watch
|
|
|
|
### 1. Reward Variance (Every Epoch)
|
|
```
|
|
Look for: "Reward std=0.152"
|
|
✅ GOOD: std > 0.1 (healthy variance)
|
|
⚠️ BAD: std < 0.01 (constant rewards - training bug)
|
|
```
|
|
|
|
### 2. Action Distribution (Every 10 Epochs)
|
|
```
|
|
Action Distribution [Epoch 10]:
|
|
BUY=28.5% (1423) | SELL=31.2% (1556) | HOLD=40.3% (2011)
|
|
|
|
✅ GOOD: Each action 20-40%
|
|
⚠️ BAD: One action < 10% or > 80%
|
|
```
|
|
|
|
### 3. Q-Value Balance (Every 10 Epochs)
|
|
```
|
|
Average Q-values [Epoch 10]:
|
|
BUY=0.1234 | SELL=0.1189 | HOLD=0.1201
|
|
|
|
✅ GOOD: All within 50% of each other
|
|
⚠️ BAD: One action > 10x another
|
|
```
|
|
|
|
### 4. Exploration Decay
|
|
```
|
|
Look for: "Final epsilon: 0.XXXX"
|
|
|
|
Epoch 1: epsilon ~0.300
|
|
Epoch 10: epsilon ~0.270
|
|
Epoch 50: epsilon ~0.150
|
|
Epoch 100: epsilon ~0.050
|
|
```
|
|
|
|
### 5. Training Progress
|
|
```
|
|
Look for checkpoint saves:
|
|
💾 Checkpoint saved: dqn_epoch_10.safetensors (12345678 bytes)
|
|
💾 Checkpoint saved: dqn_epoch_20.safetensors (12345678 bytes)
|
|
...
|
|
```
|
|
|
|
---
|
|
|
|
## S3 Checkpoint Verification
|
|
|
|
### List All Checkpoints
|
|
```bash
|
|
aws s3 ls s3://se3zdnb5o4/ml_training/dqn_fixed_reward/checkpoints/ \
|
|
--profile runpod \
|
|
--endpoint-url https://s3api-eur-is-1.runpod.io \
|
|
--recursive
|
|
```
|
|
|
|
**Expected Output**:
|
|
```
|
|
2025-11-01 10:15:32 12345678 checkpoints/dqn_epoch_10.safetensors
|
|
2025-11-01 10:25:45 12345678 checkpoints/dqn_epoch_20.safetensors
|
|
2025-11-01 10:35:58 12345678 checkpoints/dqn_epoch_30.safetensors
|
|
...
|
|
```
|
|
|
|
### Download Latest Checkpoint
|
|
```bash
|
|
# List checkpoints sorted by time
|
|
aws s3 ls s3://se3zdnb5o4/ml_training/dqn_fixed_reward/ \
|
|
--profile runpod \
|
|
--endpoint-url https://s3api-eur-is-1.runpod.io \
|
|
--recursive | sort -k1,2
|
|
|
|
# Download latest
|
|
aws s3 cp s3://se3zdnb5o4/ml_training/dqn_fixed_reward/dqn_final_epoch100.safetensors \
|
|
ml/trained_models/ \
|
|
--profile runpod \
|
|
--endpoint-url https://s3api-eur-is-1.runpod.io
|
|
```
|
|
|
|
---
|
|
|
|
## Pod Management
|
|
|
|
### Check Pod Status
|
|
```bash
|
|
curl -H "Authorization: Bearer $RUNPOD_API_KEY" \
|
|
https://rest.runpod.io/v1/pods
|
|
```
|
|
|
|
**Response**:
|
|
```json
|
|
{
|
|
"pods": [
|
|
{
|
|
"id": "abc123",
|
|
"name": "foxhunt-training",
|
|
"status": "RUNNING",
|
|
"gpuType": "RTX A4000",
|
|
"runtime": 3600 // seconds
|
|
}
|
|
]
|
|
}
|
|
```
|
|
|
|
### Stop Pod (Graceful)
|
|
```bash
|
|
POD_ID="<your_pod_id>"
|
|
|
|
curl -X POST \
|
|
-H "Authorization: Bearer $RUNPOD_API_KEY" \
|
|
https://rest.runpod.io/v1/pods/$POD_ID/stop
|
|
```
|
|
|
|
### Terminate Pod (Force)
|
|
```bash
|
|
POD_ID="<your_pod_id>"
|
|
|
|
curl -X POST \
|
|
-H "Authorization: Bearer $RUNPOD_API_KEY" \
|
|
https://rest.runpod.io/v1/pods/$POD_ID/terminate
|
|
```
|
|
|
|
---
|
|
|
|
## Training Health Checks
|
|
|
|
### Check 1: Training Started
|
|
**What to look for**: "🏋️ Starting training..."
|
|
**When**: Within 1-2 minutes of deployment
|
|
**If missing**: Check pod logs for errors
|
|
|
|
### Check 2: Replay Buffer Filling
|
|
**What to look for**: "Building replay buffer: X/104346"
|
|
**When**: First 5-10 minutes
|
|
**If stuck**: Data loading issue or OOM
|
|
|
|
### Check 3: First Checkpoint
|
|
**What to look for**: "💾 Checkpoint saved: dqn_epoch_10.safetensors"
|
|
**When**: ~10-15 minutes
|
|
**If missing**: Check S3 credentials or disk space
|
|
|
|
### Check 4: Reward Variance
|
|
**What to look for**: "Reward std > 0.1"
|
|
**When**: Every epoch after 10
|
|
**If failing**: Reward function bug (constant rewards)
|
|
|
|
### Check 5: Action Diversity
|
|
**What to look for**: "Action Distribution [Epoch X]"
|
|
**When**: Every 10 epochs
|
|
**If imbalanced**: Epsilon too low or Q-value bug
|
|
|
|
---
|
|
|
|
## Cost Monitoring
|
|
|
|
### Calculate Current Cost
|
|
```bash
|
|
# Get pod runtime in seconds
|
|
POD_RUNTIME_HOURS=$(echo "scale=2; $RUNTIME_SECONDS / 3600" | bc)
|
|
|
|
# RTX A4000 = $0.25/hr
|
|
COST=$(echo "scale=2; $POD_RUNTIME_HOURS * 0.25" | bc)
|
|
|
|
echo "Current cost: \$$COST"
|
|
```
|
|
|
|
### Set Cost Alarm (Manual)
|
|
```bash
|
|
# Check every 10 minutes
|
|
while true; do
|
|
RUNTIME=$(curl -s -H "Authorization: Bearer $RUNPOD_API_KEY" \
|
|
https://rest.runpod.io/v1/pods/$POD_ID | jq -r '.runtime')
|
|
|
|
HOURS=$(echo "scale=2; $RUNTIME / 3600" | bc)
|
|
COST=$(echo "scale=2; $HOURS * 0.25" | bc)
|
|
|
|
echo "[$(date)] Runtime: ${HOURS}h | Cost: \$${COST}"
|
|
|
|
# Alert if cost > $1.00 (4 hours)
|
|
if (( $(echo "$COST > 1.00" | bc -l) )); then
|
|
echo "⚠️ WARNING: Cost exceeded \$1.00! Consider terminating pod."
|
|
fi
|
|
|
|
sleep 600 # 10 minutes
|
|
done
|
|
```
|
|
|
|
---
|
|
|
|
## Troubleshooting Commands
|
|
|
|
### Issue: No logs appearing
|
|
```bash
|
|
# Check if S3 credentials are set
|
|
grep -E "RUNPOD_S3" .env.runpod
|
|
|
|
# Try manual S3 list
|
|
aws s3 ls s3://se3zdnb5o4/ \
|
|
--profile runpod \
|
|
--endpoint-url https://s3api-eur-is-1.runpod.io
|
|
```
|
|
|
|
### Issue: Pod stuck
|
|
```bash
|
|
# Check pod status
|
|
curl -H "Authorization: Bearer $RUNPOD_API_KEY" \
|
|
https://rest.runpod.io/v1/pods/$POD_ID
|
|
|
|
# Force restart
|
|
curl -X POST \
|
|
-H "Authorization: Bearer $RUNPOD_API_KEY" \
|
|
https://rest.runpod.io/v1/pods/$POD_ID/restart
|
|
```
|
|
|
|
### Issue: OOM errors
|
|
```bash
|
|
# Check pod logs for memory errors
|
|
python3 scripts/monitor_logs.py $POD_ID | grep -i "out of memory\|OOM\|killed"
|
|
|
|
# If OOM, reduce batch_size or buffer_size and redeploy
|
|
```
|
|
|
|
---
|
|
|
|
## Expected Timeline
|
|
|
|
| Time | Milestone | What to Check |
|
|
|------|-----------|---------------|
|
|
| 0-2 min | Pod deployed | Logs start streaming |
|
|
| 2-5 min | Training starts | "🏋️ Starting training..." |
|
|
| 5-10 min | Replay buffer fills | "Building replay buffer: 500/104346" |
|
|
| 10-15 min | Epoch 10 complete | First checkpoint saved |
|
|
| 15-30 min | Action diversity logs | "Action Distribution [Epoch 10]" |
|
|
| 30-60 min | Epoch 50 complete | Mid-training checkpoint |
|
|
| 60-90 min | Epoch 100 complete | Final checkpoint saved |
|
|
| 90-120 min | Training complete | "🎉 DQN training complete!" |
|
|
|
|
---
|
|
|
|
## Quick Checks (Copy-Paste)
|
|
|
|
```bash
|
|
# 1. Check if pod is running
|
|
curl -H "Authorization: Bearer $RUNPOD_API_KEY" \
|
|
https://rest.runpod.io/v1/pods | jq -r '.pods[] | select(.name=="foxhunt-training") | .status'
|
|
|
|
# 2. Count checkpoints saved
|
|
aws s3 ls s3://se3zdnb5o4/ml_training/dqn_fixed_reward/checkpoints/ \
|
|
--profile runpod \
|
|
--endpoint-url https://s3api-eur-is-1.runpod.io \
|
|
--recursive | wc -l
|
|
|
|
# 3. Get latest checkpoint timestamp
|
|
aws s3 ls s3://se3zdnb5o4/ml_training/dqn_fixed_reward/ \
|
|
--profile runpod \
|
|
--endpoint-url https://s3api-eur-is-1.runpod.io \
|
|
--recursive | sort -k1,2 | tail -1
|
|
|
|
# 4. Estimate current cost
|
|
POD_RUNTIME=$(curl -s -H "Authorization: Bearer $RUNPOD_API_KEY" \
|
|
https://rest.runpod.io/v1/pods | jq -r '.pods[] | select(.name=="foxhunt-training") | .runtime')
|
|
echo "Runtime: $(echo "scale=2; $POD_RUNTIME / 3600" | bc)h | Cost: \$$(echo "scale=2; $POD_RUNTIME / 3600 * 0.25" | bc)"
|
|
```
|
|
|
|
---
|
|
|
|
## When to Terminate Pod
|
|
|
|
**Terminate if**:
|
|
- ✅ Training completes successfully (100 epochs)
|
|
- ❌ Constant reward warnings persist after epoch 20
|
|
- ❌ Action diversity < 10% for any action after epoch 30
|
|
- ❌ Q-value divergence > 1000 after epoch 50
|
|
- ❌ Cost exceeds $1.00 (4 hours) without progress
|
|
- ❌ OOM errors or repeated crashes
|
|
|
|
**Keep running if**:
|
|
- Training is progressing normally
|
|
- Checkpoints saving every 10 epochs
|
|
- Reward variance > 0.1
|
|
- Action diversity 20-40% each
|
|
- Cost < $0.50
|
|
|
|
---
|
|
|
|
## References
|
|
|
|
- **Deployment Script**: `deploy_dqn_retrain.sh`
|
|
- **Validation Checklist**: `DQN_RETRAIN_VALIDATION_CHECKLIST.md`
|
|
- **Runpod Quick Ref**: `RUNPOD_DEPLOY_QUICK_REF.md`
|