# Real-Time Streaming - Implementation Roadmap **Date**: 2025-11-02 **Status**: Ready for Implementation **Estimated Total Effort**: 2-3 weeks (Priorities 1-3), 4-5 weeks (all priorities) --- ## Quick Reference ### Priority Summary | Priority | Task | Effort | Value | Status | |----------|------|--------|-------|--------| | **P1** | Fix Rust CLI nested paths | 2-4 hours | High ✅ | 🟡 Ready | | **P2** | Enhanced Python metrics + cost tracking | 1-2 days | Very High ✅ | 🟡 Ready | | **P3** | Alert system + auto-termination | 3-5 days | Medium ✅ | 🟡 Ready | | **P4** | Web dashboard (optional) | 1-2 weeks | Low ⚠️ | 🔴 Deferred | ### Expected ROI - **Priorities 1-2**: 60% of total value, 20% of total effort = **5-10x ROI** - **Priority 3**: 20% of total value (cost savings), 30% of total effort = **3x ROI** - **Priority 4**: 20% of total value, 50% of total effort = **1x ROI** (defer) --- ## Priority 1: Fix Rust CLI (IMMEDIATE - 2-4 Hours) ### Objective Enable `foxhunt-deploy monitor` to work with nested S3 directory structure. ### Current Issue **File**: `/home/jgrusewski/Work/foxhunt/foxhunt-deploy/src/s3/mod.rs:58` ```rust pub(crate) async fn list_log_files(&self, pod_id: &str) -> Result> { let prefix = format!("ml_training/{}/", pod_id); // ❌ Only searches one level deep } ``` ### Solution: Add `--run-id` Parameter #### Step 1: Update CLI Arguments (15 min) **File**: `/home/jgrusewski/Work/foxhunt/foxhunt-deploy/src/cli/monitor.rs` ```rust #[derive(Args, Debug)] pub(crate) struct MonitorArgs { /// Pod ID or Run ID to monitor #[arg(required = true)] pub id: String, /// Treat ID as run-id instead of pod-id #[arg(long)] pub run_id: bool, /// Follow logs in real-time #[arg(short, long)] pub follow: bool, /// Number of recent lines to show #[arg(short, long)] pub tail: Option, /// Filter logs by pattern (regex) #[arg(long)] pub filter: Option, /// List available log files without displaying content #[arg(short, long)] pub list: bool, } ``` #### Step 2: Add S3 Search Function (60 min) **File**: `/home/jgrusewski/Work/foxhunt/foxhunt-deploy/src/s3/mod.rs` Add after existing `list_log_files` function: ```rust /// Find log file by run ID (searches nested structure) pub(crate) async fn find_log_by_run_id(&self, run_id: &str) -> Result> { let model_types = ["mamba2", "dqn", "ppo", "tft"]; // Search pattern: ml_training/*/training_runs/{model}/{run_id}/logs/training.log for model in &model_types { // List all outer directories under ml_training/ let response = self .client .list_objects_v2() .bucket(&self.bucket) .prefix("ml_training/") .delimiter("/") .send() .await .map_err(|e| FoxhuntError::S3(format!("Failed to list ml_training: {}", e)))?; // Check each outer directory for prefix_obj in response.common_prefixes() { let outer_dir = prefix_obj.prefix().unwrap_or(""); // Construct expected log path let log_key = format!( "{}training_runs/{}/{}/logs/training.log", outer_dir, model, run_id ); // Check if this log file exists if self.object_exists(&log_key).await { return Ok(Some(log_key)); } } } Ok(None) } /// Check if an S3 object exists async fn object_exists(&self, key: &str) -> bool { self.client .head_object() .bucket(&self.bucket) .key(key) .send() .await .is_ok() } ``` #### Step 3: Update Monitor Logic (30 min) **File**: `/home/jgrusewski/Work/foxhunt/foxhunt-deploy/src/cli/monitor.rs` Replace `execute` function: ```rust pub(crate) async fn execute(config: &FoxhuntConfig, args: &MonitorArgs) -> Result<()> { let s3_client = S3LogClient::new(&config.s3).await?; // Determine log file based on --run-id flag let log_file = if args.run_id { // Search by run ID (handles nested paths) s3_client .find_log_by_run_id(&args.id) .await? .ok_or_else(|| FoxhuntError::S3(format!("No logs found for run_id: {}", args.id)))? } else { // Search by pod ID (original logic) let monitor = LogMonitor::new( s3_client.clone(), args.id.clone(), config.s3.poll_interval_secs, ); monitor .find_log_file() .await? .ok_or_else(|| FoxhuntError::S3(format!("No logs found for pod_id: {}", args.id)))? }; // Handle list mode if args.list { println!("Found log file: {}", log_file); return Ok(()); } // Create monitor for streaming let mut monitor = LogMonitor::new( s3_client, args.id.clone(), config.s3.poll_interval_secs, ); // Override log file (since we already found it) // NOTE: This requires adding a `set_log_file()` method to LogMonitor // Stream logs if args.follow { monitor.tail_logs(args.tail, args.filter.clone()).await?; } else { monitor.show_recent_logs(args.tail).await?; } Ok(()) } ``` #### Step 4: Build and Test (30 min) ```bash # Build cd foxhunt-deploy cargo build --release # Test with completed DQN run cd .. ./target/release/foxhunt-deploy monitor run_20251102_210818_hyperopt --run-id --tail 50 # Expected output: # Found log file: ml_training/dqn_hyperopt_optimized_20251102_220747/training_runs/dqn/run_20251102_210818_hyperopt/logs/training.log # [last 50 lines of logs] # Test follow mode ./target/release/foxhunt-deploy monitor run_20251102_210818_hyperopt --run-id --follow # Test list mode ./target/release/foxhunt-deploy monitor run_20251102_210818_hyperopt --run-id --list ``` ### Checklist - [ ] Update MonitorArgs struct with `run_id` boolean flag - [ ] Add `find_log_by_run_id()` function to S3LogClient - [ ] Add `object_exists()` helper function - [ ] Update `execute()` function to handle both modes - [ ] Build release binary - [ ] Test with completed run (--tail 50) - [ ] Test follow mode (--follow) - [ ] Test list mode (--list) - [ ] Update documentation - [ ] Commit changes ### Success Criteria - ✅ `foxhunt-deploy monitor --run-id` finds logs correctly - ✅ Backward compatibility maintained (pod-id mode still works) - ✅ Zero regression in existing functionality - ✅ Tests pass with real S3 data ### Estimated Time **Total**: 2-4 hours --- ## Priority 2: Enhanced Python Metrics (SHORT-TERM - 1-2 Days) ### Objective Transform `scripts/monitor_logs.py` into a feature-rich monitoring dashboard with live metrics, cost tracking, and terminal UI. ### Phase 1: Metrics Extraction (4-6 hours) #### Step 1: Create TrainingMetrics Class **File**: `scripts/monitor_logs.py` (add after imports) ```python @dataclass class TrainingMetrics: """Structured training metrics parsed from logs.""" timestamp: datetime epoch: Optional[int] = None total_epochs: Optional[int] = None loss: Optional[float] = None policy_loss: Optional[float] = None value_loss: Optional[float] = None q_buy: Optional[float] = None q_sell: Optional[float] = None q_hold: Optional[float] = None episode_reward: Optional[float] = None learning_rate: Optional[float] = None trial_number: Optional[int] = None @classmethod def parse_from_line(cls, line: str, model_type: str) -> Optional['TrainingMetrics']: """Parse metrics from log line based on model type.""" metrics = cls(timestamp=datetime.now()) if model_type == 'dqn': # Epoch parsing epoch_match = re.search(r'Epoch\s+(\d+)/(\d+)', line) if epoch_match: metrics.epoch = int(epoch_match.group(1)) metrics.total_epochs = int(epoch_match.group(2)) # Loss parsing loss_match = re.search(r'Loss:\s+([\d.]+)', line) if loss_match: metrics.loss = float(loss_match.group(1)) # Q-values parsing q_buy_match = re.search(r'Q\(buy\):\s+([-\d.]+)', line) if q_buy_match: metrics.q_buy = float(q_buy_match.group(1)) q_sell_match = re.search(r'Q\(sell\):\s+([-\d.]+)', line) if q_sell_match: metrics.q_sell = float(q_sell_match.group(1)) q_hold_match = re.search(r'Q\(hold\):\s+([-\d.]+)', line) if q_hold_match: metrics.q_hold = float(q_hold_match.group(1)) # Reward parsing reward_match = re.search(r'Reward:\s+([-\d.]+)', line) if reward_match: metrics.episode_reward = float(reward_match.group(1)) elif model_type == 'ppo': # PPO-specific parsing epoch_match = re.search(r'Epoch\s+(\d+)', line) if epoch_match: metrics.epoch = int(epoch_match.group(1)) policy_loss_match = re.search(r'Policy Loss:\s+([\d.]+)', line) if policy_loss_match: metrics.policy_loss = float(policy_loss_match.group(1)) value_loss_match = re.search(r'Value Loss:\s+([\d.]+)', line) if value_loss_match: metrics.value_loss = float(value_loss_match.group(1)) # Return only if we found at least one metric if any([metrics.epoch, metrics.loss, metrics.q_buy, metrics.policy_loss]): return metrics return None ``` #### Checklist - [ ] Create TrainingMetrics dataclass - [ ] Implement DQN parsing (epoch, loss, Q-values, reward) - [ ] Implement PPO parsing (epoch, policy_loss, value_loss) - [ ] Implement TFT parsing (epoch, loss, accuracy) - [ ] Implement MAMBA2 parsing (epoch, loss) - [ ] Add unit tests for each parser - [ ] Test with real log files ### Phase 2: Cost Tracking (2-3 hours) #### Step 1: Create CostTracker Class **File**: `scripts/monitor_logs.py` (add after TrainingMetrics) ```python class CostTracker: """Real-time GPU cost tracking.""" def __init__(self, pod_cost_per_hour: float, start_time: datetime): self.pod_cost_per_hour = pod_cost_per_hour self.start_time = start_time def get_elapsed_time(self) -> timedelta: return datetime.now() - self.start_time def get_current_cost(self) -> float: elapsed_hours = self.get_elapsed_time().total_seconds() / 3600 return self.pod_cost_per_hour * elapsed_hours def estimate_total_cost(self, trials_completed: int, total_trials: int) -> tuple[float, timedelta]: if trials_completed == 0: return 0.0, timedelta(0) elapsed = self.get_elapsed_time() progress = trials_completed / total_trials estimated_total_time = elapsed / progress estimated_remaining = estimated_total_time - elapsed total_hours = estimated_total_time.total_seconds() / 3600 estimated_total_cost = self.pod_cost_per_hour * total_hours return estimated_total_cost, estimated_remaining def format_summary(self, trials_completed: int = 0, total_trials: int = 0) -> str: current_cost = self.get_current_cost() elapsed = self.get_elapsed_time() summary = f"💰 Current Cost: ${current_cost:.4f} | ⏱️ Elapsed: {self._format_timedelta(elapsed)}" if trials_completed > 0 and total_trials > 0: est_cost, est_remaining = self.estimate_total_cost(trials_completed, total_trials) summary += f"\n Est. Total: ${est_cost:.4f} | ETA: {self._format_timedelta(est_remaining)}" return summary @staticmethod def _format_timedelta(td: timedelta) -> str: total_seconds = int(td.total_seconds()) hours, remainder = divmod(total_seconds, 3600) minutes, seconds = divmod(remainder, 60) if hours > 0: return f"{hours}h {minutes}m" elif minutes > 0: return f"{minutes}m {seconds}s" else: return f"{seconds}s" ``` #### Checklist - [ ] Create CostTracker class - [ ] Implement elapsed time calculation - [ ] Implement current cost calculation - [ ] Implement total cost estimation (based on trial progress) - [ ] Implement ETA calculation - [ ] Add formatted output - [ ] Test with mock data - [ ] Integrate with monitor script ### Phase 3: Terminal UI Dashboard (4-6 hours) #### Step 1: Create TrainingDashboard Class **File**: `scripts/monitor_logs.py` (add after CostTracker) See full implementation in `REALTIME_STREAMING_DESIGN.md` (too long to repeat here). #### Step 2: Integrate with Existing Monitor **File**: `scripts/monitor_logs.py` Update `stream_run_logs()` function to use dashboard: ```python def stream_run_logs( s3_client: S3Client, run_id: str, follow: bool = True, timeout: Optional[int] = None, poll_interval: int = 5 ) -> None: """Stream logs for a specific run with live dashboard.""" # Find the run's log path for model_type in ['mamba2', 'dqn', 'ppo', 'tft']: log_key = f"ml_training/training_runs/{model_type}/{run_id}/logs/training.log" if s3_client.object_exists(log_key): break else: console.print(f"[red]Run not found: {run_id}[/red]") return # Initialize components start_time = datetime.now() cost_tracker = CostTracker(0.25, start_time) # RTX A4000 default dashboard = TrainingDashboard(run_id, model_type, cost_tracker) # Stream with dashboard log_position = 0 with Live(dashboard.render(), refresh_per_second=2) as live: while True: # Tail new content content, log_position = s3_client.tail_log_file(log_key, start_byte=log_position) if content: text = content.decode('utf-8', errors='ignore') for line in text.splitlines(): # Parse metrics metrics = TrainingMetrics.parse_from_line(line, model_type) if metrics: dashboard.add_metrics(metrics) # Check completion if detect_completion(line): return # Update trials count trials_key = f"ml_training/training_runs/{model_type}/{run_id}/hyperopt/trials.json" try: trials_data = json.loads(s3_client.download_log(trials_key)) dashboard.trials_completed = len(trials_data) except: pass # Refresh dashboard live.update(dashboard.render()) if not follow: break time.sleep(poll_interval) ``` #### Checklist - [ ] Create TrainingDashboard class - [ ] Implement metrics table rendering (model-specific columns) - [ ] Implement header/footer panels - [ ] Integrate with CostTracker - [ ] Update stream_run_logs to use dashboard - [ ] Test with live run - [ ] Test with completed run - [ ] Add --no-dashboard flag for raw logs ### Estimated Time **Total**: 12-18 hours (1.5-2 days) --- ## Priority 3: Alert System (MEDIUM-TERM - 3-5 Days) ### Phase 1: Alert Manager (1-2 days) #### Step 1: Create Alert Classes **File**: `scripts/monitor_logs.py` (new module or separate file) See full implementation in `REALTIME_STREAMING_DESIGN.md`. #### Checklist - [ ] Create AlertSeverity enum - [ ] Create Alert dataclass with Discord formatting - [ ] Create AlertManager class - [ ] Implement error pattern detection - [ ] Implement OOM plateau detection - [ ] Implement cost overrun detection - [ ] Add Discord webhook integration - [ ] Add Slack webhook integration (optional) - [ ] Test with mock alerts ### Phase 2: Auto-Termination (1 day) #### Step 1: Create AutoTerminator Class **File**: `scripts/monitor_logs.py` See full implementation in `REALTIME_STREAMING_DESIGN.md`. #### Checklist - [ ] Create AutoTerminator class - [ ] Implement termination logic with confirmation - [ ] Add dry-run mode - [ ] Integrate with RunPodClient - [ ] Add safety checks (no force termination) - [ ] Test with test pod ### Phase 3: Integration (1 day) #### Step 1: Update Monitor Script **File**: `scripts/monitor_logs.py` Add command-line arguments: ```python parser.add_argument( '--alert-webhook', help='Discord/Slack webhook URL for alerts' ) parser.add_argument( '--auto-terminate', action='store_true', help='Automatically terminate pod on completion (requires confirmation)' ) parser.add_argument( '--cost-budget', type=float, default=1.0, help='Cost budget in USD (alert if exceeded)' ) ``` #### Checklist - [ ] Add CLI arguments for alerts and auto-termination - [ ] Integrate AlertManager with monitoring loop - [ ] Integrate AutoTerminator with completion detection - [ ] Test end-to-end with real pod - [ ] Document usage in README ### Estimated Time **Total**: 4-5 days --- ## Priority 4: Web Dashboard (LONG-TERM - 1-2 Weeks, OPTIONAL) ### Status **DEFERRED** - Only implement if Priorities 1-3 are highly successful and there's strong user demand. ### Estimated Time **Total**: 7-9 days (1-2 weeks) See `REALTIME_STREAMING_DESIGN.md` for full design. --- ## Testing Strategy ### Unit Tests ```bash # Python tests cd scripts python -m pytest test_monitor_logs.py -v # Rust tests cd foxhunt-deploy cargo test ``` ### Integration Tests ```bash # Test with completed run (no follow) python3 scripts/monitor_logs.py --run-id run_20251102_210818_hyperopt # Test with active run (follow mode) python3 scripts/monitor_logs.py --run-id --follow # Test Rust CLI ./target/release/foxhunt-deploy monitor run_20251102_210818_hyperopt --run-id --tail 50 ``` ### End-to-End Tests ```bash # Deploy test pod python3 scripts/python/runpod/runpod_deploy.py --gpu-type "RTX A4000" --image "jgrusewski/foxhunt:latest" --command "hyperopt_dqn_demo --parquet-file /runpod-volume/test_data/ES_FUT_180d.parquet --trials 5 --epochs 10 --base-dir /runpod-volume/ml_training/test_run" # Monitor with Python (extract run_id from deployment output) python3 scripts/monitor_logs.py --run-id --follow --alert-webhook --auto-terminate --cost-budget 0.10 # Monitor with Rust CLI ./target/release/foxhunt-deploy monitor --run-id --follow ``` --- ## Success Metrics ### Priority 1 Success - ✅ Rust CLI works with nested S3 paths - ✅ No regressions in existing functionality - ✅ Tests pass with real data ### Priority 2 Success - ✅ Metrics extracted correctly (90%+ accuracy) - ✅ Cost tracking within 5% accuracy - ✅ ETA within 10% accuracy - ✅ Terminal UI renders smoothly ### Priority 3 Success - ✅ Alerts trigger within 10 seconds - ✅ Auto-termination saves >20% costs - ✅ Zero false positives (no accidental terminations) --- ## Rollout Plan ### Week 1: Quick Wins (Priorities 1-2) **Monday**: - Implement Priority 1 (Rust CLI fix) - Test and commit **Tuesday-Wednesday**: - Implement metrics extraction - Implement cost tracking **Thursday-Friday**: - Implement Terminal UI dashboard - Integration testing ### Week 2: Cost Optimization (Priority 3) **Monday-Wednesday**: - Implement alert system - Implement Discord/Slack integration **Thursday-Friday**: - Implement auto-termination - End-to-end testing ### Week 3 (Optional): Web Dashboard (Priority 4) **Only proceed if Priorities 1-3 are successful and there's user demand.** --- ## Maintenance Plan ### Post-Launch Monitoring - Monitor for bugs (GitHub issues) - Collect user feedback - Track cost savings (auto-termination) ### Future Enhancements - Multi-pod monitoring (parallel runs) - Historical metrics database - Advanced cost analytics - Email alerts (SMTP) - Custom alert patterns (user-defined) --- ## Conclusion This roadmap provides a clear path from the current state to a feature-rich monitoring system with: 1. **Immediate fix** (Priority 1): 2-4 hours 2. **High-value enhancements** (Priority 2): 1-2 days 3. **Cost optimization** (Priority 3): 3-5 days 4. **Optional web dashboard** (Priority 4): 1-2 weeks (defer) **Total effort for core value (P1-P3)**: 2-3 weeks **Expected ROI**: 5-10x improvement in monitoring capabilities with 20-50% reduction in GPU costs. **Next Steps**: Begin implementation with Priority 1 (Rust CLI fix).