Parallel hyperopt trials, VRAM-aware batch bounds, 5-layer OOM protection. Scales from L4 through H100. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
5.5 KiB
GPU Utilization Optimization Design
Problem
DQN (~151K params, 1.2MB) and PPO (~70K params, 0.28MB) models run one-at-a-time on L4 24GB, using <4% of available VRAM. Three root causes:
- DQN batch_size bounds hardcoded [64, 160] — designed for RTX 3050 Ti 4GB, never updated
- Hyperopt trials strictly sequential —
Arc<Mutex<model>>serializes all PSO evaluations - No model-aware scaling — same strategy for 151K-param DQN and 15M-param TFT
Target Hardware
- L4 24GB (current CI/testing)
- L40S 48GB (planned)
- H100 80GB (planned)
Design
Three Levers
| Lever | Current | Target |
|---|---|---|
| Batch size bounds | DQN: [64, 160] hardcoded | VRAM-aware via HardwareBudget |
| Trial concurrency | 1 (Mutex) | N concurrent per model size / VRAM |
| Model-aware scaling | None | Auto: tiny models → many trials; large → big batches |
Component 1: HardwareBudget::hyperopt_strategy()
New method on existing HardwareBudget struct:
pub struct HyperoptStrategy {
pub max_concurrent_trials: usize,
pub batch_size_bounds: (f64, f64),
pub per_trial_memory_mb: f64,
pub total_reserved_mb: f64, // for VRAM watchdog
}
Formula:
- Per-trial VRAM =
model_overhead_mb * (1 + optimizer_mult + gradients + activations) + batch_overhead - Available VRAM =
total * (1 - safety_margin) max_concurrent = floor(available / per_trial_vram), capped at PSO swarm sizebatch_size_upper = (per_trial_vram - fixed_overhead) / bytes_per_sample
Examples:
- DQN on L4: per-trial ~50MB → 20 concurrent (capped at swarm size)
- TFT on L4: per-trial ~2GB → 9 concurrent
- DQN on H100: per-trial ~50MB → all particles at once
Component 2: Parallel PSO Particle Evaluation
Current (sequential):
for particle in swarm:
result = evaluate(particle)
update_personal_best(particle, result)
update_global_best()
New (chunked parallel):
chunks = swarm.chunks(max_concurrent_trials)
for chunk in chunks:
results = evaluate_parallel(chunk) // tokio::spawn per particle
for (particle, result) in zip(chunk, results):
update_personal_best(particle, result)
update_global_best()
Each parallel evaluation gets its own model instance on Device::Cuda(0).
Candle's cudarc backend is thread-safe — multiple threads submit kernels concurrently,
GPU hardware scheduler interleaves them.
Component 3: DQN continuous_bounds_for()
Copy the pattern PPO already uses — call budget.max_batch_size() to compute upper bound
dynamically instead of hardcoded 160. The PPO adapter already does this correctly.
What Stays the Same
- PSO algorithm logic (velocity/position updates unchanged)
- Model architectures (untouched)
- Training loop internals (untouched)
- Replay buffers (each trial gets its own, on CPU)
max_safe_batch_size()VRAM ceiling in trainers
Files to Modify
crates/ml/src/hyperopt/hardware_budget.rs— addhyperopt_strategy()methodcrates/ml/src/hyperopt/optimizer.rs— parallel particle evaluationcrates/ml/src/hyperopt/adapters/dqn.rs— VRAM-awarecontinuous_bounds_for()crates/ml/src/hyperopt/adapters/ppo.rs— useHyperoptStrategyfor concurrency
OOM Protection (5 layers)
Layer 1 — Conservative estimation: plan_hyperopt() uses 20% safety margin
PLUS per-trial overhead includes optimizer states (2x model for Adam), gradients
(1x model), and activation memory (1x model, 0.65x with checkpointing). Same
formula as max_safe_batch_size() but applied per-trial.
Layer 2 — Concurrency ramp-up: Don't start all trials at once. Start with
ceil(max_concurrent / 2), monitor first iteration's peak VRAM via cudarc
memory query, then scale up if headroom exists. If first batch OOMs, halve
concurrency and retry.
Layer 3 — Per-trial OOM catch: Each trial's training loop already has
reduce_batch_size() fallback. If a trial OOMs, it halves its batch_size and
retries (existing behavior). With parallel trials, one trial's OOM doesn't kill
others — Candle/cudarc allocations are independent.
Layer 4 — Global VRAM watchdog: Before spawning each new trial batch, query
cudarc::driver::CudaDevice::mem_info() for actual free VRAM. If free < 10% of
total, reduce concurrency for the next iteration. Log a warning.
Layer 5 — Graceful degradation: If max_concurrent computes to 1, fall back to current sequential behavior. Zero regression risk.
Scaling Profiles
| GPU | VRAM | DQN concurrent | PPO concurrent | TFT concurrent |
|---|---|---|---|---|
| RTX 3050 Ti | 4 GB | 3-4 | 2-3 | 1 |
| L4 | 24 GB | 20 (swarm cap) | 15-20 | 4-8 |
| L40S | 48 GB | 20 (swarm cap) | 20 (swarm cap) | 8-15 |
| H100 | 80 GB | 20 (swarm cap) | 20 (swarm cap) | 15-20 |
For tiny models (DQN/PPO), concurrency is capped by PSO swarm size (default 20), not VRAM. The GPU saturates compute via 20 concurrent forward/backward passes.
Risks
- OOM: 5-layer protection above. Worst case: graceful fallback to sequential.
- CUDA stream contention: cudarc uses default stream per thread. Multiple threads submit kernels concurrently. GPU hardware scheduler handles it.
- CPU memory: Each trial needs its own replay buffer (~10-50MB for DQN). 20 trials × 50MB = 1GB CPU RAM — acceptable on training nodes (128GB+).
- PSO convergence: Parallel evaluation doesn't change PSO math — same particles, same updates, just evaluated faster. No algorithmic change.
- Regression: When max_concurrent=1, behavior identical to current code.