Add device_pool to DQN/PPO hyperopt trainers for round-robin GPU assignment per trial. Binary detects all CUDA devices, scales VRAM budget by GPU count. Single-GPU: no behavior change (pool of 1). Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>