Approved design for moving Prioritized Experience Replay entirely to GPU — flat priority array with parallel prefix-sum sampling, GPU-resident ring buffer for experiences, async loss readback. Eliminates the last major CPU bottleneck in the DQN training loop. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>