perf: disable replay buffer VRAM auto-sizing — cap at 500K entries
Auto-sizer inflated replay buffer to 15.8M entries on H100 (45% of 80GB VRAM). The PER prefix scan inside the CUDA graph processes ALL 15.8M entries every step, dominating step time. With buffer_size=500K, the scan processes 500K entries instead. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -4,7 +4,7 @@ batch_size = 8192
|
||||
num_atoms = 52
|
||||
buffer_size = 500000
|
||||
hidden_dim_base = 256
|
||||
replay_buffer_vram_fraction = 0.45 # state_dim=96, 86 tensors + IQN 8.6GB + attention 250MB + Phase 2 128MB
|
||||
replay_buffer_vram_fraction = 0.0 # disabled — use exact buffer_size=500K (was 0.45 → 15.8M entries = slow PER prefix scan in graph)
|
||||
|
||||
[experience]
|
||||
gpu_timesteps_per_episode = 5000
|
||||
|
||||
Reference in New Issue
Block a user