The warp-parallel (32,1,1) change to c51_loss_reduce caused the same hang as isv_feature_gate: changing block_dim inside graph_mega alters CUDA Graph node structure on Hopper's TMA scheduler. Must stay (1,1,1). Also: batch_size 8192→16384, gpu_n_episodes 1024→4096, num_atoms 51→52 to align h100.toml with dqn-production.toml. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
15 lines
395 B
TOML
15 lines
395 B
TOML
# H100 PCIe/SXM (80GB VRAM) -- full production
|
||
[training]
|
||
batch_size = 16384
|
||
num_atoms = 52
|
||
buffer_size = 500000
|
||
hidden_dim_base = 256
|
||
replay_buffer_vram_fraction = 0.55 # state_dim=96, 86 weight tensors — leaves ~33GB for compute
|
||
|
||
[experience]
|
||
gpu_timesteps_per_episode = 5000
|
||
gpu_n_episodes = 4096 # 132 SMs × ~31 eps/SM — full H100 utilization
|
||
|
||
[cuda]
|
||
cuda_stack_bytes = 65536 # 64KB
|