BUG #41 kept forward pass in F32 for autograd, but the target-side tensors (reward, gamma, done, next_q) were cast to BF16 via `dtype`. The `state_action_values.sub(&target_q_values)` then hit F32-vs-BF16 mismatch on Ampere+ GPUs, causing every training step to fail silently. Fix: `.to_dtype(state_action_values.dtype())` on the detached target. Safe because target is detached (no autograd graph to break). Also: H100 runner → SXM2 pool, GPU availability checker script. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
6.3 KiB
Executable File
6.3 KiB
Executable File