Files
foxhunt/crates
jgrusewski 542be32bed perf(ml): pad DQN input dims for H100 tensor core alignment
cuBLAS requires M/N/K dimensions to be multiples of 8 for BF16 tensor
core dispatch. state_dim=51 (with OFI) caused silent fallback to scalar
FMA ops, leaving tensor cores at 0.1% utilization despite BF16 enabled.

Pad state_dim to next 8-multiple (51→56, 43→48) at network construction
and zero-pad input tensors in forward pass across all DQN network types:
Sequential, DuelingQNetwork, DistributionalDuelingQNetwork, NetworkLayers.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-08 02:10:38 +01:00
..