Files
foxhunt/crates/ml/tests
jgrusewski a4f84a253e fix: IQN action bounds + VRAM-aware DQN pipeline tests (5/5 pass)
Root causes resolved:
- IQN decode_actions_kernel: clamp flat action to valid range before
  branch decomposition (same OOB pattern as dqn_forward_loss_kernel)
- DQN pipeline tests: missing batch_size override caused 1024 > 800
  experiences, making can_sample() return false

Test infrastructure:
- scale_for_gpu(): queries actual GPU VRAM and caps buffer/batch/atoms
  on <8GB GPUs. H100 (80GB) runs full conservative() defaults unmodified:
  51 atoms, 1024 batch, 500K buffer, 256 episodes.
- Add replay buffer debug tracing (buffer_len, can_sample, batch_size)
- Remove replay_buffer_vram_fraction=0.0 (GPU PER mandatory for fused training)

All 5 DQN pipeline tests pass on RTX 3050 4GB (310s total).

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-20 21:12:36 +01:00
..