Files
foxhunt/crates
jgrusewski 4a3f6a7195 fix: action bounds check in fused DQN kernel prevents CUDA context poisoning
Root cause: dqn_forward_loss_kernel accessed online_log_probs[] out of bounds
when actions buffer contained uninitialized values. The factored action
decomposition (action/9, action%9/3, action%3) produced branch indices > array
size, causing Invalid __local__ read at the current_lp pointer dereference.

Found via compute-sanitizer with -lineinfo: crash at line 612 (current_lp[j]
read with a_d * NUM_ATOMS overflow). All threads in all blocks crashed at the
same offset +0x28d50, confirming systematic OOB rather than race condition.

Fixes:
- Clamp factored_action to [0, BRANCH_0*BRANCH_1*BRANCH_2) in both
  forward_loss and backward kernels (safe default: action 0 = hold)
- Use generic branch decomposition (BRANCH_x_SIZE) instead of hardcoded /9 /3
- Disable CUDA Graph on consumer GPUs (< 164KB shmem opt-in)
- Query CU_DEVICE_ATTRIBUTE_MAX_SHARED_MEMORY_PER_BLOCK_OPTIN for shmem limit
- Share query_max_shmem_bytes() between trainer and experience collector

This eliminates CUDA context poisoning between walk-forward folds — fold 1+
can now create new CudaStreams successfully.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-20 17:50:32 +01:00
..