Pearl 5's online_taus/target_taus/cos_features were declared as
CudaSlice<f32> (device-only), populated via upload_f32_via_pinned
which does a DtoD copy from a separate mapped-pinned staging buffer.
The DtoD inside CUDA Graph capture triggers
CUDA_ERROR_STREAM_CAPTURE_INVALIDATED and the 'continuing ungraphed'
fallback observed in smoke-test-hhr5q.
This violates feedback_no_htod_htoh_only_mapped_pinned: the rule is
mapped-pinned (cuMemHostAlloc DEVICEMAP) for ALL CPU↔GPU paths. No
DtoD copies, no HtoD copies, no exceptions.
Fix: convert all 3 buffers (online_taus, target_taus, cos_features)
to MappedF32Buffer per-branch [MappedF32Buffer; 4] arrays. Host writes
go directly to host_ptr; IQN kernel reads dev_ptr of the same memory
— no copy step at all. The mem::swap pattern is replaced with pure
selection: activate_branch_taus sets active_branch_idx; kernel launch
sites index online_taus_per_branch[active_branch_idx].dev_ptr.
Eliminates upload_f32_via_pinned calls for these buffers entirely.
Refresh becomes a host write to mapped-pinned host_ptr at fold
boundary; subsequent kernel launches see the write through the
mapped-pinned coherence guarantee after stream sync.
cargo check + cargo build --release + cargo test --lib (sp4 sp5
state_reset_registry: 13/13) all clean. Sanity grep for
upload_f32_via_pinned in gpu_iqn_head.rs returns zero.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>