Threads b_logits_dir / per_sample_support / atom_positions / n_atoms
through experience_action_select kernel launch in
gpu_backtest_evaluator.rs's chunked val pipeline (line ~1722, inside
submit_dqn_step_loop_cublas).
The evaluator sources Q-values via the QValueProvider trait
(delegates forward to the trainer's CUDA-Graphed cuBLAS), so this
change extends the trait surface rather than duplicating the
forward:
- New trait method compute_q_and_b_logits_to(states_ptr, batch,
q_out_ptr, b_logits_out_ptr) — DtoD-copies trainer's
on_b_logits_buf into caller's chunked buffer per sub-batch
iteration alongside the existing q_out copy.
- New trait accessors per_sample_support_ptr / atom_positions_ptr /
num_atoms / total_branch_atoms (stable trainer-owned buffers).
- FusedTrainingCtx implements all of the above; trainer gains
on_b_logits_buf_ptr / atom_positions_buf_ptr /
per_sample_support_ptr_get pub accessors.
- Evaluator gains chunked_b_logits_buf field (sized
[n_windows * CHUNK_SIZE, total_actions * num_atoms] + 32*3 tail
safety), allocated in ensure_action_select_ready.
Phase 6's last-step plan_params forward keeps using the original
compute_q_values_to (b_logits not consumed there).
Plan C Task 4 — evaluator companion to T3 collector wire-up. After
this commit, all production callers of experience_action_select use
the amended kernel ABI (T2 amendment 5de5e546a) end-to-end
(feedback_no_partial_refactor).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>