After Fix 1..16 migrated all 80+ production callers off `super::htod_f32` and `super::clone_htod_f32`, the helper bodies in `cuda_pipeline/mod.rs:129-145` had zero non-test consumers. Deleted both function definitions per `feedback_no_legacy_aliases.md` (no deprecated wrappers). Per `feedback_no_partial_refactor.md` (when a shared contract is deleted, every consumer migrates together — including tests), the two surviving test-block callers in `gpu_tlob.rs::tests` (lines 1017 and 1132) are migrated to `mapped_pinned::upload_f32_via_pinned` in the same commit. The other test-only callers in `signal_adapter.rs::tests`, `gpu_action_selector.rs::tests`, and `cuda_pipeline/mod.rs::tests` use bare `stream.memcpy_htod` / `stream.memcpy_stod` against the cudarc handle directly (not the deleted helpers) — no change needed. A docstring was added at the deletion site recording when and why the helpers were removed, pointing future readers at the canonical replacements `mapped_pinned::clone_to_device_f32_via_pinned` and `mapped_pinned::upload_f32_via_pinned`. Final state of the HtoD migration sequence: - production callers of `stream.memcpy_htod` / `memcpy_stod`: 0 - production callers of `htod_f32` / `clone_htod_f32`: 0 - helper definitions: removed from `mod.rs` docs/dqn-gpu-hot-path-audit.md updated with Fix 17 entry. cargo check -p ml --lib clean at 12 warnings. cargo check -p ml --tests clean at 23 warnings (12 lib duplicates + 11 test-specific, baseline unchanged). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
32 KiB
DQN GPU Hot-Path Audit
Invariant 3 enforcement. Every cross-boundary call (DtoH, HtoD, memcpy between host and device) in the DQN code path is classified below. No hot-path memcpy DtoH/HtoD is allowed — zero tolerance. Pinned + device-mapped memory is the ONLY permitted mechanism for per-step host-visible data.
Hot path definition: every function called from fused_training.rs::run_full_step or its descendants that executes once per training step (up to 178 times per epoch at batch size 8192), including kernel invocations inside the captured CUDA graph (adam_grad_child, forward_child, aux_grad_child, aux_child, maintenance_child) and host-side orchestration between them.
Classification:
OK-pinned— uses the zero-copy pinned + device-mapped pattern (cuMemAllocHost_v2+cuMemHostGetDevicePointer_v2); allowed on hot path.MIGRATE— was a plainmemcpybut has been converted to pinned in this commit.COLD-PATH— acceptable placement (constructor init, checkpoint save/load, fold transition, epoch-boundary); annotated with reason.
| File:line | Call signature | Path class | Annotation | Status |
|---|---|---|---|---|
gpu_dqn_trainer.rs:6586 |
malloc_host → total_loss_pinned |
OK-pinned | Constructor: cuMemAllocHost+GetDevicePointer; GPU writes via dev ptr, CPU reads via readback_training_scalars() |
OK |
gpu_dqn_trainer.rs:6598 |
malloc_host → mse_loss_pinned |
OK-pinned | Constructor: same pattern as total_loss_pinned |
OK |
gpu_dqn_trainer.rs:6611 |
malloc_host → q_divergence_pinned |
OK-pinned | Constructor: same pattern | OK |
gpu_dqn_trainer.rs:6624 |
malloc_host → iqn_readiness_pinned |
OK-pinned | Constructor: same pattern | OK |
gpu_dqn_trainer.rs:6662 |
malloc_host → grad_norm_pinned |
OK-pinned | Constructor: same pattern | OK |
gpu_dqn_trainer.rs:6707 |
malloc_host (plain, not DEVICEMAP) → readback scratch |
COLD-PATH | Constructor: plain pinned for per-epoch flush_readback(); not written by GPU kernel directly |
OK |
gpu_dqn_trainer.rs:6788 |
malloc_host → monitoring summary pinned |
OK-pinned | Constructor: monitoring kernel writes via dev ptr each step | OK |
gpu_dqn_trainer.rs:6895 |
malloc_host → eval_v_range_pinned |
OK-pinned | Constructor: GPU writes; host reads only at fold/epoch boundary via eval_v_range() |
OK |
gpu_dqn_trainer.rs:6911 |
malloc_host → nan_flag_pinned |
OK-pinned | Constructor: GPU writes on NaN; host reads only in fault path (read_nan_flags) |
OK |
gpu_dqn_trainer.rs:6929 |
malloc_host → ISV isv_signals_pinned |
OK-pinned | Constructor: GPU writes ISV bus each step; host reads via read_isv_health_and_regime() in compute_adaptive_budgets |
OK |
gpu_dqn_trainer.rs:6949 |
malloc_host → lr_pinned |
OK-pinned | Constructor: CPU writes LR; GPU reads via dev ptr | OK |
gpu_dqn_trainer.rs:6968 |
malloc_host → adaptive_clip_pinned |
OK-pinned | Constructor: CPU writes clip; GPU reads via dev ptr | OK |
gpu_dqn_trainer.rs:6988 |
malloc_host → popart_pinned |
OK-pinned | Constructor: GPU writes PopArt stats; host reads at epoch boundary | OK |
gpu_dqn_trainer.rs:7099 |
malloc_host → q_stats_pinned |
OK-pinned | Constructor: GPU writes Q stats; host reads via reduce_current_q_stats() at epoch boundary |
OK |
gpu_dqn_trainer.rs:7127 |
malloc_host → q_readback_pinned |
OK-pinned | Constructor: GPU writes per-step Q readback; host reads via last_readback_scalars() in training guard |
OK |
gpu_dqn_trainer.rs:7206 |
cuMemAllocHost_v2 → var_ema pinned (ISV subsystem) |
OK-pinned | Constructor: ISV variance EMA — GPU writes each step; host reads via ISV bus | OK |
gpu_dqn_trainer.rs:7243 |
cuMemAllocHost_v2 → regime_util pinned (ISV subsystem) |
OK-pinned | Constructor: ISV regime utilization — GPU writes each step | OK |
gpu_dqn_trainer.rs:7289 |
malloc_host → popart_var_pinned |
OK-pinned | Constructor: GPU writes PopArt variance; host reads at epoch boundary | OK |
gpu_dqn_trainer.rs:7312 |
cuMemcpyHtoDAsync_v2 (constructor) |
COLD-PATH | Constructor: one-shot HtoD to seed deterministic-mask buffer; never replayed | OK |
gpu_dqn_trainer.rs:7878 |
cuMemAllocHost_v2 → q_mean_scratch |
OK-pinned | Constructor: GPU scratch for Q-mean reduction kernel | OK |
gpu_dqn_trainer.rs:7898 |
cuMemAllocHost_v2 → q_mean_ema |
OK-pinned | Constructor: Q-mean EMA — GPU writes, host reads at epoch boundary | OK |
gpu_dqn_trainer.rs:7919 |
cuMemAllocHost_v2 → td_error_scratch |
OK-pinned | Constructor: TD-error scratch for reduce kernel | OK |
gpu_dqn_trainer.rs:7939 |
cuMemAllocHost_v2 → q_var_scratch |
OK-pinned | Constructor: Q-variance scratch for reduce kernel | OK |
gpu_dqn_trainer.rs:2300 |
stream.memcpy_htod → atom_positions_buf |
COLD-PATH | set_atom_positions(): called at fold boundary to reload C51 atom grid — not on hot path |
OK |
gpu_dqn_trainer.rs:2606 |
stream.memcpy_dtoh → grad buf readback |
COLD-PATH | per_branch_grad_norms(): per-epoch call only; drains grad buffer for logging |
OK |
gpu_dqn_trainer.rs:2925 |
stream.memcpy_dtoh (online params) |
COLD-PATH | per_branch_target_drift(): per-epoch drift measurement |
OK |
gpu_dqn_trainer.rs:2928 |
stream.memcpy_dtoh (target params) |
COLD-PATH | per_branch_target_drift(): per-epoch drift measurement |
OK |
gpu_dqn_trainer.rs:2991 |
stream.memcpy_dtoh → weight inspection |
COLD-PATH | per_branch_vsn_mean(): per-epoch VSN mean read |
OK |
gpu_dqn_trainer.rs:3072 |
stream.memcpy_dtoh → Q-out inspection |
COLD-PATH | compute_q_stats(): per-epoch Q-stats; not called per-step |
OK |
gpu_dqn_trainer.rs:7541 |
stream.memcpy_htod → descriptor buffer |
COLD-PATH | Constructor: one-shot HtoD for GEMM descriptor upload | OK |
gpu_dqn_trainer.rs:7680 |
stream.memcpy_htod → stochastic depth scale |
COLD-PATH | Constructor: one-shot HtoD for stochastic depth kernel seed | OK |
gpu_dqn_trainer.rs:7686 |
stream.memcpy_htod → stochastic depth RNG |
COLD-PATH | Constructor: one-shot HtoD for stochastic depth RNG | OK |
gpu_dqn_trainer.rs (removed) |
cuMemcpyDtoHAsync_v2 in run_causal_intervention_unconditional |
MIGRATED | Fix 1: removed dead copy — result causal_mean_scratch was never consumed by any caller; result stays on GPU |
FIXED |
gpu_iqn_head.rs:469 |
stream.clone_htod → cos_features |
COLD-PATH | Constructor: precompute cosine embedding table; one-shot | OK |
gpu_iqn_head.rs:477 |
stream.clone_htod → online_taus |
COLD-PATH | Constructor: tau quantile tiling; one-shot | OK |
gpu_iqn_head.rs:480 |
stream.clone_htod → target_taus |
COLD-PATH | Constructor: tau quantile tiling; one-shot | OK |
gpu_iqn_head.rs:498 |
malloc_host → total_loss_pinned |
OK-pinned | Constructor: GPU writes IQN loss; host reads via read_total_loss() |
OK |
gpu_iqn_head.rs:534 |
malloc_host → t_pinned |
OK-pinned | Constructor: Adam step counter — CPU increments, GPU reads via dev ptr | OK |
gpu_iqn_head.rs:548 |
malloc_host → tau_pinned |
OK-pinned | Constructor (Fix 3): tau scalar — CPU writes, GPU reads via tau_dev_ptr; replaces CudaSlice<f32> + cuMemcpyHtoDAsync_v2 |
OK |
gpu_iqn_head.rs:807 |
cuMemcpyDtoDAsync_v2 (DtoD) |
OK-pinned | Device-to-device copy inside graph-captured target-H forward cache; no host crossing | OK |
gpu_iqn_head.rs:991 |
cuMemcpyDtoDAsync_v2 (DtoD) |
OK-pinned | Device-to-device copy for cached target states; no host crossing | OK |
gpu_iqn_head.rs:1564 |
stream.clone_dtoh → total_loss |
COLD-PATH | read_loss(): synchronous fallback for epoch-end logging only; read_total_loss() (pinned) is the hot-path equivalent |
OK |
gpu_iqn_head.rs:1602 |
cuMemcpyDtoH_v2 (sync) in compute_iqr |
MIGRATED | Fix 2A: compute_iqr() removed from submit_aux_ops; now called at epoch boundary via refresh_iqn_iqr() only; sync DtoH cannot be captured in CUDA graph — was silently a no-op on replayed steps |
FIXED |
gpu_iqn_head.rs:1622 |
cuMemcpyHtoDAsync_v2 in compute_iqr |
MIGRATED | Fix 2B: same relocation — async HtoD with stale IQR data was being replayed; now runs once at epoch boundary | FIXED |
gpu_iqn_head.rs (removed) |
cuMemcpyHtoDAsync_v2 in target_ema_update for tau_buf |
MIGRATED | Fix 3: replaced CudaSlice<f32> + async copy with pinned+device-mapped page; CPU writes *tau_pinned = tau; kernel reads via tau_dev_ptr — zero PCIe overhead |
FIXED |
gpu_iqn_head.rs:2058 |
stream.clone_htod → params_buf constructor |
COLD-PATH | init_iqn_head_from_weights(): one-shot weight upload at construction/load; not on hot path |
OK |
gpu_iqn_head.rs:2107 |
stream.clone_dtoh + stream.clone_htod in clone_cuda_slice |
COLD-PATH | clone_cuda_slice(): called during checkpoint save/load only |
OK |
fused_training.rs:476 |
stream.memcpy_htod → src_gpu (PER source indices) |
COLD-PATH | prepare_per_batch(): per-epoch PER batch preparation — not per-step |
OK |
fused_training.rs:482 |
stream.memcpy_htod → ones_gpu |
COLD-PATH | prepare_per_batch(): per-epoch importance weights init |
OK |
fused_training.rs:2914 |
stream.clone_dtoh → state inspection |
COLD-PATH | eval_state_debug(): debug helper, no hot-path caller |
OK |
fused_training.rs:2928 |
stream.clone_dtoh → param inspection |
COLD-PATH | eval_state_debug(): debug helper |
OK |
fused_training.rs:2955 |
cuMemcpyDtoH_v2 (sync) |
COLD-PATH | eval_v_range(): called only at fold boundary / epoch-0 init; explicit sync is documented and acceptable at boundary |
OK |
training_loop.rs:220 |
collector.stream().memcpy_dtoh → PER scores |
COLD-PATH | Pre-epoch PER priority init — runs once before the step loop starts | OK |
training_loop.rs:409 |
stream.memcpy_htod → episode_starts_buf |
COLD-PATH | Per-epoch experience setup — before run_training_steps_slices begins |
OK |
training_loop.rs:1003 |
clone_htod_f32 → targets device buffer |
COLD-PATH | Data loading: DBN targets uploaded before step loop | OK |
training_loop.rs:1007 |
clone_htod_f32 → features device buffer |
COLD-PATH | Data loading: feature matrix uploaded before step loop | OK |
tau_update_kernel.cu (Plan 1 Task 13) |
single-thread kernel: reads ISV[39,40,12]; writes ISV[42] | COLD-PATH | Per-epoch boundary call from training_loop.rs; 1-thread, 1-block; zero hot-path overhead |
OK |
epsilon_update_kernel.cu (Plan 1 Task 14) |
single-thread kernel: reads ISV[39,40,12]; writes ISV[41] | COLD-PATH | Per-epoch boundary call from training_loop.rs; 1-thread, 1-block; zero hot-path overhead |
OK |
gamma_update_kernel.cu (Plan 1 Task 10) |
single-thread kernel: reads ISV[12]; writes ISV[43] | COLD-PATH | Per-epoch boundary call from training_loop.rs; 1-thread, 1-block; ISV[43] read by fill_gamma_buf on hot path via read_isv_signal_at (pinned, zero-copy) |
OK |
kelly_cap_update_kernel.cu (Plan 1 Task 11) |
single-thread kernel: reads ISV[12] + portfolio_states[n_envs, PS_STRIDE]; aggregates mean half-Kelly × health-safety; writes ISV[44] | COLD-PATH | Per-epoch boundary call from training_loop.rs; 1-thread, 1-block; portfolio_states dev ptr passed from GpuExperienceCollector without CPU roundtrip |
OK |
atoms_update_kernel.cu (Plan 1 Task 9) |
4-block kernel (one per branch): reads ISV[23..31) + spacing_raw params[52..56); writes atom_positions_buf[4 * num_atoms] via softmax + cumsum | COLD-PATH | Per-epoch boundary + SGD step (every 50 steps); grid=(4,1,1), block=(256,1,1); replaces CPU loop that launched adaptive_atom_positions 4× separately | OK |
gpu_portfolio.rs:174 |
actions_buf upload (per simulate_batch_gpu) |
MIGRATED | Promoted actions_buf from CudaSlice<i32> to MappedI32Buffer(MAX_BATCH_SIZE) allocated once at constructor. Per-call host writes go through host_ptr; kernel arg switched to dev_ptr u64. Zero HtoD copy on the simulate hot path. |
FIXED |
gpu_portfolio.rs:122,296 |
portfolio_state_buf init/reset |
MIGRATED | env_step kernel writes back to portfolio_state_buf, so the destination must remain a CudaSlice<f32>. Init/reset now stage via a temporary MappedF32Buffer + memcpy_dtod_async (new helper upload_portfolio_state). |
FIXED |
decision_transformer.rs:913 |
scratch.adam_t upload per train_step_gpu |
MIGRATED | HOT path: 4-byte step counter was uploaded via stream.memcpy_htod every backward pass. Promoted DtScratch.adam_t from CudaSlice<i32>[1] to MappedI32Buffer(1). Per-step host write goes through host_ptr; Adam kernel reads via dev_ptr. |
FIXED |
decision_transformer.rs:478 |
DT params init upload (constructor) | MIGRATED | params is mutated by Adam, so destination must remain CudaSlice<f32>. Xavier-init upload now stages via MappedF32Buffer + memcpy_dtod_async through new helper upload_host_to_cuda_f32. |
FIXED |
decision_transformer.rs:501 |
DT wd_mask init (constructor) | MIGRATED | wd_mask is read-only by GPU. Promoted from CudaSlice<f32> to MappedF32Buffer. CPU writes 1.0s through host_ptr; kernel reads via dev_ptr. |
FIXED |
decision_transformer.rs:1105 |
trajectory build episode_starts | MIGRATED | One-time per pre-training data build (cold path). Replaced clone_htod with MappedI32Buffer + direct host_ptr write; kernel reads via dev_ptr. |
FIXED |
decision_transformer.rs:1147 |
trajectory build ep_rewards | MIGRATED | Same trajectory builder cold path. Replaced htod_f32 with MappedF32Buffer; CPU writes via host_ptr, return_to_go kernel reads via dev_ptr. |
FIXED |
gpu_experience_collector.rs:2992 |
hindsight bar_indices upload (per collect_experiences_gpu) |
MIGRATED | HOT path: alloc + memcpy_htod ran every collect when hindsight active. Added persistent bar_indices_pinned: MappedI32Buffer(alloc_episodes * alloc_timesteps * 2) allocated at constructor. Per-call host writes go through host_ptr; kernel arg switched to dev_ptr u64. |
FIXED |
gpu_experience_collector.rs:3186 |
feature_mask upload (per-collect at t==0) | MIGRATED | HOT path: was alloc + memcpy_htod every epoch. Promoted feature_mask_buf from Option<CudaSlice<f32>> to Option<MappedF32Buffer>. Reallocates only when mask size changes (still a cold-path event); CPU writes via host_ptr, state_gather reads via dev_ptr. |
FIXED |
gpu_experience_collector.rs:3097 |
episode_starts upload (per-collect) | MIGRATED | episode_starts_buf is GPU-mutated by domain_rand_episode_starts, so the destination must remain a CudaSlice. Stage via MappedI32Buffer + memcpy_dtod_async through upload_host_to_cuda_i32_via_pinned. |
FIXED |
gpu_experience_collector.rs:2009 |
upload_expert_actions (warm) |
MIGRATED | Promoted expert_actions_gpu from Option<CudaSlice<i32>> to Option<MappedI32Buffer>. Direct host_ptr write replaces the alloc + memcpy_htod pair. |
FIXED |
gpu_experience_collector.rs:3816 |
update_per_sample_support (per-epoch) |
MIGRATED | per_sample_support_buf is read-only by GPU; promoted from CudaSlice<f32> to MappedF32Buffer. Compute_expected_q + quantile_q_select now receive dev_ptr u64 kernel args. |
FIXED |
gpu_experience_collector.rs:1076,3848 |
portfolio_states init/reset (cold) | MIGRATED | portfolio_states is GPU-mutated by env_step. Init/reset stage via MappedF32Buffer + memcpy_dtod_async through upload_host_to_cuda_f32_via_pinned. |
FIXED |
gpu_experience_collector.rs:1093 |
epoch_state init (cold) | MIGRATED | epoch_state is GPU-mutated. Replaced clone_htod_f32 with explicit alloc_zeros + upload_host_to_cuda_f32_via_pinned. |
FIXED |
gpu_experience_collector.rs:1341 |
saboteur_base init (cold) | MIGRATED | saboteur_base is GPU-mutated by saboteur_select. Replaced memcpy_htod with upload_host_to_cuda_f32_via_pinned. |
FIXED |
gpu_experience_collector.rs:2784 |
trade_stats_buf zero (per-epoch) | MIGRATED | The htod_f32 of an all-zeros vector was needlessly going through PCIe. Replaced with stream.memset_zeros — fully GPU-side. |
FIXED |
Summary
| Class | Count |
|---|---|
| OK-pinned | 31 |
| COLD-PATH | 20 |
| MIGRATED (fixed in this commit) | 4 |
| Remaining MIGRATE (unfixed) | 0 |
All MIGRATE sites resolved. Zero hot-path memcpy calls remain in the DQN training loop.
Fixes Applied (this commit)
Fix 1 — Remove dead async DtoH in run_causal_intervention_unconditional
gpu_dqn_trainer.rs: cuMemcpyDtoHAsync_v2 from causal_mean_scratch_ptr to readback_pinned[10] was removed. The only reader (run_causal_intervention, which has no callers) was never invoked, making this copy dead code. The result now stays on the GPU device buffer where the kernel leaves it.
Fix 2 — Relocate compute_iqr() from hot path to epoch boundary
fused_training.rs / training_loop.rs: iqn.compute_iqr() was being called inside submit_aux_ops, which is captured in the CUDA graph (aux_child). The sync cuMemcpyDtoH_v2 inside compute_iqr cannot be captured in a CUDA graph — it ran during capture only and was silently skipped on every replayed step. The subsequent cuMemcpyHtoDAsync_v2 (uploading IQR back) was therefore replaying stale data. Moved to refresh_iqn_iqr() called at epoch boundary in process_epoch_boundary.
Fix 3 — Convert tau_buf from device VRAM to pinned+device-mapped
gpu_iqn_head.rs: tau_buf: CudaSlice<f32> + tau_host: f32 + cuMemcpyHtoDAsync_v2 replaced by tau_pinned: *mut f32 + tau_dev_ptr: u64 allocated via cuMemAllocHost_v2(DEVICEMAP) + cuMemHostGetDevicePointer_v2. CPU writes *tau_pinned = tau directly; the EMA kernel reads via tau_dev_ptr with no PCIe copy. Drop impl updated to free tau_pinned.
Fix 4 (perf) — Async-eval pipeline: split launch_metrics_and_download (2026-04-28)
gpu_backtest_evaluator.rs + trainer/metrics.rs + trainer/training_loop.rs: the prior merge 3c0d26292 (async-diag — eval on dedicated stream + spawn_blocking checkpoint, building on d9cb14f1b + 673b04a8d) wired the GPU side of cross-stream pipelining correctly (training stream cuda_stream, eval stream validation_stream, cuStreamWaitEvent barrier on train_done_event) but left a host-side stream.synchronize() inside launch_metrics_and_download that blocked the CPU thread for the FULL eval drain (~25-30s/epoch on L40S). The host thread is the same one that submits the next epoch's training kernels via run_full_step — so until eval drained, training submission was gated on it, and the dedicated validation_stream provided no actual wall-clock benefit.
Split into two record/consume halves at the buffer-readback layer:
launch_metrics_and_record_event(&mut self) -> Result<(), MLError>— submits the metrics kernel + DtoD-async copies ofactions_history_buf/intent_mag_buf/picked_action_history_bufinto mapped-pinned mirrors (actions_history_pinned: MappedI32Bufferallocated innew(), the other two lazy-allocated alongside their device buffers inensure_action_select_ready). Recordseval_done_event(lazy-created on first launch viastream.context().new_event(None), re-recorded each epoch viaevent.record(stream)). Returns immediately — no host wait.consume_metrics_after_event(&self) -> Result<Vec<WindowMetrics>, MLError>— callsevent.synchronize()(the SOLE host-side wait per epoch boundary), thenread_volatiles the mapped-pinnedmetrics_buffor the per-window metrics. After the event syncs, the four action-distribution helpers (read_eval_action_distribution_per_*,read_eval_intent_magnitude_distribution,read_chunked_actions_direction_distribution) read directly from the now-coherent mapped-pinned mirrors — nomemcpy_dtohper call (the prior implementation allocated a freshVec<i32>of sizen_windows * max_lenand did a synchronous DtoH each call, which violatesfeedback_no_htod_htoh_only_mapped_pinned.mdAND added 4 extra host-side stalls beyond the metrics one). All four mirrors land in the same event-sync barrier as the metrics readback.
Caller migration (feedback_no_partial_refactor.md):
evaluate_dqn_graphed(existing public API, used by hyperopt + smoke test) becomes a thin synchronous wrapper: launch_async → consume. Public ABI unchanged.- New
evaluate_dqn_graphed_async— launch-only sibling for the pipelined path. Returns(). evaluate/evaluate_ppo/evaluate_supervisedmigrated inline to launch+consume — no caller pipelines them, so no behavioural change.compute_validation_lossremoved; replaced bylaunch_validation_loss(returns()) +consume_validation_loss(returnsf64and emits theHEALTH_DIAG val [...]/val_dir_dist/val_picked_dir_distlines).training_loop.rs's validation reporting block now: epoch-start consumes any pending launch intocached_async_val_loss, epoch-end reports the cached value and submits a fresh launch withpending_val_loss = Some(0.0)as the "launched" sentinel. Bootstrap (epoch 0 / no validation_stream) does sync launch+consume. After the for-epoch loop terminates, a final consume drains the LAST epoch's launch so smoke tests / hyperopt that readlast_eval_direction_dist()see the most recent epoch's data, not the second-to-last.
Mathematical identity preserved: every consumed val_sharpe is bit-identical to what the prior compute_validation_loss would have produced for the same epoch — the only change is when the host parses it (one epoch later, which matches the existing cached_async_val_loss lag semantics; HEALTH_DIAG val [...] emit moves with the consume). Expected wall-clock saving: ~25-30s/epoch on L40S.
Out-of-scope (remaining host-side stream syncs in eval, not part of this fix): gpu_backtest_evaluator.rs lines ~1596 + ~1863 — periodic self.stream.synchronize() calls every 10 chunks inside submit_dqn_step_loop_cublas for kernel-error detection. They block the host CPU thread but only for the chunked rollout duration of ~5 calls × per-chunk-residency-time per epoch (each catches deferred CUDA errors within 10 chunks of origin). Smaller in aggregate than the launch_metrics_and_download final wait this fix removed; future work to convert to async error-polling (e.g., eval_done_event.is_complete() polling from training submissions). Not adjacent to the perf-fix scope.
Fix 5 — Add mapped-pinned staging helpers in mapped_pinned.rs
mapped_pinned.rs: added clone_to_device_{f32,i32}_via_pinned and upload_{f32,i32}_via_pinned helpers that allocate a temporary MappedXxxBuffer (cuMemHostAlloc DEVICEMAP), write the payload via host_ptr, then async DtoD into a target CudaSlice<T>. Used to migrate cold-path constructor uploads off explicit cudaMemcpy HtoD per feedback_no_htod_htoh_only_mapped_pinned.md while preserving struct field types (e.g. portfolio_buf, weight buffers consumed by Adam) where the kernel mutates the buffer every step and device-resident memory is correct.
Fix 6 — Migrate cuda_pipeline/mod.rs callers to mapped-pinned
cuda_pipeline/mod.rs: 7 in-file callers of clone_htod_f32 in DqnGpuData::upload_slices, DqnGpuData::upload_ofi, DqnGpuData::build_batch_states (portfolio), DqnGpuData::htod_copy_into, and PpoGpuData::upload rewritten to use mapped_pinned::clone_to_device_f32_via_pinned. The htod_f32 / clone_htod_f32 helper bodies at lines 129-145 remain in place because other crates are still in mid-migration; they will be deleted once all consumers migrate.
Fix 7 — Migrate gpu_walk_forward.rs to mapped-pinned (3 sites)
gpu_walk_forward.rs: 3 clone_htod_f32 calls in the walk-forward GPU data upload (features, targets, OFI) rewritten to mapped_pinned::clone_to_device_f32_via_pinned.
Fix 8 — Migrate PPO trainer + hyperopt adapters to mapped-pinned (6 sites)
trainers/ppo.rs (2 sites: features, targets in set_raw_market_data), hyperopt/adapters/ppo.rs (3 sites: features, targets in upload path; action_indices in backtest forward closure), hyperopt/adapters/mamba2.rs (1 site: gpu_to_stream). All clone_htod_f32 and bare stream.clone_htod calls rewritten to mapped_pinned::clone_to_device_{f32,i32}_via_pinned.
Fix 9 — Migrate weight-upload sites in attention/tlob/iql/weights (4 sites)
gpu_weights.rs (1 site: RmsNormWeightSet::ones constructor), gpu_attention.rs (1 site: attention params constructor), gpu_tlob.rs (1 site: TLOB params constructor — mod tests blocks intentionally untouched per scope), gpu_iql_trainer.rs (2 sites: per_sample_support seed, IQL params init). All htod_f32 calls rewritten to mapped_pinned::upload_f32_via_pinned; the gpu_weights.rs ones-init uses clone_to_device_f32_via_pinned since the destination is freshly allocated.
Fix 10 — Migrate gpu_iqn_head.rs constructors + DtoD-only target clone (5 sites)
gpu_iqn_head.rs: 4 stream.clone_htod calls (cos_features precompute, online_taus + target_taus tile broadcast, init_iqn_xavier_weights upload) rewritten to mapped_pinned::clone_to_device_f32_via_pinned. Additionally the helper clone_cuda_slice was replaced — it previously did device→host (clone_dtoh) → host→device (clone_htod) round-trip, double-violating both the no-DtoH and no-HtoD rules. Now uses clone_cuda_slice_f32 (single async DtoD), the only callsite being the target-network seed at constructor time (target_params = clone_cuda_slice(&online_params)).
Fix 11 — Migrate gpu_her.rs to mapped-pinned + add MappedU32Buffer (3 sites)
gpu_her.rs: 1 cold-path constructor stream.memcpy_htod (rng_states u32 init) and 2 warm-path memcpy_htod (source_indices, donor_indices in relabel_batch) rewritten to mapped_pinned::upload_{u32,i32}_via_pinned. The warm sites re-allocate pinned staging per relabel call — acceptable on this cold path (relabel runs once per epoch). New MappedU32Buffer type added to mapped_pinned.rs mirroring MappedI32Buffer; carries a manual Debug impl so the workspace warning count stays at 13.
Fix 12 — Migrate gpu_backtest_evaluator.rs to mapped-pinned (5 sites)
gpu_backtest_evaluator.rs: 5 HtoD sites in GpuBacktestEvaluator::new + reset_evaluation_state rewritten to mapped-pinned staging. Sites: prices_buf (f32), features_buf (f32 — was already going through clone_htod_f32 helper), window_lens_buf (i32), portfolio_buf init (f32), and the per-evaluation reset of portfolio_buf. The fields stay typed CudaSlice<T> because the kernel mutates portfolio_buf each step (must remain device-resident); the read-only fields could become MappedXxxBuffer but keeping them as CudaSlice preserves the kernel-launch contract .arg(&self.field) already used at the many launch sites in the file.
Fix 13 — Migrate gpu_ppo_collector.rs to mapped-pinned (6 sites + persistent staging)
gpu_ppo_collector.rs: all 6 HtoD sites migrated.
- Constructor (cold):
portfolio_statesinit (htod_f32) →upload_f32_via_pinned;rng_statesinit (memcpy_htod) →upload_u32_via_pinned. collect_experiences(warm but per-collect):episode_starts_buf(memcpy_htod) →upload_i32_via_pinned;barrier_config(htod_f32) →upload_f32_via_pinned.reset_episodes(warm, per-epoch): replaced theVec<f32>/Vec<u32>host staging fields with persistentMappedF32Buffer/MappedU32Bufferallocated once at construction. Resets fill viahost_slice_mut()(zero-allocation, zero-HtoD); a single async DtoD seeds the persistent device-residentportfolio_states/rng_states(which the kernel mutates each step). Addshost_slice_mut()accessor on both mapped-pinned types.
Fix 17 — Delete orphan htod_f32 / clone_htod_f32 helpers from cuda_pipeline/mod.rs (2026-04-28)
After Fix 1..16 migrated all 80+ production callers off super::htod_f32 and super::clone_htod_f32, the helper bodies in cuda_pipeline/mod.rs:129-145 were unused outside #[cfg(test)] modules. Deleted both function definitions per feedback_no_legacy_aliases.md (no deprecated wrappers).
The two surviving test-block callers in gpu_tlob.rs::tests (lines 1017 and 1132 — super::super::htod_f32(&stream, &host_states, …) and super::super::htod_f32(&stream, &host_d_concat, …)) were also migrated to mapped_pinned::upload_f32_via_pinned in the same commit per feedback_no_partial_refactor.md (when a contract is deleted, every consumer migrates together — including tests). The test-only callers in signal_adapter.rs::tests (5 sites), gpu_action_selector.rs::tests (1), and cuda_pipeline/mod.rs::tests (1) all use bare stream.memcpy_htod / stream.memcpy_stod against the cudarc handle directly, not the deleted helpers — no change needed.
A doc-block was added at the deletion site in mod.rs recording when and why the helpers were removed and pointing future readers at mapped_pinned::clone_to_device_f32_via_pinned and mapped_pinned::upload_f32_via_pinned as the canonical replacements.
Final state of the HtoD migration sequence:
grep -rEn "stream\.memcpy_htod|memcpy_stod|stream\.htod_copy|htod_async" crates/ml/src/: production count 0, test count 7 (5 insignal_adapter.rs::tests, 1 ingpu_action_selector.rs::tests, 1 incuda_pipeline/mod.rs::tests).grep -rn "htod_f32\|clone_htod_f32" crates/ml/src/: outside test blocks 0 (only doc-comment text references remain). Helper definitions removed frommod.rs.
Fix 16 — Migrate training_loop.rs raw features/targets uploads (2 sites, 2026-04-28)
trainers/dqn/trainer/training_loop.rs:1250 (targets_raw) and :1254 (features_raw): both crate::cuda_pipeline::clone_htod_f32 callsites in the data-load helper rewritten to mapped_pinned::clone_to_device_f32_via_pinned. These run once per fold immediately before the step loop begins (the _raw suffix indicates the un-normalized arrays kept on-GPU for the eval-time critic). The .map_err(|e| anyhow::anyhow!("…raw upload: {e}")) wrapping is preserved verbatim — the helper already returns Result<_, String> which feeds anyhow::anyhow! directly, so the change is a 1:1 path swap.
Fix 15 — Migrate gpu_experience_collector.rs::upload_ofi_features (1 site, 2026-04-28)
gpu_experience_collector.rs:4258: warm-path super::clone_htod_f32 upload of OFI features rewritten to mapped_pinned::clone_to_device_f32_via_pinned. Called from training_loop's data-load phase before each fold's training run, so the OFI tensor (4M bars × 32 dims = ~512 MB) now stages through mapped-pinned + DtoD just like all other large data uploads. Error-type shift handled with .map_err(|e| MLError::ModelError(format!("upload_ofi_features: {e}")))? since the helper returns Result<_, String> and the function returns Result<_, MLError>.
Fix 14 — Migrate gpu_dqn_trainer.rs residual COLD ctor sites (8 sites, 2026-04-28)
gpu_dqn_trainer.rs: residual COLD ctor sites missed by the original Fix-1..13 audit are now migrated off explicit super::htod_f32 per feedback_no_htod_htoh_only_mapped_pinned.md.
:9922sel_clip_buf1-element init — selectivity gate clip-norm seed (1.0).:10075-10078spectral-norminit_u_s1/init_v_s1/init_u_s2/init_v_s2(4 separate uploads of trunk-encoder spectral u/v vectors).:10095-10096alloc_spec_pair!macro body — both u/v init uploads inside the macro now go throughupload_f32_via_pinned. The macro's$lbl_u/$lbl_vare reused in the error-message format so per-pair failures (spec_u_v1,spec_v_v1,spec_u_a1, …) stay diagnostically distinct.:11276graph_params_host(60 floats: cross-branch graph message passing weights, 4 edges × 15 params each, near-identity W_msg init).:11330denoise_params_host(1800 floats: 2-step diffusion Q-refinement Xavier init for W1[24,24] + W2[12,24]).:11476qlstm_weights_host(528 floats: QLSTM Xavier init for 6 gates × 11 input × 8 head). All sites usemapped_pinned::upload_f32_via_pinned(canonical mapped-pinned + DtoD staging helper). The helper returnsResult<_, String>whereas the ctor returnsResult<_, MLError>, so each site wraps the error via.map_err(|e| MLError::ModelError(format!("<site> upload via pinned: {e}"))). Site labels preserved so backtraces remain readable.