fix(gpu_dqn_trainer): migrate 11 COLD ctor HtoD sites to mapped pinned

Per `feedback_no_htod_htoh_only_mapped_pinned.md`, mapped pinned
(cuMemHostAlloc DEVICEMAP) is the only allowed CPU↔GPU path.

Adds module-level `upload_via_mapped_{f32,i32,u32,u64}` helpers and an
in-place `update_via_mapped_f32`. Each stages CPU data through a
transient mapped pinned buffer and DtoD-copies into the destination
`CudaSlice<T>`, then stream-syncs so the staging buffer is safe to
drop. Destination buffer types remain `CudaSlice<T>` so every existing
consumer (raw_ptr / kernel arg) is untouched, satisfying
`feedback_no_partial_refactor.md` for these one-shot init paths.

Migrated COLD sites in `GpuDqnTrainer::new`:
- weight_decay_mask (TOTAL_PARAMS f32)
- branch_slice_starts_dev, branch_slice_lens_dev ([i32; 4])
- branch_grad_scales_dev ([f32; 4])
- per_branch_gamma_base/max_dev ([f32; 4] each)
- q_quantile_branch_offsets/sizes_dev ([i32; 4] each)
- spectral_norm_descriptors_dev ([u64; 78])
- stochastic_depth_scale_buf ([f32; 3])
- stochastic_depth_rng_state ([u32; 1])
- vsn_group_begins_buf, vsn_group_ends_buf ([i32; num_groups] each)
- mamba2_params Xavier init (mamba2_param_count f32)

11 COLD HtoD sites eliminated. cargo check clean (15 warnings; +2 over
baseline 13 are the new MappedU32/U64 struct visibility warnings,
matching the existing MappedF32/I32 pattern).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
This commit is contained in:
jgrusewski
2026-04-28 21:06:13 +02:00
parent 0fb21c970c
commit 87ea9fa416
2 changed files with 188 additions and 67 deletions

View File

@@ -1417,6 +1417,21 @@ in constructor; kernel reads/writes via dev_ptr) and the upcoming
u64 table). All four mapped pinned variants share the same
`new`/`write_from_slice`/`read_all` API.
`gpu_dqn_trainer::new` HtoD elimination (2026-04-28): added module-level
`upload_via_mapped_{f32,i32,u32,u64}` helpers and an in-place
`update_via_mapped_f32`. Each stages CPU bytes through a transient
mapped pinned buffer and DtoD-copies into the destination
`CudaSlice<T>`, then stream-syncs so the staging buffer is safe to
drop. Migrated 11 COLD ctor sites: weight_decay_mask,
branch_slice_starts/lens, branch_grad_scales, per_branch_gamma_base/max,
q_quantile_branch_offsets/sizes, spectral_norm_descriptors (78 u64),
stochastic_depth_scale (3 f32), sd_rng_state (1 u32), vsn_group_begins,
vsn_group_ends, mamba2_params Xavier init. All consumer surfaces
unchanged — destination remains `CudaSlice<T>` with all subsequent
kernel-arg call sites intact. Per
`feedback_no_htod_htoh_only_mapped_pinned.md` and
`feedback_no_partial_refactor.md` (no consumer migration needed).
MoE moe_mixture_forward kernel + Rust wrapper (2026-04-27): first MoE
CUDA kernel landed. Single-thread-per-(b,c) kernel computes h_s2[b,c] =
Σ_k g[b,k]·expert_outputs[k,b,c]. No atomicAdd, capture-friendly. Rust