# GPU IQL + HER Porting Plan **Date**: 2026-03-16 **Goal**: Port IQL (Implicit Q-Learning) and HER (Hindsight Experience Replay) from CPU to fused CUDA kernels, wire into the existing DQN training pipeline. **Target GPUs**: H100 80GB (production), RTX 3050 Ti 4GB (local dev) --- ## Current State ### IQL (CPU): `crates/ml-dqn/src/iql.rs` **Algorithm**: Kostrikov et al. 2021 -- offline RL without querying out-of-distribution actions. Three operations, all currently CPU Candle tensor ops: 1. **ValueNetwork V(s)**: 3-layer MLP (state_dim -> hidden -> hidden -> 1). Forward pass: Linear -> ReLU -> Linear -> ReLU -> Linear -> squeeze(1). Output: `[batch]` scalar. 2. **Expectile loss** (`expectile_loss()`): Asymmetric L2 on `diff = Q(s,a) - V(s)`. - `weight = tau * 1(diff >= 0) + (1-tau) * 1(diff < 0)` - `loss = mean(weight * diff^2)` - tau > 0.5 penalizes underestimation more, extracting best in-distribution action value. 3. **Advantage-weighted action selection** (`advantage_weighted_action()`): - `A(s,a) = Q(s,a) - V(s)` per action - `pi(a|s) = softmax(beta * clamp(A, -10, 10))` with max-subtraction stability trick - Output: `[batch, num_actions]` probabilities **Config fields** (already in `DQNHyperparameters`, `crates/ml/src/trainers/dqn/config.rs:1053-1066`): - `iql_expectile_tau: f32` (default 0.7) - `iql_advantage_temperature: f32` (default 3.0, field exists at line 1062) - `use_iql: bool` (default false) **Not yet wired**: IQL code exists in `ml-dqn` but is never called from the training loop or fused training path. ### HER (CPU): `crates/ml-dqn/src/hindsight_replay.rs` **Algorithm**: Andrychowicz et al. 2017 -- learn from failures by relabeling goals. Core operations: 1. **Goal extraction**: First `goal_dim` elements of state vector are the "goal". 2. **Goal relabeling**: Replace goal portion of (state, next_state) with an achieved goal from another experience. 3. **Reward recomputation**: Sparse reward -- `+1.0` if goal achieved (distance < 0.01), `-0.01` otherwise. 4. **Batch split**: `her_ratio` fraction of batch comes from HER-relabeled experiences, rest from normal PER sampling. 5. **Strategies**: Final (use final achieved goal), Future (sample from buffer), Episode, Random. **Config fields** (already in `DQNHyperparameters`, config.rs:954-958): - `her_ratio: f64` (default 0.0 = disabled) - `her_strategy: String` (default "future") **Partially wired**: HER buffer is constructed in `constructor.rs:460-482` when `her_ratio > 0.0` and stored as `DQNTrainer::her_buffer`. But it is a CPU `HindsightReplayBuffer` wrapping a CPU `PrioritizedReplayBuffer` -- completely separate from the GPU PER path (`GpuReplayBuffer`). ### GPU Training Pipeline (Existing) **Fused CUDA trainer**: `crates/ml/src/cuda_pipeline/gpu_dqn_trainer.rs` - 4 fused kernel phases captured in CUDA Graph: forward+loss, backward, grad_norm, Adam update - Launch: `grid=(batch_size, 1, 1), block=(32, 1, 1)` -- one warp per sample - Forward kernel computes: 3 forward passes (online/states, target/next_states, online/next_states for Double DQN) + C51 distributional loss - `train_step()` signature: states, next_states, actions, rewards, dones, is_weights, online/target weight sets - Returns `FusedTrainResult { total_loss, td_errors, grad_norm }` **Fused training context**: `crates/ml/src/trainers/dqn/fused_training.rs` - `FusedTrainingCtx::run_full_step()` orchestrates: fused train_step -> GPU EMA target update -> PER priority update - Reads `BatchSample` from GPU PER buffer, extracts flat arrays, feeds to `GpuDqnTrainer::train_step()` **GPU PER buffer**: `crates/ml-dqn/src/gpu_replay_buffer.rs` - `GpuReplayBuffer` -- ring buffer with contiguous GPU tensors: states `[capacity, state_dim]`, next_states, actions, rewards, dones, priorities - Sampling: prefix-sum cumsum + GPU searchsorted (zero CPU binary search) - Returns `GpuBatch { states, next_states, actions, rewards, dones, is_weights, indices }` **Experience collector**: `crates/ml/src/cuda_pipeline/gpu_experience_collector.rs` - `dqn_full_experience_kernel` runs N episodes x L timesteps entirely on GPU - Zero CPU-GPU roundtrips per timestep - Outputs go directly into GPU PER buffer via `insert_batch()` **Training loop**: `crates/ml/src/trainers/dqn/trainer/training_loop.rs` - `run_training_steps()` pre-samples K=32 batches under READ lock, then runs GPU train steps under WRITE lock - Each step: `fused_ctx.run_full_step(batch, agent, device)` -> guard kernel check --- ## IQL GPU Porting ### What Needs to Happen IQL adds a **parallel value network** V(s) trained alongside the existing Q-network. The Q-network loss changes: instead of using `max_a Q_target(s', a')` as the bootstrap target, IQL uses `V(s')` from the value network. This eliminates querying OOD actions. The training step becomes: ``` # Standard DQN (existing): target = r + gamma * max_a Q_target(s', a') loss_q = PER_weighted_CE(Q_online(s, a), project(target)) # C51 # IQL addition: target_q = Q_online(s, a).detach() # stop gradient loss_v = expectile_loss(V(s), target_q, tau) # new kernel loss_total = loss_q + loss_v # Action selection change (inference only): A(s,a) = Q(s,a) - V(s) pi(a|s) = softmax(beta * A(s,a)) # replaces epsilon-greedy ``` ### CUDA Kernel: `iql_value_kernel.cu` (NEW) **File**: `crates/ml/src/cuda_pipeline/iql_value_kernel.cu` This kernel fuses the value network forward pass + expectile loss computation into a single launch. It runs **after** the existing forward+loss kernel (which computes Q-values and C51 loss), using the Q-values as targets for V(s). ```c /** * IQL value network forward + expectile loss kernel. * * Launch: grid=(batch_size, 1, 1), block=(32, 1, 1) * One warp per sample, matching dqn_training_kernel.cu convention. * * Network: state -> [V_H1] -> ReLU -> [V_H2] -> ReLU -> [1] (scalar V(s)) * * Inputs: * states [batch_size, STATE_DIM] -- same buffer as main forward kernel * q_values [batch_size] -- Q(s, a_taken) from main forward kernel output * v_weights flat buffer of value network parameters * * Outputs: * v_values [batch_size] -- V(s) for each sample * v_loss [1] -- scalar expectile loss (atomicAdd reduction) * v_grad [total_v_params] -- gradient buffer for Adam update * * Compile-time defines: * STATE_DIM, V_H1, V_H2, IQL_TAU, IQL_BETA (advantage temperature) */ ``` **Kernel structure** (single warp per sample): 1. Forward pass through V(s) network (3 linear layers, same warp-cooperative matmul as existing kernel) 2. Read Q(s, a_taken) from main kernel's activation buffer 3. Compute `diff = Q(s,a) - V(s)`, `weight = tau * (diff >= 0) + (1-tau) * (diff < 0)` 4. `loss_sample = weight * diff^2`, atomicAdd into scalar accumulator 5. Backward through V(s) network, accumulate gradients into flat gradient buffer **Parameter layout**: Separate flat buffer for V-network weights (6 tensors: w_v1, b_v1, w_v2, b_v2, w_vout, b_vout). Separate Adam moments (m_v, v_v). Same Adam kernel reused with different buffer pointers. ### CUDA Kernel: `iql_adam_kernel.cu` (REUSE) No new Adam kernel needed. The existing `dqn_adam_update_kernel` takes arbitrary flat parameter/gradient/moment buffers. Launch it a second time with V-network buffer pointers: ```c // Existing kernel, second launch for V-network: dqn_adam_update_kernel<<>>( v_params, v_grads, v_m, v_v, v_t, total_v_params, lr, beta1, beta2, eps, weight_decay, max_grad_norm ); ``` ### Interaction with C51 Distributional Q-Network The existing C51 forward+loss kernel (`dqn_forward_loss_kernel`) computes: - 3 forward passes: online(states), target(next_states), online(next_states) - C51 Bellman projection + cross-entropy loss - Saves activations for backward **IQL modifies the target computation**. Instead of: ``` target = r + gamma * max_a Q_target(s', a') # via C51 projection ``` It uses: ``` target = r + gamma * V(s') # V(s') from value network ``` **Implementation**: Add an IQL code path inside `dqn_forward_loss_kernel`: - After the target network forward pass, compute V(s') via the value network (same shared-memory pattern) - Replace the `max_a E[Q_target(s',a')]` expected value with `V(s')` - The C51 projection remains identical -- only the scalar `target_q` changes - Guard with `#ifdef USE_IQL` compile-time define This avoids a separate kernel launch for the target replacement. The value network forward for `s'` is fused into the same kernel that already processes `s'` through the target Q-network. ### Rust Integration: `crates/ml/src/cuda_pipeline/gpu_iql_trainer.rs` (NEW) ```rust /// GPU-accelerated IQL value network trainer. /// /// Owns V-network weight/gradient/moment CudaSlice buffers. /// Called after `GpuDqnTrainer::train_step()` with Q-values as targets. pub struct GpuIqlTrainer { // V-network weights (flat CudaSlice buffers) v_params: CudaSlice, v_grads: CudaSlice, v_m: CudaSlice, // Adam first moment v_v: CudaSlice, // Adam second moment v_t: CudaSlice, // Adam step counter // Pre-allocated output buffers v_values_buf: CudaSlice, // [batch_size] v_loss_buf: CudaSlice, // [1] // Compiled kernels forward_loss_func: CudaFunction, adam_func: CudaFunction, config: IqlGpuConfig, } pub struct IqlGpuConfig { pub state_dim: usize, pub v_h1: usize, // default: 256 (match shared_h1) pub v_h2: usize, // default: 256 (match shared_h2) pub batch_size: usize, pub tau: f32, // expectile parameter pub beta: f32, // advantage temperature pub lr: f32, pub beta1: f32, pub beta2: f32, pub epsilon: f32, pub weight_decay: f32, pub max_grad_norm: f32, } impl GpuIqlTrainer { /// Forward V(s) + expectile loss + backward + Adam update. /// /// Called after `GpuDqnTrainer::train_step()` with the same states buffer /// and Q(s,a) values from the forward kernel's activation cache. pub fn train_step( &mut self, states: &CudaSlice, // [batch_size * state_dim] q_targets: &CudaSlice, // [batch_size] Q(s, a_taken) ) -> Result; /// Get V(s') for target computation in main DQN kernel. /// /// Forward-only pass through V-network on next_states. /// Result stays on GPU for the C51 projection kernel. pub fn forward_values( &self, next_states: &CudaSlice, // [batch_size * state_dim] ) -> Result<&CudaSlice, MLError>; // [batch_size] /// EMA target update for V-network (Polyak averaging). pub fn target_ema_update(&mut self, tau: f32) -> Result<(), MLError>; } ``` ### Where in the Training Loop IQL Gets Called **File**: `crates/ml/src/trainers/dqn/fused_training.rs`, `FusedTrainingCtx::run_full_step()` Current flow: ``` 1. extract_batch_arrays(batch) 2. trainer.train_step(states, next_states, actions, rewards, dones, is_weights, ...) 3. trainer.target_ema_update(...) 4. agent.fused_post_step_no_ema(td_errors, indices) ``` IQL-augmented flow: ``` 1. extract_batch_arrays(batch) 2. if use_iql: 2a. iql_trainer.forward_values(next_states) -- V(s') for target 2b. trainer.train_step_iql(states, next_states, actions, rewards, dones, is_weights, v_next_states, ...) -- uses V(s') instead of max Q_target 2c. iql_trainer.train_step(states, q_taken_buf) -- expectile loss on V(s) else: 2. trainer.train_step(...) -- existing C51 path 3. trainer.target_ema_update(...) 4. if use_iql: iql_trainer.target_ema_update(tau) 5. agent.fused_post_step_no_ema(td_errors, indices) ``` **New field on `FusedTrainingCtx`**: ```rust pub(crate) iql_trainer: Option, ``` Initialized when `hyperparams.use_iql == true`. ### Config Wiring Already exists -- just needs to be read and propagated: 1. `DQNHyperparameters::use_iql` -> `FusedTrainingCtx::new()` -> construct `GpuIqlTrainer` if true 2. `DQNHyperparameters::iql_expectile_tau` -> `IqlGpuConfig::tau` 3. `DQNHyperparameters::iql_advantage_temperature` -> `IqlGpuConfig::beta` 4. Hyperopt adapter (`crates/ml/src/hyperopt/adapters/dqn.rs`): already sets defaults (line 2770-2772) ### CUDA Graph Implications The existing training CUDA Graph captures forward+loss+backward+Adam for the Q-network. Adding IQL means: **Option A** (simpler): Do NOT capture IQL kernels in the same graph. Launch them as separate async kernels on the same stream after graph replay. The IQL value network is much smaller than the Q-network (no 3-branch head, no C51 atoms), so launch overhead is negligible relative to computation. **Option B** (optimal): Extend the CUDA Graph capture to include IQL kernels. This requires the IQL trainer to be initialized before graph capture and its buffer pointers to be stable (same constraint as existing Q-network). More complex but eliminates 2 extra kernel launches per step. **Recommendation**: Start with Option A. The IQL value network has ~3x fewer parameters than the Q-network. At batch_size=256, the 2 extra kernel launches add ~4us overhead vs ~50us compute. Optimize to Option B only if profiling shows launch overhead is significant. --- ## HER GPU Porting ### What Needs to Happen HER operates at the **replay buffer level**, not the network level. It modifies experiences before they are sampled for training. The key insight: goal relabeling is embarrassingly parallel -- each experience can be relabeled independently. Currently, HER uses a completely separate CPU `HindsightReplayBuffer` that wraps a CPU `PrioritizedReplayBuffer`. The GPU path uses `GpuReplayBuffer` for PER. These are disjoint -- HER is dead code on the GPU path. **Goal**: Implement HER as a GPU kernel that operates on the `GpuReplayBuffer`'s tensor storage, producing relabeled experiences directly in GPU memory. ### CUDA Kernel: `her_relabel_kernel.cu` (NEW) **File**: `crates/ml/src/cuda_pipeline/her_relabel_kernel.cu` ```c /** * HER goal relabeling kernel -- operates on GPU replay buffer tensors. * * Launch: grid=(her_batch_size, 1, 1), block=(32, 1, 1) * One warp per relabeled experience. * * For each experience to relabel: * 1. Read original (state, next_state, action, reward, done) from source indices * 2. Read achieved goal from a donor experience (strategy-dependent) * 3. Replace goal portion in state and next_state * 4. Recompute reward (sparse: +1 if goal achieved, -0.01 otherwise) * 5. Write relabeled experience to output staging buffer * * Inputs: * states [capacity, STATE_DIM] -- GPU replay buffer storage (read-only) * next_states [capacity, STATE_DIM] -- GPU replay buffer storage (read-only) * actions [capacity] -- GPU replay buffer storage (read-only) * rewards [capacity] -- GPU replay buffer storage (read-only) * dones [capacity] -- GPU replay buffer storage (read-only) * source_indices [her_batch_size] -- indices into replay buffer to relabel * donor_indices [her_batch_size] -- indices of goal donors (strategy-dependent) * buffer_size int -- current replay buffer occupancy * * Outputs: * out_states [her_batch_size, STATE_DIM] -- relabeled states * out_next_states [her_batch_size, STATE_DIM] -- relabeled next_states * out_actions [her_batch_size] -- copied (unchanged) * out_rewards [her_batch_size] -- recomputed rewards * out_dones [her_batch_size] -- copied (unchanged) * * Compile-time defines: * STATE_DIM, GOAL_DIM, GOAL_THRESHOLD (default 0.01) */ __global__ void her_relabel_kernel( const float* __restrict__ states, const float* __restrict__ next_states, const int* __restrict__ actions, const float* __restrict__ rewards, const float* __restrict__ dones, const int* __restrict__ source_indices, const int* __restrict__ donor_indices, int buffer_size, float* __restrict__ out_states, float* __restrict__ out_next_states, int* __restrict__ out_actions, float* __restrict__ out_rewards, float* __restrict__ out_dones ) { int sample_idx = blockIdx.x; int lane_id = threadIdx.x; // 0-31 int src_idx = source_indices[sample_idx]; int donor_idx = donor_indices[sample_idx]; // 1. Copy full state (warp-cooperative coalesced copy) for (int d = lane_id; d < STATE_DIM; d += 32) { out_states[sample_idx * STATE_DIM + d] = states[src_idx * STATE_DIM + d]; out_next_states[sample_idx * STATE_DIM + d] = next_states[src_idx * STATE_DIM + d]; } __syncwarp(); // 2. Replace goal portion with donor's achieved goal for (int d = lane_id; d < GOAL_DIM; d += 32) { float achieved = next_states[donor_idx * STATE_DIM + d]; out_states[sample_idx * STATE_DIM + d] = achieved; out_next_states[sample_idx * STATE_DIM + d] = achieved; } // 3. Recompute reward (lane 0 does scalar work) if (lane_id == 0) { out_actions[sample_idx] = actions[src_idx]; out_dones[sample_idx] = dones[src_idx]; // Goal distance float dist_sq = 0.0f; for (int d = 0; d < GOAL_DIM; d++) { float achieved = next_states[src_idx * STATE_DIM + d]; float desired = next_states[donor_idx * STATE_DIM + d]; float diff = achieved - desired; dist_sq += diff * diff; } float dist = sqrtf(dist_sq); out_rewards[sample_idx] = (dist < GOAL_THRESHOLD) ? 1.0f : -0.01f; } } ``` **Launch config**: `grid=(her_batch_size, 1, 1), block=(32, 1, 1)` -- one warp per sample, matching existing kernel convention. For batch_size=256 and her_ratio=0.5, her_batch_size=128. ### Donor Index Generation The HER strategy determines which experiences provide the replacement goals. Donor indices must be computed **before** the relabel kernel launch. **Strategy implementations**: 1. **Final**: Donor = last experience in the same episode. Requires episode boundary tracking on GPU. - Approach: Upload episode boundary array `[num_episodes]` to GPU. Binary search to find episode end for each source index. Donor = `episode_end - 1`. - Kernel: `her_donor_final_kernel` -- one thread per source, O(log E) binary search. 2. **Future**: Donor = random future experience from same episode. - Approach: For each source index, sample uniform random index in `[source_idx+1, episode_end)`. - Kernel: `her_donor_future_kernel` -- one thread per source, cuRAND for random offset. 3. **Random**: Donor = random experience from entire buffer. - No kernel needed. Generate random indices on CPU and upload, or use cuRAND inline. - Simplest to implement first. **Recommendation**: Implement Random strategy first (no episode tracking needed). Future strategy second (requires episode boundaries on GPU). ### Rust Integration: `crates/ml/src/cuda_pipeline/gpu_her.rs` (NEW) ```rust /// GPU-accelerated Hindsight Experience Replay. /// /// Operates on GpuReplayBuffer tensor storage. Produces relabeled /// experience batches that are concatenated with normal PER samples /// before feeding to the training kernel. pub struct GpuHer { config: GpuHerConfig, // Staging buffers for relabeled experiences (pre-allocated) out_states: CudaSlice, // [max_her_batch, state_dim] out_next_states: CudaSlice, // [max_her_batch, state_dim] out_actions: CudaSlice, // [max_her_batch] out_rewards: CudaSlice, // [max_her_batch] out_dones: CudaSlice, // [max_her_batch] // Index buffers source_indices: CudaSlice, // [max_her_batch] donor_indices: CudaSlice, // [max_her_batch] // Episode tracking (for Final/Future strategies) episode_boundaries: Option>, // Compiled kernel relabel_func: CudaFunction, stream: Arc, } pub struct GpuHerConfig { pub goal_dim: usize, // default: 1 (single scalar goal) pub her_ratio: f64, // fraction of batch from HER pub strategy: HerGpuStrategy, // Final, Future, Random pub goal_threshold: f32, // default: 0.01 pub batch_size: usize, // total batch size pub state_dim: usize, } pub enum HerGpuStrategy { Final, Future, Random, } impl GpuHer { /// Generate relabeled batch from GPU replay buffer. /// /// 1. Sample source_indices from GpuReplayBuffer (PER sampling) /// 2. Generate donor_indices (strategy-dependent) /// 3. Launch relabel kernel /// 4. Return staging buffer pointers for concatenation with normal batch pub fn relabel_batch( &mut self, replay_buffer: &GpuReplayBuffer, her_batch_size: usize, ) -> Result; } /// GPU-resident relabeled batch -- pointers into staging buffers. pub struct HerBatch { pub states: CudaSlice, pub next_states: CudaSlice, pub actions: CudaSlice, pub rewards: CudaSlice, pub dones: CudaSlice, pub batch_size: usize, } ``` ### How HER Interacts with GPU PER The key integration point is **batch assembly**. Currently: ``` GpuReplayBuffer::sample(batch_size) -> GpuBatch (PER-weighted) -> extract to flat arrays -> GpuDqnTrainer::train_step(states, next_states, ...) ``` With HER: ``` normal_size = batch_size * (1 - her_ratio) # e.g., 128 her_size = batch_size * her_ratio # e.g., 128 # Sample normal PER batch normal_batch = GpuReplayBuffer::sample(normal_size) # PER-weighted # Sample source indices for HER (also PER-weighted) her_sources = GpuReplayBuffer::sample_indices(her_size) # just indices # GPU relabeling her_batch = GpuHer::relabel_batch(replay_buffer, her_size) # Concatenate on GPU (D2D copies, zero CPU) merged_states = concat_cuda_slices(normal_batch.states, her_batch.states) merged_next = concat_cuda_slices(normal_batch.next_states, her_batch.next_states) # ... etc for actions, rewards, dones # IS-weights for HER samples: uniform weight (no PER bias) merged_is_weights = concat(normal_batch.is_weights, ones(her_size)) # Train on merged batch GpuDqnTrainer::train_step(merged_states, merged_next, ..., merged_is_weights) ``` **Priority update for HER samples**: After training, TD errors are available for both normal and HER samples. Normal sample priorities update in the replay buffer. HER sample priorities are discarded (they are synthetic experiences, not stored in the buffer). **New method on `GpuReplayBuffer`**: ```rust /// Sample only indices (for HER source selection), without extracting full experiences. /// Returns PER-weighted indices and IS-weights. pub fn sample_indices(&mut self, batch_size: usize) -> Result<(CudaSlice, Tensor), MLError>; ``` ### Where in the Training Loop HER Gets Called **File**: `crates/ml/src/trainers/dqn/fused_training.rs` In `FusedTrainingCtx::run_full_step()`, **before** the existing train_step: ```rust // New HER integration point: let (states, next_states, actions, rewards, dones, is_weights) = if let Some(ref mut her) = self.gpu_her { // Split batch let her_size = (batch.batch_size as f64 * her.config.her_ratio) as usize; let normal_size = batch.batch_size - her_size; // Normal PER batch (first normal_size samples) let normal = extract_batch_arrays_partial(batch, state_dim, 0, normal_size)?; // HER relabeled batch let her_batch = her.relabel_batch(&self.replay_buffer_ref, her_size)?; // GPU concat merge_batches_gpu(normal, her_batch, &self.stream)? } else { extract_batch_arrays(batch, state_dim)? }; ``` **New field on `FusedTrainingCtx`**: ```rust pub(crate) gpu_her: Option, ``` **New field on `DQNTrainer`** (to pass replay buffer reference): ```rust // The GpuReplayBuffer reference is already accessible via agent.memory() -> buffer // but GpuHer needs raw CudaSlice access. Pass at construction time. ``` ### Config Wiring 1. `DQNHyperparameters::her_ratio` -> if > 0.0, construct `GpuHer` in `FusedTrainingCtx::new()` 2. `DQNHyperparameters::her_strategy` -> `GpuHerConfig::strategy` mapping: "final" -> Final, "future" -> Future, _ -> Random 3. `HindsightReplayConfig::goal_dim` -> `GpuHerConfig::goal_dim` (default 1) 4. Remove the CPU `HindsightReplayBuffer` from `DQNTrainer::her_buffer` when GPU path is active ### Episode Boundary Tracking on GPU For Final/Future strategies, the relabel kernel needs to know episode boundaries to select donors from the correct episode. **Approach**: The GPU experience collector already processes episodes. Add an episode boundary output buffer to the experience kernel: ```c // In dqn_experience_kernel.cu, after each episode: if (done || t == timesteps_per_episode - 1) { int ep_boundary_idx = atomicAdd(num_episodes_out, 1); episode_ends[ep_boundary_idx] = global_experience_idx; } ``` On the Rust side, maintain a GPU-resident `episode_boundaries: CudaSlice` that grows as episodes complete. The HER donor kernel reads this for binary search. --- ## Integration Order ### Phase 1: HER GPU (implement first) **Rationale**: HER operates at the data layer (replay buffer), not the network layer. It does not require modifying the existing CUDA Graph-captured training kernels. The relabel kernel is simple (just memory copies + scalar arithmetic) and can be developed and tested independently. **Steps**: 1. Write `her_relabel_kernel.cu` with Random strategy (simplest, no episode tracking) 2. Write `gpu_her.rs` with staging buffer management and kernel compilation 3. Add `GpuReplayBuffer::sample_indices()` method 4. Add GPU concat utility for merging normal + HER batches 5. Wire into `FusedTrainingCtx::run_full_step()` behind `her_ratio > 0.0` guard 6. Smoke test: verify HER-relabeled experiences have correct goal substitution **Estimated kernel complexity**: ~60 lines of CUDA (vs ~800 for the training kernel). Launch overhead: <2us at her_batch_size=128. ### Phase 2: IQL GPU (implement second) **Rationale**: IQL requires a new network (V(s)), new loss computation, and modifications to the existing forward+loss kernel's target computation. It touches the CUDA Graph-captured kernel sequence, so it requires more careful integration. **Steps**: 1. Write `iql_value_kernel.cu` with forward + expectile loss + backward 2. Write `gpu_iql_trainer.rs` with weight/gradient/moment buffer management 3. Modify `dqn_training_kernel.cu` to accept V(s') target via `#ifdef USE_IQL` 4. Modify `GpuDqnTrainer` to accept IQL target values in `train_step()` 5. Wire IQL trainer into `FusedTrainingCtx::run_full_step()` behind `use_iql` guard 6. Handle CUDA Graph invalidation (IQL changes the kernel sequence) 7. Smoke test: verify V(s) converges to Q-value range, expectile loss decreases **Estimated kernel complexity**: ~200 lines of CUDA. Two additional kernel launches per step (forward+loss+backward, Adam). ### Phase 3: Combined IQL + HER Once both work independently, enable both simultaneously. The interaction is clean: - HER produces relabeled experiences -> feeds into training step - IQL modifies the training step's target computation - No cross-dependency between HER relabeling and IQL value network --- ## Testing Strategy ### Smoke Tests (GPU-required, `#[cfg(feature = "cuda")]`) **File**: `crates/ml/src/trainers/dqn/smoke_tests/iql_her_gpu.rs` (NEW) ```rust #[test] fn test_iql_value_kernel_forward() { // Random states -> V(s) output shape [batch_size] // V(s) should be finite, in reasonable range } #[test] fn test_iql_expectile_loss_asymmetric() { // tau=0.7: loss should weight underestimation more // Compare GPU kernel output vs CPU reference implementation } #[test] fn test_her_relabel_kernel_goal_replacement() { // Fill GPU replay buffer with known experiences // Run relabel kernel with known donor indices // Verify goal portion replaced, non-goal features preserved } #[test] fn test_her_reward_recomputation() { // Source experience with goal A, donor with achieved goal B // If distance(achieved_at_source, B) < threshold: reward = +1.0 // Else: reward = -0.01 } #[test] fn test_fused_step_with_iql() { // End-to-end: batch -> fused train_step with IQL -> verify loss/grads finite } #[test] fn test_fused_step_with_her() { // End-to-end: batch with HER relabeling -> fused train_step -> verify combined batch correct } ``` ### Validation Against CPU Reference For both IQL and HER, run CPU reference implementation and GPU kernel on the same inputs, verify outputs match within BF16 tolerance (1e-2 relative error for loss, exact match for relabeled states). --- ## Performance Targets ### HER - **Walltime impact**: Near-zero. The relabel kernel processes 128 samples in <2us (pure memory copy + scalar arithmetic). The concat operation is 2 D2D copies. - **VRAM overhead**: `2 * her_batch_size * state_dim * 4 bytes` for staging buffers. At batch_size=256, her_ratio=0.5, state_dim=56: `2 * 128 * 56 * 4 = 57 KB`. Negligible. - **Data efficiency improvement**: HER literature reports 5-10x sample efficiency improvement in sparse reward settings. For HFT with shaped rewards, expect 1.5-3x. ### IQL - **Walltime impact**: ~10-15% increase per training step. The V-network is ~1/4 the size of the Q-network (no 3 branch heads, no C51 atoms, output dim 1 vs 11). The expectile loss is cheaper than C51 cross-entropy. The second Adam launch adds ~2us. - **VRAM overhead**: V-network parameters + Adam moments. At hidden_dim=256: ~400K parameters * 3 buffers (params, m, v) * 4 bytes = ~5 MB. Negligible on H100. - **Training quality improvement**: IQL eliminates OOD action querying, which is critical for offline RL with historical market data. Expected: fewer divergent Q-value episodes, more stable training in walk-forward folds with distribution shift. ### Combined - **Total walltime overhead**: ~12-18% increase per training step (IQL dominant, HER negligible) - **Total VRAM overhead**: ~5 MB (IQL) + ~60 KB (HER) = ~5 MB - **Expected quality improvement**: Better sample efficiency (HER) + more stable offline training (IQL). Target: 10-20% improvement in walk-forward validation Sharpe ratio. --- ## DSR GPU Porting (Warp Kernel Gap) ### Current State DSR (Differential Sharpe Ratio, Moody & Saffell 2001) is **partially on GPU** — the device function exists but the warp kernel doesn't call it. | Component | Status | Location | |-----------|--------|----------| | CPU struct | Done | `ml-dqn/src/reward.rs:186-275` | | GPU device function `dsr_step()` | Done | `common_device_functions.cuh:1203-1238` | | Per-thread kernel integration | Done | `dqn_experience_kernel.cu:2100` | | **Warp kernel integration (H100)** | **MISSING** | `dqn_full_experience_kernel_warp()` — no `dsr_step()` call | | Epoch state writeout | Done | `dqn_experience_kernel.cu:2230-2236` | | **Epoch state readback** | **MISSING** | `gpu_experience_collector.rs` has buffer but no CPU sync | | **CPU↔GPU state sync** | **MISSING** | CPU `DifferentialSharpeRatio` not synchronized with GPU epoch_state | ### Algorithm (per-step, inherently sequential) ``` Given: return r_t, prior EMA values A_prev and B_prev, decay rate eta delta_A = r_t - A_prev delta_B = r_t^2 - B_prev variance = B_prev - A_prev^2 numerator = B_prev * delta_A - 0.5 * A_prev * delta_B denominator = variance^1.5 DSR = clamp(numerator / denominator, -5.0, 5.0) Update: A = A_prev + eta * delta_A B = B_prev + eta * delta_B ``` DSR is inherently sequential (each step depends on previous EMA state), but parallelism comes from running **multiple episodes concurrently** — each episode has independent DSR state. ### Tasks 1. **Wire `dsr_step()` in warp kernel** (`dqn_full_experience_kernel_warp()`): - Lane 0 maintains `dsr_A`, `dsr_B`, `dsr_initialized` (thread-local) - Call `dsr_step()` at the reward computation point (matching per-thread kernel line 2100) - Synchronize at episode boundaries via `__syncwarp()` 2. **Add epoch state readback** in `GpuExperienceCollector`: ```rust pub fn read_epoch_dsr_state(&self) -> Result<(f32, f32), MLError> { // Read epoch_state[5] (dsr_A) and epoch_state[6] (dsr_B) from GPU } ``` 3. **Wire CPU sync** in `DQNTrainer::train_epoch()`: ```rust if config.use_dsr { if let Some((dsr_a, dsr_b)) = collector.read_epoch_dsr_state() { self.dsr.ema_return = dsr_a as f64; self.dsr.ema_return_sq = dsr_b as f64; self.dsr.initialized = true; } } ``` 4. **GPU kernel test**: `test_dsr_kernel_matches_cpu()` — launch kernel with known reward sequence, verify output matches CPU reference within f32 tolerance. ### Performance - Walltime: neutral (DSR is absorbed into existing warp-cooperative compute) - State readback: async memcpy, negligible cost - No additional kernel launches needed --- ## File Summary ### New Files | File | Purpose | |------|---------| | `crates/ml/src/cuda_pipeline/iql_value_kernel.cu` | IQL value network forward + expectile loss + backward | | `crates/ml/src/cuda_pipeline/gpu_iql_trainer.rs` | Rust wrapper for IQL GPU trainer | | `crates/ml/src/cuda_pipeline/her_relabel_kernel.cu` | HER goal relabeling kernel | | `crates/ml/src/cuda_pipeline/gpu_her.rs` | Rust wrapper for GPU HER | | `crates/ml/src/trainers/dqn/smoke_tests/iql_her_gpu.rs` | GPU smoke tests | ### Modified Files | File | Change | |------|--------| | `crates/ml/src/cuda_pipeline/mod.rs` | Add `pub mod gpu_iql_trainer; pub mod gpu_her;` | | `crates/ml/src/cuda_pipeline/dqn_training_kernel.cu` | Add `#ifdef USE_IQL` target replacement path | | `crates/ml/src/cuda_pipeline/gpu_dqn_trainer.rs` | Add `train_step_iql()` variant accepting V(s') targets | | `crates/ml/src/trainers/dqn/fused_training.rs` | Wire IQL + HER into `run_full_step()` | | `crates/ml/src/trainers/dqn/trainer/mod.rs` | Remove CPU `her_buffer` when GPU path active | | `crates/ml-dqn/src/gpu_replay_buffer.rs` | Add `sample_indices()` method | | `crates/ml/src/hyperopt/adapters/dqn.rs` | Add IQL hyperparams to search space (tau, temperature) |