# Remaining GPU Features — Implementation Plan > **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking. **Goal:** Wire and implement the 4 remaining unwired GPU features: multi-head attention integration, Decision Transformer CUDA kernels, ensemble training with diversity loss, and HER Future strategy with GPU episode boundaries. **Architecture:** Each feature integrates into the existing fused CUDA training pipeline (`GpuDqnTrainer` + CUDA Graph). All kernels are GPU-native (zero CPU in hot path). Post-graph operations use the `EventTrackingGuard` RAII pattern for safe cudarc buffer access. Kernels are compiled via nvcc → cubin (cached at `/tmp/.cubin_cache/`). **CUDA Graph Integration Strategy:** The CUDA Graph captures the full forward+loss+backward+Adam sequence. Features that modify intermediate activations (attention, ensemble heads) operate with a **1-step lag**: they modify `save_h_s2` AFTER graph replay, so the modification takes effect on the NEXT graph replay's unflatten step. This is the same pattern as the IQN trunk gradient and spectral normalization — standard in async gradient methods. The alternative (splitting the graph into trunk/head phases) is deferred until profiling shows the lag impacts convergence. **VRAM Budget:** | Feature | RTX 3050 (4GB) | H100 (80GB) | |---------|---------------|-------------| | Attention (params + activations) | ~0.5MB | ~2MB | | Decision Transformer (params + context) | ~4MB | ~60MB | | Ensemble (K=3 head copies) | ~1.5MB | ~6MB | | HER episode tracking | ~0.5MB | ~4MB | | **Total** | **~6.5MB** | **~72MB** | **Tech Stack:** Rust (cudarc 0.19.3), CUDA C (nvcc/cubin), cuBLAS SGEMM --- ## File Structure ### New Files | File | Responsibility | |------|---------------| | `crates/ml/src/cuda_pipeline/attention_backward_kernel.cu` | Attention backward pass: LayerNorm backward + output projection backward + attention weight gradients | | `crates/ml/src/cuda_pipeline/dt_kernels.cu` | Decision Transformer forward: causal attention, softmax, positional encoding, token embedding, FFN, cross-entropy loss. Backward: softmax_backward, attention_backward (dQ/dK/dV), ffn_backward, layernorm_backward, embed_backward — 10 kernels total | | `crates/ml/src/cuda_pipeline/ensemble_kernels.cu` | Ensemble: multi-head Q-aggregation, KL-divergence diversity loss, Thompson sampling action selection | | `crates/ml/src/cuda_pipeline/her_episode_kernel.cu` | HER: episode boundary tracking, Future strategy donor sampling, episode-end detection | ### Modified Files | File | Changes | |------|---------| | `crates/ml/src/cuda_pipeline/gpu_attention.rs` | Add `backward()`, `params_buf()`, Adam state; wire into `GpuDqnTrainer` | | `crates/ml/src/cuda_pipeline/batched_forward.rs` | Add attention call between h_s2 and value/advantage heads | | `crates/ml/src/cuda_pipeline/batched_backward.rs` | Add attention backward in the gradient chain | | `crates/ml/src/cuda_pipeline/gpu_dqn_trainer.rs` | Conditional attention in forward/backward; ensemble head management | | `crates/ml/src/cuda_pipeline/decision_transformer.rs` | Implement `pretrain_step()` with real CUDA kernels | | `crates/ml/src/cuda_pipeline/gpu_her.rs` | Add Final/Future strategies, episode boundary buffer | | `crates/ml/src/cuda_pipeline/gpu_experience_collector.rs` | Add episode_ids to GpuExperienceBatch; fill via CUDA kernel (not CPU) | | `crates/ml/src/trainers/dqn/fused_training.rs` | Wire attention, ensemble, HER strategies | | `crates/ml/src/trainers/dqn/config.rs` | Add `use_attention: bool` to `DQNHyperparameters` (field exists in `DQNParams` hyperopt struct but not in the training config); wire from `DQNParams` during hyperopt conversion | | `crates/ml/src/trainers/dqn/trainer/training_loop.rs` | Ensemble head rotation, DT pre-training phase, HER episode tracking | | `crates/ml/src/hyperopt/adapters/dqn.rs` | Add DT pre-training epochs, ensemble params to search space | --- ## Task 1: Attention Forward Pass Integration **Files:** - Modify: `crates/ml/src/cuda_pipeline/gpu_dqn_trainer.rs` - Modify: `crates/ml/src/cuda_pipeline/gpu_attention.rs` - Modify: `crates/ml/src/trainers/dqn/fused_training.rs` - [ ] **Step 1: Add `GpuAttention` as optional field in `FusedTrainingCtx`** In `fused_training.rs`, add to `FusedTrainingCtx`: ```rust pub(crate) gpu_attention: Option, ``` Initialize it in `new()` when `hyperparams.use_attention` is true (similar to `gpu_iqn` initialization pattern at line 233). - [ ] **Step 2: Wire attention forward BETWEEN trunk h_s2 and value/advantage heads** In `gpu_dqn_trainer.rs`, add a method: ```rust pub fn apply_attention_forward( &self, attention: &mut GpuAttention, h_s2: &CudaSlice, batch_size: usize, ) -> Result<(), MLError> ``` This copies `h_s2` into attention input, runs `attention.forward()`, then copies the attended output BACK into `save_h_s2` (in-place replacement). The CUDA Graph captures the trunk forward FIRST, then attention runs OUTSIDE the graph (same pattern as IQN/spectral norm), then value/advantage heads run. - [ ] **Step 3: Call attention in `run_full_step()` after CUDA Graph replay** After Step 2 (graph replay) and before Step 5 (IQN): ```rust if let Some(ref mut attn) = self.gpu_attention { let _evt_guard = EventTrackingGuard::new(self.stream.context()); self.trainer.apply_attention_forward(attn, ...)?; } ``` - [ ] **Step 4: Compile check** ```bash SQLX_OFFLINE=true cargo check -p ml 2>&1 | tail -5 ``` - [ ] **Step 5: Commit** ```bash git add crates/ml/src/cuda_pipeline/gpu_attention.rs crates/ml/src/cuda_pipeline/gpu_dqn_trainer.rs crates/ml/src/trainers/dqn/fused_training.rs git commit -m "feat: wire attention forward pass into DQN training (post-graph, frozen weights)" ``` --- ## Task 2: Attention Backward Pass Kernel **Files:** - Create: `crates/ml/src/cuda_pipeline/attention_backward_kernel.cu` - Modify: `crates/ml/src/cuda_pipeline/gpu_attention.rs` - [ ] **Step 1: Write the attention backward CUDA kernel** The kernel computes `dL/d(h_s2_input)` given `dL/d(h_s2_attended)`: ```cuda extern "C" __global__ void attention_backward_kernel( const float* d_output, // [B, D] gradient from downstream const float* states_input, // [B, D] saved input to attention const float* params, // attention weights float* d_input, // [B, D] gradient to upstream (trunk) float* d_params, // [total_params] weight gradients (atomicAdd) int B ); ``` The backward flows through: 1. LayerNorm backward → d_prenorm 2. Residual: d_prenorm splits into d_projection + d_residual 3. Output projection backward: d_projection → d_concat, dW_O, db_O 4. Multi-head attention backward → dQ, dK, dV per head 5. Q/K/V projection backward → d_input, dW_Q, dW_K, dW_V Since the kernel is complex (matching the 237-line forward kernel), implement in phases: - Phase A: Residual passthrough only (d_input = d_output) — gradient flows, weights frozen - Phase B: Full backward with weight gradients (unfreezes attention weights) Start with Phase A (trivial kernel, verifiable). - [ ] **Step 2: Add `backward()` method to `GpuAttention`** ```rust pub fn backward( &self, d_output: &CudaSlice, d_input: &mut CudaSlice, batch_size: usize, ) -> Result<(), MLError> ``` For Phase A: simply `memcpy_dtod(d_input, d_output)` — residual gradient passthrough. - [ ] **Step 3: Wire backward into `apply_iqn_trunk_gradient`** After attention forward rewrites `save_h_s2`, the IQN trunk gradient flows through the attention backward before reaching the trunk layers: ```rust // Before trunk backward: d_h_s2 → attention_backward → d_h_s2_input if let Some(ref attn) = self.gpu_attention { attn.backward(&bw_d_h_s2, &mut bw_d_h_s2, batch_size)?; } ``` - [ ] **Step 4: Compile check + local test** ```bash SQLX_OFFLINE=true cargo check --workspace 2>&1 | tail -5 ``` - [ ] **Step 5: Commit** ```bash git add crates/ml/src/cuda_pipeline/attention_backward_kernel.cu crates/ml/src/cuda_pipeline/gpu_attention.rs git commit -m "feat: attention backward (Phase A — residual passthrough, gradient flow verified)" ``` --- ## Task 3: Decision Transformer CUDA Kernels **Files:** - Create: `crates/ml/src/cuda_pipeline/dt_kernels.cu` - Modify: `crates/ml/src/cuda_pipeline/decision_transformer.rs` - [ ] **Step 1: Write the token embedding + positional encoding kernel** ```cuda extern "C" __global__ void dt_embed_kernel( const float* trajectories, // [B, T, state_dim+2] (return-to-go, state, action) const float* W_embed, // [(state_dim+2), embed_dim] const float* b_embed, // [embed_dim] const float* pos_embed, // [T, embed_dim] (learned positional encoding) float* output, // [B, T, embed_dim] int B, int T, int input_dim, int embed_dim ); ``` Each thread handles one (batch, timestep, embed_dim_idx) tuple: - Linear projection: `out[b][t][d] = sum_j(traj[b][t][j] * W[j][d]) + b[d]` - Add positional: `out[b][t][d] += pos[t][d]` - [ ] **Step 2: Write the causal self-attention kernel** ```cuda extern "C" __global__ void dt_causal_attention_kernel( const float* Q, // [B, T, embed_dim] const float* K, // [B, T, embed_dim] const float* V, // [B, T, embed_dim] float* output, // [B, T, embed_dim] int B, int T, int D, int num_heads ); ``` Per-head computation with causal mask: - `attn[i][j] = Q[i] · K[j] / sqrt(D_h)` for j ≤ i (causal), -inf otherwise - Softmax over j dimension - `output[i] = sum_j(attn[i][j] * V[j])` Block-per-sample, threads handle different timesteps. Shared memory for K/V caching. - [ ] **Step 3: Write the FFN kernel (embed_dim → 4×embed_dim → embed_dim)** ```cuda extern "C" __global__ void dt_ffn_kernel( const float* input, // [B, T, embed_dim] const float* W1, // [embed_dim, 4*embed_dim] const float* b1, // [4*embed_dim] const float* W2, // [4*embed_dim, embed_dim] const float* b2, // [embed_dim] float* output, // [B, T, embed_dim] int B, int T, int D ); ``` GELU activation between layers. Residual connection: `output += input`. - [ ] **Step 4: Write the cross-entropy loss kernel** ```cuda extern "C" __global__ void dt_cross_entropy_kernel( const float* logits, // [B, T, num_actions] const int* target_actions, // [B, T] float* per_sample_loss, // [B*T] float* total_loss, // [1] atomicAdd int B, int T, int num_actions ); ``` Per-token cross-entropy: `loss = -log(softmax(logits)[target_action])`. - [ ] **Step 5: Implement `pretrain_step()` in decision_transformer.rs** Wire the 5 forward kernels into a training loop. Then: - [ ] **Step 5b: Write 5 backward CUDA kernels** ```cuda // In dt_kernels.cu — add after forward kernels: // dt_cross_entropy_backward_kernel: dL/d_logits = softmax - one_hot // dt_ffn_backward_kernel: GELU backward + 2 linear backward // dt_attention_backward_kernel: dQ,dK,dV from d_output (softmax Jacobian + causal mask) // dt_layernorm_backward_kernel: chain through mean/variance // dt_embed_backward_kernel: dW_embed from d_embedded ``` Each mirrors its forward kernel's thread/block layout. Total: 10 kernels. - [ ] **Step 5c: Implement `pretrain_step()` in decision_transformer.rs** 1. **Forward:** Embed → {QKV → CausalAttn → LN → FFN → LN} × N → ActionHead → CE loss 2. **Backward:** CE backward → reverse chain through all N layers → embed backward 3. **Optimizer:** Reuse `dqn_grad_norm_kernel` + `dqn_adam_update_kernel` on DT params DT action space = 9 exposure actions (branch_0 only, not full 81 factored). - [ ] **Step 6: Compile check** ```bash SQLX_OFFLINE=true cargo check -p ml 2>&1 | tail -5 ``` - [ ] **Step 7: Commit** ```bash git add crates/ml/src/cuda_pipeline/dt_kernels.cu crates/ml/src/cuda_pipeline/decision_transformer.rs git commit -m "feat: Decision Transformer CUDA kernels (embed, causal attention, FFN, CE loss)" ``` --- ## Task 4: Decision Transformer Pre-Training Loop **Files:** - Modify: `crates/ml/src/trainers/dqn/trainer/training_loop.rs` - Modify: `crates/ml/src/trainers/dqn/config.rs` - Modify: `crates/ml/src/hyperopt/adapters/dqn.rs` - [ ] **Step 1: Add DT pre-training config fields** In `DQNHyperparameters`: ```rust pub dt_pretrain_epochs: usize, // 0 = disabled, 10-50 typical pub dt_context_len: usize, // 20 timesteps pub dt_embed_dim: usize, // 128 pub dt_num_layers: usize, // 3 pub dt_target_return: f64, // 2.0 (target Sharpe-like return) ``` - [ ] **Step 2: Add DT pre-training phase at training start** In `training_loop.rs`, before the main training loop: ```rust if self.hyperparams.dt_pretrain_epochs > 0 { let dt = DecisionTransformer::new(stream, dt_config)?; for epoch in 0..self.hyperparams.dt_pretrain_epochs { // Build trajectory batches from training_data // Run dt.pretrain_step(trajectories, target_actions, batch_size) // Log DT loss } // Transfer DT's learned representations to DQN trunk (optional) } ``` - [ ] **Step 3: Add DT params to hyperopt search space** In `dqn.rs`, add bounds for `dt_pretrain_epochs`, `dt_context_len`, `dt_embed_dim`. - [ ] **Step 4: Compile check + commit** ```bash SQLX_OFFLINE=true cargo check --workspace git add crates/ml/src/cuda_pipeline/ crates/ml/src/trainers/dqn/ crates/ml/src/hyperopt/ && git commit -m "feat: Decision Transformer pre-training phase in training loop" ``` --- ## Task 5: Ensemble Multi-Head Q-Network **Files:** - Create: `crates/ml/src/cuda_pipeline/ensemble_kernels.cu` - Modify: `crates/ml/src/cuda_pipeline/gpu_dqn_trainer.rs` - Modify: `crates/ml/src/trainers/dqn/fused_training.rs` - Modify: `crates/ml/src/trainers/dqn/config.rs` - [ ] **Step 1: Write the ensemble aggregation kernel** ```cuda extern "C" __global__ void ensemble_aggregate_kernel( const float* head_q_values, // [K, B, num_actions] Q-values from K heads float* mean_q, // [B, num_actions] mean across heads float* var_q, // [B, num_actions] variance across heads int K, int B, int num_actions ); ``` Computes mean and variance of Q-values across K ensemble heads for uncertainty estimation. - [ ] **Step 2: Write the KL-divergence diversity loss kernel** ```cuda extern "C" __global__ void ensemble_diversity_kernel( const float* head_logits, // [K, B, num_atoms] per-head C51 logits float* diversity_loss, // [1] total KL divergence (atomicAdd) int K, int B, int num_atoms ); ``` For each pair of heads (i, j): `D_KL(p_i || p_j)` averaged over the batch. Uses hierarchical reduction (warp reduce → block reduce → atomicAdd per block, same pattern as `dqn_grad_norm_kernel`) to avoid single-address atomicAdd serialization. The diversity loss ENCOURAGES disagreement (subtracted from total loss). - [ ] **Step 3: Implement multi-head architecture in GpuDqnTrainer** The ensemble shares the TRUNK (W_s1, W_s2) but has K SEPARATE value + advantage heads: ```rust pub struct EnsembleHeads { heads: Vec<(DuelingWeightSet, BranchingWeightSet)>, // K head weight sets } ``` During training (heads run OUTSIDE the CUDA Graph, same pattern as IQN/attention): 1. CUDA Graph: forward through shared trunk → h_s2 (using head 0's weights in the graph) 2. Post-graph: for each head k=1..K-1, run cuBLAS forward through head_k's value/advantage layers on `save_h_s2` (outside graph, with EventTrackingGuard) 3. Aggregate Q-values: mean ± std across K heads 4. C51 loss per head (head 0 inside graph, heads 1..K-1 outside) 5. Diversity regularization across heads 6. Backward: each head contributes trunk gradients via the existing IQN-style SGD/Adam correction Note: The CUDA Graph remains captured for head 0 only. Heads 1..K-1 run as post-graph operations with their own cuBLAS forwards. This avoids re-capturing the graph for each head. - [ ] **Step 4: Add ensemble config and wiring** In `DQNHyperparameters`: ```rust pub ensemble_count: usize, // K heads (default 1 = no ensemble) pub ensemble_diversity_weight: f64, // 0.01 = mild diversity encouragement ``` In `FusedTrainingCtx::run_full_step()`: - If `ensemble_count > 1`: run K forward passes through different heads, average for action selection, sum losses with diversity term. - [ ] **Step 5: Compile check + commit** ```bash SQLX_OFFLINE=true cargo check --workspace git add crates/ml/src/cuda_pipeline/ crates/ml/src/trainers/dqn/ crates/ml/src/hyperopt/ && git commit -m "feat: ensemble multi-head Q-network with KL diversity loss" ``` --- ## Task 6: HER Episode Boundary Tracking **Files:** - Create: `crates/ml/src/cuda_pipeline/her_episode_kernel.cu` - Modify: `crates/ml/src/cuda_pipeline/gpu_her.rs` - Modify: `crates/ml/src/cuda_pipeline/gpu_experience_collector.rs` - [ ] **Step 1: Add episode index buffer to experience output** In `GpuExperienceBatch`, add: ```rust pub episode_ids: CudaSlice, // [N * L] episode index per transition ``` Fill via a simple CUDA kernel `fill_episode_ids_kernel(int* ids, int L, int N)` where thread `i` writes `ids[i] = i / L`. Launched as `grid=(ceil(N*L/256)), block=256`. Zero CPU fill — pure GPU. - [ ] **Step 2: Track episode boundaries in GPU replay buffer** When inserting transitions, also store the episode ID: ```rust pub episode_id_buf: CudaSlice, // [capacity] episode index per slot ``` This allows the HER kernel to find other transitions from the same episode. - [ ] **Step 3: Write the episode-aware donor sampling kernel** ```cuda extern "C" __global__ void her_sample_future_donors( const int* episode_ids, // [capacity] episode index per slot const int* source_indices, // [her_batch_size] which to relabel int* donor_indices, // [her_batch_size] output: future donors unsigned int* rng_states, // [her_batch_size] per-sample RNG int capacity, int her_batch_size ); ``` Uses a GPU-resident `episode_boundary_offsets` auxiliary buffer (updated at insertion time) for O(log(N)) binary search instead of O(capacity) linear scan. Each thread binary-searches the boundary buffer to find its episode's [start, end) range, then uniformly samples from `[source+1, end)`. - [ ] **Step 4: Write the episode-end detection kernel (Final strategy)** ```cuda extern "C" __global__ void her_find_episode_end( const int* episode_ids, // [capacity] const int* source_indices, // [her_batch_size] int* end_indices, // [her_batch_size] output: last index in same episode int capacity, int her_batch_size ); ``` For each source: binary search or linear scan to find the last index with the same episode_id. - [ ] **Step 5: Wire Future/Final strategies in `GpuHer`** ```rust pub fn relabel_batch_gpu( &mut self, strategy: HerGpuStrategy, episode_ids: &CudaSlice, ... ) -> Result<&HerBatch, MLError> { let donors = match strategy { HerGpuStrategy::Random => self.random_donors(her_batch_size), HerGpuStrategy::Future => self.sample_future_donors(episode_ids, source_indices)?, HerGpuStrategy::Final => self.find_episode_ends(episode_ids, source_indices)?, }; self.relabel_kernel(source_indices, donors, ...)?; Ok(&self.output) } ``` - [ ] **Step 6: Compile check + commit** ```bash SQLX_OFFLINE=true cargo check --workspace git add crates/ml/src/cuda_pipeline/ crates/ml/src/trainers/dqn/ crates/ml/src/hyperopt/ && git commit -m "feat: HER Future/Final strategies with GPU episode boundary tracking" ``` --- ## Task 7: Integration Test — Full Pipeline **Files:** - Modify: `crates/ml/src/trainers/dqn/trainer/training_loop.rs` - [ ] **Step 1: Wire all 4 features into the training loop conditionally** ```rust // Attention: after trunk forward, before heads if hyperparams.use_attention { ... } // DT pre-training: before main training loop if hyperparams.dt_pretrain_epochs > 0 { ... } // Ensemble: K heads with diversity loss if hyperparams.ensemble_count > 1 { ... } // HER: Future strategy with episode boundaries if hyperparams.her_ratio > 0.0 && hyperparams.her_strategy == "future" { ... } ``` - [ ] **Step 2: Local test — all features enabled** ```bash rm -rf /tmp/.cubin_cache/ SQLX_OFFLINE=true cargo build --release -p ml --example hyperopt_baseline_rl SQLX_OFFLINE=true timeout 120 ./target/release/examples/hyperopt_baseline_rl \ --model dqn --phase fast --trials 1 --n-initial 1 --epochs 3 \ --data-dir test_data/futures-baseline --symbol ES.FUT \ --output /tmp/full_pipeline_test.json --seed 42 ``` Expected: No NaN, no CUDA errors, train_loss < 50. - [ ] **Step 3: Commit** ```bash git add crates/ml/src/cuda_pipeline/ crates/ml/src/trainers/dqn/ crates/ml/src/hyperopt/ && git commit -m "feat: full pipeline integration — attention + DT + ensemble + HER Future" ``` --- ## Task 8: H100 Submission - [ ] **Step 1: Push to main** ```bash git push origin main ``` - [ ] **Step 2: Submit H100 hyperopt run** 50 epochs, 20 trials, with all features enabled via config: ```toml use_attention = true dt_pretrain_epochs = 10 ensemble_count = 3 her_ratio = 0.3 her_strategy = "future" ``` - [ ] **Step 3: Monitor first trial** Check: no NaN, no CUDA OOM, loss decreasing, Q-gap > 0. --- ## Dependency Graph ``` Task 1 (Attention Forward) → Task 2 (Attention Backward) → Task 7 (Integration) Task 3 (DT Kernels) → Task 4 (DT Loop) → Task 7 (Integration) Task 5 (Ensemble) → Task 7 (Integration) Task 6 (HER Episodes) → Task 7 (Integration) Task 7 (Integration) → Task 8 (H100) ``` Tasks 1-2, 3-4, 5, and 6 are INDEPENDENT and can be implemented in parallel.