7 tasks: build.rs POC → parameterize #define → update launches → build all → replace runtime compilation → delete NVRTC infra → integration test. Key insight: convert #define STATE_DIM/NUM_ATOMS to kernel params so each kernel has ONE version regardless of model config. No multi-variant compilation. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
522 lines
21 KiB
Markdown
522 lines
21 KiB
Markdown
# Eliminate NVRTC Runtime Compilation — Implementation Plan
|
|
|
|
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
|
|
|
|
**Goal:** Replace all runtime nvcc/NVRTC kernel compilation with build-time precompiled cubins embedded in the binary. Zero runtime compiler invocations, zero disk cache, zero stale-cache segfaults.
|
|
|
|
**Architecture:** (1) Convert all `#define` compile-time constants to kernel parameters so each kernel has ONE version regardless of model config. (2) Add `build.rs` to `ml` crate that runs `nvcc` on all 25 `.cu` files at `cargo build` time, producing architecture-specific cubins. (3) Replace all `include_str!` + `compile_ptx_for_device()` + `OnceLock<Ptx>` patterns with `include_bytes!` + `CudaContext::load_module()`. (4) Delete `compile_ptx_for_device()`, cubin disk cache, and all runtime compilation code.
|
|
|
|
**Tech Stack:** CUDA nvcc (build-time), cudarc 0.19.3 `load_module()`, Rust `build.rs`, `include_bytes!`.
|
|
|
|
---
|
|
|
|
## Scope: 25 kernel files, 19 Rust files, 1 build.rs
|
|
|
|
### Kernel files (`.cu`) — parameterize `#define` symbols
|
|
|
|
| Kernel | Symbols to parameterize |
|
|
|--------|------------------------|
|
|
| `experience_kernels.cu` | STATE_DIM, MARKET_DIM, PORTFOLIO_DIM |
|
|
| `c51_loss_kernel.cu` | NUM_ATOMS |
|
|
| `mse_loss_kernel.cu` | NUM_ATOMS |
|
|
| `ppo_experience_kernel.cu` | STATE_DIM, MARKET_DIM, NUM_ACTIONS, CUR_INPUT, CUR_HIDDEN, CUR_OUTPUT, DIVERSITY_WINDOW |
|
|
| `curiosity_training_kernel.cu` | MARKET_DIM, DQN_NUM_ACTIONS, DQN_ORDER_ACTIONS, DQN_URGENCY_ACTIONS, CUR_INPUT, CUR_HIDDEN, CUR_OUTPUT |
|
|
| `backtest_metrics_kernel.cu` | DQN_NUM_ACTIONS, DQN_ORDER_ACTIONS, DQN_URGENCY_ACTIONS |
|
|
| `backtest_forward_ppo_kernel.cu` | STATE_DIM, NUM_ACTIONS |
|
|
| `attention_kernel.cu` | STATE_DIM |
|
|
| `attention_backward_kernel.cu` | STATE_DIM |
|
|
| `dqn_utility_kernels.cu` | STATE_DIM |
|
|
| `her_relabel_kernel.cu` | STATE_DIM |
|
|
| `iql_value_kernel.cu` | STATE_DIM |
|
|
| `iqn_dual_head_kernel.cu` | STATE_DIM |
|
|
| `nstep_kernel.cu` | STATE_DIM |
|
|
| `monitoring_kernel.cu` | DQN_NUM_ACTIONS, DQN_ORDER_ACTIONS, DQN_URGENCY_ACTIONS |
|
|
| `backtest_env_kernel.cu` | (none) |
|
|
| `backtest_forward_supervised_kernel.cu` | (none) |
|
|
| `backtest_gather_kernel.cu` | (none) |
|
|
| `dt_kernels.cu` | NUM_ACTIONS (local, not from common header) |
|
|
| `ensemble_kernels.cu` | (none) |
|
|
| `epsilon_greedy_kernel.cu` | (none) |
|
|
| `her_episode_kernel.cu` | (none) |
|
|
| `signal_adapter_kernel.cu` | (none) |
|
|
| `statistics_kernel.cu` | (none) |
|
|
| `training_guard_kernel.cu` | (none) |
|
|
|
|
### `common_device_functions.cuh` — symbols become parameters
|
|
|
|
The following `#define` symbols become kernel parameters:
|
|
- `STATE_DIM` → `int state_dim` (varies: 48/72/80)
|
|
- `MARKET_DIM` → `int market_dim` (always 42 but parameterized for flexibility)
|
|
- `PORTFOLIO_DIM` → `int portfolio_dim` (always 8)
|
|
- `NUM_ATOMS` → `int num_atoms` (varies: 51/101)
|
|
- `NUM_ACTIONS` / `DQN_NUM_ACTIONS` → `int num_actions` (always 9 for DQN, 45 for PPO)
|
|
- `DQN_ORDER_ACTIONS` → `int order_actions` (always 3)
|
|
- `DQN_URGENCY_ACTIONS` → `int urgency_actions` (always 3)
|
|
- `CUR_INPUT` / `CUR_HIDDEN` / `CUR_OUTPUT` → `int cur_input, int cur_hidden, int cur_output`
|
|
- `DIVERSITY_WINDOW` → `int diversity_window` (always 100)
|
|
|
|
Fixed constants that stay as `#define` in `common_device_functions.cuh`:
|
|
- Helper functions (`argmax_n`, `lcg_random`, `gpu_random`, etc.)
|
|
- `PORTFOLIO_STRIDE` (20) — layout constant, never changes
|
|
- `OFI_DIM` (0) — compile-time feature toggle, stays as define
|
|
|
|
---
|
|
|
|
## Task 1: Create `build.rs` infrastructure — precompile ONE kernel as proof-of-concept
|
|
|
|
**Files:**
|
|
- Modify: `crates/ml/build.rs`
|
|
- Modify: `crates/ml/Cargo.toml` (add `cc` build dependency if needed)
|
|
|
|
The `build.rs` compiles `.cu` files to cubins using nvcc and writes them to `$OUT_DIR/`. Start with `epsilon_greedy_kernel.cu` — it has NO `#define` dependencies, so it's the simplest case.
|
|
|
|
- [ ] **Step 1: Implement build.rs kernel compilation**
|
|
|
|
```rust
|
|
// crates/ml/build.rs
|
|
use std::path::{Path, PathBuf};
|
|
use std::process::Command;
|
|
|
|
fn main() {
|
|
println!("cargo:rerun-if-changed=build.rs");
|
|
|
|
// Only compile CUDA when the cuda feature is enabled
|
|
if std::env::var("CARGO_FEATURE_CUDA").is_err() {
|
|
return;
|
|
}
|
|
|
|
let out_dir = PathBuf::from(std::env::var("OUT_DIR").unwrap());
|
|
let kernel_dir = Path::new("src/cuda_pipeline");
|
|
let common_header = kernel_dir.join("common_device_functions.cuh");
|
|
|
|
// Detect GPU architecture from env or default to sm_80
|
|
let cuda_compute_cap = std::env::var("CUDA_COMPUTE_CAP").unwrap_or_else(|_| "80".to_string());
|
|
let arch = format!("sm_{cuda_compute_cap}");
|
|
|
|
// List of kernels to precompile (no #define dependencies)
|
|
let simple_kernels = [
|
|
"epsilon_greedy_kernel.cu",
|
|
"backtest_env_kernel.cu",
|
|
"backtest_forward_supervised_kernel.cu",
|
|
"dt_kernels.cu",
|
|
"ensemble_kernels.cu",
|
|
"her_episode_kernel.cu",
|
|
"signal_adapter_kernel.cu",
|
|
"statistics_kernel.cu",
|
|
"training_guard_kernel.cu",
|
|
];
|
|
|
|
for kernel_name in &simple_kernels {
|
|
let kernel_path = kernel_dir.join(kernel_name);
|
|
let cubin_name = kernel_name.replace(".cu", ".cubin");
|
|
let cubin_path = out_dir.join(&cubin_name);
|
|
|
|
println!("cargo:rerun-if-changed={}", kernel_path.display());
|
|
|
|
// Compose source: common header + kernel
|
|
let common_src = std::fs::read_to_string(&common_header)
|
|
.unwrap_or_default();
|
|
let kernel_src = std::fs::read_to_string(&kernel_path)
|
|
.expect(&format!("Failed to read {}", kernel_path.display()));
|
|
let full_source = format!("{common_src}\n{kernel_src}");
|
|
|
|
// Write composed source to temp file
|
|
let tmp_src = out_dir.join(format!("_{kernel_name}"));
|
|
std::fs::write(&tmp_src, &full_source).unwrap();
|
|
|
|
// Compile with nvcc
|
|
let status = Command::new("nvcc")
|
|
.args([
|
|
"-cubin",
|
|
&format!("-arch={arch}"),
|
|
"-O3",
|
|
"--use_fast_math",
|
|
"--ftz=true",
|
|
"--fmad=true",
|
|
"-o", cubin_path.to_str().unwrap(),
|
|
tmp_src.to_str().unwrap(),
|
|
])
|
|
.status()
|
|
.expect("nvcc not found — install CUDA toolkit");
|
|
|
|
if !status.success() {
|
|
panic!("nvcc failed to compile {kernel_name}");
|
|
}
|
|
|
|
eprintln!(" Compiled {kernel_name} → {cubin_name} ({arch})");
|
|
}
|
|
|
|
// Also rerun if common header changes
|
|
println!("cargo:rerun-if-changed={}", common_header.display());
|
|
}
|
|
```
|
|
|
|
- [ ] **Step 2: Verify epsilon_greedy_kernel.cu compiles at build time**
|
|
|
|
Run: `CUDA_COMPUTE_CAP=86 SQLX_OFFLINE=true cargo build -p ml --features cuda`
|
|
Expected: "Compiled epsilon_greedy_kernel.cu → epsilon_greedy_kernel.cubin" in build output
|
|
|
|
- [ ] **Step 3: Load precompiled cubin in `gpu_action_selector.rs`**
|
|
|
|
Replace the `OnceLock` + `compile_ptx_for_device()` pattern:
|
|
|
|
```rust
|
|
// Before:
|
|
static EPSILON_GREEDY_PTX: OnceLock<Result<Ptx, String>> = OnceLock::new();
|
|
fn compile_kernel_ptx(context: &CudaContext) -> Result<Ptx, String> { ... }
|
|
let ptx_result = EPSILON_GREEDY_PTX.get_or_init(|| compile_kernel_ptx(&context));
|
|
|
|
// After:
|
|
static EPSILON_GREEDY_CUBIN: &[u8] = include_bytes!(concat!(env!("OUT_DIR"), "/epsilon_greedy_kernel.cubin"));
|
|
// In constructor:
|
|
let module = context.load_module(EPSILON_GREEDY_CUBIN)
|
|
.map_err(|e| MLError::ModelError(format!("load epsilon_greedy cubin: {e}")))?;
|
|
```
|
|
|
|
- [ ] **Step 4: Build + test**
|
|
|
|
Run: `CUDA_COMPUTE_CAP=86 SQLX_OFFLINE=true cargo test -p ml --lib -- gpu_action_selector`
|
|
Expected: tests pass using precompiled cubin
|
|
|
|
- [ ] **Step 5: Commit**
|
|
|
|
```bash
|
|
git commit -m "feat: build.rs precompiles epsilon_greedy_kernel.cu — proof of concept for NVRTC elimination"
|
|
```
|
|
|
|
---
|
|
|
|
## Task 2: Parameterize `#define` constants in kernel sources
|
|
|
|
**Files:**
|
|
- Modify: 15 `.cu` kernel files (those with `#define` dependencies)
|
|
- Modify: `common_device_functions.cuh` (remove runtime-specific `#define` symbols)
|
|
|
|
For each kernel that uses compile-time `#define` symbols, add them as kernel parameters instead. The kernels still use the symbols internally — they just come from parameters not preprocessor.
|
|
|
|
**Pattern for each kernel:**
|
|
|
|
```cuda
|
|
// Before:
|
|
#define STATE_DIM 72
|
|
extern "C" __global__ void my_kernel(float* data, int n) {
|
|
float x = data[threadIdx.x * STATE_DIM];
|
|
}
|
|
|
|
// After:
|
|
extern "C" __global__ void my_kernel(float* data, int n, int state_dim) {
|
|
float x = data[threadIdx.x * state_dim];
|
|
}
|
|
```
|
|
|
|
- [ ] **Step 1: Parameterize `experience_kernels.cu`**
|
|
|
|
This kernel uses `STATE_DIM`, `MARKET_DIM`, `PORTFOLIO_DIM`. These are already passed partially (e.g., `market_dim` is a kernel param in `experience_env_step`). Find all usages of the `#define` versions and replace with the existing params or add new ones.
|
|
|
|
Search: `grep -n "STATE_DIM\|MARKET_DIM\|PORTFOLIO_DIM" experience_kernels.cu`
|
|
|
|
For `experience_state_gather`: add `int state_dim` param, replace `STATE_DIM` usage.
|
|
For `experience_action_select`: no `STATE_DIM` usage (already checked).
|
|
For `experience_env_step`: `state_dim` is already a param. Check if `STATE_DIM` is used directly.
|
|
|
|
- [ ] **Step 2: Parameterize `c51_loss_kernel.cu` and `mse_loss_kernel.cu`**
|
|
|
|
Replace `NUM_ATOMS` with `int num_atoms` kernel parameter. These kernels are called from `gpu_dqn_trainer.rs` which already has `num_atoms` available.
|
|
|
|
- [ ] **Step 3: Parameterize `backtest_metrics_kernel.cu`**
|
|
|
|
Replace `DQN_NUM_ACTIONS`, `DQN_ORDER_ACTIONS`, `DQN_URGENCY_ACTIONS` with kernel params. These are always 9, 3, 3 but should be parameterized for clean compilation.
|
|
|
|
- [ ] **Step 4: Parameterize remaining kernels**
|
|
|
|
Apply the same pattern to: `attention_kernel.cu`, `attention_backward_kernel.cu`, `curiosity_training_kernel.cu`, `dqn_utility_kernels.cu`, `her_relabel_kernel.cu`, `iql_value_kernel.cu`, `iqn_dual_head_kernel.cu`, `monitoring_kernel.cu`, `nstep_kernel.cu`, `ppo_experience_kernel.cu`, `backtest_forward_ppo_kernel.cu`.
|
|
|
|
For each:
|
|
1. Find `#define` symbol usages with `grep`
|
|
2. Add as kernel parameter
|
|
3. Replace macro usage with parameter name
|
|
|
|
- [ ] **Step 5: Update `common_device_functions.cuh`**
|
|
|
|
Remove the `#define` symbols that are now kernel params:
|
|
- Remove: `STATE_DIM`, `MARKET_DIM`, `PORTFOLIO_DIM` (if any default exists)
|
|
- Remove: `NUM_ATOMS`, `NUM_ACTIONS`, `DQN_NUM_ACTIONS`, etc.
|
|
- Keep: helper functions, `PORTFOLIO_STRIDE`, layout constants
|
|
|
|
Note: `common_device_functions.cuh` still gets prepended to kernel sources at build time (build.rs does `{common_src}\n{kernel_src}`). It just no longer has runtime-variable `#define` symbols.
|
|
|
|
- [ ] **Step 6: Build + test**
|
|
|
|
Run: `SQLX_OFFLINE=true cargo check -p ml --lib`
|
|
Expected: compiles (runtime compilation still active for most kernels — just testing that the `.cu` sources are valid without `#define`)
|
|
|
|
- [ ] **Step 7: Commit**
|
|
|
|
```bash
|
|
git commit -m "refactor: convert #define constants to kernel parameters in all .cu files
|
|
|
|
STATE_DIM, MARKET_DIM, NUM_ATOMS, etc. are now kernel params instead of
|
|
compile-time #define. Enables single-cubin-per-kernel precompilation."
|
|
```
|
|
|
|
---
|
|
|
|
## Task 3: Update all Rust launch sites to pass parameterized values
|
|
|
|
**Files:**
|
|
- Modify: All 19 Rust files in `crates/ml/src/cuda_pipeline/` that launch kernels
|
|
|
|
For each kernel launch site that previously prepended `#define` strings:
|
|
|
|
```rust
|
|
// Before:
|
|
let defines = format!("#define STATE_DIM {state_dim}\n#define MARKET_DIM 42\n");
|
|
let common_src = include_str!("common_device_functions.cuh");
|
|
let kernel_src = include_str!("my_kernel.cu");
|
|
let full_source = format!("{defines}{common_src}\n{kernel_src}");
|
|
let ptx = compile_ptx_for_device(&full_source, &context)?;
|
|
|
|
// After:
|
|
// (kernel loaded from precompiled cubin — no runtime compilation)
|
|
// Just pass the values as kernel args:
|
|
.arg(&state_dim_i32)
|
|
.arg(&market_dim_i32)
|
|
```
|
|
|
|
- [ ] **Step 1: Update `gpu_experience_collector.rs`**
|
|
|
|
Remove the `#define STATE_DIM ...` string composition. The experience kernels now accept `state_dim`, `market_dim`, `portfolio_dim` as parameters. Add `.arg(&state_dim_i32)` etc. to the launch calls for `experience_state_gather` and `experience_env_step`.
|
|
|
|
- [ ] **Step 2: Update `gpu_dqn_trainer.rs`**
|
|
|
|
Remove `#define STATE_DIM ...` and `#define NUM_ATOMS ...` string composition. Pass `state_dim` and `num_atoms` as kernel params to C51/MSE loss kernels.
|
|
|
|
- [ ] **Step 3: Update `gpu_backtest_evaluator.rs`**
|
|
|
|
Remove the `compile_metrics_ptx`, `compile_gather_ptx`, `compile_env_ptx` functions that compose `#define` strings. The kernels will be loaded from precompiled cubins.
|
|
|
|
- [ ] **Step 4: Update remaining launch sites**
|
|
|
|
Apply the same pattern to: `gpu_attention.rs`, `gpu_curiosity_trainer.rs`, `gpu_her.rs`, `gpu_iqn_head.rs`, `gpu_iql_trainer.rs`, `gpu_monitoring.rs`, `gpu_portfolio.rs`, `gpu_ppo_collector.rs`, `decision_transformer.rs`, `signal_adapter.rs`, `gpu_statistics.rs`, `gpu_training_guard.rs`, `batched_forward.rs`, `batched_backward.rs`.
|
|
|
|
- [ ] **Step 5: Build + test**
|
|
|
|
Run: `SQLX_OFFLINE=true cargo check -p ml --lib`
|
|
Run: `SQLX_OFFLINE=true cargo test -p ml --lib`
|
|
Expected: all tests pass
|
|
|
|
- [ ] **Step 6: Commit**
|
|
|
|
```bash
|
|
git commit -m "refactor: pass state_dim/num_atoms/etc as kernel args instead of #define strings"
|
|
```
|
|
|
|
---
|
|
|
|
## Task 4: Extend `build.rs` to precompile ALL 25 kernels
|
|
|
|
**Files:**
|
|
- Modify: `crates/ml/build.rs`
|
|
|
|
Now that all kernels accept their dimensions as parameters (no `#define` dependencies), all 25 can be compiled by the same build.rs.
|
|
|
|
- [ ] **Step 1: Add all kernels to build.rs compilation list**
|
|
|
|
```rust
|
|
let all_kernels = [
|
|
"epsilon_greedy_kernel.cu",
|
|
"backtest_env_kernel.cu",
|
|
"backtest_forward_supervised_kernel.cu",
|
|
"backtest_forward_ppo_kernel.cu",
|
|
"backtest_gather_kernel.cu",
|
|
"backtest_metrics_kernel.cu",
|
|
"dt_kernels.cu",
|
|
"ensemble_kernels.cu",
|
|
"her_episode_kernel.cu",
|
|
"her_relabel_kernel.cu",
|
|
"signal_adapter_kernel.cu",
|
|
"statistics_kernel.cu",
|
|
"training_guard_kernel.cu",
|
|
"experience_kernels.cu",
|
|
"c51_loss_kernel.cu",
|
|
"mse_loss_kernel.cu",
|
|
"curiosity_training_kernel.cu",
|
|
"dqn_utility_kernels.cu",
|
|
"attention_kernel.cu",
|
|
"attention_backward_kernel.cu",
|
|
"iql_value_kernel.cu",
|
|
"iqn_dual_head_kernel.cu",
|
|
"monitoring_kernel.cu",
|
|
"nstep_kernel.cu",
|
|
"ppo_experience_kernel.cu",
|
|
];
|
|
```
|
|
|
|
- [ ] **Step 2: Add `cargo:rerun-if-changed` for all `.cu` files**
|
|
|
|
- [ ] **Step 3: Build and verify all 25 cubins are produced**
|
|
|
|
Run: `CUDA_COMPUTE_CAP=86 SQLX_OFFLINE=true cargo build -p ml --features cuda 2>&1 | grep "Compiled"`
|
|
Expected: 25 "Compiled ... → ... .cubin" lines
|
|
|
|
- [ ] **Step 4: Commit**
|
|
|
|
```bash
|
|
git commit -m "feat: build.rs precompiles all 25 CUDA kernels at cargo build time"
|
|
```
|
|
|
|
---
|
|
|
|
## Task 5: Replace all runtime compilation with precompiled cubin loading
|
|
|
|
**Files:**
|
|
- Modify: All 19 Rust files in `crates/ml/src/cuda_pipeline/`
|
|
- Modify: `crates/ml/src/cuda_pipeline/mod.rs` (remove `compile_ptx_for_device` re-export)
|
|
|
|
For each file, replace the `OnceLock<Ptx>` + `compile_ptx_for_device()` pattern with `include_bytes!` + `load_module()`.
|
|
|
|
**Template:**
|
|
|
|
```rust
|
|
// Before (per file):
|
|
static KERNEL_PTX: OnceLock<Result<Ptx, String>> = OnceLock::new();
|
|
fn compile_kernel(ctx: &CudaContext) -> Result<Ptx, String> {
|
|
let defines = "#define STATE_DIM ...";
|
|
let common = include_str!("common_device_functions.cuh");
|
|
let src = include_str!("my_kernel.cu");
|
|
compile_ptx_for_device(&format!("{defines}{common}\n{src}"), ctx)
|
|
}
|
|
// Usage:
|
|
let ptx = KERNEL_PTX.get_or_init(|| compile_kernel(&ctx))
|
|
.as_ref().map_err(|e| ...)?;
|
|
let module = ctx.load_module(ptx.clone())?;
|
|
|
|
// After:
|
|
static KERNEL_CUBIN: &[u8] = include_bytes!(concat!(env!("OUT_DIR"), "/my_kernel.cubin"));
|
|
// Usage:
|
|
let module = ctx.load_module(KERNEL_CUBIN)
|
|
.map_err(|e| MLError::ModelError(format!("load my_kernel cubin: {e}")))?;
|
|
```
|
|
|
|
- [ ] **Step 1: Migrate `gpu_experience_collector.rs`** (most critical — training hot path)
|
|
- [ ] **Step 2: Migrate `gpu_dqn_trainer.rs`** (C51/MSE loss, ensemble, CQL kernels)
|
|
- [ ] **Step 3: Migrate `gpu_backtest_evaluator.rs`** (6 kernel loads)
|
|
- [ ] **Step 4: Migrate `gpu_action_selector.rs`** (already done in Task 1)
|
|
- [ ] **Step 5: Migrate remaining 15 files**
|
|
|
|
Apply to: `gpu_attention.rs`, `gpu_curiosity_trainer.rs`, `gpu_her.rs`, `gpu_iqn_head.rs`, `gpu_iql_trainer.rs`, `gpu_monitoring.rs`, `gpu_portfolio.rs`, `gpu_ppo_collector.rs`, `decision_transformer.rs`, `signal_adapter.rs`, `gpu_statistics.rs`, `gpu_training_guard.rs`, `batched_forward.rs`, `batched_backward.rs`, `mod.rs`.
|
|
|
|
- [ ] **Step 6: Build + test**
|
|
|
|
Run: `SQLX_OFFLINE=true cargo test -p ml --lib`
|
|
Expected: all 887+ tests pass using precompiled cubins
|
|
|
|
- [ ] **Step 7: Commit**
|
|
|
|
```bash
|
|
git commit -m "feat: all kernels loaded from precompiled cubins — zero runtime nvcc/NVRTC"
|
|
```
|
|
|
|
---
|
|
|
|
## Task 6: Delete runtime compilation infrastructure
|
|
|
|
**Files:**
|
|
- Modify: `crates/ml-core/src/cuda_compile.rs` (delete or gut)
|
|
- Modify: `crates/ml/src/cuda_pipeline/mod.rs` (remove re-export)
|
|
- Delete: `/tmp/.cubin_cache/` references
|
|
|
|
- [ ] **Step 1: Remove `compile_ptx_for_device()` function**
|
|
|
|
In `ml-core/src/cuda_compile.rs`, delete the function body. Keep the module if other utilities are used, otherwise delete the file.
|
|
|
|
- [ ] **Step 2: Remove cubin disk cache functions**
|
|
|
|
Delete `load_cached_cubin()`, `compile_cubin_with_nvcc()`, `cache_dir()`, and all SHA-256 cache key logic.
|
|
|
|
- [ ] **Step 3: Remove re-export in `mod.rs`**
|
|
|
|
In `crates/ml/src/cuda_pipeline/mod.rs`, remove:
|
|
```rust
|
|
pub use ml_core::cuda_compile::compile_ptx_for_device;
|
|
```
|
|
|
|
- [ ] **Step 4: Remove all `OnceLock<Result<Ptx, String>>` statics**
|
|
|
|
Search: `grep -rn "OnceLock.*Ptx" crates/ml/src/cuda_pipeline/ --include="*.rs"`
|
|
Remove each `static` declaration and associated `compile_*_ptx()` function.
|
|
|
|
- [ ] **Step 5: Remove `cudarc::nvrtc` imports where no longer needed**
|
|
|
|
Search: `grep -rn "nvrtc::Ptx\|use cudarc::nvrtc" crates/ml/ --include="*.rs"`
|
|
|
|
- [ ] **Step 6: Build + full test suite**
|
|
|
|
Run: `SQLX_OFFLINE=true cargo test -p ml --lib`
|
|
Run: `SQLX_OFFLINE=true cargo test -p ml-dqn --lib`
|
|
Run: `SQLX_OFFLINE=true cargo check --workspace`
|
|
Expected: everything clean, zero references to nvrtc/compile_ptx_for_device
|
|
|
|
- [ ] **Step 7: Verify no cubin cache usage**
|
|
|
|
Run: `grep -rn "cubin_cache\|\.cubin_cache" crates/ --include="*.rs"`
|
|
Expected: zero matches
|
|
|
|
- [ ] **Step 8: Commit**
|
|
|
|
```bash
|
|
git commit -m "refactor: delete all runtime NVRTC/nvcc compilation infrastructure
|
|
|
|
Removed: compile_ptx_for_device(), cubin disk cache, OnceLock<Ptx> statics.
|
|
All kernels now load from build-time precompiled cubins via include_bytes!.
|
|
Zero runtime compiler invocations. Zero disk cache. Zero stale-cache segfaults."
|
|
```
|
|
|
|
---
|
|
|
|
## Task 7: Integration test — verify zero runtime compilation
|
|
|
|
**Files:**
|
|
- No new files — validation only
|
|
|
|
- [ ] **Step 1: Full test suite**
|
|
|
|
Run: `SQLX_OFFLINE=true cargo test -p ml --lib`
|
|
Expected: 887+ pass
|
|
|
|
- [ ] **Step 2: Verify no nvcc at runtime**
|
|
|
|
Run: `strace -f -e trace=execve cargo test -p ml --lib --release -- test_metrics_ptx_compilation 2>&1 | grep nvcc`
|
|
Expected: zero nvcc executions (only at build time)
|
|
|
|
- [ ] **Step 3: Verify no cubin cache reads**
|
|
|
|
Run: `strace -f -e trace=openat cargo test -p ml --lib --release -- test_metrics_ptx_compilation 2>&1 | grep cubin_cache`
|
|
Expected: zero references
|
|
|
|
- [ ] **Step 4: Push and verify CI**
|
|
|
|
```bash
|
|
git push origin main
|
|
```
|
|
Expected: CI passes on H100 (build.rs compiles cubins at sm_90)
|
|
|
|
---
|
|
|
|
## Execution Order
|
|
|
|
```
|
|
Task 1 (build.rs POC) → Task 2 (parameterize #define) → Task 3 (update launch sites)
|
|
→ Task 4 (build all 25) → Task 5 (replace runtime with precompiled)
|
|
→ Task 6 (delete NVRTC) → Task 7 (integration test)
|
|
```
|
|
|
|
Each task produces a working, testable state. Tasks 1-3 can be done without breaking existing runtime compilation (both paths coexist). Task 5 switches over. Task 6 removes the old path.
|
|
|
|
## Risks
|
|
|
|
- **nvcc required at build time**: CI and dev machines need CUDA toolkit installed. The `build.rs` should gracefully skip compilation when nvcc is absent (non-CUDA builds).
|
|
- **Architecture mismatch**: Build machine GPU arch must match target (or use `CUDA_COMPUTE_CAP` env var). CI builds for sm_90 (H100), local dev for sm_86 (RTX 3050).
|
|
- **Binary size**: 25 embedded cubins add ~1-2 MB to the binary (cubins are typically 5-50 KB each). Negligible.
|
|
- **Shared memory sizing**: Some kernels use `STATE_DIM` in shared memory declarations (`__shared__ float buf[STATE_DIM]`). Converting to dynamic shared memory (`extern __shared__`) is required for these cases since the parameter isn't known at compile time. Use `shmem_bytes` in `LaunchConfig` to allocate dynamically.
|