# GPU Test Fixes + Feature Extraction Pipeline Fix > **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking. **Goal:** Fix all 14 failing GPU tests so they pass locally on RTX 3050, and fix OFI=false in the Argo training pipeline. **Architecture:** Six independent bug fixes targeting test assertions, GPU buffer sizes, CUDA init, training profile config, and fxcache cache-key matching. Each fix is isolated — no cross-dependencies. **Tech Stack:** Rust, CUDA (cudarc), TOML configs, cargo test --- ## File Map | File | Changes | |------|---------| | `crates/ml-core/src/gpu/profile.rs` | Update 3 stale test assertions | | `crates/ml/tests/smoke_test_real_data.rs` | Fix memcpy host vec sizes | | `crates/ml/src/trainers/dqn/smoke_tests/gpu_residency.rs` | Investigate + fix CUDA panic | | `config/training/dqn-smoketest.toml` | Change `data_source` from `mbp10` to `ohlcv`, remove `mbp10_data_dir` | | `crates/ml/tests/dqn_action_collapse_fix_test.rs` | Update `cql_alpha` assertion to match default (1.0) | | `crates/ml/src/training_profile.rs` | Update `test_load_embedded_dqn_smoketest` epochs assertion, update `test_apply_to_dqn_smoketest` | --- ### Task 1: Fix GPU profile test assertions (Bug 1) **Files:** - Modify: `crates/ml-core/src/gpu/profile.rs:291-308` - [ ] **Step 1: Update A100 test assertions** In `crates/ml-core/src/gpu/profile.rs`, change `test_embedded_a100_parses` (lines 292-298) from: ```rust #[test] fn test_embedded_a100_parses() { let profile: GpuProfile = toml::from_str(A100_TOML).unwrap(); assert_eq!(profile.training.batch_size, 512); assert_eq!(profile.training.num_atoms, 51); assert_eq!(profile.training.buffer_size, 200_000); assert_eq!(profile.cuda.cuda_stack_bytes, 65536); } ``` to: ```rust #[test] fn test_embedded_a100_parses() { let profile: GpuProfile = toml::from_str(A100_TOML).unwrap(); assert_eq!(profile.training.batch_size, 2048); assert_eq!(profile.training.num_atoms, 51); assert_eq!(profile.training.buffer_size, 0); // 0 = auto from VRAM assert_eq!(profile.experience.gpu_timesteps_per_episode, 500); assert_eq!(profile.cuda.cuda_stack_bytes, 65536); } ``` - [ ] **Step 2: Update default profile test assertion** Change `test_embedded_default_parses` (lines 301-308) — only `buffer_size` changed: ```rust #[test] fn test_embedded_default_parses() { let profile: GpuProfile = toml::from_str(DEFAULT_TOML).unwrap(); assert_eq!(profile.training.batch_size, 256); assert_eq!(profile.training.num_atoms, 21); assert_eq!(profile.training.buffer_size, 0); // 0 = auto from VRAM assert_eq!(profile.experience.gpu_timesteps_per_episode, 200); assert_eq!(profile.cuda.cuda_stack_bytes, 32768); } ``` - [ ] **Step 3: Update hardcoded fallback test** The `test_hardcoded_fallback_values` test (lines 334-341) asserts `buffer_size == 50_000`. Check if `hardcoded_fallback()` also returns `buffer_size = 0` now. If so, update: ```rust #[test] fn test_hardcoded_fallback_values() { let profile = GpuProfile::hardcoded_fallback(); assert_eq!(profile.training.batch_size, 256); assert_eq!(profile.training.num_atoms, 21); assert_eq!(profile.training.buffer_size, 0); // 0 = auto from VRAM assert_eq!(profile.experience.gpu_timesteps_per_episode, 200); assert_eq!(profile.cuda.cuda_stack_bytes, 32768); } ``` If `hardcoded_fallback()` still returns 50_000, leave the test as-is. - [ ] **Step 4: Run tests** ```bash SQLX_OFFLINE=true cargo test -p ml-core --lib -- gpu::profile -v ``` Expected: all profile tests pass. - [ ] **Step 5: Commit** ```bash git add crates/ml-core/src/gpu/profile.rs git commit -m "fix: update GPU profile test assertions to match current TOML values" ``` --- ### Task 2: Fix GPU smoke test memcpy buffer size (Bug 2) **Files:** - Modify: `crates/ml/tests/smoke_test_real_data.rs` - [ ] **Step 1: Fix `smoke_gpu_real_ohlcv_forward` memcpy** In `crates/ml/tests/smoke_test_real_data.rs`, find `smoke_gpu_real_ohlcv_forward` (around line 519). Replace lines 571-575: ```rust // Download actions and rewards from GPU to host for validation let mut actions_host = vec![0i32; expected_total]; stream.memcpy_dtoh(&batch.actions, &mut actions_host).unwrap(); let mut rewards_host = vec![0.0f32; expected_total]; stream.memcpy_dtoh(&batch.rewards, &mut rewards_host).unwrap(); ``` with: ```rust // Download actions and rewards from GPU to host for validation let mut actions_host = vec![0i32; batch.actions.len()]; stream.memcpy_dtoh(&batch.actions, &mut actions_host).unwrap(); let mut rewards_host = vec![0.0f32; batch.rewards.len()]; stream.memcpy_dtoh(&batch.rewards, &mut rewards_host).unwrap(); ``` - [ ] **Step 2: Fix `smoke_gpu_real_data_noisy_distributional` memcpy** Find `smoke_gpu_real_data_noisy_distributional` (around line 670). Apply the same pattern — replace any `vec![...; expected_total]` used as memcpy destination with `vec![...; batch..len()]`. - [ ] **Step 3: Check for other memcpy_dtoh calls in the file** Search for all `memcpy_dtoh` calls in the file. For each one, verify the host vec is sized from the GPU slice's `.len()`, not from a pre-calculated `expected_total`. Fix any others found. ```bash grep -n "memcpy_dtoh" crates/ml/tests/smoke_test_real_data.rs ``` - [ ] **Step 4: Run tests** ```bash SQLX_OFFLINE=true FOXHUNT_TEST_DATA=test_data/futures-baseline cargo test -p ml --test smoke_test_real_data -- smoke_gpu_real_ohlcv_forward smoke_gpu_real_data_noisy_distributional --nocapture ``` Expected: both tests pass without cudarc panics. - [ ] **Step 5: Commit** ```bash git add crates/ml/tests/smoke_test_real_data.rs git commit -m "fix: size memcpy host vecs from GPU buffer length (not expected_total)" ``` --- ### Task 3: Fix GPU residency test panics (Bug 3) **Files:** - Modify: `crates/ml/src/trainers/dqn/smoke_tests/gpu_residency.rs` - [ ] **Step 1: Run the tests locally to identify exact failure** ```bash SQLX_OFFLINE=true cargo test -p ml --lib -- trainers::dqn::smoke_tests::gpu_residency --nocapture 2>&1 | tail -40 ``` Observe the exact panic message. The tests use `smoke_stream()` → `cuda_device()` → `MlDevice::new_cuda(0)`. Possible outcomes: - (a) CUDA init failure → fix device init - (b) Buffer size mismatch in `sample_proportional` or `update_priorities_gpu_raw` → fix buffer allocation - (c) Kernel launch failure → fix kernel params - [ ] **Step 2: Fix based on diagnosis** If the panic is in `GpuReplayBuffer::new()` or `sample_proportional()` — read the error and fix the root cause (likely a buffer allocation size mismatch similar to Bug 2). If the panic is in `smoke_stream()` / `cuda_device()` — the fix is in `helpers.rs`. `MlDevice::new_cuda(0)` should work on RTX 3050. If it panics, check if another test has corrupted the CUDA context (the comment on line 44 warns about this). - [ ] **Step 3: Run tests again** ```bash SQLX_OFFLINE=true cargo test -p ml --lib -- trainers::dqn::smoke_tests::gpu_residency --nocapture ``` Expected: all 4 tests pass (plus `test_gpu_replay_buffer_insert_and_len` which was already passing). - [ ] **Step 4: Commit** ```bash git add crates/ml/src/trainers/dqn/smoke_tests/gpu_residency.rs git commit -m "fix: GPU residency tests — " ``` --- ### Task 4: Fix dqn-smoketest profile MBP-10 path (Bug 4) **Files:** - Modify: `config/training/dqn-smoketest.toml` - Modify: `crates/ml/src/training_profile.rs` - [ ] **Step 1: Remove MBP-10 data_source from smoketest profile** In `config/training/dqn-smoketest.toml`, change lines 23-26 from: ```toml data_source = "mbp10" mbp10_data_dir = "test_data/futures-baseline-mbp10" imbalance_bar_threshold = 1.0 imbalance_bar_ewma_alpha = 0.1 ``` to: ```toml imbalance_bar_threshold = 1.0 imbalance_bar_ewma_alpha = 0.1 ``` This removes `data_source` and `mbp10_data_dir` from the smoketest profile. The trainer will default to OHLCV data source, which is what smoketests need (fast, no MBP-10 dependency). - [ ] **Step 2: Update `test_load_embedded_dqn_smoketest` assertion** In `crates/ml/src/training_profile.rs`, the test at line 1028 asserts `epochs == 3` but the TOML has `epochs = 10`. Update: ```rust #[test] fn test_load_embedded_dqn_smoketest() { let p = DqnTrainingProfile::load("dqn-smoketest"); assert!(p.training.is_some(), "dqn-smoketest must have [training] section"); let t = p.training.unwrap(); assert_eq!(t.epochs, Some(10)); assert_eq!(t.batch_size, Some(64)); } ``` - [ ] **Step 3: Update `test_apply_to_dqn_smoketest` assertion** At line 1100-1102, the test asserts `!hp.mbp10_data_dir.is_empty()`. Since we removed `mbp10_data_dir` from the smoketest TOML, the profile no longer overrides it. Check what `DQNHyperparameters::conservative()` defaults `mbp10_data_dir` to. If it defaults to an empty string, update the assertion: ```rust // smoketest doesn't set mbp10_data_dir — uses conservative() default ``` If conservative() sets a non-empty default, the assertion still passes and needs no change. - [ ] **Step 4: Run affected tests** ```bash SQLX_OFFLINE=true cargo test -p ml --lib -- training_profile --nocapture SQLX_OFFLINE=true cargo test -p ml --test dqn_early_stopping_termination_test --nocapture ``` Expected: all training_profile tests pass, all 3 early_stopping tests pass. - [ ] **Step 5: Commit** ```bash git add config/training/dqn-smoketest.toml crates/ml/src/training_profile.rs git commit -m "fix: smoketest profile uses OHLCV (no MBP-10 dependency)" ``` --- ### Task 5: Fix CQL alpha test assertion (Bug 5) **Files:** - Modify: `crates/ml/tests/dqn_action_collapse_fix_test.rs` - [ ] **Step 1: Update CQL alpha default assertion** The `DQNHyperparameters::default()` sets `cql_alpha = 1.0` (at `crates/ml/src/trainers/dqn/config.rs:1672`), but the test at line 145 expects `0.1`. In `crates/ml/tests/dqn_action_collapse_fix_test.rs`, change lines 144-150: ```rust // use_cql and cql_alpha should exist and have correct defaults let alpha_diff = (hp.cql_alpha - 0.1).abs(); assert!( alpha_diff < 1e-6, "cql_alpha should default to 0.1, got {}", hp.cql_alpha ); ``` to: ```rust // CQL alpha defaults to 1.0 (strong conservatism for offline RL) let alpha_diff = (hp.cql_alpha - 1.0).abs(); assert!( alpha_diff < 1e-6, "cql_alpha should default to 1.0, got {}", hp.cql_alpha ); ``` - [ ] **Step 2: Run test** ```bash SQLX_OFFLINE=true cargo test -p ml --test dqn_action_collapse_fix_test -- test_cql_configurable_via_hyperparameters --nocapture ``` Expected: PASS. - [ ] **Step 3: Commit** ```bash git add crates/ml/tests/dqn_action_collapse_fix_test.rs git commit -m "fix: CQL alpha test assertion matches default (1.0)" ``` --- ### Task 6: Fix DQN training pipeline tests (Bug 7) **Files:** - Modify: potentially `crates/ml/tests/dqn_training_pipeline_test.rs` (diagnosis first) - [ ] **Step 1: Run the 5 failing tests locally** ```bash SQLX_OFFLINE=true FOXHUNT_TEST_DATA=test_data/futures-baseline cargo test -p ml --test dqn_training_pipeline_test --nocapture 2>&1 | tail -60 ``` Determine if they: - (a) Skip gracefully (return `Ok(())`) → they're not actually failing, CI had different data - (b) Fail with MBP-10 path error → fixed by Task 4's smoketest profile change - (c) Fail with other error → diagnose and fix - [ ] **Step 2: Fix if needed** If tests fail because they load `dqn-smoketest` profile internally (which Task 4 fixed), just verify they pass now. If they fail for another reason, trace the error and fix. - [ ] **Step 3: Run full integration test suite** ```bash SQLX_OFFLINE=true FOXHUNT_TEST_DATA=test_data/futures-baseline cargo test -p ml --test dqn_training_pipeline_test --nocapture ``` Expected: all 5 tests pass (not skip). - [ ] **Step 4: Commit (if changes made)** ```bash git add crates/ml/tests/dqn_training_pipeline_test.rs git commit -m "fix: DQN training pipeline tests pass with smoketest profile" ``` --- ### Task 7: Verify OFI fxcache loading (Pipeline fix) **Files:** - Possibly modify: `crates/ml/examples/train_baseline_rl.rs` (if cache key mismatch found) - [ ] **Step 1: Run precompute locally to generate fxcache with OFI** ```bash SQLX_OFFLINE=true cargo run -p ml --example precompute_features --release -- \ --data-dir test_data/futures-baseline \ --mbp10-data-dir test_data/futures-baseline-mbp10 \ --trades-data-dir test_data/futures-baseline-trades \ --output-dir /tmp/feature-cache-test \ --symbol ES.FUT \ --yes ``` Note the cache key and output path. Verify OFI is included (`OFI: 8-dim`). - [ ] **Step 2: Run training with `--feature-cache-dir` pointing to the cache** ```bash SQLX_OFFLINE=true cargo run -p ml --example train_baseline_rl --release -- \ --model dqn \ --data-dir test_data/futures-baseline \ --mbp10-data-dir test_data/futures-baseline-mbp10 \ --trades-data-dir test_data/futures-baseline-trades \ --feature-cache-dir /tmp/feature-cache-test \ --symbol ES.FUT \ --epochs 1 \ --output-dir /tmp/train-ofi-test ``` Check the output for `OFI=true` in the `init_from_fxcache` log line. - [ ] **Step 3: If OFI=false, diagnose cache key mismatch** Compare the cache key used by precompute vs train_baseline_rl: - Precompute passes `--data-dir`, `--mbp10-data-dir`, `--trades-data-dir` to `calculate_dbn_cache_key_full(symbol_dir, mbp10, trades)` - train_baseline_rl uses `args.data_dir.join(&args.symbol)` as `symbol_dir`, and `args.mbp10_data_dir` / `args.trades_data_dir` as optional paths If the paths are different (relative vs absolute, or different base dirs), canonicalize them before key calculation: ```rust let symbol_dir = args.data_dir.join(&args.symbol).canonicalize() .unwrap_or_else(|_| args.data_dir.join(&args.symbol)); ``` - [ ] **Step 4: Run smoke test with fxcache to verify OFI=true** ```bash FOXHUNT_TEST_DATA=test_data/futures-baseline FOXHUNT_FEATURE_CACHE_DIR=/tmp/feature-cache-test \ cargo test -p ml --lib -- smoke_tests --ignored --nocapture 2>&1 | grep -i ofi ``` Expected: `OFI=true` in init_from_fxcache output. - [ ] **Step 5: Commit (if changes made)** ```bash git add crates/ml/examples/train_baseline_rl.rs git commit -m "fix: canonicalize paths for fxcache key calculation (OFI=true)" ``` --- ### Task 8: Full verification pass - [ ] **Step 1: Run all GPU unit tests** ```bash SQLX_OFFLINE=true cargo test -p ml-core --lib -- gpu::profile -v SQLX_OFFLINE=true cargo test -p ml --lib -- smoke_tests -v SQLX_OFFLINE=true cargo test -p ml --lib -- training_profile -v ``` Expected: all pass. - [ ] **Step 2: Run all integration tests** ```bash SQLX_OFFLINE=true FOXHUNT_TEST_DATA=test_data/futures-baseline cargo test -p ml --test smoke_test_real_data --nocapture SQLX_OFFLINE=true FOXHUNT_TEST_DATA=test_data/futures-baseline cargo test -p ml --test dqn_early_stopping_termination_test --nocapture SQLX_OFFLINE=true cargo test -p ml --test dqn_action_collapse_fix_test --nocapture SQLX_OFFLINE=true FOXHUNT_TEST_DATA=test_data/futures-baseline cargo test -p ml --test dqn_training_pipeline_test --nocapture ``` Expected: all 14 previously-failing tests now pass. - [ ] **Step 3: Run ignored smoke tests** ```bash SQLX_OFFLINE=true FOXHUNT_TEST_DATA=test_data/futures-baseline cargo test -p ml --lib -- smoke_tests --ignored --nocapture ``` Expected: smoke tests pass on RTX 3050 with real data. - [ ] **Step 4: Push** ```bash git push origin main ```