16 KiB
GPU Test Fixes + Feature Extraction Pipeline Fix
For agentic workers: REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (
- [ ]) syntax for tracking.
Goal: Fix all 14 failing GPU tests so they pass locally on RTX 3050, and fix OFI=false in the Argo training pipeline.
Architecture: Six independent bug fixes targeting test assertions, GPU buffer sizes, CUDA init, training profile config, and fxcache cache-key matching. Each fix is isolated — no cross-dependencies.
Tech Stack: Rust, CUDA (cudarc), TOML configs, cargo test
File Map
| File | Changes |
|---|---|
crates/ml-core/src/gpu/profile.rs |
Update 3 stale test assertions |
crates/ml/tests/smoke_test_real_data.rs |
Fix memcpy host vec sizes |
crates/ml/src/trainers/dqn/smoke_tests/gpu_residency.rs |
Investigate + fix CUDA panic |
config/training/dqn-smoketest.toml |
Change data_source from mbp10 to ohlcv, remove mbp10_data_dir |
crates/ml/tests/dqn_action_collapse_fix_test.rs |
Update cql_alpha assertion to match default (1.0) |
crates/ml/src/training_profile.rs |
Update test_load_embedded_dqn_smoketest epochs assertion, update test_apply_to_dqn_smoketest |
Task 1: Fix GPU profile test assertions (Bug 1)
Files:
-
Modify:
crates/ml-core/src/gpu/profile.rs:291-308 -
Step 1: Update A100 test assertions
In crates/ml-core/src/gpu/profile.rs, change test_embedded_a100_parses (lines 292-298) from:
#[test]
fn test_embedded_a100_parses() {
let profile: GpuProfile = toml::from_str(A100_TOML).unwrap();
assert_eq!(profile.training.batch_size, 512);
assert_eq!(profile.training.num_atoms, 51);
assert_eq!(profile.training.buffer_size, 200_000);
assert_eq!(profile.cuda.cuda_stack_bytes, 65536);
}
to:
#[test]
fn test_embedded_a100_parses() {
let profile: GpuProfile = toml::from_str(A100_TOML).unwrap();
assert_eq!(profile.training.batch_size, 2048);
assert_eq!(profile.training.num_atoms, 51);
assert_eq!(profile.training.buffer_size, 0); // 0 = auto from VRAM
assert_eq!(profile.experience.gpu_timesteps_per_episode, 500);
assert_eq!(profile.cuda.cuda_stack_bytes, 65536);
}
- Step 2: Update default profile test assertion
Change test_embedded_default_parses (lines 301-308) — only buffer_size changed:
#[test]
fn test_embedded_default_parses() {
let profile: GpuProfile = toml::from_str(DEFAULT_TOML).unwrap();
assert_eq!(profile.training.batch_size, 256);
assert_eq!(profile.training.num_atoms, 21);
assert_eq!(profile.training.buffer_size, 0); // 0 = auto from VRAM
assert_eq!(profile.experience.gpu_timesteps_per_episode, 200);
assert_eq!(profile.cuda.cuda_stack_bytes, 32768);
}
- Step 3: Update hardcoded fallback test
The test_hardcoded_fallback_values test (lines 334-341) asserts buffer_size == 50_000. Check if hardcoded_fallback() also returns buffer_size = 0 now. If so, update:
#[test]
fn test_hardcoded_fallback_values() {
let profile = GpuProfile::hardcoded_fallback();
assert_eq!(profile.training.batch_size, 256);
assert_eq!(profile.training.num_atoms, 21);
assert_eq!(profile.training.buffer_size, 0); // 0 = auto from VRAM
assert_eq!(profile.experience.gpu_timesteps_per_episode, 200);
assert_eq!(profile.cuda.cuda_stack_bytes, 32768);
}
If hardcoded_fallback() still returns 50_000, leave the test as-is.
- Step 4: Run tests
SQLX_OFFLINE=true cargo test -p ml-core --lib -- gpu::profile -v
Expected: all profile tests pass.
- Step 5: Commit
git add crates/ml-core/src/gpu/profile.rs
git commit -m "fix: update GPU profile test assertions to match current TOML values"
Task 2: Fix GPU smoke test memcpy buffer size (Bug 2)
Files:
-
Modify:
crates/ml/tests/smoke_test_real_data.rs -
Step 1: Fix
smoke_gpu_real_ohlcv_forwardmemcpy
In crates/ml/tests/smoke_test_real_data.rs, find smoke_gpu_real_ohlcv_forward (around line 519). Replace lines 571-575:
// Download actions and rewards from GPU to host for validation
let mut actions_host = vec![0i32; expected_total];
stream.memcpy_dtoh(&batch.actions, &mut actions_host).unwrap();
let mut rewards_host = vec![0.0f32; expected_total];
stream.memcpy_dtoh(&batch.rewards, &mut rewards_host).unwrap();
with:
// Download actions and rewards from GPU to host for validation
let mut actions_host = vec![0i32; batch.actions.len()];
stream.memcpy_dtoh(&batch.actions, &mut actions_host).unwrap();
let mut rewards_host = vec![0.0f32; batch.rewards.len()];
stream.memcpy_dtoh(&batch.rewards, &mut rewards_host).unwrap();
- Step 2: Fix
smoke_gpu_real_data_noisy_distributionalmemcpy
Find smoke_gpu_real_data_noisy_distributional (around line 670). Apply the same pattern — replace any vec![...; expected_total] used as memcpy destination with vec![...; batch.<field>.len()].
- Step 3: Check for other memcpy_dtoh calls in the file
Search for all memcpy_dtoh calls in the file. For each one, verify the host vec is sized from the GPU slice's .len(), not from a pre-calculated expected_total. Fix any others found.
grep -n "memcpy_dtoh" crates/ml/tests/smoke_test_real_data.rs
- Step 4: Run tests
SQLX_OFFLINE=true FOXHUNT_TEST_DATA=test_data/futures-baseline cargo test -p ml --test smoke_test_real_data -- smoke_gpu_real_ohlcv_forward smoke_gpu_real_data_noisy_distributional --nocapture
Expected: both tests pass without cudarc panics.
- Step 5: Commit
git add crates/ml/tests/smoke_test_real_data.rs
git commit -m "fix: size memcpy host vecs from GPU buffer length (not expected_total)"
Task 3: Fix GPU residency test panics (Bug 3)
Files:
-
Modify:
crates/ml/src/trainers/dqn/smoke_tests/gpu_residency.rs -
Step 1: Run the tests locally to identify exact failure
SQLX_OFFLINE=true cargo test -p ml --lib -- trainers::dqn::smoke_tests::gpu_residency --nocapture 2>&1 | tail -40
Observe the exact panic message. The tests use smoke_stream() → cuda_device() → MlDevice::new_cuda(0). Possible outcomes:
-
(a) CUDA init failure → fix device init
-
(b) Buffer size mismatch in
sample_proportionalorupdate_priorities_gpu_raw→ fix buffer allocation -
(c) Kernel launch failure → fix kernel params
-
Step 2: Fix based on diagnosis
If the panic is in GpuReplayBuffer::new() or sample_proportional() — read the error and fix the root cause (likely a buffer allocation size mismatch similar to Bug 2).
If the panic is in smoke_stream() / cuda_device() — the fix is in helpers.rs. MlDevice::new_cuda(0) should work on RTX 3050. If it panics, check if another test has corrupted the CUDA context (the comment on line 44 warns about this).
- Step 3: Run tests again
SQLX_OFFLINE=true cargo test -p ml --lib -- trainers::dqn::smoke_tests::gpu_residency --nocapture
Expected: all 4 tests pass (plus test_gpu_replay_buffer_insert_and_len which was already passing).
- Step 4: Commit
git add crates/ml/src/trainers/dqn/smoke_tests/gpu_residency.rs
git commit -m "fix: GPU residency tests — <describe actual fix>"
Task 4: Fix dqn-smoketest profile MBP-10 path (Bug 4)
Files:
-
Modify:
config/training/dqn-smoketest.toml -
Modify:
crates/ml/src/training_profile.rs -
Step 1: Remove MBP-10 data_source from smoketest profile
In config/training/dqn-smoketest.toml, change lines 23-26 from:
data_source = "mbp10"
mbp10_data_dir = "test_data/futures-baseline-mbp10"
imbalance_bar_threshold = 1.0
imbalance_bar_ewma_alpha = 0.1
to:
imbalance_bar_threshold = 1.0
imbalance_bar_ewma_alpha = 0.1
This removes data_source and mbp10_data_dir from the smoketest profile. The trainer will default to OHLCV data source, which is what smoketests need (fast, no MBP-10 dependency).
- Step 2: Update
test_load_embedded_dqn_smoketestassertion
In crates/ml/src/training_profile.rs, the test at line 1028 asserts epochs == 3 but the TOML has epochs = 10. Update:
#[test]
fn test_load_embedded_dqn_smoketest() {
let p = DqnTrainingProfile::load("dqn-smoketest");
assert!(p.training.is_some(), "dqn-smoketest must have [training] section");
let t = p.training.unwrap();
assert_eq!(t.epochs, Some(10));
assert_eq!(t.batch_size, Some(64));
}
- Step 3: Update
test_apply_to_dqn_smoketestassertion
At line 1100-1102, the test asserts !hp.mbp10_data_dir.is_empty(). Since we removed mbp10_data_dir from the smoketest TOML, the profile no longer overrides it. Check what DQNHyperparameters::conservative() defaults mbp10_data_dir to. If it defaults to an empty string, update the assertion:
// smoketest doesn't set mbp10_data_dir — uses conservative() default
If conservative() sets a non-empty default, the assertion still passes and needs no change.
- Step 4: Run affected tests
SQLX_OFFLINE=true cargo test -p ml --lib -- training_profile --nocapture
SQLX_OFFLINE=true cargo test -p ml --test dqn_early_stopping_termination_test --nocapture
Expected: all training_profile tests pass, all 3 early_stopping tests pass.
- Step 5: Commit
git add config/training/dqn-smoketest.toml crates/ml/src/training_profile.rs
git commit -m "fix: smoketest profile uses OHLCV (no MBP-10 dependency)"
Task 5: Fix CQL alpha test assertion (Bug 5)
Files:
-
Modify:
crates/ml/tests/dqn_action_collapse_fix_test.rs -
Step 1: Update CQL alpha default assertion
The DQNHyperparameters::default() sets cql_alpha = 1.0 (at crates/ml/src/trainers/dqn/config.rs:1672), but the test at line 145 expects 0.1.
In crates/ml/tests/dqn_action_collapse_fix_test.rs, change lines 144-150:
// use_cql and cql_alpha should exist and have correct defaults
let alpha_diff = (hp.cql_alpha - 0.1).abs();
assert!(
alpha_diff < 1e-6,
"cql_alpha should default to 0.1, got {}",
hp.cql_alpha
);
to:
// CQL alpha defaults to 1.0 (strong conservatism for offline RL)
let alpha_diff = (hp.cql_alpha - 1.0).abs();
assert!(
alpha_diff < 1e-6,
"cql_alpha should default to 1.0, got {}",
hp.cql_alpha
);
- Step 2: Run test
SQLX_OFFLINE=true cargo test -p ml --test dqn_action_collapse_fix_test -- test_cql_configurable_via_hyperparameters --nocapture
Expected: PASS.
- Step 3: Commit
git add crates/ml/tests/dqn_action_collapse_fix_test.rs
git commit -m "fix: CQL alpha test assertion matches default (1.0)"
Task 6: Fix DQN training pipeline tests (Bug 7)
Files:
-
Modify: potentially
crates/ml/tests/dqn_training_pipeline_test.rs(diagnosis first) -
Step 1: Run the 5 failing tests locally
SQLX_OFFLINE=true FOXHUNT_TEST_DATA=test_data/futures-baseline cargo test -p ml --test dqn_training_pipeline_test --nocapture 2>&1 | tail -60
Determine if they:
-
(a) Skip gracefully (return
Ok(())) → they're not actually failing, CI had different data -
(b) Fail with MBP-10 path error → fixed by Task 4's smoketest profile change
-
(c) Fail with other error → diagnose and fix
-
Step 2: Fix if needed
If tests fail because they load dqn-smoketest profile internally (which Task 4 fixed), just verify they pass now. If they fail for another reason, trace the error and fix.
- Step 3: Run full integration test suite
SQLX_OFFLINE=true FOXHUNT_TEST_DATA=test_data/futures-baseline cargo test -p ml --test dqn_training_pipeline_test --nocapture
Expected: all 5 tests pass (not skip).
- Step 4: Commit (if changes made)
git add crates/ml/tests/dqn_training_pipeline_test.rs
git commit -m "fix: DQN training pipeline tests pass with smoketest profile"
Task 7: Verify OFI fxcache loading (Pipeline fix)
Files:
-
Possibly modify:
crates/ml/examples/train_baseline_rl.rs(if cache key mismatch found) -
Step 1: Run precompute locally to generate fxcache with OFI
SQLX_OFFLINE=true cargo run -p ml --example precompute_features --release -- \
--data-dir test_data/futures-baseline \
--mbp10-data-dir test_data/futures-baseline-mbp10 \
--trades-data-dir test_data/futures-baseline-trades \
--output-dir /tmp/feature-cache-test \
--symbol ES.FUT \
--yes
Note the cache key and output path. Verify OFI is included (OFI: 8-dim).
- Step 2: Run training with
--feature-cache-dirpointing to the cache
SQLX_OFFLINE=true cargo run -p ml --example train_baseline_rl --release -- \
--model dqn \
--data-dir test_data/futures-baseline \
--mbp10-data-dir test_data/futures-baseline-mbp10 \
--trades-data-dir test_data/futures-baseline-trades \
--feature-cache-dir /tmp/feature-cache-test \
--symbol ES.FUT \
--epochs 1 \
--output-dir /tmp/train-ofi-test
Check the output for OFI=true in the init_from_fxcache log line.
- Step 3: If OFI=false, diagnose cache key mismatch
Compare the cache key used by precompute vs train_baseline_rl:
- Precompute passes
--data-dir,--mbp10-data-dir,--trades-data-dirtocalculate_dbn_cache_key_full(symbol_dir, mbp10, trades) - train_baseline_rl uses
args.data_dir.join(&args.symbol)assymbol_dir, andargs.mbp10_data_dir/args.trades_data_diras optional paths
If the paths are different (relative vs absolute, or different base dirs), canonicalize them before key calculation:
let symbol_dir = args.data_dir.join(&args.symbol).canonicalize()
.unwrap_or_else(|_| args.data_dir.join(&args.symbol));
- Step 4: Run smoke test with fxcache to verify OFI=true
FOXHUNT_TEST_DATA=test_data/futures-baseline FOXHUNT_FEATURE_CACHE_DIR=/tmp/feature-cache-test \
cargo test -p ml --lib -- smoke_tests --ignored --nocapture 2>&1 | grep -i ofi
Expected: OFI=true in init_from_fxcache output.
- Step 5: Commit (if changes made)
git add crates/ml/examples/train_baseline_rl.rs
git commit -m "fix: canonicalize paths for fxcache key calculation (OFI=true)"
Task 8: Full verification pass
- Step 1: Run all GPU unit tests
SQLX_OFFLINE=true cargo test -p ml-core --lib -- gpu::profile -v
SQLX_OFFLINE=true cargo test -p ml --lib -- smoke_tests -v
SQLX_OFFLINE=true cargo test -p ml --lib -- training_profile -v
Expected: all pass.
- Step 2: Run all integration tests
SQLX_OFFLINE=true FOXHUNT_TEST_DATA=test_data/futures-baseline cargo test -p ml --test smoke_test_real_data --nocapture
SQLX_OFFLINE=true FOXHUNT_TEST_DATA=test_data/futures-baseline cargo test -p ml --test dqn_early_stopping_termination_test --nocapture
SQLX_OFFLINE=true cargo test -p ml --test dqn_action_collapse_fix_test --nocapture
SQLX_OFFLINE=true FOXHUNT_TEST_DATA=test_data/futures-baseline cargo test -p ml --test dqn_training_pipeline_test --nocapture
Expected: all 14 previously-failing tests now pass.
- Step 3: Run ignored smoke tests
SQLX_OFFLINE=true FOXHUNT_TEST_DATA=test_data/futures-baseline cargo test -p ml --lib -- smoke_tests --ignored --nocapture
Expected: smoke tests pass on RTX 3050 with real data.
- Step 4: Push
git push origin main