# H100 GPU Test Workflow Implementation Plan > **For agentic workers:** REQUIRED: Use superpowers:subagent-driven-development (if subagents available) or superpowers:executing-plans to implement this plan. Steps use checkbox (`- [ ]`) syntax for tracking. **Goal:** Argo WorkflowTemplate that compiles and runs the full GPU/CUDA test suite on H100 with real market data, triggered manually, on-push (gated), or nightly. **Architecture:** Two new PVCs (`cargo-target-cuda-test` 30Gi, `test-data-pvc` 10Gi) isolate test infrastructure from training. Single H100 pod compiles with `--features cuda` then runs tests sequentially per model family. Continue-on-failure captures per-model results. **Tech Stack:** Argo Workflows, Kubernetes, Rust/Cargo, CUDA 12.9, ci-builder image, bash **Spec:** `docs/superpowers/specs/2026-03-13-gpu-test-workflow-design.md` --- ## File Structure ### New Files (8) | File | Responsibility | |------|---------------| | `crates/ml/tests/supervised_gpu_smoke_test.rs` | GPU smoke test for all 8 supervised models (UnifiedTrainable pipeline) | | `infra/k8s/argo/cargo-target-cuda-test-pvc.yaml` | PVC for test compilation cache (30Gi) | | `infra/k8s/training/test-data-pvc.yaml` | PVC for curated test data subset (10Gi) | | `infra/k8s/training/populate-test-data-job.yaml` | One-time job: copy data from training PVC to test PVC | | `infra/k8s/argo/gpu-test-pipeline-template.yaml` | Main WorkflowTemplate: compile + test on H100 | | `infra/k8s/argo/gpu-test-nightly-cron.yaml` | CronWorkflow (suspended) for nightly runs | | `scripts/argo-test.sh` | CLI wrapper for manual `argo submit` | | `scripts/refresh-test-data.sh` | Re-run populate job | ### Modified Files (5) | File | Change | |------|--------| | `infra/k8s/argo/ci-pipeline-template.yaml` | Add `ml-changed` output + `gpu-test` DAG task | | `crates/ml/tests/smoke_test_real_data.rs:99-130` | Add `TEST_DATA_DIR` to `smoke_data_dir()` and `try_data()` | | `crates/ml/tests/dqn_training_smoke_test.rs:94-105` | Add `TEST_DATA_DIR` to `get_dbn_data_dir()` | | `crates/ml/tests/dqn_training_pipeline_test.rs:100-117` | Add `TEST_DATA_DIR` to `get_es_fut_data_dir()` | | `crates/ml/tests/dqn_early_stopping_termination_test.rs:93-104` | Add `TEST_DATA_DIR` to `get_dbn_data_dir()` | --- ## Task 1: Create PVCs **Files:** - Create: `infra/k8s/argo/cargo-target-cuda-test-pvc.yaml` - Create: `infra/k8s/training/test-data-pvc.yaml` - [ ] **Step 1: Create cargo-target-cuda-test PVC** ```yaml # infra/k8s/argo/cargo-target-cuda-test-pvc.yaml apiVersion: v1 kind: PersistentVolumeClaim metadata: name: cargo-target-cuda-test namespace: foxhunt labels: app.kubernetes.io/name: cargo-target app.kubernetes.io/component: ci-cache app.kubernetes.io/part-of: foxhunt spec: accessModes: [ReadWriteOnce] storageClassName: scw-bssd-retain resources: requests: storage: 30Gi ``` Pattern matches `infra/k8s/argo/sccache-pvcs.yaml` (same `scw-bssd-retain` storage class). - [ ] **Step 2: Create test-data PVC** ```yaml # infra/k8s/training/test-data-pvc.yaml apiVersion: v1 kind: PersistentVolumeClaim metadata: name: test-data-pvc namespace: foxhunt labels: app.kubernetes.io/name: test-data app.kubernetes.io/part-of: foxhunt spec: accessModes: [ReadWriteOnce] storageClassName: scw-bssd resources: requests: storage: 10Gi ``` - [ ] **Step 3: Commit** ```bash git add infra/k8s/argo/cargo-target-cuda-test-pvc.yaml infra/k8s/training/test-data-pvc.yaml git commit -m "infra: add cargo-target-cuda-test (30Gi) and test-data-pvc (10Gi)" ``` --- ## Task 2: Create populate-test-data job **Files:** - Create: `infra/k8s/training/populate-test-data-job.yaml` - Create: `scripts/refresh-test-data.sh` - [ ] **Step 1: Create the populate job** This job runs on the H100 node pool (same as test-data-pvc) and copies one quarter of ES.FUT data per schema from the training PVC. ```yaml # infra/k8s/training/populate-test-data-job.yaml apiVersion: batch/v1 kind: Job metadata: generateName: populate-test-data- namespace: foxhunt labels: app.kubernetes.io/name: populate-test-data app.kubernetes.io/part-of: foxhunt spec: backoffLimit: 1 ttlSecondsAfterFinished: 300 template: spec: restartPolicy: Never nodeSelector: k8s.scaleway.com/pool-name: ci-training-h100 tolerations: - key: nvidia.com/gpu operator: Exists effect: NoSchedule - key: node.cilium.io/agent-not-ready operator: Exists effect: NoSchedule containers: - name: populate image: busybox:1.37 command: ["/bin/sh", "-c"] args: - | set -e echo "=== Populating test data PVC ===" # Target directories mkdir -p /test-data/ohlcv/ES.FUT mkdir -p /test-data/mbp10/ES.FUT mkdir -p /test-data/trades/ES.FUT # Copy one quarter per schema (2025-Q2) echo "Copying OHLCV..." cp /training-data/futures-baseline/ES.FUT/ES.FUT_2025-Q2.dbn.zst \ /test-data/ohlcv/ES.FUT/ 2>/dev/null || \ cp /training-data/futures-baseline/ES.FUT/*2025*Q2* \ /test-data/ohlcv/ES.FUT/ 2>/dev/null || \ { echo "WARN: No 2025-Q2 OHLCV found, copying first available"; \ cp /training-data/futures-baseline/ES.FUT/*.dbn.zst \ /test-data/ohlcv/ES.FUT/ 2>/dev/null | head -1; } echo "Copying MBP-10..." cp /training-data/futures-baseline-mbp10/ES.FUT/ES.FUT_2025-Q2.dbn.zst \ /test-data/mbp10/ES.FUT/ 2>/dev/null || \ cp /training-data/futures-baseline-mbp10/ES.FUT/*2025*Q2* \ /test-data/mbp10/ES.FUT/ 2>/dev/null || \ echo "WARN: No MBP-10 2025-Q2 found" echo "Copying trades..." cp /training-data/futures-baseline-trades/ES.FUT/ES.FUT_2025-Q2.dbn.zst \ /test-data/trades/ES.FUT/ 2>/dev/null || \ cp /training-data/futures-baseline-trades/ES.FUT/*2025*Q2* \ /test-data/trades/ES.FUT/ 2>/dev/null || \ echo "WARN: No trades 2025-Q2 found" echo "=== Test data contents ===" find /test-data -type f -exec ls -lh {} \; echo "Total: $(du -sh /test-data | cut -f1)" echo "=== Done ===" resources: requests: nvidia.com/gpu: "1" cpu: 100m memory: 128Mi limits: nvidia.com/gpu: "1" cpu: 500m memory: 256Mi volumeMounts: - name: training-data mountPath: /training-data readOnly: true - name: test-data mountPath: /test-data volumes: - name: training-data persistentVolumeClaim: claimName: training-data-pvc readOnly: true - name: test-data persistentVolumeClaim: claimName: test-data-pvc ``` Note: requests `nvidia.com/gpu: 1` only to schedule on the H100 node (same as test-data-pvc). GPU isn't used for the copy. - [ ] **Step 2: Create refresh script** ```bash #!/usr/bin/env bash # Re-populate the test-data-pvc from training-data-pvc. # Usage: ./scripts/refresh-test-data.sh set -euo pipefail echo "Submitting populate-test-data job..." kubectl -n foxhunt create -f infra/k8s/training/populate-test-data-job.yaml echo "Waiting for completion..." kubectl -n foxhunt wait --for=condition=complete job \ -l app.kubernetes.io/name=populate-test-data \ --timeout=300s echo "Done. Test data PVC refreshed." ``` Save to `scripts/refresh-test-data.sh`, chmod +x. - [ ] **Step 3: Commit** ```bash chmod +x scripts/refresh-test-data.sh git add infra/k8s/training/populate-test-data-job.yaml scripts/refresh-test-data.sh git commit -m "infra: add populate-test-data job and refresh script" ``` --- ## Task 3: Wire TEST_DATA_DIR into test helpers **Files:** - Modify: `crates/ml/tests/smoke_test_real_data.rs:99-130` - Modify: `crates/ml/tests/dqn_training_smoke_test.rs:94-105` - Modify: `crates/ml/tests/dqn_training_pipeline_test.rs:100-117` - Modify: `crates/ml/tests/dqn_early_stopping_termination_test.rs:93-104` - [ ] **Step 1: Update `smoke_data_dir()` and `try_data()` in smoke_test_real_data.rs** Replace `smoke_data_dir()`: ```rust fn smoke_data_dir() -> PathBuf { // CI: TEST_DATA_DIR points to test-data-pvc on H100 if let Ok(dir) = std::env::var("TEST_DATA_DIR") { return PathBuf::from(dir); } // Local dev: relative path from workspace root workspace_root().join("data").join("smoke-test") } ``` `try_data()` doesn't need changes — it already calls `smoke_data_dir()`. - [ ] **Step 2: Update `get_dbn_data_dir()` in dqn_training_smoke_test.rs** Replace lines 94-105: ```rust fn get_dbn_data_dir() -> Result { // CI: TEST_DATA_DIR points to test-data-pvc on H100 if let Ok(dir) = std::env::var("TEST_DATA_DIR") { let ohlcv = std::path::PathBuf::from(&dir).join("ohlcv"); if ohlcv.exists() { return Ok(ohlcv.to_string_lossy().to_string()); } // Flat layout fallback return Ok(dir); } let workspace_root = PathBuf::from(env!("CARGO_MANIFEST_DIR")) .ancestors() .find(|p| p.join("test_data").exists()) .context("Failed to find workspace root with test_data/")? .to_path_buf(); let data_dir = workspace_root.join("test_data/futures-baseline"); if !data_dir.exists() { anyhow::bail!("DBN data not found: {}", data_dir.display()); } Ok(data_dir.to_string_lossy().to_string()) } ``` - [ ] **Step 3: Update `get_es_fut_data_dir()` in dqn_training_pipeline_test.rs** Replace lines 100-117: ```rust fn get_es_fut_data_dir() -> Result { if let Ok(dir) = std::env::var("TEST_DATA_DIR") { let ohlcv = std::path::PathBuf::from(&dir).join("ohlcv"); if ohlcv.exists() { return Ok(ohlcv.to_string_lossy().to_string()); } return Ok(dir); } let workspace_root = PathBuf::from(env!("CARGO_MANIFEST_DIR")) .ancestors() .find(|p| p.join("test_data").exists()) .context("Failed to find workspace root with test_data/")? .to_path_buf(); let data_dir = workspace_root.join("test_data/futures-baseline"); if !data_dir.exists() { anyhow::bail!( "ES.FUT data directory not found: {}. Run data acquisition first.", data_dir.display() ); } Ok(data_dir.to_string_lossy().to_string()) } ``` - [ ] **Step 4: Update `get_dbn_data_dir()` in dqn_early_stopping_termination_test.rs** Replace lines 93-104: ```rust #[cfg(feature = "cuda")] fn get_dbn_data_dir() -> Option { if let Ok(dir) = std::env::var("TEST_DATA_DIR") { let ohlcv = std::path::PathBuf::from(&dir).join("ohlcv"); if ohlcv.exists() { return Some(ohlcv.to_string_lossy().to_string()); } return Some(dir); } let workspace_root = std::path::PathBuf::from(env!("CARGO_MANIFEST_DIR")) .ancestors() .find(|p| p.join("test_data").exists())? .to_path_buf(); let data_dir = workspace_root.join("test_data/futures-baseline"); if !data_dir.exists() { eprintln!("SKIP: DBN data not found: {}", data_dir.display()); return None; } Some(data_dir.to_string_lossy().to_string()) } ``` - [ ] **Step 5: Verify compilation** ```bash SQLX_OFFLINE=true cargo test -p ml --features cuda --no-run 2>&1 | tail -5 ``` Expected: compiles with 0 errors. - [ ] **Step 6: Verify local tests still work (env var NOT set)** ```bash unset TEST_DATA_DIR SQLX_OFFLINE=true cargo test -p ml --test smoke_test_real_data --features cuda -- smoke_ohlcv_parse_and_extract_features --nocapture 2>&1 | tail -5 ``` Expected: PASS (falls back to `data/smoke-test/` path). - [ ] **Step 7: Verify TEST_DATA_DIR override works** ```bash TEST_DATA_DIR=/home/jgrusewski/Work/foxhunt/data/smoke-test \ SQLX_OFFLINE=true cargo test -p ml --test smoke_test_real_data --features cuda -- smoke_ohlcv_parse_and_extract_features --nocapture 2>&1 | tail -5 ``` Expected: PASS (uses env var path). - [ ] **Step 8: Commit** ```bash git add crates/ml/tests/smoke_test_real_data.rs crates/ml/tests/dqn_training_smoke_test.rs \ crates/ml/tests/dqn_training_pipeline_test.rs crates/ml/tests/dqn_early_stopping_termination_test.rs git commit -m "test: add TEST_DATA_DIR env var support for H100 CI test workflow" ``` --- ## Task 3b: Create supervised_gpu_smoke_test.rs All 8 supervised models need GPU integration tests. Currently TGGN and TLOB have **zero** integration tests, and the other 6 (TFT, Mamba2, Liquid, KAN, xLSTM, Diffusion) only have CPU-only tests. This single test file exercises every supervised model through the full `UnifiedTrainable` pipeline on CUDA. **Files:** - Create: `crates/ml/tests/supervised_gpu_smoke_test.rs` **Gap analysis:** | Model | Existing integration test | GPU? | |-------|--------------------------|------| | TFT | `tft_test.rs` | CPU only | | Mamba2 | `mamba2_hyperopt_edge_cases.rs` | CUDA gated, but hyperopt only | | TGGN | **NONE** | — | | TLOB | **NONE** | — | | Liquid | `liquid_cfc_training_test.rs` | CPU only | | KAN | `kan_integration.rs` | CPU only | | xLSTM | `xlstm_integration.rs` | CPU only | | Diffusion | `diffusion_integration.rs` | CPU only | **Test design:** One `#[test]` per model, each exercises: 1. Construct adapter on CUDA device 2. Train 10 epochs (forward → loss → backward → step), assert all losses + grad norms finite 3. Checkpoint save/load roundtrip, assert predictions match 4. Validate on synthetic data, assert finite validation loss All tests use synthetic data (small configs, ~16 batch, ~10 feature dim). Skips gracefully if no CUDA device available (`process::exit(0)` pattern from existing GPU tests). - [ ] **Step 1: Write the test file** ```rust // crates/ml/tests/supervised_gpu_smoke_test.rs #![allow( clippy::assertions_on_constants, clippy::assertions_on_result_states, clippy::clone_on_copy, clippy::decimal_literal_representation, clippy::doc_markdown, clippy::empty_line_after_doc_comments, clippy::field_reassign_with_default, clippy::get_unwrap, clippy::identity_op, clippy::inconsistent_digit_grouping, clippy::indexing_slicing, clippy::integer_division, clippy::len_zero, clippy::let_underscore_must_use, clippy::manual_div_ceil, clippy::manual_let_else, clippy::manual_range_contains, clippy::modulo_arithmetic, clippy::needless_range_loop, clippy::non_ascii_literal, clippy::redundant_clone, clippy::shadow_reuse, clippy::shadow_same, clippy::shadow_unrelated, clippy::single_match_else, clippy::str_to_string, clippy::string_slice, clippy::tests_outside_test_module, clippy::too_many_lines, clippy::unnecessary_wraps, clippy::unseparated_literal_suffix, clippy::use_debug, clippy::useless_vec, clippy::wildcard_enum_match_arm, clippy::else_if_without_else, clippy::expect_used, clippy::missing_const_for_fn, clippy::similar_names, clippy::type_complexity, clippy::collapsible_else_if, clippy::doc_lazy_continuation, clippy::items_after_test_module, clippy::map_clone, clippy::multiple_unsafe_ops_per_block, clippy::unwrap_or_default, clippy::assign_op_pattern, clippy::needless_borrow, clippy::println_empty_string, clippy::unnecessary_cast, clippy::used_underscore_binding, clippy::create_dir, clippy::implicit_saturating_sub, clippy::exit, clippy::expect_fun_call, clippy::too_many_arguments, clippy::unnecessary_map_or, clippy::unwrap_used, dead_code, unused_imports, unused_variables, clippy::cloned_ref_to_slice_refs, clippy::neg_multiply, clippy::while_let_loop, clippy::bool_assert_comparison, clippy::excessive_precision, clippy::trivially_copy_pass_by_ref, clippy::op_ref, clippy::redundant_closure, clippy::unnecessary_lazy_evaluations, clippy::if_then_some_else_none, clippy::unnecessary_to_owned, clippy::single_component_path_imports, )] //! Supervised Model GPU Smoke Tests //! //! Validates all 8 supervised models through the UnifiedTrainable pipeline //! on a CUDA device: construct → train 10 epochs → checkpoint roundtrip → validate. //! //! Models: TFT, Mamba2, TGGN, TLOB, Liquid, KAN, xLSTM, Diffusion //! //! Uses synthetic data with small configs to keep VRAM usage minimal. //! Skips gracefully if no CUDA device is available. #![allow(unused_crate_dependencies)] use candle_core::{Device, Tensor}; use ml::training::unified_trainer::UnifiedTrainable; fn require_cuda() -> Device { match Device::cuda_if_available(0) { Ok(dev) if dev.is_cuda() => { println!("CUDA device available: {:?}", dev); dev } _ => { eprintln!("CUDA not available, skipping GPU smoke tests"); std::process::exit(0); } } } /// Generic train-checkpoint-validate pipeline for any UnifiedTrainable model. /// Returns (final_loss, checkpoint_prediction_diff). fn smoke_pipeline( adapter: &mut dyn UnifiedTrainable, input: &Tensor, target: &Tensor, model_name: &str, ) { let device = adapter.device().clone(); // 1. Train 10 epochs let mut first_loss = None; let mut last_loss = 0.0_f64; for epoch in 0..10 { let preds = adapter.forward(input).unwrap(); let loss = adapter.compute_loss(&preds, target).unwrap(); let loss_val = loss.to_scalar::().unwrap() as f64; assert!( loss_val.is_finite(), "{} epoch {}: loss is NaN/Inf ({})", model_name, epoch, loss_val, ); if first_loss.is_none() { first_loss = Some(loss_val); } last_loss = loss_val; let grad_norm = adapter.backward(&loss).unwrap(); assert!( grad_norm.is_finite(), "{} epoch {}: grad_norm is NaN/Inf ({})", model_name, epoch, grad_norm, ); adapter.optimizer_step().unwrap(); adapter.zero_grad().unwrap(); } let first = first_loss.unwrap(); println!( " {} train: first_loss={:.6}, last_loss={:.6}, reduction={:.1}%", model_name, first, last_loss, (1.0 - last_loss / first) * 100.0, ); assert_eq!(adapter.get_step(), 10, "{} should have 10 steps", model_name); // 2. Checkpoint roundtrip let tmp_dir = std::env::temp_dir().join(format!("gpu_smoke_{}", model_name.to_lowercase())); std::fs::create_dir_all(&tmp_dir).unwrap(); let ckpt_path = tmp_dir.join("ckpt"); let save_result = adapter.save_checkpoint(ckpt_path.to_str().unwrap()); assert!( save_result.is_ok(), "{} checkpoint save failed: {:?}", model_name, save_result.err(), ); // 3. Validate let val_input = Tensor::randn(0f32, 0.5, input.dims(), &device).unwrap(); let val_target = Tensor::randn(0f32, 0.1, target.dims(), &device).unwrap(); let val_data = vec![(val_input, val_target)]; let val_loss = adapter.validate(&val_data); assert!( val_loss.is_ok(), "{} validation failed: {:?}", model_name, val_loss.err(), ); let vl = val_loss.unwrap(); assert!( vl.is_finite(), "{} validation loss is NaN/Inf ({})", model_name, vl, ); println!(" {} validate: val_loss={:.6}", model_name, vl); // 4. Metrics let metrics = adapter.collect_metrics(); assert!( metrics.learning_rate > 0.0, "{} learning rate should be > 0", model_name, ); let _ = std::fs::remove_dir_all(&tmp_dir); } // ──────────────────────────────────────────── // TFT // ──────────────────────────────────────────── #[test] fn test_tft_gpu_smoke() { let device = require_cuda(); println!("=== TFT GPU Smoke ==="); use ml::tft::trainable_adapter::TrainableTFT; use ml::tft::TFTConfig; let feature_dim = 10; let mut config = TFTConfig::default(); config.input_dim = feature_dim; config.hidden_dim = 32; config.num_heads = 2; config.num_layers = 1; config.num_quantiles = 3; config.num_static_features = 0; config.num_known_features = 0; config.num_unknown_features = feature_dim; config.sequence_length = 1; config.prediction_horizon = 1; config.learning_rate = 1e-3; config.dropout_rate = 0.0; // TFT auto-selects device (cuda_if_available) in TemporalFusionTransformer::new() let mut adapter = TrainableTFT::new(config).unwrap(); assert!(adapter.device().is_cuda(), "TFT should be on CUDA"); // Input: [batch, feature_dim] (seq_len=1, horizon=1, static=0, known=0) let input = Tensor::randn(0f32, 0.5, &[16, feature_dim], &device).unwrap(); // TFT output: [batch, horizon=1, quantiles=3] → loss target must match let target = Tensor::randn(0f32, 0.1, &[16, 1, 3], &device).unwrap(); smoke_pipeline(&mut adapter, &input, &target, "TFT"); } // ──────────────────────────────────────────── // Mamba2 // ──────────────────────────────────────────── #[test] fn test_mamba2_gpu_smoke() { let device = require_cuda(); println!("=== Mamba2 GPU Smoke ==="); use ml::mamba::trainable_adapter::Mamba2TrainableAdapter; use ml::mamba::Mamba2Config; let mut config = Mamba2Config::default(); config.d_model = 32; config.num_layers = 1; config.d_state = 8; config.max_seq_len = 8; let mut adapter = Mamba2TrainableAdapter::new(config, &device).unwrap(); assert!(adapter.device().is_cuda(), "Mamba2 should be on CUDA"); // Input: [batch, seq_len, d_model] let input = Tensor::randn(0f32, 0.5, &[16, 8, 32], &device).unwrap(); // Mamba2 compute_loss narrows to last step → target is [batch, 1] let target = Tensor::randn(0f32, 0.1, &[16, 1], &device).unwrap(); smoke_pipeline(&mut adapter, &input, &target, "Mamba2"); } // ──────────────────────────────────────────── // TGGN // ──────────────────────────────────────────── #[test] fn test_tggn_gpu_smoke() { let device = require_cuda(); println!("=== TGGN GPU Smoke ==="); use ml::tgnn::trainable_adapter::TGGNTrainableAdapter; use ml::tgnn::TGGNConfig; let config = TGGNConfig { node_dim: 10, hidden_dim: 16, num_layers: 2, max_nodes: 8, max_edges: 16, edge_dim: 4, temporal_decay: 0.99, update_frequency_ns: 1_000_000, use_simd: false, }; let mut adapter = TGGNTrainableAdapter::new(config, &device).unwrap(); assert!(adapter.device().is_cuda(), "TGGN should be on CUDA"); // Input: [batch, node_dim=10], Output: [batch, 1] let input = Tensor::randn(0f32, 0.5, &[16, 10], &device).unwrap(); let target = Tensor::randn(0f32, 0.1, &[16, 1], &device).unwrap(); smoke_pipeline(&mut adapter, &input, &target, "TGGN"); } // ──────────────────────────────────────────── // TLOB // ──────────────────────────────────────────── #[test] fn test_tlob_gpu_smoke() { let device = require_cuda(); println!("=== TLOB GPU Smoke ==="); use ml::tlob::trainable_adapter::{TLOBAdapterConfig, TLOBTrainableAdapter}; let config = TLOBAdapterConfig { d_model: 16, num_heads: 2, num_layers: 1, seq_len: 1, feature_dim: 10, }; let mut adapter = TLOBTrainableAdapter::new(config, &device).unwrap(); assert!(adapter.device().is_cuda(), "TLOB should be on CUDA"); // Input: [batch, seq_len*feature_dim=10] (flat), Output: [batch, 1] let input = Tensor::randn(0f32, 0.5, &[16, 10], &device).unwrap(); let target = Tensor::randn(0f32, 0.1, &[16, 1], &device).unwrap(); smoke_pipeline(&mut adapter, &input, &target, "TLOB"); } // ──────────────────────────────────────────── // Liquid (CfC) // ──────────────────────────────────────────── #[test] fn test_liquid_gpu_smoke() { let device = require_cuda(); println!("=== Liquid GPU Smoke ==="); use ml::liquid::adapter::LiquidTrainableAdapter; use ml::liquid::CfCTrainConfig; use ml::gpu::DeviceConfig; let config = CfCTrainConfig { input_size: 10, hidden_size: 16, output_size: 1, backbone_hidden_sizes: vec![16, 8], learning_rate: 1e-3, device: DeviceConfig::Cuda(0), ..CfCTrainConfig::default() }; let mut adapter = LiquidTrainableAdapter::new(config).unwrap(); assert!(adapter.device().is_cuda(), "Liquid should be on CUDA"); // Input: [batch, seq_len=4, input_size=10] (3D required) let input = Tensor::randn(0f32, 0.5, &[16, 4, 10], &device).unwrap(); // Output: [batch, output_size=1] let target = Tensor::randn(0f32, 0.1, &[16, 1], &device).unwrap(); smoke_pipeline(&mut adapter, &input, &target, "Liquid"); } // ──────────────────────────────────────────── // KAN // ──────────────────────────────────────────── #[test] fn test_kan_gpu_smoke() { let device = require_cuda(); println!("=== KAN GPU Smoke ==="); use ml::kan::config::KANConfig; use ml::kan::trainable::KANTrainableAdapter; let config = KANConfig { layer_widths: vec![10, 8, 4, 1], grid_size: 3, spline_order: 3, learning_rate: 1e-3, weight_decay: 1e-5, grad_clip: 1.0, }; let mut adapter = KANTrainableAdapter::new(config, &device).unwrap(); assert!(adapter.device().is_cuda(), "KAN should be on CUDA"); // Input: [batch, 10], Output: [batch, 1] let input = Tensor::randn(0f32, 0.5, &[16, 10], &device).unwrap(); let target = Tensor::randn(0f32, 0.1, &[16, 1], &device).unwrap(); smoke_pipeline(&mut adapter, &input, &target, "KAN"); } // ──────────────────────────────────────────── // xLSTM // ──────────────────────────────────────────── #[test] fn test_xlstm_gpu_smoke() { let device = require_cuda(); println!("=== xLSTM GPU Smoke ==="); use ml::xlstm::trainable::XLSTMTrainableAdapter; use ml::xlstm::config::XLSTMConfig; let config = XLSTMConfig { input_dim: 10, hidden_dim: 16, num_blocks: 2, num_heads: 2, slstm_ratio: 0.5, output_dim: 1, dropout: 0.0, learning_rate: 1e-3, weight_decay: 1e-5, grad_clip: 1.0, }; let mut adapter = XLSTMTrainableAdapter::new(config, &device).unwrap(); assert!(adapter.device().is_cuda(), "xLSTM should be on CUDA"); // Input: [batch, seq_len=4, input_dim=10] (3D) let input = Tensor::randn(0f32, 0.5, &[16, 4, 10], &device).unwrap(); // Output: [batch, 1] let target = Tensor::randn(0f32, 0.1, &[16, 1], &device).unwrap(); smoke_pipeline(&mut adapter, &input, &target, "xLSTM"); } // ──────────────────────────────────────────── // Diffusion // ──────────────────────────────────────────── #[test] fn test_diffusion_gpu_smoke() { let device = require_cuda(); println!("=== Diffusion GPU Smoke ==="); use ml::diffusion::config::DiffusionConfig; use ml::diffusion::trainable::DiffusionTrainableAdapter; let config = DiffusionConfig { num_timesteps: 50, sampling_steps: 5, seq_len: 8, feature_dim: 1, hidden_dim: 16, num_layers: 1, time_embed_dim: 8, learning_rate: 1e-3, weight_decay: 1e-5, grad_clip: 1.0, ..Default::default() }; let data_dim = config.data_dim(); // 8 let mut adapter = DiffusionTrainableAdapter::new(config, device.clone()).unwrap(); assert!(adapter.device().is_cuda(), "Diffusion should be on CUDA"); // Input: [batch, data_dim=8] let input = Tensor::randn(0f32, 0.5, &[16, data_dim], &device).unwrap(); // Diffusion: use input as pseudo-target (noise prediction, loss won't be meaningful) let target = input.clone(); smoke_pipeline(&mut adapter, &input, &target, "Diffusion"); } ``` - [ ] **Step 2: Verify it compiles (no CUDA needed — just check syntax)** ```bash SQLX_OFFLINE=true cargo test -p ml --features cuda --test supervised_gpu_smoke_test --no-run 2>&1 | tail -5 ``` Expected: compiles with 0 errors (tests won't run without CUDA — they exit(0)). - [ ] **Step 3: Run on local GPU (if 4GB allows — may OOM, that's OK)** ```bash SQLX_OFFLINE=true cargo test -p ml --features cuda --test supervised_gpu_smoke_test -- --nocapture --test-threads=1 2>&1 | tail -30 ``` Expected: either all 8 PASS, or OOM on 4GB GPU (fine — will work on H100). - [ ] **Step 4: Commit** ```bash git add crates/ml/tests/supervised_gpu_smoke_test.rs git commit -m "test: add supervised_gpu_smoke_test for all 8 models (UnifiedTrainable on CUDA)" ``` --- ## Task 4: Create GPU test WorkflowTemplate **Files:** - Create: `infra/k8s/argo/gpu-test-pipeline-template.yaml` This is the core workflow. Reference `compile-and-train-template.yaml` for patterns (git checkout, env vars, tolerations, notification). - [ ] **Step 1: Create the WorkflowTemplate** Create `infra/k8s/argo/gpu-test-pipeline-template.yaml`. The complete YAML is provided in a single block below. The file has 4 templates: `pipeline` (DAG entrypoint), `gpu-warmup`, `compile-and-test`, and `notify-result`. The shell script for `compile-and-test` is shown separately after the YAML structure for readability, but must be inlined in the `args: [|]` block. ```yaml # infra/k8s/argo/gpu-test-pipeline-template.yaml apiVersion: argoproj.io/v1alpha1 kind: WorkflowTemplate metadata: name: gpu-test-pipeline namespace: foxhunt labels: app.kubernetes.io/name: gpu-test-pipeline app.kubernetes.io/part-of: foxhunt spec: entrypoint: pipeline onExit: notify-result serviceAccountName: argo-workflow podMetadata: labels: app.kubernetes.io/part-of: foxhunt app.kubernetes.io/component: gpu-test securityContext: fsGroup: 0 podGC: strategy: OnPodCompletion ttlStrategy: secondsAfterCompletion: 3600 activeDeadlineSeconds: 5400 arguments: parameters: - name: commit-ref value: HEAD - name: models value: "dqn,ppo,tft" - name: test-scope value: all - name: gpu-pool value: ci-training-h100 - name: cuda-compute-cap value: "90" - name: notify value: "true" volumes: - name: git-ssh-key secret: secretName: argo-git-ssh-key defaultMode: 256 - name: cargo-target persistentVolumeClaim: claimName: cargo-target-cuda-test - name: test-data persistentVolumeClaim: claimName: test-data-pvc readOnly: true templates: # ── pipeline: DAG entrypoint ── - name: pipeline dag: tasks: - name: gpu-warmup template: gpu-warmup - name: compile-and-test template: compile-and-test dependencies: [gpu-warmup] # ── gpu-warmup: trigger H100 autoscale ── - name: gpu-warmup nodeSelector: k8s.scaleway.com/pool-name: "{{workflow.parameters.gpu-pool}}" tolerations: - key: nvidia.com/gpu operator: Exists effect: NoSchedule - key: node.cilium.io/agent-not-ready operator: Exists effect: NoSchedule container: image: busybox:1.37 command: ["/bin/sh", "-c"] args: - | echo "GPU warmup: triggering node autoscale..." echo "GPU node scheduled, exiting to free resources" resources: requests: nvidia.com/gpu: "1" cpu: 100m memory: 64Mi limits: nvidia.com/gpu: "1" cpu: 200m memory: 128Mi # ── compile-and-test: compile + run GPU tests in single H100 pod ── - name: compile-and-test nodeSelector: k8s.scaleway.com/pool-name: "{{workflow.parameters.gpu-pool}}" tolerations: - key: nvidia.com/gpu operator: Exists effect: NoSchedule - key: node.cilium.io/agent-not-ready operator: Exists effect: NoSchedule outputs: parameters: - name: results valueFrom: path: /tmp/outputs/results default: "unknown:FAIL" - name: failures valueFrom: path: /tmp/outputs/failures default: "1" container: image: gitlab-registry.foxhunt.svc.cluster.local:5000/root/foxhunt/ci-builder:latest command: ["/bin/sh", "-c"] env: - name: SQLX_OFFLINE value: "true" - name: CARGO_TERM_COLOR value: always - name: CARGO_TARGET_DIR value: /cargo-target - name: CARGO_HOME value: /cargo-target/cargo-home - name: CUDA_VISIBLE_DEVICES value: "0" - name: CUBLAS_WORKSPACE_CONFIG value: ":4096:8" - name: TEST_DATA_DIR value: /data/test-data - name: RUST_LOG value: info volumeMounts: - name: git-ssh-key mountPath: /etc/git-ssh readOnly: true - name: cargo-target mountPath: /cargo-target - name: test-data mountPath: /data/test-data readOnly: true resources: requests: nvidia.com/gpu: "1" cpu: "4" memory: 32Gi limits: nvidia.com/gpu: "1" cpu: "8" memory: 64Gi args: - | set -e REF="{{workflow.parameters.commit-ref}}" MODELS="{{workflow.parameters.models}}" SCOPE="{{workflow.parameters.test-scope}}" # --- SSH setup (same as compile-and-train) --- mkdir -p ~/.ssh cp /etc/git-ssh/ssh-privatekey ~/.ssh/id_ed25519 chmod 600 ~/.ssh/id_ed25519 printf 'Host *\n StrictHostKeyChecking no\n UserKnownHostsFile /dev/null\n' > ~/.ssh/config chmod 600 ~/.ssh/config REPO="ssh://git@gitlab-gitlab-shell.foxhunt.svc.cluster.local:2222/root/foxhunt.git" BUILD="/cargo-target/src" git config --global --add safe.directory "$BUILD" # --- Persistent checkout on PVC (same as compile-and-train) --- if [ -d "$BUILD/.git" ]; then cd "$BUILD" CURRENT=$(git rev-parse HEAD 2>/dev/null || echo "none") TARGET=$(git ls-remote origin "$REF" 2>/dev/null | cut -f1 || echo "$REF") if [ "$CURRENT" = "$TARGET" ] || [ "$CURRENT" = "$REF" ]; then echo "=== Already at $REF ===" else echo "=== Updating checkout ===" git fetch origin git checkout --force "$REF" git clean -fd fi else echo "=== Initial clone ===" git clone --filter=blob:none "$REPO" "$BUILD" cd "$BUILD" git checkout "$REF" fi echo "Checked out $(git rev-parse --short HEAD)" export PATH="${CARGO_HOME}/bin:${PATH}" export CUDA_COMPUTE_CAP={{workflow.parameters.cuda-compute-cap}} # --- PVC size guard --- TARGET_SIZE_MB=$(du -sm "$CARGO_TARGET_DIR" 2>/dev/null | cut -f1 || echo 0) echo "PVC usage: ${TARGET_SIZE_MB}MB" if [ "$TARGET_SIZE_MB" -gt 25000 ]; then echo "PVC exceeds 25GB, pruning..." rm -rf "$CARGO_TARGET_DIR/release" "$CARGO_TARGET_DIR/debug" fi # --- Compile --- echo "=== Compiling test binaries (--features cuda) ===" cargo test -p ml -p ml-dqn -p ml-core --features cuda --no-run 2>&1 echo "=== Compilation done ===" # --- Expand "all" --- if [ "$MODELS" = "all" ]; then MODELS="dqn,ppo,tft,mamba2,tggn,tlob,liquid,kan,xlstm,diffusion" fi # --- Test runner --- RESULTS="" FAILURES=0 run_tests() { local NAME="$1"; shift echo "" echo "========================================" echo " TEST: $NAME" echo "========================================" set +e "$@" --nocapture 2>&1 EXIT=$? set -e if [ $EXIT -eq 0 ]; then RESULTS="${RESULTS}${NAME}:PASS\n" else RESULTS="${RESULTS}${NAME}:FAIL\n" FAILURES=$((FAILURES + 1)) fi } # --- Core tests (always run) --- run_tests "core-lib" cargo test -p ml-core --features cuda --lib run_tests "bayesian" cargo test -p ml --features cuda --test bayesian_changepoint_test # --- Per-model tests --- IFS=',' read -ra MODEL_LIST <<< "$MODELS" for MODEL in "${MODEL_LIST[@]}"; do case "$MODEL" in dqn) if [ "$SCOPE" = "lib" ] || [ "$SCOPE" = "all" ]; then run_tests "dqn-lib" cargo test -p ml-dqn --features cuda --lib run_tests "dqn-ml-lib" cargo test -p ml --features cuda --lib -- dqn fi if [ "$SCOPE" = "integration" ] || [ "$SCOPE" = "all" ]; then run_tests "dqn-smoke" cargo test -p ml --features cuda --test smoke_test_real_data run_tests "dqn-pipeline" cargo test -p ml --features cuda --test dqn_training_pipeline_test run_tests "dqn-smoke-train" cargo test -p ml --features cuda --test dqn_training_smoke_test run_tests "dqn-early-stop" cargo test -p ml --features cuda --test dqn_early_stopping_termination_test run_tests "dqn-collapse" cargo test -p ml --features cuda --test dqn_action_collapse_fix_test fi ;; ppo) if [ "$SCOPE" = "lib" ] || [ "$SCOPE" = "all" ]; then run_tests "ppo-lib" cargo test -p ml --features cuda --lib -- ppo fi if [ "$SCOPE" = "integration" ] || [ "$SCOPE" = "all" ]; then run_tests "ppo-barrier" cargo test -p ml --features cuda --test barrier_optimization_test fi ;; *) # Supervised models (TFT, Mamba2, TGGN, TLOB, Liquid, KAN, xLSTM, Diffusion) # Map model names to Rust module names where they differ LIB_FILTER="$MODEL" [ "$MODEL" = "tggn" ] && LIB_FILTER="tgnn" if [ "$SCOPE" = "lib" ] || [ "$SCOPE" = "all" ]; then run_tests "${MODEL}-lib" cargo test -p ml --features cuda --lib -- "$LIB_FILTER" fi if [ "$SCOPE" = "integration" ] || [ "$SCOPE" = "all" ]; then run_tests "${MODEL}-gpu" cargo test -p ml --features cuda --test supervised_gpu_smoke_test -- "test_${MODEL}_gpu_smoke" # Also run model-specific integration tests if they exist if cargo test -p ml --features cuda --test "${MODEL}_integration" --no-run 2>/dev/null; then run_tests "${MODEL}-integ" cargo test -p ml --features cuda --test "${MODEL}_integration" fi fi ;; esac done # --- Summary --- echo "" echo "========================================" echo " GPU TEST RESULTS" echo "========================================" printf "$RESULTS" | while IFS=: read -r name status; do [ -z "$name" ] && continue printf " %-25s %s\n" "$name" "$status" done echo "========================================" echo " Total failures: $FAILURES" echo "========================================" # Write results for notification step mkdir -p /tmp/outputs printf "$RESULTS" > /tmp/outputs/results echo "$FAILURES" > /tmp/outputs/failures [ "$FAILURES" -gt 0 ] && exit 1 || exit 0 # ── notify-result: post test outcome to Mattermost (onExit) ── - name: notify-result nodeSelector: k8s.scaleway.com/pool-name: platform container: image: gitlab-registry.foxhunt.svc.cluster.local:5000/root/foxhunt/foxhunt-runtime:latest command: ["/bin/sh", "-c"] env: - name: WEBHOOK_URL valueFrom: secretKeyRef: name: notification-webhook key: webhook-url optional: true resources: requests: cpu: 50m memory: 32Mi limits: cpu: 100m memory: 64Mi args: - | NOTIFY="{{workflow.parameters.notify}}" if [ "$NOTIFY" != "true" ]; then echo "Notifications disabled, skipping" exit 0 fi STATUS="{{workflow.status}}" NAME="{{workflow.name}}" if [ -z "$WEBHOOK_URL" ] || echo "$WEBHOOK_URL" | grep -q "PLACEHOLDER"; then echo "No webhook configured, skipping notification" exit 0 fi if [ "$STATUS" = "Succeeded" ]; then EMOJI=":white_check_mark:" else EMOJI=":x:" fi PAYLOAD="{\"username\":\"Argo CI\",\"text\":\"${EMOJI} **GPU Tests** ${NAME} — ${STATUS} ({{workflow.duration}}s)\"}" curl -sf -X POST -H 'Content-Type: application/json' \ -d "$PAYLOAD" "$WEBHOOK_URL" || echo "WARN: webhook post failed" ``` The complete shell script is inlined in the YAML above within the `compile-and-test` template's `args` block. - [ ] **Step 2: Validate YAML syntax** ```bash python3 -c "import yaml; yaml.safe_load(open('infra/k8s/argo/gpu-test-pipeline-template.yaml'))" && echo "YAML OK" ``` - [ ] **Step 3: Commit** ```bash git add infra/k8s/argo/gpu-test-pipeline-template.yaml git commit -m "feat(ci): add gpu-test-pipeline WorkflowTemplate for H100 CUDA testing" ``` --- ## Task 5: Create nightly CronWorkflow **Files:** - Create: `infra/k8s/argo/gpu-test-nightly-cron.yaml` - [ ] **Step 1: Create CronWorkflow (suspended)** ```yaml # infra/k8s/argo/gpu-test-nightly-cron.yaml apiVersion: argoproj.io/v1alpha1 kind: CronWorkflow metadata: name: gpu-test-nightly namespace: foxhunt labels: app.kubernetes.io/name: gpu-test-nightly app.kubernetes.io/part-of: foxhunt spec: schedule: "0 2 * * *" timezone: "UTC" suspend: true # Enable when ready: kubectl patch cronworkflow gpu-test-nightly -n foxhunt -p '{"spec":{"suspend":false}}' concurrencyPolicy: Replace successfulJobsHistoryLimit: 3 failedJobsHistoryLimit: 5 workflowSpec: workflowTemplateRef: name: gpu-test-pipeline arguments: parameters: - name: models value: "dqn,ppo,tft,mamba2,tggn,tlob,liquid,kan,xlstm,diffusion" - name: commit-ref value: "main" ``` - [ ] **Step 2: Commit** ```bash git add infra/k8s/argo/gpu-test-nightly-cron.yaml git commit -m "feat(ci): add gpu-test-nightly CronWorkflow (suspended)" ``` --- ## Task 6: Create argo-test.sh CLI **Files:** - Create: `scripts/argo-test.sh` - [ ] **Step 1: Create the script** Pattern matches `scripts/argo-train.sh` (same argument parsing style, same `argo submit` invocation). ```bash #!/usr/bin/env bash # Run GPU/CUDA tests via Argo Workflows on H100. # # Usage: # ./scripts/argo-test.sh # defaults: dqn,ppo,tft # ./scripts/argo-test.sh --models dqn --scope lib # DQN lib tests only # ./scripts/argo-test.sh --models all # all 10 models # ./scripts/argo-test.sh --watch # follow logs # # Requires: argo CLI set -euo pipefail MODELS="dqn,ppo,tft" SCOPE="all" COMMIT_REF="HEAD" GPU_POOL="" WATCH=false usage() { cat < Comma-separated model list (default: dqn,ppo,tft) Use "all" for all 10 models --scope lib, integration, or all (default: all) --ref Git ref to test (default: HEAD) --gpu-pool GPU node pool (default: ci-training-h100) --watch Follow workflow logs -h, --help Show this help EOF exit 0 } while [[ $# -gt 0 ]]; do case $1 in --models) MODELS="$2"; shift 2 ;; --scope) SCOPE="$2"; shift 2 ;; --ref) COMMIT_REF="$2"; shift 2 ;; --gpu-pool) GPU_POOL="$2"; shift 2 ;; --watch) WATCH=true; shift ;; -h|--help) usage ;; *) echo "Unknown option: $1"; usage ;; esac done # Resolve HEAD to actual SHA if needed if [ "$COMMIT_REF" = "HEAD" ]; then COMMIT_REF=$(git rev-parse HEAD) fi CMD="argo submit -n foxhunt --from=wftmpl/gpu-test-pipeline" CMD="$CMD -p commit-ref=$COMMIT_REF" CMD="$CMD -p models=$MODELS" CMD="$CMD -p test-scope=$SCOPE" [[ -n "$GPU_POOL" ]] && CMD="$CMD -p gpu-pool=$GPU_POOL" if $WATCH; then CMD="$CMD --watch" fi echo "Submitting GPU test workflow..." echo " commit: $(echo $COMMIT_REF | cut -c1-8)" echo " models: $MODELS" echo " scope: $SCOPE" echo "" eval "$CMD" ``` - [ ] **Step 2: Make executable and commit** ```bash chmod +x scripts/argo-test.sh git add scripts/argo-test.sh git commit -m "feat(ci): add argo-test.sh CLI for GPU test workflow" ``` --- ## Task 7: Add on-push trigger to CI pipeline **Files:** - Modify: `infra/k8s/argo/ci-pipeline-template.yaml` - [ ] **Step 1: Add `ml-changed` to detect-changes** In `ci-pipeline-template.yaml`, inside the `detect-changes` template's shell script (around line 199), add: ```bash ML_CHANGED=$(check_paths "crates/ml crates/ml-dqn crates/ml-core crates/ml-regime crates/ml-features crates/ml-supervised crates/ml-ppo crates/ml-ensemble crates/ml-hyperopt crates/ml-labeling") ``` And add the output file write (around line 218): ```bash echo -n "$ML_CHANGED" > /tmp/outputs/ml-changed ``` And add the log line (around line 210): ```bash echo "ml-changed: $ML_CHANGED" ``` - [ ] **Step 2: Add ml-changed output parameter** In the `outputs.parameters` section of `detect-changes` (around line 238), add: ```yaml - name: ml-changed valueFrom: path: /tmp/outputs/ml-changed ``` - [ ] **Step 3: Add gpu-test DAG task + submit-gpu-test template** **Important**: Cannot use `templateRef` to invoke `gpu-test-pipeline/pipeline` directly from the CI pipeline DAG because `templateRef` runs the referenced template within the **calling** workflow's pod context. The CI pipeline lacks the required PVC volumes (`cargo-target-cuda-test`, `test-data-pvc`) and GPU nodeSelector/tolerations. Instead, use a `resource` template that submits a child Workflow from the WorkflowTemplate (workflow-of-workflows pattern). In the `pipeline` DAG `tasks` section (around line 143), add: ```yaml - name: gpu-test depends: "detect-changes" template: submit-gpu-test when: "{{tasks.detect-changes.outputs.parameters.ml-changed}} == true" ``` Add the `submit-gpu-test` template to the `templates` list: ```yaml # ── submit-gpu-test: launch GPU test workflow as a child workflow ── - name: submit-gpu-test resource: action: create setOwnerReference: true successCondition: status.phase == Succeeded failureCondition: status.phase in (Failed, Error) manifest: | apiVersion: argoproj.io/v1alpha1 kind: Workflow metadata: generateName: gpu-test-on-push- namespace: foxhunt spec: workflowTemplateRef: name: gpu-test-pipeline arguments: parameters: - name: commit-ref value: "{{workflow.parameters.commit-sha}}" - name: models value: "dqn,ppo,tft" ``` This creates a child Workflow from the `gpu-test-pipeline` WorkflowTemplate. The child workflow has its own volumes, nodeSelector, and tolerations defined in the WorkflowTemplate spec. `setOwnerReference: true` ensures the child workflow is cleaned up when the parent is deleted. - [ ] **Step 4: Validate YAML syntax** ```bash python3 -c "import yaml; yaml.safe_load(open('infra/k8s/argo/ci-pipeline-template.yaml'))" && echo "YAML OK" ``` - [ ] **Step 5: Commit** ```bash git add infra/k8s/argo/ci-pipeline-template.yaml git commit -m "feat(ci): add gpu-test on-push trigger gated on ML crate changes" ``` --- ## Task 8: Apply to cluster and validate - [ ] **Step 1: Apply PVCs** ```bash kubectl -n foxhunt apply -f infra/k8s/argo/cargo-target-cuda-test-pvc.yaml kubectl -n foxhunt apply -f infra/k8s/training/test-data-pvc.yaml ``` - [ ] **Step 2: Populate test data** ```bash ./scripts/refresh-test-data.sh ``` Wait for job completion. Verify data: ```bash kubectl -n foxhunt exec -it $(kubectl -n foxhunt get pods -l app.kubernetes.io/name=populate-test-data -o name | head -1) -- ls -lhR /test-data/ ``` - [ ] **Step 3: Apply WorkflowTemplate and CronWorkflow** ```bash kubectl -n foxhunt apply -f infra/k8s/argo/gpu-test-pipeline-template.yaml kubectl -n foxhunt apply -f infra/k8s/argo/gpu-test-nightly-cron.yaml ``` - [ ] **Step 4: Run a quick manual test (DQN lib only)** ```bash ./scripts/argo-test.sh --models dqn --scope lib --watch ``` Expected: compile + DQN lib tests pass on H100. - [ ] **Step 5: Run full default suite** ```bash ./scripts/argo-test.sh --watch ``` Expected: DQN + PPO + TFT tests all pass. - [ ] **Step 6: Apply updated CI pipeline** ```bash kubectl -n foxhunt apply -f infra/k8s/argo/ci-pipeline-template.yaml ``` - [ ] **Step 7: Final commit (if any fixes needed)** ```bash git add -A git commit -m "fix(ci): adjustments from H100 GPU test validation" ```