# Zero-Copy FxCache Training Pipeline — Implementation Plan > **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking. **Goal:** Eliminate all per-fold CPU waste by building the entire training pipeline on FxCacheData (features[42] + targets[4] + OFI[8] + timestamps). Load once, upload to GPU once, slice by index per fold. Walk-forward uses full MBP-10/trades data via precomputed OFI features. **Architecture:** FxCacheData replaces OHLCVBar as the data backbone. Walk-forward generates index ranges from timestamps. Trainer accepts `(&[[f64; 42]], &[[f64; 4]], &[[f64; 8]])` slices instead of `Vec<(FeatureVector, Vec)>`. Data uploaded to GPU once; per-fold transitions set index bounds only. **Tech Stack:** Rust 1.85, cudarc 0.17.3, CUDA --- ## Current Data Flow (per fold) ``` fxcache ─→ (features[42], targets[4], OFI[8], timestamps) # 470MB, loaded once │ ▼ OHLCVBars reconstructed from targets (close prices) # WASTE: fxcache has everything │ ▼ generate_walk_forward_windows(bars) # Clones bars into Vec per fold │ ▼ prepare_fold_data(window) # RE-EXTRACTS features from bars (!) │ ▼ features_to_trainer_format(features, bars) # 4M × Vec heap allocs │ ▼ DQNTrainer::new() per fold # Rebuilds GPU pipeline │ ▼ init_gpu_data → flatten + upload to GPU # Unpacks tuples, flattens, HtoD │ ▼ Training epochs (GPU) # The only useful work ``` ## Target Data Flow ``` fxcache ─→ FxCacheData { features[42], targets[4], OFI[8], timestamps } │ ▼ (one-time) Upload to GPU: features_cuda, targets_cuda, ofi_cuda │ ▼ (one-time) generate_walk_forward_indices(timestamps) → Vec │ ▼ (one-time) DQNTrainer::new() │ ▼ (per fold, ~0.1s) trainer.set_training_range(range.train_start, range.train_end, ...) trainer.reset_for_fold() NormStats from &features[train_start..train_end] # Slice, no copy │ ▼ Training epochs (GPU reads from offset in pre-uploaded arrays) ``` ## File Structure ### Modified Files | File | Changes | |------|---------| | `crates/ml/src/walk_forward.rs` | `generate_walk_forward_indices_from_timestamps(&[i64], config) → Vec` — date slicing on nanosecond timestamps from fxcache, no OHLCVBar dependency | | `crates/ml/src/trainers/dqn/trainer/mod.rs` | New `train_with_fxcache_data()` API accepting `&FxCacheData`. Remove `Vec<(FeatureVector, Vec)>` from the hot path. Add `upload_fxcache_to_gpu()` for one-time upload | | `crates/ml/src/trainers/dqn/trainer/training_loop.rs` | `init_gpu_raw_buffers_from_slices()` accepting `(&[[f64; 42]], &[[f64; 4]], &[[f64; 8]])`. Remove per-element flattening loops | | `crates/ml/src/cuda_pipeline/mod.rs` | `DqnGpuData::upload_slices()` accepting contiguous `&[[f64; 42]]` + `&[[f64; 4]]` — no tuple unpacking | | `crates/ml/examples/train_baseline_rl.rs` | Complete rewrite of fold loop: load fxcache → upload GPU → index ranges → per-fold: set range, reset, train. Delete `prepare_fold_data`, `FoldData`, `features_to_trainer_format`, double-buffer, prefetch thread | --- ### Task 1: Walk-forward from timestamps (already partially done) **Files:** - Modify: `crates/ml/src/walk_forward.rs` `generate_walk_forward_indices` (added in previous commit) uses `&[OHLCVBar]` for date slicing. Add a variant that works directly on `&[i64]` timestamps from fxcache — no OHLCVBar needed. - [ ] **Step 1: Add `generate_walk_forward_indices_from_timestamps`** ```rust /// Generate walk-forward index ranges from nanosecond timestamps. /// Works directly on fxcache timestamps — no OHLCVBar conversion needed. pub fn generate_walk_forward_indices_from_timestamps( timestamps_ns: &[i64], config: &WalkForwardConfig, ) -> Vec { if timestamps_ns.is_empty() { return Vec::new(); } let ts_to_date = |ts: i64| -> NaiveDate { chrono::DateTime::from_timestamp_nanos(ts).date_naive() }; let data_start = ts_to_date(timestamps_ns[0]); let data_end = ts_to_date(*timestamps_ns.last().unwrap()); // ... same fold logic as generate_walk_forward_indices // but uses partition_point on timestamps instead of bars } ``` - [ ] **Step 2: Add `compute_difficulty_from_features` that reads ADX from features[40]** The current `compute_difficulty(&[OHLCVBar])` computes ADX from raw bars. With fxcache, ADX is already at feature index 40. Add: ```rust pub fn compute_difficulty_from_features(features: &[[f64; 42]]) -> f64 { if features.is_empty() { return 0.0; } let adx_idx = 40; let sum: f64 = features.iter().map(|f| f[adx_idx]).sum(); sum / features.len() as f64 } ``` - [ ] **Step 3: Compile** ```bash SQLX_OFFLINE=true cargo check -p ml ``` - [ ] **Step 4: Commit** --- ### Task 2: DqnGpuData::upload_slices — contiguous array upload **Files:** - Modify: `crates/ml/src/cuda_pipeline/mod.rs` The current `upload()` takes `&[([f64; 42], Vec)]` and iterates element-by-element to flatten. Add `upload_slices()` that accepts pre-contiguous arrays from fxcache. - [ ] **Step 1: Add `upload_slices` method** ```rust /// Upload contiguous feature/target slices to GPU (zero per-element iteration). /// Accepts fxcache arrays directly — no tuple unpacking, no Vec allocation. pub fn upload_slices( features: &[[f64; 42]], targets: &[[f64; 4]], ofi: &[[f64; 8]], stream: &Arc, ) -> Result { let num_bars = features.len(); let feature_dim = 42; // Cast f64 → f32 in one contiguous pass (no per-element Vec push) let flat_features: Vec = features.iter() .flat_map(|f| f.iter().map(|&v| v as f32)) .collect(); let flat_targets: Vec = targets.iter() .flat_map(|t| t.iter().map(|&v| v as f32)) .collect(); let features_gpu = clone_htod_f32_to_bf16(stream, &flat_features)?; let targets_gpu = clone_htod_f32_to_bf16(stream, &flat_targets)?; let mut data = Self { features: features_gpu, targets: targets_gpu, ofi_features: None, num_bars, feature_dim, aligned_state_dim: None, }; // Upload OFI if present (non-zero) if !ofi.is_empty() && ofi.iter().any(|o| o.iter().any(|&v| v != 0.0)) { let flat_ofi: Vec = ofi.iter() .flat_map(|o| o.iter().map(|&v| v as f32)) .collect(); let ofi_gpu = clone_htod_f32_to_bf16(stream, &flat_ofi)?; data.ofi_features = Some(ofi_gpu); } Ok(data) } ``` - [ ] **Step 2: Compile** ```bash SQLX_OFFLINE=true cargo check -p ml ``` - [ ] **Step 3: Commit** --- ### Task 3: Trainer API — `init_from_fxcache` **Files:** - Modify: `crates/ml/src/trainers/dqn/trainer/mod.rs` - Modify: `crates/ml/src/trainers/dqn/trainer/training_loop.rs` Add `init_from_fxcache()` that uploads the full fxcache dataset to GPU ONCE. This replaces the per-fold `init_gpu_data` + `init_gpu_raw_buffers` pair. - [ ] **Step 1: Add `init_from_fxcache` on DQNTrainer** ```rust /// Upload full fxcache dataset to GPU ONCE. Call before the fold loop. /// Replaces per-fold init_gpu_data + init_gpu_raw_buffers. pub async fn init_from_fxcache( &mut self, features: &[[f64; 42]], targets: &[[f64; 4]], ofi: &[[f64; 8]], ) -> Result<()> { let stream = self.cuda_stream.as_ref() .ok_or_else(|| anyhow::anyhow!("CUDA stream required"))?; // Upload via DqnGpuData::upload_slices (contiguous, zero tuple unpacking) let gpu_data = DqnGpuData::upload_slices(features, targets, ofi, stream)?; let ofi_enabled = gpu_data.ofi_features.is_some(); let raw_dim = if ofi_enabled { 53 } else { 45 }; let aligned_dim = (raw_dim + 7) & !7; info!("init_from_fxcache: {} bars uploaded to GPU ({:.1} MB)", features.len(), (features.len() * (42 + 4 + 8) * 4) as f64 / 1_048_576.0); self.gpu_data = Some(gpu_data); // Also upload raw buffers for experience collector self.init_gpu_raw_buffers_from_slices(features, targets).await?; Ok(()) } ``` - [ ] **Step 2: Add `init_gpu_raw_buffers_from_slices` in training_loop.rs** ```rust /// Upload raw features + targets from contiguous slices (no tuple unpacking). pub(crate) async fn init_gpu_raw_buffers_from_slices( &mut self, features: &[[f64; 42]], targets: &[[f64; 4]], ) -> Result<()> { if self.targets_raw_cuda.is_some() { return Ok(()); } let stream = match self.cuda_stream { Some(ref s) => Arc::clone(s), None => return Ok(()), }; let num_bars = features.len(); let flat_targets: Vec = targets.iter() .flat_map(|t| t.iter().map(|&v| v as f32)) .collect(); let flat_features: Vec = features.iter() .flat_map(|f| f.iter().map(|&v| v as f32)) .collect(); self.targets_raw_cuda = Some( crate::cuda_pipeline::clone_htod_f32_to_bf16(&stream, &flat_targets)? ); self.features_raw_cuda = Some( crate::cuda_pipeline::clone_htod_f32_to_bf16(&stream, &flat_features)? ); info!("CUDA raw buffers uploaded: {} bars x 42 features + 4 targets", num_bars); Ok(()) } ``` - [ ] **Step 3: Add `set_val_data_from_slices` for validation** The trainer needs val_data for epoch-end Sharpe computation. Instead of `Vec<(FeatureVector, Vec)>`, accept slices: ```rust /// Set validation data from contiguous slices (zero Vec allocation). pub fn set_val_data_from_slices( &mut self, features: &[[f64; 42]], targets: &[[f64; 4]], ofi_val_offset: usize, ) { // Convert to the legacy format for now — val set is small (~50K bars) // Full GPU-native validation is a future task self.val_data = features.iter().zip(targets.iter()) .map(|(f, t)| (*f, t.to_vec())) .collect(); self.ofi_val_offset = ofi_val_offset; self.val_features_gpu = None; self.val_closes_gpu = None; self.val_ofi_gpu = None; } ``` - [ ] **Step 4: Compile** ```bash SQLX_OFFLINE=true cargo check -p ml ``` - [ ] **Step 5: Commit** --- ### Task 4: Rewrite train_baseline_rl fold loop **Files:** - Modify: `crates/ml/examples/train_baseline_rl.rs` This is the integration task. Replace the entire fold loop to: 1. Load fxcache ONCE 2. Create DQNTrainer ONCE 3. Upload to GPU ONCE via `init_from_fxcache` 4. Generate index ranges from timestamps 5. Per fold: slice features/targets, normalize, set_training_range, reset_for_fold, train - [ ] **Step 1: Replace data loading section (lines ~798-891)** Replace the fxcache + DBN loading + OHLCVBar reconstruction with: ```rust // Load fxcache — contains features[42] + targets[4] + OFI[8] + timestamps // MBP-10 orderbook and trades data are pre-embedded in OFI features let fxcache = load_fxcache_data(&args)?; info!("Loaded fxcache: {} bars, features[42] + targets[4] + OFI[8]", fxcache.bar_count); ``` Where `load_fxcache_data` is a helper that does the cache lookup (same as current code but returns `FxCacheData` directly, no OHLCVBar conversion). - [ ] **Step 2: Replace walk-forward window generation** ```rust let fold_ranges = generate_walk_forward_indices_from_timestamps( &fxcache.timestamps, &wf_config, ); info!("Generated {} walk-forward folds", fold_ranges.len()); ``` - [ ] **Step 3: Create DQN trainer + upload to GPU ONCE** ```rust let hp = load_hyperopt_params(&args.hyperopt_params, "dqn"); let hyperparams = build_dqn_hyperparams(args, &hp, &gpu_profile); let mut trainer = DQNTrainer::new(hyperparams)?; // One-time GPU upload of full dataset let rt = tokio::runtime::Builder::new_current_thread() .enable_all().build()?; rt.block_on(trainer.init_from_fxcache( &fxcache.features, &fxcache.targets, &fxcache.ofi, ))?; ``` - [ ] **Step 4: Replace fold loop body** ```rust for range in &fold_ranges { info!("--- Fold {} ---", range.fold); // Slice features for normalization (zero copy — just index bounds) let train_feat = &fxcache.features[range.train_start..range.train_end]; let val_feat = &fxcache.features[range.val_start..range.val_end]; let train_targets = &fxcache.targets[range.train_start..range.train_end]; let val_targets = &fxcache.targets[range.val_start..range.val_end]; // NormStats from training slice (CPU, ~10ms for 500K bars) let norm_stats = NormStats::from_features(train_feat); let train_norm = norm_stats.normalize_batch(train_feat); let val_norm = norm_stats.normalize_batch(val_feat); // Set fold range on trainer (GPU arrays stay put, just index bounds change) trainer.set_training_range( range.train_start, range.train_end, range.val_start, range.val_end, ); // Set validation data for epoch-end Sharpe trainer.set_val_data_from_slices(val_feat, val_targets, range.train_end - range.train_start); // Reset per-fold state (keep GPU infra) rt.block_on(trainer.reset_for_fold())?; // Checkpoint callback let output_dir = args.output_dir.clone(); let fold = range.fold; let checkpoint_callback = move |epoch: usize, data: Vec, is_best: bool| -> Result { let suffix = if is_best { "best" } else { &format!("epoch{}", epoch) }; let path = output_dir.join(format!("dqn_fold{}_{}.safetensors", fold, suffix)); std::fs::write(&path, &data)?; Ok(path.to_string_lossy().into_owned()) }; // Train — trainer reads from GPU arrays at fold range offset // NOTE: still passes train_norm as Vec<(FeatureVector, Vec)> for now. // Full elimination of this conversion requires changing train_with_data_full_loop // to read directly from GPU buffers — tracked as follow-up. let training_data = features_to_trainer_format_fast(&train_norm, train_targets); let val_data_fmt = features_to_trainer_format_fast(&val_norm, val_targets); match rt.block_on(trainer.train_with_preloaded_data(training_data, val_data_fmt, checkpoint_callback)) { Ok(metrics) => { info!("Fold {} complete: loss={:.6}, epochs={}", fold, metrics.loss, metrics.epochs_trained); dqn_results.push((fold, metrics.loss)); } Err(e) => error!("Fold {} failed: {:#}", fold, e), } } ``` - [ ] **Step 5: Add `features_to_trainer_format_fast` helper** Faster version that avoids per-bar `vec![open, high, low, close]`: ```rust fn features_to_trainer_format_fast( features: &[[f64; 42]], targets: &[[f64; 4]], ) -> Vec<(FeatureVector, Vec)> { features.iter().zip(targets.iter()) .map(|(f, t)| (*f, t.to_vec())) .collect() } ``` This still allocates `Vec` per bar (4 elements), but eliminates: - Bar cloning (320MB) - Feature re-extraction (30s CPU) - Trainer re-creation (3s GPU init) - Data re-upload (1s HtoD) The `Vec` allocation (~32 bytes per bar × 500K bars = 16MB) is the last remaining waste. Eliminating it requires changing `train_with_data_full_loop` to accept `&[[f64; 4]]` — tracked as Task 6. - [ ] **Step 6: Delete dead code** Remove: - `prepare_fold_data` function - `FoldData` type alias - `features_to_trainer_format` function (replaced by `_fast` variant) - `prefetched_data` and fold prefetch thread - `dqn_gpu_staged` double-buffer - `DqnGpuData::upload` from double-buffer path - [ ] **Step 7: Compile and test locally** ```bash SQLX_OFFLINE=true FOXHUNT_TEST_DATA=test_data/futures-baseline cargo run --release -p ml --example train_baseline_rl -- --model dqn --epochs 3 --data-dir test_data/futures-baseline 2>&1 | tail -20 ``` Expected: 3 folds × 3 epochs in <10s (was 159s). - [ ] **Step 8: Commit** --- ### Task 5: Eliminate last Vec allocation (trainer API) **Files:** - Modify: `crates/ml/src/trainers/dqn/trainer/training_loop.rs` - Modify: `crates/ml/src/trainers/dqn/trainer/mod.rs` Change `train_with_data_full_loop` to accept `(&[[f64; 42]], &[[f64; 4]])` instead of `&[(FeatureVector, Vec)]`. This eliminates the 4M × `Vec` allocation. - [ ] **Step 1: Add `train_fold_from_slices` on DQNTrainer** ```rust /// Train one fold from contiguous feature/target slices. /// Zero Vec allocation — reads targets directly as &[[f64; 4]]. pub async fn train_fold_from_slices( &mut self, features: &[[f64; 42]], targets: &[[f64; 4]], checkpoint_callback: F, ) -> Result where F: FnMut(usize, Vec, bool) -> Result + Send, { // GPU data already uploaded via init_from_fxcache — skip re-upload // Training loop uses self.features_raw_cuda / self.targets_raw_cuda // which are the FULL dataset on GPU. Fold range selects the slice. let num_bars = features.len(); info!("train_fold_from_slices: {} bars (GPU range {}..{})", num_bars, self.fold_train_start, self.fold_train_end); // The training loop still needs training_data for: // 1. .len() → use fold range // 2. Curriculum ADX filter → read features[bar][40] // 3. Vol normalizer → read targets for returns // Build a lightweight wrapper that avoids Vec let training_data_ref: Vec<([f64; 42], [f64; 4])> = features.iter() .zip(targets.iter()) .map(|(f, t)| (*f, *t)) .collect(); self.train_with_data_full_loop_slices(&training_data_ref, checkpoint_callback).await } ``` - [ ] **Step 2: Add `train_with_data_full_loop_slices` in training_loop.rs** Mirror of `train_with_data_full_loop` but accepts `&[([f64; 42], [f64; 4])]`. The internal code only needs `.len()`, `features[bar][adx_idx]`, and `targets[bar]` for vol normalizer — all work with fixed-size arrays. - [ ] **Step 3: Update fold loop in train_baseline_rl to use new API** Replace `features_to_trainer_format_fast` + `train_with_preloaded_data` with: ```rust match rt.block_on(trainer.train_fold_from_slices(&train_norm, train_targets, checkpoint_callback)) { ``` - [ ] **Step 4: Compile and test** - [ ] **Step 5: Commit** --- ### Task 6: Validate locally and on H100 - [ ] **Step 1: Local validation** ```bash SQLX_OFFLINE=true FOXHUNT_TEST_DATA=test_data/futures-baseline cargo run --release -p ml --example train_baseline_rl -- --model dqn --epochs 3 --data-dir test_data/futures-baseline 2>&1 | tail -20 ``` Expected: <10s for 3 folds × 3 epochs. - [ ] **Step 2: Smoketest** ```bash SQLX_OFFLINE=true FOXHUNT_TEST_DATA=test_data/futures-baseline cargo test -p ml --lib -- smoke_tests::training_stability::test_gpu_collector_auto_initializes --ignored --nocapture 2>&1 | tail -5 ``` Expected: PASS, <60s. - [ ] **Step 3: H100 submission** ```bash ./scripts/argo-train.sh dqn --epochs 50 --trials 0 --gpu-pool ci-training-h100 ``` Expected: - fxcache hit (precomputed on PVC) - One-time GPU upload: <2s - Per-fold transition: <0.5s - Per-epoch (batch=8192, 488 steps): <5s - Total 50 epochs × 3 folds: <15 minutes --- ## Performance Summary | Operation | Before | After | |-----------|--------|-------| | Data load (fxcache hit) | ~5s | ~5s (unchanged) | | Walk-forward window gen | ~200ms + 320MB cloned bars | ~1ms (index ranges only) | | Feature extraction per fold | ~30s CPU | 0s (pre-computed in fxcache) | | Trainer creation per fold | ~3s GPU init | 0s (reused) | | Data upload per fold | ~1s HtoD | 0s (stays on GPU) | | Vec allocs per fold | ~16MB (4M × 4 elements) | 0 (Task 5: fixed-size arrays) | | NormStats per fold | ~10ms (unchanged) | ~10ms (slice, no copy) | | **Total fold transition** | **~34s** | **~10ms** | | **3-fold overhead** | **~102s** | **~30ms** | ## Data Quality Improvement The walk-forward now uses the FULL precomputed feature set from fxcache: - **42 market features**: OHLCV derivatives + technical indicators + ADX + CUSUM - **8 OFI features**: from MBP-10 order book (OFI L1/L5, depth imbalance, VPIN, Kyle's Lambda, bid/ask slope, trade imbalance) - **4 targets**: raw close prices for portfolio simulation This matches production: the model trains on exactly the same feature set it will see in real-time trading, including orderbook microstructure from MBP-10 and trade flow data. ## Risks | Risk | Impact | Mitigation | |------|--------|-----------| | NormStats computed per fold on CPU | ~10ms — negligible | Future: GPU reduction kernel | | val_data still uses legacy format | Small alloc (~50K × Vec) | `set_val_data_from_slices` minimizes, full GPU val in follow-up | | OHLCVBar removal breaks PPO path | PPO uses same fold loop | Keep PPO using legacy path for now, separate refactor | | train_with_data_full_loop signature change | Touches many callers | Task 5 adds new method, doesn't change existing |