Files
foxhunt/docs/superpowers/plans/2026-03-21-h100-epoch-optimization-phase4.md
jgrusewski 5738edaa0a perf: keep GPU training data resident across epochs — init 73ms → 0ms
Remove per-epoch GPU data freeing that forced re-upload every epoch.
Training data within a walk-forward window is immutable — uploading once
and keeping it GPU-resident eliminates the init phase entirely.

Epoch time: 153ms → 78ms on RTX 3050 (epochs 2+).

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-21 16:29:32 +01:00

2.3 KiB

H100 Epoch Optimization Phase 4: Init Elimination

For agentic workers: REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (- [ ]) syntax for tracking.

Goal: Eliminate the 73ms/epoch init overhead by keeping GPU data resident across epochs, reducing total epoch time from 153ms to ~80ms.

Architecture: Remove the per-epoch GPU data freeing that forces re-upload. The training data within a walk-forward window is immutable — uploading it once and keeping it GPU-resident eliminates the largest remaining bottleneck.

Tech Stack: Rust, cudarc 0.19


Task 1: Stop freeing GPU data between epochs

Files:

  • Modify: crates/ml/src/trainers/dqn/trainer/training_loop.rs:109-116

  • Step 1: Remove the GPU data freeing block

Delete lines 109-116 in training_loop.rs:

// Free DqnGpuData when GPU experience collector is active
if gpu_experiences_collected && self.gpu_data.is_some() {
    info!("GPU experience collector active — releasing DqnGpuData to reclaim VRAM");
    self.gpu_data = None;
    if let Some(pool) = self.buffer_pool.take() {
        drop(pool);
    }
}

The training data is immutable within a walk-forward window. The experience collector has its own buffer copies. Freeing and re-uploading wastes 73ms/epoch.

  • Step 2: Verify compilation + tests

Run: SQLX_OFFLINE=true cargo check -p ml Run: SQLX_OFFLINE=true cargo test -p ml --lib

  • Step 3: GPU smoke test

Run: SQLX_OFFLINE=true FOXHUNT_TEST_DATA=test_data/futures-baseline cargo test -p ml --lib -- test_train_step_produces_finite_metrics --ignored --test-threads=1

  • Step 4: Profile

Run: SQLX_OFFLINE=true cargo run --release --example train_baseline_rl -p ml -- --model dqn --data-dir test_data/futures-baseline --symbol ES.FUT --epochs 5 --max-steps-per-epoch 50 --batch-size 64 --train-months 3 --val-months 1 --test-months 1 --step-months 3 2>&1 | grep "phase breakdown"

Expected: init < 5ms (down from 73ms), total epoch < 100ms.

  • Step 5: Commit
git add crates/ml/src/trainers/dqn/trainer/training_loop.rs
git commit -m "perf: keep GPU training data resident across epochs (eliminate 73ms/epoch re-upload)"