Files
foxhunt/docs/superpowers/plans/2026-03-26-eval-cleanup.md
jgrusewski 6414f57e34 docs: eval cleanup spec + plan — remove dead 5-action path, fix train/eval mismatch
3 fixes: (1) Delete GpuActionSelector + branching_action_select (zero callers,
hardcoded 5 actions vs production 9), (2) Add hold enforcement to backtest
env_step kernel (same Layer 2 as training), (3) Dynamic trade insufficiency
threshold = max(window_bars / (min_hold * 20), 5).

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-26 00:43:58 +01:00

7.0 KiB

Evaluation Cleanup — Implementation Plan

For agentic workers: REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (- [ ]) syntax for tracking.

Goal: Remove dead 5-action branching_action_select path, add hold enforcement to backtest evaluator, make trade penalty threshold dynamic.

Architecture: Delete gpu_action_selector.rs + epsilon_greedy_kernel.cu (zero callers). Add min_hold_bars override to backtest_env_step kernel. Dynamic trade penalty threshold = max(window_bars / (min_hold_bars * 20), 5).

Tech Stack: CUDA kernel (C), Rust, cudarc 0.19.3.


Task 1: Remove dead code — GpuActionSelector + branching_action_select

Files:

  • Delete: crates/ml/src/cuda_pipeline/gpu_action_selector.rs

  • Delete: crates/ml/src/cuda_pipeline/epsilon_greedy_kernel.cu

  • Modify: crates/ml/src/cuda_pipeline/mod.rs — remove module declaration

  • Modify: crates/ml/src/trainers/dqn/trainer/action.rs — remove select_actions_batch()

  • Modify: crates/ml/src/trainers/dqn/trainer/mod.rs — remove GpuActionSelector field + imports

  • Step 1: Remove module declaration from mod.rs

In crates/ml/src/cuda_pipeline/mod.rs, find and remove:

pub mod gpu_action_selector;
  • Step 2: Remove GpuActionSelector from DQNTrainer

In trainer/mod.rs, remove the GpuActionSelector field and any use imports referencing it. Search for gpu_action_selector and GpuActionSelector.

  • Step 3: Remove select_actions_batch() from action.rs

In trainer/action.rs, remove the entire select_actions_batch method and any helper functions only used by it. If the file becomes empty/trivial, consider removing it.

  • Step 4: Delete the files
rm crates/ml/src/cuda_pipeline/gpu_action_selector.rs
rm crates/ml/src/cuda_pipeline/epsilon_greedy_kernel.cu
  • Step 5: Fix any remaining references
grep -rn "gpu_action_selector\|GpuActionSelector\|branching_action_select\|epsilon_greedy_kernel\|select_actions_batch" crates/ml/src/ --include="*.rs"

Remove any remaining references (comments, imports, doc references).

  • Step 6: Build + test

Run: SQLX_OFFLINE=true cargo check -p ml --lib Run: SQLX_OFFLINE=true cargo test -p ml --lib Expected: compiles, all tests pass (some tests that used GpuActionSelector may need removing)

  • Step 7: Commit
git commit -m "refactor: remove dead GpuActionSelector + branching_action_select (5-action)

Zero callers in training or evaluation. The backtest evaluator uses
evaluate_dqn_graphed() with cuBLAS + experience_action_select (9-action).
Removes train/eval action space mismatch (was 5 vs 9 actions)."

Task 2: Add hold enforcement to backtest env_step kernel

Files:

  • Modify: crates/ml/src/cuda_pipeline/backtest_env_kernel.cu

  • Modify: crates/ml/src/cuda_pipeline/gpu_backtest_evaluator.rs

  • Step 1: Add min_hold_bars parameter to backtest_env_step kernel

In backtest_env_kernel.cu, add int min_hold_bars as the last parameter to the kernel signature.

  • Step 2: Add hold enforcement logic

After the position update and sign computation, add the same Layer 2 hold enforcement as the training kernel:

// Hold enforcement: block exits and reversals during minimum hold period
int wants_exit = (prev_sign != 0 && curr_sign == 0);
int wants_reversal = (prev_sign != 0 && curr_sign != 0 && prev_sign != curr_sign);
if (hold_time > 0.0f && hold_time < (float)min_hold_bars
    && (wants_exit || wants_reversal)
    && step < window_len - 1) {
    // Override: keep previous position
    position = prev_position;
    // Don't update portfolio for this bar
}

Adapt the variable names to match the backtest kernel's conventions (it may use different names than the training kernel).

  • Step 3: Pass min_hold_bars from Rust

In gpu_backtest_evaluator.rs, add min_hold_bars: i32 to GpuBacktestConfig. Pass it to the kernel launch. Default: 5.

Wire from the hyperopt adapter: when creating the evaluator config, pass hyperparams.min_hold_bars.

  • Step 4: Build + test

Run: SQLX_OFFLINE=true cargo check -p ml --lib Run: SQLX_OFFLINE=true cargo test -p ml --lib -- backtest Expected: compiles, backtest tests pass

  • Step 5: Commit
git commit -m "feat: hold enforcement in backtest_env_step — fixes train/eval mismatch

Backtest evaluator now enforces min_hold_bars during evaluation,
matching the training kernel's Layer 2 hold guard. The model is
evaluated under the same rules it learned under."

Task 3: Dynamic trade insufficiency threshold

Files:

  • Modify: crates/ml/src/hyperopt/adapters/dqn.rs

  • Step 1: Find calculate_trade_insufficiency_penalty usage

Search for the function call in extract_objective(). Currently:

let trade_penalty = calculate_trade_insufficiency_penalty(backtest.total_trades);
  • Step 2: Make threshold dynamic

Replace with:

// Dynamic threshold: adapts to window size and min_hold_bars
// With min_hold=5 and 10K window: max_trades=2000, threshold=max(100, 5)=100
// With min_hold=20 and 10K window: max_trades=500, threshold=max(25, 5)=25
let backtest_window_bars = 10_000_usize; // MAX_WINDOW_BARS from evaluator
let min_hold = params.min_hold_bars.max(1);
let max_possible_trades = backtest_window_bars / min_hold;
let dynamic_threshold = (max_possible_trades / 20).max(5);
let trade_penalty = calculate_trade_insufficiency_penalty_dynamic(
    backtest.total_trades,
    dynamic_threshold,
);
  • Step 3: Create calculate_trade_insufficiency_penalty_dynamic()

Add a new function (or modify the existing one) that accepts a threshold parameter:

fn calculate_trade_insufficiency_penalty_dynamic(total_trades: usize, min_threshold: usize) -> f64 {
    if total_trades < min_threshold {
        // Smooth ramp from 10.0 (zero trades) to 0.0 (at threshold)
        let ratio = total_trades as f64 / min_threshold as f64;
        10.0 * (1.0 - ratio)
    } else {
        0.0
    }
}
  • Step 4: Build + test

Run: SQLX_OFFLINE=true cargo test -p ml --lib -- objective Expected: objective tests pass

  • Step 5: Commit
git commit -m "feat: dynamic trade insufficiency threshold — adapts to min_hold_bars

Threshold = max(window_bars / (min_hold * 20), 5). With min_hold=5 and
10K window, model needs 100+ trades to avoid penalty (was fixed at 10).
Prevents degenerate objective when hold enforcement reduces trade count."

Task 4: Integration test

  • Step 1: SQLX_OFFLINE=true cargo test -p ml --lib — 887+ pass
  • Step 2: SQLX_OFFLINE=true cargo test -p ml-dqn --lib — 359 pass
  • Step 3: SQLX_OFFLINE=true cargo check --workspace — clean
  • Step 4: Push to main

Execution Order

Task 1 (delete dead code) → Task 2 (backtest hold enforcement) → Task 3 (dynamic threshold) → Task 4 (integration)