3 fixes: (1) Delete GpuActionSelector + branching_action_select (zero callers, hardcoded 5 actions vs production 9), (2) Add hold enforcement to backtest env_step kernel (same Layer 2 as training), (3) Dynamic trade insufficiency threshold = max(window_bars / (min_hold * 20), 5). Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
3.9 KiB
Evaluation Cleanup — Remove Dead Path + Fix Train/Eval Mismatch
Problem
-
Dead code:
GpuActionSelector,branching_action_select(5-action hardcoded),select_actions_batch()have zero callers in training or evaluation. The backtest evaluator usesevaluate_dqn_graphed()with cuBLAS forward +experience_action_select(9-action). The dead path causes confusion and was the target of Task 5 (which we correctly skipped). -
Train/eval mismatch: Training enforces
min_hold_barsvia action masking inexperience_action_select. The backtest evaluator's action selection path does NOT enforce holds. The model learns under hold constraints but is evaluated without them. -
Degenerate objective with hold enforcement: The trade insufficiency penalty threshold is hardcoded. With high
min_hold_bars, fewer trades complete, and the threshold should adapt dynamically.
Design
1. Remove dead code
Delete entirely:
crates/ml/src/cuda_pipeline/gpu_action_selector.rscrates/ml/src/cuda_pipeline/epsilon_greedy_kernel.cu
Remove from parent modules:
gpu_action_selectormodule declaration incrates/ml/src/cuda_pipeline/mod.rsselect_actions_batch()method incrates/ml/src/trainers/dqn/trainer/action.rsGpuActionSelectorfield inDQNTrainerstruct (trainer/mod.rs)- Any
useimports referencing these
2. Add hold enforcement to backtest evaluator
The evaluate_dqn_graphed() path in gpu_backtest_evaluator.rs runs its own step loop:
gather_states→ state tensor- cuBLAS forward → Q-values
compute_expected_q→ expected Q from C51 distributionbacktest_greedy_argmaxor action selection → action indicesbacktest_env_step→ portfolio update
The action selection at step 4 needs hold enforcement. The evaluator already has portfolio_buf (portfolio state per window). Two options:
(A) Add hold masking to the existing backtest_greedy_argmax kernel
(B) Pass min_hold_bars to the backtest env_step kernel and let it override (Layer 2 style)
Choose B — the backtest env_step kernel already manages portfolio state. Adding the hold override there is simpler and matches the training kernel's Layer 2 pattern. The backtest doesn't need Layer 1 (action masking) because it uses greedy argmax (no epsilon exploration to mask).
Add min_hold_bars parameter to backtest_env_step kernel. When hold_time > 0 && hold_time < min_hold_bars, override action to maintain current position (same logic as training kernel's Layer 2).
3. Dynamic trade insufficiency threshold
In extract_objective() in hyperopt/adapters/dqn.rs, the trade penalty currently uses a fixed threshold via calculate_trade_insufficiency_penalty(total_trades). Make it dynamic:
// Dynamic threshold based on window size and min_hold_bars
let max_possible_trades = backtest_window_bars / min_hold_bars.max(1);
let min_trades_threshold = (max_possible_trades / 20).max(5);
let trade_penalty = calculate_trade_insufficiency_penalty_dynamic(
backtest.total_trades,
min_trades_threshold,
);
This ensures the penalty adapts: with min_hold_bars=5 and 10K-bar windows, threshold = max(500/20, 5) = 25. With min_hold_bars=20, threshold = max(250/20, 5) = 12. Always reasonable.
Files Modified
| File | Change |
|---|---|
crates/ml/src/cuda_pipeline/gpu_action_selector.rs |
DELETE |
crates/ml/src/cuda_pipeline/epsilon_greedy_kernel.cu |
DELETE |
crates/ml/src/cuda_pipeline/mod.rs |
Remove gpu_action_selector module |
crates/ml/src/trainers/dqn/trainer/action.rs |
Remove select_actions_batch() |
crates/ml/src/trainers/dqn/trainer/mod.rs |
Remove GpuActionSelector field + imports |
crates/ml/src/cuda_pipeline/backtest_env_kernel.cu |
Add min_hold_bars hold override |
crates/ml/src/cuda_pipeline/gpu_backtest_evaluator.rs |
Pass min_hold_bars to env kernel |
crates/ml/src/hyperopt/adapters/dqn.rs |
Dynamic trade insufficiency threshold |