Files
foxhunt/docs/superpowers/specs/2026-03-26-eval-cleanup-design.md
jgrusewski 6414f57e34 docs: eval cleanup spec + plan — remove dead 5-action path, fix train/eval mismatch
3 fixes: (1) Delete GpuActionSelector + branching_action_select (zero callers,
hardcoded 5 actions vs production 9), (2) Add hold enforcement to backtest
env_step kernel (same Layer 2 as training), (3) Dynamic trade insufficiency
threshold = max(window_bars / (min_hold * 20), 5).

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-26 00:43:58 +01:00

71 lines
3.9 KiB
Markdown

# Evaluation Cleanup — Remove Dead Path + Fix Train/Eval Mismatch
## Problem
1. **Dead code**: `GpuActionSelector`, `branching_action_select` (5-action hardcoded), `select_actions_batch()` have zero callers in training or evaluation. The backtest evaluator uses `evaluate_dqn_graphed()` with cuBLAS forward + `experience_action_select` (9-action). The dead path causes confusion and was the target of Task 5 (which we correctly skipped).
2. **Train/eval mismatch**: Training enforces `min_hold_bars` via action masking in `experience_action_select`. The backtest evaluator's action selection path does NOT enforce holds. The model learns under hold constraints but is evaluated without them.
3. **Degenerate objective with hold enforcement**: The trade insufficiency penalty threshold is hardcoded. With high `min_hold_bars`, fewer trades complete, and the threshold should adapt dynamically.
## Design
### 1. Remove dead code
Delete entirely:
- `crates/ml/src/cuda_pipeline/gpu_action_selector.rs`
- `crates/ml/src/cuda_pipeline/epsilon_greedy_kernel.cu`
Remove from parent modules:
- `gpu_action_selector` module declaration in `crates/ml/src/cuda_pipeline/mod.rs`
- `select_actions_batch()` method in `crates/ml/src/trainers/dqn/trainer/action.rs`
- `GpuActionSelector` field in `DQNTrainer` struct (`trainer/mod.rs`)
- Any `use` imports referencing these
### 2. Add hold enforcement to backtest evaluator
The `evaluate_dqn_graphed()` path in `gpu_backtest_evaluator.rs` runs its own step loop:
1. `gather_states` → state tensor
2. cuBLAS forward → Q-values
3. `compute_expected_q` → expected Q from C51 distribution
4. `backtest_greedy_argmax` or action selection → action indices
5. `backtest_env_step` → portfolio update
The action selection at step 4 needs hold enforcement. The evaluator already has `portfolio_buf` (portfolio state per window). Two options:
**(A)** Add hold masking to the existing `backtest_greedy_argmax` kernel
**(B)** Pass `min_hold_bars` to the backtest env_step kernel and let it override (Layer 2 style)
**Choose B** — the backtest `env_step` kernel already manages portfolio state. Adding the hold override there is simpler and matches the training kernel's Layer 2 pattern. The backtest doesn't need Layer 1 (action masking) because it uses greedy argmax (no epsilon exploration to mask).
Add `min_hold_bars` parameter to `backtest_env_step` kernel. When `hold_time > 0 && hold_time < min_hold_bars`, override action to maintain current position (same logic as training kernel's Layer 2).
### 3. Dynamic trade insufficiency threshold
In `extract_objective()` in `hyperopt/adapters/dqn.rs`, the trade penalty currently uses a fixed threshold via `calculate_trade_insufficiency_penalty(total_trades)`. Make it dynamic:
```rust
// Dynamic threshold based on window size and min_hold_bars
let max_possible_trades = backtest_window_bars / min_hold_bars.max(1);
let min_trades_threshold = (max_possible_trades / 20).max(5);
let trade_penalty = calculate_trade_insufficiency_penalty_dynamic(
backtest.total_trades,
min_trades_threshold,
);
```
This ensures the penalty adapts: with `min_hold_bars=5` and 10K-bar windows, threshold = max(500/20, 5) = 25. With `min_hold_bars=20`, threshold = max(250/20, 5) = 12. Always reasonable.
## Files Modified
| File | Change |
|------|--------|
| `crates/ml/src/cuda_pipeline/gpu_action_selector.rs` | DELETE |
| `crates/ml/src/cuda_pipeline/epsilon_greedy_kernel.cu` | DELETE |
| `crates/ml/src/cuda_pipeline/mod.rs` | Remove gpu_action_selector module |
| `crates/ml/src/trainers/dqn/trainer/action.rs` | Remove select_actions_batch() |
| `crates/ml/src/trainers/dqn/trainer/mod.rs` | Remove GpuActionSelector field + imports |
| `crates/ml/src/cuda_pipeline/backtest_env_kernel.cu` | Add min_hold_bars hold override |
| `crates/ml/src/cuda_pipeline/gpu_backtest_evaluator.rs` | Pass min_hold_bars to env kernel |
| `crates/ml/src/hyperopt/adapters/dqn.rs` | Dynamic trade insufficiency threshold |