Files
foxhunt/docs/superpowers/specs/2026-03-22-two-phase-hyperopt-design.md
jgrusewski 43998a330a feat: two-phase hyperopt + backtest evaluator VRAM leak fix
Two-phase hyperopt splits 31D PSO search into sequential phases:
- Phase 1 (--phase fast, default): fix architecture to small network
  (hidden_dim=128, num_atoms=11), search learning dynamics (~15D).
- Phase 2 (--phase full): fix dynamics from Phase 1 JSON, search
  architecture (~5D). Halves dimensionality per phase → better convergence.
- Phase 1 output includes best_continuous_vector for Phase 2 consumption.

GpuBacktestEvaluator Drop impl: sync forked stream, destroy CUDA graph
and cuBLAS handles before CudaSlice buffers drop. Fixes 261MB/trial
VRAM leak on H100 hyperopt.

ml-core clippy fixes: hex literal, remove dead check_err drain,
unnecessary safety comment, unused OnceLock import.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-22 10:32:37 +01:00

3.6 KiB
Raw Blame History

Two-Phase Hyperopt: Search Quality over Speed

Goal

Split the 31D PSO search into two sequential phases: first optimize learning dynamics with a fixed small network, then optimize architecture with the best dynamics locked in. Better search quality from lower-dimensional spaces, with speed as a bonus.

Problem

The current 31D search mixes learning dynamics (lr, gamma, tau) with architecture (hidden_dim, num_atoms). A bad lr makes ALL network sizes look bad — PSO wastes evaluations on redundant lr×hidden_dim combinations. Halving the dimensionality at each phase dramatically improves convergence.

Architecture

Phase 1: Learning Dynamics (fast, ~30s/trial)

Fixed: hidden_dim_base=128, num_atoms=11, branch_hidden_dim=64 Search (~15D): learning_rate, gamma, tau, batch_size, buffer_size, per_alpha, per_beta_start, entropy_coefficient, cql_alpha, weight_decay, noisy_sigma_init, n_steps, kelly_fractional, dsr_eta, iqn_lambda

Config: 50 trials × 30 epochs at ~173ms/epoch = 5s training + 30s data/backtest = ~35s/trial. Total: ~30 min.

Output: Best learning dynamics config (JSON).

Phase 2: Architecture (production speed, ~2 min/trial)

Fixed: Best lr, gamma, tau, PER params from Phase 1 Search (~5D): hidden_dim_base [128-512], num_atoms [11-51], branch_hidden_dim [64-256], dueling_hidden_dim [128-512], v_range [10-50]

Config: 20 trials × 50 epochs at production speed. Total: ~40 min.

Output: Best full config (JSON) for production training.

Implementation

CLI changes to hyperopt_baseline_rl

Add --phase argument:

--phase fast      # Phase 1: fix architecture, search dynamics (default)
--phase full      # Phase 2: fix dynamics from --hyperopt-params, search architecture
--phase single    # Current behavior: search everything in one pass

Add --fix-hidden-dim <N> to override hidden_dim_base for Phase 1. Add --fix-num-atoms <N> to override num_atoms for Phase 1.

Hyperopt adapter changes

In DQNParams::continuous_bounds():

  • If phase=fast: fix hidden_dim_base, num_atoms, branch_hidden_dim to constants. Their bounds become (value, value) — single points in the search space.
  • If phase=full: fix learning dynamics from the Phase 1 JSON. Load best params, set their bounds to (best, best).

This requires NO changes to the PSO algorithm — fixed params are just single-point bounds. The optimizer trivially converges on them and focuses search on the free dimensions.

Config changes

Add to config/training/dqn-hyperopt.toml:

[phase_fast]
hidden_dim_base = 128
num_atoms = 11
branch_hidden_dim = 64

[phase_full]
# Loaded from --hyperopt-params JSON (Phase 1 output)

Workflow

# Phase 1: find best learning dynamics (~30 min)
cargo run --example hyperopt_baseline_rl -- \
  --model dqn --phase fast --trials 50 --epochs 30 \
  --data-dir data/futures-baseline --symbol ES.FUT \
  --output phase1_results.json

# Phase 2: find best architecture with locked dynamics (~40 min)
cargo run --example hyperopt_baseline_rl -- \
  --model dqn --phase full --trials 20 --epochs 50 \
  --hyperopt-params phase1_results.json \
  --data-dir data/futures-baseline --symbol ES.FUT \
  --output phase2_results.json

# Production training with best config
cargo run --example train_baseline_rl -- \
  --model dqn --epochs 100 \
  --hyperopt-params phase2_results.json \
  --data-dir data/futures-baseline --symbol ES.FUT

Non-Goals

  • Automatic phase chaining (user runs Phase 1, reviews, then Phase 2)
  • Multi-objective optimization (Sharpe + stability)
  • Architecture search beyond hidden dims (layer count, residual connections)