Replace closure-based evaluate() with evaluate_dqn_graphed() for non-OFI walk-forward backtest path. Extracts DuelingWeightSet from VarMap (branching or standard dueling) and runs hand-written warp-cooperative CUDA forward kernel with CUDA Graph capture — zero Candle dispatch overhead per step. Key changes: - GpuBacktestEvaluator::stream() getter for weight extraction on eval stream - DQNAgentType::is_using_branching() / network_dims() for CUDA kernel config - Hyperopt evaluate_gpu() non-OFI path: extract_dueling_weights_branching() → evaluate_dqn_graphed() (CUDA Graph accelerated) - OFI path: retains Candle closure for state permutation (gather kernel layout mismatch — future CUDA permutation kernel) - 66+ GPU hot-path violations hardened to hard errors across DQN/PPO/supervised - Stripped all gpu-ok suppression comments - Proper #[cfg(feature = "cuda")] gating for CUDA-only code paths 77 files, 0 errors, 0 warnings across workspace. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
7.2 KiB
Executable File
7.2 KiB
Executable File