Files
foxhunt/crates
jgrusewski 679a483de5 feat: wire CVaR action selection + implement CQL conservative loss GPU kernel
Task 4 — CVaR Action Selection:
- Add use_cvar_action_selection and cvar_alpha fields to DQNHyperparameters
- Wire from hyperparams into DQNConfig constructor (was hardcoded to false)
- Default: enabled (true) with alpha=0.05 (worst 5% quantile tail)
- Unblocks risk-aware position scaling via IQN head's compute_cvar_q()

Task 5 — Curiosity Wiring (verified active):
- GpuCuriosityTrainer trains forward model on GPU experience data
- train_curiosity_gpu() called from training_loop after experience collection
- curiosity_weight=0.05 (Task 2) gates trainer creation — active when >0
- Intrinsic reward injection into DQN kernel deferred (Phase 4+, per kernel docs)

Task 6 — CQL Conservative Loss GPU Kernel:
- Add use_cql/cql_alpha to GpuDqnTrainConfig (wired from DQNHyperparameters)
- Implement cql_logit_grad_kernel: computes dCQL/d_logits for Branching Dueling C51
  - Per-branch logsumexp penalty with softmax gradient through expectation chain
  - One thread per sample, iterates 3 branches (exposure, order, urgency)
- Add apply_cql_gradient() method: launches CQL kernel + cuBLAS backward_full
  - Accumulates CQL parameter gradients into grad_buf (beta=1.0)
  - Same injection pattern as IQN trunk gradient
- Wire into FusedTrainingCtx::run_full_step() between graph_forward and graph_adam

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-24 22:45:49 +01:00
..