feat(dqn-v2): D.3 horizon-decomposed V — widen IQL value head to 2 outputs
IQL value head FC expanded from [D_in × 1] to [D_in × 2]. v_out_buf shape
[B] → [B*2]. Consumer code sums the two outputs (V = V_short + V_long)
wherever a scalar V was previously read. Both outputs trained via the
same expectile loss (horizon-specific regression is a follow-up
enhancement — current form provides the architectural capacity for
horizon decomposition without per-horizon targeting).
Changes:
- total_params: w3 H*1+b3[1] → w3 H*2+b3[2]
- gemm_fwd_v: M=1→2; gemm_bwd_dw3: M=1→2; gemm_bwd_dh2: K=1→2
- v_out_buf, dv_buf, loss_buf: [B] → [B*2]
- iql_expectile_loss kernel: new num_heads param; q_taken[b]=q_taken[idx/num_heads]
- iql_loss_reduce kernel: new num_heads param; normalises by B (not B*num_heads)
- bias_add for b3: out_dim=1→2, N=B→B*2
- db3 bias_grad_reduce: gridDim.y=1→2 for per-head gradient accumulation
- V_W3_SIZE/V_B3_SIZE macros: H→H*2, 1→2 (used by iql_forward_kernel)
- iql_forward_kernel: updated for 2-output col-major [2,B] write
- 4 consumer kernels: v_out[b] → v_out[b*2+0] + v_out[b*2+1]
- Xavier init: w3 fan_out 1→2, w3_end H→H*2
Checkpoint compat: IQL parameter count changes. Layout fingerprint
recomputes; old checkpoints fail-fast at load per spec §4.A.2.
Retrain required.
Smoke test: training runs 28s, best Sharpe=80.59 (baseline ~80), 0 errors.
Unit tests: 889 pass / 12 fail (12 pre-existing, 0 regressions introduced).
Plan 2 Task 6B. Spec §4.D.3.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>