multi_horizon_heads_backward: sigmoid + linear chain rule. One block, HIDDEN_DIM=128 threads. Computes grad_w, grad_b, grad_h_in. No atomicAdd; per-thread accumulation only. cfc_step_backward: truncated K=1 BPTT through one CfC time step. Forward pre/decay/tanh recomputed inside the kernel; emits grad_w_in, grad_w_rec, grad_b, grad_h_old. tau is held frozen (structural log-uniform init per Hasani 2022; backprop through tau deferred to Phase A v2 if the gate needs it). Uses dynamic shared memory for the d_pre relay between threads (size = 2 * n_hid * 4 bytes). Tests (4/4 on sm_86) validate via on-GPU finite-difference: - heads grad_h vs forward(h±eps) → matches at eps=1e-3, rel<=1% - heads grad_b vs forward(b±eps) → matches at eps=1e-3, rel<=1% - cfc grad_b vs forward(b±eps) → matches at eps=1e-3, rel<=5% - cfc grad_h_old vs forward(h_old±eps) → matches at eps=1e-3, rel<=5% CPU is not the reference (per feedback_no_cpu_test_fallbacks.md). The kernel is the truth; numerical perturbation validates the analytic gradient against the kernel's own forward. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
5.8 KiB
5.8 KiB