Full backward pass for the 2D bottleneck via cuBLAS chain rule:
1. bn_tanh_backward_kernel: d_bn = d_concat[:,:bn_dim] * (1 - tanh^2)
Reads tanh values from forward pass (bn_hidden_buf), applies derivative
2. cast_dx_to_staging: f32 d_bn → bf16 for cuBLAS dW GEMM
3. launch_dw_only: dW_bn[bn_dim, market_dim] += d_bn^T @ states
Uses same cuBLAS infrastructure as all other weight gradient GEMMs
4. bn_bias_grad_kernel: db_bn = sum(d_bn, dim=0)
backward_full() modified: new s1_dx_output parameter computes
d_loss/d_bn_concat when bottleneck is active (was 0 = skip).
Gradients flow through entire bottleneck → tanh → GEMM chain.
Adam optimizer trains bottleneck weights (tensors 20-21) alongside
all other parameters — same flat grad_buf, same spectral norm,
same weight decay. No special handling needed.
3 new CUDA kernels: bn_tanh_backward, bn_bias_grad, bn_tanh_concat.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>