Add magnitude head (branch 3) to BranchingWeightSet, GpuBranchPtrs, and all
BF16 mirror structs. Update flatten/unflatten from 20 to 24 individual tensors
(indices 24-25 = bottleneck remain flat-buffer-only). Fix bottleneck index
in experience collector (was 20, now 24). Update all construction sites:
fused_training clone, gradient_budget smoke tests, backtest evaluator.
Fix monitoring comments/labels for 4-branch encoding (dir*mag 3x3), delete
dead `if false` block, delete DELETED marker comments.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>