Wires the aux supervision path parallel to BCE:
- AuxTrunk (64-hidden single-bucket CfC) consumes the same encoder output
as the main BCE trunk
- AuxHeads (linear regression long/short) maps aux_trunk output to
per-(direction, horizon) predicted outcomes
- AuxHuberLoss supervises against D-style labels from MultiHorizonLoader
Backward path with asymmetric stop-grad at encoder boundary:
- aux_trunk gets gradient signal into its OWN params at all times
- aux_trunk's encoder-boundary gradient is INITIALLY blocked
(stop_grad_aux_to_encoder = true)
- Conditional lift per E3 design: if aux_huber_ema < 0.4 AND
aux_dir_acc_ema > 0.85 within 200 steps, lift the stop-grad
- When lifted, aux_vec_add kernel folds aux's grad_x into the main
grad_h_enriched_seq slot (element-wise += per feedback_no_atomicadd)
ISV signals added: aux_huber_per_h, aux_dir_acc_per_h (per pearl).
Per-trunk scratch + reduced grad buffers (no Adam state sharing per
pearl_adam_normalizes_loss_weights — opt_aux is its own Adam group).
New helper kernel cuda/aux_vec_add.cu: position-local dst += src for
the asymmetric stop-grad lift accumulation.
New synthetic test stacked_trainer_aux_supervision_converges_on_constant_signal
validates end-to-end:
aux_huber_ema_per_h = [0.087, 0.087, 0.087] (converged)
aux_dir_acc_ema_per_h = [1.0, 1.0, 1.0] (perfect on constant)
stop_grad_aux_to_encoder = false (lift fired)
All 5 stacked_trainer tests pass on RTX 3050 (lib still converges, no
regression from parallel aux wiring).
Not yet consumed by decision policy (B7) — aux output flows through
training only.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>