- Pre-training: 50 batch supervised direction init at epoch 0 - Exposure aux targets: DtoD copy from collector to fused context - PopArt: wired behind config flag (disabled by default) - TD(λ)/hindsight/curriculum: stubs with config guards - Loss threshold 500→100K for v8 reward distribution Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>