TD(λ): self-bootstrap with rewards as Q(s') approximation, overwrites n-step Hindsight: relabel fraction of experience rewards with optimal exit PnL Curriculum: sort bars by difficulty, restrict to easy bars early, expand over training No stubs, no debug placeholders — all fully wired. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>