plan(perf): K-loop block-per-batch parallelization implementation plan
Bite-sized 22-task plan implementing the design at
docs/superpowers/specs/2026-05-17-kloop-parallelization-design.md.
Four atomic commits per the spec's Rollout section:
Commit 1: reduce_axis0 kernel + cfc_step refactor (Tasks 1-10)
Commit 2: GRN bwd refactor (Tasks 11-14)
Commit 3: VSN bwd refactor (Tasks 15-17)
Commit 4: attention_pool bwd refactor (Tasks 18-20)
Each commit covers kernel rewrite + per-batch scratch buffers +
reducer launches + memset_zeros wiring + smoke validation.
Tests added across commits:
- cfc_bwd_b1_oracle.rs (Task 6): B=1 oracle vs single-sample helper
within relative_eq 1e-5 (FP-tolerant, not bit-exact)
- stacked_trainer_loss_shrinks_at_batch_32 (Task 7): first test
that exercises the cross-batch reducer path
- cfc_bwd_scratch_clears_between_steps (Task 8): scratch-zero
regression guard
Acceptance gates 1-8 from the spec are mapped to Tasks 6, 7, 8, 9,
21, 22 (local + cluster A/B perf). Reference baseline for gate #8
is t6z89's per-epoch AUC trajectory.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
This commit is contained in:
1861
docs/superpowers/plans/2026-05-17-kloop-parallelization.md
Normal file
1861
docs/superpowers/plans/2026-05-17-kloop-parallelization.md
Normal file
File diff suppressed because it is too large
Load Diff
Reference in New Issue
Block a user