After a first dispatch attempt of monolithic Task 2 (GRN ADOPT) was
correctly refused with a thorough scope assessment, this revision
records the decomposition and the structural facts that drove it:
1. layout_fingerprint_seed() only fingerprints ISV slots, not
param-tensor layout. The "fingerprint will auto-update" assumption
in the original Task 2 spec was wrong for param-tensor reshuffles.
2. compute_param_sizes has 86 tensors (docstring saying 42 is stale).
Inserting GRN's 9 sub-tensors shifts 82 downstream tensors and
98 padded_byte_offset call sites.
3. At least 12 kernels consume h_s2 expecting ReLU non-negativity.
GRN's LayerNorm output is zero-mean (signed). Per-consumer
verification needed before swap.
4. crates/ml-supervised::tft::gated_residual is incompatible as a
port: different formula (sigmoid gate, not GLU split) and
incompatible tensor abstraction. New CUDA kernels from scratch.
Decomposition:
- Task 2a (research, no code): audit h_s2 consumers for ReLU vs
LayerNorm semantics. Output: per-consumer table.
- Task 2b (small, checkpoint break): extend fingerprint seed to
include param-tensor names + sizes. Pearl-aligned: complete
fingerprint coverage rather than partial.
- Task 2c (large, checkpoint break): GRN kernels + 98 call-site
migration + h_s2 consumer shims (per 2a). Blocked on 2a+2b.
Recommended order updated to thread 2a → 2b → 2c. Tasks 1, 6, 3
remain independent of 2c and can land in parallel where useful.
Pearl rules section added documenting the safety constraints
applied throughout the plan.
No code changes. Plan-doc only.