Plan A v1 (commit2d68ef5d8= revert Phase 2.1 on top of104fe81ca) failed at cluster (alpha-rl-4sjzw): entropy collapsed to 0.69 by step 3000. The Phase 2.0 V envelope clamp (c52282fb4) — kept in v1 — was likely the culprit: it bounds V_pred to a tight envelope, biasing advantages and breaking PPO gradient flow. Plan A v2: start fromdd049d9a4(proven working at cluster wr=0.57 +$6.3M @ step 15000 today via alpha-rl-lpbp8) and apply ONLY the kernel-only bug fixes that don't affect training dynamics: - variable_selection.cu: VSN stride 40→56 fix (104fe81ca) — prevents step-4 NaN from reading wrong-stride window_tensor - bucket_transition_kernels.cu: h_mag_per_bucket multi-warp 32→128 (7e38e46e6) — correct warp-shuffle reduction for HIDDEN_DIM=128 - compute_advantage_return.cu: branch-gate done flag (a6acc25ec) — done's terminal-state semantics applied via gate, resolves step-4 NaN NOT applied (intentionally): -c52282fb4Phase 2.0 V envelope clamp — biases V regression target -b4aadff75V envelope ±200→±10 — extends Phase 2.0 -10d4614fbatomicAdd removal — kernel non-determinism is real but intrusive (Rust changes),dd049d9a4works without this fix -db4d9a16fPhase 2.1 dueling — broken decomposition per spec analysis - 72672c9e7+ Phase 2.2/2.3/H/P1+P2+P3 — built on broken Phase 2.1 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
9.7 KiB
9.7 KiB