Per pearl_build_rs_rerun_if_env_changed, every std::env::var() must be paired with cargo:rerun-if-env-changed. Task #321 fixed crates/ml/build.rs (commite3d082968) but missed 6 sibling build.rs files. Each reads CUDA_COMPUTE_CAP without registering the rerun directive, so cargo sees no env-change between e.g. SM 89 (L40S) → SM 90 (H100) workflow re-submissions and cache-hits the wrong-arch cubin from the previous run. At runtime, the H100 fails to load the SM 89 cubin with CUDA_ERROR_NO_BINARY_FOR_GPU on rmsnorm — exactly what killed train-multi-seed-vlv8c (commit0371d6a76, post-Class-B chain). Files (all add `println!("cargo:rerun-if-env-changed=CUDA_COMPUTE_CAP")` just before the env::var() read): - crates/ml-dqn/build.rs (rmsnorm — root cause of vlv8c failure) - crates/ml-ensemble/build.rs - crates/ml-explainability/build.rs - crates/ml-ppo/build.rs - crates/ml-supervised/build.rs - crates/ml-core/build.rs Atomic single commit per feedback_no_partial_refactor. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
3.1 KiB
3.1 KiB