Files
foxhunt/docs
jgrusewski 6d0ac7beb3 fix(sp4): GPU-port calibrate_homeostatic_targets per feedback_no_cpu_compute_strict
Per-step host-side EMA loop at gpu_dqn_trainer.rs:4671 over 6 mapped-
pinned homeostatic-target slots was a feedback_no_cpu_compute_strict
violation discovered during the sweep audit (commit 6a6b58aec) but
deferred for scope. Sweep audit grid site #9.

Migrated:
  - New calibrate_homeostatic_kernel.cu — single-block, six threads
    (one thread per homeostatic slot). Reads host-passed `readiness`
    scalar (already-clamped IQN gauge from `iqn_readiness` shadow field,
    bit-for-bit match of the deleted host clamp), observations from
    `homeostatic_obs_dev_ptr`, applies adaptive α
    `0.3 × (1 - readiness) + 0.01 × readiness` to targets[k] for k=1..5;
    thread 0 forces the Q-mean invariant `targets[0] = 0.0` exactly as
    the deleted host post-loop assignment did. __threadfence_system()
    after writes for PCIe-visibility to homeostatic_kernel's dev_ptr reads.
  - build.rs cubin registration and trainer-struct wiring (cubin static,
    field, struct constructor, cubin load) mirror the C2/C3/C4 pattern
    from the prior sweep commits.
  - Host-side `for k in 0..HOMEOSTATIC_N_OBS` loop in
    `calibrate_homeostatic_targets` replaced with a single
    `launch_calibrate_homeostatic` kernel launch; chained on the
    trainer's stream so it remains graph-capture-compatible.

Preserved:
  - Same α formula, same Q-mean=0 invariant, same call ordering.
  - Mapped-pinned target buffer retained — homeostatic_kernel still
    reads via homeostatic_targets_dev_ptr unchanged.
  - No cold-start sentinel: constructor pre-initialises targets to
    `[0.0, 0.85, 0.1, 0.0, 1.0, 0.5]` (gpu_dqn_trainer.rs:14583-14588)
    so the first EMA call blends defaults with the first observation,
    same algebraic shape the deleted host loop relied on.
  - State-reset registry unchanged — deleted host loop had no fold reset
    (per-call EMA only); GPU port preserves identical per-call semantics.

cargo check clean. SP4 + state_reset_registry lib tests pass (11/11).
16/16 SP4 producer GPU tests pass on RTX 3050 Ti. No behavior change —
pure architectural fix.

Refs: feedback_no_cpu_compute_strict sweep audit grid site #9.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-01 16:14:22 +02:00
..