Phase 0 (commit 333ea7184): exhaustive inventory of every numeric value
emitted across the 7 HEALTH_DIAG sites in training_loop.rs and metrics.rs.
Classified each as GPU-already / CPU-bound / Mixed / Host-state /
ISV-already. Aggregate cost: ~50-150 MB/epoch DtoH + ~70 sec/epoch CPU
compute (the original observation that motivated the port). Inventory
lives at docs/health_diag_inventory.md.
Phase 1 (commit 60fd7de96): foundation only — no kernel, no producer
launch, no CPU code deletion yet.
* HealthDiagSnapshot #[repr(C)] POD with 147 fields (all f32 or u32)
covering every numeric value in the HEALTH_DIAG line family. Field
order matches the existing log format strings (downstream parsers
like aggregate-multi-seed-metrics.py depend on this).
* Controller fire-bools promoted from u8 to u32 to avoid #[repr(C)]
padding mismatch when followed by f32 fields. ~24-byte cost.
* MappedHealthDiagSnapshot wrapper around cuMemHostAlloc(DEVICEMAP|
PORTABLE) + cuMemHostGetDevicePointer_v2, mirroring the
MappedF32Buffer pattern.
* Default impl via MaybeUninit::zeroed() (allocator-free, valid
bit-pattern).
* 3 unit tests pinning size (588 bytes), alignment (4 bytes),
zeroed-default — all passing on local CPU host.
Phases 2-5 deferred (kernel family + wiring + CPU emit rewrite + old
code deletion). The 5 kernels each read different device buffers with
different masking semantics — getting any one wrong would silently
produce numerically-different HEALTH_DIAG output that downstream
parsers would mis-parse. Per the prompt's 'don't wedge yourself'
guidance and feedback_no_quickfixes.md, the safer move was to land
the typed wrapper cleanly so subsequent commits can iterate kernel
by kernel against the actual buffer-by-buffer reduction.