Files
foxhunt/crates
jgrusewski 6e0cb3d62c fix: CRITICAL — dense micro-reward + OFI embed read stale state positions
Task 3 refactored experience_state_gather (the WRITER) to use assemble_state()
with canonical layout (OFI at [42..62)). But env_step and ofi_embed_build_input
(the READERS) still hardcoded OFI at [66..84) — the OLD pre-refactor layout.

This meant 7 locations were reading MTF/portfolio features as if they were OFI:

1. Line 1745: dense micro-reward ofi_cur = state+66 → actually MTF[4]
2. Line 1778: book_aggression = state[82] → actually plan_isv region
3. Line 1983: ps[30..37] OFI delta storage for NEXT bar — storing MTF data
4. Lines 5833/5836/5838/5840: ofi_embed_build_input — feeds 18→10 MLP into
   Mamba2 temporal SSM and attention. Entire temporal pipeline was training
   on MTF features dressed as OFI.

Symptoms explained:
- WinRate=20.9% on validation (anti-correlated): dense micro-reward computes
  quality=sign_pos × garbage_MTF_deltas, systematically rewarding wrong direction
- mean_reward=+0.004 but Sharpe_raw=-0.0004: shaped reward exploits garbage
  signal, real portfolio loses money
- grad_norm=23560 at epoch 2: gradients chasing noise
- Q-value explosion to ±10 in one epoch: learning contradictions

Fix: replaced all hardcoded 66/74/82/83 with SL_OFI_START from state_layout.cuh.
Both reader kernels now use the same canonical layout as assemble_state().

Verified locally: smoke test passes, OFI_DIAG shows correct non-zero values
(raw_mean=-0.36, delta_mean=-0.21, log_dur=-0.23).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-20 17:30:25 +02:00
..