Commit Graph

2045 Commits

Author SHA1 Message Date
jgrusewski
b05036ab93 feat(F5/D2): Q-gap barrier gradient — piggybacks on CQL SAXPY path
Add `barrier_gradient_direction` CUDA kernel to c51_loss_kernel.cu that
computes barrier = max(0, 0.05*health - q_gap) from the direction branch
logits and ISV[12], then injects gradient via atomicAdd into the CQL
d-logit accumulator buffers. When barrier > 0, it raises Q(argmax) and
lowers Q(second_max) to widen the direction Q-gap.

Wire-up: `apply_cql_gradient` now accepts `barrier_weight` and inlines
the kernel launch after the CQL kernel but before `backward_full`, so
both CQL and barrier share one cuBLAS backward pass with no extra SAXPY.
`submit_aux_ops` passes `barrier_weight = 0.05` every step (kernel is
internally a no-op when q_gap >= min_req). D2/N2 epoch-boundary block
updated to reflect that gradient is now live.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-20 21:35:06 +02:00
jgrusewski
0bd5c68f56 feat(F3): spectral_gap via rolling Gram + power iteration (sigma_1/sigma_2)
Replace single-sample max/min proxy with mathematically proper sigma_1/sigma_2
ratio. Maintains a host-side VecDeque of up to 64 Q-value samples; once ≥ 8
samples are available computes the n_cols×n_cols Gram matrix X^T X, then
extracts the two largest singular values via power iteration + rank-one
deflation. Falls back to the coarse max/min ratio until the buffer fills.
compute_q_spectral_gap changed to &mut self; callers updated accordingly.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-20 21:25:55 +02:00
jgrusewski
bda319cfd7 feat(F2): true vector cosine grad_consistency — replaces scalar delta proxy (A3 follow-up)
Replaces HealthEmaTrackers scalar delta-ratio proxy with real vector cosine
similarity across aggregated Adam m-state buffers (q_attn, sel, denoise,
mamba2, ofi_embed). Scalar fallback retained for fused_ctx=None / readback
failure cases.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-20 21:20:30 +02:00
jgrusewski
3fa9923a39 refactor(F7): remove dead iqn/cql/ens_grad_budget config fields — superseded by adaptive budgets (B4/G5)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-20 21:16:52 +02:00
jgrusewski
edaa078a17 fix(F1): silence unsafe_code warning on snapshot alloc — SAFETY comment + allow attribute
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-20 21:14:29 +02:00
jgrusewski
782633a47e refactor(F6/D7): contrarian Q-gap uses sign-flipped Q for consistent conviction 2026-04-20 21:10:44 +02:00
jgrusewski
d9d35b6fad fix: address all 3 precommit LOW findings
1. D6 ensemble oracle no longer dormant — replaces hardcoded ens_disagreement=0.1
   with range of per-branch Q-gap EMAs (max - min over 4 branches). This gives a
   real signal that collapses to 0 when branches agree on uniform Q and expands
   to non-trivial values when branches preserve diverse action differentiation.
   D6's smoothstep(0.01, 0.1) window now actually moves.

2. HEALTH_DIAG defaults aligned with constructor inits — cql_budget default
   0.10→0.00, c51_budget 0.45→0.55. Only visible when fused_ctx is None
   (pre-first-step); eliminates the cosmetic log jump on the first real epoch.

3. reward_v8 smoke-test module changed from pub mod → mod, resolving the
   pre-existing unreachable_pub warning. No external consumers of the module.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-20 21:07:45 +02:00
jgrusewski
d5fea00108 test(E1): collapse-recovery smoke test for LearningHealth adaptive system
Trains a 10-epoch DQN run on fxcache data and asserts q_gap_ema > 0.05 and
learning_health.value > 0.3 — regression guard against the Q-uniform collapse
attractor the entire adaptive-learning-dynamics spec is designed to prevent.

Run via:
  FOXHUNT_TEST_DATA=test_data/futures-baseline \\
    cargo test -p ml --lib -- test_adaptive_learning_no_collapse --ignored --nocapture

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-20 21:00:32 +02:00
jgrusewski
8be3d05ab7 fix(D7/N7): zero out_q_gaps during contrarian mode — prevents full-conviction sizing on argmin trades
Addresses reviewer's IMPORTANT flag: the conviction Q-gap was not sign-flipped when
contrarian_active, so downstream env_step would size contrarian (argmin) positions
with maximum conviction, amplifying losses. During contrarian override, set
out_q_gaps=0 so those trades size minimally — the override is deliberate and
temporary; we don't want to compound it with aggressive position sizing.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-20 20:59:23 +02:00
jgrusewski
f0d01478ec chore(D8): derive Debug on MetaQNetwork/SimpleLcg — fixes missing_debug_implementations lint
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-20 20:58:20 +02:00
jgrusewski
d74b085191 feat(D8/N8): meta-Q network — predicts collapse K epochs ahead as leading indicator
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-20 20:56:54 +02:00
jgrusewski
598c1d57f7 feat(D7/N7 Part B): wire contrarian Q-negation into Boltzmann action selection
Add `contrarian_active` int parameter to `experience_action_select` kernel.
When non-zero, a `q_sign = -1.0f` multiplier is applied to Q values across
all 4 branches (direction/magnitude/order/urgency) before Boltzmann softmax,
converting argmax-favoring sampling to argmin-favoring without touching
temperature, epsilon, conviction filter, masking, or sampling logic.
When zero, q_sign = +1.0f — behavior is bit-identical to before.

Wire: GpuExperienceCollector gains `contrarian_active_cache: u8` field plus
`set_contrarian_active()` / `contrarian_active()` accessors. The flag is
appended as the last arg at the action_select launch site. training_loop.rs
propagates `self.contrarian_active` to the collector immediately after the
D7 Part A state machine updates.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-20 20:51:11 +02:00
jgrusewski
ac58b9d402 feat(D7/N7): contrarian override — argmin flag when WinRate<40% for 5+ epochs during collapse
Part A fully implemented: state machine tracks low_winrate_count across epochs,
activates contrarian_active for 2 epochs when WinRate<40% × 5 consecutive AND
health<0.3, deactivates on WinRate>=45% recovery. last_epoch_win_rate carries
financials.win_rate (f32 cast) from log_epoch_metrics_and_financials into next
process_epoch_boundary. HEALTH_DIAG already consumes last_contrarian_active as
contrarian=on/off. Part B (kernel argmin flip) deferred: direction selection uses
Boltzmann softmax, not simple argmax, making inversion non-trivial.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-20 20:45:56 +02:00
jgrusewski
af9373ced9 feat(D6/N6): ensemble as collapse oracle — pairwise Q-gap triggers plasticity
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-20 20:41:25 +02:00
jgrusewski
087c21050a feat(D5/N5): information bottleneck — scalar penalty visible in HEALTH_DIAG
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-20 20:39:12 +02:00
jgrusewski
02f2f75f5e feat(D4/N4): counterfactual curriculum — cf_ratio scales with (1-health)
Replace hardcoded 0.5f flip probability in experience_env_step kernel with
dynamic cf_ratio = clamp(0.5 + 0.3*(1-health), 0.0, 1.0). Healthy (h=1)
keeps cf_ratio=0.5; collapsed (h=0) raises to 0.8 for more counterfactual
exposure to break collapse. Propagated via set_learning_health() each epoch.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-20 20:36:43 +02:00
jgrusewski
7e345118f4 feat(D3/N3): plasticity injection — health-triggered shrink_perturb replaces periodic
Track consecutive epochs with health < 0.3 in `unhealthy_epoch_count`; after 3
consecutive unhealthy epochs, trigger shrink_and_perturb(alpha, sigma) and reset
the counter. Remove the periodic interval-based trigger (every N epochs). The
Phase 3 boundary trigger is intentionally kept as an independent mechanism.
last_plasticity_ready: Some(true) = accumulating, Some(false) = just triggered.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-20 20:31:53 +02:00
jgrusewski
ebc1d9e9e2 feat(D2/N2): Q-gap barrier constraint — scalar loss visible in HEALTH_DIAG
Compute barrier_loss = 0.5 × max(0, 0.05×health − q_gap)² on the host
from cached q_gap_ema. Written to last_barrier_loss (already declared by
A4) and surfaced in HEALTH_DIAG `barrier=...`. Scalar-only / no gradient.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-20 20:29:01 +02:00
jgrusewski
4911563b2a feat(D1/N1): temporal self-distillation — snapshot at high health, pull toward best during collapse
- Add q_snapshot.rs: SnapshotRing ring buffer (MAX_SNAPSHOTS=5), health/q_gap admission gate
- Add GpuDqnTrainer::maybe_snapshot_params() — DtoD copy of params_buf into snapshot slot
- Add GpuDqnTrainer::apply_distillation_gradient() — two ungraphed saxpy_f32_aux calls:
  grad += alpha * (params - best_snapshot), alpha = 0.1 * (1 - health), skipped when health >= 0.99
- Wire FusedTrainingCtx::maybe_snapshot_qnet() / apply_distillation() / last_distill_active()
- Call from process_epoch_boundary: snapshot when health >= 0.7, distill when health < 0.4
- No new CUDA kernel — reuses existing dqn_saxpy_f32_kernel (saxpy_f32_aux handle)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-20 20:26:36 +02:00
jgrusewski
8669b74d1f feat(C4/P4): temporal-coupled gamma — regime × health scales discount factor
Apply gamma_base + 0.005×(regime_stability−0.5) − 0.05×(1−health), clamped
to [0.9, 0.995], so stable regimes get longer-horizon credit assignment while
collapse (health=0) forces short-horizon local credit to break the collapse.

- GpuDqnTrainer: add last_gamma_eff field + apply_adaptive_gamma() method
- FusedTrainingCtx: add apply_adaptive_gamma() wrapper + last_gamma_eff() accessor
- training_loop: replace both set_adaptive_gamma call sites with apply_adaptive_gamma
- process_epoch_boundary: propagate last_gamma_eff into DQNTrainer for HEALTH_DIAG

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-20 20:18:06 +02:00
jgrusewski
8d088cbc38 feat(B4/G5): adaptive gradient budget — IQN/CQL/C51 scale with health & regime
IQN and CQL SAXPY contributions into grad_buf are now scaled by adaptive
budget factors derived from learning_health (ISV[12]) and regime_stability
(ISV[11]):
  iqn_budget = 0.10 + 0.30 × health  (0.10..0.40)
  cql_budget = 0.10 × (1−regime) × health  (0 at collapse, volatile+healthy → 0.10)
  ens_budget = 0.05  (constant)
  c51_budget = 1 − iqn − cql − ens  (C51 absorbs headroom at collapse → 0.85)

At collapse (health=0): IQN backed off to 10%, CQL off, C51 takes 85% —
stable directional learning when distributional components are unreliable.

Changes:
- GpuDqnTrainer: add 4 last_*_budget_eff fields (initialized to health=1 defaults)
- GpuDqnTrainer: add read_isv_health_and_regime() public accessor
- apply_iqn_trunk_gradient(): add iqn_budget param; scale = iqn_lambda × readiness × iqn_budget
- apply_cql_saxpy(): add cql_budget param; SAXPY alpha = cql_budget (was 1.0)
- FusedTrainingCtx: add compute_adaptive_budgets() — reads ISV, computes all 4, caches to trainer
- FusedTrainingCtx: add last_{iqn,cql,c51,ens}_budget_eff() accessors
- submit_aux_ops(): call compute_adaptive_budgets() once per step, thread to IQN/CQL sites
- training_loop.rs: B4/G5 propagation block for HEALTH_DIAG logging

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-20 20:13:04 +02:00
jgrusewski
aaf6e70f7c feat(C1/P1): health-weighted PER priorities — boost diverse-action experiences during collapse
When learning_health < 0.8, the PER priority update switches from the
standard per_update_pa kernel to pow_alpha_diverse_f32, which multiplies
each priority by (1 + 2*(1-health)*|action - mean_action|). This rescues
the replay buffer's diversity signal during Q-collapse by surfacing
experiences whose action deviates from the batch mean.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-20 20:04:38 +02:00
jgrusewski
e07c82976e feat(B3/G4): health-scaled Expected SARSA temperature — breaks collapse attractor
When health=1 (healthy): tau unchanged → sharp softmax → near-argmax target (deterministic).
When health=0 (collapsed): tau scales 6× → wide softmax → stochastic sampling breaks Q-collapse attractor.

- c51_loss_kernel.cu: add get_learning_health() helper, isv_signals as last param, tau_base → tau * factor
- gpu_dqn_trainer.rs: add last_sarsa_tau_factor field + init, launch_c51_loss → &mut self, host-side mirror, append isv_signals_dev_ptr arg
- fused_training.rs: add last_sarsa_tau_factor() accessor
- training_loop.rs: propagate SARSA tau factor from fused ctx for HEALTH_DIAG logging

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-20 20:02:43 +02:00
jgrusewski
edc59ed6bd feat(B2/G3): health-coupled tau — floor at 0.01×(1-health) for collapse recovery
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-20 19:57:11 +02:00
jgrusewski
46666afef9 feat(B1/G2): uncertainty-gated CQL — cql_alpha = base × (1-regime) × health
Replace static cql_alpha with cql_alpha_eff computed from ISV signals:
health (ISV[12]) × (1 − regime_stability (ISV[11])) × base.
Collapse→0 (CQL off); volatile+healthy→full; stable+healthy→0.
Adds last_cql_alpha_eff field to GpuDqnTrainer (f32, init 0.0),
accessor on FusedTrainingCtx, and propagation into DQNTrainer for
HEALTH_DIAG logging.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-20 19:52:01 +02:00
jgrusewski
1734ae7e34 feat(A4): extend HEALTH_DIAG with all effective values and novel placeholders
Adds 15 new last_* fields (all None) to DQNTrainer for B/C/D tasks to populate,
and replaces the A3 HEALTH_DIAG log line with the full format covering components,
effective hyperparams, and novel mechanism states.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-20 19:46:09 +02:00
jgrusewski
3e9a21d52b fix(A3): clarify docs and guard q_var against infinite sentinels 2026-04-20 19:43:24 +02:00
jgrusewski
6622d90906 feat(A3): compute LearningHealth per epoch, broadcast via ISV[12], HEALTH_DIAG logging
- Add HealthEmaTrackers struct (metrics.rs): EMA for q_gap/q_var/grad_norm with
  scalar grad_consistency proxy (successive grad_norm delta ratio)
- Add LearningHealth + HealthEmaTrackers + 5 last_* fields to DQNTrainer (mod.rs)
- Initialize new fields in constructor.rs
- Add write_isv_signal_at, read_atom_utilization, compute_q_spectral_gap to
  GpuDqnTrainer (gpu_dqn_trainer.rs) with coarse max/min spectral-gap proxy
- Forward same three methods on FusedTrainingCtx (fused_training.rs)
- Compute LearningHealth in process_epoch_boundary (training_loop.rs): reads
  per_branch_q_gap_ema, epoch min/max, atom_util, spectral_gap; emits HEALTH_DIAG log;
  broadcasts health_value to ISV[12] via write_isv_signal_at

Known proxy substitutions (documented in training_loop.rs comment):
  - grad_consistency: scalar proxy instead of full Adam vector cosine (per-component buffers)
  - spectral_gap: coarse max/min ratio on q_readback_pinned instead of SVD
  - ens_disagreement: 0.1 placeholder until D6/N6

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-20 19:36:58 +02:00
jgrusewski
31f4f6314c docs(A2): add rationale for smoothstep thresholds and clamp bounds 2026-04-20 19:29:10 +02:00
jgrusewski
586f5c1058 feat(A2): LearningHealth module with 7-component composition + EMA 2026-04-20 19:25:15 +02:00
jgrusewski
1ad09f63c9 chore(A1): refresh ISV_DIM annotation comments 12→13
Follow-up to 9f3cbc77d — field-doc comments referenced the old value.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-20 19:20:40 +02:00
jgrusewski
9f3cbc77d7 feat(A1): extend ISV_DIM 12→13, add LEARNING_HEALTH_INDEX constant 2026-04-20 19:18:22 +02:00
jgrusewski
6e0cb3d62c fix: CRITICAL — dense micro-reward + OFI embed read stale state positions
Task 3 refactored experience_state_gather (the WRITER) to use assemble_state()
with canonical layout (OFI at [42..62)). But env_step and ofi_embed_build_input
(the READERS) still hardcoded OFI at [66..84) — the OLD pre-refactor layout.

This meant 7 locations were reading MTF/portfolio features as if they were OFI:

1. Line 1745: dense micro-reward ofi_cur = state+66 → actually MTF[4]
2. Line 1778: book_aggression = state[82] → actually plan_isv region
3. Line 1983: ps[30..37] OFI delta storage for NEXT bar — storing MTF data
4. Lines 5833/5836/5838/5840: ofi_embed_build_input — feeds 18→10 MLP into
   Mamba2 temporal SSM and attention. Entire temporal pipeline was training
   on MTF features dressed as OFI.

Symptoms explained:
- WinRate=20.9% on validation (anti-correlated): dense micro-reward computes
  quality=sign_pos × garbage_MTF_deltas, systematically rewarding wrong direction
- mean_reward=+0.004 but Sharpe_raw=-0.0004: shaped reward exploits garbage
  signal, real portfolio loses money
- grad_norm=23560 at epoch 2: gradients chasing noise
- Q-value explosion to ±10 in one epoch: learning contradictions

Fix: replaced all hardcoded 66/74/82/83 with SL_OFI_START from state_layout.cuh.
Both reader kernels now use the same canonical layout as assemble_state().

Verified locally: smoke test passes, OFI_DIAG shows correct non-zero values
(raw_mean=-0.36, delta_mean=-0.21, log_dur=-0.23).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-20 17:30:25 +02:00
jgrusewski
904185004c feat: L40S GPU profile + auto-derive cuda-compute-cap from GPU pool
argo-train.sh now auto-selects cuda-compute-cap based on --gpu-pool:
  - ci-training-h100* → sm_90 (Hopper)
  - ci-training-l40s  → sm_89 (Ada Lovelace)

Added config/gpu/l40s.toml:
  - batch_size=4096 (between H100's 8192 and A100's 2048)
  - buffer_size=300K (scaled for 48GB VRAM)
  - gpu_timesteps_per_episode=2000 (bandwidth-limited)
  - gpu_n_episodes=2048 (scaled from H100's 4096)

GPU profile loader maps "L40S" → "l40s" (was "a100" fallback).

Also fixed pre-existing test drift: num_atoms=52 in h100.toml/a100.toml
was 51 in test expectations (padding alignment for C51 kernels).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-20 16:59:06 +02:00
jgrusewski
733b2c32ec test: add state layout match integration test
Verifies the canonical state layout at all 5 sections (market/OFI/MTF/portfolio/padding).
If any section drifts to the wrong offset, this test fails. Since training and
backtest share assemble_state() in state_layout.cuh, a passing test guarantees
both paths produce identical layouts.

Result: ✓ market=41 ofi=9 mtf=9 portfolio=5 padding=all-zero

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-20 15:35:41 +02:00
jgrusewski
882497caa4 fix: OFI_DIAG reads new canonical indices [42..62), remove state_dim from evaluate_baseline
- OFI_DIAG now reads positions [42..62) using OFI_START constant instead of
  hardcoded [66..74). Verified: raw_mean=0.0891, delta_mean=-0.0370,
  book_agg=0.4500, log_dur=-0.2303 (was all zeros before).
- Removed 3 stale state_dim field initializers from evaluate_baseline.rs.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-20 15:33:42 +02:00
jgrusewski
66bc8d12e5 refactor: remove configurable state_dim — use STATE_DIM constant everywhere
Remove pub state_dim field from DQNConfig and GpuReplayBufferConfig; remove the
state_dim field from GpuExperienceCollector. Replace all reads with
ml_core::state_layout::STATE_DIM (and STATE_DIM_PADDED for cuBLAS-padded
strides). Checkpoint loading now validates saved state_dim against the
constant and hard-errors on mismatch. GpuAttentionConfig.state_dim is a
distinct attention-feature dim and is left untouched.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-20 15:29:05 +02:00
jgrusewski
b2fc2bbe40 fix: pre-existing smoke_test_real_data missing curiosity_weight arg
3 call sites to GpuExperienceCollector::new() were missing the 12th
argument (curiosity_weight). Pre-existing test issue surfaced during
Task 4 verification. Fixed by passing 0.0 (disabled in tests).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-20 15:02:51 +02:00
jgrusewski
10b90268d6 feat: delete backtest_gather_kernel, use shared assemble_state() for validation 2026-04-20 15:01:03 +02:00
jgrusewski
65ad9debbb refactor: experience_state_gather uses assemble_state() — fixes OFI/MTF collision
The kernel now writes to local arrays (market[], portfolio[], plan_isv[],
mtf[], ofi[]) then calls assemble_state() from state_layout.cuh to produce
the canonical layout. This eliminates the hardcoded ofi_start=66 that
collided with MTF features and ensures training/validation use identical
state vectors.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-20 14:49:03 +02:00
jgrusewski
62650aa53b feat: add state_layout.cuh — CUDA header with layout constants and assemble_state()
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-20 14:41:16 +02:00
jgrusewski
c60cd98a35 feat: add state_layout constants module — single source of truth for STATE_DIM
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-20 14:39:30 +02:00
jgrusewski
a7ef847922 cleanup: remove OFI diagnostics + dead bf16 scatter_insert kernel
Diagnostics served their purpose — confirmed OFI flows through full
pipeline on H100 (state_gather → PER → trainer). Remove:
- OFI host verify readback (upload_ofi_features)
- STATE_GATHER_DIAG pinned memory readback (timestep loop)
- Dead bf16 scatter_insert kernel (states are f32, was never called)

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-20 14:18:04 +02:00
jgrusewski
7af0a228fb fix: PER insert used 1D kernel for 2D state matrices — model trained on garbage
scatter_insert_f32 is a 1D scalar kernel (5 args: dst, src, cursor, cap,
batch_size). insert_batch called it with 6 args for state matrices, passing
state_dim as batch_size. CUDA silently dropped the 6th arg (actual batch_size).

Result: only 96 floats (1 state row) inserted per experience batch into PER.
The model was training on ~99.99% uninitialized GPU memory. This bug affected
ALL state features, not just OFI — market features and portfolio were also
garbage in PER-sampled training batches.

Fix: added scatter_insert_f32_rows kernel (2D-aware, 6 args: dst, src,
cursor, cap, state_dim, batch_size) matching the existing scatter_insert
(bf16) pattern. States and next_states now use the row-aware kernel.

Verified locally: OFI_DIAG shows non-zero values through the full chain
(state_gather → env_step → PER insert → PER sample → trainer states_buf).

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-20 13:57:23 +02:00
jgrusewski
bd0d90482d diag: pinned memory readback of batch_states after state_gather
Uses PinnedHostBuf + cuMemcpyDtoHAsync for state_gather diagnostic.
Reads batch_states[0..state_dim] to verify OFI at positions [66..84).

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-20 13:31:16 +02:00
jgrusewski
4ddd6e1ffa diag: verify OFI host data before GPU upload + state_gather readback
Host-side check (zero-cost) before clone_htod to confirm data isn't
zero before it reaches the GPU. Fixes crash from previous diagnostic
(memcpy_dtoh size mismatch).

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-20 13:19:26 +02:00
jgrusewski
2f7828aefb diag: GPU readback after OFI upload + state_gather to trace H100 zero OFI
Temporary diagnostic: readback ofi_gpu[0..20] after upload and
batch_states[66..84] after first state_gather. OFI works locally on
RTX 3050 but shows zeros on H100 with identical v4 cache.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-20 13:10:16 +02:00
jgrusewski
602628f698 fix: reject legacy v2/v3 fxcache — force v4 rebuild with correct OFI
The PVC had a v2 cache that passed has_ofi and nonzero checks but contained
stale OFI data computed by an old binary. load_fxcache accepted v2/v3 for
backward compat, keeping the stale cache alive across every deploy.

v4 is the only valid version. Legacy loading code removed (-50 lines).
Argo ensure-fxcache will delete the v2 cache and rebuild from 148GB MBP-10.

Verified locally: v4 cache produces OFI_DIAG raw_mean=-0.2391 (non-zero).

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-20 12:49:07 +02:00
jgrusewski
a240b7a8c4 fix: remove Option<> from ofi_gpu, remove default total_bars=10000
ofi_gpu is unconditional — wrapping in Option allowed silent NULL pointer
fallback to kernel. Now a plain CudaSlice<f32>.

total_bars default was 10,000 (a lie) — must come from actual data length.
Default changed to 0 with hard error if not set before collect.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-20 12:23:07 +02:00
jgrusewski
e41c4909f1 fix: validate OFI data content, not just has_ofi flag — v2 cache had all-zero OFI
The v2 fxcache on PVC passed has_ofi=true validation (mbp10_dir was present
when built) but contained all-zero OFI data. The old binary set the flag
based on directory existence, not actual computed values.

Now counts non-zero OFI rows before accepting a cache — forces rebuild
when MBP-10 data is available but OFI content is all zeros.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-20 12:08:11 +02:00