Commit Graph

4940 Commits

Author SHA1 Message Date
jgrusewski
d1638959d3 fix(sp21): Return v3.1 — drop short-rollout guard for volume bars (atomic)
Single-file fix to `compute_epoch_financials`: remove the
`if n_returns_f >= bars_per_year` short-rollout fallback that
left v2 semantics in place for sub-year rollouts. For Foxhunt's
volume bars, `bars_per_day ≈ 34_496` → `bars_per_year ≈ 8.69M`,
while a training epoch produces `n_returns ≈ 4.10M`. The guard
fired on every production epoch, so the v3 CAGR fix was a no-op.

Diagnosis chain:
- v3 commit (2937da889) merged the n_returns >= bars_per_year guard
- Smoke v4 (commit 62b5a50e8, workflow train-frv8x) epoch 1 showed
  Return=+2.963e2% — bit-identical to v1's pre-fix output
- Hypothesis 1 (cache poisoning): ruled out — ensure-binary log
  shows "Cache MISS: compiling binaries for 62b5a50e8" and ml crate
  was recompiled fresh
- Hypothesis 2 (different commit): ruled out — workflow params confirm
  commit-sha = 62b5a50e8 = current HEAD
- Hypothesis 3 (bars_per_year mismatch): confirmed — v4 log emits
  "Bars per day (from data): 34496" which makes bars_per_year > n_returns
  and triggers the v2 fallback inside the v3 branch

Fix: unconditional CAGR. The log-space clamp [-23, +20] bounds
the display in all edge cases (tests with tiny n_returns
extrapolate aggressively; the clamp caps at exp(20) - 1 ≈ +4.85e8%).

Expected v5 epoch 1 Return: ~+1.770e3% (was +2.963e2% under v4).
The new value is the *actual annualized* projection: 1377% over
a 0.47-year rollout. Overfit cycles cap at +4.85e8% (was +e19%).

Tests:
- cargo test -p ml --lib financials → 7/7
- Sign-only assertions in test cases (all > 0.0) — no regressions

Files changed:
- crates/ml/src/trainers/dqn/financials.rs: 1 conditional removed,
  comment block updated with v3 → v3.1 history
- docs/dqn-wire-up-audit.md: diagnosis + fix entry for 2026-05-11

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-11 09:12:56 +02:00
jgrusewski
62b5a50e8b fix(eval): shape-mismatch on checkpoint load — read arch from safetensors metadata (atomic)
Smoke v1 (train-grfcw) evaluate phase failed with "Failed to load DQN
checkpoint" for fold 0 and fold 1. MinIO log inspection confirmed
checkpoints WERE saved (1431144 bytes each) — the failure was
eval-side shape mismatch.

Root cause:
- Training uses STATE_DIM=128 (ml_core::state_layout), num_actions=108
  (factored b0*b1*b2*b3=4*3*3*3), num_order_types=3,
  num_urgency_levels=3.
- evaluate_baseline CLI defaults: --feature-dim=54, --num-actions=5
  (legacy from pre-branching DQN era).
- Loading 128-state-dim 108-action checkpoint into 54-feature 5-action
  net → tensor shape mismatch → `load_from_safetensors` returned
  parse error → `with_context(...)` wrapped it as the generic "Failed
  to load DQN checkpoint" message, hiding the actual shape error.
- Both GPU and CPU eval paths hit the same root cause.

Fix:
Both eval paths now call `DQNConfig::from_safetensors_file(&ckpt_path)`
to read architecture-critical fields from the checkpoint's embedded
metadata (state_dim, num_actions, hidden_dims, num_order_types,
num_urgency_levels, dueling_hidden_dim, num_atoms, gamma). Eval-time
fields (LR, epsilon, buffer caps) overridden; hyperopt-derived gamma/
v_min/v_max applied if present in hyperopt config.

Older checkpoints without embedded metadata fall back to CLI-args-built
config + warn! log. All production SP21+ checkpoints embed metadata
via the existing DQNConfig::checkpoint_metadata path.

Files changed:
- crates/ml/examples/evaluate_baseline.rs: shape-aware config for both
  dqn_eval_gpu_path (line ~1238) and dqn_eval_cpu_path (line ~1029)
- docs/dqn-wire-up-audit.md: 2026-05-11 audit entry

Verification:
- cargo check -p ml --examples --features cuda: 0 errors
- cargo test -p ml --lib financials: 7/7 (unchanged)
- cargo test -p ml --lib sp21_isv_slots: 4/4 (unchanged)

Behavioral gate: smoke v3 (train-psf86, in-flight on 2937da889) won't
have this fix; smoke v4 dispatch on this commit will validate
evaluate phase succeeds for all folds. Look for
"[DQN GPU] Architecture from checkpoint: ..." log line per fold.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-11 08:28:53 +02:00
jgrusewski
2937da8898 fix(sp21): inflated return + missing best.safetensors at training end (atomic)
Two smoke-driven fixes for issues surfaced by smoke v1 (train-grfcw).

Fix 1: inflated Return (financials.rs)
- Prior `exp(sum(log(1+r_i))) - 1` produced `Return=+8.730e19%` on
  ~4M step_returns/epoch. Math correct but metric meaningless.
- Fix: ANNUALIZED compounded return (CAGR). For n_returns ≥
  bars_per_year, scale log_growth by `bars_per_year / n_returns`
  then exp. Short rollouts (tests/warmup) fall back to total
  compounded. Clamped to log-space `[-23, +20]` for display sanity.
- Smoke v1 epoch 3 with fix: 24.7% annualized (was +e19%).
- Documented v1→v2→v3 history in comment.

Fix 2: missing dqn_fold{N}_best.safetensors (training_loop.rs)
- Smoke v1 evaluate phase failed with "Failed to load DQN
  checkpoint" for fold 0 and fold 1.
- Root cause: async best-worker swallowed errors non-fatally; with
  3 epochs and checkpoint_frequency=10, periodic never fired; if
  async failed, NO checkpoint existed.
- Fix: guaranteed final save at training end. After async drain,
  restore_best_gpu_params + serialize + sync callback(is_best=true).
  Idempotent if async already wrote; authoritative if it failed.
  All errors here non-fatal (training succeeded; eval reports its
  own missing-ckpt at proper boundary).

Files changed:
- crates/ml/src/trainers/dqn/financials.rs: v3 annualized return
- crates/ml/src/trainers/dqn/trainer/training_loop.rs: final save
- docs/dqn-wire-up-audit.md: 2026-05-11 audit entry

Verification (passing):
- cargo check -p ml --tests --features cuda: 0 errors
- cargo test -p ml --lib financials: 7/7
- cargo test -p ml --lib sp21_isv_slots: 4/4
- sp20_aggregate_inputs_test: 12/12
- sp20_phase1_4_wireup_test: 2/2
- sp20_emas_compute_test: 4/4
- sp20_controllers_compute_test: 7/7
- sp21_per_trade_predicted_q_test: 3/3
Total: 39 tests, 0 failures.

Smoke v2 (train-rl5x2) is on commit ad99b79e0 and won't have these
fixes. A v3 dispatch after this commit validates both fixes end-
to-end.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-11 08:18:11 +02:00
jgrusewski
ad99b79e07 feat(sp21): T2.2 Phase 7.5 — E8 curriculum_weights → per-segment PER insert priority boost (atomic)
Closes the "true E8 per-segment PER sampling" deferral from Phase 7.
Phase 7 wired E8's SCALAR concentration to per_update_pa's alpha
boost; Phase 7.5 wires the FULL Vec<f32> of per-segment weights to
per_insert_pa's priority boost.

What lands:
1. 8 new ISV slots [528..536): CURRICULUM_WEIGHT_{0..8}_INDEX.
2. ISV_TOTAL_DIM 528 → 536 (bus extension); fingerprint adds 8 SLOT
   entries; CURRICULUM_N_SEGMENTS=8 const + curriculum_weight_index
   accessor.
3. per_insert_pa kernel reads isv[528 + seg_id] where seg_id = i % 8
   (round-robin segment tag); effective priority × N_SEGMENTS ×
   weight[seg_id]. Uniform weights → no-op (× 1.0); cold-start
   sentinel → no-op; 0.1× floor against pathological zero-weight
   segments preventing sticky exclusion.
4. Producer in training_loop writes 8 ISV slots from
   result.curriculum_weights[0..8].

Segment tagging rationale:
- Naïve approach (tag tuples by val-curriculum-segment id) is
  infeasible — val and training have separate coordinate systems
  (same problem documented in Phase 5+6 audit re E6 winner indices).
- Round-robin via `i % 8` distributes experience-collector's typical
  512+ tuple batch evenly across 8 segments. Over time buffer has
  equal representation per segment; E8 weights redirect sampling
  pressure toward "hard" segments at insert time.
- HEURISTIC mapping (doesn't preserve val-segment semantics) but
  consumes the curriculum_weights vector for real PER priority
  redistribution — Phase 7.5's stated goal.

Files changed:
- crates/ml/src/cuda_pipeline/sp21_isv_slots.rs: 8 new slot consts
  + N_SEGMENTS + curriculum_weight_index accessor + tests
- crates/ml/src/cuda_pipeline/gpu_dqn_trainer.rs: ISV_TOTAL_DIM bump
  + fingerprint
- crates/ml-dqn/src/per_kernels.cu: per_insert_pa per-segment boost
- crates/ml/src/trainers/dqn/trainer/training_loop.rs: producer wireup
- docs/dqn-wire-up-audit.md: 2026-05-11 audit entry

Verification (passing):
- cargo check -p ml --tests --features cuda: 0 errors
- cargo test -p ml --lib sp21_isv_slots: 4/4 (new curriculum_weight_
  index test)
- sp20_aggregate_inputs_test: 12/12
- sp20_phase1_4_wireup_test: 2/2
- sp20_emas_compute_test: 4/4
- sp20_controllers_compute_test: 7/7
- sp21_per_trade_predicted_q_test: 3/3
Total: 35 tests, 0 failures.

SP21 T2.2 cascade — TRULY fully complete (13 atomic commits). Every
enrichment output E1-E8 wires to a real consumer. No remaining
deferrals or hardcoded controller anchors in SP21 T2.2 scope.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-11 07:50:02 +02:00
jgrusewski
39d4577b77 feat(sp21): T2.2 Phase 8.1 — signal-drive agree_thr clamp bounds via val_sharpe_std (atomic)
Closes the last hardcoded anchor in compute_agreement_threshold per
pearl_controller_anchors_isv_driven. Smoke-driven motivation: the
9-cycle smoke run of commit 1d2dd38a1 (Phases 1.5..8) produced
monotonic agree_thr loosening from 1.30 → 10.00, hitting the
hardcoded upper clamp on cycle 9.

Bound formula:
  scale = (1.0 + val_sharpe_std × 2.0).clamp(1.0, 5.0)
  lo = 0.01 / scale
  hi = 10.0 × scale

Bound behaviour:
  val_sharpe_std=0 (cold)    → scale=1.00 → [0.01, 10.0]  (= pre-8.1 baseline)
  val_sharpe_std=0.05 (mild) → scale=1.10 → [0.009, 11.0]
  val_sharpe_std=0.30 (noisy)→ scale=1.60 → [0.006, 16.0]
  val_sharpe_std≥2.0 (extreme) → scale=5.00 → [0.002, 50.0]  (Invariant 1 ceiling)

The 2.0× multiplier and [1.0, 5.0] scale clamp are themselves
hardcoded but explicitly Invariant 1 carve-outs (numerical-
stability bounds on the bound formula, NOT controller anchors).
The recursion terminates at structural floors/ceilings per
pearl_wiener_alpha_floor_for_nonstationary's canonical pattern —
making meta-meta-meta-bounds signal-driven gains nothing.

Cold-start preservation: prior special-case short-circuit
returned current.clamp(0.01, 10.0). New formula reduces to that
exact behaviour when std=0 (scale=1, lo=0.01, hi=10.0). The
short-circuit is retained for explicit "no update on cold start"
semantics. No behavioural regression at cold-start.

Files changed:
- crates/ml/src/trainers/dqn/trainer/enrichment.rs: compute_agreement_
  threshold clamp refactor (single-function change, no ABI churn)
- docs/dqn-wire-up-audit.md: 2026-05-11 audit entry with full
  smoke cycle table

Verification (passing):
- cargo check -p ml --tests --features cuda: 0 errors
- cargo test -p ml --lib sp21_isv_slots: 3/3
- sp20_aggregate_inputs_test: 12/12
- sp20_phase1_4_wireup_test: 2/2
- sp20_emas_compute_test: 4/4
- sp20_controllers_compute_test: 7/7
- sp21_per_trade_predicted_q_test: 3/3
Total: 34 tests, 0 failures. Behavioral gate: a repeat smoke
should show agree_thr breaking past 10.0 as val_sharpe_std
drives bounds outward.

SP21 T2.2 cascade — FULLY COMPLETE after this commit. 12 atomic
commits, no hardcoded anchors remaining in enrichment controllers
(only Invariant 1 stability carve-outs on bound-on-bound formulas,
which terminate the recursion).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-11 07:44:25 +02:00
jgrusewski
1077f1e165 feat(sp21): T2.2 Phase 6.5b — hindsight synthetic injection producer wireup (atomic)
Closes Phase 6.5. The consumer-side infrastructure landed in 6.5a (val
state retention + accessor); this commit adds the producer.

What lands:
1. GpuReplayBuffer::insert_synthetic_via_pinned (ml-dqn) — raw-u64
   dev_ptr mirror of insert_batch's scatter pipeline + per_insert_pa.
   Takes 7 device pointers + count; bridges from mapped-pinned
   scratch (in ml crate) to PER's scatter kernels.
2. HindsightScratch struct + MAX_SYNTHETIC_HINDSIGHT=32 constant in
   enrichment.rs. Holds 7 mapped-pinned buffers (states + next_states
   + actions + rewards + dones + aux_sign + aux_conf), lazy-allocated.
3. hindsight_scratch: Option<HindsightScratch> trainer field.
4. async fn inject_hindsight_experiences trainer helper: looks up val
   state at (window_index, bar_index) via 6.5a's read_retained_state,
   encodes factored action (dir × 27 + mag × 9 + 0 × 3 + 1 for
   Market/Normal defaults), maps optimal_direction to aux_sign
   (-1/0/+1), writes reward = counterfactual_pnl, done = 1.0
   (terminal), aux_conf = 0.0, calls insert_synthetic_via_pinned.
5. Hook in post-enrichment block — non-fatal warn on infrastructure
   errors per feedback_kill_runs_on_anomaly_quickly.

Design choices:
- Terminal done=1: Bellman target reduces to target_q = reward.
  Avoids synthesizing a valid next_state (optimal counterfactual
  action would produce a DIFFERENT next state, unsimulable from val
  data alone). Pure value-target injection at (state, action).
- Cap at 32 synthetic per epoch: prevents domination of PER buffer.
  Scratch alloc ≈ 30 KB pinned host RAM total.
- Reward in pnl units: counterfactual_pnl is fraction-of-equity;
  training reward kernel handles natively (PopArt normalizes).
  Future scaling via ISV[PNL_REWARD_MAGNITUDE_EMA_INDEX=359] is a
  one-line follow-up if smoke surfaces gradient outliers.
- Mapped-pinned bridge: ml-dqn doesn't have MappedF32Buffer (in ml
  crate). Raw-u64 API takes dev_ptrs directly — clean cross-crate
  boundary, no type duplication.

Files changed:
- crates/ml-dqn/src/gpu_replay_buffer.rs: insert_synthetic_via_pinned API
- crates/ml/src/trainers/dqn/trainer/enrichment.rs: HindsightScratch struct + cap const
- crates/ml/src/trainers/dqn/trainer/mod.rs: hindsight_scratch field
- crates/ml/src/trainers/dqn/trainer/constructor.rs: hindsight_scratch: None init
- crates/ml/src/trainers/dqn/trainer/training_loop.rs: inject_hindsight_experiences
  helper + post-enrichment hook
- docs/dqn-wire-up-audit.md: 2026-05-11 audit entry

Verification (passing):
- cargo check -p ml --tests --features cuda: 0 errors
- cargo test -p ml --lib sp21_isv_slots: 3/3
- sp20_aggregate_inputs_test: 12/12
- sp20_phase1_4_wireup_test: 2/2
- sp20_emas_compute_test: 4/4
- sp20_controllers_compute_test: 7/7
- sp21_per_trade_predicted_q_test: 3/3
Total: 34 tests, 0 failures.

SP21 T2.2 cascade FULLY COMPLETE (Phases 1.5, 2, 3, 4, 4.5, 5+6, 6.5a,
6.5b, 7, 8). All enrichment outputs (E1-E8) wire to real consumers.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-11 00:20:08 +02:00
jgrusewski
f29dcc47c9 feat(sp21): T2.2 Phase 6.5a — val state retention infrastructure (mapped-pinned, atomic)
Lays the consumer-side infrastructure for true E7 hindsight synthetic
injection. Phase 6.5b will follow with the producer wireup.

What lands:
1. EvalTrade.window_index + HindsightExperience.window_index fields
   (host-side only; set in read_per_trade_tape from the w loop var
   and propagated by compute_hindsight_labels).
2. GpuBacktestEvaluator.retained_states_buf — mapped-pinned
   MappedF32Buffer sized [max_len × n_windows × state_dim_padded].
   Populated by a DtoD copy after every launch_gather_chunk inside
   submit_dqn_step_loop_cublas; layout matches chunked_states_buf
   so the copy is a single contiguous block per chunk (no
   transpose).
3. pub fn read_retained_state(window_idx, bar_idx) — zero-copy
   host read via std::ptr::read_volatile on host_ptr (no
   memcpy_dtoh per feedback_no_htod_htoh_only_mapped_pinned).

Mapped-pinned decision (jgrusewski review):
- Initial draft used CudaSlice<f32> + memcpy_dtoh for host read,
  caught at review: violates feedback_no_htod_htoh_only_mapped_pinned.
- Refactored to MappedF32Buffer (cuMemHostAlloc DEVICEMAP). The
  DtoD copy remains (rule forbids HtoD/DtoH, not DtoD; kernel
  writes via dev_ptr aliasing pinned host memory). Caller must
  sync eval stream before read_retained_state — production path's
  consume_metrics_after_event already does this.

Memory cost at production cfg (max_len=200_000, n_windows=5,
state_dim_padded≈128): ~512 MB pinned host RAM. Substantial but
feasible on L40S host (192 GB+).

Files changed:
- crates/ml/src/cuda_pipeline/gpu_backtest_evaluator.rs: +retained
  buffer + accessor; EvalTrade.window_index field
- crates/ml/src/trainers/dqn/trainer/enrichment.rs: HindsightExperience
  .window_index field
- docs/dqn-wire-up-audit.md: 2026-05-11 audit entry

Verification (passing):
- cargo check -p ml --tests --features cuda: 0 errors
- cargo test -p ml --lib sp21_isv_slots: 3/3
- sp20_aggregate_inputs_test: 12/12
- sp20_phase1_4_wireup_test: 2/2
- sp20_emas_compute_test: 4/4
- sp20_controllers_compute_test: 7/7
- sp21_per_trade_predicted_q_test: 3/3
Total: 34 tests, 0 failures. Infrastructure works without exercising
it (accessor returns None until eval populates the retained buffer
— graceful degradation for test scaffolds bypassing the full eval).

After this commit (Phase 6.5b):
- Mapped-pinned synthetic-tuple scratch on GpuReplayBuffer
- New insert_synthetic_via_pinned API (raw dev_ptrs, no HtoD)
- training_loop hindsight injection wireup (~300 LOC)

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-11 00:09:10 +02:00
jgrusewski
1d2dd38a10 feat(sp21): T2.2 Phase 8 — signal-drive E2+E5 controller gains via val_sharpe_std (atomic)
Eliminates remaining hardcoded controller GAINS in enrichment.rs per
pearl_controller_anchors_isv_driven. Both E2 (compute_adaptive_epsilon)
and E5 (compute_agreement_threshold) now derive gain magnitudes
from val_sharpe_std = √ISV[VAL_SHARPE_VAR_EMA_INDEX=351] — same
signal source as the early-stopping pipeline. Phase 2 already
signal-drove the anchors; this commit closes the GAIN half.

NO new ISV slots. NO kernel changes. NO ISV_TOTAL_DIM bump. Pure
value-driven refactor of two enrichment functions.

E2 transformation:
- Bracket anchors 2.0/0.5/-0.5 → 2.0×std / 0.5×std / -0.5×std
- Multiplicative gains 0.8/0.95/1.2 → (1 ± gain_mag) and (1 - 0.5×gain_mag)
- gain_mag = val_sharpe_std.clamp(0.05, 0.30) (Invariant 1 carve-out)
- Cold-start (var_ema==0) → pass-through

E5 transformation:
- Tighten step 0.9 → (1 - gain_mag)
- Loosen step 1.1 → (1 + gain_mag)
- Same gain_mag formula as E2 (consistency)

Invariant 1 carve-outs explicitly retained (project-wide priors):
- [0.05, 0.30] gain_mag stability clamp (mirrors Wiener-α floor)
- [0.85, 0.98] E3 gamma support range (trading-frequency prior)
- [0.5, 2.0] E4 per-branch LR multiplier (collapse/divergence guard)

Files changed:
- crates/ml/src/trainers/dqn/trainer/enrichment.rs: E2 takes new
  val_sharpe_var_ema arg; both E2 and E5 derive gains from
  val_sharpe_std; run_enrichments call-site arg added
- docs/dqn-wire-up-audit.md: 2026-05-11 audit entry

Verification (passing):
- cargo check -p ml --tests --features cuda: 0 errors
- cargo test -p ml --lib sp21_isv_slots: 3/3
- sp20_aggregate_inputs_test: 12/12
- sp20_phase1_4_wireup_test: 2/2
- sp20_emas_compute_test: 4/4
- sp20_controllers_compute_test: 7/7
- sp21_per_trade_predicted_q_test: 3/3
Total: 34 tests, 0 failures.

SP21 T2.2 cascade COMPLETE — all 8 atomic phases landed (1.5, 2, 3,
4, 4.5, 5+6, 7, 8). Remaining future work out of T2.2 scope:
- Phase 6.5 (deferred): true E7 hindsight synthetic injection
- Phase 7.5 (deferred): true E8 per-segment PER sampling
Next operational step: dispatch L40S smoke training run.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-10 23:58:59 +02:00
jgrusewski
34d19955ff feat(sp21): T2.2 Phase 7 — E8 curriculum_concentration → PER alpha boost (atomic)
Wires E8's curriculum_weights distribution into the per_update_pa
alpha boost composition via a new scalar signal:
`compute_curriculum_concentration(weights) = 1 - entropy/log(n)`.
Same pattern as Phases 5+6 (E6 winner_concentration + E7
hindsight_magnitude); all three feed the same boost_delta sum.

Plan revision (mirrors Phase 5+6 pattern):
- Original plan: "wire E8 (curriculum weights) → segment sampling
  weights." Literal interpretation requires new curriculum-segment
  abstraction in PER (segments don't exist — flat ring buffer
  today).
- Resolution: signal-driven from the SHAPE of the weights
  distribution (entropy concentration), not the CONTENT (per-segment
  weighted sampling). Feeds existing per_update_pa kernel; no new
  segment-sampling kernel.
- True per-segment PER sampling deferred to Phase 7.5 (mirrors
  Phase 6.5 deferral for E7's literal injection consumer).

Signal semantics (compute_curriculum_concentration):
- Input: Vec<f32> of E8's per-segment weights (sum=1).
- Output: 1 - entropy/log(n) ∈ [0, 1].
  - 0.0 = uniform (segments equally hard)
  - 1.0 = one segment dominates (concentrated difficulty)
- Single-segment trivially → 1.0
- Cold start (empty input) → 0.0 sentinel
- Zero-weight segments skipped (p log p → 0)

ISV slot allocation:
- CURRICULUM_CONCENTRATION_INDEX = 527
- ISV_TOTAL_DIM 527 → 528 (bus extension)
- Layout fingerprint adds SLOT_527 entry

Kernel change (4 lines, no ABI churn):
- per_update_pa reads isv_signals[527] (already-existing arg from
  Phase 5+6)
- curriculum_term = clamp(curric_conc, 0, 1) × 0.1 ∈ [0, 0.1]
- boost_delta upper bound 0.4 → 0.5 to accommodate new term
- Cold-start short-circuit predicate extended to all 3 signals

Producer wireup:
- enrichment.rs: new helper compute_curriculum_concentration;
  EnrichmentResult.curriculum_concentration field;
  run_enrichments populates it; log line extended
- training_loop.rs: post-enrichment block writes ISV[527]

Files changed:
- crates/ml/src/cuda_pipeline/sp21_isv_slots.rs: +1 slot const
- crates/ml/src/cuda_pipeline/gpu_dqn_trainer.rs: ISV_TOTAL_DIM
  bump + fingerprint
- crates/ml-dqn/src/per_kernels.cu: 4-line boost composition
  extension
- crates/ml/src/trainers/dqn/trainer/enrichment.rs: helper + field
- crates/ml/src/trainers/dqn/trainer/training_loop.rs: write ISV
- docs/dqn-wire-up-audit.md: 2026-05-11 audit entry

Verification (passing):
- cargo check -p ml --tests --features cuda: 0 errors
- cargo test -p ml --lib sp21_isv_slots: 3/3
- sp20_aggregate_inputs_test: 12/12
- sp20_phase1_4_wireup_test: 2/2
- sp20_emas_compute_test: 4/4
- sp20_controllers_compute_test: 7/7
- sp21_per_trade_predicted_q_test: 3/3
Total: 34 tests, 0 failures.

After this commit (Phase 8 + 6.5 + 7.5):
- Phase 8: signal-drive remaining controller GAINS in enrichment
- Phase 6.5: true E7 hindsight synthetic injection (deferred)
- Phase 7.5: true E8 per-segment PER sampling (deferred)

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-10 23:55:46 +02:00
jgrusewski
a6087a0b23 feat(sp21): T2.2 Phase 5+6 combined — E6/E7 signal-driven PER alpha scaling (atomic)
Wires E6 (compute_winner_indices) + E7 (compute_hindsight_labels)
outputs into the PER priority update kernel via signal-driven ISV
slots. Phases 5 and 6 combined into one atomic commit per
feedback_no_deferrals_for_complementary_fixes — both attack the
same consumer kernel (per_update_pa) with non-overlapping refactor
scopes (E6 contributes one ISV-mediated scalar, E7 contributes
another, kernel composes both into alpha_boost).

Plan revision (briefed and accepted 2026-05-11):
- E6's literal Vec<usize> of val bar indices CANNOT directly bump
  training-PER priorities — val and training have separate coord
  systems (no bar-index → buffer-slot mapping).
- E7's true synthetic injection requires val state retention
  infrastructure (n_windows × max_len × state_dim scratch buffer)
  — significant scope expansion deferred to Phase 6.5.
- The MEANINGFUL signal in both phases is scalar aggregations of
  the per-trade tape: E6 winner concentration (top-decile-mean
  P&L / all-mean P&L) and E7 hindsight magnitude (mean
  |counterfactual_pnl|). Both compose multiplicatively into
  per_update_pa's alpha_eff.

ISV slot allocation:
- WINNER_CONCENTRATION_INDEX = 525
- HINDSIGHT_MAGNITUDE_INDEX = 526
- ISV_TOTAL_DIM 525 → 527 (bus extension)
- Layout fingerprint adds 2 SLOT entries
- Compile-time test asserts SP21_SLOT_END (527) ≤ ISV_TOTAL_DIM

Signal semantics:
- winner_concentration: top_decile_mean_pnl / max(all_mean_pnl, EPS).
  > 1.0 = top decile dominates; ≈ 1.0 = uniform; ≤ 0 → 0.0 sentinel
  (losing strategy, signal ill-defined).
- hindsight_magnitude: mean(|counterfactual_pnl|). Higher = larger
  missed opportunities.
- Cold start: both at 0.0 sentinel; kernel short-circuits to
  alpha_eff = alpha (no-op).

ABI surgery (minimal):
- per_update_pa: ONE new arg `const float* __restrict__ isv_signals`
  at end. NULL-tolerant. Reads slots 525/526 device-side, composes
  boost_delta (bounded [0, 0.4]), applies alpha_eff = alpha + delta.
- update_priorities_gpu launcher: passes existing
  self.isv_signals_dev_ptr (already on struct, settable via
  set_isv_signals_ptr — same infra per_insert_pa uses for
  recovery_oversample). Zero new public API.

Producer wireup:
- enrichment.rs: 2 new helpers (compute_winner_concentration,
  compute_hindsight_magnitude); EnrichmentResult gains 2 fields
  (winner_concentration, hindsight_magnitude); run_enrichments
  populates both; log line extended.
- training_loop.rs: post-enrichment block writes both ISV slots
  via fused.trainer().write_isv_signal_at.

Files changed:
- crates/ml/src/cuda_pipeline/sp21_isv_slots.rs: +2 slot consts
- crates/ml/src/cuda_pipeline/gpu_dqn_trainer.rs: ISV_TOTAL_DIM bump
  + fingerprint
- crates/ml-dqn/src/per_kernels.cu: per_update_pa ABI + boost logic
- crates/ml-dqn/src/gpu_replay_buffer.rs: launcher passes ISV ptr
- crates/ml/src/trainers/dqn/trainer/enrichment.rs: 2 helpers + 2
  fields + populate
- crates/ml/src/trainers/dqn/trainer/training_loop.rs: write 2 slots
- docs/dqn-wire-up-audit.md: 2026-05-11 audit entry

Verification (passing):
- cargo check -p ml --tests --features cuda: 0 errors
- cargo test -p ml --lib sp21_isv_slots: 3/3
- sp20_aggregate_inputs_test: 12/12
- sp20_phase1_4_wireup_test: 2/2
- sp20_emas_compute_test: 4/4
- sp20_controllers_compute_test: 7/7
- sp21_per_trade_predicted_q_test: 3/3
Total: 34 tests, 0 failures. Behavioral gate (alpha_eff lift on
post-warmup val passes) is the upcoming smoke training run.

Deferred follow-up (Phase 6.5): true E7 hindsight synthetic
experience injection via val state retention + per_insert_pa
extension. Significant infrastructure (n_windows × max_len ×
state_dim scratch + insert API extension). Documented in audit.

After this commit (T2.2 Phases 7-8 + 6.5 deferred):
- Phase 7: E8 curriculum weights → segment sampling
- Phase 8: signal-drive remaining controller GAINS
- Phase 6.5: synthetic hindsight injection (deferred — needs val
  state buffer infrastructure)

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-10 23:07:12 +02:00
jgrusewski
47e67011c9 feat(sp21): T2.2 Phase 4.5 — re-instate Pearl C engagement tracking for branches (atomic)
Closes the Phase 4 deferral. The 4 DqnBranches sub-launches now write
per-block engagement counts to 4 distinct ranges in
clamp_engage_per_block_buf; pearl_c_post_adam_engagement_check
aggregates all 4 ranges into one per-group rate-deficit EMA —
semantically equivalent to the pre-Phase-4 single-launch tracking.

Design: Option (b) sub-block offsetting, mirroring the existing
Curiosity pattern exactly. SP4_ENGAGE_EXTRA_BRANCHES_SUBLAUNCHES = 3
reserves 3 extra slots beyond Curiosity's tail; sub-launch 0 (Dir)
reuses the canonical DqnBranches slot 2; sub-launches 1/2/3 (Mag,
Order, Urgency) get extra slots 11/12/13. This keeps ParamGroup at
8 entries (no taxonomy growth, no 28 new ISV slot allocations).

SP4_ENGAGE_BUF_LEN: 11 → 14 × MAX_BLOCKS_PER_ADAM (45056 → 57344
i32 slots, +12 KiB on a non-hot-path mapped-pinned buffer).

Branch-to-offset mapping:
| Sub-launch | Offset constant                    | Slot | Buffer offset |
|------------|------------------------------------|------|---------------|
| Dir (b0)   | SP4_ENGAGE_OFFSET_BRANCH_DIR       | 2    | 8 192         |
| Mag (b1)   | SP4_ENGAGE_OFFSET_BRANCH_MAG       | 11   | 45 056        |
| Order (b2) | SP4_ENGAGE_OFFSET_BRANCH_ORDER     | 12   | 49 152        |
| Urgency(b3)| SP4_ENGAGE_OFFSET_BRANCH_URGENCY   | 13   | 53 248        |

pearl_c_post_adam_engagement_check generalized: a `multi_range` flag
covers both Curiosity (4 sub-launches W1/b1/W2/b2) and DqnBranches
(4 sub-launches Dir/Mag/Order/Urgency). Read AND zero-out paths use
the same flag — no duplication of branching logic.

Files changed:
- crates/ml/src/cuda_pipeline/sp4_isv_slots.rs: SP4_ENGAGE_EXTRA_
  BRANCHES_SUBLAUNCHES const; SP4_ENGAGE_BUF_LEN formula extended;
  SP4_ENGAGE_OFFSET_BRANCH_{DIR,MAG,ORDER,URGENCY}; layout test
  updated to 14× MAX_BLOCKS_PER_ADAM
- crates/ml/src/cuda_pipeline/gpu_dqn_trainer.rs: launch_adam_update
  branch sub-launches use new offsets; pearl_c_post_adam_engagement_
  check aggregates DqnBranches multi-range
- docs/dqn-wire-up-audit.md: 2026-05-11 audit entry

Verification (passing):
- cargo check -p ml --tests --features cuda: 0 errors
- cargo test -p ml --lib sp4_isv_slots: 2/2 (engage buf layout asserts
  14× + branch offsets correct)
- cargo test -p ml --lib sp21_isv_slots: 3/3
- sp20_aggregate_inputs_test: 12/12
- sp20_phase1_4_wireup_test: 2/2
- sp20_emas_compute_test: 4/4
- sp20_controllers_compute_test: 7/7
- sp21_per_trade_predicted_q_test: 3/3
Total: 33 tests, 0 failures. Behavioral gate: HEALTH_DIAG emits
per-branch engagement-rate-deficit EMAs in upcoming smoke run.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-10 22:52:13 +02:00
jgrusewski
b7c4f84ea0 feat(sp21): T2.2 Phase 4 — wire E4 branch_lr_scale → per-branch Adam (atomic)
Wires EnrichmentResult::branch_lr_scale (E4's [f32; 4] clamped [0.5,
2.0] per-branch LR multiplier) into the DQN Adam optimizer. Splits
the existing single DqnBranches Adam sub-launch into 4 per-branch
sub-launches, each consuming its own LR scale from ISV[521..525).

Plan amendment from on-paper design:
- Plan said "via existing per-group Adam infrastructure" but per-group
  operates at PARAM_GROUP granularity (8 groups, all 4 action branches
  lumped into DqnBranches). E4 needs per-action-BRANCH granularity.
- Resolution: split DqnBranches sub-launch into 4 per-branch sub-launches
  using the already-canonical branch byte ranges (4 param tensors per
  branch: Dir [17..21), Mag [21..25), Order [25..29), Urgency [29..33)).
  Coverage invariant updated to 7-way (was 4-way).
- Pearl C engagement tracking on branches DEFERRED to Phase 4.5: the
  shared DqnBranches engagement counter offset would collide on writes
  if 4 sub-launches use the same offset. Pearl C is a diagnostic system
  (not load-bearing) so its temporary unavailability for branches is
  acceptable; Trunk + Value + Trunk-extras Pearl C still active.
  Phase 4.5 follow-up will re-instate via ParamGroup expansion or
  sub-block offsetting scheme.

ISV slot allocation:
- BRANCH_LR_SCALE_{DIR,MAG,ORDER,URGENCY}_INDEX = 521..525
- ISV_TOTAL_DIM 521 → 525 (bus extension)
- Layout fingerprint adds 4 SLOT entries
- New branch_lr_scale_index(branch_idx) accessor for clean mapping
- Cold-start floor: launcher reads ISV, floors at 1.0 if at sentinel
  0.0 (per pearl_first_observation_bootstrap — first emit replaces
  directly with no intermediate state)

ABI surgery:
- dqn_adam_update_kernel: ONE new arg float lr_scale at end of
  signature; lr = *lr_ptr * lr_scale inside kernel
- 5 callers migrated atomically per feedback_no_partial_refactor:
  - launch_adam_update (main DQN): 7 sub-launches with per-branch
    lr_scale (Trunk + Value + 4 branch + Trunk-extras)
  - 4 post-aux launchers (ofi_embed/aux_trunk/denoise/sel) pass 1.0
  - decision_transformer Adam launch passes 1.0

Producer wireup (training_loop.rs post-enrichment block):
```rust
for (branch_idx, &scale) in result.branch_lr_scale.iter().enumerate() {
    fused.trainer().write_isv_signal_at(
        branch_lr_scale_index(branch_idx),
        scale,
    );
}
```

Files changed:
- crates/ml/src/cuda_pipeline/sp21_isv_slots.rs: +4 slots + accessor + tests
- crates/ml/src/cuda_pipeline/gpu_dqn_trainer.rs: ISV_TOTAL_DIM bump +
  fingerprint + 7-way Adam split with per-branch lr_scale
- crates/ml/src/cuda_pipeline/dqn_utility_kernels.cu: lr_scale arg + apply
- crates/ml/src/cuda_pipeline/decision_transformer.rs: lr_scale=1.0
- crates/ml/src/trainers/dqn/trainer/training_loop.rs: write 4 ISV slots
- docs/dqn-wire-up-audit.md: 2026-05-11 audit entry

Verification (passing):
- cargo check -p ml --tests --features cuda: 0 errors
- cargo test -p ml --lib sp21_isv_slots --features cuda: 3/3 (bus bounds
  + slot uniqueness + branch index mapping)
- sp20_aggregate_inputs_test: 12/12
- sp20_phase1_4_wireup_test: 2/2
- sp20_emas_compute_test: 4/4
- sp20_controllers_compute_test: 7/7
- sp21_per_trade_predicted_q_test: 3/3
Total: 31 tests, 0 failures. Behavioral gate (per-branch LR divergence)
is the upcoming smoke training run.

After this commit (T2.2 Phases 5-7 + 8 + 4.5):
- Phase 5: E6 winner indices → PER priority bumps
- Phase 6: E7 hindsight → replay buffer injection
- Phase 7: E8 curriculum weights → segment sampling
- Phase 8: signal-drive remaining controller GAINS
- Phase 4.5: re-instate Pearl C engagement tracking for branches

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-10 22:44:43 +02:00
jgrusewski
21f911151a feat(sp21): T2.2 Phase 3 — wire E1 q_correction → Bellman target offset (atomic)
Wires `EnrichmentResult::q_correction` (Phase 1.5+2's E1
clamp(-mean(predicted_q − pnl), ±10.0) from the per-trade tape) into
the training loss kernels' Bellman target computation. Both
mse_loss_batched and c51_loss_kernel::block_bellman_project_f apply
q_correction as additive shift on the Bellman target — blended
(1-α)·MSE + α·C51 sees identical Q-target shifts (symmetric
application per feedback_no_partial_refactor).

ISV slot allocation (per pearl_controller_anchors_isv_driven +
feedback_isv_for_adaptive_bounds — adaptive Q-target shift IS
adaptive bound, lives in ISV bus):
- New Q_CORRECTION_INDEX = 520 in cuda_pipeline::sp21_isv_slots
- ISV_TOTAL_DIM 520 → 521 (bus extension)
- Layout fingerprint adds SLOT_520_Q_CORRECTION
- New module registered in cuda_pipeline/mod.rs
- Compile-time test verifies SP21_SLOT_END <= ISV_TOTAL_DIM

Sign convention: q_correction is the additive correction (E1 already
negates the bias). predicted_q > pnl → q_correction < 0 → "lower
Q-targets". predicted_q < pnl → q_correction > 0 → "raise Q-targets".
Cold-start sentinel 0.0 (no trades yet) is no-op until first val
pass writes (per pearl_first_observation_bootstrap).

ABI surgery (minimal):
- mse_loss_batched: ONE new scalar arg (float q_correction) at end
- c51_loss_batched: ZERO new args (block_bellman_project_f already
  takes isv_signals; reads slot 520 from there)
- launch_mse_loss: reads ISV[520] host-side via existing
  read_isv_signal_at, passes scalar
- training_loop.rs (post-enrichment): writes result.q_correction
  to ISV[520] via existing write_isv_signal_at

Files changed:
- crates/ml/src/cuda_pipeline/sp21_isv_slots.rs (NEW): slot const +
  bounds-check tests
- crates/ml/src/cuda_pipeline/mod.rs: pub mod sp21_isv_slots
- crates/ml/src/cuda_pipeline/gpu_dqn_trainer.rs: ISV_TOTAL_DIM bump +
  fingerprint update + launch_mse_loss launcher arg
- crates/ml/src/cuda_pipeline/mse_loss_kernel.cu: q_correction arg +
  application to target_q
- crates/ml/src/cuda_pipeline/c51_loss_kernel.cu: read isv_signals[520]
  in block_bellman_project_f, add to t_z
- crates/ml/src/trainers/dqn/trainer/training_loop.rs: write_isv_signal_at(
  Q_CORRECTION_INDEX, result.q_correction) post-enrichment
- docs/dqn-wire-up-audit.md: 2026-05-11 audit entry

Verification:
- cargo check -p ml --tests --features cuda: 0 errors
- cargo test -p ml --lib sp21_isv_slots --features cuda: 2/2 (bus
  bounds + slot uniqueness)
- sp20_aggregate_inputs_test: 12/12
- sp20_phase1_4_wireup_test: 2/2
- sp20_emas_compute_test: 4/4
- sp20_controllers_compute_test: 7/7
- sp21_per_trade_predicted_q_test: 3/3
Total: 30 tests, 0 failures. Behavioral gate (q_correction → non-
zero ISV[520] → Q-target shift) is the upcoming smoke training run.

After this commit (T2.2 Phases 4-8):
- Phase 4: E4 per-branch LR scaling via per-group Adam
- Phase 5: E6 winner indices → PER priority bumps
- Phase 6: E7 hindsight → replay buffer injection
- Phase 7: E8 curriculum weights → segment sampling
- Phase 8: signal-drive remaining controller GAINS

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-10 22:32:07 +02:00
jgrusewski
fb01906ddc docs(claude): foxhunt agents & skills rollout — phase 1 close-out
Phase 1 of the foxhunt specialized agents/skills rollout is complete: 14
commits, 5 auditor agents, 7 workflow/maintenance skills, 2 helper scripts,
warn-only PostToolUse hook router. All 8 acceptance criteria from the spec
verified. Hook latency 11-13 ms per call (target <200 ms). Memory-write
invariant held: only pearl-distiller and memory-curator are authorized
writers, no agent-driven memory edits during rollout.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-10 22:19:28 +02:00
jgrusewski
7f065e4d02 feat(claude): sp-critical-reviewer agent (foxhunt)
Replicates the user's second/third critical-review iteration cycle. Cites
every claim against a memory file. Enforces no-deferrals-for-complementary-
fixes, ISV-slot enumeration, smoke kill conditions, anti-pattern callouts.
Composes the four foxhunt auditor agents for domain-specific verification.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-10 22:19:28 +02:00
jgrusewski
10eab64d56 feat(claude): stale-worktree-cleaner skill (foxhunt)
Classifies worktrees (NO-COMMITS / GONE / MERGED / UNCOMMITTED / ACTIVE) and
proposes cleanup. User confirms each removal individually; never --force,
never -D. Composes commit-commands:clean_gone where applicable.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-10 22:19:28 +02:00
jgrusewski
478f70d3ed feat(claude): memory-curator skill (foxhunt)
Audits memory/ corpus: stale references (gone files/symbols/flags/commits),
duplicates, superseded pearls, orphans (file vs MEMORY.md index drift).
Proposes archive/merge/update; never auto-deletes. Archive ≠ delete.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-10 22:19:28 +02:00
jgrusewski
529845c016 feat(claude): smoke-pilot skill (foxhunt)
Monitors argo smoke training; kills on first useful anomaly signal per
stop-on-anomaly + kill-quickly discipline. Distinguishes anomaly from
metric-pipeline inflation. Streams via argo logs -f, no run_in_background
Monitor layering.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-10 22:19:28 +02:00
jgrusewski
34cb57aee5 feat(claude): argo-deploy-helper skill (foxhunt)
Wraps scripts/argo-train.sh with deploy discipline pre-flight: push-before-
deploy gate, mandatory --mbp10-data-dir/--trades-data-dir, --per-enabled,
default L40S unless overridden, H100 max-utilization when chosen.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-10 22:19:28 +02:00
jgrusewski
65559bde07 feat(claude): pearl-distiller skill (foxhunt)
One of two authorized writers into memory/ (the other is memory-curator).
Templatizes new-pearl creation: structured body (pattern / detection signal /
fix / canonical commit:file:line / related pearls) + MEMORY.md index entry
in the right section, ≤150 chars.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-10 22:19:28 +02:00
jgrusewski
9acda51329 feat(claude): isv-slot-scaffolder skill (foxhunt)
Templatizes the ISV-slot scaffolding step. Reproduces the c146c4fff commit
shape: spN_isv_slots.rs with named constants, ISV_TOTAL_DIM bump at canonical
site, fold-boundary reset registration. Enforces wire-everything-up discipline.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-10 22:19:28 +02:00
jgrusewski
20b1163ffc feat(claude): sp-spec-writer skill (foxhunt)
Templatizes the SP design spec format used in docs/superpowers/specs/.
Enforces pearl/feedback grounding, ISV-slot enumeration, smoke kill conditions,
sister-fix check (no-deferrals-for-complementary-fixes pearl).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-10 22:19:28 +02:00
jgrusewski
127bcba599 feat(claude): code-hygiene-auditor agent (foxhunt)
Pattern-cluster auditor for general code-quality rules: no stubs, no TODO,
no hiding, no legacy aliases, no feature flags, no partial refactors,
v7-gem methodology, invariant tests not observed-value tests, no deferrals
of complementary fixes. Bundles 12 feedback rules and 3 pearls.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-10 22:19:28 +02:00
jgrusewski
b3480fe0e3 feat(claude): reward-controller-auditor agent (foxhunt)
Pattern-cluster auditor for reward composition and controller mechanics:
bilateral clamps, structural activations, asymmetric bounded clamps, event-
driven density alignment, one-unbounded-multiplicand, Thompson selector,
trail-on-pre_mag, Adam-normalized loss weights, per-bar vs segment PnL,
imbalance-bar EWMA wash-out, aux trunk separation, magnitude-trap.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-10 22:19:28 +02:00
jgrusewski
e18f97a474 feat(claude): isv-discipline-auditor agent (foxhunt)
Pattern-cluster auditor for ISV-driven adaptive bounds, controller anchors,
EMA bootstrap, Wiener-α floors, blend formulas, per-branch budgets, per-group
Adam, Kelly floors, trail-stop. Bundles 2 feedback rules and 16 pearls.
Read-only; cites memory file by name on every finding.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-10 22:19:28 +02:00
jgrusewski
1b1b997be1 feat(claude): gpu-contract-auditor agent (foxhunt)
Pattern-cluster auditor for CUDA/GPU host code. Bundles 6 feedback rules and
7 pearls (no atomicAdd, mapped-pinned, no nvrtc, f64/f32 ABI, no host branches
in graph capture, fused per-group stats, build.rs rerun, canary launch order).
Read-only; cites memory file by name on every finding.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-10 22:19:28 +02:00
jgrusewski
8517776be5 fix(claude): foxhunt audit router — plans path + REPO_ROOT consistency
Two Important review findings:
1. docs/superpowers/plans/*.md missing from sp-critical-reviewer routing
   table — sp-critical-reviewer would never self-suggest on plan files
   (its primary use case).
2. STATE_FILE and REPO_ROOT diverged when CLAUDE_PROJECT_DIR was unset
   (one used "." relative, the other "$(pwd)" absolute). Compute REPO_ROOT
   first and derive STATE_FILE from it; remove the duplicate later
   assignment.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-10 22:19:28 +02:00
jgrusewski
856e89da1f chore(claude): foxhunt audit hook router (warn-only PostToolUse)
Adds .claude/helpers/foxhunt-audit-router.sh that suggests foxhunt-* auditor
agents based on edited file path. Warn-only, <200ms, never blocks. Dedup
state file .claude/.foxhunt-audit-state cleared at SessionStart by sibling
script. Settings.json additive merge — existing hooks preserved.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-10 22:19:28 +02:00
jgrusewski
bc3ebb5fb0 docs(claude): foxhunt agents & skills implementation plan — 14 tasks
14-task plan covering: foundation hook router (Task 1), 5 auditor agents
(Tasks 2-5, 13), 5 workflow skills (Tasks 6-10), 2 maintenance skills
(Tasks 11-12), end-to-end acceptance check (Task 14). Tasks 2-5, 6-8, 9-10,
11-12 fan out in parallel. Task 13 (sp-critical-reviewer) composes the four
domain auditors and is built last.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-10 22:19:28 +02:00
jgrusewski
9e6637c065 docs(claude): foxhunt-aware agents & skills design — phase 1 (12 items)
Pattern-cluster auditors + workflow-skill hybrid that internalize accumulated
project wisdom (15 feedback_*, 30+ pearl_*, 12+ project_* memory files; 30 SP
specs) into foxhunt-namespaced agents and skills under .claude/agents/foxhunt/
and .claude/skills/foxhunt/.

Phase 1: 5 auditor agents (isv-discipline, gpu-contract, reward-controller,
code-hygiene, sp-critical-reviewer) + 5 workflow skills (sp-spec-writer,
isv-slot-scaffolder, pearl-distiller, argo-deploy-helper, smoke-pilot) +
2 maintenance skills (memory-curator, stale-worktree-cleaner). Hook integration
is warn-only PostToolUse via a single thin router; agents read memory but only
pearl-distiller and memory-curator may write into memory/.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-10 22:19:28 +02:00
jgrusewski
ea12c172e6 test(sp21): T2.2 Phase 1.5 — kernel-direct GPU oracle for per-trade predicted_q
Promotes the Phase 2.5 follow-up from the plan to a load-bearing test.
Three #[ignore = "requires GPU"] tests exercise backtest_env_step_batch
directly with controlled inputs:

  1. predicted_q_populated_on_close_with_real_q_values — open Long Full
     at step 0 with q_values_per_window[w=0, a=Long_Full]=2.5, hold for
     two bars, close Flat at step 3. Asserts the per-trade tape's
     predicted_q[0] ≈ 2.5 (entry_q captured at open, persisted in
     entry_q_state across holds, snapshotted before close emit).

  2. predicted_q_stays_zero_when_q_values_is_null — same scenario with
     q_values_per_window=NULL (single-step evaluate() semantics). The
     kernel's open-time write is gated on the NULL guard, so
     entry_q_state[w] stays at buffer-init 0.0 and the tape's
     predicted_q is 0.0. Proves the NULL-tolerance contract is
     kernel-enforced, not just launcher convention.

  3. no_trades_no_predicted_q_emission — all-Flat action sequence
     produces zero trades. Path-coverage check that entry_q's open
     guard fires only on actual opens / reverses.

Test design:
- Kernel-direct (loads ENV_CUBIN via include_bytes!), no eval-pipeline
  scaffolding (no QValueProvider, no cuBLAS forward, no chunked state
  gather). Avoids the heavy mock infrastructure the plan flagged.
- Flat-market 4-bar single-window window: every OHLC=100, zero costs,
  initial_capital=100k. Isolates the entry_q signal from P&L noise
  so the assert-predicted_q is exact under IEEE-754 (kernel writes
  the Q-value verbatim with no math).
- NULL fallback for isv_signals/conviction/exploration_scale matches
  the existing single-step launcher pattern; FEATURE_DIM=4 (<41) ⇒
  compute_regime_trail_scales takes the fixed-width fallback (no
  feature deref needed).

Verification:
- SQLX_OFFLINE=true cargo test -p ml --test sp21_per_trade_predicted_q_test
  --features cuda -- --ignored --nocapture
- 3/3 pass on RTX 3050 Ti (sm_86).
- All 4 prior SP20 GPU oracle suites still pass (25 tests).
- Total: 28 GPU oracle tests, 0 failures.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-10 22:12:07 +02:00
jgrusewski
26deaa5004 feat(sp21): T2.2 Phase 1.5 + Phase 2 — entry_q tracking + enrichment real-tape wireup (atomic)
Closes the multi-step T2.2 work. Phase 1.5 captures entry_q (predicted
Q at trade open) on the per-trade tape; Phase 2 wires
enrichment::run_enrichments to the real GpuBacktestEvaluator
per-trade tape and deletes the fake-trade synthesizer
(extract_eval_trades_from_metrics) per feedback_no_stubs.

Plan amendment from on-paper design:
- Audit caught plan's "portfolio_state[ps+6] is unused" claim was
  wrong — slot 6 is in active use as cum_return. Replaced with a
  dedicated entry_q_state_buf [n_windows] separate from
  portfolio_state. Doesn't touch shmem layout, gather kernel, or
  model state-dim.
- E5 design question resolved as option (2): quartile-spread Sharpe
  with ISV-driven significance anchor read from
  ISV[VAL_SHARPE_VAR_EMA_INDEX=351] (per
  pearl_controller_anchors_isv_driven). Hardcoded 1.5σ/0.5σ
  thresholds eliminated; the noise floor is the val_sharpe variance
  EMA already produced per epoch by the early-stopping pipeline.
- Single-step evaluate() launcher passes NULL q_values_per_window
  (forward_fn closure exposes only action indices, not Q-values);
  same NULL-tolerant pattern as exploration_scale_ptr et al. The
  production val pipeline (chunked path) DOES wire real Q-values.

Files (atomic per feedback_no_partial_refactor):
- backtest_env_kernel.cu: 4 new args at end of both kernels
  (entry_q_state, q_values_per_window, num_actions,
  per_trade_predicted_q_out); pre_entry_q snapshot + close emit +
  open/reverse capture.
- gpu_backtest_evaluator.rs: entry_q_state_buf + per_trade_predicted_q_buf
  allocated; both reset in reset_evaluation_state; both launchers
  migrated; EvalTrade.predicted_q field added; read_per_trade_tape
  populates it.
- trainer/enrichment.rs: EvalTrade is now pub(crate) use re-export
  from gpu_backtest_evaluator (drops ensemble_var); E5 refactored
  to quartile-spread Sharpe with ISV-driven anchor;
  extract_eval_trades_from_metrics deleted; HindsightExperience
  retained (Phase 6 wiring).
- trainer/training_loop.rs: enrichment block replaced with
  evaluator.read_per_trade_tape() + ISV slot 351 read; val_bars /
  real_trade_count / real_total_pnl / real_win_rate plumbing
  dropped.
- trainer/{mod,constructor,metrics}.rs: last_val_metrics field
  removed (last consumer gone — feedback_no_hiding).
- docs/dqn-wire-up-audit.md: 2026-05-11 audit entry.

Verification (passing):
- SQLX_OFFLINE=true cargo check -p ml --tests --features cuda
- sp20_aggregate_inputs_test (12/12)
- sp20_phase1_4_wireup_test (2/2)
- sp20_emas_compute_test (4/4)
- sp20_controllers_compute_test (7/7)

After-this scope (Phase 3-7 + Phase 8 in T2.2 multi-phase):
- E1 q_correction → ISV slot consumer
- E4 per-branch LR scaling via per-group Adam
- E6 winner indices → PER priority bumps
- E7 hindsight → replay buffer injection
- E8 curriculum weights → segment sampling
- Signal-drive remaining controller GAINS (0.9/1.1 in E5; 2.0/0.5/-0.5 in E2)

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-10 22:02:17 +02:00
jgrusewski
c274b99ea9 docs(sp21): T2.2 Phase 1.5 + Phase 2 continuation plan + amendment
Adds the continuation plan for SP21 T2.2 Phase 1.5 (entry_q tracking
in portfolio_state slot 6) + Phase 2 (wire enrichment to real
per-trade tape from gpu_backtest_evaluator). Self-contained plan
that a fresh session can pick up cold:

  - State-at-session-start summary with all 8 prior commits
  - Phase 1.5 design (storage slot, capture site, plumbing,
    EvalTrade extension)
  - Phase 2 design (training_loop wire-up, enrichment refactor,
    E5 alternate-signal options with recommendation)
  - Hard rules carried from prior session (no NULLs, no stubs,
    atomic per feedback_no_partial_refactor)
  - Verification plan (5 GPU oracle test suites)
  - Open design question (E5 alternate signal) with recommendation
  - Multi-phase continuation map (Phases 3-7 + Phase 8)
  - Final cascade verification (smoke run criteria)

Also amends the parent SP21 plan
(2026-05-10-sp21-train-eval-coherence-isv-defrost.md) with the
T2.2 multi-phase scope section that documents the full per-trade-tape
expansion the user committed to (option 3 — full wiring across
multiple sessions instead of the original recommended option (b)
aggregate-stats refactor).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-10 21:35:58 +02:00
jgrusewski
7d538d9304 feat(sp21): T2.2 Phase 1 Step B — per-trade tape buffers + readback (atomic)
Replaces Step A's NULL launcher passes with real device buffers.
Both kernel launch sites in gpu_backtest_evaluator.rs (backtest_env_step
and backtest_env_step_batch) now pass real per-trade tape pointers.
The kernel's per-trade emission block fires unconditionally on close
events — single-threaded per-window writes preserve event ordering and
enable race-free counter increment without atomicAdd.

New constants + types:
  - MAX_TRADES_PER_WINDOW = 200_000 (typical eval window bar count;
    per-window memory: 4 SoA buffers × 4 bytes × 200k = 3.2MB)
  - pub struct EvalTrade with 5 fields: bar_index, pnl, holding_bars,
    direction, magnitude. Does NOT include predicted_q / ensemble_var
    — those need entry-time captures (entry_q, entry_var in portfolio
    state) deferred to Phase 1.5.

New struct fields on GpuBacktestEvaluator:
  - per_trade_pnl_buf:          CudaSlice<f32> [n_windows × MAX_TRADES]
  - per_trade_holding_bars_buf: CudaSlice<u32>
  - per_trade_bar_index_buf:    CudaSlice<u32>
  - per_trade_dir_mag_buf:      CudaSlice<u32> (packed dir/mag)
  - per_trade_count_buf:        CudaSlice<u32> [n_windows]

New methods:
  - reset_per_trade_tape (folded into reset_evaluation_state): zeros
    the count buffer at the start of each eval window. SoA value
    buffers don't need zeroing — read up to count[w] only.
  - pub fn read_per_trade_tape(&self) -> Result<Vec<EvalTrade>, MLError>:
    reads count buffer first (cheap), early-returns empty if no trades,
    else reads 4 SoA buffers (~16MB DtoH at PCIe ≈ 1ms) and flattens
    window-major into chronological Vec<EvalTrade>.

Phase 2 follow-up (next commit) — wire read_per_trade_tape to the
enrichment caller in training_loop.rs:1510-1568, replacing
extract_eval_trades_from_metrics (the fake-trade synthesizer).

Phase 1.5 follow-up (if Phase 2 keeps E1+E5) — add entry_q + entry_var
to portfolio_state at trade open, extend per-trade tape with 6th/7th
SoA buffers.

Affected files:
  - crates/ml/src/cuda_pipeline/gpu_backtest_evaluator.rs
    (constants, struct fields, alloc, construction, 2 launcher sites,
    reset_evaluation_state addition, read_per_trade_tape method)

Verification:
  - cargo check -p ml --tests: passes (warnings only)
  - GPU oracle tests: behavior preserved by construction (existing
    WindowMetrics aggregator unaffected — separate kernel)

Plan reference: docs/plans/2026-05-10-sp21-train-eval-coherence-isv-defrost.md
T2.2 multi-phase scope; Phase 1 Step B closure.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-10 21:26:40 +02:00
jgrusewski
5d01190e19 feat(sp21): T2.2 Phase 1 Step A — per-trade tape kernel ABI (atomic)
Extends backtest_env_step + backtest_env_step_batch kernel signatures
with 6 new NULL-tolerant args for per-trade tape emission:
  - float* per_trade_pnl_out
  - unsigned int* per_trade_holding_bars_out
  - unsigned int* per_trade_bar_index_out
  - unsigned int* per_trade_dir_mag_out (packed hi-16 dir, lo-16 mag)
  - unsigned int* per_trade_count
  - int max_trades_per_window

Each kernel snapshots entry_price + hold_time BEFORE the
unified_env_step_core call, then post-call mirrors
record_kelly_trade_outcome's close-detection predicate
(is_exit || is_reversal with positive entry_price + valid pre-trade
equity guards) to recompute the realized return for per-trade tape
emission. Single thread per window writes its own slot — race-free
counter increment without atomicAdd per feedback_no_atomicadd.

Step A only — Step B (host alloc + readback + enrichment integration)
follows in a separate commit. This commit is the kernel ABI extension
with NULL launchers, bit-identical to pre-Phase-1 behavior.

NULL-tolerant ABI extension is the codebase's canonical phased-rollout
pattern. Same shape as alpha_per_env / is_win_per_env additive contracts
in sp20_aggregate_inputs_kernel. Per feedback_no_partial_refactor's
"consumer migrates with contract change" rule, the consumer (launcher)
MIGRATES here by passing NULL — output behavior preserved bit-identically.

Step B (next commit):
  - Allocate per-trade buffers in GpuBacktestEvaluator constructor
  - Pass real device pointers from launcher
  - Reset per-window counter at fold start
  - Host-side readback into Vec<EvalTrade>
  - Wire to enrichment caller, replace extract_eval_trades_from_metrics

Affected files:
  - crates/ml/src/cuda_pipeline/backtest_env_kernel.cu (both kernels)
  - crates/ml/src/cuda_pipeline/gpu_backtest_evaluator.rs (both launchers)

Verification:
  - cargo check -p ml --tests: passes (warnings only)
  - GPU oracle tests: NULL-tolerance contract preserved by construction
    (per_trade_count == NULL gates the entire emission block)

Plan reference: docs/plans/2026-05-10-sp21-train-eval-coherence-isv-defrost.md
T2.2 multi-phase scope section.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-10 21:19:25 +02:00
jgrusewski
4ab1c132e8 feat(sp21): T2.1+T2.4 — Q-value early-stop + MIN_HOLD zombies deleted
Closes two architectural-debt items from SP21 Tier 2.

T2.1 — check_early_stopping(avg_q_value) deleted entirely:
  - Combined two failed mechanisms: (a) Q-value floor — not a learning
    signal (high Q can mean edge OR value explosion, indistinguishable);
    (b) Sharpe plateau with hardcoded `improvement < 0.01` threshold,
    structurally meaningless against typical val-sharpe deltas O(1-10).
  - Both subsumed by the SP21 T1.1a+T1.1b val-loss patience early-stop
    with signal-driven min_delta from VAL_SHARPE_VAR_EMA.
  - Legacy `old_should_stop` branch + function body deleted.
  - Per feedback_no_legacy_aliases.

T2.4 — MIN_HOLD_TARGET / MIN_HOLD_PENALTY_MAX #defines deleted:
  - Investigation: macros referenced ONLY in comments and the defining
    line itself — no actual code use. The SP12 v3 production callers
    were removed in SP20 Phase 2 Task 2.2.
  - HEALTH_DIAG line at training_loop.rs:5159 updated to drop the dead
    30.0/3.0 literals.
  - Scope boundary: MIN_HOLD_TEMPERATURE_* chain is NOT a zombie —
    actively wired (kernel producer + SP16 controller consumer).
  - Per feedback_no_legacy_aliases.

T2.5 — PER hyperparams disposition (no code change):
  - per_alpha=0.6, per_beta_start=0.6 are paper-canonical (Schaul et al.).
    Per the SP21 plan recommendation, kept fixed for SP21. Filed for a
    separate SP if later identified as a leverage point.

Affected files:
  - crates/ml/src/trainers/dqn/trainer/metrics.rs:435-488
    (check_early_stopping body deleted)
  - crates/ml/src/trainers/dqn/trainer/training_loop.rs (7253 caller +
    7291-7322 old_should_stop branch + 5159-5167 HEALTH_DIAG line)
  - crates/ml/src/cuda_pipeline/state_layout.cuh:317-318
    (#defines deleted)

Verification:
  - cargo check -p ml --tests: passes (warnings only)

Cumulative SP21 Tier 2 status: T2.1 ✓, T2.4 ✓, T2.5 ✓ (deferred-doc).
T2.2+T1.3 (enrichment.rs constants soup, ~400 LOC) remaining.

Plan reference: docs/plans/2026-05-10-sp21-train-eval-coherence-isv-defrost.md

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-10 20:52:38 +02:00
jgrusewski
c90553953c feat(sp21): T1.2+T1.4 — enrichment real metrics + backtracking signal-driven (atomic)
Closes the remaining SP21 Tier 1 hardcoded-constant items.

T1.2 — enrichment fed real metrics, not placeholders:
  - was: extract_eval_trades_from_metrics(_, 60000.0, 0.0, 0.5, ...)
         with hardcoded trade_count=60000, total_pnl=0.0, win_rate=0.5
  - now: reads from self.last_val_metrics: Option<[f32; 14]> populated
         by val backtest pass at metrics.rs:868. Layout [2]=win_rate,
         [4]=total_trades, [7]=total_pnl. Cold-start fallback (None)
         is (0.0, 0.0, 0.0) — preferable to fabricated 60000-trade
         signal that biased E2/gamma/ensemble from epoch 0.
  - Per feedback_no_todo_fixme + feedback_no_stubs.

T1.4 — backtracking thresholds signal-driven:
  - Three hardcoded thresholds in run_backtracking_epoch_end replaced
    with sigma = sqrt(ISV[VAL_SHARPE_VAR_EMA_INDEX=351]) derivatives:

    a) Save trigger (improvement_rate > 0.01) → > 0.5σ.
       The 0.01 fired on every epoch (any tiny change > 0.01);
       0.5σ requires a meaningful move (typical sigma O(1-10)).
    b) Plateau-detection frozen check (abs(delta) < 0.01) → < 0.5σ.
       The 0.01 ~never fired; 0.5σ correctly identifies stagnation.
    c) Route acceptance (>= min_improvement_rate=0.1) → >= 1.5σ.
       Stricter than save-trigger as designed.

  - BacktrackingState::min_improvement_rate field deleted — replaced
    by per-call signal-driven computation. Floor 0.5 covers cold-start
    before var_ema bootstraps from sentinel per
    pearl_blend_formulas_must_have_permanent_floor.
  - Per feedback_isv_for_adaptive_bounds + feedback_adaptive_not_tuned.

Affected files:
  - crates/ml/src/trainers/dqn/trainer/training_loop.rs:1510-1530
    (T1.2 enrichment) + :7458-7530 (T1.4 sigma + 3 threshold sites)
  - crates/ml/src/trainers/dqn/trainer/mod.rs:108,142
    (T1.4 min_improvement_rate field deletion)

Verification:
  - cargo check -p ml --tests: passes (warnings only)
  - cargo test -p ml --lib early_stopping: 8/8 pass

Cumulative SP21 Tier 1 status: T1.1a ✓, T1.1b ✓, T1.2 ✓, T1.4 ✓,
T2.3 ✓ — Tier 1 closed. Tier 2 (check_early_stopping(avg_q_value)
deletion + enrichment.rs constants soup) is next.

Plan reference: docs/plans/2026-05-10-sp21-train-eval-coherence-isv-defrost.md

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-10 20:45:00 +02:00
jgrusewski
74c7a80114 feat(sp21): T1.1a+T1.1b+T2.3 — signal-driven early-stopping (atomic)
Closes the patience-based early-stopping bug that ran xmd6b 30 epochs
past peak val performance (epoch 2: val_Sharpe=90, total_pnl=0.44 →
epoch 30: val_Sharpe=26, total_pnl=0.16) — a 70% loss of alpha to
training-induced overfitting.

Two intertwined bugs, fixed atomically:

T1.1a (wrong-source) at training_loop.rs:7234:
  - was: self.early_stopping.should_stop(-log_output.epoch_sharpe, epoch)
         reading the TRAINING ROLLOUT Sharpe (Thompson-noisy, in-sample,
         oscillates even when the model is frozen)
  - now: self.early_stopping.should_stop(log_output.val_loss, min_delta, epoch)
         reading the deterministic-backtest val_loss
  - The comment 5733 lines earlier (line 1499) explicitly says "Use
    val_Sharpe (deterministic backtest), NOT epoch_sharpe" — patience
    path was the inconsistency, backtracking already honored it.

T1.1b (hardcoded threshold) in early_stopping.rs:
  - was: EarlyStopping::new(patience, min_delta) with min_delta=0.001
         constructor-set, struct field, structurally meaningless
         against the val_loss noise floor (typical val_sharpe deltas
         are O(1-10), so 0.001 essentially never gates)
  - now: EarlyStopping::new(patience), should_stop(val_loss, min_delta,
         epoch) with min_delta computed per-call from
         ISV[VAL_SHARPE_VAR_EMA_INDEX=351] as
         sqrt(var_ema).max(0.5)
  - Floor 0.5 covers cold-start before var_ema bootstraps from sentinel
    per pearl_blend_formulas_must_have_permanent_floor.
  - Per feedback_isv_for_adaptive_bounds + feedback_adaptive_not_tuned.

T2.3 (test signature update) absorbed:
  - 6 existing unit tests migrated to new should_stop signature.
  - 1 NEW test (test_min_delta_can_change_per_call) verifying per-call
    threshold change works correctly.
  - EarlyStopping::min_delta struct field deleted.
  - Atomic per feedback_no_partial_refactor.

Affected files:
  - crates/ml/src/trainers/dqn/early_stopping.rs (struct + tests)
  - crates/ml/src/trainers/dqn/trainer/constructor.rs (new() arg)
  - crates/ml/src/trainers/dqn/trainer/training_loop.rs (call site)

Verification:
  - cargo check -p ml --tests: passes
  - cargo test -p ml --lib early_stopping: 8/8 pass

Behavioral expectation post-fix: xmd6b-shape runs (val_Sharpe rising
31→90 epochs 0-2, declining 90→26 epochs 3-30) will trigger early-stop
near the peak. With patience=5 and var_ema bootstrapping by epoch 2-3,
the controller detects "no improvement of ≥ 1σ for 5 consecutive
epochs" by ~epoch 7-8 and stops, saving ~22 epochs of overfitting.

Plan reference: docs/plans/2026-05-10-sp21-train-eval-coherence-isv-defrost.md
Tier 1 status: T1.1a ✓, T1.1b ✓, T2.3 ✓ (this commit). T1.2 + T1.4 next.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-10 20:40:54 +02:00
jgrusewski
4d4cd996db feat(sp21): T3.3 ISV defrost — hold_reward_ema (atomic Phase 3.2 wireup)
Closes the Phase 3.2 forward-reference loop in the SP20 aggregator.
Previously `out_inputs->per_bar_hold_reward = 0.0f` was hardcoded; the
per-bar Hold opportunity-cost producer existed
(experience_kernels.cu:3823 — `per_bar_opp_cost = -aux_conf × cost_scale`)
and wrote to `hold_baseline_buffer` and `r_micro` directly, but never
reached HOLD_REWARD_EMA. Result: HOLD_REWARD_EMA frozen at sentinel 0.0
across all observed epochs; the SP20 reward centering loop
(`r_micro += per_bar_opp_cost - HOLD_REWARD_EMA`) stayed uncentered,
biasing the policy's reward signal away from zero-mean.

T3.3 wireup mirrors the T2.2 alpha refactor: a per-env scratch buffer
that the producer writes at every step (alongside the existing
`hold_baseline_buffer` write), and the aggregator reads at the same
step. Sums opp_cost over Hold-direction envs only; emits the mean as
`per_bar_hold_reward`. The HOLD_REWARD_EMA's gate (`hold_fraction > 0.5f`
from T3.2) preserves the strict-majority semantic from before.

Atomic across producer site, kernel signature, aggregator, launcher,
collector, and tests (per feedback_no_partial_refactor):

  - experience_kernels.cu — new `float* per_bar_opp_cost_per_env`
    kernel arg, NULL-tolerant write at line ~3823.
  - sp20_aggregate_inputs_kernel.cu — new arg, 6th shmem stripe
    (`sh_opp_cost_sum`), per-thread Hold-gated accumulation, tree
    reduction extended to 6 stripes, output `per_bar_hold_reward =
    opp_cost_sum / hold_count` when hold_count > 0 else 0.
  - sp20_aggregate_inputs.rs — launcher signature, dynamic_shmem_bytes
    5→6 stripes, doc table, internal test.
  - gpu_experience_collector.rs — new `per_bar_opp_cost_per_env:
    CudaSlice<f32>` field, alloc, struct construction, kernel-arg
    pass at both experience_env_step and sp20_aggregate_inputs
    launches (raw_ptr).
  - sp20_aggregate_inputs_test.rs — helper renamed
    `run_kernel_with_is_win_and_opp_cost`, 4 existing sites pass
    NULL, 2 NEW oracle tests verifying real producer + NULL fallback.
  - sp20_phase1_4_wireup_test.rs — NULL fallback at the wireup site;
    HOLD_REWARD_EMA-stays-at-zero assertion remains valid.

Verification:
  - cargo check -p ml --tests: passes (warnings only)
  - cargo test -p ml --test sp20_aggregate_inputs_test --features cuda
    -- --ignored: 12/12 GPU oracle tests pass on RTX 3050 Ti, including
    both new T3.3 tests (per_bar_hold_reward_means_over_hold_envs_only,
    null_per_bar_opp_cost_emits_zero).
  - cargo test -p ml --test sp20_phase1_4_wireup_test --features cuda
    -- --ignored: 2/2 pass under the new NULL-tolerant contract.

Plan reference: docs/plans/2026-05-10-sp21-train-eval-coherence-isv-defrost.md
Tier 3 status: T3.1 ✓, T3.2 ✓, T3.3 ✓ (this commit), T3.4 withdrawn,
T3.5 cascade-pending, T3.6 withdrawn.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-10 20:35:39 +02:00
jgrusewski
1790a31b66 feat(sp21): T3.1+T3.2 ISV defrost — wr_ema + hold_pct_ema (atomic)
SP21 Tier-3 foundation: defrost two production-frozen ISV slots that
pinned at 0 across every observed training epoch (d7bj7, xmd6b). Both
bugs in the same aggregator kernel, same structural shape: per-step
binary majority-vote indicators feeding fractional EMAs.

T3.1 — WR_EMA defrost (sp20_aggregate_inputs_kernel.cu):
  - was: is_win_out = (2 * wins_count >= closed_count) ? 1 : 0
         binary "majority won this step" indicator
  - now: win_fraction_out = (float)wins_count / (float)closed_count
         fractional in [0, 1]
  - With actual val WR≈0.46, the binary signal was 0 most of the time,
    pinning WR_EMA at 0 across 30+ epochs in xmd6b. EMA now converges
    to population win rate.

T3.2 — HOLD_PCT_EMA defrost (same kernel, same pattern):
  - was: action_is_hold_out = (hold_count * 2 > n_envs) ? 1 : 0
         strict-majority indicator
  - now: hold_fraction_out = (float)hold_count / (float)n_envs
  - Cascade: HOLD_COST_SCALE controller compared hold_pct_ema=0 to
    tgt±0.05, always saw < lower, ramped × 0.95 → clamped at floor
    0.01 (the hold_cost_scale=0.0100 observation in d7bj7 logs).
    T3.5 expected to cascade-fix in next training run.
  - HOLD_REWARD_EMA gate preserves strict-majority semantic via
    hold_fraction > 0.5f test in the consumer kernel.

Atomic across struct fields, kernel logic, Rust mirror, byte
serialization, doc tables, and 4 test files (per
feedback_no_partial_refactor):
  - sp20_aggregate_inputs_kernel.cu (struct + 2 computations)
  - sp20_emas_compute_kernel.cu (struct + reader + gate)
  - sp20_aggregate_inputs.rs (doc table)
  - sp20_emas_compute.rs (Rust struct + serialize + 3 tests)
  - sp20_aggregate_inputs_test.rs (reader sig + 4 test assertions)
  - sp20_emas_compute_test.rs (5 struct literal updates)
  - sp20_phase1_4_wireup_test.rs (HOLD_PCT_EMA expected 0.625)

Verification:
  - cargo check -p ml --tests: passes (warnings only)
  - cargo test -p ml --lib sp20_emas: 6/6 unit tests pass

Pearl candidate: binary-majority aggregator over a fractional
underlying signal cannot serve as input to a fractional-target EMA.
The previous fix (commit 64bbbe418) addressed the per-bar vs segment
predicate at the producer site, but didn't notice the aggregator's
binarization step still collapsed the fraction to {0, 1}. Two bugs
in series, both now resolved.

Plan reference: docs/plans/2026-05-10-sp21-train-eval-coherence-isv-defrost.md
Tiers 3.3-3.6 remaining; T3.5 expected to cascade-fix.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-10 20:13:13 +02:00
jgrusewski
6df44e4c6d docs(sp21): plan + revert sp20-aux-h-fixed experiment
Reverts the forced H=30 diagnostic in aux_horizon_update_kernel.cu
(commit c78c4766c) — the experiment confirmed the data ceiling
hypothesis (aux_dir_acc dropped 0.46→0.50 with pred_tanh collapsing
to ~0, opposite of the bootstrap-failure prediction). Restores the
adaptive Pearl-A + Wiener-α floor logic.

Adds SP21 plan: train/eval coherence + ISV defrost. Catalogs 26
findings across 5 tiers from a pair-audit of MinIO-archived training
logs (xmd6b 30 epochs, d7bj7 2 epochs).

Cross-cutting principle: every threshold or bound in training
control flow must be signal-driven from an ISV slot, not hardcoded
(per feedback_isv_for_adaptive_bounds + feedback_adaptive_not_tuned).

Canonical pattern documented: ISV[CURIOSITY_PRESSURE_INDEX=346]
(SP11 Fix 39) is the reference implementation.

Meta-finding: across SP3-SP20 we've been monitoring training-rollout
metrics (Thompson-noisy) instead of val metrics (deterministic
backtest). val_PF=1.18-1.33 with WR=46-48% across 30 epochs of
xmd6b shows the policy DOES extract asymmetric-payoff alpha — we
just couldn't see it through the meter inflation.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-10 20:12:20 +02:00
jgrusewski
14bafb5b58 fix(argo-train): apply train-template.yaml before single-job submission
The multi-seed path already does `kubectl apply` before `argo submit`,
so cluster template stays in sync with source. The single-job path used
`argo submit --from=wftmpl/train` directly, expecting the cluster's
template to already match — which silently drifts when defaults change.

Caused workflow train-jpxvn (2026-05-10) to dispatch with stale
imbalance-bar-threshold=0.5 default (the cluster's old value) when the
source had been bumped to 20.0. Triggered near-OOM in feature extraction.

One-line fix: apply the template before submission. Mirrors what
multi-seed already does. No behavior change for users who pass explicit
flags; just makes implicit defaults track source.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-10 17:47:44 +02:00
jgrusewski
c78c4766ca experiment(sp20-aux-h-fixed): force H=30 to test bootstrap failure hypothesis
Throwaway diagnostic edit on aux_horizon_update_kernel.cu. Tests whether
the adaptive horizon's self-defeating loop (WR=50% → H~2 → noise horizon →
53.5% acc → WR=50%) is the actual bottleneck. If aux_dir_acc rises to
58%+ at fixed H=30, bootstrap failure confirmed and the proper fix is
constraint-based (min hold time during exploration). Audit-doc updated.

DO NOT merge to mainline — branch is throwaway after the experiment.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-10 17:33:44 +02:00
jgrusewski
d6bfad7033 docs(sp20): consolidate Phase 5 audit-doc + close-out
Replaces the per-commit audit-doc entries from commits 1-3 with a single
consolidated Phase 5 close-out entry. The consolidated entry covers:

  - The full design rationale: gate the REWARD (not the target); avoids
    needing q_mean_a entirely. At low aux confidence, r_used → 0 ⇒
    Bellman target collapses to gamma * Q(s', a').
  - All 5 components: trainer buffer, FusedTrainerCtx accessor, training
    loop wire-up, kernel signature + gate, launcher arg.
  - NULL-tolerance contract: aux_conf_at_state == NULL OR isv_signals
    == NULL ⇒ gate = 1.0 (identity).
  - Default-state semantics: alloc_zeros 0.0 sentinel → gate ≈ 0.12 →
    reward mostly suppressed pre-population (graceful degradation).
  - reward_bias interaction (composes cleanly — gate damps reward
    pre-projection, reward_bias lifts target Q-mean per-branch).
  - Plan accuracy errata: the user spec's "add aux_conf_at_state_buf
    field to GpuBatch struct" was unnecessary — GpuBatch doesn't
    carry the SP13 B1.1b aux_sign_labels_ptr either; both follow the
    "trainer-only buffer + direct-gather" pattern.
  - Test coverage: 3 CPU math tests + 1 GPU behavioral integration test.
  - Confidence: medium-high that the gate fires correctly on real data;
    end-to-end smoke validation deferred (no smokes dispatched per
    controller instruction).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-10 14:54:41 +02:00
jgrusewski
ab4a7db33c feat(sp20): c51_loss launcher aux_conf arg + Phase 5 gate tests
Threads `self.aux_conf_at_state_buf` into the `c51_loss_batched` launch
in `GpuDqnTrainer::launch_c51_loss`. Position matches the kernel's
appended trailing arg from the previous commit.

Tests added in `crates/ml-dqn/src/gpu_replay_buffer.rs::tests`:

  - `aux_gate_high_confidence_passes_full_target` (CPU pure-math):
    gate(aux_conf=0.5, threshold=0.10, temp=0.05) > 0.99 proves
    high-confidence reward pass-through.
  - `aux_gate_low_confidence_attenuates_reward` (CPU pure-math):
    gate(aux_conf=0.02, threshold=0.10, temp=0.05) < 0.20 proves
    the uncertain-state neutralizer semantic.
  - `aux_gate_temp_floor_keeps_gate_finite` (CPU pure-math):
    sweeps {temp, aux_conf, threshold} and asserts finite gate ∈ [0,1]
    across the ISV-controllable parameter range — proves the
    fmaxf(temp, 1e-3) floor keeps the kernel numerically safe.
  - `aux_conf_direct_to_trainer_gather_populates_destination` (GPU
    behavioral): wires a fresh CudaSlice<f32> as the trainer
    destination, inserts 8 transitions with strictly-positive distinct
    aux_conf values, samples 1, asserts the trainer destination
    buffer post-sample holds a value from the inserted set (NOT the
    alloc_zeros sentinel) — proves the direct-gather wiring actually
    populates the trainer buffer with non-trivial data.

All 3 CPU math tests + 1 GPU integration test pass on RTX 3050.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-10 14:54:41 +02:00
jgrusewski
96b76d9298 feat(sp20): c51_loss_batched aux_conf_at_state reward gate
Adds the Phase 5 consumer kernel-side gate. New kernel arg
`const float* __restrict__ aux_conf_at_state` appended to
`c51_loss_batched`'s signature. Gate computation runs once per sample
at the kernel-entry reward-setup site (after the #27 ensemble-
disagreement adjustment), then the gated `reward` propagates through
every branch's `block_bellman_project_f` call without per-branch changes.

Formula:
    gate    = sigmoid((aux_conf - threshold) / temp)
    reward  = gate * reward
where:
    threshold = ISV[AUX_CONF_THRESHOLD_INDEX=518]
    temp      = max(ISV[AUX_GATE_TEMP_INDEX=519], 1e-3)

Mathematical interpretation: at low aux confidence (gate→0),
`r_used → 0`, so the Bellman target becomes `gamma * Q(s', a')`. The
Q value at the current state collapses toward `gamma * Q(s', a')` —
model gets no reward feedback on uncertain transitions. Effectively
"don't update Q on uncertain transitions" — the "uncertain-state
neutralizer" semantic from the Phase 3 Task 3.4 audit doc spec §4.4.

NULL-tolerant: `aux_conf_at_state == NULL` OR `isv_signals == NULL`
⇒ gate skipped (identity, no-op = pre-Phase-5 behaviour). Test
scaffolds without a wired aux head still work.

Out of scope: `iqn_dual_head_kernel.cu` — IQN is the auxiliary loss,
C51 is production. Gating IQN is more complexity for marginal gain.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-10 14:54:41 +02:00
jgrusewski
d3a057af4f feat(sp20): allocate trainer aux_conf_at_state_buf + wire PER direct-gather
Phase 5 plumbing — consumer-side wire-up. Allocates `aux_conf_at_state_buf:
CudaSlice<f32>` ([batch_size]) on `GpuDqnTrainer`, exposes the raw_ptr via
`aux_conf_at_state_buf_ptr()` and the `FusedTrainerCtx` delegating accessor
`trainer_aux_conf_at_state_buf_ptr()`, and invokes
`GpuReplayBuffer::set_trainer_aux_conf_ptr` at both fused_ctx init sites in
training_loop.rs (init + re-init, atomic per `feedback_no_partial_refactor`).

The PER `gather_f32_scalar` now writes the SAMPLED bar's per-batch aux_conf
directly into the trainer's f32 buffer on every step — same direct-to-trainer
pattern as the SP13 B1.1b `aux_nb_label_buf` (i32) wire-up immediately above
the new call. The c51_loss_batched reward gate (lands in the next commit)
reads this buffer to compute `gate = sigmoid((aux_conf - ISV[AUX_CONF_THRESHOLD])
/ ISV[AUX_GATE_TEMP])` and applies `r_used = gate * reward` at the Bellman
projection.

`alloc_zeros` cold-start: 0.0 sentinel → at threshold ≈ 0.10 and temp ≈ 0.05
the gate is `sigmoid(-2) ≈ 0.12` → reward is mostly suppressed pre-population.
This is the "graceful degradation" semantic from the Phase 3 Task 3.4 audit
doc spec §4.4. Once PER's direct-gather populates from the producer ring on
the first sample step, the per-bar aux_conf values drive the gate as designed.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-10 14:54:41 +02:00
jgrusewski
d1a8ec206f fix(sp20): bump MAX_UPLOAD_BYTES 2GB → 8GB + audit doc
L40S (48GB) and H100 (80GB) have plenty of headroom; the 2GB cap was a
conservative leftover that tripped on workflow zgjgc with 17.8M imbalance
bars (3.2GB). 8GB budget leaves ~28GB free on L40S after model + activations
+ workspace.

Updates both call sites:
- DqnGpuData::upload_slices (training data, the failing one on zgjgc)
- PpoGpuData::upload (market data, same constant)

Error message format strings updated 2.0 GB → 8.0 GB.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-10 13:45:17 +02:00
jgrusewski
db9fd1b2da feat(sp20): parallelise imbalance-bar OHLCV reduction via two-pass decomposition
Bottleneck B of two: replaces the sequential `ImbalanceBarSampler` walk
in `mbp10_to_imbalance_bars` with a two-pass parallel decomposition.
Bit-identical to sequential — verified by 6/6 passing tests in
`crates/ml-features/tests/imbalance_bars_parallel_bit_equiv_test.rs` at
both dense (T=20, ~23k bars) and sparse (T=500, ~300 bars) emission
regimes, exact 0.0 diff on every f64 field.

Why two-pass (NOT time-bucket sharding)
=======================================
Time-bucket sharding (the SP20 OFI pattern in
`compute_ofi_per_bar_parallel`) bit-equates to sequential because OFI
features derive from BOUNDED rolling windows (VPIN ≤50, Kyle ≤100,
trade-imb ≤100, OFI-stats ≤300). After ≥window-size warmup updates,
two rolling buffers started from different initial states converge
bit-identically.

`ImbalanceBarSampler` is fundamentally different: its `cumulative_imbalance`
is a path integral that resets only when crossing ±threshold. There is
NO bounded-lookback window. Two replays of the same trade tape starting
from different cum_imb offsets emit at DIFFERENT trade indices, and the
phase offset is NOT guaranteed to vanish at any future trade — even
after many emissions, a bounded phase difference can persist
indefinitely. An earlier time-bucket-sharded prototype with WARMUP=10000
trades passed K=4/K=8 at threshold=20 (dense emissions, ~30 emissions
per shard's warmup window) but FAILED at threshold=500 with a 1-bar
count drift (golden=303, parallel=304) — exactly the kind of phase-
offset divergence predicted by the path-integral analysis.

Two-pass decomposition
======================
Pass 1 (sequential, lightweight): walk the trade tape ONCE tracking only
cum_imb, prev_price, last_direction. Record the trade index of every
emission-triggering trade as the end of a segment. No OHLCV state, no
Vec<OHLCVBar> allocation, no per-trade conditional bar construction.
O(N) simple arithmetic — for 50k-2M trades it runs in 1-15ms.

Pass 2 (parallel rayon): each emission segment [boundary[i-1]+1,
boundary[i]] produces exactly one bar via independent OHLCV reduction
over its trade slice. Segments are fully independent; `par_iter()` over
segment ranges is trivially correct.

Bit-equivalence guarantee
=========================
Pass 1 mirrors `ImbalanceBarSampler::update` arithmetic verbatim:
- zero-volume early-return (alternative_bars.rs:484-486)
- direction tie-break (alternative_bars.rs:495-506)
- cum_imb update (alternative_bars.rs:508-510)
- emission threshold check (alternative_bars.rs:522-523)
- prev_price/last_direction kept across emissions (reset() at :558-566)
- NO trailing-partial-bar flush (matches sequential exactly)

Pass 2's per-segment OHLCV reduction matches the sampler's per-bar OHLCV
update verbatim: open = first non-zero-volume trade in segment, close =
last non-zero-volume trade, high/low = max/min over non-zero-volume
trades, volume = sum, timestamp = open's timestamp.

Performance (release-mode bench, 16-thread rayon pool)
======================================================
n=100k:   seq=1.39ms,  par=1.88ms,  speedup=0.74x  (rayon overhead dominates)
n=500k:   seq=8.04ms,  par=3.35ms,  speedup=2.40x
n=2M:     seq=27.59ms, par=12.39ms, speedup=2.23x

Speedup is bounded by Pass 1 (sequential, ~5-7ms at 2M trades) since
Pass 2 (parallel, ~2-3ms at 2M trades on 16 threads) is ~3x faster
than Pass 1's sequential floor. At production scale (50k-500k trades
per `mbp10_to_imbalance_bars` call) we get 2-2.4x speedup over the
all-sequential baseline.

The `min_bars_per_task` cutoff falls back to a sequential Pass 2 when
segment count < 256 (the rayon spawn overhead exceeds the parallel
benefit). Tunable via the second function arg.

Files changed
=============
- crates/ml-features/src/alternative_bars.rs (+1 -1):
    derive `Clone` on `ImbalanceBarSampler` for parallel-shard use
    (carried through this commit even though Pass 1's hand-rolled walk
    in `mbp10_loader.rs` doesn't need it — keeps the type clonable for
    future test/bench infrastructure).
- crates/ml-features/src/lib.rs (+3 -2):
    re-export `imbalance_bars_parallel`, `imbalance_bars_sequential`,
    `DEFAULT_IMBALANCE_BAR_MIN_BARS_PER_TASK`.
- crates/ml-features/src/mbp10_loader.rs (+253 -15):
    new `imbalance_bars_parallel`, `imbalance_bars_sequential`,
    `DEFAULT_IMBALANCE_BAR_MIN_BARS_PER_TASK` const. Replace the inline
    sampler-walk loop in `mbp10_to_imbalance_bars` with a call into
    `imbalance_bars_parallel`. Module note explains why time-bucket
    sharding doesn't apply.
- crates/ml-features/tests/imbalance_bars_parallel_bit_equiv_test.rs
    (+272 NEW): hermetic synthetic-trade bit-equivalence tests at:
        * default cutoff (par_iter Pass 2)
        * forced par_iter (min_bars_per_task=1)
        * forced sequential Pass 2 (min_bars_per_task=usize::MAX)
        * empty input
        * small input (Pass 2 below par cutoff)
        * high-threshold sparse emissions (the case that broke the
          time-bucket sharding prototype). All pass exact 0.0 diff.

Also note: 293/293 ml-features lib tests pass (no regression).

Pairs with the file-level par_iter trade extraction in
8f5c64e10 (Bottleneck A). Together they parallelise both halves of
`mbp10_to_imbalance_bars`.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-10 12:53:51 +02:00
jgrusewski
9ccc37749a feat(sp20): parallelise per-file MBP-10 trade extraction in mbp10_to_imbalance_bars
Bottleneck A of two: the loop in `mbp10_to_imbalance_bars` that calls
`extract_trades_from_dbn_file` + `filter_front_month_mbp10` per file
ran serially across the 9 quarterly DBN files. Each file is independent
(different contract universe per quarter), the front-month filter is
purely intra-file, and the final `all_trades.sort_by(|a, b| timestamp)`
re-sequences across files — so reduce order across files is irrelevant
to correctness.

Switched the loop to `dbn_files.par_iter().filter_map(...).collect()`,
mirroring the trades_loader-side pattern in
`precompute_features.rs:551-565`. Per-file logging (raw count → filtered
front-month count) preserved verbatim. On the `ci-compile-cpu` 28-core
node decoding 9 .dbn.zst files, the file-decoder/zstd-decompress phase
should drop from ~9× single-file latency to ~1× — bounded by the slowest
single file.

This is the trivial half of the parallelisation. Bottleneck B (the
sequential `ImbalanceBarSampler` pass over the concatenated trade list)
follows in the next commit with time-bucket sharding + warmup overlap +
bit-equivalence test, mirroring the SP20 OFI pattern at
`crates/ml-features/src/ofi_calculator.rs::compute_ofi_per_bar_parallel`.

Build: `cargo check -p ml-features` clean.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-10 12:53:51 +02:00