705d6c156b4afc732da2ce26e89bc2caa4d41baf
66 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
705d6c156b |
audit: ISV-ify 10 more design constants — Schulman + bootstraps + streaming α
Per `feedback_isv_for_adaptive_bounds` + user "do all except floors
and clamp bounds": 10 more constants moved from kernel-side `#define`s
into ISV slots (78 slots total now).
## Slot additions (468-477)
RL_SCHULMAN_TOLERANCE_INDEX (468, =1.5) — shared by 4 controllers
RL_SCHULMAN_ADJUST_RATE_INDEX (469, =1.5) — shared by 4 controllers
RL_STREAM_ALPHA_INDEX (470, =0.05) — shared by var + kurt streaming
RL_KURT_GAUSSIAN_INDEX (471, =3.0)
RL_KURT_NOISE_FLOOR_INDEX (472, =1.0)
RL_TAU_BOOTSTRAP_INDEX (473, =0.005)
RL_EPS_BOOTSTRAP_INDEX (474, =0.2)
RL_ROLLOUT_BOOTSTRAP_INDEX (475, =2048)
RL_REWARD_SCALE_BOOTSTRAP_INDEX (476, =1.0)
RL_PPO_RATIO_CLAMP_BOOTSTRAP_INDEX (477, =10.0)
## Skipped (per user "do all except floors and clamp bounds")
* `*_MIN`/`*_MAX` clamp bounds (algebraic domain — risk γ=1.5 nonsense)
* Numerical floors: ABS_MEAN_FLOOR=1e-6, M2_SQ_FLOOR=1e-12, EPS_PNL=1e-3
(risk div-by-zero if mis-tuned)
* C51 atom layout (V_MIN/V_MAX) — architecture, not config
## Wiring
* Shared Schulman pattern: 4 controllers (ppo_clip, target_tau,
rollout_steps, plus per_α independent KURT slots) now read TOLERANCE
+ ADJUST_RATE from the same 2 ISV slots. Single source of truth.
* Each controller's bootstrap (1st-emit on sentinel-zero) reads
isv[*_BOOTSTRAP_INDEX] instead of #define value. The `prev ==
BOOTSTRAP` first-observation replace-direct check also reads from
ISV.
* 2 streaming kernels (var + kurt) share RL_STREAM_ALPHA_INDEX.
## Diag bake-in
JSONL `isv_config` block grows by 10 new fields: schulman_tolerance,
schulman_adjust_rate, stream_alpha, kurt_gaussian, kurt_noise_floor,
tau_bootstrap, eps_bootstrap, rollout_bootstrap,
reward_scale_bootstrap, ppo_ratio_clamp_bootstrap. Total isv_config
fields: 26.
Also includes windowed action_entropy fix (was structurally 0 at
b_size=1) — accumulates EMA-smoothed action distribution over
~1k-step window, computes entropy on the windowed dist. Makes the
exploration metric meaningful at b_size=1.
## Slot total
RL_SLOTS_END: 468 → 478. **78 total ISV slots.**
## Verified gates (local sm_86)
G1 isv_bootstrap ✅ (with 10 new assertions)
G3 controllers ✅
G4 target_update ✅
integrated_smoke ✅
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
|
||
|
|
827a0e9416 |
fix(rl): ISV-ify ALL remaining tunable design constants (10 new slots)
Per `feedback_isv_for_adaptive_bounds`: every controller design knob
that's genuinely tunable now lives in ISV instead of as a kernel-side
`#define`. Tuning is a re-seed (kernel launch with new arg) rather
than a recompile.
## New ISV slots (10 design constants)
RL_REWARD_CLAMP_WIN_INDEX (452, =1.0) apply_reward_scale
RL_REWARD_CLAMP_LOSS_INDEX (453, =3.0) apply_reward_scale
RL_KL_TARGET_INDEX (454, =0.01) rl_ppo_clip_controller
RL_IMPROVEMENT_THRESHOLD_INDEX (455, =0.99) rl_lr_controller
RL_PLATEAU_PATIENCE_INDEX (456, =1000.0) rl_lr_controller
RL_DIV_TARGET_INDEX (457, =0.01) rl_target_tau_controller
RL_ENTROPY_TARGET_FRAC_INDEX (458, =0.7) rl_entropy_coef_controller
RL_KURT_LIFT_SCALE_INDEX (459, =7.0) rl_per_alpha_controller
RL_PPO_CLAMP_MARGIN_INDEX (460, =10.0) rl_ppo_ratio_clamp_controller
RL_LR_WARMUP_STEPS_INDEX (461, =2000.0) rl_lr_controller
RL_SLOTS_END: 452 → 462.
## Constants NOT converted (truly fundamental)
* All `*_INDEX` (ABI)
* All `*_MIN`/`*_MAX` clamp bounds (algebraic domain)
* All `*_BOOTSTRAP` (one-shot init)
* `WIENER_ALPHA_FLOOR` (per pearl_wiener_alpha_floor_for_nonstationary)
* Schulman pattern parameters (`*_TOLERANCE`/`*_ADJUST_RATE`)
* C51 (`Q_N_ATOMS`, `V_MIN/MAX`, `N_ACTIONS`)
* Kernel numerics (`STREAM_ALPHA`, `ABS_MEAN_FLOOR`, `EPS_PNL`)
* `KURT_GAUSSIAN` (statistical constant = 3.0 for Gaussian)
* `KURT_NOISE_FLOOR` (defensive)
* `LR_BOOTSTRAP`/`LR_MIN`/`LR_MAX`/`LR_LOSS_EMA_ALPHA`/`DECAY_FACTOR`
## New infrastructure
New CUDA kernel `rl_isv_write.cu` — generic single-thread device-side
seeder taking `(int slot, float value)`. Trainer loops calling it
once per design constant at init. Replaces the prior pattern of
extending `rl_streaming_clamp_init`'s arg list every time a new
constant was added.
## Ordering fix
Design constants must be seeded BEFORE controllers bootstrap — the
controllers' bootstrap paths read these slots (e.g.
`rl_entropy_coef_controller` reads `RL_ENTROPY_TARGET_FRAC_INDEX`
to derive its target). Without correct ordering, controllers see
sentinel 0.0 and bootstrap to wrong values (caught by failing G1
test before commit). Seed loop runs at TOP of
`with_controllers_bootstrapped`.
## Diag bake-in
JSONL gains `isv_config` block exposing all 10 design constants per
step:
isv_config.{reward_clamp_win, reward_clamp_loss, kl_target,
improvement_threshold, plateau_patience, div_target,
entropy_target_frac, kurt_lift_scale, ppo_clamp_margin,
lr_warmup_steps}
Post-hoc analysis can correlate any controller's behaviour with the
exact design constants it saw, without grepping the source for
`#define` defaults.
## Test updates
G1 (isv_bootstrap) + G3 (r5_controllers) — skip 10 new design-
constant slots in sentinel-zero loop, assert seeded values
separately.
## Verified gates (local sm_86)
G1 isv_bootstrap ✅ (with 10 new assertions)
G3 controllers ✅
G4 target_update ✅
integrated_smoke ✅
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
|
||
|
|
644fbe0348 |
fix(rl): ISV-driven K-loop divisor + max ceiling (slot 450, 451)
f2ggr confirmed K-loop wiring works mechanically but K=8 firing on
22 % of steps over-trained at b_size=1: KL excursions to 12.44
(vs prior 3.4e-4), policy overshoot, reward/trade -$0.585 → -$0.723.
Per `feedback_isv_for_adaptive_bounds` the K-loop config must live
in ISV, not as hardcoded values in the trainer. Two new slots:
RL_K_LOOP_DIVISOR_INDEX (450) — divides n_rollout_steps to get K
Default 2048 (matches ROLLOUT_BOOTSTRAP
so K=1 at controller bootstrap)
RL_K_LOOP_MAX_INDEX (451) — clamp ceiling on K
Default 4 (was hardcoded 8; halved
to prevent gradient overtraining)
K computation in step_with_lobsim now reads both from ISV:
K = clamp(isv[404] / isv[450], 1, isv[451])
Halves worst-case overtraining while preserving the controller
cascade activation (KL above noise floor, ε actively adapting,
ratio_clamp firing). Distribution shifts from K=8 @ 22% → K=4 @ 22%
(half the gradient updates in the high-noise case).
## Wiring
`rl_streaming_clamp_init.cu` extended to seed 5 ISV-resident design
constants (was 3): adv_var_clamp, td_kurt_clamp, adv_var_target,
k_loop_divisor, k_loop_max. Still one kernel call, no HtoD.
## Diag bake-in
JSONL `k_updates` field replaced with `k_loop` block:
k_loop.k_updates — actual K used this step
k_loop.divisor — current divisor (reads isv[450])
k_loop.max — current max (reads isv[451])
Post-hoc analysis can verify the K-computation by independently
recomputing K from isv[404] / k_loop.divisor.
## Slot allocation
RL_SLOTS_END: 450 → 452 (+2 new config slots).
## Test updates
G1 + G3 skip slots 450, 451 in sentinel-zero loop and assert seeded
values (2048.0 + 4.0) separately.
## Verified gates (local sm_86)
G1 isv_bootstrap ✅
G3 controllers ✅
G4 target_update ✅
integrated_smoke ✅
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
|
||
|
|
1d8ef94848 |
fix(rl): wire n_rollout_steps as K-loop + raise LR_MIN to 1e-4
Two coordinated fixes for the alpha-rl-frt7s findings:
## Issue 1: n_rollout_steps controller was write-only
ISV consumer audit confirmed: 7 of 8 RL controllers had a non-
controller consumer in the per-step path; n_rollout_steps had ZERO.
The controller adapted its output between 256-8192 but nothing read
it. Bit-identical losses between cvf86 and frt7s confirmed: even
fixing the target (0.1 → 5.0) and putting the controller into
healthy HOLD/SHRINK/WIDEN distribution had zero behavioral impact
because no downstream code gated on the emitted value.
### Fix: wire as DQN-replay + PPO+V K-loop multiplier
step_with_lobsim now wraps (sample_and_gather + step_synthetic +
PER priority update) in a K-loop where:
K = clamp(isv[RL_N_ROLLOUT_STEPS_INDEX] / 1024, 1, 8)
Mapping:
* isv[404] = 256 (MIN) → K = 1 (current behavior)
* isv[404] = 2048 (BOOTSTRAP) → K = 2
* isv[404] = 8192 (MAX) → K = 8
Each iteration re-samples PER (different transitions per Adam step)
and runs full Q + π + V forward + backward + Adam. Adapts the
training:env ratio so noisy-advantages regimes get more gradient
samples per env step without slowing env stepping. Directly
addresses the b_size=1 gradient starvation that left l_q stuck at
2.82 in frt7s.
Semantic fit: n_rollout_steps's design intent ("noisy advantages →
need more samples per update") now drives "more training updates
per env step" — equivalent semantics, fits the b_size=1
architecture without requiring a PPO rollout buffer refactor.
`last_k_updates` field tracks the per-step K value for diag.
## Issue 2: LR plateau-decay Q-lock
frt7s deep dive showed:
* Q best=2.3230 locked at step ~783 from a brief downward
excursion during early-training noise
* loss_ema range across 50k steps: [2.323, 3.113]; mean 2.819,
std 0.104
* ZERO steps had loss_ema < best in entire run (let alone <
best × 0.99 = 2.30 threshold)
* 7 LR halvings drove all heads to LR_MIN = 1e-5 by step 7783
* At 1e-5, Q's per-step Adam update is too small to escape;
l_q stayed at ~2.82 for 42k more steps
The plateau-decay is CORRECTLY identifying "model has stopped
improving" — the fix isn't to make plateau detection less
sensitive (loosening threshold to 0.95/0.90 still finds zero
improvements). The fix is to raise the floor LR so the model
has enough learning rate to escape the noise-locked best.
### Fix: LR_MIN 1e-5 → 1e-4 + WARMUP_STEPS 500 → 2000
* LR_MIN raised 10× — even at the plateau-decay floor the model
gets meaningful gradient. Still 10× below LR_BOOTSTRAP=1e-3
so the controller has full dynamic range.
* WARMUP_STEPS raised 4× — gives loss_ema 2000 observations
(≈145 EMA half-lives at α=0.05) to settle BEFORE best is
locked. Prevents the "lucky early excursion locks unreachable
bar" failure mode.
## Diag bake-in
JSONL gains `k_updates` field (per-step K value from the n_rollout
loop) so post-hoc analysis can correlate the K-multiplier with
loss trajectories.
## Verified gates (local sm_86)
G1 isv_bootstrap ✅
G3 controllers ✅
G4 target_update ✅
integrated_smoke ✅
## Quality-first scope decision
User requested "quality over speed". Considered alternatives:
* Building a proper PPO rollout buffer (Issue 1) — significant
refactor, ~1-2 days. K-loop interpretation chosen instead
because it (a) matches the controller's design intent, (b)
requires no buffer/gradient-accumulation infrastructure, (c)
directly addresses Q learning starvation by giving more
gradient samples per env step.
* Encoder LR decoupling (Issue 2) — encoder receives gradient
from all head backward kernels with their own LRs; treating
the encoder separately would require restructuring all
backward kernels. LR_MIN raise + WARMUP extension gives the
same benefit at the head level without that scope.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
|
||
|
|
95dcc4e312 |
fix(rl): ISV-driven ADV_VAR_RATIO_TARGET for rl_rollout_steps_controller
cvf86 controller_branch diag (commit
|
||
|
|
708c121f20 |
fix(rl): bounded multiplicative step + noise-floor on rollout_steps + per_α
kc2h9 confirmed: clamping streaming-kernel outputs to [≤100, ≤30]
had ZERO behavioral impact because rl_rollout_steps_controller's
prior design used `scale = clamp(input/target, 0.5, 2.0)` — the
scale saturated to ±2× on the SIGN of (input − target), not the
magnitude. With target=0.1 and typical input=1–10 the controller
slammed to MAX in ≤4 steps regardless of whether input was 4 or
3e5. Bit-identical losses between gxhr8 and kc2h9 confirmed the
saturation.
## Fix 1: rl_rollout_steps_controller — same Schulman pattern as ppo_clip
* input > TARGET × 1.5 → scale = 1.5 (widen)
* input < TARGET / 1.5 → scale = 1/1.5 (shrink)
* in-band → scale = 1.0 (hold)
* input < TARGET × 0.01 → return (noise floor — hold prev)
Per-step adjustment bounded at 1.5×, so rollout_steps drifts
smoothly toward MIN/MAX rather than slamming there. The noise-floor
gate matches the pattern from
`pearl_multiplicative_controllers_need_bounded_step_and_noise_floor`
applied to the ε and τ controllers earlier in R9.
## Fix 2: rl_per_alpha_controller — noise-floor gate (defensive)
per_α uses a LINEAR lift `0.4 + 0.2·(kurt-3)/7` (not multiplicative),
so it doesn't have the saturation bug. But added a noise-floor gate
at KURT_NOISE_FLOOR = 1.0 so a sub-Gaussian kurtosis reading from
the streaming estimator's startup window (when per-step batch-mean
deviations are small before tails develop) doesn't drag α toward
PER_ALPHA_MIN on cold-start.
## Diag bake-in (per user request "bake in diags")
JSONL gains a `controller_branch` block exposing the
multiplicative-controller inputs alongside their design targets:
controller_branch: {
rollout_steps_input: isv[421], rollout_steps_target: 0.1,
ppo_clip_input: isv[419], ppo_clip_target: 0.01,
target_tau_input: isv[418], target_tau_target: 0.01,
per_alpha_input: isv[422], per_alpha_target: 0.6,
}
Post-hoc analysis can compute the branch each step (WIDEN / HOLD /
SHRINK / NOISE) by comparing input/target against the ±33%
tolerance band, revealing whether each controller is being driven
by real signal or sitting in the in-band hold zone. Targets are
reflected from the kernel #defines (synchronised by code review at
the controller-cu file level — there's no ISV slot for these
design constants because they're fundamental to the controller's
behaviour, not adaptive).
## Verified gates (local sm_86)
G1 isv_bootstrap ✅
G3 controllers ✅
G4 target_update ✅
integrated_smoke ✅
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
|
||
|
|
66115007ab |
fix(rl): ISV-driven output clamp on streaming var/kurtosis kernels
gxhr8 confirmed the streaming kernels work — both formerly-dead
controllers (rl_rollout_steps, rl_per_alpha) now adapt instead of
pegging at MIN. But the unclamped streaming outputs reached
advantage_var_ratio = 3e5 (when streaming-mean passed through zero
and `var/|mean|` blew up under the 1e-6 denominator floor) and
td_kurtosis = 50.6, pegging both downstream controllers at MAX
instead. Per_α at MAX over-concentrates PER sampling on outliers,
which hurts distributional Q learning (best l_q window regressed
from 2.41 → 2.69 between pdgxn and gxhr8).
## Fix: ISV-resident output clamp ceilings
Two new ISV slots hold the streaming-kernel output ceilings:
RL_ADV_VAR_RATIO_CLAMP_INDEX = 447 (default 100.0)
RL_TD_KURTOSIS_CLAMP_INDEX = 448 (default 30.0)
* 100.0 for var_ratio = 1000× ADV_VAR_RATIO_TARGET (= 0.1) — wide
enough that healthy signal (typical 1-10) passes through, tight
enough that 3e5 outliers don't peg rollout_steps.
* 30.0 for kurtosis = 3× (KURT_GAUSSIAN + KURT_LIFT_SCALE) — lets
the full per_α response range engage on heavy-tailed signal
(≤ 10), bounds runaway above that.
Per `feedback_isv_for_adaptive_bounds`: the clamps live in ISV
(visible in diag, modifiable at runtime via re-launching the init
kernel or a future adaptive controller) rather than as kernel-side
`#define`s.
## Seeding (no HtoD per feedback_no_htod_htoh_only_mapped_pinned)
New device kernel `rl_streaming_clamp_init.cu` — single thread,
writes both clamp ceilings directly to ISV. Launched once at the
end of `with_controllers_bootstrapped` alongside the 8 existing
controller-bootstrap launches. Zero host→device transfer.
## Diag bake-in (per user request "ensure to bake in diags")
JSONL gains a new `streaming` block exposing:
* `streaming.adv_var.{mean, m2, clamp}`
* `streaming.td_kurt.{mean, m2, m4, clamp}`
Cross-check: when consumer-input slot (RL_ADVANTAGE_VAR_RATIO_EMA_INDEX
or RL_TD_KURTOSIS_EMA_INDEX) reads exactly the same value as
`streaming.*.clamp`, the clamp fired this step.
## Test updates
G1 (isv_bootstrap) + G3 (r5_controllers) blanket-assert that
ISV[417..END] is sentinel-zero at bootstrap. Both new slots are
seeded to non-zero values by rl_streaming_clamp_init during
bootstrap, so both tests skip these slots in the loop and assert
the seeded values separately.
## Verified gates (local sm_86)
G1 isv_bootstrap ✅
G3 controllers ✅
G4 target_update ✅
integrated_smoke ✅
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
|
||
|
|
53aeef099b |
feat(rl): ISV-driven PPO importance-ratio clamp + log-ratio diagnostic
pt67l confirmed reward-scale + V-target clamp eliminate V regression
spikes — but exposed a residual: |l_pi| max=586 with mean 0.22. Root
cause: PPO's clip(r, 1-ε, 1+ε) bounds the loss only when surr2 is
the active min. The unclipped branch IS active when A<0,r>1+ε
(surr1=A·r is then more negative than surr2=A·(1+ε), so min selects
surr1) and when A>0,r<1-ε. In the first case `r` can blow up: we've
seen r reach 1e10 from policy drift over a multi-step rollout
producing l_pi=O(1e10) spikes that contaminate the loss-balance
controller and the LR controller's plateau detection.
## Fix: ISV-driven ratio clamp
Per `feedback_isv_for_adaptive_bounds` and
`pearl_controller_anchors_isv_driven`: the clamp ceiling lives in
ISV[RL_PPO_RATIO_CLAMP_MAX_INDEX = 440], not as a hardcoded #define.
New controller `rl_ppo_ratio_clamp_controller.cu`:
* Anchors on the (already KL-adaptive) PPO clip ε at ISV[402]
* target = (1 + ε) × PPO_CLAMP_MARGIN (MARGIN = 10.0)
* Wiener-α blend with floor 0.4 per
pearl_wiener_alpha_floor_for_nonstationary (ε is non-stationary)
* Permanent floor 2.0 / ceiling 1000 per
pearl_blend_formulas_must_have_permanent_floor
* Bootstrap 10.0, replace-directly on first non-bootstrap ε
observation per pearl_first_observation_bootstrap
When ε is small (rl_ppo_clip_controller seeing low KL → tight clip
band), the ratio clamp tightens — outliers should be rare anomalies.
When ε widens (large KL → wide clip band), the clamp widens
proportionally — outliers are expected so we permit more
magnitude before bounding.
## Wiring
ppo_clipped_surrogate_fwd and _bwd both read
isv[RL_PPO_RATIO_CLAMP_MAX_INDEX] and clamp ratio to
[1/ratio_max, ratio_max] before forming surr1/surr2. The clamp is
forward-only in effect (bwd gates pg_grad inside [1-ε, 1+ε] anyway
so gradients were already bounded), but bounding the FORWARD ratio
keeps l_pi sane for the controllers downstream.
The new controller is wired into both:
* `with_controllers_bootstrapped` — bootstrap launch alongside
the other 7 R1 controllers
* `launch_rl_controllers_per_step` — per-step refresh alongside
the other 7 R5 controllers
## Diagnostic: per-step max |log_ratio|
New kernel `ppo_log_ratio_abs_max_b.cu` (same tree-reduce shape as
rl_kl_approx_b) writes per-batch max(|log π_new − log π_old|) to
ISV[RL_PPO_LOG_RATIO_ABS_MAX_INDEX = 441]. Launched right after
rl_kl_approx_b (uses the same log_pi_old_d + pi_log_prob_d inputs).
Surfaces in diag JSONL as:
"ppo": {
"ratio_clamp_max": isv[440], # adaptive ceiling
"log_ratio_abs_max": isv[441] # per-step observed max
}
The clamp fires when log_ratio_abs_max > ln(ratio_clamp_max).
For ratio_clamp_max = 10, ln = 2.30. Healthy training has
log_ratio_abs_max well below this most steps; outliers touch or
exceed it on rare excursions which the clamp bounds before they
pollute l_pi.
## Slot allocation
RL_PPO_RATIO_CLAMP_MAX_INDEX = 440 (controller output)
RL_PPO_LOG_RATIO_ABS_MAX_INDEX = 441 (per-step diag)
RL_SLOTS_END = 442 (was 440)
## Test updates
G1 (isv_bootstrap) + G3 (r5_controllers) blanket-assert ISV[417..END]
== 0.0 to catch slot-wiring bugs. Slot 440 is now a controller
OUTPUT bootstrapped to 10.0, so both tests skip it in the loop and
assert == 10.0 separately.
## Verified gates (local sm_86)
G1 isv_bootstrap ✅ (with new slot-440 assertion)
G3 controllers ✅
G4 target_update ✅
G6 r7d_per_wiring ✅
integrated_smoke ✅
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
|
||
|
|
20c7852b66 |
fix(rl): asymmetric clamp on scaled reward + pre-clamp |max| diag
The xv66n smoke (commit
|
||
|
|
d5c29fb4fa |
fix(rl): warmup window in plateau-decay LR controller fixes V cold-start
`alpha-rl-rzltn` exposed a bug in the plateau-decay design: V head's
`best` got bootstrapped to 7.12e-10 (machine epsilon) at step 1
because V regression had no reward signal yet — no trade had closed,
the bootstrap V target was 0, so the first V loss was effectively 0.
Every subsequent V loss EMA was orders of magnitude higher (4.07
at step 100, 1.15 at step 1000), so the improvement check
`loss_ema < best * 0.99` evaluated false FOREVER. The controller
then decayed lr_v every 1000 steps purely on the patience clock,
not because the model genuinely plateaued.
Cross-check across the 50k-step rzltn run:
* V best unique values: {0.0, 7.12e-10} — ONLY 2 across 50000 rows
* V best max: 7.12e-10
* V best-improvements: 0 (Q: 12, π: 12)
* V decays still fired: 7 (one every 1000 steps from step 1001)
The plateau-decay mechanics worked correctly — the controller counted
to 999 then halved LR exactly as designed. The bug was that "first
observation defines best forever" is degenerate for sparse-signal
heads whose first loss is a cold-start artifact.
## Fix: LR_WARMUP_STEPS
Three new ISV slots (one per head — Q, π, V at 436/437/438) hold a
monotonic warmup counter clamped at LR_WARMUP_STEPS = 500. During
warmup the controller:
* always overwrites `best` with current loss_ema (tracks the EMA
as it converges)
* holds the plateau counter at 0 (no decay fires during warmup)
* increments warmup_counter
Once warmup_counter >= LR_WARMUP_STEPS, the controller switches to
standard plateau detection — `best` then locks in at the
post-warmup loss_ema value (representative of the head's converged
loss scale), and patience counting begins.
At α=0.05 the EMA half-life is ~14 steps; 500 updates leaves ~35
half-lives, well past convergence. This gives V time to see its
first actual losses after trades start closing.
## Slot allocation
RL_SLOTS_END: 436 → 439 (adds 3 warmup counter slots).
## Wiring
* rl_lr_controller.cu — adds warmup_slot param to
plateau_decay_head, kernel takes 12
slot ints (was 9)
* isv_slots.rs — 3 new constants, RL_SLOTS_END += 3
* integrated.rs — launch_rl_lr_controller passes 12
slot ints
* alpha_rl_train.rs — diag JSONL emits new
lr_plateau.{head}.warmup field
## Verified gates (local sm_86)
G1 isv_bootstrap ✅
G3 controllers ✅
G4 target_update ✅
G6 r7d_per_wiring ✅
integrated_smoke ✅
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
|
||
|
|
13d81dc5e6 |
diag(rl): emit grad_norm_ema + lr_plateau state in alpha_rl_train JSONL
Adds two new top-level keys to each diag.jsonl row:
"grad_norm_ema": {q, pi, v} — slots 424-426
"lr_plateau": {q,pi,v} × {loss_ema, best, stale} — slots 427-435
With these in place we can independently verify each plateau-decay
event in `mjgsj`'s diag (and all future runs):
* `loss_ema` traces the controller's slow EMA of head loss
(α=0.05); confirms the EMA actually moves and isn't stuck on the
bootstrap zero
* `best` shows the rolling minimum the controller compares against;
confirms it improves early then plateaus
* `stale` is the steps-since-best counter; should hit
PLATEAU_PATIENCE = 1000 exactly when an LR halving fires; reset to
0 after every decay event or every improvement
The `grad_norm_ema` block is kept because the grad-norm producers are
still wired (commit
|
||
|
|
87a22d12c9 |
feat(rl): walk-forward G8 eval phase + fold split (MVP, manual fan-out)
Adds the minimum-viable implementation of the R9 multi-fold G8 gate
per `pearl_single_window_oos_is_not_oos` ("a single window is NOT
out-of-sample"). The trainer can now:
1. Slice the MBP-10 file list into K equal-sized blocks
(`--n-folds K --fold-idx k`).
2. Train on blocks [0..=k] (passed to MultiHorizonLoader).
3. Run a separate eval phase of `--n-eval-steps` on block [k+1]
using a second loader instance.
4. Drain LobSim trade records gated by a pre-eval head checkpoint
so train-phase trades don't contaminate the eval summary.
5. Compute profit_factor + sharpe + drawdown via existing
`ml_backtesting::artifacts::compute_summary`.
6. Write `eval_summary.json` alongside `alpha_rl_train_summary.json`.
## Manual fan-out (this MVP)
The dispatcher (`scripts/argo-alpha-rl.sh`) gains three new flags
that thread through the Argo template into the CLI: `--fold-idx`,
`--n-folds`, `--n-eval-steps`. To run a 3-fold G8:
./scripts/argo-alpha-rl.sh --n-folds 3 --fold-idx 0 --n-eval-steps 200
./scripts/argo-alpha-rl.sh --n-folds 3 --fold-idx 1 --n-eval-steps 200
(With n_folds=3 the valid fold indices are 0 and 1 — the third block
is the eval window for fold 1. n_folds=K accepts fold_idx ∈ [0, K-2].)
Each submission produces one `eval_summary.json` at the resolved
output dir; the per-fold profit_factor is the value to aggregate.
Manual aggregation for now — automated DAG matrix fan-out + an
in-cluster aggregator pod is a follow-up commit. The aggregator
will mean ± SD the per-fold PFs and gate on `PF > 1.0`.
## What's NOT pure eval
The eval loop calls `step_with_lobsim` (same as train) — Adam steps,
PER updates, controller adaptations all still fire during eval. At
b_size=1 the per-step learning effect is small relative to the
train-phase-accumulated policy, so the eval PF approximates the
OOS performance of the train-end policy. A clean pure-eval mode
(forward + LobSim step only, no backward/Adam/PER) is a follow-up
architectural change; documented inline at the eval phase block.
## Default behaviour unchanged
`--n-folds=1` (default) skips the eval split entirely and uses all
files for training — identical to the prior single-window smoke.
The R9 prior smokes ran in this mode. Default `--fold-idx=0` and
`--n-eval-steps=0` keep prior smoke runs binary-compatible.
## Template + dispatcher changes
* `alpha-rl-template.yaml`: adds 3 new workflow parameters
(`fold-idx`, `n-folds`, `n-eval-steps`) and threads them into
the train container's `alpha_rl_train` invocation.
* `argo-alpha-rl.sh`: adds matching CLI flags with explicit
documentation of the multi-fold dispatching pattern.
## Verified gates
Local sm_86 build + dispatcher syntax clean. Tests unchanged
(the walk-forward path is exercised by cluster smokes, not unit
tests — the loader-slicing logic is straightforward index math).
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
|
||
|
|
ce1e13519b |
fix(rl): mapped-pinned for all R7d/R8 CPU↔GPU paths + diag JSONL + guard
Two concerns in one commit since they're entangled:
## 1. feedback_no_htod_htoh_only_mapped_pinned violations
R7d (PER push/sample) + R8 (CLI binary) + the new per-step diag dump
shipped with 8 raw `stream.memcpy_htod` / `stream.memcpy_dtoh` calls.
The rule is explicit: "mapped-pinned only for CPU↔GPU; tests not
exempt." A raw `stream.memcpy_*` on a regular `&[T]` / `&mut [T]` is
NOT mapped-pinned — the source/dest slice isn't page-locked, so the
CUDA driver does an internal blocking HtoD/DtoH that stalls the
stream.
Refactored all 8 violations to use the mapped-pinned + DtoD pattern
(cuMemHostAlloc DEVICEMAP — host writes via `host_ptr`, kernel reads
`dev_ptr`, DtoD between them via `cudarc::driver::result::memcpy_dtod_async`).
New shared helpers in `trainer/integrated.rs`:
* `read_slice_i32_d` — DtoH for `i32` device buffers via
`MappedI32Buffer` staging. Counterpart to the existing
`read_slice_d` (f32 version).
* `write_slice_f32_d` — CPU→GPU upload for `f32` via
`MappedF32Buffer.write_from_slice` + DtoD into destination.
* `write_slice_i32_d` — CPU→GPU upload for `i32` via
`MappedI32Buffer.host_slice_mut().copy_from_slice` + DtoD.
`pub fn` wrappers (`read_slice_*_d_pub`) expose the f32/i32 helpers
to the CLI binary so the per-step diag DtoH uses the same canonical
pattern.
Call-site refactors:
* `push_to_replay`: 4× `stream.memcpy_dtoh` → `read_slice_*_d`.
* `sample_and_gather`: 3× `stream.memcpy_htod` → `write_slice_*_d`.
* `step_with_lobsim` pre-PER ISV refresh: raw `memcpy_dtoh` →
`read_slice_d` (424 floats per step).
* `step_with_lobsim` post-Q PER priority TD readback: raw
`memcpy_dtoh` → `read_slice_d` (b_size floats per step).
* `step_synthetic` ISV mirror refresh: raw `memcpy_dtoh` →
`read_slice_d` (pre-existing pre-R9 violation; fixed in the
same commit since it's the same pattern in the same file).
* Init-time (one-shot) `prng_state` upload: raw `memcpy_htod` →
inline mapped-pinned DtoD (custom because cast through i32 for
the u32 buffer).
* Init-time (one-shot) `atom_supports` upload: raw `memcpy_htod`
→ `write_slice_f32_d`.
* `examples/alpha_rl_train.rs` per-step diag DtoH (3 calls) →
`read_slice_*_d_pub`.
## 2. Pre-commit guard gap — diff-aware HtoD/DtoH check
The existing GPU hot-path guard (`scripts/gpu-hotpath-guard.sh`)
EXPLICITLY skips memcpy_htod/dtoh on the assumption that such calls
only appear in `cuda_pipeline/` (where mapped-pinned is the
convention). That assumption was falsified by R7d/R8 — the guard
shipped 8 violations green.
Added `check_no_raw_htod_dtoh` to `scripts/pre-commit-hook.sh` (the
real file behind the `.git/hooks/pre-commit` symlink). The check is
DIFF-AWARE: it greps only the `+` lines of `git diff --cached -U0`,
so pre-existing violations elsewhere (143 sites across the codebase)
don't block commits touching unrelated files. NEW additions of
`\.memcpy_(htod|dtoh)\(` are flagged with a clear error pointing at
the mapped-pinned alternative. Suppress per-line with `// gpu-ok:
<reason>` (same convention as the existing guards).
Pre-existing violations in `ml-alpha/src/aux_heads.rs`,
`mamba2_block.rs`, `cfc/`, `data/`, etc. are a separate cleanup —
not blocked by this commit's check because the diff-aware filter
ignores anything that was already on `HEAD~1`.
## Verified gates (post-fix, local sm_86)
G1 isv_bootstrap ✅ unchanged
G3 controllers_emit ✅ unchanged
G4 target_soft_update ✅ unchanged
G6 r7d_per_wiring ✅ unchanged (PER round-trips all
mapped-pinned now)
R3 ema/advantage (3 tests) ✅ unchanged
R4 action kernels (3 tests) ✅ unchanged
end integrated_trainer_smoke ✅ unchanged
Mapped-pinned is semantically equivalent to raw memcpy_htod/dtoh —
just routed through page-locked staging so the driver doesn't have
to do its own internal pinning. Behaviour identical; the cost shifts
from "driver hidden HtoD per call" to "mapped-pinned alloc + DtoD
per call." For the smoke (b_size=1, 1000 steps), the cost difference
is in the microseconds.
## Per-step diag JSONL (separate concern, same commit)
Added `--diag-jsonl <PATH>` flag to `alpha_rl_train.rs` (default:
`<out>/diag.jsonl`). After each `step_with_lobsim`, writes one JSON
record capturing:
* step number, elapsed wall time
* all 5 head losses + λs
* all 7 RL controller outputs (γ τ ε coef n_roll per_α scale)
* all 5 per-head learning rates (lr_bce/q/pi/v/aux)
* all 7 EMA inputs the controllers consume
* replay buffer length
* per-step reward stats (sum, max, min, abs_max)
* per-step done count
* per-step action histogram (9 action classes)
Critical for cluster smoke debugging — the prior CLI only flushed
an `eprintln` progress line every N steps (default 100), making
in-flight controller drift / replay stagnation / reward explosion
invisible until they produced a NaN abort. The JSONL is line-
buffered + flushed every `log_every` steps so `tail -f` shows
progress live.
The stderr progress line is also beefed up to include γ / ε / per_α /
reward_scale / dones / rew_sum at each tick so a casual `argo logs`
inspection sees the controller behaviour without parsing JSONL.
## Why R9 cluster submission needs this
Without the diag dump, an R9 1000-step smoke is "blind" — only the
final summary tells us what happened. With the dump, post-hoc
analysis can answer:
* Did the controllers adapt or stay at bootstrap?
* Did the reward scale stabilise or saturate?
* Did the PER buffer fill?
* Was the action histogram dominated by any one action?
* Where did the per-head losses converge to?
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
|
||
|
|
1168f3ea83 |
feat(rl): R8 — alpha_rl_train CLI + Argo template + dispatcher
Closes the rebuild plan's R8 scope: production runner shape for the
integrated RL trainer. Three artifacts wired end-to-end:
1. `crates/ml-alpha/examples/alpha_rl_train.rs` — clap CLI driving
`IntegratedTrainer::step_with_lobsim` against MBP-10 windows
loaded via `MultiHorizonLoader::next_sequence_pair` (R2) for
true `(s_t, s_{t+1})` adjacency. Per `feedback_mbp10_mandatory`,
`--mbp10-data-dir` is required — no synthetic-data fallback in
the production path.
2. `infra/k8s/argo/alpha-rl-template.yaml` — WorkflowTemplate
mirroring alpha-perception's DAG (check-cache → ensure-binary →
train; warmup-gpu parallel). Binary cache slot is
`/data/bin/<sha>/alpha_rl_train` (distinct from `alpha_train`
so the two binaries coexist at the same SHA).
3. `scripts/argo-alpha-rl.sh` — dispatcher with three rebuild-plan
guards baked in.
## Dispatcher guards (per the rebuild plan's feedback list)
`feedback_default_to_l40s_pool` (2026-05-09): default `--gpu-pool`
is `ci-training-l40s` (sm_89). H100 (sm_90) is opt-in for production
scale-up only. Cubins must match the device, so the dispatcher
derives `cuda-compute-cap` from the pool name and threads it into
the workflow params.
`feedback_argo_template_must_apply` (2026-05-21 canonical incident):
`argo submit --from=wftmpl/<name>` reads the cluster CRD, NOT the
on-disk YAML; unknown `-p` parameters silently no-op without a prior
`kubectl apply`. Dispatcher applies the local template BEFORE every
submission (overrideable via `--skip-template-apply` for the rare
case where you've already applied manually).
`feedback_push_before_deploy` (2026-05-20 canonical incident): the
in-cluster `ensure-binary` pod fetches source from `origin/<branch>`,
NOT the local working tree. Submitting before `git push` deploys the
last-pushed SHA, which can lag local diff by N commits. Dispatcher
verifies `git rev-parse HEAD == git rev-parse origin/<branch>` and
hard-errors with the explicit push command otherwise. Bypass via
`--skip-push-check` (only when intentionally deploying a previously-
pushed SHA via `--sha`).
## CLI: gate G8 (NaN abort)
Per `feedback_stop_on_anomaly` + `feedback_kill_runs_on_anomaly_quickly`,
the CLI checks every per-head loss (l_bce / l_q / l_pi / l_v / l_aux
/ l_total) for finiteness after each `step_with_lobsim` call.
Non-finite at any step → write summary with `nan_abort_step` set →
`process::exit(2)`. R9's cluster smoke tail-watcher kills the
workflow on the non-zero exit code, satisfying gate G8 from the
rebuild plan.
## CLI: knobs that ARE on the CLI
Structural / boundary parameters only (per
`pearl_controller_anchors_isv_driven`: every adaptive knob lives in
ISV, not CLI flags):
* `--mbp10-data-dir / --predecoded-dir / --out` — I/O paths.
* `--n-steps` — wall-budget control (1000 R9 smoke / 50k+ prod).
* `--seq-len / --n-backtests / --per-capacity` — structural
sizing. seq_len threads into the loader's multi-resolution
`1:<seq_len>` config; n_backtests into both LobSimCuda and
PerceptionTrainerConfig.n_batch.
* `--seed` — reproducibility per
`pearl_scoped_init_seed_for_reproducibility` (forks deterministic
sub-seeds for dqn / ppo / per).
* `--instrument-mode` — MBP-10 filter (all / front-month / id=N).
* `--gpu-idx` — CUDA device selection.
What's NOT on the CLI: γ / τ / ε / entropy_coef / per_α / reward_scale
/ per-head LRs — all live in ISV[400..417] and are driven by R5's
controllers from EMA-tracked diagnostics. Per the rebuild plan
A1: "every adaptive bound is signal-driven, not tuned."
## Cluster smoke entry point
```bash
# R9 validation smoke (after pre-cluster local CUDA tests green).
./scripts/argo-alpha-rl.sh --n-steps 1000 --instrument-mode front-month
# Production scale-up (gated by R9's multi-fold pass).
./scripts/argo-alpha-rl.sh --n-steps 50000 --n-backtests 32 \
--per-capacity 100000 # GPU sum-tree R-future when capacity > 4096
```
## What's NOT in this commit
The R9 cluster smoke run itself is out of band — this commit ships
the entry points. R9 will execute the pre-cluster validation
checklist + first 1000-step smoke + multi-fold walk-forward G8 gate
per the rebuild plan §"Cluster smoke discipline".
The summary JSON's schema is intentionally narrow (final-step losses
+ replay len + completion state + NaN abort marker). R-future may
add per-epoch breakdowns + per-ISV-slot snapshots once the cluster
smoke tells us which diagnostics are actually load-bearing for kill
decisions.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
|
||
|
|
f68e0a1d0d |
revert(loader): multi-resolution default '1:32' (single-scale) after htpp6 falsification
The Phase 1 multi-resolution layout (10 raw + 10 agg@30 + 12 agg@100) regressed ALL horizons in alpha-perception-htpp6 (2026-05-22): - auc_h100: 0.681 -> 0.512 (-0.169) - auc_h300: 0.617 -> 0.506 (-0.111) - auc_h1000: 0.576 -> 0.526 (-0.050) Hypothesis falsified. Root cause: Mamba2+CfC SSM encoder already used all 32 raw ticks effectively via state-recurrence; replacing 22 raw ticks with arithmetic-mean aggregates destroyed within-window microstructure variance (the actual h100 signal) AND broke temporal continuity that the recurrence relies on. Δt Fourier encoder couldn't compensate. Architectural pearl: SSM/RNN/CfC + multi-resolution input is incompatible without separate-encoder-per-scale or explicit scale tokens. Transformer-style positional encoding tolerates scale-mixing; recurrent state updates assume consecutive positions. Reverts default to '1:32'. Adds explicit single_scale_32() constructor for callers (harness, tests). Keeps default_three_scale() in code with deprecation note for future sub-variant experiments. Production defaults across alpha_train CLI, Argo template, dispatcher script, ml-backtesting harness now match the proven baseline. |
||
|
|
7e438aeba0 |
feat(alpha_train): --multi-resolution CLI flag
Replaces --seq-len. Default '1:10,30:10,100:12' = 32 positions covering 1510 ticks of context, the Phase 1 fix for temporal receptive field mismatch. Logs the active config at preload start so Argo logs surface the per-run choice. |
||
|
|
78a9e08358 |
feat(loader): InstrumentFilter::FrontMonth for cross-quarter ES.FUT data
Replaces Option<u32> instrument_id_filter with InstrumentFilter enum {All,
Id(u32), FrontMonth}. FrontMonth runs a two-pass detect over the DBN
stream: pass 1 counts instrument_ids and collects SymbolMapping records,
picks the dominant id, validates it resolves to an ES contract via regex
ES[FGHJKMNQUVXZ]\d{1,2}; pass 2 streams the filtered records.
Motivated by alpha-perception-k54wd: a single-id filter on parent-symbol
ES.FUT data caught Q1 2024 (kept=73M) but kept=0 for Q2-Q9 because ES
front-month rolls quarterly (ESH4 -> ESM4 -> ESU4 -> ESZ4 ...). FrontMonth
self-tunes across the rolls without needing a per-file id table.
Sidecar keys distinguish modes: mbp10 / mbp10_instr<id> / mbp10_front_month.
CLI flag renamed --instrument-id -> --instrument-mode {all,id=N,front-month}
with matching parameter rename in argo-alpha-perception.sh + template.
|
||
|
|
783297e002 | feat(loader): instrument_id filter + outdated test fix + sp18 fingerprint | ||
|
|
f065cfbaa2 | diag(aux): log pos_fraction once at epoch=0, step=0 (diagnose pw bit-identity) | ||
|
|
0f5d5c7b4a |
feat(aux-supervision): wire BCE + conditional-Huber actual calls (CB5)
CB3+CB4 shipped the kernels with a holding-pattern (grad-zero) call site.
CB5 wires the real calls + per-loss ISV signals.
perception.rs:
- step_batched body: pos_weight computed host-side from
self.last_pos_fraction (n_neg/n_pos clamped to [1.0, 50.0] per E3),
uploaded mapped-pinned to stg_aux_pos_weight_{long,short}
- 4 kernel calls per step (slab mode, K*B*N_AUX_HORIZONS):
* aux_bce_loss_gpu × 2 (prof_long, prof_short) — class-weighted
* aux_huber_masked_loss_gpu × 2 (size_long, size_short) — NaN-mask
from CB1's y_size=NaN at y_prof=0 gives conditional-Huber for free
- Per-loss EMAs replace single aux_huber_ema:
* aux_prof_bce_ema_per_h (BCE EMA per horizon)
* aux_size_huber_ema_per_h (Huber EMA per horizon)
* aux_dir_acc_ema_per_h (unchanged)
- Stop-grad lift condition: aux_prof_bce_ema < 0.4 AND aux_dir_acc > 0.85
for ALL horizons (uses BCE not Huber per E3 — size Huber is
observability-only since the regression scale varies more than the
binary classification quality)
- aux_lift_huber_threshold renamed to aux_lift_prof_bce_threshold
alpha_train.rs:
- AlphaTrainSummary: final_aux_huber_ema_per_h split into
final_aux_prof_bce_ema_per_h + final_aux_size_huber_ema_per_h
- Per-epoch tracing: aux_prof_bce_h{100,300,1000} + aux_size_huber_h{...}
+ aux_dir_acc_h{...} + stop_grad_aux_to_encoder
perception_overfit.rs: synthetic test asserts BCE finite + below ln(2)
chance baseline, size Huber finite, dir_acc >= 0.5.
Synthetic test on RTX 3050:
aux_prof_bce_ema = [0.0142, 0.0142, 0.0142] (30× below threshold)
aux_size_huber_ema = [0.1213, 0.1213, 0.1213] (small)
aux_dir_acc_ema = [1.0, 1.0, 1.0] (perfect)
stop_grad lifted = false (lift fired)
Known limitation (follow-up CB6): both kernels return single joint scalar
over the [K × B × N_AUX_HORIZONS] slab — all per-horizon EMA entries
carry the broadcast joint mean. Per-h gating would require kernel split.
cargo check --workspace --all-targets clean.
cargo test -p ml-alpha --lib: 43 passed.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
9ed882e740 |
feat(aux-labels): replace D with A+B paired labels + loader/trainer migration (CB1+CB2)
Smoke 1 v2 empirically falsified the D-style labels: 99.99% of K=100 D-labels
were negative on real ES MBP-10 (cost+1.5×MaxDD dominates every typical
100-tick move). Aux head couldn't escape predicting the mean.
CB1 — replace label generator:
- generate_outcome_labels_d → generate_outcome_labels_ab returning
OutcomeLabelsAB { y_prof_{long,short}, y_size_{long,short}, sigma_k,
cost_price_units, pos_fraction }
- y_prof = 1 iff signed_pnl > 2×cost (binary, class-weighted BCE target)
- y_size = signed_pnl / σ_K (σ-normalized regression target,
conditional-masked when y_prof=0)
- σ_K: rolling 1000-bar Welford std of K-step log-returns, floored at
cost/4 per pearl_trade_level_vol_for_stop_distance
- pos_fraction: per-(direction, horizon) positive class fraction for
downstream BCE pos_weight balancing
- 10 unit tests validating sign-correctness, NaN edges, balance,
cost-threshold, sigma-floor, per-horizon independence, error paths
CB2 — atomic caller migration:
- LabeledSequence + LoadedFile: outcome_long/short (2 arrays) → 5 arrays
(prof_long, prof_short, size_long, size_short, sigma_k) + pos_fraction
- Loader splits new generator's outputs + propagates pos_fraction
file-level into each yielded LabeledSequence
- step / step_batched signatures widened to 7 params + pos_fraction
- alpha_train.rs: 4 separate batches + per-snapshot row build for each
- last_pos_fraction stashed on PerceptionTrainer
- aux test fixture (synthetic_aux_outcomes) updated to emit 4 arrays +
pos_fraction matching the constant-up-ramp test signal
- HOLDING PATTERN: step_batched body validates new arrays + stages
prof-binary through the existing D-era Huber kernel so training
produces a real gradient. CB3 rewrites the kernel; CB4 widens head
outputs; CB5 wires BCE+Huber proper loss with class weighting.
cargo check --workspace --all-targets: clean (only pre-existing cudarc
cupti + ml insert_batch unrelated errors).
cargo test -p ml-alpha --lib: 43 passed (15 multi_horizon_labels).
cargo test -p ml-backtesting --lib: 33 passed.
stacked_trainer_aux_supervision RTX 3050: pass (huber=0.0004, dir_acc=1.0,
stop_grad lifted).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
2e06297671 |
feat(aux-supervision): per-epoch aux observability in alpha_train (B6.5)
Smoke 1 at HEAD
|
||
|
|
21e7dfd63c |
feat(aux-supervision): wire AuxTrunk + AuxHeads + Huber loss into PerceptionTrainer (B5)
Wires the aux supervision path parallel to BCE: - AuxTrunk (64-hidden single-bucket CfC) consumes the same encoder output as the main BCE trunk - AuxHeads (linear regression long/short) maps aux_trunk output to per-(direction, horizon) predicted outcomes - AuxHuberLoss supervises against D-style labels from MultiHorizonLoader Backward path with asymmetric stop-grad at encoder boundary: - aux_trunk gets gradient signal into its OWN params at all times - aux_trunk's encoder-boundary gradient is INITIALLY blocked (stop_grad_aux_to_encoder = true) - Conditional lift per E3 design: if aux_huber_ema < 0.4 AND aux_dir_acc_ema > 0.85 within 200 steps, lift the stop-grad - When lifted, aux_vec_add kernel folds aux's grad_x into the main grad_h_enriched_seq slot (element-wise += per feedback_no_atomicadd) ISV signals added: aux_huber_per_h, aux_dir_acc_per_h (per pearl). Per-trunk scratch + reduced grad buffers (no Adam state sharing per pearl_adam_normalizes_loss_weights — opt_aux is its own Adam group). New helper kernel cuda/aux_vec_add.cu: position-local dst += src for the asymmetric stop-grad lift accumulation. New synthetic test stacked_trainer_aux_supervision_converges_on_constant_signal validates end-to-end: aux_huber_ema_per_h = [0.087, 0.087, 0.087] (converged) aux_dir_acc_ema_per_h = [1.0, 1.0, 1.0] (perfect on constant) stop_grad_aux_to_encoder = false (lift fired) All 5 stacked_trainer tests pass on RTX 3050 (lib still converges, no regression from parallel aux wiring). Not yet consumed by decision policy (B7) — aux output flows through training only. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
859d0c738b |
feat(aux-labels): wire D-style outcome labels into MultiHorizonLoader
Extends MultiHorizonLoaderConfig with `outcome_label_cost` (default DEFAULT_OUTCOME_LABEL_COST_ES = 0.5 = 2 ticks for ES round-trip). LabeledSequence + LoadedFile gain outcome_long + outcome_short arrays of [Vec<f32>; N_HORIZONS] populated by generate_outcome_labels_d during LoadedFile construction (same !inference_only branch as BCE labels). CLI exposure: alpha_train.rs gains --outcome-label-cost flag with the ES default. All 6 MultiHorizonLoaderConfig consumers updated atomically per feedback_no_partial_refactor: alpha_train (train + val loaders), in-file test fixtures (×2), multi_horizon_loader.rs (×2), harness.rs, trainer_parity.rs, ring3_replay.rs. cargo check --workspace clean; ml-alpha lib tests 41 passed. B3+ tasks not yet touched: outcome labels are computed and propagated to LabeledSequence but no consumer reads them yet (aux trunk Tasks B3-B5 will wire that). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
81e9970d4c |
refactor(per-horizon): N_HORIZONS 5→3 — alpha_train example
Renames all *_h6000 references to *_h1000 (new longest horizon): - AlphaTrainSummary struct fields: best_auc_h6000_* → best_auc_h1000_*, best_h6000_ckpt_path → best_h1000_ckpt_path - State vars: best_auc_h6000*, h6000_no_improvement → h1000 equivalents - Checkpoint filename: trunk_best_h6000.bin → trunk_best_h1000.bin - CLI early-stop metric: "auc_h6000" → "auc_h1000" - Tracing log fields: w_h30..w_h6000 → w_h10/w_h100/w_h1000 - 4 hardcoded [seq.labels[0..4][k]] literals → std::array::from_fn - MultiHorizonLoaderConfig.horizons + AlphaTrainSummary.horizons: literal [30,100,300,1000,6000] → HORIZONS constant - Doc comments: "5 horizons", "h6000-snapshot" → "N_HORIZONS", "longest-horizon" KNOWN FOLLOW-UP (Task 7): 5 sweep YAMLs reference trunk_best_h6000.bin filename — must update atomically when applying workflow template changes. KNOWN FOLLOW-UP (Task 8): crates/ml-alpha/src/gpu_log.rs:176-192 still decodes 5-horizon JSON fields (raw_h30..h6000, ema_h30..h6000, etc.). cargo check --example alpha_train: PASS. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
2f69bb1fc0 |
feat(per-horizon-cfc): wire Controller A + Phase 1→2 transition in trainer
Per spec §2.3 + §3.1 and Task 9 of the plan.
Adds two device kernels to bucket_transition_kernels.cu:
- tau_change_frobenius_kernel: ‖tau_t − tau_{t-1}‖_F² via block-tree
reduction (no atomicAdd per feedback_no_atomicadd)
- slack_factor_apply_kernel: (Q1, Q3) → (Q1/slack, Q3*slack) where
slack = sqrt(Q3/Q1), ISV-derived per feedback_isv_for_adaptive_bounds
Adds to PerceptionTrainer:
- TrainingPhase enum (Phase1Warmup / Phase2Routed)
- ControllerA state machine instance
- prev_tau_d device buffer + tau_change_d + tau_change_staged mapped-pinned
- bucket_warmup_cap_override CLI diagnostic field
- bucket_routing_metadata stored after transition fires
- heads_w_skip_compact_d (HIDDEN_DIM floats) populated by transition
- Cached function handles for the two new kernels
Per-step in step_batched (Phase 1 only):
- Launch tau_change_frobenius_kernel → mapped-pinned scalar shadow
- DtoD prev_tau ← tau_all_d for next step
- Sync + host scalar read
- controller_a.update(tau_change, bucket_warmup_cap_override)
- On trigger: stage tau into scratch buffer (avoids in/out aliasing in
tau_reorder_kernel), execute_transition writes routing metadata +
reorders tau_all_d + populates heads_w_skip_compact_d,
slack_factor_apply widens IQR bounds, invalidate CUDA Graph,
latch phase = Phase2Routed, log via tracing::info.
Phase 2 dispatch + Controllers B/C/D wiring deferred to Tasks 10–12.
In Task 9's transient state, the trainer enters Phase 2 but continues
Phase 1 dispatch path; the reordered tau_all_d still produces a valid
forward pass since CfC's per-channel decay math is order-invariant.
Per pearl_no_host_branches_in_captured_graph: transition fires OUTSIDE
the captured graph (cached graph invalidated → recaptured next step).
Per pearl_cudarc_disable_event_tracking_for_graph_capture: event tracking
is already disabled trainer-wide; recapture on next step is safe.
Per feedback_no_htod_htoh_only_mapped_pinned: host scalar read goes via
mapped-pinned DtoD shadow, not bulk DtoH.
CLI: --bucket-warmup-cap-steps added to alpha_train (Option<u64>,
diagnostic override of Controller A's ISV-derived cap).
Tests: 7 GPU oracle tests pass on RTX 3050 (5 existing + 2 new for
Frobenius + slack_factor). 33 ml-alpha lib tests pass. Workspace
cargo check clean modulo pre-existing cupti / sp15 / gpu_per_integration
errors unrelated to this task.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
|
||
|
|
f13d6c6dcf |
refactor(per-horizon-cfc): atomically remove MTER scaffolding
Per spec §2.5 and feedback_no_partial_refactor. Removes: - LoaderMode enum (single-variant after Sequential removal → deleted entirely) - next_sequence_sequential + sequential_file_idx + sequential_anchor_idx + last_call_was_file_boundary + is_file_boundary - last_seen_file_boundary, train_graph_boundary_state fields - notify_file_boundary method + boundary-state graph recapture logic - need_attn_pool_bootstrap match — always true now (random mode always bootstraps) - --loader-mode CLI flag + notify_file_boundary() call - loader-mode Argo template param Additional consumers migrated atomically (beyond the 4 files listed in the plan): ml-alpha/tests/perception_overfit.rs, ml-alpha/tests/multi_horizon_loader.rs, ml-backtesting/src/harness.rs, ml-backtesting/tests/trainer_parity.rs, ml-backtesting/tests/ring3_replay.rs — all referenced LoaderMode or the removed config fields. Workspace builds clean at this commit (pre-existing cudarc-cupti example and ml-crate test errors are unrelated to MTER removal — they fail at HEAD too). New per-horizon CfC arch lands in subsequent tasks of the same plan. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
7dd953aa2b |
feat(intervention-b): add LoaderMode::Sequential + attn_pool gate
Loader can now serve sequences in temporal order (Sequential mode); when active, the trainer skips attn_pool reset at non-file-boundary sequence boundaries (next commit will use this for stateful CfC + MTER). Random mode preserved as default; ZERO behavioral change for existing runs. Phased commit 1/8 of intervention B per docs/superpowers/specs/2026-05-21-crt-train-intervention-b-multi-timescale-readout.md |
||
|
|
2d5a66f6e4 |
feat(kernel-step-trace): runtime --kernel-step-trace CLI flag + JSONL drain
The compile-time feature kernel-step-trace gates inclusion of the ring
code. NEW: --kernel-step-trace <PATH> CLI flag on alpha_train gates
RUNTIME activation:
- Path provided -> ring allocated, drain spawns, JSONL records written
to <PATH> (truncated). One record per kernel emission.
- Path omitted -> ring not allocated, drain not spawned, zero overhead
even with the feature compiled in.
JSONL schema: {"step": u32, "kid": u8, "kname": str, "rt": u8,
"rt_name": str, "payload": {field: f32, ...}}. The payload field names
come from the existing per-(kid, rt) decoders in gpu_log.rs (ported to
serde_json::json! in decode_to_json).
Trainer ring fields are now Option<_>; populated only when the runtime
trace path is Some. Tick kernel, step-counter shadow, smoothness
controller pointer-passing, and Drop all become path-conditional.
The tracing::info!-based drain variant is removed (file-writer is
strictly more useful; tests migrated to assert JSONL file contents).
Argo template:
- New parameter kernel-step-trace-enable (build-time feature opt-in)
- New parameter kernel-step-trace-path (runtime CLI value)
- ensure-binary cache-busts when feature toggles
- training step passes --kernel-step-trace flag conditionally
Per feedback_no_feature_flags: compile-time gate retained because the
ring carries real memory cost (~2 MiB pinned) and per-step write overhead;
the specific name kernel-step-trace narrows scope to this mechanism.
Runtime gate is Option<PathBuf>, not a boolean enable_*.
|
||
|
|
60d0b62b8d | feat(crt-train): add --smoothness-base-lambda CLI flag + telemetry | ||
|
|
4a32186a5f |
fix(crt-train): migrate all PerceptionTrainerConfig literals to new smoothness_base_lambda field
Task 5 added smoothness_base_lambda to PerceptionTrainerConfig but only updated the struct definition + Default impl. Eight test sites in perception_overfit.rs and one site in alpha_train.rs construct the config via explicit field list (no ..Default::default() spread), so they broke with E0063 missing-field errors under cargo check --all-targets. Per feedback_no_partial_refactor.md: when a contract changes, every consumer migrates atomically. This commit adds smoothness_base_lambda: 0.0 to all 9 sites so the workspace compiles cleanly. The alpha_train.rs value of 0.0 is a placeholder — Task 7 will replace it with cli.smoothness_base_lambda once the CLI flag is added. |
||
|
|
045850e8f3 |
arch(crt-a): delete decision_stride field — greenfields atomic refactor
Per spec 2026-05-20-continuous-reasoning-trader-design.md §3.3 and §8:
decision_stride is REMOVED, not deprecated. No backwards-compat shim,
no fallback. Every consumer migrates in this commit per
feedback_no_partial_refactor.
Removed from:
- bin/fxt-backtest: RunArgs CLI flag, SweepBase field, SweepCell
override field, default_decision_stride() function, all three
BacktestHarnessConfig and RunArgs construction sites
- crates/ml-backtesting/src/harness.rs: BacktestHarnessConfig field,
MultiHorizonLoaderConfig decision_stride initializer, `let stride`
local, `if event_count % stride == 0` gate around
step_decision_with_latency; forward_step_into + step_decision now
share a single window-full guard (merged into one `if` block)
- crates/ml-alpha/src/data/loader.rs: MultiHorizonLoaderConfig field,
next_sequence stride logic simplified to stride=1 (consecutive
snapshots only)
- crates/ml-alpha/src/trainer/perception.rs: PerceptionTrainerConfig
field and Default impl; all four dt_s locals replaced with 1.0_f32
(training K-loop, graph-capture K-loop, forward_step_into CfC step,
eval K-loop)
- crates/ml-alpha/examples/alpha_train.rs: CLI flag, trainer_cfg and
both loader configs
- crates/ml/examples/alpha_baseline.rs: CLI flag, train + eval stride
gates replaced with unconditional read_all()
- config/ml/*.yaml: decision_stride: lines removed from
sweep_smoke, sweep_threshold_tuning, sweep_deployability,
sweep_decision_stride_example (file repurposed as generic example)
- tests: forward_step_golden, perception_overfit (×7 structs including
the stride=4 smoke repurposed as a second convergence check),
multi_horizon_loader (stride=4 spacing test repurposed as
ts_ns monotonicity check), ring3_replay, trainer_parity
Harness loop now invokes BOTH forward_step_into AND
step_decision_with_latency on every event whenever the snapshot window
is full. forward_step_into advances SSM state and writes alpha_probs_d;
step_decision_with_latency reads alpha_probs_d immediately after —
no CPU roundtrip, no stride gate.
n_decisions ≈ events_processed - seq_len + 1 after this commit
(vs ~9999 at stride=200 in the S2 baseline).
cargo check --workspace: clean
cargo test -p ml-backtesting --lib: 33 passed
cargo test -p ml-alpha --lib: 33 passed (6 ignored)
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
87b8303950 |
chore(ml-alpha): drop 'v2' naming — greenfield, no backward compat
Per project_ml_alpha_starting_capital greenfield posture: there is no V1 to differentiate from (the V1 trunk forward was dead code, no V1 checkpoint files exist in the wild). The 'v2' prefix on every identifier was historical baggage from the migration period. Renames: - CheckpointV2 -> Checkpoint (also drops the version: u32 field — bincode either deserialises a current envelope or errors; no migration path needed) - CheckpointVersionProbe removed (was only for V1 rejection) - LAYER_NORM_CUBIN_V2 / VARIABLE_SELECTION_CUBIN_V2 / ATTENTION_POOL_CUBIN_V2 -> LAYER_NORM_CUBIN / VARIABLE_SELECTION_CUBIN / ATTENTION_POOL_CUBIN - _ln_module_v2 / _vsn_module_v2 / _attn_module_v2 -> drop _v2 suffix - smoke_load_v2_checkpoint test -> smoke_load_checkpoint - config/ml/sweep_v2_*.yaml -> config/ml/sweep_*.yaml - migration-era 'V2 weight skeleton' / 'V2 fields' / etc. comments cleaned to remove the v2 prefix Pre-existing 'v2' references in ml-backtesting CUDA files (decision_policy.cu, pnl_track.cu) are NOT touched — those refer to future planned 'v2' refinements (Portfolio mode, multi-fill averaging) from the C1-C19 commits and reflect aspirational features unrelated to this session's trunk-grows work. Verification: ml-alpha + ml-backtesting + fxt-backtest all build clean. perception_forward_golden bit-exact (max_diff = 0.000000). |
||
|
|
7545651bce |
feat(ml-alpha): PerceptionTrainer.save_checkpoint + alpha_train wiring (X13+X14)
X13: Adds PerceptionTrainer::save_checkpoint as a thin delegate to self.trunk.save_checkpoint. Inference-only serialization — grads + AdamW state aren't included. X14: Inside the existing auc_h6000_improved block in alpha_train.rs, calls trainer.save_checkpoint(out_dir / 'trunk_best_h6000.bin') so the trained trunk lands alongside alpha_train_summary.json. Extends AlphaTrainSummary with best_h6000_ckpt_path (Option<String>) so downstream tooling (fxt-backtest --checkpoint) can locate the file without re-deriving the path. After this commit, every alpha-perception Argo workflow run produces a CheckpointV2 file at every new-best-h6000 epoch, ready for backtest consumption. Verification: - ml-alpha lib tests: 34 pass - alpha_train example builds clean (release) Per spec §1.1 (X13+X14). |
||
|
|
004b662c80 |
obs(ml-alpha): per-N-step train_loss log line for liveness + intra-epoch monitor
Adds a single tracing::info!("step") emitted every 500 training steps:
if epoch_train_steps % 500 == 0 {
tracing::info!(epoch, step = epoch_train_steps, loss = loss, "step");
}
Consumed by /tmp/alpha_monitor.py (v2: per-epoch trajectories + ISV +
liveness) to render intra-epoch loss trajectory and detect stalls
faster than the once-per-epoch granularity allowed.
At 8000 steps/epoch and ~36s/epoch on L40S → ~16 step lines per epoch
per fold, well below the cluster log volume budget.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
|
||
|
|
e9c4acfb12 |
feat(ml-alpha): MultiHorizonLoader inference-only mode
Add inference_only flag to MultiHorizonLoaderConfig that skips per-file forward-label precomputation (~half the file-load cost), plus peek_first() and next_inference_input() chronological-streaming methods for the ml-backtesting LOB harness. - min_size relaxed to cfg.seq_len when inference_only=true (training still requires seq_len + max_horizon + 1 for label generation) - New cursor fields (inference_file_idx, inference_snap_idx) walk every loaded snapshot in chronological order; reset() zeros both - peek_first() seeds CfcTrunk::capture_graph_a with cur==prev semantics (prev_ts_ns==ts_ns, trade_signed_vol=0) — natural stream-start - next_inference_input() errors if cfg.inference_only=false (guard against accidental mixing of training/inference paths) - All trainer call-sites (alpha_train example + multi_horizon_loader tests) updated with inference_only: false (zero behaviour change) - Inline test module exercises both modes; tests skip gracefully when fixture data isn't populated rather than panicking See docs/superpowers/specs/2026-05-18-real-lob-integration-design.md §1 (trainer parity) + §7 (orchestrator). Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> |
||
|
|
70d5fc29cf |
feat(ml-alpha): decision-stride loader + CLI + Mamba2 dt_s scaling (Phase 2A)
Decision-stride S lets a length-K sequence span ((K-1)*S + 1) raw snapshots instead of K consecutive ones — expands the effective time-window covered by each sequence at the same K-positions compute cost. With K=64 and S=4, the window covers 256 ticks (~5s on ES MBP-10 at 20ms-tick) instead of 64 ticks (~1.3s). Loader (crates/ml-alpha/src/data/loader.rs): - `MultiHorizonLoaderConfig.decision_stride: usize` (default 1, must pre-existing call sites add the new field). - `next_sequence` reads snapshot at `anchor + k * stride`; labels at the same indices (labels stay in absolute-snapshot horizons regardless of stride, e.g. h=6000 always means "predict 6000 raw snapshots forward"). - `prev` snapshot for microstructure features (prev_mid, prev_ts_ns) now points to the prior K-position (`anchor + (k-1)*stride`), NOT the consecutive-snapshot prior, so `Δt = ts_ns - prev_ts_ns` carries the actual elapsed time between K-positions (consumed by Mamba2's dt_s and the planned Phase 2C TGN Fourier features). - New `#[ignore]` real-data test: `loader_stride_4_yields_correct_spacing` asserts Δt monotonicity at stride=4. Mamba2 dt_s (crates/ml-alpha/src/trainer/perception.rs): - `PerceptionTrainerConfig.decision_stride: usize` plumbs the stride through. dispatch_train_step + evaluate_batched now use `dt_s = decision_stride as f32` so Mamba2's selective scan `exp(-dt * sigmoid(a))` reflects the real elapsed time. With stride=1 the behaviour is identical to before. CLI (crates/ml-alpha/examples/alpha_train.rs): - `--decision-stride <S>` flag (default 1) wired into both train and val loaders + PerceptionTrainerConfig. Argo workflow: - `decision-stride` parameter on the template (default "1") + `--decision-stride` script flag + propagation into the train pod's alpha_train invocation. Synthetic smoke (tests/perception_overfit.rs): - `stacked_trainer_loss_shrinks_with_stride_4` proves the trainer-level dt_s=4.0 keeps the Mamba2+LN+CfC+GRN chain numerically stable. Converges 0.32 → 0.0000 (matches stride=1 smoke trajectory — dt_s scaling didn't break the SSM dynamics). Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> |
||
|
|
00da163078 |
feat(ml-alpha): h6000-aligned ISV — uniform BCE + z-score lambda + K=64
Three correlated fixes addressing the architectural inconsistency
surfaced by the 3-fold ISV CV: we built a horizon-aware gradient
controller (ISV) but suppressed its target horizon (h6000) to 0.36%
of the loss via auto-horizon-weights, then used a ratio formula
that never approached its own clamp ceiling. ISV's lambda was
operating on rounding error.
(1) Uniform BCE weights as auto-default
trainer/perception.rs: `auto_horizon_weights` now returns
[1.0; 5] regardless of seq_len. Prior schedule `min(1, K/h)`
gave h6000 weight 0.0053 at K=32 — combined with lambda ~1.04,
h6000's effective loss contribution was ~0.37%, indistinguishable
from zero. With uniform weights, each horizon contributes 20% and
ISV's lambda actually has something to scale.
(2) Z-score lambda derivation
cuda/horizon_lambda.cu: replace `ratio = ema_h / mean(ema)` with
`z_h = (ema_h - mean) / std(ema); lambda = clamp(1.0 + 0.5*z, 1.0, 2.0)`.
Per `pearl_zscore_normalization_for_magnitude_asymmetric_signals.md`
z-score makes lambda spread scale-invariant of the absolute EMA
level. The ratio formula gave lambdas ≤ 1.04 in our data because
per-horizon BCE clusters tightly (range ~0.04) while mean is
~0.65. Z-score fills the [1.0, 2.0] envelope: 1σ → 1.5, 2σ →
ceiling. Boost-only asymmetric clamp preserved.
Test verification on the existing smoke (after 5 steps):
ema = [0.526, 0.522, 0.608, 0.641, 0.553]
lambda = [1.00, 1.00, 1.40, 1.76, 1.00]
Previously with ratio formula, max lambda on the same data
would have been ~1.05. h1000 (1.5σ above mean BCE here) now
gets a 76% trunk-gradient boost vs uniform.
(3) Default --seq-len 32 → 64
examples/alpha_train.rs: K=32 gives the model 0.5% of the
h6000 prediction window as in-window context. K=64 doubles
that, giving Mamba2's SSM state more material to build
long-horizon predictions. Within the kernel's MAMBA2_KERNEL_SEQ_MAX
cap of 96. Per-epoch wall scales ~K (more K-loop launches in
the captured graph, ~2× wall at K=64 vs K=32 for the K-loop
portion of dispatch).
Cache-bust v5 in build.rs to force nvcc recompile against the new
horizon_lambda.cu formula on the cluster's /cargo-target PVC. Old
cubins compute a numerically different lambda; running them against
the new Rust loop would silently apply the wrong gradient scaler.
Validation: 7 perception_overfit tests + 26 lib + 23 integration
ml-alpha tests pass. Synthetic overfit still converges to 0.0006.
horizon_ema_and_lambda_track_after_training observes the new
lambda spread (1.0-1.76) and asserts the asymmetric clamp envelope.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
|
||
|
|
a45fd85986 |
feat(ml-alpha): raise Mamba2 state cap 16→32 + add auc_h6000 early-stop
Three correlated changes for the next CV round: 1. Mamba2 state_dim cap: 16 → 32 cuda/mamba2_alpha_kernel.cu: MAMBA2_ALPHA_MAX_STATE_D 16 → 32. Per-thread state register `float x[32]` (128 B/thread) and per-thread x_hist replay cache `float x_hist[K*32]` (up to 12 KiB/thread of local memory at K=96). L40S/H100 register file (256 KiB/SM) absorbs this without occupancy collapse for our block dims (32-128 threads). Update Rust-side MAMBA2_KERNEL_STATE_MAX + validation message + test name. Kernel header doc updated. 2. New early-stop option: auc_h6000 examples/alpha_train.rs: add the long-horizon AUC as a third early-stop metric. The ISV CV ( |
||
|
|
d220755309 |
obs(alpha_train): per-epoch ISV controller snapshot (EMA + lambda)
Fold-2 of the asymmetric-ISV 3-fold CV regressed -2.0pt mean_auc
vs no-ISV (0.696 vs 0.716) while folds 0 and 1 gained +1.3pt and
+2.5pt respectively. To understand why the controller hurts that
specific regime, log the per-horizon BCE EMA and lambda multiplier
each epoch via the trainer's `loss_ema_snapshot()` /
`lambda_snapshot()` accessors (mapped-pinned reads; per-epoch
budget, not hot-path).
Single fold-2 re-run on `--cv-fold 2 --cv-n-folds 3
--cv-train-window 4` will surface:
- which horizon's BCE the controller flagged as hardest each epoch
- whether lambda saturated at the [1.0, 2.0] ceiling for one horizon
- whether the lambda trajectory has more variance vs the
stable-fold runs (suggests the regime drifts during training)
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
|
||
|
|
eb51c0f9cd |
feat(ml-alpha): walk-forward CV via file-list-driven loader
mhzs7 reported val mean_auc=0.726 on
|
||
|
|
3f089210f3 |
perf(ml-alpha): preload all MBP-10 files into RAM, reset loaders per epoch
cnjfl wall-time analysis: 80% of training time was disk IO. The MultiHorizonLoader was constructed fresh per epoch in the CLI loop, forcing 9 file deserializations from bincode (~50s/file × 9 = 7-8 min) per epoch. For a 6-epoch run that was ~45 min of pure IO out of ~50 min total. Now loaders preload ALL files into RAM at construction (one-time ~7 min startup) and expose a `reset(seed: u64)` method that re-seeds anchor sampling per epoch — no disk IO between epochs. Memory cost: ~13-15 GB for the 9-quarter ES.FUT dataset (45M snapshots × ~280 bytes). Well under the training pod's 64 GB limit; current 16 GB request remains sufficient since the resident set fits. Expected wall-time impact at K=96, B=8, 16K seqs: 6 epochs: ~50 min → ~13 min (~3.7×) 15 epochs: ~120 min → ~23 min (~5×) The internal LoadedFile-cache-with-cycling logic is gone — the loader now holds Vec<LoadedFile> with all files resident. Per-call `next_sequence` picks a uniformly random file + uniformly random anchor inside it (instead of cycling files with a per-file budget). Distribution is equivalent: each file contributes ~n_max_sequences / n_files samples per epoch in expectation. 77 ml-alpha tests pass. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> |
||
|
|
26d91a816c |
feat(alpha_train): configurable early-stop metric (default mean_auc)
cnjfl evidence: val_loss and mean_auc disagree.
val_loss best at e3 (0.5592)
mean_auc best at e4 (0.7670 — new h300 + h6000 peaks)
For downstream trading, ranking quality (AUC) matters more than
probability calibration (BCE loss). New default is mean_auc-based
early stopping, but val_loss/none remain selectable.
AUC is noisier than loss epoch-to-epoch (1-2pt bounces are common
even when long-horizon AUCs are still drifting up under
auto-horizon-weights), so patience defaults bump from 3 → 5.
CLI: --early-stop-metric {val_loss|mean_auc|none} default mean_auc
--early-stop-patience N default 5
Argo template parameters added with matching defaults.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
|
||
|
|
894188d34f |
feat(alpha_train): track best-by-mean-AUC alongside best-by-val_loss
cnjfl run showed val_loss and mean-AUC peak at different epochs: epoch 3: val_loss=0.5592 (best) mean_auc=0.7608 epoch 4: val_loss=0.5609 (worse) mean_auc=0.7670 (best — new h300 + h6000 peaks) val_loss tracks probability calibration; AUC tracks ranking quality. For downstream trading the ranking profile matters more — so we now publish both bests in alpha_train_summary.json and log a "new best mean_auc" line whenever a new mean-AUC peak lands. Early stopping still gates on val_loss (the two policies stay decoupled — mean-AUC is reported-only). New summary fields: best_mean_auc_epoch best_mean_auc best_mean_auc_per_horizon Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> |
||
|
|
c3ee5e165a |
feat(ml-alpha): wire batched kernels through trainer + CLI (#8)
Plumbs the batched cfc + heads kernels added in
|
||
|
|
85ce295773 |
feat(ml-alpha): per-horizon BCE weighting (fixes label-correlation inflation)
For seq_len K and horizon h with h ≫ K, the K position-supervised labels in a single sequence are near-identical (sequential positions' forward windows overlap by ~(h-1)/h). Per-position BCE therefore treats ~K highly-correlated labels as independent samples, inflating gradient pressure on long horizons by a factor of K. Concretely at K=96: h=30 → ~3 effective samples per seq (forward windows overlap ~97%) h=100 → ~1 (~99%) h=6000 → ~1 (~99.98%) Per-position supervision was paying 96× the natural signal density on h=6000, pulling the model toward fitting noise at long horizons. Fix: the fused BCE kernel now accepts an optional `loss_weights[N_HORIZONS]` (nullptr → uniform = no-op). Each (k, h) loss + grad contribution is multiplied by w_h; the normaliser is the sum of weighted valid entries instead of the raw valid count. `auto_horizon_weights(K, horizons)` computes `w_h = min(1.0, K/h)` so short horizons stay at full weight and long horizons collapse to their independent-sample density. Exposed via CLI: --auto-horizon-weights # K/h auto-derived --horizon-weights "1,1,0.5,0.1,0.02" # explicit floats Default behaviour is uniform (1.0) — apples-to-apples with the in-flight qf5mj baseline. Synthetic overfit still 0.6268 → 0.1144 in 250 steps (82% drop). 77 tests pass. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> |
||
|
|
248d8fe510 |
feat(ml-alpha): full architectural pass — recurrent CfC, GPU BCE, regime features, training discipline
Comprehensive fix for the issues identified after the BPTT-unroll cluster
run plateaued at val_loss ~0.692 with oscillating AUCs:
ARCHITECTURE
- CfC h_old is now RECURRENT across positions. Previously reset to
zero every step → CfC degenerated to a per-cell tanh-FC layer.
New: h_old at step k IS h_new at step k-1. Heads still operate
on h_new_k, but now the CfC actually carries state. Reverse-order
backward through the K positions accumulates grad_h_old → grad_h_new
via the new optional `grad_h_carry` arg on multi_horizon_heads_backward.
- tau is TRAINED. cfc_step_backward now writes grad_tau (per-cell decay
constant derivative), trainer gets a 7th AdamW group at 0.1× cfc lr.
- 6 NEW regime features (EMA cascade computed loader-side per file)
fill slots out[20..26] of snap_features. Gives the model multi-minute
trend / volatility / liquidity context that is structurally unreachable
inside the K-snapshot BPTT window. Slots: mid-z (med/slow), trend
signal, log-vol slow, log-spread med, log-trade-rate med. All bounded
via log1p / signed-log so no tuned constants leak in.
PERFORMANCE (NVIDIA-style)
- GPU-fused multi-horizon BCE for the entire [K, N_HORIZONS] grid in
ONE launch (was K host roundtrips). Native NaN-label masking.
- K-loop is fully GPU-resident: pre-allocated per-K scratch
(h_new_per_k, probs_per_k, labels_per_k, grad_probs_per_k), zero
device allocs inside step(). Only TWO syncs per sequence (after
forward, after backward) vs previously 2K+1.
- Stream-ordered kernel launches with pointer-offset addressing into
per-K buffers — host doesn't wait between K iterations.
- cfc_step_backward / multi_horizon_heads_backward both use += grad
semantics; trainer pre-zeroes accumulators once per step().
- MAMBA2_ALPHA_MAX_K capped at 96 (was temporarily at 256). 96 covers
h=30/100/300 with room; regime features handle h=1000/h=6000.
TRAINING DISCIPLINE
- LR schedule: linear warmup (default 200 steps) + cosine decay to
lr * lr_min_factor (default 0.1). Applied per training step to both
CfC and Mamba2 AdamW groups via new set_lr_cfc/set_lr_mamba2.
- Best-checkpoint tracking by val_loss; recorded in summary
(best_epoch, best_val_loss, best_val_auc).
- Early stopping on val_loss plateau (default patience = 3).
- CRITICAL BUG FIX: validation now uses new `evaluate()` method
(forward-only) instead of `step()`. Previous CLI called step()
on val data, which ran the full backward + AdamW update on the
validation set. With per-step BPTT that's ~K× more pressure than
the old comment ("statistically negligible") assumed.
Synthetic overfit: 0.6442 → 0.1233 in 250 steps (81% drop, sharper
than the previous 70%). 77 ml-alpha tests pass.
Local 2Q smoke (seq_len=64, 600 train seqs/epoch, 4 epochs):
val_loss 0.7011 → 0.6990, best epoch=1, h300 AUC 0.565 in epoch 0.
Phase E.3 callers (ml/examples/alpha_baseline.rs,
alpha_dqn_h600_smoke.rs) use the LEGACY Mamba2 forward_train +
backward_from_h_enriched path — unaffected by these changes (their
kernels are pre-zeroed via alloc_zeros, so the += grad semantics
remain correct in single-call mode).
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
|
||
|
|
16f5febf27 |
feat(ml-alpha): per-step supervision unrolls BPTT through full sequence
The final-step-only trainer (one BCE prediction per 32-snapshot
window) trained flat at chance on real ES data despite working on
synthetic overfit: train_loss=0.6953, val_loss=0.6943 across 40k
gradient steps. Gradient density was the bottleneck — one supervised
position per sequence × ~8K seqs/epoch isn't enough signal for the
SSM to find the alpha.
This commit supervises the model at EVERY position in the sequence:
mamba2_alpha_scan_fwd_seq — emits h_enriched at every t step
([N, K, sh2] instead of [N, sh2])
mamba2_alpha_scan_bwd_seq — accepts d_h_enriched_seq, injects
gradient at each t before propagating
d_state through the gate chain.
d_w_c and d_h_s2 accumulate across t.
PerceptionTrainer.step() — loop k=0..K; cfc + heads + BCE at
each valid label; cfc/heads grads
accumulate via += in kernel writes.
One Mamba2 backward call consumes the
full grad_h_enriched_seq.
cfc_step_backward — grad_w_in/w_rec/b writes changed
to += (callers MUST pre-zero).
multi_horizon_heads_backward — grad_w/grad_b writes changed to +=.
alpha_train.rs — passes per-position label rows to
step(); AUC still scored from
last-position predictions.
Phase E.3 callers (alpha_baseline.rs, alpha_dqn_h600_smoke.rs) use
the LEGACY Mamba2 forward_train + backward path with `alloc_zeros`
grad buffers — unaffected.
Synthetic overfit still converges 0.6664 → 0.1976 in 250 steps.
Local 2-quarter ES.FUT smoke shows the val AUC at h300 climbing
0.513 → 0.566 over 3 epochs (was flat-at-chance before). First
gradient signal we've gotten through the new architecture.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
|
||
|
|
ef35b10f3e |
refactor(ml-alpha): drop competitive-gate machinery — stacked is default
Per user direction "no gating, this is the new default": the stacked
Mamba2 -> CfC -> heads design is THE production architecture. There's
no competing-baseline comparison to run. Validation reduces to normal
training metrics (per-horizon val AUC, train loss curve, sanity floor
of >0.5 AUC).
Deletions:
- crates/ml-alpha/src/gate/cfc_vs_mamba2.rs (gate verdict logic)
- crates/ml-alpha/src/gate/mod.rs
- crates/ml-alpha/examples/alpha_gate.rs (gate runner binary)
Renames:
- crates/ml-alpha/src/gate/auc.rs -> crates/ml-alpha/src/eval/auc.rs
- lib.rs: pub mod gate -> pub mod eval (gate implied comparison;
eval doesn't)
Spec amendments:
- Drop the "Gate baseline strategy" amendment (committed earlier
this session)
- Reframe the stacked-architecture amendment as a "decision" not a
"gate"; production path is unambiguous
- Reframe Section 4 "Validation gate: CfC must meet Mamba2" -> just
"Validation: per-horizon val AUC" with the >0.5 sanity floor
Doc cleanups: stale "Mamba2 gate baseline" mentions in build.rs and
pinned_mem.rs replaced with neutral wording. The Argo template
comment about "downstream gate consumption" becomes "for monitoring".
Test status: all 26+ ml-alpha tests pass. AUC tests (6/6) still pass
under the eval:: namespace.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
|
||
|
|
1dc1f3563c |
refactor(ml-alpha): consolidate trainers — stacked is THE perception trainer
The merged Mamba2 -> CfC -> heads design from the 2026-05-16 spec
amendment supersedes the CfC-alone path. Removing the old CfC-only
PerceptionTrainer and renaming the stacked Mamba2CfcTrainer to
PerceptionTrainer (one trainer, clean naming).
Deletions:
- src/trainer/perception.rs (the OLD CfC-only trainer)
- tests/perception_overfit.rs (CfC-only smoke)
- tests/perception_debug_dump.rs (CfC-only trajectory print)
- tests/stacked_overfit.rs (replaced by perception_overfit pointing
at the renamed module)
Renames:
- src/trainer/stacked.rs -> src/trainer/perception.rs
- Mamba2CfcTrainer -> PerceptionTrainer
- Mamba2CfcTrainerConfig -> PerceptionTrainerConfig
- tests/stacked_overfit.rs content -> tests/perception_overfit.rs
CLI rewrite:
examples/alpha_train.rs now drives the stacked PerceptionTrainer.
Per-step inputs are sequences (Vec<Mbp10RawInput>) of length
seq_len; labels come from the LAST position of the window
(per-horizon). Flags: --seq-len, --mamba2-state-dim, --lr-cfc,
--lr-mamba2 (no more --n-hid since hidden_dim is fixed at 128 to
match Mamba2 and CfC by design).
Test status:
- 64 GPU tests pass on local sm_86 (31 lib unit + 33 integration)
- synthetic-overfit (rebranded perception_overfit): 250 steps,
initial=0.5951 -> final=0.1917 (68% drop, well above 40% gate)
- All bit-equiv / finite-diff / invariant tests still PASS
- One ignored test: the gate_artifact integration (waiting for
cluster-trained summary inputs)
The cluster gate (Task 18) now compares stacked-trained AUC vs a
Mamba2-baseline AUC (TBD: stacked vs a simpler "Mamba2 only" config
or an external reference baseline).
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
|