Three root-cause bugs made the collapse-recovery distillation mechanism
a silent no-op. Diagnosed via 50-epoch smoke test trajectory: Q-gap
peaked at 1.2 in epoch 2 then collapsed irreversibly by epoch 6, with
HEALTH_DIAG reporting `distill=off` throughout despite the trigger
conditions firing every epoch. Fixed each bug in turn and re-ran:
Q-gap now stabilizes above 0.18 from epoch 33 onwards, final 1.18.
Bug 1: timing
apply_distillation_gradient ran at epoch-boundary and SAXPY'd into
grad_buf. But the next training step's graph_forward.replay() starts
with `cuMemsetD32Async(grad_buf, 0, total_params)` — wiping the
contribution before any Adam update could see it. Moved SAXPY into
the per-step aux-op phase (between graph_forward and graph_adam),
matching the cadence of CQL/IQN/ensemble gradients.
Bug 2: CUDA Graph scalar baking
Moving SAXPY to per-step hit a deeper issue: graph capture bakes
kernel scalar args at capture time. The alpha value was captured
at 0.0 (initial) and never updated across replays, regardless of
per-epoch recomputation. Fix: new dqn_distill_saxpy_kernel reads
health from isv_signals[LEARNING_HEALTH_INDEX=12] directly and
computes alpha in-kernel. `distill_best_buf` is a stable device
pointer; its contents are DtoD-refreshed at epoch boundary when
maybe_snapshot_params accepts a new best. Zero CPU writes on any
path — pure GPU dataflow.
Bug 3: snapshot gate using wrong signal
The snapshot gate passed `self.last_q_gap` (an EMA that was stuck
at 0 due to broken propagation — see companion commit). Gate was
`health ≥ 0.65 OR winrate_fallback`, neither of which opened in
runs where high-q_gap epochs and high-winrate epochs don't overlap.
Replaced with q_gap-primary gate: `epoch_q_gap ≥ dynamic_floor`,
where the floor is `0.5 × decaying_peak` scaled by a per-epoch 0.99
decay. Adapts to network size automatically (production peaks at
~0.5 → floor 0.25; smoke test peaks at ~0.05 → floor 0.025).
Also drops the winrate-fallback "inflate health to 0.75" hack from
training_loop — q_gap is the direct measure of what distillation
preserves, no proxies needed.
Verified: local E1 smoke test (RTX 3050 Ti, 50 epochs) shows
distillation engaging from epoch 2 onwards and keeping Q-gap above
0.18 for epochs 33-50. Production L40S 50-epoch run pending deploy.
Files changed:
- dqn_utility_kernels.cu: new dqn_distill_saxpy_kernel (numerically
unchanged from saxpy_f32_kernel; alpha computed from ISV per-thread)
- gpu_dqn_trainer.rs: distill_saxpy_aux kernel handle, distill_best_buf
stable device buffer initialized from params at construction,
apply_distillation_gradient() rewritten, mirror_best_snapshot_to_distill_buf()
invoked on snapshot acceptance, maybe_snapshot_params uses dynamic
q_gap floor via SnapshotRing::observe_q_gap/dynamic_q_gap_floor
- q_snapshot.rs: SnapshotRing grows max_q_gap_observed decaying-peak
tracker, observe_q_gap() + dynamic_q_gap_floor() helpers,
MIN_SNAPSHOT_Q_GAP constant dropped in favor of relative floor,
module docstring rewritten to explain q_gap-primary gate
- fused_training.rs: set_distill_alpha + distill_alpha_per_step field
dropped (no longer needed — kernel reads ISV directly),
submit_aux_ops calls apply_distillation_gradient() unconditionally
- training_loop.rs: snapshot call simplified — passes epoch_q_gap
(raw, not the stuck EMA), drops winrate fallback and distill alpha
plumbing; also calls fused.update_eval_v_range() in the epoch-end
Q-stats block (the path previously writing to per_branch_q_gap_ema
was disabled via `if false` guard, leaving the health EMA frozen
at zero)
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
DQNConfig.enable_q_value_clipping was true at all 4 construction sites;
the else branch returned the pre-clamp tensor unchanged (dead identity
path). Q-value clipping is a BUG #37 fix — clipping happens after
forward pass so gradients still flow, and prevents Q-value explosions
from poisoning action selection. Per feedback_no_feature_flags.md, the
flag is gone — clipping is now unconditional. Kept q_value_clip_min /
q_value_clip_max since they're still passed to clamp().
QNetworkConfig had three write-only fields with zero workspace readers:
use_spectral_norm, spectral_norm_iterations, use_residual. Defaults set
at the single Default impl, never checked in forward/backward. Deleted
all three + their Default init. use_gpu kept — line 138 gates CUDA
construction and errors without it (genuine sanity check).
RainbowNetworkConfig.use_spectral_norm + spectral_norm_iterations were
declared but never read anywhere in the workspace. Defaulted false, set
false at the sole construction site. Pure write-only pair. Deleted both
fields + 2 construction sites. Serde default (no deny_unknown_fields)
keeps old JSON configs deserializable.
Tuning pass on the adaptive mechanisms. Changes:
1. F5 barrier weight raised 0.05 → 0.20 base, amplified 1×..2× by meta-Q
collapse prediction (proactive, not reactive). Old 0.05 couldn't escape
the Q-uniform attractor locally.
2. last_meta_q_pred field added on GpuDqnTrainer with set/get accessors,
wired from DQNTrainer's meta_q.predict() each epoch boundary. Aux-op
kernels now have per-step access to temporal collapse prediction.
3. DISTILL_HEALTH_THRESHOLD raised 0.4 → 0.55 (fire earlier). Additional
temporal trigger: distill also fires when meta_q_pred > 0.5.
4. SNAPSHOT_HEALTH_THRESHOLD lowered 0.7 → 0.65.
5. Snapshot-on-winrate fallback: when last_epoch_win_rate >= 0.45, inflate
effective health to 0.75 so a snapshot IS taken even if the health EMA
is stuck in the 0.48 trough. Without this, the good moments (WinRate 56%,
49%) are never captured → distillation has nothing to pull toward.
Result on local E1: distill=on every epoch (was permanently off), D6
fires only on bad-outcome epochs (was firing on convergence). Q-gap
still collapsed — tuning alone won't fix the underlying attractor;
root cause investigation next.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
DQNConfig.use_soft_updates was true at all 4 construction sites; the
else branch (legacy hard update) was unreachable. Deleted field + 4
setters + collapsed update_target_networks() to the unconditional
cosine-annealed EMA path. Hard-copy mode removed per
feedback_no_feature_flags.md.
DQNConfig.use_regime_conditioning was declared, defaulted true at 3
construction sites, but never read anywhere in the workspace. Classic
write-only flag (like use_different_seeds in DEAD-001). Deleted field +
docstring + 3 assignment sites. Regime conditioning is already
unconditional per the trainer — the flag was vestigial.
Previous D6 logic fired plasticity on any ensemble_collapse_score > 0.8,
which destroyed models at the exact moment they converged to a consistent
policy (low per-branch Q-gap variance means either collapse OR convergence —
the signal alone cannot distinguish them). Result: WinRate=46% epoch killed
by plasticity → WinRate=0.9% next epoch.
Added temporal outcome gate: require last_epoch_win_rate < 0.45 OR
last_meta_q_pred > 0.5 alongside the ensemble-collapse signal. Local
smoke-test run shows D6 correctly fires on WinRate 0% / 0.2% epochs but
spares WinRate 49% / 46% epochs.
The underlying Q-collapse attractor still wins on short (10-epoch) runs,
but the adaptive system is no longer destroying the rare good epochs.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
LoggingConfig.enable_network_diagnostics was declared and defaulted but
never read anywhere (only test round-trip assertions existed). Deleted
field + Default init + two test assertions that validated the round-trip.
Zero production callers per rg.
Three sites compared f32 q_gap values via partial_cmp().unwrap() — any
NaN q_gap would panic the snapshot swap-out / best-pointer paths on the
training hot path. Replaced with partial_cmp().unwrap_or(Ordering::Equal)
and tightened the outer Option unwrap to expect() with the invariant text
(len >= MAX_SNAPSHOTS guarantees a worst element exists).
F6 added contrarian_active as the final parameter to experience_action_select
but gpu_backtest_evaluator.rs's launch site at line 980-1002 wasn't updated,
so CUDA read garbage for the new int arg and crashed the eval forward pass.
Added .arg(&0i32) as the final arg in the backtest launch — contrarian is
always off during deterministic backtest evaluation.
Verified locally via test_adaptive_learning_no_collapse: 10 epochs + 10
validation backtests now complete without segfault.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
The field was gated by validate() to always be true (absolute-dollar mode
caused Q-value explosion ±50,000). All call sites already set true. Deleted:
- RewardConfig.use_percentage_pnl field + Default init
- Absolute-dollar branch in calculate_pnl_reward (dead)
- Validation guard against false
- RewardConfigBuilder field + builder method + build() init
- Constructor call site
Percentage-based P&L is now unconditional per feedback_no_feature_flags.md.
EnsembleConfig.use_different_seeds was set but never read anywhere in
the workspace. rg confirmed zero readers. Field removed from struct,
Default impl, and new() constructor (3 sites).
Root cause of L40S segfault (train-tgdlq workflow, commit 0d639dfe9):
1. F4 (IB) and F5 (barrier) gradient kernels indexed cql_d_adv_logits and
on_b_logits_buf with stride b0_size (4) instead of total_actions (13).
The buffer is sized [B, total_actions, num_atoms]. Writes for batch i>0
landed in other actions'/batches' gradient slots — corrupting CQL gradient
for every batch beyond the first. After cuBLAS backward, weights diverged
in undefined ways; validation forward then segfaulted reading NaN-laced
parameters in GpuBacktestEvaluator init.
Fix: add `total_actions` kernel parameter (int), use it as the full stride
for adv_row and d_adv_a pointer arithmetic; keep b0_size as the loop
bound (direction-branch-only update). All three launch sites updated:
inlined F5 launch in apply_cql_gradient, inlined F4 launch alongside,
and the standalone inject_barrier_into_cql_d_logits method.
2. D6 ensemble oracle fired at epoch 0 because per-branch Q-gaps are all 0
at random init (range = 0 → score = 1.0 → plasticity trigger).
Shrink-and-perturb ran immediately, then again next epoch, etc. Gate
behind `learning_health.epoch > 5` (3 warmup + 2 buffer epochs for Q to
move) so the oracle only fires on post-warmup real collapse, not
untrained networks.
Both bugs are regressions from today's work — F4/F5 introduced yesterday,
D6 became live after the ens_disagreement real-signal fix in d9d35b6fa.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Add GpuDqnTrainer::apply_c51_budget_scale() which runs scale_f32_ungraphed
on grad_buf immediately after submit_forward_ops_main (C51 backward) and
before CQL/IQN SAXPYs in submit_aux_ops. Final master gradient is now a
weighted sum: c51_budget×C51 + cql_budget×CQL + iqn_budget×IQN.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Adds ib_gradient_direction CUDA kernel that fires when population variance
of Q(a) across b0 actions falls below min_var=0.01, pushing each Q(a) away
from mean_q. Piggybacked on cql_d_adv_logits / cql_d_value_logits (same
atomicAdd path as F5 barrier). ib_weight = 0.05 × (1−health) → zero-op when
network is healthy (health≈1). Wired inside apply_cql_gradient immediately
after the F5 barrier launch, flows through the existing CQL SAXPY + backward.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
LearningHealth::new() now uses α=0.1 (the production-appropriate ~10-sample
EMA window for 80-100 epoch runs). Tests that need to reach threshold
assertions inside a 5-iteration loop (warmup=3 + 2 EMA steps) use
LearningHealth::with_alpha(0.3) instead of bending production behaviour
to satisfy test convenience.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Add `barrier_gradient_direction` CUDA kernel to c51_loss_kernel.cu that
computes barrier = max(0, 0.05*health - q_gap) from the direction branch
logits and ISV[12], then injects gradient via atomicAdd into the CQL
d-logit accumulator buffers. When barrier > 0, it raises Q(argmax) and
lowers Q(second_max) to widen the direction Q-gap.
Wire-up: `apply_cql_gradient` now accepts `barrier_weight` and inlines
the kernel launch after the CQL kernel but before `backward_full`, so
both CQL and barrier share one cuBLAS backward pass with no extra SAXPY.
`submit_aux_ops` passes `barrier_weight = 0.05` every step (kernel is
internally a no-op when q_gap >= min_req). D2/N2 epoch-boundary block
updated to reflect that gradient is now live.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Replace single-sample max/min proxy with mathematically proper sigma_1/sigma_2
ratio. Maintains a host-side VecDeque of up to 64 Q-value samples; once ≥ 8
samples are available computes the n_cols×n_cols Gram matrix X^T X, then
extracts the two largest singular values via power iteration + rank-one
deflation. Falls back to the coarse max/min ratio until the buffer fills.
compute_q_spectral_gap changed to &mut self; callers updated accordingly.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
1. D6 ensemble oracle no longer dormant — replaces hardcoded ens_disagreement=0.1
with range of per-branch Q-gap EMAs (max - min over 4 branches). This gives a
real signal that collapses to 0 when branches agree on uniform Q and expands
to non-trivial values when branches preserve diverse action differentiation.
D6's smoothstep(0.01, 0.1) window now actually moves.
2. HEALTH_DIAG defaults aligned with constructor inits — cql_budget default
0.10→0.00, c51_budget 0.45→0.55. Only visible when fused_ctx is None
(pre-first-step); eliminates the cosmetic log jump on the first real epoch.
3. reward_v8 smoke-test module changed from pub mod → mod, resolving the
pre-existing unreachable_pub warning. No external consumers of the module.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Trains a 10-epoch DQN run on fxcache data and asserts q_gap_ema > 0.05 and
learning_health.value > 0.3 — regression guard against the Q-uniform collapse
attractor the entire adaptive-learning-dynamics spec is designed to prevent.
Run via:
FOXHUNT_TEST_DATA=test_data/futures-baseline \\
cargo test -p ml --lib -- test_adaptive_learning_no_collapse --ignored --nocapture
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Addresses reviewer's IMPORTANT flag: the conviction Q-gap was not sign-flipped when
contrarian_active, so downstream env_step would size contrarian (argmin) positions
with maximum conviction, amplifying losses. During contrarian override, set
out_q_gaps=0 so those trades size minimally — the override is deliberate and
temporary; we don't want to compound it with aggressive position sizing.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Add `contrarian_active` int parameter to `experience_action_select` kernel.
When non-zero, a `q_sign = -1.0f` multiplier is applied to Q values across
all 4 branches (direction/magnitude/order/urgency) before Boltzmann softmax,
converting argmax-favoring sampling to argmin-favoring without touching
temperature, epsilon, conviction filter, masking, or sampling logic.
When zero, q_sign = +1.0f — behavior is bit-identical to before.
Wire: GpuExperienceCollector gains `contrarian_active_cache: u8` field plus
`set_contrarian_active()` / `contrarian_active()` accessors. The flag is
appended as the last arg at the action_select launch site. training_loop.rs
propagates `self.contrarian_active` to the collector immediately after the
D7 Part A state machine updates.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Part A fully implemented: state machine tracks low_winrate_count across epochs,
activates contrarian_active for 2 epochs when WinRate<40% × 5 consecutive AND
health<0.3, deactivates on WinRate>=45% recovery. last_epoch_win_rate carries
financials.win_rate (f32 cast) from log_epoch_metrics_and_financials into next
process_epoch_boundary. HEALTH_DIAG already consumes last_contrarian_active as
contrarian=on/off. Part B (kernel argmin flip) deferred: direction selection uses
Boltzmann softmax, not simple argmax, making inversion non-trivial.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Replace hardcoded 0.5f flip probability in experience_env_step kernel with
dynamic cf_ratio = clamp(0.5 + 0.3*(1-health), 0.0, 1.0). Healthy (h=1)
keeps cf_ratio=0.5; collapsed (h=0) raises to 0.8 for more counterfactual
exposure to break collapse. Propagated via set_learning_health() each epoch.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Track consecutive epochs with health < 0.3 in `unhealthy_epoch_count`; after 3
consecutive unhealthy epochs, trigger shrink_and_perturb(alpha, sigma) and reset
the counter. Remove the periodic interval-based trigger (every N epochs). The
Phase 3 boundary trigger is intentionally kept as an independent mechanism.
last_plasticity_ready: Some(true) = accumulating, Some(false) = just triggered.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Compute barrier_loss = 0.5 × max(0, 0.05×health − q_gap)² on the host
from cached q_gap_ema. Written to last_barrier_loss (already declared by
A4) and surfaced in HEALTH_DIAG `barrier=...`. Scalar-only / no gradient.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
When learning_health < 0.8, the PER priority update switches from the
standard per_update_pa kernel to pow_alpha_diverse_f32, which multiplies
each priority by (1 + 2*(1-health)*|action - mean_action|). This rescues
the replay buffer's diversity signal during Q-collapse by surfacing
experiences whose action deviates from the batch mean.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Replace static cql_alpha with cql_alpha_eff computed from ISV signals:
health (ISV[12]) × (1 − regime_stability (ISV[11])) × base.
Collapse→0 (CQL off); volatile+healthy→full; stable+healthy→0.
Adds last_cql_alpha_eff field to GpuDqnTrainer (f32, init 0.0),
accessor on FusedTrainingCtx, and propagation into DQNTrainer for
HEALTH_DIAG logging.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Adds 15 new last_* fields (all None) to DQNTrainer for B/C/D tasks to populate,
and replaces the A3 HEALTH_DIAG log line with the full format covering components,
effective hyperparams, and novel mechanism states.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>