The grad_norm_finalize kernel writes both f32 (sum-of-squares) and bf16
(L2 norm). The IQN and ensemble trunk gradient computations were passing
the MAIN grad_norm_buf for the bf16 output, overwriting it mid-pipeline.
While compute_grad_norm_outside_graph() overwrites it again before Adam,
this was still a correctness hazard. Now uses a dedicated 1-element bf16
scratch buffer (aux_norm_bf16_scratch) for IQN/ensemble finalize calls.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
compute_grad_norm_outside_graph() called launch_grad_norm() which uses
the graph-captured CUfunction handle. After graph_mega capture, this
handle cannot be launched outside the graph — CUDA silently corrupts
kernel state on H100 (SM_90). The corrupted grad_norm_f32_buf caused
Adam to read bogus norms, zeroing all gradient updates.
Fix: use launch_grad_norm_standalone() which has its own separately
loaded kernel handle, safe for post-graph launches.
This was the ROOT CAUSE of the H100 gradient collapse at epoch 3-4.
The collapse happened after mega-graph capture (step 2 of each epoch),
and the corrupted norm persisted for all subsequent steps.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- iqn_lambda restored to 0.25 (was 0.0 for NaN isolation)
- STEP_DIAG per-step logging removed (was temporary)
- FOXHUNT_GRAD_DIAG env var removed from Argo template
IQN NaN root cause fixed in 8cbabcef5 (f32-as-bf16 type mismatch).
Ready for H100 baseline deployment.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
The IQN backward kernel writes f32 via atomicAdd into d_h_s2_buf
(allocated as CudaSlice<f32>). But apply_iqn_trunk_gradient called
bf16_to_f32_kernel on this buffer, reinterpreting f32 bits as bf16.
This produced garbage/NaN that poisoned the trunk gradient on H100.
Replace with direct f32 DtoD copy — no cast needed.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Graph recapture ruled out as NaN cause (NaN persists without graph
invalidation). Now testing: does NaN disappear when IQN is disabled?
If yes → IQN backward/Adam produces NaN on H100.
If no → NaN is in the main C51/MSE pipeline.
Also reverts graph invalidation hack from previous diagnostic commit.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
TEMPORARY: graphs keep stale alpha (wrong training, but isolates NaN source).
If NaN disappears → recapture path is the bug.
If NaN persists → something else causes it.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
The f32→bf16 cast in cast_d_logits_to_bf16() reads
b*(b0+b1+b2+b3)*na + 32*4 elements, but d_adv_logits_buf and
d_adv_logits_bf16 were allocated with only +32*3 padding (96 vs 128).
The 32 extra f32 reads went past the allocation into adjacent GPU
memory. On H100 with 80GB and different memory layout, this grabbed
NaN/garbage values. The NaN propagated through the cuBLAS backward
pass into grad_buf, killing training at step ~720 (non-deterministic
depending on what lives in adjacent memory).
Fix: allocate +32*4 padding in d_adv_logits_buf, d_adv_logits_bf16,
and d_adv_logits_mse to match the cast kernel's access pattern.
Also reverts c51_alpha_max default to 1.0 — the gradient collapse
was from this buffer bug, not C51 premature convergence. The
c51_alpha_max parameter is kept for hyperopt exploration.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Logs grad_norm, loss, c51_alpha every 50 steps via existing async readback
scalars. No stream sync or extra GPU work. Works with mega-graph path.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Temporary diagnostic: logs per-stage gradient norms (pre-clip, post-clip,
post-IQN, final) to identify where H100 gradient collapse occurs.
Remove after root cause is found.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
H100 baseline (20 epochs) showed gradient collapse at epoch 10: C51
cross-entropy converges its distributional fit before the policy converges,
leaving zero gradient signal. The collapse happened 5 epochs after C51
reached alpha=1.0 (pure C51, zero MSE).
Fix: cap the C51 alpha ramp at c51_alpha_max (default 0.5) so MSE always
contributes (1 - alpha_max) of the primary gradient. MSE loss measures
Q-error directly and only goes to zero when Q-values are correct, not
just when the distribution shape is right.
- c51_alpha_max added to DQNHyperparameters (default 0.5)
- Added to PSO search space as 15th dimension (range [0.3, 0.9])
- Added to TOML profiles: smoketest, production, hyperopt
- Training loop caps alpha ramp at alpha_max instead of 1.0
- All 907 ml tests + 300 ml-core tests + 6 smoke tests pass
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Three root causes found and fixed:
1. SUM-reduced gradients without 1/N: all CUDA loss gradient kernels
(C51, MSE, IQN backward, CQL) accumulated per-sample gradients as
raw SUM. At batch=16384 (H100) the raw norm was 282x larger than
batch=58, causing gradient clipping to destroy signal-to-noise ratio
and collapse training at epoch 2-3. Now all kernels multiply by
1/batch_size, making gradient scale batch-invariant.
2. ExposureLevel::target_exposure() used a flat 9-level scale that did
not match the 4-branch dir*mag encoding. The Rust backtest evaluator
computed wrong position sizes (e.g. 4x oversize for Short+Small).
Now uses dir x mag formula. Also fixed is_buy/is_sell/is_hold and
from_trading_action for 4-branch semantics.
3. Rust epsilon-greedy only explored 5/9 exposure combos (0..5 instead
of dir*3+mag), ignored the magnitude branch on greedy, and used wrong
indices for order/urgency (get(1)/get(2) instead of get(2)/get(3)).
LR recalibrated: old gradient_clip_norm was accidentally a batch-size-
dependent LR reducer (~60x at batch=58, ~16000x at batch=16384). With
mean-reduced gradients the clip rarely fires, so LR is now the sole
training speed control. Smoketest 1e-4 -> 2e-6, production 1e-4 -> 1e-5,
hyperopt range [1e-5,3e-4] -> [1e-7,1e-4].
Diagnostics: FOXHUNT_GRAD_DIAG=1 enables per-stage gradient norm logging.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
RUN 4 (1540c0287) survived 9 epochs with stable grad_norm=0.44.
That commit INCLUDED reward normalization (÷pos_frac). The revert
in 09fb80baa removed it, and the resulting H100 run collapsed at
epoch 7 (grad_norm=0.000).
The reward normalization is part of the stable configuration.
experience_kernels.cu restored to exact RUN 4 state.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Both destabilized H100 training:
- Reward normalization (÷pos_frac): collapsed gradients at epoch 2-3
- Magnitude gradient restore (beta*MSE): Q-value explosion at epoch 25
(tested beta=0.10, 0.05, 0.02 — ALL cause Q explosion >50)
The magnitude MSE gradient is structurally unstable post-warmup.
ANY non-zero beta feeds a Q-overestimation loop that CQL can't counter.
The magnitude branch head learns during MSE warmup (epochs 1-5), then
the trunk continues improving via IQN (60% budget). This is the ONLY
stable configuration proven on H100 (RUN 4: stable through 9 epochs,
grad_norm=0.44, Q=2.03).
Reverts 76478559b (reward norm) and 22b7bc083 (gradient restore).
Keeps 1540c0287 (CUDA Graph fixes — the critical root cause).
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
The alpha blend (grad = alpha*C51 + (1-alpha)*MSE) progressively killed
magnitude gradient: at alpha=1.0 post-warmup, magnitude got 0*C51 + 0*MSE
= zero gradient. The magnitude branch head stopped learning at epoch 5.
H100 confirmed: grad_norm stabilized at 0.44 (epochs 6-9) — other
components alive but magnitude frozen. Q-values plateaued at 2.03.
Fix: after the alpha blend, SAXPY d_adv[d==1] += alpha * d_adv_mse[d==1]
to restore the MSE portion removed by the (1-alpha) scaling. This gives
magnitude full MSE gradient at all alpha values.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Two critical bugs causing gradient collapse at epoch 3-4 on H100
(batch=16384) while smoke test (batch=58) works:
1. c51_alpha baked as literal scalar in CUDA Graph at capture time.
set_c51_alpha() updated the CPU field but the graphed SAXPY blend
kernels kept using the frozen capture-time value (~0.01). The C51/MSE
gradient blend NEVER ramped — model trained with pure MSE forever.
Fix: invalidate graph_mega + graph_forward when alpha changes.
Graph re-captures on next step with the correct blend.
2. grad_norm_standalone loaded but never wired in. clip_grad_buf_inplace
used the CAPTURED grad_norm_kernel CUfunction outside graph context.
CUDA forbids launching a captured CUfunction outside its graph —
silently corrupts kernel state. grad_norm buffer contains garbage,
Adam reads garbage clip scale, gradients effectively zero.
Fix: use grad_norm_standalone (separate CUfunction handle) for
all non-graph gradient norm computations.
Root cause of every H100 gradient collapse this session.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Divide segment_return by position fraction (|pos|/max_pos) so the
reward measures directional skill per unit of risk, not absolute PnL.
Root cause: Quarter (0.25) produces 4× lower PnL variance than Full
(1.00). MSE Bellman target for Quarter has 16× lower variance →
gradient 16× more consistent → Quarter Q converges first → Boltzmann
locks in (82%) → model never tries Full enough to learn its true Q.
Same mechanism as C51 distributional collapse but through reward
variance instead of distributional shape.
After normalization, all magnitudes produce identical reward variance
for the same directional skill. The model selects magnitude based on
risk-adjusted return per unit position, not convergence speed.
Applied after Kelly stats (which need raw returns) and after reversal
override. NOT applied to backtest evaluator (measures actual PnL).
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Mean-advantage argmax + zero C51 gradient for direction (d==0) caused
all 9 per-action Q-values to converge to 0.85 by epoch 6 on H100.
Gradients collapsed to 0.000 at epoch 4. Root cause: mean-advantage
gives nearly equal values for all direction actions after centering,
causing degenerate argmax → same target action every step → all Q
converge to same Bellman target.
Direction does NOT need the magnitude treatment:
- Flat bias from C51 is smaller relative to genuine directional Q-gaps
- C51 provides essential distributional signal for risk-aware direction
- Boltzmann + 30% Flat floor already prevents Flat lock-in
Reverted to d==1 only (magnitude) for:
- C51 gradient zeroing
- Mean-advantage Bellman argmax
- Mean-logit in compute_expected_q
- 2× MSE amplification
Direction (d==0) uses full C51 softmax everywhere.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Mean-logit for online_eq/target_eq caused Q-value stagnation on H100
(5.7M bars): Q froze at 1.8 after epoch 3, action distribution identical
every epoch, Sharpe drifted +1.12→-0.36 over 11 epochs.
Root cause: mean-logit gradient is uniform across 51 atoms (1/51 each).
Too diffuse for continued learning — model converges to local minimum
in 3 epochs then can't fine-tune. C51 softmax focuses gradient on
high-probability atoms, enabling continued learning.
The distributional bias (Small/Flat learn faster) is countered by:
- C51 gradient zeroed for d<=1 (no cross-entropy pull)
- Bellman argmax uses mean-advantage (unbiased action selection)
- Boltzmann + Flat floor (diverse experience collection)
- compute_expected_q uses mean-logit (unbiased action selection/eval)
Softmax bias is in gradient EFFICIENCY (convergence speed), not in
the converged Q-VALUE (both sides of TD use same softmax → unbiased).
Result: Q-values 0.71→0.83 (progressing), Sharpe +7.41 at epoch 10
(was +1.98 with mean-logit), 7/7 diversity maintained.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Three fixes for production readiness:
1. Direction Flat floor: 30% unconditional Flat probability prevents
over-trading (was 7% Flat = 93% directional = massive cost drag).
Hard constraint, not Q-dependent, can't snowball. Combined with
hold enforcement (min_hold_bars=10), actual Flat is ~13%.
2. MSE loss online_eq + target_eq: consistent mean-logit for d<=1.
Both sides of TD error use the same representation — no mismatch.
Removes the last C51 softmax bias from magnitude gradient path.
Half grows 2.7%→5.7%, Full grows 2.7%→5.3% across epochs.
3. compute_expected_q: mean-logit for d<=1 (action selection + eval).
Safe — not in training loss path.
Result: 7/7 diversity sustained, Flat stable at 13%, Sharpe positive
at 3/4 epochs. 19/19 smoke tests pass.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Two remaining C51 softmax paths were causing Quarter (smallest position)
to dominate at 94%:
1. MSE loss online_eq: C51 softmax biased towards Quarter via gradient
d(td²)/d(weights) flowing through d(softmax_expected_Q)/d(weights).
Fix: mean-logit for d<=1, matching target_eq representation.
2. compute_expected_q: C51 softmax Q-values fed Boltzmann action
selection, inflating Quarter's Q. Fix: mean-logit for d<=1.
Safe — this kernel is NOT in the training loss path.
Both online_eq and target_eq now consistently use mean-logit for
direction+magnitude. The TD error is unbiased: no representation
mismatch. C51 softmax retained for order+urgency (d>=2) where
C51 gradient is active.
Result: Half grows 2.7%→5.7%, Full grows 2.7%→5.3% across epochs.
No collapse. 19/19 smoke tests pass.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Root cause: C51 distributional softmax structurally favors zero/low-variance
actions (Flat for direction, Small for magnitude). This created irrecoverable
feedback loops through FOUR paths: gradient, Bellman target argmax, action
selection, and the Q-gap conviction filter.
Direction branch (new):
- Zero C51 gradient for direction (d==0) — same treatment as magnitude
- Boltzmann softmax replaces argmax for direction selection (tau=2×Q_range)
- 2× MSE gradient amplification for direction branch head
- Mean advantage (not C51 softmax) for Bellman target argmax at d==0
- Q-gap conviction filter REMOVED from training path (redundant with
Boltzmann, harmful after high-Sharpe epochs where Bellman max bootstrap
raises Q(Flat) towards Q(directional), shrinking the gap)
Magnitude branch (Bellman target fix):
- Mean advantage for Bellman target argmax at d==1 in both MSE + C51 kernels
- Proper softmax retained for online_eq and target_eq (learning objective)
Metric fixes:
- Diversity denominator: 9 → 7 (Flat forces mag=Half, max reachable is 7)
- Threshold: 0.5% of total → 1% of directional (Flat-dominant policies
mechanically killed diversity under the old metric)
Result: 7/7 dir×mag diversity sustained epochs 7-10, Flat stable at 7-9%
(was 60-87% growing). 19/19 smoke tests pass.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
The MSE and C51 loss kernels computed Bellman targets using C51's
distributional softmax expected Q for the argmax over next-state actions.
For magnitude (d==1), this structurally favored Small (tight distribution
→ higher softmax expected Q), creating an irrecoverable target feedback
loop even when C51 gradient was zeroed.
Fix: for magnitude branch (d==1), use mean-logit Q (average of V+A
across atoms without softmax weighting) for both argmax selection AND
target Q computation. This is variance-neutral — only the average
advantage level matters, not the distributional shape.
Applied to all 3 kernels:
- mse_loss_kernel.cu: argmax + target_eq
- c51_loss_kernel.cu: argmax
- expected_q_kernel.cu: expected Q for backtest evaluator
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
S&P on direction branch kills Q-gap conviction → forces Flat →
magnitude diversity collapses mechanically. Protect all 4 branch
heads [8..24], perturb only trunk + value head + bottleneck.
Note: smoke test (3200 bars, 10 epochs) naturally converges to Flat
regardless of S&P — insufficient data for sustained directional
conviction. H100 production run (5.7M bars) showed 7/9 diversity
sustained through all 7 epochs with positive Sharpe at epoch 6.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Two issues found in the H100 production run:
1. Primary gradient budget dropped from 1.0 to 0.10 when C51 alpha
ramped to 1.0 (c51_frac=0.10 with IQN at 60%). Branch heads
depend entirely on primary gradient — starved → grad_norm=0.000
at epoch 7. Fix: floor c51_frac at 0.30 so branch heads always
get meaningful gradient.
2. val_loss identical at epochs 1-2 (6.228316) because the backtest
evaluator's CUDA Graph was never invalidated between epochs.
Fix: call invalidate_dqn_graph() before each evaluation.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
The forward pass (action selection) must apply the same advantage
standardization as the loss kernels. Otherwise the network's raw
advantages can encode stale C51 preferences that don't match the
standardized training signal.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
C51 cross-entropy structurally prefers low-variance actions (Small positions
have tighter return distributions). MSE is variance-neutral. By reducing
C51's gradient to 20% and amplifying MSE to 4× for the magnitude branch,
MSE becomes the dominant loss signal for position sizing.
Direction/order/urgency keep full C51 gradient for distributional risk awareness.
Replaces the previous mag_amp=1.5 which addressed a symptom (Flat gradient
starvation) rather than the root cause (C51 variance bias).
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
The test proves that global filter_front_month on merged trades drops
entire quarters (Q1 lost when Q2 has higher volume). Per-file filtering
keeps both quarters. This is the regression test for the bug that
produced "Need at least 18 months of data" on H100.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Global filter selected only ONE quarter's instrument_id, producing 3 months
of data instead of 2+ years. Now filters within each file so quarterly
rolls are handled correctly.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Three fixes:
1. spread_cost removed contract_multiplier ($50) — was computing in dollars
but PnL is in points, making spread 12.5× actual. The model saw every
trade as guaranteed loss → preferred Flat.
2. Mirror universe now negates ALL 17 directional features (returns, MACD,
Bollinger, SMA ratios, regression slope, CUSUM) + flips RSI. Was only
negating 4, teaching wrong direction on 50% of epochs.
3. tx_cost_multiplier updated to 0.18 bps (IBKR ES RT=$4.50/contract).
min_hold_bars=10 for volume bars (~26s at 23 bars/min).
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Replace extract_ohlcv_bars_from_dbn() with trades-based volume bar pipeline
(load_trades_sync -> filter_front_month -> build_volume_bars) in both the
precompute binary and the runtime fallback path, eliminating fake price jumps
from interleaved contract months in OHLCV DBN files.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
The mirror universe feature negated only features[0..3] (4 return-like
features) but inverted the direction action (Short↔Long). The remaining
38 directional features (momentum, trend, patterns) were NOT negated.
This taught the model the WRONG direction on 50% of epochs, cancelling
out the directional signal and producing worse-than-random performance.
Disabled until proper full-feature mirroring is implemented.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- Delete 3 unused bf16 scratch buffers (exp_h_b0/b1/b2) from
GpuExperienceCollector — allocated VRAM but never read.
The f32 versions (exp_h_b0_f32 etc.) remain in use.
- Fix stale doc comments: 45 actions -> 81 actions, index range 0-44 -> 0-80
- Extend round-trip tests to cover all 81 action indices
- Add TODO for future 4-branch struct refactor (direction + magnitude fields)
— 64 callers make it too invasive for this commit.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
The single-step backtest_env_step kernel still called decode_exposure_index()
and compute_target_position() (old linear 3-branch mapping). Updated to use
decode_direction_4b/decode_magnitude_4b/compute_target_position_4branch.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
CQL grad kernel had identical bug to expected_q: global mean across
ALL atoms instead of per-atom dueling centering. This disabled CQL's
conservative penalty by destroying per-atom C51 structure.
Also: backtest_env_step single-step kernel was missing b3_size param
(only had b0,b1,b2). Added b3_size and fixed capital floor breach
calls that passed b2_size twice instead of b2_size,b3_size.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
The Q-diagnostics was reading 9 values from a 12-element per-branch
buffer, mislabeling branch outputs as dir×mag combos. The Q-gap was
computed across ALL branches instead of within direction only.
Fixed: gap measures direction conviction (best_dir - 2nd_dir), and
per-action Q display combines dir+mag Q-values into 9 dir×mag combos.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
The expected_q kernel computed mean_adv as a SINGLE scalar averaged
across ALL atoms AND actions, then subtracted from every logit.
This destroyed per-atom advantage structure, making all actions
produce nearly identical expected Q-values.
Fixed to compute per-atom mean: mean_a(A[a,j]) separately for each
atom j, matching the C51/MSE loss kernels' correct implementation.
This bug affected: backtest evaluation action selection, Q-gap
conviction filter (always 0 → floor 0.25), and Q-value monitoring.
Training loss kernels were NOT affected (they had correct per-atom mean).
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
CVaR, Q-gap conviction, and Kelly were scaling target_position AFTER
compute_target_position_4branch, which collapsed all magnitudes to the
same effective size (conviction floor 0.25 made Full ≈ Small). Now they
scale effective_max_pos BEFORE the mapping, so the magnitude branch
controls relative sizing while risk management controls absolute scale.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>