attn_sdp_fwd and attn_layer_norm_fwd were loaded from attention_kernel.cubin
(which only contains the legacy multihead_feature_attention) instead of
attention_backward_kernel.cubin where they actually live. This caused
CUDA_ERROR_NOT_FOUND on H100, silently disabling attention.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
IQN read_total_loss() called cuStreamSynchronize + cuMemcpyDtoH every training
step, blocking the CPU until all GPU work completed. This serialized the entire
async pipeline, causing 3537ms/step instead of <10ms.
Fix: pinned device-mapped memory for IQN total_loss (same pattern as DQN trainer).
GPU writes via device pointer, CPU reads via host pointer, zero sync.
Also re-applies the cuGraphClone elimination (was lost when agents modified
fused_training.rs) and adds recursive_confidence_reduce for deterministic
gradient accumulation.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- capture_child_graph: use std::mem::forget to take ownership of cudarc's
CudaGraph handles directly. cuGraphClone corrupts cublasLtMatmul HMMA tile
scheduling state on Hopper, causing graph replay to hang with temporal ops.
- create_eval_forward_exec: instantiate directly from original graph (no clone).
Multiple execs from the same CUgraph is safe per NVIDIA docs.
- recursive_confidence_backward: replace atomicAdd into grad_buf with per-block
partial output + deterministic reduce kernel. Zero atomicAdd on gradient path.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Split into two kernels: kan_gate_backward writes per-element basis
contributions to intermediate buffers (no cross-element reduction),
then kan_grad_reduce deterministically accumulates them with one
thread per output parameter — fully eliminating atomicAdd from the
KAN spline gradient path.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Three atomicAdd anti-patterns removed from Decision Transformer kernels:
1. dt_linear_backward_kernel (critical): Split into dt_linear_backward_kernel
(input gradient only, per-sample deterministic) + dt_linear_grad_kernel
(one thread per weight, loop over samples — zero atomics).
2. dt_causal_attention_kernel: Replace cross-head atomicAdd output projection
with per-head output buffer [B, H, T, E] + separate dt_sum_heads_kernel
that sums heads in fixed order.
3. dt_cross_entropy_kernel: Remove atomicAdd for total_loss, add
dt_reduce_loss_kernel (single-thread sequential sum for full determinism).
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- gpu_experience_collector: add curiosity_weight param to new(); conditionally
initialize GpuCuriosityTrainer when weight > 0.001 (was always None)
- training_loop: pass hyperparams.curiosity_weight to GpuExperienceCollector::new()
- gpu_experience_collector: replace misleading 'curiosity disabled' log with
weight-conditional message; zero-init comment explains trainer populates weights
- attention_backward_kernel.cu: remove dead per-sample attention_backward_kernel
and attn_weight_grad_reduce (not loaded by any Rust code since cuBLAS rewrite);
remove helpers (f32_shfl, f32_block_max/sum_bwd) and defines only used by them;
update file header to describe current cuBLAS element-wise kernel contents
- c51_loss_kernel.cu: fix stale 'atomicAdd accumulator' comment on q_divergence
param — it is written by the reduction kernel, not via atomicAdd
Audit (task 4): curiosity_training_kernel.cu uses warp-reduced atomicAdd in
curiosity_forward_backward and curiosity_fused_zero_fwd_bwd_adam. These run
outside CUDA graphs (experience collection between epochs), so atomicAdd is
acceptable; existing DETERMINISM NOTE comments document this accurately.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
mamba2_step: replace memcpy_dtod_async (not captured in CUDA Graph) with
a mamba2_copy_enriched kernel that IS captured — fixes stale data on replay.
predictive_coding_loss: replace cuMemsetD8Async + atomicAdd with per-sample
output + c51_loss_reduce — eliminates both the uncaptured memset and the
non-deterministic atomicAdd, matching the existing loss pipeline pattern.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Pod logs are now archived to MinIO before pod GC. Historical logs
accessible via `argo logs <workflow>` after pod completion.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
cudaGraphAddChildGraphNode parent replay hangs on sm_90.
Individual child launches (5 per step, ~10µs overhead) work.
Each child is instantiated (cuGraphInstantiateWithFlags) and
launched (cuGraphLaunch) sequentially on the training stream.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
CLI arg was removed in cost-driven hold timing but Argo template
still passed it, causing exit code 2 on H100 deploy.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Tests had hardcoded state_dim=45/53 from before OFI was always enabled.
PPO tests use raw 45-dim data (no OFI), DQN tests use 65-dim.
899 unit tests pass, 20 GPU smoke tests pass.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
17 commits: IQL/IQN/Attention rewritten to batched cuBLAS GEMMs,
child graph architecture (5 sub-graphs in single parent launch),
temporal pipeline wired into training forward, ISV signal update,
cost-driven hold timing (no min_hold_bars), Q-spread opportunity cost,
hardcoded dimension cleanup, dead code removal.
Target: <80s epochs (from 685s), all features active.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
# Conflicts:
# docs/superpowers/plans/2026-04-17-unified-cublas-training-graph.md
# docs/superpowers/specs/2026-04-17-unified-cublas-training-graph-design.md
Adds opportunity cost to experience_env_step reward: when the model is
flat, subtract |q_gap| * opp_cost_scale from the reward. q_gap is the
model's own conviction signal (max_long_Q - max_short_Q), already
computed by experience_action_select and stored in q_gaps_buf[N].
- experience_kernels.cu: new opp_cost_scale kernel param; apply penalty
after churn penalty, before plan conviction scaling
- backtest_env_kernel.cu: matching opp_cost_scale param in both kernels;
set to 0.0 during backtest (no Q-gap available post-hoc)
- DQNHyperparameters: opportunity_cost_scale field, default 0.001
- ExperienceCollectorConfig: opportunity_cost_scale field, wired through
training_loop.rs; backtest configs default to 0.0
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Removed: enforce_hold(), min_hold_bars from config/kernels/backtest.
Added: holding_cost_rate (inventory penalty), churn_threshold_bars +
churn_penalty_scale (graduated flip penalty) to reward in
experience_env_step and backtest kernels.
The model learns optimal hold timing from cost signals:
- Per-trade tx cost prevents churning (existing)
- Inventory penalty makes large positions expensive to hold
- Churn penalty graduates cost for rapid flips
- Temporal attention learns when holding cost > expected profit
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Inventory penalty makes flat optimal — model learns to do nothing.
Instead: penalize inaction proportional to model's own predicted edge
(Q-value spread). Creates virtuous cycle: better temporal attention →
higher self-imposed penalty for missed trades → more trading on signal.
No hindsight bias (uses predicted edge, not actual price change).
Micro-reward (already exists) rewards correct positioning.
Churn penalty (new) prevents rapid flips.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Temporal ops (mamba2, predictive coding, regime dropout, ISV temporal
routing, risk budget) now run inside submit_forward_ops_main() as step
1c, captured in the cuBLAS training graph. Previously these only ran in
monitoring via reduce_current_q_stats. Also adds submit_isv_signal_update
wrapper for graph-capturable ISV signal updates.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
forward_child: 4 branch streams fork after h_s2, join before loss
aux_child: IQL/IQN/attention on parallel streams
Uses CU_STREAM_CAPTURE_MODE_RELAXED with fork-join events as graph edges
Existing branch_streams[4] + events in batched_forward.rs ready to activate
Target: <40s epochs (from <80s Phase 1)
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Add 6 new CUDA kernel functions to attention_backward_kernel.cu for the
batched cuBLAS attention path: attn_sdp_fwd, attn_sdp_bwd (sigmoid-gated
feature attention), attn_layer_norm_fwd, attn_layer_norm_bwd (per-sample
LayerNorm with residual), attn_bias_add, and attn_bias_grad_reduce.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Append 7 new CUDA kernels to iqn_dual_head_kernel.cu for the batched
cuBLAS IQN rewrite: iqn_hadamard_sigmoid, iqn_relu_fwd, iqn_relu_bwd,
iqn_quantile_huber_loss, iqn_hadamard_sigmoid_bwd, iqn_bias_add,
iqn_bias_grad_reduce.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Replace per-sample CUDA kernels (16,384 launches per GEMM) with batched
cublasLtMatmul — ONE call per GEMM layer. Forward pass uses 3 GEMMs +
SiLU + bias_add + expectile_loss. Backward pass uses 5 GEMMs + SiLU
backward + bias_grad_reduce. Eliminates grads_per_sample buffer (was
27MB), tiled backward loop, and weight_grad_reduce kernel.
Constructor now accepts Arc<SharedCublasHandle> instead of bare stream,
sharing the cuBLAS handle with the trunk forward/backward pipeline.
Eight IqlGemmDesc descriptors are cached at init with heuristic algo
selection for CUDA Graph compatibility.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Appends 5 new kernels to iql_value_kernel.cu for the cuBLAS batched GEMM
rewrite: iql_silu_fwd, iql_silu_bwd, iql_expectile_loss, iql_bias_add,
iql_bias_grad_reduce. Zero atomicAdd, fully deterministic, extern "C" for
cubin export.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
A: train_step_gpu must not capture its own graphs (rename to submit_training_ops_ungraphed)
B: Mamba2 backward must be inside unified graph (forward+backward same capture)
C: submit_adam_ops signature verified, needs weight set refs
D: cuBLAS workspace sharing — add verification logging, fail if >32MB needed
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Single cudarc-captured graph replaces 6 separate graphs.
IQL/IQN/Attention rewritten from per-sample kernels to batched cuBLAS.
Temporal ops inside unified graph. Target: <80s epochs from 685s.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Temporal ops (mamba2, regime_dropout, isv_temporal_route, risk_budget)
between graph_forward and graph_aux replay corrupt graph_aux on Hopper.
ISV signal update + hold enforcement + aux_frequency still active.
Temporal pipeline will be properly addressed in the cuBLAS aux rewrite.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Temporal ops between graph_forward and graph_aux replay use shared
buffers (h_s2). cudarc event recording on these buffers corrupts
the graph_aux replay state on Hopper.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Adding mamba2/predictive_coding/regime_dropout/isv_temporal_route/
risk_budget to submit_forward_ops_main broke graph_forward replay
on Hopper (same graph structure issue). Moved to run_full_step
between graph_forward replay and aux ops — ungraphed but still
runs every step, enriching h_s2 before IQL/IQN/attention.
graph_forward keeps the working kernel set (cuBLAS + ISV + plan).
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
min_hold=10 allowed ~1200 trades/day — costs destroyed the edge.
min_hold=50 (~6.5 min) limits to ~230 trades/day max.
Adaptive extension now scales 0-2× base (was hardcoded 0-8 bars).
With base=50: range is 50-150 bars depending on ISV stability.
Losing positions (negative reward_ema) get base only (50 bars).
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
1. Temporal pipeline in training forward (submit_forward_ops_main):
- mamba2_step: temporal scan enriches h_s2 with history
- compute_predictive_coding_loss: temporal smoothness
- apply_regime_dropout: regime-conditioned dropout
- launch_isv_temporal_route: per-feature temporal weights
- risk_budget_forward: risk budget from h_s2
These were only in the monitoring path (reduce_current_q_stats),
never in the actual training forward. The model trained without
any temporal context.
2. ISV signal update after each adam step:
update_isv_signals() called after mamba2 backward, reads pinned
loss/grad_norm/Q-mean and updates the 12-element ISV vector.
Was never called during training — ISV signals stayed at zero.
3. Backtest min_hold fixed:
eval_min_hold was hardcoded 0 — no hold enforcement in validation.
Now uses config.min_hold_bars.
4. Backtest ISV signals:
Passes frozen ISV signals from last training step to validation
backtest for adaptive hold enforcement.
5. aux_frequency parameter (default 4):
IQL, IQN, attention, CQL run every 4th step instead of every step.
~4x faster epochs. graph_forward still runs every step.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
graph_mega captures spectral+forward+aux in one graph. Replay hangs
on Hopper even with cudarc wrapper. Split graphs work: graph_forward
(from GpuDqnTrainer) + graph_aux (cudarc wrapper). Both graphed,
2 launches per step instead of 1.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Rewrote capture_graph_mega to use cudarc's begin_capture/end_capture
instead of raw cuda_sys. Same fix that resolved graph_aux hang.
All graph types (graph_mega, graph_aux, graph_spectral) now use
cudarc wrapper consistently. Falls back to split graphs if mega fails.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
compute_validation_loss only returned Sharpe, discarding all other
backtest metrics. Now logs the complete picture on every epoch so we
can verify if the learning system actually works on out-of-sample data.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
bars_per_day was hardcoded 390 (1-minute candles) but we use MBP-10
imbalance bars (~7500/day). All Sharpe/Sortino/return annualization
was off by sqrt(7500/390) ≈ 4.4×.
Now computed from actual fxcache timestamps:
bars_per_day = total_bars / (trading_days from timestamp span)
Propagates to: val_Sharpe, train_Sharpe, financials, backtest evaluator.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Root cause: capture_graph_aux and capture_graph_spectral used raw
cuda_sys::cuStreamBeginCapture_v2 / cuStreamEndCapture, bypassing
cudarc's internal state tracking. Mixing raw FFI with cudarc's
managed context corrupts the captured graph, causing replay hangs.
capture_training_graphs (GpuDqnTrainer) always used cudarc's
stream.begin_capture() / stream.end_capture() and worked. Now all
graph captures use the same cudarc wrapper.
All graphs re-enabled: graph_forward, graph_forward_ddqn, graph_adam,
graph_spectral, graph_aux. Full graphed pipeline.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
graph_aux was still being captured (then not replayed). The capture
process uses cuStreamBeginCapture which corrupts cudarc's internal
state, causing the NEXT graph_forward replay to hang on step 2.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
graph_forward replays fine (step 1 completes in 3.8s).
graph_aux replay hangs on step 2. All aux ops work ungraphed
(confirmed by diagnostic run). Keep graph_forward graphed for
speed, run aux ops directly on stream until root cause found.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>