Commit Graph

3820 Commits

Author SHA1 Message Date
jgrusewski
1d3a533f99 fix: capture PER sampling inside parent graph — zero ungraphed PER kernels per step
Move sample_proportional (prefix scan, binary search, gather x6, IS weights)
from training_loop.rs into run_full_step, captured as per_sample child graph
node. Steps 1+ replay all 13 children via single cuGraphLaunch with no
ungraphed PER kernel dispatches.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-19 01:14:53 +02:00
jgrusewski
ead6d195fa feat: true single-graph — ONE cuGraphLaunch per step, zero ungraphed launches
Composes 11 child graphs into single parent via cuGraphAddChildGraphNode.
Counter increments (GPU-side), IQL modulate, PER priority update all
captured as child graph nodes. PhaseEvents removed — single parent launch
makes per-child event profiling irrelevant.

Step 0 runs ungraphed + captures all 11 children + composes parent.
Steps 1+ replay via single cuGraphLaunch on the training stream.
After graph launch: only pinned scalar reads (nanoseconds, no GPU ops).

Child graph order:
  counters -> spectral -> forward -> ddqn -> aux -> post_aux ->
  adam_grad -> adam -> maintenance -> iql_modulate -> per_priority

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-19 00:53:46 +02:00
jgrusewski
b2135d01e0 feat: direct-to-trainer gather — eliminates upload_batch_gpu + 6 DtoD copies per step
PER gather now writes directly to GpuDqnTrainer's padded buffers via
gather_f32_rows_padded, gather_f32_scalar, and gather_i32_scalar kernels
compiled into the replay buffer cubin. set_trainer_buffers() wires stable
device pointers at init; sample_proportional uses them when available,
falling back to intermediate buffers otherwise. memset_zeros calls in
the sampling hot path converted to raw cuMemsetD8Async.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-19 00:34:10 +02:00
jgrusewski
5b197151db feat: graph_utility_kernels.cu — gather_padded + increment_step_counters for true single-graph
Adds 4 GPU-native kernels replacing host-side operations on the training
hot path: gather_f32_rows_padded (row gather + 128-byte zero-pad),
gather_f32_scalar, gather_i32_scalar, and increment_step_counters (atomic
counter bumps + cosine-annealed tau — zero CPU sync per step).
Wired into build.rs (nvcc cubin compile) and gpu_dqn_trainer.rs
(GRAPH_UTILITY_CUBIN include_bytes!).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-19 00:22:45 +02:00
jgrusewski
26edbb6d8f fix: isolate shrink_perturb and scale_adam_momentum from graphed CUmodule
shrink_and_perturb() is called at epoch boundaries after child graphs are
captured. shrink_perturb_kernel lives in the same CUmodule as scale_f32,
saxpy_f32, spectral_norm, adam_update, and other graphed CUfunctions.
On Hopper (sm_90), launching any CUfunction from a graphed CUmodule
ungraphed corrupts the child graph kernel state → 3100ms replay.

scale_adam_momentum() is called at cosine LR warm restarts (also after
graph capture). scale_f32_kernel is captured in forward_child — same
CUmodule violation.

Fix: add shrink_perturb_ungraphed and scale_f32_ungraphed loaded from
the existing ungraphed_module (separate CUmodule instance of the same
DQN_UTILITY_CUBIN). Both ungraphed callers now use isolated handles.
CUmodule count: 5 → 5 (ungraphed_module already existed, reused).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-19 00:16:29 +02:00
jgrusewski
4f7b1f4804 plan: true single-graph — 7 tasks, zero CPU on hot path
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-19 00:06:05 +02:00
jgrusewski
e8d382fc2f fix: isolate ungraphed pad_states + stochastic_depth_rng into separate CUmodule
ROOT CAUSE: pad_states_kernel and stochastic_depth_rng_kernel were loaded from
the SAME CUmodule as graph-captured kernels (adam_update, grad_norm, saxpy, etc.)
but launched UNGRAPHED every step. On Hopper, launching ANY CUfunction from a
CUmodule that also has graph-captured CUfunctions corrupts the graph's kernel
state — causing 3100ms replay instead of ~30ms.

Fix: load from a separate ungraphed_module. Zero CUmodule contamination between
graphed and ungraphed execution contexts.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-19 00:02:23 +02:00
jgrusewski
b0b06383ad spec: true single-graph — zero CPU on hot path, zero ungraphed launches
Single cuGraphExecLaunch per step. PER sampling, gather, training,
priority update ALL as child graph nodes in one parent. Direct-to-trainer
gather eliminates DtoD copies. GPU-side counters eliminate host writes.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-18 23:59:35 +02:00
jgrusewski
ad83f8cc89 infra: add nsight-systems-cli to ci-builder Docker image
Enables nsys profiling for CUDA graph node-level traces on H100.
Use --sanitizer nsys in argo-train.sh to activate.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-18 23:54:57 +02:00
jgrusewski
4bd26c7361 feat: nsys profiling support in argo train template
Use --sanitizer nsys to wrap training binary with nsys profile.
Captures CUDA graph node-level traces + GPU metrics.
Output: /workspace/output/nsys_profile.nsys-rep

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-18 23:37:47 +02:00
jgrusewski
19b6b5af47 fix: complete CUfunction isolation — zero cross-child sharing + raw memset + pinned memory
Three changes to eliminate the 3100ms replay regression on Hopper:

1. CUfunction isolation: saxpy_f32_kernel was shared across 3 child graphs
   (forward, aux, adam_grad). Added saxpy_f32_adam_grad and saxpy_f32_aux from
   separate CUmodules. Also isolated grad_norm_standalone for post_aux_child
   (was shared with forward_child's d_logits clipping path).

2. Raw cuMemsetD8Async: replaced all cudarc memset_zeros in graph-captured
   functions (submit_forward_ops_main, apply_cql_gradient,
   run_causal_intervention_unconditional) with raw cuMemsetD8Async which is
   properly captured by CUDA Graph. 6 call sites fixed.

3. DtoD memcpy audit: all memcpy_dtod_async calls verified — large buffer
   copies (grad snapshot 2.6MB, multi-horizon blend) are correct for DtoD;
   no scalar copies found that should use pinned memory.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-18 23:24:52 +02:00
jgrusewski
79fce72afb fix: CUfunction isolation for post_aux child + raw memset in IQN/IQL/Attention
ROOT CAUSE of 3100ms adam_child: adam_update_kernel and grad_norm_finalize_kernel
were shared between post_aux_child and adam_update/forward children. On Hopper,
CUfunction sharing across child graphs corrupts kernel state.

Fix: load utility cubin from separate CUmodule for post_aux child, giving it
isolated adam_update_post_aux, grad_norm_finalize_post_aux, and scale_f32_post_aux
handles. Zero cross-child CUfunction sharing for these kernels.

Also: convert memset_zeros to raw cuMemsetD8Async in IQN (3 calls), IQL (1),
Attention (1) — eliminates cudarc device_ptr_mut() overhead during graph capture.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-18 23:11:15 +02:00
jgrusewski
c7185d2483 perf: convert all CudaSlice kernel args to raw u64 — eliminate cudarc overhead in graph
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-18 23:01:51 +02:00
jgrusewski
4b6d9d2e82 perf: fused RELU_BIAS epilogue for target, collector, and ensemble forward paths
Adds try-fused-fallback-to-separate pattern to forward_target_raw,
forward_online_f32, and forward_value_head — matching the existing
pattern in forward_online_raw. Eliminates ~10-14 separate bias+ReLU
kernel launches by fusing them into the preceding cuBLAS GEMM epilogue.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-18 22:40:46 +02:00
jgrusewski
76cfc0d185 perf: IQN quantiles 64→32 — saves 4.3GB VRAM, halves IQN GEMMs
Academic literature shows diminishing returns past 32 quantiles.
64 quantiles allocated 8.6GB for tiled intermediates (B×Q×hidden).
32 quantiles: 4.3GB (half), ~13 GEMMs halved in aux_child.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-18 22:28:23 +02:00
jgrusewski
1481ba2699 feat: unified single-graph — 7+ children, zero ungraphed kernel launches
Add post_aux_child (selectivity + denoise + risk_sgd + multi-horizon) and
maintenance_child (causal intervention + gradient vaccine) to the CUDA graph
pipeline. All kernel launches now run inside child graphs; only kernel-free
ops (pinned readbacks, host counters, IQL modulate_td_errors on its own
CUfunction) remain outside the graph.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-18 22:08:54 +02:00
jgrusewski
ebc4e66578 spec: add Phase 3 performance optimizations (8 items, prioritized)
3.1 Fuse bias+ReLU into cuBLAS epilogue (highest impact)
3.2 IQN quantiles 64→32 (4.3GB VRAM + 15-20% aux speedup)
3.3 Double-buffered experience collection
3.4 cuGraphExecUpdate for hyperparameter changes
3.5 Tensor core dimension padding (adv_h 32→128)
3.6 Batch size 16384→8192 (unlocks multi-stream benefit)
3.7 Persistent Adam kernel
3.8 Memory pool (cuMemAllocAsync)

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-18 22:04:41 +02:00
jgrusewski
2334d1d6ce feat: unconditional submit_post_aux_ops + submit_maintenance_ops + prev_grad_buf
Add child-graph-capturable submit methods that run post-aux and
maintenance ops without conditional guards (graphs_captured, step %
interval). This enables moving all kernel launches into CUDA child
graphs, eliminating ungraphed ops that cause CUfunction corruption on
Hopper during 3100ms adam replays.

- prev_grad_buf: snapshot of grad_buf after Adam for next-step vaccine
- submit_post_aux_ops: selectivity, denoise, risk SGD, multi-horizon
- submit_maintenance_ops: causal intervention + gradient vaccine
- run_causal_intervention_unconditional: no step/graph guards
- apply_gradient_vaccine_from_prev: uses prev_grad_buf instead of
  separate forward+backward pass

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-18 21:52:51 +02:00
jgrusewski
c4c7303e82 audit: 21 CUfunction conflicts — vaccine uses 15+ forward_child kernels ungraphed 2026-04-18 21:45:27 +02:00
jgrusewski
c054b4e1d8 plan: unified single-graph — 5 tasks, zero ungraphed kernel launches
Task 1: CUfunction audit (map every kernel to its child graph)
Task 2: Expose submit methods (post_aux + maintenance + prev_grad_buf)
Task 3: Add new child graphs (capture + launch 7 children)
Task 4: Verify remaining outside-graph code is kernel-free
Task 5: Deploy and validate (target: adam ~30ms, epoch ~64s)

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-18 21:26:40 +02:00
jgrusewski
55f26e0f51 spec: unified single-graph architecture — zero ungraphed kernel launches
Root cause: CUfunction sharing between graphed children and ungraphed
outside-graph ops corrupts kernel state on Hopper → 3100ms adam replay.

Fix: capture EVERYTHING unconditionally in child graphs. Selectivity,
denoise, causal intervention, vaccine, Q-stats all become graphed children.
Zero ungraphed launches = zero CUfunction conflicts.

7 children, ~190 kernel nodes, one capture, one replay per step.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-18 21:22:34 +02:00
jgrusewski
075ed8ba4f fix: disable selectivity+denoise outside graph — CUfunction shared with graphed children
step_denoise_adam uses adam_update_kernel + grad_norm_finalize_kernel UNGRAPHED
after child graph replays. Same CUfunctions are captured in adam_child and
forward_child. On Hopper, launching a captured CUfunction ungraphed corrupts
graph kernel state → 3100ms adam_child replay on next step.

Fix: skip selectivity+denoise ungraphed training. These are auxiliary optimizers
with non-fatal error handling. To re-enable: load separate CUfunction instances
for ungraphed paths or move into child graphs.

Expected: adam drops from 3100ms to ~30ms. Epoch from ~690s to ~60s.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-18 21:00:05 +02:00
jgrusewski
e28add3768 fix: adam_child 3100ms — grad_norm_finalize_kernel shared between forward+adam children
ROOT CAUSE: grad_norm_finalize_kernel (CUfunction) was captured in BOTH
forward_child (d_logits clipping at lines 7572/7821) AND adam_child
(compute_grad_norm_for_adam via launch_grad_norm_finalize at line 6346).
Sharing a CUfunction between two captured graphs corrupts kernel state
on Hopper (sm_90), causing the second graph to replay at 3100ms instead
of <1ms. This is the SAME class of bug as the earlier grad_norm_standalone
fix — we fixed one shared CUfunction but missed the finalize kernel.

FIX: Load grad_norm_finalize from a SEPARATE CUmodule (new cubin load)
for the adam path. Each CUfunction instance is now captured in exactly
one child graph. Also split adam into adam_grad + adam_update children
for per-operation timing visibility.

Expected: adam drops from 3100ms to ~30ms. Total step ~350ms. Epoch ~62s.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-18 20:09:19 +02:00
jgrusewski
7b0280a187 debug: split adam into adam_grad + adam_update to isolate 3.1s bottleneck
adam_child = 3113ms even after CUfunction fix. Split into:
- adam_grad: mamba2_backward + step_mamba2_adam + pruning + grad_norm
- adam_update: Adam kernel + unflatten + ISV

If adam_grad is fast and adam_update is slow, the bottleneck is Adam/unflatten.
If adam_grad is slow, grad_norm_kernel has an issue.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-18 19:13:32 +02:00
jgrusewski
b4cf279b73 fix: reduce H100 replay buffer VRAM fraction 55% → 45% — OOM from IQN 8.6GB + Phase 2
IQN with 64 quantiles × batch=16384 allocates 8.6GB for tiled intermediates.
Combined with attention (250MB), Phase 2 aux workspaces (64MB), backward
branch scratch (16MB), the 55% replay fraction left only 349MB free —
not enough for experience collector's 192MB next_states allocation.

45% frees ~8GB headroom: replay buffer 19.4M → ~15.9M (still 60× batch).

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-18 18:42:06 +02:00
jgrusewski
32151ba8bc fix: backward branches reuse forward workspace — saves 128MB (OOM fix)
The backward branch parallelism allocated 4 × 32MB = 128MB for per-branch
cuBLAS workspaces, pushing H100 VRAM over 80GB → OOM. Forward and backward
are in the same child graph (sequential) — workspaces never conflict.

Now backward receives forward's branch_workspace_ptrs as a parameter.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-18 16:57:34 +02:00
jgrusewski
80d871ae1d feat: IQN prepare_buffers extraction for parallel stream dispatch
Separates buffer prep (decode_actions, DtoD copies) from training pipeline
so prep runs on main stream and execute_training_pipeline runs on iqn_stream.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-18 16:37:54 +02:00
jgrusewski
108a0993c7 perf: parallel backward branches using fork-join on branch_streams
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-18 16:35:27 +02:00
jgrusewski
9fb330da7e fix: adam_child 3120ms — CUfunction conflict between forward_child and adam_child
grad_norm_standalone was captured in BOTH forward_child (d_logits clipping)
and adam_child (gradient norm). Same CUfunction in two captured graphs
corrupts kernel state on Hopper — the GPU replayed a corrupted node for
3120ms per step (88% of total step time).

Fix: adam_child uses grad_norm_kernel (the regular handle, not captured
elsewhere) instead of grad_norm_standalone. Each CUfunction is now captured
in exactly one child graph.

Expected: adam_child drops from 3120ms to ~30ms. Total step from ~3550ms
to ~460ms. Epoch from ~700s to ~80s.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-18 16:34:12 +02:00
jgrusewski
16f5147356 feat: IQN/Attention accept workspace+stream override for multi-stream
Add optional override_stream and override_workspace parameters to
execute_training_pipeline (GpuIqnHead), and forward/backward/adam_step
(GpuAttention). All existing callers pass None, None — behaviour is
unchanged. The parallel path can now dispatch these ops to dedicated
CUDA streams with separate cuBLAS workspaces.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-18 16:20:31 +02:00
jgrusewski
dc9f25df63 feat: allocate aux streams, workspaces, events for Phase 2 multi-stream
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-18 16:15:38 +02:00
jgrusewski
f851119e2d plan: Phase 2 multi-stream parallelism — 7 tasks
Task 1: Allocate aux streams + workspaces + events
Task 2: IQN/Attention accept workspace+stream override
Task 3: Fork-join in submit_aux_ops (IQN+Attention parallel)
Task 4: Switch capture mode GLOBAL → RELAXED
Task 5: Forward backward branch parallelism
Task 6: Fix adam child event timing (verify)
Task 7: Deploy + verify child graph breakdown

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-18 16:07:06 +02:00
jgrusewski
f1ebf78b5a spec: Phase 2 multi-stream parallelism design
Hybrid approach based on H100 measurements:
- aux_child: grouped GEMM for IQL high+low, 2 streams for IQN+Attention
- forward_child: backward branch parallelism (forward already multi-stream)
- Target: 4.0s/step → 2.0s/step

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-18 16:03:45 +02:00
jgrusewski
8ebb65a942 fix: add adam_child_end event — previous timing used wrong event (adam_end from old phase system)
The child graph breakdown showed adam=-1.0ms because it was measuring
elapsed(backward_end, adam_end) where adam_end is from the OLD phase
timing system, not the adam child graph. Added proper adam_child_end
event recorded after the adam_child launch.

This will reveal the actual adam_child GPU time — currently suspected
to be 3117ms (73% of step time) based on total - measured children.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-18 16:00:06 +02:00
jgrusewski
4295cbde79 perf: add per-child graph CUDA event timing breakdown
Records cuEventRecord between each of the 5 child graph launches.
At epoch end, cuEventElapsedTime reports per-child GPU time:
spectral, forward, ddqn, aux, adam.

This tells us exactly which child dominates the 4s/step graph replay
cost, informing Phase 2 multi-stream parallelism priorities.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-18 14:59:57 +02:00
jgrusewski
61e4d8730c docs: update unified graph spec with Phase 1 H100 measurements
Phase 1 results: 4.0s/step GPU compute (graph replay = ungraphed, cuBLAS
replaced manual matmul but feature count increased). CPU pipeline fully
async (0.1ms/step). Total ~718s/epoch — 100% GPU bound.

Updated Phase 2 targets: aux parallelism (2b) is highest priority —
80/177 kernels with 4 independent trainers. nsys profiling should run
first to validate SM vs memory-bound assumptions.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-18 14:59:57 +02:00
jgrusewski
c8d87626d3 perf: move Q-stats to epoch boundary only — zero sync in training loop
Per-step reduce_current_q_stats created GPU pipeline bubbles every 50 steps.
Each sync waited for the GPU to drain all pending graph replays before
the CPU could continue. Q-stats are now computed at epoch boundary only
(process_epoch_boundary already calls reduce_current_q_stats).

Training loop is now fully async: 178 graph launches queue in <1ms,
GPU executes back-to-back with zero pipeline stalls.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-18 14:36:36 +02:00
jgrusewski
ff2b599965 fix: Q-stats readback via pinned device-mapped memory — zero cuStreamSynchronize
q_stats_kernel now writes directly to q_readback_dev_ptr (pinned device-mapped).
CPU reads previous step's values from q_readback_pinned — one-step lag, zero sync.

This eliminates the LAST cuStreamSynchronize from the per-step training path.
The training loop is now fully async: graph launch → CPU returns immediately →
next step starts while GPU still executes → zero pipeline bubbles.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-18 14:15:04 +02:00
jgrusewski
a171eaa710 fix: reduce_current_q_stats now just reads graph-computed q_out_buf
The function re-ran 15+ kernels (expected_q, q_mean_center, q_anchoring,
risk_budget, branch_confidence, epistemic_gate, temporal_consistency,
branch_independence, q_attention, graph_message_pass, q_denoise) outside
the graph every step — duplicating work the graph already did.

Now: q_out_buf is populated by graph replay (populate_q_out in aux_child).
reduce_current_q_stats just runs q_stats_kernel (1 kernel) + DtoH sync.
Runs every step since cost is now ~1ms, not ~4s.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-18 14:11:10 +02:00
jgrusewski
ba1c79763e fix: reduce_current_q_stats every 50 steps, not every step — 15 kernels + sync per call
reduce_current_q_stats ran 15+ GPU kernels (expected_q, q_mean_center,
q_anchoring, risk_budget, branch_confidence, epistemic_gate, temporal_consistency,
branch_independence, q_attention, graph_message_pass, q_denoise, q_stats)
PLUS cuStreamSynchronize for DtoH readback — EVERY training step.

This added ~4s per step (matching graph replay cost), making epochs 700s
instead of ~35s. Running every 50 steps reduces Q-stats overhead to <1%
of epoch time while still tracking Q-range changes within 200 steps.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-18 14:07:27 +02:00
jgrusewski
a156b48844 perf: move ensemble into aux_child graph — eliminate ungraphed per-step overhead
Ensemble diversity (2 extra head forward passes + KL gradient + trunk backward)
was running OUTSIDE the CUDA graph on every training step. Now runs inside
submit_aux_ops, captured in the aux_child graph for zero launch overhead.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-18 13:50:51 +02:00
jgrusewski
144e70ca29 fix: remove redundant temporal/ISV re-execution in reduce_current_q_stats
reduce_current_q_stats() re-ran 10 kernels (mamba2_step, predictive_coding,
regime_dropout, ISV forward/gate/route, recursive_confidence, trade_plan,
plan_noise, risk_budget) every training step OUTSIDE the graph replay.

These ops already executed inside the graph — the logit buffers are current.
The redundant execution added ~4s of GPU compute per step (matching the
graph replay cost), doubling total epoch time.

Fix: compute expected_q directly from the existing logit buffers.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-18 13:42:35 +02:00
jgrusewski
c3cdb39081 fix: acquire write lock once per epoch, not per training step
The async write lock (self.agent.write().await) was acquired inside the
per-step training loop, causing 696s/epoch overhead from async yield on
every iteration — 178 steps × ~4s per lock acquisition. Actual GPU compute
was only 3.6s (graph replay at 2ms/step).

Fix: acquire the lock once before the loop. The training loop is the only
writer during this phase — no contention, no need to release between steps.

Expected: epoch time drops from 700s to ~5s.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-18 13:15:25 +02:00
jgrusewski
8d88c63434 debug: add per-step timing logs to identify hang vs slow compute
Logs step_ms for first 5 steps and every 50th step thereafter.
If steps aren't completing, the graph is hanging.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-18 12:56:25 +02:00
jgrusewski
93ce3e118b fix: attention LayerNorm kernel arg mismatch — extra arg in fwd, missing 2 args in bwd
Forward: removed spurious ln_save_xnorm arg (kernel has 9 params, code passed 10).
The extra arg shifted D and B, causing CUDA_ERROR_ILLEGAL_ADDRESS on H100.

Backward: added missing residual (projected_buf) and save_mean args (kernel has
11 params, code passed 9). Without these, gamma/rstd/d_x/d_gamma/d_beta were all
shifted to wrong positions.

Both bugs existed since the cuBLAS attention rewrite but were latent because
attn_sdp_fwd was loaded from the wrong cubin (attention_kernel.cubin instead of
attention_backward_kernel.cubin), silently disabling attention.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-18 12:41:41 +02:00
jgrusewski
3ff5bd696d debug: add attention tracing + sync barriers to isolate H100 segfault
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-18 12:23:05 +02:00
jgrusewski
f8f1147de7 fix: attention kernel load from wrong cubin + build.rs bf16 label
attn_sdp_fwd and attn_layer_norm_fwd were loaded from attention_kernel.cubin
(which only contains the legacy multihead_feature_attention) instead of
attention_backward_kernel.cubin where they actually live. This caused
CUDA_ERROR_NOT_FOUND on H100, silently disabling attention.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-18 12:07:13 +02:00
jgrusewski
69ace546aa fix: eliminate cuStreamSynchronize in IQN loss readback — was serializing every step
IQN read_total_loss() called cuStreamSynchronize + cuMemcpyDtoH every training
step, blocking the CPU until all GPU work completed. This serialized the entire
async pipeline, causing 3537ms/step instead of <10ms.

Fix: pinned device-mapped memory for IQN total_loss (same pattern as DQN trainer).
GPU writes via device pointer, CPU reads via host pointer, zero sync.

Also re-applies the cuGraphClone elimination (was lost when agents modified
fused_training.rs) and adds recursive_confidence_reduce for deterministic
gradient accumulation.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-18 11:35:10 +02:00
jgrusewski
6960a02af4 fix: eliminate cuGraphClone + recursive_confidence atomicAdd — deterministic graph replay
- capture_child_graph: use std::mem::forget to take ownership of cudarc's
  CudaGraph handles directly. cuGraphClone corrupts cublasLtMatmul HMMA tile
  scheduling state on Hopper, causing graph replay to hang with temporal ops.
- create_eval_forward_exec: instantiate directly from original graph (no clone).
  Multiple execs from the same CUgraph is safe per NVIDIA docs.
- recursive_confidence_backward: replace atomicAdd into grad_buf with per-block
  partial output + deterministic reduce kernel. Zero atomicAdd on gradient path.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-18 10:50:46 +02:00
jgrusewski
80769abb9c cleanup: purge all bf16 naming remnants — pure f32/TF32 pipeline
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-18 10:30:04 +02:00