Commit Graph

3046 Commits

Author SHA1 Message Date
jgrusewski
61477bc360 feat(cuda): add DSR + n-step to warp experience kernel
Mirror the DSR (Differential Sharpe Ratio) and n-step return
accumulation added by Task 4 to the scalar kernel into the warp
kernel (dqn_full_experience_kernel_warp). DSR/n-step state runs
on lane 0 only, matching the existing reward computation pattern.
All accumulators are reset on episode boundaries.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-03-10 13:05:54 +01:00
jgrusewski
aee3871788 feat(cuda): add DSR + n-step to scalar experience kernel
Adds Differential Sharpe Ratio reward shaping and n-step return
accumulation to dqn_full_experience_kernel (scalar path). DSR replaces
EMA normalization when use_dsr=1; n-step ring buffer accumulates
discounted returns for effective_n>1. Both are reset on episode
boundaries.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-03-10 13:03:01 +01:00
jgrusewski
923fab9919 test(cuda): add DSR + n-step CPU parity tests for GPU device functions
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-03-10 12:56:03 +01:00
jgrusewski
b11669b17c feat(cuda): add GpuStatistics wrapper with host-side mean/variance computation
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-03-10 12:54:25 +01:00
jgrusewski
626255d507 feat(cuda): add select_actions_branching() with allocate-then-DtoD pattern
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-03-10 12:53:47 +01:00
jgrusewski
a168fa86ad feat(cuda): add nstep_push_and_sum() device function for GPU n-step returns
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-03-10 12:52:37 +01:00
jgrusewski
4c9dc52fbf feat(cuda): add branching_action_select kernel + composition tests
Append third CUDA kernel to epsilon_greedy_kernel.cu for Branching DQN:
3 independent epsilon-greedy heads (exposure/5, order/3, urgency/3)
compose to factored action index 0-44 (exposure*9 + order*3 + urgency).

Add Rust composition tests verifying full 45-action coverage and
round-trip decomposition correctness.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-10 12:45:09 +01:00
jgrusewski
ef7b336818 feat(cuda): add batch_statistics single-block parallel reduction kernel
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-03-10 12:44:55 +01:00
jgrusewski
61a9f45af5 feat(cuda): add dsr_step() device function for GPU DSR reward shaping
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-03-10 12:44:01 +01:00
jgrusewski
ccf9e22a55 fix(clippy): resolve misc warnings across ML crates
- Implement Display trait instead of inherent to_string() (versioning.rs)
- Replace iter().nth() with .get(), .get(0) with .first()
- Collapse else-if chains, derive Default for enums
- Fix deref warnings, redundant let bindings, push_str('\n')
- Convert match-to-if-let, use saturating_sub, RangeInclusive::contains
- Replace vec![push;push] with vec![...] literal (feature_extraction.rs)
- Remove redundant redefinitions, unnecessary parentheses, casts
- Fix &Vec<_> to &[_] in function parameters

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-10 12:12:14 +01:00
jgrusewski
fca2495a73 fix(clippy): add doc backticks and const fn across ML crates
- Add backticks to type names in doc comments (doc_markdown)
- Mark eligible functions as const fn (missing_const_for_fn)
No behavior changes.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-10 11:51:31 +01:00
jgrusewski
278fd3673f fix(clippy): allow module_name_repetitions and integer_division in ML crates
module_name_repetitions: PpoConfig in ppo module is conventional ML naming.
Renaming breaks every import across the workspace.

integer_division: Basis point calculations, batch size math, combinatorial
formulas — truncation is intentional. Float conversion would introduce bugs.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-10 11:27:01 +01:00
jgrusewski
7ef92983f9 fix(clippy): apply cargo clippy --fix across workspace
Mechanical auto-fixes: redundant borrows, clone on Copy, or_insert_with,
single-char push_str, get(0) → first(), needless borrow, let_and_return.
150 files, no behavior changes.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-10 11:17:51 +01:00
jgrusewski
c2b6a290e2 ci: trigger warm-cache benchmark on Ubuntu 24.04 images
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-10 10:46:35 +01:00
jgrusewski
e6e31305a2 fix(ml-dqn): suppress unused grad_norm_tensor warning without cuda
Prefix with underscore since it's only consumed in #[cfg(feature = "cuda")] block.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-10 10:33:18 +01:00
jgrusewski
2662d5cda2 fix(build): gate GPU PER calls behind #[cfg(feature = "cuda")]
is_gpu_prioritized() and grad_norm squeeze are only available with CUDA.
Without the gate, compile-services (CPU-only) fails with E0599.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-10 10:25:52 +01:00
jgrusewski
9a96691001 perf(dqn): eliminate all GPU→CPU roundtrips from training hot path
Replace per-step to_scalar()/to_vec1() readbacks with GPU-resident
tensor accumulation and single epoch-boundary sync. On H100 this
removes ~12-15 pipeline flushes per training step (~5μs each),
enabling full GPU saturation with zero CPU sync in the inner loop.

Key changes:
- GpuTrainResult: train_step() returns GPU scalar tensors (loss_gpu,
  grad_norm_gpu) instead of f32 — zero readback per step
- Regime conditional: train all 3 heads unconditionally with
  zero-masked weights (mathematical no-op) instead of 3-6
  to_scalar() mask count checks per step
- GPU PER: max_priority as GPU tensor with flush_max_priority(),
  delta-based priority update via index_add (no CPU dedup)
- Deferred diagnostics: NaN detection, CQL logging, Q-value
  estimation all moved to epoch boundary where pipeline is
  already synced
- Legacy CPU-readback wrappers removed entirely

9 files changed, +393/-295 lines. 520 DQN tests passing.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-10 10:05:18 +01:00
jgrusewski
18166a9e6f perf(gpu): zero-roundtrip DQN experience collection via cuMemcpyDtoDAsync
Eliminate GPU→CPU→GPU roundtrip in the experience collection hot path.
Kernel outputs (states, rewards, actions, dones) now stay on GPU via
device-to-device copy into candle Tensors. Only rewards + actions are
downloaded for monitoring metrics (~0.5MB vs ~3.14GB at 32K episodes).

- Extract launch_kernel() helper from collect_experiences() (DRY)
- Add collect_experiences_gpu() — DtoD path for GPU PER
- Add cuda_slice_to_tensor_f32/i32_to_u32 DtoD copy utilities
- GPU next_states via tensor narrow/cat/reshape (no CPU loop)
- Trainer selects GPU vs CPU path based on is_gpu_prioritized()
- stream.synchronize() barrier ensures cross-stream data visibility

0 warnings, 1266 tests pass (874 ml + 392 ml-dqn)

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-10 10:05:18 +01:00
jgrusewski
21462d5f0b perf(dqn): eliminate GPU→CPU roundtrips from training hot path
Forward-port from worktree-gpu-hotpath-audit (2 commits):

1. Zero-roundtrip DQN experience collection via cuMemcpyDtoDAsync:
   - GpuExperienceCollector uses device-to-device copies for state tensors
   - Shared memory weight caching in GPU replay buffer
   - NaN priority clamping on GPU (no CPU readback)

2. Eliminate all GPU→CPU roundtrips from training loop:
   - GpuTrainResult: loss + grad_norm stay as GPU scalar tensors
   - Single to_scalar() readback at epoch boundary (not per step)
   - GPU-resident loss accumulation across training steps
   - RegimeConditional head selection via GPU tensor ops
   - Removed per-step NaN diagnostic checks (now epoch-level)

Also removes dead `states_tensor` field from ComputeLossResult and
fixes redundant field name clippy warning.

Expected: ~15-20% training throughput improvement on H100 by
eliminating synchronous GPU→CPU transfers in the inner loop.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-10 10:04:36 +01:00
jgrusewski
645e251fbb fix(docker): add passwd package for useradd on ubuntu:24.04
Ubuntu 24.04 minimal doesn't include passwd (useradd/groupadd).
Debian bookworm-slim had it by default. Required for foxhunt user.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-10 09:53:27 +01:00
jgrusewski
76bf7d9df1 infra(ci): set CARGO_BUILD_JOBS=14 to match cgroup CPU request
Cargo defaults to /proc/cpuinfo CPU count (32 host cores) but each
compile pod is cgroup-limited to 14 CPU request. Without this, both
pods spawn 32 threads each (64 total) causing heavy throttling on
the 32-core POP2 node. Setting CARGO_BUILD_JOBS=14 per pod avoids
context-switch overhead and uses cgroup-aware parallelism.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-10 09:43:50 +01:00
jgrusewski
e4a911941b chore: remove benchmark comment, trigger Docker rebuild + warm cache
Removes temporary benchmark comment. Docker changes from previous commit
will trigger image rebuilds. This run serves as warm-cache benchmark
after cold-cache run completes.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-10 09:35:36 +01:00
jgrusewski
63d8619379 infra(docker): upgrade all images to Ubuntu 24.04 + CUDA 12.9
Align all 4 Docker images with local dev environment:
- ci-builder: CUDA 12.4.1/Ubuntu 22.04 → 12.9.1/Ubuntu 24.04
- ci-builder-cpu: Debian bookworm → Ubuntu 24.04 (+ explicit rustup)
- foxhunt-runtime: Debian bookworm → Ubuntu 24.04
- foxhunt-training-runtime: CUDA 12.6.3 → 12.9.1, nvrtc 12-6 → 12-9

All images now have glibc 2.39, matching the local build machine.
Binaries compiled locally or in CI will run in any of these containers
without GLIBC_2.3x version mismatch errors.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-10 09:34:47 +01:00
jgrusewski
2933e0f014 chore: trigger full CI compile (cold PVC sccache benchmark)
Temporary comment to trigger detect-changes for both services and
training compile steps. Will be removed after benchmarking.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-10 09:30:51 +01:00
jgrusewski
bcde1c2a95 infra(ci): replace MinIO sccache with local RWO PVCs
Switch sccache backend from S3 (MinIO) to persistent local storage.
Two 20Gi PVCs (sccache-cpu, sccache-cuda) on scw-bssd-retain eliminate
network I/O for cache reads/writes. Both compile steps always run on
the same node so RWO is sufficient.

Removes 13 S3 env vars per compile step, replaces with SCCACHE_DIR=/sccache.
Updates ci-pipeline-template, compile-and-train-template, kustomization.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-10 09:25:44 +01:00
jgrusewski
ff7d79b03a fix(gpu): resolve 3 pre-existing issues flagged by zen code review
1. PER duplicate index accumulation (HIGH): update_priorities_gpu()
   used index_add delta trick which accumulates deltas for duplicate
   indices, overshooting target priority. Now deduplicates the small
   indices tensor (batch_size=256, ~1KB) via HashMap before delta
   computation. Fast path (no dupes) reuses original tensors.

2. RegimeConditional silent experience drop (MEDIUM):
   insert_batch_tensors() returned Ok(()) for CPU replay buffers,
   silently discarding all GPU-collected experiences. Now converts
   tensors→Vec<Experience> via extracted helper (lazy, only on first
   CPU head encountered) and inserts into CPU buffer.

3. Unsafe direct indexing (LOW): gpu_experience_collector.rs used
   &states[s..e] in fallback paths, violating deny(indexing_slicing).
   Replaced with .get() safe bounds checking.

392/392 ml-dqn + 874/874 ml tests pass, 0 clippy warnings.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-10 00:53:16 +01:00
jgrusewski
5149c71444 feat(gpu): warp-cooperative DQN experience kernel for H100 (sm_90+)
Add warp-cooperative CUDA kernel for DQN experience collection that
exploits H100's improved warp shuffle throughput. On sm_90+, 32 threads
(1 warp) cooperate on each dot product via __shfl_xor_sync butterfly
reduction, cutting per-thread register pressure from ~5.5KB to ~200B
and enabling full occupancy on 132 SMs.

Key additions:
- 6 warp-cooperative device functions in common_device_functions.cuh:
  distributed/broadcast matvec (clean + NoisyNet), warp RMSNorm,
  warp_reduce_sum_all
- 4 TILE_LAYER_WARP_* macros using __syncwarp() for warp-level sync
- 3 warp forward passes: standard dueling, NoisyNet, C51 distributional
  (distributed heavy layers + broadcast atom layers for softmax)
- Full warp kernel: dqn_full_experience_kernel_warp with lane-0
  simulation, strided state scatter, warp-shuffle curiosity gather
- Host-side compute capability detection (sm_major >= 9) with
  automatic fallback to standard 256-thread kernel on older GPUs
- cuCtxSetLimit(STACK_SIZE, 16KB) for standard kernel safety

874/874 tests pass, 0 clippy warnings.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-10 00:44:59 +01:00
jgrusewski
f9b6cdb923 perf(gpu): H100 optimizations — dynamic NOISY_MAX_DIM, 32768 episodes, scratch aliasing
Three zen-validated optimizations for H100 GPU experience collection:

1. Dynamic NOISY_MAX_DIM: inject via NVRTC as max(state_dim, shared_h1,
   shared_h2). On H100 with SHARED_H1=512, noise now covers all 512 input
   dims instead of truncating at 256. Fixes exploration correctness bug.

2. MAX_EPISODES_LIMIT raised 4096→32768: H100 (132 SMs) can now run up to
   32768 episodes per kernel launch = 128 blocks = ~97% SM utilization
   (was 16 blocks = 12%). All 3 trainer.rs caps updated to match.

3. Scratch array aliasing: scratch_v and scratch_a now alias scratch1
   memory after shared layers complete, saving 1024 bytes per thread.
   Per-thread local memory reduced from ~6.5KB to ~5.5KB.

Zen validation confirms: kernel is safe and correct for H100 deployment.
Remaining architectural improvements (warp-cooperative matvec, batched
GEMM) deferred to future optimization phase.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-09 23:44:35 +01:00
jgrusewski
6b3bb26c05 fix(gpu-per): clamp priorities after index_add to prevent NaN weights
index_add with duplicate indices double-applies deltas — e.g. if index i
is sampled twice with delta -0.5, priority goes 1.0 + (-0.5) + (-0.5) = 0.0,
causing NaN in importance sampling weights. Post-clamp to [epsilon, 1e6]
ensures priorities stay positive.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-09 23:34:29 +01:00
jgrusewski
9502c1e22c fix(gpu): dynamic SHMEM_MAX_IN_DIM prevents shared memory overflow on H100
Root cause: SHMEM_MAX_IN_DIM was hardcoded to 256, but on H100 with
hidden_dim_base=2048, SHARED_H1=SHARED_H2=512. The cooperative tile
loader wrote 64×512=32768 floats into shmem allocated for 64×256=16384,
causing CUDA_ERROR_INVALID_VALUE on kernel launch.

Fixes:
- Make SHMEM_MAX_IN_DIM injectable via NVRTC #define (was hardcoded)
- Compute max(state_dim, shared_h1, shared_h2) and inject at compile time
- Use dynamic value for host-side shmem_bytes allocation
- Shrink eps_out[NOISY_MAX_DIM] → eps_out[SHMEM_TILE_ROWS] (saves 768B/thread)
- Force smoke tests to Device::Cpu (CUDA driver sensitivity to binary layout)

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-09 23:31:35 +01:00
jgrusewski
dd075ce02a fix(gpu): cap auto-scaled episode count at GPU buffer limit (4096)
optimal_n_episodes() returns 8192 on H100 (132 SMs) but the GPU
experience collector pre-allocates buffers for MAX_EPISODES_LIMIT=4096.
Exceeding that caused "Episode count exceeds allocated buffer size"
and silent fallback to CPU. Now all 3 auto-scaling sites clamp at 4096.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-09 22:43:34 +01:00
jgrusewski
6a8f8d7b50 fix(train): correct CLI args for hyperopt vs training binaries
hyperopt_baseline_rl uses --output (JSON file path), while
train_baseline_rl uses --output-dir (checkpoint directory).
Split arg generation by preset.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-09 22:39:02 +01:00
jgrusewski
88c83c17ce fix(train): use correct Argo Events label selector for CI workflow lookup
Argo Events labels workflows with events.argoproj.io/trigger, not
workflows.argoproj.io/workflow-template. Fixed find_ci_workflow() so
ci-train preset can actually find CI pipeline runs.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-09 22:35:16 +01:00
jgrusewski
d67f696971 feat(train): add ci-train preset for CI-gated auto-training
New `ci-train` preset monitors the Argo CI workflow for a commit,
polls until completion, then auto-submits the training job (defaults
to hyperopt). Prints Prometheus/log links during monitoring.

Usage: ./infra/scripts/train.sh ci-train --model dqn [--commit SHA]

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-09 22:30:59 +01:00
jgrusewski
befa7e2f55 perf(gpu): shared memory weight caching, GPU reductions, fix staging flush dim
CUDA kernel: cooperative_load_tile() + matvec_leaky_relu_shmem() cache weight
matrices in shared memory (SRAM), reducing HBM traffic 32x (1.76TB→55GB).
Block size increased to 256 threads for better SHMEM utilization.

Trainer hot-path: replace to_vec2 full-tensor CPU readbacks with GPU-native
flatten_all/min/max/mean_all (4 scalar readbacks). Merge two Q-diagnostic
forward passes into one. Deferred max_priority flush (1 GPU→CPU sync/epoch
instead of ~8/batch).

Fix StagedGpuBuffer::flush() dimension mismatch: used raw experience
state.len() (45) instead of gpu.state_dim() (48 aligned), causing
slice_scatter crash [500,48] vs [n,45] when CPU experiences flush into
GPU-aligned replay buffer.

Fix hardcoded state_dim=48 in optimal_n_episodes() — now uses dynamic
state_dim from network config (correct for OFI-enabled 56-dim states).

Dynamic episode auto-scaling guard: only auto-scale when gpu_n_episodes≥128
(production), respecting explicit test overrides (gpu_n_episodes=2).

Tests: 874 ml + 392 ml-dqn = 1266 pass, 0 fail.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-09 22:24:33 +01:00
jgrusewski
3865354d8f perf(gpu): dynamic episode scaling + BF16 cross-entropy stabilization
GPU experience collector:
- Add GpuHardwareInfo with SM count detection from device name lookup
  (H100=132, A100=108, L40S=142, RTX 4090=128, etc.)
- optimal_n_episodes() scales to GPU: sm_count × 2 warps × 32 threads,
  capped by 15% free VRAM budget, 256-aligned for block scheduling
- Remove hardcoded .min(256) cap in trainer — auto-scales from 128 to 8192
- MAX_EPISODES_LIMIT raised from 256 to 4096

Numerical stability:
- BF16 cross-entropy epsilon: 1e-8 → 1e-4 (below BF16 precision floor
  1e-8 rounds to zero, making log-stabilization a no-op → -inf → NaN)
- Add log_probs.clamp(-20, 0) guard against -inf × 0 = NaN in loss

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-09 21:16:47 +01:00
jgrusewski
356f1d295f fix(train): remove duplicate annotations block from rebase artifact
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-09 21:08:02 +01:00
jgrusewski
fb18e0f1dc fix(dqn): prevent BF16 softmax overflow, increase batch size for H100
BF16's 7-bit mantissa overflows on exp(50+) in softmax → Inf → NaN.
Cast distributional logits to F32 before softmax in both dueling and
rainbow network heads, then cast back to original dtype.

Increase default batch_size 128→1024 (standard), 128→256 (conservative),
512→2048 (aggressive) to saturate H100 tensor cores. AutoBatchSizer
already caps to VRAM ceiling for smaller GPUs.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-09 21:05:42 +01:00
jgrusewski
e89fbc2b4d fix(metrics): wire training pod scraping to Grafana dashboard
Three root causes for "no metrics" on the training dashboard:

1. Dashboard template variables ($model, $fold) sourced from
   foxhunt_training_current_epoch which isn't emitted until the first
   epoch completes. Switch to foxhunt_training_step which fires from
   step 500 onward.

2. train.sh pod template missing Prometheus annotations
   (prometheus.io/scrape, port, path). Also add the
   app.kubernetes.io/component label to the eval manifest so
   evaluation pods are discoverable too.

3. DQN and PPO trainers only called set_epoch() at the END of each
   epoch. Move the call to the TOP of the epoch loop so the gauge
   exists from the first training iteration.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-09 21:02:29 +01:00
jgrusewski
9d6a04ab4a fix(grafana): add 20 missing metric panels to training dashboard
Remove phantom foxhunt_training_total_epochs reference (never registered)
and replace with foxhunt_training_step from the recent metrics commit.
Add full coverage for all 52 Tier-1 Prometheus metrics across 5 new rows:
RL diagnostics (Q-overestimation, PPO entropy/KL/advantage), training
health (NaN/grad explosion/feature errors, checkpoint ops, data load
latency), epoch returns, supervised eval (accuracy/precision/recall/F1),
and hyperopt details (mode, trial epoch, failed trials).

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-09 21:02:29 +01:00
jgrusewski
47d4959588 fix(train): correct --output to --output-dir arg, add Prometheus annotations
The train_baseline_rl binary expects --output-dir, not --output.
Also adds prometheus.io scrape annotations to job pod templates
so training metrics appear in Grafana.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-09 20:53:36 +01:00
jgrusewski
ee12f1d9e4 perf(gpu): zero-sync gradient clipping — eliminate all GPU→CPU barriers from hot path
clip_grad_norm now returns a GPU-resident Tensor instead of (f64, f64),
keeping backward → clip → optimizer.step fully pipelined on GPU with
zero cuStreamSynchronize stalls. The unconditional multiply trick
(scale = min(max_norm/(norm+eps), 1.0)) avoids the conditional branch
that previously required reading the norm to CPU.

Key changes:
- gradient_utils::clip_grad_norm: return Tensor, unconditional GPU multiply
- Handle BF16 mixed-precision via per-gradient to_dtype cast
- adam.rs: backward_step_with_monitoring returns Tensor (zero sync)
- gradient_accumulation: delegate to gradient_utils (DRY, same GPU path)
- dqn.rs: single sync boundary after ALL GPU work queued

1746 tests passing (ml-core 286, ml-dqn 388, ml-ppo 198, ml 874).

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-09 20:26:45 +01:00
jgrusewski
db65eb56a5 perf(gpu): eliminate GPU→CPU sync barriers from training hot path
clip_grad_norm: accumulate squared norms on GPU-resident scalar, single
to_scalar() at end (was 16-20 per-param syncs per step × 2917 steps/epoch).

check_gradients_finite: same GPU-accumulation pattern, single sync.

gpu_replay_buffer sample_proportional/rank_based: generate random targets
via rand(0,1)*total_sum on GPU, normalize weights via broadcast_div
(eliminates 2 to_vec0 syncs per sample call).

gpu_replay_buffer update_priorities_gpu: replace CPU loop of 50 individual
slice_scatter calls with single batched index_add delta trick.

dqn NaN detection (every 500 steps): accumulate 3 NaN counts on GPU,
single to_scalar for total; detailed breakdown only if NaN found.

dqn dead neuron detection (every 1000 steps): accumulate dead count on
GPU-resident scalar (was to_vec0 per parameter tensor, ~16-20 syncs).

Net: ~22 GPU→CPU syncs + 50 micro-kernels per training step → 2 syncs.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-09 20:02:16 +01:00
jgrusewski
0355fb17bc perf(metrics): expose step-level DQN metrics to Prometheus gauges
Move Q-value stats, gradient norm, and training step from tracing::info!
(stdout-only, block-buffered, invisible mid-epoch) to Prometheus gauges
(atomic set, scrapable every 15s). Downgrade formatted step logs to
debug! level to eliminate allocation/serialization overhead in the hot path.

New gauge: foxhunt_training_step (current step within epoch)
Updated: foxhunt_training_q_value_mean/max, foxhunt_training_gradient_norm
now update every 500 steps instead of only at epoch boundaries.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-09 18:44:29 +01:00
jgrusewski
7751f7615a fix(ci): unblock CPU service builds and fix H100 BF16 regime classification
Three fixes validated by 20/20 hyperopt trials on H100 (zero OOM):

1. Workspace default-features: ml-core, ml-dqn, ml-ppo, ml-supervised
   workspace deps now have default-features=false. Prevents cudarc
   (which requires nvcc) from leaking into CPU service builds via
   Cargo feature unification. CI compile-services was failing with
   "Failed to execute nvcc: No such file or directory" (exit 101).

2. BF16 comparison fix: Candle's gt()/le() don't support BF16 operands.
   Cast ADX/CUSUM features to F32 before threshold comparison in
   regime classification. Previous approach (cast threshold to BF16)
   failed due to Candle broadcast_as reverting dtype.

3. CI pipeline: expand ML change detection to all 14 sub-crates,
   add component:compile labels for sccache network policy matching,
   bump training runtime to CUDA 12.6 + Ubuntu 24.04 (glibc 2.39).

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-09 18:11:04 +01:00
jgrusewski
379c0bee37 fix(dqn): cast F32 constants to input dtype for BF16 mixed precision on H100
On H100 with BF16 mixed precision enabled, Candle's comparison ops (gt,
lt, le, ge) and arithmetic ops require matching dtypes. F32 threshold
constants compared against BF16 state tensors caused every training step
to fail with "dtype mismatch in cmp, lhs: BF16, rhs: F32".

Fixes:
- regime_conditional.rs: ADX/CUSUM thresholds cast to adx.dtype()
- quantile_regression.rs: Huber kappa, zero, and half tensors match
  input dtype (both quantile_huber_loss variants)
- rainbow_network.rs: LeakyReLU slope and ELU alpha/one tensors cast
  to x.dtype()
- continuous_ppo.rs: Huber delta/half tensors cast to abs_diff.dtype()

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-09 17:13:46 +01:00
jgrusewski
fe223b3843 fix(dqn): eliminate OOM in hyperopt by gating GPU features on small GPUs
Three root causes fixed:
1. GPU PER + experience collector (cudarc) created dual CUDA allocator
   fragmentation on GPUs ≤8 GB, making training impossible even at
   batch_size=1. Added `use_gpu_replay_buffer` config flag; both GPU PER
   and experience collector now disabled when VRAM ≤8192 MB.

2. Search space bounds were WIDENED instead of capped — max_batch_size
   returning 4096 replaced the original 512 upper bound, and
   max_hidden_dim_base_full returning 3072 replaced the original 1024.
   Fixed with min() to only narrow, never widen.

3. VRAM estimator assumed GPU features always active, overcharging when
   they're disabled on small GPUs. Now conditional: when
   replay_buffer_capacity=0 (proxy for GPU PER disabled), collector/cudarc
   costs are zero and fragmentation multiplier drops from 3× to 1.5×.

Additional small-GPU guard: GPUs ≤8 GB get clamped search space
(batch≤128, hidden≤512, atoms≤51, buffer≤50K) to fit 3 regime heads
+ C51 + noisy nets + dueling in limited VRAM.

Validated: 7/7 trials complete on RTX 3050 Ti 4 GB, zero OOM, best trial
Sharpe 9.8 with 50.6% win rate. Previous runs had 100% OOM failure rate.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-09 16:35:38 +01:00
jgrusewski
06d0a8d618 Merge branch 'feature/dqn-branching'
perf(dqn): GPU-native branching action decomposition (eliminates cudaSync stall)

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-09 13:36:01 +01:00
jgrusewski
1b9f09fce1 perf(dqn): eliminate GPU→CPU roundtrip in branching action decomposition
Replace `actions_tensor.to_vec1()` → CPU loop → `Tensor::from_vec` with
GPU-native `decompose_actions_batch_gpu` using F32 affine+floor arithmetic.
Removes the last `cudaDeviceSynchronize` stall in the branching training
hot path.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-09 13:26:44 +01:00
jgrusewski
3a37454ea5 Merge branch 'feature/dqn-branching' 2026-03-09 13:17:50 +01:00