CUDA infrastructure in ml-core (cuda_compile, gpu/capabilities, gpu/l2_cache)
previously accessed cudarc types through candle_core::cuda_backend::cudarc --
a transitive re-export that coupled low-level CUDA driver calls to Candle.
Changes:
- Add cudarc 0.17 as direct optional dep (gated behind cuda feature)
- Replace all candle_core::cuda_backend::cudarc paths with direct cudarc::
- Generalize OOM detection to accept &dyn Debug (was &CandleError)
- Add generic computation_error_to_common_error() alongside legacy alias
Candle references: 163 -> 72 (remaining are autograd-dependent: Tensor,
Var, GradStore, VarBuilder -- cannot be replaced with cudarc raw ops)
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- Remove ALL #[cfg(feature = "cuda")] guards (~400+ occurrences)
- Remove ALL #[cfg_attr(not(feature = "cuda"), ignore)] test annotations (~250)
- Make cuda default feature in 9 ML crates (ml, ml-core, ml-dqn, ml-ppo, etc.)
- Convert nvrtc JIT compilation to precompiled nvcc (searchsorted, prefix_sum)
- Move compile_ptx_for_device() to ml-core for shared access
- Delete dead CPU code: multi_step.rs, self_supervised_pretraining.rs,
training_guard_gpu_tests.rs, CPU PER buffer paths, CPU Q-diagnostics
- Replace unwrap_or(Device::Cpu) with hard errors everywhere
- Remove dead is_cuda() else branches in DQN/PPO/hyperopt trainers
- Change config defaults from "cpu" to "cuda" (rainbow, tlob, pipeline)
- Port IQL value network to GPU kernel (5 CUDA entry points)
- Port HER goal relabeling to GPU kernel (warp-per-sample)
- Wire DSR GPU-to-CPU sync in training loop
- cfg!(feature = "cuda") → true in inference_validator
Zero warnings, zero errors across entire workspace.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- backtest_forward_kernel: dim_overrides must precede common_device_functions.cuh
which has #error guards requiring STATE_DIM/MARKET_DIM/PORTFOLIO_DIM to be
defined before inclusion. Experience collector already had correct ordering.
- Cap MAX_EPISODES from 8192→4096 (diminishing returns above 4096, wastes walltime)
- Cap trainer .min() from 0x8000 (32768) → 4096 to match
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- Cast input to weight dtype in DQN residual, rmsnorm, noisy_layers
- Set use_gpu=true in QNetworkConfig defaults and all config sites
- Resolve BF16 boundary mismatches in attention, curiosity, branching,
distributional_dueling across ml-dqn
- GPU-resident regime ops with BF16 boundary casts, eliminate .expect() in CUDA paths
- Eliminate all Device::Cpu fallbacks — GPU-only across 10 ML crates
- PPO: cast logits to F32 before softmax, cast batch tensors to training dtype
- Gradient collapse detection for RegimeConditionalDQN
- Wire halt_grad_collapse from CUDA guard kernel to halt training
- Dead neuron detection uses active network VarMap + squeeze factored readback
- Increment gradient_logging_step in GPU PER path
- Gradient collapse warmup guards use original buffer_size
- Cap training steps per epoch + tracing migration
- Replace Tensor::all() with sum_all() for pinned Candle compatibility
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
The GPU PER replay buffer had a hardcoded 4 GB MAX_BYTES limit that
rejected the auto-sizer's 10M-entry proposal on H100 (80 GB VRAM).
Now per_max_buffer_bytes() computes 20% of total VRAM (min 1 GB) and
flows through OptimalReplayConfig → DQNConfig → GpuReplayBufferConfig
so both subsystems agree on the budget.
Also fixes misleading regime detection log (indices 211/219 → 40/41)
and renames dqn_config_2025 → dqn_default_config.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Prevents experience collector output tensors from starving backward
graph and GPU PER on VRAM-constrained GPUs like RTX 3050 Ti 4GB.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
free_memory_mb is sampled after model/optimizer load, so subtract only
backward-graph + fragmentation + data upload headroom (400 MB), not the
full model/PER/data stack. Increase collector share to 50%.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- Staging/GPU replay buffer: truncate batch when it exceeds ring buffer
capacity (CPU fallback can flush 100K+ experiences at once)
- GpuExperienceCollector: compute shmem tile rows dynamically to stay
under 48 KB default limit instead of hardcoded 64-row constant
- Auto batch size: subtract concurrent VRAM consumers (~530 MB model +
optimizer + PER + data) before budgeting experience collector at 40%
- Hyperopt DQN adapter: GPU PER always on, include replay buffer in
VRAM estimate for small GPUs (was excluded assuming CPU-only PER)
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
GPU experience collector:
- Add GpuHardwareInfo with SM count detection from device name lookup
(H100=132, A100=108, L40S=142, RTX 4090=128, etc.)
- optimal_n_episodes() scales to GPU: sm_count × 2 warps × 32 threads,
capped by 15% free VRAM budget, 256-aligned for block scheduling
- Remove hardcoded .min(256) cap in trainer — auto-scales from 128 to 8192
- MAX_EPISODES_LIMIT raised from 256 to 4096
Numerical stability:
- BF16 cross-entropy epsilon: 1e-8 → 1e-4 (below BF16 precision floor
1e-8 rounds to zero, making log-stabilization a no-op → -inf → NaN)
- Add log_probs.clamp(-20, 0) guard against -inf × 0 = NaN in loss
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- Move memory_optimization/ (4.6K lines, 7 files) from ml to ml-core
- Move batch_size_resolver.rs to ml-core (now both deps in same crate)
- Add tempfile to ml-core dev-dependencies (qat tests)
- Re-export both modules from ml facade
- Keep checkpoint/ in ml (model_implementations.rs has model-specific deps)
185 ml-core tests + 2573 ml tests = 2758 total, 0 failures
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>