Per-component gradient clipping: - 2 new CUDA kernels (dqn_clipped_saxpy, dqn_clip_grad) - CQL gradient isolated into separate scratch buffer - Budget allocation: C51=70%, CQL=15%, IQN=10%, Ens=5% - Dynamic budget — inactive components' share goes to C51 Spectral norm extended to all 10 weight matrices: - Was trunk-only (W_s1, W_s2), now covers all heads - 16 u/v power iteration buffers, batched GOFF sync-back - Uniform sigma_max, end-to-end Lipschitz bounded Bug fixes from 9-agent audit: - Backtest episode reset: all 8 fields (was 5), max_equity updated before floor check - gradient_clip_norm unified: 10.0 everywhere (was 10.0 vs 1.0) - entropy_coefficient: Option<f64> → f64, single default 0.001 - Close price: unwrap_or(1.0) → direct indexing (no silent wrong rewards) - step_count: only increments during training (was incrementing on inference) - Quantile loss: single CUDA context (was 3×), silent fallbacks removed - MaybeNoisyLinear: single-variant enum removed → direct NoisyLinear - Budget fractions: exported as pub(crate) constants, referenced not hardcoded Monitoring: - Per-component gradient Prometheus gauges (C51 raw, combined) - Diagnostic logging every 1000 steps Tests: 7 new gradient budget tests, 1 gradient bounds smoke test All 488 tests pass (359 ml-dqn + 129 ml) Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
20 KiB
Spec C: Code Hygiene -- Dead Code Deletion, Hardcoded Value Elimination, Config Unification
Date: 2026-03-27 Spec: C of 3 (DQN CUDA training pipeline overhaul) Execution order: LAST -- after Spec A (BF16 + perf) and Spec B (bugs) Estimated LOC removed: ~400 Estimated hardcoded values eliminated: ~30
1. Motivation
The DQN CUDA training pipeline has accumulated dead code paths from iterative
development (WAVE/Phase labeling, abandoned BF16 per-struct mirrors, single-variant
enums). Hardcoded magic numbers -- CUDA block sizes, monitoring limits, expert demo
parameters, kernel #define constants -- are scattered across 12+ files, making the
codebase fragile and difficult to tune via hyperopt. This spec removes all of it in a
single coordinated pass.
2. Dead Code Deletion
2.1 BF16 Per-Struct Mirrors in gpu_weights.rs
File: crates/ml/src/cuda_pipeline/gpu_weights.rs
The flat-buffer BF16 path (lines 2140-2251 in gpu_dqn_trainer.rs) superseded
the per-struct BF16 mirror approach. The following are unused:
| Item | Lines | LOC |
|---|---|---|
DuelingWeightSetBf16 struct + alloc_from + sync_from_f32 |
473-593 | ~120 |
BranchingWeightSetBf16 struct + alloc_from + sync_from_f32 |
490-627 | ~55 |
CuriosityWeightSetBf16 struct + alloc_from + sync_from_f32 |
503-653 | ~45 |
KernelWeightPackBf16 struct + build() + DeviceRepr impl |
670-719+ | ~80 |
alloc_bf16_mirror() helper |
511-519 | ~10 |
convert_f32_to_bf16() helper |
527-552 | ~25 |
raw_device_ptr_bf16() helper |
657-661 | ~5 |
Total: ~340 LOC deleted.
Spec A will introduce a replacement flat-buffer BF16 conversion path. These per-struct types are dead regardless -- they are not called from any production code path today. Coordinate with Spec A to ensure the new BF16 conversion lands before these are removed.
Action: Delete all items listed above from gpu_weights.rs. Remove any pub use
re-exports from crates/ml/src/cuda_pipeline/mod.rs referencing these types.
2.2 Uncalled BF16 Methods in gpu_dqn_trainer.rs
File: crates/ml/src/cuda_pipeline/gpu_dqn_trainer.rs
| Method | Line | LOC | Reason |
|---|---|---|---|
ensure_bf16_mirrors() |
2144 | ~25 | Defined but never called; replaced by flat-buffer path |
sync_online_bf16() |
2237 | ~10 | Delegates to launch_bf16_convert_online() directly; Spec A replaces |
Both methods accept _online_d, _online_b, _target_d, _target_b parameters
(note the leading underscores -- the compiler already flagged them unused). The actual
BF16 conversion uses launch_segmented_bf16_convert() which operates on flat buffers.
Action: Delete both methods. Verify no call sites exist (grep confirms zero callers).
2.3 MaybeNoisyLinear Single-Variant Enum in branching.rs
File: crates/ml-dqn/src/branching.rs, lines 252-330
MaybeNoisyLinear is an enum with exactly one variant:
enum MaybeNoisyLinear {
Noisy(NoisyLinear),
}
Every match arm is Self::Noisy(n) => ... -- pure ceremony. The Deterministic
variant was removed when NoisyNet became mandatory.
Action: Replace MaybeNoisyLinear with NoisyLinear directly. Update all field
types in BranchingDuelingQNetwork (lines 345-350: value_fc, value_out,
branch_fcs, branch_outs) and the make_noisy_linear() factory (line 458) to
return NoisyLinear directly. Remove the Debug impl (line 326-330) since
NoisyLinear derives or implements Debug. Remove all wrapper delegation methods
(lines 256-324) -- callers invoke NoisyLinear methods directly. Update
copy_noisy_params() (line 1116-1121) to accept &mut NoisyLinear / &NoisyLinear.
2.4 Dead Code Path in config.rs
File: crates/ml/src/trainers/dqn/config.rs, lines 105-107
if !true {
return Ok(None);
}
This is unreachable. The condition !true is always false.
Action: Delete these three lines.
2.5 Empty reset_episode() in agent.rs
File: crates/ml-dqn/src/agent.rs, line 553
pub const fn reset_episode(&mut self) {
// Reset any per-episode tracking if needed
// Network weights and replay buffer are preserved
}
Empty no-op function. No callers.
Action: Delete the method and its doc comment (lines 552-556).
3. Hardcoded Value Elimination
3.1 CUDA Grid Dimension Helper
Problem: 50+ occurrences of ((n + 255) / 256) scattered across 13 files in
crates/ml/src/cuda_pipeline/. Each instance re-derives the ceiling division and
hardcodes the block dimension 256.
Files affected (sample of 50+ sites):
gpu_dqn_trainer.rs-- 20 occurrences (lines 751, 793, 840, 956, 995, 1036, 1081, 1131, 1202, 1312, 1527, 1633, 2190, 2550, 2956, 2957, 3422, 3590, 3684, 3721, 3983, 4402)gpu_backtest_evaluator.rs-- 5 occurrences (lines 684, 1074, 1298, 1621, 1661)gpu_iqn_head.rs-- 6 occurrences (lines 349, 511, 534, 583, 666, 719)gpu_experience_collector.rs-- 3 occurrences (lines 1065, 1178, 1743)gpu_her.rs-- 1 occurrence (line 409)gpu_curiosity_trainer.rs-- 3 occurrences (lines 130, 344, 387)gpu_attention.rs-- 2 occurrences (lines 346, 364)gpu_iql_trainer.rs-- 2 occurrences (lines 308, 328)batched_backward.rs-- 2 occurrences (lines 316, 757)batched_forward.rs-- 3 occurrences (lines 572, 605, 651)decision_transformer.rs-- 5 occurrences (lines 652, 753, 853, 878, 999)signal_adapter.rs-- 3 occurrences (lines 75, 133, 202)
Solution: Add to crates/ml/src/cuda_pipeline/mod.rs:
/// Standard CUDA block dimension for 1D kernels.
pub(crate) const CUDA_BLOCK_DIM: u32 = 256;
/// Compute the 1D grid dimension for `n` elements at `CUDA_BLOCK_DIM` threads per block.
#[inline(always)]
pub(crate) const fn cuda_grid_1d(n: usize) -> u32 {
((n + (CUDA_BLOCK_DIM as usize) - 1) / (CUDA_BLOCK_DIM as usize)) as u32
}
Action: Replace all 50+ manual grid computations with cuda_grid_1d(n).
Replace all hardcoded 256 in block_dim with CUDA_BLOCK_DIM.
3.2 Monitoring Limits
File: crates/ml/src/trainers/dqn/monitoring.rs
| Constant | Current value | Line(s) | Usage |
|---|---|---|---|
| Reward history limit | 1000 |
64 | reward_history.len() > 1000 |
| Reward drain amount | 500 |
65 | reward_history.drain(0..500) |
| Q-value history limit | 1000 |
130 | q_value_history.len() > 1000 |
| Q-value drain amount | 500 |
131 | q_value_history.drain(0..500) |
| Constant reward threshold | 0.01 |
162 | std < 0.01 |
| Max constant epochs | 5 |
171 | consecutive_constant_epochs >= 5 |
| Q-value divergence threshold | 1000.0 |
235 | (max_q - min_q).abs() > 1000.0 |
Solution: Add a MonitoringConfig struct:
pub(crate) struct MonitoringConfig {
pub history_capacity: usize, // default: 1000
pub history_drain_count: usize, // default: 500
pub constant_reward_std_threshold: f32, // default: 0.01
pub max_constant_epochs: usize, // default: 5
pub q_divergence_threshold: f64, // default: 1000.0
}
Pass MonitoringConfig into TrainingMonitor::new(). Default values are preserved
but become tunable. Alternatively, embed these in DQNHyperparameters if hyperopt
should search over them.
3.3 Factored Action Space Array Sizes
File: crates/ml/src/trainers/dqn/monitoring.rs
| Constant | Current value | Line(s) |
|---|---|---|
| Factored action count array | [usize; 81] |
22, 45 |
| Exposure action array | [usize; 9] |
15-17, 40 |
| Q-value arrays | [f64; 9], [usize; 9] |
16-17, 41 |
The array size 81 = 9 * 3 * 3 (exposure * order * urgency). The size 9 is
branch_0_size (exposure levels).
Solution: Make TrainingMonitor accept branch sizes at construction time and
use Vec instead of fixed arrays, or derive the constant from BranchingConfig:
impl TrainingMonitor {
pub(crate) fn new(epoch: usize, branch_sizes: &[usize]) -> Self {
let total_actions: usize = branch_sizes.iter().product();
let exposure_size = branch_sizes.first().copied().unwrap_or(9);
Self {
factored_action_counts: vec![0; total_actions],
action_counts: vec![0; exposure_size],
q_value_sums: vec![0.0; exposure_size],
q_value_counts: vec![0; exposure_size],
..
}
}
}
3.4 Expert Demo Parameters
File: crates/ml/src/trainers/dqn/expert_demos.rs
| Parameter | Current value | Line |
|---|---|---|
| Fast EMA period | 20 |
46 |
| Slow EMA period | 50 |
47 |
| ADX threshold | 25.0 |
48 |
| ADX period | 14 |
100, 200 |
These are already fields on ExpertDemoGenerator (lines 34-41) with defaults in
the Default impl (lines 43-51). The ADX period is the only one hardcoded at line
100 (let adx_period: usize = 14;) and line 200 (let period: usize = 14;).
Action: Add adx_period: usize field to ExpertDemoGenerator. Update
Default impl to set adx_period: 14. Replace the two hardcoded 14 values
with self.adx_period. Update ExpertDemoGenerator::new() signature to accept
adx_period.
3.5 CUDA Kernel #define Constants
Files:
crates/ml/src/cuda_pipeline/c51_loss_kernel.cucrates/ml/src/cuda_pipeline/mse_loss_kernel.cu
| Define | Value | File(s) | Issue |
|---|---|---|---|
MAX_PER_SAMPLE_CE |
50.0f |
c51_loss_kernel.cu:25 | Should derive from v_range |
LABEL_SMOOTHING_EPS |
0.01f |
c51_loss_kernel.cu:28 | Should be a kernel parameter |
MAX_ATOMS |
128 |
c51_loss_kernel.cu:32 | Should be validated against num_atoms at compile time |
MAX_BRANCH_SIZE |
9 |
c51_loss_kernel.cu:35, mse_loss_kernel.cu:19 | Should derive from branch config |
NUM_BRANCHES |
3 |
c51_loss_kernel.cu:36, mse_loss_kernel.cu:20 | Should be a kernel parameter |
Solution: Pass these as NVRTC compile-time defines rather than hardcoding them in
the .cu source. The Rust NVRTC compilation step already accepts --define-macro
options.
// In the NVRTC compilation step:
let defines = format!(
"-DMAX_PER_SAMPLE_CE={:.1}f -DLABEL_SMOOTHING_EPS={:.4}f \
-DMAX_ATOMS={} -DMAX_BRANCH_SIZE={} -DNUM_BRANCHES={}",
v_range * 2.0, // derived from config
label_smoothing_eps, // from DQNHyperparameters
num_atoms.next_power_of_two(), // validated upper bound
branch_sizes.iter().max().unwrap(), // max branch size
branch_sizes.len(), // number of branches
);
The .cu files already guard with #ifndef, so this is backwards-compatible.
Add a runtime assertion at trainer construction time:
assert!(
config.num_atoms <= MAX_ATOMS,
"num_atoms ({}) exceeds kernel MAX_ATOMS ({})",
config.num_atoms, MAX_ATOMS
);
3.6 Training Loop Dimensions
File: crates/ml/src/trainers/dqn/trainer/training_loop.rs
| Constant | Value | Line | Issue |
|---|---|---|---|
STATIC_MAX_BATCH_SIZE |
8192 |
(not found as named const, used implicitly) | Should be config |
| Auto-batch safety margin | 0.15 |
(not found as named const, used implicitly) | Should be config |
Note: Grep did not find STATIC_MAX_BATCH_SIZE or AutoBatchSizer as named items
in the current codebase. These may have been refactored or renamed. Verify before
implementation. If they exist as inline literals, extract to named constants in
GpuDqnTrainerConfig or a new BatchSizerConfig.
4. Stale Comment Cleanup
4.1 WAVE / Phase / BUG Labels
Remove all development-era labels from comments across these files:
| File | Count | Examples |
|---|---|---|
monitoring.rs |
4 | WAVE 9-11 production monitoring, WAVE P2: Episode length tracking |
constructor.rs |
15+ | WAVE 26 P2.2, WAVE 16, BUG #37 FIX, WAVE 23 P0 Fix #1, WAVE 1.1, WAVE 16S, WAVE 24, WAVE 26 P1, WAVE 3 FIX #2 |
agent.rs |
2 | BUG #7 FIX (lines 615, 669) |
Action: Replace with descriptive comments that explain what and why without referencing internal development phases. Examples:
| Before | After |
|---|---|
// WAVE 9-11 production monitoring |
// Q-value range tracking for divergence detection |
// BUG #37 FIX: Q-value clipping |
// Clip Q-values to prevent step-level explosions |
// WAVE 26 P2.2: Validate gradient_accumulation_steps > 0 |
// Validate gradient accumulation config |
// WAVE 1.1: Initialize triple barrier engine |
// Initialize triple barrier engine (max 1000 active trackers) |
// WAVE 3 FIX #2: Start with None |
// Start with None; collect feature stats during initial epochs |
4.2 MEMORY LEAK FIX Labels
Lines 61, 129 in monitoring.rs have // MEMORY LEAK FIX: comments. Once the
limits are extracted to MonitoringConfig (section 3.2), these comments become
self-documenting and should be shortened to explain the bounded-history invariant
rather than referencing the original bug.
5. Config Unification
5.1 entropy_coefficient: Remove Option Wrapper
File: crates/ml/src/trainers/dqn/config.rs, line 960
Currently pub entropy_coefficient: Option<f64>. Every usage site unwraps with
.unwrap_or(0.01) or .unwrap_or(0.001) -- and these defaults disagree:
| Call site | Default |
|---|---|
constructor.rs:314 |
0.01 |
constructor.rs:399 |
0.01 |
gpu_dqn_trainer.rs:213 |
0.001 |
DQNHyperparameters::default() config.rs:1431 |
Some(0.001) |
Action: Make the field pub entropy_coefficient: f64 with default 0.001.
Remove all .unwrap_or() calls. Update serialization/deserialization to handle
the schema migration (serde default attribute).
5.2 gradient_clip_norm: Unify Default
File: crates/ml/src/trainers/dqn/config.rs, line 867
Currently pub gradient_clip_norm: Option<f64> with inconsistent defaults:
| Call site | Default |
|---|---|
fused_training.rs:194 (main DQN) |
10.0 |
fused_training.rs:263 (IQN) |
1.0 |
fused_training.rs:298 (ensemble) |
1.0 |
constructor.rs:257 |
10.0 |
DQNHyperparameters::default() config.rs:1382 |
Some(10.0) |
Action: Make the field pub gradient_clip_norm: f64 with default 10.0.
For IQN and ensemble paths that need a different norm, add dedicated fields
(e.g., iqn_gradient_clip_norm, ensemble_gradient_clip_norm) or use the
same unified value. Remove all .unwrap_or() calls.
5.3 noisy_epsilon_floor: Wire Through Consistently
Currently pub noisy_epsilon_floor: Option<f64> with disagreeing defaults:
| Call site | Default |
|---|---|
constructor.rs:315 |
0.0 |
training_loop.rs:520 |
0.05 |
hyperopt/adapters/dqn.rs:476 |
0.10 |
DQNHyperparameters::default() config.rs:1432 |
Some(0.02) |
| Test assertions (tests.rs:582) | assert 0.10 |
Action: Make the field pub noisy_epsilon_floor: f64 with default 0.10
(matching the test assertions and hyperopt default). Remove all .unwrap_or()
calls and ensure a single source of truth.
5.4 count_bonus_coefficient: Wire Through Consistently
Currently pub count_bonus_coefficient: Option<f64> with defaults:
| Call site | Default |
|---|---|
constructor.rs:316-317 |
0.0 |
gpu_experience_collector.rs:278 |
0.0 |
DQNHyperparameters::default() config.rs:1433 |
Some(0.1) |
| Test assertions (tests.rs:600) | assert 0.05 |
Action: Make the field pub count_bonus_coefficient: f64 with default 0.05
(matching test expectations). Remove all .unwrap_or() calls.
6. Files Changed Summary
| File | Deletions | Additions | Net |
|---|---|---|---|
crates/ml/src/cuda_pipeline/gpu_weights.rs |
~340 LOC | 0 | -340 |
crates/ml/src/cuda_pipeline/gpu_dqn_trainer.rs |
~35 LOC methods | ~5 LOC (use helper) | -30 |
crates/ml/src/cuda_pipeline/mod.rs |
0 | ~10 LOC (const + fn) | +10 |
crates/ml-dqn/src/branching.rs |
~80 LOC (enum + impls) | ~5 LOC (type aliases) | -75 |
crates/ml-dqn/src/agent.rs |
~5 LOC | 0 | -5 |
crates/ml/src/trainers/dqn/config.rs |
~3 LOC dead path + Option wrappers | ~10 LOC (defaults) | +7 |
crates/ml/src/trainers/dqn/monitoring.rs |
~5 LOC (stale comments) | ~15 LOC (MonitoringConfig) | +10 |
crates/ml/src/trainers/dqn/expert_demos.rs |
0 | ~5 LOC (adx_period field) | +5 |
crates/ml/src/trainers/dqn/trainer/constructor.rs |
~15 LOC (stale comments) | ~5 LOC (simplified unwraps) | -10 |
crates/ml/src/trainers/dqn/fused_training.rs |
0 | ~5 LOC (remove unwrap_or) | +5 |
crates/ml/src/cuda_pipeline/c51_loss_kernel.cu |
0 | ~5 LOC (guard comments) | +5 |
crates/ml/src/cuda_pipeline/mse_loss_kernel.cu |
0 | ~3 LOC (guard comments) | +3 |
13 files for cuda_grid_1d migration |
~50 LOC (manual math) | ~50 LOC (helper calls) | 0 |
| TOTAL | ~530 | ~120 | -415 |
7. Success Criteria
-
Zero dead code warnings. Compile with
#[deny(dead_code)]ongpu_weights.rs,branching.rs,agent.rs, andconfig.rs-- zero warnings. -
Zero magic numbers in kernel launches. Every
grid_dimcomputation incrates/ml/src/cuda_pipeline/usescuda_grid_1d()orCUDA_BLOCK_DIM. Grep for+ 255) / 256returns zero results. -
All config values traceable. Every training parameter flows from
DQNHyperparametersthroughGpuDqnTrainerConfigto the GPU kernel. No.unwrap_or()calls with differing defaults for the same field. -
No stale labels. Grep for
WAVE \d,Phase \d,BUG #\dacrosscrates/ml/andcrates/ml-dqn/returns zero results. -
All 129+ tests pass.
SQLX_OFFLINE=true cargo test -p ml --libandSQLX_OFFLINE=true cargo test -p ml-dqn --libboth green. -
Clean compile.
SQLX_OFFLINE=true cargo check --workspaceproduces zero warnings related to unused imports, dead code, or unreachable patterns from these changes.
8. Non-Goals
- Performance optimization -- covered by Spec A (BF16, cuBLAS SGEMM, CUDA Graphs).
- Bug fixes -- covered by Spec B.
- New features -- no new training capabilities are introduced.
- Kernel algorithm changes --
#definevalues become parameters but the kernel algorithms are unchanged.
9. Risks and Mitigations
| Risk | Likelihood | Impact | Mitigation |
|---|---|---|---|
| Spec A lands new BF16 types that reference deleted items | Medium | Build break | Execute Spec C after Spec A merges; coordinate on gpu_weights.rs |
Removing Option wrappers breaks JSON deserialization of saved trials |
Medium | Hyperopt regression | Add #[serde(default)] on all unwrapped fields; migration test |
MaybeNoisyLinear removal misses a usage site |
Low | Build break | Grep for MaybeNoisyLinear across entire workspace before deletion |
NVRTC --define-macro changes break kernel compilation on older CUDA |
Low | Runtime crash | Gate behind feature flag; keep #ifndef guards as fallback |
10. Execution Checklist
This is the recommended execution order within Spec C. Each step is independently testable.
- C.1 Add
cuda_grid_1d()andCUDA_BLOCK_DIMtocuda_pipeline/mod.rs - C.2 Migrate all 50+ grid computations to use the helper (13 files)
- C.3 Delete BF16 per-struct types from
gpu_weights.rs(~340 LOC) - C.4 Delete
ensure_bf16_mirrors()andsync_online_bf16()fromgpu_dqn_trainer.rs - C.5 Replace
MaybeNoisyLinearwithNoisyLinearinbranching.rs - C.6 Delete
if !truedead path fromconfig.rs - C.7 Delete
reset_episode()fromagent.rs - C.8 Extract monitoring limits to
MonitoringConfig - C.9 Add
adx_periodfield toExpertDemoGenerator - C.10 Pass CUDA kernel defines via NVRTC compile-time options
- C.11 Make
entropy_coefficientnon-optional with default0.001 - C.12 Make
gradient_clip_normnon-optional with default10.0 - C.13 Make
noisy_epsilon_floornon-optional with default0.10 - C.14 Make
count_bonus_coefficientnon-optional with default0.05 - C.15 Clean all WAVE/Phase/BUG labels from comments
- C.16 Make
TrainingMonitoraccept dynamic branch sizes - C.17 Run full test suite and verify zero warnings