Files
foxhunt/docs/superpowers/specs/2026-03-27-code-hygiene-design.md
jgrusewski be8f3310d5 feat: gradient stability + 9-agent audit bug fixes
Per-component gradient clipping:
- 2 new CUDA kernels (dqn_clipped_saxpy, dqn_clip_grad)
- CQL gradient isolated into separate scratch buffer
- Budget allocation: C51=70%, CQL=15%, IQN=10%, Ens=5%
- Dynamic budget — inactive components' share goes to C51

Spectral norm extended to all 10 weight matrices:
- Was trunk-only (W_s1, W_s2), now covers all heads
- 16 u/v power iteration buffers, batched GOFF sync-back
- Uniform sigma_max, end-to-end Lipschitz bounded

Bug fixes from 9-agent audit:
- Backtest episode reset: all 8 fields (was 5), max_equity updated before floor check
- gradient_clip_norm unified: 10.0 everywhere (was 10.0 vs 1.0)
- entropy_coefficient: Option<f64> → f64, single default 0.001
- Close price: unwrap_or(1.0) → direct indexing (no silent wrong rewards)
- step_count: only increments during training (was incrementing on inference)
- Quantile loss: single CUDA context (was 3×), silent fallbacks removed
- MaybeNoisyLinear: single-variant enum removed → direct NoisyLinear
- Budget fractions: exported as pub(crate) constants, referenced not hardcoded

Monitoring:
- Per-component gradient Prometheus gauges (C51 raw, combined)
- Diagnostic logging every 1000 steps

Tests: 7 new gradient budget tests, 1 gradient bounds smoke test
All 488 tests pass (359 ml-dqn + 129 ml)

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-27 19:52:44 +01:00

491 lines
20 KiB
Markdown

# Spec C: Code Hygiene -- Dead Code Deletion, Hardcoded Value Elimination, Config Unification
**Date:** 2026-03-27
**Spec:** C of 3 (DQN CUDA training pipeline overhaul)
**Execution order:** LAST -- after Spec A (BF16 + perf) and Spec B (bugs)
**Estimated LOC removed:** ~400
**Estimated hardcoded values eliminated:** ~30
---
## 1. Motivation
The DQN CUDA training pipeline has accumulated dead code paths from iterative
development (WAVE/Phase labeling, abandoned BF16 per-struct mirrors, single-variant
enums). Hardcoded magic numbers -- CUDA block sizes, monitoring limits, expert demo
parameters, kernel `#define` constants -- are scattered across 12+ files, making the
codebase fragile and difficult to tune via hyperopt. This spec removes all of it in a
single coordinated pass.
---
## 2. Dead Code Deletion
### 2.1 BF16 Per-Struct Mirrors in `gpu_weights.rs`
**File:** `crates/ml/src/cuda_pipeline/gpu_weights.rs`
The flat-buffer BF16 path (lines 2140-2251 in `gpu_dqn_trainer.rs`) superseded
the per-struct BF16 mirror approach. The following are unused:
| Item | Lines | LOC |
|------|-------|-----|
| `DuelingWeightSetBf16` struct + `alloc_from` + `sync_from_f32` | 473-593 | ~120 |
| `BranchingWeightSetBf16` struct + `alloc_from` + `sync_from_f32` | 490-627 | ~55 |
| `CuriosityWeightSetBf16` struct + `alloc_from` + `sync_from_f32` | 503-653 | ~45 |
| `KernelWeightPackBf16` struct + `build()` + `DeviceRepr` impl | 670-719+ | ~80 |
| `alloc_bf16_mirror()` helper | 511-519 | ~10 |
| `convert_f32_to_bf16()` helper | 527-552 | ~25 |
| `raw_device_ptr_bf16()` helper | 657-661 | ~5 |
**Total: ~340 LOC deleted.**
Spec A will introduce a replacement flat-buffer BF16 conversion path. These per-struct
types are dead regardless -- they are not called from any production code path today.
Coordinate with Spec A to ensure the new BF16 conversion lands before these are removed.
**Action:** Delete all items listed above from `gpu_weights.rs`. Remove any `pub use`
re-exports from `crates/ml/src/cuda_pipeline/mod.rs` referencing these types.
### 2.2 Uncalled BF16 Methods in `gpu_dqn_trainer.rs`
**File:** `crates/ml/src/cuda_pipeline/gpu_dqn_trainer.rs`
| Method | Line | LOC | Reason |
|--------|------|-----|--------|
| `ensure_bf16_mirrors()` | 2144 | ~25 | Defined but never called; replaced by flat-buffer path |
| `sync_online_bf16()` | 2237 | ~10 | Delegates to `launch_bf16_convert_online()` directly; Spec A replaces |
Both methods accept `_online_d`, `_online_b`, `_target_d`, `_target_b` parameters
(note the leading underscores -- the compiler already flagged them unused). The actual
BF16 conversion uses `launch_segmented_bf16_convert()` which operates on flat buffers.
**Action:** Delete both methods. Verify no call sites exist (grep confirms zero callers).
### 2.3 `MaybeNoisyLinear` Single-Variant Enum in `branching.rs`
**File:** `crates/ml-dqn/src/branching.rs`, lines 252-330
`MaybeNoisyLinear` is an enum with exactly one variant:
```rust
enum MaybeNoisyLinear {
Noisy(NoisyLinear),
}
```
Every match arm is `Self::Noisy(n) => ...` -- pure ceremony. The `Deterministic`
variant was removed when NoisyNet became mandatory.
**Action:** Replace `MaybeNoisyLinear` with `NoisyLinear` directly. Update all field
types in `BranchingDuelingQNetwork` (lines 345-350: `value_fc`, `value_out`,
`branch_fcs`, `branch_outs`) and the `make_noisy_linear()` factory (line 458) to
return `NoisyLinear` directly. Remove the `Debug` impl (line 326-330) since
`NoisyLinear` derives or implements `Debug`. Remove all wrapper delegation methods
(lines 256-324) -- callers invoke `NoisyLinear` methods directly. Update
`copy_noisy_params()` (line 1116-1121) to accept `&mut NoisyLinear` / `&NoisyLinear`.
### 2.4 Dead Code Path in `config.rs`
**File:** `crates/ml/src/trainers/dqn/config.rs`, lines 105-107
```rust
if !true {
return Ok(None);
}
```
This is unreachable. The condition `!true` is always `false`.
**Action:** Delete these three lines.
### 2.5 Empty `reset_episode()` in `agent.rs`
**File:** `crates/ml-dqn/src/agent.rs`, line 553
```rust
pub const fn reset_episode(&mut self) {
// Reset any per-episode tracking if needed
// Network weights and replay buffer are preserved
}
```
Empty no-op function. No callers.
**Action:** Delete the method and its doc comment (lines 552-556).
---
## 3. Hardcoded Value Elimination
### 3.1 CUDA Grid Dimension Helper
**Problem:** 50+ occurrences of `((n + 255) / 256)` scattered across 13 files in
`crates/ml/src/cuda_pipeline/`. Each instance re-derives the ceiling division and
hardcodes the block dimension 256.
**Files affected (sample of 50+ sites):**
- `gpu_dqn_trainer.rs` -- 20 occurrences (lines 751, 793, 840, 956, 995, 1036, 1081, 1131, 1202, 1312, 1527, 1633, 2190, 2550, 2956, 2957, 3422, 3590, 3684, 3721, 3983, 4402)
- `gpu_backtest_evaluator.rs` -- 5 occurrences (lines 684, 1074, 1298, 1621, 1661)
- `gpu_iqn_head.rs` -- 6 occurrences (lines 349, 511, 534, 583, 666, 719)
- `gpu_experience_collector.rs` -- 3 occurrences (lines 1065, 1178, 1743)
- `gpu_her.rs` -- 1 occurrence (line 409)
- `gpu_curiosity_trainer.rs` -- 3 occurrences (lines 130, 344, 387)
- `gpu_attention.rs` -- 2 occurrences (lines 346, 364)
- `gpu_iql_trainer.rs` -- 2 occurrences (lines 308, 328)
- `batched_backward.rs` -- 2 occurrences (lines 316, 757)
- `batched_forward.rs` -- 3 occurrences (lines 572, 605, 651)
- `decision_transformer.rs` -- 5 occurrences (lines 652, 753, 853, 878, 999)
- `signal_adapter.rs` -- 3 occurrences (lines 75, 133, 202)
**Solution:** Add to `crates/ml/src/cuda_pipeline/mod.rs`:
```rust
/// Standard CUDA block dimension for 1D kernels.
pub(crate) const CUDA_BLOCK_DIM: u32 = 256;
/// Compute the 1D grid dimension for `n` elements at `CUDA_BLOCK_DIM` threads per block.
#[inline(always)]
pub(crate) const fn cuda_grid_1d(n: usize) -> u32 {
((n + (CUDA_BLOCK_DIM as usize) - 1) / (CUDA_BLOCK_DIM as usize)) as u32
}
```
**Action:** Replace all 50+ manual grid computations with `cuda_grid_1d(n)`.
Replace all hardcoded `256` in `block_dim` with `CUDA_BLOCK_DIM`.
### 3.2 Monitoring Limits
**File:** `crates/ml/src/trainers/dqn/monitoring.rs`
| Constant | Current value | Line(s) | Usage |
|----------|--------------|---------|-------|
| Reward history limit | `1000` | 64 | `reward_history.len() > 1000` |
| Reward drain amount | `500` | 65 | `reward_history.drain(0..500)` |
| Q-value history limit | `1000` | 130 | `q_value_history.len() > 1000` |
| Q-value drain amount | `500` | 131 | `q_value_history.drain(0..500)` |
| Constant reward threshold | `0.01` | 162 | `std < 0.01` |
| Max constant epochs | `5` | 171 | `consecutive_constant_epochs >= 5` |
| Q-value divergence threshold | `1000.0` | 235 | `(max_q - min_q).abs() > 1000.0` |
**Solution:** Add a `MonitoringConfig` struct:
```rust
pub(crate) struct MonitoringConfig {
pub history_capacity: usize, // default: 1000
pub history_drain_count: usize, // default: 500
pub constant_reward_std_threshold: f32, // default: 0.01
pub max_constant_epochs: usize, // default: 5
pub q_divergence_threshold: f64, // default: 1000.0
}
```
Pass `MonitoringConfig` into `TrainingMonitor::new()`. Default values are preserved
but become tunable. Alternatively, embed these in `DQNHyperparameters` if hyperopt
should search over them.
### 3.3 Factored Action Space Array Sizes
**File:** `crates/ml/src/trainers/dqn/monitoring.rs`
| Constant | Current value | Line(s) |
|----------|--------------|---------|
| Factored action count array | `[usize; 81]` | 22, 45 |
| Exposure action array | `[usize; 9]` | 15-17, 40 |
| Q-value arrays | `[f64; 9]`, `[usize; 9]` | 16-17, 41 |
The array size 81 = 9 * 3 * 3 (exposure * order * urgency). The size 9 is
`branch_0_size` (exposure levels).
**Solution:** Make `TrainingMonitor` accept branch sizes at construction time and
use `Vec` instead of fixed arrays, or derive the constant from `BranchingConfig`:
```rust
impl TrainingMonitor {
pub(crate) fn new(epoch: usize, branch_sizes: &[usize]) -> Self {
let total_actions: usize = branch_sizes.iter().product();
let exposure_size = branch_sizes.first().copied().unwrap_or(9);
Self {
factored_action_counts: vec![0; total_actions],
action_counts: vec![0; exposure_size],
q_value_sums: vec![0.0; exposure_size],
q_value_counts: vec![0; exposure_size],
..
}
}
}
```
### 3.4 Expert Demo Parameters
**File:** `crates/ml/src/trainers/dqn/expert_demos.rs`
| Parameter | Current value | Line |
|-----------|--------------|------|
| Fast EMA period | `20` | 46 |
| Slow EMA period | `50` | 47 |
| ADX threshold | `25.0` | 48 |
| ADX period | `14` | 100, 200 |
These are already fields on `ExpertDemoGenerator` (lines 34-41) with defaults in
the `Default` impl (lines 43-51). The ADX period is the only one hardcoded at line
100 (`let adx_period: usize = 14;`) and line 200 (`let period: usize = 14;`).
**Action:** Add `adx_period: usize` field to `ExpertDemoGenerator`. Update
`Default` impl to set `adx_period: 14`. Replace the two hardcoded `14` values
with `self.adx_period`. Update `ExpertDemoGenerator::new()` signature to accept
`adx_period`.
### 3.5 CUDA Kernel `#define` Constants
**Files:**
- `crates/ml/src/cuda_pipeline/c51_loss_kernel.cu`
- `crates/ml/src/cuda_pipeline/mse_loss_kernel.cu`
| Define | Value | File(s) | Issue |
|--------|-------|---------|-------|
| `MAX_PER_SAMPLE_CE` | `50.0f` | c51_loss_kernel.cu:25 | Should derive from `v_range` |
| `LABEL_SMOOTHING_EPS` | `0.01f` | c51_loss_kernel.cu:28 | Should be a kernel parameter |
| `MAX_ATOMS` | `128` | c51_loss_kernel.cu:32 | Should be validated against `num_atoms` at compile time |
| `MAX_BRANCH_SIZE` | `9` | c51_loss_kernel.cu:35, mse_loss_kernel.cu:19 | Should derive from branch config |
| `NUM_BRANCHES` | `3` | c51_loss_kernel.cu:36, mse_loss_kernel.cu:20 | Should be a kernel parameter |
**Solution:** Pass these as NVRTC compile-time defines rather than hardcoding them in
the `.cu` source. The Rust NVRTC compilation step already accepts `--define-macro`
options.
```rust
// In the NVRTC compilation step:
let defines = format!(
"-DMAX_PER_SAMPLE_CE={:.1}f -DLABEL_SMOOTHING_EPS={:.4}f \
-DMAX_ATOMS={} -DMAX_BRANCH_SIZE={} -DNUM_BRANCHES={}",
v_range * 2.0, // derived from config
label_smoothing_eps, // from DQNHyperparameters
num_atoms.next_power_of_two(), // validated upper bound
branch_sizes.iter().max().unwrap(), // max branch size
branch_sizes.len(), // number of branches
);
```
The `.cu` files already guard with `#ifndef`, so this is backwards-compatible.
Add a runtime assertion at trainer construction time:
```rust
assert!(
config.num_atoms <= MAX_ATOMS,
"num_atoms ({}) exceeds kernel MAX_ATOMS ({})",
config.num_atoms, MAX_ATOMS
);
```
### 3.6 Training Loop Dimensions
**File:** `crates/ml/src/trainers/dqn/trainer/training_loop.rs`
| Constant | Value | Line | Issue |
|----------|-------|------|-------|
| `STATIC_MAX_BATCH_SIZE` | `8192` | (not found as named const, used implicitly) | Should be config |
| Auto-batch safety margin | `0.15` | (not found as named const, used implicitly) | Should be config |
Note: Grep did not find `STATIC_MAX_BATCH_SIZE` or `AutoBatchSizer` as named items
in the current codebase. These may have been refactored or renamed. Verify before
implementation. If they exist as inline literals, extract to named constants in
`GpuDqnTrainerConfig` or a new `BatchSizerConfig`.
---
## 4. Stale Comment Cleanup
### 4.1 WAVE / Phase / BUG Labels
Remove all development-era labels from comments across these files:
| File | Count | Examples |
|------|-------|---------|
| `monitoring.rs` | 4 | `WAVE 9-11 production monitoring`, `WAVE P2: Episode length tracking` |
| `constructor.rs` | 15+ | `WAVE 26 P2.2`, `WAVE 16`, `BUG #37 FIX`, `WAVE 23 P0 Fix #1`, `WAVE 1.1`, `WAVE 16S`, `WAVE 24`, `WAVE 26 P1`, `WAVE 3 FIX #2` |
| `agent.rs` | 2 | `BUG #7 FIX` (lines 615, 669) |
**Action:** Replace with descriptive comments that explain *what* and *why* without
referencing internal development phases. Examples:
| Before | After |
|--------|-------|
| `// WAVE 9-11 production monitoring` | `// Q-value range tracking for divergence detection` |
| `// BUG #37 FIX: Q-value clipping` | `// Clip Q-values to prevent step-level explosions` |
| `// WAVE 26 P2.2: Validate gradient_accumulation_steps > 0` | `// Validate gradient accumulation config` |
| `// WAVE 1.1: Initialize triple barrier engine` | `// Initialize triple barrier engine (max 1000 active trackers)` |
| `// WAVE 3 FIX #2: Start with None` | `// Start with None; collect feature stats during initial epochs` |
### 4.2 MEMORY LEAK FIX Labels
Lines 61, 129 in `monitoring.rs` have `// MEMORY LEAK FIX:` comments. Once the
limits are extracted to `MonitoringConfig` (section 3.2), these comments become
self-documenting and should be shortened to explain the bounded-history invariant
rather than referencing the original bug.
---
## 5. Config Unification
### 5.1 `entropy_coefficient`: Remove `Option` Wrapper
**File:** `crates/ml/src/trainers/dqn/config.rs`, line 960
Currently `pub entropy_coefficient: Option<f64>`. Every usage site unwraps with
`.unwrap_or(0.01)` or `.unwrap_or(0.001)` -- and these defaults **disagree**:
| Call site | Default |
|-----------|---------|
| `constructor.rs:314` | `0.01` |
| `constructor.rs:399` | `0.01` |
| `gpu_dqn_trainer.rs:213` | `0.001` |
| `DQNHyperparameters::default()` config.rs:1431 | `Some(0.001)` |
**Action:** Make the field `pub entropy_coefficient: f64` with default `0.001`.
Remove all `.unwrap_or()` calls. Update serialization/deserialization to handle
the schema migration (serde default attribute).
### 5.2 `gradient_clip_norm`: Unify Default
**File:** `crates/ml/src/trainers/dqn/config.rs`, line 867
Currently `pub gradient_clip_norm: Option<f64>` with inconsistent defaults:
| Call site | Default |
|-----------|---------|
| `fused_training.rs:194` (main DQN) | `10.0` |
| `fused_training.rs:263` (IQN) | `1.0` |
| `fused_training.rs:298` (ensemble) | `1.0` |
| `constructor.rs:257` | `10.0` |
| `DQNHyperparameters::default()` config.rs:1382 | `Some(10.0)` |
**Action:** Make the field `pub gradient_clip_norm: f64` with default `10.0`.
For IQN and ensemble paths that need a different norm, add dedicated fields
(e.g., `iqn_gradient_clip_norm`, `ensemble_gradient_clip_norm`) or use the
same unified value. Remove all `.unwrap_or()` calls.
### 5.3 `noisy_epsilon_floor`: Wire Through Consistently
Currently `pub noisy_epsilon_floor: Option<f64>` with disagreeing defaults:
| Call site | Default |
|-----------|---------|
| `constructor.rs:315` | `0.0` |
| `training_loop.rs:520` | `0.05` |
| `hyperopt/adapters/dqn.rs:476` | `0.10` |
| `DQNHyperparameters::default()` config.rs:1432 | `Some(0.02)` |
| Test assertions (tests.rs:582) | assert `0.10` |
**Action:** Make the field `pub noisy_epsilon_floor: f64` with default `0.10`
(matching the test assertions and hyperopt default). Remove all `.unwrap_or()`
calls and ensure a single source of truth.
### 5.4 `count_bonus_coefficient`: Wire Through Consistently
Currently `pub count_bonus_coefficient: Option<f64>` with defaults:
| Call site | Default |
|-----------|---------|
| `constructor.rs:316-317` | `0.0` |
| `gpu_experience_collector.rs:278` | `0.0` |
| `DQNHyperparameters::default()` config.rs:1433 | `Some(0.1)` |
| Test assertions (tests.rs:600) | assert `0.05` |
**Action:** Make the field `pub count_bonus_coefficient: f64` with default `0.05`
(matching test expectations). Remove all `.unwrap_or()` calls.
---
## 6. Files Changed Summary
| File | Deletions | Additions | Net |
|------|-----------|-----------|-----|
| `crates/ml/src/cuda_pipeline/gpu_weights.rs` | ~340 LOC | 0 | -340 |
| `crates/ml/src/cuda_pipeline/gpu_dqn_trainer.rs` | ~35 LOC methods | ~5 LOC (use helper) | -30 |
| `crates/ml/src/cuda_pipeline/mod.rs` | 0 | ~10 LOC (const + fn) | +10 |
| `crates/ml-dqn/src/branching.rs` | ~80 LOC (enum + impls) | ~5 LOC (type aliases) | -75 |
| `crates/ml-dqn/src/agent.rs` | ~5 LOC | 0 | -5 |
| `crates/ml/src/trainers/dqn/config.rs` | ~3 LOC dead path + Option wrappers | ~10 LOC (defaults) | +7 |
| `crates/ml/src/trainers/dqn/monitoring.rs` | ~5 LOC (stale comments) | ~15 LOC (MonitoringConfig) | +10 |
| `crates/ml/src/trainers/dqn/expert_demos.rs` | 0 | ~5 LOC (adx_period field) | +5 |
| `crates/ml/src/trainers/dqn/trainer/constructor.rs` | ~15 LOC (stale comments) | ~5 LOC (simplified unwraps) | -10 |
| `crates/ml/src/trainers/dqn/fused_training.rs` | 0 | ~5 LOC (remove unwrap_or) | +5 |
| `crates/ml/src/cuda_pipeline/c51_loss_kernel.cu` | 0 | ~5 LOC (guard comments) | +5 |
| `crates/ml/src/cuda_pipeline/mse_loss_kernel.cu` | 0 | ~3 LOC (guard comments) | +3 |
| 13 files for `cuda_grid_1d` migration | ~50 LOC (manual math) | ~50 LOC (helper calls) | 0 |
| **TOTAL** | **~530** | **~120** | **-415** |
---
## 7. Success Criteria
1. **Zero dead code warnings.** Compile with `#[deny(dead_code)]` on
`gpu_weights.rs`, `branching.rs`, `agent.rs`, and `config.rs` -- zero warnings.
2. **Zero magic numbers in kernel launches.** Every `grid_dim` computation in
`crates/ml/src/cuda_pipeline/` uses `cuda_grid_1d()` or `CUDA_BLOCK_DIM`.
Grep for `+ 255) / 256` returns zero results.
3. **All config values traceable.** Every training parameter flows from
`DQNHyperparameters` through `GpuDqnTrainerConfig` to the GPU kernel. No
`.unwrap_or()` calls with differing defaults for the same field.
4. **No stale labels.** Grep for `WAVE \d`, `Phase \d`, `BUG #\d` across
`crates/ml/` and `crates/ml-dqn/` returns zero results.
5. **All 129+ tests pass.** `SQLX_OFFLINE=true cargo test -p ml --lib` and
`SQLX_OFFLINE=true cargo test -p ml-dqn --lib` both green.
6. **Clean compile.** `SQLX_OFFLINE=true cargo check --workspace` produces zero
warnings related to unused imports, dead code, or unreachable patterns from
these changes.
---
## 8. Non-Goals
- **Performance optimization** -- covered by Spec A (BF16, cuBLAS SGEMM, CUDA Graphs).
- **Bug fixes** -- covered by Spec B.
- **New features** -- no new training capabilities are introduced.
- **Kernel algorithm changes** -- `#define` values become parameters but the
kernel algorithms are unchanged.
---
## 9. Risks and Mitigations
| Risk | Likelihood | Impact | Mitigation |
|------|-----------|--------|------------|
| Spec A lands new BF16 types that reference deleted items | Medium | Build break | Execute Spec C after Spec A merges; coordinate on `gpu_weights.rs` |
| Removing `Option` wrappers breaks JSON deserialization of saved trials | Medium | Hyperopt regression | Add `#[serde(default)]` on all unwrapped fields; migration test |
| `MaybeNoisyLinear` removal misses a usage site | Low | Build break | Grep for `MaybeNoisyLinear` across entire workspace before deletion |
| NVRTC `--define-macro` changes break kernel compilation on older CUDA | Low | Runtime crash | Gate behind feature flag; keep `#ifndef` guards as fallback |
---
## 10. Execution Checklist
This is the recommended execution order within Spec C. Each step is independently
testable.
- [ ] **C.1** Add `cuda_grid_1d()` and `CUDA_BLOCK_DIM` to `cuda_pipeline/mod.rs`
- [ ] **C.2** Migrate all 50+ grid computations to use the helper (13 files)
- [ ] **C.3** Delete BF16 per-struct types from `gpu_weights.rs` (~340 LOC)
- [ ] **C.4** Delete `ensure_bf16_mirrors()` and `sync_online_bf16()` from `gpu_dqn_trainer.rs`
- [ ] **C.5** Replace `MaybeNoisyLinear` with `NoisyLinear` in `branching.rs`
- [ ] **C.6** Delete `if !true` dead path from `config.rs`
- [ ] **C.7** Delete `reset_episode()` from `agent.rs`
- [ ] **C.8** Extract monitoring limits to `MonitoringConfig`
- [ ] **C.9** Add `adx_period` field to `ExpertDemoGenerator`
- [ ] **C.10** Pass CUDA kernel defines via NVRTC compile-time options
- [ ] **C.11** Make `entropy_coefficient` non-optional with default `0.001`
- [ ] **C.12** Make `gradient_clip_norm` non-optional with default `10.0`
- [ ] **C.13** Make `noisy_epsilon_floor` non-optional with default `0.10`
- [ ] **C.14** Make `count_bonus_coefficient` non-optional with default `0.05`
- [ ] **C.15** Clean all WAVE/Phase/BUG labels from comments
- [ ] **C.16** Make `TrainingMonitor` accept dynamic branch sizes
- [ ] **C.17** Run full test suite and verify zero warnings