Add live training metrics monitor CLI command (streaming & one-shot) using the monitoring gRPC service. Update DQN tests to match post-fix defaults: IQN disabled, CQL alpha=0.1, v_min/v_max widened, 26D search space. - train.rs: `fxt train monitor [--once] [--model X] [--interval N]` - Rewrite gradient collapse test for BF16 mixed precision awareness - Update inference test config to match trainer defaults (IQN off, CQL on) - Update production smoke test for 26D parameter space - Add dqn_action_collapse_fix_test.rs verifying all 6 root cause fixes - Add planning docs for monitoring service and epoch financial metrics Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
13 KiB
DQN Action Collapse Fix — Implementation Plan
For Claude: REQUIRED SUB-SKILL: Use superpowers:executing-plans to implement this plan task-by-task.
Goal: Fix DQN action collapse (2/45 actions, Trades=0, Q-values stuck at [-0.48, -0.47]) by correcting broken defaults, reward bias, GPU metrics gap, and exploration deficiencies.
Architecture: 4 files modified. Config defaults fixed (CQL alpha 1.0→0.1 + configurable, IQN disabled, v_min/v_max widened). Reward hold bias eliminated. GPU pnl_history populated. Count bonus wired into batch selection. cql_alpha added to 26D hyperopt search space.
Tech Stack: Rust, Candle ML framework, CUDA experience collector
Task 1: Fix DQNConfig Defaults
Files:
- Modify:
crates/ml/src/dqn/dqn.rs:238-259
Step 1: Edit the defaults
Change lines 238-259 in DQNConfig::default():
// FROM:
v_min: -2.0, // Bug #5: Corrected from -10.0
v_max: 2.0, // Bug #5: Corrected from 10.0
// ...
entropy_coefficient: 0.01,
// ...
use_cql: true,
cql_alpha: 1.0,
// ...
use_iqn: true,
// TO:
v_min: -10.0, // Widened: [-2,2] compressed Q-values, [-10,10] gives C51 room to separate 45 actions
v_max: 10.0,
// ...
entropy_coefficient: 0.05, // 5x stronger anti-collapse for 45-action space
// ...
use_cql: true,
cql_alpha: 0.1, // Reduced: 1.0 added ~ln(45)=3.8 loss penalty, crushing Q-value differentiation
// ...
use_iqn: false, // Disabled: IQN trains base q_network but inference uses dist_dueling (zero gradients)
Step 2: Run tests to verify compilation
Run: SQLX_OFFLINE=true cargo check -p ml --lib 2>&1 | head -20
Expected: Compiles (may show warnings, no errors)
Step 3: Commit
git add crates/ml/src/dqn/dqn.rs
git commit -m "fix(ml): correct DQNConfig defaults — widen v_min/v_max, reduce cql_alpha, disable IQN"
Task 2: Add cql_alpha to DQNHyperparameters and Wire Through Trainer
Files:
- Modify:
crates/ml/src/trainers/dqn/config.rs:413+(DQNHyperparameters struct) - Modify:
crates/ml/src/trainers/dqn/trainer.rs:443-444(hardcoded CQL config)
Step 1: Add fields to DQNHyperparameters
In crates/ml/src/trainers/dqn/config.rs, add two new fields to DQNHyperparameters after the existing fields (near the risk management section):
/// Enable Conservative Q-Learning (default: true, alpha controls strength)
pub use_cql: bool,
/// CQL regularization strength (default: 0.1, range 0.0-1.0)
/// 0.0 = disabled, 0.1 = mild conservatism, 1.0 = full offline-RL strength
pub cql_alpha: f64,
Find every place DQNHyperparameters is constructed with struct literals and add the new fields with defaults use_cql: true, cql_alpha: 0.1. Key locations:
DQNHyperparameters::default()impl- Any test fixtures creating DQNHyperparameters
Step 2: Wire through trainer
In crates/ml/src/trainers/dqn/trainer.rs:443-444, change:
// FROM:
use_cql: true,
cql_alpha: 1.0,
// TO:
use_cql: hyperparams.use_cql,
cql_alpha: hyperparams.cql_alpha as f32,
Step 3: Run tests
Run: SQLX_OFFLINE=true cargo check -p ml --lib 2>&1 | head -20
Expected: Compiles. Fix any struct literal missing-field errors by adding use_cql: true, cql_alpha: 0.1.
Step 4: Commit
git add crates/ml/src/trainers/dqn/config.rs crates/ml/src/trainers/dqn/trainer.rs
git commit -m "feat(ml): make CQL configurable via DQNHyperparameters (default alpha=0.1)"
Task 3: Fix Hold Reward Bias
Files:
- Modify:
crates/ml/src/dqn/reward.rs:929-935(hold_penalty /1000 divisor) - Modify:
crates/ml/src/trainers/dqn/trainer.rs:492(hold_reward default)
Step 1: Remove /1000 divisor from hold_penalty
In crates/ml/src/dqn/reward.rs:929-935, change:
// FROM:
// BUG FIX: Scale hold_penalty_weight to percentage units
// Config value: 0.5-2.0 (raw scalar from hyperopt)
// Scaled value: 0.0005-0.002 (50-200 basis points = 0.05-0.2%)
// This makes hold penalty comparable to transaction costs (0.05-0.15%), not 30x weaker
let hold_penalty_scale = Decimal::try_from(1000.0)
.unwrap_or(Decimal::ONE);
let hold_penalty_pct = self.config.hold_penalty_weight / hold_penalty_scale;
// TO:
// Hold penalty weight used directly (0.01-2.0 range from hyperopt).
// Previously divided by 1000, making it 0.00001 — negligible vs tx costs (0.05-0.15%).
let hold_penalty_pct = self.config.hold_penalty_weight;
Step 2: Zero hold_reward in trainer RewardConfig
In crates/ml/src/trainers/dqn/trainer.rs:492, change:
// FROM:
hold_reward: Decimal::try_from(0.001).unwrap_or(Decimal::ZERO),
// TO:
hold_reward: Decimal::ZERO, // Flat position = no edge = zero reward (was +0.001, 20x trade PnL)
Step 3: Run tests
Run: SQLX_OFFLINE=true cargo test -p ml --lib -- reward 2>&1 | tail -20
Expected: PASS (reward tests should still pass since they test relative behavior, not absolute values)
Step 4: Commit
git add crates/ml/src/dqn/reward.rs crates/ml/src/trainers/dqn/trainer.rs
git commit -m "fix(ml): eliminate hold reward bias — zero hold_reward, remove /1000 penalty divisor"
Task 4: Populate pnl_history from GPU Experience Collection
Files:
- Modify:
crates/ml/src/trainers/dqn/trainer.rs:1762-1777(after GPU batch Ok)
Step 1: Add pnl_history population
In crates/ml/src/trainers/dqn/trainer.rs, inside the Ok(batch) => arm (after line 1770, before let experiences = gpu_batch_to_experiences), add:
// Populate pnl_history from GPU-collected rewards for financial metrics.
// Without this, compute_epoch_financials() sees an empty deque and
// reports Trades=0, Sharpe=0.00 even when the agent is trading.
for &reward in &batch.rewards {
self.pnl_history.push_back(reward as f64);
if self.pnl_history.len() > 1000 {
self.pnl_history.pop_front();
}
}
Step 2: Verify compilation
Run: SQLX_OFFLINE=true cargo check -p ml --lib 2>&1 | head -10
Expected: Compiles
Step 3: Commit
git add crates/ml/src/trainers/dqn/trainer.rs
git commit -m "fix(ml): populate pnl_history from GPU experience collector for financial metrics"
Task 5: Wire Count Bonus into select_actions_batch
Files:
- Modify:
crates/ml/src/trainers/dqn/trainer.rs:3238-3254(select_actions_batch)
Step 1: Add count bonus to batch Q-values
In select_actions_batch(), after the forward pass (line 3241) and before drop(agent) (line 3243), add count bonus application:
// Apply count bonus (UCB exploration) to Q-values before argmax.
// This was previously only applied in DQN::select_action() (single-sample path),
// meaning the batch training path had no UCB exploration — the count bonus was dead code.
let batch_q_values = if agent.config.use_count_bonus {
let bonuses = agent.get_count_bonuses(); // Vec<f32> of length num_actions
let bonus_tensor = Tensor::from_vec(bonuses, (1, agent.config.num_actions), batch_q_values.device())
.map_err(|e| anyhow::anyhow!("Failed to create bonus tensor: {}", e))?
.to_dtype(batch_q_values.dtype())
.map_err(|e| anyhow::anyhow!("Failed to cast bonus tensor: {}", e))?;
batch_q_values.broadcast_add(&bonus_tensor)
.map_err(|e| anyhow::anyhow!("Failed to add count bonus to Q-values: {}", e))?
} else {
batch_q_values
};
Note: You'll need to check if get_count_bonuses() exists on the DQN struct. If it only has count_bonus.compute_bonuses(num_actions), expose a method:
// In crates/ml/src/dqn/dqn.rs, add to DQN impl:
pub fn get_count_bonuses(&self) -> Vec<f32> {
self.count_bonus.compute_bonuses(self.config.num_actions)
}
Step 2: Verify compilation
Run: SQLX_OFFLINE=true cargo check -p ml --lib 2>&1 | head -10
Expected: Compiles
Step 3: Commit
git add crates/ml/src/dqn/dqn.rs crates/ml/src/trainers/dqn/trainer.rs
git commit -m "feat(ml): wire count bonus (UCB) into batch action selection for exploration"
Task 6: Add cql_alpha to Hyperopt Search Space (25D → 26D)
Files:
- Modify:
crates/ml/src/hyperopt/adapters/dqn.rs(multiple functions)
Step 1: Expand continuous_bounds_for (25→26 params)
In continuous_bounds_for() (~line 504), add after the last bound:
// CQL conservatism (1D)
(0.0, 0.5), // 25: cql_alpha (linear, 0.0=disabled, 0.1=mild, 0.5=moderate)
Step 2: Update from_continuous (25→26 params)
Change the length check from 25 to 26:
if x.len() != 26 {
return Err(MLError::ConfigError {
reason: format!("Expected 26 continuous parameters, got {}", x.len()),
});
}
Add parsing:
let cql_alpha = x[25].clamp(0.0, 0.5);
Add to the returned struct:
cql_alpha,
Step 3: Update to_continuous (add emission)
Add at end of the vec:
self.cql_alpha, // 25
Step 4: Update param_names (add name)
Add at end of the vec:
"cql_alpha", // 25
Step 5: Update v_min/v_max search ranges
Change bounds:
// FROM:
(-3.0, -1.0), // 11: v_min (linear) - Bug #5 fix: center -2.0
(1.0, 3.0), // 12: v_max (linear) - Bug #5 fix: center +2.0
// TO:
(-15.0, -3.0), // 11: v_min (linear) - widened for C51 action separation
(3.0, 15.0), // 12: v_max (linear) - widened for C51 action separation
Step 6: Add cql_alpha field to DQNParams struct
pub cql_alpha: f64,
Step 7: Add cql_alpha to DQNParams::default()
cql_alpha: 0.1,
Step 8: Wire cql_alpha into DQNHyperparameters construction
In the function that builds DQNHyperparameters from DQNParams (~line 2246+):
use_cql: true,
cql_alpha: params.cql_alpha,
Step 9: Fix IQN default in hyperopt adapter
In the adapter defaults (~line 647):
// FROM:
use_qr_dqn: true,
// TO:
use_qr_dqn: false, // IQN disabled: conflicts with C51 (train-test mismatch)
Step 10: Run tests and fix assertions
Run: SQLX_OFFLINE=true cargo test -p ml --lib -- dqn 2>&1 | tail -40
Fix the tests that assert on the old values:
test_continuous_bounds: Updateassert_eq!(bounds.len(), 26),bounds[11]to(-15.0, -3.0),bounds[12]to(3.0, 15.0), addbounds[25]to(0.0, 0.5)test_param_names: Updateassert_eq!(names.len(), 26), addnames[25]="cql_alpha"test_per_params_always_enabledand similar: Add0.1(cql_alpha) to end of continuous vectors (now 26 elements)test_validate_rejects_v_min_gte_v_max: v_min/v_max test values may need adjustmenttest_qr_dqn_params_in_default: Now asserts!params.use_qr_dqn
Step 11: Commit
git add crates/ml/src/hyperopt/adapters/dqn.rs
git commit -m "feat(ml): add cql_alpha to 26D hyperopt search space, widen v_min/v_max ranges"
Task 7: Run Full Test Suite and Fix Remaining Failures
Files:
- All 4 modified files
Step 1: Run all ml tests
Run: SQLX_OFFLINE=true cargo test -p ml --lib 2>&1 | tail -30
Expected: PASS (all tests, fix any remaining failures from struct changes)
Step 2: Run clippy
Run: SQLX_OFFLINE=true cargo clippy -p ml --lib -- -D warnings 2>&1 | tail -20
Expected: 0 errors, 0 warnings
Step 3: Run workspace check
Run: SQLX_OFFLINE=true cargo check --workspace 2>&1 | tail -10
Expected: Compiles (check that trading_service, fxt, etc. still compile with new DQNHyperparameters fields)
Step 4: Commit any test fixes
git add -A
git commit -m "test(ml): update assertions for new DQN defaults and 26D search space"
Task 8: Final Validation Commit
Step 1: Verify git status is clean
Run: git status
Step 2: Squash or keep commits as-is
Keep individual commits for traceability. The final commit history should be:
fix(ml): correct DQNConfig defaultsfeat(ml): make CQL configurable via DQNHyperparametersfix(ml): eliminate hold reward biasfix(ml): populate pnl_history from GPU experience collectorfeat(ml): wire count bonus (UCB) into batch action selectionfeat(ml): add cql_alpha to 26D hyperopt search spacetest(ml): update assertions for new DQN defaults
Summary of Expected Behavioral Changes
| Metric | Before (job #8885) | After (expected) |
|---|---|---|
| Action diversity | 2/45 (4.4%) | >9/45 (>20%) |
| Entropy | 0.120 | >0.5 |
| Q-value range | [-0.48, -0.47] (0.01 spread) | >0.5 spread |
| Trades | 0 | >0 |
| Sharpe | 0.00 | Non-zero |
| Hold reward per bar | +0.001 | 0.0 |
| Hold penalty per bar | -0.00001 | -0.01 to -2.0 (direct from hyperopt) |
| CQL loss penalty | ~3.8 per step | ~0.38 per step |
| Count bonus in batch | Dead code | Active |
| GPU financial metrics | Dark (empty pnl_history) | Reporting |