Root causes: entropy regularizer dead code, noisy nets disable epsilon, diversity penalty calibrated for 3 actions not 45. Fixes: wire entropy into loss, fix normalization, epsilon floor for noisy nets, sigma scheduling, UCB count-based bonus. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
12 KiB
DQN Action Collapse Fix — Design & Implementation Plan
For Claude: REQUIRED SUB-SKILL: Use superpowers:executing-plans to implement this plan task-by-task.
Goal: Fix 45-action DQN collapse where the model converges to a single action (A38 = Long100/Market/Aggressive) across all hyperopt trials, despite existing anti-collapse infrastructure.
Root Causes:
- Entropy regularizer is dead code (initialized, never called)
- Noisy nets disable epsilon-greedy but provide insufficient exploration for 45 actions
- Diversity penalty entropy normalized for 3 actions, not 45
Approach: Fix all three existing mechanisms + add two new ones (epsilon floor for noisy nets, count-based exploration bonus).
Task 1: Wire entropy regularizer into TD loss
Files:
- Modify:
crates/ml/src/dqn/dqn.rs(add entropy bonus tocompute_lossor loss path) - Modify:
crates/ml/src/trainers/dqn/trainer.rs(pass entropy regularizer into training step)
Step 1: Add entropy bonus computation to the Q-learning loss
The entropy regularizer at crates/ml/src/dqn/entropy_regularization.rs is fully implemented but never called. Wire it into the training loop.
In the training step where loss is computed (trainer.rs, the train_step / gradient update section), after computing Q-values for the current batch:
// Compute softmax policy from Q-values: π(a|s) = softmax(Q(s,a) / τ)
// Then entropy: H(π) = -Σ π(a|s) * log(π(a|s))
// Add to loss: loss = td_loss - entropy_coefficient * mean_entropy
Specifically:
- After the forward pass that produces
q_valuestensor (shape [batch, 45]) - Compute
log_softmax(q_values, dim=1)andsoftmax(q_values, dim=1) - Entropy per sample:
H = -(softmax * log_softmax).sum(dim=1)→ shape [batch] - Mean entropy:
mean_H = H.mean() - Final loss:
loss = td_loss - entropy_coefficient * mean_H
The entropy_coefficient field already exists in DQNHyperparameters (default 0.01, hyperopt range [0.0, 0.1]).
Step 2: Remove unused entropy_regularizer field from trainer struct
The Arc<EntropyRegularizer> stored at trainer.rs:754 is no longer needed — entropy is computed directly on Q-value tensors. Remove the field and the initialization at lines 531-535.
Step 3: Add unit test
Test that with entropy_coefficient > 0, training on a batch where all Q-values are identical (uniform policy) produces lower loss than a batch with one dominant Q-value (collapsed policy).
Step 4: Build and test
SQLX_OFFLINE=true cargo test -p ml --lib -- trainers::dqn
Step 5: Commit
git add crates/ml/src/dqn/ crates/ml/src/trainers/dqn/
git commit -m "fix(dqn): wire entropy regularization into TD loss (was dead code)"
Task 2: Fix diversity penalty entropy normalization
Files:
- Modify:
crates/ml/src/dqn/reward.rs(fixcalculate_entropy()normalization)
Step 1: Fix entropy normalization from log(3) to log(45)
In reward.rs, find calculate_entropy() (around line 364-395). The function computes Shannon entropy of recent actions. Currently it normalizes by log(3) (3 legacy action categories). Change to log(45) to match the 45 factored actions.
Also change the entropy threshold from 0.5 to 0.3. With 45 actions, entropy of 0.3 (normalized) means ~4 distinct actions — a reasonable collapse threshold.
// Before:
let max_entropy = (3.0_f64).ln(); // log(3) = 1.099
let entropy_threshold = 0.5;
// After:
let max_entropy = (45.0_f64).ln(); // log(45) = 3.807
let entropy_threshold = 0.3;
Step 2: Update the action counting to use 45 action indices
If the entropy calculation currently maps actions to 3 categories (Buy/Sell/Hold), change it to use the full to_index() (0-44) for the frequency distribution.
Step 3: Add test
Test that:
- Uniform distribution over 45 actions → entropy ≈ 1.0 (normalized) → no penalty
- Single-action collapse (all A38) → entropy ≈ 0.0 → penalty fires (-0.1)
- 5 distinct actions → entropy ≈ 0.42 → penalty fires (< 0.3 threshold? depends on distribution)
Step 4: Build and test
SQLX_OFFLINE=true cargo test -p ml --lib -- dqn::reward
Step 5: Commit
git add crates/ml/src/dqn/reward.rs
git commit -m "fix(dqn): normalize diversity entropy for 45 actions (was calibrated for 3)"
Task 3: Add epsilon floor for noisy nets
Files:
- Modify:
crates/ml/src/trainers/dqn/config.rs(addnoisy_epsilon_floorfield) - Modify:
crates/ml/src/dqn/dqn.rs(use floor inselect_action()) - Modify:
crates/ml/src/hyperopt/adapters/dqn.rs(add to search space)
Step 1: Add config field
In config.rs, add to DQNHyperparameters:
/// Minimum epsilon when noisy nets are enabled (guarantees some random exploration)
pub noisy_epsilon_floor: f64, // default: 0.05
Add default value in the Default impl: noisy_epsilon_floor: 0.05.
Step 2: Use floor in select_action()
In dqn.rs, change the noisy nets epsilon logic:
// Before:
let effective_epsilon = if self.config.use_noisy_nets {
0.0 // Learned exploration via parameter noise
} else {
self.epsilon
};
// After:
let effective_epsilon = if self.config.use_noisy_nets {
self.config.noisy_epsilon_floor // Minimum random exploration even with noisy nets
} else {
self.epsilon
};
Step 3: Add to hyperopt search space
In dqn.rs hyperopt adapter, add noisy_epsilon_floor to:
DQNParamsstructcontinuous_bounds()— range [0.02, 0.10]from_continuous()/to_continuous()to_hyperparams()mapping
Step 4: Skip epsilon decay when noisy nets enabled
In trainer.rs, guard the agent.update_epsilon() call:
// Before:
agent.update_epsilon();
// After:
if !self.hyperparams.use_noisy_nets {
agent.update_epsilon();
}
Step 5: Build and test
SQLX_OFFLINE=true cargo check -p ml
SQLX_OFFLINE=true cargo test -p ml --lib -- dqn::dqn::tests
Step 6: Commit
git add crates/ml/src/trainers/dqn/config.rs crates/ml/src/dqn/dqn.rs crates/ml/src/hyperopt/adapters/dqn.rs crates/ml/src/trainers/dqn/trainer.rs
git commit -m "feat(dqn): add epsilon floor for noisy nets (0.05 default, hyperopt-tunable)"
Task 4: Enable noisy sigma scheduling with higher initial sigma
Files:
- Modify:
crates/ml/src/trainers/dqn/config.rs(change defaults) - Modify:
crates/ml/src/dqn/dqn.rs(add sigma floor)
Step 1: Change defaults for sigma scheduling
In config.rs, update the defaults:
// Before:
enable_noisy_sigma_scheduler: false,
noisy_sigma_initial: 0.6,
noisy_sigma_final: 0.4,
// After:
enable_noisy_sigma_scheduler: true, // Enable by default
noisy_sigma_initial: 0.8, // Wider initial exploration
noisy_sigma_final: 0.4,
Step 2: Add sigma floor
In the sigma scheduling logic (wherever sigma is computed/annealed), add a minimum:
let sigma = computed_sigma.max(0.3); // Never let sigma drop below 0.3
Step 3: Update hyperopt adapter defaults
Make sure the hyperopt adapter uses the new defaults when constructing DQNHyperparameters.
Step 4: Build and test
SQLX_OFFLINE=true cargo check -p ml
Step 5: Commit
git add crates/ml/src/trainers/dqn/config.rs crates/ml/src/dqn/dqn.rs
git commit -m "feat(dqn): enable noisy sigma scheduling (0.8→0.4, floor=0.3)"
Task 5: Add count-based exploration bonus (UCB)
Files:
- Create:
crates/ml/src/dqn/count_bonus.rs - Modify:
crates/ml/src/dqn/mod.rs(add module) - Modify:
crates/ml/src/trainers/dqn/config.rs(add config fields) - Modify:
crates/ml/src/dqn/dqn.rs(apply bonus in action selection) - Modify:
crates/ml/src/hyperopt/adapters/dqn.rs(add to search space)
Step 1: Create count_bonus.rs
/// UCB-style count-based exploration bonus for action selection.
///
/// Adds bonus = β * sqrt(log(N) / (1 + n_a)) to Q-values during action selection,
/// where N = total actions taken this epoch, n_a = count of action a.
/// This gives direct exploration incentive to under-visited actions.
pub struct CountBonus {
action_counts: [u64; 45],
total_count: u64,
coefficient: f64,
}
impl CountBonus {
pub fn new(coefficient: f64) -> Self { ... }
/// Record that action was taken
pub fn record_action(&mut self, action_idx: usize) { ... }
/// Compute bonus for each of 45 actions, returns [f64; 45]
pub fn bonuses(&self) -> [f64; 45] {
if self.total_count == 0 {
return [self.coefficient; 45]; // Max bonus when no data
}
let log_n = (self.total_count as f64).ln();
let mut result = [0.0; 45];
for a in 0..45 {
result[a] = self.coefficient * (log_n / (1.0 + self.action_counts[a] as f64)).sqrt();
}
result
}
/// Reset counts (call at epoch boundary)
pub fn reset(&mut self) { ... }
}
Step 2: Add config fields
In config.rs:
pub use_count_bonus: bool, // default: true
pub count_bonus_coefficient: f64, // default: 0.1
Step 3: Wire into action selection
In dqn.rs select_action(), when choosing the greedy action:
// After computing q_values from forward pass:
if self.config.use_count_bonus {
let bonuses = self.count_bonus.bonuses();
let bonus_tensor = Tensor::from_vec(bonuses.to_vec(), 45, device)?;
let q_with_bonus = q_values.broadcast_add(&bonus_tensor)?;
// Use q_with_bonus for argmax instead of q_values
}
Record action after selection:
if self.config.use_count_bonus {
self.count_bonus.record_action(action.to_index());
}
Step 4: Reset at epoch boundary
In trainer.rs, at the start of each epoch:
agent.reset_count_bonus(); // Resets per-epoch exploration pressure
Step 5: Add to hyperopt search space
count_bonus_coefficient— range [0.01, 0.5]use_count_bonus— always true (not tuned, it's a safety mechanism)
Step 6: Add tests
- Test that bonuses are high for unvisited actions and low for frequent actions
- Test that after recording 100 of action 38, action 38's bonus is near zero while others are high
- Test reset clears all counts
Step 7: Build and test
SQLX_OFFLINE=true cargo test -p ml --lib -- dqn::count_bonus
SQLX_OFFLINE=true cargo check -p ml
Step 8: Commit
git add crates/ml/src/dqn/count_bonus.rs crates/ml/src/dqn/mod.rs crates/ml/src/trainers/dqn/config.rs crates/ml/src/dqn/dqn.rs crates/ml/src/hyperopt/adapters/dqn.rs
git commit -m "feat(dqn): add UCB count-based exploration bonus for 45-action diversity"
Task 6: Update monitoring to report exploration diagnostics
Files:
- Modify:
crates/ml/src/trainers/dqn/trainer.rs(log exploration source stats)
Step 1: Track exploration source in action diversity log
In the per-epoch action diversity logging, add:
- Count of random (epsilon) vs greedy vs noisy-net actions
- Effective entropy over 45 actions (normalized by log(45))
- Count bonus range (min/max/mean)
Step 2: Add Prometheus metrics
In crates/common/src/metrics/training_metrics.rs, add:
foxhunt_training_action_entropy // Normalized entropy (0-1) of epoch action distribution
foxhunt_training_epsilon_actions // Count of epsilon-random actions this epoch
foxhunt_training_greedy_actions // Count of greedy actions this epoch
Step 3: Commit
git add crates/ml/src/trainers/dqn/trainer.rs crates/common/src/metrics/training_metrics.rs
git commit -m "feat(monitoring): add exploration diagnostics (entropy, epsilon/greedy split)"
Task 7: Full build verification
Step 1: SQLX_OFFLINE=true cargo check --workspace
Step 2: SQLX_OFFLINE=true cargo test -p ml --lib
Step 3: SQLX_OFFLINE=true cargo test -p common --lib
Step 4: SQLX_OFFLINE=true cargo clippy --workspace -- -D warnings
Expected: 0 errors, 0 warnings, all tests pass.