DQN Hyperparameters Analysis
Complete Field Documentation and Categorization
File: /home/jgrusewski/Work/foxhunt/ml/src/trainers/dqn/config.rs
Struct: DQNHyperparameters (lines 264-535)
Total Fields: 83
FIXED FLAGS (Architectural Decisions - Should NOT Be Tuned)
These are Rainbow DQN architectural features that should remain constant for model architecture consistency:
Rainbow DQN Core Components (ALWAYS ON)
| Field |
Default |
Rationale |
use_double_dqn |
true |
FIXED ON: Double DQN eliminates overestimation bias. Core Rainbow component. |
use_dueling |
true |
FIXED ON: Dueling architecture separates value/advantage streams. Core Rainbow component. |
use_per |
true |
FIXED ON: Prioritized Experience Replay improves sample efficiency. Core Rainbow component. |
use_noisy_nets |
true |
FIXED ON: Noisy Networks replace epsilon-greedy for state-dependent exploration. Core Rainbow component. |
n_steps |
3 |
FIXED: 3-step returns balance bias/variance. Rainbow DQN standard. |
Rainbow DQN Disabled Components
| Field |
Default |
Rationale |
use_distributional |
true |
SHOULD BE FALSE: C51 distributional RL has Candle bugs. Comment says "OFF (Candle bug)" but default is true - CONFLICT! |
Fixed Architectural Parameters
| Field |
Default |
Rationale |
dueling_hidden_dim |
128 |
FIXED: Network architecture parameter. Changing breaks checkpoint compatibility. |
num_atoms |
51 |
FIXED: Rainbow DQN standard for distributional RL (if enabled). |
per_alpha |
0.6 |
FIXED: PER prioritization exponent. Rainbow standard = 0.6. |
per_beta_start |
0.4 |
FIXED: PER importance sampling start. Rainbow standard = 0.4 → 1.0. |
Target Update Strategy (Architectural Choice)
| Field |
Default |
Rationale |
target_update_mode |
Soft |
FIXED: Soft (Polyak) updates are Rainbow DQN standard for stability. |
tau |
0.001 |
FIXED: Polyak coefficient. Rainbow standard = 0.001 (693-step half-life). |
Data Pipeline Configuration
| Field |
Default |
Rationale |
enable_preprocessing |
true |
FIXED ON: Preprocessing (log returns + normalization) is critical for stability. |
preprocessing_window |
50 |
FIXED: 50-bar rolling window for feature normalization. |
preprocessing_clip_sigma |
5.0 |
FIXED: Clip outliers at ±5σ. |
Loss Function Configuration
| Field |
Default |
Rationale |
use_huber_loss |
true |
FIXED ON: Huber loss is more robust to outliers than MSE. |
huber_delta |
100.0 |
FIXED: Scaled 100x for gradient explosion prevention (BUG #12 fix). |
Feature Engineering Pipeline
| Field |
Default |
Rationale |
feature_stats_collection_ratio |
0.3 |
FIXED: 30% of epochs for feature stats collection (WAVE 23). |
max_feature_stats_epochs |
Some(10) |
FIXED: Cap stats collection at 10 epochs. |
Default-Disabled Experimental Features
These should remain false/0.0 unless explicitly researching:
| Field |
Default |
Rationale |
enable_triple_barrier |
false |
FIXED OFF: Multi-step reward labeling - experimental feature. |
use_ensemble_uncertainty |
false |
FIXED OFF: Conflicts with use_noisy_nets. Mutually exclusive. |
enable_dropout_scheduler |
false |
FIXED OFF: Adaptive dropout - experimental feature. |
enable_gae |
false |
FIXED OFF: GAE is for actor-critic (PPO/A2C), not DQN. |
enable_noisy_sigma_scheduler |
false |
FIXED OFF: Sigma annealing - experimental feature. |
sharpe_weight |
0.0 |
FIXED OFF: Sharpe ratio reward component - experimental. |
curiosity_weight |
0.0 |
FIXED OFF: Curiosity-driven exploration - experimental. |
her_ratio |
0.0 |
FIXED OFF: Hindsight Experience Replay - experimental. |
TUNABLE HYPERPARAMETERS (Should Be Optimized)
These are true hyperparameters that should be tuned via hyperopt:
Learning Rate & Optimization
| Field |
Default |
Valid Range |
Rationale |
learning_rate |
0.0001 |
[1e-5, 1e-3] |
TUNE: Core hyperparameter. Trial 19 used 1e-4, typical range 1e-5 to 1e-3. |
batch_size |
128 |
[64, 512] |
TUNE: Memory-constrained by GPU. RTX 3050 Ti limit ≤230. Trial 19 used 256. |
gradient_clip_norm |
Some(10.0) |
[1.0, 100.0] or None |
TUNE: Prevents gradient explosion. Production uses 10.0, Trial 19 used 100.0. |
Learning Rate Scheduling (WAVE 26)
| Field |
Default |
Valid Range |
Rationale |
lr_decay_type |
Constant |
{Constant, Linear, Exponential, Cosine, Step} |
TUNE: LR schedule type. |
lr_decay_steps |
1000 |
[500, 10000] |
TUNE: Steps between LR decay updates. |
lr_decay_rate |
0.99 |
[0.9, 0.999] |
TUNE: Exponential decay rate (1% per step). |
min_learning_rate |
1e-6 |
[1e-7, 1e-5] |
TUNE: Minimum LR floor for Cosine decay. |
lr_min |
1e-6 |
[1e-7, 1e-5] |
TUNE: Absolute minimum LR. |
Discount Factor
| Field |
Default |
Valid Range |
Rationale |
gamma |
0.99 |
[0.95, 0.999] |
TUNE: Discount factor. Higher γ = longer planning horizon. |
Exploration Parameters
| Field |
Default |
Valid Range |
Rationale |
epsilon_start |
1.0 |
[0.5, 1.0] |
TUNE: Initial ε for ε-greedy. With noisy nets, ε can start lower. |
epsilon_end |
0.01 |
[0.001, 0.05] |
TUNE: Final ε. Rainbow uses 0.01. |
epsilon_decay |
0.995 |
[0.99, 0.9999] |
TUNE: Decay rate. Reaches epsilon_end after ~1000 steps (0.995) or ~100K steps (0.9999). |
noisy_sigma_init |
0.5 |
[0.3, 0.7] |
TUNE: Initial noise std for noisy networks. Rainbow standard = 0.5. |
Noisy Network Sigma Scheduling (Experimental)
| Field |
Default |
Valid Range |
Rationale |
noisy_sigma_initial |
0.6 |
[0.4, 0.8] |
TUNE: Initial sigma when scheduler enabled. |
noisy_sigma_final |
0.4 |
[0.2, 0.6] |
TUNE: Final sigma when scheduler enabled. |
noisy_sigma_anneal_steps |
10000 |
[5000, 50000] |
TUNE: Steps for sigma annealing. |
Replay Buffer Configuration
| Field |
Default |
Valid Range |
Rationale |
buffer_size |
500000 |
[100000, 1000000] |
TUNE: WAVE 24 increased to 500K for diversity. Larger = better sample diversity. |
min_replay_size |
1000 |
[100, 10000] |
TUNE: Min samples before training starts. |
Target Network Updates
| Field |
Default |
Valid Range |
Rationale |
target_update_frequency |
500 |
[100, 10000] |
TUNE: Hard update frequency (if mode=Hard). BUG #9 fix: 500 steps optimal. |
warmup_steps |
0 |
[0, 80000] |
TUNE: Random exploration warmup. Adaptive in CLI: 0 (<200K steps), 5% (200K-500K), 8% (500K-1M), 80K (>1M). |
Training Schedule
| Field |
Default |
Valid Range |
Rationale |
epochs |
100 |
[50, 500] |
TUNE: Total training epochs. Production uses 100-500. |
checkpoint_frequency |
10 |
[5, 50] |
TUNE: Checkpoint save interval. |
Early Stopping Configuration
| Field |
Default |
Valid Range |
Rationale |
early_stopping_enabled |
true |
{true, false} |
TUNE: Enable early stopping. |
q_value_floor |
-5.0 |
[-10.0, -2.0] |
TUNE: Min Q-value before stopping. Wave 3 fix: allow normal negative Q-values. |
min_loss_improvement_pct |
2.0 |
[1.0, 5.0] |
TUNE: Min loss improvement % over plateau window. |
plateau_window |
30 |
[10, 50] |
TUNE: Window size for plateau detection. |
min_epochs_before_stopping |
50 |
[20, 100] |
TUNE: Min epochs before early stop can trigger. |
Gradient Collapse Detection (WAVE 23)
| Field |
Default |
Valid Range |
Rationale |
gradient_collapse_multiplier |
100.0 |
[10.0, 1000.0] |
TUNE: Adaptive threshold = LR × multiplier. WAVE 23 replaces hardcoded 0.1 threshold. |
gradient_collapse_patience |
5 |
[3, 10] |
TUNE: Consecutive epochs before early stop. |
Penalty & Reward Shaping
| Field |
Default |
Valid Range |
Rationale |
hold_penalty |
-0.001 |
[-0.01, 0.0] |
TUNE: Small negative penalty for HOLD action (Bug #3 fix). |
hold_penalty_weight |
0.01 |
[0.0, 0.1] |
TUNE: HOLD penalty during large price movements. |
movement_threshold |
0.02 |
[0.01, 0.05] |
TUNE: Price movement % threshold (2% default). |
transaction_cost_multiplier |
1.0 |
[0.5, 2.0] |
TUNE: Multiplier for transaction costs in reward. |
Portfolio & Risk Management
| Field |
Default |
Valid Range |
Rationale |
initial_capital |
100000.0 |
[1000.0, 1000000.0] |
TUNE: Initial trading capital. Min $1K validated at CLI. |
cash_reserve_percent |
0.0 |
[0.0, 50.0] |
TUNE: Cash reserve % (0 = no reserve, backward compatible). |
max_position_absolute |
2.0 |
[1.0, 10.0] |
TUNE: Max absolute position size for action masking. |
Kelly Criterion Position Sizing
| Field |
Default |
Valid Range |
Rationale |
enable_kelly_sizing |
true |
{true, false} |
TUNE: Enable Kelly criterion position sizing. |
kelly_fractional |
0.5 |
[0.25, 1.0] |
TUNE: Kelly multiplier (0.5 = half-Kelly, conservative). |
kelly_max_fraction |
0.25 |
[0.1, 0.5] |
TUNE: Max Kelly fraction (25% = max 25% of portfolio). |
kelly_min_trades |
20 |
[10, 50] |
TUNE: Min trades for Kelly calculation. |
Volatility & Risk Metrics
| Field |
Default |
Valid Range |
Rationale |
volatility_window |
20 |
[10, 50] |
TUNE: Rolling window for volatility calculation. |
enable_volatility_epsilon |
true |
{true, false} |
TUNE: Volatility-adjusted exploration. |
enable_risk_adjusted_rewards |
true |
{true, false} |
TUNE: Sharpe-based rewards. |
Advanced Risk Features (WAVE 16S)
| Field |
Default |
Valid Range |
Rationale |
enable_drawdown_monitoring |
true |
{true, false} |
TUNE: 15% max drawdown early stop. |
enable_position_limits |
true |
{true, false} |
TUNE: 3-tier position limits (absolute, notional, concentration). |
enable_circuit_breaker |
true |
{true, false} |
TUNE: 5-failure trip mechanism. |
enable_action_masking |
true |
{true, false} |
TUNE: Filter invalid actions based on position limits. |
enable_entropy_regularization |
true |
{true, false} |
TUNE: Prevent policy collapse. |
enable_stress_testing |
true |
{true, false} |
TUNE: Robustness validation. |
Advanced Features (WAVE 35)
| Field |
Default |
Valid Range |
Rationale |
enable_regime_qnetwork |
true |
{true, false} |
TUNE: Regime-conditional Q-networks (3 heads: Trending, Ranging, Volatile). |
enable_compliance |
true |
{true, false} |
TUNE: Real-time regulatory validation. |
Entropy Regularization (WAVE 17)
| Field |
Default |
Valid Range |
Rationale |
entropy_coefficient |
None |
[0.0, 0.1] |
TUNE: Entropy bonus for policy diversity. None = disabled. |
Triple Barrier (Experimental - Default OFF)
| Field |
Default |
Valid Range |
Rationale |
triple_barrier_profit_target_bps |
100 |
[50, 500] |
TUNE: Profit target in bps (if enabled). |
triple_barrier_stop_loss_bps |
50 |
[25, 200] |
TUNE: Stop loss in bps (if enabled). |
triple_barrier_max_holding_seconds |
3600 |
[300, 7200] |
TUNE: Max holding period (if enabled). |
Distributional RL (C51) - SHOULD BE DISABLED
| Field |
Default |
Valid Range |
Rationale |
v_min |
-2.0 |
[-10.0, -1.0] |
TUNE (if C51 enabled): Min value for distribution. BUG #5 fix: align with reward range ±2. |
v_max |
2.0 |
[1.0, 10.0] |
TUNE (if C51 enabled): Max value for distribution. BUG #5 fix: align with reward range ±2. |
WAVE 26 Advanced Features (Experimental - Default OFF)
Ensemble Uncertainty (Conflicts with Noisy Nets)
| Field |
Default |
Valid Range |
Rationale |
ensemble_size |
5 |
[3, 10] |
TUNE (if enabled): Number of Q-network heads. |
beta_variance |
0.5 |
[0.1, 1.0] |
TUNE (if enabled): Variance penalty weight. |
beta_disagreement |
0.5 |
[0.1, 1.0] |
TUNE (if enabled): Disagreement penalty weight. |
beta_entropy |
0.1 |
[0.05, 0.5] |
TUNE (if enabled): Entropy bonus weight. |
variance_cap |
1.0 |
[0.1, 2.0] |
TUNE (if enabled): Variance cap to prevent over-penalization. |
Gradient Accumulation
| Field |
Default |
Valid Range |
Rationale |
gradient_accumulation_steps |
1 |
[1, 8] |
TUNE: Mini-batches per optimizer step. Effective batch = batch_size × steps. |
Sharpe Ratio Reward (Experimental)
| Field |
Default |
Valid Range |
Rationale |
sharpe_window |
20 |
[10, 50] |
TUNE (if sharpe_weight > 0): Rolling window for Sharpe calculation. |
Adaptive Dropout (Experimental)
| Field |
Default |
Valid Range |
Rationale |
dropout_initial |
0.5 |
[0.3, 0.7] |
TUNE (if scheduler enabled): Initial dropout rate. |
dropout_final |
0.1 |
[0.0, 0.3] |
TUNE (if scheduler enabled): Final dropout rate. |
dropout_anneal_steps |
10000 |
[5000, 50000] |
TUNE (if scheduler enabled): Annealing steps. |
Hindsight Experience Replay (Experimental)
| Field |
Default |
Valid Range |
Rationale |
her_strategy |
"future" |
{"final", "future"} |
TUNE (if her_ratio > 0): HER strategy. "future" > "final". |
GAE (Not Recommended for DQN)
| Field |
Default |
Valid Range |
Rationale |
gae_lambda |
0.95 |
[0.9, 0.99] |
TUNE (if GAE enabled): GAE λ parameter. Note: GAE is for actor-critic, not DQN. |
CRITICAL ISSUES & CONFLICTS
🚨 BUG: Distributional RL Default Mismatch
Field: use_distributional
Default: true
Documentation Says: "C51/Distributional: OFF (Candle bug)"
Problem: Default contradicts stated requirement
Required Fix: Change default to false in line 638:
⚠️ WARNING: Mutually Exclusive Features
Conflict 1: use_noisy_nets vs use_ensemble_uncertainty
- Both provide exploration bonuses
- Should not be enabled simultaneously
- Default:
use_noisy_nets=true, use_ensemble_uncertainty=false ✅ CORRECT
Conflict 2: epsilon_decay vs use_noisy_nets
- Noisy networks replace ε-greedy exploration
- With
use_noisy_nets=true, ε should be set to 0
- Current: ε-greedy still active alongside noisy nets
- Recommendation: When
use_noisy_nets=true, set epsilon_start=0.0
RAINBOW DQN CONFIGURATION SUMMARY
✅ CORRECT: Enabled by Default (Rainbow Core)
- Double DQN (
use_double_dqn=true)
- Dueling Networks (
use_dueling=true)
- Prioritized Experience Replay (
use_per=true)
- Noisy Networks (
use_noisy_nets=true)
- 3-Step Returns (
n_steps=3)
- Soft Target Updates (
tau=0.001, target_update_mode=Soft)
❌ INCORRECT: Should Be Disabled (Candle Bug)
- Distributional RL (
use_distributional=true → should be false)
Rainbow DQN Compliance Score
6/7 components correct (85.7%)
Action Required: Disable use_distributional to reach 100% Rainbow compliance.
HYPEROPT SEARCH SPACE RECOMMENDATIONS
Priority Tier 1 (Highest Impact)
learning_rate - Range: [1e-5, 1e-3], log-scale
batch_size - Range: [64, 256], categorical {64, 128, 256}
gamma - Range: [0.95, 0.999]
epsilon_decay - Range: [0.99, 0.9999], log-scale
gradient_clip_norm - Range: [10.0, 100.0] or None
Priority Tier 2 (Moderate Impact)
buffer_size - Range: [100000, 1000000], categorical {100K, 250K, 500K, 1M}
warmup_steps - Range: [0, 80000], adaptive based on total steps
kelly_fractional - Range: [0.25, 1.0]
max_position_absolute - Range: [1.0, 10.0]
noisy_sigma_init - Range: [0.3, 0.7]
Priority Tier 3 (Fine-Tuning)
- Early stopping parameters (
q_value_floor, plateau_window, min_epochs_before_stopping)
- Penalty weights (
hold_penalty, hold_penalty_weight, transaction_cost_multiplier)
- Learning rate schedule parameters (if
lr_decay_type != Constant)
- Risk management toggles (Kelly, volatility epsilon, risk-adjusted rewards)
VALIDATION CHECKLIST
Before Hyperopt
After Hyperopt
TOTAL FIELD COUNT
- Total Fields: 83
- Fixed Flags: 28 (33.7%)
- Tunable Hyperparameters: 55 (66.3%)
Field Breakdown:
- Rainbow Core (Fixed): 7
- Architectural (Fixed): 14
- Experimental Disabled (Fixed): 7
- Learning/Optimization (Tunable): 12
- Exploration (Tunable): 9
- Replay Buffer (Tunable): 2
- Early Stopping (Tunable): 7
- Portfolio/Risk (Tunable): 25
- Advanced Features (Experimental, Tunable if enabled): 21
REFERENCES
- Rainbow DQN Paper: Hessel et al. (2018)
- Trial 19 Hyperopt Results:
DQN_HYPEROPT_RESULTS_SUMMARY.md
- BUG #5 Fix: v_min/v_max alignment with reward range
- BUG #9 Fix: target_update_frequency = 500 steps optimal
- BUG #12 Fix: huber_delta scaled 100x
- WAVE 16 (Agent 36): Soft target updates
- WAVE 23: Feature caching + early stopping
- WAVE 24: Buffer size increase to 500K
- WAVE 26: Advanced features (LR scheduling, ensemble, etc.)
- WAVE 35: Regime-conditional Q-networks + compliance