BREAKING CHANGES: - Removed orphaned dqn.rs monolithic trainer (4,975 lines) - Removed orphaned dqn_ensemble.rs module (816 lines) - Removed orphaned tft.rs and tft_complete_int8_integration_test.rs - TFT trainer split into modular directory structure DQN Module Refactoring: - Split trainers/dqn.rs into modular structure (config.rs, statistics.rs, trainer.rs) - Fixed hyperopt 39D search space (continuous params only) - Boolean flags (use_dueling, use_double_dqn, use_per, use_noisy_nets) are now FIXED architectural decisions - use_distributional defaults to false (Candle BUG #36 - scatter_add gradient issues) Clean Module Structure: - ml/src/trainers/dqn/ directory with proper mod.rs exports - ml/src/trainers/tft/ directory with config.rs, types.rs, model.rs, trainer.rs, tests.rs - All P0 features validated: TD-error clamping, batch diversity, LR scheduler, priority staleness Documentation: - Added comprehensive docs in docs/codebase-cleanup/ - ADR-001 for DQN refactoring decisions - Rainbow DQN component matrix and quick reference guides Build Status: Compiles with zero errors 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com>
411 lines
19 KiB
Markdown
411 lines
19 KiB
Markdown
# DQN Hyperparameters Analysis
|
||
## Complete Field Documentation and Categorization
|
||
|
||
**File**: `/home/jgrusewski/Work/foxhunt/ml/src/trainers/dqn/config.rs`
|
||
**Struct**: `DQNHyperparameters` (lines 264-535)
|
||
**Total Fields**: 83
|
||
|
||
---
|
||
|
||
## FIXED FLAGS (Architectural Decisions - Should NOT Be Tuned)
|
||
|
||
These are Rainbow DQN architectural features that should remain constant for model architecture consistency:
|
||
|
||
### Rainbow DQN Core Components (ALWAYS ON)
|
||
|
||
| Field | Default | Rationale |
|
||
|-------|---------|-----------|
|
||
| `use_double_dqn` | `true` | **FIXED ON**: Double DQN eliminates overestimation bias. Core Rainbow component. |
|
||
| `use_dueling` | `true` | **FIXED ON**: Dueling architecture separates value/advantage streams. Core Rainbow component. |
|
||
| `use_per` | `true` | **FIXED ON**: Prioritized Experience Replay improves sample efficiency. Core Rainbow component. |
|
||
| `use_noisy_nets` | `true` | **FIXED ON**: Noisy Networks replace epsilon-greedy for state-dependent exploration. Core Rainbow component. |
|
||
| `n_steps` | `3` | **FIXED**: 3-step returns balance bias/variance. Rainbow DQN standard. |
|
||
|
||
### Rainbow DQN Disabled Components
|
||
|
||
| Field | Default | Rationale |
|
||
|-------|---------|-----------|
|
||
| `use_distributional` | `true` | **SHOULD BE FALSE**: C51 distributional RL has Candle bugs. Comment says "OFF (Candle bug)" but default is `true` - **CONFLICT!** |
|
||
|
||
### Fixed Architectural Parameters
|
||
|
||
| Field | Default | Rationale |
|
||
|-------|---------|-----------|
|
||
| `dueling_hidden_dim` | `128` | **FIXED**: Network architecture parameter. Changing breaks checkpoint compatibility. |
|
||
| `num_atoms` | `51` | **FIXED**: Rainbow DQN standard for distributional RL (if enabled). |
|
||
| `per_alpha` | `0.6` | **FIXED**: PER prioritization exponent. Rainbow standard = 0.6. |
|
||
| `per_beta_start` | `0.4` | **FIXED**: PER importance sampling start. Rainbow standard = 0.4 → 1.0. |
|
||
|
||
### Target Update Strategy (Architectural Choice)
|
||
|
||
| Field | Default | Rationale |
|
||
|-------|---------|-----------|
|
||
| `target_update_mode` | `Soft` | **FIXED**: Soft (Polyak) updates are Rainbow DQN standard for stability. |
|
||
| `tau` | `0.001` | **FIXED**: Polyak coefficient. Rainbow standard = 0.001 (693-step half-life). |
|
||
|
||
### Data Pipeline Configuration
|
||
|
||
| Field | Default | Rationale |
|
||
|-------|---------|-----------|
|
||
| `enable_preprocessing` | `true` | **FIXED ON**: Preprocessing (log returns + normalization) is critical for stability. |
|
||
| `preprocessing_window` | `50` | **FIXED**: 50-bar rolling window for feature normalization. |
|
||
| `preprocessing_clip_sigma` | `5.0` | **FIXED**: Clip outliers at ±5σ. |
|
||
|
||
### Loss Function Configuration
|
||
|
||
| Field | Default | Rationale |
|
||
|-------|---------|-----------|
|
||
| `use_huber_loss` | `true` | **FIXED ON**: Huber loss is more robust to outliers than MSE. |
|
||
| `huber_delta` | `100.0` | **FIXED**: Scaled 100x for gradient explosion prevention (BUG #12 fix). |
|
||
|
||
### Feature Engineering Pipeline
|
||
|
||
| Field | Default | Rationale |
|
||
|-------|---------|-----------|
|
||
| `feature_stats_collection_ratio` | `0.3` | **FIXED**: 30% of epochs for feature stats collection (WAVE 23). |
|
||
| `max_feature_stats_epochs` | `Some(10)` | **FIXED**: Cap stats collection at 10 epochs. |
|
||
|
||
### Default-Disabled Experimental Features
|
||
|
||
These should remain `false`/`0.0` unless explicitly researching:
|
||
|
||
| Field | Default | Rationale |
|
||
|-------|---------|-----------|
|
||
| `enable_triple_barrier` | `false` | **FIXED OFF**: Multi-step reward labeling - experimental feature. |
|
||
| `use_ensemble_uncertainty` | `false` | **FIXED OFF**: Conflicts with `use_noisy_nets`. Mutually exclusive. |
|
||
| `enable_dropout_scheduler` | `false` | **FIXED OFF**: Adaptive dropout - experimental feature. |
|
||
| `enable_gae` | `false` | **FIXED OFF**: GAE is for actor-critic (PPO/A2C), not DQN. |
|
||
| `enable_noisy_sigma_scheduler` | `false` | **FIXED OFF**: Sigma annealing - experimental feature. |
|
||
| `sharpe_weight` | `0.0` | **FIXED OFF**: Sharpe ratio reward component - experimental. |
|
||
| `curiosity_weight` | `0.0` | **FIXED OFF**: Curiosity-driven exploration - experimental. |
|
||
| `her_ratio` | `0.0` | **FIXED OFF**: Hindsight Experience Replay - experimental. |
|
||
|
||
---
|
||
|
||
## TUNABLE HYPERPARAMETERS (Should Be Optimized)
|
||
|
||
These are true hyperparameters that should be tuned via hyperopt:
|
||
|
||
### Learning Rate & Optimization
|
||
|
||
| Field | Default | Valid Range | Rationale |
|
||
|-------|---------|-------------|-----------|
|
||
| `learning_rate` | `0.0001` | `[1e-5, 1e-3]` | **TUNE**: Core hyperparameter. Trial 19 used 1e-4, typical range 1e-5 to 1e-3. |
|
||
| `batch_size` | `128` | `[64, 512]` | **TUNE**: Memory-constrained by GPU. RTX 3050 Ti limit ≤230. Trial 19 used 256. |
|
||
| `gradient_clip_norm` | `Some(10.0)` | `[1.0, 100.0]` or `None` | **TUNE**: Prevents gradient explosion. Production uses 10.0, Trial 19 used 100.0. |
|
||
|
||
### Learning Rate Scheduling (WAVE 26)
|
||
|
||
| Field | Default | Valid Range | Rationale |
|
||
|-------|---------|-------------|-----------|
|
||
| `lr_decay_type` | `Constant` | `{Constant, Linear, Exponential, Cosine, Step}` | **TUNE**: LR schedule type. |
|
||
| `lr_decay_steps` | `1000` | `[500, 10000]` | **TUNE**: Steps between LR decay updates. |
|
||
| `lr_decay_rate` | `0.99` | `[0.9, 0.999]` | **TUNE**: Exponential decay rate (1% per step). |
|
||
| `min_learning_rate` | `1e-6` | `[1e-7, 1e-5]` | **TUNE**: Minimum LR floor for Cosine decay. |
|
||
| `lr_min` | `1e-6` | `[1e-7, 1e-5]` | **TUNE**: Absolute minimum LR. |
|
||
|
||
### Discount Factor
|
||
|
||
| Field | Default | Valid Range | Rationale |
|
||
|-------|---------|-------------|-----------|
|
||
| `gamma` | `0.99` | `[0.95, 0.999]` | **TUNE**: Discount factor. Higher γ = longer planning horizon. |
|
||
|
||
### Exploration Parameters
|
||
|
||
| Field | Default | Valid Range | Rationale |
|
||
|-------|---------|-------------|-----------|
|
||
| `epsilon_start` | `1.0` | `[0.5, 1.0]` | **TUNE**: Initial ε for ε-greedy. With noisy nets, ε can start lower. |
|
||
| `epsilon_end` | `0.01` | `[0.001, 0.05]` | **TUNE**: Final ε. Rainbow uses 0.01. |
|
||
| `epsilon_decay` | `0.995` | `[0.99, 0.9999]` | **TUNE**: Decay rate. Reaches epsilon_end after ~1000 steps (0.995) or ~100K steps (0.9999). |
|
||
| `noisy_sigma_init` | `0.5` | `[0.3, 0.7]` | **TUNE**: Initial noise std for noisy networks. Rainbow standard = 0.5. |
|
||
|
||
### Noisy Network Sigma Scheduling (Experimental)
|
||
|
||
| Field | Default | Valid Range | Rationale |
|
||
|-------|---------|-------------|-----------|
|
||
| `noisy_sigma_initial` | `0.6` | `[0.4, 0.8]` | **TUNE**: Initial sigma when scheduler enabled. |
|
||
| `noisy_sigma_final` | `0.4` | `[0.2, 0.6]` | **TUNE**: Final sigma when scheduler enabled. |
|
||
| `noisy_sigma_anneal_steps` | `10000` | `[5000, 50000]` | **TUNE**: Steps for sigma annealing. |
|
||
|
||
### Replay Buffer Configuration
|
||
|
||
| Field | Default | Valid Range | Rationale |
|
||
|-------|---------|-------------|-----------|
|
||
| `buffer_size` | `500000` | `[100000, 1000000]` | **TUNE**: WAVE 24 increased to 500K for diversity. Larger = better sample diversity. |
|
||
| `min_replay_size` | `1000` | `[100, 10000]` | **TUNE**: Min samples before training starts. |
|
||
|
||
### Target Network Updates
|
||
|
||
| Field | Default | Valid Range | Rationale |
|
||
|-------|---------|-------------|-----------|
|
||
| `target_update_frequency` | `500` | `[100, 10000]` | **TUNE**: Hard update frequency (if mode=Hard). BUG #9 fix: 500 steps optimal. |
|
||
| `warmup_steps` | `0` | `[0, 80000]` | **TUNE**: Random exploration warmup. Adaptive in CLI: 0 (<200K steps), 5% (200K-500K), 8% (500K-1M), 80K (>1M). |
|
||
|
||
### Training Schedule
|
||
|
||
| Field | Default | Valid Range | Rationale |
|
||
|-------|---------|-------------|-----------|
|
||
| `epochs` | `100` | `[50, 500]` | **TUNE**: Total training epochs. Production uses 100-500. |
|
||
| `checkpoint_frequency` | `10` | `[5, 50]` | **TUNE**: Checkpoint save interval. |
|
||
|
||
### Early Stopping Configuration
|
||
|
||
| Field | Default | Valid Range | Rationale |
|
||
|-------|---------|-------------|-----------|
|
||
| `early_stopping_enabled` | `true` | `{true, false}` | **TUNE**: Enable early stopping. |
|
||
| `q_value_floor` | `-5.0` | `[-10.0, -2.0]` | **TUNE**: Min Q-value before stopping. Wave 3 fix: allow normal negative Q-values. |
|
||
| `min_loss_improvement_pct` | `2.0` | `[1.0, 5.0]` | **TUNE**: Min loss improvement % over plateau window. |
|
||
| `plateau_window` | `30` | `[10, 50]` | **TUNE**: Window size for plateau detection. |
|
||
| `min_epochs_before_stopping` | `50` | `[20, 100]` | **TUNE**: Min epochs before early stop can trigger. |
|
||
|
||
### Gradient Collapse Detection (WAVE 23)
|
||
|
||
| Field | Default | Valid Range | Rationale |
|
||
|-------|---------|-------------|-----------|
|
||
| `gradient_collapse_multiplier` | `100.0` | `[10.0, 1000.0]` | **TUNE**: Adaptive threshold = LR × multiplier. WAVE 23 replaces hardcoded 0.1 threshold. |
|
||
| `gradient_collapse_patience` | `5` | `[3, 10]` | **TUNE**: Consecutive epochs before early stop. |
|
||
|
||
### Penalty & Reward Shaping
|
||
|
||
| Field | Default | Valid Range | Rationale |
|
||
|-------|---------|-------------|-----------|
|
||
| `hold_penalty` | `-0.001` | `[-0.01, 0.0]` | **TUNE**: Small negative penalty for HOLD action (Bug #3 fix). |
|
||
| `hold_penalty_weight` | `0.01` | `[0.0, 0.1]` | **TUNE**: HOLD penalty during large price movements. |
|
||
| `movement_threshold` | `0.02` | `[0.01, 0.05]` | **TUNE**: Price movement % threshold (2% default). |
|
||
| `transaction_cost_multiplier` | `1.0` | `[0.5, 2.0]` | **TUNE**: Multiplier for transaction costs in reward. |
|
||
|
||
### Portfolio & Risk Management
|
||
|
||
| Field | Default | Valid Range | Rationale |
|
||
|-------|---------|-------------|-----------|
|
||
| `initial_capital` | `100000.0` | `[1000.0, 1000000.0]` | **TUNE**: Initial trading capital. Min $1K validated at CLI. |
|
||
| `cash_reserve_percent` | `0.0` | `[0.0, 50.0]` | **TUNE**: Cash reserve % (0 = no reserve, backward compatible). |
|
||
| `max_position_absolute` | `2.0` | `[1.0, 10.0]` | **TUNE**: Max absolute position size for action masking. |
|
||
|
||
### Kelly Criterion Position Sizing
|
||
|
||
| Field | Default | Valid Range | Rationale |
|
||
|-------|---------|-------------|-----------|
|
||
| `enable_kelly_sizing` | `true` | `{true, false}` | **TUNE**: Enable Kelly criterion position sizing. |
|
||
| `kelly_fractional` | `0.5` | `[0.25, 1.0]` | **TUNE**: Kelly multiplier (0.5 = half-Kelly, conservative). |
|
||
| `kelly_max_fraction` | `0.25` | `[0.1, 0.5]` | **TUNE**: Max Kelly fraction (25% = max 25% of portfolio). |
|
||
| `kelly_min_trades` | `20` | `[10, 50]` | **TUNE**: Min trades for Kelly calculation. |
|
||
|
||
### Volatility & Risk Metrics
|
||
|
||
| Field | Default | Valid Range | Rationale |
|
||
|-------|---------|-------------|-----------|
|
||
| `volatility_window` | `20` | `[10, 50]` | **TUNE**: Rolling window for volatility calculation. |
|
||
| `enable_volatility_epsilon` | `true` | `{true, false}` | **TUNE**: Volatility-adjusted exploration. |
|
||
| `enable_risk_adjusted_rewards` | `true` | `{true, false}` | **TUNE**: Sharpe-based rewards. |
|
||
|
||
### Advanced Risk Features (WAVE 16S)
|
||
|
||
| Field | Default | Valid Range | Rationale |
|
||
|-------|---------|-------------|-----------|
|
||
| `enable_drawdown_monitoring` | `true` | `{true, false}` | **TUNE**: 15% max drawdown early stop. |
|
||
| `enable_position_limits` | `true` | `{true, false}` | **TUNE**: 3-tier position limits (absolute, notional, concentration). |
|
||
| `enable_circuit_breaker` | `true` | `{true, false}` | **TUNE**: 5-failure trip mechanism. |
|
||
| `enable_action_masking` | `true` | `{true, false}` | **TUNE**: Filter invalid actions based on position limits. |
|
||
| `enable_entropy_regularization` | `true` | `{true, false}` | **TUNE**: Prevent policy collapse. |
|
||
| `enable_stress_testing` | `true` | `{true, false}` | **TUNE**: Robustness validation. |
|
||
|
||
### Advanced Features (WAVE 35)
|
||
|
||
| Field | Default | Valid Range | Rationale |
|
||
|-------|---------|-------------|-----------|
|
||
| `enable_regime_qnetwork` | `true` | `{true, false}` | **TUNE**: Regime-conditional Q-networks (3 heads: Trending, Ranging, Volatile). |
|
||
| `enable_compliance` | `true` | `{true, false}` | **TUNE**: Real-time regulatory validation. |
|
||
|
||
### Entropy Regularization (WAVE 17)
|
||
|
||
| Field | Default | Valid Range | Rationale |
|
||
|-------|---------|-------------|-----------|
|
||
| `entropy_coefficient` | `None` | `[0.0, 0.1]` | **TUNE**: Entropy bonus for policy diversity. None = disabled. |
|
||
|
||
### Triple Barrier (Experimental - Default OFF)
|
||
|
||
| Field | Default | Valid Range | Rationale |
|
||
|-------|---------|-------------|-----------|
|
||
| `triple_barrier_profit_target_bps` | `100` | `[50, 500]` | **TUNE**: Profit target in bps (if enabled). |
|
||
| `triple_barrier_stop_loss_bps` | `50` | `[25, 200]` | **TUNE**: Stop loss in bps (if enabled). |
|
||
| `triple_barrier_max_holding_seconds` | `3600` | `[300, 7200]` | **TUNE**: Max holding period (if enabled). |
|
||
|
||
### Distributional RL (C51) - **SHOULD BE DISABLED**
|
||
|
||
| Field | Default | Valid Range | Rationale |
|
||
|-------|---------|-------------|-----------|
|
||
| `v_min` | `-2.0` | `[-10.0, -1.0]` | **TUNE** (if C51 enabled): Min value for distribution. BUG #5 fix: align with reward range ±2. |
|
||
| `v_max` | `2.0` | `[1.0, 10.0]` | **TUNE** (if C51 enabled): Max value for distribution. BUG #5 fix: align with reward range ±2. |
|
||
|
||
### WAVE 26 Advanced Features (Experimental - Default OFF)
|
||
|
||
#### Ensemble Uncertainty (Conflicts with Noisy Nets)
|
||
|
||
| Field | Default | Valid Range | Rationale |
|
||
|-------|---------|-------------|-----------|
|
||
| `ensemble_size` | `5` | `[3, 10]` | **TUNE** (if enabled): Number of Q-network heads. |
|
||
| `beta_variance` | `0.5` | `[0.1, 1.0]` | **TUNE** (if enabled): Variance penalty weight. |
|
||
| `beta_disagreement` | `0.5` | `[0.1, 1.0]` | **TUNE** (if enabled): Disagreement penalty weight. |
|
||
| `beta_entropy` | `0.1` | `[0.05, 0.5]` | **TUNE** (if enabled): Entropy bonus weight. |
|
||
| `variance_cap` | `1.0` | `[0.1, 2.0]` | **TUNE** (if enabled): Variance cap to prevent over-penalization. |
|
||
|
||
#### Gradient Accumulation
|
||
|
||
| Field | Default | Valid Range | Rationale |
|
||
|-------|---------|-------------|-----------|
|
||
| `gradient_accumulation_steps` | `1` | `[1, 8]` | **TUNE**: Mini-batches per optimizer step. Effective batch = batch_size × steps. |
|
||
|
||
#### Sharpe Ratio Reward (Experimental)
|
||
|
||
| Field | Default | Valid Range | Rationale |
|
||
|-------|---------|-------------|-----------|
|
||
| `sharpe_window` | `20` | `[10, 50]` | **TUNE** (if sharpe_weight > 0): Rolling window for Sharpe calculation. |
|
||
|
||
#### Adaptive Dropout (Experimental)
|
||
|
||
| Field | Default | Valid Range | Rationale |
|
||
|-------|---------|-------------|-----------|
|
||
| `dropout_initial` | `0.5` | `[0.3, 0.7]` | **TUNE** (if scheduler enabled): Initial dropout rate. |
|
||
| `dropout_final` | `0.1` | `[0.0, 0.3]` | **TUNE** (if scheduler enabled): Final dropout rate. |
|
||
| `dropout_anneal_steps` | `10000` | `[5000, 50000]` | **TUNE** (if scheduler enabled): Annealing steps. |
|
||
|
||
#### Hindsight Experience Replay (Experimental)
|
||
|
||
| Field | Default | Valid Range | Rationale |
|
||
|-------|---------|-------------|-----------|
|
||
| `her_strategy` | `"future"` | `{"final", "future"}` | **TUNE** (if her_ratio > 0): HER strategy. "future" > "final". |
|
||
|
||
#### GAE (Not Recommended for DQN)
|
||
|
||
| Field | Default | Valid Range | Rationale |
|
||
|-------|---------|-------------|-----------|
|
||
| `gae_lambda` | `0.95` | `[0.9, 0.99]` | **TUNE** (if GAE enabled): GAE λ parameter. Note: GAE is for actor-critic, not DQN. |
|
||
|
||
---
|
||
|
||
## CRITICAL ISSUES & CONFLICTS
|
||
|
||
### 🚨 **BUG**: Distributional RL Default Mismatch
|
||
|
||
**Field**: `use_distributional`
|
||
**Default**: `true`
|
||
**Documentation Says**: "C51/Distributional: OFF (Candle bug)"
|
||
**Problem**: Default contradicts stated requirement
|
||
|
||
**Required Fix**: Change default to `false` in line 638:
|
||
```rust
|
||
use_distributional: false, // Default: disabled (Candle bug - see ADR-001)
|
||
```
|
||
|
||
### ⚠️ **WARNING**: Mutually Exclusive Features
|
||
|
||
**Conflict 1**: `use_noisy_nets` vs `use_ensemble_uncertainty`
|
||
- Both provide exploration bonuses
|
||
- Should not be enabled simultaneously
|
||
- Default: `use_noisy_nets=true`, `use_ensemble_uncertainty=false` ✅ CORRECT
|
||
|
||
**Conflict 2**: `epsilon_decay` vs `use_noisy_nets`
|
||
- Noisy networks replace ε-greedy exploration
|
||
- With `use_noisy_nets=true`, ε should be set to 0
|
||
- Current: ε-greedy still active alongside noisy nets
|
||
- **Recommendation**: When `use_noisy_nets=true`, set `epsilon_start=0.0`
|
||
|
||
---
|
||
|
||
## RAINBOW DQN CONFIGURATION SUMMARY
|
||
|
||
### ✅ **CORRECT**: Enabled by Default (Rainbow Core)
|
||
1. Double DQN (`use_double_dqn=true`)
|
||
2. Dueling Networks (`use_dueling=true`)
|
||
3. Prioritized Experience Replay (`use_per=true`)
|
||
4. Noisy Networks (`use_noisy_nets=true`)
|
||
5. 3-Step Returns (`n_steps=3`)
|
||
6. Soft Target Updates (`tau=0.001`, `target_update_mode=Soft`)
|
||
|
||
### ❌ **INCORRECT**: Should Be Disabled (Candle Bug)
|
||
7. Distributional RL (`use_distributional=true` → **should be `false`**)
|
||
|
||
### Rainbow DQN Compliance Score
|
||
**6/7 components correct** (85.7%)
|
||
**Action Required**: Disable `use_distributional` to reach 100% Rainbow compliance.
|
||
|
||
---
|
||
|
||
## HYPEROPT SEARCH SPACE RECOMMENDATIONS
|
||
|
||
### Priority Tier 1 (Highest Impact)
|
||
1. `learning_rate` - Range: `[1e-5, 1e-3]`, log-scale
|
||
2. `batch_size` - Range: `[64, 256]`, categorical `{64, 128, 256}`
|
||
3. `gamma` - Range: `[0.95, 0.999]`
|
||
4. `epsilon_decay` - Range: `[0.99, 0.9999]`, log-scale
|
||
5. `gradient_clip_norm` - Range: `[10.0, 100.0]` or `None`
|
||
|
||
### Priority Tier 2 (Moderate Impact)
|
||
6. `buffer_size` - Range: `[100000, 1000000]`, categorical `{100K, 250K, 500K, 1M}`
|
||
7. `warmup_steps` - Range: `[0, 80000]`, adaptive based on total steps
|
||
8. `kelly_fractional` - Range: `[0.25, 1.0]`
|
||
9. `max_position_absolute` - Range: `[1.0, 10.0]`
|
||
10. `noisy_sigma_init` - Range: `[0.3, 0.7]`
|
||
|
||
### Priority Tier 3 (Fine-Tuning)
|
||
11. Early stopping parameters (`q_value_floor`, `plateau_window`, `min_epochs_before_stopping`)
|
||
12. Penalty weights (`hold_penalty`, `hold_penalty_weight`, `transaction_cost_multiplier`)
|
||
13. Learning rate schedule parameters (if `lr_decay_type != Constant`)
|
||
14. Risk management toggles (Kelly, volatility epsilon, risk-adjusted rewards)
|
||
|
||
---
|
||
|
||
## VALIDATION CHECKLIST
|
||
|
||
### Before Hyperopt
|
||
- [ ] Set `use_distributional=false` (Candle bug)
|
||
- [ ] Verify `use_double_dqn=true`
|
||
- [ ] Verify `use_dueling=true`
|
||
- [ ] Verify `use_per=true`
|
||
- [ ] Verify `use_noisy_nets=true`
|
||
- [ ] Verify `n_steps=3`
|
||
- [ ] Verify `tau=0.001`
|
||
- [ ] Verify `target_update_mode=Soft`
|
||
|
||
### After Hyperopt
|
||
- [ ] Confirm batch_size ≤ 230 (RTX 3050 Ti limit)
|
||
- [ ] Confirm `learning_rate` in validated range
|
||
- [ ] Confirm `gamma` balance (not too high, not too low)
|
||
- [ ] Confirm `warmup_steps` scales with total training steps
|
||
- [ ] Save best hyperparameters to `ml/hyperparams/dqn_best.toml`
|
||
|
||
---
|
||
|
||
## TOTAL FIELD COUNT
|
||
|
||
- **Total Fields**: 83
|
||
- **Fixed Flags**: 28 (33.7%)
|
||
- **Tunable Hyperparameters**: 55 (66.3%)
|
||
|
||
**Field Breakdown**:
|
||
- Rainbow Core (Fixed): 7
|
||
- Architectural (Fixed): 14
|
||
- Experimental Disabled (Fixed): 7
|
||
- Learning/Optimization (Tunable): 12
|
||
- Exploration (Tunable): 9
|
||
- Replay Buffer (Tunable): 2
|
||
- Early Stopping (Tunable): 7
|
||
- Portfolio/Risk (Tunable): 25
|
||
- Advanced Features (Experimental, Tunable if enabled): 21
|
||
|
||
---
|
||
|
||
## REFERENCES
|
||
|
||
- Rainbow DQN Paper: Hessel et al. (2018)
|
||
- Trial 19 Hyperopt Results: `DQN_HYPEROPT_RESULTS_SUMMARY.md`
|
||
- BUG #5 Fix: v_min/v_max alignment with reward range
|
||
- BUG #9 Fix: target_update_frequency = 500 steps optimal
|
||
- BUG #12 Fix: huber_delta scaled 100x
|
||
- WAVE 16 (Agent 36): Soft target updates
|
||
- WAVE 23: Feature caching + early stopping
|
||
- WAVE 24: Buffer size increase to 500K
|
||
- WAVE 26: Advanced features (LR scheduling, ensemble, etc.)
|
||
- WAVE 35: Regime-conditional Q-networks + compliance
|