BREAKING CHANGES: - Removed orphaned dqn.rs monolithic trainer (4,975 lines) - Removed orphaned dqn_ensemble.rs module (816 lines) - Removed orphaned tft.rs and tft_complete_int8_integration_test.rs - TFT trainer split into modular directory structure DQN Module Refactoring: - Split trainers/dqn.rs into modular structure (config.rs, statistics.rs, trainer.rs) - Fixed hyperopt 39D search space (continuous params only) - Boolean flags (use_dueling, use_double_dqn, use_per, use_noisy_nets) are now FIXED architectural decisions - use_distributional defaults to false (Candle BUG #36 - scatter_add gradient issues) Clean Module Structure: - ml/src/trainers/dqn/ directory with proper mod.rs exports - ml/src/trainers/tft/ directory with config.rs, types.rs, model.rs, trainer.rs, tests.rs - All P0 features validated: TD-error clamping, batch diversity, LR scheduler, priority staleness Documentation: - Added comprehensive docs in docs/codebase-cleanup/ - ADR-001 for DQN refactoring decisions - Rainbow DQN component matrix and quick reference guides Build Status: Compiles with zero errors 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com>
156 lines
4.4 KiB
Markdown
156 lines
4.4 KiB
Markdown
# WAVE 28.11: DQNConfig Default Implementation
|
|
|
|
## Summary
|
|
|
|
Added `Default` implementation for `DQNConfig` struct in `/home/jgrusewski/Work/foxhunt/ml/src/dqn/dqn.rs`.
|
|
|
|
## Changes Made
|
|
|
|
### File: `/home/jgrusewski/Work/foxhunt/ml/src/dqn/dqn.rs`
|
|
|
|
Added comprehensive `Default` implementation between the struct definition (line 146) and the existing `impl DQNConfig` block (line 195).
|
|
|
|
**Location**: Lines 148-193
|
|
|
|
### Default Values Chosen
|
|
|
|
The defaults are optimized for trading DQN with sensible production-ready parameters:
|
|
|
|
#### Core Architecture
|
|
- `state_dim: 54` - Standard feature dimension
|
|
- `num_actions: 45` - FactoredAction space
|
|
- `hidden_dims: vec![256, 256]` - Two hidden layers
|
|
|
|
#### Learning Parameters
|
|
- `learning_rate: 1e-4` - Conservative learning rate
|
|
- `gamma: 0.99` - Standard discount factor
|
|
- `gradient_clip_norm: 1.0` - Gradient clipping threshold
|
|
|
|
#### Exploration
|
|
- `epsilon_start: 1.0` - Start with full exploration
|
|
- `epsilon_end: 0.01` - Minimal exploration at end
|
|
- `epsilon_decay: 0.995` - Gradual decay
|
|
|
|
#### Replay Buffer
|
|
- `replay_buffer_capacity: 100_000` - Large capacity
|
|
- `batch_size: 64` - Standard batch size
|
|
- `min_replay_size: 1000` - Minimum before training
|
|
|
|
#### Target Updates
|
|
- `target_update_freq: 1000` - Hard update frequency
|
|
- `tau: 0.005` - Soft update coefficient
|
|
- `use_soft_updates: true` - Enable Polyak averaging
|
|
|
|
#### Rainbow DQN Features
|
|
- `use_double_dqn: true` - Enable Double DQN
|
|
- `use_huber_loss: true` - Use Huber loss
|
|
- `huber_delta: 1.0` - Huber loss threshold
|
|
- `warmup_steps: 1000` - Warmup before training
|
|
- `n_steps: 1` - Single-step returns (conservative)
|
|
|
|
#### Prioritized Experience Replay (PER)
|
|
- `use_per: true` - Enable PER
|
|
- `per_alpha: 0.6` - Prioritization exponent
|
|
- `per_beta_start: 0.4` - Initial importance sampling weight
|
|
- `per_beta_max: 1.0` - Maximum beta value
|
|
- `per_beta_annealing_steps: 100_000` - Annealing schedule
|
|
|
|
#### Dueling Networks
|
|
- `use_dueling: true` - Enable dueling architecture
|
|
- `dueling_hidden_dim: 128` - Advantage stream hidden dim
|
|
|
|
#### Distributional RL (C51)
|
|
- `use_distributional: false` - Disabled by default
|
|
- `num_atoms: 51` - Distribution atoms
|
|
- `v_min: -10.0` - Minimum value
|
|
- `v_max: 10.0` - Maximum value
|
|
|
|
#### Noisy Networks
|
|
- `use_noisy_nets: false` - Disabled by default
|
|
- `noisy_sigma_init: 0.5` - Initial noise std
|
|
|
|
#### Q-Value Clipping (BUG #37 Fix)
|
|
- `enable_q_value_clipping: true` - Enable clipping
|
|
- `q_value_clip_min: -100.0` - Minimum Q-value
|
|
- `q_value_clip_max: 100.0` - Maximum Q-value
|
|
|
|
#### Early Stopping (WAVE 23)
|
|
- `gradient_collapse_multiplier: 2.0` - Learning-rate aware threshold
|
|
- `gradient_collapse_patience: 100` - Epochs before stopping
|
|
|
|
#### Trading Parameters
|
|
- `initial_capital: 100_000.0` - Starting capital
|
|
- `leaky_relu_alpha: 0.01` - LeakyReLU negative slope
|
|
|
|
## Verification
|
|
|
|
All 41 struct fields are covered in the Default implementation:
|
|
- state_dim
|
|
- num_actions
|
|
- hidden_dims
|
|
- learning_rate
|
|
- gamma
|
|
- epsilon_start
|
|
- epsilon_end
|
|
- epsilon_decay
|
|
- replay_buffer_capacity
|
|
- batch_size
|
|
- min_replay_size
|
|
- target_update_freq
|
|
- use_double_dqn
|
|
- use_huber_loss
|
|
- huber_delta
|
|
- leaky_relu_alpha
|
|
- gradient_clip_norm
|
|
- tau
|
|
- use_soft_updates
|
|
- warmup_steps
|
|
- n_steps
|
|
- initial_capital
|
|
- use_per
|
|
- per_alpha
|
|
- per_beta_start
|
|
- per_beta_max
|
|
- per_beta_annealing_steps
|
|
- use_dueling
|
|
- dueling_hidden_dim
|
|
- use_distributional
|
|
- num_atoms
|
|
- v_min
|
|
- v_max
|
|
- use_noisy_nets
|
|
- noisy_sigma_init
|
|
- enable_q_value_clipping
|
|
- q_value_clip_min
|
|
- q_value_clip_max
|
|
- gradient_collapse_multiplier
|
|
- gradient_collapse_patience
|
|
|
|
## Notes
|
|
|
|
1. The defaults are **conservative and production-ready**
|
|
2. Rainbow features are selectively enabled (Double DQN, Huber, PER, Dueling)
|
|
3. More experimental features (Distributional, Noisy Nets) are disabled by default
|
|
4. Q-value clipping is enabled to prevent explosions (BUG #37 fix)
|
|
5. Early stopping with gradient collapse detection is configured (WAVE 23)
|
|
|
|
## Usage
|
|
|
|
```rust
|
|
// Create config with sensible defaults
|
|
let config = DQNConfig::default();
|
|
|
|
// Or customize specific fields
|
|
let config = DQNConfig {
|
|
learning_rate: 1e-3,
|
|
use_distributional: true,
|
|
..DQNConfig::default()
|
|
};
|
|
```
|
|
|
|
## Compilation Status
|
|
|
|
The Default implementation itself is syntactically correct and complete.
|
|
|
|
Note: There are pre-existing compilation errors in the ml crate related to type mismatches between `ml::dqn::dqn::DQNConfig` and `ml::dqn::agent::DQNConfig`. These are separate issues not introduced by this change.
|