Files
foxhunt/DQN_FACTORED_ACTIONS_BUG_REPORT.md
jgrusewski 00ef9e2866 Wave 15: Complete FactoredAction migration to 45-action system
Major Changes:
- Migrated from 3-action TradingAction to 45-action FactoredAction
- 45 actions: 5 exposure × 3 order types × 3 urgency levels
- Absolute exposure model (target positions -1.0 to +1.0)
- Transaction cost differentiation (Market 0.15%, LimitMaker 0.05%, IoC 0.10%)
- Fixed action diversity threshold (1.11% → 0.5% for 45-action space)

Bug Fixes:
- Bug #15: Incomplete FactoredAction integration (code existed but unused)
- Bug #16: Runtime crash in action diversity checking (hardcoded 3-action match)

Code Changes (13 files, ~464 lines):
- ml/src/dqn/action_space.rs: Core FactoredAction + 4 helper methods
- ml/src/trainers/dqn.rs: Action diversity refactored (3→45 dynamic)
- ml/src/dqn/reward.rs: calculate_reward() signature updated
- ml/src/dqn/portfolio_tracker.rs: execute_action() absolute exposure
- ml/src/dqn/dqn.rs: WorkingDQN action selection migrated
- ml/tests/*.rs: 9 test files updated with FactoredAction assertions

Test Results:
- 1-epoch smoke test: 100% action diversity (45/45 actions, 80.2s)
- 10-epoch production: 87.8% readiness (79/90 scorecard, 14.0 min)
- Loss convergence: 96.9% reduction (119K → 3.6K)
- Action diversity: 100% → 44% (healthy specialization)
- Checkpoint reliability: 12/12 files saved (100%)
- DQN tests: 195/195 passing (100%)
- ML baseline: 1,514/1,515 passing (99.93%)

Production Status:  CERTIFIED (87.8% readiness)
Go/No-Go:  GO FOR 100-EPOCH PRODUCTION TRAINING

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
2025-11-11 23:27:02 +01:00

349 lines
11 KiB
Markdown

# DQN Factored Actions Bug Report - CRITICAL DISCOVERY
**Status**: CRITICAL BUG FOUND - Only 3 out of 45 factored actions are selectable
**Date**: 2025-11-11
**Severity**: CATASTROPHIC (45-action space is completely non-functional)
**Impact**: All training uses 3-action space (Buy/Sell/Hold) instead of 45-action factored space
---
## Executive Summary
The DQN is hardcoded to use **only 3 actions** regardless of the feature flag setting for factored actions. The 45-action factored space infrastructure exists but is **never activated** because:
1. **`num_actions` always defaults to 3** in the WorkingDQNConfig initialization
2. **Feature flag compilation is broken**: The condition checks compile `num_actions: 45` OR `num_actions: 3` but trainer defaults to `num_actions: 3` regardless of flag
3. **FactoredQNetwork is implemented correctly** but completely bypassed (initialized as `None` in trainer, never used)
4. **Action selection still uses TradingAction enum** (Buy/Sell/Hold) instead of FactoredAction
---
## Root Cause Analysis
### Bug #1: Trainer Always Creates 3-Action Config (Line 614-619)
**File**: `/home/jgrusewski/Work/foxhunt/ml/src/trainers/dqn.rs`
```rust
// Lines 614-619
let config = WorkingDQNConfig {
state_dim: 128,
#[cfg(feature = "factored-actions")]
num_actions: 45, // ← SET TO 45 WHEN FEATURE FLAG ENABLED
#[cfg(not(feature = "factored-actions"))]
num_actions: 3, // ← SET TO 3 WHEN FEATURE FLAG DISABLED
// ... rest of config
};
```
**PROBLEM**: This code is CORRECT! The feature flag properly sets `num_actions` to either 3 or 45. **BUT** the trainer initialization always uses 3-action mode.
### Bug #2: CLI Never Actually Enables Feature Flag
The `--use-factored-actions` CLI flag doesn't enable the `factored-actions` feature at **compile time**. The binary needs to be compiled with:
```bash
cargo build --features factored-actions
```
Without this compile-time flag, the code compiles with `#[cfg(not(feature = "factored-actions"))]`, forcing `num_actions: 3`.
### Bug #3: FactoredQNetwork Created But Never Used (Lines 728-731)
```rust
// Lines 728-731
#[cfg(feature = "factored-actions")]
factored_network: None, // ← ALWAYS INITIALIZED AS NONE!
#[cfg(feature = "factored-actions")]
use_factored_actions: false, // ← ALWAYS FALSE!
```
**PROBLEM**: Even if the feature flag was enabled:
- `factored_network` is initialized as `None` and never created
- `use_factored_actions` is hardcoded to `false`
- The trainer never calls FactoredQNetwork methods for action selection
- Instead, it continues using TradingAction (Buy/Sell/Hold)
### Bug #4: Action Selection Still Uses 3-Action TradingAction (Lines 263-268)
```rust
// In TrainingMonitor (lines 263-268)
fn track_action(&mut self, action: &TradingAction) {
let idx = match action {
TradingAction::Buy => 0,
TradingAction::Sell => 1,
TradingAction::Hold => 2,
};
self.action_counts[idx] += 1;
}
```
**PROBLEM**: This only tracks 3 actions. When factored actions are enabled, we should be tracking FactoredAction with indices 0-44.
---
## Evidence: Hardcoded Constants
**File**: `/home/jgrusewski/Work/foxhunt/ml/src/trainers/dqn.rs` lines 228-232
```rust
// Action count depends on feature flag
#[cfg(feature = "factored-actions")]
const NUM_ACTIONS: usize = 45;
#[cfg(not(feature = "factored-actions"))]
const NUM_ACTIONS: usize = 3;
```
This is correct at **compile time**, but training data shows only 3 actions used, meaning the binary was compiled WITHOUT the `factored-actions` feature flag.
---
## Why Only 3 Actions Are Selected
### Scenario A: Feature Flag NOT Enabled (Current State)
If compiled without `--features factored-actions`:
1. `NUM_ACTIONS = 3`
2. `num_actions = 3` (line 619)
3. Q-network outputs 3 Q-values (one per action)
4. Action selection argmax picks from 3 indices: [0, 1, 2] = [BUY, SELL, HOLD]
5. FactoredQNetwork never instantiated
6. Training produces 3-action distribution
**Result**: Only 3 actions available. ✅ Explains observed behavior.
### Scenario B: Feature Flag Enabled But CLI Flag Not Propagated
If compiled WITH `--features factored-actions` but CLI `--use-factored-actions` not activated:
1. `NUM_ACTIONS = 45`
2. `num_actions = 45` (line 617)
3. Q-network outputs 45 Q-values
4. BUT `use_factored_actions = false` (line 731)
5. Action selection still uses TradingAction enum (only 3 variants)
6. Argmax on 45 Q-values returns indices 0-44
7. BUT code tries to convert to TradingAction (only 3 valid)
**Result**: Runtime error or silent fallback to actions 0-2.
---
## FactoredQNetwork Implementation Status
**File**: `/home/jgrusewski/Work/foxhunt/ml/src/dqn/factored_q_network.rs`
The FactoredQNetwork is **fully implemented and correct**:
**Structure** (lines 44-59):
- `shared_encoder`: 128 → 64
- `exposure_head`: 64 → 5
- `order_head`: 64 → 3
- `urgency_head`: 64 → 3
**Forward Pass** (lines 123-154):
- Computes 3 separate heads
- Returns (5, 3, 3) tensors
**Joint Q-Values** (lines 161-201):
- Combines via additive factorization: Q(s,a) = Q_exp + Q_ord + Q_urg
- Returns [batch, 45] tensor ✅
**Action Selection** (lines 204-260):
- `select_greedy_action()`: Takes argmax per head
- `select_epsilon_greedy()`: Random factored action exploration
- Both return FactoredAction (not TradingAction)
**Problem**: This network is created but **NEVER INSTANTIATED** in the trainer.
---
## Action Space Mapping (Correct Implementation)
**File**: `/home/jgrusewski/Work/foxhunt/ml/src/dqn/action_space.rs`
Action mapping is **fully correct**:
```
Index = exposure_idx * 9 + order_idx * 3 + urgency_idx
Exposure (5 options): 0=Short100, 1=Short50, 2=Flat, 3=Long50, 4=Long100
Order (3 options): 0=Market, 1=LimitMaker, 2=IoC
Urgency (3 options): 0=Patient, 1=Normal, 2=Aggressive
Example: (Flat=2, Market=0, Normal=1) → 2*9 + 0*3 + 1 = 19 ✅
Example: (Long100=4, Market=0, Aggressive=2) → 4*9 + 0*3 + 2 = 38 ✅
```
All 45 combinations are unique and valid (verified by round-trip tests).
---
## How to Fix
### Short-term Fix: Enable Feature Flag at Compile Time
```bash
# Currently broken:
cargo build --release
# Must use:
cargo build --release --features factored-actions
# Or in Cargo.toml:
cargo run --features factored-actions --example train_dqn --release
```
**Problem**: This only addresses compilation. Action selection still broken (Bug #3).
### Long-term Fix: Implement 45-Action Selection in Trainer
The trainer needs to be refactored to:
1. **Actually create FactoredQNetwork** instead of passing `None`
```rust
#[cfg(feature = "factored-actions")]
let factored_network = if use_factored_actions {
Some(Arc::new(RwLock::new(
FactoredQNetwork::new(128, &device)?
)))
} else {
None
};
```
2. **Use FactoredQNetwork for action selection** when enabled
```rust
if self.use_factored_actions {
// Use FactoredQNetwork.select_epsilon_greedy()
let factored_action = self.factored_network
.as_ref()
.unwrap()
.read()
.await
.select_epsilon_greedy(&state_tensor, epsilon)?;
// Convert to action index for storage
} else {
// Use standard 3-action selection (current)
}
```
3. **Refactor action tracking** to support both 3 and 45 actions
```rust
// Current: hardcoded for 3 actions
action_counts: vec![0; NUM_ACTIONS],
// Already correct via feature flag!
// But tracking logic needs to handle FactoredAction
```
4. **Update action-to-reward mapping** for factored actions
- Current code maps TradingAction → reward
- Need to map FactoredAction → exposure, order, urgency → reward
---
## Test Cases Affected
Files that expect 3 actions but would break with 45:
| File | Issue | Impact |
|------|-------|--------|
| `ml/tests/dqn_factored_smoke_tests.rs` | Tests factored-actions feature | Will fail with 45 actions until trainer is fixed |
| `ml/src/dqn/tests/factored_integration_tests.rs` | Integration tests | Needs updated action selection logic |
| `ml/examples/train_dqn.rs` | CLI training example | Works but uses 3-action fallback |
| `ml/src/trainers/dqn.rs` lines 263-281 | TrainingMonitor | Hard-coded 3-action tracking |
---
## Current Training Status
**What's Happening**:
1. Binary compiled without `factored-actions` feature
2. `num_actions = 3` (forced by #[cfg(not(feature = "factored-actions"))])
3. Q-network has 3 outputs (Buy, Sell, Hold)
4. Argmax selects from [0, 1, 2]
5. Actions stored as TradingAction variants
**Result**: Only 3 actions available ❌
---
## Verification Commands
```bash
# Check if binary compiled with factored-actions feature
grep "const NUM_ACTIONS: usize = " ml/src/trainers/dqn.rs
# Expected: Should show NUM_ACTIONS = 45 if compiled with feature
# Check training logs
grep "Action Distribution" target/release/examples/train_dqn.log
# Current output: BUY=XX% SELL=XX% HOLD=XX%
# Expected with fix: Top 10 actions with index 0-44
# Compile with factored-actions (doesn't fully fix, but required step)
cargo build --release --features factored-actions
```
---
## Recommendations
### Priority 1: Implement Full 45-Action Support
- **Effort**: 2-4 hours
- **Steps**:
1. Create FactoredQNetwork in trainer when feature enabled
2. Route action selection to FactoredQNetwork.select_epsilon_greedy()
3. Update TrainingMonitor to track 45 actions
4. Update reward calculation for FactoredAction
### Priority 2: Add --use-factored-actions CLI Flag
- **Effort**: 30 minutes
- **Steps**:
1. Add `--use-factored-actions` flag to train_dqn.rs
2. Pass flag to DQNTrainer::new_with_factored_actions()
3. Set `use_factored_actions = true` in trainer
### Priority 3: Validation Tests
- **Effort**: 1 hour
- **Steps**:
1. Create test that verifies 45 actions are selectable
2. Verify action-to-exposure-order-urgency mapping
3. Validate that all combinations (0-44) can be reached
---
## Files to Modify
```
ml/src/trainers/dqn.rs
- Line 728: Initialize FactoredQNetwork properly
- Line 731: Set use_factored_actions from CLI flag
- Lines 263-281: Update track_action() for 45 actions
- Lines 1600+: Update action selection logic
ml/examples/train_dqn.rs
- Add --use-factored-actions flag
- Pass to DQNTrainer initialization
ml/src/dqn/dqn.rs
- Verify Q-network output dimension matches num_actions (should be automatic)
```
---
## Conclusion
**The 45-action factored space is fully implemented but completely disconnected from the training pipeline.** The trainer:
1. ✅ Sets `num_actions = 45` when feature flag enabled
2. ✅ FactoredQNetwork is fully functional
3. ❌ **Never instantiates FactoredQNetwork**
4. ❌ **Still uses TradingAction for selection** (3 variants only)
5. ❌ **CLI flag --use-factored-actions doesn't exist**
**Result**: Training always uses 3 actions, regardless of feature flag or infrastructure availability.
**Expected behavior after fix**: With `--features factored-actions --use-factored-actions`, should see all 45 actions selected with proper exposure/order/urgency combinations.