Files
foxhunt/GAMMA_0.90_TEST_RESULTS.md
jgrusewski 00ef9e2866 Wave 15: Complete FactoredAction migration to 45-action system
Major Changes:
- Migrated from 3-action TradingAction to 45-action FactoredAction
- 45 actions: 5 exposure × 3 order types × 3 urgency levels
- Absolute exposure model (target positions -1.0 to +1.0)
- Transaction cost differentiation (Market 0.15%, LimitMaker 0.05%, IoC 0.10%)
- Fixed action diversity threshold (1.11% → 0.5% for 45-action space)

Bug Fixes:
- Bug #15: Incomplete FactoredAction integration (code existed but unused)
- Bug #16: Runtime crash in action diversity checking (hardcoded 3-action match)

Code Changes (13 files, ~464 lines):
- ml/src/dqn/action_space.rs: Core FactoredAction + 4 helper methods
- ml/src/trainers/dqn.rs: Action diversity refactored (3→45 dynamic)
- ml/src/dqn/reward.rs: calculate_reward() signature updated
- ml/src/dqn/portfolio_tracker.rs: execute_action() absolute exposure
- ml/src/dqn/dqn.rs: WorkingDQN action selection migrated
- ml/tests/*.rs: 9 test files updated with FactoredAction assertions

Test Results:
- 1-epoch smoke test: 100% action diversity (45/45 actions, 80.2s)
- 10-epoch production: 87.8% readiness (79/90 scorecard, 14.0 min)
- Loss convergence: 96.9% reduction (119K → 3.6K)
- Action diversity: 100% → 44% (healthy specialization)
- Checkpoint reliability: 12/12 files saved (100%)
- DQN tests: 195/195 passing (100%)
- ML baseline: 1,514/1,515 passing (99.93%)

Production Status:  CERTIFIED (87.8% readiness)
Go/No-Go:  GO FOR 100-EPOCH PRODUCTION TRAINING

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
2025-11-11 23:27:02 +01:00

6.4 KiB
Raw Blame History

Gamma Reduction Test Results

Test Configuration

  • Dataset: ES_FUT_180d.parquet (174,053 bars, +23.9% return)
  • Epochs: 10 (5 completed epochs analyzed)
  • Device: CUDA GPU (RTX 3050 Ti)
  • Gamma Modified: 0.9626 → 0.90 (56% noise amplification reduction)

Critical Findings: Gamma DID NOT Help

Problem Persists with Gamma=0.90

Observation: Gradient collapse still occurs at step 20 onwards, identical to gamma=0.9626 runs.

Evidence from Logs:

  • Step 10: grad=12.4702 (healthy)
  • Step 20: grad=0.0000 ← GRADIENT COLLAPSE
  • Step 30-21700: grad=0.0000 (continuous collapse)

Epochs 3-5:

  • Q-value: -333.3333 (stuck at constant)
  • Q_std: 942.81 (constant)
  • Q_range: 2000.00 (constant, clamped at ±1000 limit)
  • grad_norm: 0.000000 (complete collapse)

Gamma 0.90 Results (Current Test)

Epoch Q-value Train Loss Grad Norm Q_range Status
1 5.74 9.405 0.225 633.60 Learning
2 -282.76 9.414 0.232 1442.79 ⚠️ Degrading
3 -333.33 9.398 0.000 2000.00 COLLAPSED
4 -333.33 9.413 0.000 2000.00 COLLAPSED
5 -333.33 9.420 0.000 2000.00 COLLAPSED

Q-Value Progression (Every 10 Steps, Epoch 1)

Step 10:   BUY=-128.58, SELL=-156.74, HOLD=-373.16  (grad=12.47)
Step 20:   BUY=-72.82,  SELL=-142.84, HOLD=-270.09  (grad=0.00) ← COLLAPSE
Step 30:   BUY=-49.30,  SELL=-134.15, HOLD=-222.33  (grad=0.00)
...
Step 260:  BUY=79.29,   SELL=-151.78, HOLD=-157.49  (grad=0.00)
Step 270:  BUY=280.69,  SELL=-211.58, HOLD=-121.32  (grad=0.00)
Step 280:  BUY=351.17,  SELL=-234.52, HOLD=-110.06  (grad=0.00)
Step 400:  BUY=393.62,  SELL=-247.88, HOLD=-104.34  (grad=0.00)

Pattern: Q-values drift wildly (±600 range) with ZERO gradients after step 20.


Comparative Analysis: Gamma 0.9626 vs 0.90

Gradient Health

Metric Gamma 0.9626 Gamma 0.90 Improvement
Gradient collapse step 20 20 IDENTICAL
Zero gradient occurrences 100% (steps 20+) 100% (steps 20+) NO CHANGE
Epochs with grad=0 3-10 3-5+ NO CHANGE

Q-Value Stability

Metric Gamma 0.9626 Gamma 0.90 Improvement
Epoch 1 Q-value +5.74 +5.74 Same
Epoch 3 Q-value -333.33 (stuck) -333.33 (stuck) IDENTICAL
Q-value range (epoch 3) 2000.00 2000.00 IDENTICAL
Q-value convergence NO NO NO IMPROVEMENT

Loss Convergence

Metric Gamma 0.9626 Gamma 0.90 Improvement
Loss pattern Stuck 9.4-9.42 Stuck 9.4-9.42 IDENTICAL
Loss decreasing trend NO NO NO CHANGE

Key Finding: Problem is NOT Gamma

Evidence

  1. Gradient collapse timing: Identical (step 20)
  2. Collapse pattern: Identical (0.000 gradient from step 20 onwards)
  3. Q-value behavior: Identical (-333.33 stuck value after epoch 2)
  4. Loss stagnation: Identical (9.4-9.42 range)

Theoretical vs. Actual

  • Theory: γ=0.90 reduces noise amplification by 56% (10× vs 22× over 50 steps)
  • Actual: γ=0.90 produces IDENTICAL gradient collapse at same step as γ=0.9626

Conclusion: Gamma is NOT the root cause of gradient collapse.


Root Cause Assessment

What Gamma Reduction Ruled Out

Discount factor amplifying noise over long horizons Future value estimation instability
Bootstrapping feedback loop from distant rewards

What Remains (Actual Root Causes)

  1. Reward Scale Mismatch HIGHEST PRIORITY

    • Reward normalization scale: 3197.23× (calculated from 0.0313% typical move)
    • May not match actual P&L variance
    • Elite reward system compounds scaling issues
  2. Network Architecture Issues LIKELY

    • Dead neurons: 0.00% reported (may be false negative)
    • Activation function (LeakyReLU 0.01) may saturate
    • Hidden dims [512, 256, 128, 64] may be over-parameterized
  3. Optimizer Instability CONFIRMED

    • Adam epsilon: 1.5e-4 (Rainbow DQN standard)
    • Gradient clipping: max_norm=10.0 (not preventing collapse)
    • Learning rate: 0.0001 (may be too high for unstable gradients)
  4. TD-Error Explosion LIKELY

    • TD-error clipping: ±10.0 (may be insufficient)
    • Target network: Hard updates every 10K steps (sudden shifts)
    • Huber loss delta: 10.0 (may need tighter bound)

1. Reward System Analysis (IMMEDIATE - 30 MIN)

Test SimplePnL reward (no Elite multi-component) to isolate reward scaling issues:

# Remove Elite reward complexity
--reward-system SimplePnL --epochs 10

Expected: If SimplePnL shows healthier gradients, Elite reward is compounding instability.

2. Learning Rate Reduction (HIGH PRIORITY - 15 MIN)

Test LR 10× lower (1e-5 instead of 1e-4):

--learning-rate 0.00001 --epochs 10

Expected: Slower gradient changes may prevent collapse at step 20.

3. Reward Normalization Tuning (HIGH PRIORITY - 30 MIN)

Override adaptive reward_scale (currently 3197.23×):

# Test 10× smaller scale
--reward-scale 300.0 --epochs 10

Expected: Reduced scaling may prevent Q-value explosions.

4. TD-Error Clipping Tightening (MEDIUM PRIORITY - 15 MIN)

Reduce TD-error clip from ±10.0 to ±1.0:

--td-error-clip 1.0 --epochs 10

Expected: Tighter clipping prevents Bellman update explosions.

5. Network Architecture Simplification (LOW PRIORITY - 45 MIN)

Reduce hidden dims from [512, 256, 128, 64] to [128, 64]:

  • Modify emergency_safe_defaults() in dqn.rs
  • Smaller network may stabilize gradients

Success Criteria

Gradient health: No collapse before epoch 5, grad_norm >1.0 throughout training Q-value stability: Q-values stay in ±100 range, smooth convergence
Loss decreasing: 50%+ reduction from epoch 1 to 10 Action diversity: No action >50% at any epoch

Current Status: ALL CRITERIA FAILED with gamma=0.90


Deployment Recommendation

DO NOT deploy gamma=0.90 - provides zero improvement over baseline.

Priority Investigation: Reward system (SimplePnL test) + learning rate reduction (1e-5).