Files
foxhunt/crates
jgrusewski caf0c07121 fix: adaptive lambda scaling for trade-level reward system
Homeostatic lambda_base: 0.01 (fixed) → adaptive 0.01/q_gap (scales
inversely with Q-value range). Trade-level reward has 300× smaller
reward std → Q-values are proportionally smaller → fixed lambda too weak.
budget_max scales with lambda for consistent budget ratio.

C51 grad drift penalty: 0.01 → 0.1 (10× stronger to match smaller
Q-value scale from trade-level rewards).

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-16 09:39:35 +02:00
..