jgrusewski
caf0c07121
fix: adaptive lambda scaling for trade-level reward system
Homeostatic lambda_base: 0.01 (fixed) → adaptive 0.01/q_gap (scales
inversely with Q-value range). Trade-level reward has 300× smaller
reward std → Q-values are proportionally smaller → fixed lambda too weak.
budget_max scales with lambda for consistent budget ratio.
C51 grad drift penalty: 0.01 → 0.1 (10× stronger to match smaller
Q-value scale from trade-level rewards).
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-16 09:39:35 +02:00
..
2026-04-11 11:54:09 +02:00
2026-04-02 16:29:55 +02:00
2026-03-14 11:35:15 +01:00
2026-03-13 10:18:35 +01:00
2026-03-13 10:18:35 +01:00
2026-04-16 09:39:35 +02:00
2026-03-13 10:18:35 +01:00
2026-03-13 10:18:35 +01:00
2026-03-13 10:18:35 +01:00
2026-04-13 17:09:34 +02:00
2026-03-13 10:18:35 +01:00
2026-04-13 19:28:00 +02:00
2026-03-29 22:35:37 +02:00
2026-03-29 22:35:37 +02:00
2026-04-08 21:31:46 +02:00
2026-03-29 19:18:47 +02:00
2026-03-19 00:39:03 +01:00
2026-03-13 10:18:35 +01:00
2026-03-13 10:18:35 +01:00
2026-04-13 17:09:34 +02:00
2026-03-16 21:01:28 +01:00
2026-03-13 10:18:35 +01:00
2026-03-13 10:18:35 +01:00
2026-03-13 10:18:35 +01:00
2026-04-10 18:54:14 +02:00
2026-03-13 10:18:35 +01:00
2026-03-13 10:18:35 +01:00
2026-03-13 10:18:35 +01:00
2026-03-18 00:53:47 +01:00
2026-03-14 11:35:15 +01:00
2026-03-15 11:59:31 +01:00