Extends the reward inventory with a review pass that preserves the
*ideas* behind behavioral terms while deleting the wrong-level mechanics.
Gems (relocate, don't delete):
- Probabilistic fill model (real physics, not shaping) → apply in both
modes via shared Philox RNG + exploration_scale
- Kelly-as-physics-constraint → hard position-size cap in trade_physics,
not a soft reward
- Path quality via per-step drawdown integration → no separate term,
just run drawdown penalty every bar
- Saboteur relocation → both modes, scaled by exploration_scale (training
1.0, validation 0.5, diagnostic 0.0) — avoids creating a new clean-vs-
noisy mismatch in the opposite direction
Pearls:
- Sparse reward is fine if TD propagation works — diagnose before
deleting micro_reward_scale; measure cov(Q(s_entry,a), trade_return)
- Every behavioral term maps to one of four failure modes: double-count,
misplaced physics, wrong-level regularization, or compensating for a
downstream bug. None are semantically "neutral" shaping.
Novel architectural moves:
- Diagnostic-only entropy tracking (no reward, just HEALTH_DIAG
logging)
- Single exploration_scale scalar replacing all mode toggles
(feedback_no_feature_flags compliant)
- Layered reward with per-term budget caps — makes the train/val Sharpe
gap computable instead of uncomputable magic
- Q-target smoothing replacing reward_noise_scale (regularization at
gradient level, not reward level)
Updated disposition table: 4 DELETE, 2 MOVE, 2 RELOCATE-to-physics,
1 DIAGNOSE-then-DELETE, 4 delete-dead-plumbing.
Three relocations (Kelly cap, saboteur scaling, Q-target smoothing) can
land as independent PRs before the unified_env_kernel rewrite, shrinking
Phase 2's change surface.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>