Two fixes that decouple entropy stability from loss aversion:
1. Thompson sampling probability floor (RL_THOMPSON_FLOOR_INDEX=588,
bootstrap 0.05): 5% of steps pick uniform random action. Prevents
any action from reaching π=0 (δ-function attractor). ISV-driven.
Per pearl_pi_actor_collapses_without_entropy_floor.
2. LOSS bootstrap restored to 3.0 (was 1.5). LOSS=1.5 caused entropy
collapse to 0.94 despite max SAC. LOSS=3.0 keeps entropy stable at
1.87 (proven over 58k steps). The done-gated EMAs (slots 585/586)
will adapt LOSS toward ~1.45 AFTER entropy stabilizes.
Together: entropy stays stable (LOSS=3.0 + 5% floor) AND PnL improves
(adaptive LOSS lowers toward real L/W ratio after warmup).
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>