A) TD-propagation diagnostic smoke test (sparse rewards)
- New smoke_tests/td_propagation.rs captures training_sharpe_ema across
20 epochs with micro_reward_scale=0 to answer: "can sparse-reward
Q-learning extract policy from trade-completion P&L alone?"
- Asserts: finite sharpe_ema, no monotonic degradation (tolerance 0.1
over first-3 vs last-3 epoch avg), q_gap > 0.05
- First run answered YES: val Sharpe peaks at +12.49 at epoch 8 with
pure sparse rewards, confirming the objective is learnable. The
remaining issue is stability/overfitting, not TD propagation.
B) Entropy → plasticity gate in training_loop
- D3/N3 shrink_perturb trigger now fires on `last_action_entropy < 0.3`
OR `health_value < 0.3` (OR semantics, 3 consecutive epochs).
- Catches action-collapse cases the generic health metric misses: Q-values
separate cleanly but argmax stays pinned to a single branch.
- tracing::info now logs both signals.
C) commitment_lambda coverage verified
- Almgren-Chriss sqrt market-impact already in compute_tx_cost
(trade_physics.cuh:168-171). commitment_lambda was pure duplication.
- Inventory doc updated: P2 marked satisfied, table row status updated.
Files touched:
- crates/ml/src/trainers/dqn/smoke_tests/mod.rs (+2)
- crates/ml/src/trainers/dqn/smoke_tests/td_propagation.rs (new)
- crates/ml/src/trainers/dqn/trainer/training_loop.rs (+15/-3)
- docs/superpowers/specs/2026-04-21-phase1-reward-inventory.md (+6/-3)
Smoke test PASSED locally (RTX 3050 Ti, 30.99s end-to-end).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>