s_max_dd, s_trades, s_wins parallel reductions were made redundant by the exact sequential drawdown scan and upcoming boundary stitching. Reduces shared memory from 25KB to 22KB and removes 3 dead ops/bar from the per-thread loop. Trade count/win rate temporarily zeroed — restored by boundary stitching in next commit.