Add ml/examples/train_baseline.rs that trains DQN and PPO models using expanding walk-forward windows on real Databento OHLCV data. Features: - CLI args via clap (--model, --epochs, --batch-size, --data-dir, etc.) - Recursive .dbn.zst file discovery and OHLCV bar loading - 51-dim feature extraction via extract_ml_features() - Walk-forward window generation with NormStats per fold - DQN training loop with epsilon-greedy, experience replay, early stopping - PPO training loop with GAE, trajectory collection, early stopping - PnL-based reward (BUY/SELL/HOLD) - Safetensors checkpoint saving per fold - NormStats JSON export for evaluation reproducibility Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
28 KiB
28 KiB