Files
foxhunt/docs/superpowers/specs/2026-03-23-profitability-roadmap.md
jgrusewski 33d91cce5b spec: profitability roadmap — 11 findings from deep research
4 critical root causes fixed this session (stop-loss, dense shaping,
action aliasing, HFT objective weight). 7 remaining improvements
documented with priority order: three-phase hyperopt, behavioral
cloning warm start, curriculum learning, ensemble, multi-timeframe,
attention, Decision Transformer.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-23 12:45:06 +01:00

1.7 KiB
Raw Blame History

Profitability Roadmap — From 0.12% Win Rate to Production Trading

Implemented (this session)

  1. Stop-loss 0.3%→1%, TP 0.5%→2% — noise was killing 99.88% of trades
  2. Dense shaping 0.1x→0.01x — trade completion now dominates
  3. Action aliasing fix — 5-bar hold override corrupted Q-values
  4. HFT activity weight 25%→5% — hyperopt now rewards selectivity

Next priorities

P5: Three-phase hyperopt (5D→8D→5D)

  • Phase 1 (5D): lr, gamma, batch_size, w_dsr, w_pnl
  • Phase 2 (8D): kelly, dd_threshold, loss_aversion, sl/tp levels, time_decay, entropy, cql_alpha
  • Phase 3 (5D): hidden_dim, num_atoms, branch_hidden, n_steps, tau
  • PSO 100 iterations × 30 particles per phase = 3000 evaluations

P6: Behavioral cloning warm start

  • Generate expert demonstrations from MA crossover (20/50 EMA, ADX>25 filter)
  • Pre-fill replay buffer: 50% expert, 50% online for first 10 epochs
  • CQL pre-training with cql_alpha=1.0 for 10 epochs, then 0.1 online

P7: Curriculum learning

  • Phase 1: only windows where ADX>30 (strong trends, easy)
  • Phase 2: all windows, trending weighted 2x
  • Phase 3: equal weight (full generalization)

P8: Ensemble of 3-5 agents

  • Different random seeds, average Q-values
  • KL-divergence diversity loss (weight 0.01)
  • Infrastructure exists in ensemble_network.rs

P9: Multi-timeframe features

  • 5-min and 15-min: close_return, volume_ratio, spread_ratio, RSI
  • +8 features → state_dim stays 64 (reuse padding)

P10: Temporal attention

  • Self-attention over last N state embeddings
  • use_attention: bool already in DQN config

P11: Decision Transformer pre-training

  • Return-conditioned sequence model on historical data
  • LoRA fine-tuning for online adaptation