jgrusewski 5c325f8947 spec(dqn): pivot to Thompson sampling — distributional RL action selection
Revises the C51-bias spec after deeper review surfaced 12 design gaps,
of which 4 were critical:

  1. Train-only vs train+eval ambiguity — UCB at eval would conflate
     "model recommends Long" with "model is uncertain about Long",
     inflating reported edge. CRITICAL for trading where eval drives
     real capital decisions.
  2. Thompson sampling is more principled than UCB:
       - parameter-free (no κ to tune)
       - uses distribution directly without scalar reduction
       - naturally explore-exploit balanced via distribution shape
  3. c51_alpha is the wrong blend weight (it's C51-vs-MSE-warmup, not
     C51-vs-IQN). Equal-weight average of C51 and IQN samples is the
     structural choice — no tuned blend weight needed.
  4. The bias might be CORRECT BEHAVIOUR — model rationally choosing
     Flat when no edge has been discovered. Phase 0 must include a
     synthetic-edge test (controlled MDP with KNOWN positive Long
     expected value) to verify Thompson can discover edge if it exists.

Other gaps fixed:
  - Eval at argmax E[Q] (not Boltzmann, not Thompson)
  - Pearl wording broadened to cover ensembles + future methods
  - Ensemble Q-head added to aggregation contract table
  - Explicit caveat: NEVER extend Thompson to magnitude branch (would
    worsen existing magnitude saturation)
  - Phase 0.F uses CONVERGED checkpoint (≥30 epochs), not 2-epoch run
  - L4 long smoke (30 epochs, ~1 hour) added — Thompson edge discovery
    needs longer feedback loop than 5 epochs
  - Phase 3 explicitly removes eps-floor + tau-floor band-aids
    (Thompson replaces direction Boltzmann; band-aids become dead code)
  - Conviction stays E[Q]-based, not sample-based (avoid Kelly cap
    jitter from stochastic samples)

Architecture (Thompson only, no UCB):
  TRAINING: dir_idx = argmax(0.5 × (sample_C51(d) + sample_IQN(d)))
            magnitude/order/urgency: existing Boltzmann + ε-greedy
  EVAL:     dir_idx = argmax(0.5 × (E[Q_C51] + E[Q_IQN]))
            magnitude/order/urgency: existing Boltzmann (eval mode)

Direction-branch ε-greedy + Boltzmann are REMOVED — Thompson is the
exploration mechanism. No new GPU buffers; existing C51 atoms + IQN
quantiles passed to action_select.

5-7 days active work across 4 sub-plans; each gets its own
writing-plans cycle.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-26 23:35:40 +02:00

Foxhunt

Production HFT trading system in Rust.

Architecture

The workspace contains 32 crates organized as follows:

Core Libraries (16)

Crate Purpose
trading_engine Order processing, FIX 4.4, IB TWS, SIMD, RDTSC timing
risk VaR, Kelly, circuit breakers, kill switches, compliance
risk-data Risk data types and shared structures
trading-data Trading data types
ml DQN Rainbow, PPO, TFT, Mamba2, ensemble inference
ml-data ML data types and feature definitions
data Market data ingestion and storage
backtesting Replay engine, strategy tester
adaptive-strategy Ensemble execution, microstructure analysis
common Shared types, resilience, error handling
storage S3 and local model storage
model_loader Model serialization and loading
market-data Market data feed handlers
database PostgreSQL access layer (SQLx)
config Configuration management
tli CLI commands and tooling

Services (8)

Service Purpose
backtesting_service gRPC backtesting service
broker_gateway_service FIX routing, broker connectivity
trading_service Core trading operations
ml_training_service Model training orchestration
data_acquisition_service Market data acquisition
trading_agent_service Autonomous trading agents
api_gateway gRPC API gateway with auth
web-gateway Axum REST + WebSocket gateway

Frontend

web-dashboard/ -- React 19 + TypeScript + Vite + TradingView charts.

Building

# Check compilation (no PostgreSQL required)
SQLX_OFFLINE=true cargo check --workspace

# Run tests for a specific crate
SQLX_OFFLINE=true cargo test -p <crate> --lib

# Clippy
SQLX_OFFLINE=true cargo clippy --workspace

ML Models

Four production model architectures on Candle v0.9.1 with CUDA:

  • DQN Rainbow -- Deep Q-Network with prioritized replay, dueling heads, noisy nets
  • PPO -- Proximal Policy Optimization with GAE and LSTM policies
  • TFT -- Temporal Fusion Transformer for multi-horizon forecasting
  • Mamba2 -- State space model for sequence prediction

Each model has a standalone trainer and a UnifiedTrainable adapter for the hyperopt pipeline.

Infrastructure

  • Git: Gitea at git.fxhnt.ai (Tailscale-only), Scaleway DEV1-S
  • Observability: OpenTelemetry OTLP (env OTEL_EXPORTER_OTLP_ENDPOINT)
  • Database: PostgreSQL with SQLx offline mode for CI

License

Proprietary. All rights reserved.

Description
No description provided
Readme 849 MiB
Languages
Rust 88.2%
Cuda 7.7%
Python 1.3%
Shell 1.1%
PLpgSQL 0.8%
Other 0.8%