b21c6d5cff35d0c3be2ad0100b884a663e3a203e
Major architectural fixes discovered through systematic investigation: Q-gap measurement (3 bugs): - compute_expected_q never ran during training → q_out_buf was zeros - epoch_q_gap reset in process_epoch_boundary before logging read it - flush_q_stats_readback drained async readback before in-loop read Zero-copy pinned memory (4 hot-path scalars): - t_buf, tau_buf, v_range_buf, adaptive_clip_buf → pinned device-mapped - GPU reads via cuMemHostGetDevicePointer, host writes directly, no HtoD Q-stats-driven adaptive v_range: - v_range = q_mean ± 3σ + Bellman headroom (was fixed ±1.0) - Adaptive MIN_RANGE scales with |Q_mean| (was fixed 0.02) - Per-step adaptation (was every 50 steps) - 100× finer atom resolution from epoch 1 Gradient stability: - IS-weight clamp at 10.0 in all loss/grad kernels (PER spike prevention) - 3 power iterations in spectral norm (was 1 — underestimated sigma) - Bottleneck w_bn added as 13th spectral-normed matrix (was missing) - Pre-Adam grad_buf clip via clip_grad_buf_inplace (activation amplification) - EMA-based adaptive gradient clipping (pinned device buffer) - Consolidated grad_norm to single buffer (was 2 — eliminated grad_norm_f32_buf) Adaptive tau from online-target Q-divergence: - C51 loss kernel accumulates (E[Q_online] - E[Q_target])² per batch - Tau scales with sqrt(divergence/baseline), clamped [0.5×, 10×] base - Accelerates target convergence during discovery, stabilizes during plateau Deterministic evaluation: - eval_mode in action_select kernel: pure greedy argmax, no Boltzmann/RNG - Eliminated ±40 val_Sharpe noise from near-uniform Boltzmann sampling - Backtest evaluator uses adaptive v_range (was config v_min/v_max — 1500× mismatch) Atom utilization metrics: - compute_expected_q accumulates entropy + utilization per step - q_stats_kernel extended to 7 outputs (was 5) - Logged per epoch: atoms=98%ent/92%util Pessimistic Q-init removed — incompatible with adaptive v_range (bias was 255× outside ±0.01 support, causing 5-epoch cold-start and late Q-value drift). 903/903 tests passing. val_Sharpe positive from epoch 1 with greedy eval. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Foxhunt
Production HFT trading system in Rust.
Architecture
The workspace contains 32 crates organized as follows:
Core Libraries (16)
| Crate | Purpose |
|---|---|
trading_engine |
Order processing, FIX 4.4, IB TWS, SIMD, RDTSC timing |
risk |
VaR, Kelly, circuit breakers, kill switches, compliance |
risk-data |
Risk data types and shared structures |
trading-data |
Trading data types |
ml |
DQN Rainbow, PPO, TFT, Mamba2, ensemble inference |
ml-data |
ML data types and feature definitions |
data |
Market data ingestion and storage |
backtesting |
Replay engine, strategy tester |
adaptive-strategy |
Ensemble execution, microstructure analysis |
common |
Shared types, resilience, error handling |
storage |
S3 and local model storage |
model_loader |
Model serialization and loading |
market-data |
Market data feed handlers |
database |
PostgreSQL access layer (SQLx) |
config |
Configuration management |
tli |
CLI commands and tooling |
Services (8)
| Service | Purpose |
|---|---|
backtesting_service |
gRPC backtesting service |
broker_gateway_service |
FIX routing, broker connectivity |
trading_service |
Core trading operations |
ml_training_service |
Model training orchestration |
data_acquisition_service |
Market data acquisition |
trading_agent_service |
Autonomous trading agents |
api_gateway |
gRPC API gateway with auth |
web-gateway |
Axum REST + WebSocket gateway |
Frontend
web-dashboard/ -- React 19 + TypeScript + Vite + TradingView charts.
Building
# Check compilation (no PostgreSQL required)
SQLX_OFFLINE=true cargo check --workspace
# Run tests for a specific crate
SQLX_OFFLINE=true cargo test -p <crate> --lib
# Clippy
SQLX_OFFLINE=true cargo clippy --workspace
ML Models
Four production model architectures on Candle v0.9.1 with CUDA:
- DQN Rainbow -- Deep Q-Network with prioritized replay, dueling heads, noisy nets
- PPO -- Proximal Policy Optimization with GAE and LSTM policies
- TFT -- Temporal Fusion Transformer for multi-horizon forecasting
- Mamba2 -- State space model for sequence prediction
Each model has a standalone trainer and a UnifiedTrainable adapter for the hyperopt pipeline.
Infrastructure
- Git: Gitea at
git.fxhnt.ai(Tailscale-only), Scaleway DEV1-S - Observability: OpenTelemetry OTLP (env
OTEL_EXPORTER_OTLP_ENDPOINT) - Database: PostgreSQL with SQLx offline mode for CI
License
Proprietary. All rights reserved.
Description
Languages
Rust
88.2%
Cuda
7.7%
Python
1.3%
Shell
1.1%
PLpgSQL
0.8%
Other
0.8%