jgrusewski e8ecb2f626 fix(training_loop): surface FusedTrainingCtx init errors instead of hiding them
Both lazy-init sites (train_with_data_full_loop_slices @ L344 and
run_training_steps_slices @ L1511) previously logged FusedTrainingCtx::new
failures via tracing::error! and continued with fused_ctx = None. On CUDA
builds there is no CPU fallback, so the next stage (GPU experience collector)
then fails ~100ms later with a misleading "GPU experience collector MUST be
active for CUDA training" error that obscures the real CUDA root cause (OOM,
driver error, etc.).

Both sites now propagate the original error via map_err → anyhow::anyhow! so
callers see the actual failure. The post-init wiring (PER buffer pointers,
RNG dev ptr) is refactored from `if let Some(ref mut fused)` — which was
silently-no-op on the hidden-error path — to an unconditional
as_mut().expect() since fused_ctx is guaranteed Some after the `?`.

The two remaining `fused_ctx = None` assignments are legitimate:
  - mod.rs:655 in Drop (ordered GPU teardown)
  - training_loop.rs:1508 batch-size-change recreate (immediately reassigned
    by the ? on the next line)

No new API, no threshold changes, no feature flags.
2026-04-22 08:18:43 +02:00

Foxhunt

Production HFT trading system in Rust.

Architecture

The workspace contains 32 crates organized as follows:

Core Libraries (16)

Crate Purpose
trading_engine Order processing, FIX 4.4, IB TWS, SIMD, RDTSC timing
risk VaR, Kelly, circuit breakers, kill switches, compliance
risk-data Risk data types and shared structures
trading-data Trading data types
ml DQN Rainbow, PPO, TFT, Mamba2, ensemble inference
ml-data ML data types and feature definitions
data Market data ingestion and storage
backtesting Replay engine, strategy tester
adaptive-strategy Ensemble execution, microstructure analysis
common Shared types, resilience, error handling
storage S3 and local model storage
model_loader Model serialization and loading
market-data Market data feed handlers
database PostgreSQL access layer (SQLx)
config Configuration management
tli CLI commands and tooling

Services (8)

Service Purpose
backtesting_service gRPC backtesting service
broker_gateway_service FIX routing, broker connectivity
trading_service Core trading operations
ml_training_service Model training orchestration
data_acquisition_service Market data acquisition
trading_agent_service Autonomous trading agents
api_gateway gRPC API gateway with auth
web-gateway Axum REST + WebSocket gateway

Frontend

web-dashboard/ -- React 19 + TypeScript + Vite + TradingView charts.

Building

# Check compilation (no PostgreSQL required)
SQLX_OFFLINE=true cargo check --workspace

# Run tests for a specific crate
SQLX_OFFLINE=true cargo test -p <crate> --lib

# Clippy
SQLX_OFFLINE=true cargo clippy --workspace

ML Models

Four production model architectures on Candle v0.9.1 with CUDA:

  • DQN Rainbow -- Deep Q-Network with prioritized replay, dueling heads, noisy nets
  • PPO -- Proximal Policy Optimization with GAE and LSTM policies
  • TFT -- Temporal Fusion Transformer for multi-horizon forecasting
  • Mamba2 -- State space model for sequence prediction

Each model has a standalone trainer and a UnifiedTrainable adapter for the hyperopt pipeline.

Infrastructure

  • Git: Gitea at git.fxhnt.ai (Tailscale-only), Scaleway DEV1-S
  • Observability: OpenTelemetry OTLP (env OTEL_EXPORTER_OTLP_ENDPOINT)
  • Database: PostgreSQL with SQLx offline mode for CI

License

Proprietary. All rights reserved.

Description
No description provided
Readme 849 MiB
Languages
Rust 88.2%
Cuda 7.7%
Python 1.3%
Shell 1.1%
PLpgSQL 0.8%
Other 0.8%