jgrusewski 6d0ac7beb3 fix(sp4): GPU-port calibrate_homeostatic_targets per feedback_no_cpu_compute_strict
Per-step host-side EMA loop at gpu_dqn_trainer.rs:4671 over 6 mapped-
pinned homeostatic-target slots was a feedback_no_cpu_compute_strict
violation discovered during the sweep audit (commit 6a6b58aec) but
deferred for scope. Sweep audit grid site #9.

Migrated:
  - New calibrate_homeostatic_kernel.cu — single-block, six threads
    (one thread per homeostatic slot). Reads host-passed `readiness`
    scalar (already-clamped IQN gauge from `iqn_readiness` shadow field,
    bit-for-bit match of the deleted host clamp), observations from
    `homeostatic_obs_dev_ptr`, applies adaptive α
    `0.3 × (1 - readiness) + 0.01 × readiness` to targets[k] for k=1..5;
    thread 0 forces the Q-mean invariant `targets[0] = 0.0` exactly as
    the deleted host post-loop assignment did. __threadfence_system()
    after writes for PCIe-visibility to homeostatic_kernel's dev_ptr reads.
  - build.rs cubin registration and trainer-struct wiring (cubin static,
    field, struct constructor, cubin load) mirror the C2/C3/C4 pattern
    from the prior sweep commits.
  - Host-side `for k in 0..HOMEOSTATIC_N_OBS` loop in
    `calibrate_homeostatic_targets` replaced with a single
    `launch_calibrate_homeostatic` kernel launch; chained on the
    trainer's stream so it remains graph-capture-compatible.

Preserved:
  - Same α formula, same Q-mean=0 invariant, same call ordering.
  - Mapped-pinned target buffer retained — homeostatic_kernel still
    reads via homeostatic_targets_dev_ptr unchanged.
  - No cold-start sentinel: constructor pre-initialises targets to
    `[0.0, 0.85, 0.1, 0.0, 1.0, 0.5]` (gpu_dqn_trainer.rs:14583-14588)
    so the first EMA call blends defaults with the first observation,
    same algebraic shape the deleted host loop relied on.
  - State-reset registry unchanged — deleted host loop had no fold reset
    (per-call EMA only); GPU port preserves identical per-call semantics.

cargo check clean. SP4 + state_reset_registry lib tests pass (11/11).
16/16 SP4 producer GPU tests pass on RTX 3050 Ti. No behavior change —
pure architectural fix.

Refs: feedback_no_cpu_compute_strict sweep audit grid site #9.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-01 16:14:22 +02:00

Foxhunt

Production HFT trading system in Rust.

Architecture

The workspace contains 32 crates organized as follows:

Core Libraries (16)

Crate Purpose
trading_engine Order processing, FIX 4.4, IB TWS, SIMD, RDTSC timing
risk VaR, Kelly, circuit breakers, kill switches, compliance
risk-data Risk data types and shared structures
trading-data Trading data types
ml DQN Rainbow, PPO, TFT, Mamba2, ensemble inference
ml-data ML data types and feature definitions
data Market data ingestion and storage
backtesting Replay engine, strategy tester
adaptive-strategy Ensemble execution, microstructure analysis
common Shared types, resilience, error handling
storage S3 and local model storage
model_loader Model serialization and loading
market-data Market data feed handlers
database PostgreSQL access layer (SQLx)
config Configuration management
tli CLI commands and tooling

Services (8)

Service Purpose
backtesting_service gRPC backtesting service
broker_gateway_service FIX routing, broker connectivity
trading_service Core trading operations
ml_training_service Model training orchestration
data_acquisition_service Market data acquisition
trading_agent_service Autonomous trading agents
api_gateway gRPC API gateway with auth
web-gateway Axum REST + WebSocket gateway

Frontend

web-dashboard/ -- React 19 + TypeScript + Vite + TradingView charts.

Building

# Check compilation (no PostgreSQL required)
SQLX_OFFLINE=true cargo check --workspace

# Run tests for a specific crate
SQLX_OFFLINE=true cargo test -p <crate> --lib

# Clippy
SQLX_OFFLINE=true cargo clippy --workspace

ML Models

Four production model architectures on Candle v0.9.1 with CUDA:

  • DQN Rainbow -- Deep Q-Network with prioritized replay, dueling heads, noisy nets
  • PPO -- Proximal Policy Optimization with GAE and LSTM policies
  • TFT -- Temporal Fusion Transformer for multi-horizon forecasting
  • Mamba2 -- State space model for sequence prediction

Each model has a standalone trainer and a UnifiedTrainable adapter for the hyperopt pipeline.

Infrastructure

  • Git: Gitea at git.fxhnt.ai (Tailscale-only), Scaleway DEV1-S
  • Observability: OpenTelemetry OTLP (env OTEL_EXPORTER_OTLP_ENDPOINT)
  • Database: PostgreSQL with SQLx offline mode for CI

License

Proprietary. All rights reserved.

Description
No description provided
Readme 849 MiB
Languages
Rust 88.2%
Cuda 7.7%
Python 1.3%
Shell 1.1%
PLpgSQL 0.8%
Other 0.8%