jgrusewski 10d4614fb4 fix(cuda): remove atomicAdd violations in ppo/dqn loss accumulators
Three atomicAdd sites violated feedback_no_atomicadd (deferred Phase C/E
fixes per file headers):
- ppo_clipped_surrogate.cu lines 270-271: atomicAdd(loss_pi/loss_entropy)
- dqn_distributional_q.cu line 285: atomicAdd(loss_out, ce/B)

Float addition is non-associative; atomicAdd across blocks produces
order-dependent sums. When inputs were extreme (l_v=6+ early in training
driving wide PPO ratios), the order-dependence flipped finite→NaN
non-deterministically — same SHA + same seed + same params produced
different NaN outcomes (smoke bisect 2026-05-29).

Fix is structural per the codebase pattern (compute_advantage_rms,
rl_q_bias_correction): per-batch sole-writer outputs + a deterministic
single-warp block-tree reducer. New kernel `mean_b_reduce.cu` performs
the [B]→[1] mean via grid-stride loop + warp-shuffle reduce.

Changes:
- cuda/mean_b_reduce.cu (NEW): single-block single-warp deterministic
  mean reducer mirroring compute_advantage_rms.cu pattern
- cuda/dqn_distributional_q.cu: remove atomicAdd; loss_per_batch[]
  already sole-writer, caller invokes mean_b_reduce after
- cuda/ppo_clipped_surrogate.cu: replace `loss_pi`/`loss_entropy` scalar
  args with `l_pi_per_batch`/`l_ent_per_batch` per-batch sole-writer
  outputs; caller invokes mean_b_reduce twice after kernel
- build.rs: register mean_b_reduce cubin
- src/rl/dqn.rs: DqnHead loads mean_b_reduce_fn; backward_logits invokes
  it on loss_per_batch → loss_out_dev_ptr
- src/rl/ppo.rs: PolicyHead loads mean_b_reduce_fn; surrogate_forward
  signature gains l_pi_per_batch + l_ent_per_batch buffer args; invokes
  mean_b_reduce on each → scalar dev ptrs
- src/trainer/integrated.rs: IntegratedTrainer gains
  ss_pi_l_pi_per_batch_d + ss_pi_l_ent_per_batch_d scratch fields;
  allocated at b_size; passed to surrogate_forward

Smoke status: l_q/l_v values now reproducible bit-for-bit across runs
(verified by comparing two runs of the same seed). NaN at step 4
still reproduces — proximate cause is elsewhere (likely in V head
forward dynamics, not the loss summation). The atomicAdd removal is
still required: it was a real correctness violation independent of the
NaN symptom, and the deterministic loss values are now a precondition
for diagnosing the remaining instability.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-29 01:25:26 +02:00

Foxhunt

Production HFT trading system in Rust.

Architecture

The workspace contains 32 crates organized as follows:

Core Libraries (16)

Crate Purpose
trading_engine Order processing, FIX 4.4, IB TWS, SIMD, RDTSC timing
risk VaR, Kelly, circuit breakers, kill switches, compliance
risk-data Risk data types and shared structures
trading-data Trading data types
ml DQN Rainbow, PPO, TFT, Mamba2, ensemble inference
ml-data ML data types and feature definitions
data Market data ingestion and storage
backtesting Replay engine, strategy tester
adaptive-strategy Ensemble execution, microstructure analysis
common Shared types, resilience, error handling
storage S3 and local model storage
model_loader Model serialization and loading
market-data Market data feed handlers
database PostgreSQL access layer (SQLx)
config Configuration management
tli CLI commands and tooling

Services (8)

Service Purpose
backtesting_service gRPC backtesting service
broker_gateway_service FIX routing, broker connectivity
trading_service Core trading operations
ml_training_service Model training orchestration
data_acquisition_service Market data acquisition
trading_agent_service Autonomous trading agents
api_gateway gRPC API gateway with auth
web-gateway Axum REST + WebSocket gateway

Frontend

web-dashboard/ -- React 19 + TypeScript + Vite + TradingView charts.

Building

# Check compilation (no PostgreSQL required)
SQLX_OFFLINE=true cargo check --workspace

# Run tests for a specific crate
SQLX_OFFLINE=true cargo test -p <crate> --lib

# Clippy
SQLX_OFFLINE=true cargo clippy --workspace

ML Models

Four production model architectures on Candle v0.9.1 with CUDA:

  • DQN Rainbow -- Deep Q-Network with prioritized replay, dueling heads, noisy nets
  • PPO -- Proximal Policy Optimization with GAE and LSTM policies
  • TFT -- Temporal Fusion Transformer for multi-horizon forecasting
  • Mamba2 -- State space model for sequence prediction

Each model has a standalone trainer and a UnifiedTrainable adapter for the hyperopt pipeline.

Infrastructure

  • Git: Gitea at git.fxhnt.ai (Tailscale-only), Scaleway DEV1-S
  • Observability: OpenTelemetry OTLP (env OTEL_EXPORTER_OTLP_ENDPOINT)
  • Database: PostgreSQL with SQLx offline mode for CI

License

Proprietary. All rights reserved.

Description
No description provided
Readme 849 MiB
Languages
Rust 88.2%
Cuda 7.7%
Python 1.3%
Shell 1.1%
PLpgSQL 0.8%
Other 0.8%