jgrusewski ed4b30b493 fix(smoke): controller_activity — V7 intervention-based fire detection
Redefine the "fire" semantic for adaptive controllers per the V7 audit
(policy-quality-design spec §5.3): a controller fires iff it made an
ADAPTIVE INTERVENTION this epoch, not merely because its observable
output value changed.

Previously fire detection was absolute-delta on the output:
- grad_clip fired in 98.3% of epochs because the adaptive clip threshold
  is an EMA recomputed every training step. The EMA drifts by > 1e-3
  every epoch regardless of whether the clip actually clamped a gradient.
- cost_anneal fired in 98.3% of epochs because it is a deterministic
  sigmoid of current_epoch (1/(1+exp(-(epoch-10)/3))) with no adaptive
  or reactive component. Every epoch moves it by > 1e-4 by design.

Neither was "load-bearing" in the V7 sense — one was a per-step EMA
tracker, the other a pure curriculum schedule. The prior test output
"controller 'grad_clip' fires in 98.3%" was a false positive from
measuring the wrong signal.

New semantics:
- anti_lr / tau / gamma / cql_alpha: unchanged — absolute delta vs
  prior epoch on the effective output value (real adaptive controllers).
- grad_clip: intervention-based latch `grad_clip_kicked_this_epoch`,
  set in run_training_steps_slices iff raw_grad_norm > active clip at
  any training step this epoch. Reset in reset_epoch_state.
- cost_anneal: pure deterministic schedule → never load-bearing →
  always fires=false. The value is still tracked in prev_controller_values
  and the HEALTH_DIAG line still emits it for observability, but it
  cannot trip the 50% load-bearing gate.

After fix, controller_activity smoke reports:
  anti_lr=0.000 tau=0.033 gamma=0.017 clip=0.233 cql=0.033 cost=0.000
All 6 rates ≤ 0.5. Test passes.

Touched:
- crates/ml/src/trainers/dqn/trainer/mod.rs (add grad_clip_kicked_this_epoch)
- crates/ml/src/trainers/dqn/trainer/constructor.rs (init new field)
- crates/ml/src/trainers/dqn/trainer/training_loop.rs (latch kick per step,
  redefine fire_clip + fire_cost in HEALTH_DIAG block)

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-22 01:16:45 +02:00

Foxhunt

Production HFT trading system in Rust.

Architecture

The workspace contains 32 crates organized as follows:

Core Libraries (16)

Crate Purpose
trading_engine Order processing, FIX 4.4, IB TWS, SIMD, RDTSC timing
risk VaR, Kelly, circuit breakers, kill switches, compliance
risk-data Risk data types and shared structures
trading-data Trading data types
ml DQN Rainbow, PPO, TFT, Mamba2, ensemble inference
ml-data ML data types and feature definitions
data Market data ingestion and storage
backtesting Replay engine, strategy tester
adaptive-strategy Ensemble execution, microstructure analysis
common Shared types, resilience, error handling
storage S3 and local model storage
model_loader Model serialization and loading
market-data Market data feed handlers
database PostgreSQL access layer (SQLx)
config Configuration management
tli CLI commands and tooling

Services (8)

Service Purpose
backtesting_service gRPC backtesting service
broker_gateway_service FIX routing, broker connectivity
trading_service Core trading operations
ml_training_service Model training orchestration
data_acquisition_service Market data acquisition
trading_agent_service Autonomous trading agents
api_gateway gRPC API gateway with auth
web-gateway Axum REST + WebSocket gateway

Frontend

web-dashboard/ -- React 19 + TypeScript + Vite + TradingView charts.

Building

# Check compilation (no PostgreSQL required)
SQLX_OFFLINE=true cargo check --workspace

# Run tests for a specific crate
SQLX_OFFLINE=true cargo test -p <crate> --lib

# Clippy
SQLX_OFFLINE=true cargo clippy --workspace

ML Models

Four production model architectures on Candle v0.9.1 with CUDA:

  • DQN Rainbow -- Deep Q-Network with prioritized replay, dueling heads, noisy nets
  • PPO -- Proximal Policy Optimization with GAE and LSTM policies
  • TFT -- Temporal Fusion Transformer for multi-horizon forecasting
  • Mamba2 -- State space model for sequence prediction

Each model has a standalone trainer and a UnifiedTrainable adapter for the hyperopt pipeline.

Infrastructure

  • Git: Gitea at git.fxhnt.ai (Tailscale-only), Scaleway DEV1-S
  • Observability: OpenTelemetry OTLP (env OTEL_EXPORTER_OTLP_ENDPOINT)
  • Database: PostgreSQL with SQLx offline mode for CI

License

Proprietary. All rights reserved.

Description
No description provided
Readme 849 MiB
Languages
Rust 88.2%
Cuda 7.7%
Python 1.3%
Shell 1.1%
PLpgSQL 0.8%
Other 0.8%