jgrusewski ce1e13519b fix(rl): mapped-pinned for all R7d/R8 CPU↔GPU paths + diag JSONL + guard
Two concerns in one commit since they're entangled:

## 1. feedback_no_htod_htoh_only_mapped_pinned violations

R7d (PER push/sample) + R8 (CLI binary) + the new per-step diag dump
shipped with 8 raw `stream.memcpy_htod` / `stream.memcpy_dtoh` calls.
The rule is explicit: "mapped-pinned only for CPU↔GPU; tests not
exempt." A raw `stream.memcpy_*` on a regular `&[T]` / `&mut [T]` is
NOT mapped-pinned — the source/dest slice isn't page-locked, so the
CUDA driver does an internal blocking HtoD/DtoH that stalls the
stream.

Refactored all 8 violations to use the mapped-pinned + DtoD pattern
(cuMemHostAlloc DEVICEMAP — host writes via `host_ptr`, kernel reads
`dev_ptr`, DtoD between them via `cudarc::driver::result::memcpy_dtod_async`).
New shared helpers in `trainer/integrated.rs`:

  * `read_slice_i32_d` — DtoH for `i32` device buffers via
    `MappedI32Buffer` staging. Counterpart to the existing
    `read_slice_d` (f32 version).
  * `write_slice_f32_d` — CPU→GPU upload for `f32` via
    `MappedF32Buffer.write_from_slice` + DtoD into destination.
  * `write_slice_i32_d` — CPU→GPU upload for `i32` via
    `MappedI32Buffer.host_slice_mut().copy_from_slice` + DtoD.

`pub fn` wrappers (`read_slice_*_d_pub`) expose the f32/i32 helpers
to the CLI binary so the per-step diag DtoH uses the same canonical
pattern.

Call-site refactors:

  * `push_to_replay`: 4× `stream.memcpy_dtoh` → `read_slice_*_d`.
  * `sample_and_gather`: 3× `stream.memcpy_htod` → `write_slice_*_d`.
  * `step_with_lobsim` pre-PER ISV refresh: raw `memcpy_dtoh` →
    `read_slice_d` (424 floats per step).
  * `step_with_lobsim` post-Q PER priority TD readback: raw
    `memcpy_dtoh` → `read_slice_d` (b_size floats per step).
  * `step_synthetic` ISV mirror refresh: raw `memcpy_dtoh` →
    `read_slice_d` (pre-existing pre-R9 violation; fixed in the
    same commit since it's the same pattern in the same file).
  * Init-time (one-shot) `prng_state` upload: raw `memcpy_htod` →
    inline mapped-pinned DtoD (custom because cast through i32 for
    the u32 buffer).
  * Init-time (one-shot) `atom_supports` upload: raw `memcpy_htod`
    → `write_slice_f32_d`.
  * `examples/alpha_rl_train.rs` per-step diag DtoH (3 calls) →
    `read_slice_*_d_pub`.

## 2. Pre-commit guard gap — diff-aware HtoD/DtoH check

The existing GPU hot-path guard (`scripts/gpu-hotpath-guard.sh`)
EXPLICITLY skips memcpy_htod/dtoh on the assumption that such calls
only appear in `cuda_pipeline/` (where mapped-pinned is the
convention). That assumption was falsified by R7d/R8 — the guard
shipped 8 violations green.

Added `check_no_raw_htod_dtoh` to `scripts/pre-commit-hook.sh` (the
real file behind the `.git/hooks/pre-commit` symlink). The check is
DIFF-AWARE: it greps only the `+` lines of `git diff --cached -U0`,
so pre-existing violations elsewhere (143 sites across the codebase)
don't block commits touching unrelated files. NEW additions of
`\.memcpy_(htod|dtoh)\(` are flagged with a clear error pointing at
the mapped-pinned alternative. Suppress per-line with `// gpu-ok:
<reason>` (same convention as the existing guards).

Pre-existing violations in `ml-alpha/src/aux_heads.rs`,
`mamba2_block.rs`, `cfc/`, `data/`, etc. are a separate cleanup —
not blocked by this commit's check because the diff-aware filter
ignores anything that was already on `HEAD~1`.

## Verified gates (post-fix, local sm_86)

  G1  isv_bootstrap                 unchanged
  G3  controllers_emit              unchanged
  G4  target_soft_update            unchanged
  G6  r7d_per_wiring                unchanged (PER round-trips all
                                     mapped-pinned now)
  R3  ema/advantage (3 tests)       unchanged
  R4  action kernels (3 tests)      unchanged
  end integrated_trainer_smoke      unchanged

Mapped-pinned is semantically equivalent to raw memcpy_htod/dtoh —
just routed through page-locked staging so the driver doesn't have
to do its own internal pinning. Behaviour identical; the cost shifts
from "driver hidden HtoD per call" to "mapped-pinned alloc + DtoD
per call." For the smoke (b_size=1, 1000 steps), the cost difference
is in the microseconds.

## Per-step diag JSONL (separate concern, same commit)

Added `--diag-jsonl <PATH>` flag to `alpha_rl_train.rs` (default:
`<out>/diag.jsonl`). After each `step_with_lobsim`, writes one JSON
record capturing:

  * step number, elapsed wall time
  * all 5 head losses + λs
  * all 7 RL controller outputs (γ τ ε coef n_roll per_α scale)
  * all 5 per-head learning rates (lr_bce/q/pi/v/aux)
  * all 7 EMA inputs the controllers consume
  * replay buffer length
  * per-step reward stats (sum, max, min, abs_max)
  * per-step done count
  * per-step action histogram (9 action classes)

Critical for cluster smoke debugging — the prior CLI only flushed
an `eprintln` progress line every N steps (default 100), making
in-flight controller drift / replay stagnation / reward explosion
invisible until they produced a NaN abort. The JSONL is line-
buffered + flushed every `log_every` steps so `tail -f` shows
progress live.

The stderr progress line is also beefed up to include γ / ε / per_α /
reward_scale / dones / rew_sum at each tick so a casual `argo logs`
inspection sees the controller behaviour without parsing JSONL.

## Why R9 cluster submission needs this

Without the diag dump, an R9 1000-step smoke is "blind" — only the
final summary tells us what happened. With the dump, post-hoc
analysis can answer:
  * Did the controllers adapt or stay at bootstrap?
  * Did the reward scale stabilise or saturate?
  * Did the PER buffer fill?
  * Was the action histogram dominated by any one action?
  * Where did the per-head losses converge to?

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-23 14:14:55 +02:00

Foxhunt

Production HFT trading system in Rust.

Architecture

The workspace contains 32 crates organized as follows:

Core Libraries (16)

Crate Purpose
trading_engine Order processing, FIX 4.4, IB TWS, SIMD, RDTSC timing
risk VaR, Kelly, circuit breakers, kill switches, compliance
risk-data Risk data types and shared structures
trading-data Trading data types
ml DQN Rainbow, PPO, TFT, Mamba2, ensemble inference
ml-data ML data types and feature definitions
data Market data ingestion and storage
backtesting Replay engine, strategy tester
adaptive-strategy Ensemble execution, microstructure analysis
common Shared types, resilience, error handling
storage S3 and local model storage
model_loader Model serialization and loading
market-data Market data feed handlers
database PostgreSQL access layer (SQLx)
config Configuration management
tli CLI commands and tooling

Services (8)

Service Purpose
backtesting_service gRPC backtesting service
broker_gateway_service FIX routing, broker connectivity
trading_service Core trading operations
ml_training_service Model training orchestration
data_acquisition_service Market data acquisition
trading_agent_service Autonomous trading agents
api_gateway gRPC API gateway with auth
web-gateway Axum REST + WebSocket gateway

Frontend

web-dashboard/ -- React 19 + TypeScript + Vite + TradingView charts.

Building

# Check compilation (no PostgreSQL required)
SQLX_OFFLINE=true cargo check --workspace

# Run tests for a specific crate
SQLX_OFFLINE=true cargo test -p <crate> --lib

# Clippy
SQLX_OFFLINE=true cargo clippy --workspace

ML Models

Four production model architectures on Candle v0.9.1 with CUDA:

  • DQN Rainbow -- Deep Q-Network with prioritized replay, dueling heads, noisy nets
  • PPO -- Proximal Policy Optimization with GAE and LSTM policies
  • TFT -- Temporal Fusion Transformer for multi-horizon forecasting
  • Mamba2 -- State space model for sequence prediction

Each model has a standalone trainer and a UnifiedTrainable adapter for the hyperopt pipeline.

Infrastructure

  • Git: Gitea at git.fxhnt.ai (Tailscale-only), Scaleway DEV1-S
  • Observability: OpenTelemetry OTLP (env OTEL_EXPORTER_OTLP_ENDPOINT)
  • Database: PostgreSQL with SQLx offline mode for CI

License

Proprietary. All rights reserved.

Description
No description provided
Readme 849 MiB
Languages
Rust 88.2%
Cuda 7.7%
Python 1.3%
Shell 1.1%
PLpgSQL 0.8%
Other 0.8%