Two same-seed runs now produce bit-equal eval_summary.json, alpha_rl_train_summary.json,
and diag.jsonl (modulo wall-clock elapsed_s). The 5-phase falsification chain landed:
Phase 2 PER tree-rebuild: __threadfence is NOT a grid-wide barrier; multiple blocks
raced across sum-tree levels. Fix: Grid=(1) Block=(1024) + __syncthreads
in rl_per_tree_rebuild.cu.
Phase 2.3 cuBLAS GEMM_DFALT + TF32 default-math allowed split-K non-deterministic
accumulation at 3 sites. New crates/ml-alpha/src/cublas_determinism.rs
applies CUBLAS_PEDANTIC_MATH via FOXHUNT_DETERMINISTIC env toggle
(0=TF32 prod, 1=PEDANTIC dev default, 2=DEFAULT_MATH control).
Phase 2.6 Two bugs surfaced sequentially in the backward kernel chain:
(1) rl_iqn_tau_cos_features had a multi-block r/w race on prng_state[batch]
— all N_TAU=32 blocks read seed; only tau_idx==0 wrote back; no
inter-block barrier. Fix: split into READ-ONLY rl_iqn_tau_cos_features
+ new sibling rl_iqn_advance_prng_state launched on same stream
(kernel-launch ordering = grid-wide barrier).
(2) OutcomeHead::new called near_zero_xavier without scoped_init_seed,
falling back to time+thread-id RNG. Stayed dormant until first done
event activated non-sentinel labels and divergent weights flowed via
grad_h_t_outcome into encoder gradient. Fix: add seed param + install
scoped_init_seed(dqn_seed.wrapping_add(0x0CE0)) guard.
Validation (./scripts/determinism-check.sh --quick, RTX 3050, b=128, 200+50 steps):
- All 200 rows of checksums.* leaves match (rel-tol 1e-5, abs-tol 1e-7)
- eval_summary.json, alpha_rl_train_summary.json byte-equal between runs
- diag.jsonl byte-equal modulo elapsed_s
- Eval pnl identical run-A vs run-B at seed 42
Pre-fix baseline (Phase 2.5 measurement): same-seed eval pnl spread $450k
($187k vs -$261k). Post-fix: $0 spread.
Speed cost: ~1.5ms/step amortised; ~10-15% slower than TF32 production
(PEDANTIC tax — acceptable in dev, toggle to FOXHUNT_DETERMINISTIC=0 for prod).
Mapped-pinned discipline: all 11 NEW memcpy_dtoh sites in diagnostic dump methods
+ per-step checksum readback use a new pub(crate) helper
read_slice_d_into<T: Copy>(stream, src, dst) — MappedRecordBuffer + raw
memcpy_dtod_async + raw_stream_sync + volatile read. Generic over T (f32, f64,
i32, u32, u8). Satisfies feedback_no_htod_htoh_only_mapped_pinned + hook guard.
Bundled Tier 1.5 fast-dev-cycle infrastructure (spec
docs/superpowers/specs/2026-06-02-fast-dev-cycle.md):
- scripts/local-mid-smoke.sh b=128, 2000+500, ~10min on RTX 3050
- scripts/determinism-check.sh runs mid-smoke twice, diffs checksums
- scripts/tier1_5_verdict.py behavioral kill verdict
- AdamW checkpoint save/load (crates/ml-alpha/src/trainer/optim.rs)
- IntegratedTrainer checkpoint save/load (resume from checkpoint)
- 15 Phase 1 checksum leaves in build_diag_value
- Env-gated dump methods (FOXHUNT_DETERMINISM_DEBUG_PER/MAMBA2/RL/BACKWARD)
for future divergence-chasing — never run in production
Documentation:
- docs/superpowers/specs/2026-06-02-determinism-foundation.md
- docs/superpowers/specs/2026-06-02-fast-dev-cycle.md
- docs/superpowers/plans/2026-06-02-determinism-foundation-implementation.md
- docs/superpowers/notes/2026-06-02-determinism-phase{1,2,2.2,2.5,2.6}-*.md
- Adjacent specs/plans/notes from the analytical chain that surfaced determinism
as the load-bearing blocker (eval-summary, eval-boundary, regime-observer,
multi-head policy, regime-invariance, Phase 3 IQN-complement post-mortem)
Unlocks: every controller / architecture / reward-shaping A/B from this commit
onward attributes outcome differences to the change, not random-init kernel-race
drift cascading through training x eval LOB-sim trajectories. The eval-collapse
investigation (pearl_reward_signal_anti_aligned_with_pnl, multi-head spec,
regime-invariance spec) is now testable with trustworthy verdicts.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
207 lines
3.4 KiB
Plaintext
207 lines
3.4 KiB
Plaintext
# Build artifacts
|
|
/target/
|
|
target/
|
|
|
|
# Feature cache (precomputed .fxcache binaries, ~470MB each)
|
|
test_data/feature-cache/
|
|
/data/feature-cache/
|
|
*.fxcache
|
|
|
|
# Tier 1.5 mid-smoke local data (symlinks to test_data/futures-baseline-mbp10/)
|
|
test_data/futures-baseline-mid/
|
|
|
|
# IDE files
|
|
.vscode/
|
|
.idea/
|
|
*.swp
|
|
*.swo
|
|
|
|
# OS files
|
|
.DS_Store
|
|
Thumbs.db
|
|
|
|
# Logs
|
|
*.log
|
|
|
|
# Environment variables and secrets
|
|
.env
|
|
.env.*
|
|
!.env.example
|
|
|
|
|
|
# Secret files and directories
|
|
/config/secrets/
|
|
secrets/
|
|
*.key
|
|
*.pem
|
|
*.p12
|
|
*.pfx
|
|
*.crt
|
|
*.cert
|
|
|
|
# Credential files
|
|
credentials.json
|
|
credentials.toml
|
|
auth.json
|
|
auth.toml
|
|
ibkr.txt
|
|
|
|
# Certificate security files
|
|
certs/security.env
|
|
certs/production.env.template
|
|
certs/*.serial
|
|
certs/**/*.serial
|
|
|
|
# API keys and tokens
|
|
*api_key*
|
|
*token*
|
|
*secret*
|
|
!*secret*.example
|
|
!infra/modules/secrets/
|
|
!infra/live/production/secrets/
|
|
!infra/k8s/secrets/
|
|
# Allow ml-alpha horizon-token kernel sources (no secrets, just the v2 attention path)
|
|
!crates/ml-alpha/cuda/horizon_token_*
|
|
!crates/ml-alpha/src/horizon_token_*
|
|
!crates/ml-alpha/tests/horizon_token_*
|
|
!infra/k8s/secrets/*.yaml
|
|
!scripts/deploy-secrets.sh
|
|
|
|
# Database credentials
|
|
database.conf
|
|
db_config.json
|
|
|
|
# Temporary files
|
|
*.tmp
|
|
*.temp
|
|
|
|
# GPU test artifacts
|
|
gpu_test.rs
|
|
simd_bench
|
|
simd_bench.rs
|
|
temp_script.sh
|
|
|
|
# Build artifacts (additional)
|
|
**/*.rs.bk
|
|
*.pdb
|
|
Cargo.lock
|
|
|
|
# Coverage reports
|
|
tarpaulin-report.html
|
|
cobertura.xml
|
|
lcov.info
|
|
|
|
# Benchmark results
|
|
target/criterion/
|
|
target/bench/
|
|
|
|
# CI artifacts
|
|
.training-generated.yml
|
|
audit-results.json
|
|
geiger-report.md
|
|
security-report.md
|
|
outdated.json*.profraw
|
|
|
|
# Python virtual environments and cache
|
|
venv/
|
|
.venv/
|
|
.venv_databento/
|
|
venv_databento/
|
|
__pycache__/
|
|
*.py[cod]
|
|
*.so
|
|
.Python
|
|
*.egg-info/
|
|
.coverage
|
|
coverage.xml
|
|
htmlcov/
|
|
.python-version
|
|
|
|
# ML model checkpoints and training artifacts
|
|
diagnostic_data/
|
|
checkpoints/
|
|
ml/checkpoints/
|
|
ml/trained_models/*.safetensors
|
|
ml/trained_models/norm_stats_*.json
|
|
ml/trained_models/ppo_fold*.safetensors
|
|
ml/trained_models/ppo_fold*.json
|
|
ml/ml/
|
|
|
|
# Removed Windows wrapper files per user request
|
|
hive-mind-prompt-*.txt
|
|
|
|
# Node.js
|
|
node_modules/
|
|
web-dashboard/dist/
|
|
|
|
# Git worktrees
|
|
.worktrees/
|
|
.claude/worktrees/
|
|
|
|
# Serena (IDE agent)
|
|
.serena/
|
|
|
|
# Claude local state
|
|
.claude/settings.local.json
|
|
.claude/ralph-loop.local.md
|
|
.mcp.json
|
|
claude-flow.config.json
|
|
.swarm/
|
|
.hive-mind/
|
|
.claude-flow/
|
|
memory/
|
|
coordination/
|
|
memory/claude-flow-data.json
|
|
memory/sessions/*
|
|
!memory/sessions/README.md
|
|
memory/agents/*
|
|
!memory/agents/README.md
|
|
coordination/memory_bank/*
|
|
coordination/subtasks/*
|
|
coordination/orchestration/*
|
|
*.db
|
|
*.db-journal
|
|
*.db-wal
|
|
*.sqlite
|
|
*.sqlite-journal
|
|
*.sqlite-wal
|
|
claude-flow
|
|
|
|
# Terragrunt cache and generated files
|
|
.terragrunt-cache/
|
|
**/.terragrunt-cache/
|
|
.terraform.lock.hcl
|
|
**/.terraform.lock.hcl
|
|
|
|
# Data cache (downloaded market data, not checked in)
|
|
data/cache/
|
|
|
|
# Generated reports and validation artifacts in test_data
|
|
test_data/**/*REPORT*
|
|
test_data/**/*SUMMARY*
|
|
test_data/**/*METADATA*
|
|
test_data/**/README.md
|
|
test_data/databento/samples/
|
|
|
|
# Large binary files (prevent repo bloat - use LFS or external storage)
|
|
*.dbn
|
|
*.dbn.zst
|
|
*.safetensors
|
|
*.onnx
|
|
*.ot
|
|
*.pt
|
|
*.pth
|
|
*.bin
|
|
# Exception: deterministic test fixtures (golden bytes for bit-equivalence)
|
|
!crates/ml-alpha/tests/fixtures/*.bin
|
|
# Load test results
|
|
services/*/load_tests/results/
|
|
services/*/load_tests/*.json
|
|
!services/*/load_tests/package.json
|
|
.playwright-mcp/
|
|
crates/ml/ml/
|
|
|
|
# Foxhunt audit hook dedup state (cleared at SessionStart)
|
|
.claude/.foxhunt-audit-state
|
|
/config/ml/alpha_logits_cache.bin
|