Compare commits
105 Commits
ml-alpha-r
...
chore/gite
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
79f2ad94a4 | ||
|
|
c3aec12f9a | ||
|
|
b21d9b671d | ||
|
|
08cee0de0d | ||
|
|
80f3fbecef | ||
|
|
042fcd76e5 | ||
|
|
af3815ab6a | ||
|
|
3a18a348ac | ||
|
|
2eb8cd8333 | ||
|
|
2c08b49b4f | ||
|
|
b1a60ca134 | ||
|
|
f4f54c352b | ||
|
|
a0181431b3 | ||
|
|
feb0ae2541 | ||
|
|
4afaa5e248 | ||
|
|
e21a9a5da4 | ||
|
|
b2efc58940 | ||
|
|
fbb7b429e3 | ||
|
|
f4bd1e2432 | ||
|
|
6878aac950 | ||
|
|
bccd814716 | ||
|
|
cfa28368d5 | ||
|
|
f9b62d108c | ||
|
|
54e3d50db8 | ||
|
|
d82439219d | ||
|
|
3f321bb4bf | ||
|
|
02b851ca8b | ||
|
|
a3f309dd7e | ||
|
|
db8f5434a0 | ||
|
|
1cdbc81390 | ||
|
|
7662d82232 | ||
|
|
2a468e4b5e | ||
|
|
8ebb65f27e | ||
|
|
f320c5f09f | ||
|
|
dc2b7214f7 | ||
|
|
4bbf07bddf | ||
|
|
66e37955ce | ||
|
|
63d4f7f61d | ||
|
|
ba543ea85e | ||
|
|
e35b79a531 | ||
|
|
40c6e2f2ab | ||
|
|
4817bd4e7c | ||
|
|
a9a1c3c511 | ||
|
|
247e469a31 | ||
|
|
ac01db39fa | ||
|
|
9a90720500 | ||
|
|
ce8f8cbd46 | ||
|
|
ba7e7dca79 | ||
|
|
60e900e43a | ||
|
|
be084b5154 | ||
|
|
764fd99480 | ||
|
|
e731c2fc2f | ||
|
|
2b36929b5f | ||
|
|
ee2215f4d3 | ||
|
|
027d73a504 | ||
|
|
4f73e99225 | ||
|
|
9bf67e731d | ||
|
|
f203998613 | ||
|
|
107bcc6648 | ||
|
|
24ac921cd3 | ||
|
|
42e9621c47 | ||
|
|
04e7a61320 | ||
|
|
902eb1c85f | ||
|
|
df5c591441 | ||
|
|
9c2c38eb5d | ||
|
|
c10ebe0257 | ||
|
|
0fc3eac2ab | ||
|
|
244ccaaf0b | ||
|
|
7457ef679c | ||
|
|
6c113b0df2 | ||
|
|
d43fb5e61f | ||
|
|
0c5af2d9a0 | ||
|
|
30f944f015 | ||
|
|
05a81c5e5f | ||
|
|
2065c98c25 | ||
|
|
5ed94ddc97 | ||
|
|
a7d6671dea | ||
|
|
b664673df3 | ||
|
|
a9d1af9ec8 | ||
|
|
bc4cc46775 | ||
|
|
74b9d092ad | ||
|
|
f986510099 | ||
|
|
ad7feb61ff | ||
|
|
9642ad3156 | ||
|
|
dbbe8f7b22 | ||
|
|
ec308346b2 | ||
|
|
d6eb5b2ada | ||
|
|
4336a71e26 | ||
|
|
99b747ad2c | ||
|
|
a2cbed4a83 | ||
|
|
7edeec9106 | ||
|
|
ca3877c328 | ||
|
|
7495eaa286 | ||
|
|
0d6d58428d | ||
|
|
0f1a04d16d | ||
|
|
af5e6d523b | ||
|
|
cb402ab81f | ||
|
|
1283812c54 | ||
|
|
1cf38232e8 | ||
|
|
fa700c7b5f | ||
|
|
2bb2b50fd1 | ||
|
|
8df1c7eea2 | ||
|
|
e07c409950 | ||
|
|
565511c5f5 | ||
|
|
ce13a72ba1 |
1
.gitignore
vendored
1
.gitignore
vendored
@@ -204,3 +204,4 @@ crates/ml/ml/
|
||||
# Foxhunt audit hook dedup state (cleared at SessionStart)
|
||||
.claude/.foxhunt-audit-state
|
||||
/config/ml/alpha_logits_cache.bin
|
||||
data/surfer/
|
||||
|
||||
@@ -0,0 +1,192 @@
|
||||
# Surfer Phase 0 — Diversified-Trend Floor + Validation (CUDA/Rust) — Implementation Plan
|
||||
|
||||
> **For agentic workers:** REQUIRED SUB-SKILL: superpowers:executing-plans (or subagent-driven-development).
|
||||
> Steps use checkbox (`- [ ]`). TDD throughout — **GPU-oracle tests only, NO CPU reference oracle**
|
||||
> (`feedback_no_cpu_test_fallbacks`). **CUDA-only, CPU read-only** (`feedback_cpu_is_read_only`,
|
||||
> `pearl_cold_path_no_exception_to_gpu_drives`): every formula is a CUDA kernel; CPU reads only final gate scalars
|
||||
> from mapped-pinned buffers and only enumerates split index-sets as control flow.
|
||||
|
||||
**Goal:** Build the deterministic diversified-trend FLOOR and the CPCV/PBO/Deflated-Sharpe validation, **entirely as
|
||||
CUDA kernels + Rust orchestration in `crates/ml-alpha`**, then measure whether the floor has a real OOS edge after
|
||||
costs. The cheap, falsifiable gate before any ML (Phases 1-3). Runs on the local GPU (data is tiny: ~30 instruments
|
||||
× ~5000 days = ~150k floats).
|
||||
|
||||
**Architecture:** Reuse ml-alpha's CUDA pipeline — `build.rs` cubin precompile (`KERNELS` list), `pinned_mem.rs`
|
||||
mapped-pinned buffers, cudarc `load_cubin`/`launch_builder`, the determinism foundation (`FOXHUNT_DETERMINISTIC`,
|
||||
single-thread-per-series sequential sweeps, no `atomicAdd`, no nvrtc). GPU-drives / CPU-reads, uniformly.
|
||||
|
||||
**Tech Stack:** Rust 1.85 (`crates/ml-alpha`), CUDA 12.4 (pre-compiled cubins via build.rs, NO nvrtc), cudarc,
|
||||
mapped-pinned host buffers. Data via `crates/data` dbn decode (the L0-fixed `dbn_parser`). NO Python in the built
|
||||
system. NO GPU↔CPU roundtrips except the final gate-scalar readback.
|
||||
|
||||
**Spec:** `docs/superpowers/specs/2026-06-05-surfer-uncorrelated-universe-design.md`
|
||||
|
||||
---
|
||||
|
||||
## GPU-contract compliance (apply to EVERY task)
|
||||
- **No `atomicAdd`** (`feedback_no_atomicadd`) → reductions are tree-reduce or single-thread-per-series sequential.
|
||||
- **Mapped-pinned only** for any host visibility (`feedback_no_htod_htoh_only_mapped_pinned`); no raw HtoD/DtoH for
|
||||
compute results. Data UPLOAD uses the existing mapped-pinned dataset path.
|
||||
- **Determinism**: per-series sequential sweeps (EWMA vol, portfolio carry) are single-thread-per-instrument with a
|
||||
sequential time loop — mirror `gae_backward_sweep.cu` (deterministic, no inter-thread race). Verify with
|
||||
`determinism-check`-style same-seed bit-equality.
|
||||
- **No CPU compute**: signals, vol, sizing, returns, Sharpe, PBO, Deflated-Sharpe are kernels. CPU may only:
|
||||
(a) enumerate combinatorial split index-sets (control flow, like `epoch_idx`), uploading them as int arrays;
|
||||
(b) read the FINAL gate scalars from mapped-pinned at the end.
|
||||
- **No CPU test oracle**: tests assert kernel outputs against HAND-DERIVED analytic constants read from mapped-pinned
|
||||
(e.g. "rising series → TSMOM signal == +1.0"), never a CPU re-implementation of the algorithm.
|
||||
|
||||
## File structure
|
||||
| File | Change | Responsibility |
|
||||
|---|---|---|
|
||||
| `crates/ml-alpha/cuda/surfer_continuous_adjust.cu` | new | element-wise backward ratio-adjust of stitched closes |
|
||||
| `crates/ml-alpha/cuda/surfer_tsmom_signal.cu` | new | per (instrument,day) 1/3/12-mo sign-ensemble signal |
|
||||
| `crates/ml-alpha/cuda/surfer_ewma_vol.cu` | new | per-instrument EWMA vol (single-thread-per-instrument seq) |
|
||||
| `crates/ml-alpha/cuda/surfer_floor_portfolio.cu` | new | inverse-vol size + vol-target + no-trade band → daily port. returns |
|
||||
| `crates/ml-alpha/cuda/surfer_cpcv_sharpe.cu` | new | per (path,config) Sharpe over returns matrix (parallel) |
|
||||
| `crates/ml-alpha/cuda/surfer_overfit_stats.cu` | new | PBO + Deflated-Sharpe scalars on GPU (erff cdf) |
|
||||
| `crates/ml-alpha/build.rs` | edit `KERNELS` | register the 6 cubins |
|
||||
| `crates/ml-alpha/src/surfer/mod.rs` | new | Rust orchestration (load→upload→launch→read gates) |
|
||||
| `crates/ml-alpha/src/surfer/continuous.rs` | new | roll detection + stitch (data-assembly; decode via `data` crate) |
|
||||
| `crates/ml-alpha/examples/surfer_phase0.rs` | new | the verdict-run binary |
|
||||
| `crates/ml-alpha/tests/surfer_floor_invariants.rs` | new | GPU-oracle invariant tests |
|
||||
|
||||
---
|
||||
|
||||
## Task A: Continuous-contract series (data-assembly + ratio-adjust kernel)
|
||||
|
||||
**Files:** `crates/ml-alpha/src/surfer/continuous.rs`; `crates/ml-alpha/cuda/surfer_continuous_adjust.cu`; build.rs.
|
||||
|
||||
- [ ] **Step 1:** Add `"surfer_continuous_adjust"` to `KERNELS` in `crates/ml-alpha/build.rs`.
|
||||
- [ ] **Step 2: kernel** `surfer_continuous_adjust.cu` — given stitched raw closes `[N_days]` and a roll-ratio array
|
||||
`ratio[N_days]` (cumulative back-adjust factor per day, =1.0 after the last roll), output `adj[i]=raw[i]*ratio[i]`.
|
||||
Element-wise, one thread per day. (Roll DETECTION — which contract is active per day by max volume, and the per-roll
|
||||
ratio — is data-assembly control flow in `continuous.rs`: it decides WHAT was loaded, not a compute-over-results.
|
||||
It reads per-expiry volume from the decoded dbn, picks argmax-volume contract per day, and builds the cumulative
|
||||
`ratio` array, which is then UPLOADED mapped-pinned; the multiply is the kernel.)
|
||||
- [ ] **Step 3: `continuous.rs`** — `pub fn build_continuous(root, decoder) -> ContinuousSeries { ts, raw_close_d,
|
||||
adj_close_d, volume… }` decoding per-expiry OHLCV via the `data` crate, picking active contract per day by volume,
|
||||
computing cumulative ratio, uploading raw+ratio mapped-pinned, launching the adjust kernel → `adj_close_d` on GPU.
|
||||
Keep BOTH `raw_close_d` (fills/reward USD) and `adj_close_d` (signals/returns). Per spec §4: NEVER Panama; reward on raw.
|
||||
- [ ] **Step 4: GPU-oracle test** (`surfer_floor_invariants.rs`): upload raw=[100,101,102,202,204], roll at day3,
|
||||
ratio=[r,r,r,1,1] with r=202/102; launch; read `adj` mapped-pinned; assert `adj[3]==202`, `adj[2]==102*r`, and the
|
||||
seam log-return `ln(adj[3]/adj[2])≈0`. (Analytic constants, no CPU re-impl.)
|
||||
- [ ] **Step 5:** `SQLX_OFFLINE=true cargo test -p ml-alpha --test surfer_floor_invariants continuous` → PASS.
|
||||
- [ ] **Step 6: Commit** — `feat(surfer): continuous-contract assembly + ratio-adjust kernel`.
|
||||
|
||||
---
|
||||
|
||||
## Task B: Floor kernels (TSMOM signal, EWMA vol, portfolio backtest)
|
||||
|
||||
**Files:** `surfer_tsmom_signal.cu`, `surfer_ewma_vol.cu`, `surfer_floor_portfolio.cu`; `src/surfer/mod.rs`; build.rs.
|
||||
|
||||
- [ ] **Step 1:** Register the 3 kernels in build.rs `KERNELS`.
|
||||
- [ ] **Step 2: `surfer_tsmom_signal.cu`** — input `logc[N_inst × N_days]`; per (inst, day) output
|
||||
`sig = clip( mean_{L∈{21,63,252}} sign(logc[end]-logc[end-L]), -1, 1)`, where for L=252 `end = day-21` (skip recent),
|
||||
else `end=day`; 0 before warmup. Grid over (inst×day), branch-free. (Sign-inversion for yield-based rates handled
|
||||
by a per-instrument `sign_flip[inst]` multiplier uploaded as config.)
|
||||
- [ ] **Step 3: `surfer_ewma_vol.cu`** — per instrument, single thread, sequential over days:
|
||||
`v[t] = a·r[t]^2 + (1-a)·v[t-1]`, `a=1-exp(ln0.5/halflife)`, `sigma[t]=sqrt(v[t]·252)`. One thread per instrument
|
||||
(deterministic sequential carry — mirror `gae_backward_sweep.cu`'s single-thread-per-series pattern). No atomics.
|
||||
- [ ] **Step 4: `surfer_floor_portfolio.cu`** — per day (sequential single-thread for the portfolio vol-target carry;
|
||||
per-instrument inner loop): target weight `w_i = sig_i · (risk_budget / sigma_i)`; apply no-trade band (skip change
|
||||
if |w_i - w_i_prev| < band); portfolio gross scaled to `vol_target/realized_port_vol` (realized from a trailing
|
||||
window kernel-side); daily portfolio return `R[t] = Σ_i w_i[t-1]·r_i[t]`. Output `port_ret_d[N_days]`,
|
||||
`turnover_d[N_days]` on GPU.
|
||||
- [ ] **Step 5: `src/surfer/mod.rs`** — `pub struct Floor` loads the 3 cubins; `pub fn run_floor(&mut self,
|
||||
series: &[ContinuousSeries], cfg: FloorCfg) -> FloorOutput` launches signal→vol→portfolio, returns GPU handles
|
||||
(`port_ret_d`, `turnover_d`). No host math.
|
||||
- [ ] **Step 6: GPU-oracle tests** (analytic): (a) monotonic-rising logc → `tsmom_signal` last value == +1.0;
|
||||
monotonic-falling → −1.0; (b) constant-vol synthetic → `ewma_vol` converges to the known σ (hand value); (c) a
|
||||
2-instrument hand-built case → assert `port_ret_d[t]` equals the hand-derived `Σ w·r` for one step. Read all via
|
||||
mapped-pinned; no CPU re-impl of the formulas.
|
||||
- [ ] **Step 7:** Run tests → PASS. **Smoke**: `run_floor` on the local 4 instruments (6E/ES/NQ/ZN), read
|
||||
`port_ret_d` mapped-pinned, print annualized Sharpe (expect noisy/underpowered on 2y — proves the machinery).
|
||||
- [ ] **Step 8: Commit** — `feat(surfer): floor kernels (TSMOM signal + EWMA vol + portfolio backtest)`.
|
||||
|
||||
---
|
||||
|
||||
## Task C: Validation kernels (CPCV Sharpe, PBO, Deflated Sharpe) + cost
|
||||
|
||||
**Files:** `surfer_cpcv_sharpe.cu`, `surfer_overfit_stats.cu`; `src/surfer/validation.rs`; build.rs.
|
||||
|
||||
- [ ] **Step 1:** Register both kernels in build.rs `KERNELS`.
|
||||
- [ ] **Step 2: cost in the portfolio kernel** — extend `surfer_floor_portfolio.cu` to subtract per-rebalance cost:
|
||||
`cost[t] = turnover[t]·(per_contract_usd/notional + half_spread_bps/1e4)` (+ roll cost on roll days), producing a
|
||||
NET `port_ret_d`. Costs are config scalars uploaded mapped-pinned. (No host arithmetic.)
|
||||
- [ ] **Step 3: `surfer_cpcv_sharpe.cu`** — inputs: net `port_ret_d[N_days]` (or `[N_cfg × N_days]` if sweeping
|
||||
configs), and CPU-enumerated `test_mask[N_path × N_days]` (uint8, uploaded). Per (path[, cfg]) compute mean/std →
|
||||
Sharpe over the path's OOS days via tree-reduce (no atomics). Output `oos_sharpe[N_path(× N_cfg)]` on GPU.
|
||||
- [ ] **Step 4: `surfer_overfit_stats.cu`** — on GPU: (a) Deflated Sharpe `DSR = Φ((SR−SR0)√(T−1)/√(1−γ3·SR+
|
||||
(γ4−1)/4·SR²))` with `Φ` via `0.5·erfcf(-x/√2)`, `SR0` from `n_trials` (Euler-γ order-statistic formula), `n_trials`
|
||||
uploaded; (b) PBO from the per-split IS-argmax / OOS-rank arrays (the CSCV split sets enumerated CPU-side, the
|
||||
per-split argmax+rank reduced GPU-side). Output `{dsr, pbo, oos_sharpe_5pct, rank_consistency}` to a mapped-pinned
|
||||
gate-scalar struct.
|
||||
- [ ] **Step 5: `src/surfer/validation.rs`** — `pub fn validate(net_ret_d, cfg_grid) -> Gates`: CPU ENUMERATES the
|
||||
CPCV combinations and CSCV splits (control flow), uploads index masks mapped-pinned, launches the two kernels,
|
||||
reads the final `Gates` struct from mapped-pinned. The ONLY host readback.
|
||||
- [ ] **Step 6: GPU-oracle tests**: (a) feed a returns series with known mean/std → assert `oos_sharpe` equals the
|
||||
hand-derived value; (b) Deflated-Sharpe: same SR, n_trials=1 vs 50 → assert `dsr` decreases (monotonic, hand-checked
|
||||
direction); (c) PBO: construct IS-best=OOS-worst returns → assert `pbo > 0.5`. All via mapped-pinned; no CPU oracle.
|
||||
- [ ] **Step 7:** Run tests → PASS. **Determinism**: run `validate` twice same inputs → gate scalars bit-equal.
|
||||
- [ ] **Step 8: Commit** — `feat(surfer): CUDA validation (CPCV Sharpe + PBO + Deflated Sharpe + cost)`.
|
||||
|
||||
---
|
||||
|
||||
## Task D: Acquire broad-universe deep-history data
|
||||
|
||||
**Files:** reuse `services/data_acquisition` / the existing Databento pipeline; cache decoded series.
|
||||
|
||||
- [ ] **Step 1:** Target ≥15 FULL-SIZE roots across classes (equity ES/NQ/YM/RTY; rates ZN/ZB/ZF/ZT; FX 6E/6J/6B/6A/6C;
|
||||
metals GC/SI/HG; energy CL/NG/RB; ags ZC/ZS/ZW), max history via Databento `GLBX.MDP3 ohlcv-1d`. (Edge is identical to
|
||||
micros; micros are a deployment-sizing concern only.)
|
||||
- [ ] **Step 2:** Fetch per-expiry daily OHLCV (existing data pipeline), land on the PVC/local; the surfer's
|
||||
`continuous.rs` builds continuous series at load.
|
||||
- [ ] **Step 3:** Validate each root: ≥10y, decode clean via the L0-fixed `dbn_parser`, ratio seam-returns sane.
|
||||
- [ ] **STOP-if:** Databento history/cost prohibitive → document and proceed with the largest free continuous set,
|
||||
noting the data-quality caveat in the verdict.
|
||||
|
||||
---
|
||||
|
||||
## Task E: Phase-0 verdict run (the decisive gate)
|
||||
|
||||
**Files:** `crates/ml-alpha/examples/surfer_phase0.rs`.
|
||||
|
||||
- [ ] **Step 1:** Wire: load all continuous series → upload → `run_floor` (net of cost) → `validate(net_ret_d, grid)`
|
||||
over the lookback/vol-target config grid → read the `Gates` struct (mapped-pinned, the only readback) → print a
|
||||
verdict block (per `tier1_5_verdict.py` STYLE, but the binary is Rust reading GPU scalars):
|
||||
- SV-G1: CPCV (≥50 paths) 5th-pct OOS Sharpe after costs > 0.
|
||||
- SV-G3: Deflated Sharpe > 0.95, `n_trials` = ALL trials ever (record the count: 64 commits + session harnesses +
|
||||
every floor lookback/vol-target variant in the grid).
|
||||
- SV-G4: IS↔OOS config rank-consistency > 0.
|
||||
- SV-G5: edge survives 2× cost + roll (re-run with doubled cost scalars).
|
||||
- [ ] **Step 2:** Determinism: same-inputs run twice → identical verdict.
|
||||
- [ ] **Step 3:** Record verdict in spec §9 + a memory pearl.
|
||||
- [ ] **Step 4: Commit.**
|
||||
|
||||
**VERDICT BRANCH:**
|
||||
- **PASS** (floor clears SV-G1/G3/G5 OOS net of costs) → diversified-trend premium is real & capturable on this
|
||||
universe → proceed to Phase 1 (regime overlay). The floor becomes the live benchmark.
|
||||
- **FAIL** → even the floor has no powered OOS edge → do NOT build ML. Reassess universe/history/thesis. Cost: a few
|
||||
days of local-GPU dev, $0 cluster.
|
||||
|
||||
---
|
||||
|
||||
## Self-review
|
||||
- **CUDA-only / CPU-read-only**: every formula is a kernel (Tasks A-C); CPU only enumerates split masks + reads the
|
||||
final `Gates` struct (Task C Step 5, Task E Step 1). Complies with `feedback_cpu_is_read_only`,
|
||||
`pearl_cold_path_no_exception_to_gpu_drives`. ✓
|
||||
- **No CPU test oracle**: all tests assert kernel outputs vs hand-derived analytic constants read from mapped-pinned
|
||||
(`feedback_no_cpu_test_fallbacks`). ✓
|
||||
- **GPU-contract**: no atomicAdd (tree-reduce / single-thread-per-series), mapped-pinned only, cubins via build.rs
|
||||
(no nvrtc), determinism via sequential-per-series sweeps + same-input bit-equality. ✓
|
||||
- **Spec coverage**: floor §3.1 → Task B ✓ | validation gates §5 → Task C+E ✓ | universe+pipeline §4 → Task A+D ✓ |
|
||||
Phase-0 STOP §6 → Task E verdict ✓ | edge-validation on full-size (not micros) → Task D ✓ | reuse ml-alpha infra §3.5 → file structure ✓.
|
||||
- **Placeholders**: kernel bodies are given as signature + exact formula + determinism pattern + the existing kernel to
|
||||
mirror (gae_backward_sweep / ema_update_per_step / the reduce kernels) — a competent engineer following ml-alpha's
|
||||
`cuda/` conventions writes them directly. Not silent gaps.
|
||||
|
||||
## Execution note
|
||||
Tasks A-C run on the local GPU against the 4 local instruments to prove the kernels + machinery; D-E need the broad
|
||||
universe. No cluster, no Python in the built system. The floor + CUDA validation are the reusable core the ML surfer
|
||||
(Phases 1-3) sits on.
|
||||
263
docs/superpowers/plans/2026-06-21-decommission-rust-infra.md
Normal file
263
docs/superpowers/plans/2026-06-21-decommission-rust-infra.md
Normal file
@@ -0,0 +1,263 @@
|
||||
# Decommission Dead Rust Infra — Implementation Plan (Phase 1)
|
||||
|
||||
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:executing-plans (INLINE, with checkpoints).
|
||||
> **Do NOT run this subagent-driven** — every task performs irreversible production deletes that need
|
||||
> human confirmation at each gate. Steps use checkbox (`- [ ]`) syntax.
|
||||
|
||||
**Goal:** Remove the now-unused Rust-specific infra (GPU pools, GPU CI runners, build-cache PVCs, Rust
|
||||
artifact buckets, training/GPU manifests) — preserving ALL `.dbn` market data and every fxhnt/platform resource.
|
||||
|
||||
**Architecture:** Verify-then-delete, gated. Terraform destroys the GPU/precompute pools (plan reviewed
|
||||
for exactly-3); helm uninstalls the GPU runners; kubectl deletes the unmounted build-cache PVCs; mc deletes
|
||||
the no-`.dbn` Rust buckets; repo + cluster Argo templates cleaned up. Cockpit/dagster/fund verified healthy.
|
||||
|
||||
**Tech Stack:** terragrunt + OpenTofu (Scaleway provider), helm, kubectl, mc (MinIO), Scaleway Kapsule.
|
||||
|
||||
**Spec:** `docs/superpowers/specs/2026-06-21-decommission-rust-infra-design.md`
|
||||
|
||||
---
|
||||
|
||||
## Shared env (export at the start of each shell session that runs terragrunt/scw)
|
||||
```bash
|
||||
cd /home/jgrusewski/Work/foxhunt
|
||||
export TG_TF_PATH=tofu TERRAGRUNT_TFPATH=tofu
|
||||
export SCW_ACCESS_KEY=$(kubectl get secret scaleway-credentials -n foxhunt -o jsonpath='{.data.access-key}'|base64 -d)
|
||||
export SCW_SECRET_KEY=$(kubectl get secret scaleway-credentials -n foxhunt -o jsonpath='{.data.secret-key}'|base64 -d)
|
||||
export SCW_DEFAULT_PROJECT_ID=$(kubectl get secret scaleway-credentials -n foxhunt -o jsonpath='{.data.project-id}'|base64 -d)
|
||||
export SCW_DEFAULT_REGION=fr-par SCW_DEFAULT_ZONE=fr-par-2
|
||||
export TF_HTTP_USERNAME=root TF_HTTP_PASSWORD=$(kubectl get secret gitlab-pat -n foxhunt -o jsonpath='{.data.token}'|base64 -d)
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Task 0: Branch
|
||||
- [ ] **Step 1**
|
||||
```bash
|
||||
cd /home/jgrusewski/Work/foxhunt && git checkout -b chore/decommission-rust-infra
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Task 1: Destroy the GPU + precompute pools (Terraform)
|
||||
|
||||
**Files:** Modify `infra/live/production/kapsule/terragrunt.hcl`
|
||||
|
||||
- [ ] **Step 1: Set the three enable flags to false**
|
||||
|
||||
In `infra/live/production/kapsule/terragrunt.hcl`, change:
|
||||
```
|
||||
enable_ci_compile_cpu_hm_pool = true
|
||||
```
|
||||
to `false`; and
|
||||
```
|
||||
enable_ci_training_l40s_pool = true
|
||||
```
|
||||
to `false`; and
|
||||
```
|
||||
enable_ci_training_h100_pool = true
|
||||
```
|
||||
to `false`. (Leave `enable_ci_compile_cpu_pool = true` — Python cockpit builds use it.)
|
||||
|
||||
- [ ] **Step 2: Plan and GATE on exactly-3-destroys**
|
||||
|
||||
Run (with shared env, in `infra/live/production/kapsule`):
|
||||
```bash
|
||||
cd /home/jgrusewski/Work/foxhunt/infra/live/production/kapsule
|
||||
terragrunt plan -input=false -no-color 2>&1 | sed 's/.*tofu: //' | grep -iE "^Plan:|will be destroyed|# scaleway"
|
||||
```
|
||||
Expected: `Plan: 0 to add, 0 to change, 3 to destroy.` and the 3 destroyed are EXACTLY
|
||||
`scaleway_k8s_pool.ci_compile_cpu_hm[0]`, `scaleway_k8s_pool.ci_training_l40s[0]`,
|
||||
`scaleway_k8s_pool.ci_training_h100[0]`.
|
||||
**STOP if the count ≠ 3 or any other resource is destroyed.**
|
||||
|
||||
- [ ] **Step 3: Apply**
|
||||
```bash
|
||||
terragrunt apply -input=false -no-color -auto-approve 2>&1 | sed 's/.*tofu: //' | grep -iE "Apply complete|Destroy complete|Error"
|
||||
```
|
||||
Expected: `Apply complete! Resources: 0 added, 0 changed, 3 destroyed.`
|
||||
|
||||
- [ ] **Step 4: Verify pools gone**
|
||||
```bash
|
||||
scw k8s pool list cluster-id=34a1e3c4-ac35-48c8-ab49-5f6ec4df32c1 -o json 2>/dev/null | python3 -c "import sys,json;print([p['name'] for p in json.load(sys.stdin)])"
|
||||
```
|
||||
Expected: `['ci-compile-cpu', 'platform']` (no l40s/h100/hm).
|
||||
|
||||
- [ ] **Step 5: Commit**
|
||||
```bash
|
||||
cd /home/jgrusewski/Work/foxhunt
|
||||
git add infra/live/production/kapsule/terragrunt.hcl
|
||||
git commit -m "chore(infra): destroy unused Rust GPU + precompute pools (L40S/H100/hm)
|
||||
|
||||
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>"
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Task 2: Uninstall the GPU CI runners (helm)
|
||||
|
||||
- [ ] **Step 1: Uninstall the 3 GPU runners**
|
||||
```bash
|
||||
for r in gitlab-runner-h100 gitlab-runner-h100-sxm gitlab-runner-h100x2; do
|
||||
helm uninstall "$r" -n foxhunt 2>&1 | tail -1
|
||||
done
|
||||
```
|
||||
Expected: `release "<r>" uninstalled` for each.
|
||||
|
||||
- [ ] **Step 2: Decide the main `gitlab-runner`**
|
||||
|
||||
Check whether anything still uses it (fxhnt has no `.gitlab-ci.yml`; cockpit deploys via Argo):
|
||||
```bash
|
||||
kubectl get pods -n foxhunt | grep gitlab-runner # expect: no running runner pods
|
||||
# Inspect what the main runner is registered for (manual judgement):
|
||||
helm get values gitlab-runner -n foxhunt 2>/dev/null | grep -iE "tags|runUntagged|description" | head
|
||||
```
|
||||
- [ ] **Step 3: If confirmed dead, uninstall it; else keep (document the decision):**
|
||||
```bash
|
||||
helm uninstall gitlab-runner -n foxhunt 2>&1 | tail -1 # ONLY if Step 2 confirms no consumer
|
||||
```
|
||||
|
||||
- [ ] **Step 4: Verify**
|
||||
```bash
|
||||
helm list -n foxhunt | grep -i runner # expect: none (or only a deliberately-kept one)
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Task 3: Delete the build-cache PVCs (re-verify unmounted first)
|
||||
|
||||
- [ ] **Step 1: Re-confirm each PVC is unmounted (safety)**
|
||||
```bash
|
||||
for p in cargo-target-cpu cargo-target-cuda cargo-target-cuda-test sccache-cpu sccache-cuda; do
|
||||
echo -n "$p: "; kubectl describe pvc $p -n foxhunt 2>/dev/null | grep -A1 "Used By" | tr '\n' ' '; echo
|
||||
done
|
||||
```
|
||||
Expected: every line shows `Used By: <none>`. **STOP on any PVC that lists a pod.**
|
||||
|
||||
- [ ] **Step 2: Delete them**
|
||||
```bash
|
||||
kubectl delete pvc cargo-target-cpu cargo-target-cuda cargo-target-cuda-test sccache-cpu sccache-cuda -n foxhunt 2>&1 | tail
|
||||
```
|
||||
Expected: `persistentvolumeclaim "<name>" deleted` ×5.
|
||||
|
||||
- [ ] **Step 3: Verify reclaim**
|
||||
```bash
|
||||
kubectl get pvc -n foxhunt | grep -E "cargo-target|sccache" || echo "all build-cache PVCs gone (~175Gi reclaimed)"
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Task 4: Delete the Rust artifact buckets (re-verify 0 `.dbn` first)
|
||||
|
||||
- [ ] **Step 1: Start a MinIO port-forward + mc alias**
|
||||
```bash
|
||||
pkill -f "port-forward svc/minio" 2>/dev/null; sleep 1
|
||||
kubectl port-forward svc/minio -n foxhunt 19000:9000 >/tmp/minio-pf.log 2>&1 &
|
||||
sleep 4
|
||||
ak=$(kubectl get secret minio-credentials -n foxhunt -o jsonpath='{.data.access-key}'|base64 -d)
|
||||
sk=$(kubectl get secret minio-credentials -n foxhunt -o jsonpath='{.data.secret-key}'|base64 -d)
|
||||
mc alias set fxl http://localhost:19000 "$ak" "$sk"
|
||||
```
|
||||
|
||||
- [ ] **Step 2: Re-verify the 3 target buckets contain ZERO `.dbn`/market data, and the KEPT one does**
|
||||
```bash
|
||||
for b in foxhunt-binaries foxhunt-training-results foxhunt-models; do
|
||||
echo -n "$b dbn-count="; mc ls --recursive fxl/$b 2>/dev/null | grep -icE "\.dbn|\.zst|databento|mbp"
|
||||
done
|
||||
echo -n "KEEP foxhunt-training-data dbn-count="; mc ls --recursive fxl/foxhunt-training-data 2>/dev/null | grep -icE "\.dbn|\.zst|databento|mbp"
|
||||
```
|
||||
Expected: the 3 targets show `0`; `foxhunt-training-data` shows `>0`. **STOP if any target shows >0.**
|
||||
|
||||
- [ ] **Step 3: Delete the 3 Rust buckets**
|
||||
```bash
|
||||
for b in foxhunt-binaries foxhunt-training-results foxhunt-models; do
|
||||
mc rb --force fxl/$b 2>&1 | tail -1
|
||||
done
|
||||
```
|
||||
Expected: `Removed '<bucket>' successfully.` ×3.
|
||||
|
||||
- [ ] **Step 4: Verify market data intact**
|
||||
```bash
|
||||
mc du fxl/foxhunt-training-data # expect: still ~38GiB, 54+ objects
|
||||
kubectl get pvc training-data-pvc test-data-pvc feature-cache-pvc -n foxhunt # expect: all 3 still Bound
|
||||
pkill -f "port-forward svc/minio" 2>/dev/null
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Task 5: Remove dead manifests + Argo templates (repo + cluster)
|
||||
|
||||
**Files:** Remove `infra/k8s/training/`, `infra/k8s/gpu-overlays/`, Rust Argo templates, the databento job, GPU-runner helm values.
|
||||
|
||||
- [ ] **Step 1: Identify the Rust Argo WorkflowTemplates in the cluster (keep fxhnt-cockpit!)**
|
||||
```bash
|
||||
kubectl get wftmpl -n foxhunt -o name 2>/dev/null
|
||||
```
|
||||
Note the Rust ones (e.g. train-multi-seed, lob-backtest-sweep, ci-pipeline, alpha-rl-*). **Do NOT touch `fxhnt-cockpit`.**
|
||||
|
||||
- [ ] **Step 2: Delete the dead WorkflowTemplates from the cluster** (substitute the exact names from Step 1):
|
||||
```bash
|
||||
# example — replace with the actual Rust template names from Step 1:
|
||||
kubectl delete wftmpl -n foxhunt <rust-template-1> <rust-template-2> ... 2>&1 | tail
|
||||
```
|
||||
|
||||
- [ ] **Step 3: Remove dead manifests + the now-unused Terraform pool blocks from the repo**
|
||||
```bash
|
||||
cd /home/jgrusewski/Work/foxhunt
|
||||
git rm -r infra/k8s/training infra/k8s/gpu-overlays 2>/dev/null
|
||||
git rm infra/k8s/jobs/download-trades-job.yaml infra/k8s/argo/train-multi-seed-template.yaml infra/k8s/argo/lob-backtest-sweep-template.yaml infra/k8s/argo/ci-pipeline-template.yaml 2>/dev/null
|
||||
# (also git rm any GPU-runner helm value files + alpha-rl argo templates found under infra/)
|
||||
ls infra/k8s/argo infra/k8s/jobs # eyeball what remains; keep cockpit/fxhnt/fund + cert-manager etc.
|
||||
```
|
||||
- [ ] **Step 4: Remove the dead pool resource blocks + variables from the kapsule module** (optional tidy):
|
||||
in `infra/modules/kapsule/main.tf` remove the `scaleway_k8s_pool.ci_training_l40s`,
|
||||
`...ci_training_h100`, `...ci_compile_cpu_hm` resource blocks (now count=0); in `variables.tf` remove
|
||||
their `enable_*`/`*_type`/`*_max_size` vars; in `terragrunt.hcl` remove the 3 dead `enable_*`/`*_type`
|
||||
lines. Then re-plan to confirm still clean:
|
||||
```bash
|
||||
cd infra/live/production/kapsule && terragrunt plan -input=false -no-color 2>&1 | sed 's/.*tofu: //' | grep -iE "^Plan:|No changes"
|
||||
```
|
||||
Expected: `No changes` (removing count=0 blocks is a no-op against the cluster).
|
||||
|
||||
- [ ] **Step 5: Commit**
|
||||
```bash
|
||||
cd /home/jgrusewski/Work/foxhunt
|
||||
git add -A infra
|
||||
git commit -m "chore(infra): remove dead Rust training/GPU manifests, Argo templates, pool config
|
||||
|
||||
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>"
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Task 6: Final verification (cluster + fund still healthy)
|
||||
|
||||
- [ ] **Step 1: Platform/fund health**
|
||||
```bash
|
||||
curl -sS -m 12 -o /dev/null -w "dashboard HTTP %{http_code}\n" https://dashboard.fxhnt.ai/
|
||||
kubectl get pods -n foxhunt | grep -E "dagster|fxhnt-dashboard|gitlab-webservice" | grep -vE "Completed"
|
||||
kubectl get pvc training-data-pvc test-data-pvc feature-cache-pvc fxhnt-surfer-data fxhnt-backtest-data -n foxhunt
|
||||
```
|
||||
Expected: dashboard 200; dagster/dashboard/gitlab Running; all kept PVCs Bound.
|
||||
|
||||
- [ ] **Step 2: Terraform clean**
|
||||
```bash
|
||||
cd /home/jgrusewski/Work/foxhunt/infra/live/production/kapsule && terragrunt plan -input=false -no-color 2>&1 | sed 's/.*tofu: //' | grep -iE "^Plan:|No changes"
|
||||
```
|
||||
Expected: `No changes`.
|
||||
|
||||
- [ ] **Step 3: Merge to main**
|
||||
```bash
|
||||
cd /home/jgrusewski/Work/foxhunt
|
||||
git checkout main && git merge --ff-only chore/decommission-rust-infra && git branch -d chore/decommission-rust-infra
|
||||
git push origin main
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Notes for the executor
|
||||
- **Data safety is paramount:** never delete `training-data-pvc`, `test-data-pvc`, `feature-cache-pvc`, the
|
||||
`foxhunt-training-data` bucket, or any fxhnt/platform resource. Re-verify (`Used By`, `.dbn` count) at each
|
||||
delete gate; STOP on any surprise.
|
||||
- **Terraform gate:** Task 1 Step 2 must show exactly the 3 GPU/precompute pools destroyed — stop otherwise.
|
||||
- `pkill -f "port-forward svc/minio"` after Task 4 to clean up the forward.
|
||||
- If the main `gitlab-runner` is ambiguous, KEEP it (a kept idle runner is harmless; a wrongly-removed one breaks CI).
|
||||
472
docs/superpowers/plans/2026-06-21-gitea-replace-gitlab.md
Normal file
472
docs/superpowers/plans/2026-06-21-gitea-replace-gitlab.md
Normal file
@@ -0,0 +1,472 @@
|
||||
# Gitea Replaces GitLab + Cockpit Registry to Scaleway — Implementation Plan (Phase 2B)
|
||||
|
||||
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:executing-plans (INLINE, with checkpoints).
|
||||
> **Do NOT run subagent-driven** — this touches live git access, the container registry, and the cockpit
|
||||
> deploy; each cutover step needs a human-confirmed gate. Steps use checkbox (`- [ ]`) syntax.
|
||||
|
||||
**Goal:** Stand up Gitea (chart + existing Postgres) as the git host behind the unchanged `git.fxhnt.ai`,
|
||||
migrate both repos, and move the cockpit image to Scaleway Container Registry — so GitLab can be removed in 2C.
|
||||
|
||||
**Architecture:** Gitea comes up internal-only; repos migrate + validate over a port-forward; the cockpit
|
||||
image moves to Scaleway registry and is build+rollout tested; THEN `git.fxhnt.ai` (HTTPS + SSH `:2222`)
|
||||
cuts over to Gitea in one switch; finally a Gitea webhook → Argo Events auto-deploys the cockpit. GitLab
|
||||
stays running (unaddressed) as the rollback net until 2C.
|
||||
|
||||
**Tech Stack:** Gitea Helm chart (`gitea-charts/gitea`), existing in-cluster PostgreSQL, nginx+socat
|
||||
Tailscale proxy, Scaleway Container Registry, Argo Events, kaniko.
|
||||
|
||||
**Spec:** `docs/superpowers/specs/2026-06-21-gitea-replace-gitlab-design.md`
|
||||
|
||||
**TWO REPOS:** infra changes → `foxhunt` (`~/Work/foxhunt`); cockpit build/deploy → `fxhnt`
|
||||
(`~/Work/fxhnt`). Each gets its own branch + commits.
|
||||
|
||||
---
|
||||
|
||||
## Shared env
|
||||
```bash
|
||||
export SCW_ACCESS_KEY=$(kubectl get secret scaleway-credentials -n foxhunt -o jsonpath='{.data.access-key}'|base64 -d)
|
||||
export SCW_SECRET_KEY=$(kubectl get secret scaleway-credentials -n foxhunt -o jsonpath='{.data.secret-key}'|base64 -d)
|
||||
export SCW_DEFAULT_PROJECT_ID=$(kubectl get secret scaleway-credentials -n foxhunt -o jsonpath='{.data.project-id}'|base64 -d)
|
||||
export SCW_DEFAULT_REGION=fr-par SCW_DEFAULT_ZONE=fr-par-2
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Task 0: Branches
|
||||
- [ ] **Step 1**
|
||||
```bash
|
||||
cd /home/jgrusewski/Work/foxhunt && git checkout -b chore/gitea-replace-gitlab
|
||||
cd /home/jgrusewski/Work/fxhnt && git checkout -b chore/gitea-cockpit-registry
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Task 1: Postgres `gitea` DB + secrets
|
||||
|
||||
**Files:** none committed (cluster secrets). Record secret names only.
|
||||
|
||||
- [ ] **Step 1: Read existing postgres superuser creds**
|
||||
```bash
|
||||
PGU=$(kubectl get deploy postgres -n foxhunt -o jsonpath='{.spec.template.spec.containers[0].env[?(@.name=="POSTGRES_USER")].value}')
|
||||
PGPOD=$(kubectl get pod -n foxhunt -l app=postgres -o jsonpath='{.items[0].metadata.name}' 2>/dev/null || kubectl get pod -n foxhunt | grep '^postgres-' | grep -v backup | awk '{print $1}' | head -1)
|
||||
echo "postgres user=$PGU pod=$PGPOD"
|
||||
```
|
||||
Expected: a username + a running pod name. (If `POSTGRES_USER` is via secretRef not value, read it from the referenced secret instead.)
|
||||
|
||||
- [ ] **Step 2: Generate a Gitea DB password + create role/db**
|
||||
```bash
|
||||
GITEA_DB_PW=$(openssl rand -base64 24 | tr -d '/+=' | head -c 24)
|
||||
kubectl exec -n foxhunt "$PGPOD" -- psql -U "$PGU" -v ON_ERROR_STOP=1 -c \
|
||||
"CREATE ROLE gitea LOGIN PASSWORD '$GITEA_DB_PW';" -c \
|
||||
"CREATE DATABASE gitea OWNER gitea ENCODING 'UTF8';" 2>&1 | grep -iE "CREATE|already exists|Error"
|
||||
echo "$GITEA_DB_PW" # capture for the next step
|
||||
```
|
||||
Expected: `CREATE ROLE` + `CREATE DATABASE` (or `already exists` if re-run).
|
||||
|
||||
- [ ] **Step 3: Create the k8s secrets Gitea needs** (DB creds + admin)
|
||||
```bash
|
||||
GITEA_ADMIN_PW=$(openssl rand -base64 24 | tr -d '/+=' | head -c 24)
|
||||
kubectl create secret generic gitea-db -n foxhunt \
|
||||
--from-literal=password="$GITEA_DB_PW" --dry-run=client -o yaml | kubectl apply -f -
|
||||
kubectl create secret generic gitea-admin -n foxhunt \
|
||||
--from-literal=username=gitadmin --from-literal=password="$GITEA_ADMIN_PW" \
|
||||
--from-literal=email=jeroen@bizworx.nl --dry-run=client -o yaml | kubectl apply -f -
|
||||
echo "ADMIN PW (save in your password manager): $GITEA_ADMIN_PW"
|
||||
```
|
||||
Expected: both secrets `created`/`configured`. **Save the admin password** — needed for Gitea login.
|
||||
|
||||
---
|
||||
|
||||
## Task 2: Gitea Helm values + install (internal-only)
|
||||
|
||||
**Files:** Create `infra/k8s/gitea/values.yaml` (foxhunt repo).
|
||||
|
||||
- [ ] **Step 1: Add the Gitea Helm repo**
|
||||
```bash
|
||||
helm repo add gitea-charts https://dl.gitea.com/charts/ && helm repo update gitea-charts
|
||||
helm search repo gitea-charts/gitea --versions | head -3
|
||||
```
|
||||
Expected: lists chart versions (use the latest stable in Step 3).
|
||||
|
||||
- [ ] **Step 2: Write `infra/k8s/gitea/values.yaml`**
|
||||
```yaml
|
||||
# Gitea — lightweight git host replacing GitLab (Phase 2B). External Postgres (existing in-cluster
|
||||
# `postgres`), no bundled DB/redis/memcached, Actions off (Argo does CI). Internal-only until cutover.
|
||||
replicaCount: 1
|
||||
image:
|
||||
rootless: true
|
||||
|
||||
# Disable all bundled subcharts — reuse the existing in-cluster postgres
|
||||
postgresql:
|
||||
enabled: false
|
||||
postgresql-ha:
|
||||
enabled: false
|
||||
redis-cluster:
|
||||
enabled: false
|
||||
redis:
|
||||
enabled: false
|
||||
|
||||
persistence:
|
||||
enabled: true
|
||||
size: 5Gi
|
||||
storageClass: sbs-default-retain
|
||||
|
||||
resources:
|
||||
requests: { cpu: 100m, memory: 128Mi }
|
||||
limits: { cpu: "1", memory: 512Mi }
|
||||
|
||||
service:
|
||||
http: { type: ClusterIP, port: 3000 }
|
||||
ssh: { type: ClusterIP, port: 22 }
|
||||
|
||||
actions:
|
||||
enabled: false
|
||||
|
||||
gitea:
|
||||
admin:
|
||||
existingSecret: gitea-admin
|
||||
config:
|
||||
server:
|
||||
ROOT_URL: https://git.fxhnt.ai/
|
||||
DOMAIN: git.fxhnt.ai
|
||||
SSH_DOMAIN: git.fxhnt.ai
|
||||
SSH_PORT: "22" # clean git@git.fxhnt.ai clone URLs. Port 22 is free (nothing host-SSHes
|
||||
# the git node); :2222 retired, local remotes updated in Task 5 Step 2b.
|
||||
DISABLE_SSH: "false"
|
||||
database:
|
||||
DB_TYPE: postgres
|
||||
HOST: postgres.foxhunt.svc.cluster.local:5432
|
||||
NAME: gitea
|
||||
USER: gitea
|
||||
service:
|
||||
DISABLE_REGISTRATION: "true"
|
||||
cache:
|
||||
ADAPTER: memory
|
||||
additionalConfigFromEnvs:
|
||||
- name: GITEA__database__PASSWD
|
||||
valueFrom:
|
||||
secretKeyRef: { name: gitea-db, key: password }
|
||||
```
|
||||
|
||||
- [ ] **Step 3: Install (internal-only — no public hostname yet)**
|
||||
```bash
|
||||
cd /home/jgrusewski/Work/foxhunt
|
||||
helm install gitea gitea-charts/gitea -n foxhunt -f infra/k8s/gitea/values.yaml --wait --timeout 5m 2>&1 | tail -5
|
||||
```
|
||||
Expected: `STATUS: deployed`. (If `--wait` times out, continue; verify pod in Step 4.)
|
||||
|
||||
- [ ] **Step 4: Verify Gitea is healthy**
|
||||
```bash
|
||||
kubectl get pods -n foxhunt | grep gitea
|
||||
kubectl run gitea-probe --rm -i --restart=Never -n foxhunt --image=curlimages/curl -- \
|
||||
curl -s http://gitea-http.foxhunt.svc.cluster.local:3000/api/healthz
|
||||
```
|
||||
Expected: a JSON health blob with `"status": "pass"`. Pod `Running`.
|
||||
|
||||
- [ ] **Step 5: Commit values (foxhunt repo)**
|
||||
```bash
|
||||
cd /home/jgrusewski/Work/foxhunt && git add infra/k8s/gitea/values.yaml
|
||||
git commit -m "feat(infra): Gitea Helm values (external postgres, internal-only) — Phase 2B"
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Task 3: Migrate both repos via port-forward + validate
|
||||
|
||||
**Files:** none committed (git data operations).
|
||||
|
||||
- [ ] **Step 1: Port-forward Gitea + create the repos via API**
|
||||
```bash
|
||||
kubectl port-forward -n foxhunt svc/gitea-http 3000:3000 >/tmp/gitea-pf.log 2>&1 & echo $! >/tmp/gitea-pf.pid
|
||||
sleep 3
|
||||
GA=gitadmin; GP=$(kubectl get secret gitea-admin -n foxhunt -o jsonpath='{.data.password}'|base64 -d)
|
||||
for r in fxhnt foxhunt; do
|
||||
curl -s -u "$GA:$GP" -X POST http://localhost:3000/api/v1/user/repos \
|
||||
-H 'Content-Type: application/json' -d "{\"name\":\"$r\",\"private\":true}" -o /dev/null -w "$r: %{http_code}\n"
|
||||
done
|
||||
```
|
||||
Expected: `fxhnt: 201` and `foxhunt: 201`.
|
||||
|
||||
- [ ] **Step 2: Mirror-push both repos from local clones**
|
||||
```bash
|
||||
GA=gitadmin; GP=$(kubectl get secret gitea-admin -n foxhunt -o jsonpath='{.data.password}'|base64 -d)
|
||||
( cd /home/jgrusewski/Work/fxhnt && git push --mirror "http://$GA:$GP@localhost:3000/$GA/fxhnt.git" 2>&1 | tail -2 )
|
||||
( cd /home/jgrusewski/Work/foxhunt && git push --mirror "http://$GA:$GP@localhost:3000/$GA/foxhunt.git" 2>&1 | tail -2 )
|
||||
```
|
||||
Expected: both report refs pushed (no error). foxhunt is 876 MB — may take a minute.
|
||||
|
||||
- [ ] **Step 3: VALIDATE — clone back + diff top SHAs vs GitLab**
|
||||
```bash
|
||||
GA=gitadmin; GP=$(kubectl get secret gitea-admin -n foxhunt -o jsonpath='{.data.password}'|base64 -d)
|
||||
for r in fxhnt foxhunt; do
|
||||
gitea_sha=$(git ls-remote "http://$GA:$GP@localhost:3000/$GA/$r.git" HEAD | awk '{print $1}')
|
||||
gitlab_sha=$(cd /home/jgrusewski/Work/$r && git ls-remote origin HEAD | awk '{print $1}')
|
||||
[ "$gitea_sha" = "$gitlab_sha" ] && echo "$r: MATCH ($gitea_sha)" || echo "$r: MISMATCH gitea=$gitea_sha gitlab=$gitlab_sha"
|
||||
done
|
||||
```
|
||||
Expected: `fxhnt: MATCH` and `foxhunt: MATCH`. **STOP if any MISMATCH.**
|
||||
|
||||
- [ ] **Step 4: Mark foxhunt archived (read-only)**
|
||||
```bash
|
||||
GA=gitadmin; GP=$(kubectl get secret gitea-admin -n foxhunt -o jsonpath='{.data.password}'|base64 -d)
|
||||
curl -s -u "$GA:$GP" -X PATCH "http://localhost:3000/api/v1/repos/$GA/foxhunt" \
|
||||
-H 'Content-Type: application/json' -d '{"archived":true}' -o /dev/null -w "archive foxhunt: %{http_code}\n"
|
||||
kill $(cat /tmp/gitea-pf.pid) 2>/dev/null
|
||||
```
|
||||
Expected: `archive foxhunt: 200`.
|
||||
|
||||
---
|
||||
|
||||
## Task 4: Move cockpit image to Scaleway Container Registry
|
||||
|
||||
**Files:** Modify `~/Work/fxhnt/infra/argo/cockpit-build-deploy.yaml`,
|
||||
`~/Work/fxhnt/infra/k8s/orchestration/dagster.yaml` (+ dashboard manifest if separate).
|
||||
|
||||
- [ ] **Step 1: Create the Scaleway registry pull/push secret**
|
||||
```bash
|
||||
kubectl create secret docker-registry scw-registry -n foxhunt \
|
||||
--docker-server=rg.fr-par.scw.cloud \
|
||||
--docker-username=nologin --docker-password="$SCW_SECRET_KEY" \
|
||||
--dry-run=client -o yaml | kubectl apply -f -
|
||||
```
|
||||
Expected: secret `created`/`configured`. (Scaleway registry auth: username `nologin`, password = a
|
||||
Scaleway secret key.)
|
||||
|
||||
- [ ] **Step 2: Repoint kaniko build → Scaleway registry**
|
||||
|
||||
In `~/Work/fxhnt/infra/argo/cockpit-build-deploy.yaml`, change the kaniko args:
|
||||
```
|
||||
--destination=rg.fr-par.scw.cloud/bizworx/fxhnt-cockpit:latest \
|
||||
--cache-repo=rg.fr-par.scw.cloud/bizworx/cache \
|
||||
```
|
||||
Remove the `--insecure-registry=…gitlab-registry…:5000` and `--skip-tls-verify-registry=…` lines
|
||||
(Scaleway is TLS). Mount the `scw-registry` dockerconfig for kaniko (replace the `gitlab-registry`
|
||||
dockerconfig volume): set the kaniko `DOCKER_CONFIG`/volume to the `scw-registry` secret at
|
||||
`/kaniko/.docker/config.json`.
|
||||
|
||||
- [ ] **Step 3: Repoint the deploy image refs**
|
||||
|
||||
In `~/Work/fxhnt/infra/k8s/orchestration/dagster.yaml` (and the dashboard manifest), change every
|
||||
`image: gitlab-registry.foxhunt.svc.cluster.local:5000/root/foxhunt/fxhnt-cockpit:latest` →
|
||||
`image: rg.fr-par.scw.cloud/bizworx/fxhnt-cockpit:latest`, and add under each pod spec:
|
||||
```yaml
|
||||
imagePullSecrets:
|
||||
- name: scw-registry
|
||||
```
|
||||
|
||||
- [ ] **Step 4: Run one cockpit build → Scaleway + verify the image exists**
|
||||
```bash
|
||||
cd /home/jgrusewski/Work/fxhnt && ./scripts/argo-deploy-cockpit.sh 2>&1 | tail -8
|
||||
# verify the pushed image
|
||||
scw registry image list namespace-id=6561a90a-f4ba-4f44-b43f-be6e27cbdfca 2>/dev/null | grep -i fxhnt-cockpit || echo "(verify via Scaleway console)"
|
||||
```
|
||||
Expected: workflow `Succeeded`; `fxhnt-cockpit` image listed in the `bizworx` namespace.
|
||||
|
||||
- [ ] **Step 5: Verify a rollout pulls from Scaleway**
|
||||
```bash
|
||||
kubectl rollout restart deploy/dagster deploy/fxhnt-dashboard -n foxhunt 2>/dev/null
|
||||
kubectl rollout status deploy/dagster -n foxhunt --timeout=180s
|
||||
kubectl get pods -n foxhunt -o jsonpath='{range .items[*]}{.spec.containers[*].image}{"\n"}{end}' | grep -i "rg.fr-par.scw.cloud/bizworx/fxhnt-cockpit" | head
|
||||
```
|
||||
Expected: rollout succeeds; pods now reference the `rg.fr-par.scw.cloud/bizworx/...` image (no `ImagePullBackOff`).
|
||||
|
||||
- [ ] **Step 6: Commit (fxhnt repo)**
|
||||
```bash
|
||||
cd /home/jgrusewski/Work/fxhnt
|
||||
git add infra/argo/cockpit-build-deploy.yaml infra/k8s/orchestration/dagster.yaml
|
||||
git commit -m "feat(infra): cockpit image -> Scaleway Container Registry (Phase 2B)"
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Task 5: Cut `git.fxhnt.ai` + SSH `:2222` over to Gitea
|
||||
|
||||
**Files:** Modify `infra/k8s/gitlab/tailscale-proxy.yaml` (foxhunt repo). **CHECKPOINT before applying.**
|
||||
|
||||
- [ ] **Step 1: Repoint nginx `git.fxhnt.ai:443` → Gitea**
|
||||
|
||||
In `infra/k8s/gitlab/tailscale-proxy.yaml`, in the `server_name git.fxhnt.ai;` (`listen 443`) block,
|
||||
change:
|
||||
```
|
||||
proxy_pass http://gitlab-webservice-default.foxhunt.svc.cluster.local:8181;
|
||||
```
|
||||
to:
|
||||
```
|
||||
proxy_pass http://gitea-http.foxhunt.svc.cluster.local:3000;
|
||||
```
|
||||
Delete the entire `listen 5050 ssl; server_name git.fxhnt.ai;` registry block (Scaleway registry is
|
||||
external — not proxied).
|
||||
|
||||
- [ ] **Step 2: Repoint the socat SSH proxy to Gitea SSH on `:22` (drop `:2222`)**
|
||||
|
||||
Nothing host-SSHes the git node, so use the standard port `22` and retire `:2222` entirely (clean
|
||||
`git@git.fxhnt.ai` URLs). In the `ssh-proxy` container args, change the listen port to `22` and the target
|
||||
to Gitea:
|
||||
```
|
||||
- "TCP-LISTEN:22,fork,reuseaddr,sndbuf=1048576,rcvbuf=1048576"
|
||||
- "TCP:gitea-ssh.foxhunt.svc.cluster.local:22,sndbuf=1048576,rcvbuf=1048576"
|
||||
```
|
||||
Update the container `ports:` entry from `containerPort: 2222` → `containerPort: 22`. Ensure the Tailscale
|
||||
node/proxy advertises port `22` (update the tailnet sidecar's exposed ports / `serve` config: `22`,
|
||||
replacing `2222`).
|
||||
|
||||
- [ ] **Step 2b: Update local git remotes from `:2222` → default (port 22)**
|
||||
```bash
|
||||
cd /home/jgrusewski/Work/foxhunt && git remote set-url origin ssh://git@git.fxhnt.ai/gitadmin/foxhunt.git
|
||||
cd /home/jgrusewski/Work/fxhnt && git remote set-url origin ssh://git@git.fxhnt.ai/gitadmin/fxhnt.git
|
||||
cd /home/jgrusewski/Work/foxhunt && git remote -v | grep fetch
|
||||
```
|
||||
Expected: remotes now show `ssh://git@git.fxhnt.ai/...` (no `:2222`).
|
||||
|
||||
- [ ] **Step 3: Apply + restart the proxy**
|
||||
```bash
|
||||
cd /home/jgrusewski/Work/foxhunt
|
||||
kubectl apply -f infra/k8s/gitlab/tailscale-proxy.yaml
|
||||
kubectl rollout restart deploy/tailscale-gitlab-proxy -n foxhunt
|
||||
kubectl rollout status deploy/tailscale-gitlab-proxy -n foxhunt --timeout=120s
|
||||
```
|
||||
Expected: rollout complete.
|
||||
|
||||
- [ ] **Step 4: GATE — verify `git.fxhnt.ai` serves Gitea**
|
||||
```bash
|
||||
curl -sk https://git.fxhnt.ai/api/healthz | head -c 200; echo
|
||||
echo "--- SSH on :22 (git@git.fxhnt.ai) ---"; git ls-remote ssh://git@git.fxhnt.ai/gitadmin/fxhnt.git HEAD 2>&1 | head -1
|
||||
```
|
||||
Expected: Gitea healthz JSON; the `:22` ssh ls-remote returns the HEAD sha (proves Gitea web + SSH on the
|
||||
canonical name). **STOP + rollback (revert this file, re-apply) if either fails.**
|
||||
|
||||
- [ ] **Step 5: Commit (foxhunt repo)**
|
||||
```bash
|
||||
cd /home/jgrusewski/Work/foxhunt && git add infra/k8s/gitlab/tailscale-proxy.yaml
|
||||
git commit -m "feat(infra): cut git.fxhnt.ai + SSH :2222 over to Gitea; drop :5050 registry (Phase 2B)"
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Task 6: Wire Gitea push webhook → Argo Events auto-deploy
|
||||
|
||||
**Files:** Create `infra/k8s/argo/events/gitea-push-eventsource.yaml` +
|
||||
`infra/k8s/argo/events/gitea-deploy-sensor.yaml` (foxhunt repo).
|
||||
|
||||
- [ ] **Step 1: Create the webhook eventsource**
|
||||
|
||||
Create `infra/k8s/argo/events/gitea-push-eventsource.yaml`:
|
||||
```yaml
|
||||
apiVersion: argoproj.io/v1alpha1
|
||||
kind: EventSource
|
||||
metadata:
|
||||
name: gitea-push
|
||||
namespace: foxhunt
|
||||
spec:
|
||||
service:
|
||||
ports:
|
||||
- port: 12000
|
||||
targetPort: 12000
|
||||
webhook:
|
||||
fxhnt-push:
|
||||
port: "12000"
|
||||
endpoint: /push
|
||||
method: POST
|
||||
```
|
||||
|
||||
- [ ] **Step 2: Create the sensor that submits the cockpit workflow**
|
||||
|
||||
Create `infra/k8s/argo/events/gitea-deploy-sensor.yaml`:
|
||||
```yaml
|
||||
apiVersion: argoproj.io/v1alpha1
|
||||
kind: Sensor
|
||||
metadata:
|
||||
name: gitea-deploy
|
||||
namespace: foxhunt
|
||||
spec:
|
||||
dependencies:
|
||||
- name: push
|
||||
eventSourceName: gitea-push
|
||||
eventName: fxhnt-push
|
||||
filters:
|
||||
data:
|
||||
- path: body.ref
|
||||
type: string
|
||||
value: ["refs/heads/main", "refs/heads/master"]
|
||||
triggers:
|
||||
- template:
|
||||
name: submit-cockpit
|
||||
argoWorkflow:
|
||||
operation: submit
|
||||
source:
|
||||
resource:
|
||||
apiVersion: argoproj.io/v1alpha1
|
||||
kind: Workflow
|
||||
metadata:
|
||||
generateName: fxhnt-cockpit-
|
||||
namespace: foxhunt
|
||||
spec:
|
||||
workflowTemplateRef:
|
||||
name: fxhnt-cockpit
|
||||
```
|
||||
|
||||
- [ ] **Step 3: Apply + register the webhook in Gitea**
|
||||
```bash
|
||||
cd /home/jgrusewski/Work/foxhunt
|
||||
kubectl apply -f infra/k8s/argo/events/gitea-push-eventsource.yaml -f infra/k8s/argo/events/gitea-deploy-sensor.yaml
|
||||
kubectl get pods -n foxhunt | grep -iE "gitea-push|gitea-deploy"
|
||||
# register webhook (in-cluster URL) on the fxhnt repo
|
||||
GA=gitadmin; GP=$(kubectl get secret gitea-admin -n foxhunt -o jsonpath='{.data.password}'|base64 -d)
|
||||
kubectl port-forward -n foxhunt svc/gitea-http 3000:3000 >/tmp/gitea-pf.log 2>&1 & echo $! >/tmp/gitea-pf.pid; sleep 3
|
||||
curl -s -u "$GA:$GP" -X POST "http://localhost:3000/api/v1/repos/$GA/fxhnt/hooks" -H 'Content-Type: application/json' -d '{
|
||||
"type":"gitea","active":true,"events":["push"],
|
||||
"config":{"url":"http://gitea-push-eventsource-svc.foxhunt.svc.cluster.local:12000/push","content_type":"json"}
|
||||
}' -o /dev/null -w "hook: %{http_code}\n"
|
||||
kill $(cat /tmp/gitea-pf.pid) 2>/dev/null
|
||||
```
|
||||
Expected: eventsource + sensor pods `Running`; `hook: 201`.
|
||||
|
||||
- [ ] **Step 4: GATE — test push triggers a deploy**
|
||||
```bash
|
||||
cd /home/jgrusewski/Work/fxhnt
|
||||
git commit --allow-empty -m "test: trigger Gitea->Argo cockpit deploy"
|
||||
git push origin main # origin already = git.fxhnt.ai (now Gitea)
|
||||
sleep 10
|
||||
kubectl get wf -n foxhunt --sort-by=.metadata.creationTimestamp 2>/dev/null | grep fxhnt-cockpit | tail -2
|
||||
```
|
||||
Expected: a new `fxhnt-cockpit-*` workflow appears (triggered by the push). **If none fires**, check the
|
||||
eventsource pod logs; fallback is the manual `./scripts/argo-deploy-cockpit.sh` (still works).
|
||||
|
||||
- [ ] **Step 5: Commit (foxhunt repo)**
|
||||
```bash
|
||||
cd /home/jgrusewski/Work/foxhunt && git add infra/k8s/argo/events/gitea-push-eventsource.yaml infra/k8s/argo/events/gitea-deploy-sensor.yaml
|
||||
git commit -m "feat(infra): Gitea push webhook -> Argo Events cockpit auto-deploy (Phase 2B)"
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Task 7: Final verification
|
||||
|
||||
- [ ] **Step 1: Footprint + health summary**
|
||||
```bash
|
||||
echo "=== Gitea (new) ==="; kubectl top pods -n foxhunt 2>/dev/null | grep -i gitea
|
||||
echo "=== cockpit healthy on Scaleway image ==="; curl -sk -o /dev/null -w "dashboard.fxhnt.ai: %{http_code}\n" https://dashboard.fxhnt.ai
|
||||
echo "=== git.fxhnt.ai serves Gitea ==="; curl -sk https://git.fxhnt.ai/api/healthz | head -c 80; echo
|
||||
echo "=== GitLab still running (fallback for 2C) ==="; kubectl get pods -n foxhunt | grep -c gitlab | xargs echo "gitlab pods:"
|
||||
```
|
||||
Expected: Gitea ≤~300 MiB; dashboard 200; git.fxhnt.ai healthz pass; GitLab pods still present.
|
||||
|
||||
---
|
||||
|
||||
## Rollback
|
||||
- **Cutover (Task 5):** `git checkout infra/k8s/gitlab/tailscale-proxy.yaml` + `kubectl apply` + rollout
|
||||
restart → `git.fxhnt.ai`/SSH back to GitLab.
|
||||
- **Registry (Task 4):** revert the fxhnt manifests + rollout restart → pods pull the GitLab image again
|
||||
(still present until 2C).
|
||||
- **Gitea:** `helm uninstall gitea -n foxhunt`; drop the `gitea` DB. GitLab was never touched.
|
||||
|
||||
---
|
||||
|
||||
## Acceptance criteria (from spec)
|
||||
- Gitea running on existing postgres, healthz pass, ≤~300 MiB. ✅ Task 2, Task 7
|
||||
- Both repos in Gitea, SHAs match GitLab, foxhunt archived. ✅ Task 3
|
||||
- Cockpit builds to + deploys from Scaleway registry; dashboard 200. ✅ Task 4, Task 7
|
||||
- `git.fxhnt.ai` (HTTPS + SSH :2222) serves Gitea; remotes unchanged. ✅ Task 5
|
||||
- Push to fxhnt main auto-triggers a cockpit deploy. ✅ Task 6
|
||||
- GitLab still running as fallback. ✅ Task 7
|
||||
```
|
||||
264
docs/superpowers/plans/2026-06-21-tfstate-to-scaleway.md
Normal file
264
docs/superpowers/plans/2026-06-21-tfstate-to-scaleway.md
Normal file
@@ -0,0 +1,264 @@
|
||||
# Migrate Terraform State to Scaleway Object Storage — Implementation Plan (Phase 2A)
|
||||
|
||||
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:executing-plans (INLINE, with checkpoints).
|
||||
> **Do NOT run subagent-driven** — state migration is irreversible-adjacent and each module needs a
|
||||
> human-confirmed `plan = No changes` gate. Steps use checkbox (`- [ ]`) syntax.
|
||||
|
||||
**Goal:** Move the 4 production modules' Terraform state from the self-hosted GitLab http backend to
|
||||
external Scaleway Object Storage (bucket `foxhunt-tfstate`), with native S3 lockfile locking.
|
||||
|
||||
**Architecture:** Back up all 4 states from GitLab first; flip the single `remote_state` block in
|
||||
`infra/live/production/root.hcl` from `http`→`s3`; migrate each module in-place with `terragrunt init
|
||||
-migrate-state`; gate each on `terragrunt plan = No changes`. GitLab state is left intact as a fallback
|
||||
(deleted later in 2C).
|
||||
|
||||
**Tech Stack:** terragrunt v0.77.20 + OpenTofu v1.11.5, Scaleway Object Storage (S3-compatible), AWS s3
|
||||
backend.
|
||||
|
||||
**Spec:** `docs/superpowers/specs/2026-06-21-tfstate-to-scaleway-design.md`
|
||||
|
||||
---
|
||||
|
||||
## Shared env (export in EVERY shell that runs terragrunt during this plan)
|
||||
The s3 backend reads **`AWS_*`** env vars; the Scaleway provider reads **`SCW_*`**; reading the OLD
|
||||
GitLab state during `-migrate-state` needs **`TF_HTTP_*`**. Export all three sets:
|
||||
```bash
|
||||
cd /home/jgrusewski/Work/foxhunt
|
||||
export TG_TF_PATH=tofu TERRAGRUNT_TFPATH=tofu
|
||||
# Scaleway provider creds
|
||||
export SCW_ACCESS_KEY=$(kubectl get secret scaleway-credentials -n foxhunt -o jsonpath='{.data.access-key}'|base64 -d)
|
||||
export SCW_SECRET_KEY=$(kubectl get secret scaleway-credentials -n foxhunt -o jsonpath='{.data.secret-key}'|base64 -d)
|
||||
export SCW_DEFAULT_PROJECT_ID=$(kubectl get secret scaleway-credentials -n foxhunt -o jsonpath='{.data.project-id}'|base64 -d)
|
||||
export SCW_DEFAULT_REGION=fr-par SCW_DEFAULT_ZONE=fr-par-2
|
||||
# s3 backend creds (SAME Scaleway keys, AWS-style names)
|
||||
export AWS_ACCESS_KEY_ID="$SCW_ACCESS_KEY"
|
||||
export AWS_SECRET_ACCESS_KEY="$SCW_SECRET_KEY"
|
||||
export AWS_REGION=fr-par
|
||||
# OLD GitLab http backend creds (needed to READ existing state during migrate)
|
||||
export TF_HTTP_USERNAME=root
|
||||
export TF_HTTP_PASSWORD=$(kubectl get secret gitlab-pat -n foxhunt -o jsonpath='{.data.token}'|base64 -d)
|
||||
echo "env: SCW=${SCW_ACCESS_KEY:0:6}… AWS=${AWS_ACCESS_KEY_ID:0:6}… PAT=${TF_HTTP_PASSWORD:0:4}…"
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Task 0: Branch
|
||||
- [ ] **Step 1**
|
||||
```bash
|
||||
cd /home/jgrusewski/Work/foxhunt && git checkout -b chore/tfstate-to-scaleway
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Task 1: Verify prerequisites + back up all 4 GitLab states
|
||||
|
||||
**Files:** none (read-only + /tmp backups)
|
||||
|
||||
- [ ] **Step 1: Confirm the target bucket exists + is empty-ish**
|
||||
```bash
|
||||
scw object bucket list 2>/dev/null | grep foxhunt-tfstate || echo "MISSING — create: scw object bucket create name=foxhunt-tfstate region=fr-par"
|
||||
```
|
||||
Expected: a line containing `foxhunt-tfstate`. If MISSING, run the create shown, then re-check.
|
||||
|
||||
- [ ] **Step 2: Confirm tooling versions**
|
||||
```bash
|
||||
tofu version | head -1; terragrunt --version | head -1
|
||||
```
|
||||
Expected: OpenTofu `v1.11.x` (≥1.10 for `use_lockfile`), terragrunt `v0.77.x`.
|
||||
|
||||
- [ ] **Step 3: Back up each module's CURRENT (GitLab) state** — run with shared env exported
|
||||
```bash
|
||||
for m in kapsule public-gateway dns block-storage; do
|
||||
( cd infra/live/production/$m && terragrunt state pull > /tmp/tfstate-backup-$m.tfstate 2>/dev/null )
|
||||
echo "$m: $(python3 -c "import json;d=json.load(open('/tmp/tfstate-backup-$m.tfstate'));print('serial',d.get('serial'),'resources',len(d.get('resources',[])))" 2>/dev/null || echo 'EMPTY/ERROR')"
|
||||
done
|
||||
```
|
||||
Expected: each module prints a non-zero `resources` count (e.g. kapsule ~5+, dns/public-gateway/block-storage ≥1). **STOP if any backup is EMPTY/ERROR** — that module's GitLab state isn't readable; do not migrate it blind.
|
||||
|
||||
- [ ] **Step 4: Record current state serials for post-migration comparison**
|
||||
```bash
|
||||
grep -H '"serial"' /tmp/tfstate-backup-*.tfstate 2>/dev/null || for m in kapsule public-gateway dns block-storage; do echo "$m $(python3 -c "import json;print(json.load(open('/tmp/tfstate-backup-$m.tfstate'))['serial'])")"; done
|
||||
```
|
||||
Expected: a serial number per module. Keep this output — the migrated state should have the same resource set.
|
||||
|
||||
---
|
||||
|
||||
## Task 2: Switch the backend in root.hcl (http → s3)
|
||||
|
||||
**Files:** Modify `infra/live/production/root.hcl`
|
||||
|
||||
- [ ] **Step 1: Replace the `remote_state` block**
|
||||
|
||||
In `infra/live/production/root.hcl`, replace the entire existing `remote_state { ... }` block (the `backend = "http"`
|
||||
GitLab block) with:
|
||||
```hcl
|
||||
# Remote state in Scaleway Object Storage (bucket foxhunt-tfstate, fr-par) — migrated off GitLab 2026-06-21.
|
||||
# s3 backend reads AWS_* env vars (set to the Scaleway access/secret keys); SCW_* drives the provider.
|
||||
remote_state {
|
||||
backend = "s3"
|
||||
generate = {
|
||||
path = "backend.tf"
|
||||
if_exists = "overwrite"
|
||||
}
|
||||
config = {
|
||||
bucket = "foxhunt-tfstate"
|
||||
key = "${path_relative_to_include()}/terraform.tfstate"
|
||||
region = "fr-par"
|
||||
|
||||
endpoints = {
|
||||
s3 = "https://s3.fr-par.scw.cloud"
|
||||
}
|
||||
|
||||
# Scaleway S3-compat: skip AWS-specific preflight calls
|
||||
skip_credentials_validation = true
|
||||
skip_region_validation = true
|
||||
skip_requesting_account_id = true
|
||||
skip_metadata_api_check = true
|
||||
|
||||
# OpenTofu-native lock (no DynamoDB)
|
||||
use_lockfile = true
|
||||
}
|
||||
}
|
||||
```
|
||||
Leave the `generate "provider"` block and `inputs`/`locals` untouched.
|
||||
|
||||
- [ ] **Step 2: Sanity-check the file parses**
|
||||
```bash
|
||||
cd /home/jgrusewski/Work/foxhunt
|
||||
grep -A2 'backend = "s3"' infra/live/production/root.hcl && grep -c 'backend = "http"' infra/live/production/root.hcl
|
||||
```
|
||||
Expected: shows the `s3` backend line; the `backend = "http"` count is `0`.
|
||||
|
||||
---
|
||||
|
||||
## Task 3: Migrate `kapsule` (canary) + lock smoke test
|
||||
|
||||
**Files:** none (state operation)
|
||||
|
||||
- [ ] **Step 1: Migrate state (with shared env exported)**
|
||||
```bash
|
||||
cd /home/jgrusewski/Work/foxhunt/infra/live/production/kapsule
|
||||
terragrunt init -migrate-state -force-copy -no-color 2>&1 | sed 's/.*tofu: //' | grep -iE "Successfully configured|migrat|Error|s3" | head
|
||||
```
|
||||
Expected: a line like `Terraform has been successfully migrated to the "s3" backend!` (or `Successfully
|
||||
configured the backend "s3"`). **If it errors on checksum** (Scaleway quirk), add `skip_s3_checksum = true`
|
||||
to the `config` block in `infra/live/production/root.hcl` and re-run this step.
|
||||
|
||||
- [ ] **Step 2: GATE — plan must show No changes**
|
||||
```bash
|
||||
terragrunt plan -input=false -no-color 2>&1 | sed 's/.*tofu: //' | grep -iE "No changes|^Plan:|Error" | head
|
||||
```
|
||||
Expected: `No changes. Your infrastructure matches the configuration.`
|
||||
**STOP if it shows any add/change/destroy** — the migrated state doesn't match reality. Revert (Task 7
|
||||
rollback) and investigate before continuing.
|
||||
|
||||
- [ ] **Step 3: Confirm the state object landed in the bucket**
|
||||
```bash
|
||||
scw object bucket list 2>/dev/null >/dev/null # ensure scw configured
|
||||
aws --endpoint-url https://s3.fr-par.scw.cloud s3 ls s3://foxhunt-tfstate/kapsule/ 2>/dev/null || \
|
||||
echo "(aws cli n/a — verify via: mc ls on a scw alias, or scw object)"
|
||||
```
|
||||
Expected: `terraform.tfstate` listed under `kapsule/`. (If `aws` CLI is absent, this is best-effort;
|
||||
the `plan = No changes` in Step 2 already proves the s3 backend is the live source.)
|
||||
|
||||
- [ ] **Step 4: Lock smoke test** — confirm `use_lockfile` works on Scaleway
|
||||
```bash
|
||||
cd /home/jgrusewski/Work/foxhunt/infra/live/production/kapsule
|
||||
# A refresh acquires + releases the lock; success + no leftover .tflock proves locking works
|
||||
terragrunt apply -refresh-only -auto-approve -input=false -no-color 2>&1 | sed 's/.*tofu: //' | grep -iE "Apply complete|Error|lock" | head
|
||||
aws --endpoint-url https://s3.fr-par.scw.cloud s3 ls s3://foxhunt-tfstate/kapsule/ 2>/dev/null | grep -i tflock && echo "WARN: orphan lockfile" || echo "OK: no orphan lockfile"
|
||||
```
|
||||
Expected: `Apply complete!` and `OK: no orphan lockfile`. **If locking errors**, fall back to
|
||||
`use_lockfile = false` (single-maintainer is safe) — document the change in the commit.
|
||||
|
||||
---
|
||||
|
||||
## Task 4: Migrate the remaining 3 modules (`public-gateway`, `dns`, `block-storage`)
|
||||
|
||||
**Files:** none (state operations). Same procedure as Task 3 Steps 1-2, per module, each gated.
|
||||
|
||||
- [ ] **Step 1: Migrate + gate `public-gateway`**
|
||||
```bash
|
||||
cd /home/jgrusewski/Work/foxhunt/infra/live/production/public-gateway
|
||||
terragrunt init -migrate-state -force-copy -no-color 2>&1 | sed 's/.*tofu: //' | grep -iE "migrat|successfully|Error" | head
|
||||
terragrunt plan -input=false -no-color 2>&1 | sed 's/.*tofu: //' | grep -iE "No changes|^Plan:|Error" | head
|
||||
```
|
||||
Expected: migration success line, then `No changes`. **STOP on any drift.**
|
||||
|
||||
- [ ] **Step 2: Migrate + gate `dns`**
|
||||
```bash
|
||||
cd /home/jgrusewski/Work/foxhunt/infra/live/production/dns
|
||||
terragrunt init -migrate-state -force-copy -no-color 2>&1 | sed 's/.*tofu: //' | grep -iE "migrat|successfully|Error" | head
|
||||
terragrunt plan -input=false -no-color 2>&1 | sed 's/.*tofu: //' | grep -iE "No changes|^Plan:|Error" | head
|
||||
```
|
||||
Expected: migration success line, then `No changes`. **STOP on any drift.**
|
||||
|
||||
- [ ] **Step 3: Migrate + gate `block-storage`**
|
||||
```bash
|
||||
cd /home/jgrusewski/Work/foxhunt/infra/live/production/block-storage
|
||||
terragrunt init -migrate-state -force-copy -no-color 2>&1 | sed 's/.*tofu: //' | grep -iE "migrat|successfully|Error" | head
|
||||
terragrunt plan -input=false -no-color 2>&1 | sed 's/.*tofu: //' | grep -iE "No changes|^Plan:|Error" | head
|
||||
```
|
||||
Expected: migration success line, then `No changes`. **STOP on any drift.**
|
||||
|
||||
---
|
||||
|
||||
## Task 5: Final verification + commit
|
||||
|
||||
**Files:** Modify (commit) `infra/live/production/root.hcl`; also commit generated `backend.tf` only if not gitignored.
|
||||
|
||||
- [ ] **Step 1: Re-plan all 4 modules from a clean shell** (proves the s3 backend is authoritative)
|
||||
```bash
|
||||
cd /home/jgrusewski/Work/foxhunt
|
||||
for m in kapsule public-gateway dns block-storage; do
|
||||
echo "=== $m ==="
|
||||
( cd infra/live/production/$m && terragrunt plan -input=false -no-color 2>&1 | sed 's/.*tofu: //' | grep -iE "No changes|^Plan:|Error" | head -1 )
|
||||
done
|
||||
```
|
||||
Expected: every module prints `No changes.`
|
||||
|
||||
- [ ] **Step 2: Confirm all 4 state objects exist in the bucket**
|
||||
```bash
|
||||
for m in kapsule public-gateway dns block-storage; do
|
||||
aws --endpoint-url https://s3.fr-par.scw.cloud s3 ls s3://foxhunt-tfstate/$m/ 2>/dev/null | grep -q terraform.tfstate && echo "$m: state present" || echo "$m: NOT FOUND (verify manually)"
|
||||
done
|
||||
```
|
||||
Expected: `state present` for all 4 (or manual confirmation if `aws` CLI absent — `plan = No changes` is the real proof).
|
||||
|
||||
- [ ] **Step 3: Verify GitLab state still intact (fallback preserved)** — should NOT be deleted in 2A
|
||||
```bash
|
||||
echo "GitLab TF state is intentionally left in place as a 2A fallback; deleted in Phase 2C."
|
||||
```
|
||||
|
||||
- [ ] **Step 4: Commit**
|
||||
```bash
|
||||
cd /home/jgrusewski/Work/foxhunt
|
||||
git add infra/live/production/root.hcl
|
||||
# backend.tf is terragrunt-generated; add only if tracked (usually gitignored)
|
||||
git status --porcelain infra/live/production/*/backend.tf 2>/dev/null
|
||||
git commit -m "chore(infra): migrate Terraform state GitLab -> Scaleway Object Storage (Phase 2A)
|
||||
|
||||
All 4 modules (kapsule/public-gateway/dns/block-storage) migrated to s3 backend
|
||||
(bucket foxhunt-tfstate, fr-par, use_lockfile). plan=No changes verified per module.
|
||||
GitLab http state left intact as fallback (removed in 2C).
|
||||
|
||||
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>"
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Rollback (if any module's post-migration plan shows drift)
|
||||
1. Restore the `http` backend block in `infra/live/production/root.hcl` (git checkout the file).
|
||||
2. Re-attach GitLab state: in each affected module, `terragrunt init -migrate-state -force-copy` (copies
|
||||
back s3→http) OR `terragrunt init -reconfigure` (the GitLab state was never deleted).
|
||||
3. If state is corrupted, restore from `/tmp/tfstate-backup-<module>.tfstate` via `terragrunt state push`.
|
||||
4. The GitLab copy is untouched throughout 2A, so reverting is always possible.
|
||||
|
||||
---
|
||||
|
||||
## Acceptance criteria (from spec)
|
||||
- `infra/live/production/root.hcl` uses the Scaleway s3 backend; committed. ✅ Task 2 + Task 5.4
|
||||
- All 4 modules: `terragrunt plan = No changes` on the new backend. ✅ Task 3.2, Task 4, Task 5.1
|
||||
- State objects at `foxhunt-tfstate/<module>/terraform.tfstate`. ✅ Task 3.3, Task 5.2
|
||||
- Lock smoke test passes. ✅ Task 3.4
|
||||
- GitLab http state intact (fallback). ✅ Task 5.3
|
||||
@@ -0,0 +1,144 @@
|
||||
# The Surfer over an Uncorrelated Universe — Design Spec
|
||||
|
||||
> **For agentic workers:** This spec defines WHY and WHAT for foxhunt's strategic pivot from
|
||||
> single-instrument seconds-horizon order-book RL to a diversified, days-to-weeks, multi-asset
|
||||
> ML trend system. It does NOT implement; an implementation plan follows (Phase 0 first).
|
||||
|
||||
**Status:** Draft 1 (2026-06-05)
|
||||
**Supersedes (direction):** the seconds-horizon ES order-book RL thesis (measured structurally unprofitable).
|
||||
**Branch target:** new branch `surfer-universe` off the current branch (keeps the session's validation harnesses).
|
||||
**Capital stance:** $35k is a PILOT to prove the pipeline + edge (paper / tiny live); real book targets ~$100k.
|
||||
|
||||
**Linked pearls (load-bearing):**
|
||||
- `pearl_surfer_thesis_is_uncorrelated_trend_recognition` — the thesis
|
||||
- `pearl_surfer_universe_build_synthesis` — the 8-agent research synthesis behind every choice here
|
||||
- `pearl_ofi_edge_uncapturable_by_crossing` / `pearl_passive_mm_knife_edge_ofi_conditioning_promising` — why seconds-horizon ES is dead
|
||||
- `pearl_lowfreq_no_robust_edge_on_2y_es` — why breadth/universe is the bottleneck; validation discipline
|
||||
- `pearl_omnisearch_synthesis_signal_first_and_mtm_reward` — dense-MtM → portfolio Differential Sharpe reward
|
||||
|
||||
---
|
||||
|
||||
## §1. Why this, why now
|
||||
|
||||
Measured this session (zero cluster runs, ~5 cheap harnesses): on ES, seconds-horizon crossing edge is ~100× smaller
|
||||
than the spread; passive MM is adversely-selected negative (winner's curse); intraday has no IC; daily/cross-asset
|
||||
edges sign-flip IS↔OOS (overfit). **The bottleneck was the configuration — single instrument × seconds-horizon ×
|
||||
order-book microstructure — not the model.** That is the worst possible config for a non-colocated ML participant:
|
||||
most-efficient market, most speed-dependent horizon, zero diversification.
|
||||
|
||||
The fix is to invert all three: **many uncorrelated instruments × days-to-weeks horizon × trend/cross-sectional
|
||||
recognition.** This is also foxhunt's actual "surfer" thesis ("no wave is the same") finally pointed at an ocean
|
||||
with enough waves to ride. Breadth = opportunity sourcing: across many uncorrelated markets, some subset always
|
||||
has a rideable trend, and uncorrelated rides don't sink together.
|
||||
|
||||
## §2. The thesis
|
||||
|
||||
A fixed trend rule applies identical logic to every wave. The ML "surfer" reads each unique wave — is it a real
|
||||
trend or chop? building or exhausting? what's the sea-state (regime)? — and adjusts. That recognition is the
|
||||
legitimate ML edge ON TOP of the trend risk premium. **Breadth gives the floor; ML earns its keep above it.**
|
||||
|
||||
## §3. Architecture
|
||||
|
||||
**Core design principle (the eval-collapse cure):** the ML surfer outputs **bounded DEVIATIONS from a deterministic
|
||||
diversified-trend floor**: `w_i = floor_i + clamp(ML_delta_i, ±δ)`. The model can only add value above a proven
|
||||
baseline and degrades gracefully to the floor when it has no edge. It structurally cannot run off a cliff.
|
||||
|
||||
### §3.1 The floor (deterministic)
|
||||
Diversified time-series momentum: `signal_i = mean over h∈{1,3,12}mo of sign(return_{i,h})` (12mo skips last ~21d);
|
||||
inverse-vol size with 63-day EWMA vol; portfolio vol-target 10% annual (rescale weekly on realized portfolio vol);
|
||||
weekly rebalance + ≥10%-change no-trade band. Realistic net Sharpe 0.4–0.5 (small universe), DD 15–25%.
|
||||
|
||||
### §3.2 The ML surfer (bounded deltas), in confidence order
|
||||
1. **Regime / change-point overlay (build FIRST)** — recognizes trend-vs-chop, throttles gross, cuts drawdowns.
|
||||
ML's most defensible edge (DD reduction, NOT return prediction). Lowest overfit surface.
|
||||
2. **Cross-sectional tilt (panel)** — ONE shared encoder over all instruments (N×T samples) + per-instrument
|
||||
embedding → continuous dollar-neutral tilt on the floor; gated on rank-IC / ICIR. NOT discrete ranks (a <20-name
|
||||
universe is statistically starved; breadth is recovered through time + pooling).
|
||||
3. **Direct return prediction → DEFERRED** (the 64-commit trap; cost-cliff at 2–3 bps = turnover noise).
|
||||
|
||||
### §3.3 Reward / objective
|
||||
Portfolio **Differential Sharpe Ratio** (Moody-Saffell) + drawdown penalty + turnover/cost penalty — the
|
||||
portfolio-level, risk-adjusted generalization of the dense per-step mark-to-market reward (Φ=unrealized-PnL).
|
||||
|
||||
### §3.4 Risk / sizing
|
||||
Vol-target 10%; inverse-vol within asset class + ERC across classes; fractional Kelly (half/quarter, reuse the
|
||||
existing Kelly controller clamped [0,0.5]); drawdown throttle θ; reuse the CMDP overlay as the tail kill-switch.
|
||||
Discreteness is binding at micro size: target <0.5 contract → hold 0, carry the residual; $35k sustains ~3–5
|
||||
simultaneous 1-lot positions.
|
||||
|
||||
### §3.5 Reuse of existing foxhunt assets + architecture discipline
|
||||
CfC/Mamba2 encoder (shared across instruments + instrument embedding + cross-sectional attention head); RL stack;
|
||||
Kelly + CMDP controllers; the determinism foundation; ml-alpha's CUDA build pipeline (build.rs cubins, mapped-pinned
|
||||
buffers, cudarc launch). **All compute is CUDA-only, CPU read-only** (`feedback_cpu_is_read_only`,
|
||||
`pearl_cold_path_no_exception_to_gpu_drives`, `feedback_no_cpu_test_fallbacks`): every formula — signals, vol,
|
||||
sizing, backtest returns, and the validation statistics (CPCV per-path Sharpe, PBO, Deflated Sharpe) — is a CUDA
|
||||
kernel; CPU only reads final gate scalars from mapped-pinned buffers and only enumerates split index-sets as control
|
||||
flow. No CPU compute, no cold-path exception, no CPU test oracle. The session's Python `scripts/measure_*.py` were
|
||||
one-off edge AUDITS, not part of the built system; the system is Rust+CUDA. The new surfer lives in `crates/ml-alpha`
|
||||
(reuses its build.rs/cuda/mapped-pinned/determinism infra and, in Phases 1-3, its encoder + RL).
|
||||
|
||||
## §4. Universe + data pipeline
|
||||
|
||||
**Universe (~8 micros, maximal asset-class diversification):** MES, M2K (equity); 10Y micro-yield (rates — note
|
||||
YIELD-based = sign inversion vs price); M6E, M6A (FX); MGC (metals); MCL (energy); MBT (crypto).
|
||||
**Data:** Databento `GLBX.MDP3`, `ohlcv-1d` (+ `ohlcv-1h`). **Continuous contracts built in-house:** volume-roll
|
||||
(validate vs OI) + **backward RATIO adjustment for the encoder's returns**; keep an **UNADJUSTED** series for
|
||||
actual fills + USD reward. NEVER Panama-adjust for a trend learner; NEVER compute reward on adjusted prices.
|
||||
|
||||
## §5. Validation gates (the guardrail — non-negotiable)
|
||||
|
||||
| Gate | Metric | Pass |
|
||||
|---|---|---|
|
||||
| SV-G0 | Leakage audit: purge+embargo covers full hold + feature lookback | verified |
|
||||
| SV-G1 | CPCV (≥50 paths): 5th-percentile OOS pnl after costs | > 0 |
|
||||
| SV-G2 | PBO (CSCV) | < 0.20 |
|
||||
| SV-G3 | Deflated Sharpe, deflated by N = ALL trials ever (64-commit + session + this build) | > 0.95 |
|
||||
| SV-G4 | IS↔OOS config rank-consistency (Spearman) — the sign-flip detector | > 0 |
|
||||
| SV-G5 | Edge survives 2× cost + roll cost | net OOS pnl > 0 |
|
||||
| SV-G6 | Capacity/turnover: cost × turnover < fixed fraction of gross | pass |
|
||||
|
||||
Every config / seed / hyperparameter is a logged trial feeding N. The 64 failed commits prove N is large.
|
||||
|
||||
## §6. Phases (signal-first, cheapest-first, zero cluster until validated)
|
||||
|
||||
- **Phase 0 — Floor + validation harness (CUDA kernels + Rust orchestration, CPU read-only, no ML, no cluster — per foxhunt GPU-drives discipline; `feedback_cpu_is_read_only`, `pearl_cold_path_no_exception_to_gpu_drives`). Runs on the local GPU. FULLY SPECIFIED in the plan.**
|
||||
Build the deterministic floor on real Databento micro data; build the CPCV/PBO/DSR harness; measure whether the
|
||||
floor itself clears the gates OOS after costs. **STOP-if:** floor fails SV-G1/G3/G5 → the universe has no
|
||||
capturable trend premium at this scale; reassess (more instruments / more history) before any ML.
|
||||
- **Phase 1 — Regime overlay.** Must beat the floor OOS (DD reduction first). Gates SV-G1..G6 vs floor.
|
||||
- **Phase 2 — Cross-sectional panel tilt** (shared encoder). Must beat floor+regime OOS via rank-IC/ICIR.
|
||||
- **Phase 3 — Full RL surfer** (floor + bounded delta, DSR reward). Cluster only here, gated by all above.
|
||||
|
||||
## §7. Scope boundary
|
||||
|
||||
**In (this spec):** the floor, the bounded-delta architecture, the 3 ML layers (regime → cross-sectional → RL),
|
||||
the universe + continuous-contract pipeline, the validation gates, the phased build. **Phase 0 is the immediate
|
||||
deliverable.**
|
||||
**Out:** seconds-horizon / order-book microstructure (measured dead); direct return-prediction sequence models
|
||||
(deferred until floor+regime+cross-sectional prove out); live-capital deployment beyond a pilot (until ~$100k and
|
||||
gates pass); broker/execution automation (separate spec once a validated strategy exists).
|
||||
|
||||
## §8. Risks
|
||||
|
||||
1. **Capital ($35k below viable minimum)** — under-diversified → trend's drawdowns without its smoothing; micro-cost
|
||||
drag eats Sharpe. MITIGATION: treat as a pilot to prove pipeline/edge; scale to ~$100k for the real book.
|
||||
2. **ML adds nothing over the floor** — entirely possible; the bounded-delta design makes that a graceful
|
||||
no-op (you keep the floor), and the gates catch it before any spend.
|
||||
3. **Floor itself has no edge at this scale/universe** — Phase 0 STOP catches this cheaply.
|
||||
4. **Re-overfitting (commit #65)** — the SV gates + deflation-by-all-trials are the explicit defense; honor them.
|
||||
5. **Roll / continuous-contract artifacts** — ratio-adjust + volume-roll + reward-on-unadjusted; validated in Phase 0.
|
||||
|
||||
## §9. Decision log
|
||||
|
||||
```
|
||||
Decision: Pivot foxhunt to "the surfer over an uncorrelated universe" — diversified days-to-weeks ML trend across
|
||||
~8 uncorrelated micro futures, ML as bounded deviations from a deterministic trend floor, gated by CPCV/PBO/Deflated-
|
||||
Sharpe. Phase 0 (build+validate the floor, CPU-only) is the immediate, decisive, near-free test.
|
||||
Rejected: continuing seconds-horizon ES order-book RL (measured structurally unprofitable for a non-colocated ML
|
||||
participant — crossing edge 100× < spread; passive MM adversely-selected; daily/cross-asset edges sign-flip).
|
||||
Capital: $35k = pilot to prove the approach; real book ~$100k (Carver retail floor; under-diversification = main risk).
|
||||
ML scope: regime overlay first (DD reduction = ML's defensible edge); cross-sectional tilt second; direct return
|
||||
prediction deferred (the 64-commit trap).
|
||||
Date: 2026-06-05
|
||||
Falsification (Phase 0): floor fails SV-G1/G3/G5 OOS → no capturable trend premium at this scale; reassess before ML.
|
||||
```
|
||||
@@ -0,0 +1,226 @@
|
||||
# Adaptive Multi-Strat Book — Deployable Spec
|
||||
|
||||
**Date:** 2026-06-07
|
||||
**Status:** Validated design, pre-deployment (paper-forward running)
|
||||
**Branch:** surfer-universe
|
||||
**Engine:** `scripts/surfer/multistrat_book_v3.py` (validated), `sixtyforty_paper.py` (baseline, live)
|
||||
|
||||
---
|
||||
|
||||
## 1. Executive summary
|
||||
|
||||
A **diversified, edge-decay-adaptive, adaptively-risk-managed, unlevered** multi-asset book — the
|
||||
honest endpoint of an exhaustive search. It is *not* alpha; it is **premium harvesting with superior
|
||||
risk management** (the foxhunt engine's genuine, validated strength). Realistic **~0.5–0.7 Sharpe,
|
||||
~−7% max drawdown**, ~$2–3k/yr on $35k, scales linearly with capital.
|
||||
|
||||
**Why this and not the alternatives** (all tested to OOS exhaustion, all in memory):
|
||||
- Single predictive edges (equities, futures, ML, AI4Finance, PEAD): no OOS alpha for non-colocated retail.
|
||||
- Crypto funding (single-venue): hedgeability-blocked. Cross-venue: marginal/breakeven OOS (the +14.9 Sharpe was a max-min artifact; honest fixed-pair = ~breakeven).
|
||||
- Leverage: **doesn't pay at retail financing** (6–7% margin vs prime-brokerage SOFR+1–2%) — adaptive 1x Sharpe +0.48 vs 2x +0.14. The hedge-fund moat is *cheap financing*, not the strategy.
|
||||
|
||||
**What we keep** (validated, measurable): the **edge-decay-adaptive allocation** (+0.11 Sharpe, lower
|
||||
DD vs static risk-parity) and the **adaptive risk layer** (beat the static one decisively; halves
|
||||
drawdown). The engine's value is *risk management + adaptive allocation*, and that is real.
|
||||
|
||||
---
|
||||
|
||||
## 2. Strategy overview
|
||||
|
||||
Hold a small set of **uncorrelated return streams**, each vol-normalized to equal risk, weighted by
|
||||
a **per-stream edge-health trust** (down-weight decaying streams, resurrect recovered ones), combined
|
||||
and run through an **adaptive risk-management layer**, **unlevered (~1x)**, rebalanced on a schedule.
|
||||
|
||||
Return source = the structural premia (equity, term, gold/inflation, commodity, trend/crisis-alpha,
|
||||
a small crypto sleeve). The *alpha* is the disciplined, adaptive risk allocation — not prediction.
|
||||
|
||||
---
|
||||
|
||||
## 3. Instruments (deployable at $35k via fractional-share ETFs + small crypto)
|
||||
|
||||
| Stream | ETF | Role |
|
||||
|---|---|---|
|
||||
| Equity | **SPY** (or VTI) | equity risk premium |
|
||||
| Bonds | **IEF** (7–10y Treasuries) | term premium, equity diversifier |
|
||||
| Gold | **GLD** (or IAU) | inflation/crisis diversifier |
|
||||
| Commodity | **PDBC** (or DBC) | inflation/commodity premium |
|
||||
| Trend / managed futures | **DBMF** (or KMLM) | crisis-alpha, uncorrelated (use the *real* CTA ETF — our DIY trend was ~0 Sharpe) |
|
||||
| Crypto (small) | **BTC spot** (or IBIT) | high-return uncorrelated sleeve; trust-layer auto-caps it (currently down-weighted in deleverage) |
|
||||
|
||||
All ETFs are marginable, fractional-shareable, ~zero commission. **Unlevered** (cash account is fine;
|
||||
no margin needed). Crypto held separately (spot/IBIT), sized small.
|
||||
|
||||
---
|
||||
|
||||
## 4. Allocation engine — edge-decay-adaptive trust (foxhunt idea, validated)
|
||||
|
||||
For each stream `i`, maintain a **trust** `θᵢ ∈ [0.1, 1]`:
|
||||
- Compute trailing-126d risk-adjusted return (Sharpe) of the stream.
|
||||
- Map: Sharpe ≥ +0.5 → θ=1.0; ≤ −0.5 → θ=0.1 (floored — **resurrection-capable**); linear between.
|
||||
- EMA-smooth (α≈0.06) so it adapts gradually, not jumpily.
|
||||
|
||||
Weight each (vol-normalized) stream by `θᵢ`, normalize to sum 1. This **down-weights decaying premia
|
||||
and re-weights recovered ones** automatically — currently: equity 0.73, commod 1.00, trend 0.79
|
||||
(healthy) vs bond 0.11, crypto 0.10 (correctly de-emphasized). Validated: +0.11 Sharpe, lower DD vs
|
||||
static equal-risk. Source: `pearl_edge_decay_detection_is_a_missing_abstraction_layer`,
|
||||
`pearl_dead_signal_resurrection_discipline`.
|
||||
|
||||
---
|
||||
|
||||
## 5. Risk-management layer — adaptive controllers (the foxhunt port; the load-bearing piece)
|
||||
|
||||
Static thresholds **crush returns** (one-way latch, `pearl_cmdp_consec_loss_counter_is_one_way_latch`).
|
||||
The adaptive layer (validated: beat static +0.03→+0.14 Sharpe, maxDD −18.7%→−14.5%). Daily leverage
|
||||
`L = clip( min(L_vol, L_kelly) × dd_mult × corr_mult , LEV_FLOOR, MAXLEV )`, applied to the combination:
|
||||
|
||||
1. **EMA online vol** (α≈0.03) → `L_vol = TARGET_VOL / realized_vol` (responsive, no window edges).
|
||||
2. **Kelly with floor + bootstrap** → `L_kelly = clip(EMA_mean·252 / EMA_var·252, KELLY_FLOOR=0.5, MAXLEV)` — adapts to edge strength, **never dies** (`pearl_bootstrap_must_respect_clamp_range`).
|
||||
3. **Drawdown de-lever, CONTINUOUS + self-recovering** → `dd_mult = clip(1 − 3·max(0, −dd − 0.05), 0.40, 1.0)` — scales down as drawdown deepens, **recovers immediately as it heals** (NOT a latch — this is the fix).
|
||||
4. **Correlation de-risk, z-scored** → when avg cross-stream correlation is unusually high vs its own trailing distribution (diversification breaking in a crisis), reduce: `corr_mult = clip(1 − 0.2·max(0, z), 0.5, 1.0)`.
|
||||
5. **Leverage floor** `LEV_FLOOR=0.3`, **MAXLEV=1.0** (unlevered — see §8).
|
||||
|
||||
Defaults: `TARGET_VOL=10%`. The whole layer is the engine's true job — *survive and adapt*, not maximize.
|
||||
|
||||
---
|
||||
|
||||
## 6. Rebalance & execution
|
||||
|
||||
- **Weekly rebalance** (fixed-phase): recompute vol-normalization, trust weights, and `L`; trade to target.
|
||||
- **Hysteresis / no-churn:** only trade a stream if its target weight moved > ~2% (cut turnover/cost).
|
||||
- **Cost budget:** ETFs ~0 commission; the only friction is spread/slippage (minimal on SPY/IEF/GLD/DBMF). Crypto sleeve: small, spot.
|
||||
- **Cash account, unlevered** — no financing cost, no liquidation risk.
|
||||
|
||||
---
|
||||
|
||||
## 7. Expected performance (honest)
|
||||
|
||||
**Validated** on the real ETFs (SPY/IEF/GLD/PDBC/DBMF + BTC) over 2019-05..2026-06, exact live
|
||||
pipeline (`multistrat_etf_backtest.py`), multiple regimes (COVID, 2022 bear, bulls):
|
||||
|
||||
| Metric | Backtest 2019–26 | Realistic forward |
|
||||
|---|---|---|
|
||||
| Sharpe (unlevered) | **+1.20** | ~0.8–1.0 (haircut: favorable period + low-vol-flatter) |
|
||||
| Annual return | +6.1% | ~5–7% |
|
||||
| Realized vol | 5.1% | ~5–8% |
|
||||
| Max drawdown | **−5.4%** | ~−8 to −12% |
|
||||
| Per-year | **positive every year** incl 2022 (+0.1) | — |
|
||||
| vs 60/40 (SPY/IEF) | +0.87 Sharpe, −21% DD | book wins on Sharpe AND drawdown |
|
||||
| vs equity buy-hold | +0.85 Sharpe, −34% DD | — |
|
||||
| On $35k | ~$2.1k/yr, very smooth (−5% DD) | — |
|
||||
| On $500k | ~$30k/yr (Sharpe scale-invariant — capital is the lever) | — |
|
||||
|
||||
The robust signals are the **−5.4% max drawdown and every-year-positive** (not just the low-vol
|
||||
Sharpe). The book genuinely beats 60/40 on risk-adjusted terms with ~1/4 the drawdown.
|
||||
|
||||
Honest caveats: it's **long beta** (falls in everything-down 2022-style, though trend + adaptive
|
||||
de-lever cushion it); returns are *modest* in dollars at small capital (the lever is capital, not
|
||||
Sharpe); crypto sleeve validated on one window — keep it small and let the trust layer cap it.
|
||||
|
||||
---
|
||||
|
||||
## 8. Why unlevered (the key finding)
|
||||
|
||||
Leverage scales return but **at retail financing (6–7%) it lowers risk-adjusted return**: adaptive 1x
|
||||
Sharpe +0.48 vs 2x +0.14 — the financing drag eats the leverage benefit, while drawdown grows. Hedge
|
||||
funds lever ~0.7-Sharpe books profitably **only because prime brokerage finances at ~SOFR+1–2%.** That
|
||||
financing access is the institutional moat, and it's structural — not replicable at retail. **So: run
|
||||
it 1x.** Re-evaluate leverage only if you ever access institutional-rate financing.
|
||||
|
||||
---
|
||||
|
||||
## 9. Validation status
|
||||
|
||||
**Validated:** edge-decay-adaptive allocation (>static), adaptive risk layer (>static, halves DD),
|
||||
combination Sharpe ~0.5–0.72 across windows, the unlevered-is-best financing finding. Scripts:
|
||||
`multistrat.py`, `multistrat_book.py`, `multistrat_book_v2.py`, `multistrat_book_v3.py`.
|
||||
**Live baseline:** `sixtyforty_paper.py` paper-forward running (the 60/40 core).
|
||||
**Not yet:** live execution of the full 6-stream adaptive book (paper-forward harness for it = next),
|
||||
real ETF fills/spreads (minor), the crypto sleeve operationally.
|
||||
|
||||
---
|
||||
|
||||
## 10. Phased rollout
|
||||
|
||||
1. **Paper-forward (4–8 wk):** extend the live harness to the full 6-stream adaptive book (ETF closes
|
||||
via Yahoo + BTC), logging target weights + realized — confirm the design forward. (60/40 core
|
||||
already runs via `sixtyforty_paper.py`.)
|
||||
2. **Micro-live (small):** deploy the ETF book unlevered at small size; validate rebalance discipline,
|
||||
spreads, the trust/risk layer operating live.
|
||||
3. **Scale:** full capital, same unlevered book. Returns scale with capital, Sharpe unchanged.
|
||||
|
||||
Gate each phase: forward Sharpe ~consistent, drawdown controlled, trust/risk layer behaving.
|
||||
|
||||
---
|
||||
|
||||
## 11. Go / no-go + risk limits
|
||||
|
||||
- **Go:** you accept a modest (~0.5–0.7 Sharpe, ~$2–3k/yr on $35k) but *robust, well-risk-managed,
|
||||
unlevered* diversified book whose return scales with capital — and that the engine's role is risk
|
||||
management, not alpha.
|
||||
- **No-go / reconsider:** if you need higher absolute return at $35k (then the lever is *more capital*,
|
||||
not more strategy), or if you want market-neutral (this is long-beta — accept the equity-correlated
|
||||
drawdowns).
|
||||
- **Risk limits:** unlevered (MAXLEV 1.0); crypto sleeve ≤ ~15% target (trust-capped); weekly rebalance;
|
||||
the adaptive layer de-risks automatically in vol spikes / correlation crises / drawdowns.
|
||||
|
||||
**The honest bottom line:** after exhausting every alpha avenue (incl. the engine's own advanced
|
||||
ideas), this is what's real and deployable — a simple, diversified, adaptively-risk-managed harvest.
|
||||
The engine earns its keep as the risk/allocation brain, not as an oracle. The path to *meaningful*
|
||||
money is capital, not a better signal.
|
||||
|
||||
---
|
||||
|
||||
## Appendix A — Micro-live runbook (Phase 2, concrete)
|
||||
|
||||
**Validation so far (historical, exhaustive):** in-sample Sharpe +1.20; holdout-OOS +1.41; 20-year
|
||||
(2006–26, conservative no-DBMF/crypto version) +0.96 surviving 2008/2020/2022 with −10% maxDD;
|
||||
bootstrap CI p5 +0.62 / median +0.99, P(Sharpe>0.5)=99%. The *only* thing history can't give is
|
||||
forward/live confirmation — that's what micro-live provides.
|
||||
|
||||
**Objective:** validate *execution* — fills, spreads, the weekly/monthly rebalance, the live adaptive
|
||||
weighting — and start the real-money clock. **NOT to make money** (on $1k the P&L is pennies). It
|
||||
de-risks the scale step; that's its whole job.
|
||||
|
||||
**Capital:** $500–1,000, fully at-risk, **unlevered**. Worst case ≈ the book's maxDD (~−7%) ≈ −$70.
|
||||
Trivial by design — that's the point of micro.
|
||||
|
||||
**Instruments (all in ONE brokerage, fractional shares):**
|
||||
| Stream | Ticker | Note |
|
||||
|---|---|---|
|
||||
| equity | SPY | |
|
||||
| bonds | IEF | |
|
||||
| gold | GLD | |
|
||||
| commodity | PDBC | |
|
||||
| trend / CTA | DBMF | the real managed-futures ETF (our DIY trend was ~0) |
|
||||
| crypto | **IBIT** | iShares Bitcoin ETF — use instead of BTC spot: in-brokerage, fractional, **NO crypto-exchange counterparty tail** |
|
||||
|
||||
Broker: any fractional-share, ~zero-commission (IBKR / Fidelity / Schwab / Robinhood). IBKR or
|
||||
Fidelity recommended (clean fractional fills).
|
||||
|
||||
**Step 1 — today's target:**
|
||||
```
|
||||
python3 scripts/surfer/multistrat_paper.py weights 1000
|
||||
```
|
||||
→ exact % and $ per instrument (the adaptive trust + risk-layer output). Map the crypto line → IBIT.
|
||||
|
||||
**Step 2 — place orders manually** (micro scale → eyeball the fills). Marketable-limit or market on
|
||||
these liquid ETFs. Record actual fill prices.
|
||||
|
||||
**Step 3 — rebalance MONTHLY (not weekly) at micro scale.** At $1k, weekly deltas are ~$3 —
|
||||
rounding-dominated and operationally silly. Rebalance monthly, and only trade an instrument if its
|
||||
target weight drifted >~5% (hysteresis). The adaptive layer's value shows over months, not days.
|
||||
|
||||
**Step 4 — track actual vs intended.** Each rebalance, log the harness's intended weights vs your
|
||||
actual fills → measure (a) tracking error, (b) realized cost (spread/slippage), (c) live-vs-paper match.
|
||||
|
||||
**Gate to scale (~4–8 weeks):**
|
||||
- Live cumulative ≈ paper-forward cumulative (small tracking error).
|
||||
- Realized cost ≤ ~10–20 bp/rebalance (should be tiny on these ETFs).
|
||||
- Rebalance executes cleanly; adaptive weights behave sensibly; no operational surprises.
|
||||
- → all green: **scale to full capital, same unlevered book** (Sharpe is scale-invariant; capital is the lever).
|
||||
- → tracking error large / cost high: diagnose before scaling.
|
||||
|
||||
**Risk limits:** unlevered (no margin); $500–1k at-risk; adaptive layer auto-de-risks in vol/DD;
|
||||
IBIT removes the counterparty tail. **What it does NOT do:** make meaningful money — it proves the
|
||||
machine runs cleanly with real fills before you commit real capital. The honest bridge from
|
||||
validated backtest to scaled deployment.
|
||||
@@ -0,0 +1,183 @@
|
||||
# Crypto Funding Harvest — Deployable Strategy Spec
|
||||
|
||||
**Date:** 2026-06-07
|
||||
**Status:** Validated (4 gates passed), pre-deployment
|
||||
**Branch:** surfer-universe
|
||||
**Author:** research campaign synthesis
|
||||
|
||||
---
|
||||
|
||||
## 1. Executive summary
|
||||
|
||||
A **delta-neutral crypto perpetual-funding harvest**: hold `long spot + short perp` on coins
|
||||
whose funding rate is reliably positive, collecting the funding payment as carry with **no price
|
||||
exposure**. No prediction, no ML — a pure cross-sectional carry harvest gated by a regime filter.
|
||||
|
||||
**Why this and not everything else:** a multi-week, multi-market search (efficient equities/futures
|
||||
= no alpha, leak-free ML IC 0.004; simple 60/40 = ~0.7 Sharpe ceiling; energy = capital not
|
||||
algorithm) found this is the **only edge that breaks the ~0.7 retail Sharpe ceiling**, and it does
|
||||
so because crypto funding is structurally inefficient (retail perp longs over-pay) and the harvest
|
||||
is *latency-insensitive* (no colocation needed — a carry trade, not a race).
|
||||
|
||||
**Validated metrics (this dataset, 135 coins, 2019–2026):**
|
||||
|
||||
| Gate | Result |
|
||||
|---|---|
|
||||
| Structural | funding positive 75.5% of the time; 78.8% of coins pay carry on an avg day |
|
||||
| Delta-neutral | worst monthly bucket −1.5% (price risk genuinely hedged) |
|
||||
| Liveness | 2025–26 weakness is *regime* (deleverage), not decay; positive-funding yield intact (~3bp, = 2019/2023 levels) |
|
||||
| **Clean OOS** | filter chosen on 2019–24, applied blind to 2025–26 → **OOS Sharpe +5.1 raw** (IS-best); 9/12 strong-IS filters also OOS-positive; naive no-filter FAILS OOS (−5.8) |
|
||||
|
||||
**Honest expected performance (realistic, after basis-vol haircut ÷~2.5):**
|
||||
- **Sharpe ~2–3.6** (raw funding-only model shows 6–9; realistic accounts for basis/tracking-error vol we cannot model with single-price data)
|
||||
- **APR ~15–22%** on deployed capital (sensible-hurdle filter), regime-dependent
|
||||
- **Max drawdown ~−7 to −15%** (price-neutral; drawdowns are funding-flip + churn, not crashes)
|
||||
- **The Sharpe does NOT show the real risk: counterparty/exchange failure (−100% tail).**
|
||||
|
||||
---
|
||||
|
||||
## 2. The edge (why it exists, why it persists)
|
||||
|
||||
- **Perpetual futures funding** is the mechanism that tethers perp price to spot. When perps trade
|
||||
at a premium (bullish retail leverage), **longs pay shorts** a periodic funding rate (typically
|
||||
8-hourly on Binance/Bybit/OKX).
|
||||
- **Net structural bias is positive**: retail over-leverages long → funding is positive ~75% of
|
||||
the time. A `long spot + short perp` position is **delta-neutral** (spot and perp price moves
|
||||
cancel) and **collects the funding** the short perp leg receives.
|
||||
- **Why it persists** (not arbitraged to zero): capturing it requires capital *pre-positioned on
|
||||
each venue*, active management, and acceptance of counterparty risk — operational/risk frictions,
|
||||
not informational ones. It is the crypto analogue of an insurance premium.
|
||||
- **Latency-insensitive:** funding accrues over 8h periods; seconds of execution delay are
|
||||
immaterial. This is the single reason it is reachable by a non-colocated retail trader, unlike
|
||||
cross-exchange *latency* arb (which needs colocation and is NOT this strategy).
|
||||
|
||||
---
|
||||
|
||||
## 3. Strategy rules (precise, implementable)
|
||||
|
||||
### 3.1 Universe
|
||||
- Perpetual contracts on a **tier-1 venue** (see §4) that offers both spot and perp for the coin.
|
||||
- **Liquidity floor:** trailing-30d mean quote volume **> $5M/day** (validated threshold).
|
||||
- Exclude stablecoin-pair anomalies, delisting candidates, and coins without a spot leg.
|
||||
|
||||
### 3.2 Signal (the regime filter — the load-bearing piece)
|
||||
- For each eligible coin, compute **trailing-30-day mean funding rate** `tf30`.
|
||||
- **Qualify a coin iff `tf30 > 5 bp/day`** (≈ persistently well-paid carry). This is the validated
|
||||
filter: strong both IS (+7.9) and OOS (+9.1), and it sits the book out of the deleverage regime
|
||||
that whipsaws the naive "harvest anything positive" version (which FAILED OOS).
|
||||
- Causality: position decisions use **yesterday's** funding (`tf30` through t−1) to harvest day t.
|
||||
Funding is persistent, so this is both realistic and leak-free.
|
||||
|
||||
### 3.3 Positioning
|
||||
- For each qualifying coin: **long spot notional X + short perp notional X** (delta-neutral).
|
||||
- **Equal-weight** across qualifying coins (validated; funding-weighting did not improve risk-adj).
|
||||
- **Cash when nothing qualifies** — in a deleverage regime few/no coins clear the 5bp hurdle; the
|
||||
book correctly de-risks to stablecoin (do NOT force-harvest a dead regime).
|
||||
|
||||
### 3.4 Rebalance & cost
|
||||
- **Daily** check: enter coins crossing above the hurdle, exit coins falling below.
|
||||
- Cost budget: strategy validated net of **10 bp round-trip** (perp + spot, both legs); survives to
|
||||
~20 bp. Use **maker/limit orders** on entry/exit where possible to stay inside budget. Avoid
|
||||
rebalancing on marginal hurdle-crossings (add hysteresis: enter >5bp, exit <3bp, to cut churn).
|
||||
|
||||
### 3.5 Sizing
|
||||
- Per-coin notional = (deployed capital) / (number of qualifying coins), capped by §4 per-venue
|
||||
limits.
|
||||
- **Leverage:** the short-perp leg uses exchange margin; keep effective leverage low (≤2–3×) so a
|
||||
funding-flip + basis move cannot trigger liquidation. The spot leg is the hedge — never let the
|
||||
perp get liquidated while holding spot (that converts delta-neutral into naked long).
|
||||
|
||||
---
|
||||
|
||||
## 4. Risk management — counterparty is THE risk (read this twice)
|
||||
|
||||
Delta-neutrality removes *price* risk. It does **not** remove **counterparty/exchange risk**, which
|
||||
is the actual way this strategy produces a −100% (FTX, Mt. Gox, QuadrigaCX). The Sharpe ratio is
|
||||
blind to it. **This section matters more than the backtest.**
|
||||
|
||||
1. **Venue selection:** only tier-1 venues with **proof-of-reserves**, deep liquidity, and a
|
||||
solvency track record. No yield-farming protocols, no obscure CEXs chasing higher funding.
|
||||
2. **Collateral spreading:** spread capital across **≥2–3 venues** so no single failure is fatal.
|
||||
*Caveat at $35k:* small capital makes spreading hard (per-venue minimums + cost). Below ~$50k,
|
||||
counterparty concentration is unavoidable — treat the whole strategy as at-risk capital you can
|
||||
lose entirely, and start tiny (§7).
|
||||
3. **Withdrawal discipline:** keep only working collateral on exchanges; sweep profits to
|
||||
self-custody on a schedule. Never let the on-exchange balance grow unmonitored.
|
||||
4. **Liquidation guard:** low leverage (≤2–3×), automated margin-top-up alerts; the spot hedge must
|
||||
always survive a perp-leg margin call.
|
||||
5. **Deleverage kill-switch:** if fleet-wide funding goes broadly negative (the 2022/2025-26 signal:
|
||||
`frac_pos < 0.5` or `mean_all < 0`), **go fully to cash** — the filter does this automatically,
|
||||
but add a hard override.
|
||||
6. **Stablecoin risk:** the cash leg sits in stablecoins (USDT/USDC) — itself a (smaller)
|
||||
counterparty/depeg risk; diversify stable holdings.
|
||||
7. **Position cap per venue:** no more than (venue risk budget) on any one exchange.
|
||||
|
||||
---
|
||||
|
||||
## 5. Expected performance (honest)
|
||||
|
||||
| Metric | Raw model (funding-only) | Realistic (after basis vol) |
|
||||
|---|---|---|
|
||||
| Sharpe | 6–9 | **~2–3.6** |
|
||||
| APR (deployed) | ~17–22% | ~15–22% |
|
||||
| Max drawdown | −7 to −15% | similar (price-neutral) |
|
||||
| Worst month | −1.5% | basis noise adds some |
|
||||
| Regime behavior | cash in deleverage | cash in deleverage |
|
||||
|
||||
- **The realistic Sharpe (~2–3.6) is still 3–5× the 0.7 retail ceiling** for everything else found.
|
||||
- **Dollar reality at $35k:** ~15–20% APR ≈ **$5–7k/yr** expected, market-neutral — but with full
|
||||
counterparty-loss tail. Scales linearly with capital (Sharpe is scale-invariant; capacity is far
|
||||
above retail size).
|
||||
- **Unmodeled / open:** basis-vol (need spot+perp tick pairs to measure precisely), and the
|
||||
counterparty tail (un-backtestable — managed via §4, not measured).
|
||||
|
||||
---
|
||||
|
||||
## 6. Validation status
|
||||
|
||||
**Passed:**
|
||||
- Structural (funding positive 75.5%), delta-neutral (worst month −1.5%), liveness (regime not
|
||||
decay), clean OOS (IS-chosen filter holds OOS +5.1, 9/12 robust, naive no-filter fails OOS).
|
||||
- Scripts: `scripts/surfer/crypto_funding_harvest.py`, `crypto_funding_liveness.py`,
|
||||
`crypto_funding_oos.py`.
|
||||
|
||||
**Not yet validated (must close before scaling):**
|
||||
- **Basis vol** — current Sharpe is funding-only; get spot+perp price pairs and remodel realized P&L
|
||||
including tracking error → confirm realistic Sharpe.
|
||||
- **OOS length** — only ~1.5y OOS (single deleverage regime); held through the *hard* case, but
|
||||
extend as data accrues.
|
||||
- **Counterparty** — un-backtestable; validated only by §4 design discipline.
|
||||
- **Live execution** — fills, slippage, funding-timestamp capture, margin mechanics on the real
|
||||
venue (the paper-forward test, §7).
|
||||
|
||||
---
|
||||
|
||||
## 7. Phased rollout (gate each phase before the next)
|
||||
|
||||
1. **Paper-forward (4–8 weeks):** run the daily algorithm against *live* funding feeds, record
|
||||
intended positions + realized funding, **no capital**. Confirms the edge persists on truly
|
||||
unseen forward data (mirrors the surfer PoC's cron paper-test). Gate: forward funding-APR
|
||||
positive and consistent with backtest.
|
||||
2. **Micro-live ($1–3k, 1 venue):** smallest real size to validate execution, fills, funding
|
||||
capture, margin mechanics, cost budget. Gate: realized net ≈ paper, cost ≤ 15bp round-trip.
|
||||
3. **Small-live ($5–15k, 2 venues):** add collateral spreading, confirm counterparty controls and
|
||||
withdrawal discipline operate. Gate: clean operation through a funding-flip / mini-deleverage.
|
||||
4. **Scale ($35k+):** deploy full capital across ≥3 venues, full §4 risk stack engaged.
|
||||
|
||||
**Stop at any gate that fails.** This is the same diagnostic-first discipline that caught every
|
||||
false positive in the search (the inflated 5.68 Sharpe, the carry trap, the equity factors).
|
||||
|
||||
---
|
||||
|
||||
## 8. Go / no-go
|
||||
|
||||
**Go** if: paper-forward (§7.1) confirms positive forward funding-APR AND you accept that this is
|
||||
**at-risk crypto capital with a counterparty −100% tail** in exchange for a realistic ~2–3.6 Sharpe
|
||||
market-neutral return.
|
||||
|
||||
**No-go** if: you cannot accept the counterparty tail, OR paper-forward shows the edge has decayed
|
||||
(funding compression to ~cost), OR you cannot spread across ≥2 reputable venues at your capital.
|
||||
|
||||
**The honest framing:** this is the one validated way past 0.7 the whole search found. It is real,
|
||||
alive, and reachable — and the price of admission is crypto + counterparty risk, not market risk.
|
||||
That trade-off is a values decision, now made with complete information.
|
||||
86
docs/superpowers/specs/2026-06-08-cagr-reference-plan.md
Normal file
86
docs/superpowers/specs/2026-06-08-cagr-reference-plan.md
Normal file
@@ -0,0 +1,86 @@
|
||||
# CAGR-Maximizer — Honest Reference Plan
|
||||
|
||||
**Date:** 2026-06-08
|
||||
**Purpose:** the honest answer to "how do we build the most wealth?" after exhaustively testing strategies,
|
||||
the adaptive multi-strat book, leverage, and risk-management. Comparison tool: `scripts/surfer/strategy_compare.py`.
|
||||
|
||||
---
|
||||
|
||||
## The core truth (measured, not assumed)
|
||||
|
||||
1. **CAGR builds terminal wealth, not Sharpe.** Over a long horizon with contributions, the highest-CAGR
|
||||
asset wins — even at lower Sharpe. (SPY 0.64 Sharpe but 12.4% CAGR beats the book's 0.96 Sharpe / 5.1% CAGR on money.)
|
||||
2. **More CAGR = more risk. Always. No exception.** Every risk-managed variant (the book, the overlay,
|
||||
blends) gives up CAGR for lower drawdown. Risk management is *insurance you pay for*, not a free improvement.
|
||||
3. **The only "free" CAGR is drag reduction** — taxes, fees, and (biggest) not selling at the bottom.
|
||||
4. **The book's apparent Sharpe edge (0.96 vs 0.64) was an rf=0 artifact.** Measured as excess-over-financing
|
||||
(what matters for leverage), equities (0.49) actually beat the diversified book (0.40) in this era —
|
||||
so levering the book does NOT beat buy-hold equity. The institutional risk-parity edge needs the book's
|
||||
*excess* Sharpe > equity's, which it wasn't.
|
||||
5. **The leverage moat is cheap financing.** Futures/box ≈ SOFR+spread (~3%); retail margin ~6.5% destroys it.
|
||||
|
||||
---
|
||||
|
||||
## The candidates (measured 2006–2026, incl. 2008 −55%)
|
||||
|
||||
| Strategy | CAGR | vol | Sharpe | maxDD | $360k+$8k/mo → 20y |
|
||||
|---|---|---|---|---|---|
|
||||
| **SPY buy-hold (1.0x)** | ~10.6% | 19% | 0.64 | **−55%** | ~$7.8M |
|
||||
| SPY 1.2–1.3x (cheap futures) | ~12–13% | 23–25% | ~0.6 | **−63 to −67%** | ~$8.8–9.2M |
|
||||
| SPY 1.5x | ~14.5% | 28% | — | −73% | ~$10.1M |
|
||||
| 60/40 | ~8% | 14% | 0.73 | −20 to −30% | ~$6.5M |
|
||||
| 70/30 SPY+book | ~10% | 14% | 0.73 | −40% | ~$9.0M |
|
||||
| SPY + adaptive overlay | ~8% | 11% | 0.77 | −18% | ~$6.7M |
|
||||
| **Adaptive multi-strat book** | ~5–6% | 5% | **0.96** | **−10%** | ~$4.3–4.9M |
|
||||
|
||||
(SPY ≥2x: CAGR peaks ~2x then volatility-drag falls; −84%+ drawdown = ruin/margin-call territory. Not viable.)
|
||||
|
||||
---
|
||||
|
||||
## CAGR levers, ranked by sense
|
||||
|
||||
| Lever | Effect | Honesty |
|
||||
|---|---|---|
|
||||
| **Max equity allocation, held** | base ~10–11%/yr | highest-return asset; reliable |
|
||||
| **Reduce drag** (tax-advantaged account, cheap index funds, never sell, stay invested) | **+1–3%/yr** | **free + reliable — most people leave this on the table** |
|
||||
| **Time + contributions ($8k/mo)** | dominant terminal driver | your strongest card |
|
||||
| **Modest cheap leverage (~1.2–1.3x via futures)** | +~1.5–2%/yr (+$1–1.4M/20y) | amplifies drawdown to −63–67% + margin-call risk |
|
||||
| Factor tilts (small-cap value, momentum, quality) | hist. +1–2%/yr | less reliable forward |
|
||||
| ~~Risk-managed book / overlay~~ | **−4%/yr CAGR** | a risk *reducer*, not a CAGR maximizer |
|
||||
|
||||
---
|
||||
|
||||
## The leverage reality (the only knob with real upside — handle with care)
|
||||
|
||||
- 1.2–1.3x boosts CAGR ~+1.5–2%/yr (~$1–1.4M over 20y) but turns the −55% crash into **−63–67%**.
|
||||
- **−67% on a $2M pot = $660k at the 2009 bottom**, while still contributing into the abyss.
|
||||
- **Margin calls:** fixed-notional futures in a −55% crash force liquidation at the bottom → *realized* ruin
|
||||
(the backtest "survives" only because it's daily-rebalanced math). Lived experience is worse than backtest.
|
||||
- Above ~1.5x: reckless (−73%+); above ~2x: ruin.
|
||||
- **The real constraint is not the math — it's whether you survive a −65% drawdown + margin without forced/panic selling.**
|
||||
|
||||
---
|
||||
|
||||
## Recommendation
|
||||
|
||||
For a disciplined accumulator with a strong savings rate (~$8k/mo), long horizon, and stable income:
|
||||
|
||||
1. **Max equity (100%), in a tax-advantaged account, cheap index funds, contribute monthly, never sell.**
|
||||
Base case ~$7.8M over 20y. Simple, cheap, the most for the disciplined.
|
||||
2. **Optional ~1.2–1.3x via futures (cheap financing)** — only if you can survive (financially + emotionally)
|
||||
a −65% drawdown + margin calls without forced selling. Adds ~$1–1.4M over 20y.
|
||||
3. **Always: reduce drag** (tax-efficiency, low fees, stay invested) — the free CAGR.
|
||||
4. **The adaptive multi-strat book is NOT the wealth-maximizer** — it's the *capital-preservation / low-stress / withdrawal-phase* tool. Use it if you'd panic-sell equities, are near withdrawal, or value sleep over ~4%/yr.
|
||||
|
||||
**The decision hinges on one honest question: do you sell at the bottom of a crash, or hold?**
|
||||
- Hold → high equity (±1.3x), simplest, most money.
|
||||
- Unsure → 70/30 or the overlay (pay return for a tolerable ride).
|
||||
- Need capital preservation → the book.
|
||||
|
||||
---
|
||||
|
||||
## What's deployed
|
||||
|
||||
- **Adaptive multi-strat book:** live on IBKR paper via the K8s CronJob (autonomous, daily eval) — the *low-drawdown* reference.
|
||||
- **Comparison tool:** `scripts/surfer/strategy_compare.py` — re-runnable side-by-side of all candidates (backtest + trajectory).
|
||||
- **Honest bottom line:** there is no secret signal and no clever leverage trick that beats "high-return assets, cheap, held long, funded heavily." We tested them all. The wealth levers are savings rate, time, low costs, and the discipline to not sell.
|
||||
@@ -0,0 +1,80 @@
|
||||
# Decommission Dead Rust Infra — Design (Phase 1 of infra consolidation)
|
||||
|
||||
**Date:** 2026-06-21
|
||||
**Status:** Design (pending approval)
|
||||
**Repo:** foxhunt (where the platform IaC currently lives)
|
||||
|
||||
## Motivation
|
||||
|
||||
The active trading work is now the Python **fxhnt** fund; the Rust **foxhunt** ML/HFT system is dormant.
|
||||
Its dedicated infra (GPU training pools, Rust build/training caches, GPU CI runners, training artifact
|
||||
buckets, training manifests) is unused but still declared/provisioned. Phase 1 decommissions it — for
|
||||
cost/clutter savings and to shrink the surface before Phase 2 (relocating the remaining *platform* IaC
|
||||
into `fxhnt/infra`, a separate spec).
|
||||
|
||||
**Verified unused (2026-06-21):** L40S + H100 pools have 0 nodes; `cargo-target-{cpu,cuda,cuda-test}` +
|
||||
`feature-cache-pvc` are `Used By: <none>`; no gitlab-runner pods running.
|
||||
|
||||
## Scope — what gets removed (all verified unused)
|
||||
|
||||
1. **Terraform (kapsule module + terragrunt.hcl):** set to `false` →
|
||||
- `enable_ci_training_l40s_pool` (L40S-1-48G GPU pool)
|
||||
- `enable_ci_training_h100_pool` (H100-1-80G GPU pool)
|
||||
- `enable_ci_compile_cpu_hm_pool` (POP2-HM-32C-256G precompute pool — Rust `precompute_features` only)
|
||||
`terragrunt apply` destroys exactly these 3 pools. Also delete their now-dead var blocks/resources.
|
||||
2. **Helm uninstall** (foxhunt ns): `gitlab-runner-h100`, `gitlab-runner-h100-sxm`, `gitlab-runner-h100x2`
|
||||
(Rust GPU CI runners). **Main `gitlab-runner`:** verify nothing non-Rust uses it (fxhnt has no
|
||||
`.gitlab-ci.yml`; cockpit deploys via Argo) → uninstall if confirmed dead, else keep. (confirm step)
|
||||
3. **PVCs delete** (foxhunt ns — PURE BUILD CACHES only, unmounted-verified, ~175 GB reclaimed):
|
||||
`cargo-target-cpu` (60Gi), `cargo-target-cuda` (45Gi), `cargo-target-cuda-test` (30Gi),
|
||||
`sccache-cpu` (20Gi), `sccache-cuda` (20Gi). These hold zero data — regenerated on any build.
|
||||
4. **MinIO buckets delete** (NO market data — verified 0 `.dbn` objects): `foxhunt-binaries` (2.3 GiB
|
||||
compiled binaries), `foxhunt-training-results` (4.4 GiB run logs/outputs), `foxhunt-models` (empty).
|
||||
5. **Repo manifests/templates remove** (foxhunt repo): `infra/k8s/training/`, `infra/k8s/gpu-overlays/`,
|
||||
the Rust Argo workflow templates (`train-multi-seed-template.yaml`, `lob-backtest-sweep-template.yaml`,
|
||||
`ci-pipeline-template.yaml`, alpha-rl train templates), `infra/k8s/jobs/download-trades-job.yaml`
|
||||
(databento), the 3 GPU-runner helm value files. Also remove the dead Argo `WorkflowTemplate`s from the
|
||||
cluster (`kubectl delete wftmpl`) for the Rust train/backtest pipelines.
|
||||
|
||||
## 🛑 EXPLICITLY PRESERVE (do NOT delete — market data + reusable assets)
|
||||
- **`training-data-pvc` (500 GiB)** — the raw Databento MBP-10 `.dbn` market data. KEEP.
|
||||
- **`foxhunt-training-data` bucket (38 GiB, 54 `.dbn`/`.zst`)** — Databento market data (expensive to
|
||||
re-acquire; the fund's `.dbn` backtests read it). KEEP.
|
||||
- **`test-data-pvc` (50 GiB)** — `.dbn` test subsets (tier-1.5 smoke uses test_data). KEEP (verify, don't delete).
|
||||
- **`feature-cache-pvc` (100 GiB)** — derived ML features; regenerable but compute-costly. KEEP (conservative).
|
||||
- All fxhnt/platform data: `fxhnt-backtest-data`, `fxhnt-surfer-data`, `multistrat-state`, `questdb-pvc`,
|
||||
`tempo-data`, `netbird-data`, `foxhunt-gitlab-*` + `foxhunt-backups` buckets.
|
||||
General rule: delete only **pure compute/build artifacts** (caches, compiled binaries, run logs, GPU pools);
|
||||
**never** anything holding `.dbn`/market data or any potentially-reusable dataset.
|
||||
|
||||
## Out of scope (KEEP — shared platform / Python fund)
|
||||
`ci-compile-cpu` pool (Python cockpit builds), platform pool, GitLab, MinIO, monitoring, Mattermost,
|
||||
Stalwart, Kanidm, NetBird, DNS, databases, cert-manager, tailscale proxy, public-gateway, dagster/cockpit,
|
||||
the fxhnt forward-track jobs. Phase 2 (IaC relocation to `fxhnt/infra` + TF-state move) is a separate spec.
|
||||
|
||||
## Execution order (each destructive step verified + confirmed)
|
||||
1. **Terraform first**: edit terragrunt.hcl (3 `enable_*=false`), `terragrunt plan` → **STOP unless the
|
||||
plan shows ONLY the 3 pools destroyed + 0 other destroys**; then `apply`.
|
||||
2. **Helm uninstalls** (GPU runners; main runner only after the use-check confirms dead).
|
||||
3. **PVC deletes** (re-confirm `Used By: <none>` immediately before each delete — irreversible).
|
||||
4. **Bucket deletes** (list contents first; irreversible — explicit confirm; `mc rb --force` via the
|
||||
port-forward + minio creds).
|
||||
5. **Repo cleanup**: remove the dead manifests/templates + module var blocks, `kubectl delete wftmpl` the
|
||||
dead templates, commit.
|
||||
|
||||
## Risks / safety
|
||||
- **Irreversible**: PVC + bucket deletes, GPU-pool destroys. Mitigation: verify-unused immediately before
|
||||
each; Terraform plan reviewed for 0-unexpected-destroys; explicit confirm on PVC/bucket deletes.
|
||||
- **Mis-scope risk**: a kept resource accidentally listed. Mitigation: the Terraform plan gate (only the 3
|
||||
pools) + the `Used By` re-check + bucket content listing before delete.
|
||||
- **Main gitlab-runner**: don't uninstall without confirming no active consumer (could break a CI path).
|
||||
- Cluster health unaffected: nothing here touches the cockpit/dagster/fund/platform services.
|
||||
|
||||
## Acceptance criteria
|
||||
- `terragrunt plan` (kapsule) shows the 3 GPU/precompute pools gone, then clean (no changes).
|
||||
- GPU-runner helm releases uninstalled; main runner resolved (kept or uninstalled per check).
|
||||
- The 5 build-cache PVCs deleted (cargo×3 + sccache×2, ~175 GB reclaimed); 3 Rust buckets deleted
|
||||
(binaries, training-results, models). **`.dbn` market data untouched** (training-data-pvc + the
|
||||
foxhunt-training-data bucket + test-data-pvc still present, byte-for-byte).
|
||||
- Dead training/gpu manifests + Argo templates removed from repo + cluster; committed.
|
||||
- Cockpit (`dashboard.fxhnt.ai` 200), dagster, and the fund tracks still healthy post-cleanup.
|
||||
114
docs/superpowers/specs/2026-06-21-gitea-replace-gitlab-design.md
Normal file
114
docs/superpowers/specs/2026-06-21-gitea-replace-gitlab-design.md
Normal file
@@ -0,0 +1,114 @@
|
||||
# Stand up Gitea + migrate repos + move cockpit registry — Design (Phase 2B)
|
||||
|
||||
**Date:** 2026-06-21
|
||||
**Status:** Design (approved)
|
||||
**Repo:** foxhunt (IaC current home; relocation to fxhnt is Phase 2D)
|
||||
**Part of:** Phase 2 (GitLab → Gitea consolidation, [[project_phase2_gitlab_to_gitea]]). 2A (TF state →
|
||||
Scaleway) DONE. This is **2B**. 2C (GitLab removal) and 2D (IaC relocation) follow.
|
||||
|
||||
## Motivation
|
||||
|
||||
Self-hosted GitLab uses ~6.5 GiB RAM across 17 pods (webservice 1.96 GiB + 3× sidekiq ~3 GiB + gitaly +
|
||||
registry + prometheus + …). The active work is the Python `fxhnt` fund; GitLab is Rust-era overkill.
|
||||
Replace it with **Gitea** (~150–250 MiB, chart + existing Postgres) for git hosting, and move the
|
||||
**cockpit container image** to **Scaleway Container Registry** so GitLab can be removed in 2C without
|
||||
breaking the live cockpit deploy.
|
||||
|
||||
## Decisions (from brainstorming)
|
||||
|
||||
- **Gitea install:** `gitea/gitea` Helm chart. Subcharts `postgresql`/`redis-cluster`/`redis`/`memcached`
|
||||
**disabled**. Built-in **Actions disabled** (Argo does CI/CD). Cache = memory (single replica).
|
||||
- **Database:** reuse the **existing in-cluster `postgres`** (svc `postgres:5432`, already has a backup
|
||||
cronjob). Create a dedicated `gitea` database + role.
|
||||
- **Repos:** migrate **both** — `fxhnt` (active) and `foxhunt` (876 MB, flagged **archived/read-only**) —
|
||||
via `git push --mirror` from existing local clones.
|
||||
- **Canonical hostname stays `git.fxhnt.ai`** (no permanent `gitea.fxhnt.ai`, no remote-URL churn).
|
||||
Gitea comes up internal-only; validate over a port-forward; then cut `git.fxhnt.ai` (HTTPS) + SSH over
|
||||
to Gitea in one switch. **SSH on `:22`** (nothing host-SSHes the git node, so the standard port is free);
|
||||
clone URLs become the clean `git@git.fxhnt.ai` and `:2222` is retired (local remotes updated once).
|
||||
GitLab keeps running (unaddressed) as the 2C-removal fallback.
|
||||
- **Container registry:** cockpit image → **Scaleway Container Registry** (`rg.fr-par.scw.cloud/bizworx/`,
|
||||
namespace exists). Retire the in-cluster GitLab registry. External/managed, no registry pod.
|
||||
- **Argo integration:** Gitea push webhook → Argo Events webhook eventsource → sensor → submit
|
||||
`wftmpl/fxhnt-cockpit`. Replaces the old `gitlab-push-eventsource`.
|
||||
|
||||
## Components
|
||||
|
||||
1. **`infra/k8s/gitea/` (new):** Helm values `values.yaml` + a thin install (release `gitea`, ns
|
||||
`foxhunt`). Storage: 1 PVC `sbs-default-retain` ~5 Gi (foxhunt 876 MB + fxhnt + headroom). Resources:
|
||||
requests 128 Mi/100m, limits 512 Mi. Admin user + secrets from a k8s secret (NOT committed).
|
||||
Gitea config: `server.SSH_PORT`/`SSH_LISTEN_PORT`, `server.ROOT_URL=https://git.fxhnt.ai/`,
|
||||
`server.DISABLE_SSH=false`, DB type `postgres` → `postgres:5432/gitea`.
|
||||
2. **Postgres prep:** one-time `psql` on the existing `postgres` pod — `CREATE ROLE gitea …; CREATE
|
||||
DATABASE gitea OWNER gitea;`. Credentials in a k8s secret consumed by the Gitea release.
|
||||
3. **Exposure:**
|
||||
- HTTPS: extend the nginx `*.fxhnt.ai` proxy (`infra/k8s/gitlab/tailscale-proxy.yaml`) — at cutover,
|
||||
repoint the `git.fxhnt.ai` server block `proxy_pass` → `gitea-http.foxhunt.svc:3000`.
|
||||
- SSH: expose Gitea SSH on `:22` via the Tailscale proxy (socat) at cutover; retire `:2222`; update the
|
||||
two local remotes to `ssh://git@git.fxhnt.ai/...`.
|
||||
- Registry `:5050` server block: removed (Scaleway registry is external, not proxied here).
|
||||
4. **Registry migration:**
|
||||
- `fxhnt/infra/argo/cockpit-build-deploy.yaml`: kaniko `--destination` + `--cache-repo` →
|
||||
`rg.fr-par.scw.cloud/bizworx/fxhnt-cockpit:latest` (+ cache repo); drop the `--insecure-registry`/
|
||||
`--skip-tls-verify` GitLab flags (Scaleway is TLS); auth via a Scaleway registry dockerconfig secret.
|
||||
- `fxhnt/infra/k8s/orchestration/dagster.yaml` + dashboard: image ref →
|
||||
`rg.fr-par.scw.cloud/bizworx/fxhnt-cockpit:latest`; `imagePullSecrets` → Scaleway registry secret.
|
||||
- Create the Scaleway registry pull/push secret (dockerconfigjson) in `foxhunt` ns.
|
||||
5. **Argo Events wiring:** `infra/k8s/argo/events/` — a Gitea webhook eventsource + sensor that submits
|
||||
`fxhnt-cockpit` on push to `fxhnt` default branch (reuses the surviving eventbus + workflow-trigger
|
||||
machinery). Webhook secret in a k8s secret; configured in Gitea repo settings.
|
||||
6. **Deploy script:** `fxhnt/scripts/argo-deploy-cockpit.sh` + the `fxhnt-cockpit` wftmpl clone URL →
|
||||
Gitea (`git.fxhnt.ai`). After cutover the remote name is unchanged, so this is a clone-source confirm.
|
||||
|
||||
## Data flow
|
||||
|
||||
Dev pushes to `git.fxhnt.ai` (Gitea) → webhook → Argo Events sensor → `fxhnt-cockpit` workflow → kaniko
|
||||
builds + pushes to Scaleway registry → `kubectl apply` + rollout → dagster/dashboard pull the new image
|
||||
from Scaleway. Terraform unaffected (state external since 2A).
|
||||
|
||||
## Migration & cutover sequence (each step verified)
|
||||
|
||||
1. Create `gitea` DB on existing postgres + secrets.
|
||||
2. `helm install gitea` (internal svc only). Verify `/api/healthz` 200, admin login via port-forward.
|
||||
3. Port-forward Gitea; `git push --mirror` `fxhnt` and `foxhunt` to it; mark `foxhunt` archived.
|
||||
4. **Validate (pre-cutover):** clone both repos back from the port-forward, diff top commit SHAs vs
|
||||
GitLab — must match.
|
||||
5. **Registry migration:** create Scaleway registry secret; update kaniko build (+cache) + deploy image
|
||||
refs; run one cockpit build → confirm image in Scaleway registry + a test rollout pulls it OK.
|
||||
6. **Cutover:** repoint nginx `git.fxhnt.ai` → Gitea + SSH `:2222` → Gitea; remove the `:5050` block.
|
||||
Verify `git.fxhnt.ai` serves Gitea (web + clone + SSH push).
|
||||
7. **Wire Argo Events:** add Gitea webhook → sensor; test push to `fxhnt` main fires a cockpit deploy.
|
||||
|
||||
## Safety / reversibility
|
||||
|
||||
- GitLab stays fully running throughout 2B (just unaddressed after cutover) — rollback = repoint nginx/SSH
|
||||
back to GitLab. Removal is **2C** only.
|
||||
- The Scaleway registry change is additive until the deploy refs flip; the GitLab registry image stays as
|
||||
fallback until 2C.
|
||||
- All repo data validated (SHA diff) before the public name moves.
|
||||
- Secrets (DB creds, admin, webhook, registry dockerconfig) are k8s secrets — **never committed**.
|
||||
|
||||
## Risks
|
||||
|
||||
- **Registry auth:** Scaleway registry needs a valid API-key dockerconfig for both kaniko push and pod
|
||||
pull. Mitigation: test the build+rollout (step 5) before cutover.
|
||||
- **Gitea↔Postgres coupling:** Gitea shares the existing `postgres`. If that instance is later removed,
|
||||
Gitea breaks. Acceptable (postgres is a kept platform service with backups); noted for 2C/2D.
|
||||
- **SSH port collision at cutover:** GitLab + Gitea can't both own `:2222` on the Tailscale proxy. Cutover
|
||||
flips it atomically; brief push unavailability during the switch (single maintainer — fine).
|
||||
- **Webhook reachability:** Gitea (in-cluster) must reach the Argo Events webhook svc in-cluster — same
|
||||
ns, no Tailscale hop needed.
|
||||
|
||||
## Out of scope
|
||||
|
||||
GitLab helm uninstall + GitLab buckets/state cleanup (**2C**); `git mv` IaC foxhunt → fxhnt (**2D**).
|
||||
|
||||
## Acceptance criteria
|
||||
|
||||
- Gitea running (1 pod, existing postgres), `/api/healthz` 200, ~≤300 MiB.
|
||||
- Both repos in Gitea; clone SHAs match GitLab; `foxhunt` archived.
|
||||
- Cockpit image builds to + deploys from Scaleway Container Registry; `dashboard.fxhnt.ai` 200 on the
|
||||
Scaleway-sourced image.
|
||||
- `git.fxhnt.ai` (HTTPS + SSH `:2222`) serves Gitea; git remotes unchanged.
|
||||
- Push to `fxhnt` main auto-triggers a cockpit deploy via Argo Events.
|
||||
- GitLab still running as fallback (removed in 2C).
|
||||
@@ -0,0 +1,94 @@
|
||||
# Migrate Terraform State to Scaleway Object Storage — Design (Phase 2A)
|
||||
|
||||
**Date:** 2026-06-21
|
||||
**Status:** Design (approved)
|
||||
**Repo:** foxhunt (IaC current home; relocation to fxhnt is Phase 2D)
|
||||
**Part of:** Phase 2 (GitLab → Gitea consolidation). Sub-projects: **2A state migration (this)**, 2B Gitea
|
||||
stand-up + repo migration + Argo wiring, 2C GitLab decommission, 2D IaC relocation into fxhnt.
|
||||
|
||||
## Motivation
|
||||
|
||||
Terraform state for the whole production cluster currently lives in the **self-hosted GitLab** http
|
||||
backend (project ID=1), whose object data sits in **in-cluster MinIO**. That is a chicken-and-egg
|
||||
fragility: the state describing the cluster lives *inside* the cluster it manages. GitLab is also slated
|
||||
for decommission (Phase 2C, replaced by lightweight Gitea). Moving state to **external Scaleway Object
|
||||
Storage** decouples it from both GitLab and the cluster — a prerequisite for removing GitLab and a
|
||||
durability win on its own (state survives cluster loss).
|
||||
|
||||
## Decision (keystone)
|
||||
|
||||
**Target backend: Scaleway Object Storage (native S3), bucket `foxhunt-tfstate`** (already exists,
|
||||
fr-par, created ~3 months ago but never wired). Reuse it (no rename — avoids a needless create; the
|
||||
`foxhunt-` prefix is cosmetic and harmless). Locking via **OpenTofu native S3 lockfile** (`use_lockfile`,
|
||||
OpenTofu v1.11.5 present) — no DynamoDB.
|
||||
|
||||
## Scope — what changes
|
||||
|
||||
1. **`infra/live/production/root.hcl`** — replace the single `remote_state` block: `backend = "http"` (GitLab) →
|
||||
`backend = "s3"` (Scaleway). New config:
|
||||
```hcl
|
||||
remote_state {
|
||||
backend = "s3"
|
||||
generate = { path = "backend.tf", if_exists = "overwrite" }
|
||||
config = {
|
||||
bucket = "foxhunt-tfstate"
|
||||
key = "${path_relative_to_include()}/terraform.tfstate"
|
||||
region = "fr-par"
|
||||
endpoints = { s3 = "https://s3.fr-par.scw.cloud" }
|
||||
skip_credentials_validation = true
|
||||
skip_region_validation = true
|
||||
skip_requesting_account_id = true
|
||||
skip_metadata_api_check = true
|
||||
use_lockfile = true
|
||||
}
|
||||
}
|
||||
```
|
||||
Credentials: Scaleway access/secret keys exported as `AWS_ACCESS_KEY_ID` / `AWS_SECRET_ACCESS_KEY`
|
||||
(the s3 backend reads AWS-style env vars) from the existing `scaleway-credentials` k8s secret.
|
||||
2. **State migration, per module, sequential** (`kapsule`, `public-gateway`, `dns`, `block-storage`):
|
||||
- Back up first: `terragrunt state pull > /tmp/tfstate-backup-<module>.tfstate` (from GitLab, BEFORE
|
||||
changing the backend for that module).
|
||||
- `terragrunt init -migrate-state -force-copy` → copies GitLab state → Scaleway S3.
|
||||
- `terragrunt plan` → **must report `No changes`** before proceeding to the next module.
|
||||
3. **Lock smoke test** (once, on one module): confirm `use_lockfile` works against Scaleway — a lock
|
||||
object is created during apply/plan-with-lock and released after (verify no orphan `.tflock` left).
|
||||
|
||||
## Migration order & gates
|
||||
|
||||
`kapsule` → `public-gateway` → `dns` → `block-storage`. After EACH: backup taken, `-migrate-state`
|
||||
succeeded, `plan = No changes` confirmed. **STOP and revert if any module's post-migration plan shows
|
||||
drift** (revert = restore `http` backend in root.hcl + `terragrunt init` to re-attach GitLab state; the
|
||||
GitLab copy is never deleted in 2A).
|
||||
|
||||
## Safety / reversibility
|
||||
|
||||
- **Old GitLab state is NOT deleted in 2A** — it stays as a live fallback until GitLab is removed in 2C.
|
||||
So the entire migration is reversible by flipping `root.hcl` back to the http backend.
|
||||
- `/tmp` state backups taken before each module (belt-and-suspenders; never committed — contains
|
||||
resource IDs, not secrets, but treat as sensitive).
|
||||
- One module at a time with a `plan = No changes` gate prevents a bad backend config from cascading.
|
||||
- `endpoints.s3` (not deprecated top-level `endpoint`) + the four `skip_*` flags are required for the
|
||||
AWS s3 backend to talk to Scaleway's S3-compatible API without AWS-specific preflight calls.
|
||||
|
||||
## Risks
|
||||
|
||||
- **Scaleway S3 quirks:** some provider/backend versions need `skip_s3_checksum = true` for Scaleway. If
|
||||
`-migrate-state` errors on checksum, add it. (Documented as a known fallback in the plan.)
|
||||
- **Lock semantics:** Scaleway Object Storage must honor conditional-write for `use_lockfile`. The lock
|
||||
smoke test catches this; fallback is to run with locking disabled (single maintainer) if unsupported.
|
||||
- **Credential confusion:** the s3 backend reads `AWS_*` env vars, NOT `SCW_*`. Plan must export both
|
||||
(SCW_* for the provider, AWS_* for the backend) — a common foot-gun.
|
||||
|
||||
## Out of scope (later sub-projects)
|
||||
|
||||
Gitea stand-up + repo migration + Argo wiring (2B); GitLab helm removal + GitLab state/bucket cleanup
|
||||
(2C); `git mv` of IaC from foxhunt → fxhnt + path updates (2D). Files stay in foxhunt for 2A; only the
|
||||
backend changes.
|
||||
|
||||
## Acceptance criteria
|
||||
|
||||
- `infra/live/production/root.hcl` uses the Scaleway s3 backend; committed.
|
||||
- All 4 modules migrated: each shows `terragrunt plan = No changes` against the new backend.
|
||||
- State objects present in `foxhunt-tfstate` at `<module>/terraform.tfstate` (verified via `scw`/`mc`).
|
||||
- Lock smoke test passes (lock acquired + released, no orphan lockfile).
|
||||
- GitLab http state still intact (untouched fallback) — not yet deleted.
|
||||
@@ -1,161 +0,0 @@
|
||||
# Deps Cache Image — Cold-start Compile Acceleration
|
||||
|
||||
This README documents the **`ci-builder-cpu-with-deps:nightly`** image
|
||||
and the Argo plumbing that uses it. Created as part of the CI cold-start
|
||||
optimization sweep (April 2026).
|
||||
|
||||
## What it is
|
||||
|
||||
A Docker image that carries a pre-built `target/` directory at
|
||||
`/cargo-target-prebuilt/` containing the workspace's third-party
|
||||
dependency rlibs. It's pulled by the `compile-services` step of
|
||||
`compile-and-deploy-template.yaml` as an **initContainer** that rsyncs
|
||||
its payload into the `cargo-target-cpu` PVC.
|
||||
|
||||
## Why it exists
|
||||
|
||||
On a fresh node (autoscaler scaled up, fresh PVC, or after a PVC purge),
|
||||
`cargo build --release` would otherwise spend 3-5 minutes compiling
|
||||
~1500 third-party crates before touching workspace member code. With
|
||||
the deps cache image, the cold-start compile path is:
|
||||
|
||||
1. **initContainer pulls image** from in-cluster registry (~30-60s)
|
||||
2. **rsync prebuilt rlibs into PVC** (~30-60s, ~6-8 GB transfer over local FS)
|
||||
3. **cargo build delta-compiles only workspace member code** (~1-2 min)
|
||||
|
||||
Net cold-start saving: **2-4 min per compile pod**.
|
||||
|
||||
## Files
|
||||
|
||||
| File | Role |
|
||||
|------|------|
|
||||
| `infra/docker/Dockerfile.ci-deps-cache` | Builds the deps-cache image (multi-stage). Stage 1 atop ci-builder-cpu, runs `cargo build --release --workspace`. Stage 2 thin Ubuntu + rsync layer carrying just `/cargo-target-prebuilt/`. |
|
||||
| `infra/k8s/argo/refresh-deps-cache-template.yaml` | WorkflowTemplate that drives a Kaniko build of the Dockerfile + push to GitLab registry. Includes a CronWorkflow that runs nightly at 03:00 UTC (suspended by default). |
|
||||
| `infra/k8s/argo/compile-and-deploy-template.yaml` | The `compile-services` step has a `seed-deps-cache` initContainer that rsyncs `/cargo-target-prebuilt/` into the PVC, gated by a stamp file. |
|
||||
|
||||
## How to refresh manually
|
||||
|
||||
```bash
|
||||
# Trigger a one-shot rebuild against current main HEAD.
|
||||
argo submit --from workflowtemplate/refresh-deps-cache -n foxhunt
|
||||
```
|
||||
|
||||
To target a specific commit or bump the cache version (forces re-seed
|
||||
on every compile pod):
|
||||
|
||||
```bash
|
||||
argo submit --from workflowtemplate/refresh-deps-cache -n foxhunt \
|
||||
-p commit-sha=$(git rev-parse main) \
|
||||
-p deps-version=2 \
|
||||
-p image-tag=nightly
|
||||
```
|
||||
|
||||
## Enabling the nightly cron
|
||||
|
||||
The CronWorkflow ships **suspended** so you can validate manually first.
|
||||
To enable:
|
||||
|
||||
```bash
|
||||
kubectl -n foxhunt patch cronworkflow refresh-deps-cache-nightly \
|
||||
--type=merge -p '{"spec":{"suspend":false}}'
|
||||
```
|
||||
|
||||
To disable (return to manual-only):
|
||||
|
||||
```bash
|
||||
kubectl -n foxhunt patch cronworkflow refresh-deps-cache-nightly \
|
||||
--type=merge -p '{"spec":{"suspend":true}}'
|
||||
```
|
||||
|
||||
## Bumping `DEPS_VERSION`
|
||||
|
||||
The compile-services initContainer skips re-seeding when
|
||||
`/cargo-target/cpu_deps_v${DEPS_VERSION}.stamp` exists on the PVC.
|
||||
Bump `DEPS_VERSION` to force every running compile pod to re-seed
|
||||
its PVC from a fresh image. Bump in two places (must match):
|
||||
|
||||
1. `refresh-deps-cache-template.yaml` workflow `deps-version` parameter
|
||||
default value
|
||||
2. `compile-and-deploy-template.yaml` initContainer `DEPS_VERSION` env var
|
||||
|
||||
Bump when:
|
||||
- Cargo.lock has churned significantly (most prebuilt rlibs
|
||||
fingerprint-mismatch cargo's freshness check)
|
||||
- The Rust toolchain version changes (e.g. 1.89 -> 1.90)
|
||||
- A workspace member is renamed or split (rare)
|
||||
|
||||
You do **not** need to bump for routine commits — cargo's own
|
||||
fingerprint check will discard stale rlibs and rebuild as needed.
|
||||
|
||||
## Debugging
|
||||
|
||||
**Symptom: compile-services is still slow on a fresh PVC.**
|
||||
|
||||
Check the seed-deps-cache initContainer logs first:
|
||||
|
||||
```bash
|
||||
argo logs -n foxhunt <workflow-name> -c seed-deps-cache
|
||||
```
|
||||
|
||||
Look for:
|
||||
- `WARN: /cargo-target-prebuilt missing in image` — the image was built
|
||||
but the snapshot dir is empty. The inner `cargo build --workspace`
|
||||
failed in the deps-cache image build. Check the
|
||||
`refresh-deps-cache` workflow logs.
|
||||
- `PVC already seeded for deps v1 (stamp present), skipping.` —
|
||||
expected steady-state on warm pods.
|
||||
- `=== Seeding cargo-target-cpu PVC from prebuilt deps cache ===` +
|
||||
`=== Seed complete ===` — first run on this PVC, working as designed.
|
||||
|
||||
**Symptom: cargo recompiles many third-party deps anyway despite the
|
||||
seed completing.**
|
||||
|
||||
Cargo's fingerprint check is sensitive to: rustc version, the entire
|
||||
target.rustflags list, `RUSTFLAGS` env var, source code mtimes (on
|
||||
build.rs), feature flags. The commit-SHA used to build the deps cache
|
||||
image must match the live compile workflow's flag set. The most common
|
||||
cause of full-rebuild is a `RUSTFLAGS` mismatch between the deps-cache
|
||||
image build and the compile pod (e.g. mold vs wild linker swap).
|
||||
|
||||
To diagnose, look at the cargo build output: lines like
|
||||
`Compiling tokio v1.x.x` for crates that were obviously in the cache
|
||||
indicate fingerprint mismatch.
|
||||
|
||||
**Symptom: PVC fills up faster than before.**
|
||||
|
||||
Each seed adds ~6-8 GB to the PVC. The compile-services step has a
|
||||
prune-at-25GB guard that wipes `release/` and `debug/` if usage
|
||||
exceeds 25GB — when that fires, the next run re-seeds (stamp file is
|
||||
in PVC root, survives the prune). To raise the threshold, edit the
|
||||
`TARGET_SIZE_MB -gt 25000` check in compile-and-deploy-template.yaml.
|
||||
|
||||
## Assumptions that break the cache
|
||||
|
||||
- **rustc toolchain version mismatch** between deps-cache image and
|
||||
compile pod — cargo will full-rebuild. Resolution: rebuild the
|
||||
deps-cache image (refresh-deps-cache).
|
||||
- **Different `RUSTFLAGS` / `target.<triple>.rustflags`** between
|
||||
image build and live compile (e.g. switching CPU target features)
|
||||
— full rebuild. Resolution: rebuild deps-cache image after the
|
||||
rustflags change is committed.
|
||||
- **Cargo.lock SHA divergence** > a few weeks old — most rlibs no
|
||||
longer match. Resolution: rebuild deps-cache image (refresh nightly,
|
||||
or trigger manually).
|
||||
- **Linker swap (e.g. mold -> wild)** — link-output is cached at the
|
||||
fingerprint level; changing linker doesn't trigger rebuild but may
|
||||
produce different final binaries. Not a cache invalidation, just a
|
||||
link-time difference.
|
||||
|
||||
## When to delete this entirely
|
||||
|
||||
If the underlying compile time drops to <30s in cold-start (e.g. via a
|
||||
much faster registry or sccache covering 100% of deps), the
|
||||
deps-cache complexity is no longer worth maintaining. To remove:
|
||||
|
||||
1. Remove the `seed-deps-cache` initContainer block from
|
||||
`compile-and-deploy-template.yaml` (the compile script tolerates
|
||||
an unseeded PVC — cargo just compiles deps from scratch).
|
||||
2. Delete `infra/docker/Dockerfile.ci-deps-cache` and
|
||||
`infra/k8s/argo/refresh-deps-cache-template.yaml`.
|
||||
3. Delete the published image:
|
||||
`crane delete gitlab-registry.foxhunt.svc.cluster.local:5000/root/foxhunt/ci-builder-cpu-with-deps:nightly`.
|
||||
@@ -1,495 +0,0 @@
|
||||
# Alpha walk-forward CV workflow.
|
||||
#
|
||||
# Trains the Mamba2+stacker alpha logits cache on the full 9Q fxcache,
|
||||
# then runs `alpha_baseline` 9 times across disjoint 1.9M-bar windows
|
||||
# (one per quarter). Each fold trains its own execution-policy DQN
|
||||
# from scratch on the front of the window and backtests on the back
|
||||
# (purged train/eval split via `--train-frac`).
|
||||
#
|
||||
# Sequential by design: 9 folds chain in the DAG so a single L40S node
|
||||
# stays warm across the run. With --decision-stride 200 and --horizon
|
||||
# 1200, each fold's DQN training is ~5-10 min on L40S; total wall ~60-90
|
||||
# min for stacker + 9 folds.
|
||||
#
|
||||
# DAG:
|
||||
# ensure-binary ──> ensure-fxcache ──> stacker-train ──> fold-0 ──> ... ──> fold-8
|
||||
#
|
||||
# Usage:
|
||||
# argo submit -n foxhunt --from=wftmpl/alpha-cv \
|
||||
# -p commit-sha=HEAD -p git-branch=main -p decision-stride=200
|
||||
apiVersion: argoproj.io/v1alpha1
|
||||
kind: WorkflowTemplate
|
||||
metadata:
|
||||
name: alpha-cv
|
||||
namespace: foxhunt
|
||||
labels:
|
||||
app.kubernetes.io/name: alpha-cv
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
app.kubernetes.io/component: train
|
||||
spec:
|
||||
entrypoint: pipeline
|
||||
serviceAccountName: argo-workflow
|
||||
archiveLogs: true
|
||||
podMetadata:
|
||||
labels:
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
app.kubernetes.io/component: train
|
||||
securityContext:
|
||||
fsGroup: 0
|
||||
ttlStrategy:
|
||||
secondsAfterCompletion: 3600
|
||||
activeDeadlineSeconds: 28800
|
||||
|
||||
arguments:
|
||||
parameters:
|
||||
- name: commit-sha
|
||||
value: HEAD
|
||||
- name: git-branch
|
||||
value: main
|
||||
- name: cuda-compute-cap
|
||||
value: "89"
|
||||
- name: gpu-pool
|
||||
value: ci-training-l40s
|
||||
- name: symbol
|
||||
value: ES.FUT
|
||||
# Match the fxcache build params (cluster's 9Q fxcache was built
|
||||
# with these on 2026-05-16; see commit e7ce4395e).
|
||||
- name: imbalance-bar-threshold
|
||||
value: "20.0"
|
||||
- name: imbalance-bar-ewma-alpha
|
||||
value: "0.1"
|
||||
- name: volume-bar-size
|
||||
value: "100"
|
||||
- name: data-source
|
||||
value: "mbp10"
|
||||
# Stacker training params (mirror local 2Q smoke).
|
||||
- name: stacker-horizon
|
||||
value: "6000"
|
||||
- name: stacker-seq-len
|
||||
value: "32"
|
||||
- name: stacker-hidden-dim
|
||||
value: "64"
|
||||
- name: stacker-state-dim
|
||||
value: "16"
|
||||
- name: stacker-epochs
|
||||
value: "5"
|
||||
- name: stacker-batch-size
|
||||
value: "128"
|
||||
- name: stacker-lr
|
||||
value: "3e-3"
|
||||
- name: stacker-train-frac
|
||||
value: "0.8"
|
||||
# CV / alpha_baseline params — defaults match the validated 2Q
|
||||
# config that produced +1.78 Sharpe at quarter-tick cost.
|
||||
- name: decision-stride
|
||||
value: "200"
|
||||
- name: fold-window
|
||||
value: "1900000"
|
||||
- name: train-frac
|
||||
value: "0.6"
|
||||
- name: window-k
|
||||
value: "16"
|
||||
- name: horizon
|
||||
value: "1200"
|
||||
- name: n-train-par
|
||||
value: "25"
|
||||
- name: n-train-episodes
|
||||
value: "8000"
|
||||
|
||||
volumes:
|
||||
- name: git-ssh-key
|
||||
secret:
|
||||
secretName: argo-git-ssh-key
|
||||
defaultMode: 256
|
||||
- name: training-data
|
||||
persistentVolumeClaim:
|
||||
claimName: training-data-pvc
|
||||
- name: cargo-target-cuda
|
||||
persistentVolumeClaim:
|
||||
claimName: cargo-target-cuda
|
||||
- name: feature-cache
|
||||
persistentVolumeClaim:
|
||||
claimName: feature-cache-pvc
|
||||
|
||||
templates:
|
||||
# ── DAG ──
|
||||
- name: pipeline
|
||||
dag:
|
||||
tasks:
|
||||
- name: ensure-binary
|
||||
template: ensure-binary
|
||||
- name: stacker-train
|
||||
template: stacker-train
|
||||
dependencies: [ensure-binary]
|
||||
arguments:
|
||||
parameters:
|
||||
- name: sha
|
||||
value: "{{tasks.ensure-binary.outputs.parameters.sha}}"
|
||||
- name: fold-0
|
||||
template: alpha-cv-fold
|
||||
dependencies: [stacker-train]
|
||||
arguments:
|
||||
parameters:
|
||||
- name: sha
|
||||
value: "{{tasks.ensure-binary.outputs.parameters.sha}}"
|
||||
- name: fold-idx
|
||||
value: "0"
|
||||
- name: offset
|
||||
value: "0"
|
||||
- name: fold-1
|
||||
template: alpha-cv-fold
|
||||
dependencies: [fold-0]
|
||||
arguments:
|
||||
parameters:
|
||||
- name: sha
|
||||
value: "{{tasks.ensure-binary.outputs.parameters.sha}}"
|
||||
- name: fold-idx
|
||||
value: "1"
|
||||
- name: offset
|
||||
value: "1900000"
|
||||
- name: fold-2
|
||||
template: alpha-cv-fold
|
||||
dependencies: [fold-1]
|
||||
arguments:
|
||||
parameters:
|
||||
- name: sha
|
||||
value: "{{tasks.ensure-binary.outputs.parameters.sha}}"
|
||||
- name: fold-idx
|
||||
value: "2"
|
||||
- name: offset
|
||||
value: "3800000"
|
||||
- name: fold-3
|
||||
template: alpha-cv-fold
|
||||
dependencies: [fold-2]
|
||||
arguments:
|
||||
parameters:
|
||||
- name: sha
|
||||
value: "{{tasks.ensure-binary.outputs.parameters.sha}}"
|
||||
- name: fold-idx
|
||||
value: "3"
|
||||
- name: offset
|
||||
value: "5700000"
|
||||
- name: fold-4
|
||||
template: alpha-cv-fold
|
||||
dependencies: [fold-3]
|
||||
arguments:
|
||||
parameters:
|
||||
- name: sha
|
||||
value: "{{tasks.ensure-binary.outputs.parameters.sha}}"
|
||||
- name: fold-idx
|
||||
value: "4"
|
||||
- name: offset
|
||||
value: "7600000"
|
||||
- name: fold-5
|
||||
template: alpha-cv-fold
|
||||
dependencies: [fold-4]
|
||||
arguments:
|
||||
parameters:
|
||||
- name: sha
|
||||
value: "{{tasks.ensure-binary.outputs.parameters.sha}}"
|
||||
- name: fold-idx
|
||||
value: "5"
|
||||
- name: offset
|
||||
value: "9500000"
|
||||
- name: fold-6
|
||||
template: alpha-cv-fold
|
||||
dependencies: [fold-5]
|
||||
arguments:
|
||||
parameters:
|
||||
- name: sha
|
||||
value: "{{tasks.ensure-binary.outputs.parameters.sha}}"
|
||||
- name: fold-idx
|
||||
value: "6"
|
||||
- name: offset
|
||||
value: "11400000"
|
||||
- name: fold-7
|
||||
template: alpha-cv-fold
|
||||
dependencies: [fold-6]
|
||||
arguments:
|
||||
parameters:
|
||||
- name: sha
|
||||
value: "{{tasks.ensure-binary.outputs.parameters.sha}}"
|
||||
- name: fold-idx
|
||||
value: "7"
|
||||
- name: offset
|
||||
value: "13300000"
|
||||
- name: fold-8
|
||||
template: alpha-cv-fold
|
||||
dependencies: [fold-7]
|
||||
arguments:
|
||||
parameters:
|
||||
- name: sha
|
||||
value: "{{tasks.ensure-binary.outputs.parameters.sha}}"
|
||||
- name: fold-idx
|
||||
value: "8"
|
||||
- name: offset
|
||||
value: "15200000"
|
||||
|
||||
# ── ensure-binary: compile alpha_train_stacker + alpha_baseline ──
|
||||
- name: ensure-binary
|
||||
outputs:
|
||||
parameters:
|
||||
- name: sha
|
||||
valueFrom:
|
||||
path: /tmp/sha
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: ci-compile-cpu
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
tolerations:
|
||||
- key: node.cilium.io/agent-not-ready
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
container:
|
||||
image: gitlab-registry.foxhunt.svc.cluster.local:5000/root/foxhunt/ci-builder:latest
|
||||
imagePullPolicy: Always
|
||||
command: ["/bin/bash", "-c"]
|
||||
env:
|
||||
- name: SQLX_OFFLINE
|
||||
value: "true"
|
||||
- name: CARGO_TERM_COLOR
|
||||
value: always
|
||||
- name: CARGO_TARGET_DIR
|
||||
value: /cargo-target
|
||||
- name: CARGO_HOME
|
||||
value: /cargo-target/cargo-home
|
||||
- name: RUSTC_WRAPPER
|
||||
value: sccache
|
||||
- name: SCCACHE_DIR
|
||||
value: /cargo-target/sccache
|
||||
- name: SCCACHE_CACHE_SIZE
|
||||
value: "40G"
|
||||
- name: CARGO_INCREMENTAL
|
||||
value: "0"
|
||||
- name: GITLAB_PAT
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: gitlab-pat
|
||||
key: token
|
||||
resources:
|
||||
requests:
|
||||
cpu: "14"
|
||||
memory: 32Gi
|
||||
limits:
|
||||
cpu: "30"
|
||||
memory: 64Gi
|
||||
volumeMounts:
|
||||
- name: git-ssh-key
|
||||
mountPath: /etc/git-ssh
|
||||
readOnly: true
|
||||
- name: cargo-target-cuda
|
||||
mountPath: /cargo-target
|
||||
- name: training-data
|
||||
mountPath: /data
|
||||
args:
|
||||
- |
|
||||
set -e
|
||||
SHA="{{workflow.parameters.commit-sha}}"
|
||||
BRANCH="{{workflow.parameters.git-branch}}"
|
||||
mkdir -p ~/.ssh
|
||||
cp /etc/git-ssh/ssh-privatekey ~/.ssh/id_ed25519
|
||||
chmod 600 ~/.ssh/id_ed25519
|
||||
printf 'Host *\n StrictHostKeyChecking no\n UserKnownHostsFile /dev/null\n' > ~/.ssh/config
|
||||
chmod 600 ~/.ssh/config
|
||||
REPO="ssh://git@gitlab-gitlab-shell.foxhunt.svc.cluster.local:2222/root/foxhunt.git"
|
||||
BUILD="/cargo-target/src"
|
||||
git config --global --add safe.directory "$BUILD"
|
||||
|
||||
if [ "$SHA" = "HEAD" ]; then
|
||||
if [ -d "$BUILD/.git" ]; then
|
||||
cd "$BUILD"; git fetch origin; SHA=$(git rev-parse "origin/$BRANCH"); cd /
|
||||
else
|
||||
git clone --filter=blob:none "$REPO" "$BUILD"; cd "$BUILD"
|
||||
git checkout "origin/$BRANCH"; SHA=$(git rev-parse HEAD); cd /
|
||||
fi
|
||||
fi
|
||||
SHORT_SHA=$(echo "$SHA" | cut -c1-9)
|
||||
echo "Resolved SHA: $SHA (short: $SHORT_SHA)"
|
||||
|
||||
BIN_DIR="/data/bin/$SHORT_SHA"
|
||||
BINARIES="alpha_train_stacker alpha_baseline"
|
||||
ALL_CACHED=true
|
||||
for bin in $BINARIES; do
|
||||
[ ! -x "$BIN_DIR/$bin" ] && ALL_CACHED=false && break
|
||||
done
|
||||
|
||||
if [ "$ALL_CACHED" = "true" ]; then
|
||||
echo "=== Cache HIT ==="
|
||||
ls -lh "$BIN_DIR/"
|
||||
echo "$SHORT_SHA" > /tmp/sha
|
||||
exit 0
|
||||
fi
|
||||
echo "=== Cache MISS: compiling alpha binaries for $SHORT_SHA ==="
|
||||
|
||||
if [ -d "$BUILD/.git" ]; then
|
||||
cd "$BUILD"; git fetch origin
|
||||
CURRENT=$(git rev-parse HEAD 2>/dev/null || echo "none")
|
||||
if [ "$CURRENT" != "$SHA" ]; then
|
||||
git checkout --force "$SHA"; git clean -fd
|
||||
fi
|
||||
else
|
||||
git clone --filter=blob:none "$REPO" "$BUILD"; cd "$BUILD"; git checkout "$SHA"
|
||||
fi
|
||||
|
||||
export PATH="${CARGO_HOME}/bin:${PATH}"
|
||||
export CUDA_COMPUTE_CAP={{workflow.parameters.cuda-compute-cap}}
|
||||
|
||||
# alpha_train_stacker lives in ml-alpha, alpha_baseline in ml.
|
||||
echo "Building alpha_train_stacker (ml-alpha)..."
|
||||
cargo build --release -p ml-alpha --example alpha_train_stacker
|
||||
echo "Building alpha_baseline (ml)..."
|
||||
cargo build --release -p ml --features ml/cuda --example alpha_baseline
|
||||
|
||||
mkdir -p "$BIN_DIR"
|
||||
cp "$CARGO_TARGET_DIR/release/examples/alpha_train_stacker" "$BIN_DIR/"
|
||||
cp "$CARGO_TARGET_DIR/release/examples/alpha_baseline" "$BIN_DIR/"
|
||||
strip "$BIN_DIR/"*
|
||||
# alpha_baseline reads config/ml/alpha_fill_coeffs.json as a
|
||||
# required --fill-coeffs input. The source tree's copy travels
|
||||
# with the binaries so the GPU fold pods don't need to clone.
|
||||
cp "$BUILD/config/ml/alpha_fill_coeffs.json" "$BIN_DIR/" || true
|
||||
ls -lh "$BIN_DIR/"
|
||||
|
||||
cd /data/bin
|
||||
ls -1t | tail -n +6 | while read -r old; do
|
||||
echo "Pruning old cache: $old"; rm -rf "$old"
|
||||
done
|
||||
echo "$SHORT_SHA" > /tmp/sha
|
||||
|
||||
# ── stacker-train: train Mamba2+stacker on the 9Q fxcache ──
|
||||
- name: stacker-train
|
||||
inputs:
|
||||
parameters:
|
||||
- name: sha
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: "{{workflow.parameters.gpu-pool}}"
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
tolerations:
|
||||
- key: nvidia.com/gpu
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
- key: node.cilium.io/agent-not-ready
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
container:
|
||||
image: gitlab-registry.foxhunt.svc.cluster.local:5000/root/foxhunt/training-runtime:latest
|
||||
imagePullPolicy: Always
|
||||
command: ["/bin/bash", "-c"]
|
||||
env:
|
||||
- name: RUST_LOG
|
||||
value: info
|
||||
- name: SQLX_OFFLINE
|
||||
value: "true"
|
||||
resources:
|
||||
requests:
|
||||
cpu: "8"
|
||||
memory: 32Gi
|
||||
nvidia.com/gpu: "1"
|
||||
limits:
|
||||
cpu: "16"
|
||||
memory: 80Gi
|
||||
nvidia.com/gpu: "1"
|
||||
volumeMounts:
|
||||
- name: training-data
|
||||
mountPath: /data
|
||||
- name: feature-cache
|
||||
mountPath: /feature-cache
|
||||
args:
|
||||
- |
|
||||
set -e
|
||||
SHA="{{inputs.parameters.sha}}"
|
||||
BIN="/data/bin/$SHA/alpha_train_stacker"
|
||||
FXCACHE=$(ls -t /feature-cache/*.fxcache | head -1)
|
||||
ALPHA_OUT="/feature-cache/alpha_logits_cache_9q.bin"
|
||||
echo "Using fxcache: $FXCACHE"
|
||||
echo "Output: $ALPHA_OUT"
|
||||
|
||||
"$BIN" \
|
||||
--fxcache-path "$FXCACHE" \
|
||||
--horizon {{workflow.parameters.stacker-horizon}} \
|
||||
--seq-len {{workflow.parameters.stacker-seq-len}} \
|
||||
--hidden-dim {{workflow.parameters.stacker-hidden-dim}} \
|
||||
--state-dim {{workflow.parameters.stacker-state-dim}} \
|
||||
--epochs {{workflow.parameters.stacker-epochs}} \
|
||||
--batch-size {{workflow.parameters.stacker-batch-size}} \
|
||||
--lr {{workflow.parameters.stacker-lr}} \
|
||||
--train-frac {{workflow.parameters.stacker-train-frac}} \
|
||||
--alpha-cache-out "$ALPHA_OUT"
|
||||
ls -lh "$ALPHA_OUT"
|
||||
|
||||
# ── alpha-cv-fold: run alpha_baseline on one window ──
|
||||
- name: alpha-cv-fold
|
||||
inputs:
|
||||
parameters:
|
||||
- name: sha
|
||||
- name: fold-idx
|
||||
- name: offset
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: "{{workflow.parameters.gpu-pool}}"
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
tolerations:
|
||||
- key: nvidia.com/gpu
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
- key: node.cilium.io/agent-not-ready
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
container:
|
||||
image: gitlab-registry.foxhunt.svc.cluster.local:5000/root/foxhunt/training-runtime:latest
|
||||
imagePullPolicy: Always
|
||||
command: ["/bin/bash", "-c"]
|
||||
env:
|
||||
- name: RUST_LOG
|
||||
value: info
|
||||
- name: SQLX_OFFLINE
|
||||
value: "true"
|
||||
resources:
|
||||
requests:
|
||||
cpu: "8"
|
||||
memory: 32Gi
|
||||
nvidia.com/gpu: "1"
|
||||
limits:
|
||||
cpu: "16"
|
||||
memory: 80Gi
|
||||
nvidia.com/gpu: "1"
|
||||
volumeMounts:
|
||||
- name: training-data
|
||||
mountPath: /data
|
||||
- name: feature-cache
|
||||
mountPath: /feature-cache
|
||||
args:
|
||||
- |
|
||||
set -e
|
||||
SHA="{{inputs.parameters.sha}}"
|
||||
FOLD={{inputs.parameters.fold-idx}}
|
||||
OFFSET={{inputs.parameters.offset}}
|
||||
BIN="/data/bin/$SHA/alpha_baseline"
|
||||
FXCACHE=$(ls -t /feature-cache/*.fxcache | head -1)
|
||||
ALPHA="/feature-cache/alpha_logits_cache_9q.bin"
|
||||
# alpha_fill_coeffs.json travels with the source tree; the
|
||||
# binary needs it next to itself or via explicit --fill-coeffs.
|
||||
# The ci-builder image bakes it under /opt/foxhunt/config/ml.
|
||||
FILL="/data/bin/$SHA/alpha_fill_coeffs.json"
|
||||
if [ ! -f "$FILL" ]; then
|
||||
# Fall back to copying from the cloned source on the
|
||||
# cargo-target PVC mounted in ensure-binary; we don't have
|
||||
# that PVC here, so embed via fxcache PVC pre-stage.
|
||||
FILL="/feature-cache/alpha_fill_coeffs.json"
|
||||
fi
|
||||
OUT="/feature-cache/cv_results_9fold/fold_${FOLD}.json"
|
||||
mkdir -p /feature-cache/cv_results_9fold
|
||||
echo "Fold $FOLD: offset=$OFFSET out=$OUT"
|
||||
|
||||
"$BIN" \
|
||||
--fxcache-path "$FXCACHE" \
|
||||
--alpha-cache "$ALPHA" \
|
||||
--fill-coeffs "$FILL" \
|
||||
--data-start-offset $OFFSET \
|
||||
--max-snapshots {{workflow.parameters.fold-window}} \
|
||||
--train-frac {{workflow.parameters.train-frac}} \
|
||||
--window-k {{workflow.parameters.window-k}} \
|
||||
--decision-stride {{workflow.parameters.decision-stride}} \
|
||||
--horizon {{workflow.parameters.horizon}} \
|
||||
--n-train-par {{workflow.parameters.n-train-par}} \
|
||||
--n-train-episodes {{workflow.parameters.n-train-episodes}} \
|
||||
--out-path "$OUT"
|
||||
ls -lh "$OUT"
|
||||
@@ -1,484 +0,0 @@
|
||||
# Alpha perception trainer workflow.
|
||||
#
|
||||
# Stacked Mamba2 -> CfC -> heads supervised pretrain on MBP-10 from the
|
||||
# training-data PVC. Emits `alpha_train_summary.json` to MinIO-backed
|
||||
# feature-cache PVC with per-horizon val AUC for monitoring. Runs on
|
||||
# a single L40S; ~30-90 min wall depending on
|
||||
# n-train-seqs.
|
||||
#
|
||||
# DAG:
|
||||
# ensure-binary ──> train
|
||||
#
|
||||
# Usage:
|
||||
# argo submit -n foxhunt --from=wftmpl/alpha-perception \
|
||||
# -p commit-sha=HEAD -p git-branch=ml-alpha-phase-a
|
||||
apiVersion: argoproj.io/v1alpha1
|
||||
kind: WorkflowTemplate
|
||||
metadata:
|
||||
name: alpha-perception
|
||||
namespace: foxhunt
|
||||
labels:
|
||||
app.kubernetes.io/name: alpha-perception
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
app.kubernetes.io/component: train
|
||||
spec:
|
||||
entrypoint: pipeline
|
||||
serviceAccountName: argo-workflow
|
||||
archiveLogs: true
|
||||
podMetadata:
|
||||
labels:
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
app.kubernetes.io/component: train
|
||||
securityContext:
|
||||
fsGroup: 0
|
||||
ttlStrategy:
|
||||
secondsAfterCompletion: 3600
|
||||
activeDeadlineSeconds: 14400
|
||||
|
||||
arguments:
|
||||
parameters:
|
||||
- name: commit-sha
|
||||
value: HEAD
|
||||
- name: git-branch
|
||||
value: ml-alpha-phase-a
|
||||
- name: cuda-compute-cap
|
||||
value: "89"
|
||||
- name: gpu-pool
|
||||
value: ci-training-l40s
|
||||
# Trainer hyperparameters (mirror the validated smoke config).
|
||||
- name: epochs
|
||||
value: "5"
|
||||
- name: multi-resolution
|
||||
value: "1:32"
|
||||
- name: mamba2-state-dim
|
||||
value: "16"
|
||||
- name: lr-cfc
|
||||
value: "3e-3"
|
||||
- name: lr-mamba2
|
||||
value: "1e-3"
|
||||
- name: n-train-seqs
|
||||
value: "8000"
|
||||
- name: n-val-seqs
|
||||
value: "1000"
|
||||
- name: seed
|
||||
value: "16962"
|
||||
- name: batch-size
|
||||
value: "1"
|
||||
- name: auto-horizon-weights
|
||||
value: "false"
|
||||
- name: instrument-mode
|
||||
value: "all"
|
||||
- name: early-stop-metric
|
||||
value: "mean_auc"
|
||||
- name: early-stop-patience
|
||||
value: "5"
|
||||
- name: cv-fold
|
||||
value: "0"
|
||||
- name: cv-n-folds
|
||||
value: "1"
|
||||
- name: cv-train-window
|
||||
value: "0"
|
||||
- name: smoothness-base-lambda
|
||||
value: "0.0"
|
||||
- name: kernel-step-trace-enable
|
||||
value: "false"
|
||||
- name: kernel-step-trace-path
|
||||
value: ""
|
||||
|
||||
volumes:
|
||||
- name: git-ssh-key
|
||||
secret:
|
||||
secretName: argo-git-ssh-key
|
||||
defaultMode: 256
|
||||
- name: training-data
|
||||
persistentVolumeClaim:
|
||||
claimName: training-data-pvc
|
||||
- name: cargo-target-cuda
|
||||
persistentVolumeClaim:
|
||||
claimName: cargo-target-cuda
|
||||
- name: feature-cache
|
||||
persistentVolumeClaim:
|
||||
claimName: feature-cache-pvc
|
||||
|
||||
templates:
|
||||
# ── DAG ──
|
||||
#
|
||||
# check-cache (tiny alpine pod on platform pool, ~3 sec) inspects
|
||||
# the training-data PVC for /data/bin/$SHA/alpha_train. If present
|
||||
# (cache hit), ensure-binary is SKIPPED via `when:` — the
|
||||
# ~4.8GB ci-builder image pull + sccache compile cycle are avoided
|
||||
# entirely. train still runs, sourcing the binary path from
|
||||
# check-cache's SHA output.
|
||||
#
|
||||
# warmup-gpu fires unconditionally in parallel — triggers the L40S
|
||||
# autoscaler so the node is warm by the time train needs it,
|
||||
# whether or not we compile.
|
||||
- name: pipeline
|
||||
dag:
|
||||
tasks:
|
||||
- name: check-cache
|
||||
template: check-cache
|
||||
- name: ensure-binary
|
||||
template: ensure-binary
|
||||
dependencies: [check-cache]
|
||||
when: "{{tasks.check-cache.outputs.parameters.cache}} == miss"
|
||||
arguments:
|
||||
parameters:
|
||||
- name: sha
|
||||
value: "{{tasks.check-cache.outputs.parameters.sha}}"
|
||||
- name: warmup-gpu
|
||||
template: warmup-gpu
|
||||
- name: train
|
||||
template: train
|
||||
dependencies: [check-cache, ensure-binary]
|
||||
arguments:
|
||||
parameters:
|
||||
- name: sha
|
||||
value: "{{tasks.check-cache.outputs.parameters.sha}}"
|
||||
|
||||
# ── check-cache: probe the training-data PVC for an existing binary ──
|
||||
#
|
||||
# Runs on the platform pool (always up, no autoscaler delay). Tiny
|
||||
# alpine pod, ~3 sec end-to-end. Emits two output parameters:
|
||||
# - sha = the short SHA used for binary cache keying
|
||||
# - cache = "hit" if /data/bin/$SHA/alpha_train exists else "miss"
|
||||
#
|
||||
# Resolves `commit-sha=HEAD` by reading /tmp/head-sha from a
|
||||
# ConfigMap... actually no, simpler: the submission script
|
||||
# (scripts/argo-alpha-perception.sh) pre-resolves HEAD via local
|
||||
# git so commit-sha is always a real SHA in the workflow. This
|
||||
# pod just trims to short SHA and stats the binary.
|
||||
- name: check-cache
|
||||
outputs:
|
||||
parameters:
|
||||
- name: sha
|
||||
valueFrom:
|
||||
path: /tmp/sha
|
||||
- name: cache
|
||||
valueFrom:
|
||||
path: /tmp/cache
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: platform
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
tolerations:
|
||||
- key: node.cilium.io/agent-not-ready
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
container:
|
||||
image: alpine:3.20
|
||||
command: ["/bin/sh", "-c"]
|
||||
resources:
|
||||
requests:
|
||||
cpu: "50m"
|
||||
memory: 32Mi
|
||||
limits:
|
||||
cpu: "100m"
|
||||
memory: 64Mi
|
||||
volumeMounts:
|
||||
- name: training-data
|
||||
mountPath: /data
|
||||
readOnly: true
|
||||
args:
|
||||
- |
|
||||
set -e
|
||||
SHA="{{workflow.parameters.commit-sha}}"
|
||||
if [ "$SHA" = "HEAD" ]; then
|
||||
echo "ERROR: commit-sha cannot be HEAD inside the workflow."
|
||||
echo " Resolve via the submission script (scripts/argo-alpha-perception.sh)."
|
||||
exit 1
|
||||
fi
|
||||
SHORT_SHA=$(echo "$SHA" | cut -c1-9)
|
||||
echo "$SHORT_SHA" > /tmp/sha
|
||||
|
||||
BIN="/data/bin/$SHORT_SHA/alpha_train"
|
||||
if [ -x "$BIN" ]; then
|
||||
SIZE=$(stat -c %s "$BIN")
|
||||
echo "Cache HIT: $BIN ($SIZE bytes)"
|
||||
echo "hit" > /tmp/cache
|
||||
else
|
||||
echo "Cache MISS: $BIN not present, ensure-binary will compile"
|
||||
ls -lh "/data/bin/" 2>/dev/null | head -10 || echo " (no /data/bin/ directory yet)"
|
||||
echo "miss" > /tmp/cache
|
||||
fi
|
||||
|
||||
# ── ensure-binary: compile `alpha_train` example via sccache ──
|
||||
# Only runs when check-cache returned cache=miss. Outputs the SHA
|
||||
# for consistency, but the DAG sources SHA from check-cache directly.
|
||||
- name: ensure-binary
|
||||
inputs:
|
||||
parameters:
|
||||
- name: sha
|
||||
outputs:
|
||||
parameters:
|
||||
- name: sha
|
||||
valueFrom:
|
||||
path: /tmp/sha
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: ci-compile-cpu
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
tolerations:
|
||||
- key: node.cilium.io/agent-not-ready
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
container:
|
||||
image: gitlab-registry.foxhunt.svc.cluster.local:5000/root/foxhunt/ci-builder:latest
|
||||
imagePullPolicy: Always
|
||||
command: ["/bin/bash", "-c"]
|
||||
env:
|
||||
- name: SQLX_OFFLINE
|
||||
value: "true"
|
||||
- name: CARGO_TERM_COLOR
|
||||
value: always
|
||||
- name: CARGO_TARGET_DIR
|
||||
value: /cargo-target
|
||||
- name: CARGO_HOME
|
||||
value: /cargo-target/cargo-home
|
||||
- name: RUSTC_WRAPPER
|
||||
value: sccache
|
||||
- name: SCCACHE_DIR
|
||||
value: /cargo-target/sccache
|
||||
- name: SCCACHE_CACHE_SIZE
|
||||
value: "40G"
|
||||
- name: CARGO_INCREMENTAL
|
||||
value: "0"
|
||||
- name: CUDA_COMPUTE_CAP
|
||||
value: "{{workflow.parameters.cuda-compute-cap}}"
|
||||
resources:
|
||||
requests:
|
||||
cpu: "14"
|
||||
memory: 32Gi
|
||||
limits:
|
||||
cpu: "30"
|
||||
memory: 64Gi
|
||||
volumeMounts:
|
||||
- name: git-ssh-key
|
||||
mountPath: /etc/git-ssh
|
||||
readOnly: true
|
||||
- name: cargo-target-cuda
|
||||
mountPath: /cargo-target
|
||||
- name: training-data
|
||||
mountPath: /data
|
||||
args:
|
||||
- |
|
||||
set -e
|
||||
SHA="{{workflow.parameters.commit-sha}}"
|
||||
BRANCH="{{workflow.parameters.git-branch}}"
|
||||
mkdir -p ~/.ssh
|
||||
cp /etc/git-ssh/ssh-privatekey ~/.ssh/id_ed25519
|
||||
chmod 600 ~/.ssh/id_ed25519
|
||||
printf 'Host *\n StrictHostKeyChecking no\n UserKnownHostsFile /dev/null\n' > ~/.ssh/config
|
||||
chmod 600 ~/.ssh/config
|
||||
REPO="ssh://git@gitlab-gitlab-shell.foxhunt.svc.cluster.local:2222/root/foxhunt.git"
|
||||
BUILD="/cargo-target/src"
|
||||
git config --global --add safe.directory "$BUILD"
|
||||
|
||||
if [ "$SHA" = "HEAD" ]; then
|
||||
if [ -d "$BUILD/.git" ]; then
|
||||
cd "$BUILD"; git fetch origin; SHA=$(git rev-parse "origin/$BRANCH"); cd /
|
||||
else
|
||||
git clone --filter=blob:none "$REPO" "$BUILD"; cd "$BUILD"
|
||||
git checkout "origin/$BRANCH"; SHA=$(git rev-parse HEAD); cd /
|
||||
fi
|
||||
fi
|
||||
SHORT_SHA=$(echo "$SHA" | cut -c1-9)
|
||||
echo "Resolved SHA: $SHA (short: $SHORT_SHA)"
|
||||
|
||||
BIN_DIR="/data/bin/$SHORT_SHA"
|
||||
if [ -x "$BIN_DIR/alpha_train" ]; then
|
||||
echo "=== Cache HIT: $BIN_DIR/alpha_train ==="
|
||||
ls -lh "$BIN_DIR/"
|
||||
echo "$SHORT_SHA" > /tmp/sha
|
||||
exit 0
|
||||
fi
|
||||
echo "=== Cache MISS: compiling alpha_train for $SHORT_SHA ==="
|
||||
|
||||
if [ -d "$BUILD/.git" ]; then
|
||||
cd "$BUILD"; git fetch origin
|
||||
CURRENT=$(git rev-parse HEAD 2>/dev/null || echo "none")
|
||||
if [ "$CURRENT" != "$SHA" ]; then
|
||||
git checkout --force "$SHA"; git clean -fd
|
||||
fi
|
||||
else
|
||||
git clone --filter=blob:none "$REPO" "$BUILD"; cd "$BUILD"; git checkout "$SHA"
|
||||
fi
|
||||
|
||||
export PATH="${CARGO_HOME}/bin:${PATH}"
|
||||
|
||||
FEATURES_FLAG=""
|
||||
if [ "{{workflow.parameters.kernel-step-trace-enable}}" = "true" ]; then
|
||||
FEATURES_FLAG="--features kernel-step-trace"
|
||||
echo "Building with --features kernel-step-trace"
|
||||
# Bust cache: feature-on binary differs from feature-off at same SHA.
|
||||
if [ -x "$BIN_DIR/alpha_train" ]; then
|
||||
echo "Removing cached binary (feature change requires rebuild)"
|
||||
rm -f "$BIN_DIR/alpha_train"
|
||||
fi
|
||||
fi
|
||||
|
||||
echo "Building alpha_train (ml-alpha)..."
|
||||
cargo build --release $FEATURES_FLAG -p ml-alpha --example alpha_train
|
||||
|
||||
mkdir -p "$BIN_DIR"
|
||||
cp "$CARGO_TARGET_DIR/release/examples/alpha_train" "$BIN_DIR/"
|
||||
strip "$BIN_DIR/alpha_train"
|
||||
ls -lh "$BIN_DIR/"
|
||||
|
||||
# Prune old cache (keep last 5)
|
||||
cd /data/bin
|
||||
ls -1t | tail -n +6 | while read -r old; do
|
||||
echo "Pruning old cache: $old"; rm -rf "$old"
|
||||
done
|
||||
echo "$SHORT_SHA" > /tmp/sha
|
||||
|
||||
# ── warmup-gpu: trigger L40S autoscaler scale-up in parallel ──
|
||||
#
|
||||
# Schedules a tiny CPU-only pod on the gpu-pool's nodeSelector.
|
||||
# Kubernetes sees an unschedulable pod (no node currently in the
|
||||
# pool), autoscaler scales the pool 0 → 1. The pod exits the
|
||||
# instant it lands; the node enters scaledown-grace.
|
||||
#
|
||||
# Cluster autoscaler config (see infra/modules/kapsule/main.tf):
|
||||
# scale_down_delay_after_add = "10m" -- 10m before any scaledown consideration
|
||||
# scale_down_unneeded_time = "10m" -- 10m empty before action
|
||||
# Effective grace window is up to ~20m, which always spans
|
||||
# ensure-binary compile (sccache-hit ~10s through cold ~15m).
|
||||
# train lands on a hot node with zero provisioning latency.
|
||||
#
|
||||
# No GPU resource request — that would compete with `train` for
|
||||
# the single GPU and serialise the steps. The nodeSelector +
|
||||
# nvidia.com/gpu toleration are enough to force placement on the
|
||||
# right pool.
|
||||
- name: warmup-gpu
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: "{{workflow.parameters.gpu-pool}}"
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
tolerations:
|
||||
- key: nvidia.com/gpu
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
- key: node.cilium.io/agent-not-ready
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
container:
|
||||
image: alpine:3.20
|
||||
command: ["/bin/sh", "-c"]
|
||||
resources:
|
||||
requests:
|
||||
cpu: "50m"
|
||||
memory: 32Mi
|
||||
limits:
|
||||
cpu: "100m"
|
||||
memory: 64Mi
|
||||
args:
|
||||
- |
|
||||
echo "GPU warmup pod scheduled on $(hostname) — autoscaler scale-up triggered, node now in scaledown-grace window."
|
||||
|
||||
# ── train: run alpha_train on L40S ──
|
||||
- name: train
|
||||
inputs:
|
||||
parameters:
|
||||
- name: sha
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: "{{workflow.parameters.gpu-pool}}"
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
tolerations:
|
||||
- key: nvidia.com/gpu
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
- key: node.cilium.io/agent-not-ready
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
container:
|
||||
image: gitlab-registry.foxhunt.svc.cluster.local:5000/root/foxhunt/foxhunt-training-runtime:latest
|
||||
imagePullPolicy: Always
|
||||
command: ["/bin/bash", "-c"]
|
||||
env:
|
||||
- name: RUST_LOG
|
||||
value: info
|
||||
- name: SQLX_OFFLINE
|
||||
value: "true"
|
||||
- name: CUDA_COMPUTE_CAP
|
||||
value: "{{workflow.parameters.cuda-compute-cap}}"
|
||||
resources:
|
||||
# L40S-1-48G has 8 vCPU (7800m allocatable, after kubelet
|
||||
# overhead) and ~91Gi memory. Request <7800m so the pod
|
||||
# fits alongside the system daemonsets (cilium, csi-node,
|
||||
# nvidia drivers/dcgm/feature-discovery) on the same node.
|
||||
requests:
|
||||
cpu: "6"
|
||||
memory: 16Gi
|
||||
nvidia.com/gpu: "1"
|
||||
limits:
|
||||
cpu: "7"
|
||||
memory: 64Gi
|
||||
nvidia.com/gpu: "1"
|
||||
volumeMounts:
|
||||
- name: training-data
|
||||
mountPath: /data
|
||||
- name: feature-cache
|
||||
mountPath: /feature-cache
|
||||
args:
|
||||
- |
|
||||
set -e
|
||||
SHA="{{inputs.parameters.sha}}"
|
||||
BIN="/data/bin/$SHA/alpha_train"
|
||||
MBP10_DIR="/data/futures-baseline-mbp10/ES.FUT"
|
||||
PREDECODED_DIR="/feature-cache/predecoded"
|
||||
OUT_DIR="/feature-cache/alpha-perception-runs/$SHA"
|
||||
mkdir -p "$OUT_DIR" "$PREDECODED_DIR"
|
||||
|
||||
if [ ! -x "$BIN" ]; then
|
||||
echo "ERROR: binary not found at $BIN"
|
||||
ls -lh /data/bin/ || true
|
||||
exit 1
|
||||
fi
|
||||
if [ ! -d "$MBP10_DIR" ]; then
|
||||
echo "ERROR: MBP-10 data directory not found at $MBP10_DIR"
|
||||
ls -lh /data/ || true
|
||||
exit 1
|
||||
fi
|
||||
echo "Running alpha_train (stacked Mamba2 -> CfC -> heads)"
|
||||
echo " binary: $BIN"
|
||||
echo " mbp10: $MBP10_DIR ($(find "$MBP10_DIR" -name '*.dbn.zst' | wc -l) files)"
|
||||
echo " predecoded: $PREDECODED_DIR"
|
||||
echo " out: $OUT_DIR"
|
||||
|
||||
EXTRA_FLAGS=""
|
||||
if [ "{{workflow.parameters.auto-horizon-weights}}" = "true" ]; then
|
||||
EXTRA_FLAGS="$EXTRA_FLAGS --auto-horizon-weights"
|
||||
fi
|
||||
if [ -n "{{workflow.parameters.instrument-mode}}" ]; then
|
||||
EXTRA_FLAGS="$EXTRA_FLAGS --instrument-mode {{workflow.parameters.instrument-mode}}"
|
||||
fi
|
||||
|
||||
TRACE_FLAG=""
|
||||
if [ -n "{{workflow.parameters.kernel-step-trace-path}}" ]; then
|
||||
TRACE_FLAG="--kernel-step-trace {{workflow.parameters.kernel-step-trace-path}}"
|
||||
fi
|
||||
|
||||
"$BIN" \
|
||||
--mbp10-data-dir "$MBP10_DIR" \
|
||||
--predecoded-dir "$PREDECODED_DIR" \
|
||||
--out "$OUT_DIR" \
|
||||
--epochs {{workflow.parameters.epochs}} \
|
||||
--multi-resolution "{{workflow.parameters.multi-resolution}}" \
|
||||
--mamba2-state-dim {{workflow.parameters.mamba2-state-dim}} \
|
||||
--lr-cfc {{workflow.parameters.lr-cfc}} \
|
||||
--lr-mamba2 {{workflow.parameters.lr-mamba2}} \
|
||||
--n-train-seqs {{workflow.parameters.n-train-seqs}} \
|
||||
--n-val-seqs {{workflow.parameters.n-val-seqs}} \
|
||||
--seed {{workflow.parameters.seed}} \
|
||||
--batch-size {{workflow.parameters.batch-size}} \
|
||||
--early-stop-metric {{workflow.parameters.early-stop-metric}} \
|
||||
--early-stop-patience {{workflow.parameters.early-stop-patience}} \
|
||||
--cv-fold {{workflow.parameters.cv-fold}} \
|
||||
--cv-n-folds {{workflow.parameters.cv-n-folds}} \
|
||||
--cv-train-window {{workflow.parameters.cv-train-window}} \
|
||||
--smoothness-base-lambda {{workflow.parameters.smoothness-base-lambda}} \
|
||||
${TRACE_FLAG} \
|
||||
$EXTRA_FLAGS
|
||||
|
||||
echo "=== Training complete ==="
|
||||
ls -lh "$OUT_DIR/"
|
||||
echo "--- alpha_train_summary.json ---"
|
||||
cat "$OUT_DIR/alpha_train_summary.json"
|
||||
@@ -1,207 +0,0 @@
|
||||
# alpha-rl: single-pod compile + train on GPU.
|
||||
#
|
||||
# Compiles alpha_rl_train incrementally on the GPU node (~3s warm,
|
||||
# ~90s cold) using the cargo-target-cuda PVC, then runs training
|
||||
# immediately. No separate compile node, no binary transfer, no
|
||||
# fxcache step. Predecoded MBP-10 sidecars live on feature-cache PVC.
|
||||
#
|
||||
# Usage:
|
||||
# argo submit -n foxhunt --from=wftmpl/alpha-rl \
|
||||
# -p git-branch=ml-alpha-phase-a
|
||||
# argo submit -n foxhunt --from=wftmpl/alpha-rl \
|
||||
# -p git-branch=ml-alpha-phase-a -p n-steps=50000 -p n-backtests=128
|
||||
apiVersion: argoproj.io/v1alpha1
|
||||
kind: WorkflowTemplate
|
||||
metadata:
|
||||
name: alpha-rl
|
||||
namespace: foxhunt
|
||||
labels:
|
||||
app.kubernetes.io/name: alpha-rl
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
app.kubernetes.io/component: train
|
||||
spec:
|
||||
entrypoint: compile-and-train
|
||||
serviceAccountName: argo-workflow
|
||||
archiveLogs: true
|
||||
podMetadata:
|
||||
labels:
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
app.kubernetes.io/component: train
|
||||
securityContext:
|
||||
fsGroup: 0
|
||||
ttlStrategy:
|
||||
secondsAfterCompletion: 3600
|
||||
activeDeadlineSeconds: 14400
|
||||
|
||||
arguments:
|
||||
parameters:
|
||||
- name: git-branch
|
||||
value: ml-alpha-phase-a
|
||||
- name: gpu-pool
|
||||
value: ci-training-l40s
|
||||
- name: n-steps
|
||||
value: "50000"
|
||||
- name: n-backtests
|
||||
value: "128"
|
||||
- name: seed
|
||||
value: "16962"
|
||||
- name: seq-len
|
||||
value: "32"
|
||||
- name: per-capacity
|
||||
value: "32768"
|
||||
- name: instrument-mode
|
||||
value: "all"
|
||||
- name: log-every
|
||||
value: "5000"
|
||||
- name: fold-idx
|
||||
value: "0"
|
||||
- name: n-folds
|
||||
value: "1"
|
||||
- name: n-eval-steps
|
||||
value: "0"
|
||||
- name: nsys-profile
|
||||
value: "false"
|
||||
- name: use-multi-head-policy
|
||||
value: "0"
|
||||
- name: band-enabled
|
||||
value: "0"
|
||||
|
||||
volumes:
|
||||
- name: git-ssh-key
|
||||
secret:
|
||||
secretName: argo-git-ssh-key
|
||||
defaultMode: 256
|
||||
- name: training-data
|
||||
persistentVolumeClaim:
|
||||
claimName: training-data-pvc
|
||||
- name: cargo-target-cuda
|
||||
persistentVolumeClaim:
|
||||
claimName: cargo-target-cuda
|
||||
- name: feature-cache
|
||||
persistentVolumeClaim:
|
||||
claimName: feature-cache-pvc
|
||||
|
||||
templates:
|
||||
- name: compile-and-train
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: "{{workflow.parameters.gpu-pool}}"
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
tolerations:
|
||||
- key: nvidia.com/gpu
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
- key: node.cilium.io/agent-not-ready
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
container:
|
||||
image: gitlab-registry.foxhunt.svc.cluster.local:5000/root/foxhunt/ci-builder:latest
|
||||
imagePullPolicy: Always
|
||||
command: ["/bin/bash", "-c"]
|
||||
env:
|
||||
- name: SQLX_OFFLINE
|
||||
value: "true"
|
||||
- name: CARGO_TARGET_DIR
|
||||
value: /cargo-target
|
||||
- name: CARGO_HOME
|
||||
value: /cargo-target/cargo-home
|
||||
- name: CARGO_INCREMENTAL
|
||||
value: "1"
|
||||
- name: FOXHUNT_USE_MULTI_HEAD_POLICY
|
||||
value: "{{workflow.parameters.use-multi-head-policy}}"
|
||||
- name: FOXHUNT_BAND_ENABLED
|
||||
value: "{{workflow.parameters.band-enabled}}"
|
||||
resources:
|
||||
requests:
|
||||
cpu: "4"
|
||||
memory: 16Gi
|
||||
nvidia.com/gpu: "1"
|
||||
limits:
|
||||
cpu: "7"
|
||||
memory: 40Gi
|
||||
nvidia.com/gpu: "1"
|
||||
volumeMounts:
|
||||
- name: git-ssh-key
|
||||
mountPath: /etc/git-ssh
|
||||
readOnly: true
|
||||
- name: training-data
|
||||
mountPath: /data
|
||||
readOnly: true
|
||||
- name: feature-cache
|
||||
mountPath: /feature-cache
|
||||
- name: cargo-target-cuda
|
||||
mountPath: /cargo-target
|
||||
args:
|
||||
- |
|
||||
set -e
|
||||
export PATH="${CARGO_HOME}/bin:${PATH}"
|
||||
export LD_LIBRARY_PATH=$(echo "$LD_LIBRARY_PATH" | tr ':' '\n' | grep -v stubs | tr '\n' ':' | sed 's/:$//')
|
||||
|
||||
# SSH
|
||||
mkdir -p ~/.ssh
|
||||
cp /etc/git-ssh/ssh-privatekey ~/.ssh/id_ed25519
|
||||
chmod 600 ~/.ssh/id_ed25519
|
||||
printf 'Host *\n StrictHostKeyChecking no\n UserKnownHostsFile /dev/null\n' > ~/.ssh/config
|
||||
chmod 600 ~/.ssh/config
|
||||
|
||||
nvidia-smi
|
||||
|
||||
# Git
|
||||
BRANCH="{{workflow.parameters.git-branch}}"
|
||||
BUILD="/cargo-target/src"
|
||||
REPO="ssh://git@gitlab-gitlab-shell.foxhunt.svc.cluster.local:2222/root/foxhunt.git"
|
||||
git config --global --add safe.directory "$BUILD"
|
||||
|
||||
if [ -d "$BUILD/.git" ]; then
|
||||
cd "$BUILD"
|
||||
git fetch origin
|
||||
git checkout --force "origin/$BRANCH"
|
||||
git clean -fd
|
||||
else
|
||||
git clone --filter=blob:none "$REPO" "$BUILD"
|
||||
cd "$BUILD"
|
||||
git checkout "origin/$BRANCH"
|
||||
fi
|
||||
SHA=$(git rev-parse --short=9 HEAD)
|
||||
echo "=== Branch: $BRANCH SHA: $SHA ==="
|
||||
|
||||
# Auto-detect GPU compute capability for all crates' build.rs
|
||||
export CUDA_COMPUTE_CAP=$(nvidia-smi --query-gpu=compute_cap --format=csv,noheader | head -1 | tr -d '.')
|
||||
echo "=== GPU arch: sm_${CUDA_COMPUTE_CAP} ==="
|
||||
|
||||
# Clear stale build artifacts when arch changes
|
||||
rm -rf ${CARGO_TARGET_DIR}/build/ml-alpha-* ${CARGO_TARGET_DIR}/build/ml-core-*
|
||||
|
||||
# Compile (~90s cold, ~3s warm after first run)
|
||||
echo "=== Compile ==="
|
||||
time cargo build --release -p ml-alpha --example alpha_rl_train 2>&1
|
||||
|
||||
# Train
|
||||
OUT="/feature-cache/alpha-rl-runs/$SHA/fold{{workflow.parameters.fold-idx}}"
|
||||
mkdir -p "$OUT"
|
||||
echo "=== Train on $(nvidia-smi --query-gpu=name --format=csv,noheader) ==="
|
||||
TRAIN_BIN="${CARGO_TARGET_DIR}/release/examples/alpha_rl_train"
|
||||
if [ "{{workflow.parameters.nsys-profile}}" = "true" ]; then
|
||||
TRAIN_CMD="nsys profile -o $OUT/nsys_trace --stats=true --force-overwrite=true $TRAIN_BIN"
|
||||
else
|
||||
TRAIN_CMD="stdbuf -oL $TRAIN_BIN"
|
||||
fi
|
||||
$TRAIN_CMD \
|
||||
--mbp10-data-dir /data/futures-baseline-mbp10/ES.FUT \
|
||||
--predecoded-dir /feature-cache/predecoded \
|
||||
--out "$OUT" \
|
||||
--n-steps {{workflow.parameters.n-steps}} \
|
||||
--seq-len {{workflow.parameters.seq-len}} \
|
||||
--n-backtests {{workflow.parameters.n-backtests}} \
|
||||
--per-capacity {{workflow.parameters.per-capacity}} \
|
||||
--seed {{workflow.parameters.seed}} \
|
||||
--instrument-mode "{{workflow.parameters.instrument-mode}}" \
|
||||
--fold-idx {{workflow.parameters.fold-idx}} \
|
||||
--n-folds {{workflow.parameters.n-folds}} \
|
||||
--n-eval-steps {{workflow.parameters.n-eval-steps}} \
|
||||
--log-every {{workflow.parameters.log-every}}
|
||||
|
||||
echo "=== Complete: $OUT ==="
|
||||
ls -lh "$OUT/"
|
||||
if [ -f "$OUT/alpha_rl_train_summary.json" ]; then
|
||||
cat "$OUT/alpha_rl_train_summary.json"
|
||||
fi
|
||||
@@ -1,121 +0,0 @@
|
||||
apiVersion: argoproj.io/v1alpha1
|
||||
kind: WorkflowTemplate
|
||||
metadata:
|
||||
name: build-ci-image
|
||||
namespace: foxhunt
|
||||
labels:
|
||||
app.kubernetes.io/name: build-ci-image
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
spec:
|
||||
activeDeadlineSeconds: 3600
|
||||
serviceAccountName: argo-workflow
|
||||
entrypoint: build
|
||||
podMetadata:
|
||||
labels:
|
||||
app.kubernetes.io/component: ci-pipeline
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
ttlStrategy:
|
||||
secondsAfterCompletion: 3600
|
||||
arguments:
|
||||
parameters:
|
||||
- name: commit-sha
|
||||
value: HEAD
|
||||
- name: dockerfile
|
||||
value: Dockerfile.ci-builder
|
||||
- name: image-name
|
||||
value: ci-builder
|
||||
|
||||
templates:
|
||||
# Single-pod build: init container clones repo, kaniko builds image.
|
||||
# Uses emptyDir — no PVC needed, works when called via templateRef.
|
||||
- name: build
|
||||
inputs:
|
||||
parameters:
|
||||
- name: commit-sha
|
||||
value: "{{workflow.parameters.commit-sha}}"
|
||||
- name: dockerfile
|
||||
value: "{{workflow.parameters.dockerfile}}"
|
||||
- name: image-name
|
||||
value: "{{workflow.parameters.image-name}}"
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: ci-compile-cpu
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
tolerations:
|
||||
- key: node.cilium.io/agent-not-ready
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
volumes:
|
||||
- name: workspace
|
||||
emptyDir: {}
|
||||
- name: git-ssh-key
|
||||
secret:
|
||||
secretName: argo-git-ssh-key
|
||||
defaultMode: 256
|
||||
- name: registry-auth
|
||||
secret:
|
||||
secretName: gitlab-registry
|
||||
items:
|
||||
- key: .dockerconfigjson
|
||||
path: config.json
|
||||
initContainers:
|
||||
- name: git-clone
|
||||
image: alpine/git:latest
|
||||
command: ["/bin/sh", "-c"]
|
||||
resources:
|
||||
requests:
|
||||
cpu: 200m
|
||||
memory: 256Mi
|
||||
limits:
|
||||
cpu: "1"
|
||||
memory: 512Mi
|
||||
volumeMounts:
|
||||
- name: git-ssh-key
|
||||
mountPath: /etc/git-ssh
|
||||
readOnly: true
|
||||
- name: workspace
|
||||
mountPath: /workspace
|
||||
args:
|
||||
- |
|
||||
set -ex
|
||||
mkdir -p /root/.ssh
|
||||
cp /etc/git-ssh/ssh-privatekey /root/.ssh/id_ed25519
|
||||
chmod 600 /root/.ssh/id_ed25519
|
||||
printf 'Host *\n StrictHostKeyChecking no\n UserKnownHostsFile /dev/null\n' > /root/.ssh/config
|
||||
chmod 600 /root/.ssh/config
|
||||
|
||||
SHA="{{inputs.parameters.commit-sha}}"
|
||||
REPO="ssh://git@gitlab-gitlab-shell.foxhunt.svc.cluster.local:2222/root/foxhunt.git"
|
||||
git clone --no-checkout --filter=blob:none "$REPO" /workspace/src
|
||||
cd /workspace/src
|
||||
git checkout "$SHA"
|
||||
echo "Checked out $(git rev-parse --short HEAD)"
|
||||
container:
|
||||
image: gcr.io/kaniko-project/executor:debug
|
||||
command: ["/busybox/sh", "-c"]
|
||||
env:
|
||||
- name: DOCKER_CONFIG
|
||||
value: /kaniko/.docker
|
||||
resources:
|
||||
requests:
|
||||
cpu: "4"
|
||||
memory: 8Gi
|
||||
limits:
|
||||
cpu: "8"
|
||||
memory: 16Gi
|
||||
volumeMounts:
|
||||
- name: registry-auth
|
||||
mountPath: /kaniko/.docker
|
||||
readOnly: true
|
||||
- name: workspace
|
||||
mountPath: /workspace
|
||||
args:
|
||||
- |
|
||||
/kaniko/executor \
|
||||
--context=/workspace/src \
|
||||
--dockerfile=/workspace/src/infra/docker/{{inputs.parameters.dockerfile}} \
|
||||
--destination=gitlab-registry.foxhunt.svc.cluster.local:5000/root/foxhunt/{{inputs.parameters.image-name}}:latest \
|
||||
--insecure-registry=gitlab-registry.foxhunt.svc.cluster.local:5000 \
|
||||
--skip-tls-verify-registry=gitlab-registry.foxhunt.svc.cluster.local:5000 \
|
||||
--cache=true \
|
||||
--cache-repo=gitlab-registry.foxhunt.svc.cluster.local:5000/root/foxhunt/cache \
|
||||
--snapshot-mode=redo
|
||||
@@ -1,16 +0,0 @@
|
||||
# infra/k8s/argo/cargo-target-cuda-test-pvc.yaml
|
||||
apiVersion: v1
|
||||
kind: PersistentVolumeClaim
|
||||
metadata:
|
||||
name: cargo-target-cuda-test
|
||||
namespace: foxhunt
|
||||
labels:
|
||||
app.kubernetes.io/name: cargo-target
|
||||
app.kubernetes.io/component: ci-cache
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
spec:
|
||||
accessModes: [ReadWriteOnce]
|
||||
storageClassName: scw-bssd-retain
|
||||
resources:
|
||||
requests:
|
||||
storage: 30Gi
|
||||
@@ -1,43 +0,0 @@
|
||||
# Persistent cargo target directories for incremental compilation.
|
||||
#
|
||||
# Why: sccache can't cache workspace rlib crates (109 non-cacheable per build).
|
||||
# Persisting target/ lets cargo's incremental compilation skip unchanged crates,
|
||||
# reducing typical CI builds from ~20 min (full rebuild) to ~2-3 min.
|
||||
#
|
||||
# Two separate PVCs because compile-services (no cuda feature) and
|
||||
# compile-training (cuda feature) produce incompatible artifacts.
|
||||
#
|
||||
# NOTE: Only one workflow should use each PVC at a time. ci-pipeline and
|
||||
# compile-and-train must not run their compile steps concurrently.
|
||||
---
|
||||
apiVersion: v1
|
||||
kind: PersistentVolumeClaim
|
||||
metadata:
|
||||
name: cargo-target-cpu
|
||||
namespace: foxhunt
|
||||
labels:
|
||||
app.kubernetes.io/name: cargo-target
|
||||
app.kubernetes.io/component: ci-cache
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
spec:
|
||||
accessModes: [ReadWriteOnce]
|
||||
storageClassName: scw-bssd-retain
|
||||
resources:
|
||||
requests:
|
||||
storage: 30Gi
|
||||
---
|
||||
apiVersion: v1
|
||||
kind: PersistentVolumeClaim
|
||||
metadata:
|
||||
name: cargo-target-cuda
|
||||
namespace: foxhunt
|
||||
labels:
|
||||
app.kubernetes.io/name: cargo-target
|
||||
app.kubernetes.io/component: ci-cache
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
spec:
|
||||
accessModes: [ReadWriteOnce]
|
||||
storageClassName: scw-bssd-retain
|
||||
resources:
|
||||
requests:
|
||||
storage: 30Gi
|
||||
@@ -7,18 +7,25 @@ metadata:
|
||||
labels:
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
rules:
|
||||
# apps: roll + apply Deployments. watch is required by `kubectl rollout status`.
|
||||
- apiGroups: ["apps"]
|
||||
resources: [deployments]
|
||||
verbs: [get, list, patch]
|
||||
verbs: [get, list, watch, patch]
|
||||
- apiGroups: ["argoproj.io"]
|
||||
resources: [workflowtemplates, eventsources, sensors, eventbus]
|
||||
verbs: [get, list, create, update, patch]
|
||||
# core: services/configmaps, plus serviceaccounts (tailscale-dashboard, dagster) and the forward-track PVC
|
||||
# — added so the fxhnt-cockpit deploy can `kubectl apply` its full manifest set without a partial-apply failure.
|
||||
- apiGroups: [""]
|
||||
resources: [services, configmaps]
|
||||
resources: [services, configmaps, serviceaccounts, persistentvolumeclaims]
|
||||
verbs: [get, list, create, update, patch]
|
||||
- apiGroups: ["networking.k8s.io"]
|
||||
resources: [networkpolicies]
|
||||
verbs: [get, list, create, update, patch]
|
||||
# batch: the fxhnt-forward CronJob
|
||||
- apiGroups: ["batch"]
|
||||
resources: [cronjobs]
|
||||
verbs: [get, list, create, update, patch]
|
||||
- apiGroups: ["rbac.authorization.k8s.io"]
|
||||
resources: [roles, rolebindings]
|
||||
verbs: [get, list, create, update, patch, bind, escalate]
|
||||
|
||||
@@ -1,685 +0,0 @@
|
||||
apiVersion: argoproj.io/v1alpha1
|
||||
kind: WorkflowTemplate
|
||||
metadata:
|
||||
name: ci-pipeline
|
||||
namespace: foxhunt
|
||||
labels:
|
||||
app.kubernetes.io/name: ci-pipeline
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
spec:
|
||||
activeDeadlineSeconds: 7200
|
||||
serviceAccountName: argo-workflow
|
||||
entrypoint: pipeline
|
||||
onExit: notify-result
|
||||
podMetadata:
|
||||
labels:
|
||||
app.kubernetes.io/component: ci-pipeline
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
archiveLogs: true
|
||||
podGC:
|
||||
strategy: OnPodCompletion
|
||||
ttlStrategy:
|
||||
secondsAfterCompletion: 3600
|
||||
volumes:
|
||||
- name: git-ssh-key
|
||||
secret:
|
||||
secretName: argo-git-ssh-key
|
||||
defaultMode: 256
|
||||
- name: registry-auth
|
||||
secret:
|
||||
secretName: gitlab-registry
|
||||
optional: true
|
||||
items:
|
||||
- key: .dockerconfigjson
|
||||
path: config.json
|
||||
- name: cargo-target-cpu
|
||||
persistentVolumeClaim:
|
||||
claimName: cargo-target-cpu
|
||||
arguments:
|
||||
parameters:
|
||||
- name: commit-sha
|
||||
value: HEAD
|
||||
- name: commits-json
|
||||
value: "[]"
|
||||
templates:
|
||||
# ── DAG: orchestrate pipeline ──
|
||||
# Policy (user directive): push-triggered runs are limited to infrastructure
|
||||
# sync work — docker image rebuilds + argo-template apply + terragrunt apply.
|
||||
# These run automatically when the relevant paths change:
|
||||
# - rebuild-ci-builder* / rebuild-runtime / rebuild-training-runtime: gated on
|
||||
# `detect-changes.docker-images == true` (infra/docker/)
|
||||
# - apply-argo-templates: gated on `detect-changes.needs-argo-templates` (infra/k8s/argo/)
|
||||
# - terragrunt-apply: gated on `detect-changes.needs-infra` (infra/live/, infra/modules/)
|
||||
# All other tasks (test-gate, build-web-dashboard, gpu-test) are `when: "false"` —
|
||||
# they run via manual script invocation (`argo submit` / `./scripts/argo-*.sh`)
|
||||
# or the manual `/compile-deploy` webhook. Training + compile/deploy workflows
|
||||
# are separate templates, always manual.
|
||||
- name: pipeline
|
||||
dag:
|
||||
tasks:
|
||||
- name: detect-changes
|
||||
template: detect-changes
|
||||
|
||||
- name: build-web-dashboard
|
||||
depends: "detect-changes"
|
||||
template: build-web-dashboard
|
||||
# Policy: push-triggered runs are docker image rebuilds only.
|
||||
# Dashboard builds run via manual `argo submit` (scripts).
|
||||
when: "false"
|
||||
|
||||
- name: rebuild-ci-builder
|
||||
depends: "detect-changes"
|
||||
templateRef:
|
||||
name: build-ci-image
|
||||
template: build
|
||||
arguments:
|
||||
parameters:
|
||||
- name: commit-sha
|
||||
value: "{{workflow.parameters.commit-sha}}"
|
||||
- name: dockerfile
|
||||
value: Dockerfile.ci-builder
|
||||
- name: image-name
|
||||
value: ci-builder
|
||||
when: "{{tasks.detect-changes.outputs.parameters.docker-images}} == true"
|
||||
|
||||
- name: rebuild-ci-builder-cpu
|
||||
depends: "detect-changes"
|
||||
templateRef:
|
||||
name: build-ci-image
|
||||
template: build
|
||||
arguments:
|
||||
parameters:
|
||||
- name: commit-sha
|
||||
value: "{{workflow.parameters.commit-sha}}"
|
||||
- name: dockerfile
|
||||
value: Dockerfile.ci-builder-cpu
|
||||
- name: image-name
|
||||
value: ci-builder-cpu
|
||||
when: "{{tasks.detect-changes.outputs.parameters.docker-images}} == true"
|
||||
|
||||
- name: rebuild-runtime
|
||||
depends: "detect-changes"
|
||||
templateRef:
|
||||
name: build-ci-image
|
||||
template: build
|
||||
arguments:
|
||||
parameters:
|
||||
- name: commit-sha
|
||||
value: "{{workflow.parameters.commit-sha}}"
|
||||
- name: dockerfile
|
||||
value: Dockerfile.foxhunt-runtime
|
||||
- name: image-name
|
||||
value: foxhunt-runtime
|
||||
when: "{{tasks.detect-changes.outputs.parameters.docker-images}} == true"
|
||||
|
||||
- name: rebuild-training-runtime
|
||||
depends: "detect-changes"
|
||||
templateRef:
|
||||
name: build-ci-image
|
||||
template: build
|
||||
arguments:
|
||||
parameters:
|
||||
- name: commit-sha
|
||||
value: "{{workflow.parameters.commit-sha}}"
|
||||
- name: dockerfile
|
||||
value: Dockerfile.foxhunt-training-runtime
|
||||
- name: image-name
|
||||
value: foxhunt-training-runtime
|
||||
when: "{{tasks.detect-changes.outputs.parameters.docker-images}} == true"
|
||||
|
||||
- name: test-gate
|
||||
depends: "detect-changes && (rebuild-ci-builder-cpu.Succeeded || rebuild-ci-builder-cpu.Skipped)"
|
||||
template: test-gate
|
||||
arguments:
|
||||
parameters:
|
||||
- name: commit-sha
|
||||
value: "{{workflow.parameters.commit-sha}}"
|
||||
# Policy: push-triggered runs are docker image rebuilds only.
|
||||
# Cargo test + clippy quality gate runs via manual `argo submit`
|
||||
# (scripts/argo-test.sh or CI-driven `argo submit --from=wftmpl/ci-pipeline`).
|
||||
when: "false"
|
||||
|
||||
- name: apply-argo-templates
|
||||
depends: "detect-changes.Succeeded && (rebuild-ci-builder-cpu.Succeeded || rebuild-ci-builder-cpu.Skipped)"
|
||||
template: apply-argo-templates
|
||||
arguments:
|
||||
parameters:
|
||||
- name: commit-sha
|
||||
value: "{{workflow.parameters.commit-sha}}"
|
||||
when: "{{tasks.detect-changes.outputs.parameters.needs-argo-templates}} == true"
|
||||
|
||||
- name: terragrunt-apply
|
||||
depends: "detect-changes.Succeeded && (rebuild-ci-builder-cpu.Succeeded || rebuild-ci-builder-cpu.Skipped)"
|
||||
template: terragrunt-apply
|
||||
arguments:
|
||||
parameters:
|
||||
- name: commit-sha
|
||||
value: "{{workflow.parameters.commit-sha}}"
|
||||
when: "{{tasks.detect-changes.outputs.parameters.needs-infra}} == true"
|
||||
|
||||
# GPU tests: manual only (argo submit --from=wftmpl/gpu-test-pipeline).
|
||||
# Auto-trigger disabled — H100 nodes are expensive, run on demand.
|
||||
- name: gpu-test
|
||||
depends: "detect-changes"
|
||||
template: submit-gpu-test
|
||||
when: "false"
|
||||
|
||||
# ── detect-changes: classify changed files to gate downstream steps ──
|
||||
- name: detect-changes
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: platform
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
container:
|
||||
image: gitlab-registry.foxhunt.svc.cluster.local:5000/root/foxhunt/foxhunt-runtime:latest
|
||||
command: ["/bin/sh", "-c"]
|
||||
env:
|
||||
- name: COMMITS_JSON
|
||||
value: "{{workflow.parameters.commits-json}}"
|
||||
resources:
|
||||
requests:
|
||||
cpu: 100m
|
||||
memory: 128Mi
|
||||
limits:
|
||||
cpu: 500m
|
||||
memory: 256Mi
|
||||
args:
|
||||
- |
|
||||
set -e
|
||||
cat <<'SCRIPT' > /tmp/detect.sh
|
||||
#!/bin/sh
|
||||
set -e
|
||||
|
||||
# Extract changed file paths from webhook JSON using grep+sed (no jq dependency).
|
||||
CHANGED_FILES=$(echo "$COMMITS_JSON" \
|
||||
| grep -oE '"(added|modified|removed)":\[[^]]*\]' \
|
||||
| sed 's/"added"://;s/"modified"://;s/"removed"://' \
|
||||
| tr ',' '\n' | tr -d '[]"' | sed '/^$/d' | sort -u \
|
||||
|| echo "")
|
||||
|
||||
if [ -z "$CHANGED_FILES" ]; then
|
||||
echo "No changed files detected — triggering full rebuild"
|
||||
CHANGED_FILES="Cargo.toml"
|
||||
fi
|
||||
|
||||
echo "=== Changed files ==="
|
||||
echo "$CHANGED_FILES"
|
||||
echo "===================="
|
||||
|
||||
# --- Helper: check if any changed file matches a set of path prefixes ---
|
||||
check_paths() {
|
||||
local patterns="$1"
|
||||
for file in $CHANGED_FILES; do
|
||||
for pattern in $patterns; do
|
||||
case "$file" in
|
||||
${pattern}*) echo "true"; return ;;
|
||||
esac
|
||||
done
|
||||
done
|
||||
echo "false"
|
||||
}
|
||||
|
||||
# --- Boolean gates for downstream DAG tasks ---
|
||||
NEEDS_CODE=$(check_paths "Cargo.toml Cargo.lock crates/ services/ bin/fxt/")
|
||||
NEEDS_DASHBOARD=$(check_paths "web-dashboard/")
|
||||
DOCKER_IMAGES=$(check_paths "infra/docker/")
|
||||
NEEDS_INFRA=$(check_paths "infra/live/ infra/modules/")
|
||||
NEEDS_ARGO_TEMPLATES=$(check_paths "infra/k8s/argo/")
|
||||
ML_CHANGED=$(check_paths "crates/ml crates/ml-dqn crates/ml-core crates/ml-regime crates/ml-features crates/ml-supervised crates/ml-ppo crates/ml-ensemble crates/ml-hyperopt crates/ml-labeling config/training config/gpu")
|
||||
|
||||
echo "=== Build decisions ==="
|
||||
echo "needs-code: $NEEDS_CODE"
|
||||
echo "needs-dashboard: $NEEDS_DASHBOARD"
|
||||
echo "docker-images: $DOCKER_IMAGES"
|
||||
echo "needs-infra: $NEEDS_INFRA"
|
||||
echo "needs-argo: $NEEDS_ARGO_TEMPLATES"
|
||||
echo "ml-changed: $ML_CHANGED"
|
||||
echo "======================"
|
||||
|
||||
mkdir -p /tmp/outputs
|
||||
echo -n "$NEEDS_CODE" > /tmp/outputs/needs-code
|
||||
echo -n "$NEEDS_DASHBOARD" > /tmp/outputs/needs-dashboard
|
||||
echo -n "$DOCKER_IMAGES" > /tmp/outputs/docker-images
|
||||
echo -n "$NEEDS_INFRA" > /tmp/outputs/needs-infra
|
||||
echo -n "$NEEDS_ARGO_TEMPLATES" > /tmp/outputs/needs-argo-templates
|
||||
echo -n "$ML_CHANGED" > /tmp/outputs/ml-changed
|
||||
SCRIPT
|
||||
chmod +x /tmp/detect.sh
|
||||
/tmp/detect.sh
|
||||
outputs:
|
||||
parameters:
|
||||
- name: needs-code
|
||||
valueFrom:
|
||||
path: /tmp/outputs/needs-code
|
||||
- name: needs-dashboard
|
||||
valueFrom:
|
||||
path: /tmp/outputs/needs-dashboard
|
||||
- name: docker-images
|
||||
valueFrom:
|
||||
path: /tmp/outputs/docker-images
|
||||
- name: needs-infra
|
||||
valueFrom:
|
||||
path: /tmp/outputs/needs-infra
|
||||
- name: needs-argo-templates
|
||||
valueFrom:
|
||||
path: /tmp/outputs/needs-argo-templates
|
||||
- name: ml-changed
|
||||
valueFrom:
|
||||
path: /tmp/outputs/ml-changed
|
||||
|
||||
# ── test-gate: clippy + cargo test quality gate ──
|
||||
# Uses ci-builder (with CUDA toolkit) because cudarc 0.19 requires nvcc at build time.
|
||||
- name: test-gate
|
||||
inputs:
|
||||
parameters:
|
||||
- name: commit-sha
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: ci-compile-cpu
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
tolerations:
|
||||
- key: k8s.scaleway.com/pool-name
|
||||
operator: Equal
|
||||
value: ci-compile-cpu
|
||||
effect: NoSchedule
|
||||
sidecars:
|
||||
- name: redis
|
||||
image: redis:7-alpine
|
||||
command: ["redis-server", "--save", "", "--appendonly", "no"]
|
||||
ports:
|
||||
- containerPort: 6379
|
||||
resources:
|
||||
requests:
|
||||
cpu: 100m
|
||||
memory: 64Mi
|
||||
limits:
|
||||
cpu: 500m
|
||||
memory: 128Mi
|
||||
readinessProbe:
|
||||
tcpSocket:
|
||||
port: 6379
|
||||
initialDelaySeconds: 1
|
||||
periodSeconds: 1
|
||||
container:
|
||||
image: gitlab-registry.foxhunt.svc.cluster.local:5000/root/foxhunt/ci-builder:latest
|
||||
imagePullPolicy: Always
|
||||
command: ["/bin/sh", "-c"]
|
||||
env:
|
||||
- name: SQLX_OFFLINE
|
||||
value: "true"
|
||||
- name: CARGO_TERM_COLOR
|
||||
value: always
|
||||
- name: CARGO_TARGET_DIR
|
||||
value: /cargo-target
|
||||
- name: CARGO_HOME
|
||||
value: /cargo-target/cargo-home
|
||||
- name: REDIS_URL
|
||||
value: "redis://localhost:6379"
|
||||
resources:
|
||||
requests:
|
||||
cpu: "14"
|
||||
memory: 16Gi
|
||||
limits:
|
||||
cpu: "30"
|
||||
memory: 32Gi
|
||||
volumeMounts:
|
||||
- name: git-ssh-key
|
||||
mountPath: /etc/git-ssh
|
||||
readOnly: true
|
||||
- name: cargo-target-cpu
|
||||
mountPath: /cargo-target
|
||||
args:
|
||||
- |
|
||||
set -e
|
||||
SHA="{{inputs.parameters.commit-sha}}"
|
||||
|
||||
# SSH setup
|
||||
mkdir -p ~/.ssh
|
||||
cp /etc/git-ssh/ssh-privatekey ~/.ssh/id_ed25519
|
||||
chmod 600 ~/.ssh/id_ed25519
|
||||
printf 'Host *\n StrictHostKeyChecking no\n UserKnownHostsFile /dev/null\n' > ~/.ssh/config
|
||||
chmod 600 ~/.ssh/config
|
||||
|
||||
REPO="ssh://git@gitlab-gitlab-shell.foxhunt.svc.cluster.local:2222/root/foxhunt.git"
|
||||
WORKSPACE="/cargo-target/src"
|
||||
|
||||
git config --global --add safe.directory "$WORKSPACE"
|
||||
|
||||
if [ -d "$WORKSPACE/.git" ]; then
|
||||
cd "$WORKSPACE"
|
||||
CURRENT=$(git rev-parse HEAD 2>/dev/null || echo "none")
|
||||
if [ "$CURRENT" = "$SHA" ]; then
|
||||
echo "=== Already at $SHA ==="
|
||||
else
|
||||
git fetch origin
|
||||
git checkout --force "$SHA"
|
||||
git clean -fd
|
||||
fi
|
||||
else
|
||||
git clone --filter=blob:none "$REPO" "$WORKSPACE"
|
||||
cd "$WORKSPACE"
|
||||
git checkout "$SHA"
|
||||
fi
|
||||
|
||||
export PATH="${CARGO_HOME}/bin:${PATH}"
|
||||
|
||||
echo "=== Waiting for Redis sidecar ==="
|
||||
for i in 1 2 3 4 5; do
|
||||
if printf 'PING\r\n' | nc -w1 127.0.0.1 6379 2>/dev/null | grep -q PONG; then
|
||||
echo "Redis ready"
|
||||
break
|
||||
fi
|
||||
sleep 1
|
||||
done
|
||||
|
||||
echo "=== Running clippy (lib targets) ==="
|
||||
cargo clippy --workspace --lib -- -D warnings 2>&1 | tee /cargo-target/clippy.log
|
||||
|
||||
echo "=== Running tests (lib only, no integration) ==="
|
||||
set +e
|
||||
cargo test --workspace --lib 2>&1 | tee /cargo-target/test-output.log
|
||||
TEST_EXIT=$?
|
||||
set -e
|
||||
if [ "$TEST_EXIT" -ne 0 ]; then
|
||||
echo "=== TEST FAILURES ==="
|
||||
grep -A5 'FAILED\|panicked\|test result: FAILED\|error\[' /cargo-target/test-output.log || true
|
||||
exit "$TEST_EXIT"
|
||||
fi
|
||||
|
||||
echo "=== Test gate passed ==="
|
||||
|
||||
# ── build-web-dashboard: npm build + upload to MinIO ──
|
||||
- name: build-web-dashboard
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: platform
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
container:
|
||||
image: node:22-alpine
|
||||
command: ["/bin/sh", "-c"]
|
||||
env:
|
||||
- name: MINIO_ACCESS_KEY
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: minio-credentials
|
||||
key: access-key
|
||||
- name: MINIO_SECRET_KEY
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: minio-credentials
|
||||
key: secret-key
|
||||
resources:
|
||||
requests:
|
||||
cpu: "1"
|
||||
memory: 1Gi
|
||||
limits:
|
||||
cpu: "2"
|
||||
memory: 2Gi
|
||||
volumeMounts:
|
||||
- name: git-ssh-key
|
||||
mountPath: /etc/git-ssh
|
||||
readOnly: true
|
||||
args:
|
||||
- |
|
||||
set -e
|
||||
SHA="{{workflow.parameters.commit-sha}}"
|
||||
|
||||
# Git clone
|
||||
apk add --no-cache git openssh rclone
|
||||
mkdir -p ~/.ssh
|
||||
cp /etc/git-ssh/ssh-privatekey ~/.ssh/id_ed25519
|
||||
chmod 600 ~/.ssh/id_ed25519
|
||||
printf 'Host *\n StrictHostKeyChecking no\n UserKnownHostsFile /dev/null\n' > ~/.ssh/config
|
||||
chmod 600 ~/.ssh/config
|
||||
|
||||
REPO="ssh://git@gitlab-gitlab-shell.foxhunt.svc.cluster.local:2222/root/foxhunt.git"
|
||||
WORKSPACE="/tmp/workspace"
|
||||
git clone --no-checkout --filter=blob:none "$REPO" "$WORKSPACE"
|
||||
cd "$WORKSPACE"
|
||||
git checkout "$SHA"
|
||||
|
||||
cd "$WORKSPACE/web-dashboard"
|
||||
echo "=== Building web dashboard ==="
|
||||
npm ci
|
||||
npm run build
|
||||
|
||||
echo "=== Uploading to MinIO ==="
|
||||
rclone copy dist/ :s3:foxhunt-binaries/web-dashboard/ \
|
||||
--s3-provider=Minio \
|
||||
--s3-endpoint=http://minio.foxhunt.svc.cluster.local:9000 \
|
||||
--s3-access-key-id="${MINIO_ACCESS_KEY}" \
|
||||
--s3-secret-access-key="${MINIO_SECRET_KEY}" \
|
||||
--s3-no-check-bucket
|
||||
|
||||
echo "=== Web dashboard build + upload done ==="
|
||||
|
||||
# ── terragrunt-apply: apply infra changes on main push ──
|
||||
- name: terragrunt-apply
|
||||
inputs:
|
||||
parameters:
|
||||
- name: commit-sha
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: platform
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
serviceAccountName: argo-workflow
|
||||
container:
|
||||
image: gitlab-registry.foxhunt.svc.cluster.local:5000/root/foxhunt/ci-builder-cpu:latest
|
||||
command: ["/bin/sh", "-c"]
|
||||
env:
|
||||
- name: GITLAB_TOKEN
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: gitlab-pat
|
||||
key: token
|
||||
- name: TF_HTTP_USERNAME
|
||||
value: root
|
||||
- name: TF_HTTP_PASSWORD
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: gitlab-pat
|
||||
key: token
|
||||
- name: SCW_ACCESS_KEY
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: scaleway-credentials
|
||||
key: access-key
|
||||
- name: SCW_SECRET_KEY
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: scaleway-credentials
|
||||
key: secret-key
|
||||
- name: SCW_DEFAULT_PROJECT_ID
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: scaleway-credentials
|
||||
key: project-id
|
||||
- name: GITLAB_TF_STATE_URL
|
||||
value: "http://gitlab-webservice-default.foxhunt.svc.cluster.local:8181"
|
||||
resources:
|
||||
requests:
|
||||
cpu: 200m
|
||||
memory: 256Mi
|
||||
limits:
|
||||
cpu: "1"
|
||||
memory: 512Mi
|
||||
volumeMounts:
|
||||
- name: git-ssh-key
|
||||
mountPath: /etc/git-ssh
|
||||
readOnly: true
|
||||
args:
|
||||
- |
|
||||
set -e
|
||||
|
||||
# Clone repo
|
||||
mkdir -p /root/.ssh
|
||||
cp /etc/git-ssh/ssh-privatekey /root/.ssh/id_ed25519
|
||||
chmod 600 /root/.ssh/id_ed25519
|
||||
printf 'Host *\n StrictHostKeyChecking no\n UserKnownHostsFile /dev/null\n' > /root/.ssh/config
|
||||
chmod 600 /root/.ssh/config
|
||||
|
||||
SHA="{{inputs.parameters.commit-sha}}"
|
||||
REPO="ssh://git@gitlab-gitlab-shell.foxhunt.svc.cluster.local:2222/root/foxhunt.git"
|
||||
git clone --no-checkout --filter=blob:none "$REPO" /workspace/src
|
||||
cd /workspace/src
|
||||
git checkout "$SHA"
|
||||
echo "Checked out $(git rev-parse --short HEAD)"
|
||||
|
||||
tofu --version
|
||||
terragrunt --version
|
||||
|
||||
# Apply each terragrunt module
|
||||
for module in block-storage dns public-gateway kapsule; do
|
||||
MODULE_DIR="infra/live/production/${module}"
|
||||
if [ -d "$MODULE_DIR" ]; then
|
||||
echo "=== Terragrunt plan: ${module} ==="
|
||||
cd "/workspace/src/${MODULE_DIR}"
|
||||
terragrunt init --non-interactive -reconfigure
|
||||
OUTPUT=$(terragrunt plan --non-interactive -detailed-exitcode 2>&1) || EXITCODE=$?
|
||||
EXITCODE=${EXITCODE:-0}
|
||||
|
||||
if [ "$EXITCODE" -eq 2 ]; then
|
||||
echo "=== Terragrunt plan output: ${module} ==="
|
||||
echo "$OUTPUT"
|
||||
echo "=== Terragrunt apply: ${module} ==="
|
||||
terragrunt apply --non-interactive -auto-approve
|
||||
elif [ "$EXITCODE" -eq 0 ]; then
|
||||
echo "=== No changes for ${module} ==="
|
||||
else
|
||||
echo "=== ERROR planning ${module} ==="
|
||||
echo "$OUTPUT"
|
||||
exit 1
|
||||
fi
|
||||
cd /workspace/src
|
||||
fi
|
||||
done
|
||||
|
||||
echo "=== Terragrunt apply complete ==="
|
||||
|
||||
# ── apply-argo-templates: self-apply WorkflowTemplates, EventSources, Sensors ──
|
||||
- name: apply-argo-templates
|
||||
inputs:
|
||||
parameters:
|
||||
- name: commit-sha
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: platform
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
serviceAccountName: argo-workflow
|
||||
container:
|
||||
image: gitlab-registry.foxhunt.svc.cluster.local:5000/root/foxhunt/ci-builder-cpu:latest
|
||||
command: ["/bin/sh", "-c"]
|
||||
resources:
|
||||
requests:
|
||||
cpu: 100m
|
||||
memory: 128Mi
|
||||
limits:
|
||||
cpu: 500m
|
||||
memory: 256Mi
|
||||
volumeMounts:
|
||||
- name: git-ssh-key
|
||||
mountPath: /etc/git-ssh
|
||||
readOnly: true
|
||||
args:
|
||||
- |
|
||||
set -e
|
||||
|
||||
# Clone repo at commit
|
||||
mkdir -p /root/.ssh
|
||||
cp /etc/git-ssh/ssh-privatekey /root/.ssh/id_ed25519
|
||||
chmod 600 /root/.ssh/id_ed25519
|
||||
printf 'Host *\n StrictHostKeyChecking no\n UserKnownHostsFile /dev/null\n' > /root/.ssh/config
|
||||
chmod 600 /root/.ssh/config
|
||||
|
||||
SHA="{{inputs.parameters.commit-sha}}"
|
||||
REPO="ssh://git@gitlab-gitlab-shell.foxhunt.svc.cluster.local:2222/root/foxhunt.git"
|
||||
git clone --no-checkout --filter=blob:none "$REPO" /workspace/src
|
||||
cd /workspace/src
|
||||
git checkout "$SHA"
|
||||
echo "Checked out $(git rev-parse --short HEAD)"
|
||||
|
||||
# Apply all Argo WorkflowTemplates
|
||||
for f in infra/k8s/argo/*-template.yaml; do
|
||||
if [ -f "$f" ]; then
|
||||
echo "=== Applying $(basename $f) ==="
|
||||
kubectl -n foxhunt apply -f "$f"
|
||||
fi
|
||||
done
|
||||
|
||||
# Apply EventSources and Sensors
|
||||
for f in infra/k8s/argo/events/*.yaml; do
|
||||
if [ -f "$f" ]; then
|
||||
echo "=== Applying $(basename $f) ==="
|
||||
kubectl -n foxhunt apply -f "$f"
|
||||
fi
|
||||
done
|
||||
|
||||
# Apply RBAC and NetworkPolicies
|
||||
kubectl -n foxhunt apply -f infra/k8s/argo/ci-deploy-rbac.yaml
|
||||
kubectl -n foxhunt apply -f infra/k8s/argo/argo-workflow-netpol.yaml
|
||||
|
||||
echo "=== Argo templates applied ==="
|
||||
|
||||
# ── submit-gpu-test: launch GPU test workflow as a child workflow ──
|
||||
- name: submit-gpu-test
|
||||
# Argo executor (main + wait) needs 256Mi to track large child workflow status JSON.
|
||||
# Default 64Mi causes OOMKilled when child workflow has many parameters.
|
||||
podSpecPatch: '{"containers":[{"name":"main","resources":{"requests":{"memory":"128Mi"},"limits":{"memory":"256Mi"}}},{"name":"wait","resources":{"requests":{"memory":"128Mi"},"limits":{"memory":"256Mi"}}}]}'
|
||||
resource:
|
||||
action: create
|
||||
setOwnerReference: true
|
||||
successCondition: status.phase == Succeeded
|
||||
failureCondition: status.phase in (Failed, Error)
|
||||
manifest: |
|
||||
apiVersion: argoproj.io/v1alpha1
|
||||
kind: Workflow
|
||||
metadata:
|
||||
generateName: gpu-test-on-push-
|
||||
namespace: foxhunt
|
||||
spec:
|
||||
workflowTemplateRef:
|
||||
name: gpu-test-pipeline
|
||||
arguments:
|
||||
parameters:
|
||||
- name: commit-ref
|
||||
value: "{{workflow.parameters.commit-sha}}"
|
||||
- name: models
|
||||
value: "dqn,ppo,tft"
|
||||
|
||||
# ── notify-result: post workflow outcome to Mattermost ──
|
||||
- name: notify-result
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: platform
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
container:
|
||||
image: gitlab-registry.foxhunt.svc.cluster.local:5000/root/foxhunt/foxhunt-runtime:latest
|
||||
command: ["/bin/sh", "-c"]
|
||||
env:
|
||||
- name: WEBHOOK_URL
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: notification-webhook
|
||||
key: webhook-url
|
||||
optional: true
|
||||
resources:
|
||||
requests:
|
||||
cpu: 50m
|
||||
memory: 32Mi
|
||||
limits:
|
||||
cpu: 100m
|
||||
memory: 64Mi
|
||||
args:
|
||||
- |
|
||||
STATUS="{{workflow.status}}"
|
||||
NAME="{{workflow.name}}"
|
||||
|
||||
if [ -z "$WEBHOOK_URL" ] || echo "$WEBHOOK_URL" | grep -q "PLACEHOLDER"; then
|
||||
echo "No webhook configured, skipping notification"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
if [ "$STATUS" = "Succeeded" ]; then
|
||||
EMOJI=":white_check_mark:"
|
||||
else
|
||||
EMOJI=":x:"
|
||||
fi
|
||||
|
||||
PAYLOAD="{\"username\":\"Argo CI\",\"text\":\"${EMOJI} **${NAME}** — ${STATUS} ({{workflow.duration}}s)\"}"
|
||||
|
||||
curl -sf -X POST -H 'Content-Type: application/json' \
|
||||
-d "$PAYLOAD" "$WEBHOOK_URL" || echo "WARN: webhook post failed"
|
||||
@@ -1,614 +0,0 @@
|
||||
# Compile-and-Deploy: manual workflow for building and deploying service binaries.
|
||||
#
|
||||
# DAG:
|
||||
# create-tag ──> compile-services ──> upload-release ──> deploy-services
|
||||
#
|
||||
# Image builds handled by ci-pipeline, not here.
|
||||
# Training compilation lives in compile-and-train-template.yaml.
|
||||
---
|
||||
apiVersion: argoproj.io/v1alpha1
|
||||
kind: WorkflowTemplate
|
||||
metadata:
|
||||
name: compile-and-deploy
|
||||
namespace: foxhunt
|
||||
labels:
|
||||
app.kubernetes.io/name: compile-and-deploy
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
spec:
|
||||
entrypoint: pipeline
|
||||
onExit: notify-result
|
||||
serviceAccountName: argo-workflow
|
||||
podMetadata:
|
||||
labels:
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
app.kubernetes.io/component: compile-and-deploy
|
||||
archiveLogs: true
|
||||
podGC:
|
||||
strategy: OnPodCompletion
|
||||
ttlStrategy:
|
||||
secondsAfterCompletion: 3600
|
||||
activeDeadlineSeconds: 3600 # 1 hour (compile + deploy)
|
||||
|
||||
arguments:
|
||||
parameters:
|
||||
- name: commit-sha
|
||||
value: HEAD
|
||||
- name: service-packages
|
||||
value: "api trading-service backtesting-service trading-agent-service broker-gateway data-acquisition-service"
|
||||
|
||||
volumes:
|
||||
- name: git-ssh-key
|
||||
secret:
|
||||
secretName: argo-git-ssh-key
|
||||
defaultMode: 256
|
||||
- name: cargo-target-cpu
|
||||
persistentVolumeClaim:
|
||||
claimName: cargo-target-cpu
|
||||
- name: sccache
|
||||
persistentVolumeClaim:
|
||||
claimName: sccache-cpu
|
||||
|
||||
templates:
|
||||
- name: pipeline
|
||||
dag:
|
||||
tasks:
|
||||
- name: create-tag
|
||||
template: create-tag
|
||||
- name: compile-services
|
||||
depends: "create-tag"
|
||||
template: compile-services
|
||||
arguments:
|
||||
parameters:
|
||||
- name: tag
|
||||
value: "{{tasks.create-tag.outputs.parameters.tag}}"
|
||||
- name: service-packages
|
||||
value: "{{workflow.parameters.service-packages}}"
|
||||
- name: upload-release
|
||||
depends: "compile-services"
|
||||
template: upload-release
|
||||
arguments:
|
||||
parameters:
|
||||
- name: tag
|
||||
value: "{{tasks.create-tag.outputs.parameters.tag}}"
|
||||
- name: deploy-services
|
||||
depends: "upload-release"
|
||||
template: deploy-services
|
||||
arguments:
|
||||
parameters:
|
||||
- name: tag
|
||||
value: "{{tasks.create-tag.outputs.parameters.tag}}"
|
||||
- name: deploy-list
|
||||
value: "{{workflow.parameters.service-packages}}"
|
||||
|
||||
# ── create-tag: CalVer auto-tag on code changes ──
|
||||
- name: create-tag
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: platform
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
container:
|
||||
image: curlimages/curl:8.12.1
|
||||
command: ["/bin/sh", "-c"]
|
||||
env:
|
||||
- name: GITLAB_PAT
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: gitlab-pat
|
||||
key: token
|
||||
resources:
|
||||
requests:
|
||||
cpu: 100m
|
||||
memory: 64Mi
|
||||
limits:
|
||||
cpu: 200m
|
||||
memory: 128Mi
|
||||
args:
|
||||
- |
|
||||
set -e
|
||||
GITLAB="http://gitlab-webservice-default.foxhunt.svc.cluster.local:8181"
|
||||
PROJECT_ID=1
|
||||
SHA="{{workflow.parameters.commit-sha}}"
|
||||
|
||||
# Compute CalVer prefix: vYYYY.MM
|
||||
PREFIX="v$(date +%Y.%m)"
|
||||
|
||||
# Query existing tags for this month
|
||||
TAGS=$(curl -sf \
|
||||
-H "PRIVATE-TOKEN: ${GITLAB_PAT}" \
|
||||
"${GITLAB}/api/v4/projects/${PROJECT_ID}/repository/tags?search=${PREFIX}" \
|
||||
|| echo "[]")
|
||||
|
||||
# Parse highest N from vYYYY.MM.N tags
|
||||
LAST_N=$(echo "$TAGS" | grep -oP "\"name\":\"${PREFIX}\.\K[0-9]+" | sort -n | tail -1)
|
||||
if [ -z "$LAST_N" ]; then
|
||||
NEXT_N=1
|
||||
else
|
||||
NEXT_N=$((LAST_N + 1))
|
||||
fi
|
||||
|
||||
TAG="${PREFIX}.${NEXT_N}"
|
||||
echo "Creating tag: ${TAG} at ${SHA}"
|
||||
|
||||
# Create the tag (tolerate failure if tag already exists)
|
||||
HTTP_CODE=$(curl -s -o /tmp/tag_response -w "%{http_code}" -X POST \
|
||||
-H "PRIVATE-TOKEN: ${GITLAB_PAT}" \
|
||||
-d "tag_name=${TAG}" \
|
||||
-d "ref=${SHA}" \
|
||||
"${GITLAB}/api/v4/projects/${PROJECT_ID}/repository/tags")
|
||||
|
||||
RESULT=$(cat /tmp/tag_response 2>/dev/null || echo "{}")
|
||||
|
||||
if [ "$HTTP_CODE" = "201" ]; then
|
||||
echo "Tag ${TAG} created successfully"
|
||||
elif [ "$HTTP_CODE" = "400" ] && echo "$RESULT" | grep -q "already exists"; then
|
||||
echo "Tag ${TAG} already exists — reusing"
|
||||
else
|
||||
echo "WARN: Tag creation returned HTTP ${HTTP_CODE}: ${RESULT}"
|
||||
echo "Falling back to dev tag"
|
||||
TAG="dev-$(echo $SHA | cut -c1-8)"
|
||||
fi
|
||||
|
||||
mkdir -p /tmp/outputs
|
||||
echo -n "$TAG" > /tmp/outputs/tag
|
||||
outputs:
|
||||
parameters:
|
||||
- name: tag
|
||||
valueFrom:
|
||||
path: /tmp/outputs/tag
|
||||
|
||||
# ── compile-services: selective per-binary build, incremental on local RWO PVC ──
|
||||
- name: compile-services
|
||||
metadata:
|
||||
labels:
|
||||
app.kubernetes.io/component: compile
|
||||
inputs:
|
||||
parameters:
|
||||
- name: tag
|
||||
- name: service-packages
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: ci-compile-cpu
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
tolerations:
|
||||
- key: node.cilium.io/agent-not-ready
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
# ── seed-deps-cache: rsync prebuilt third-party rlibs into the PVC ──
|
||||
# On a true cold PVC (e.g. fresh node, first run, or after a purge),
|
||||
# this saves 2-4 min that would otherwise be spent compiling the
|
||||
# ~1500 third-party crates in the workspace dep graph. The image is
|
||||
# rebuilt nightly by refresh-deps-cache (see
|
||||
# refresh-deps-cache-template.yaml). On a warm PVC, the stamp file
|
||||
# short-circuits the whole step in ~50ms.
|
||||
#
|
||||
# Why rsync --ignore-existing instead of cp/overwrite: a warmer PVC
|
||||
# may already have NEWER artifacts from a prior compile; we don't
|
||||
# want to clobber them with potentially-stale prebuilt rlibs. Cargo
|
||||
# will discard rlibs whose fingerprint mismatches anyway.
|
||||
#
|
||||
# If the deps-cache image isn't published yet (first deploy of this
|
||||
# template), set imagePullPolicy: IfNotPresent + the initContainer
|
||||
# is non-blocking-on-failure (`|| true`) so the compile-services
|
||||
# step still runs without the cache prelude.
|
||||
initContainers:
|
||||
- name: seed-deps-cache
|
||||
image: gitlab-registry.foxhunt.svc.cluster.local:5000/root/foxhunt/ci-builder-cpu-with-deps:nightly
|
||||
imagePullPolicy: IfNotPresent
|
||||
env:
|
||||
- name: DEPS_VERSION
|
||||
value: "1"
|
||||
resources:
|
||||
requests:
|
||||
cpu: 500m
|
||||
memory: 512Mi
|
||||
limits:
|
||||
cpu: "2"
|
||||
memory: 1Gi
|
||||
volumeMounts:
|
||||
- name: cargo-target-cpu
|
||||
mountPath: /cargo-target
|
||||
command: ["/bin/sh", "-c"]
|
||||
args:
|
||||
- |
|
||||
set -e
|
||||
STAMP="/cargo-target/cpu_deps_v${DEPS_VERSION}.stamp"
|
||||
if [ -f "$STAMP" ]; then
|
||||
echo "PVC already seeded for deps v${DEPS_VERSION} (stamp present), skipping."
|
||||
exit 0
|
||||
fi
|
||||
if [ ! -d /cargo-target-prebuilt ]; then
|
||||
echo "WARN: /cargo-target-prebuilt missing in image — image not yet published?"
|
||||
echo " Compile will run without prebuilt deps cache (slower cold-start)."
|
||||
exit 0
|
||||
fi
|
||||
echo "=== Seeding cargo-target-cpu PVC from prebuilt deps cache (v${DEPS_VERSION}) ==="
|
||||
du -sh /cargo-target-prebuilt 2>/dev/null || true
|
||||
# --ignore-existing: never clobber newer artifacts already on PVC.
|
||||
# -a: preserve mtimes, perms, links — cargo's freshness check needs accurate mtimes.
|
||||
rsync -a --ignore-existing /cargo-target-prebuilt/ /cargo-target/
|
||||
touch "$STAMP"
|
||||
echo "=== Seed complete ==="
|
||||
du -sh /cargo-target/release 2>/dev/null || true
|
||||
container:
|
||||
image: gitlab-registry.foxhunt.svc.cluster.local:5000/root/foxhunt/ci-builder-cpu:latest
|
||||
imagePullPolicy: Always
|
||||
command: ["/bin/sh", "-c"]
|
||||
env:
|
||||
- name: SQLX_OFFLINE
|
||||
value: "true"
|
||||
- name: CARGO_TERM_COLOR
|
||||
value: always
|
||||
- name: CARGO_TARGET_DIR
|
||||
value: /cargo-target
|
||||
- name: CARGO_HOME
|
||||
value: /cargo-target/cargo-home
|
||||
# sccache: rustc wrapper that content-hashes compile inputs and reuses
|
||||
# object output across pods/commits via the sccache-cpu PVC. Survives
|
||||
# cargo's incremental-cache invalidation from git checkout mtime touches.
|
||||
- name: RUSTC_WRAPPER
|
||||
value: sccache
|
||||
- name: SCCACHE_DIR
|
||||
value: /sccache-cache
|
||||
- name: SCCACHE_CACHE_SIZE
|
||||
value: 20G
|
||||
# Required: workspace `.cargo/config.toml` sets [build] incremental = true,
|
||||
# which forces incremental compilation even on --release. sccache cannot
|
||||
# cache incremental rustc output (documented limitation). Override here
|
||||
# so sccache actually catches. Local dev keeps incremental via config.toml.
|
||||
- name: CARGO_INCREMENTAL
|
||||
value: "0"
|
||||
# Match the cgroup cpu limit below (limits.cpu: "30").
|
||||
# Without this, cargo asks num_cpus::get_physical() — the host's
|
||||
# whole CPU count — and over-subscribes vs the cgroup throttle,
|
||||
# producing scheduling waste under load.
|
||||
- name: CARGO_BUILD_JOBS
|
||||
value: "30"
|
||||
- name: GITLAB_PAT
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: gitlab-pat
|
||||
key: token
|
||||
resources:
|
||||
requests:
|
||||
cpu: "14"
|
||||
memory: 16Gi
|
||||
limits:
|
||||
cpu: "30"
|
||||
memory: 32Gi
|
||||
volumeMounts:
|
||||
- name: git-ssh-key
|
||||
mountPath: /etc/git-ssh
|
||||
readOnly: true
|
||||
- name: cargo-target-cpu
|
||||
mountPath: /cargo-target
|
||||
- name: sccache
|
||||
mountPath: /sccache-cache
|
||||
args:
|
||||
- |
|
||||
set -e
|
||||
SHA="{{workflow.parameters.commit-sha}}"
|
||||
SERVICE_PKGS="{{inputs.parameters.service-packages}}"
|
||||
|
||||
# SSH setup
|
||||
mkdir -p ~/.ssh
|
||||
cp /etc/git-ssh/ssh-privatekey ~/.ssh/id_ed25519
|
||||
chmod 600 ~/.ssh/id_ed25519
|
||||
printf 'Host *\n StrictHostKeyChecking no\n UserKnownHostsFile /dev/null\n' > ~/.ssh/config
|
||||
chmod 600 ~/.ssh/config
|
||||
|
||||
REPO="ssh://git@gitlab-gitlab-shell.foxhunt.svc.cluster.local:2222/root/foxhunt.git"
|
||||
WORKSPACE="/cargo-target/src"
|
||||
|
||||
# PVC may be owned by a different UID from a previous run
|
||||
git config --global --add safe.directory "$WORKSPACE"
|
||||
|
||||
# Persistent checkout on PVC — only changed files get new mtimes,
|
||||
# so cargo skips recompiling unchanged workspace crates.
|
||||
if [ -d "$WORKSPACE/.git" ]; then
|
||||
cd "$WORKSPACE"
|
||||
CURRENT=$(git rev-parse HEAD 2>/dev/null || echo "none")
|
||||
if [ "$CURRENT" = "$SHA" ]; then
|
||||
echo "=== Already at $SHA, skipping checkout ==="
|
||||
else
|
||||
echo "=== Updating checkout: $(echo $CURRENT | cut -c1-8) -> $(echo $SHA | cut -c1-8) ==="
|
||||
git fetch origin
|
||||
git checkout --force "$SHA"
|
||||
git clean -fd
|
||||
fi
|
||||
else
|
||||
echo "=== Initial clone (first run) ==="
|
||||
git clone --filter=blob:none "$REPO" "$WORKSPACE"
|
||||
cd "$WORKSPACE"
|
||||
git checkout "$SHA"
|
||||
fi
|
||||
echo "Checked out $(git rev-parse --short HEAD)"
|
||||
|
||||
# Ensure cargo home registry is on PVC
|
||||
export PATH="${CARGO_HOME}/bin:${PATH}"
|
||||
|
||||
export FOXHUNT_BUILD_VERSION="{{inputs.parameters.tag}}"
|
||||
|
||||
# Prune build artifacts if PVC exceeds 25GB (prevents unbounded growth)
|
||||
TARGET_SIZE_MB=$(du -sm "$CARGO_TARGET_DIR" 2>/dev/null | cut -f1 || echo 0)
|
||||
echo "PVC usage: ${TARGET_SIZE_MB}MB"
|
||||
if [ "$TARGET_SIZE_MB" -gt 25000 ]; then
|
||||
echo "PVC exceeds 25GB limit, pruning build artifacts..."
|
||||
rm -rf "$CARGO_TARGET_DIR/release" "$CARGO_TARGET_DIR/debug"
|
||||
fi
|
||||
|
||||
# Guard: empty package list would build entire workspace
|
||||
if [ -z "$SERVICE_PKGS" ]; then
|
||||
echo "ERROR: service-packages is empty, refusing to build entire workspace"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
# Build only the affected service packages (incremental via persistent target dir)
|
||||
CARGO_ARGS=""
|
||||
for pkg in $SERVICE_PKGS; do
|
||||
CARGO_ARGS="$CARGO_ARGS -p $pkg"
|
||||
done
|
||||
|
||||
# CI-only linker swap: prefer wild over mold. wild is a Rust-native
|
||||
# linker, typically 10-30% faster than mold on release-LTO links;
|
||||
# we have ~5 service binaries each doing release-LTO, so this is
|
||||
# a meaningful 30-90 sec wall-time saving. Both linkers ship in
|
||||
# Dockerfile.ci-builder-cpu — to revert, drop this sed and rebuild.
|
||||
# Sed is idempotent (no-op if wild already substituted from a prior
|
||||
# run on the same PVC checkout) and surgical (single line in
|
||||
# .cargo/config.toml). If wild is missing for any reason, the build
|
||||
# falls back to mold once we revert this hunk.
|
||||
if command -v wild >/dev/null 2>&1; then
|
||||
echo "=== Swapping linker mold -> wild for this CI run ==="
|
||||
sed -i 's|-fuse-ld=mold|-fuse-ld=wild|' .cargo/config.toml
|
||||
grep -n 'fuse-ld' .cargo/config.toml || true
|
||||
else
|
||||
echo "WARN: wild not on PATH, sticking with mold"
|
||||
fi
|
||||
|
||||
echo "=== Building service binaries: $SERVICE_PKGS (incremental) ==="
|
||||
# --locked: skip Cargo.lock resolver work (~5-15s saved) and fail fast
|
||||
# if the lock has drifted (catches accidental Cargo.toml edits without
|
||||
# a re-resolve commit).
|
||||
cargo build --locked --release $CARGO_ARGS
|
||||
|
||||
# Collect built binaries
|
||||
mkdir -p "$WORKSPACE/bin/services"
|
||||
for pkg in $SERVICE_PKGS; do
|
||||
bin_name=$(echo "$pkg" | tr '-' '_')
|
||||
cp "$CARGO_TARGET_DIR/release/$pkg" "$WORKSPACE/bin/services/" 2>/dev/null \
|
||||
|| cp "$CARGO_TARGET_DIR/release/$bin_name" "$WORKSPACE/bin/services/" 2>/dev/null \
|
||||
|| { echo "Binary not found for $pkg"; ls "$CARGO_TARGET_DIR/release/"; exit 1; }
|
||||
done
|
||||
strip "$WORKSPACE/bin/services/"*
|
||||
|
||||
echo "=== Service binaries ==="
|
||||
ls -lh "$WORKSPACE/bin/services/"
|
||||
|
||||
echo "=== Uploading service binaries to GitLab packages ==="
|
||||
GITLAB="http://gitlab-webservice-default.foxhunt.svc.cluster.local:8181"
|
||||
TAG="${FOXHUNT_BUILD_VERSION}"
|
||||
for bin in "$WORKSPACE/bin/services/"*; do
|
||||
BIN_NAME=$(basename "$bin")
|
||||
echo "Uploading ${BIN_NAME} (${TAG})..."
|
||||
curl -f --upload-file "$bin" \
|
||||
-H "PRIVATE-TOKEN: ${GITLAB_PAT}" \
|
||||
"${GITLAB}/api/v4/projects/1/packages/generic/foxhunt-services/${TAG}/${BIN_NAME}"
|
||||
done
|
||||
|
||||
# Update 'latest' per-file (preserves unbuilt binaries from prior runs)
|
||||
echo "=== Updating 'latest' package ==="
|
||||
LATEST_PKG=$(curl -sf -H "PRIVATE-TOKEN: ${GITLAB_PAT}" \
|
||||
"${GITLAB}/api/v4/projects/1/packages?package_name=foxhunt-services&package_version=latest" \
|
||||
| grep -oP '"id":\K[0-9]+' | head -1)
|
||||
for bin in "$WORKSPACE/bin/services/"*; do
|
||||
BIN_NAME=$(basename "$bin")
|
||||
# Delete existing file by name before uploading replacement
|
||||
if [ -n "$LATEST_PKG" ]; then
|
||||
FILE_ID=$(curl -sf -H "PRIVATE-TOKEN: ${GITLAB_PAT}" \
|
||||
"${GITLAB}/api/v4/projects/1/packages/${LATEST_PKG}/package_files" \
|
||||
| grep -oP "\"id\":([0-9]+),\"package_id\":${LATEST_PKG}[^}]*\"file_name\":\"${BIN_NAME}\"" \
|
||||
| grep -oP '"id":\K[0-9]+' | head -1)
|
||||
if [ -n "$FILE_ID" ]; then
|
||||
curl -sf -X DELETE -H "PRIVATE-TOKEN: ${GITLAB_PAT}" \
|
||||
"${GITLAB}/api/v4/projects/1/packages/${LATEST_PKG}/package_files/${FILE_ID}" || true
|
||||
fi
|
||||
fi
|
||||
curl -f --upload-file "$bin" \
|
||||
-H "PRIVATE-TOKEN: ${GITLAB_PAT}" \
|
||||
"${GITLAB}/api/v4/projects/1/packages/generic/foxhunt-services/latest/${BIN_NAME}" || true
|
||||
done
|
||||
|
||||
echo "=== sccache stats ==="
|
||||
sccache --show-stats || true
|
||||
|
||||
echo "=== Service compile + upload done ($SERVICE_PKGS) ==="
|
||||
|
||||
# ── upload-release: create GitLab Release with package links ──
|
||||
- name: upload-release
|
||||
inputs:
|
||||
parameters:
|
||||
- name: tag
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: platform
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
container:
|
||||
image: curlimages/curl:8.12.1
|
||||
command: ["/bin/sh", "-c"]
|
||||
env:
|
||||
- name: GITLAB_PAT
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: gitlab-pat
|
||||
key: token
|
||||
resources:
|
||||
requests:
|
||||
cpu: 100m
|
||||
memory: 64Mi
|
||||
limits:
|
||||
cpu: 200m
|
||||
memory: 128Mi
|
||||
args:
|
||||
- |
|
||||
set -e
|
||||
GITLAB="http://gitlab-webservice-default.foxhunt.svc.cluster.local:8181"
|
||||
PROJECT_ID=1
|
||||
TAG="{{inputs.parameters.tag}}"
|
||||
|
||||
echo "=== Creating GitLab Release ${TAG} ==="
|
||||
|
||||
# Get commits since previous tag for release notes
|
||||
PREV_TAG=$(curl -sf \
|
||||
-H "PRIVATE-TOKEN: ${GITLAB_PAT}" \
|
||||
"${GITLAB}/api/v4/projects/${PROJECT_ID}/repository/tags?per_page=2" \
|
||||
| grep -oP '"name":"\K[^"]+' | sed -n '2p')
|
||||
|
||||
if [ -n "$PREV_TAG" ]; then
|
||||
COMMITS=$(curl -sf \
|
||||
-H "PRIVATE-TOKEN: ${GITLAB_PAT}" \
|
||||
"${GITLAB}/api/v4/projects/${PROJECT_ID}/repository/compare?from=${PREV_TAG}&to=${TAG}" \
|
||||
| grep -oP '"title":"\K[^"]+' | head -20 \
|
||||
| sed 's/^/- /' || echo "- Release ${TAG}")
|
||||
DESCRIPTION="## Changes since ${PREV_TAG}\n\n${COMMITS}"
|
||||
else
|
||||
DESCRIPTION="## Initial release\n\nFirst CalVer release."
|
||||
fi
|
||||
|
||||
# Create the release
|
||||
RESPONSE=$(curl -sf -X POST \
|
||||
-H "PRIVATE-TOKEN: ${GITLAB_PAT}" \
|
||||
-H "Content-Type: application/json" \
|
||||
-d "{
|
||||
\"tag_name\": \"${TAG}\",
|
||||
\"name\": \"${TAG}\",
|
||||
\"description\": \"$(printf '%s' "$DESCRIPTION" | sed 's/"/\\"/g')\"
|
||||
}" \
|
||||
"${GITLAB}/api/v4/projects/${PROJECT_ID}/releases") || {
|
||||
echo "WARN: Release creation failed (may already exist)"
|
||||
}
|
||||
|
||||
echo "Release ${TAG} created"
|
||||
echo "$RESPONSE" | head -5
|
||||
|
||||
# Add package links as release assets
|
||||
for pkg_name in foxhunt-services foxhunt-training; do
|
||||
PKG_CHECK=$(curl -sf \
|
||||
-H "PRIVATE-TOKEN: ${GITLAB_PAT}" \
|
||||
"${GITLAB}/api/v4/projects/${PROJECT_ID}/packages?package_name=${pkg_name}&package_version=${TAG}" \
|
||||
|| echo "[]")
|
||||
|
||||
if echo "$PKG_CHECK" | grep -q "$TAG"; then
|
||||
curl -sf -X POST \
|
||||
-H "PRIVATE-TOKEN: ${GITLAB_PAT}" \
|
||||
-H "Content-Type: application/json" \
|
||||
-d "{
|
||||
\"name\": \"${pkg_name}\",
|
||||
\"url\": \"${GITLAB}/-/packages?type=generic&search=${pkg_name}&version=${TAG}\",
|
||||
\"link_type\": \"package\"
|
||||
}" \
|
||||
"${GITLAB}/api/v4/projects/${PROJECT_ID}/releases/${TAG}/assets/links" || true
|
||||
echo "Linked ${pkg_name} package to release"
|
||||
fi
|
||||
done
|
||||
|
||||
echo "=== Release ${TAG} complete ==="
|
||||
|
||||
# ── deploy-services: selective rolling restart for affected deployments ──
|
||||
- name: deploy-services
|
||||
inputs:
|
||||
parameters:
|
||||
- name: tag
|
||||
- name: deploy-list
|
||||
serviceAccountName: argo-workflow
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: platform
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
container:
|
||||
image: curlimages/curl:8.12.1
|
||||
command: ["/bin/sh", "-c"]
|
||||
resources:
|
||||
requests:
|
||||
cpu: 100m
|
||||
memory: 64Mi
|
||||
limits:
|
||||
cpu: 200m
|
||||
memory: 128Mi
|
||||
args:
|
||||
- |
|
||||
set -e
|
||||
if ! command -v kubectl >/dev/null 2>&1; then
|
||||
echo "=== Installing kubectl ==="
|
||||
curl -sLo /tmp/kubectl "https://dl.k8s.io/release/v1.31.4/bin/linux/amd64/kubectl"
|
||||
chmod +x /tmp/kubectl
|
||||
export PATH="/tmp:$PATH"
|
||||
fi
|
||||
|
||||
TAG="{{inputs.parameters.tag}}"
|
||||
DEPLOY_LIST="{{inputs.parameters.deploy-list}}"
|
||||
|
||||
if [ -z "$DEPLOY_LIST" ]; then
|
||||
echo "=== No services to deploy ==="
|
||||
exit 0
|
||||
fi
|
||||
|
||||
echo "=== Deploying release ${TAG}: $DEPLOY_LIST ==="
|
||||
|
||||
for svc in $DEPLOY_LIST; do
|
||||
echo "Patching $svc with FOXHUNT_RELEASE=${TAG}..."
|
||||
kubectl -n foxhunt patch deployment "$svc" -p "{
|
||||
\"spec\":{\"template\":{
|
||||
\"metadata\":{\"annotations\":{\"foxhunt.io/release\":\"${TAG}\"}},
|
||||
\"spec\":{\"initContainers\":[{
|
||||
\"name\":\"fetch-binary\",
|
||||
\"env\":[{\"name\":\"FOXHUNT_RELEASE\",\"value\":\"${TAG}\"}]
|
||||
}]}
|
||||
}}
|
||||
}" || echo "WARN: $svc patch failed (may not exist)"
|
||||
done
|
||||
|
||||
echo "=== Waiting for rollouts ==="
|
||||
for svc in $DEPLOY_LIST; do
|
||||
kubectl -n foxhunt rollout status deployment "$svc" --timeout=120s || echo "WARN: $svc rollout timeout"
|
||||
done
|
||||
|
||||
echo "=== Deploy ${TAG} complete ($DEPLOY_LIST) ==="
|
||||
|
||||
# ── notify-result: post workflow outcome to Mattermost ──
|
||||
- name: notify-result
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: platform
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
container:
|
||||
image: curlimages/curl:8.12.1
|
||||
command: ["/bin/sh", "-c"]
|
||||
env:
|
||||
- name: WEBHOOK_URL
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: notification-webhook
|
||||
key: webhook-url
|
||||
optional: true
|
||||
resources:
|
||||
requests:
|
||||
cpu: 50m
|
||||
memory: 32Mi
|
||||
limits:
|
||||
cpu: 100m
|
||||
memory: 64Mi
|
||||
args:
|
||||
- |
|
||||
STATUS="{{workflow.status}}"
|
||||
NAME="{{workflow.name}}"
|
||||
|
||||
if [ -z "$WEBHOOK_URL" ] || echo "$WEBHOOK_URL" | grep -q "PLACEHOLDER"; then
|
||||
echo "No webhook configured, skipping notification"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
if [ "$STATUS" = "Succeeded" ]; then
|
||||
EMOJI=":white_check_mark:"
|
||||
else
|
||||
EMOJI=":x:"
|
||||
fi
|
||||
|
||||
PAYLOAD="{\"username\":\"Argo CI\",\"text\":\"${EMOJI} **${NAME}** — ${STATUS} ({{workflow.duration}}s)\"}"
|
||||
|
||||
curl -sf -X POST -H 'Content-Type: application/json' \
|
||||
-d "$PAYLOAD" "$WEBHOOK_URL" || echo "WARN: webhook post failed"
|
||||
@@ -1,12 +0,0 @@
|
||||
apiVersion: v1
|
||||
kind: ConfigMap
|
||||
metadata:
|
||||
name: auto-compile-config
|
||||
namespace: foxhunt
|
||||
labels:
|
||||
app.kubernetes.io/name: auto-compile-config
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
data:
|
||||
# Set to "true" to enable auto-compile on push to main.
|
||||
# kubectl -n foxhunt patch configmap auto-compile-config -p '{"data":{"enabled":"true"}}'
|
||||
enabled: "false"
|
||||
@@ -1,70 +0,0 @@
|
||||
apiVersion: argoproj.io/v1alpha1
|
||||
kind: Sensor
|
||||
metadata:
|
||||
name: ci-pipeline-trigger
|
||||
namespace: foxhunt
|
||||
labels:
|
||||
app.kubernetes.io/name: ci-pipeline-trigger
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
spec:
|
||||
template:
|
||||
serviceAccountName: argo-workflow
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: platform
|
||||
eventBusName: default
|
||||
dependencies:
|
||||
- name: gitlab-push-dep
|
||||
eventSourceName: gitlab-push
|
||||
eventName: gitlab-push
|
||||
filters:
|
||||
data:
|
||||
- path: body.ref
|
||||
type: string
|
||||
value:
|
||||
- "refs/heads/main"
|
||||
# Policy (user directive): push-triggered runs are infra-sync only.
|
||||
# The ci-pipeline DAG runs three conditional task groups on push:
|
||||
# - docker image rebuilds (gated on infra/docker/ changes)
|
||||
# - apply-argo-templates (gated on infra/k8s/argo/ changes)
|
||||
# - terragrunt-apply (gated on infra/live/ or infra/modules/ changes)
|
||||
# Non-infra tasks (test-gate, build-web-dashboard, gpu-test) have `when: "false"`
|
||||
# in the WorkflowTemplate. Training + compile/deploy are separate templates,
|
||||
# always manual.
|
||||
#
|
||||
# All non-push workflows (cargo tests, training, compile/deploy, GPU tests)
|
||||
# run via scripts/manual `argo submit`.
|
||||
|
||||
triggers:
|
||||
- template:
|
||||
name: ci-pipeline
|
||||
argoWorkflow:
|
||||
operation: submit
|
||||
source:
|
||||
resource:
|
||||
apiVersion: argoproj.io/v1alpha1
|
||||
kind: Workflow
|
||||
metadata:
|
||||
generateName: ci-pipeline-
|
||||
namespace: foxhunt
|
||||
spec:
|
||||
serviceAccountName: argo-workflow
|
||||
podGC:
|
||||
strategy: OnPodCompletion
|
||||
ttlStrategy:
|
||||
secondsAfterCompletion: 3600
|
||||
workflowTemplateRef:
|
||||
name: ci-pipeline
|
||||
arguments:
|
||||
parameters:
|
||||
- name: commit-sha
|
||||
- name: commits-json
|
||||
parameters:
|
||||
- src:
|
||||
dependencyName: gitlab-push-dep
|
||||
dataKey: body.checkout_sha
|
||||
dest: spec.arguments.parameters.0.value
|
||||
- src:
|
||||
dependencyName: gitlab-push-dep
|
||||
dataKey: body.commits
|
||||
dataTemplate: "{{ toJson .Input }}"
|
||||
dest: spec.arguments.parameters.1.value
|
||||
@@ -1,15 +0,0 @@
|
||||
apiVersion: argoproj.io/v1alpha1
|
||||
kind: EventSource
|
||||
metadata:
|
||||
name: gitlab-push
|
||||
namespace: foxhunt
|
||||
labels:
|
||||
app.kubernetes.io/name: gitlab-push-eventsource
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
spec:
|
||||
eventBusName: default
|
||||
webhook:
|
||||
gitlab-push:
|
||||
port: "12000"
|
||||
endpoint: /push
|
||||
method: POST
|
||||
@@ -1,93 +0,0 @@
|
||||
apiVersion: networking.k8s.io/v1
|
||||
kind: NetworkPolicy
|
||||
metadata:
|
||||
name: argo-gpu-test-workflow
|
||||
namespace: foxhunt
|
||||
labels:
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
spec:
|
||||
podSelector:
|
||||
matchLabels:
|
||||
app.kubernetes.io/component: gpu-test
|
||||
policyTypes:
|
||||
- Egress
|
||||
egress:
|
||||
# DNS
|
||||
- ports:
|
||||
- port: 53
|
||||
protocol: UDP
|
||||
- port: 53
|
||||
protocol: TCP
|
||||
# Kubernetes API + internal services (service CIDR)
|
||||
- ports:
|
||||
- port: 443
|
||||
protocol: TCP
|
||||
- port: 6443
|
||||
protocol: TCP
|
||||
to:
|
||||
- ipBlock:
|
||||
cidr: 10.32.0.0/16
|
||||
- ipBlock:
|
||||
cidr: 172.16.0.4/32
|
||||
# HTTPS egress (crates.io, etc.)
|
||||
- ports:
|
||||
- port: 443
|
||||
protocol: TCP
|
||||
# Git SSH (GitLab)
|
||||
- ports:
|
||||
- port: 2222
|
||||
protocol: TCP
|
||||
to:
|
||||
- ipBlock:
|
||||
cidr: 100.90.76.85/32
|
||||
- podSelector:
|
||||
matchLabels:
|
||||
app: gitlab-shell
|
||||
# GitLab webservice (git-http)
|
||||
- ports:
|
||||
- port: 8181
|
||||
protocol: TCP
|
||||
to:
|
||||
- podSelector:
|
||||
matchLabels:
|
||||
app: webservice
|
||||
# MinIO (artifact/log storage)
|
||||
- ports:
|
||||
- port: 9000
|
||||
protocol: TCP
|
||||
to:
|
||||
- podSelector:
|
||||
matchLabels:
|
||||
app.kubernetes.io/name: minio
|
||||
# Pushgateway (Prometheus metrics)
|
||||
- ports:
|
||||
- port: 9091
|
||||
protocol: TCP
|
||||
to:
|
||||
- podSelector:
|
||||
matchLabels:
|
||||
app.kubernetes.io/name: pushgateway
|
||||
# OTLP (Tempo)
|
||||
- ports:
|
||||
- port: 4317
|
||||
protocol: TCP
|
||||
to:
|
||||
- podSelector:
|
||||
matchLabels:
|
||||
app.kubernetes.io/name: tempo
|
||||
# Mattermost (notifications)
|
||||
- ports:
|
||||
- port: 8065
|
||||
protocol: TCP
|
||||
to:
|
||||
- podSelector:
|
||||
matchLabels:
|
||||
app.kubernetes.io/name: mattermost
|
||||
# GitLab container registry
|
||||
- ports:
|
||||
- port: 5000
|
||||
protocol: TCP
|
||||
to:
|
||||
- podSelector:
|
||||
matchLabels:
|
||||
app: registry
|
||||
@@ -1,25 +0,0 @@
|
||||
# infra/k8s/argo/gpu-test-nightly-cron.yaml
|
||||
apiVersion: argoproj.io/v1alpha1
|
||||
kind: CronWorkflow
|
||||
metadata:
|
||||
name: gpu-test-nightly
|
||||
namespace: foxhunt
|
||||
labels:
|
||||
app.kubernetes.io/name: gpu-test-nightly
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
spec:
|
||||
schedule: "0 2 * * *"
|
||||
timezone: "UTC"
|
||||
suspend: true # Enable when ready: kubectl patch cronworkflow gpu-test-nightly -n foxhunt -p '{"spec":{"suspend":false}}'
|
||||
concurrencyPolicy: Replace
|
||||
successfulJobsHistoryLimit: 3
|
||||
failedJobsHistoryLimit: 5
|
||||
workflowSpec:
|
||||
workflowTemplateRef:
|
||||
name: gpu-test-pipeline
|
||||
arguments:
|
||||
parameters:
|
||||
- name: models
|
||||
value: "dqn,ppo,tft,mamba2,tggn,tlob,liquid,kan,xlstm,diffusion"
|
||||
- name: commit-ref
|
||||
value: "main"
|
||||
@@ -1,529 +0,0 @@
|
||||
# infra/k8s/argo/gpu-test-pipeline-template.yaml
|
||||
apiVersion: argoproj.io/v1alpha1
|
||||
kind: WorkflowTemplate
|
||||
metadata:
|
||||
name: gpu-test-pipeline
|
||||
namespace: foxhunt
|
||||
labels:
|
||||
app.kubernetes.io/name: gpu-test-pipeline
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
spec:
|
||||
entrypoint: pipeline
|
||||
onExit: notify-result
|
||||
serviceAccountName: argo-workflow
|
||||
podMetadata:
|
||||
labels:
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
app.kubernetes.io/component: gpu-test
|
||||
securityContext:
|
||||
fsGroup: 0
|
||||
archiveLogs: true
|
||||
podGC:
|
||||
strategy: OnPodCompletion
|
||||
ttlStrategy:
|
||||
secondsAfterCompletion: 3600
|
||||
activeDeadlineSeconds: 7200
|
||||
arguments:
|
||||
parameters:
|
||||
- name: commit-ref
|
||||
value: HEAD
|
||||
- name: models
|
||||
value: "dqn,ppo,tft"
|
||||
- name: test-scope
|
||||
value: all
|
||||
- name: gpu-pool
|
||||
value: ci-training-h100
|
||||
- name: cuda-compute-cap
|
||||
value: "90"
|
||||
- name: notify
|
||||
value: "true"
|
||||
volumes:
|
||||
- name: git-ssh-key
|
||||
secret:
|
||||
secretName: argo-git-ssh-key
|
||||
defaultMode: 256
|
||||
- name: cargo-target
|
||||
persistentVolumeClaim:
|
||||
claimName: cargo-target-cuda-test
|
||||
- name: test-data
|
||||
persistentVolumeClaim:
|
||||
claimName: test-data-pvc
|
||||
readOnly: true
|
||||
|
||||
templates:
|
||||
# ── pipeline: DAG entrypoint ──
|
||||
# compile-and-test runs on GPU node (RWO PVC can't be shared cross-node).
|
||||
# gpu-warmup ensures H100 is scaled up before compile starts.
|
||||
- name: pipeline
|
||||
dag:
|
||||
tasks:
|
||||
- name: gpu-warmup
|
||||
template: gpu-warmup
|
||||
- name: compile-and-test
|
||||
template: compile-and-test
|
||||
dependencies: [gpu-warmup]
|
||||
- name: perf-benchmark
|
||||
template: perf-benchmark
|
||||
dependencies: [compile-and-test]
|
||||
|
||||
# ── gpu-warmup: trigger H100 autoscale ──
|
||||
# Requests GPU to force autoscaler to provision node, then releases it.
|
||||
- name: gpu-warmup
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: "{{workflow.parameters.gpu-pool}}"
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
tolerations:
|
||||
- key: nvidia.com/gpu
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
- key: node.cilium.io/agent-not-ready
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
container:
|
||||
image: alpine:3.21
|
||||
command: ["/bin/sh", "-c"]
|
||||
args:
|
||||
- |
|
||||
echo "GPU warmup: triggering node autoscale..."
|
||||
nvidia-smi --query-gpu=name,memory.total --format=csv,noheader || true
|
||||
echo "GPU node ready, releasing for compile-and-test"
|
||||
resources:
|
||||
requests:
|
||||
nvidia.com/gpu: "1"
|
||||
cpu: 100m
|
||||
memory: 64Mi
|
||||
limits:
|
||||
nvidia.com/gpu: "1"
|
||||
cpu: 200m
|
||||
memory: 128Mi
|
||||
|
||||
# ── compile-and-test: compile + run GPU tests in single H100 pod ──
|
||||
- name: compile-and-test
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: "{{workflow.parameters.gpu-pool}}"
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
tolerations:
|
||||
- key: nvidia.com/gpu
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
- key: node.cilium.io/agent-not-ready
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
outputs:
|
||||
parameters:
|
||||
- name: results
|
||||
valueFrom:
|
||||
path: /tmp/outputs/results
|
||||
default: "unknown:FAIL"
|
||||
- name: failures
|
||||
valueFrom:
|
||||
path: /tmp/outputs/failures
|
||||
default: "1"
|
||||
container:
|
||||
image: gitlab-registry.foxhunt.svc.cluster.local:5000/root/foxhunt/ci-builder:latest
|
||||
command: ["/bin/bash", "-c"]
|
||||
env:
|
||||
- name: SQLX_OFFLINE
|
||||
value: "true"
|
||||
- name: CARGO_TERM_COLOR
|
||||
value: always
|
||||
- name: CARGO_TARGET_DIR
|
||||
value: /cargo-target
|
||||
- name: CARGO_HOME
|
||||
value: /cargo-target/cargo-home
|
||||
- name: CUDA_VISIBLE_DEVICES
|
||||
value: "0"
|
||||
- name: CUBLAS_WORKSPACE_CONFIG
|
||||
value: ":4096:8"
|
||||
- name: TEST_DATA_DIR
|
||||
value: /data/test-data
|
||||
- name: RUST_LOG
|
||||
value: info
|
||||
- name: LD_LIBRARY_PATH
|
||||
value: /usr/local/nvidia/lib64:/usr/local/cuda/lib64:/usr/local/cuda/extras/CUPTI/lib64
|
||||
- name: CARGO_PROFILE_TEST_OPT_LEVEL
|
||||
value: "2"
|
||||
volumeMounts:
|
||||
- name: git-ssh-key
|
||||
mountPath: /etc/git-ssh
|
||||
readOnly: true
|
||||
- name: cargo-target
|
||||
mountPath: /cargo-target
|
||||
- name: test-data
|
||||
mountPath: /data/test-data
|
||||
readOnly: true
|
||||
resources:
|
||||
requests:
|
||||
nvidia.com/gpu: "1"
|
||||
cpu: "4"
|
||||
memory: 32Gi
|
||||
limits:
|
||||
nvidia.com/gpu: "1"
|
||||
cpu: "8"
|
||||
memory: 64Gi
|
||||
args:
|
||||
- |
|
||||
set -e
|
||||
|
||||
REF="{{workflow.parameters.commit-ref}}"
|
||||
MODELS="{{workflow.parameters.models}}"
|
||||
SCOPE="{{workflow.parameters.test-scope}}"
|
||||
|
||||
# --- SSH setup (same as compile-and-train) ---
|
||||
mkdir -p ~/.ssh
|
||||
cp /etc/git-ssh/ssh-privatekey ~/.ssh/id_ed25519
|
||||
chmod 600 ~/.ssh/id_ed25519
|
||||
printf 'Host *\n StrictHostKeyChecking no\n UserKnownHostsFile /dev/null\n' > ~/.ssh/config
|
||||
chmod 600 ~/.ssh/config
|
||||
|
||||
REPO="ssh://git@gitlab-gitlab-shell.foxhunt.svc.cluster.local:2222/root/foxhunt.git"
|
||||
BUILD="/cargo-target/src"
|
||||
git config --global --add safe.directory "$BUILD"
|
||||
|
||||
# --- Persistent checkout on PVC ---
|
||||
if [ -d "$BUILD/.git" ]; then
|
||||
cd "$BUILD"
|
||||
echo "=== Fetching latest refs ==="
|
||||
git fetch origin
|
||||
# Resolve REF after fetch so we always get the latest commit.
|
||||
# Try origin/$REF (branch), then $REF directly (tag or SHA).
|
||||
if git rev-parse --verify "origin/$REF" >/dev/null 2>&1; then
|
||||
TARGET=$(git rev-parse "origin/$REF")
|
||||
else
|
||||
TARGET=$(git rev-parse --verify "$REF" 2>/dev/null || echo "$REF")
|
||||
fi
|
||||
CURRENT=$(git rev-parse HEAD 2>/dev/null || echo "none")
|
||||
if [ "$CURRENT" = "$TARGET" ]; then
|
||||
echo "=== Already at $REF ($TARGET) ==="
|
||||
else
|
||||
echo "=== Updating checkout to $REF ($TARGET) ==="
|
||||
git checkout --force --detach "$TARGET"
|
||||
git clean -fd
|
||||
fi
|
||||
else
|
||||
echo "=== Initial clone ==="
|
||||
git clone --filter=blob:none "$REPO" "$BUILD"
|
||||
cd "$BUILD"
|
||||
git fetch origin
|
||||
if git rev-parse --verify "origin/$REF" >/dev/null 2>&1; then
|
||||
git checkout --force --detach "origin/$REF"
|
||||
else
|
||||
git checkout "$REF"
|
||||
fi
|
||||
fi
|
||||
echo "Checked out $(git rev-parse --short HEAD)"
|
||||
|
||||
export PATH="${CARGO_HOME}/bin:${PATH}"
|
||||
export CUDA_COMPUTE_CAP={{workflow.parameters.cuda-compute-cap}}
|
||||
|
||||
# --- PTX cache invalidation ---
|
||||
# Purge stale cached PTX if any CUDA kernel source changed since last run.
|
||||
bash scripts/ptx-cache-invalidate.sh "${CARGO_TARGET_DIR}/.ptx_cache"
|
||||
|
||||
# --- PVC size guard ---
|
||||
TARGET_SIZE_MB=$(du -sm "$CARGO_TARGET_DIR" 2>/dev/null | cut -f1 || echo 0)
|
||||
echo "PVC usage: ${TARGET_SIZE_MB}MB"
|
||||
if [ "$TARGET_SIZE_MB" -gt 25000 ]; then
|
||||
echo "PVC exceeds 25GB, pruning..."
|
||||
rm -rf "$CARGO_TARGET_DIR/release" "$CARGO_TARGET_DIR/debug"
|
||||
fi
|
||||
|
||||
# --- Compile ---
|
||||
echo "=== Compiling test binaries (--features cuda) ==="
|
||||
cargo test -p ml -p ml-dqn -p ml-core --features cuda --no-run 2>&1 | tee /cargo-target/gpu-compile.log
|
||||
echo "=== Compilation done ==="
|
||||
|
||||
# --- Expand "all" ---
|
||||
if [ "$MODELS" = "all" ]; then
|
||||
MODELS="dqn,ppo,tft,mamba2,tggn,tlob,liquid,kan,xlstm,diffusion"
|
||||
fi
|
||||
|
||||
# --- Test runner ---
|
||||
RESULTS=""
|
||||
FAILURES=0
|
||||
|
||||
# Reset CUDA context between test binaries to prevent cuBLAS
|
||||
# CUBLAS_STATUS_NOT_INITIALIZED cascades. Each test binary creates
|
||||
# and destroys hundreds of cuBLAS handles; without a reset, the
|
||||
# driver fails to re-init for the next binary.
|
||||
gpu_context_drain() {
|
||||
nvidia-smi -rgc >/dev/null 2>&1 || true
|
||||
sleep 1
|
||||
}
|
||||
|
||||
run_tests() {
|
||||
local NAME="$1"; shift
|
||||
echo ""
|
||||
echo "========================================"
|
||||
echo " TEST: $NAME"
|
||||
echo "========================================"
|
||||
# --nocapture is a test-binary flag, must come after --
|
||||
# If args already contain --, append after it; otherwise add -- first
|
||||
local HAS_SEP=false
|
||||
for arg in "$@"; do
|
||||
[ "$arg" = "--" ] && HAS_SEP=true && break
|
||||
done
|
||||
set +e
|
||||
if $HAS_SEP; then
|
||||
"$@" --nocapture 2>&1
|
||||
else
|
||||
"$@" -- --nocapture 2>&1
|
||||
fi
|
||||
EXIT=$?
|
||||
set -e
|
||||
if [ $EXIT -eq 0 ]; then
|
||||
RESULTS="${RESULTS}${NAME}:PASS\n"
|
||||
else
|
||||
RESULTS="${RESULTS}${NAME}:FAIL\n"
|
||||
FAILURES=$((FAILURES + 1))
|
||||
fi
|
||||
# Drain CUDA context after each GPU test suite
|
||||
gpu_context_drain
|
||||
}
|
||||
|
||||
# --- Core tests (always run) ---
|
||||
# --test-threads=1 for all GPU lib tests: prevents concurrent cuBLAS
|
||||
# handle creation that causes CUBLAS_STATUS_NOT_INITIALIZED cascades.
|
||||
run_tests "core-lib" cargo test -p ml-core --features cuda --lib -- --test-threads=1
|
||||
run_tests "bayesian" cargo test -p ml --features cuda --test bayesian_changepoint_test -- --test-threads=1
|
||||
|
||||
# --- Per-model tests ---
|
||||
IFS=',' read -ra MODEL_LIST <<< "$MODELS"
|
||||
for MODEL in "${MODEL_LIST[@]}"; do
|
||||
case "$MODEL" in
|
||||
dqn)
|
||||
if [ "$SCOPE" = "lib" ] || [ "$SCOPE" = "all" ]; then
|
||||
# --test-threads=1: GPU lib tests must run serially — each test
|
||||
# creates a cuBLAS handle via Device::new_cuda(0). Under parallel
|
||||
# execution, concurrent cuBLAS init races cause
|
||||
# CUBLAS_STATUS_NOT_INITIALIZED failures (51 test cascade).
|
||||
run_tests "dqn-lib" cargo test -p ml-dqn --features cuda --lib -- --test-threads=1
|
||||
run_tests "dqn-ml-lib" cargo test -p ml --features cuda --lib -- --test-threads=1 dqn
|
||||
fi
|
||||
if [ "$SCOPE" = "integration" ] || [ "$SCOPE" = "all" ]; then
|
||||
# --test-threads=1 for ALL GPU integration tests: parallel execution
|
||||
# corrupts the CUDA primary context (cuDevicePrimaryCtxRetain fails
|
||||
# when multiple threads race on context init/teardown).
|
||||
run_tests "dqn-smoke" cargo test -p ml --features cuda --test smoke_test_real_data -- --test-threads=1
|
||||
# Run each pipeline test in its own cargo test process.
|
||||
# CUDA Graph capture corrupts the async memory pool, making
|
||||
# cuMemAllocAsync fail with CUDA_ERROR_INVALID_VALUE in
|
||||
# subsequent DQNTrainer instances within the same process.
|
||||
run_tests "dqn-pipeline-train" cargo test -p ml --features cuda --test dqn_training_pipeline_test test_dqn_trains_on_es_fut -- --test-threads=1 --exact
|
||||
run_tests "dqn-pipeline-loss" cargo test -p ml --features cuda --test dqn_training_pipeline_test test_dqn_loss_decreases -- --test-threads=1 --exact
|
||||
run_tests "dqn-pipeline-ckpt" cargo test -p ml --features cuda --test dqn_training_pipeline_test test_dqn_checkpoint_save_load -- --test-threads=1 --exact
|
||||
run_tests "dqn-pipeline-qval" cargo test -p ml --features cuda --test dqn_training_pipeline_test test_dqn_q_value_predictions -- --test-threads=1 --exact
|
||||
run_tests "dqn-pipeline-eps" cargo test -p ml --features cuda --test dqn_training_pipeline_test test_dqn_epsilon_greedy -- --test-threads=1 --exact
|
||||
run_tests "dqn-smoke-train" cargo test -p ml --features cuda --test dqn_training_smoke_test -- --test-threads=1
|
||||
run_tests "dqn-early-stop" cargo test -p ml --features cuda --test dqn_early_stopping_termination_test -- --test-threads=1
|
||||
run_tests "dqn-collapse" cargo test -p ml --features cuda --test dqn_action_collapse_fix_test -- --test-threads=1
|
||||
fi
|
||||
;;
|
||||
ppo)
|
||||
if [ "$SCOPE" = "lib" ] || [ "$SCOPE" = "all" ]; then
|
||||
run_tests "ppo-lib" cargo test -p ml --features cuda --lib -- --test-threads=1 ppo
|
||||
fi
|
||||
if [ "$SCOPE" = "integration" ] || [ "$SCOPE" = "all" ]; then
|
||||
run_tests "ppo-barrier" cargo test -p ml --features cuda --test barrier_optimization_test -- --test-threads=1
|
||||
fi
|
||||
;;
|
||||
*)
|
||||
# Supervised models (TFT, Mamba2, TGGN, TLOB, Liquid, KAN, xLSTM, Diffusion)
|
||||
# Map model names to Rust module names where they differ
|
||||
LIB_FILTER="$MODEL"
|
||||
[ "$MODEL" = "tggn" ] && LIB_FILTER="tgnn"
|
||||
if [ "$SCOPE" = "lib" ] || [ "$SCOPE" = "all" ]; then
|
||||
run_tests "${MODEL}-lib" cargo test -p ml --features cuda --lib -- --test-threads=1 "$LIB_FILTER"
|
||||
fi
|
||||
if [ "$SCOPE" = "integration" ] || [ "$SCOPE" = "all" ]; then
|
||||
run_tests "${MODEL}-gpu" cargo test -p ml --features cuda --test supervised_gpu_smoke_test -- --test-threads=1 "test_${MODEL}_gpu_smoke"
|
||||
# Also run model-specific integration tests if they exist
|
||||
if cargo test -p ml --features cuda --test "${MODEL}_integration" --no-run 2>/dev/null; then
|
||||
run_tests "${MODEL}-integ" cargo test -p ml --features cuda --test "${MODEL}_integration" -- --test-threads=1
|
||||
fi
|
||||
fi
|
||||
;;
|
||||
esac
|
||||
done
|
||||
|
||||
# --- Summary ---
|
||||
echo ""
|
||||
echo "========================================"
|
||||
echo " GPU TEST RESULTS"
|
||||
echo "========================================"
|
||||
printf "$RESULTS" | while IFS=: read -r name status; do
|
||||
[ -z "$name" ] && continue
|
||||
printf " %-25s %s\n" "$name" "$status"
|
||||
done
|
||||
echo "========================================"
|
||||
echo " Total failures: $FAILURES"
|
||||
echo "========================================"
|
||||
|
||||
# Write results for notification step
|
||||
mkdir -p /tmp/outputs
|
||||
printf "$RESULTS" > /tmp/outputs/results
|
||||
echo "$FAILURES" > /tmp/outputs/failures
|
||||
|
||||
[ "$FAILURES" -gt 0 ] && exit 1 || exit 0
|
||||
|
||||
# ── perf-benchmark: DQN epoch/s on 3Q data (performance regression guard) ──
|
||||
# Runs after tests pass. Trains DQN on 3Q of ES.FUT data and reports epoch time.
|
||||
# Uses the same binary compiled by compile-and-test (shared cargo-target PVC).
|
||||
# Fails the pipeline if epoch time exceeds 500ms (regression threshold for H100).
|
||||
- name: perf-benchmark
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: "{{workflow.parameters.gpu-pool}}"
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
tolerations:
|
||||
- key: nvidia.com/gpu
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
- key: node.cilium.io/agent-not-ready
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
outputs:
|
||||
parameters:
|
||||
- name: epoch-ms
|
||||
valueFrom:
|
||||
path: /tmp/outputs/epoch-ms
|
||||
default: "9999"
|
||||
container:
|
||||
image: gitlab-registry.foxhunt.svc.cluster.local:5000/root/foxhunt/ci-builder:latest
|
||||
command: ["/bin/bash", "-c"]
|
||||
env:
|
||||
- name: SQLX_OFFLINE
|
||||
value: "true"
|
||||
- name: CARGO_TARGET_DIR
|
||||
value: /cargo-target
|
||||
- name: CARGO_HOME
|
||||
value: /cargo-target/cargo-home
|
||||
- name: CUDA_VISIBLE_DEVICES
|
||||
value: "0"
|
||||
- name: LD_LIBRARY_PATH
|
||||
value: /usr/local/nvidia/lib64:/usr/local/cuda/lib64:/usr/local/cuda/extras/CUPTI/lib64
|
||||
volumeMounts:
|
||||
- name: cargo-target
|
||||
mountPath: /cargo-target
|
||||
- name: test-data
|
||||
mountPath: /data/test-data
|
||||
readOnly: true
|
||||
resources:
|
||||
requests:
|
||||
nvidia.com/gpu: "1"
|
||||
cpu: "4"
|
||||
memory: 16Gi
|
||||
limits:
|
||||
nvidia.com/gpu: "1"
|
||||
cpu: "8"
|
||||
memory: 32Gi
|
||||
args:
|
||||
- |
|
||||
set -e
|
||||
cd /cargo-target/src
|
||||
export PATH="${CARGO_HOME}/bin:${PATH}"
|
||||
export CUDA_COMPUTE_CAP={{workflow.parameters.cuda-compute-cap}}
|
||||
|
||||
echo "========================================"
|
||||
echo " PERF BENCHMARK: DQN epoch/s (3Q ES.FUT)"
|
||||
echo "========================================"
|
||||
|
||||
# Run 5 epochs × 100 steps on 3Q data (--train-months 3 fits in 1 quarter)
|
||||
# Use --step-months 3 to avoid walk-forward window splits
|
||||
OUTPUT=$(cargo run --release --example train_baseline_rl -p ml -- \
|
||||
--model dqn \
|
||||
--data-dir /data/test-data/ohlcv \
|
||||
--mbp10-data-dir /data/test-data/mbp10 \
|
||||
--trades-data-dir /data/test-data/trades \
|
||||
--symbol ES.FUT \
|
||||
--epochs 5 \
|
||||
--train-months 3 --val-months 1 --test-months 1 --step-months 3 \
|
||||
2>&1)
|
||||
|
||||
# Extract epoch times (skip epoch 1 which includes init)
|
||||
EPOCH_TIMES=$(echo "$OUTPUT" | grep "phase breakdown" | grep -v "Epoch 1/" | \
|
||||
sed 's/.*total=\([0-9]*\)ms.*/\1/' | head -4)
|
||||
|
||||
if [ -z "$EPOCH_TIMES" ]; then
|
||||
echo "ERROR: No phase breakdown output found"
|
||||
echo "$OUTPUT" | tail -20
|
||||
mkdir -p /tmp/outputs
|
||||
echo "9999" > /tmp/outputs/epoch-ms
|
||||
exit 1
|
||||
fi
|
||||
|
||||
# Compute average epoch time (epochs 2-5)
|
||||
SUM=0
|
||||
COUNT=0
|
||||
for T in $EPOCH_TIMES; do
|
||||
SUM=$((SUM + T))
|
||||
COUNT=$((COUNT + 1))
|
||||
done
|
||||
AVG=$((SUM / COUNT))
|
||||
|
||||
echo ""
|
||||
echo " Epoch times (ms, excl. epoch 1): $EPOCH_TIMES"
|
||||
echo " Average: ${AVG}ms"
|
||||
echo ""
|
||||
|
||||
mkdir -p /tmp/outputs
|
||||
echo "$AVG" > /tmp/outputs/epoch-ms
|
||||
|
||||
# Regression guard: fail if avg epoch > 500ms on H100
|
||||
THRESHOLD=500
|
||||
if [ "$AVG" -gt "$THRESHOLD" ]; then
|
||||
echo "PERF REGRESSION: ${AVG}ms > ${THRESHOLD}ms threshold"
|
||||
echo "========================================"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
echo " PASS: ${AVG}ms <= ${THRESHOLD}ms threshold"
|
||||
echo "========================================"
|
||||
|
||||
# ── notify-result: post test outcome to Mattermost (onExit) ──
|
||||
- name: notify-result
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: platform
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
container:
|
||||
image: curlimages/curl:8.12.1
|
||||
command: ["/bin/sh", "-c"]
|
||||
env:
|
||||
- name: WEBHOOK_URL
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: notification-webhook
|
||||
key: webhook-url
|
||||
optional: true
|
||||
resources:
|
||||
requests:
|
||||
cpu: 50m
|
||||
memory: 32Mi
|
||||
limits:
|
||||
cpu: 100m
|
||||
memory: 64Mi
|
||||
args:
|
||||
- |
|
||||
NOTIFY="{{workflow.parameters.notify}}"
|
||||
if [ "$NOTIFY" != "true" ]; then
|
||||
echo "Notifications disabled, skipping"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
STATUS="{{workflow.status}}"
|
||||
NAME="{{workflow.name}}"
|
||||
|
||||
if [ -z "$WEBHOOK_URL" ] || echo "$WEBHOOK_URL" | grep -q "PLACEHOLDER"; then
|
||||
echo "No webhook configured, skipping notification"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
if [ "$STATUS" = "Succeeded" ]; then
|
||||
EMOJI=":white_check_mark:"
|
||||
else
|
||||
EMOJI=":x:"
|
||||
fi
|
||||
|
||||
PAYLOAD="{\"username\":\"Argo CI\",\"text\":\"${EMOJI} **GPU Tests** ${NAME} — ${STATUS} ({{workflow.duration}}s)\"}"
|
||||
|
||||
curl -sf -X POST -H 'Content-Type: application/json' \
|
||||
-d "$PAYLOAD" "$WEBHOOK_URL" || echo "WARN: webhook post failed"
|
||||
@@ -1,15 +1,8 @@
|
||||
apiVersion: kustomize.config.k8s.io/v1beta1
|
||||
kind: Kustomization
|
||||
# Rust train/CI/cache templates removed 2026-06-21 (decommission-rust-infra).
|
||||
# Remaining: platform RBAC + network policy used by the active fxhnt-cockpit deploys.
|
||||
resources:
|
||||
- sccache-pvcs.yaml
|
||||
- ci-pipeline-template.yaml
|
||||
- compile-and-deploy-template.yaml
|
||||
- train-template.yaml
|
||||
- build-ci-image-template.yaml
|
||||
- sanitizer-test-template.yaml
|
||||
- nsys-test-template.yaml
|
||||
- smoke-test-template.yaml
|
||||
- refresh-deps-cache-template.yaml
|
||||
- argo-workflow-netpol.yaml
|
||||
- ci-deploy-rbac.yaml
|
||||
- archive-rbac.yaml
|
||||
|
||||
@@ -1,459 +0,0 @@
|
||||
# Real-LOB backtest sweep workflow.
|
||||
#
|
||||
# Fans out a parameter grid across N parallel GPU pods, each running
|
||||
# `fxt-backtest run` against the same MBP-10 data + checkpoint with one
|
||||
# cell's worth of overrides. A final CPU pod runs `fxt-backtest aggregate`
|
||||
# to produce aggregate.parquet + pareto_frontier.json at the sweep root.
|
||||
#
|
||||
# The `# __SWEEP_CELLS__` marker on the dag.tasks line is replaced by
|
||||
# scripts/argo-lob-sweep.sh with N generated WorkflowTask stanzas before
|
||||
# submission. Same convention as train-multi-seed-template.yaml's
|
||||
# # __MATRIX_TASKS__ marker.
|
||||
#
|
||||
# DAG:
|
||||
# ensure-binary ──> [N parallel run-cell-<i> tasks on ci-training-l40s,
|
||||
# each with its own outputDir on the shared PVC]
|
||||
# │
|
||||
# └──> aggregate (CPU node, runs fxt-backtest aggregate
|
||||
# against the sweep root)
|
||||
#
|
||||
# Per `feedback_default_to_l40s_pool.md` (2026-05-09): defaults to
|
||||
# ci-training-l40s + sm_89. Override via --gpu-pool ci-training-h100
|
||||
# for sm_90 / 80 GB.
|
||||
#
|
||||
# Usage:
|
||||
# ./scripts/argo-lob-sweep.sh --grid config/ml/sweep_decision_stride_example.yaml
|
||||
# ./scripts/argo-lob-sweep.sh --grid <path> --sha abc1234 --gpu-pool ci-training-h100
|
||||
# ./scripts/argo-lob-sweep.sh --grid <path> --watch
|
||||
# ./scripts/argo-lob-sweep.sh --grid <path> --dry-run > /tmp/wf.yaml
|
||||
|
||||
apiVersion: argoproj.io/v1alpha1
|
||||
kind: WorkflowTemplate
|
||||
metadata:
|
||||
name: lob-backtest-sweep
|
||||
namespace: foxhunt
|
||||
labels:
|
||||
app.kubernetes.io/name: lob-backtest-sweep
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
app.kubernetes.io/component: train # reuses argo-train-workflow NetworkPolicy egress (port 2222 to gitlab-shell)
|
||||
spec:
|
||||
entrypoint: sweep-matrix
|
||||
onExit: notify-result
|
||||
serviceAccountName: argo-workflow
|
||||
archiveLogs: true
|
||||
podMetadata:
|
||||
labels:
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
app.kubernetes.io/component: train
|
||||
securityContext:
|
||||
fsGroup: 0
|
||||
ttlStrategy:
|
||||
secondsAfterCompletion: 3600
|
||||
# LOB sweep cells are much faster than training (~minutes vs hours);
|
||||
# 2h walltime is generous even for N=128 cells.
|
||||
activeDeadlineSeconds: 7200
|
||||
|
||||
arguments:
|
||||
parameters:
|
||||
- name: commit-sha
|
||||
value: HEAD
|
||||
- name: git-branch
|
||||
value: main
|
||||
- name: cuda-compute-cap
|
||||
value: "89"
|
||||
- name: gpu-pool
|
||||
value: ci-training-l40s
|
||||
- name: data-root
|
||||
value: /mnt/training-data/futures-baseline/ES.FUT
|
||||
- name: predecoded-dir
|
||||
value: /mnt/training-data/futures-baseline/ES.FUT
|
||||
- name: checkpoint
|
||||
value: "" # empty = trunk runs with --seed (random init);
|
||||
# set to a path on the data PVC for trained weights
|
||||
- name: sweep-root
|
||||
value: /mnt/training-data/sweeps/lob-backtest
|
||||
- name: sweep-tag
|
||||
value: default # subdirectory: <sweep-root>/<sweep-tag>/
|
||||
# Operational knobs. argo-lob-sweep.sh sets these per-grid.
|
||||
- name: n-parallel
|
||||
value: "1"
|
||||
- name: latency-ns
|
||||
value: "100000000"
|
||||
- name: target-annual-vol-units
|
||||
value: "50.0"
|
||||
- name: annualisation-factor
|
||||
value: "825.0"
|
||||
- name: max-lots
|
||||
value: "5"
|
||||
- name: max-events
|
||||
value: "0"
|
||||
# Per-step kernel-state JSONL trace path. Empty (default) = disabled:
|
||||
# ensure-binary compiles fxt-backtest WITHOUT --features kernel-step-trace
|
||||
# and run-cell omits the --kernel-step-trace flag. When set to a
|
||||
# non-empty path, ensure-binary rebuilds with the feature enabled
|
||||
# and run-cell passes --kernel-step-trace <path> through to the
|
||||
# binary. Trace is written to that absolute path on the pod (must
|
||||
# resolve to a mounted PVC — usually /feature-cache or /mnt/training-data).
|
||||
- name: kernel-step-trace
|
||||
value: ""
|
||||
|
||||
volumes:
|
||||
- name: training-data
|
||||
persistentVolumeClaim:
|
||||
claimName: training-data-pvc
|
||||
- name: git-ssh-key
|
||||
secret:
|
||||
secretName: argo-git-ssh-key
|
||||
defaultMode: 0400
|
||||
- name: cargo-target-cuda
|
||||
persistentVolumeClaim:
|
||||
claimName: cargo-target-cuda
|
||||
- name: feature-cache
|
||||
persistentVolumeClaim:
|
||||
claimName: feature-cache-pvc
|
||||
|
||||
templates:
|
||||
# ── DAG entrypoint ──────────────────────────────────────────────
|
||||
- name: sweep-matrix
|
||||
dag:
|
||||
tasks:
|
||||
- name: ensure-binary
|
||||
template: ensure-binary
|
||||
# __SWEEP_CELLS__ — replaced by argo-lob-sweep.sh
|
||||
- name: aggregate
|
||||
template: aggregate
|
||||
dependencies: [ensure-binary] # plus all run-cell-* via shell-injected list
|
||||
arguments:
|
||||
parameters:
|
||||
- name: sha
|
||||
value: "{{tasks.ensure-binary.outputs.parameters.sha}}"
|
||||
|
||||
# ── ensure-binary: cache-or-compile fxt-backtest by commit SHA ─
|
||||
# Same shape as train-multi-seed-template.yaml's ensure-binary but
|
||||
# for the fxt-backtest binary only.
|
||||
- name: ensure-binary
|
||||
outputs:
|
||||
parameters:
|
||||
- name: sha
|
||||
valueFrom:
|
||||
path: /tmp/sha
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: ci-compile-cpu
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
tolerations:
|
||||
- key: node.cilium.io/agent-not-ready
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
container:
|
||||
image: gitlab-registry.foxhunt.svc.cluster.local:5000/root/foxhunt/ci-builder:latest
|
||||
imagePullPolicy: Always
|
||||
command: ["/bin/bash", "-c"]
|
||||
env:
|
||||
- name: SQLX_OFFLINE
|
||||
value: "true"
|
||||
- name: CARGO_TARGET_DIR
|
||||
value: /cargo-target
|
||||
- name: CARGO_HOME
|
||||
value: /cargo-target/cargo-home
|
||||
- name: RUSTC_WRAPPER
|
||||
value: sccache
|
||||
- name: SCCACHE_DIR
|
||||
value: /cargo-target/sccache
|
||||
- name: SCCACHE_CACHE_SIZE
|
||||
value: "40G"
|
||||
- name: CARGO_INCREMENTAL
|
||||
value: "0"
|
||||
- name: FOXHUNT_CUDA_ARCH
|
||||
value: "sm_{{workflow.parameters.cuda-compute-cap}}"
|
||||
- name: GITLAB_PAT
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: gitlab-pat
|
||||
key: token
|
||||
resources:
|
||||
requests:
|
||||
cpu: "14"
|
||||
memory: 32Gi
|
||||
limits:
|
||||
cpu: "30"
|
||||
memory: 64Gi
|
||||
volumeMounts:
|
||||
- name: git-ssh-key
|
||||
mountPath: /etc/git-ssh
|
||||
readOnly: true
|
||||
- name: cargo-target-cuda
|
||||
mountPath: /cargo-target
|
||||
- name: training-data
|
||||
mountPath: /mnt/training-data
|
||||
args:
|
||||
- |
|
||||
set -e
|
||||
SHA="{{workflow.parameters.commit-sha}}"
|
||||
BRANCH="{{workflow.parameters.git-branch}}"
|
||||
KERNEL_TRACE="{{workflow.parameters.kernel-step-trace}}"
|
||||
|
||||
mkdir -p ~/.ssh
|
||||
cp /etc/git-ssh/ssh-privatekey ~/.ssh/id_ed25519
|
||||
chmod 600 ~/.ssh/id_ed25519
|
||||
printf 'Host *\n StrictHostKeyChecking no\n UserKnownHostsFile /dev/null\n' > ~/.ssh/config
|
||||
chmod 600 ~/.ssh/config
|
||||
|
||||
REPO="ssh://git@gitlab-gitlab-shell.foxhunt.svc.cluster.local:2222/root/foxhunt.git"
|
||||
BUILD="/cargo-target/src"
|
||||
git config --global --add safe.directory "$BUILD"
|
||||
|
||||
if [ "$SHA" = "HEAD" ]; then
|
||||
if [ -d "$BUILD/.git" ]; then
|
||||
cd "$BUILD"; git fetch origin
|
||||
SHA=$(git rev-parse "origin/$BRANCH"); cd /
|
||||
else
|
||||
git clone --filter=blob:none "$REPO" "$BUILD"
|
||||
cd "$BUILD"; git checkout "origin/$BRANCH"
|
||||
SHA=$(git rev-parse HEAD); cd /
|
||||
fi
|
||||
fi
|
||||
SHORT_SHA=$(echo "$SHA" | cut -c1-9)
|
||||
echo "Resolved SHA: $SHA (short: $SHORT_SHA)"
|
||||
|
||||
# Feature variant: when kernel-step-trace is enabled, build
|
||||
# with the Cargo feature and cache under a distinct subdir
|
||||
# so the default + diagnostic binaries don't clobber each other.
|
||||
FEATURE_FLAGS=""
|
||||
VARIANT="default"
|
||||
if [ -n "$KERNEL_TRACE" ]; then
|
||||
FEATURE_FLAGS="--features kernel-step-trace"
|
||||
VARIANT="kstrace"
|
||||
fi
|
||||
BIN_DIR="/mnt/training-data/bin/$SHORT_SHA-$VARIANT"
|
||||
mkdir -p "$BIN_DIR"
|
||||
if [ -x "$BIN_DIR/fxt-backtest" ]; then
|
||||
echo "=== Cache HIT: fxt-backtest ($VARIANT) present in $BIN_DIR ==="
|
||||
ls -lh "$BIN_DIR/fxt-backtest"
|
||||
echo "$SHORT_SHA-$VARIANT" > /tmp/sha
|
||||
exit 0
|
||||
fi
|
||||
|
||||
echo "=== Cache MISS: compiling fxt-backtest ($VARIANT) for $SHORT_SHA ==="
|
||||
if [ -d "$BUILD/.git" ]; then
|
||||
# Same pattern as alpha-perception-template.yaml: prior
|
||||
# cargo build mutates Cargo.lock (and sometimes other
|
||||
# generated files); --force checkout overwrites them,
|
||||
# `git clean -fd` drops untracked. Without this, the
|
||||
# second SHA's checkout fails on dirty working tree.
|
||||
cd "$BUILD"; git fetch origin
|
||||
git checkout --force "$SHA"; git clean -fd
|
||||
else
|
||||
git clone --filter=blob:none "$REPO" "$BUILD"
|
||||
cd "$BUILD"; git checkout "$SHA"
|
||||
fi
|
||||
|
||||
cargo build -p fxt-backtest --release $FEATURE_FLAGS
|
||||
cp "$CARGO_TARGET_DIR/release/fxt-backtest" "$BIN_DIR/fxt-backtest"
|
||||
echo "$SHORT_SHA-$VARIANT" > /tmp/sha
|
||||
ls -lh "$BIN_DIR/fxt-backtest"
|
||||
|
||||
# ── run-cell: one sweep cell against a GPU pod ─────────────────
|
||||
# Inputs: cell name + every Run arg that varies across the grid.
|
||||
# The shell-rendered DAG tasks supply these per-cell.
|
||||
- name: run-cell
|
||||
inputs:
|
||||
parameters:
|
||||
- name: sha
|
||||
- name: cell-name
|
||||
- name: decision-stride
|
||||
value: "4"
|
||||
- name: latency-ns
|
||||
value: "{{workflow.parameters.latency-ns}}"
|
||||
- name: target-annual-vol-units
|
||||
value: "{{workflow.parameters.target-annual-vol-units}}"
|
||||
- name: annualisation-factor
|
||||
value: "{{workflow.parameters.annualisation-factor}}"
|
||||
- name: max-lots
|
||||
value: "{{workflow.parameters.max-lots}}"
|
||||
- name: max-events
|
||||
value: "{{workflow.parameters.max-events}}"
|
||||
- name: seed
|
||||
value: "0xC0FFEE"
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: "{{workflow.parameters.gpu-pool}}"
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
tolerations:
|
||||
- key: nvidia.com/gpu
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
container:
|
||||
image: gitlab-registry.foxhunt.svc.cluster.local:5000/root/foxhunt/foxhunt-training-runtime:latest
|
||||
imagePullPolicy: Always
|
||||
command: ["/bin/bash", "-c"]
|
||||
resources:
|
||||
requests:
|
||||
cpu: "4"
|
||||
memory: 16Gi
|
||||
nvidia.com/gpu: "1"
|
||||
limits:
|
||||
cpu: "8"
|
||||
memory: 32Gi
|
||||
nvidia.com/gpu: "1"
|
||||
volumeMounts:
|
||||
- name: training-data
|
||||
mountPath: /mnt/training-data
|
||||
- name: feature-cache
|
||||
mountPath: /feature-cache
|
||||
args:
|
||||
- |
|
||||
set -e
|
||||
SHA="{{inputs.parameters.sha}}"
|
||||
CELL="{{inputs.parameters.cell-name}}"
|
||||
BIN="/mnt/training-data/bin/$SHA/fxt-backtest"
|
||||
CKPT="{{workflow.parameters.checkpoint}}"
|
||||
KERNEL_TRACE="{{workflow.parameters.kernel-step-trace}}"
|
||||
OUT="{{workflow.parameters.sweep-root}}/{{workflow.parameters.sweep-tag}}/$CELL"
|
||||
mkdir -p "$OUT"
|
||||
echo "=== sweep cell $CELL on $(hostname) → $OUT ==="
|
||||
|
||||
CKPT_FLAG=""
|
||||
if [ -n "$CKPT" ]; then
|
||||
CKPT_FLAG="--checkpoint $CKPT"
|
||||
fi
|
||||
KERNEL_TRACE_FLAG=""
|
||||
if [ -n "$KERNEL_TRACE" ]; then
|
||||
# Per-cell trace path: append cell name so concurrent run-cell
|
||||
# pods don't clobber each other's JSONL files. Operators get
|
||||
# one trace per cell; aggregate by reading <root>/<sweep>/<cell>/
|
||||
# at analysis time.
|
||||
KERNEL_TRACE_FLAG="--kernel-step-trace $OUT/kernel_step_trace.jsonl"
|
||||
fi
|
||||
|
||||
"$BIN" run \
|
||||
--data "{{workflow.parameters.data-root}}" \
|
||||
--predecoded-dir "{{workflow.parameters.predecoded-dir}}" \
|
||||
--n-parallel "{{workflow.parameters.n-parallel}}" \
|
||||
--decision-stride "{{inputs.parameters.decision-stride}}" \
|
||||
--latency-ns "{{inputs.parameters.latency-ns}}" \
|
||||
--target-annual-vol-units "{{inputs.parameters.target-annual-vol-units}}" \
|
||||
--annualisation-factor "{{inputs.parameters.annualisation-factor}}" \
|
||||
--max-lots "{{inputs.parameters.max-lots}}" \
|
||||
--max-events "{{inputs.parameters.max-events}}" \
|
||||
--seed "{{inputs.parameters.seed}}" \
|
||||
$CKPT_FLAG \
|
||||
$KERNEL_TRACE_FLAG \
|
||||
--out "$OUT"
|
||||
|
||||
echo "=== cell $CELL done ==="
|
||||
ls -lh "$OUT"
|
||||
|
||||
# ── run-sweep: P6 batched flow. Invokes `fxt-backtest sweep` ───
|
||||
# against the FULL grid YAML inside a single GPU pod. When the YAML
|
||||
# carries sim_variants, each cell expands into n_parallel=variants.len()
|
||||
# backtests sharing one forward pass. argo-lob-sweep.sh emits ONE
|
||||
# run-sweep task (instead of N fan-out run-cell tasks) when --batched
|
||||
# is set or sim_variants is detected in the YAML.
|
||||
- name: run-sweep
|
||||
inputs:
|
||||
parameters:
|
||||
- name: sha
|
||||
- name: grid-yaml-b64 # base64-encoded grid YAML, written to a file in-pod
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: "{{workflow.parameters.gpu-pool}}"
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
tolerations:
|
||||
- key: nvidia.com/gpu
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
container:
|
||||
image: gitlab-registry.foxhunt.svc.cluster.local:5000/root/foxhunt/foxhunt-training-runtime:latest
|
||||
imagePullPolicy: Always
|
||||
command: ["/bin/bash", "-c"]
|
||||
resources:
|
||||
requests:
|
||||
cpu: "4"
|
||||
memory: 16Gi
|
||||
nvidia.com/gpu: "1"
|
||||
limits:
|
||||
cpu: "8"
|
||||
memory: 32Gi
|
||||
nvidia.com/gpu: "1"
|
||||
volumeMounts:
|
||||
- name: training-data
|
||||
mountPath: /mnt/training-data
|
||||
- name: feature-cache
|
||||
mountPath: /feature-cache
|
||||
args:
|
||||
- |
|
||||
set -e
|
||||
SHA="{{inputs.parameters.sha}}"
|
||||
BIN="/mnt/training-data/bin/$SHA/fxt-backtest"
|
||||
OUT="{{workflow.parameters.sweep-root}}/{{workflow.parameters.sweep-tag}}"
|
||||
mkdir -p "$OUT"
|
||||
# Write the grid YAML to a known path. base64 in/out keeps Argo
|
||||
# parameter encoding stable across YAML special characters
|
||||
# (curly braces from data_template + sim_variants).
|
||||
GRID_YAML=/tmp/sweep-grid.yaml
|
||||
echo "{{inputs.parameters.grid-yaml-b64}}" | base64 -d > "$GRID_YAML"
|
||||
echo "=== running fxt-backtest sweep on $(hostname) → $OUT ==="
|
||||
echo "=== grid yaml: ==="
|
||||
head -20 "$GRID_YAML"
|
||||
echo "=== ... ==="
|
||||
|
||||
"$BIN" sweep --grid "$GRID_YAML" --out "$OUT"
|
||||
|
||||
echo "=== sweep complete ==="
|
||||
ls -lh "$OUT"
|
||||
|
||||
# ── aggregate: runs fxt-backtest aggregate on the GPU pool. ───
|
||||
# The binary is dynamically linked against libcuda.so.1, which the
|
||||
# ci-compile-cpu pool's host doesn't expose; the aggregate workload
|
||||
# itself is CPU-only (~seconds) but the binary needs the driver libs.
|
||||
# Reusing the GPU pool is cheap (single fast pod) and avoids needing
|
||||
# a separate CPU-only binary.
|
||||
- name: aggregate
|
||||
inputs:
|
||||
parameters:
|
||||
- name: sha
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: "{{workflow.parameters.gpu-pool}}"
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
tolerations:
|
||||
- key: nvidia.com/gpu
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
container:
|
||||
image: gitlab-registry.foxhunt.svc.cluster.local:5000/root/foxhunt/foxhunt-training-runtime:latest
|
||||
imagePullPolicy: Always
|
||||
command: ["/bin/bash", "-c"]
|
||||
resources:
|
||||
requests:
|
||||
cpu: "2"
|
||||
memory: 4Gi
|
||||
# Request a GPU only to make the Scaleway L40S device plugin
|
||||
# mount libcuda.so.1 into the container — aggregate logic is
|
||||
# CPU-only (~seconds). Without the request, the binary's
|
||||
# dynamic loader fails before main(). Cheapest fix; the GPU
|
||||
# is held for ~5s of CPU work, then released.
|
||||
nvidia.com/gpu: "1"
|
||||
limits:
|
||||
cpu: "4"
|
||||
memory: 8Gi
|
||||
nvidia.com/gpu: "1"
|
||||
volumeMounts:
|
||||
- name: training-data
|
||||
mountPath: /mnt/training-data
|
||||
args:
|
||||
- |
|
||||
set -e
|
||||
SHA="{{inputs.parameters.sha}}"
|
||||
BIN="/mnt/training-data/bin/$SHA/fxt-backtest"
|
||||
SWEEP_DIR="{{workflow.parameters.sweep-root}}/{{workflow.parameters.sweep-tag}}"
|
||||
echo "=== aggregate sweep at $SWEEP_DIR ==="
|
||||
"$BIN" aggregate "$SWEEP_DIR"
|
||||
ls -lh "$SWEEP_DIR"
|
||||
|
||||
# ── notify-result: exit hook (placeholder; real impl emits to Slack/MinIO) ─
|
||||
- name: notify-result
|
||||
container:
|
||||
image: alpine:3.20
|
||||
command: ["sh", "-c"]
|
||||
args:
|
||||
- |
|
||||
echo "lob-backtest-sweep workflow {{workflow.name}} finished with status {{workflow.status}}"
|
||||
echo "results under {{workflow.parameters.sweep-root}}/{{workflow.parameters.sweep-tag}}"
|
||||
@@ -1,318 +0,0 @@
|
||||
# infra/k8s/argo/nsys-test-template.yaml
|
||||
#
|
||||
# One-shot Nsight Systems profiling run on L40S. Wraps a smoke test under
|
||||
# `nsys profile` and uploads the .nsys-rep to MinIO at
|
||||
# foxhunt-training-results/profiles/smoke/<short-sha>/.
|
||||
#
|
||||
# Mirrors the nsys integration in train-multi-seed-template (Plan 5 Task 3,
|
||||
# A.4.1) but for smoke tests. Use this for rapid kernel-timing iteration
|
||||
# (per-step bottleneck hunts, SOL %, occupancy diffs) without the full
|
||||
# multi-seed Argo cycle.
|
||||
#
|
||||
# Why L40S not local: laptop RTX 3050 Ti has 4 GB VRAM; nsys on L40S (48 GB)
|
||||
# avoids any contention and runs at near-native speed (typical ~5-10%
|
||||
# slowdown vs uninstrumented).
|
||||
#
|
||||
# Usage:
|
||||
# argo submit --watch -n foxhunt nsys-test-template.yaml \
|
||||
# -p commit-ref=main \
|
||||
# -p test-name=performance::test_real_data_single_epoch
|
||||
#
|
||||
# Or via the wrapper script: scripts/argo-nsys.sh
|
||||
|
||||
apiVersion: argoproj.io/v1alpha1
|
||||
kind: WorkflowTemplate
|
||||
metadata:
|
||||
name: nsys-test
|
||||
namespace: foxhunt
|
||||
labels:
|
||||
app.kubernetes.io/name: nsys-test
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
spec:
|
||||
entrypoint: nsys-run
|
||||
serviceAccountName: argo-workflow
|
||||
podMetadata:
|
||||
labels:
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
app.kubernetes.io/component: nsys-test
|
||||
archiveLogs: true
|
||||
podGC:
|
||||
strategy: OnPodCompletion
|
||||
ttlStrategy:
|
||||
secondsAfterCompletion: 7200
|
||||
activeDeadlineSeconds: 3600 # nsys overhead is small — 1h cap is generous
|
||||
arguments:
|
||||
parameters:
|
||||
- name: commit-ref
|
||||
value: HEAD
|
||||
- name: gpu-pool
|
||||
value: ci-training-l40s
|
||||
- name: cuda-compute-cap
|
||||
value: "89"
|
||||
- name: test-name
|
||||
value: "performance::test_real_data_single_epoch"
|
||||
# nsys capture options. Defaults capture the full kernel timeline.
|
||||
# GPU metrics counters (--gpu-metrics-devices) require elevated CUDA
|
||||
# performance-counter privileges (NVGPUCTRPERM) which the container
|
||||
# lacks; CUDA/NVTX/osrt traces alone still give per-kernel timing and
|
||||
# call-graph context — sufficient for the bottleneck-hunt use case.
|
||||
# Override via --extra-args if running on a node with relaxed perms.
|
||||
- name: nsys-extra-args
|
||||
value: "--trace=cuda,nvtx,osrt"
|
||||
|
||||
volumes:
|
||||
- name: git-ssh-key
|
||||
secret:
|
||||
secretName: argo-git-ssh-key
|
||||
defaultMode: 256
|
||||
- name: cargo-target
|
||||
persistentVolumeClaim:
|
||||
claimName: cargo-target-cuda-test
|
||||
# Use the same training-data-pvc as production training (full 27 months
|
||||
# of OHLCV + MBP-10 + trades) instead of the smaller test-data-pvc.
|
||||
- name: training-data
|
||||
persistentVolumeClaim:
|
||||
claimName: training-data-pvc
|
||||
readOnly: true
|
||||
|
||||
templates:
|
||||
- name: nsys-run
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: "{{workflow.parameters.gpu-pool}}"
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
tolerations:
|
||||
- key: nvidia.com/gpu
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
- key: node.cilium.io/agent-not-ready
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
container:
|
||||
image: gitlab-registry.foxhunt.svc.cluster.local:5000/root/foxhunt/ci-builder:latest
|
||||
command: ["/bin/bash", "-c"]
|
||||
env:
|
||||
- name: SQLX_OFFLINE
|
||||
value: "true"
|
||||
- name: CARGO_TERM_COLOR
|
||||
value: always
|
||||
- name: CARGO_TARGET_DIR
|
||||
value: /cargo-target
|
||||
- name: CARGO_HOME
|
||||
value: /cargo-target/cargo-home
|
||||
- name: CUDA_VISIBLE_DEVICES
|
||||
value: "0"
|
||||
- name: CUBLAS_WORKSPACE_CONFIG
|
||||
value: ":4096:8"
|
||||
- name: FOXHUNT_TEST_DATA
|
||||
value: /data/futures-baseline
|
||||
- name: TEST_DATA_DIR
|
||||
value: /data/futures-baseline
|
||||
# MBP-10 + trades are mandatory per feedback_mbp10_mandatory.md
|
||||
# (OFI features are part of the production state vector).
|
||||
- name: FOXHUNT_MBP10_DATA
|
||||
value: /data/futures-baseline-mbp10
|
||||
- name: FOXHUNT_TRADES_DATA
|
||||
value: /data/futures-baseline-trades
|
||||
- name: RUST_LOG
|
||||
value: info
|
||||
- name: LD_LIBRARY_PATH
|
||||
value: /usr/local/nvidia/lib64:/usr/local/cuda/lib64:/usr/local/cuda/extras/CUPTI/lib64
|
||||
# MinIO upload — same secret + optional pattern as train-multi-seed.
|
||||
# Both refs are `optional: true` so env mount succeeds on clusters
|
||||
# that do not pre-create `minio-credentials` (upload then warn-fails).
|
||||
- name: MINIO_ACCESS_KEY
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: minio-credentials
|
||||
key: access-key
|
||||
optional: true
|
||||
- name: MINIO_SECRET_KEY
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: minio-credentials
|
||||
key: secret-key
|
||||
optional: true
|
||||
volumeMounts:
|
||||
- name: git-ssh-key
|
||||
mountPath: /etc/git-ssh
|
||||
readOnly: true
|
||||
- name: cargo-target
|
||||
mountPath: /cargo-target
|
||||
- name: training-data
|
||||
mountPath: /data
|
||||
readOnly: true
|
||||
resources:
|
||||
requests:
|
||||
nvidia.com/gpu: "1"
|
||||
cpu: "4"
|
||||
memory: 32Gi
|
||||
limits:
|
||||
nvidia.com/gpu: "1"
|
||||
cpu: "8"
|
||||
memory: 64Gi
|
||||
args:
|
||||
- |
|
||||
set -euo pipefail
|
||||
|
||||
REF="{{workflow.parameters.commit-ref}}"
|
||||
TEST_NAME="{{workflow.parameters.test-name}}"
|
||||
EXTRA_ARGS="{{workflow.parameters.nsys-extra-args}}"
|
||||
|
||||
echo "==================================="
|
||||
echo " nsys profile L40S run"
|
||||
echo "==================================="
|
||||
echo " Ref: $REF"
|
||||
echo " Test: $TEST_NAME"
|
||||
echo " Extra: $EXTRA_ARGS"
|
||||
echo " Pool: {{workflow.parameters.gpu-pool}}"
|
||||
echo "==================================="
|
||||
|
||||
# --- SSH setup ---
|
||||
mkdir -p ~/.ssh
|
||||
cp /etc/git-ssh/ssh-privatekey ~/.ssh/id_ed25519
|
||||
chmod 600 ~/.ssh/id_ed25519
|
||||
printf 'Host *\n StrictHostKeyChecking no\n UserKnownHostsFile /dev/null\n' > ~/.ssh/config
|
||||
chmod 600 ~/.ssh/config
|
||||
|
||||
REPO="ssh://git@gitlab-gitlab-shell.foxhunt.svc.cluster.local:2222/root/foxhunt.git"
|
||||
BUILD="/cargo-target/src"
|
||||
git config --global --add safe.directory "$BUILD"
|
||||
|
||||
# --- Persistent checkout on PVC ---
|
||||
if [ -d "$BUILD/.git" ]; then
|
||||
cd "$BUILD"
|
||||
git fetch origin
|
||||
if git rev-parse --verify "origin/$REF" >/dev/null 2>&1; then
|
||||
TARGET=$(git rev-parse "origin/$REF")
|
||||
else
|
||||
TARGET=$(git rev-parse --verify "$REF" 2>/dev/null || echo "$REF")
|
||||
fi
|
||||
CURRENT=$(git rev-parse HEAD 2>/dev/null || echo "none")
|
||||
if [ "$CURRENT" != "$TARGET" ]; then
|
||||
echo "=== Updating checkout to $REF ($TARGET) ==="
|
||||
git checkout --force --detach "$TARGET"
|
||||
git clean -fd
|
||||
fi
|
||||
else
|
||||
git clone --filter=blob:none "$REPO" "$BUILD"
|
||||
cd "$BUILD"
|
||||
git fetch origin
|
||||
if git rev-parse --verify "origin/$REF" >/dev/null 2>&1; then
|
||||
git checkout --force --detach "origin/$REF"
|
||||
else
|
||||
git checkout "$REF"
|
||||
fi
|
||||
fi
|
||||
SHA=$(git rev-parse --short HEAD)
|
||||
echo "Checked out $SHA"
|
||||
|
||||
export PATH="${CARGO_HOME}/bin:/usr/local/cuda/bin:${PATH}"
|
||||
export CUDA_COMPUTE_CAP={{workflow.parameters.cuda-compute-cap}}
|
||||
|
||||
bash scripts/ptx-cache-invalidate.sh "${CARGO_TARGET_DIR}/.ptx_cache"
|
||||
|
||||
# --- Verify nsys ---
|
||||
NSYS_BIN=$(which nsys 2>/dev/null || echo /usr/local/cuda/bin/nsys)
|
||||
if [ ! -x "$NSYS_BIN" ]; then
|
||||
echo "ERROR: nsys not found at $NSYS_BIN"
|
||||
exit 2
|
||||
fi
|
||||
echo "=== nsys version ==="
|
||||
"$NSYS_BIN" --version | head -3
|
||||
|
||||
# --- Compile test binary + train_baseline_rl ---
|
||||
echo "=== Compiling (--release --no-run + train_baseline_rl) ==="
|
||||
cargo test -p ml --release --lib --no-run 2>&1 | tee /cargo-target/nsys-compile.log
|
||||
cargo build -p ml --release --example train_baseline_rl 2>&1 | tee -a /cargo-target/nsys-compile.log
|
||||
|
||||
# Honour CARGO_TARGET_DIR — outputs go to ${CARGO_TARGET_DIR}/release/deps,
|
||||
# NOT the repo-relative target/. (CI workflow sets CARGO_TARGET_DIR=/cargo-target.)
|
||||
# Subshell + || true wraps the head -1 SIGPIPE so pipefail doesn't fire.
|
||||
TARGET_DIR="${CARGO_TARGET_DIR:-target}"
|
||||
TEST_BIN=$( (ls -t "$TARGET_DIR/release/deps/ml-"* 2>/dev/null \
|
||||
| grep -vE '\.(d|rcgu|rmeta|rlib|so|json)$' \
|
||||
| head -1) || true )
|
||||
if [ -z "$TEST_BIN" ] || [ ! -x "$TEST_BIN" ]; then
|
||||
echo "ERROR: could not locate compiled test binary in $TARGET_DIR/release/deps/"
|
||||
ls -la "$TARGET_DIR/release/deps/" 2>&1 | head -20
|
||||
exit 3
|
||||
fi
|
||||
echo "=== Test binary: $TEST_BIN ==="
|
||||
|
||||
# grep -q + pipefail = SIGPIPE on early exit. Capture once, grep
|
||||
# without -q so the consumer reads all input.
|
||||
TEST_LIST=$("$TEST_BIN" --list 2>&1 || true)
|
||||
if ! echo "$TEST_LIST" | grep -F "$TEST_NAME" > /dev/null; then
|
||||
echo "ERROR: test '$TEST_NAME' not found"
|
||||
echo "$TEST_LIST" | grep smoke_tests | (head -30 || true)
|
||||
exit 4
|
||||
fi
|
||||
|
||||
# --- Run under nsys profile ---
|
||||
mkdir -p /tmp/nsys-output
|
||||
POD_NAME="${HOSTNAME:-pod}"
|
||||
NSYS_OUT="/tmp/nsys-output/profile-${POD_NAME}.nsys-rep"
|
||||
|
||||
echo ""
|
||||
echo "==================================="
|
||||
echo " Launching nsys profile..."
|
||||
echo "==================================="
|
||||
START_TS=$(date +%s)
|
||||
set +e
|
||||
"$NSYS_BIN" profile \
|
||||
--output="$NSYS_OUT" \
|
||||
--force-overwrite=true \
|
||||
--stats=true \
|
||||
${EXTRA_ARGS} \
|
||||
-- "$TEST_BIN" "$TEST_NAME" --ignored --nocapture
|
||||
EXIT_CODE=$?
|
||||
set -e
|
||||
END_TS=$(date +%s)
|
||||
ELAPSED=$((END_TS - START_TS))
|
||||
|
||||
echo ""
|
||||
echo "==================================="
|
||||
echo " nsys run complete"
|
||||
echo "==================================="
|
||||
echo " Exit code: $EXIT_CODE"
|
||||
echo " Elapsed: ${ELAPSED}s"
|
||||
ls -la "$NSYS_OUT" 2>/dev/null && \
|
||||
echo " Profile: $(du -h "$NSYS_OUT" | cut -f1)"
|
||||
echo ""
|
||||
|
||||
# --- Upload to MinIO (same pattern as train-multi-seed) ---
|
||||
if [ -f "$NSYS_OUT" ]; then
|
||||
echo "=== Uploading nsys profile to MinIO ==="
|
||||
MC_BIN=$(which mc 2>/dev/null || echo "")
|
||||
if [ -z "$MC_BIN" ]; then
|
||||
MC_BIN=/tmp/mc
|
||||
curl -fsSL https://dl.min.io/client/mc/release/linux-amd64/mc -o "$MC_BIN" || {
|
||||
echo "WARN: failed to download mc — skipping upload"
|
||||
echo "Profile available locally at: $NSYS_OUT (PVC: cargo-target)"
|
||||
exit 0
|
||||
}
|
||||
chmod +x "$MC_BIN"
|
||||
fi
|
||||
"$MC_BIN" alias set foxhunt http://minio.foxhunt.svc.cluster.local:9000 \
|
||||
"${MINIO_ACCESS_KEY:-}" "${MINIO_SECRET_KEY:-}" 2>/dev/null || true
|
||||
# Smoke profiles go to a separate prefix to keep the
|
||||
# production training profile bucket clean.
|
||||
UPLOAD_PATH="foxhunt/foxhunt-training-results/profiles/smoke/${SHA}/profile-${POD_NAME}.nsys-rep"
|
||||
if "$MC_BIN" cp "$NSYS_OUT" "$UPLOAD_PATH"; then
|
||||
echo "=== nsys profile uploaded to ${UPLOAD_PATH} ==="
|
||||
echo ""
|
||||
echo "Download with:"
|
||||
echo " mc cp ${UPLOAD_PATH} ./profile.nsys-rep"
|
||||
echo "Open with: nsys-ui profile.nsys-rep"
|
||||
else
|
||||
echo "WARN: nsys upload failed — profile remains on PVC at $NSYS_OUT"
|
||||
fi
|
||||
fi
|
||||
|
||||
if [ "$EXIT_CODE" -ne 0 ]; then
|
||||
echo ""
|
||||
echo "FAIL: test exited with code $EXIT_CODE"
|
||||
exit "$EXIT_CODE"
|
||||
fi
|
||||
echo ""
|
||||
echo "PASS: nsys profile captured, test passed."
|
||||
@@ -1,181 +0,0 @@
|
||||
# refresh-deps-cache: rebuild ci-builder-cpu-with-deps:nightly image
|
||||
#
|
||||
# Triggered:
|
||||
# - nightly via CronWorkflow (see CronWorkflow at the bottom of this file)
|
||||
# - manually via: argo submit --from workflowtemplate/refresh-deps-cache -n foxhunt
|
||||
#
|
||||
# What it does: clones the repo at HEAD of main, then runs Kaniko to build
|
||||
# infra/docker/Dockerfile.ci-deps-cache and pushes to
|
||||
# gitlab-registry.foxhunt.svc.cluster.local:5000/root/foxhunt/ci-builder-cpu-with-deps:nightly.
|
||||
#
|
||||
# The compile-and-deploy workflow's seed-deps-cache initContainer pulls
|
||||
# this image and rsyncs its /cargo-target-prebuilt/ into the cargo-target-cpu PVC.
|
||||
---
|
||||
apiVersion: argoproj.io/v1alpha1
|
||||
kind: WorkflowTemplate
|
||||
metadata:
|
||||
name: refresh-deps-cache
|
||||
namespace: foxhunt
|
||||
labels:
|
||||
app.kubernetes.io/name: refresh-deps-cache
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
spec:
|
||||
activeDeadlineSeconds: 7200 # 2 hours; full workspace cargo build is slow
|
||||
serviceAccountName: argo-workflow
|
||||
entrypoint: build
|
||||
podMetadata:
|
||||
labels:
|
||||
# Reuse ci-pipeline label so the pod inherits argo-ci-pipeline netpol egress
|
||||
# clone, gitlab-registry:5000 for image push, crates.io HTTPS for cargo
|
||||
# NetworkPolicy clone — same egress targets, no functional difference.
|
||||
app.kubernetes.io/component: ci-pipeline
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
ttlStrategy:
|
||||
secondsAfterCompletion: 7200
|
||||
arguments:
|
||||
parameters:
|
||||
- name: commit-sha
|
||||
value: HEAD
|
||||
- name: deps-version
|
||||
value: "1"
|
||||
- name: image-tag
|
||||
value: nightly
|
||||
|
||||
templates:
|
||||
- name: build
|
||||
inputs:
|
||||
parameters:
|
||||
- name: commit-sha
|
||||
value: "{{workflow.parameters.commit-sha}}"
|
||||
- name: deps-version
|
||||
value: "{{workflow.parameters.deps-version}}"
|
||||
- name: image-tag
|
||||
value: "{{workflow.parameters.image-tag}}"
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: ci-compile-cpu
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
tolerations:
|
||||
- key: node.cilium.io/agent-not-ready
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
volumes:
|
||||
- name: workspace
|
||||
emptyDir:
|
||||
sizeLimit: 30Gi # full workspace + target/release ~10GB
|
||||
- name: git-ssh-key
|
||||
secret:
|
||||
secretName: argo-git-ssh-key
|
||||
defaultMode: 256
|
||||
- name: registry-auth
|
||||
secret:
|
||||
secretName: gitlab-registry
|
||||
items:
|
||||
- key: .dockerconfigjson
|
||||
path: config.json
|
||||
initContainers:
|
||||
- name: git-clone
|
||||
image: alpine/git:latest
|
||||
command: ["/bin/sh", "-c"]
|
||||
resources:
|
||||
requests:
|
||||
cpu: 200m
|
||||
memory: 256Mi
|
||||
limits:
|
||||
cpu: "1"
|
||||
memory: 512Mi
|
||||
volumeMounts:
|
||||
- name: git-ssh-key
|
||||
mountPath: /etc/git-ssh
|
||||
readOnly: true
|
||||
- name: workspace
|
||||
mountPath: /workspace
|
||||
args:
|
||||
- |
|
||||
set -ex
|
||||
mkdir -p /root/.ssh
|
||||
cp /etc/git-ssh/ssh-privatekey /root/.ssh/id_ed25519
|
||||
chmod 600 /root/.ssh/id_ed25519
|
||||
printf 'Host *\n StrictHostKeyChecking no\n UserKnownHostsFile /dev/null\n' > /root/.ssh/config
|
||||
chmod 600 /root/.ssh/config
|
||||
|
||||
SHA="{{inputs.parameters.commit-sha}}"
|
||||
REPO="ssh://git@gitlab-gitlab-shell.foxhunt.svc.cluster.local:2222/root/foxhunt.git"
|
||||
git clone --no-checkout --filter=blob:none "$REPO" /workspace/src
|
||||
cd /workspace/src
|
||||
git checkout "$SHA"
|
||||
echo "Checked out $(git rev-parse --short HEAD)"
|
||||
container:
|
||||
image: gcr.io/kaniko-project/executor:debug
|
||||
command: ["/busybox/sh", "-c"]
|
||||
env:
|
||||
- name: DOCKER_CONFIG
|
||||
value: /kaniko/.docker
|
||||
# Larger than build-ci-image because the inner cargo build is heavy.
|
||||
resources:
|
||||
requests:
|
||||
cpu: "8"
|
||||
memory: 16Gi
|
||||
limits:
|
||||
cpu: "16"
|
||||
memory: 32Gi
|
||||
volumeMounts:
|
||||
- name: registry-auth
|
||||
mountPath: /kaniko/.docker
|
||||
readOnly: true
|
||||
- name: workspace
|
||||
mountPath: /workspace
|
||||
args:
|
||||
- |
|
||||
/kaniko/executor \
|
||||
--context=/workspace/src \
|
||||
--dockerfile=/workspace/src/infra/docker/Dockerfile.ci-deps-cache \
|
||||
--destination=gitlab-registry.foxhunt.svc.cluster.local:5000/root/foxhunt/ci-builder-cpu-with-deps:{{inputs.parameters.image-tag}} \
|
||||
--build-arg=DEPS_VERSION={{inputs.parameters.deps-version}} \
|
||||
--build-arg=BASE_TAG=latest \
|
||||
--insecure-registry=gitlab-registry.foxhunt.svc.cluster.local:5000 \
|
||||
--skip-tls-verify-registry=gitlab-registry.foxhunt.svc.cluster.local:5000 \
|
||||
--cache=true \
|
||||
--cache-repo=gitlab-registry.foxhunt.svc.cluster.local:5000/root/foxhunt/cache \
|
||||
--snapshot-mode=redo
|
||||
|
||||
---
|
||||
# Nightly cron — rebuilds the deps cache image at 03:00 UTC.
|
||||
# Suspended by default; enable with:
|
||||
# kubectl -n foxhunt patch cronworkflow refresh-deps-cache-nightly \
|
||||
# -p '{"spec":{"suspend":false}}' --type=merge
|
||||
apiVersion: argoproj.io/v1alpha1
|
||||
kind: CronWorkflow
|
||||
metadata:
|
||||
name: refresh-deps-cache-nightly
|
||||
namespace: foxhunt
|
||||
labels:
|
||||
app.kubernetes.io/name: refresh-deps-cache-nightly
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
spec:
|
||||
schedules:
|
||||
- "0 3 * * *"
|
||||
timezone: UTC
|
||||
concurrencyPolicy: Forbid
|
||||
successfulJobsHistoryLimit: 2
|
||||
failedJobsHistoryLimit: 3
|
||||
suspend: true
|
||||
workflowSpec:
|
||||
entrypoint: trigger
|
||||
serviceAccountName: argo-workflow
|
||||
ttlStrategy:
|
||||
secondsAfterCompletion: 7200
|
||||
templates:
|
||||
- name: trigger
|
||||
steps:
|
||||
- - name: refresh
|
||||
templateRef:
|
||||
name: refresh-deps-cache
|
||||
template: build
|
||||
arguments:
|
||||
parameters:
|
||||
- name: commit-sha
|
||||
value: HEAD
|
||||
- name: deps-version
|
||||
value: "1"
|
||||
- name: image-tag
|
||||
value: nightly
|
||||
@@ -1,311 +0,0 @@
|
||||
# infra/k8s/argo/sanitizer-test-template.yaml
|
||||
#
|
||||
# One-shot compute-sanitizer run on L40S. Wraps a smoke test under
|
||||
# `compute-sanitizer --tool memcheck` to validate no CUDA memory errors
|
||||
# (out-of-bounds, leaks, sync violations, race conditions).
|
||||
#
|
||||
# Why L40S not local: the laptop RTX 3050 Ti (4 GB VRAM) cannot fit
|
||||
# compute-sanitizer's instrumentation metadata alongside the production
|
||||
# training workload — the sanitizer falls back to "didn't track the
|
||||
# launch" with 60k+ internal-allocation errors. L40S (48 GB) has ample
|
||||
# headroom for memcheck's 2-3× shadow-memory overhead.
|
||||
#
|
||||
# Usage:
|
||||
# argo submit --watch -n foxhunt sanitizer-test-template.yaml \
|
||||
# -p commit-ref=main \
|
||||
# -p test-name=iqn_quantile_monotonicity::iqn_multi_quantile_heads_produce_monotonic_estimates \
|
||||
# -p sanitizer-tool=memcheck
|
||||
#
|
||||
# Or via the wrapper script: scripts/argo-sanitizer.sh
|
||||
|
||||
apiVersion: argoproj.io/v1alpha1
|
||||
kind: WorkflowTemplate
|
||||
metadata:
|
||||
name: sanitizer-test
|
||||
namespace: foxhunt
|
||||
labels:
|
||||
app.kubernetes.io/name: sanitizer-test
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
spec:
|
||||
entrypoint: sanitizer-run
|
||||
serviceAccountName: argo-workflow
|
||||
podMetadata:
|
||||
labels:
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
app.kubernetes.io/component: sanitizer-test
|
||||
archiveLogs: true
|
||||
podGC:
|
||||
strategy: OnPodCompletion
|
||||
ttlStrategy:
|
||||
secondsAfterCompletion: 7200
|
||||
activeDeadlineSeconds: 7200 # 2h cap — sanitizer is slow but bounded
|
||||
arguments:
|
||||
parameters:
|
||||
- name: commit-ref
|
||||
value: HEAD
|
||||
- name: gpu-pool
|
||||
value: ci-training-l40s
|
||||
- name: cuda-compute-cap
|
||||
value: "89" # L40S Ada Lovelace
|
||||
- name: test-name
|
||||
value: "iqn_quantile_monotonicity::iqn_multi_quantile_heads_produce_monotonic_estimates"
|
||||
- name: sanitizer-tool
|
||||
value: memcheck # memcheck | racecheck | synccheck | initcheck
|
||||
- name: sanitizer-extra-args
|
||||
value: ""
|
||||
|
||||
volumes:
|
||||
- name: git-ssh-key
|
||||
secret:
|
||||
secretName: argo-git-ssh-key
|
||||
defaultMode: 256
|
||||
- name: cargo-target
|
||||
persistentVolumeClaim:
|
||||
claimName: cargo-target-cuda-test
|
||||
# Use the same training-data-pvc as production training (full 27 months
|
||||
# of OHLCV + MBP-10 + trades) instead of the smaller test-data-pvc which
|
||||
# only has 3-4 months of MBP-10 — walk-forward needs ≥10 months.
|
||||
- name: training-data
|
||||
persistentVolumeClaim:
|
||||
claimName: training-data-pvc
|
||||
readOnly: true
|
||||
|
||||
templates:
|
||||
- name: sanitizer-run
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: "{{workflow.parameters.gpu-pool}}"
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
tolerations:
|
||||
- key: nvidia.com/gpu
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
- key: node.cilium.io/agent-not-ready
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
container:
|
||||
image: gitlab-registry.foxhunt.svc.cluster.local:5000/root/foxhunt/ci-builder:latest
|
||||
command: ["/bin/bash", "-c"]
|
||||
env:
|
||||
- name: SQLX_OFFLINE
|
||||
value: "true"
|
||||
- name: CARGO_TERM_COLOR
|
||||
value: always
|
||||
- name: CARGO_TARGET_DIR
|
||||
value: /cargo-target
|
||||
- name: CARGO_HOME
|
||||
value: /cargo-target/cargo-home
|
||||
- name: CUDA_VISIBLE_DEVICES
|
||||
value: "0"
|
||||
- name: CUBLAS_WORKSPACE_CONFIG
|
||||
value: ":4096:8"
|
||||
- name: FOXHUNT_TEST_DATA
|
||||
value: /data/futures-baseline
|
||||
- name: TEST_DATA_DIR
|
||||
value: /data/futures-baseline
|
||||
# MBP-10 + trades are mandatory per feedback_mbp10_mandatory.md
|
||||
# (OFI features are part of the production state vector).
|
||||
- name: FOXHUNT_MBP10_DATA
|
||||
value: /data/futures-baseline-mbp10
|
||||
- name: FOXHUNT_TRADES_DATA
|
||||
value: /data/futures-baseline-trades
|
||||
- name: RUST_LOG
|
||||
value: info
|
||||
- name: LD_LIBRARY_PATH
|
||||
value: /usr/local/nvidia/lib64:/usr/local/cuda/lib64:/usr/local/cuda/extras/CUPTI/lib64
|
||||
volumeMounts:
|
||||
- name: git-ssh-key
|
||||
mountPath: /etc/git-ssh
|
||||
readOnly: true
|
||||
- name: cargo-target
|
||||
mountPath: /cargo-target
|
||||
- name: training-data
|
||||
mountPath: /data
|
||||
readOnly: true
|
||||
resources:
|
||||
requests:
|
||||
nvidia.com/gpu: "1"
|
||||
cpu: "4"
|
||||
memory: 32Gi
|
||||
limits:
|
||||
nvidia.com/gpu: "1"
|
||||
cpu: "8"
|
||||
memory: 64Gi
|
||||
args:
|
||||
- |
|
||||
set -euo pipefail
|
||||
|
||||
REF="{{workflow.parameters.commit-ref}}"
|
||||
TEST_NAME="{{workflow.parameters.test-name}}"
|
||||
SANITIZER_TOOL="{{workflow.parameters.sanitizer-tool}}"
|
||||
EXTRA_ARGS="{{workflow.parameters.sanitizer-extra-args}}"
|
||||
|
||||
echo "==================================="
|
||||
echo " compute-sanitizer L40S run"
|
||||
echo "==================================="
|
||||
echo " Ref: $REF"
|
||||
echo " Test: $TEST_NAME"
|
||||
echo " Tool: $SANITIZER_TOOL"
|
||||
echo " Pool: {{workflow.parameters.gpu-pool}}"
|
||||
echo " Compute: {{workflow.parameters.cuda-compute-cap}}"
|
||||
echo "==================================="
|
||||
|
||||
# --- SSH setup (same as gpu-test-pipeline) ---
|
||||
mkdir -p ~/.ssh
|
||||
cp /etc/git-ssh/ssh-privatekey ~/.ssh/id_ed25519
|
||||
chmod 600 ~/.ssh/id_ed25519
|
||||
printf 'Host *\n StrictHostKeyChecking no\n UserKnownHostsFile /dev/null\n' > ~/.ssh/config
|
||||
chmod 600 ~/.ssh/config
|
||||
|
||||
REPO="ssh://git@gitlab-gitlab-shell.foxhunt.svc.cluster.local:2222/root/foxhunt.git"
|
||||
BUILD="/cargo-target/src"
|
||||
git config --global --add safe.directory "$BUILD"
|
||||
|
||||
# --- Persistent checkout on PVC ---
|
||||
if [ -d "$BUILD/.git" ]; then
|
||||
cd "$BUILD"
|
||||
git fetch origin
|
||||
if git rev-parse --verify "origin/$REF" >/dev/null 2>&1; then
|
||||
TARGET=$(git rev-parse "origin/$REF")
|
||||
else
|
||||
TARGET=$(git rev-parse --verify "$REF" 2>/dev/null || echo "$REF")
|
||||
fi
|
||||
CURRENT=$(git rev-parse HEAD 2>/dev/null || echo "none")
|
||||
if [ "$CURRENT" != "$TARGET" ]; then
|
||||
echo "=== Updating checkout to $REF ($TARGET) ==="
|
||||
git checkout --force --detach "$TARGET"
|
||||
git clean -fd
|
||||
fi
|
||||
else
|
||||
echo "=== Initial clone ==="
|
||||
git clone --filter=blob:none "$REPO" "$BUILD"
|
||||
cd "$BUILD"
|
||||
git fetch origin
|
||||
if git rev-parse --verify "origin/$REF" >/dev/null 2>&1; then
|
||||
git checkout --force --detach "origin/$REF"
|
||||
else
|
||||
git checkout "$REF"
|
||||
fi
|
||||
fi
|
||||
echo "Checked out $(git rev-parse --short HEAD)"
|
||||
|
||||
export PATH="${CARGO_HOME}/bin:/usr/local/cuda/bin:${PATH}"
|
||||
export CUDA_COMPUTE_CAP={{workflow.parameters.cuda-compute-cap}}
|
||||
|
||||
# --- PTX cache invalidation ---
|
||||
bash scripts/ptx-cache-invalidate.sh "${CARGO_TARGET_DIR}/.ptx_cache"
|
||||
|
||||
# --- Verify compute-sanitizer is available ---
|
||||
if ! command -v compute-sanitizer >/dev/null 2>&1; then
|
||||
echo "ERROR: compute-sanitizer not found in PATH. CUDA toolkit incomplete."
|
||||
exit 2
|
||||
fi
|
||||
echo "=== compute-sanitizer version ==="
|
||||
compute-sanitizer --version | head -3
|
||||
|
||||
# --- Compile test binary + train_baseline_rl example ---
|
||||
# train_baseline_rl is needed for multi_fold_convergence test which
|
||||
# spawns it as a subprocess via `cargo run --example`. Pre-building
|
||||
# avoids cargo doing it inside the sanitizer-instrumented run.
|
||||
echo "=== Compiling test binary + train_baseline_rl (--release) ==="
|
||||
cargo test -p ml --release --lib --no-run 2>&1 | tee /cargo-target/sanitizer-compile.log
|
||||
cargo build -p ml --release --example train_baseline_rl 2>&1 | tee -a /cargo-target/sanitizer-compile.log
|
||||
|
||||
# Locate the freshest ml-* test binary
|
||||
# Honour CARGO_TARGET_DIR — outputs go to ${CARGO_TARGET_DIR}/release/deps,
|
||||
# NOT the repo-relative target/. (CI workflow sets CARGO_TARGET_DIR=/cargo-target.)
|
||||
# Subshell + || true wraps the head -1 SIGPIPE so pipefail doesn't fire.
|
||||
TARGET_DIR="${CARGO_TARGET_DIR:-target}"
|
||||
TEST_BIN=$( (ls -t "$TARGET_DIR/release/deps/ml-"* 2>/dev/null \
|
||||
| grep -vE '\.(d|rcgu|rmeta|rlib|so|json)$' \
|
||||
| head -1) || true )
|
||||
if [ -z "$TEST_BIN" ] || [ ! -x "$TEST_BIN" ]; then
|
||||
echo "ERROR: could not locate compiled test binary in $TARGET_DIR/release/deps/"
|
||||
ls -la "$TARGET_DIR/release/deps/" 2>&1 | head -20
|
||||
exit 3
|
||||
fi
|
||||
echo "=== Test binary: $TEST_BIN ==="
|
||||
|
||||
# --- Verify the test exists in the binary ---
|
||||
# grep -q + pipefail = SIGPIPE on early exit. Capture once, grep
|
||||
# without -q so the consumer reads all input. Same below for nsys.
|
||||
TEST_LIST=$("$TEST_BIN" --list 2>&1 || true)
|
||||
if ! echo "$TEST_LIST" | grep -F "$TEST_NAME" > /dev/null; then
|
||||
echo "ERROR: test '$TEST_NAME' not found in binary."
|
||||
echo "Available smoke tests:"
|
||||
echo "$TEST_LIST" | grep smoke_tests | (head -30 || true)
|
||||
exit 4
|
||||
fi
|
||||
|
||||
# --- Run under compute-sanitizer ---
|
||||
mkdir -p /tmp/sanitizer-output
|
||||
LOG_FILE=/tmp/sanitizer-output/${SANITIZER_TOOL}.log
|
||||
|
||||
echo ""
|
||||
echo "==================================="
|
||||
echo " Launching compute-sanitizer..."
|
||||
echo "==================================="
|
||||
START_TS=$(date +%s)
|
||||
# --target-processes all so multi_fold_convergence's spawned
|
||||
# `train_baseline_rl` subprocess is also instrumented (the test
|
||||
# itself does no GPU work — the child binary does).
|
||||
set +e
|
||||
compute-sanitizer \
|
||||
--tool "$SANITIZER_TOOL" \
|
||||
--target-processes all \
|
||||
--launch-timeout 1200 \
|
||||
--error-exitcode 99 \
|
||||
--print-limit 200 \
|
||||
--log-file "$LOG_FILE" \
|
||||
${EXTRA_ARGS} \
|
||||
-- "$TEST_BIN" "$TEST_NAME" --ignored --nocapture
|
||||
EXIT_CODE=$?
|
||||
set -e
|
||||
END_TS=$(date +%s)
|
||||
ELAPSED=$((END_TS - START_TS))
|
||||
|
||||
echo ""
|
||||
echo "==================================="
|
||||
echo " Sanitizer run complete"
|
||||
echo "==================================="
|
||||
echo " Exit code: $EXIT_CODE"
|
||||
echo " Elapsed: ${ELAPSED}s"
|
||||
echo " Log: $LOG_FILE ($(wc -l <"$LOG_FILE") lines)"
|
||||
echo ""
|
||||
|
||||
# --- Triage report ---
|
||||
echo "=== Sanitizer error summary ==="
|
||||
grep -E "ERROR SUMMARY|Internal Sanitizer Error" "$LOG_FILE" | tail -5 || echo " (no summary line)"
|
||||
echo ""
|
||||
|
||||
# Categorise errors. compute-sanitizer prints "ERROR SUMMARY: N errors" at the
|
||||
# end. "Internal Sanitizer Error" is sanitizer-internal (allocation failure,
|
||||
# tracking gap) — NOT a real bug in the application code.
|
||||
INTERNAL_ERRORS=$(grep -c "Internal Sanitizer Error" "$LOG_FILE" || true)
|
||||
REAL_ERROR_LINE=$(grep "ERROR SUMMARY:" "$LOG_FILE" | tail -1 || echo "")
|
||||
REAL_ERRORS=$(echo "$REAL_ERROR_LINE" | grep -oE "[0-9]+ errors" | head -1 | grep -oE "[0-9]+" || echo "0")
|
||||
|
||||
echo "=== Triage ==="
|
||||
echo " Internal Sanitizer Errors (instrumentation gaps): $INTERNAL_ERRORS"
|
||||
echo " Reported errors total: $REAL_ERRORS"
|
||||
echo " Real (non-internal) errors: $((REAL_ERRORS - INTERNAL_ERRORS))"
|
||||
echo ""
|
||||
|
||||
# Show first 50 real (non-internal) errors with context
|
||||
echo "=== First non-internal error excerpts ==="
|
||||
grep -vE "Internal Sanitizer Error|^=========\s*$" "$LOG_FILE" \
|
||||
| grep -E "^=========" \
|
||||
| head -50 || echo " (none)"
|
||||
|
||||
# Final pass/fail
|
||||
if [ "$EXIT_CODE" -eq 99 ]; then
|
||||
echo ""
|
||||
echo "FAIL: compute-sanitizer detected real memory errors (--error-exitcode 99 fired)."
|
||||
exit 1
|
||||
elif [ "$EXIT_CODE" -ne 0 ]; then
|
||||
echo ""
|
||||
echo "FAIL: test process exited with code $EXIT_CODE (test failure or sanitizer abort)."
|
||||
exit "$EXIT_CODE"
|
||||
else
|
||||
echo ""
|
||||
echo "PASS: zero real memory errors detected, test passed."
|
||||
fi
|
||||
@@ -1,32 +0,0 @@
|
||||
---
|
||||
apiVersion: v1
|
||||
kind: PersistentVolumeClaim
|
||||
metadata:
|
||||
name: sccache-cpu
|
||||
namespace: foxhunt
|
||||
labels:
|
||||
app.kubernetes.io/name: sccache
|
||||
app.kubernetes.io/component: ci-cache
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
spec:
|
||||
accessModes: [ReadWriteOnce]
|
||||
storageClassName: scw-bssd-retain
|
||||
resources:
|
||||
requests:
|
||||
storage: 20Gi
|
||||
---
|
||||
apiVersion: v1
|
||||
kind: PersistentVolumeClaim
|
||||
metadata:
|
||||
name: sccache-cuda
|
||||
namespace: foxhunt
|
||||
labels:
|
||||
app.kubernetes.io/name: sccache
|
||||
app.kubernetes.io/component: ci-cache
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
spec:
|
||||
accessModes: [ReadWriteOnce]
|
||||
storageClassName: scw-bssd-retain
|
||||
resources:
|
||||
requests:
|
||||
storage: 20Gi
|
||||
@@ -1,308 +0,0 @@
|
||||
# infra/k8s/argo/smoke-test-template.yaml
|
||||
#
|
||||
# Plain smoke-test runner on L40S — no compute-sanitizer, no nsys profile,
|
||||
# no MinIO artefact upload. Mirrors the compile + checkout + GPU + data-mount
|
||||
# layout of nsys-test-template.yaml but executes the test binary directly.
|
||||
#
|
||||
# Use this for fast multi-fold / single-epoch validation runs against the
|
||||
# full training-data PVC (27 months) when local laptop data is too short
|
||||
# (laptop only has 1 quarter of MBP-10/trades).
|
||||
#
|
||||
# Why a separate template: the nsys/sanitizer wrappers add overhead and
|
||||
# extract artefacts not needed for a straight "did the test pass on the
|
||||
# real dataset" gate. Keeping the smoke variant lean keeps the iteration
|
||||
# cycle short (no .nsys-rep upload, no sanitizer instrumentation slowdown).
|
||||
#
|
||||
# Usage:
|
||||
# argo submit --watch -n foxhunt --from=wftmpl/smoke-test \
|
||||
# -p commit-ref=main \
|
||||
# -p test-name=multi_fold_convergence::test_multi_fold_convergence
|
||||
#
|
||||
# Or via the wrapper script: scripts/argo-smoke.sh
|
||||
|
||||
apiVersion: argoproj.io/v1alpha1
|
||||
kind: WorkflowTemplate
|
||||
metadata:
|
||||
name: smoke-test
|
||||
namespace: foxhunt
|
||||
labels:
|
||||
app.kubernetes.io/name: smoke-test
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
spec:
|
||||
entrypoint: smoke-run
|
||||
serviceAccountName: argo-workflow
|
||||
podMetadata:
|
||||
labels:
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
app.kubernetes.io/component: smoke-test
|
||||
archiveLogs: true
|
||||
podGC:
|
||||
strategy: OnPodCompletion
|
||||
ttlStrategy:
|
||||
secondsAfterCompletion: 7200
|
||||
activeDeadlineSeconds: 3600
|
||||
arguments:
|
||||
parameters:
|
||||
- name: commit-ref
|
||||
value: HEAD
|
||||
- name: gpu-pool
|
||||
value: ci-training-l40s
|
||||
- name: cuda-compute-cap
|
||||
value: "89"
|
||||
- name: test-name
|
||||
value: "multi_fold_convergence::test_multi_fold_convergence"
|
||||
# Wipe `/cargo-target/release` for the ml + ml-dqn crates before
|
||||
# compile so the smoke binary is built from a known-clean state. The
|
||||
# PVC's persistent target dir accumulates rmeta/object artefacts
|
||||
# across probes; file deletions (e.g., `regime_conditional.rs` in
|
||||
# ff00af68a) can leave dangling references that perturb downstream
|
||||
# codegen between bisect runs. Defaults to "false" — only set when
|
||||
# bisecting suspected build-cache contamination.
|
||||
- name: clean-cache
|
||||
value: "false"
|
||||
|
||||
volumes:
|
||||
- name: git-ssh-key
|
||||
secret:
|
||||
secretName: argo-git-ssh-key
|
||||
defaultMode: 256
|
||||
- name: cargo-target
|
||||
persistentVolumeClaim:
|
||||
claimName: cargo-target-cuda-test
|
||||
# Same training-data-pvc as production training (full 27 months of
|
||||
# OHLCV + MBP-10 + trades) — laptop's 1-quarter MBP-10 truncates the
|
||||
# fxcache below the smoke's 10-month minimum.
|
||||
#
|
||||
# RW (not readOnly) so the smoke can populate `/data/bin/$SHORT_SHA/`
|
||||
# with the train binaries it just built, letting the next train run's
|
||||
# `ensure-binary` cache check (train-multi-seed-template.yaml:235-246)
|
||||
# hit and skip the 6-min compile. Read-side (fxcache + raw market data)
|
||||
# is unchanged.
|
||||
- name: training-data
|
||||
persistentVolumeClaim:
|
||||
claimName: training-data-pvc
|
||||
|
||||
templates:
|
||||
- name: smoke-run
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: "{{workflow.parameters.gpu-pool}}"
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
tolerations:
|
||||
- key: nvidia.com/gpu
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
- key: node.cilium.io/agent-not-ready
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
container:
|
||||
image: gitlab-registry.foxhunt.svc.cluster.local:5000/root/foxhunt/ci-builder:latest
|
||||
command: ["/bin/bash", "-c"]
|
||||
env:
|
||||
- name: SQLX_OFFLINE
|
||||
value: "true"
|
||||
- name: CARGO_TERM_COLOR
|
||||
value: always
|
||||
- name: CARGO_TARGET_DIR
|
||||
value: /cargo-target
|
||||
- name: CARGO_HOME
|
||||
value: /cargo-target/cargo-home
|
||||
- name: CUDA_VISIBLE_DEVICES
|
||||
value: "0"
|
||||
- name: CUBLAS_WORKSPACE_CONFIG
|
||||
value: ":4096:8"
|
||||
- name: FOXHUNT_TEST_DATA
|
||||
value: /data/futures-baseline
|
||||
- name: TEST_DATA_DIR
|
||||
value: /data/futures-baseline
|
||||
# MBP-10 + trades are mandatory per feedback_mbp10_mandatory.md
|
||||
# (OFI features are part of the production state vector).
|
||||
- name: FOXHUNT_MBP10_DATA
|
||||
value: /data/futures-baseline-mbp10
|
||||
- name: FOXHUNT_TRADES_DATA
|
||||
value: /data/futures-baseline-trades
|
||||
- name: RUST_LOG
|
||||
value: info
|
||||
- name: LD_LIBRARY_PATH
|
||||
value: /usr/local/nvidia/lib64:/usr/local/cuda/lib64
|
||||
volumeMounts:
|
||||
- name: git-ssh-key
|
||||
mountPath: /etc/git-ssh
|
||||
readOnly: true
|
||||
- name: cargo-target
|
||||
mountPath: /cargo-target
|
||||
- name: training-data
|
||||
mountPath: /data
|
||||
# RW so the post-PASS cache-population step can write to
|
||||
# /data/bin/$SHORT_SHA/. Read-side paths (fxcache, MBP-10,
|
||||
# trades) are unaffected.
|
||||
resources:
|
||||
requests:
|
||||
nvidia.com/gpu: "1"
|
||||
cpu: "4"
|
||||
memory: 32Gi
|
||||
limits:
|
||||
nvidia.com/gpu: "1"
|
||||
cpu: "8"
|
||||
memory: 64Gi
|
||||
args:
|
||||
- |
|
||||
set -euo pipefail
|
||||
|
||||
REF="{{workflow.parameters.commit-ref}}"
|
||||
TEST_NAME="{{workflow.parameters.test-name}}"
|
||||
|
||||
echo "==================================="
|
||||
echo " Plain smoke L40S run"
|
||||
echo "==================================="
|
||||
echo " Ref: $REF"
|
||||
echo " Test: $TEST_NAME"
|
||||
echo " Pool: {{workflow.parameters.gpu-pool}}"
|
||||
echo "==================================="
|
||||
|
||||
# --- SSH setup ---
|
||||
mkdir -p ~/.ssh
|
||||
cp /etc/git-ssh/ssh-privatekey ~/.ssh/id_ed25519
|
||||
chmod 600 ~/.ssh/id_ed25519
|
||||
printf 'Host *\n StrictHostKeyChecking no\n UserKnownHostsFile /dev/null\n' > ~/.ssh/config
|
||||
chmod 600 ~/.ssh/config
|
||||
|
||||
REPO="ssh://git@gitlab-gitlab-shell.foxhunt.svc.cluster.local:2222/root/foxhunt.git"
|
||||
BUILD="/cargo-target/src"
|
||||
git config --global --add safe.directory "$BUILD"
|
||||
|
||||
# --- Persistent checkout on PVC ---
|
||||
if [ -d "$BUILD/.git" ]; then
|
||||
cd "$BUILD"
|
||||
git fetch origin
|
||||
if git rev-parse --verify "origin/$REF" >/dev/null 2>&1; then
|
||||
TARGET=$(git rev-parse "origin/$REF")
|
||||
else
|
||||
TARGET=$(git rev-parse --verify "$REF" 2>/dev/null || echo "$REF")
|
||||
fi
|
||||
CURRENT=$(git rev-parse HEAD 2>/dev/null || echo "none")
|
||||
if [ "$CURRENT" != "$TARGET" ]; then
|
||||
echo "=== Updating checkout to $REF ($TARGET) ==="
|
||||
git checkout --force --detach "$TARGET"
|
||||
git clean -fd
|
||||
fi
|
||||
else
|
||||
git clone --filter=blob:none "$REPO" "$BUILD"
|
||||
cd "$BUILD"
|
||||
git fetch origin
|
||||
if git rev-parse --verify "origin/$REF" >/dev/null 2>&1; then
|
||||
git checkout --force --detach "origin/$REF"
|
||||
else
|
||||
git checkout "$REF"
|
||||
fi
|
||||
fi
|
||||
SHA=$(git rev-parse --short HEAD)
|
||||
echo "Checked out $SHA"
|
||||
|
||||
export PATH="${CARGO_HOME}/bin:/usr/local/cuda/bin:${PATH}"
|
||||
export CUDA_COMPUTE_CAP={{workflow.parameters.cuda-compute-cap}}
|
||||
|
||||
bash scripts/ptx-cache-invalidate.sh "${CARGO_TARGET_DIR}/.ptx_cache"
|
||||
|
||||
# --- Optional clean: ml + ml-dqn before compile ---
|
||||
CLEAN_CACHE="{{workflow.parameters.clean-cache}}"
|
||||
if [ "$CLEAN_CACHE" = "true" ]; then
|
||||
echo "=== clean-cache=true — wiping ml + ml-dqn build artefacts ==="
|
||||
cargo clean -p ml --release 2>&1 | tee /cargo-target/smoke-clean.log
|
||||
cargo clean -p ml-dqn --release 2>&1 | tee -a /cargo-target/smoke-clean.log
|
||||
echo "=== Cleaned. Forced fresh compile of ml/ml-dqn (deps stay cached) ==="
|
||||
fi
|
||||
|
||||
# --- Compile test binary + the 3 train binaries ---
|
||||
#
|
||||
# The 3 examples mirror exactly what `ensure-binary` produces
|
||||
# (see train-multi-seed-template.yaml line ~268). Compiling
|
||||
# them here lets the post-PASS step populate the
|
||||
# `/data/bin/$SHORT_SHA/` cache so the next train run at the
|
||||
# same SHA skips its own ensure-binary compile.
|
||||
echo "=== Compiling (--release --no-run + train binaries) ==="
|
||||
cargo test -p ml --release --lib --no-run 2>&1 | tee /cargo-target/smoke-compile.log
|
||||
cargo build -p ml --release \
|
||||
--example train_baseline_rl \
|
||||
--example evaluate_baseline \
|
||||
--example precompute_features \
|
||||
2>&1 | tee -a /cargo-target/smoke-compile.log
|
||||
|
||||
# Honour CARGO_TARGET_DIR — outputs go to ${CARGO_TARGET_DIR}/release/deps,
|
||||
# NOT the repo-relative target/. Subshell + || true wraps the head -1
|
||||
# SIGPIPE so pipefail doesn't fire.
|
||||
TARGET_DIR="${CARGO_TARGET_DIR:-target}"
|
||||
TEST_BIN=$( (ls -t "$TARGET_DIR/release/deps/ml-"* 2>/dev/null \
|
||||
| grep -vE '\.(d|rcgu|rmeta|rlib|so|json)$' \
|
||||
| head -1) || true )
|
||||
if [ -z "$TEST_BIN" ] || [ ! -x "$TEST_BIN" ]; then
|
||||
echo "ERROR: could not locate compiled test binary in $TARGET_DIR/release/deps/"
|
||||
ls -la "$TARGET_DIR/release/deps/" 2>&1 | head -20
|
||||
exit 3
|
||||
fi
|
||||
echo "=== Test binary: $TEST_BIN ==="
|
||||
|
||||
# grep -q + pipefail = SIGPIPE on early exit. Capture once, grep
|
||||
# without -q so the consumer reads all input.
|
||||
TEST_LIST=$("$TEST_BIN" --list 2>&1 || true)
|
||||
if ! echo "$TEST_LIST" | grep -F "$TEST_NAME" > /dev/null; then
|
||||
echo "ERROR: test '$TEST_NAME' not found"
|
||||
echo "$TEST_LIST" | grep smoke_tests | (head -30 || true)
|
||||
exit 4
|
||||
fi
|
||||
|
||||
# --- Run the test ---
|
||||
echo ""
|
||||
echo "==================================="
|
||||
echo " Launching test..."
|
||||
echo "==================================="
|
||||
START_TS=$(date +%s)
|
||||
set +e
|
||||
"$TEST_BIN" "$TEST_NAME" --ignored --nocapture
|
||||
EXIT_CODE=$?
|
||||
set -e
|
||||
END_TS=$(date +%s)
|
||||
ELAPSED=$((END_TS - START_TS))
|
||||
|
||||
echo ""
|
||||
echo "==================================="
|
||||
echo " Smoke run complete"
|
||||
echo "==================================="
|
||||
echo " Exit code: $EXIT_CODE"
|
||||
echo " Elapsed: ${ELAPSED}s"
|
||||
echo " Commit: $SHA"
|
||||
echo " Test: $TEST_NAME"
|
||||
echo "==================================="
|
||||
|
||||
if [ "$EXIT_CODE" -ne 0 ]; then
|
||||
echo ""
|
||||
echo "FAIL: test exited with code $EXIT_CODE"
|
||||
exit "$EXIT_CODE"
|
||||
fi
|
||||
echo ""
|
||||
echo "PASS: smoke completed."
|
||||
|
||||
# --- Populate /data/bin/$SHORT_SHA/ for ensure-binary cache ---
|
||||
# On PASS only — broken binaries must not be cached. Mirrors
|
||||
# train-multi-seed-template.yaml:270-274 exactly: same
|
||||
# destination layout, same strip step. Idempotent: writes are
|
||||
# to a SHA-keyed dir so concurrent smokes at different SHAs
|
||||
# don't collide; same-SHA writes are deterministic compile
|
||||
# output and overwrite-safe.
|
||||
FULL_SHA=$(git rev-parse HEAD)
|
||||
SHORT_SHA=$(echo "$FULL_SHA" | cut -c1-9)
|
||||
BIN_DIR="/data/bin/$SHORT_SHA"
|
||||
BINARIES="train_baseline_rl evaluate_baseline precompute_features"
|
||||
echo ""
|
||||
echo "=== Populating ensure-binary cache at $BIN_DIR ==="
|
||||
mkdir -p "$BIN_DIR"
|
||||
for bin in $BINARIES; do
|
||||
src="${CARGO_TARGET_DIR}/release/examples/$bin"
|
||||
if [ ! -x "$src" ]; then
|
||||
echo "WARN: $src not found (cache pop skipped for $bin)"
|
||||
continue
|
||||
fi
|
||||
cp "$src" "$BIN_DIR/"
|
||||
strip "$BIN_DIR/$bin" 2>/dev/null || true
|
||||
done
|
||||
ls -lh "$BIN_DIR/"
|
||||
echo "=== Cache populated; next train@$SHORT_SHA hits ensure-binary cache ==="
|
||||
@@ -1,605 +0,0 @@
|
||||
# Multi-seed training workflow — Plan 5 Task 5 Phase B (one-job-per-seed).
|
||||
#
|
||||
# Renders an Argo DAG that fans out N seeds into N parallel `train-single` task
|
||||
# instances. Each task receives `seed` via inputs.parameters and runs a
|
||||
# walk-forward training that internally sweeps all K folds via the binary's
|
||||
# `--max-folds K` arg. Per-job runtime is K× longer than the original (seed,
|
||||
# fold) matrix but fanout drops from N*K to N — a better fit for the L40S pool
|
||||
# (5-GPU capacity vs 30 jobs queueing) and a simpler binary contract
|
||||
# (`train_baseline_rl` is a multi-fold walk-forward executor; it does NOT
|
||||
# accept `--fold K`).
|
||||
#
|
||||
# The `# __MATRIX_TASKS__` marker on the dag.tasks line is replaced by
|
||||
# scripts/argo-train.sh with N generated WorkflowTask stanzas before
|
||||
# submission. This avoids hand-writing the matrix and keeps the template
|
||||
# human-readable.
|
||||
#
|
||||
# Usage:
|
||||
# ./scripts/argo-train.sh dqn --multi-seed 5 --folds 6 --tag plan5-final
|
||||
#
|
||||
# DAG (per task):
|
||||
# ensure-binary ──┐
|
||||
# gpu-warmup ─────┼──> ensure-fxcache ──> [N parallel train-single tasks,
|
||||
# one per seed, each runs all K folds]
|
||||
# │
|
||||
# └──> aggregate (manual via
|
||||
# scripts/gather-multi-seed-metrics.sh)
|
||||
---
|
||||
apiVersion: argoproj.io/v1alpha1
|
||||
kind: WorkflowTemplate
|
||||
metadata:
|
||||
name: train-multi-seed
|
||||
namespace: foxhunt
|
||||
labels:
|
||||
app.kubernetes.io/name: train-multi-seed
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
app.kubernetes.io/component: train
|
||||
spec:
|
||||
entrypoint: multi-seed-matrix
|
||||
onExit: notify-result
|
||||
serviceAccountName: argo-workflow
|
||||
archiveLogs: true
|
||||
podMetadata:
|
||||
labels:
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
app.kubernetes.io/component: train
|
||||
securityContext:
|
||||
fsGroup: 0
|
||||
ttlStrategy:
|
||||
secondsAfterCompletion: 3600
|
||||
# Multi-seed runs: N jobs in parallel (one per seed), each running all K folds
|
||||
# in walk-forward sequence. Allow 12h walltime — 6-fold runs are ~6× longer
|
||||
# than the original per-(seed,fold) jobs but easily fit in 12h on L40S.
|
||||
activeDeadlineSeconds: 43200
|
||||
|
||||
arguments:
|
||||
parameters:
|
||||
- name: commit-sha
|
||||
value: HEAD
|
||||
- name: git-branch
|
||||
value: main
|
||||
# Default sm_89 / ci-training-l40s as of 2026-05-14 per
|
||||
# `feedback_default_to_l40s_pool.md`. argo-train.sh's
|
||||
# CUDA_COMPUTE_CAP derivation passes "89" by default to match;
|
||||
# explicit --gpu-pool ci-training-h100 overrides both back to
|
||||
# sm_90 + H100.
|
||||
- name: cuda-compute-cap
|
||||
value: "89"
|
||||
- name: model
|
||||
value: dqn
|
||||
- name: gpu-pool
|
||||
value: ci-training-l40s
|
||||
- name: hyperopt-trials
|
||||
value: "0" # Multi-seed runs typically skip hyperopt (re-use baseline params)
|
||||
- name: hyperopt-epochs
|
||||
value: "8"
|
||||
- name: train-epochs
|
||||
value: "50"
|
||||
- name: symbol
|
||||
value: ES.FUT
|
||||
- name: initial-capital
|
||||
value: "35000"
|
||||
- name: tx-cost-bps
|
||||
value: "0.1"
|
||||
- name: tick-size
|
||||
value: "0.25"
|
||||
- name: spread-ticks
|
||||
value: "1.0"
|
||||
- name: sanitizer
|
||||
value: "none"
|
||||
- name: multi-seed
|
||||
value: "5"
|
||||
- name: folds
|
||||
value: "6"
|
||||
# Plan 5 Task 3 (A.4.1): nsys profile harness toggle. When "true", each
|
||||
# train-single job wraps the training binary under `nsys profile` and
|
||||
# uploads the resulting .nsys-rep to MinIO bucket
|
||||
# foxhunt-training-artifacts/profiles/<short-sha>/. Default off.
|
||||
- name: profile
|
||||
value: "false"
|
||||
# Bar formation params — passed to BOTH precompute_features (when
|
||||
# building fxcache) AND train_baseline_rl (when looking up fxcache).
|
||||
# MUST match between the two for fxcache HIT (cache key includes them
|
||||
# as of 2026-05-09 architectural fix). Defaults match dqn-production.toml.
|
||||
#
|
||||
# imbalance-bar-threshold default: 20.0 (set 2026-05-10). At threshold=0.5
|
||||
# the sampler produced 209M bars from 209M trade ticks (1:1 ratio,
|
||||
# essentially per-tick) on workflow f5wnd, near-OOM in feature extraction
|
||||
# (54Gi/56Gi limit). To match the volume-bar density baseline (~5.74M
|
||||
# bars at 100 contracts), threshold ~20 is the right scale. Override via
|
||||
# argo-train.sh --imbalance-bar-threshold for resolution-sweep experiments.
|
||||
- name: imbalance-bar-threshold
|
||||
value: "20.0"
|
||||
- name: imbalance-bar-ewma-alpha
|
||||
value: "0.1"
|
||||
# Volume bar size (contracts/bar). Used when data-source != "mbp10".
|
||||
# Default 100 matches DEFAULT_VOLUME_BAR_SIZE. Cache key includes this.
|
||||
- name: volume-bar-size
|
||||
value: "100"
|
||||
# Data source mode: "mbp10" (imbalance bars from MBP-10) or "ohlcv"
|
||||
# (volume bars from trades). Threaded into both precompute_features and
|
||||
# the trainer so cache keys align.
|
||||
- name: data-source
|
||||
value: "mbp10"
|
||||
|
||||
volumes:
|
||||
- name: git-ssh-key
|
||||
secret:
|
||||
secretName: argo-git-ssh-key
|
||||
defaultMode: 256
|
||||
- name: training-data
|
||||
persistentVolumeClaim:
|
||||
claimName: training-data-pvc
|
||||
- name: cargo-target-cuda
|
||||
persistentVolumeClaim:
|
||||
claimName: cargo-target-cuda
|
||||
- name: feature-cache
|
||||
persistentVolumeClaim:
|
||||
claimName: feature-cache-pvc
|
||||
|
||||
volumeClaimTemplates:
|
||||
- metadata:
|
||||
name: workspace
|
||||
spec:
|
||||
accessModes: ["ReadWriteOnce"]
|
||||
storageClassName: scw-bssd
|
||||
resources:
|
||||
requests:
|
||||
storage: 5Gi
|
||||
|
||||
templates:
|
||||
# ── DAG: fan out to N train-single tasks (one per seed) ──
|
||||
# The `# __MATRIX_TASKS__` marker is replaced by argo-train.sh with the
|
||||
# generated per-seed WorkflowTask stanzas. Each task sweeps all K folds
|
||||
# via the binary's `--max-folds {{workflow.parameters.folds}}` argument.
|
||||
# The marker MUST stay on its own line for the awk substitution to work.
|
||||
- name: multi-seed-matrix
|
||||
dag:
|
||||
tasks:
|
||||
- name: ensure-binary
|
||||
template: ensure-binary
|
||||
- name: gpu-warmup
|
||||
template: gpu-warmup
|
||||
- name: ensure-fxcache
|
||||
template: ensure-fxcache
|
||||
dependencies: [ensure-binary]
|
||||
arguments:
|
||||
parameters:
|
||||
- name: sha
|
||||
value: "{{tasks.ensure-binary.outputs.parameters.sha}}"
|
||||
# __MATRIX_TASKS__
|
||||
|
||||
# ── ensure-binary: cache-or-compile training binaries by commit SHA ──
|
||||
# (Identical to train-template.yaml. Kept inline rather than templated
|
||||
# via wftmpl-cross-ref to avoid Argo's reluctance to chase template
|
||||
# references across WorkflowTemplates at submission time.)
|
||||
- name: ensure-binary
|
||||
outputs:
|
||||
parameters:
|
||||
- name: sha
|
||||
valueFrom:
|
||||
path: /tmp/sha
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: ci-compile-cpu
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
tolerations:
|
||||
- key: node.cilium.io/agent-not-ready
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
container:
|
||||
image: gitlab-registry.foxhunt.svc.cluster.local:5000/root/foxhunt/ci-builder:latest
|
||||
imagePullPolicy: Always
|
||||
command: ["/bin/bash", "-c"]
|
||||
env:
|
||||
- name: SQLX_OFFLINE
|
||||
value: "true"
|
||||
- name: CARGO_TERM_COLOR
|
||||
value: always
|
||||
- name: CARGO_TARGET_DIR
|
||||
value: /cargo-target
|
||||
- name: CARGO_HOME
|
||||
value: /cargo-target/cargo-home
|
||||
- name: CARGO_INCREMENTAL
|
||||
value: "1"
|
||||
- name: GITLAB_PAT
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: gitlab-pat
|
||||
key: token
|
||||
resources:
|
||||
requests:
|
||||
cpu: "14"
|
||||
memory: 32Gi
|
||||
limits:
|
||||
cpu: "30"
|
||||
memory: 64Gi
|
||||
volumeMounts:
|
||||
- name: git-ssh-key
|
||||
mountPath: /etc/git-ssh
|
||||
readOnly: true
|
||||
- name: cargo-target-cuda
|
||||
mountPath: /cargo-target
|
||||
- name: training-data
|
||||
mountPath: /data
|
||||
args:
|
||||
- |
|
||||
set -e
|
||||
SHA="{{workflow.parameters.commit-sha}}"
|
||||
BRANCH="{{workflow.parameters.git-branch}}"
|
||||
MODEL="{{workflow.parameters.model}}"
|
||||
|
||||
mkdir -p ~/.ssh
|
||||
cp /etc/git-ssh/ssh-privatekey ~/.ssh/id_ed25519
|
||||
chmod 600 ~/.ssh/id_ed25519
|
||||
printf 'Host *\n StrictHostKeyChecking no\n UserKnownHostsFile /dev/null\n' > ~/.ssh/config
|
||||
chmod 600 ~/.ssh/config
|
||||
|
||||
REPO="ssh://git@gitlab-gitlab-shell.foxhunt.svc.cluster.local:2222/root/foxhunt.git"
|
||||
BUILD="/cargo-target/src"
|
||||
git config --global --add safe.directory "$BUILD"
|
||||
|
||||
if [ "$SHA" = "HEAD" ]; then
|
||||
if [ -d "$BUILD/.git" ]; then
|
||||
cd "$BUILD"; git fetch origin
|
||||
SHA=$(git rev-parse "origin/$BRANCH"); cd /
|
||||
else
|
||||
git clone --filter=blob:none "$REPO" "$BUILD"
|
||||
cd "$BUILD"; git checkout "origin/$BRANCH"
|
||||
SHA=$(git rev-parse HEAD); cd /
|
||||
fi
|
||||
fi
|
||||
SHORT_SHA=$(echo "$SHA" | cut -c1-9)
|
||||
echo "Resolved SHA: $SHA (short: $SHORT_SHA)"
|
||||
|
||||
case "$MODEL" in
|
||||
alpha-rl) BINARIES="alpha_rl_train" ;;
|
||||
dqn|ppo) BINARIES="train_baseline_rl evaluate_baseline precompute_features" ;;
|
||||
*) BINARIES="train_baseline_supervised evaluate_supervised precompute_features" ;;
|
||||
esac
|
||||
|
||||
BIN_DIR="/data/bin/$SHORT_SHA"
|
||||
ALL_CACHED=true
|
||||
for bin in $BINARIES; do
|
||||
if [ ! -x "$BIN_DIR/$bin" ]; then ALL_CACHED=false; break; fi
|
||||
done
|
||||
|
||||
if [ "$ALL_CACHED" = "true" ]; then
|
||||
echo "=== Cache HIT: all binaries present in $BIN_DIR ==="
|
||||
ls -lh "$BIN_DIR/"
|
||||
echo "$SHORT_SHA" > /tmp/sha
|
||||
exit 0
|
||||
fi
|
||||
|
||||
echo "=== Cache MISS: compiling binaries for $SHORT_SHA ==="
|
||||
if [ -d "$BUILD/.git" ]; then
|
||||
cd "$BUILD"; git fetch origin
|
||||
CURRENT=$(git rev-parse HEAD 2>/dev/null || echo "none")
|
||||
if [ "$CURRENT" != "$SHA" ]; then
|
||||
git checkout --force "$SHA"; git clean -fd
|
||||
fi
|
||||
else
|
||||
git clone --filter=blob:none "$REPO" "$BUILD"
|
||||
cd "$BUILD"; git checkout "$SHA"
|
||||
fi
|
||||
|
||||
export PATH="${CARGO_HOME}/bin:${PATH}"
|
||||
export CUDA_COMPUTE_CAP={{workflow.parameters.cuda-compute-cap}}
|
||||
|
||||
if [ "$MODEL" = "alpha-rl" ]; then
|
||||
cargo build --release -p ml-alpha --example alpha_rl_train
|
||||
else
|
||||
ML_EXAMPLE_ARGS=""
|
||||
for ex in $BINARIES; do
|
||||
ML_EXAMPLE_ARGS="$ML_EXAMPLE_ARGS --example $ex"
|
||||
done
|
||||
cargo build --release -p ml --features ml/cuda $ML_EXAMPLE_ARGS
|
||||
fi
|
||||
|
||||
mkdir -p "$BIN_DIR"
|
||||
for bin in $BINARIES; do
|
||||
cp "$CARGO_TARGET_DIR/release/examples/$bin" "$BIN_DIR/"
|
||||
done
|
||||
strip "$BIN_DIR/"*
|
||||
echo "$SHORT_SHA" > /tmp/sha
|
||||
|
||||
# ── gpu-warmup: trigger GPU node autoscale during compile ──
|
||||
- name: gpu-warmup
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: "{{workflow.parameters.gpu-pool}}"
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
tolerations:
|
||||
- key: nvidia.com/gpu
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
- key: node.cilium.io/agent-not-ready
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
container:
|
||||
image: busybox:1.37
|
||||
command: ["/bin/sh", "-c"]
|
||||
args:
|
||||
- |
|
||||
echo "GPU warmup: triggering node autoscale..."
|
||||
resources:
|
||||
requests:
|
||||
nvidia.com/gpu: "1"
|
||||
cpu: 100m
|
||||
memory: 64Mi
|
||||
limits:
|
||||
nvidia.com/gpu: "1"
|
||||
cpu: 200m
|
||||
memory: 128Mi
|
||||
|
||||
# ── ensure-fxcache: precompute feature cache (shared across all seeds/folds) ──
|
||||
- name: ensure-fxcache
|
||||
inputs:
|
||||
parameters:
|
||||
- name: sha
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: ci-compile-cpu
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
tolerations:
|
||||
- key: node.cilium.io/agent-not-ready
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
container:
|
||||
image: ubuntu:24.04
|
||||
command: ["/bin/bash", "-c"]
|
||||
env:
|
||||
- name: RUST_LOG
|
||||
value: info
|
||||
resources:
|
||||
requests:
|
||||
cpu: "4"
|
||||
memory: 32Gi
|
||||
limits:
|
||||
cpu: "28"
|
||||
memory: 96Gi
|
||||
volumeMounts:
|
||||
- name: training-data
|
||||
mountPath: /data
|
||||
readOnly: false
|
||||
- name: feature-cache
|
||||
mountPath: /feature-cache
|
||||
args:
|
||||
- |
|
||||
set -e
|
||||
SHA="{{inputs.parameters.sha}}"
|
||||
BINARY="/data/bin/$SHA/precompute_features"
|
||||
if [ ! -x "$BINARY" ]; then
|
||||
echo "ERROR: precompute_features not found at $BINARY"; exit 1
|
||||
fi
|
||||
export RAYON_NUM_THREADS=20
|
||||
if $BINARY \
|
||||
--data-dir /data/futures-baseline \
|
||||
--mbp10-data-dir /data/futures-baseline-mbp10 \
|
||||
--trades-data-dir /data/futures-baseline-trades \
|
||||
--output-dir /feature-cache \
|
||||
--symbol {{workflow.parameters.symbol}} \
|
||||
--data-source {{workflow.parameters.data-source}} \
|
||||
--imbalance-bar-threshold {{workflow.parameters.imbalance-bar-threshold}} \
|
||||
--imbalance-bar-ewma-alpha {{workflow.parameters.imbalance-bar-ewma-alpha}} \
|
||||
--volume-bar-size {{workflow.parameters.volume-bar-size}} \
|
||||
--yes; then
|
||||
echo "=== Feature cache ready ==="
|
||||
else
|
||||
echo "=== Cache stale or missing — regenerating ==="
|
||||
# NOTE: don't blanket-rm /feature-cache/*.fxcache anymore. With
|
||||
# bar-params now in the cache key, multiple valid caches can
|
||||
# coexist (different threshold experiments). The early-exit
|
||||
# check above already validates the specific key — fall through
|
||||
# to regen only the one we need.
|
||||
$BINARY \
|
||||
--data-dir /data/futures-baseline \
|
||||
--mbp10-data-dir /data/futures-baseline-mbp10 \
|
||||
--trades-data-dir /data/futures-baseline-trades \
|
||||
--output-dir /feature-cache \
|
||||
--symbol {{workflow.parameters.symbol}} \
|
||||
--data-source {{workflow.parameters.data-source}} \
|
||||
--imbalance-bar-threshold {{workflow.parameters.imbalance-bar-threshold}} \
|
||||
--imbalance-bar-ewma-alpha {{workflow.parameters.imbalance-bar-ewma-alpha}} \
|
||||
--volume-bar-size {{workflow.parameters.volume-bar-size}} \
|
||||
--yes
|
||||
fi
|
||||
|
||||
# ── train-single: one seed, all folds (training instance) ──
|
||||
# Invoked once per matrix entry. Reads SEED from inputs.parameters and
|
||||
# forwards it via `--seed`; the binary's `--max-folds K` arg drives the
|
||||
# walk-forward sweep over all K folds inside this single process.
|
||||
# Plan 5 Task 5 Phase B pivot: was per-(seed, fold) on the failed deploy
|
||||
# because train_baseline_rl does not accept `--fold N` (it is a multi-fold
|
||||
# executor, not a single-fold one). One-job-per-seed matches the binary's
|
||||
# actual contract and the L40S pool's capacity.
|
||||
- name: train-single
|
||||
inputs:
|
||||
parameters:
|
||||
- name: seed
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: "{{workflow.parameters.gpu-pool}}"
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
tolerations:
|
||||
- key: nvidia.com/gpu
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
- key: node.cilium.io/agent-not-ready
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
container:
|
||||
image: gitlab-registry.foxhunt.svc.cluster.local:5000/root/foxhunt/ci-builder:latest
|
||||
imagePullPolicy: IfNotPresent
|
||||
command: ["/bin/sh", "-c"]
|
||||
env:
|
||||
- name: RUST_LOG
|
||||
value: info
|
||||
- name: SQLX_OFFLINE
|
||||
value: "true"
|
||||
- name: CUBLAS_WORKSPACE_CONFIG
|
||||
value: ":4096:8"
|
||||
- name: FOXHUNT_FEATURE_CACHE_DIR
|
||||
value: /feature-cache
|
||||
- name: SEED
|
||||
value: "{{inputs.parameters.seed}}"
|
||||
# Plan 5 Task 3 (A.4.1): MinIO creds for the optional `mc cp` of
|
||||
# the .nsys-rep artefact at the end of train-single. Both refs are
|
||||
# `optional: true` so the env mount succeeds on clusters that do
|
||||
# not pre-create `minio-credentials` (the upload then warn-fails
|
||||
# gracefully — training itself doesn't depend on these).
|
||||
- name: MINIO_ACCESS_KEY
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: minio-credentials
|
||||
key: access-key
|
||||
optional: true
|
||||
- name: MINIO_SECRET_KEY
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: minio-credentials
|
||||
key: secret-key
|
||||
optional: true
|
||||
resources:
|
||||
requests:
|
||||
nvidia.com/gpu: "1"
|
||||
cpu: "4"
|
||||
memory: 32Gi
|
||||
limits:
|
||||
nvidia.com/gpu: "1"
|
||||
cpu: "8"
|
||||
memory: 64Gi
|
||||
volumeMounts:
|
||||
- name: workspace
|
||||
mountPath: /workspace
|
||||
- name: training-data
|
||||
mountPath: /data
|
||||
readOnly: true
|
||||
- name: feature-cache
|
||||
mountPath: /feature-cache
|
||||
readOnly: true
|
||||
args:
|
||||
- |
|
||||
set -e
|
||||
# Resolve SHA via the binary cache directory layout the
|
||||
# ensure-binary task established. multi-seed-matrix passes the
|
||||
# SHA implicitly via the workflow-scoped /data/bin tree.
|
||||
SHA=$(ls -1t /data/bin | head -1)
|
||||
export PATH="/data/bin/$SHA:$PATH"
|
||||
export LD_LIBRARY_PATH=$(echo "$LD_LIBRARY_PATH" | tr ':' '\n' | grep -v stubs | tr '\n' ':' | sed 's/:$//')
|
||||
nvidia-smi
|
||||
|
||||
MODEL="{{workflow.parameters.model}}"
|
||||
case "$MODEL" in
|
||||
alpha-rl) BINARY=alpha_rl_train ;;
|
||||
dqn|ppo) BINARY=train_baseline_rl ;;
|
||||
*) BINARY=train_baseline_supervised ;;
|
||||
esac
|
||||
|
||||
mkdir -p /workspace/output
|
||||
|
||||
# Plan 5 Task 3 (A.4.1): optional nsys profile wrapper.
|
||||
# When workflow.parameters.profile=="true", wrap the training
|
||||
# binary under `nsys profile --capture-range=cudaProfilerApi`.
|
||||
# The .nsys-rep is uploaded to MinIO bucket
|
||||
# foxhunt-training-artifacts/profiles/<short-sha>/ in the
|
||||
# post-training upload block below.
|
||||
PROFILE="{{workflow.parameters.profile}}"
|
||||
POD_NAME="${HOSTNAME:-pod}"
|
||||
NSYS_OUT="/workspace/output/profile-${POD_NAME}.nsys-rep"
|
||||
NSYS_PREFIX=""
|
||||
if [ "$PROFILE" = "true" ]; then
|
||||
NSYS_BIN=$(which nsys 2>/dev/null || echo "/usr/local/cuda/bin/nsys")
|
||||
if [ -x "$NSYS_BIN" ]; then
|
||||
NSYS_PREFIX="$NSYS_BIN profile --capture-range=cudaProfilerApi --output=${NSYS_OUT} --force-overwrite=true"
|
||||
echo "=== nsys profile enabled: ${NSYS_OUT} ==="
|
||||
else
|
||||
echo "WARN: --profile requested but nsys not found at $NSYS_BIN — running without"
|
||||
fi
|
||||
fi
|
||||
|
||||
echo "=== Training: $MODEL seed=$SEED folds={{workflow.parameters.folds}} ==="
|
||||
if [ "$MODEL" = "alpha-rl" ]; then
|
||||
stdbuf -oL $NSYS_PREFIX ${BINARY} \
|
||||
--mbp10-data-dir /data/futures-baseline-mbp10/{{workflow.parameters.symbol}} \
|
||||
--predecoded-dir /feature-cache/predecoded \
|
||||
--out /workspace/output \
|
||||
--n-steps 50000 \
|
||||
--n-backtests 16 \
|
||||
--seed "$SEED" \
|
||||
--instrument-mode all
|
||||
else
|
||||
stdbuf -oL $NSYS_PREFIX ${BINARY} \
|
||||
--model "$MODEL" \
|
||||
--symbol {{workflow.parameters.symbol}} \
|
||||
--tx-cost-bps {{workflow.parameters.tx-cost-bps}} \
|
||||
--tick-size {{workflow.parameters.tick-size}} \
|
||||
--spread-ticks {{workflow.parameters.spread-ticks}} \
|
||||
--initial-capital {{workflow.parameters.initial-capital}} \
|
||||
--data-dir /data/futures-baseline \
|
||||
--mbp10-data-dir /data/futures-baseline-mbp10 \
|
||||
--trades-data-dir /data/futures-baseline-trades \
|
||||
--output-dir /workspace/output \
|
||||
--epochs {{workflow.parameters.train-epochs}} \
|
||||
--seed "$SEED" \
|
||||
--max-folds {{workflow.parameters.folds}} \
|
||||
--imbalance-bar-threshold {{workflow.parameters.imbalance-bar-threshold}} \
|
||||
--imbalance-bar-ewma-alpha {{workflow.parameters.imbalance-bar-ewma-alpha}} \
|
||||
--volume-bar-size {{workflow.parameters.volume-bar-size}}
|
||||
fi
|
||||
|
||||
echo "=== Training complete: seed=$SEED folds={{workflow.parameters.folds}} ==="
|
||||
|
||||
# Plan 5 Task 3 (A.4.1): upload .nsys-rep to MinIO if profile run.
|
||||
# mc is fetched on-demand (~25 MB single static binary) — the
|
||||
# ci-builder training image does not bundle it. Bucket:
|
||||
# foxhunt-training-artifacts/profiles/<short-sha>/.
|
||||
if [ "$PROFILE" = "true" ] && [ -f "$NSYS_OUT" ]; then
|
||||
echo "=== Uploading nsys profile to MinIO ==="
|
||||
MC_BIN=$(which mc 2>/dev/null || echo "")
|
||||
if [ -z "$MC_BIN" ]; then
|
||||
MC_BIN=/tmp/mc
|
||||
curl -fsSL https://dl.min.io/client/mc/release/linux-amd64/mc -o "$MC_BIN" || {
|
||||
echo "WARN: failed to download mc — skipping upload"
|
||||
exit 0
|
||||
}
|
||||
chmod +x "$MC_BIN"
|
||||
fi
|
||||
"$MC_BIN" alias set foxhunt http://minio.foxhunt.svc.cluster.local:9000 \
|
||||
"$MINIO_ACCESS_KEY" "$MINIO_SECRET_KEY" 2>/dev/null || true
|
||||
"$MC_BIN" cp "$NSYS_OUT" \
|
||||
"foxhunt/foxhunt-training-artifacts/profiles/$SHA/profile-seed${SEED}-${POD_NAME}.nsys-rep" || \
|
||||
echo "WARN: nsys upload failed for seed=$SEED"
|
||||
echo "=== nsys profile upload complete ==="
|
||||
fi
|
||||
|
||||
# ── notify-result: post workflow outcome to Mattermost (onExit) ──
|
||||
- name: notify-result
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: platform
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
container:
|
||||
image: curlimages/curl:8.12.1
|
||||
command: ["/bin/sh", "-c"]
|
||||
env:
|
||||
- name: WEBHOOK_URL
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: notification-webhook
|
||||
key: webhook-url
|
||||
optional: true
|
||||
resources:
|
||||
requests:
|
||||
cpu: 50m
|
||||
memory: 32Mi
|
||||
limits:
|
||||
cpu: 100m
|
||||
memory: 64Mi
|
||||
args:
|
||||
- |
|
||||
STATUS="{{workflow.status}}"
|
||||
NAME="{{workflow.name}}"
|
||||
if [ -z "$WEBHOOK_URL" ] || echo "$WEBHOOK_URL" | grep -q "PLACEHOLDER"; then
|
||||
echo "No webhook configured, skipping notification"; exit 0
|
||||
fi
|
||||
EMOJI=":x:"
|
||||
[ "$STATUS" = "Succeeded" ] && EMOJI=":white_check_mark:"
|
||||
PAYLOAD="{\"username\":\"Argo CI\",\"text\":\"${EMOJI} **${NAME}** (multi-seed) — ${STATUS}\"}"
|
||||
curl -sf -X POST -H 'Content-Type: application/json' \
|
||||
-d "$PAYLOAD" "$WEBHOOK_URL" || echo "WARN: webhook post failed"
|
||||
@@ -1,838 +0,0 @@
|
||||
# Unified training workflow — replaces all model-specific training templates.
|
||||
#
|
||||
# Key innovation: binary caching by commit SHA on PVC.
|
||||
# Check /data/bin/$SHA first; compile only on cache miss.
|
||||
#
|
||||
# DAG:
|
||||
# ensure-binary ──┐
|
||||
# gpu-warmup ─────┼──> ensure-fxcache ──> hyperopt ──> train-best ──> evaluate ──> upload-results
|
||||
#
|
||||
# When hyperopt-trials=0 (the `scripts/argo-precompute.sh` path),
|
||||
# gpu-warmup / hyperopt / train-best / evaluate all skip and the
|
||||
# workflow runs ensure-binary → ensure-fxcache only — no GPU node
|
||||
# gets provisioned for nothing.
|
||||
#
|
||||
# Usage:
|
||||
# argo submit -n foxhunt --from=wftmpl/train
|
||||
# argo submit -n foxhunt --from=wftmpl/train -p model=ppo -p train-epochs=100
|
||||
# argo submit -n foxhunt --from=wftmpl/train -p hyperopt-trials=0 # fxcache-only path
|
||||
apiVersion: argoproj.io/v1alpha1
|
||||
kind: WorkflowTemplate
|
||||
metadata:
|
||||
name: train
|
||||
namespace: foxhunt
|
||||
labels:
|
||||
app.kubernetes.io/name: train
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
app.kubernetes.io/component: train
|
||||
spec:
|
||||
entrypoint: pipeline
|
||||
onExit: notify-result
|
||||
serviceAccountName: argo-workflow
|
||||
archiveLogs: true
|
||||
podMetadata:
|
||||
labels:
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
app.kubernetes.io/component: train
|
||||
securityContext:
|
||||
fsGroup: 0
|
||||
ttlStrategy:
|
||||
secondsAfterCompletion: 3600
|
||||
activeDeadlineSeconds: 28800 # 8 hours
|
||||
|
||||
arguments:
|
||||
parameters:
|
||||
- name: commit-sha
|
||||
value: HEAD
|
||||
- name: git-branch
|
||||
value: main
|
||||
- name: cuda-compute-cap
|
||||
value: "90"
|
||||
- name: model
|
||||
value: dqn
|
||||
# gpu-pool default: ci-training-l40s (set 2026-05-14 per
|
||||
# `feedback_default_to_l40s_pool.md` — SP-chain training has been
|
||||
# standardising on L40S since 2026-05-09; the prior ci-training-h100
|
||||
# default required every SP-run invocation to pass an explicit
|
||||
# `--gpu-pool ci-training-l40s` override. H100 remains opt-in via
|
||||
# `--gpu-pool ci-training-h100` for runs that genuinely need
|
||||
# 80 GB VRAM or sm_90 features). Compute-cap derivation in
|
||||
# `argo-train.sh` matches this default to sm_89 (Ada Lovelace).
|
||||
- name: gpu-pool
|
||||
value: ci-training-l40s
|
||||
- name: hyperopt-trials
|
||||
value: "20"
|
||||
- name: hyperopt-epochs
|
||||
value: "8"
|
||||
- name: train-epochs
|
||||
value: "50"
|
||||
- name: symbol
|
||||
value: ES.FUT
|
||||
- name: initial-capital
|
||||
value: "35000"
|
||||
- name: tx-cost-bps
|
||||
value: "0.1"
|
||||
- name: tick-size
|
||||
value: "0.25"
|
||||
- name: spread-ticks
|
||||
value: "1.0"
|
||||
- name: sanitizer
|
||||
value: "none" # "none", "memcheck", "racecheck", "synccheck"
|
||||
# Bar formation params — passed to BOTH precompute_features (when
|
||||
# building fxcache) AND train_baseline_rl (when looking up fxcache).
|
||||
# MUST match between the two for fxcache HIT (cache key includes them
|
||||
# as of 2026-05-09 architectural fix).
|
||||
#
|
||||
# imbalance-bar-threshold default: 20.0 (set 2026-05-10). At threshold=0.5
|
||||
# the sampler produced 209M bars from 209M trade ticks (1:1 ratio,
|
||||
# essentially per-tick) on workflow f5wnd, near-OOM in feature extraction.
|
||||
# threshold=20 produces ~5-6M bars matching the volume-bar density baseline.
|
||||
- name: imbalance-bar-threshold
|
||||
value: "20.0"
|
||||
- name: imbalance-bar-ewma-alpha
|
||||
value: "0.1"
|
||||
- name: volume-bar-size
|
||||
value: "100"
|
||||
- name: data-source
|
||||
value: "mbp10"
|
||||
|
||||
volumes:
|
||||
- name: git-ssh-key
|
||||
secret:
|
||||
secretName: argo-git-ssh-key
|
||||
defaultMode: 256
|
||||
- name: training-data
|
||||
persistentVolumeClaim:
|
||||
claimName: training-data-pvc
|
||||
- name: cargo-target-cuda
|
||||
persistentVolumeClaim:
|
||||
claimName: cargo-target-cuda
|
||||
- name: feature-cache
|
||||
persistentVolumeClaim:
|
||||
claimName: feature-cache-pvc
|
||||
|
||||
volumeClaimTemplates:
|
||||
- metadata:
|
||||
name: workspace
|
||||
spec:
|
||||
accessModes: ["ReadWriteOnce"]
|
||||
storageClassName: scw-bssd
|
||||
resources:
|
||||
requests:
|
||||
storage: 5Gi
|
||||
|
||||
templates:
|
||||
# ── DAG: orchestrate all steps ──
|
||||
- name: pipeline
|
||||
dag:
|
||||
tasks:
|
||||
- name: ensure-binary
|
||||
template: ensure-binary
|
||||
- name: gpu-warmup
|
||||
template: gpu-warmup
|
||||
# gpu-warmup pre-provisions an L40S node so hyperopt /
|
||||
# train-best don't pay autoscaler latency at start. When
|
||||
# hyperopt-trials==0 (precompute-only path used by
|
||||
# `scripts/argo-precompute.sh`), both downstream consumers
|
||||
# are skipped and the GPU node would sit idle until the
|
||||
# workflow ends. Skip the warmup in that case to avoid
|
||||
# provisioning expensive L40S capacity for nothing.
|
||||
when: "{{workflow.parameters.hyperopt-trials}} != 0"
|
||||
- name: ensure-fxcache
|
||||
template: ensure-fxcache
|
||||
dependencies: [ensure-binary]
|
||||
arguments:
|
||||
parameters:
|
||||
- name: sha
|
||||
value: "{{tasks.ensure-binary.outputs.parameters.sha}}"
|
||||
- name: hyperopt
|
||||
template: hyperopt
|
||||
dependencies: [ensure-binary, gpu-warmup, ensure-fxcache]
|
||||
when: "{{workflow.parameters.hyperopt-trials}} != 0"
|
||||
arguments:
|
||||
parameters:
|
||||
- name: sha
|
||||
value: "{{tasks.ensure-binary.outputs.parameters.sha}}"
|
||||
- name: train-best
|
||||
template: train-best
|
||||
dependencies: [ensure-binary, gpu-warmup, ensure-fxcache, hyperopt]
|
||||
arguments:
|
||||
parameters:
|
||||
- name: sha
|
||||
value: "{{tasks.ensure-binary.outputs.parameters.sha}}"
|
||||
- name: evaluate
|
||||
template: evaluate
|
||||
dependencies: [train-best]
|
||||
arguments:
|
||||
parameters:
|
||||
- name: sha
|
||||
value: "{{tasks.ensure-binary.outputs.parameters.sha}}"
|
||||
- name: upload-results
|
||||
template: upload-results
|
||||
dependencies: [evaluate]
|
||||
|
||||
# ── ensure-binary: cache-or-compile training binaries by commit SHA ──
|
||||
- name: ensure-binary
|
||||
outputs:
|
||||
parameters:
|
||||
- name: sha
|
||||
valueFrom:
|
||||
path: /tmp/sha
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: ci-compile-cpu
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
tolerations:
|
||||
- key: node.cilium.io/agent-not-ready
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
container:
|
||||
image: gitlab-registry.foxhunt.svc.cluster.local:5000/root/foxhunt/ci-builder:latest
|
||||
imagePullPolicy: Always
|
||||
command: ["/bin/bash", "-c"]
|
||||
env:
|
||||
- name: SQLX_OFFLINE
|
||||
value: "true"
|
||||
- name: CARGO_TERM_COLOR
|
||||
value: always
|
||||
- name: CARGO_TARGET_DIR
|
||||
value: /cargo-target
|
||||
- name: CARGO_HOME
|
||||
value: /cargo-target/cargo-home
|
||||
# sccache: per-crate compile-output cache. Disk-local on the same
|
||||
# cargo-target-cuda PVC so no network cost on hit; survives across
|
||||
# pods and commits (cargo's incremental cache is invalidated by
|
||||
# git checkout mtime touches — sccache is content-hash keyed so it
|
||||
# isn't). Binary assumed present in ci-builder image.
|
||||
#
|
||||
# CARGO_INCREMENTAL=0 is required to make sccache effective:
|
||||
# the project's Cargo.toml / .cargo/config.toml set
|
||||
# `incremental = true`, which emits per-query save-analysis
|
||||
# artifacts that sccache does not cache. Turning incremental off
|
||||
# lets rustc emit pure object output that sccache hashes
|
||||
# uniformly — same source + same args → cache hit.
|
||||
- name: RUSTC_WRAPPER
|
||||
value: sccache
|
||||
- name: SCCACHE_DIR
|
||||
value: /cargo-target/sccache
|
||||
- name: SCCACHE_CACHE_SIZE
|
||||
value: "40G"
|
||||
- name: CARGO_INCREMENTAL
|
||||
value: "0"
|
||||
- name: GITLAB_PAT
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: gitlab-pat
|
||||
key: token
|
||||
resources:
|
||||
requests:
|
||||
cpu: "14"
|
||||
memory: 32Gi
|
||||
limits:
|
||||
cpu: "30"
|
||||
memory: 64Gi
|
||||
volumeMounts:
|
||||
- name: git-ssh-key
|
||||
mountPath: /etc/git-ssh
|
||||
readOnly: true
|
||||
- name: cargo-target-cuda
|
||||
mountPath: /cargo-target
|
||||
- name: training-data
|
||||
mountPath: /data
|
||||
args:
|
||||
- |
|
||||
set -e
|
||||
SHA="{{workflow.parameters.commit-sha}}"
|
||||
BRANCH="{{workflow.parameters.git-branch}}"
|
||||
MODEL="{{workflow.parameters.model}}"
|
||||
|
||||
# ── SSH setup ──
|
||||
mkdir -p ~/.ssh
|
||||
cp /etc/git-ssh/ssh-privatekey ~/.ssh/id_ed25519
|
||||
chmod 600 ~/.ssh/id_ed25519
|
||||
printf 'Host *\n StrictHostKeyChecking no\n UserKnownHostsFile /dev/null\n' > ~/.ssh/config
|
||||
chmod 600 ~/.ssh/config
|
||||
|
||||
REPO="ssh://git@gitlab-gitlab-shell.foxhunt.svc.cluster.local:2222/root/foxhunt.git"
|
||||
BUILD="/cargo-target/src"
|
||||
|
||||
git config --global --add safe.directory "$BUILD"
|
||||
|
||||
# ── Resolve SHA ──
|
||||
if [ "$SHA" = "HEAD" ]; then
|
||||
if [ -d "$BUILD/.git" ]; then
|
||||
cd "$BUILD"
|
||||
git fetch origin
|
||||
SHA=$(git rev-parse "origin/$BRANCH")
|
||||
cd /
|
||||
else
|
||||
git clone --filter=blob:none "$REPO" "$BUILD"
|
||||
cd "$BUILD"
|
||||
git checkout "origin/$BRANCH"
|
||||
SHA=$(git rev-parse HEAD)
|
||||
cd /
|
||||
fi
|
||||
fi
|
||||
SHORT_SHA=$(echo "$SHA" | cut -c1-9)
|
||||
echo "Resolved SHA: $SHA (short: $SHORT_SHA)"
|
||||
|
||||
# ── Derive needed binaries from model ──
|
||||
case "$MODEL" in
|
||||
dqn|ppo)
|
||||
BINARIES="hyperopt_baseline_rl train_baseline_rl evaluate_baseline precompute_features"
|
||||
;;
|
||||
*)
|
||||
BINARIES="hyperopt_baseline_supervised train_baseline_supervised evaluate_supervised precompute_features"
|
||||
;;
|
||||
esac
|
||||
|
||||
# ── Cache check ──
|
||||
BIN_DIR="/data/bin/$SHORT_SHA"
|
||||
ALL_CACHED=true
|
||||
for bin in $BINARIES; do
|
||||
if [ ! -x "$BIN_DIR/$bin" ]; then
|
||||
ALL_CACHED=false
|
||||
break
|
||||
fi
|
||||
done
|
||||
|
||||
if [ "$ALL_CACHED" = "true" ]; then
|
||||
echo "=== Cache HIT: all binaries present in $BIN_DIR ==="
|
||||
ls -lh "$BIN_DIR/"
|
||||
echo "$SHORT_SHA" > /tmp/sha
|
||||
exit 0
|
||||
fi
|
||||
|
||||
echo "=== Cache MISS: compiling binaries for $SHORT_SHA ==="
|
||||
|
||||
# ── Clone / checkout ──
|
||||
if [ -d "$BUILD/.git" ]; then
|
||||
cd "$BUILD"
|
||||
git fetch origin
|
||||
CURRENT=$(git rev-parse HEAD 2>/dev/null || echo "none")
|
||||
if [ "$CURRENT" != "$SHA" ]; then
|
||||
echo "Updating checkout: $(echo $CURRENT | cut -c1-8) -> $SHORT_SHA"
|
||||
git checkout --force "$SHA"
|
||||
git clean -fd
|
||||
fi
|
||||
else
|
||||
git clone --filter=blob:none "$REPO" "$BUILD"
|
||||
cd "$BUILD"
|
||||
git checkout "$SHA"
|
||||
fi
|
||||
|
||||
# ── Build ──
|
||||
export PATH="${CARGO_HOME}/bin:${PATH}"
|
||||
export CUDA_COMPUTE_CAP={{workflow.parameters.cuda-compute-cap}}
|
||||
|
||||
ML_EXAMPLE_ARGS=""
|
||||
for ex in $BINARIES; do
|
||||
ML_EXAMPLE_ARGS="$ML_EXAMPLE_ARGS --example $ex"
|
||||
done
|
||||
|
||||
echo "Building: $BINARIES"
|
||||
echo " CUDA_COMPUTE_CAP=$CUDA_COMPUTE_CAP"
|
||||
cargo build --release -p ml --features ml/cuda $ML_EXAMPLE_ARGS
|
||||
|
||||
# ── Install to cache dir ──
|
||||
mkdir -p "$BIN_DIR"
|
||||
for bin in $BINARIES; do
|
||||
cp "$CARGO_TARGET_DIR/release/examples/$bin" "$BIN_DIR/"
|
||||
done
|
||||
strip "$BIN_DIR/"*
|
||||
|
||||
echo "=== Cached binaries ==="
|
||||
ls -lh "$BIN_DIR/"
|
||||
|
||||
# ── Prune old SHAs: keep last 5 ──
|
||||
cd /data/bin
|
||||
ls -1t | tail -n +6 | while read -r old; do
|
||||
echo "Pruning old cache: $old"
|
||||
rm -rf "$old"
|
||||
done
|
||||
|
||||
echo "$SHORT_SHA" > /tmp/sha
|
||||
|
||||
# ── ensure-fxcache: precompute feature cache if needed ──
|
||||
#
|
||||
# Pinned to the high-memory ci-compile-cpu-hm pool (POP2-HM-32C-256G,
|
||||
# 32 vCPU + 256GB RAM, min_size=0). The 9-quarter precompute_features
|
||||
# peaks ~50-60GB during the post-OFI alpha_trades conversion which
|
||||
# OOM-killed the standard 64GB ci-compile-cpu pool twice on
|
||||
# 2026-05-16 (workflows train-wq8b8 + train-2l6p4 both died at
|
||||
# exitCode 137 right after "OFI computed"). The HM pool autoscales
|
||||
# to zero when idle so this pin only costs anything during an
|
||||
# actual fxcache rebuild.
|
||||
- name: ensure-fxcache
|
||||
inputs:
|
||||
parameters:
|
||||
- name: sha
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: ci-compile-cpu-hm
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
tolerations:
|
||||
- key: node.cilium.io/agent-not-ready
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
container:
|
||||
image: ubuntu:24.04
|
||||
command: ["/bin/bash", "-c"]
|
||||
env:
|
||||
- name: RUST_LOG
|
||||
value: info
|
||||
resources:
|
||||
requests:
|
||||
cpu: "4"
|
||||
memory: 16Gi
|
||||
limits:
|
||||
cpu: "28"
|
||||
memory: 200Gi
|
||||
volumeMounts:
|
||||
- name: training-data
|
||||
mountPath: /data
|
||||
readOnly: true
|
||||
- name: feature-cache
|
||||
mountPath: /feature-cache
|
||||
args:
|
||||
- |
|
||||
set -e
|
||||
SHA="{{inputs.parameters.sha}}"
|
||||
BINARY="/data/bin/$SHA/precompute_features"
|
||||
|
||||
if [ ! -x "$BINARY" ]; then
|
||||
echo "ERROR: precompute_features not found at $BINARY"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
# fxcache has built-in version validation (FXCACHE_VERSION in header).
|
||||
# The precompute binary checks existing cache — if version matches,
|
||||
# it skips regeneration. If version mismatches, it fails and we
|
||||
# delete + regenerate. No unconditional rm — cache is reused when valid.
|
||||
echo "=== Running precompute_features (SHA: $SHA) ==="
|
||||
if $BINARY \
|
||||
--data-dir /data/futures-baseline \
|
||||
--mbp10-data-dir /data/futures-baseline-mbp10 \
|
||||
--trades-data-dir /data/futures-baseline-trades \
|
||||
--output-dir /feature-cache \
|
||||
--symbol {{workflow.parameters.symbol}} \
|
||||
--data-source {{workflow.parameters.data-source}} \
|
||||
--imbalance-bar-threshold {{workflow.parameters.imbalance-bar-threshold}} \
|
||||
--imbalance-bar-ewma-alpha {{workflow.parameters.imbalance-bar-ewma-alpha}} \
|
||||
--volume-bar-size {{workflow.parameters.volume-bar-size}} \
|
||||
--yes; then
|
||||
echo "=== Feature cache ready ==="
|
||||
else
|
||||
echo "=== Cache stale or missing — regenerating ==="
|
||||
# NOTE: don't blanket-rm /feature-cache/*.fxcache anymore. With
|
||||
# bar-params now in the cache key, multiple valid caches can
|
||||
# coexist (different threshold experiments).
|
||||
$BINARY \
|
||||
--data-dir /data/futures-baseline \
|
||||
--mbp10-data-dir /data/futures-baseline-mbp10 \
|
||||
--trades-data-dir /data/futures-baseline-trades \
|
||||
--output-dir /feature-cache \
|
||||
--symbol {{workflow.parameters.symbol}} \
|
||||
--data-source {{workflow.parameters.data-source}} \
|
||||
--imbalance-bar-threshold {{workflow.parameters.imbalance-bar-threshold}} \
|
||||
--imbalance-bar-ewma-alpha {{workflow.parameters.imbalance-bar-ewma-alpha}} \
|
||||
--volume-bar-size {{workflow.parameters.volume-bar-size}} \
|
||||
--yes
|
||||
echo "=== Feature cache regenerated ==="
|
||||
fi
|
||||
|
||||
# ── gpu-warmup: trigger GPU node autoscale during compile ──
|
||||
- name: gpu-warmup
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: "{{workflow.parameters.gpu-pool}}"
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
tolerations:
|
||||
- key: nvidia.com/gpu
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
- key: node.cilium.io/agent-not-ready
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
container:
|
||||
image: busybox:1.37
|
||||
command: ["/bin/sh", "-c"]
|
||||
args:
|
||||
- |
|
||||
echo "GPU warmup: triggering node autoscale..."
|
||||
echo "GPU node scheduled, exiting to free resources"
|
||||
resources:
|
||||
requests:
|
||||
nvidia.com/gpu: "1"
|
||||
cpu: 100m
|
||||
memory: 64Mi
|
||||
limits:
|
||||
nvidia.com/gpu: "1"
|
||||
cpu: 200m
|
||||
memory: 128Mi
|
||||
|
||||
# ── hyperopt: PSO/TPE hyperparameter optimization on GPU ──
|
||||
- name: hyperopt
|
||||
inputs:
|
||||
parameters:
|
||||
- name: sha
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: "{{workflow.parameters.gpu-pool}}"
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
tolerations:
|
||||
- key: nvidia.com/gpu
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
- key: node.cilium.io/agent-not-ready
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
container:
|
||||
image: gitlab-registry.foxhunt.svc.cluster.local:5000/root/foxhunt/ci-builder:latest
|
||||
imagePullPolicy: IfNotPresent
|
||||
command: ["/bin/sh", "-c"]
|
||||
env:
|
||||
- name: RUST_LOG
|
||||
value: info
|
||||
- name: SQLX_OFFLINE
|
||||
value: "true"
|
||||
- name: CUBLAS_WORKSPACE_CONFIG
|
||||
value: ":4096:8"
|
||||
- name: FOXHUNT_FEATURE_CACHE_DIR
|
||||
value: /feature-cache
|
||||
resources:
|
||||
requests:
|
||||
nvidia.com/gpu: "1"
|
||||
cpu: "4"
|
||||
memory: 32Gi
|
||||
limits:
|
||||
nvidia.com/gpu: "1"
|
||||
cpu: "8"
|
||||
memory: 64Gi
|
||||
volumeMounts:
|
||||
- name: workspace
|
||||
mountPath: /workspace
|
||||
- name: training-data
|
||||
mountPath: /data
|
||||
readOnly: true
|
||||
- name: feature-cache
|
||||
mountPath: /feature-cache
|
||||
readOnly: true
|
||||
args:
|
||||
- |
|
||||
set -e
|
||||
SHA="{{inputs.parameters.sha}}"
|
||||
export PATH="/data/bin/$SHA:$PATH"
|
||||
export LD_LIBRARY_PATH=$(echo "$LD_LIBRARY_PATH" | tr ':' '\n' | grep -v stubs | tr '\n' ':' | sed 's/:$//')
|
||||
nvidia-smi
|
||||
|
||||
mkdir -p /workspace/output/hyperopt
|
||||
|
||||
MODEL="{{workflow.parameters.model}}"
|
||||
case "$MODEL" in
|
||||
dqn|ppo) BINARY=hyperopt_baseline_rl ;;
|
||||
*) BINARY=hyperopt_baseline_supervised ;;
|
||||
esac
|
||||
|
||||
echo "=== Running hyperopt: $MODEL ($BINARY) ==="
|
||||
echo " Trials: {{workflow.parameters.hyperopt-trials}}, Epochs: {{workflow.parameters.hyperopt-epochs}}"
|
||||
|
||||
${BINARY} \
|
||||
--model "$MODEL" \
|
||||
--phase fast \
|
||||
--trials {{workflow.parameters.hyperopt-trials}} \
|
||||
--epochs {{workflow.parameters.hyperopt-epochs}} \
|
||||
--symbol {{workflow.parameters.symbol}} \
|
||||
--tx-cost-bps {{workflow.parameters.tx-cost-bps}} \
|
||||
--tick-size {{workflow.parameters.tick-size}} \
|
||||
--spread-ticks {{workflow.parameters.spread-ticks}} \
|
||||
--initial-capital {{workflow.parameters.initial-capital}} \
|
||||
--data-dir /data/futures-baseline \
|
||||
--mbp10-data-dir /data/futures-baseline-mbp10 \
|
||||
--trades-data-dir /data/futures-baseline-trades \
|
||||
--base-dir /workspace/output/hyperopt \
|
||||
--output /workspace/output/${MODEL}_hyperopt_results.json
|
||||
|
||||
echo "=== Hyperopt complete ==="
|
||||
cat /workspace/output/${MODEL}_hyperopt_results.json 2>/dev/null || echo "No results file"
|
||||
|
||||
# ── train-best: full training with best hyperparams ──
|
||||
- name: train-best
|
||||
inputs:
|
||||
parameters:
|
||||
- name: sha
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: "{{workflow.parameters.gpu-pool}}"
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
tolerations:
|
||||
- key: nvidia.com/gpu
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
- key: node.cilium.io/agent-not-ready
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
container:
|
||||
image: gitlab-registry.foxhunt.svc.cluster.local:5000/root/foxhunt/ci-builder:latest
|
||||
imagePullPolicy: IfNotPresent
|
||||
command: ["/bin/sh", "-c"]
|
||||
env:
|
||||
- name: RUST_LOG
|
||||
value: info
|
||||
- name: SQLX_OFFLINE
|
||||
value: "true"
|
||||
- name: CUBLAS_WORKSPACE_CONFIG
|
||||
value: ":4096:8"
|
||||
- name: FOXHUNT_FEATURE_CACHE_DIR
|
||||
value: /feature-cache
|
||||
resources:
|
||||
requests:
|
||||
nvidia.com/gpu: "1"
|
||||
cpu: "4"
|
||||
memory: 32Gi
|
||||
limits:
|
||||
nvidia.com/gpu: "1"
|
||||
cpu: "8"
|
||||
memory: 64Gi
|
||||
volumeMounts:
|
||||
- name: workspace
|
||||
mountPath: /workspace
|
||||
- name: training-data
|
||||
mountPath: /data
|
||||
readOnly: true
|
||||
- name: feature-cache
|
||||
mountPath: /feature-cache
|
||||
readOnly: true
|
||||
args:
|
||||
- |
|
||||
set -e
|
||||
SHA="{{inputs.parameters.sha}}"
|
||||
export PATH="/data/bin/$SHA:$PATH"
|
||||
export LD_LIBRARY_PATH=$(echo "$LD_LIBRARY_PATH" | tr ':' '\n' | grep -v stubs | tr '\n' ':' | sed 's/:$//')
|
||||
nvidia-smi
|
||||
|
||||
MODEL="{{workflow.parameters.model}}"
|
||||
case "$MODEL" in
|
||||
dqn|ppo) BINARY=train_baseline_rl ;;
|
||||
*) BINARY=train_baseline_supervised ;;
|
||||
esac
|
||||
|
||||
HYPEROPT_FLAG=""
|
||||
if [ -f "/workspace/output/${MODEL}_hyperopt_results.json" ]; then
|
||||
HYPEROPT_FLAG="--hyperopt-params /workspace/output/${MODEL}_hyperopt_results.json"
|
||||
echo " Using hyperopt results: ${MODEL}_hyperopt_results.json"
|
||||
else
|
||||
echo " No hyperopt results — training with default hyperparams"
|
||||
fi
|
||||
|
||||
echo "=== Training: $MODEL ({{workflow.parameters.train-epochs}} epochs) ==="
|
||||
|
||||
# compute-sanitizer / nsys: optional GPU debugging/profiling
|
||||
SANITIZER="{{workflow.parameters.sanitizer}}"
|
||||
SANITIZER_PREFIX=""
|
||||
if [ "$SANITIZER" = "nsys" ]; then
|
||||
NSYS_BIN=$(which nsys 2>/dev/null || echo "/usr/local/cuda/bin/nsys")
|
||||
if [ -x "$NSYS_BIN" ]; then
|
||||
NSYS_OUT="/feature-cache/nsys_$(date +%Y%m%d_%H%M%S)"
|
||||
mkdir -p /feature-cache
|
||||
SANITIZER_PREFIX="$NSYS_BIN profile -o $NSYS_OUT --cuda-graph-trace=node --stats=true --show-output=true -f true --duration=60"
|
||||
echo " nsys profiling enabled — output: ${NSYS_OUT}.nsys-rep (persistent PVC)"
|
||||
echo " --cuda-graph-trace=node: per-kernel timing inside parent graph"
|
||||
echo " --cudabacktrace=kernel: kernel source attribution"
|
||||
echo " --capture-range=cudaProfilerApi: use cudaProfilerStart/Stop to limit capture"
|
||||
echo " WARNING: ~2x slower — use with 1-2 epochs only"
|
||||
else
|
||||
echo " WARNING: nsys not found at $NSYS_BIN — running without"
|
||||
fi
|
||||
elif [ "$SANITIZER" != "none" ] && [ -n "$SANITIZER" ]; then
|
||||
SANITIZER_BIN=$(which compute-sanitizer 2>/dev/null || echo "/usr/local/cuda/bin/compute-sanitizer")
|
||||
if [ -x "$SANITIZER_BIN" ]; then
|
||||
SANITIZER_PREFIX="$SANITIZER_BIN --tool $SANITIZER --print-limit 20 --error-exitcode 1"
|
||||
echo " compute-sanitizer enabled: --tool $SANITIZER"
|
||||
echo " WARNING: 10-100x slower — use with 1-2 epochs only"
|
||||
else
|
||||
echo " WARNING: compute-sanitizer not found at $SANITIZER_BIN — running without"
|
||||
fi
|
||||
fi
|
||||
|
||||
stdbuf -oL $SANITIZER_PREFIX ${BINARY} \
|
||||
--model "$MODEL" \
|
||||
--symbol {{workflow.parameters.symbol}} \
|
||||
--tx-cost-bps {{workflow.parameters.tx-cost-bps}} \
|
||||
--tick-size {{workflow.parameters.tick-size}} \
|
||||
--spread-ticks {{workflow.parameters.spread-ticks}} \
|
||||
--initial-capital {{workflow.parameters.initial-capital}} \
|
||||
--data-dir /data/futures-baseline \
|
||||
--mbp10-data-dir /data/futures-baseline-mbp10 \
|
||||
--trades-data-dir /data/futures-baseline-trades \
|
||||
--output-dir /workspace/output \
|
||||
--epochs {{workflow.parameters.train-epochs}} \
|
||||
--imbalance-bar-threshold {{workflow.parameters.imbalance-bar-threshold}} \
|
||||
--imbalance-bar-ewma-alpha {{workflow.parameters.imbalance-bar-ewma-alpha}} \
|
||||
--volume-bar-size {{workflow.parameters.volume-bar-size}} \
|
||||
$HYPEROPT_FLAG
|
||||
|
||||
echo "=== Training complete ==="
|
||||
|
||||
# ── evaluate: run evaluation on trained model ──
|
||||
- name: evaluate
|
||||
inputs:
|
||||
parameters:
|
||||
- name: sha
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: "{{workflow.parameters.gpu-pool}}"
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
tolerations:
|
||||
- key: nvidia.com/gpu
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
- key: node.cilium.io/agent-not-ready
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
container:
|
||||
image: gitlab-registry.foxhunt.svc.cluster.local:5000/root/foxhunt/ci-builder:latest
|
||||
imagePullPolicy: IfNotPresent
|
||||
command: ["/bin/sh", "-c"]
|
||||
env:
|
||||
- name: RUST_LOG
|
||||
value: info
|
||||
- name: SQLX_OFFLINE
|
||||
value: "true"
|
||||
resources:
|
||||
requests:
|
||||
nvidia.com/gpu: "1"
|
||||
cpu: "2"
|
||||
memory: 16Gi
|
||||
limits:
|
||||
nvidia.com/gpu: "1"
|
||||
cpu: "4"
|
||||
memory: 32Gi
|
||||
volumeMounts:
|
||||
- name: workspace
|
||||
mountPath: /workspace
|
||||
- name: training-data
|
||||
mountPath: /data
|
||||
readOnly: true
|
||||
- name: feature-cache
|
||||
mountPath: /feature-cache
|
||||
readOnly: true
|
||||
args:
|
||||
- |
|
||||
set -e
|
||||
SHA="{{inputs.parameters.sha}}"
|
||||
export PATH="/data/bin/$SHA:$PATH"
|
||||
export LD_LIBRARY_PATH=$(echo "$LD_LIBRARY_PATH" | tr ':' '\n' | grep -v stubs | tr '\n' ':' | sed 's/:$//')
|
||||
|
||||
MODEL="{{workflow.parameters.model}}"
|
||||
case "$MODEL" in
|
||||
dqn|ppo) BINARY=evaluate_baseline ;;
|
||||
*) BINARY=evaluate_supervised ;;
|
||||
esac
|
||||
|
||||
echo "=== Evaluating: $MODEL ==="
|
||||
# Args must match evaluate_baseline's Args struct — see crates/ml/examples/evaluate_baseline.rs
|
||||
# `--models-dir` (not --checkpoint-dir), `--output` is a FILE path (not --output-dir).
|
||||
mkdir -p /workspace/output/eval
|
||||
${BINARY} \
|
||||
--model "$MODEL" \
|
||||
--symbol {{workflow.parameters.symbol}} \
|
||||
--data-dir /data/futures-baseline \
|
||||
--models-dir /workspace/output \
|
||||
--output /workspace/output/eval/evaluation_report.json \
|
||||
--initial-capital {{workflow.parameters.initial-capital}} \
|
||||
--tx-cost-bps {{workflow.parameters.tx-cost-bps}} \
|
||||
--tick-size {{workflow.parameters.tick-size}} \
|
||||
--spread-ticks {{workflow.parameters.spread-ticks}} || {
|
||||
echo "WARN: Evaluation failed, continuing"
|
||||
}
|
||||
|
||||
echo "=== Evaluation complete ==="
|
||||
|
||||
# ── upload-results: push artifacts to GitLab packages ──
|
||||
- name: upload-results
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: platform
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
container:
|
||||
image: curlimages/curl:8.12.1
|
||||
command: ["/bin/sh", "-c"]
|
||||
env:
|
||||
- name: GITLAB_PAT
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: gitlab-pat
|
||||
key: token
|
||||
resources:
|
||||
requests:
|
||||
cpu: 100m
|
||||
memory: 128Mi
|
||||
limits:
|
||||
cpu: 500m
|
||||
memory: 256Mi
|
||||
volumeMounts:
|
||||
- name: workspace
|
||||
mountPath: /workspace
|
||||
args:
|
||||
- |
|
||||
set -e
|
||||
MODEL="{{workflow.parameters.model}}"
|
||||
SYMBOL="{{workflow.parameters.symbol}}"
|
||||
TIMESTAMP=$(date +%Y%m%d-%H%M%S)
|
||||
GITLAB_API="http://gitlab-webservice-default.foxhunt.svc.cluster.local:8181/api/v4"
|
||||
PKG_NAME="foxhunt-training-results"
|
||||
PKG_VERSION="${MODEL}-${SYMBOL}-${TIMESTAMP}"
|
||||
|
||||
echo "Uploading training artifacts: ${PKG_NAME}/${PKG_VERSION}"
|
||||
|
||||
UPLOADED=0
|
||||
find /workspace/output -type f | while read -r file; do
|
||||
REL_PATH="${file#/workspace/output/}"
|
||||
SAFE_NAME=$(echo "$REL_PATH" | tr '/' '--')
|
||||
echo " Uploading ${REL_PATH} as ${SAFE_NAME}..."
|
||||
curl -f --upload-file "$file" \
|
||||
-H "PRIVATE-TOKEN: ${GITLAB_PAT}" \
|
||||
"${GITLAB_API}/projects/1/packages/generic/${PKG_NAME}/${PKG_VERSION}/${SAFE_NAME}" && \
|
||||
UPLOADED=$((UPLOADED + 1)) || \
|
||||
echo " WARN: Failed to upload ${REL_PATH}"
|
||||
done
|
||||
|
||||
echo "=== Upload complete (${UPLOADED} files) ==="
|
||||
echo "Package: ${PKG_NAME}/${PKG_VERSION}"
|
||||
|
||||
# ── notify-result: post workflow outcome to Mattermost (onExit) ──
|
||||
- name: notify-result
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: platform
|
||||
topology.kubernetes.io/zone: fr-par-2
|
||||
container:
|
||||
image: curlimages/curl:8.12.1
|
||||
command: ["/bin/sh", "-c"]
|
||||
env:
|
||||
- name: WEBHOOK_URL
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: notification-webhook
|
||||
key: webhook-url
|
||||
optional: true
|
||||
resources:
|
||||
requests:
|
||||
cpu: 50m
|
||||
memory: 32Mi
|
||||
limits:
|
||||
cpu: 100m
|
||||
memory: 64Mi
|
||||
args:
|
||||
- |
|
||||
STATUS="{{workflow.status}}"
|
||||
NAME="{{workflow.name}}"
|
||||
|
||||
if [ -z "$WEBHOOK_URL" ] || echo "$WEBHOOK_URL" | grep -q "PLACEHOLDER"; then
|
||||
echo "No webhook configured, skipping notification"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
if [ "$STATUS" = "Succeeded" ]; then
|
||||
EMOJI=":white_check_mark:"
|
||||
else
|
||||
EMOJI=":x:"
|
||||
fi
|
||||
|
||||
PAYLOAD="{\"username\":\"Argo CI\",\"text\":\"${EMOJI} **${NAME}** — ${STATUS} ({{workflow.duration}}s)\"}"
|
||||
|
||||
curl -sf -X POST -H 'Content-Type: application/json' \
|
||||
-d "$PAYLOAD" "$WEBHOOK_URL" || echo "WARN: webhook post failed"
|
||||
72
infra/k8s/cert-manager/fxhnt-acme.yaml
Normal file
72
infra/k8s/cert-manager/fxhnt-acme.yaml
Normal file
@@ -0,0 +1,72 @@
|
||||
# ACME (Let's Encrypt) wildcard cert for *.fxhnt.ai via cert-manager + Scaleway DNS-01.
|
||||
# Replaces the prior MANUAL, unrenewed LE cert that expired 2026-05-26 (the tailscale-proxy
|
||||
# fell back to the GitLab self-signed cert). DNS-01 is required: the services are tailnet-only
|
||||
# (*.fxhnt.ai -> 100.x Tailscale CGNAT) so HTTP-01 is unreachable, and wildcards need DNS-01.
|
||||
#
|
||||
# Solver creds: cert-manager/scaleway-dns-credentials (SCW_ACCESS_KEY/SCW_SECRET_KEY, copied from
|
||||
# foxhunt/scaleway-credentials). Webhook: scaleway-certmanager-webhook (helm, cert-manager ns).
|
||||
# The issued secret gitlab-tls-cert is mounted by infra/k8s/gitlab/tailscale-proxy.yaml.
|
||||
---
|
||||
apiVersion: cert-manager.io/v1
|
||||
kind: ClusterIssuer
|
||||
metadata:
|
||||
name: letsencrypt-staging
|
||||
spec:
|
||||
acme:
|
||||
email: jeroen@grusewski.nl
|
||||
server: https://acme-staging-v02.api.letsencrypt.org/directory
|
||||
privateKeySecretRef:
|
||||
name: letsencrypt-staging-account-key
|
||||
solvers:
|
||||
- dns01:
|
||||
webhook:
|
||||
groupName: acme.scaleway.com
|
||||
solverName: scaleway
|
||||
config:
|
||||
accessKeySecretRef:
|
||||
name: scaleway-dns-credentials
|
||||
key: SCW_ACCESS_KEY
|
||||
secretKeySecretRef:
|
||||
name: scaleway-dns-credentials
|
||||
key: SCW_SECRET_KEY
|
||||
---
|
||||
apiVersion: cert-manager.io/v1
|
||||
kind: ClusterIssuer
|
||||
metadata:
|
||||
name: letsencrypt-prod
|
||||
spec:
|
||||
acme:
|
||||
email: jeroen@grusewski.nl
|
||||
server: https://acme-v02.api.letsencrypt.org/directory
|
||||
privateKeySecretRef:
|
||||
name: letsencrypt-prod-account-key
|
||||
solvers:
|
||||
- dns01:
|
||||
webhook:
|
||||
groupName: acme.scaleway.com
|
||||
solverName: scaleway
|
||||
config:
|
||||
accessKeySecretRef:
|
||||
name: scaleway-dns-credentials
|
||||
key: SCW_ACCESS_KEY
|
||||
secretKeySecretRef:
|
||||
name: scaleway-dns-credentials
|
||||
key: SCW_SECRET_KEY
|
||||
---
|
||||
# Wildcard cert for the tailscale-proxy. issuerRef flips staging->prod after staging validates;
|
||||
# secretName flips to gitlab-tls-cert (the secret the proxy mounts) for the prod issuance.
|
||||
apiVersion: cert-manager.io/v1
|
||||
kind: Certificate
|
||||
metadata:
|
||||
name: fxhnt-wildcard
|
||||
namespace: foxhunt
|
||||
spec:
|
||||
secretName: gitlab-tls-cert # takes over the secret the tailscale-proxy mounts
|
||||
privateKey:
|
||||
rotationPolicy: Always # old manual key had a mismatching algorithm; regenerate on issue/renew
|
||||
dnsNames:
|
||||
- fxhnt.ai
|
||||
- "*.fxhnt.ai"
|
||||
issuerRef:
|
||||
name: letsencrypt-prod
|
||||
kind: ClusterIssuer
|
||||
40
infra/k8s/gitea/clusterip-services.yaml
Normal file
40
infra/k8s/gitea/clusterip-services.yaml
Normal file
@@ -0,0 +1,40 @@
|
||||
# The gitea chart's gitea-http/gitea-ssh services are HEADLESS (clusterIP: None), so they resolve to the
|
||||
# pod IP (100.64.x.x). That range OVERLAPS the Tailscale CGNAT range (100.64.0.0/10), so the tailscale
|
||||
# sidecar in the tailscale-gitlab-proxy pod swallows traffic to the pod IP. These normal ClusterIP services
|
||||
# give Gitea a service IP in the service CIDR (10.32.x.x, outside the tailscale range) — the proxy nginx +
|
||||
# socat target these instead. (GitLab worked because its webservice svc was a normal ClusterIP.)
|
||||
apiVersion: v1
|
||||
kind: Service
|
||||
metadata:
|
||||
name: gitea-web
|
||||
namespace: foxhunt
|
||||
labels:
|
||||
app.kubernetes.io/name: gitea
|
||||
spec:
|
||||
type: ClusterIP
|
||||
selector:
|
||||
app.kubernetes.io/name: gitea
|
||||
app.kubernetes.io/instance: gitea
|
||||
ports:
|
||||
- name: http
|
||||
port: 3000
|
||||
targetPort: 3000
|
||||
protocol: TCP
|
||||
---
|
||||
apiVersion: v1
|
||||
kind: Service
|
||||
metadata:
|
||||
name: gitea-sshd
|
||||
namespace: foxhunt
|
||||
labels:
|
||||
app.kubernetes.io/name: gitea
|
||||
spec:
|
||||
type: ClusterIP
|
||||
selector:
|
||||
app.kubernetes.io/name: gitea
|
||||
app.kubernetes.io/instance: gitea
|
||||
ports:
|
||||
- name: ssh
|
||||
port: 22
|
||||
targetPort: 22
|
||||
protocol: TCP
|
||||
61
infra/k8s/gitea/networkpolicy.yaml
Normal file
61
infra/k8s/gitea/networkpolicy.yaml
Normal file
@@ -0,0 +1,61 @@
|
||||
# Gitea NetworkPolicy (foxhunt ns has default-deny-all). Egress: postgres + DNS + in-cluster webhook
|
||||
# target + general HTTPS. Ingress: from platform/argo pods + the tailscale proxy → :3000 (web) and :22 (ssh).
|
||||
apiVersion: networking.k8s.io/v1
|
||||
kind: NetworkPolicy
|
||||
metadata:
|
||||
name: gitea
|
||||
namespace: foxhunt
|
||||
spec:
|
||||
podSelector:
|
||||
matchLabels:
|
||||
app.kubernetes.io/name: gitea
|
||||
policyTypes:
|
||||
- Ingress
|
||||
- Egress
|
||||
ingress:
|
||||
# web (3000) + ssh (22) from platform + argo-workflow pods (cockpit build clones over HTTP)
|
||||
- from:
|
||||
- podSelector:
|
||||
matchExpressions:
|
||||
- key: app.kubernetes.io/part-of
|
||||
operator: In
|
||||
values: ["foxhunt", "argo-workflows"]
|
||||
ports:
|
||||
- { port: 3000, protocol: TCP }
|
||||
- { port: 22, protocol: TCP }
|
||||
# web + ssh from the tailscale proxy (nginx + socat) that fronts git.fxhnt.ai
|
||||
- from:
|
||||
- podSelector:
|
||||
matchLabels:
|
||||
app.kubernetes.io/name: tailscale-gitlab-proxy
|
||||
ports:
|
||||
- { port: 3000, protocol: TCP }
|
||||
- { port: 22, protocol: TCP }
|
||||
egress:
|
||||
# PostgreSQL
|
||||
- to:
|
||||
- podSelector:
|
||||
matchLabels:
|
||||
app.kubernetes.io/name: postgres
|
||||
ports:
|
||||
- { port: 5432, protocol: TCP }
|
||||
# DNS
|
||||
- ports:
|
||||
- { port: 53, protocol: UDP }
|
||||
- { port: 53, protocol: TCP }
|
||||
# Argo Events webhook target (gitea push hook → eventsource :12000, in-cluster)
|
||||
- to:
|
||||
- podSelector:
|
||||
matchExpressions:
|
||||
- key: app.kubernetes.io/part-of
|
||||
operator: In
|
||||
values: ["foxhunt", "argo-workflows"]
|
||||
ports:
|
||||
- { port: 12000, protocol: TCP }
|
||||
# general HTTPS (avatars, external clones) — external only, exclude cluster CIDRs
|
||||
- to:
|
||||
- ipBlock:
|
||||
cidr: 0.0.0.0/0
|
||||
except: ["10.32.0.0/16", "172.16.0.0/16"]
|
||||
ports:
|
||||
- { port: 443, protocol: TCP }
|
||||
64
infra/k8s/gitea/values.yaml
Normal file
64
infra/k8s/gitea/values.yaml
Normal file
@@ -0,0 +1,64 @@
|
||||
# Gitea — lightweight git host replacing GitLab (Phase 2B). External Postgres (existing in-cluster
|
||||
# `postgres`), no bundled DB/redis/memcached, Actions off (Argo does CI). Internal-only until cutover.
|
||||
replicaCount: 1
|
||||
image:
|
||||
rootless: true
|
||||
|
||||
# Disable all bundled subcharts — reuse the existing in-cluster postgres, no redis/memcached.
|
||||
postgresql:
|
||||
enabled: false
|
||||
postgresql-ha:
|
||||
enabled: false
|
||||
redis-cluster:
|
||||
enabled: false
|
||||
redis:
|
||||
enabled: false
|
||||
# chart 12.x replaced redis with valkey — disable both, gitea uses embedded queue/cache/session
|
||||
valkey-cluster:
|
||||
enabled: false
|
||||
valkey:
|
||||
enabled: false
|
||||
|
||||
persistence:
|
||||
enabled: true
|
||||
size: 5Gi
|
||||
storageClass: sbs-default-retain
|
||||
|
||||
resources:
|
||||
requests: { cpu: 100m, memory: 128Mi }
|
||||
limits: { cpu: "1", memory: 512Mi }
|
||||
|
||||
service:
|
||||
http: { type: ClusterIP, port: 3000 }
|
||||
ssh: { type: ClusterIP, port: 22 }
|
||||
|
||||
gitea:
|
||||
admin:
|
||||
existingSecret: gitea-admin
|
||||
config:
|
||||
actions:
|
||||
ENABLED: "false" # Argo does CI/CD; Gitea Actions off (subchart removed in chart 12.x)
|
||||
server:
|
||||
ROOT_URL: https://git.fxhnt.ai/
|
||||
DOMAIN: git.fxhnt.ai
|
||||
SSH_DOMAIN: git.fxhnt.ai
|
||||
SSH_PORT: "22" # clean git@git.fxhnt.ai clone URLs; :2222 retired (Task 5).
|
||||
DISABLE_SSH: "false"
|
||||
database:
|
||||
DB_TYPE: postgres
|
||||
HOST: postgres.foxhunt.svc.cluster.local:5432
|
||||
NAME: gitea
|
||||
USER: gitea
|
||||
service:
|
||||
DISABLE_REGISTRATION: "true"
|
||||
# No redis/memcached — use embedded adapters
|
||||
cache:
|
||||
ADAPTER: memory
|
||||
session:
|
||||
PROVIDER: db
|
||||
queue:
|
||||
TYPE: level
|
||||
additionalConfigFromEnvs:
|
||||
- name: GITEA__database__PASSWD
|
||||
valueFrom:
|
||||
secretKeyRef: { name: gitea-db, key: password }
|
||||
@@ -1,73 +0,0 @@
|
||||
# GitLab Runner — H100 GPU training workloads (hyperopt, walk-forward)
|
||||
# Runs on ci-training-h100 pool (H100 PCIe 1x80GB)
|
||||
# SXM pools have zero Scaleway quota — use PCIe until quota is granted.
|
||||
|
||||
gitlabUrl: http://gitlab-webservice-default.foxhunt.svc.cluster.local:8181
|
||||
# runnerToken set via --set at install time
|
||||
|
||||
replicas: 1
|
||||
|
||||
# Reuse the existing gitlab-runner SA (has pods/secrets/configmaps RBAC)
|
||||
rbac:
|
||||
create: false
|
||||
serviceAccount:
|
||||
create: false
|
||||
name: gitlab-runner
|
||||
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: infra
|
||||
|
||||
runners:
|
||||
cloneUrl: http://gitlab-webservice-default.foxhunt.svc.cluster.local:8181
|
||||
config: |
|
||||
[[runners]]
|
||||
clone_url = "http://gitlab-webservice-default.foxhunt.svc.cluster.local:8181"
|
||||
tag_list = ["kapsule", "h100"]
|
||||
[runners.kubernetes]
|
||||
namespace = "foxhunt"
|
||||
service_account = "gitlab-runner"
|
||||
image = "rust:1.89-slim"
|
||||
privileged = false
|
||||
node_selector_overwrite_allowed = ".*"
|
||||
cpu_request_overwrite_max_allowed = "24000m"
|
||||
cpu_limit_overwrite_max_allowed = "24000m"
|
||||
memory_request_overwrite_max_allowed = "200Gi"
|
||||
memory_limit_overwrite_max_allowed = "200Gi"
|
||||
poll_timeout = 600
|
||||
runtime_class_name = "nvidia"
|
||||
pod_annotations_overwrite_allowed = ".*"
|
||||
# Default resources for H100 training
|
||||
cpu_request = "2000m"
|
||||
cpu_limit = "3800m"
|
||||
memory_request = "4Gi"
|
||||
memory_limit = "8Gi"
|
||||
helper_cpu_request = "100m"
|
||||
helper_cpu_limit = "500m"
|
||||
helper_memory_request = "128Mi"
|
||||
helper_memory_limit = "512Mi"
|
||||
image_pull_secrets = ["gitlab-registry"]
|
||||
[runners.kubernetes.node_selector]
|
||||
"k8s.scaleway.com/pool-name" = "ci-training-h100"
|
||||
[runners.kubernetes.node_tolerations]
|
||||
"nvidia.com/gpu" = "NoSchedule"
|
||||
"node.cilium.io/agent-not-ready" = "NoSchedule"
|
||||
[runners.kubernetes.pod_labels]
|
||||
"app.kubernetes.io/part-of" = "foxhunt-ci"
|
||||
# Request GPU via K8s scheduler so only one training pod runs per GPU
|
||||
[[runners.kubernetes.pod_spec]]
|
||||
name = "build"
|
||||
patch_type = "strategic"
|
||||
patch = '{"containers":[{"name":"build","resources":{"requests":{"nvidia.com/gpu":"1"},"limits":{"nvidia.com/gpu":"1"}}}]}'
|
||||
# H100 PCIe PVCs (separate from L40S to avoid RWO conflicts)
|
||||
[[runners.kubernetes.volumes.pvc]]
|
||||
name = "training-data-h100-pvc"
|
||||
mount_path = "/mnt/training-data"
|
||||
read_only = true
|
||||
[[runners.kubernetes.volumes.pvc]]
|
||||
name = "sccache-h100-pvc"
|
||||
mount_path = "/mnt/sccache"
|
||||
read_only = false
|
||||
|
||||
tags: "kapsule,h100"
|
||||
|
||||
concurrent: 2
|
||||
@@ -1,69 +0,0 @@
|
||||
# GitLab Runner — 2×H100 GPU training workloads (multi-GPU hyperopt)
|
||||
# Runs on ci-training-h100x2 pool (2× H100 PCIe: 160GB total VRAM)
|
||||
# Multi-GPU auto-detected by MultiGpuConfig::detect() in ml crate.
|
||||
|
||||
gitlabUrl: http://gitlab-webservice-default.foxhunt.svc.cluster.local:8181
|
||||
|
||||
replicas: 1
|
||||
|
||||
rbac:
|
||||
create: false
|
||||
serviceAccount:
|
||||
create: false
|
||||
name: gitlab-runner
|
||||
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: infra
|
||||
|
||||
runners:
|
||||
cloneUrl: http://gitlab-webservice-default.foxhunt.svc.cluster.local:8181
|
||||
config: |
|
||||
[[runners]]
|
||||
clone_url = "http://gitlab-webservice-default.foxhunt.svc.cluster.local:8181"
|
||||
tag_list = ["kapsule", "h100x2"]
|
||||
[runners.kubernetes]
|
||||
namespace = "foxhunt"
|
||||
service_account = "gitlab-runner"
|
||||
image = "rust:1.89-slim"
|
||||
privileged = false
|
||||
node_selector_overwrite_allowed = ".*"
|
||||
cpu_request_overwrite_max_allowed = "16000m"
|
||||
cpu_limit_overwrite_max_allowed = "16000m"
|
||||
memory_request_overwrite_max_allowed = "192Gi"
|
||||
memory_limit_overwrite_max_allowed = "192Gi"
|
||||
poll_timeout = 600
|
||||
runtime_class_name = "nvidia"
|
||||
pod_annotations_overwrite_allowed = ".*"
|
||||
cpu_request = "4000m"
|
||||
cpu_limit = "8000m"
|
||||
memory_request = "8Gi"
|
||||
memory_limit = "16Gi"
|
||||
helper_cpu_request = "100m"
|
||||
helper_cpu_limit = "500m"
|
||||
helper_memory_request = "128Mi"
|
||||
helper_memory_limit = "512Mi"
|
||||
image_pull_secrets = ["gitlab-registry"]
|
||||
[runners.kubernetes.node_selector]
|
||||
"k8s.scaleway.com/pool-name" = "ci-training-h100x2"
|
||||
[runners.kubernetes.node_tolerations]
|
||||
"nvidia.com/gpu" = "NoSchedule"
|
||||
"node.cilium.io/agent-not-ready" = "NoSchedule"
|
||||
[runners.kubernetes.pod_labels]
|
||||
"app.kubernetes.io/part-of" = "foxhunt-ci"
|
||||
# Request both GPUs via K8s scheduler for exclusive access
|
||||
[[runners.kubernetes.pod_spec]]
|
||||
name = "build"
|
||||
patch_type = "strategic"
|
||||
patch = '{"containers":[{"name":"build","resources":{"requests":{"nvidia.com/gpu":"2"},"limits":{"nvidia.com/gpu":"2"}}}]}'
|
||||
[[runners.kubernetes.volumes.pvc]]
|
||||
name = "training-data-h100x2-pvc"
|
||||
mount_path = "/mnt/training-data"
|
||||
read_only = true
|
||||
[[runners.kubernetes.volumes.pvc]]
|
||||
name = "sccache-h100x2-pvc"
|
||||
mount_path = "/mnt/sccache"
|
||||
read_only = false
|
||||
|
||||
tags: "kapsule,h100x2"
|
||||
|
||||
concurrent: 2
|
||||
@@ -1,76 +0,0 @@
|
||||
# GitLab Runner — RL training workloads (DQN, PPO)
|
||||
# Runs on ci-training pool (L40S: 48GB VRAM)
|
||||
# Mounts separate PVCs to avoid RWO conflicts with main runner
|
||||
|
||||
gitlabUrl: http://gitlab-webservice-default.foxhunt.svc.cluster.local:8181
|
||||
# runnerToken set via --set at install time
|
||||
|
||||
replicas: 1
|
||||
|
||||
# Reuse the existing gitlab-runner SA (has pods/secrets/configmaps RBAC)
|
||||
rbac:
|
||||
create: false
|
||||
serviceAccount:
|
||||
create: false
|
||||
name: gitlab-runner
|
||||
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: infra
|
||||
|
||||
runners:
|
||||
cloneUrl: http://gitlab-webservice-default.foxhunt.svc.cluster.local:8181
|
||||
config: |
|
||||
[[runners]]
|
||||
clone_url = "http://gitlab-webservice-default.foxhunt.svc.cluster.local:8181"
|
||||
tag_list = ["kapsule", "l40s"]
|
||||
[runners.kubernetes]
|
||||
namespace = "foxhunt"
|
||||
service_account = "gitlab-runner"
|
||||
image = "rust:1.89-slim"
|
||||
privileged = false
|
||||
node_selector_overwrite_allowed = ".*"
|
||||
# Hyperopt RL jobs need 6000m CPU + 20Gi mem (PSO parallel trials)
|
||||
cpu_request_overwrite_max_allowed = "8000m"
|
||||
cpu_limit_overwrite_max_allowed = "8000m"
|
||||
memory_request_overwrite_max_allowed = "48Gi"
|
||||
memory_limit_overwrite_max_allowed = "48Gi"
|
||||
poll_timeout = 600
|
||||
# All RL runner jobs target GPU nodes → set nvidia runtime globally
|
||||
runtime_class_name = "nvidia"
|
||||
# Allow CI jobs to set pod annotations (training jobs expose Prometheus metrics)
|
||||
pod_annotations_overwrite_allowed = ".*"
|
||||
# Default resources for RL training on L40S
|
||||
cpu_request = "2000m"
|
||||
cpu_limit = "3800m"
|
||||
memory_request = "4Gi"
|
||||
memory_limit = "8Gi"
|
||||
helper_cpu_request = "100m"
|
||||
helper_cpu_limit = "500m"
|
||||
helper_memory_request = "128Mi"
|
||||
helper_memory_limit = "512Mi"
|
||||
image_pull_secrets = ["gitlab-registry"]
|
||||
[runners.kubernetes.node_selector]
|
||||
"k8s.scaleway.com/pool-name" = "ci-training"
|
||||
[runners.kubernetes.node_tolerations]
|
||||
"nvidia.com/gpu" = "NoSchedule"
|
||||
"node.cilium.io/agent-not-ready" = "NoSchedule"
|
||||
[runners.kubernetes.pod_labels]
|
||||
"app.kubernetes.io/part-of" = "foxhunt-ci"
|
||||
# Request GPU via K8s scheduler so only one training pod runs per GPU
|
||||
[[runners.kubernetes.pod_spec]]
|
||||
name = "build"
|
||||
patch_type = "strategic"
|
||||
patch = '{"containers":[{"name":"build","resources":{"requests":{"nvidia.com/gpu":"1"},"limits":{"nvidia.com/gpu":"1"}}}]}'
|
||||
# Training PVCs — separate from main runner to avoid RWO Multi-Attach errors
|
||||
[[runners.kubernetes.volumes.pvc]]
|
||||
name = "training-data-l4-pvc"
|
||||
mount_path = "/mnt/training-data"
|
||||
read_only = true
|
||||
[[runners.kubernetes.volumes.pvc]]
|
||||
name = "sccache-l4-pvc"
|
||||
mount_path = "/mnt/sccache"
|
||||
read_only = false
|
||||
|
||||
tags: "kapsule,l40s"
|
||||
|
||||
concurrent: 2
|
||||
@@ -1,86 +0,0 @@
|
||||
# GitLab Runner — Kubernetes executor
|
||||
# Runner manager pod lives on gitlab node pool
|
||||
# Default: build pods spawn on ci-compile-cpu pool (POP2-32C-128G)
|
||||
# CI jobs override node_selector for their target pool
|
||||
|
||||
gitlabUrl: http://gitlab-webservice-default.foxhunt.svc.cluster.local:8181
|
||||
# runnerToken set via --set at install time
|
||||
|
||||
replicas: 1
|
||||
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: platform
|
||||
|
||||
runners:
|
||||
# Override clone URL to internal service (pods can't reach Tailscale IPs)
|
||||
cloneUrl: http://gitlab-webservice-default.foxhunt.svc.cluster.local:8181
|
||||
config: |
|
||||
[[runners]]
|
||||
clone_url = "http://gitlab-webservice-default.foxhunt.svc.cluster.local:8181"
|
||||
tag_list = ["kapsule", "rust", "docker", "gpu"]
|
||||
[runners.kubernetes]
|
||||
namespace = "foxhunt"
|
||||
image = "rust:1.89-slim"
|
||||
privileged = false
|
||||
# Allow CI jobs to override node selector via KUBERNETES_NODE_SELECTOR_* vars
|
||||
# Format: KUBERNETES_NODE_SELECTOR_<LABEL>: "key=value"
|
||||
node_selector_overwrite_allowed = ".*"
|
||||
# Allow CI jobs to override resources via KUBERNETES_*_{REQUEST,LIMIT} vars
|
||||
# Values are maximums (not regexes) — training can request up to 80Gi
|
||||
cpu_request_overwrite_max_allowed = "24000m"
|
||||
cpu_limit_overwrite_max_allowed = "24000m"
|
||||
memory_request_overwrite_max_allowed = "80Gi"
|
||||
memory_limit_overwrite_max_allowed = "80Gi"
|
||||
# Scale-to-zero pools need ~3-5 min to provision; default 180s times out
|
||||
poll_timeout = 600
|
||||
# Allow CI jobs to override runtimeClassName via KUBERNETES_RUNTIME_CLASS_NAME
|
||||
# GPU jobs set "nvidia"; compile/kaniko jobs leave unset (default runc)
|
||||
runtime_class_name_overwrite_allowed = ".*"
|
||||
# Allow CI jobs to set pod annotations (training jobs expose Prometheus metrics)
|
||||
pod_annotations_overwrite_allowed = ".*"
|
||||
# Default resource limits for ci-compile-cpu (POP2-32C-128G)
|
||||
cpu_request = "3500m"
|
||||
cpu_limit = "7800m"
|
||||
memory_request = "12Gi"
|
||||
memory_limit = "28Gi"
|
||||
helper_cpu_request = "100m"
|
||||
helper_cpu_limit = "500m"
|
||||
helper_memory_request = "128Mi"
|
||||
helper_memory_limit = "512Mi"
|
||||
image_pull_secrets = ["gitlab-registry"]
|
||||
# Sub-tables must come AFTER all scalar values (TOML rule)
|
||||
[runners.kubernetes.node_selector]
|
||||
"k8s.scaleway.com/pool-name" = "ci-compile-cpu"
|
||||
[runners.kubernetes.node_tolerations]
|
||||
"nvidia.com/gpu" = "NoSchedule"
|
||||
# Cilium CNI agent takes ~30s to initialize on fresh scale-from-zero nodes
|
||||
"node.cilium.io/agent-not-ready" = "NoSchedule"
|
||||
[runners.kubernetes.pod_labels]
|
||||
"app.kubernetes.io/part-of" = "foxhunt-ci"
|
||||
# Mount training data PVC (Databento futures .dbn.zst files)
|
||||
# Path must NOT be under /data — Redis image's WORKDIR is /data and entrypoint chowns it
|
||||
[[runners.kubernetes.volumes.pvc]]
|
||||
name = "training-data-pvc"
|
||||
mount_path = "/mnt/training-data"
|
||||
read_only = true
|
||||
|
||||
# Runner tags for job matching
|
||||
tags: "kapsule,rust,docker,gpu"
|
||||
|
||||
# Concurrency — 10 allows all 9 Kaniko builds + test to run in parallel
|
||||
concurrent: 10
|
||||
|
||||
# RBAC for runner to spawn pods
|
||||
rbac:
|
||||
create: true
|
||||
rules:
|
||||
- apiGroups: [""]
|
||||
resources: ["pods", "pods/exec", "pods/log", "secrets", "configmaps"]
|
||||
verbs: ["get", "list", "watch", "create", "delete", "update", "patch"]
|
||||
- apiGroups: [""]
|
||||
resources: ["pods/attach"]
|
||||
verbs: ["create", "get"]
|
||||
|
||||
serviceAccount:
|
||||
create: true
|
||||
name: gitlab-runner
|
||||
@@ -75,10 +75,10 @@ spec:
|
||||
- name: ssh-proxy
|
||||
image: alpine/socat:latest
|
||||
args:
|
||||
- "TCP-LISTEN:2222,fork,reuseaddr,sndbuf=1048576,rcvbuf=1048576"
|
||||
- "TCP:gitlab-gitlab-shell.foxhunt.svc.cluster.local:2222,sndbuf=1048576,rcvbuf=1048576"
|
||||
- "TCP-LISTEN:22,fork,reuseaddr,sndbuf=1048576,rcvbuf=1048576"
|
||||
- "TCP:gitea-sshd.foxhunt.svc.cluster.local:22,sndbuf=1048576,rcvbuf=1048576"
|
||||
ports:
|
||||
- containerPort: 2222
|
||||
- containerPort: 22
|
||||
resources:
|
||||
requests:
|
||||
cpu: 25m
|
||||
@@ -152,7 +152,7 @@ data:
|
||||
return 301 https://$host$request_uri;
|
||||
}
|
||||
|
||||
# GitLab — git.fxhnt.ai
|
||||
# Gitea — git.fxhnt.ai (replaced GitLab, Phase 2B cutover)
|
||||
server {
|
||||
listen 443 ssl;
|
||||
server_name git.fxhnt.ai;
|
||||
@@ -164,7 +164,7 @@ data:
|
||||
client_max_body_size 0;
|
||||
|
||||
location / {
|
||||
proxy_pass http://gitlab-webservice-default.foxhunt.svc.cluster.local:8181;
|
||||
proxy_pass http://gitea-web.foxhunt.svc.cluster.local:3000;
|
||||
proxy_set_header Host $http_host;
|
||||
proxy_set_header X-Real-IP $remote_addr;
|
||||
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
|
||||
@@ -226,7 +226,11 @@ data:
|
||||
}
|
||||
}
|
||||
|
||||
# Web Dashboard — dashboard.fxhnt.ai
|
||||
# fxhnt cockpit — dashboard.fxhnt.ai
|
||||
# Proxies to the cockpit's Tailscale node (peer-to-peer over the tailnet) rather than the cluster Service:
|
||||
# the pod CIDR 100.64.0.0/15 overlaps Tailscale CGNAT, so this kernel-mode proxy can't reach platform-pool
|
||||
# pods via ClusterIP, but it CAN reach another tailnet node. IP is stable while the cockpit's TS state
|
||||
# secret (fxhnt-dashboard-ts-state) persists. (Was web-dashboard.foxhunt.svc — replaced by the fxhnt cockpit.)
|
||||
server {
|
||||
listen 443 ssl;
|
||||
server_name dashboard.fxhnt.ai;
|
||||
@@ -236,7 +240,7 @@ data:
|
||||
ssl_protocols TLSv1.2 TLSv1.3;
|
||||
|
||||
location / {
|
||||
proxy_pass http://web-dashboard.foxhunt.svc.cluster.local:80;
|
||||
proxy_pass http://100.81.150.18:80; # fxhnt-dashboard tailnet node (TCP:80 -> cockpit :8080)
|
||||
proxy_set_header Host $http_host;
|
||||
proxy_set_header X-Real-IP $remote_addr;
|
||||
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
|
||||
|
||||
@@ -57,6 +57,16 @@ global:
|
||||
connection:
|
||||
secret: gitlab-s3-credentials
|
||||
key: connection
|
||||
# Terraform/OpenTofu state object storage. MUST be present — without it the
|
||||
# Terraform::StateUploader has no object store and every TF-state API call 403s
|
||||
# ("Object Storage is not enabled for Terraform::StateUploader"). State files already
|
||||
# live in foxhunt-gitlab-artifacts (6b/86/<sha256(project_id)>/<state>/<ver>.tfstate).
|
||||
terraformState:
|
||||
enabled: true
|
||||
bucket: foxhunt-gitlab-artifacts
|
||||
connection:
|
||||
secret: gitlab-s3-credentials
|
||||
key: connection
|
||||
gitlab_kas:
|
||||
enabled: false
|
||||
|
||||
|
||||
@@ -1,169 +0,0 @@
|
||||
# GPU-enabled overlay for ml-training-service
|
||||
# Apply manually: kubectl apply -f infra/k8s/gpu-overlays/ml-training-service-gpu.yaml
|
||||
# Revert to CPU: kubectl apply -f infra/k8s/services/ml-training-service.yaml
|
||||
#
|
||||
# Binary fetched from MinIO at pod startup — works on any node pool.
|
||||
apiVersion: apps/v1
|
||||
kind: Deployment
|
||||
metadata:
|
||||
name: ml-training-service
|
||||
namespace: foxhunt
|
||||
labels:
|
||||
app.kubernetes.io/name: ml-training-service
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
spec:
|
||||
replicas: 1
|
||||
strategy:
|
||||
type: RollingUpdate
|
||||
rollingUpdate:
|
||||
maxSurge: 0
|
||||
maxUnavailable: 1
|
||||
selector:
|
||||
matchLabels:
|
||||
app.kubernetes.io/name: ml-training-service
|
||||
template:
|
||||
metadata:
|
||||
annotations:
|
||||
prometheus.io/scrape: "true"
|
||||
prometheus.io/port: "9094"
|
||||
prometheus.io/path: "/metrics"
|
||||
labels:
|
||||
app.kubernetes.io/name: ml-training-service
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
spec:
|
||||
serviceAccountName: ml-training-service
|
||||
securityContext:
|
||||
runAsNonRoot: true
|
||||
runAsUser: 1000
|
||||
runAsGroup: 1000
|
||||
fsGroup: 1000
|
||||
seccompProfile:
|
||||
type: RuntimeDefault
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: gpu-inference
|
||||
tolerations:
|
||||
- key: nvidia.com/gpu
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
imagePullSecrets:
|
||||
- name: gitlab-registry
|
||||
initContainers:
|
||||
- name: fetch-binary
|
||||
image: gitlab-registry.foxhunt.svc.cluster.local:5000/root/foxhunt/foxhunt-runtime:latest
|
||||
command: ["/bin/sh", "-c"]
|
||||
args:
|
||||
- |
|
||||
set -e
|
||||
rclone copyto \
|
||||
":s3:foxhunt-binaries/services/ml-training-service" \
|
||||
"/binaries/ml-training-service" \
|
||||
--s3-provider=Minio \
|
||||
--s3-endpoint=http://minio.foxhunt.svc.cluster.local:9000 \
|
||||
--s3-access-key-id="${MINIO_ACCESS_KEY}" \
|
||||
--s3-secret-access-key="${MINIO_SECRET_KEY}" \
|
||||
--s3-no-check-bucket
|
||||
chmod +x /binaries/ml-training-service
|
||||
echo "Fetched ml-training-service ($(stat -c%s /binaries/ml-training-service) bytes)"
|
||||
env:
|
||||
- name: MINIO_ACCESS_KEY
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: minio-credentials
|
||||
key: access-key
|
||||
- name: MINIO_SECRET_KEY
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: minio-credentials
|
||||
key: secret-key
|
||||
volumeMounts:
|
||||
- name: binaries
|
||||
mountPath: /binaries
|
||||
resources:
|
||||
requests:
|
||||
cpu: 100m
|
||||
memory: 64Mi
|
||||
limits:
|
||||
cpu: 500m
|
||||
memory: 128Mi
|
||||
containers:
|
||||
- name: ml-training-service
|
||||
image: gitlab-registry.foxhunt.svc.cluster.local:5000/root/foxhunt/foxhunt-runtime:latest
|
||||
securityContext:
|
||||
allowPrivilegeEscalation: false
|
||||
readOnlyRootFilesystem: true
|
||||
capabilities:
|
||||
drop: ["ALL"]
|
||||
command: ["/binaries/ml-training-service", "serve"]
|
||||
ports:
|
||||
- containerPort: 50053
|
||||
name: grpc
|
||||
- containerPort: 9094
|
||||
name: metrics
|
||||
env:
|
||||
- name: DATABASE_PASSWORD
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: db-credentials
|
||||
key: password
|
||||
- name: DATABASE_URL
|
||||
value: "postgresql://foxhunt:$(DATABASE_PASSWORD)@postgres:5432/foxhunt"
|
||||
- name: REDIS_URL
|
||||
value: "redis://redis:6379"
|
||||
- name: JWT_SECRET
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: jwt-secret
|
||||
key: secret
|
||||
- name: JWT_ISSUER
|
||||
value: foxhunt-api
|
||||
- name: JWT_AUDIENCE
|
||||
value: foxhunt-services
|
||||
- name: S3_ENDPOINT
|
||||
value: "http://minio.foxhunt.svc.cluster.local:9000"
|
||||
- name: S3_BUCKET
|
||||
value: foxhunt-models
|
||||
- name: ENABLE_GPU
|
||||
value: "true"
|
||||
- name: RUST_LOG
|
||||
value: info
|
||||
- name: OTEL_EXPORTER_OTLP_ENDPOINT
|
||||
value: "http://tempo.foxhunt.svc.cluster.local:4317"
|
||||
volumeMounts:
|
||||
- name: binaries
|
||||
mountPath: /binaries
|
||||
readOnly: true
|
||||
- name: tls-certs
|
||||
mountPath: /app/certs/ml-training-service
|
||||
readOnly: true
|
||||
- name: tmp
|
||||
mountPath: /tmp
|
||||
readinessProbe:
|
||||
tcpSocket:
|
||||
port: 50053
|
||||
initialDelaySeconds: 15
|
||||
periodSeconds: 10
|
||||
livenessProbe:
|
||||
tcpSocket:
|
||||
port: 50053
|
||||
initialDelaySeconds: 30
|
||||
periodSeconds: 15
|
||||
failureThreshold: 5
|
||||
resources:
|
||||
requests:
|
||||
nvidia.com/gpu: "1"
|
||||
cpu: "1"
|
||||
memory: 2Gi
|
||||
limits:
|
||||
nvidia.com/gpu: "1"
|
||||
cpu: "4"
|
||||
memory: 8Gi
|
||||
volumes:
|
||||
- name: binaries
|
||||
emptyDir:
|
||||
sizeLimit: 200Mi
|
||||
- name: tls-certs
|
||||
secret:
|
||||
secretName: ml-training-tls
|
||||
- name: tmp
|
||||
emptyDir:
|
||||
sizeLimit: 50Mi
|
||||
@@ -1,156 +0,0 @@
|
||||
# GPU-enabled overlay for trading-service
|
||||
# Apply manually: kubectl apply -f infra/k8s/gpu-overlays/trading-service-gpu.yaml
|
||||
# Revert to CPU: kubectl apply -f infra/k8s/services/trading-service.yaml
|
||||
#
|
||||
# Binary fetched from MinIO at pod startup — works on any node pool.
|
||||
apiVersion: apps/v1
|
||||
kind: Deployment
|
||||
metadata:
|
||||
name: trading-service
|
||||
namespace: foxhunt
|
||||
labels:
|
||||
app.kubernetes.io/name: trading-service
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
spec:
|
||||
replicas: 1
|
||||
strategy:
|
||||
type: RollingUpdate
|
||||
rollingUpdate:
|
||||
maxSurge: 0
|
||||
maxUnavailable: 1
|
||||
selector:
|
||||
matchLabels:
|
||||
app.kubernetes.io/name: trading-service
|
||||
template:
|
||||
metadata:
|
||||
annotations:
|
||||
prometheus.io/scrape: "true"
|
||||
prometheus.io/port: "9092"
|
||||
prometheus.io/path: "/metrics"
|
||||
labels:
|
||||
app.kubernetes.io/name: trading-service
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
spec:
|
||||
securityContext:
|
||||
runAsNonRoot: true
|
||||
runAsUser: 1000
|
||||
runAsGroup: 1000
|
||||
fsGroup: 1000
|
||||
seccompProfile:
|
||||
type: RuntimeDefault
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: l40s
|
||||
tolerations:
|
||||
- key: nvidia.com/gpu
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
imagePullSecrets:
|
||||
- name: gitlab-registry
|
||||
initContainers:
|
||||
- name: fetch-binary
|
||||
image: gitlab-registry.foxhunt.svc.cluster.local:5000/root/foxhunt/foxhunt-runtime:latest
|
||||
command: ["/bin/sh", "-c"]
|
||||
args:
|
||||
- |
|
||||
set -e
|
||||
rclone copyto \
|
||||
":s3:foxhunt-binaries/services/trading-service" \
|
||||
"/binaries/trading-service" \
|
||||
--s3-provider=Minio \
|
||||
--s3-endpoint=http://minio.foxhunt.svc.cluster.local:9000 \
|
||||
--s3-access-key-id="${MINIO_ACCESS_KEY}" \
|
||||
--s3-secret-access-key="${MINIO_SECRET_KEY}" \
|
||||
--s3-no-check-bucket
|
||||
chmod +x /binaries/trading-service
|
||||
echo "Fetched trading-service ($(stat -c%s /binaries/trading-service) bytes)"
|
||||
env:
|
||||
- name: MINIO_ACCESS_KEY
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: minio-credentials
|
||||
key: access-key
|
||||
- name: MINIO_SECRET_KEY
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: minio-credentials
|
||||
key: secret-key
|
||||
volumeMounts:
|
||||
- name: binaries
|
||||
mountPath: /binaries
|
||||
resources:
|
||||
requests:
|
||||
cpu: 100m
|
||||
memory: 64Mi
|
||||
limits:
|
||||
cpu: 500m
|
||||
memory: 128Mi
|
||||
containers:
|
||||
- name: trading-service
|
||||
image: gitlab-registry.foxhunt.svc.cluster.local:5000/root/foxhunt/foxhunt-runtime:latest
|
||||
securityContext:
|
||||
allowPrivilegeEscalation: false
|
||||
readOnlyRootFilesystem: true
|
||||
capabilities:
|
||||
drop: ["ALL"]
|
||||
command: ["/binaries/trading-service"]
|
||||
ports:
|
||||
- containerPort: 50051
|
||||
name: grpc
|
||||
- containerPort: 9092
|
||||
name: metrics
|
||||
env:
|
||||
- name: DATABASE_PASSWORD
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: db-credentials
|
||||
key: password
|
||||
- name: DATABASE_URL
|
||||
value: "postgresql://foxhunt:$(DATABASE_PASSWORD)@postgres:5432/foxhunt"
|
||||
- name: REDIS_URL
|
||||
value: "redis://redis:6379"
|
||||
- name: JWT_SECRET
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: jwt-secret
|
||||
key: secret
|
||||
- name: JWT_ISSUER
|
||||
value: foxhunt-api
|
||||
- name: JWT_AUDIENCE
|
||||
value: foxhunt-services
|
||||
- name: QUESTDB_ILP_HOST
|
||||
value: "questdb:9009"
|
||||
- name: GRPC_PORT
|
||||
value: "50051"
|
||||
- name: ENABLE_GPU_INFERENCE
|
||||
value: "true"
|
||||
- name: RUST_LOG
|
||||
value: info
|
||||
volumeMounts:
|
||||
- name: binaries
|
||||
mountPath: /binaries
|
||||
readOnly: true
|
||||
- name: tmp
|
||||
mountPath: /tmp
|
||||
readinessProbe:
|
||||
exec:
|
||||
command:
|
||||
- grpc_health_probe
|
||||
- -addr=localhost:50051
|
||||
initialDelaySeconds: 10
|
||||
periodSeconds: 10
|
||||
resources:
|
||||
requests:
|
||||
nvidia.com/gpu: "1"
|
||||
cpu: "1"
|
||||
memory: 2Gi
|
||||
limits:
|
||||
nvidia.com/gpu: "1"
|
||||
cpu: "4"
|
||||
memory: 8Gi
|
||||
volumes:
|
||||
- name: binaries
|
||||
emptyDir:
|
||||
sizeLimit: 200Mi
|
||||
- name: tmp
|
||||
emptyDir:
|
||||
sizeLimit: 50Mi
|
||||
@@ -18,6 +18,9 @@ spec:
|
||||
- podSelector:
|
||||
matchLabels:
|
||||
app.kubernetes.io/name: broker-gateway
|
||||
- podSelector:
|
||||
matchLabels:
|
||||
app: multistrat-rebalancer
|
||||
ports:
|
||||
- protocol: TCP
|
||||
port: 4002
|
||||
|
||||
@@ -1,4 +1,4 @@
|
||||
# Postgres: accepts connections from foxhunt app pods, GitLab, and Grafana
|
||||
# Postgres: accepts connections from foxhunt app pods, GitLab, Grafana, and Gitea
|
||||
apiVersion: networking.k8s.io/v1
|
||||
kind: NetworkPolicy
|
||||
metadata:
|
||||
@@ -23,6 +23,9 @@ spec:
|
||||
- podSelector:
|
||||
matchLabels:
|
||||
app.kubernetes.io/name: grafana
|
||||
- podSelector: # Gitea (Phase 2B) reuses the in-cluster postgres
|
||||
matchLabels:
|
||||
app.kubernetes.io/name: gitea
|
||||
ports:
|
||||
- protocol: TCP
|
||||
port: 5432
|
||||
|
||||
123
infra/k8s/services/multistrat-rebalancer.yaml
Normal file
123
infra/k8s/services/multistrat-rebalancer.yaml
Normal file
@@ -0,0 +1,123 @@
|
||||
# Multi-strat rebalancer — weekly CronJob running the adaptive book on IBKR (paper) via the
|
||||
# in-cluster ib-gateway. Production deployment of scripts/surfer/multistrat_bot.py.
|
||||
#
|
||||
# Code is delivered via ConfigMap (no image build needed — lean). Create/refresh it with:
|
||||
# kubectl create configmap multistrat-bot-code -n foxhunt \
|
||||
# --from-file=scripts/surfer/multistrat_bot.py \
|
||||
# --from-file=scripts/surfer/multistrat_paper.py \
|
||||
# --dry-run=client -o yaml | kubectl apply -f -
|
||||
#
|
||||
# SAFETY: MULTISTRAT_EXECUTE defaults to "false" (cluster dry-run). After validating a few dry-run
|
||||
# CronJob logs, flip to "true" to place paper orders. Account is paper (ib-gateway TRADING_MODE=paper);
|
||||
# the bot additionally refuses any non-DU (live) account unless MULTISTRAT_ALLOW_LIVE_CONFIRMED is set.
|
||||
---
|
||||
apiVersion: v1
|
||||
kind: PersistentVolumeClaim
|
||||
metadata:
|
||||
name: multistrat-state
|
||||
namespace: foxhunt
|
||||
labels:
|
||||
app: multistrat-rebalancer
|
||||
spec:
|
||||
accessModes: ["ReadWriteOnce"] # sequential CronJob runs (Forbid concurrency) — RWO is fine
|
||||
resources:
|
||||
requests:
|
||||
storage: 1Gi
|
||||
---
|
||||
apiVersion: networking.k8s.io/v1
|
||||
kind: NetworkPolicy
|
||||
metadata:
|
||||
name: multistrat-rebalancer
|
||||
namespace: foxhunt
|
||||
spec:
|
||||
podSelector:
|
||||
matchLabels:
|
||||
app: multistrat-rebalancer
|
||||
policyTypes: [Egress]
|
||||
egress:
|
||||
- to: [] # DNS
|
||||
ports:
|
||||
- {protocol: UDP, port: 53}
|
||||
- {protocol: TCP, port: 53}
|
||||
- to: # in-cluster ib-gateway (4002 API + 4004 socat bridge for pod-to-pod)
|
||||
- podSelector:
|
||||
matchLabels:
|
||||
app: ib-gateway
|
||||
ports:
|
||||
- {protocol: TCP, port: 4002}
|
||||
- {protocol: TCP, port: 4004}
|
||||
- to: # internet 443: PyPI (pip) + Yahoo (prices), no internal ranges
|
||||
- ipBlock:
|
||||
cidr: 0.0.0.0/0
|
||||
except: ["10.0.0.0/8", "172.16.0.0/12", "192.168.0.0/16"]
|
||||
ports:
|
||||
- {protocol: TCP, port: 443}
|
||||
---
|
||||
apiVersion: batch/v1
|
||||
kind: CronJob
|
||||
metadata:
|
||||
name: multistrat-rebalancer
|
||||
namespace: foxhunt
|
||||
labels:
|
||||
app: multistrat-rebalancer
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
spec:
|
||||
schedule: "35 14 * * 1-5" # weekdays 14:35 UTC (~1h after US open) — daily risk re-eval; hysteresis prevents churn
|
||||
concurrencyPolicy: Forbid
|
||||
startingDeadlineSeconds: 3600
|
||||
successfulJobsHistoryLimit: 5
|
||||
failedJobsHistoryLimit: 10
|
||||
jobTemplate:
|
||||
spec:
|
||||
backoffLimit: 1
|
||||
activeDeadlineSeconds: 600
|
||||
template:
|
||||
metadata:
|
||||
labels:
|
||||
app: multistrat-rebalancer
|
||||
spec:
|
||||
restartPolicy: Never
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: platform
|
||||
securityContext:
|
||||
seccompProfile:
|
||||
type: RuntimeDefault
|
||||
containers:
|
||||
- name: rebalancer
|
||||
image: python:3.12-slim
|
||||
command: ["/bin/sh", "-c"]
|
||||
args:
|
||||
- >-
|
||||
pip install --quiet --no-cache-dir ib_async==2.1.0 numpy &&
|
||||
python /app/multistrat_bot.py run
|
||||
env:
|
||||
- {name: IB_HOST, value: "ib-gateway"}
|
||||
- {name: IB_PORT, value: "4004"} # socat bridge: relays via localhost so IB Gateway trusts it (4002 = pod-IP, untrusted -> handshake timeout)
|
||||
- {name: IB_CLIENT_ID, value: "11"}
|
||||
- {name: MULTISTRAT_MAXLEV, value: "1.0"}
|
||||
- {name: MULTISTRAT_HYST, value: "0.03"}
|
||||
- {name: MULTISTRAT_REBALANCE_DAYS, value: "1"} # daily eval (weekday cron); hysteresis (3%) gates actual trades
|
||||
- {name: MULTISTRAT_DD_HALT, value: "0.20"}
|
||||
- {name: MULTISTRAT_MAX_ORDER, value: "0.30"}
|
||||
- {name: MULTISTRAT_EXECUTE, value: "true"} # ARMED: places paper orders (account DU* paper-guarded; live needs MULTISTRAT_ALLOW_LIVE_CONFIRMED)
|
||||
- {name: MULTISTRAT_STATE, value: "/data/state.json"}
|
||||
- {name: MULTISTRAT_LOG, value: "/data/bot.log"}
|
||||
- {name: PIP_DISABLE_PIP_VERSION_CHECK, value: "1"}
|
||||
- {name: PYTHONUNBUFFERED, value: "1"}
|
||||
securityContext:
|
||||
allowPrivilegeEscalation: false
|
||||
capabilities:
|
||||
drop: ["ALL"]
|
||||
resources:
|
||||
requests: {memory: "256Mi", cpu: "100m"}
|
||||
limits: {memory: "512Mi", cpu: "500m"}
|
||||
volumeMounts:
|
||||
- {mountPath: /app, name: code, readOnly: true}
|
||||
- {mountPath: /data, name: state}
|
||||
volumes:
|
||||
- name: code
|
||||
configMap:
|
||||
name: multistrat-bot-code
|
||||
- name: state
|
||||
persistentVolumeClaim:
|
||||
claimName: multistrat-state
|
||||
@@ -1,60 +0,0 @@
|
||||
# DaemonSet image pre-puller — keeps training images cached on GPU nodes
|
||||
# Runs on ci-training (L40S) pool so training jobs skip the pull.
|
||||
# Init containers pull :latest tags, then the main container sleeps forever.
|
||||
---
|
||||
apiVersion: apps/v1
|
||||
kind: DaemonSet
|
||||
metadata:
|
||||
name: image-prepuller
|
||||
namespace: foxhunt
|
||||
labels:
|
||||
app: image-prepuller
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
spec:
|
||||
selector:
|
||||
matchLabels:
|
||||
app: image-prepuller
|
||||
template:
|
||||
metadata:
|
||||
labels:
|
||||
app: image-prepuller
|
||||
spec:
|
||||
affinity:
|
||||
nodeAffinity:
|
||||
requiredDuringSchedulingIgnoredDuringExecution:
|
||||
nodeSelectorTerms:
|
||||
- matchExpressions:
|
||||
- key: k8s.scaleway.com/pool-name
|
||||
operator: In
|
||||
values:
|
||||
- ci-training
|
||||
tolerations:
|
||||
- key: nvidia.com/gpu
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
- key: node.cilium.io/agent-not-ready
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
imagePullSecrets:
|
||||
- name: gitlab-registry
|
||||
initContainers:
|
||||
- name: pull-training-runtime
|
||||
image: gitlab-registry.foxhunt.svc.cluster.local:5000/root/foxhunt/foxhunt-training-runtime:latest
|
||||
command: ["echo", "foxhunt-training-runtime image pulled"]
|
||||
resources:
|
||||
requests:
|
||||
cpu: 10m
|
||||
memory: 16Mi
|
||||
limits:
|
||||
cpu: 10m
|
||||
memory: 16Mi
|
||||
containers:
|
||||
- name: pause
|
||||
image: registry.k8s.io/pause:3.10
|
||||
resources:
|
||||
requests:
|
||||
cpu: 10m
|
||||
memory: 16Mi
|
||||
limits:
|
||||
cpu: 10m
|
||||
memory: 16Mi
|
||||
@@ -1,130 +0,0 @@
|
||||
apiVersion: batch/v1
|
||||
kind: Job
|
||||
metadata:
|
||||
generateName: training-
|
||||
namespace: foxhunt
|
||||
labels:
|
||||
app.kubernetes.io/name: training
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
foxhunt/job-type: training
|
||||
spec:
|
||||
backoffLimit: 1
|
||||
activeDeadlineSeconds: 21600 # 6 hours — hyperopt runs 20 trials × 8 epochs
|
||||
ttlSecondsAfterFinished: 600
|
||||
template:
|
||||
metadata:
|
||||
annotations:
|
||||
prometheus.io/scrape: "true"
|
||||
prometheus.io/port: "9094"
|
||||
prometheus.io/path: "/metrics"
|
||||
labels:
|
||||
app.kubernetes.io/name: training
|
||||
app.kubernetes.io/component: training-workflow
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
foxhunt/job-type: training
|
||||
spec:
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: ci-training-h100
|
||||
tolerations:
|
||||
- key: nvidia.com/gpu
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
# Cilium CNI agent takes ~30s to initialize on fresh scale-from-zero nodes
|
||||
- key: node.cilium.io/agent-not-ready
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
imagePullSecrets:
|
||||
- name: gitlab-registry
|
||||
restartPolicy: Never
|
||||
initContainers:
|
||||
# 1. Fetch training binary from GitLab Generic Package Registry
|
||||
- name: fetch-binary
|
||||
image: gitlab-registry.foxhunt.svc.cluster.local:5000/root/foxhunt/foxhunt-training-runtime:latest
|
||||
command: ["/bin/sh", "-c"]
|
||||
args:
|
||||
- |
|
||||
set -e
|
||||
BINARY="$(TRAINING_BINARY)"
|
||||
curl -fSL -o "/binaries/${BINARY}" \
|
||||
--header "DEPLOY-TOKEN: ${GITLAB_DEPLOY_TOKEN}" \
|
||||
"${GITLAB_API}/projects/1/packages/generic/foxhunt-training/${FOXHUNT_RELEASE}/${BINARY}"
|
||||
chmod +x "/binaries/${BINARY}"
|
||||
echo "Fetched ${BINARY} ${FOXHUNT_RELEASE} ($(stat -c%s /binaries/${BINARY}) bytes)"
|
||||
env:
|
||||
- name: TRAINING_BINARY
|
||||
value: train_baseline_supervised
|
||||
- name: GITLAB_DEPLOY_TOKEN
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: gitlab-deploy-token
|
||||
key: token
|
||||
- name: GITLAB_API
|
||||
value: "http://gitlab-webservice-default.foxhunt.svc.cluster.local:8181/api/v4"
|
||||
- name: FOXHUNT_RELEASE
|
||||
value: "latest"
|
||||
volumeMounts:
|
||||
- name: binaries
|
||||
mountPath: /binaries
|
||||
resources:
|
||||
requests:
|
||||
cpu: 100m
|
||||
memory: 64Mi
|
||||
limits:
|
||||
cpu: 500m
|
||||
memory: 128Mi
|
||||
containers:
|
||||
- name: training
|
||||
image: gitlab-registry.foxhunt.svc.cluster.local:5000/root/foxhunt/foxhunt-training-runtime:latest
|
||||
# Available binaries (copied by initContainer):
|
||||
# train_baseline_rl (for dqn, ppo)
|
||||
# train_baseline_supervised (for tft, mamba2, tggn, tlob, liquid, kan, xlstm, diffusion)
|
||||
# evaluate_baseline
|
||||
# hyperopt_baseline_rl (for dqn, ppo)
|
||||
# hyperopt_baseline_supervised (for tft, mamba2)
|
||||
command: ["/binaries/$(TRAINING_BINARY)"]
|
||||
args:
|
||||
- "--symbol=ES.FUT"
|
||||
- "--data-dir=/data/futures-baseline"
|
||||
- "--mbp10-data-dir=/data/futures-baseline-mbp10"
|
||||
- "--trades-data-dir=/data/futures-baseline-trades"
|
||||
- "--training-profile=$(TRAINING_PROFILE)"
|
||||
- "--output=/output"
|
||||
env:
|
||||
- name: TRAINING_BINARY
|
||||
value: train_baseline_supervised
|
||||
- name: TRAINING_PROFILE
|
||||
value: "dqn-production"
|
||||
- name: RUST_LOG
|
||||
value: info
|
||||
- name: SQLX_OFFLINE
|
||||
value: "true"
|
||||
- name: OTEL_EXPORTER_OTLP_ENDPOINT
|
||||
value: "http://tempo.foxhunt.svc.cluster.local:4317"
|
||||
volumeMounts:
|
||||
- name: training-data
|
||||
mountPath: /data
|
||||
readOnly: true
|
||||
- name: output
|
||||
mountPath: /output
|
||||
- name: binaries
|
||||
mountPath: /binaries
|
||||
readOnly: true
|
||||
resources:
|
||||
requests:
|
||||
nvidia.com/gpu: "1"
|
||||
cpu: "4"
|
||||
memory: 16Gi
|
||||
limits:
|
||||
nvidia.com/gpu: "1"
|
||||
cpu: "8"
|
||||
memory: 32Gi
|
||||
volumes:
|
||||
- name: training-data
|
||||
persistentVolumeClaim:
|
||||
claimName: training-data-pvc
|
||||
- name: output
|
||||
emptyDir:
|
||||
sizeLimit: 2Gi
|
||||
- name: binaries
|
||||
emptyDir:
|
||||
sizeLimit: 500Mi
|
||||
@@ -1,110 +0,0 @@
|
||||
# infra/k8s/training/populate-test-data-job.yaml
|
||||
apiVersion: batch/v1
|
||||
kind: Job
|
||||
metadata:
|
||||
generateName: populate-test-data-
|
||||
namespace: foxhunt
|
||||
labels:
|
||||
app.kubernetes.io/name: populate-test-data
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
spec:
|
||||
backoffLimit: 1
|
||||
ttlSecondsAfterFinished: 300
|
||||
template:
|
||||
spec:
|
||||
restartPolicy: Never
|
||||
nodeSelector:
|
||||
k8s.scaleway.com/pool-name: ci-training-h100
|
||||
tolerations:
|
||||
- key: nvidia.com/gpu
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
- key: node.cilium.io/agent-not-ready
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
containers:
|
||||
- name: populate
|
||||
image: busybox:1.37
|
||||
command: ["/bin/sh", "-c"]
|
||||
args:
|
||||
- |
|
||||
set -e
|
||||
echo "=== Populating test data PVC ==="
|
||||
|
||||
# All symbols used in CI tests (4 symbols from futures-baseline)
|
||||
SYMBOLS="ES.FUT 6E.FUT ZN.FUT NQ.FUT"
|
||||
|
||||
for SYM in $SYMBOLS; do
|
||||
echo "--- $SYM ---"
|
||||
|
||||
# OHLCV (used by DQN/PPO training tests and walk-forward validation)
|
||||
# Copy 3 quarters per symbol: enough for --train-months 3 --val-months 1 --test-months 1
|
||||
# (5 months window), perf benchmarks, and walk-forward tests. ~6MB/symbol compressed.
|
||||
mkdir -p /test-data/ohlcv/$SYM
|
||||
if ls /training-data/futures-baseline/$SYM/*.dbn.zst >/dev/null 2>&1; then
|
||||
COUNT=0
|
||||
for F in $(ls /training-data/futures-baseline/$SYM/*.dbn.zst 2>/dev/null | head -3); do
|
||||
cp "$F" /test-data/ohlcv/$SYM/
|
||||
echo " OHLCV: copied $(basename $F)"
|
||||
COUNT=$((COUNT + 1))
|
||||
done
|
||||
echo " OHLCV: $COUNT quarters copied"
|
||||
else
|
||||
echo " OHLCV: no source data for $SYM"
|
||||
fi
|
||||
|
||||
# MBP-10 (used by OFI feature loading)
|
||||
mkdir -p /test-data/mbp10/$SYM
|
||||
if ls /training-data/futures-baseline-mbp10/$SYM/*.dbn.zst >/dev/null 2>&1; then
|
||||
COUNT=0
|
||||
for F in $(ls /training-data/futures-baseline-mbp10/$SYM/*.dbn.zst 2>/dev/null | head -3); do
|
||||
cp "$F" /test-data/mbp10/$SYM/
|
||||
echo " MBP-10: copied $(basename $F)"
|
||||
COUNT=$((COUNT + 1))
|
||||
done
|
||||
else
|
||||
echo " MBP-10: no source data for $SYM"
|
||||
fi
|
||||
|
||||
# Trades
|
||||
mkdir -p /test-data/trades/$SYM
|
||||
if ls /training-data/futures-baseline-trades/$SYM/*.dbn.zst >/dev/null 2>&1; then
|
||||
COUNT=0
|
||||
for F in $(ls /training-data/futures-baseline-trades/$SYM/*.dbn.zst 2>/dev/null | head -3); do
|
||||
cp "$F" /test-data/trades/$SYM/
|
||||
echo " Trades: copied $(basename $F)"
|
||||
COUNT=$((COUNT + 1))
|
||||
done
|
||||
else
|
||||
echo " Trades: no source data for $SYM"
|
||||
fi
|
||||
done
|
||||
|
||||
echo ""
|
||||
echo "=== Test data contents ==="
|
||||
find /test-data -type f -exec ls -lh {} \;
|
||||
echo "Total: $(du -sh /test-data | cut -f1)"
|
||||
echo "=== Done ==="
|
||||
resources:
|
||||
requests:
|
||||
nvidia.com/gpu: "1"
|
||||
cpu: 100m
|
||||
memory: 128Mi
|
||||
limits:
|
||||
nvidia.com/gpu: "1"
|
||||
cpu: 500m
|
||||
memory: 256Mi
|
||||
volumeMounts:
|
||||
- name: training-data
|
||||
mountPath: /training-data
|
||||
readOnly: true
|
||||
- name: test-data
|
||||
mountPath: /test-data
|
||||
volumes:
|
||||
- name: training-data
|
||||
persistentVolumeClaim:
|
||||
claimName: training-data-pvc
|
||||
readOnly: true
|
||||
- name: test-data
|
||||
persistentVolumeClaim:
|
||||
claimName: test-data-pvc
|
||||
@@ -1,24 +0,0 @@
|
||||
# MinIO credentials for rclone output sync in training jobs.
|
||||
# No longer needed as a standalone secret — training jobs now use minio-credentials
|
||||
# (deployed via infra/k8s/minio/minio.yaml).
|
||||
#
|
||||
# If you need to recreate manually:
|
||||
# kubectl -n foxhunt create secret generic minio-credentials \
|
||||
# --from-literal=access-key=<MINIO_ACCESS_KEY> \
|
||||
# --from-literal=secret-key=<MINIO_SECRET_KEY> \
|
||||
# --from-literal=root-user=<MINIO_ACCESS_KEY> \
|
||||
# --from-literal=root-password=<MINIO_SECRET_KEY>
|
||||
apiVersion: v1
|
||||
kind: Secret
|
||||
metadata:
|
||||
name: minio-credentials
|
||||
namespace: foxhunt
|
||||
labels:
|
||||
app.kubernetes.io/name: minio-credentials
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
type: Opaque
|
||||
stringData:
|
||||
access-key: REPLACE_ME
|
||||
secret-key: REPLACE_ME
|
||||
root-user: REPLACE_ME
|
||||
root-password: REPLACE_ME
|
||||
@@ -1,15 +0,0 @@
|
||||
apiVersion: v1
|
||||
kind: PersistentVolumeClaim
|
||||
metadata:
|
||||
name: training-output-pvc
|
||||
namespace: foxhunt
|
||||
labels:
|
||||
app.kubernetes.io/name: training-output
|
||||
app.kubernetes.io/part-of: foxhunt
|
||||
spec:
|
||||
accessModes:
|
||||
- ReadWriteOnce
|
||||
resources:
|
||||
requests:
|
||||
storage: 50Gi
|
||||
storageClassName: scw-bssd
|
||||
@@ -8,6 +8,7 @@ terraform {
|
||||
|
||||
inputs = {
|
||||
dns_zone = "fxhnt.ai"
|
||||
# Tailscale IP of foxhunt-gitlab proxy pod
|
||||
git_ip = "100.90.76.85"
|
||||
# Tailscale IP of the foxhunt-gitlab proxy pod (live node foxhunt-gitlab; the old 100.90.76.85
|
||||
# was a stale prior IP — its duplicate A records across *.fxhnt.ai were removed 2026-06-21).
|
||||
git_ip = "100.95.225.27"
|
||||
}
|
||||
|
||||
@@ -14,9 +14,12 @@ inputs = {
|
||||
platform_type = "DEV1-L"
|
||||
platform_max_size = 3
|
||||
|
||||
# CPU compile pool (POP2-HC-32C-64G — 32 vCPU, 64GB RAM, high clock)
|
||||
# CI build pool — right-sized for the Python fxhnt cockpit/deploy builds (kaniko + pip, ~2-6 vCPU).
|
||||
# Was POP2-HC-32C-64G (32 vCPU/64GB) for the now-dormant Rust cargo compile + precompute_features;
|
||||
# downsized 2026-06-21 to POP2-4C-16G (in stock in fr-par-2; the 32C type hit a zone stock-out that
|
||||
# jammed the autoscaler — "Resource POP2-HC-32C-64G is out of stock"). Far cheaper, scale-to-zero.
|
||||
enable_ci_compile_cpu_pool = true
|
||||
ci_compile_cpu_type = "POP2-HC-32C-64G"
|
||||
ci_compile_cpu_type = "POP2-4C-16G"
|
||||
ci_compile_cpu_max_size = 4
|
||||
|
||||
# High-memory CPU pool (POP2-HM-32C-256G — 32 vCPU, 256GB RAM).
|
||||
@@ -24,17 +27,17 @@ inputs = {
|
||||
# the 64GB ci-compile-cpu pool (9-quarter precompute_features peaks
|
||||
# ~50-60GB during the post-OFI alpha_trades conversion).
|
||||
# min_size=0 + size=0 initial so this pool costs nothing when idle.
|
||||
enable_ci_compile_cpu_hm_pool = true
|
||||
enable_ci_compile_cpu_hm_pool = false # decommissioned 2026-06-21 — Rust precompute_features retired
|
||||
ci_compile_cpu_hm_type = "POP2-HM-32C-256G"
|
||||
ci_compile_cpu_hm_max_size = 1
|
||||
|
||||
# L40S training pool (48GB VRAM, CUDA CC 89)
|
||||
enable_ci_training_l40s_pool = true
|
||||
enable_ci_training_l40s_pool = false # decommissioned 2026-06-21 — Rust ML training retired
|
||||
ci_training_l40s_type = "L40S-1-48G"
|
||||
ci_training_l40s_max_size = 1
|
||||
|
||||
# H100 training pool (80GB VRAM, CUDA CC 90)
|
||||
enable_ci_training_h100_pool = true
|
||||
enable_ci_training_h100_pool = false # decommissioned 2026-06-21 — Rust RL training retired
|
||||
ci_training_h100_type = "H100-1-80G"
|
||||
ci_training_h100_max_size = 1
|
||||
|
||||
|
||||
@@ -13,4 +13,5 @@ dependency "kapsule" {
|
||||
inputs = {
|
||||
private_network_id = dependency.kapsule.outputs.private_network_id
|
||||
gateway_type = "VPC-GW-S"
|
||||
bastion_enabled = true # matches the live gateway (carried over from the old kapsule gw config)
|
||||
}
|
||||
|
||||
@@ -7,21 +7,31 @@ locals {
|
||||
project_id = get_env("SCW_DEFAULT_PROJECT_ID")
|
||||
}
|
||||
|
||||
# Remote state in GitLab Terraform state backend (project root/foxhunt, ID=1)
|
||||
# Token stored in k8s secret gitlab-terraform-state-token
|
||||
# Set TF_HTTP_USERNAME and TF_HTTP_PASSWORD env vars before running terragrunt.
|
||||
# Remote state in Scaleway Object Storage (bucket foxhunt-tfstate, fr-par) — migrated off GitLab 2026-06-21.
|
||||
# The s3 backend reads AWS_* env vars (set them to the Scaleway access/secret keys); SCW_* drives the provider.
|
||||
remote_state {
|
||||
backend = "http"
|
||||
backend = "s3"
|
||||
generate = {
|
||||
path = "backend.tf"
|
||||
if_exists = "overwrite"
|
||||
}
|
||||
config = {
|
||||
address = "${get_env("GITLAB_TF_STATE_URL", "https://tf.fxhnt.ai")}/api/v4/projects/1/terraform/state/${path_relative_to_include()}"
|
||||
lock_address = "${get_env("GITLAB_TF_STATE_URL", "https://tf.fxhnt.ai")}/api/v4/projects/1/terraform/state/${path_relative_to_include()}/lock"
|
||||
unlock_address = "${get_env("GITLAB_TF_STATE_URL", "https://tf.fxhnt.ai")}/api/v4/projects/1/terraform/state/${path_relative_to_include()}/lock"
|
||||
lock_method = "POST"
|
||||
unlock_method = "DELETE"
|
||||
bucket = "foxhunt-tfstate"
|
||||
key = "${path_relative_to_include()}/terraform.tfstate"
|
||||
region = "fr-par"
|
||||
|
||||
endpoints = {
|
||||
s3 = "https://s3.fr-par.scw.cloud"
|
||||
}
|
||||
|
||||
# Scaleway S3-compat: skip AWS-specific preflight calls
|
||||
skip_credentials_validation = true
|
||||
skip_region_validation = true
|
||||
skip_requesting_account_id = true
|
||||
skip_metadata_api_check = true
|
||||
|
||||
# OpenTofu-native lock (no DynamoDB)
|
||||
use_lockfile = true
|
||||
}
|
||||
}
|
||||
|
||||
|
||||
@@ -3,31 +3,11 @@ resource "scaleway_vpc_private_network" "foxhunt" {
|
||||
region = var.region
|
||||
}
|
||||
|
||||
# VPC Public Gateway — provides NAT (masquerade) for private nodes
|
||||
# and DHCP with default route propagation to prevent DNS deadlock.
|
||||
# See incident_dns_deadlock.md for why this is critical.
|
||||
resource "scaleway_vpc_public_gateway" "foxhunt" {
|
||||
name = "${var.cluster_name}-gw"
|
||||
type = "VPC-GW-S"
|
||||
zone = "${var.region}-2"
|
||||
bastion_enabled = true
|
||||
bastion_port = 61000
|
||||
}
|
||||
|
||||
resource "scaleway_vpc_public_gateway_dhcp" "foxhunt" {
|
||||
subnet = "172.16.0.0/22"
|
||||
push_default_route = true
|
||||
push_dns_server = true
|
||||
zone = "${var.region}-2"
|
||||
}
|
||||
|
||||
resource "scaleway_vpc_gateway_network" "foxhunt" {
|
||||
gateway_id = scaleway_vpc_public_gateway.foxhunt.id
|
||||
private_network_id = scaleway_vpc_private_network.foxhunt.id
|
||||
dhcp_id = scaleway_vpc_public_gateway_dhcp.foxhunt.id
|
||||
enable_masquerade = true
|
||||
zone = "${var.region}-2"
|
||||
}
|
||||
# NAT gateway (foxhunt-gw) is owned by the dedicated `public-gateway` module (IPAM mode +
|
||||
# dedicated IP), which depends on this module's private_network_id output. The legacy
|
||||
# gateway/dhcp/gateway_network resources that used to live here were a pre-refactor duplicate
|
||||
# (never in this module's state) and were removed 2026-06-21 to stop `terragrunt apply` here
|
||||
# from trying to create a second gateway. See reference_ci_pool_and_tfstate / incident_dns_deadlock.
|
||||
|
||||
resource "scaleway_k8s_cluster" "foxhunt" {
|
||||
name = var.cluster_name
|
||||
|
||||
41
scripts/install_torch_gpu.sh
Normal file
41
scripts/install_torch_gpu.sh
Normal file
@@ -0,0 +1,41 @@
|
||||
#!/usr/bin/env bash
|
||||
# Install GPU PyTorch for the surfer experiments (local RTX 3050 Ti, driver 580 → CUDA 12.x).
|
||||
#
|
||||
# RECOMMENDED (no sudo — installs to ~/.local, matching your existing numpy/scipy/databento):
|
||||
# bash scripts/install_torch_gpu.sh
|
||||
#
|
||||
# System-wide (only if you really want it):
|
||||
# sudo bash scripts/install_torch_gpu.sh
|
||||
#
|
||||
# Pick a different CUDA wheel tag if cu124 ever 404s (cu126 / cu128 also work on driver 580):
|
||||
# bash scripts/install_torch_gpu.sh cu126
|
||||
set -euo pipefail
|
||||
|
||||
CUDA_TAG="${1:-cu124}"
|
||||
TORCH_INDEX="https://download.pytorch.org/whl/${CUDA_TAG}"
|
||||
|
||||
if [ "$(id -u)" -eq 0 ]; then
|
||||
echo "[install] running as ROOT → system-wide site-packages (--break-system-packages)"
|
||||
FLAGS="--break-system-packages"
|
||||
else
|
||||
echo "[install] running as USER → ~/.local (matches existing packages; no sudo needed)"
|
||||
FLAGS="--user --break-system-packages"
|
||||
fi
|
||||
|
||||
echo "[install] torch from ${TORCH_INDEX}"
|
||||
python3 -m pip install ${FLAGS} torch --index-url "${TORCH_INDEX}"
|
||||
|
||||
echo "[verify] importing torch + checking CUDA on the local GPU"
|
||||
python3 - <<'PY'
|
||||
import torch
|
||||
print(" torch:", torch.__version__)
|
||||
print(" cuda available:", torch.cuda.is_available())
|
||||
if torch.cuda.is_available():
|
||||
print(" device:", torch.cuda.get_device_name(0),
|
||||
"| capability:", torch.cuda.get_device_capability(0))
|
||||
x = torch.randn(1_000_000, device="cuda")
|
||||
print(" GPU tensor op OK, sum =", float(x.sum()))
|
||||
else:
|
||||
raise SystemExit("CUDA NOT available to torch — check driver / try a different CUDA_TAG (cu126/cu128)")
|
||||
print(" ✅ GPU PyTorch ready")
|
||||
PY
|
||||
161
scripts/surfer/cross_exchange_funding.py
Normal file
161
scripts/surfer/cross_exchange_funding.py
Normal file
@@ -0,0 +1,161 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Frontier (a) — cross-exchange funding dispersion: edge distinct from momentum?
|
||||
|
||||
Fetch funding from Bybit + OKX (free) to pair with Binance funding + prices (crypto_pit).
|
||||
Build cross-sectional signals among multi-venue coins: multi-venue-avg funding (robust carry),
|
||||
cross-venue dispersion (positioning stress), Binance-premium. Validate each + correlation to
|
||||
the full crypto momentum book + marginal-alpha regression. Realistic: Reff = return - Binance
|
||||
funding (traded venue), death-excl, 10bp. Honest prior: low (cross-venue spreads arbitraged).
|
||||
"""
|
||||
import json
|
||||
import math
|
||||
import os
|
||||
import sys
|
||||
import time
|
||||
import urllib.request
|
||||
|
||||
import numpy as np
|
||||
|
||||
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
|
||||
import pit_sweep # noqa: E402
|
||||
from surfer_poc import compute_weights, CFG # noqa: E402
|
||||
from signal_sweep import xs_weights, pnl_w, validate, sharpe_t # noqa: E402
|
||||
import torch # noqa: E402
|
||||
|
||||
DEV = "cuda" if torch.cuda.is_available() else "cpu"
|
||||
DAY_MS = 86_400_000
|
||||
XF = "data/surfer/xfund"
|
||||
COINS = ["BTC", "ETH", "SOL", "XRP", "DOGE", "ADA", "AVAX", "LINK", "LTC", "DOT",
|
||||
"NEAR", "ATOM", "FIL", "ETC", "XLM", "UNI", "AAVE", "BNB"]
|
||||
|
||||
|
||||
def _get(u):
|
||||
try:
|
||||
return json.load(urllib.request.urlopen(urllib.request.Request(u, headers={"User-Agent": "curl/8"}), timeout=30))
|
||||
except Exception:
|
||||
return None
|
||||
|
||||
|
||||
def bybit(coin):
|
||||
cache = f"{XF}/bybit_{coin}.json"
|
||||
if os.path.exists(cache):
|
||||
return {int(k): v for k, v in json.load(open(cache)).items()}
|
||||
out, end = {}, int(time.time() * 1000)
|
||||
for _ in range(40):
|
||||
r = _get(f"https://api.bybit.com/v5/market/funding/history?category=linear&symbol={coin}USDT&endTime={end}&limit=200")
|
||||
lst = (r or {}).get("result", {}).get("list", [])
|
||||
if not lst:
|
||||
break
|
||||
for x in lst:
|
||||
t = int(x["fundingRateTimestamp"]); out.setdefault(t // DAY_MS, 0.0)
|
||||
out[t // DAY_MS] += float(x["fundingRate"])
|
||||
end = min(int(x["fundingRateTimestamp"]) for x in lst) - 1
|
||||
if len(lst) < 200:
|
||||
break
|
||||
time.sleep(0.06)
|
||||
os.makedirs(XF, exist_ok=True); json.dump({str(k): v for k, v in out.items()}, open(cache, "w"))
|
||||
return out
|
||||
|
||||
|
||||
def okx(coin):
|
||||
cache = f"{XF}/okx_{coin}.json"
|
||||
if os.path.exists(cache):
|
||||
return {int(k): v for k, v in json.load(open(cache)).items()}
|
||||
out, before = {}, ""
|
||||
for _ in range(60):
|
||||
u = f"https://www.okx.com/api/v5/public/funding-rate-history?instId={coin}-USDT-SWAP&limit=100"
|
||||
if before:
|
||||
u += f"&after={before}"
|
||||
r = _get(u); data = (r or {}).get("data", [])
|
||||
if not data:
|
||||
break
|
||||
for x in data:
|
||||
t = int(x["fundingTime"]); out.setdefault(t // DAY_MS, 0.0)
|
||||
out[t // DAY_MS] += float(x["fundingRate"])
|
||||
before = min(int(x["fundingTime"]) for x in data)
|
||||
if len(data) < 100:
|
||||
break
|
||||
time.sleep(0.06)
|
||||
os.makedirs(XF, exist_ok=True); json.dump({str(k): v for k, v in out.items()}, open(cache, "w"))
|
||||
return out
|
||||
|
||||
|
||||
def main():
|
||||
syms, days, close, qv, fund = pit_sweep.load()
|
||||
idx = {s: j for j, s in enumerate(syms)}
|
||||
coins = [c for c in COINS if c + "USDT" in idx]
|
||||
print(f"fetching Bybit+OKX funding for {len(coins)} coins...")
|
||||
by = {c: bybit(c) for c in coins}
|
||||
ok = {c: okx(c) for c in coins}
|
||||
nby = sum(1 for c in coins if len(by[c]) > 100); nok = sum(1 for c in coins if len(ok[c]) > 100)
|
||||
print(f" Bybit covered {nby}/{len(coins)}, OKX covered {nok}/{len(coins)}")
|
||||
|
||||
N = len(coins); T = len(days)
|
||||
di = {int(days[t]): t for t in range(T)}
|
||||
bn_f = np.full((T, N), np.nan); by_f = np.full((T, N), np.nan); ok_f = np.full((T, N), np.nan)
|
||||
cl = np.full((T, N), np.nan)
|
||||
for j, c in enumerate(coins):
|
||||
col = idx[c + "USDT"]
|
||||
cl[:, j] = close[:, col]; bn_f[:, j] = fund[:, col]
|
||||
for d, v in by[c].items():
|
||||
if d in di:
|
||||
by_f[di[d], j] = v
|
||||
for d, v in ok[c].items():
|
||||
if d in di:
|
||||
ok_f[di[d], j] = v
|
||||
|
||||
R = np.zeros((T, N)); R[1:] = np.log(cl)[1:] - np.log(cl)[:-1]; R = np.where(np.isfinite(R), R, 0.0)
|
||||
bnf = np.where(np.isfinite(bn_f), bn_f, 0.0)
|
||||
Reff = R - bnf # trade on Binance -> pay Binance funding
|
||||
# cross-venue stack
|
||||
stack = np.stack([bn_f, by_f, ok_f]) # [3,T,N]
|
||||
avg_f = np.nanmean(stack, axis=0) # multi-venue avg funding
|
||||
disp_f = np.nanstd(stack, axis=0) # cross-venue dispersion
|
||||
prem = bn_f - np.nanmean(np.stack([by_f, ok_f]), axis=0) # Binance premium vs others
|
||||
|
||||
def roll_mean(A, L):
|
||||
out = np.full_like(A, np.nan)
|
||||
for t in range(L, len(A)):
|
||||
out[t] = np.nanmean(A[t - L:t], axis=0)
|
||||
return out
|
||||
|
||||
sigs = {
|
||||
"carry_BINANCE_only": -roll_mean(np.where(np.isfinite(bn_f), bn_f, np.nan), 7), # control: is it just major-coin carry?
|
||||
"carry_multivenue": -roll_mean(avg_f, 7), # long low avg funding (robust carry)
|
||||
"venue_dispersion": -roll_mean(disp_f, 7), # low disagreement? test
|
||||
"venue_dispersion+": roll_mean(disp_f, 7), # high disagreement
|
||||
"binance_premium": -roll_mean(prem, 7), # fade Binance-crowded
|
||||
}
|
||||
valid = np.isfinite(cl) & np.isfinite(avg_f)
|
||||
year = (1970 + days / 365.25).astype(int)
|
||||
|
||||
def book(sig):
|
||||
s = sig.copy(); s[~valid] = np.nan
|
||||
return pnl_w(xs_weights(s), Reff, cost_bp=10)
|
||||
|
||||
# full-universe momentum book for correlation
|
||||
cw, _ = compute_weights(close, qv, days, CFG)
|
||||
cR = np.zeros_like(close); cR[1:] = np.log(close)[1:] - np.log(close)[:-1]; cR = np.where(np.isfinite(cR), cR, 0.0)
|
||||
cf = np.where(np.isfinite(fund), fund, 0.0)
|
||||
mom = np.sum(cw[:-1] * (cR - cf)[1:], axis=1)
|
||||
mom_by = {int(days[1:][i]): mom[i] for i in range(len(mom))}
|
||||
|
||||
print(f"\n===== CROSS-EXCHANGE FUNDING — {len(coins)} multi-venue coins, death-excl, 10bp, deflate N=55 =====")
|
||||
print(f"{'signal':>18} {'full':>6} {'IS':>6} {'OOS':>6} {'CPCVmed':>8} {'DSR':>5} {'corr_mom':>8} {'marg_t':>7}")
|
||||
pdays = days[1:]
|
||||
for nm, sg in sigs.items():
|
||||
p = book(sg); v = validate(p, days, 55)
|
||||
pb = {int(pdays[i]): p[i] for i in range(len(p))}
|
||||
common = sorted(set(pb) & set(mom_by))
|
||||
a = np.array([mom_by[d] for d in common]); b = np.array([pb[d] for d in common])
|
||||
m = np.isfinite(a) & np.isfinite(b); a, b = a[m], b[m]
|
||||
corr = float(np.corrcoef(a, b)[0, 1]) if len(a) > 100 else float("nan")
|
||||
beta1 = np.cov(a, b)[0, 1] / (np.var(a) + 1e-12); res = b - beta1 * a
|
||||
mt = float(res.mean() / (res.std() / math.sqrt(len(res)) + 1e-12))
|
||||
print(f"{nm:>18} {v['full']:>+6.2f} {v['is_']:>+6.2f} {v['oos']:>+6.2f} {v['med']:>+8.2f} {v['dsr']:>5.2f} {corr:>+8.2f} {mt:>+7.2f}")
|
||||
print("\nVERDICT: any signal with full+OOS+CPCVmed>0, DSR>0.5, low |corr_mom|, AND marg_t>2 = real distinct edge.")
|
||||
print("(marg_t = t-stat of marginal alpha vs momentum book). Honest prior: cross-venue spreads arbitraged -> expect fail.")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
160
scripts/surfer/cross_venue_funding.py
Normal file
160
scripts/surfer/cross_venue_funding.py
Normal file
@@ -0,0 +1,160 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Cross-venue crypto funding-arb scanner + paper-forward (price-neutral, no spot leg needed).
|
||||
|
||||
Funding differs across venues. For a coin on >=2 liquid venues: SHORT the highest-funding venue
|
||||
perp + LONG the lowest-funding venue perp (same coin) -> price-neutral (both perps, opposite),
|
||||
collect the funding DIFFERENCE (max-min). No spot leg -> solves the single-venue hedgeability block.
|
||||
|
||||
Funding intervals differ (Binance/Bybit 8h, Hyperliquid 1h) -> normalize to DAILY before comparing.
|
||||
No agents, no key, no capital. Venues: Binance, Bybit, Hyperliquid (all bulk-fetch).
|
||||
|
||||
python3 cross_venue_funding.py scan live top cross-venue spreads
|
||||
python3 cross_venue_funding.py run book daily carry on the top-K book, persist (cron)
|
||||
python3 cross_venue_funding.py status cumulative paper track record
|
||||
(alias: paper == run)
|
||||
"""
|
||||
import datetime
|
||||
import json
|
||||
import math
|
||||
import os
|
||||
import sys
|
||||
import urllib.request
|
||||
|
||||
_REPO = os.path.dirname(os.path.dirname(os.path.dirname(os.path.abspath(__file__))))
|
||||
STATE = os.path.join(_REPO, "data/surfer/crossvenue_state.json")
|
||||
LIQ = 10e6 # >$10M/day on BOTH legs (clean, fungible)
|
||||
MAXF = 0.005 # exclude legs with |daily funding| > 50bp/day = distress/artifact (un-tradeable)
|
||||
TOPK = 10 # book the top-K spreads
|
||||
COST_RT = 0.0010 # ~10bp round-trip (2 perp legs, maker)
|
||||
HURDLE = 0.0005 # 5bp/day spread to show in scan
|
||||
ENTRY = 0.0010 # hysteresis: only ENTER a new pair above 10bp/day
|
||||
EXIT = 0.0005 # hysteresis: HOLD a pair until its spread decays below 5bp/day (cuts turnover)
|
||||
|
||||
|
||||
def get(url, post=None):
|
||||
data = json.dumps(post).encode() if post else None
|
||||
h = {"User-Agent": "Mozilla/5.0", "Accept": "application/json"}
|
||||
if post:
|
||||
h["Content-Type"] = "application/json"
|
||||
for a in range(3):
|
||||
try:
|
||||
return json.loads(urllib.request.urlopen(urllib.request.Request(url, data=data, headers=h), timeout=25).read())
|
||||
except Exception:
|
||||
if a == 2:
|
||||
raise
|
||||
return None
|
||||
|
||||
|
||||
def base(sym):
|
||||
for q in ("USDT", "USDC"):
|
||||
if sym.endswith(q):
|
||||
return sym[:-len(q)]
|
||||
return sym
|
||||
|
||||
|
||||
def binance(): # {coin: (daily_funding, daily_vol)}
|
||||
fund = {x["symbol"]: float(x["lastFundingRate"]) for x in get("https://fapi.binance.com/fapi/v1/premiumIndex")}
|
||||
try:
|
||||
itv = {x["symbol"]: float(x.get("fundingIntervalHours", 8)) for x in get("https://fapi.binance.com/fapi/v1/fundingInfo")}
|
||||
except Exception:
|
||||
itv = {}
|
||||
vol = {x["symbol"]: float(x["quoteVolume"]) for x in get("https://fapi.binance.com/fapi/v1/ticker/24hr")}
|
||||
return {base(s): (f * (24.0 / itv.get(s, 8)), vol.get(s, 0.0)) for s, f in fund.items() if s.endswith("USDT")}
|
||||
|
||||
|
||||
def bybit():
|
||||
r = get("https://api.bybit.com/v5/market/tickers?category=linear")["result"]["list"]
|
||||
return {base(x["symbol"]): (float(x["fundingRate"]) * 3, float(x.get("turnover24h", 0.0)))
|
||||
for x in r if x["symbol"].endswith("USDT") and x.get("fundingRate")}
|
||||
|
||||
|
||||
def hyperliquid():
|
||||
r = get("https://api.hyperliquid.xyz/info", post={"type": "metaAndAssetCtxs"})
|
||||
meta, ctxs = r[0], r[1]
|
||||
out = {}
|
||||
for u, c in zip(meta["universe"], ctxs):
|
||||
if c.get("funding") is not None:
|
||||
out[u["name"]] = (float(c["funding"]) * 24, float(c.get("dayNtlVlm", 0.0)))
|
||||
return out
|
||||
|
||||
|
||||
def spreads():
|
||||
V = {}
|
||||
for nm, fn in [("Binance", binance), ("Bybit", bybit), ("HL", hyperliquid)]:
|
||||
try:
|
||||
V[nm] = fn()
|
||||
except Exception as e:
|
||||
print(f" ({nm} fetch failed: {str(e)[:40]})")
|
||||
coins = set().union(*[set(v) for v in V.values()])
|
||||
rows = []
|
||||
for c in coins:
|
||||
pts = {nm: V[nm][c] for nm in V if c in V[nm] and V[nm][c][1] > LIQ and abs(V[nm][c][0]) <= MAXF}
|
||||
if len(pts) < 2:
|
||||
continue
|
||||
f = {nm: pts[nm][0] for nm in pts}
|
||||
hi = max(f, key=f.get); lo = min(f, key=f.get)
|
||||
rows.append({"coin": c, "spread": f[hi] - f[lo], "short": hi, "long": lo, "f": f})
|
||||
rows.sort(key=lambda r: -r["spread"])
|
||||
return rows, list(V)
|
||||
|
||||
|
||||
def load_state():
|
||||
if os.path.exists(STATE):
|
||||
return json.load(open(STATE))
|
||||
return {"positions": {}, "cum_gross": 0.0, "cum_net": 0.0, "days": 0, "last_run_date": ""}
|
||||
|
||||
|
||||
def cmd_scan():
|
||||
rows, venues = spreads()
|
||||
print(f"cross-venue funding scan {datetime.date.today()} (venues: {', '.join(venues)}; liquid both legs >${LIQ/1e6:.0f}M)")
|
||||
n_ok = sum(1 for r in rows if r["spread"] > HURDLE)
|
||||
print(f" {len(rows)} coins on >=2 venues | {n_ok} with spread > {HURDLE*1e4:.0f}bp/day")
|
||||
print(f" {'coin':>8} {'spread/day':>11} {'ann%':>7} short -> long (daily funding)")
|
||||
for r in rows[:15]:
|
||||
leg = " ".join(f"{k}{1e4*v:+.1f}" for k, v in sorted(r["f"].items(), key=lambda kv: -kv[1]))
|
||||
print(f" {r['coin']:>8} {1e4*r['spread']:>9.1f}bp {100*r['spread']*365:>6.0f}% short {r['short']}->long {r['long']} [{leg}]")
|
||||
|
||||
|
||||
def cmd_run():
|
||||
st = load_state()
|
||||
today = datetime.datetime.now(datetime.timezone.utc).date().isoformat()
|
||||
if st.get("last_run_date") == today:
|
||||
print(f"already booked {today} (day {st['days']}); skipping"); return
|
||||
rows, _ = spreads()
|
||||
cur = {r["coin"]: r["spread"] for r in rows}
|
||||
prev = st["positions"]
|
||||
realized = sum(w * cur.get(c, 0.0) for c, w in prev.items()) # carry on yesterday's book at today's spreads
|
||||
# HYSTERESIS (the deployable, low-turnover version): hold winners until they decay, only enter strong fresh
|
||||
held = [c for c in prev if cur.get(c, 0.0) > EXIT]
|
||||
for r in sorted([r for r in rows if r["spread"] > ENTRY], key=lambda r: -r["spread"]):
|
||||
if len(held) >= TOPK:
|
||||
break
|
||||
if r["coin"] not in held:
|
||||
held.append(r["coin"])
|
||||
qual = [r for r in rows if r["coin"] in held]
|
||||
newpos = {c: 1.0 / len(held) for c in held} if held else {}
|
||||
turn = sum(abs(newpos.get(c, 0) - prev.get(c, 0)) for c in set(newpos) | set(prev))
|
||||
net = realized - turn * (COST_RT / 2)
|
||||
st.update(positions=newpos, days=st["days"] + 1, last_run_date=today)
|
||||
st["cum_gross"] += realized; st["cum_net"] += net
|
||||
os.makedirs(os.path.dirname(STATE), exist_ok=True); json.dump(st, open(STATE, "w"))
|
||||
top = " ".join(f"{r['coin']}:{1e4*r['spread']:.0f}bp({r['short'][:2]}>{r['long'][:2]})" for r in qual[:5])
|
||||
print(f"{today} day={st['days']} | qualifying={len(qual)} | realized24h gross={100*realized:+.3f}% net={100*net:+.3f}% "
|
||||
f"| cum net={100*st['cum_net']:+.2f}% | top: {top}")
|
||||
|
||||
|
||||
def cmd_status():
|
||||
st = load_state()
|
||||
print(f"cross-venue funding paper — day {st['days']} (through {st.get('last_run_date') or 'n/a'}), "
|
||||
f"{len(st['positions'])} pairs, cum gross {100*st['cum_gross']:+.2f}% net {100*st['cum_net']:+.2f}%")
|
||||
for c, w in sorted(st["positions"].items(), key=lambda kv: -kv[1]):
|
||||
print(f" {c:>8} w={w:.3f}")
|
||||
|
||||
|
||||
def main():
|
||||
cmd = sys.argv[1] if len(sys.argv) > 1 else "scan"
|
||||
{"scan": cmd_scan, "run": cmd_run, "paper": cmd_run, "status": cmd_status}.get(cmd, cmd_scan)()
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
116
scripts/surfer/crypto_cascade_reversion.py
Normal file
116
scripts/surfer/crypto_cascade_reversion.py
Normal file
@@ -0,0 +1,116 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Cascade reversion — do crypto coins BOUNCE after an extreme forced-flow down-move?
|
||||
|
||||
Liquidation cascades are forced deleveraging that overshoots; the harvestable consequence is the
|
||||
mean-reversion bounce. Binance killed the historical liquidation feed (live websocket only), so we
|
||||
test the PROXY already on disk: extreme negative trailing-H return = a cascade. Pre-registered:
|
||||
|
||||
formation: trailing-H log return per coin (cross-sectional)
|
||||
signal: LONG the bottom-q fraction (biggest losers) — long-only on crashed names
|
||||
hold: next H hours, NON-OVERLAPPING rebalance
|
||||
cost: 5 bp / leg on actual weight turnover
|
||||
two builds: raw_long (equal-weight losers, has crypto beta)
|
||||
hedged (long losers − short equal-weight UNIVERSE basket; beta-neutral, short leg
|
||||
is the diversified basket NOT a single coin → dodges the short-melt-up ruin
|
||||
that killed generic XS reversal)
|
||||
KILL: net annualized SR < 0.5 OR not positive in most years (chrono) → cascade reversion closed.
|
||||
|
||||
LIMITATION: npz panel has close only (no volume) → cannot condition on the volume spike that
|
||||
confirms a true liquidation cascade; the extreme-return tail is the proxy. Flagged, not hidden.
|
||||
"""
|
||||
import math
|
||||
import os
|
||||
|
||||
import numpy as np
|
||||
|
||||
OUT = "data/surfer/crypto1h"
|
||||
COST_BP = 5.0
|
||||
SYMS = ["BTCUSDT", "ETHUSDT", "SOLUSDT", "XRPUSDT", "BNBUSDT", "DOGEUSDT", "ADAUSDT", "AVAXUSDT",
|
||||
"LINKUSDT", "LTCUSDT", "DOTUSDT", "TRXUSDT", "BCHUSDT", "ETCUSDT", "FILUSDT", "ATOMUSDT"]
|
||||
HOUR_MS = 3_600_000
|
||||
HOLDS = [1, 2, 4, 8, 12, 24] # formation == hold, hours
|
||||
Q = 0.20 # bottom quintile = biggest losers
|
||||
|
||||
|
||||
def build_panel():
|
||||
series = {}
|
||||
for s in SYMS:
|
||||
cache = f"{OUT}/{s}.npz"
|
||||
if not os.path.exists(cache):
|
||||
continue
|
||||
d = np.load(cache)
|
||||
if len(d["ts"]) > 24 * 60:
|
||||
series[s] = (d["ts"], d["close"])
|
||||
allts = sorted(set().union(*[set((ts // HOUR_MS).tolist()) for ts, _ in series.values()]))
|
||||
tindex = {t: i for i, t in enumerate(allts)}
|
||||
syms = sorted(series)
|
||||
P = np.full((len(allts), len(syms)), np.nan)
|
||||
for j, s in enumerate(syms):
|
||||
ts, c = series[s]
|
||||
for t, px in zip(ts // HOUR_MS, c):
|
||||
P[tindex[int(t)], j] = px
|
||||
yrs = np.array([1970 + int(t) // int(365.25 * 24) for t in allts])
|
||||
return P, syms, yrs
|
||||
|
||||
|
||||
def ann(rets, per_year_factor):
|
||||
rets = np.asarray(rets)
|
||||
if len(rets) < 20 or rets.std() == 0:
|
||||
return float("nan"), float("nan")
|
||||
sr = rets.mean() / rets.std()
|
||||
t = rets.mean() / (rets.std() / math.sqrt(len(rets)))
|
||||
return sr * per_year_factor, t
|
||||
|
||||
|
||||
def run():
|
||||
P, syms, yrs = build_panel()
|
||||
T, N = P.shape
|
||||
logP = np.log(P)
|
||||
print(f"\n===== CRYPTO CASCADE REVERSION (long the crashed names, net {COST_BP}bp/leg) =====")
|
||||
print(f"panel: {T} hours x {N} coins years {int(yrs.min())}-{int(yrs.max())} q={Q} (bottom quintile)")
|
||||
print("LONG bottom-q losers; hedged = long losers - short equal-wt universe basket\n")
|
||||
hdr = f"{'hold':>5} {'periods':>8} {'build':>8} {'net_SR':>8} {'t(net)':>7} {'hit%':>6} {'per-year net-SR (chrono)':>30}"
|
||||
print(hdr); print("-" * len(hdr))
|
||||
|
||||
for H in HOLDS:
|
||||
pf = math.sqrt(24 * 365 / H)
|
||||
idx = np.arange(H, T - H, H) # need t-H for formation, t+H for fwd
|
||||
raw, hed, yr_raw, yr_hed = [], [], [], []
|
||||
w_raw_prev = np.zeros(N); w_hed_prev = np.zeros(N)
|
||||
for t in idx:
|
||||
past = logP[t] - logP[t - H]
|
||||
fwd = logP[t + H] - logP[t]
|
||||
ok = np.isfinite(past) & np.isfinite(fwd)
|
||||
n_ok = int(ok.sum())
|
||||
if n_ok < 6:
|
||||
continue
|
||||
okidx = np.where(ok)[0]
|
||||
k = max(1, int(round(Q * n_ok)))
|
||||
losers = okidx[np.argsort(past[okidx])[:k]] # most-negative trailing return
|
||||
# raw long
|
||||
w_raw = np.zeros(N); w_raw[losers] = 1.0 / k
|
||||
# hedged: long losers - short equal-weight universe basket
|
||||
w_hed = np.zeros(N); w_hed[losers] += 1.0 / k; w_hed[okidx] -= 1.0 / n_ok
|
||||
r_raw = float(np.nansum(w_raw * fwd)) - np.abs(w_raw - w_raw_prev).sum() * COST_BP / 1e4
|
||||
r_hed = float(np.nansum(w_hed * fwd)) - np.abs(w_hed - w_hed_prev).sum() * COST_BP / 1e4
|
||||
raw.append(r_raw); hed.append(r_hed)
|
||||
yr_raw.append(int(yrs[t])); yr_hed.append(int(yrs[t]))
|
||||
w_raw_prev = w_raw; w_hed_prev = w_hed
|
||||
if len(raw) < 20:
|
||||
print(f"{H:>5} (too few periods)"); continue
|
||||
for name, series, yrl in (("raw_long", raw, yr_raw), ("hedged", hed, yr_hed)):
|
||||
sr, t = ann(series, pf)
|
||||
hit = 100.0 * np.mean(np.asarray(series) > 0)
|
||||
py = {}
|
||||
for r, y in zip(series, yrl):
|
||||
py.setdefault(y, []).append(r)
|
||||
pystr = " ".join(f"{y}:{(np.mean(v)/(np.std(v)+1e-12)):+.2f}"
|
||||
for y, v in sorted(py.items()) if len(v) >= 10)
|
||||
tag = f"{H}h" if name == "raw_long" else ""
|
||||
print(f"{tag:>5} {len(series):>8} {name:>8} {sr:>+8.2f} {t:>+7.2f} {hit:>6.1f} {pystr}")
|
||||
print("-" * len(hdr))
|
||||
print("PASS: hedged net SR > 0.5, t>=2, positive in most years. Else cascade reversion closed.")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
run()
|
||||
91
scripts/surfer/crypto_funding_harvest.py
Normal file
91
scripts/surfer/crypto_funding_harvest.py
Normal file
@@ -0,0 +1,91 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Delta-neutral crypto funding HARVEST (cash-and-carry) — the one retail path to higher Sharpe.
|
||||
|
||||
When perp funding > 0 (longs pay shorts), go short-perp + long-spot (DELTA-NEUTRAL: price cancels)
|
||||
and collect the funding rate as carry. No price prediction. Causal: position[t] from funding[t-1]
|
||||
(funding is persistent), collect funding[t]. Net of realistic round-trip cost. Stress: worst months
|
||||
(crashes flip funding negative). Variants: broad vs majors, threshold, cost sensitivity.
|
||||
Treats `fund` as per-day funding (conservative; if per-8h the real APR/Sharpe is higher).
|
||||
"""
|
||||
import math
|
||||
import os
|
||||
import sys
|
||||
|
||||
import numpy as np
|
||||
|
||||
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
|
||||
import pit_sweep # noqa: E402
|
||||
|
||||
|
||||
def metrics(r):
|
||||
r = np.asarray(r); r = r[np.isfinite(r)]
|
||||
if len(r) < 50 or r.std() == 0:
|
||||
return (float("nan"),) * 4
|
||||
ann = r.mean() * 365; vol = r.std() * math.sqrt(365)
|
||||
eq = np.cumprod(1 + r); dd = float((eq / np.maximum.accumulate(eq) - 1).min())
|
||||
return ann, vol, ann / vol, dd
|
||||
|
||||
|
||||
def harvest(fund, liq, thr, c_round, year, label, topk=None):
|
||||
T, N = fund.shape
|
||||
elig = liq & np.isfinite(fund)
|
||||
sig = np.zeros((T, N))
|
||||
sig[1:] = np.where(elig[1:] & (fund[:-1] > thr), 1.0, 0.0) # causal: position from yesterday's funding
|
||||
if topk: # restrict to top-K best-funded eligible
|
||||
for t in range(1, T):
|
||||
idx = np.where(sig[t] > 0)[0]
|
||||
if len(idx) > topk:
|
||||
keep = idx[np.argsort(-fund[t - 1, idx])[:topk]]
|
||||
m = np.zeros(N); m[keep] = 1.0; sig[t] = m
|
||||
w = sig / np.maximum(sig.sum(1, keepdims=True), 1) # equal-weight positioned
|
||||
f = np.nan_to_num(fund)
|
||||
gross = np.sum(w[:-1] * f[1:], axis=1)
|
||||
turn = np.sum(np.abs(w[1:] - w[:-1]), axis=1) # fraction of book traded
|
||||
net = gross - turn * (c_round / 2) # one-way cost each side
|
||||
ann, vol, sr, dd = metrics(net)
|
||||
npos = sig.sum(1); avgpos = npos[npos > 0].mean()
|
||||
print(f"{label:>26} {100*ann:>6.1f} {100*vol:>6.1f} {sr:>+7.2f} {100*dd:>+7.1f} {avgpos:>6.0f} {100*turn.mean():>6.1f}")
|
||||
return net
|
||||
|
||||
|
||||
def main():
|
||||
syms, days, close, qv, fund = pit_sweep.load()
|
||||
T, N = fund.shape
|
||||
year = (1970 + days / 365.25).astype(int)
|
||||
qv30 = np.full_like(qv, np.nan)
|
||||
for t in range(30, T):
|
||||
qv30[t] = np.nanmean(qv[t - 30:t], axis=0)
|
||||
liq = np.isfinite(qv30) & (qv30 > 5e6) # >$5M/day quote volume
|
||||
majors = np.array([s in ("BTCUSDT", "ETHUSDT", "BNBUSDT", "SOLUSDT", "XRPUSDT", "ADAUSDT",
|
||||
"DOGEUSDT", "AVAXUSDT", "LINKUSDT", "MATICUSDT") for s in syms])
|
||||
liq_maj = liq & majors[None, :]
|
||||
|
||||
print(f"\n===== CRYPTO FUNDING HARVEST (delta-neutral, {N} coins, {T} days) =====")
|
||||
print(f"{'variant':>26} {'APR%':>6} {'vol%':>6} {'Sharpe':>7} {'maxDD%':>7} {'#pos':>6} {'turn%':>6}")
|
||||
base = harvest(fund, liq, 0.0, 0.0010, year, "broad, thr=0, 10bp")
|
||||
harvest(fund, liq, 0.0003, 0.0010, year, "broad, thr=3bp, 10bp")
|
||||
harvest(fund, liq, 0.0, 0.0010, year, "broad top-20, 10bp", topk=20)
|
||||
harvest(fund, liq_maj, 0.0, 0.0010, year, "majors only, 10bp")
|
||||
print(" -- cost sensitivity (broad, thr=0) --")
|
||||
harvest(fund, liq, 0.0, 0.0005, year, "broad, 5bp")
|
||||
harvest(fund, liq, 0.0, 0.0020, year, "broad, 20bp")
|
||||
harvest(fund, liq, 0.0, 0.0040, year, "broad, 40bp (harsh)")
|
||||
|
||||
# per-year + worst-month stress on the base case
|
||||
print("\n per-year Sharpe (broad, thr=0, 10bp):")
|
||||
print(" " + " ".join(f"{y}:{metrics(base[year[1:]==y])[2]:+.1f}" for y in range(2019, 2027) if (year[1:] == y).sum() > 60))
|
||||
# monthly buckets
|
||||
mo = (days[1:] / 30.4).astype(int)
|
||||
worst = sorted(set(mo), key=lambda m: base[mo == m].sum())[:5]
|
||||
print(" worst 5 ~monthly buckets (crash stress — does delta-neutral hold?):")
|
||||
for m in worst:
|
||||
seg = base[mo == m]
|
||||
d0 = int(days[1:][mo == m][0])
|
||||
import datetime
|
||||
print(f" ~{datetime.date.fromordinal(d0+719163)}: {100*seg.sum():+.1f}% over {len(seg)}d")
|
||||
print("\nVERDICT: if net Sharpe >> 1 AND survives crash months (delta-neutral so price-neutral) = the genuine")
|
||||
print("retail higher-Sharpe edge. Watch: turnover cost (funding flips), and whether tail months go deeply negative.")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
97
scripts/surfer/crypto_funding_liveness.py
Normal file
97
scripts/surfer/crypto_funding_liveness.py
Normal file
@@ -0,0 +1,97 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Decay/liveness: is the crypto funding carry still ALIVE now, or competed to death?
|
||||
|
||||
The harvest was +12 Sharpe 2020-24 but negative 2022, 2025, 2026. Decisive question: is that
|
||||
(a) STRUCTURAL COMPRESSION (funding yield shrinking year-over-year = arbitraged away = dead) or
|
||||
(b) REGIME (funding flips negative in deleverage, recovers = dormant)? And does a regime filter
|
||||
(only harvest reliably-positive funding, else sit in cash) rescue the recent period?
|
||||
|
||||
Diagnostics: (1) funding-yield decay table by year; (2) quarterly net-harvest liveness timeline;
|
||||
(3) regime-filtered vs unfiltered in the recent window.
|
||||
"""
|
||||
import datetime
|
||||
import math
|
||||
import os
|
||||
import sys
|
||||
|
||||
import numpy as np
|
||||
|
||||
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
|
||||
import pit_sweep # noqa: E402
|
||||
|
||||
|
||||
def sharpe(r):
|
||||
r = r[np.isfinite(r)]
|
||||
return float("nan") if len(r) < 20 or r.std() == 0 else r.mean() / r.std() * math.sqrt(365)
|
||||
|
||||
|
||||
def harvest(fund, elig, hurdle, c_round, trail=None):
|
||||
T, N = fund.shape
|
||||
cond = fund > hurdle
|
||||
if trail: # regime filter: trailing-mean funding > hurdle
|
||||
tf = np.full_like(fund, np.nan)
|
||||
for t in range(trail, T):
|
||||
tf[t] = np.nanmean(fund[t - trail:t], axis=0)
|
||||
cond = tf > hurdle
|
||||
sig = np.zeros((T, N))
|
||||
sig[1:] = np.where(elig[1:] & cond[:-1], 1.0, 0.0) # causal
|
||||
w = sig / np.maximum(sig.sum(1, keepdims=True), 1)
|
||||
f = np.nan_to_num(fund)
|
||||
gross = np.sum(w[:-1] * f[1:], axis=1)
|
||||
turn = np.sum(np.abs(w[1:] - w[:-1]), axis=1)
|
||||
net = gross - turn * (c_round / 2)
|
||||
inmkt = sig.sum(1)[1:] > 0 # days actually positioned
|
||||
return net, inmkt
|
||||
|
||||
|
||||
def main():
|
||||
syms, days, close, qv, fund = pit_sweep.load()
|
||||
T, N = fund.shape
|
||||
year = (1970 + days / 365.25).astype(int)
|
||||
qv30 = np.full_like(qv, np.nan)
|
||||
for t in range(30, T):
|
||||
qv30[t] = np.nanmean(qv[t - 30:t], axis=0)
|
||||
elig = np.isfinite(qv30) & (qv30 > 5e6) & np.isfinite(fund)
|
||||
|
||||
print("\n===== (1) FUNDING-YIELD DECAY by year (compression vs regime) =====")
|
||||
print(f"{'year':>6} {'#coins':>7} {'frac_pos':>9} {'mean_all(bp)':>13} {'mean_pos(bp)':>13}")
|
||||
for y in range(2019, 2027):
|
||||
m = year == y
|
||||
e = elig[m]
|
||||
fy = fund[m][e]
|
||||
if len(fy) < 100:
|
||||
continue
|
||||
ncoin = int(e.sum(1).mean()) if e.ndim == 2 else int(e.mean() * N)
|
||||
fpos = float((fy > 0).mean())
|
||||
print(f"{y:>6} {ncoin:>7} {fpos:>9.2f} {1e4*fy.mean():>13.2f} {1e4*fy[fy > 0].mean():>13.2f}")
|
||||
|
||||
print("\n===== (2) QUARTERLY net-harvest liveness (broad, thr=0, 10bp) =====")
|
||||
net, _ = harvest(fund, elig, 0.0, 0.0010)
|
||||
q = ((days[1:] - days[1]) / 91.3).astype(int)
|
||||
dd = days[1:]
|
||||
print(f"{'quarter':>10} {'netAPR%':>8} {'Sharpe':>7}")
|
||||
for qq in sorted(set(q)):
|
||||
seg = net[q == qq]
|
||||
d0 = datetime.date.fromordinal(int(dd[q == qq][0]) + 719163)
|
||||
if d0.year >= 2024: # focus recent
|
||||
print(f"{str(d0):>10} {100 * seg.sum() * (365 / max(len(seg), 1)) / 1:>8.1f} {sharpe(seg):>7.1f}")
|
||||
|
||||
print("\n===== (3) REGIME FILTER rescue test (recent: 2025-2026) =====")
|
||||
recent = year[1:] >= 2025
|
||||
variants = [("unfiltered thr=0 10bp", harvest(fund, elig, 0.0, 0.0010)[0]),
|
||||
("filtered tf14>3bp 10bp", harvest(fund, elig, 0.0003, 0.0010, trail=14)[0]),
|
||||
("filtered tf14>5bp 10bp", harvest(fund, elig, 0.0005, 0.0010, trail=14)[0]),
|
||||
("filtered tf30>5bp 10bp", harvest(fund, elig, 0.0005, 0.0010, trail=30)[0])]
|
||||
print(f"{'variant':>24} {'2025-26 APR%':>13} {'2025-26 Sharpe':>15} {'full Sharpe':>12}")
|
||||
for nm, r in variants:
|
||||
rr = r[recent]
|
||||
apr = 100 * np.nansum(rr) * 365 / max(np.isfinite(rr).sum(), 1)
|
||||
print(f"{nm:>24} {apr:>13.1f} {sharpe(rr):>15.1f} {sharpe(r):>12.1f}")
|
||||
|
||||
print("\nVERDICT: if mean_pos(bp) stable across years but frac_pos dipped 2025-26 = REGIME (dormant, recoverable).")
|
||||
print("If mean_pos shrinks monotonically = COMPRESSION (competed to death). If a filter turns 2025-26 positive")
|
||||
print("= alive-but-needs-regime-timing. If even filtered 2025-26 is negative/zero = the edge is GONE now.")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
92
scripts/surfer/crypto_funding_oos.py
Normal file
92
scripts/surfer/crypto_funding_oos.py
Normal file
@@ -0,0 +1,92 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Clean OOS validation of the funding-harvest regime filter.
|
||||
|
||||
Protocol (no peeking): choose the regime-filter params on IN-SAMPLE 2019-2024 by IS Sharpe, then
|
||||
apply that EXACT filter BLIND to OUT-OF-SAMPLE 2025-2026. If the IS-best filter still works OOS,
|
||||
the edge is validated, not curve-fit. Also reports the full grid's OOS so we can see if it's
|
||||
robust across reasonable filters (real) or only the lucky one (suspicious). Strategy structure
|
||||
fixed; only the regime filter (trailing window, hurdle) is the tuned degree of freedom.
|
||||
Raw Sharpe is funding-only (vol understated ~2-3x); realistic ~ raw / 2.5.
|
||||
"""
|
||||
import math
|
||||
import os
|
||||
import sys
|
||||
|
||||
import numpy as np
|
||||
|
||||
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
|
||||
import pit_sweep # noqa: E402
|
||||
|
||||
|
||||
def sharpe(r):
|
||||
r = r[np.isfinite(r)]
|
||||
return float("nan") if len(r) < 20 or r.std() == 0 else r.mean() / r.std() * math.sqrt(365)
|
||||
|
||||
|
||||
def apr(r):
|
||||
r = r[np.isfinite(r)]
|
||||
return 100 * r.sum() * 365 / max(len(r), 1)
|
||||
|
||||
|
||||
def main():
|
||||
syms, days, close, qv, fund = pit_sweep.load()
|
||||
T, N = fund.shape
|
||||
year = (1970 + days / 365.25).astype(int)
|
||||
qv30 = np.full_like(qv, np.nan)
|
||||
for t in range(30, T):
|
||||
qv30[t] = np.nanmean(qv[t - 30:t], axis=0)
|
||||
elig = np.isfinite(qv30) & (qv30 > 5e6) & np.isfinite(fund)
|
||||
f = np.nan_to_num(fund)
|
||||
|
||||
# precompute trailing means
|
||||
trailmean = {}
|
||||
for tr in (14, 30):
|
||||
tm = np.full_like(fund, np.nan)
|
||||
for t in range(tr, T):
|
||||
tm[t] = np.nanmean(fund[t - tr:t], axis=0)
|
||||
trailmean[tr] = tm
|
||||
|
||||
def harvest(trail, hurdle, c=0.0010):
|
||||
cond = (fund > hurdle) if trail == 0 else (trailmean[trail] > hurdle)
|
||||
sig = np.zeros((T, N))
|
||||
sig[1:] = np.where(elig[1:] & cond[:-1], 1.0, 0.0) # causal
|
||||
w = sig / np.maximum(sig.sum(1, keepdims=True), 1)
|
||||
gross = np.sum(w[:-1] * f[1:], axis=1)
|
||||
turn = np.sum(np.abs(w[1:] - w[:-1]), axis=1)
|
||||
net = gross - turn * (c / 2)
|
||||
inmkt = (sig.sum(1)[1:] > 0)
|
||||
return net, inmkt
|
||||
|
||||
yr = year[1:]
|
||||
IS = yr <= 2024
|
||||
OOS = yr >= 2025
|
||||
|
||||
grid = [(tr, h) for tr in (0, 14, 30) for h in (0.0, 0.0002, 0.0003, 0.0005)]
|
||||
rows = []
|
||||
for tr, h in grid:
|
||||
net, inmkt = harvest(tr, h)
|
||||
rows.append(dict(tr=tr, h=h, net=net, inmkt=inmkt,
|
||||
is_sr=sharpe(net[IS]), oos_sr=sharpe(net[OOS]),
|
||||
oos_apr=apr(net[OOS]), oos_inmkt=float(inmkt[OOS].mean())))
|
||||
|
||||
print(f"\n===== CLEAN OOS VALIDATION — IS 2019-2024 / OOS 2025-2026 =====")
|
||||
print(f"{'filter':>18} {'IS Sharpe':>10} {'OOS Sharpe':>11} {'OOS APR%':>9} {'OOS in-mkt%':>11}")
|
||||
for r in rows:
|
||||
nm = "naive thr=0" if (r["tr"] == 0 and r["h"] == 0) else (f"inst>{r['h']*1e4:.0f}bp" if r["tr"] == 0 else f"tf{r['tr']}>{r['h']*1e4:.0f}bp")
|
||||
print(f"{nm:>18} {r['is_sr']:>+10.1f} {r['oos_sr']:>+11.1f} {r['oos_apr']:>+9.1f} {100*r['oos_inmkt']:>10.0f}")
|
||||
|
||||
best = max(rows, key=lambda r: r["is_sr"])
|
||||
bnm = f"tf{best['tr']}>{best['h']*1e4:.0f}bp" if best["tr"] else f"inst>{best['h']*1e4:.0f}bp"
|
||||
print(f"\n IS-BEST filter (chosen blind to OOS): {bnm} -> IS Sharpe {best['is_sr']:+.1f}")
|
||||
print(f" ITS OOS RESULT: Sharpe {best['oos_sr']:+.1f} APR {best['oos_apr']:+.1f}% in-market {100*best['oos_inmkt']:.0f}% of days")
|
||||
print(f" (realistic Sharpe after basis vol ~ raw / 2.5 = {best['oos_sr']/2.5:+.1f})")
|
||||
# robustness: how many filters with IS Sharpe>2 also have OOS Sharpe>1
|
||||
good = [r for r in rows if r["is_sr"] > 2]
|
||||
robust = [r for r in good if r["oos_sr"] > 1]
|
||||
print(f"\n robustness: of {len(good)} filters with IS Sharpe>2, {len(robust)} also have OOS Sharpe>1")
|
||||
print("\nVERDICT: IS-best filter OOS Sharpe>1 (realistic) + broad robustness = VALIDATED, deployable.")
|
||||
print("If IS-best collapses OOS or only 1 filter works = curve-fit, not real.")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
242
scripts/surfer/crypto_funding_paper.py
Normal file
242
scripts/surfer/crypto_funding_paper.py
Normal file
@@ -0,0 +1,242 @@
|
||||
#!/usr/bin/env python3
|
||||
"""fundcli — local CLI for the crypto funding-harvest strategy (Phase-1 paper-forward).
|
||||
|
||||
No agents, no cloud. A self-contained local tool over Binance's public API (no key, no capital).
|
||||
Delta-neutral funding harvest: hold long-spot/short-perp on CRYPTO-NATIVE coins whose trailing-30d
|
||||
mean daily funding > 5bp; collect funding as market-neutral carry. Validated OOS (realistic
|
||||
Sharpe ~2-3.6); this tool tracks the no-capital forward record.
|
||||
|
||||
python3 crypto_funding_paper.py snapshot live liveness check (qualifying coins; no state change)
|
||||
python3 crypto_funding_paper.py run daily step: book funding on prior positions, log, persist
|
||||
python3 crypto_funding_paper.py status current book + cumulative track record
|
||||
python3 crypto_funding_paper.py gate Phase-1 -> Phase-2 assessment vs backtest
|
||||
python3 crypto_funding_paper.py orders [USD] Phase-2: current book -> exact delta-neutral orders
|
||||
python3 crypto_funding_paper.py log [N] last N run-log lines (default 25)
|
||||
(alias: 'paper' == 'run', for the cron)
|
||||
"""
|
||||
import datetime
|
||||
import json
|
||||
import os
|
||||
import sys
|
||||
import time
|
||||
import urllib.request
|
||||
|
||||
_REPO = os.path.dirname(os.path.dirname(os.path.dirname(os.path.abspath(__file__))))
|
||||
STATE = os.path.join(_REPO, "data/surfer/funding_paper_state.json")
|
||||
LOG = os.path.join(_REPO, "data/surfer/funding_paper_runs.log")
|
||||
FAPI = "https://fapi.binance.com"
|
||||
|
||||
HURDLE = 0.0005 # 5 bp/day trailing-30d mean funding (validated filter)
|
||||
EXIT_HURDLE = 0.0003 # hysteresis: keep held coins until they fall below 3 bp/day
|
||||
LIQ_USD = 5e6 # >$5M/day quote volume
|
||||
COST_RT = 0.0010 # 10 bp round-trip (net-of-cost awareness)
|
||||
INTERVALS_PER_DAY = 3 # Binance funding settles every 8h
|
||||
BACKTEST_APR = (15, 22) # validated APR band on deployed capital
|
||||
|
||||
|
||||
def get(url, tries=4):
|
||||
for a in range(tries):
|
||||
try:
|
||||
req = urllib.request.Request(url, headers={"User-Agent": "curl/8"})
|
||||
return json.loads(urllib.request.urlopen(req, timeout=30).read())
|
||||
except Exception:
|
||||
if a == tries - 1:
|
||||
raise
|
||||
time.sleep(2 * (a + 1))
|
||||
|
||||
|
||||
def crypto_native():
|
||||
info = get(f"{FAPI}/fapi/v1/exchangeInfo")
|
||||
return {s["symbol"] for s in info["symbols"] if s.get("underlyingType") == "COIN"}
|
||||
|
||||
|
||||
def spot_symbols():
|
||||
"""Symbols with a Binance SPOT market — required to build the delta-neutral hedge.
|
||||
Perp-only coins can't be cash-and-carry harvested (no spot leg); their high funding survives
|
||||
PRECISELY because the arb is impossible, so they must be excluded."""
|
||||
return {x["symbol"] for x in get("https://api.binance.com/api/v3/ticker/price")}
|
||||
|
||||
|
||||
def universe():
|
||||
"""Liquid, crypto-native, HEDGEABLE (spot+perp) USDT perps — the only ones that can be
|
||||
delta-neutral harvested."""
|
||||
native = crypto_native()
|
||||
spot = spot_symbols()
|
||||
t = get(f"{FAPI}/fapi/v1/ticker/24hr")
|
||||
return {x["symbol"]: float(x["quoteVolume"]) for x in t
|
||||
if x["symbol"].endswith("USDT") and x["symbol"] in native and x["symbol"] in spot
|
||||
and float(x["quoteVolume"]) > LIQ_USD}
|
||||
|
||||
|
||||
def funding_hist(sym, limit=90):
|
||||
h = get(f"{FAPI}/fapi/v1/fundingRate?symbol={sym}&limit={limit}")
|
||||
return [float(x["fundingRate"]) for x in h]
|
||||
|
||||
|
||||
def scan(prev=None):
|
||||
"""Fetch universe + funding; return (liq, tf30, last24) for union of universe and prev positions."""
|
||||
liq = universe()
|
||||
need = set(liq) | set(prev or {})
|
||||
tf30, last24 = {}, {}
|
||||
for i, sym in enumerate(sorted(need)):
|
||||
try:
|
||||
h = funding_hist(sym)
|
||||
except Exception:
|
||||
continue
|
||||
if len(h) < 30:
|
||||
continue
|
||||
tf30[sym] = (sum(h) / len(h)) * INTERVALS_PER_DAY
|
||||
last24[sym] = sum(h[-INTERVALS_PER_DAY:])
|
||||
if i % 50 == 49:
|
||||
time.sleep(0.5)
|
||||
return liq, tf30, last24
|
||||
|
||||
|
||||
def qualify(liq, tf30, prev):
|
||||
q = {}
|
||||
for c in liq:
|
||||
t = tf30.get(c)
|
||||
if t is not None and (t > HURDLE or (c in prev and t > EXIT_HURDLE)):
|
||||
q[c] = t
|
||||
return q
|
||||
|
||||
|
||||
def regime(n):
|
||||
return "ALIVE (healthy)" if n >= 15 else ("THIN / deleverage (mostly cash, by design)" if n >= 5 else "COMPRESSED (edge thin/decaying)")
|
||||
|
||||
|
||||
def load_state():
|
||||
if os.path.exists(STATE):
|
||||
return json.load(open(STATE))
|
||||
return {"positions": {}, "cum_gross": 0.0, "cum_net": 0.0, "days": 0}
|
||||
|
||||
|
||||
# ---- commands ----
|
||||
def cmd_snapshot():
|
||||
liq, tf30, _ = scan()
|
||||
q = qualify(liq, tf30, {})
|
||||
top = sorted(q.items(), key=lambda kv: -kv[1])
|
||||
med = sorted(q.values())[len(q) // 2] if q else 0.0
|
||||
print(f"funding snapshot {datetime.date.today()} (crypto-native, live)")
|
||||
print(f" liquid universe: {len(liq)} perps (>${LIQ_USD/1e6:.0f}M/day)")
|
||||
print(f" qualifying (tf30 > {HURDLE*1e4:.0f}bp/day): {len(q)} regime: {regime(len(q))}")
|
||||
print(f" median qualifying funding: {med*1e4:.1f} bp/day")
|
||||
for c, t in top[:15]:
|
||||
print(f" {c:>16} {t*1e4:5.1f} bp/day")
|
||||
|
||||
|
||||
def cmd_run():
|
||||
st = load_state()
|
||||
today = datetime.datetime.now(datetime.timezone.utc).date().isoformat()
|
||||
if st.get("last_run_date") == today: # idempotent: one booking per UTC day
|
||||
print(f"already booked {today} (day {st['days']}); skipping")
|
||||
return
|
||||
prev = st["positions"]
|
||||
liq, tf30, last24 = scan(prev)
|
||||
realized = sum(w * last24.get(c, 0.0) for c, w in prev.items())
|
||||
q = qualify(liq, tf30, prev)
|
||||
n = len(q)
|
||||
newpos = {c: 1.0 / n for c in q} if n else {}
|
||||
turnover = sum(abs(newpos.get(c, 0) - prev.get(c, 0)) for c in set(newpos) | set(prev))
|
||||
net = realized - turnover * (COST_RT / 2)
|
||||
st.update(positions=newpos, days=st["days"] + 1, last_run_date=today)
|
||||
st["cum_gross"] += realized
|
||||
st["cum_net"] += net
|
||||
os.makedirs(os.path.dirname(STATE), exist_ok=True)
|
||||
json.dump(st, open(STATE, "w"))
|
||||
top = " ".join(f"{c}:{t*1e4:.0f}bp" for c, t in sorted(q.items(), key=lambda kv: -kv[1])[:5])
|
||||
line = (f"{datetime.date.today()} day={st['days']} | qualifying={n} liquid={len(liq)} | "
|
||||
f"realized24h gross={100*realized:+.3f}% net={100*net:+.3f}% turn={turnover:.2f} | "
|
||||
f"cum gross={100*st['cum_gross']:+.2f}% net={100*st['cum_net']:+.2f}% | top: {top}")
|
||||
print(line)
|
||||
|
||||
|
||||
def cmd_status():
|
||||
st = load_state()
|
||||
print(f"funding-harvest paper state: day {st['days']}, {len(st['positions'])} positions, "
|
||||
f"cum gross {100*st['cum_gross']:+.2f}% cum net {100*st['cum_net']:+.2f}%")
|
||||
for c, w in sorted(st["positions"].items(), key=lambda kv: -kv[1]):
|
||||
print(f" {c:>16} w={w:.3f}")
|
||||
|
||||
|
||||
def cmd_gate():
|
||||
st = load_state()
|
||||
d = st["days"]
|
||||
if d == 0:
|
||||
print("no data yet — run 'run' (or wait for cron) to start the track record."); return
|
||||
apr = st["cum_net"] * 365 / d * 100
|
||||
print(f"=== Phase-1 gate assessment (day {d}, {d/7:.1f} weeks) ===")
|
||||
print(f" cumulative net: {100*st['cum_net']:+.2f}% annualized: {apr:+.1f}% APR")
|
||||
print(f" backtest target band: {BACKTEST_APR[0]}-{BACKTEST_APR[1]}% APR (gross carry, deployed capital)")
|
||||
if d < 28:
|
||||
rec = f"KEEP ACCUMULATING — need >=4 weeks ({28-d} days to go) before the gate is meaningful."
|
||||
elif apr >= 10:
|
||||
rec = "PROCEED-candidate -> Phase 2 (micro-live $1-3k, ONE tier-1 venue, low leverage). Carry holding forward."
|
||||
elif apr > 0:
|
||||
rec = "BORDERLINE -> extend paper-forward to 8 weeks; carry positive but below backtest band (thin regime)."
|
||||
else:
|
||||
rec = "DO NOT GO LIVE -> carry flat/negative forward; extend or stop. No capital risked."
|
||||
print(f" recommendation: {rec}")
|
||||
print(" caveats: GROSS carry only (realistic Sharpe ~ raw/2.5 after basis vol); counterparty/exchange")
|
||||
print(" tail is the real -100% risk (un-modeled); current qualifiers may be small/niche coins.")
|
||||
|
||||
|
||||
def cmd_orders(capital):
|
||||
"""Phase-2 bridge: turn the current paper book into exact delta-neutral orders for `capital`,
|
||||
and flag coins that are perp-only (no spot leg = can't hedge cleanly on Binance)."""
|
||||
st = load_state()
|
||||
book = st.get("positions", {})
|
||||
if not book:
|
||||
print("no current book — run 'snapshot'/'run' first."); return
|
||||
try:
|
||||
spot_px = {x["symbol"]: float(x["price"]) for x in get("https://api.binance.com/api/v3/ticker/price")}
|
||||
except Exception:
|
||||
spot_px = {}
|
||||
perp_px = {x["symbol"]: float(x["price"]) for x in get(f"{FAPI}/fapi/v1/ticker/price")}
|
||||
LEV = 2.0 # conservative perp leverage (avoid liquidation)
|
||||
n = len(book)
|
||||
X = capital / (n * (1 + 1 / LEV)) # notional per leg per coin
|
||||
print(f"=== Phase-2 delta-neutral orders for ${capital:.0f} across {n} coins (perp {LEV:.0f}x) ===")
|
||||
print(f" per coin: ~${X:.0f} notional/leg, ~${X/LEV:.0f} perp margin, ~${X*(1+1/LEV):.0f} capital")
|
||||
print(f"{'coin':>14} {'hedge':>6} {'SPOT buy (units @ px)':>26} {'PERP short (units @ px)':>26}")
|
||||
ok = 0
|
||||
for c in sorted(book):
|
||||
if c in spot_px and c in perp_px:
|
||||
print(f"{c:>14} {'YES':>6} {X/spot_px[c]:>14.4f} @ {spot_px[c]:<9.5g} {X/perp_px[c]:>14.4f} @ {perp_px[c]:<9.5g}")
|
||||
ok += 1
|
||||
else:
|
||||
why = "no-spot" if c not in spot_px else "no-perp"
|
||||
print(f"{c:>14} {'NO':>6} ({why}) cannot delta-neutral hedge on Binance — skip / alt-venue")
|
||||
print(f"\n {ok}/{n} coins hedgeable on Binance (spot+perp both exist).")
|
||||
print(" Place SPOT + PERP legs together to stay delta-neutral. Keep perp leverage low; never let the")
|
||||
print(" perp leg liquidate while holding spot. API keys: ENV ONLY, never commit. Place manually for")
|
||||
print(" micro-live to validate fills before any automation. DEPLOY ONLY AFTER the Phase-1 gate passes.")
|
||||
|
||||
|
||||
def cmd_log(n=25):
|
||||
if not os.path.exists(LOG):
|
||||
print(f"(no log yet at {LOG} — cron writes it nightly; 'run' appends when redirected)"); return
|
||||
lines = open(LOG).read().splitlines()
|
||||
print("\n".join(lines[-n:]))
|
||||
|
||||
|
||||
def main():
|
||||
cmd = sys.argv[1] if len(sys.argv) > 1 else "snapshot"
|
||||
if cmd in ("run", "paper"):
|
||||
cmd_run()
|
||||
elif cmd == "snapshot":
|
||||
cmd_snapshot()
|
||||
elif cmd == "status":
|
||||
cmd_status()
|
||||
elif cmd == "gate":
|
||||
cmd_gate()
|
||||
elif cmd == "orders":
|
||||
cmd_orders(float(sys.argv[2]) if len(sys.argv) > 2 else 2000.0)
|
||||
elif cmd == "log":
|
||||
cmd_log(int(sys.argv[2]) if len(sys.argv) > 2 else 25)
|
||||
else:
|
||||
print(__doc__)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
129
scripts/surfer/crypto_mft_xsec.py
Normal file
129
scripts/surfer/crypto_mft_xsec.py
Normal file
@@ -0,0 +1,129 @@
|
||||
#!/usr/bin/env python3
|
||||
"""MFT crypto pulse — cross-sectional momentum at HOUR holds, net of turnover cost, hardened-lite.
|
||||
|
||||
The proven crypto edge is DAILY cross-sectional momentum (Sharpe ~1.4). micro_gate showed MINUTE-horizon
|
||||
(per-coin) is dead. This tests the gap: does cross-sectional momentum work at HOUR holds (1h..48h)? If a
|
||||
pulse survives 5bp/leg turnover cost with stable per-year sign, it justifies a full multi-year hardened
|
||||
build; if not, intraday/MFT is closed and the edge is purely daily.
|
||||
|
||||
Hourly Binance klines (free, timestamped) for liquid perps → aligned panel. Pre-registered: lookback=hold=H,
|
||||
weights = dollar-neutral cross-sectional z-score of trailing-H return (gross 1), rebalance every H hours
|
||||
(NON-OVERLAPPING), forward-H return, cost = turnover · 5bp. Report gross/net annualized Sharpe + t + per-year.
|
||||
"""
|
||||
import json
|
||||
import math
|
||||
import os
|
||||
import time
|
||||
import urllib.request
|
||||
|
||||
import numpy as np
|
||||
|
||||
OUT = "data/surfer/crypto1h"
|
||||
COST_BP = 5.0
|
||||
HOLDS = [1, 2, 4, 8, 12, 24, 48]
|
||||
SYMS = ["BTCUSDT", "ETHUSDT", "SOLUSDT", "XRPUSDT", "BNBUSDT", "DOGEUSDT", "ADAUSDT", "AVAXUSDT",
|
||||
"LINKUSDT", "LTCUSDT", "DOTUSDT", "TRXUSDT", "BCHUSDT", "ETCUSDT", "FILUSDT", "ATOMUSDT"]
|
||||
HOUR_MS = 3_600_000
|
||||
|
||||
|
||||
def _get(u):
|
||||
return json.load(urllib.request.urlopen(urllib.request.Request(u, headers={"User-Agent": "curl/8"}), timeout=30))
|
||||
|
||||
|
||||
def fetch(sym):
|
||||
cache = f"{OUT}/{sym}.npz"
|
||||
if os.path.exists(cache):
|
||||
d = np.load(cache); return d["ts"], d["close"]
|
||||
end = int(time.time() * 1000); start = end - 5 * 365 * 86_400_000 # up to ~5y
|
||||
rows, cur = [], start
|
||||
for _ in range(400):
|
||||
k = _get(f"https://fapi.binance.com/fapi/v1/klines?symbol={sym}&interval=1h&startTime={cur}&limit=1500")
|
||||
if not k:
|
||||
break
|
||||
rows += k
|
||||
cur = k[-1][0] + 1
|
||||
if len(k) < 1500 or cur >= end:
|
||||
break
|
||||
time.sleep(0.05)
|
||||
if not rows:
|
||||
return np.array([]), np.array([])
|
||||
d = {int(r[0]): float(r[4]) for r in rows}
|
||||
ts = np.array(sorted(d)); close = np.array([d[t] for t in ts])
|
||||
os.makedirs(OUT, exist_ok=True)
|
||||
np.savez(cache, ts=ts, close=close)
|
||||
return ts, close
|
||||
|
||||
|
||||
def build_panel():
|
||||
series = {}
|
||||
for s in SYMS:
|
||||
ts, c = fetch(s)
|
||||
if len(ts) > 24 * 60: # need >~2 months
|
||||
series[s] = (ts, c)
|
||||
# align on the union of hourly timestamps (floored to the hour)
|
||||
allts = sorted(set().union(*[set((ts // HOUR_MS).tolist()) for ts, _ in series.values()]))
|
||||
tindex = {t: i for i, t in enumerate(allts)}
|
||||
syms = sorted(series)
|
||||
P = np.full((len(allts), len(syms)), np.nan)
|
||||
for j, s in enumerate(syms):
|
||||
ts, c = series[s]
|
||||
for t, px in zip(ts // HOUR_MS, c):
|
||||
P[tindex[int(t)], j] = px
|
||||
yrs = np.array([1970 + int(t) // (365.25 * 24) for t in allts])
|
||||
return P, syms, yrs
|
||||
|
||||
|
||||
def ann_sharpe(x):
|
||||
x = x[np.isfinite(x)]
|
||||
if len(x) < 20 or x.std() == 0:
|
||||
return float("nan"), float("nan"), 0
|
||||
return float(x.mean() / x.std()), x.mean() / (x.std() / math.sqrt(len(x))), len(x)
|
||||
|
||||
|
||||
def run():
|
||||
P, syms, yrs = build_panel()
|
||||
T, N = P.shape
|
||||
logP = np.log(P)
|
||||
print(f"\n========== MFT CRYPTO CROSS-SECTIONAL MOMENTUM (hour holds, net {COST_BP}bp/turnover) ==========")
|
||||
print(f"panel: {T} hours × {N} coins ({syms[0]}..{syms[-1]}) years {int(yrs.min())}-{int(yrs.max())}")
|
||||
print("dollar-neutral xs z-score momentum, lookback=hold=H, non-overlapping rebalance, cost=turnover·5bp\n")
|
||||
print(f"{'hold':>6} {'periods':>8} {'gross_SR':>9} {'net_SR':>8} {'t(net)':>7} {'pos%':>6} {'per-year net-SR (chrono)':>26}")
|
||||
print("-" * 92)
|
||||
for H in HOLDS:
|
||||
idx = np.arange(0, T - H, H) # non-overlapping rebalance points
|
||||
rets, yr_of = [], []
|
||||
w_prev = np.zeros(N)
|
||||
per_year = {}
|
||||
for t in idx:
|
||||
past = logP[t] - logP[t - H] if t - H >= 0 else np.full(N, np.nan)
|
||||
fwd = logP[t + H] - logP[t]
|
||||
ok = np.isfinite(past) & np.isfinite(fwd)
|
||||
if ok.sum() < 6:
|
||||
continue
|
||||
z = np.zeros(N)
|
||||
pv = past[ok]; zz = (pv - pv.mean()) / (pv.std() + 1e-12)
|
||||
z[ok] = zz
|
||||
g = np.abs(z).sum()
|
||||
w = z / g if g > 0 else z # dollar-neutral, gross 1
|
||||
turn = np.abs(w - w_prev).sum()
|
||||
r = float(np.nansum(w * fwd)) - turn * COST_BP / 1e4
|
||||
rets.append(r); yr_of.append(int(yrs[t]))
|
||||
per_year.setdefault(int(yrs[t]), []).append(r)
|
||||
w_prev = w
|
||||
rets = np.array(rets)
|
||||
if len(rets) < 20:
|
||||
print(f"{H:>6} (too few periods)"); continue
|
||||
gross_sr, _, _ = ann_sharpe(rets + 0) # gross approximated below
|
||||
# recompute gross (no cost) quickly
|
||||
per_yr_str = " ".join(f"{y}:{(np.mean(v)/ (np.std(v)+1e-12)):+.2f}" for y, v in sorted(per_year.items()) if len(v) >= 10)
|
||||
per_period = math.sqrt(24 * 365 / H) # annualization factor for H-hour periods
|
||||
nsr, nt, n = ann_sharpe(rets)
|
||||
# gross series
|
||||
# (recompute gross by adding back mean turnover cost is approximate; instead report net only + per-year)
|
||||
print(f"{H:>6}h {len(rets):>8} {'':>9} {nsr*per_period:>+8.2f} {nt:>+7.2f} {100*np.mean(rets>0):>6.1f} {per_yr_str}")
|
||||
print("-" * 92)
|
||||
print("PASS: net annualized SR > 0 with t≥2 AND positive in most years. Else MFT-crypto closed → daily only.")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
run()
|
||||
119
scripts/surfer/crypto_stablecoin_dislocation.py
Normal file
119
scripts/surfer/crypto_stablecoin_dislocation.py
Normal file
@@ -0,0 +1,119 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Stablecoin peg-reversion existence test — the cleanest crypto-FX dislocation.
|
||||
|
||||
Stablecoins are supposed to be worth $1; mint/redeem arbitrage (forced flow) pulls a deviation back
|
||||
to parity. This tests whether that reversion is harvestable NET OF COST on free Binance spot data, and
|
||||
— critically — how bad the DEATH-SPIRAL tail is (some "depegs" never revert: UST->0, BUSD wind-down).
|
||||
|
||||
Pre-registered: pair STABLE/USDT, deviation d=close-1. Reversion bet when |d|>thresh: position =
|
||||
-sign(d) (rich->short, cheap->long), hold H days, pnl_net = -sign(d)*(close[t+H]/close[t]-1) - cost_rt.
|
||||
Report per-pair + pooled: trade count, mean net pnl (bp), hit%, WORST trade (bp, the death-spiral tail),
|
||||
and a daily-strategy Sharpe. Cost levels bracket Binance stablecoin fees {0,2,10}bp round-trip.
|
||||
KILL: pooled net mean<=0 OR edge entirely from one terminal name OR worst-trade tail dwarfs mean edge.
|
||||
|
||||
LIMITATION: overlapping forward windows (first-look existence test, not a hardened backtest); daily
|
||||
granularity (intraday depeg spikes under-sampled). Flagged, not hidden.
|
||||
"""
|
||||
import json
|
||||
import math
|
||||
import os
|
||||
import time
|
||||
import urllib.request
|
||||
|
||||
import numpy as np
|
||||
|
||||
OUT = "data/surfer/stablecoin"
|
||||
PAIRS = ["USDCUSDT", "FDUSDUSDT", "TUSDUSDT", "DAIUSDT", "BUSDUSDT", "USDPUSDT"]
|
||||
DAY_MS = 86_400_000
|
||||
THRESHOLDS_BP = [10, 25, 50]
|
||||
HOLDS = [1, 2, 5]
|
||||
COSTS_BP = [0.0, 2.0, 10.0]
|
||||
|
||||
|
||||
def _get(u):
|
||||
return json.load(urllib.request.urlopen(urllib.request.Request(u, headers={"User-Agent": "curl/8"}), timeout=30))
|
||||
|
||||
|
||||
def fetch(sym):
|
||||
cache = f"{OUT}/{sym}.npz"
|
||||
if os.path.exists(cache):
|
||||
d = np.load(cache); return d["ts"], d["close"]
|
||||
start = 1_546_300_800_000 # 2019-01-01
|
||||
end = int(time.time() * 1000)
|
||||
rows, cur = [], start
|
||||
for _ in range(400):
|
||||
k = _get(f"https://api.binance.com/api/v3/klines?symbol={sym}&interval=1d&startTime={cur}&limit=1000")
|
||||
if not k:
|
||||
break
|
||||
rows += k
|
||||
cur = k[-1][0] + DAY_MS
|
||||
if len(k) < 1000 or cur >= end:
|
||||
break
|
||||
time.sleep(0.05)
|
||||
if not rows:
|
||||
return np.array([]), np.array([])
|
||||
d = {int(r[0]): float(r[4]) for r in rows}
|
||||
ts = np.array(sorted(d)); close = np.array([d[t] for t in ts])
|
||||
os.makedirs(OUT, exist_ok=True)
|
||||
np.savez(cache, ts=ts, close=close)
|
||||
return ts, close
|
||||
|
||||
|
||||
def reversion(close, thresh_bp, H, cost_bp):
|
||||
"""Event reversion bet: enter when |d|>thresh, pnl=-sign(d)*(fwd_ret)-cost_rt. Returns net-pnl array (bp)."""
|
||||
d = close - 1.0
|
||||
th = thresh_bp / 1e4
|
||||
pnls = []
|
||||
for t in range(len(close) - H):
|
||||
if abs(d[t]) > th and close[t] > 0:
|
||||
fwd = close[t + H] / close[t] - 1.0
|
||||
pnl = -np.sign(d[t]) * fwd - cost_bp / 1e4
|
||||
pnls.append(pnl * 1e4) # in bp
|
||||
return np.array(pnls)
|
||||
|
||||
|
||||
def run():
|
||||
series = {}
|
||||
for s in PAIRS:
|
||||
ts, c = fetch(s)
|
||||
if len(ts) > 60:
|
||||
series[s] = c
|
||||
print("\n========== STABLECOIN PEG-REVERSION (crypto-FX dislocation, free Binance spot daily) ==========")
|
||||
print(f"pairs: {', '.join(f'{s}({len(c)}d)' for s, c in series.items())}\n")
|
||||
|
||||
# 1) dislocation frequency per pair
|
||||
print("--- dislocation magnitude (|close-1|) per pair ---")
|
||||
print(f"{'pair':>9} {'days':>5} {'med_bp':>7} {'p95_bp':>7} {'max_bp':>8} {'>10bp%':>7} {'>50bp%':>7}")
|
||||
for s, c in series.items():
|
||||
d = np.abs(c - 1.0) * 1e4
|
||||
print(f"{s:>9} {len(c):>5} {np.median(d):>7.1f} {np.percentile(d,95):>7.1f} {d.max():>8.0f} "
|
||||
f"{100*np.mean(d>10):>6.1f}% {100*np.mean(d>50):>6.1f}%")
|
||||
|
||||
# 2) reversion edge, pooled across pairs, by (thresh, H, cost)
|
||||
print("\n--- reversion edge (pooled all pairs); pnl in bp/trade, net of round-trip cost ---")
|
||||
print(f"{'thr_bp':>6} {'H':>2} {'cost':>5} {'trades':>7} {'mean_bp':>8} {'hit%':>6} {'worst_bp':>9} {'daily_SR':>9}")
|
||||
for th in THRESHOLDS_BP:
|
||||
for H in HOLDS:
|
||||
for cost in COSTS_BP:
|
||||
allp = np.concatenate([reversion(c, th, H, cost) for c in series.values()]) if series else np.array([])
|
||||
if len(allp) < 10:
|
||||
continue
|
||||
sr = (allp.mean() / allp.std() * math.sqrt(365 / H)) if allp.std() > 0 else float("nan")
|
||||
print(f"{th:>6} {H:>2} {cost:>4.0f}b {len(allp):>7} {allp.mean():>+8.1f} "
|
||||
f"{100*np.mean(allp>0):>5.1f}% {allp.min():>+9.0f} {sr:>+9.2f}")
|
||||
print(" " + "-" * 60)
|
||||
|
||||
# 3) per-pair edge at a fixed mid setting (thr=25bp, H=2, cost=2bp) — is it one terminal name?
|
||||
print("\n--- per-pair edge @ thr=25bp H=2 cost=2bp (is the edge concentrated/terminal?) ---")
|
||||
print(f"{'pair':>9} {'trades':>7} {'mean_bp':>8} {'hit%':>6} {'worst_bp':>9}")
|
||||
for s, c in series.items():
|
||||
p = reversion(c, 25, 2, 2.0)
|
||||
if len(p) >= 5:
|
||||
print(f"{s:>9} {len(p):>7} {p.mean():>+8.1f} {100*np.mean(p>0):>5.1f}% {p.min():>+9.0f}")
|
||||
else:
|
||||
print(f"{s:>9} {len(p):>7} (too few events)")
|
||||
print("\nKILL if: pooled mean<=0, OR edge is one terminal name, OR |worst_bp| dwarfs mean (death-spiral tail).")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
run()
|
||||
145
scripts/surfer/crypto_stablecoin_harden.py
Normal file
145
scripts/surfer/crypto_stablecoin_harden.py
Normal file
@@ -0,0 +1,145 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Stablecoin peg-reversion — HARDENED. Make-or-break test of the crypto-FX dislocation edge.
|
||||
|
||||
The spike (crypto_stablecoin_dislocation.py) found a real reversion edge (SR 3-4) but was blind to:
|
||||
(1) survivorship — all sampled stables reverted; the true tail is buying a cheap stable that goes to
|
||||
ZERO (UST -75%->delist). (2) overlap inflation. (3) stress slippage. (4) long/short asymmetry.
|
||||
|
||||
This hardens all four:
|
||||
* UST included (the real death-spiral). FRAX excluded (garbage prints). USTC excluded (post-death).
|
||||
* NON-OVERLAPPING event walk with explicit exit (revert / max_hold / stop-loss).
|
||||
* deviation-SCALED slippage: cost_rt = base_bp + slip_k*|deviation| (big depeg = thin book).
|
||||
* LONG-cheap (unbounded -100% tail) vs SHORT-rich (bounded) decomposed separately.
|
||||
* collapse filters compared: raw / stop-loss / skip-falling-knife(+stop).
|
||||
* per-year sign + BTC-stress correlation proxy (does it lose when crypto crashes?).
|
||||
KILL: long-side net<=0 after UST+filters, OR edge only pre-cost, OR loses concentrated in BTC crashes.
|
||||
"""
|
||||
import math
|
||||
import os
|
||||
|
||||
import numpy as np
|
||||
|
||||
OUT = "data/surfer/stablecoin"
|
||||
HEALTHY = ["USDCUSDT", "FDUSDUSDT", "TUSDUSDT", "BUSDUSDT", "USDPUSDT", "SUSDUSDT", "DAIUSDT"]
|
||||
TERMINAL = ["USTUSDT"]
|
||||
DAY_MS = 86_400_000
|
||||
THRESH_BP = 50
|
||||
EXIT_BP = 10
|
||||
MAX_HOLD = 5
|
||||
BASE_BP = 2.0
|
||||
SLIP_K = 0.15 # slippage = 15% of the entry deviation (stress thinness)
|
||||
STOP = 0.03 # 3% stop-loss
|
||||
|
||||
|
||||
def fetch(sym):
|
||||
import json, time, urllib.request
|
||||
cache = f"{OUT}/{sym}.npz"
|
||||
if os.path.exists(cache):
|
||||
d = np.load(cache); return d["ts"], d["close"]
|
||||
g = lambda u: json.load(urllib.request.urlopen(urllib.request.Request(u, headers={"User-Agent": "curl/8"}), timeout=30))
|
||||
start, end = 1_546_300_800_000, int(time.time() * 1000)
|
||||
rows, cur = [], start
|
||||
for _ in range(400):
|
||||
k = g(f"https://api.binance.com/api/v3/klines?symbol={sym}&interval=1d&startTime={cur}&limit=1000")
|
||||
if not k:
|
||||
break
|
||||
rows += k; cur = k[-1][0] + DAY_MS
|
||||
if len(k) < 1000 or cur >= end:
|
||||
break
|
||||
time.sleep(0.05)
|
||||
if not rows:
|
||||
return np.array([]), np.array([])
|
||||
d = {int(r[0]): float(r[4]) for r in rows}
|
||||
ts = np.array(sorted(d)); close = np.array([d[t] for t in ts])
|
||||
os.makedirs(OUT, exist_ok=True); np.savez(cache, ts=ts, close=close)
|
||||
return ts, close
|
||||
|
||||
|
||||
def walk(ts, close, mode):
|
||||
"""Non-overlapping reversion trades. mode: 'raw'|'stop'|'skip'. Returns list of (dir, year, net_bp, entry_dev_bp)."""
|
||||
n = len(close); d = close - 1.0; th = THRESH_BP / 1e4; ex = EXIT_BP / 1e4
|
||||
trades = []; i = 1
|
||||
while i < n - 1:
|
||||
if abs(d[i]) <= th:
|
||||
i += 1; continue
|
||||
direction = -1.0 if d[i] > 0 else 1.0 # rich->short(-1), cheap->long(+1)
|
||||
# skip-falling-knife: only buy a cheap stable if it is NOT still dropping (bounce confirmation)
|
||||
if mode == "skip" and direction > 0 and close[i] < close[i - 1]:
|
||||
i += 1; continue
|
||||
entry = close[i]; j = i + 1; stop = STOP if mode in ("stop", "skip") else 9.9
|
||||
while j < n and (j - i) <= MAX_HOLD:
|
||||
pnl = direction * (close[j] / entry - 1.0)
|
||||
if pnl <= -stop or abs(close[j] - 1.0) < ex:
|
||||
break
|
||||
j += 1
|
||||
j = min(j, n - 1)
|
||||
gross = direction * (close[j] / entry - 1.0)
|
||||
slip = (BASE_BP + SLIP_K * abs(d[i]) * 1e4) / 1e4
|
||||
yr = 1970 + int(ts[i]) // (365 * DAY_MS)
|
||||
trades.append((direction, yr, (gross - slip) * 1e4, abs(d[i]) * 1e4))
|
||||
i = j + 1
|
||||
return trades
|
||||
|
||||
|
||||
def stats(pnls):
|
||||
p = np.array(pnls)
|
||||
if len(p) < 3:
|
||||
return f"n={len(p)} (too few)"
|
||||
sr = p.mean() / p.std() * math.sqrt(365 / ((MAX_HOLD + 1) / 2)) if p.std() > 0 else float("nan")
|
||||
return f"n={len(p):>4} mean={p.mean():>+7.1f}bp hit={100*np.mean(p>0):>4.1f}% worst={p.min():>+7.0f}bp SR~{sr:>+5.2f}"
|
||||
|
||||
|
||||
def run():
|
||||
data = {}
|
||||
for s in HEALTHY + TERMINAL:
|
||||
ts, c = fetch(s)
|
||||
if len(ts) > 30:
|
||||
data[s] = (ts, c)
|
||||
print("\n===== STABLECOIN PEG-REVERSION — HARDENED (non-overlap, UST-in, slip-scaled, long/short split) =====")
|
||||
print(f"pairs: {', '.join(data)} thr={THRESH_BP}bp hold<={MAX_HOLD}d exit<{EXIT_BP}bp slip={BASE_BP}+{SLIP_K}*|dev| stop={STOP*100:.0f}%\n")
|
||||
|
||||
for mode in ("raw", "stop", "skip"):
|
||||
allt = []
|
||||
for s in data:
|
||||
allt += [(s, *t) for t in walk(*data[s], mode)]
|
||||
longs = [t[3] for t in allt if t[1] > 0]
|
||||
shorts = [t[3] for t in allt if t[1] < 0]
|
||||
ust = [t[3] for t in allt if t[0] == "USTUSDT"]
|
||||
print(f"[{mode:>4}] ALL : {stats([t[3] for t in allt])}")
|
||||
print(f" LONG(cheap, -100% tail): {stats(longs)}")
|
||||
print(f" SHORT(rich, bounded) : {stats(shorts)}")
|
||||
print(f" UST trades only : {stats(ust)}")
|
||||
print()
|
||||
|
||||
# per-year (SHORT-only, skip mode = the candidate deployable core)
|
||||
print("--- per-year SHORT-rich net mean bp (the bounded-risk core), skip mode ---")
|
||||
by_yr = {}
|
||||
for s in data:
|
||||
for t in walk(*data[s], "skip"):
|
||||
if t[0] < 0:
|
||||
by_yr.setdefault(t[1], []).append(t[2])
|
||||
print(" ".join(f"{y}:{np.mean(v):+.0f}({len(v)})" for y, v in sorted(by_yr.items()) if v))
|
||||
|
||||
# BTC-stress proxy: does the LONG side lose during BTC crashes?
|
||||
bts, btc = fetch("BTCUSDT")
|
||||
if len(btc) > 100:
|
||||
bret = {int(bts[k]) // DAY_MS: (btc[k] / btc[k - 1] - 1.0) for k in range(1, len(btc))}
|
||||
crash, calm = [], []
|
||||
for s in data:
|
||||
ts, c = data[s]
|
||||
for dirn, yr, net, dev in walk(ts, c, "skip"):
|
||||
pass
|
||||
# simpler: tag each long trade entry day with BTC 1d return
|
||||
for s in data:
|
||||
ts, c = data[s]; d = c - 1.0; th = THRESH_BP / 1e4
|
||||
for i in range(1, len(c) - 1):
|
||||
if abs(d[i]) > th and d[i] < 0: # long-cheap entries
|
||||
br = bret.get(int(ts[i]) // DAY_MS, 0.0)
|
||||
(crash if br < -0.05 else calm).append(1)
|
||||
print(f"\nLONG-cheap entries on BTC-crash days (>-5%): {len(crash)} vs calm: {len(calm)} "
|
||||
f"({100*len(crash)/max(1,len(crash)+len(calm)):.0f}% in crashes = tail-correlation flag)")
|
||||
print("\nREAD: deployable core = SHORT-rich (bounded). LONG-cheap viable only if filters tame the UST tail.")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
run()
|
||||
120
scripts/surfer/crypto_stablecoin_intraday.py
Normal file
120
scripts/surfer/crypto_stablecoin_intraday.py
Normal file
@@ -0,0 +1,120 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Stablecoin peg-reversion — INTRADAY EXECUTION validation (the make-or-break-on-fills test).
|
||||
|
||||
Daily test found short-rich reversion ~+25bp/SR4. This checks it survives REALISTIC intraday execution:
|
||||
* 1h bars, REALISTIC fills: enter at NEXT bar's OPEN (reaction lag — you can't transact at the peak),
|
||||
exit at the reverting bar's close; pay round-trip spread (calibrated from live book: ~1bp/side).
|
||||
* deviation HALF-LIFE: of rich events at hour t, what % still rich at t+1/+2/+4/+8 (do you have time?).
|
||||
* short-rich (validated, bounded) focus; long-cheap with skip-falling-knife for contrast.
|
||||
Universe = live-tradeable liquid stables only (USDC/FDUSD/TUSD/USDP; DAI/BUSD books are empty).
|
||||
KILL: edge gone after next-open fills + spread, OR deviations vanish within 1h (no reaction window).
|
||||
"""
|
||||
import math
|
||||
import os
|
||||
|
||||
import numpy as np
|
||||
|
||||
OUT = "data/surfer/stable1h"
|
||||
PAIRS = ["USDCUSDT", "FDUSDUSDT", "TUSDUSDT", "USDPUSDT"]
|
||||
HOUR_MS = 3_600_000
|
||||
THRESHES = [25, 50]
|
||||
EXIT_BP = 10
|
||||
MAX_HOLD_H = 48
|
||||
SPREAD_BP = 1.0 # per side, from live book; round-trip = 2*SPREAD_BP
|
||||
|
||||
|
||||
def fetch(sym):
|
||||
import json, time, urllib.request
|
||||
cache = f"{OUT}/{sym}.npz"
|
||||
if os.path.exists(cache):
|
||||
d = np.load(cache); return d["ts"], d["open"], d["close"]
|
||||
g = lambda u: json.load(urllib.request.urlopen(urllib.request.Request(u, headers={"User-Agent": "curl/8"}), timeout=30))
|
||||
start, end = int(time.time() * 1000) - 3 * 365 * 86_400_000, int(time.time() * 1000)
|
||||
rows, cur = [], start
|
||||
for _ in range(800):
|
||||
k = g(f"https://api.binance.com/api/v3/klines?symbol={sym}&interval=1h&startTime={cur}&limit=1000")
|
||||
if not k:
|
||||
break
|
||||
rows += k; cur = k[-1][0] + HOUR_MS
|
||||
if len(k) < 1000 or cur >= end:
|
||||
break
|
||||
time.sleep(0.04)
|
||||
if not rows:
|
||||
return np.array([]), np.array([]), np.array([])
|
||||
o = {int(r[0]): (float(r[1]), float(r[4])) for r in rows}
|
||||
ts = np.array(sorted(o)); op = np.array([o[t][0] for t in ts]); cl = np.array([o[t][1] for t in ts])
|
||||
os.makedirs(OUT, exist_ok=True); np.savez(cache, ts=ts, open=op, close=cl)
|
||||
return ts, op, cl
|
||||
|
||||
|
||||
def walk(op, cl, thresh, side, skip=False):
|
||||
"""Non-overlapping. side='short'(rich) or 'long'(cheap). enter next-bar OPEN, exit revert close, pay spread."""
|
||||
n = len(cl); d = cl - 1.0; th = thresh / 1e4; ex = EXIT_BP / 1e4; rt = 2 * SPREAD_BP / 1e4
|
||||
out = []; i = 1
|
||||
while i < n - 2:
|
||||
rich = d[i] > th; cheap = d[i] < -th
|
||||
take = (side == "short" and rich) or (side == "long" and cheap)
|
||||
if not take:
|
||||
i += 1; continue
|
||||
if skip and side == "long" and cl[i] < cl[i - 1]: # don't catch a still-falling knife
|
||||
i += 1; continue
|
||||
direction = -1.0 if rich else 1.0
|
||||
entry = op[i + 1] # realistic: fill at next bar open
|
||||
j = i + 2
|
||||
while j < n and (j - (i + 1)) <= MAX_HOLD_H:
|
||||
if abs(cl[j] - 1.0) < ex:
|
||||
break
|
||||
j += 1
|
||||
j = min(j, n - 1)
|
||||
gross = direction * (cl[j] / entry - 1.0)
|
||||
out.append((gross - rt) * 1e4)
|
||||
i = j + 1
|
||||
return np.array(out)
|
||||
|
||||
|
||||
def half_life(cl, thresh):
|
||||
d = cl - 1.0; th = thresh / 1e4
|
||||
ev = np.where(np.abs(d) > th)[0]; ev = ev[ev < len(cl) - 9]
|
||||
if len(ev) == 0:
|
||||
return None
|
||||
return {h: 100 * np.mean(np.abs(d[ev + h]) > th) for h in (1, 2, 4, 8)}
|
||||
|
||||
|
||||
def st(p):
|
||||
if len(p) < 3:
|
||||
return f"n={len(p)} (few)"
|
||||
sr = p.mean() / p.std() * math.sqrt(365 * 24 / (MAX_HOLD_H / 2)) if p.std() > 0 else float("nan")
|
||||
return f"n={len(p):>4} mean={p.mean():>+7.1f}bp hit={100*np.mean(p>0):>4.1f}% worst={p.min():>+7.0f}bp SR~{sr:>+5.2f}"
|
||||
|
||||
|
||||
def run():
|
||||
data = {}
|
||||
for s in PAIRS:
|
||||
ts, op, cl = fetch(s)
|
||||
if len(ts) > 500:
|
||||
data[s] = (op, cl, ts)
|
||||
print("\n===== STABLECOIN INTRADAY EXECUTION (1h, next-open fills, spread 2bp rt) =====")
|
||||
print(f"pairs: {', '.join(f'{s}({len(v[1])}h)' for s,v in data.items())}\n")
|
||||
|
||||
print("--- deviation HALF-LIFE: % of events still beyond threshold after H hours (reaction window) ---")
|
||||
for s, (op, cl, ts) in data.items():
|
||||
for th in THRESHES:
|
||||
hl = half_life(cl, th)
|
||||
if hl:
|
||||
print(f" {s:>9} thr{th}: t+1={hl[1]:.0f}% t+2={hl[2]:.0f}% t+4={hl[4]:.0f}% t+8={hl[8]:.0f}%")
|
||||
print()
|
||||
|
||||
for th in THRESHES:
|
||||
sh = np.concatenate([walk(d[0], d[1], th, "short") for d in data.values()]) if data else np.array([])
|
||||
lo = np.concatenate([walk(d[0], d[1], th, "long", skip=True) for d in data.values()]) if data else np.array([])
|
||||
print(f"thr={th}bp SHORT-rich : {st(sh)}")
|
||||
print(f" LONG-cheap*: {st(lo)} (*skip-falling-knife)")
|
||||
# per-pair short
|
||||
for s, (op, cl, ts) in data.items():
|
||||
print(f" {s:>9} short: {st(walk(op, cl, th, 'short'))}")
|
||||
print()
|
||||
print("READ: edge survives if SHORT-rich stays +ve net of next-open+spread AND half-life>~2h (time to act).")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
run()
|
||||
126
scripts/surfer/crypto_sweep.py
Normal file
126
scripts/surfer/crypto_sweep.py
Normal file
@@ -0,0 +1,126 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Deflated signal sweep — Batch 2 (crypto perps): broaden + robustness-check.
|
||||
|
||||
Cross-sectional, market-neutral, unit-gross signals on major Binance USDT perps. Funding
|
||||
baked into the return (Reff = log-return − daily funding; long pays positive funding). ~10bp
|
||||
taker cost on turnover. Deflated Sharpe deflated by the CUMULATIVE search (Batch1 17 + these).
|
||||
For any DEVELOP-grade signal (CPCVmed>0 & IS>0 & OOS>0): also report 2× cost + per-year Sharpe
|
||||
(regime robustness). Survivorship caveat: v1 = currently-liquid majors over history.
|
||||
"""
|
||||
import glob
|
||||
import math
|
||||
import os
|
||||
import sys
|
||||
|
||||
import numpy as np
|
||||
|
||||
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
|
||||
from signal_sweep import xs_weights, pnl_w, validate, sharpe_t # noqa: E402
|
||||
import torch # noqa: E402
|
||||
|
||||
PRIOR_TRIALS = 17 # Batch 1 (futures factor zoo)
|
||||
DEV = "cuda" if torch.cuda.is_available() else "cpu"
|
||||
|
||||
|
||||
MIN_FUNDCOV = float(os.environ.get("MIN_FUNDCOV", "0.0")) # filter coins by funding coverage
|
||||
MIN_DAYS = int(os.environ.get("MIN_DAYS", "0")) # ...and minimum history
|
||||
|
||||
|
||||
def load_crypto():
|
||||
syms, data = [], {}
|
||||
for p in sorted(glob.glob("data/surfer/crypto/*.npz")):
|
||||
d = np.load(p); s = p.split("/")[-1][:-4]
|
||||
fund = d["funding"].astype(float)
|
||||
if len(d["day"]) < MIN_DAYS or np.mean(fund != 0) < MIN_FUNDCOV:
|
||||
continue
|
||||
data[s] = (d["day"].astype(np.int64), d["close"].astype(float), fund)
|
||||
syms.append(s)
|
||||
syms = sorted(syms)
|
||||
days = np.array(sorted(set().union(*[set(data[s][0].tolist()) for s in syms])))
|
||||
di = {d: i for i, d in enumerate(days.tolist())}
|
||||
T, N = len(days), len(syms)
|
||||
close = np.full((T, N), np.nan); fund = np.full((T, N), np.nan)
|
||||
for j, s in enumerate(syms):
|
||||
dd, cc, ff = data[s]
|
||||
for k in range(len(dd)):
|
||||
r = di[int(dd[k])]; close[r, j] = cc[k]; fund[r, j] = ff[k]
|
||||
return syms, days, close, fund
|
||||
|
||||
|
||||
def trailing(lc, L):
|
||||
out = np.full_like(lc, np.nan); out[L:] = lc[L:] - lc[:-L]; return out
|
||||
|
||||
|
||||
def zc(x):
|
||||
mu = np.nanmean(x, axis=1, keepdims=True); sd = np.nanstd(x, axis=1, keepdims=True)
|
||||
return np.nan_to_num((x - mu) / np.where(sd > 0, sd, 1))
|
||||
|
||||
|
||||
def main():
|
||||
syms, days, close, fund = load_crypto()
|
||||
T, N = close.shape
|
||||
lc = np.log(close)
|
||||
R = np.zeros((T, N)); R[1:] = lc[1:] - lc[:-1]; R = np.where(np.isfinite(R), R, 0.0)
|
||||
fund = np.where(np.isfinite(fund), fund, 0.0)
|
||||
Reff = R - fund
|
||||
vol30 = np.full_like(lc, np.nan)
|
||||
for t in range(30, T):
|
||||
vol30[t] = np.nanstd(R[t - 30:t], axis=0)
|
||||
|
||||
def tmf(K): # trailing mean funding over K days
|
||||
out = np.full_like(fund, np.nan)
|
||||
for t in range(K, T):
|
||||
out[t] = np.nanmean(fund[t - K:t], axis=0)
|
||||
return out
|
||||
|
||||
f7, f3, f14 = tmf(7), tmf(3), tmf(14)
|
||||
W = {} # name -> weights[T,N]
|
||||
for K in [3, 7, 14, 30]:
|
||||
W[f"XS_carry_{K}"] = xs_weights(-tmf(K))
|
||||
for L in [7, 30, 90]:
|
||||
W[f"XS_mom_{L}"] = xs_weights(trailing(lc, L))
|
||||
W["XS_rev_3"] = xs_weights(-trailing(lc, 3))
|
||||
W["XS_fundmom"] = xs_weights(f3 - f14) # rising funding (positioning building)
|
||||
W["XS_lowvol"] = xs_weights(-vol30)
|
||||
W["XS_carry+mom30"] = xs_weights(zc(-f7) + zc(trailing(lc, 30)))
|
||||
W["XS_carry+rev3"] = xs_weights(zc(-f7) + zc(-trailing(lc, 3)))
|
||||
W["XS_carry+lowvol"] = xs_weights(zc(-f7) + zc(-vol30))
|
||||
|
||||
NT = PRIOR_TRIALS + len(W)
|
||||
year = (1970 + days / 365.25).astype(int)
|
||||
rows = []
|
||||
pnls = {}
|
||||
for nm, w in W.items():
|
||||
pnl = pnl_w(w, Reff, cost_bp=10)
|
||||
pnls[nm] = (w, pnl)
|
||||
rows.append((nm, validate(pnl, days, NT)))
|
||||
rows.sort(key=lambda r: -(r[1]["oos"] if not math.isnan(r[1]["oos"]) else -9))
|
||||
print(f"\n===== CRYPTO SWEEP (broadened) — {N} perps, {T}d ({days.min()}..{days.max()}), fundcov={np.mean(fund!=0):.2f} =====")
|
||||
print(f"N_trials(cumulative deflation) = {NT}")
|
||||
print(f"{'signal':>16} {'full':>6} {'IS':>6} {'OOS':>6} {'CPCVmed':>8} {'5%':>6} {'DSR':>5}")
|
||||
for nm, v in rows:
|
||||
print(f"{nm:>16} {v['full']:>+6.2f} {v['is_']:>+6.2f} {v['oos']:>+6.2f} {v['med']:>+8.2f} {v['p5']:>+6.2f} {v['dsr']:>5.2f}")
|
||||
|
||||
print("\n----- ROBUSTNESS for develop-grade signals (CPCVmed>0 & IS>0 & OOS>0) -----")
|
||||
yrs = sorted(set(year[1:].tolist()))
|
||||
print(f"{'signal':>16} {'full10bp':>9} {'full20bp':>9} | per-year Sharpe: " + " ".join(f"{y}" for y in yrs))
|
||||
for nm, v in rows:
|
||||
if v["med"] > 0 and v["is_"] > 0 and v["oos"] > 0:
|
||||
w, pnl = pnls[nm]
|
||||
pnl2 = pnl_w(w, Reff, cost_bp=20)
|
||||
s10 = sharpe_t(torch.tensor(pnl, device=DEV, dtype=torch.float64))
|
||||
s20 = sharpe_t(torch.tensor(pnl2, device=DEV, dtype=torch.float64))
|
||||
py = []
|
||||
for y in yrs:
|
||||
m = year[1:] == y
|
||||
py.append(sharpe_t(torch.tensor(pnl[m], device=DEV, dtype=torch.float64)) if m.sum() > 30 else float("nan"))
|
||||
print(f"{nm:>16} {s10:>+9.2f} {s20:>+9.2f} | " + " ".join(f"{p:>+5.2f}" if not math.isnan(p) else " n/a" for p in py))
|
||||
|
||||
surv = [nm for nm, v in rows if v["dsr"] > 0.95 and v["oos"] > 0 and v["med"] > 0]
|
||||
print(f"\nDEPLOY survivors (DSR>0.95 & OOS>0 & CPCVmed>0): {surv if surv else 'NONE'}")
|
||||
print("Develop-grade = robust real edge worth building; deploy-grade = DSR>0.95 (harsh, deflated by all trials).")
|
||||
print("Caveat: survivorship (current majors); confirm point-in-time next.")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
120
scripts/surfer/crypto_trend_sizing.py
Normal file
120
scripts/surfer/crypto_trend_sizing.py
Normal file
@@ -0,0 +1,120 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Crypto TS-trend sizing — reconfirm the validated return engine + find the deployable CAGR.
|
||||
|
||||
The TSMOM crypto sleeve is VALIDATED (memory: Sharpe ~1.23, CAGR 75.7% @ 61% vol, corr~0 to funding).
|
||||
The 76% CAGR is just 61%-vol sizing, not a free lunch. This re-runs the SAME pre-registered config
|
||||
(LONG/FLAT, lookbacks 20/60/120, on crypto_pit daily close, liquidity floor on qvol) and sweeps a
|
||||
vol-target overlay to answer the #2 question: at a SANE vol target, what is the deployable CAGR and
|
||||
does the Sharpe hold (it should — vol-target is a rescale, but per-period sizing adds timing).
|
||||
|
||||
NOT a search: the 20/60/120 long/flat config is replicated verbatim from the validated harness.
|
||||
"""
|
||||
import glob
|
||||
import math
|
||||
|
||||
import numpy as np
|
||||
|
||||
PIT = "data/surfer/crypto_pit"
|
||||
LOOKBACKS = [20, 60, 120]
|
||||
COST_BP = 5.0
|
||||
ADV_FLOOR = 5e6 # $5M trailing quote-volume liquidity floor
|
||||
VOL_TARGETS = [0.10, 0.15, 0.20]
|
||||
LEV_CAP = 3.0
|
||||
PPY = 365 # crypto trades every day
|
||||
|
||||
|
||||
def build():
|
||||
days = set()
|
||||
raw = {}
|
||||
for f in sorted(glob.glob(f"{PIT}/*.npz")):
|
||||
d = np.load(f)
|
||||
if len(d["day"]) < 130:
|
||||
continue
|
||||
sym = f.split("/")[-1][:-4]
|
||||
raw[sym] = {int(dd): (c, q) for dd, c, q in zip(d["day"], d["close"], d["qvol"])}
|
||||
days.update(int(x) for x in d["day"])
|
||||
days = sorted(days)
|
||||
di = {d: i for i, d in enumerate(days)}
|
||||
syms = sorted(raw)
|
||||
C = np.full((len(days), len(syms)), np.nan)
|
||||
V = np.full((len(days), len(syms)), np.nan)
|
||||
for j, s in enumerate(syms):
|
||||
for dd, (c, q) in raw[s].items():
|
||||
C[di[dd], j] = c; V[di[dd], j] = q
|
||||
return np.array(days), syms, C, V
|
||||
|
||||
|
||||
def maxdd(equity):
|
||||
peak = np.maximum.accumulate(equity)
|
||||
return float((equity / peak - 1).min())
|
||||
|
||||
|
||||
def run():
|
||||
days, syms, C, V = build()
|
||||
T, N = C.shape
|
||||
logC = np.log(C)
|
||||
ret = np.full((T, N), np.nan)
|
||||
ret[1:] = C[1:] / C[:-1] - 1.0
|
||||
# long/flat signal = mean over lookbacks of 1[trailing-L log return > 0]
|
||||
sig = np.zeros((T, N))
|
||||
cnt = np.zeros((T, N))
|
||||
for L in LOOKBACKS:
|
||||
tr = np.full((T, N), np.nan)
|
||||
tr[L:] = logC[L:] - logC[:-L]
|
||||
on = np.where(np.isfinite(tr), (tr > 0).astype(float), np.nan)
|
||||
m = np.isfinite(on)
|
||||
sig[m] += on[m]; cnt[m] += 1
|
||||
sig = np.where(cnt > 0, sig / np.maximum(cnt, 1), 0.0) # 0..1 long/flat conviction
|
||||
|
||||
print(f"\n===== CRYPTO TS-TREND SIZING (long/flat {LOOKBACKS}, ${ADV_FLOOR/1e6:.0f}M ADV floor, {COST_BP}bp/leg) =====")
|
||||
print(f"panel: {T} days x {N} coins (survivorship-free incl. dead)\n")
|
||||
|
||||
# base book (gross 1, daily rebalance, equal-weight by conviction among liquid+on coins)
|
||||
w_prev = np.zeros(N)
|
||||
book = np.zeros(T)
|
||||
for t in range(1, T):
|
||||
eligible = np.isfinite(ret[t]) & np.isfinite(V[t - 1]) & (V[t - 1] > ADV_FLOOR)
|
||||
s = sig[t - 1] * eligible
|
||||
g = s.sum()
|
||||
w = s / g if g > 0 else np.zeros(N)
|
||||
turn = np.abs(w - w_prev).sum()
|
||||
book[t] = float(np.nansum(w * ret[t])) - turn * COST_BP / 1e4
|
||||
w_prev = w
|
||||
|
||||
valid = np.arange(T) > max(LOOKBACKS)
|
||||
base = book[valid]
|
||||
base_sr = base.mean() / base.std() * math.sqrt(PPY) if base.std() > 0 else float("nan")
|
||||
base_vol = base.std() * math.sqrt(PPY)
|
||||
base_cagr = float(np.prod(1 + base) ** (PPY / len(base)) - 1)
|
||||
print(f"BASE (un-vol-targeted): Sharpe {base_sr:+.2f} vol {base_vol*100:.0f}% CAGR {base_cagr*100:+.1f}% maxDD {maxdd(np.cumprod(1+base))*100:.1f}%")
|
||||
print(f" (memory reference: Sharpe ~1.23, CAGR ~76% @ ~61% vol — replicating)\n")
|
||||
|
||||
print(f"{'vol_tgt':>8} {'Sharpe':>7} {'CAGR':>7} {'real_vol':>9} {'maxDD':>7} {'avg_lev':>8} {'per-year SR':>30}")
|
||||
print("-" * 86)
|
||||
# vol-target overlay on the base book (trailing 30d realized, lagged, leverage-capped)
|
||||
rv = np.full(T, np.nan)
|
||||
for t in range(31, T):
|
||||
w = book[t - 30:t]
|
||||
rv[t] = w.std() * math.sqrt(PPY)
|
||||
yrs = 1970 + days // 365
|
||||
for vt in VOL_TARGETS:
|
||||
lev = np.where(np.isfinite(rv) & (rv > 0), np.clip(vt / rv, 0, LEV_CAP), 0.0)
|
||||
tgt = book * np.concatenate([[0], lev[:-1]]) # lag leverage by 1 day
|
||||
s = tgt[valid]
|
||||
sr = s.mean() / s.std() * math.sqrt(PPY) if s.std() > 0 else float("nan")
|
||||
vol = s.std() * math.sqrt(PPY)
|
||||
cagr = float(np.prod(1 + s) ** (PPY / len(s)) - 1)
|
||||
dd = maxdd(np.cumprod(1 + s))
|
||||
avglev = float(np.mean(lev[valid]))
|
||||
yv = {}
|
||||
for r, y in zip(s, yrs[valid]):
|
||||
yv.setdefault(int(y), []).append(r)
|
||||
ystr = " ".join(f"{y}:{(np.mean(v)/(np.std(v)+1e-12)*math.sqrt(PPY)):+.1f}"
|
||||
for y, v in sorted(yv.items()) if len(v) >= 30)
|
||||
print(f"{vt*100:>6.0f}% {sr:>+7.2f} {cagr*100:>+6.1f}% {vol*100:>8.0f}% {dd*100:>+6.1f}% {avglev:>8.2f} {ystr}")
|
||||
print("-" * 86)
|
||||
print("Deployable read: pick the vol target whose maxDD you can stomach; Sharpe should ~hold across targets.")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
run()
|
||||
93
scripts/surfer/energy_battery_gate.py
Normal file
93
scripts/surfer/energy_battery_gate.py
Normal file
@@ -0,0 +1,93 @@
|
||||
#!/usr/bin/env python3
|
||||
"""First energy gate: is battery arbitrage a real, persistent, engine-suited edge?
|
||||
|
||||
Free DE-LU day-ahead hourly prices (Fraunhofer Energy-Charts, no key). For a 4h/1MW battery
|
||||
(1 cycle/day, 85% round-trip): daily arbitrage value under PERFECT FORESIGHT (upper bound) vs a
|
||||
NAIVE fixed-hours heuristic (charge night / discharge evening). The gap = the value forecasting/RL
|
||||
could add. Annualize vs ~e50-80k/yr/MW needed to justify a battery. Trend = is it being competed away?
|
||||
"""
|
||||
import datetime
|
||||
import json
|
||||
import math
|
||||
import os
|
||||
import time
|
||||
import urllib.request
|
||||
|
||||
import numpy as np
|
||||
|
||||
CACHE = "data/surfer/energy"
|
||||
ETA = 0.85 # round-trip efficiency
|
||||
HRS = 4 # 4h battery, 1 MW -> 4 MWh / cycle
|
||||
ed = math.sqrt(ETA)
|
||||
|
||||
|
||||
def get(u):
|
||||
for attempt in range(6):
|
||||
try:
|
||||
return json.load(urllib.request.urlopen(urllib.request.Request(u, headers={"User-Agent": "curl/8"}), timeout=60))
|
||||
except urllib.error.HTTPError as e:
|
||||
if e.code == 429:
|
||||
time.sleep(8 * (attempt + 1)); continue
|
||||
raise
|
||||
raise RuntimeError("rate-limited after retries")
|
||||
|
||||
|
||||
def fetch():
|
||||
os.makedirs(CACHE, exist_ok=True)
|
||||
ts, px = [], []
|
||||
for yr in range(2019, 2026):
|
||||
f = f"{CACHE}/de_{yr}.json"
|
||||
if os.path.exists(f):
|
||||
d = json.load(open(f))
|
||||
else:
|
||||
d = get(f"https://api.energy-charts.info/price?bzn=DE-LU&start={yr}-01-01&end={yr}-12-31")
|
||||
json.dump({"unix_seconds": d.get("unix_seconds", []), "price": d.get("price", [])}, open(f, "w"))
|
||||
time.sleep(5) # space requests (avoid 429)
|
||||
ts += d["unix_seconds"]; px += d["price"]
|
||||
return np.array(ts), np.array(px, dtype=float)
|
||||
|
||||
|
||||
def main():
|
||||
ts, px = fetch()
|
||||
ok = np.isfinite(px)
|
||||
ts, px = ts[ok], px[ok]
|
||||
day = np.array([datetime.datetime.utcfromtimestamp(int(t)).date().toordinal() for t in ts])
|
||||
udays = np.unique(day)
|
||||
rows, dord = [], []
|
||||
for d in udays:
|
||||
h = px[day == d]
|
||||
if len(h) == 24: # drop DST-transition days (23/25h)
|
||||
rows.append(h); dord.append(d)
|
||||
P = np.array(rows); D = np.array(dord) # [days, 24]
|
||||
yr = np.array([datetime.date.fromordinal(int(x)).year for x in D])
|
||||
print(f"DE day-ahead: {len(P)} clean days ({datetime.date.fromordinal(int(D[0]))}..{datetime.date.fromordinal(int(D[-1]))})")
|
||||
|
||||
# perfect-foresight 4h arbitrage per day (1 MW): discharge top-4 hrs, charge bottom-4 hrs
|
||||
srt = np.sort(P, axis=1)
|
||||
fore = srt[:, -HRS:].sum(1) * ed - srt[:, :HRS].sum(1) / ed # EUR/day per MW
|
||||
# naive fixed-hours heuristic: charge 02-05h, discharge 18-21h (typical shape)
|
||||
naive = P[:, 18:22].sum(1) * ed - P[:, 2:6].sum(1) / ed
|
||||
spread = P.max(1) - P.min(1)
|
||||
|
||||
def stats(name, v):
|
||||
ann = np.mean(v) * 365
|
||||
print(f" {name:>16}: e{np.mean(v):6.1f}/day/MW -> e{ann/1000:6.1f}k/yr/MW (median e{np.median(v):.1f}/day)")
|
||||
|
||||
print(f"\n===== BATTERY ARBITRAGE ({HRS}h/1MW, {int(ETA*100)}% round-trip) =====")
|
||||
print(f"daily price spread: mean e{spread.mean():.1f}/MWh median e{np.median(spread):.1f} (negative-price days: {100*np.mean(P.min(1)<0):.0f}%)")
|
||||
stats("perfect-foresight", fore)
|
||||
stats("naive fixed-hours", naive)
|
||||
print(f" forecasting/RL gap (foresight-naive): e{(fore.mean()-naive.mean())*365/1000:.1f}k/yr/MW "
|
||||
f"= {100*(fore.mean()-naive.mean())/fore.mean():.0f}% of the value")
|
||||
print(f"\n viability: a 4h/1MW battery ~e400-600k capex; needs ~e50-80k/yr/MW arbitrage over 10y.")
|
||||
print(f" per-year perfect-foresight arb value (e k/yr/MW):")
|
||||
for y in range(2019, 2026):
|
||||
m = yr == y
|
||||
if m.sum() > 100:
|
||||
print(f" {y}: e{np.mean(fore[m])*365/1000:5.1f}k (naive e{np.mean(naive[m])*365/1000:.1f}k)")
|
||||
print("\nVERDICT: foresight value >> e50-80k threshold = battery arb economically real; big foresight-naive gap")
|
||||
print("= forecasting/RL adds real value (the engine's role); persistent across years = not yet competed away.")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
102
scripts/surfer/energy_capture_gate.py
Normal file
102
scripts/surfer/energy_capture_gate.py
Normal file
@@ -0,0 +1,102 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Realistic-capture gate: how much of perfect-foresight battery value does a real day-ahead
|
||||
forecast + dispatch capture? (The make-or-break for the energy pivot.)
|
||||
|
||||
Day-ahead model: commit tomorrow's charge/discharge schedule from a FORECAST of tomorrow's 24
|
||||
prices, realize P&L against ACTUAL prices. Ladder of forecasters: persistence -> last-week ->
|
||||
seasonal climatology -> walk-forward ML (gradient boosting). Report each one's realized capture
|
||||
% of perfect foresight, EUR k/yr/MW, per-year. Leak-free (ML trained only on past, predict forward).
|
||||
"""
|
||||
import datetime
|
||||
import math
|
||||
import os
|
||||
import sys
|
||||
|
||||
import numpy as np
|
||||
|
||||
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
|
||||
from energy_battery_gate import fetch # cached DE prices # noqa: E402
|
||||
from sklearn.ensemble import HistGradientBoostingRegressor # noqa: E402
|
||||
|
||||
ETA = 0.85
|
||||
HRS = 4
|
||||
ed = math.sqrt(ETA)
|
||||
|
||||
|
||||
def dispatch_value(yhat_day, actual_day):
|
||||
"""Schedule from forecast (charge HRS cheapest, discharge HRS priciest), realize on ACTUAL."""
|
||||
ch = np.argsort(yhat_day)[:HRS]
|
||||
dis = np.argsort(yhat_day)[-HRS:]
|
||||
return actual_day[dis].sum() * ed - actual_day[ch].sum() / ed
|
||||
|
||||
|
||||
def main():
|
||||
ts, px = fetch()
|
||||
ok = np.isfinite(px); ts, px = ts[ok], px[ok]
|
||||
day = np.array([datetime.datetime.utcfromtimestamp(int(t)).date().toordinal() for t in ts])
|
||||
rows, dord = [], []
|
||||
for d in np.unique(day):
|
||||
h = px[day == d]
|
||||
if len(h) == 24:
|
||||
rows.append(h); dord.append(d)
|
||||
P = np.array(rows); D = np.array(dord); ND = len(P)
|
||||
dow = np.array([datetime.date.fromordinal(int(x)).weekday() for x in D])
|
||||
month = np.array([datetime.date.fromordinal(int(x)).month for x in D])
|
||||
yr = np.array([datetime.date.fromordinal(int(x)).year for x in D])
|
||||
foresight = np.array([np.sort(P[d])[-HRS:].sum() * ed - np.sort(P[d])[:HRS].sum() / ed for d in range(ND)])
|
||||
|
||||
# ---- forecasters (each -> yhat[ND,24]; only valid from d>=7) ----
|
||||
yh = {}
|
||||
yh["persistence"] = np.roll(P, 1, axis=0)
|
||||
yh["last_week"] = np.roll(P, 7, axis=0)
|
||||
clim = np.full_like(P, np.nan) # trailing 28d climatology by weekend/weekday
|
||||
for d in range(14, ND):
|
||||
wk = dow[d] >= 5
|
||||
past = [k for k in range(max(0, d - 28), d) if (dow[k] >= 5) == wk]
|
||||
if past:
|
||||
clim[d] = P[past].mean(0)
|
||||
yh["climatology"] = clim
|
||||
|
||||
# ---- ML forecaster (walk-forward, leak-free) ----
|
||||
roll7 = np.full_like(P, np.nan)
|
||||
for d in range(7, ND):
|
||||
roll7[d] = P[d - 7:d].mean(0)
|
||||
# feature builder per (day d, hour h), causal
|
||||
def feats(d):
|
||||
y1, y7, y2 = P[d - 1], P[d - 7], P[d - 2]
|
||||
base = np.array([dow[d], month[d], 1.0 if dow[d] >= 5 else 0.0,
|
||||
y1.mean(), y1.max() - y1.min(), y7.mean(), roll7[d].mean()])
|
||||
X = np.zeros((24, 7 + 4))
|
||||
for h in range(24):
|
||||
X[h] = np.concatenate([[h], base[:1], base[1:], [y1[h], y7[h], y2[h], roll7[d, h]]])[:11]
|
||||
return X
|
||||
ml = np.full_like(P, np.nan)
|
||||
INIT, STEP = 365, 90
|
||||
Xcache = {d: feats(d) for d in range(7, ND)}
|
||||
for s in range(INIT, ND, STEP):
|
||||
e = min(s + STEP, ND)
|
||||
Xtr = np.vstack([Xcache[d] for d in range(7, s)])
|
||||
ytr = np.concatenate([P[d] for d in range(7, s)])
|
||||
gb = HistGradientBoostingRegressor(max_depth=4, max_iter=200, learning_rate=0.05,
|
||||
min_samples_leaf=50).fit(Xtr, ytr)
|
||||
for d in range(s, e):
|
||||
ml[d] = gb.predict(Xcache[d])
|
||||
yh["ML walk-fwd"] = ml
|
||||
|
||||
print(f"\n===== REALISTIC-CAPTURE GATE (DE 4h/1MW day-ahead, {ND} days) =====")
|
||||
print(f"perfect-foresight: e{foresight.mean()*365/1000:.1f}k/yr/MW")
|
||||
print(f"{'forecaster':>14} {'capture%':>8} {'EURk/yr/MW':>11} | per-year capture%")
|
||||
valid = np.arange(INIT, ND) # compare all on the ML-valid window
|
||||
for nm, yhat in yh.items():
|
||||
vv = np.array([dispatch_value(yhat[d], P[d]) for d in valid])
|
||||
fv = foresight[valid]
|
||||
cap = vv.sum() / fv.sum()
|
||||
py = " ".join(f"{y}:{100*np.array([dispatch_value(yhat[d],P[d]) for d in valid[yr[valid]==y]]).sum()/foresight[valid[yr[valid]==y]].sum():.0f}%"
|
||||
for y in range(2020, 2026) if (yr[valid] == y).sum() > 100)
|
||||
print(f"{nm:>14} {100*cap:>7.0f}% {vv.mean()*365/1000:>10.1f} | {py}")
|
||||
print("\nVERDICT: ML capture >> simple baselines AND >=60% of foresight (~e55-80k/yr/MW) = engine earns its keep.")
|
||||
print("If even ML captures little, or simple climatology already gets most, the engine adds little here.")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
142
scripts/surfer/energy_capture_gate2.py
Normal file
142
scripts/surfer/energy_capture_gate2.py
Normal file
@@ -0,0 +1,142 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Harder gate: does the engine beat the climatology heuristic when given the REAL price drivers
|
||||
(weather -> renewables)? And specifically on the high-volatility days where the value concentrates?
|
||||
|
||||
Adds free Open-Meteo weather (wind@100m, solar radiation, temp) for 3 German points to the
|
||||
day-ahead forecaster. Compares: climatology (the 83% champ) vs ML-price-only (76%) vs
|
||||
ML-with-weather. Capture % overall AND on the top-20% highest-spread days (where wind-drought
|
||||
spikes live and fundamentals should matter). Leak-free walk-forward. Weather@day-d is the
|
||||
forecast you'd have at decision time (day-ahead weather forecasts are ~90%+ accurate).
|
||||
"""
|
||||
import datetime
|
||||
import json
|
||||
import math
|
||||
import os
|
||||
import sys
|
||||
import time
|
||||
import urllib.request
|
||||
|
||||
import numpy as np
|
||||
|
||||
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
|
||||
from energy_battery_gate import fetch # cached prices # noqa: E402
|
||||
from sklearn.ensemble import HistGradientBoostingRegressor # noqa: E402
|
||||
|
||||
ETA = 0.85; HRS = 4; ed = math.sqrt(ETA)
|
||||
WCACHE = "data/surfer/energy/weather"
|
||||
PTS = {"north": (53.55, 9.99), "central": (50.1, 8.68), "south": (48.14, 11.58)}
|
||||
|
||||
|
||||
def wget(u):
|
||||
for a in range(6):
|
||||
try:
|
||||
return json.load(urllib.request.urlopen(urllib.request.Request(u, headers={"User-Agent": "curl/8"}), timeout=90))
|
||||
except urllib.error.HTTPError as e:
|
||||
if e.code == 429:
|
||||
time.sleep(10 * (a + 1)); continue
|
||||
raise
|
||||
raise RuntimeError("rate-limited")
|
||||
|
||||
|
||||
def weather():
|
||||
os.makedirs(WCACHE, exist_ok=True)
|
||||
acc = {}
|
||||
for nm, (la, lo) in PTS.items():
|
||||
f = f"{WCACHE}/{nm}.json"
|
||||
if os.path.exists(f):
|
||||
d = json.load(open(f))
|
||||
else:
|
||||
u = (f"https://archive-api.open-meteo.com/v1/archive?latitude={la}&longitude={lo}"
|
||||
f"&start_date=2019-01-01&end_date=2025-09-29"
|
||||
f"&hourly=wind_speed_100m,shortwave_radiation,temperature_2m&timezone=UTC")
|
||||
d = wget(u)["hourly"]; json.dump(d, open(f, "w")); time.sleep(3)
|
||||
acc[nm] = d
|
||||
return acc
|
||||
|
||||
|
||||
def dispatch(yhat, actual):
|
||||
ch = np.argsort(yhat)[:HRS]; dis = np.argsort(yhat)[-HRS:]
|
||||
return actual[dis].sum() * ed - actual[ch].sum() / ed
|
||||
|
||||
|
||||
def main():
|
||||
ts, px = fetch()
|
||||
ok = np.isfinite(px); ts, px = ts[ok], px[ok]
|
||||
pday = np.array([datetime.datetime.utcfromtimestamp(int(t)).date().toordinal() for t in ts])
|
||||
phour = np.array([datetime.datetime.utcfromtimestamp(int(t)).hour for t in ts])
|
||||
# weather aligned to (ordinal_day, hour), averaged over 3 points
|
||||
W = weather()
|
||||
wind = {}; sol = {}; tmp = {}
|
||||
for nm, d in W.items():
|
||||
for i, tstr in enumerate(d["time"]):
|
||||
dt = datetime.datetime.strptime(tstr, "%Y-%m-%dT%H:%M")
|
||||
k = (dt.date().toordinal(), dt.hour)
|
||||
wind.setdefault(k, []).append(d["wind_speed_100m"][i] or 0.0)
|
||||
sol.setdefault(k, []).append(d["shortwave_radiation"][i] or 0.0)
|
||||
tmp.setdefault(k, []).append(d["temperature_2m"][i] or 0.0)
|
||||
|
||||
rows, dord = [], []
|
||||
for dd in np.unique(pday):
|
||||
h = px[pday == dd]
|
||||
if len(h) == 24 and all((dd, hr) in wind for hr in range(24)):
|
||||
rows.append(h); dord.append(dd)
|
||||
P = np.array(rows); D = np.array(dord); ND = len(P)
|
||||
Wd = np.array([[np.mean(wind[(d, h)]) for h in range(24)] for d in D])
|
||||
Sd = np.array([[np.mean(sol[(d, h)]) for h in range(24)] for d in D])
|
||||
Td = np.array([[np.mean(tmp[(d, h)]) for h in range(24)] for d in D])
|
||||
dow = np.array([datetime.date.fromordinal(int(x)).weekday() for x in D])
|
||||
month = np.array([datetime.date.fromordinal(int(x)).month for x in D])
|
||||
yr = np.array([datetime.date.fromordinal(int(x)).year for x in D])
|
||||
fore = np.array([np.sort(P[d])[-HRS:].sum() * ed - np.sort(P[d])[:HRS].sum() / ed for d in range(ND)])
|
||||
spread = P.max(1) - P.min(1)
|
||||
print(f"DE prices+weather aligned: {ND} days")
|
||||
|
||||
roll7 = np.full_like(P, np.nan)
|
||||
for d in range(7, ND):
|
||||
roll7[d] = P[d - 7:d].mean(0)
|
||||
|
||||
def feats(d, with_w):
|
||||
y1, y7 = P[d - 1], P[d - 7]
|
||||
X = np.zeros((24, 12 if with_w else 8))
|
||||
for h in range(24):
|
||||
f = [h, dow[d], month[d], y1[h], y7[h], y1.mean(), y7.mean(), roll7[d, h]]
|
||||
if with_w:
|
||||
f += [Wd[d, h], Sd[d, h], Td[d, h], Wd[d].mean()]
|
||||
X[h] = f
|
||||
return X
|
||||
|
||||
INIT, STEP = 365, 90
|
||||
yhat = {"price_only": np.full_like(P, np.nan), "with_weather": np.full_like(P, np.nan)}
|
||||
for with_w, key in [(False, "price_only"), (True, "with_weather")]:
|
||||
Xc = {d: feats(d, with_w) for d in range(7, ND)}
|
||||
for s in range(INIT, ND, STEP):
|
||||
e = min(s + STEP, ND)
|
||||
Xtr = np.vstack([Xc[d] for d in range(7, s)]); ytr = np.concatenate([P[d] for d in range(7, s)])
|
||||
gb = HistGradientBoostingRegressor(max_depth=4, max_iter=200, learning_rate=0.05, min_samples_leaf=50).fit(Xtr, ytr)
|
||||
for d in range(s, e):
|
||||
yhat[key][d] = gb.predict(Xc[d])
|
||||
# climatology champ
|
||||
clim = np.full_like(P, np.nan)
|
||||
for d in range(14, ND):
|
||||
wk = dow[d] >= 5
|
||||
past = [k for k in range(max(0, d - 28), d) if (dow[k] >= 5) == wk]
|
||||
if past:
|
||||
clim[d] = P[past].mean(0)
|
||||
yhat["climatology"] = clim
|
||||
|
||||
valid = np.arange(INIT, ND)
|
||||
hi = valid[spread[valid] >= np.quantile(spread[valid], 0.80)] # top-20% volatile days
|
||||
print(f"\n===== HARDER CAPTURE GATE (fundamentals + volatile-day focus), {len(valid)} days =====")
|
||||
print(f"perfect-foresight e{fore.mean()*365/1000:.1f}k/yr/MW | top-20%-vol days hold {100*fore[hi].sum()/fore[valid].sum():.0f}% of value")
|
||||
print(f"{'forecaster':>14} {'capture%':>8} {'EURk/yr':>8} {'capture% on HI-VOL days':>24}")
|
||||
for nm in ["climatology", "price_only", "with_weather"]:
|
||||
yh = yhat[nm]
|
||||
vv = np.array([dispatch(yh[d], P[d]) for d in valid]); cap = vv.sum() / fore[valid].sum()
|
||||
vh = np.array([dispatch(yh[d], P[d]) for d in hi]); caph = vh.sum() / fore[hi].sum()
|
||||
print(f"{nm:>14} {100*cap:>7.0f}% {vv.mean()*365/1000:>7.1f} {100*caph:>23.0f}%")
|
||||
print("\nVERDICT: with_weather > climatology (esp on HI-VOL days) = engine+fundamentals earns its keep.")
|
||||
print("If climatology still wins, even fundamentals don't beat the heuristic for day-ahead -> engine's home is elsewhere (intraday/multi-market).")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
98
scripts/surfer/energy_intraday_gate.py
Normal file
98
scripts/surfer/energy_intraday_gate.py
Normal file
@@ -0,0 +1,98 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Intraday gate: on CAISO real-time prices (less seasonal, forecast-error-driven), does the
|
||||
ENGINE finally beat the heuristics — unlike day-ahead, where climatology won?
|
||||
|
||||
Tests: (1) is RT more volatile than DA (more battery value)? (2) RT-price forecastability:
|
||||
baseline "RT=DA" vs climatology vs ML(walk-fwd, features incl. DA price + DART lags) ->
|
||||
capture % of perfect-foresight RT battery value. If ML >> baselines on RT, the engine earns
|
||||
its keep in the less-seasonal market. If nothing beats "RT=DA", intraday deviations are noise.
|
||||
"""
|
||||
import datetime
|
||||
import json
|
||||
import math
|
||||
import os
|
||||
import sys
|
||||
|
||||
import numpy as np
|
||||
|
||||
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
|
||||
from sklearn.ensemble import HistGradientBoostingRegressor # noqa: E402
|
||||
|
||||
ETA = 0.85; HRS = 4; ed = math.sqrt(ETA)
|
||||
OUT = "data/surfer/caiso"
|
||||
|
||||
|
||||
def load():
|
||||
dam = json.load(open(f"{OUT}/dam.json")); rtm = json.load(open(f"{OUT}/rtm.json"))
|
||||
keys = sorted(set(dam) & set(rtm)) # "YYYY-MM-DDHHH"
|
||||
# group into days with full 24h
|
||||
byday = {}
|
||||
for k in keys:
|
||||
d, h = k[:10], int(k[-2:])
|
||||
byday.setdefault(d, {})[h] = (dam[k], rtm[k])
|
||||
rows, dord = [], []
|
||||
for d in sorted(byday):
|
||||
if len(byday[d]) == 24:
|
||||
da = [byday[d][h][0] for h in range(1, 25)]
|
||||
rt = [byday[d][h][1] for h in range(1, 25)]
|
||||
rows.append((da, rt)); dord.append(d)
|
||||
DA = np.array([r[0] for r in rows]); RT = np.array([r[1] for r in rows])
|
||||
D = [datetime.date.fromisoformat(x) for x in dord]
|
||||
return DA, RT, D
|
||||
|
||||
|
||||
def dispatch(yhat, actual):
|
||||
ch = np.argsort(yhat)[:HRS]; dis = np.argsort(yhat)[-HRS:]
|
||||
return actual[dis].sum() * ed - actual[ch].sum() / ed
|
||||
|
||||
|
||||
def main():
|
||||
DA, RT, D = load()
|
||||
ND = len(DA)
|
||||
dow = np.array([d.weekday() for d in D]); month = np.array([d.month for d in D])
|
||||
sprDA = DA.max(1) - DA.min(1); sprRT = RT.max(1) - RT.min(1)
|
||||
foreDA = np.array([np.sort(DA[d])[-HRS:].sum() * ed - np.sort(DA[d])[:HRS].sum() / ed for d in range(ND)])
|
||||
foreRT = np.array([np.sort(RT[d])[-HRS:].sum() * ed - np.sort(RT[d])[:HRS].sum() / ed for d in range(ND)])
|
||||
print(f"CAISO NP15 {ND} days ({D[0]}..{D[-1]})")
|
||||
print(f"daily spread: DA mean ${sprDA.mean():.0f}/MWh RT mean ${sprRT.mean():.0f}/MWh (RT/DA = {sprRT.mean()/sprDA.mean():.1f}x)")
|
||||
print(f"perfect-foresight battery: DA ${foreDA.mean()*365/1000:.0f}k/yr/MW RT ${foreRT.mean()*365/1000:.0f}k/yr/MW")
|
||||
|
||||
# RT forecasters
|
||||
roll7 = np.full_like(RT, np.nan)
|
||||
for d in range(7, ND):
|
||||
roll7[d] = RT[d - 7:d].mean(0)
|
||||
dart = RT - DA
|
||||
def feats(d):
|
||||
X = np.zeros((24, 9))
|
||||
for h in range(24):
|
||||
X[h] = [h, dow[d], month[d], DA[d, h], RT[d - 1, h], RT[d - 7, h],
|
||||
dart[d - 1, h], DA[d].mean(), roll7[d, h]]
|
||||
return X
|
||||
ml = np.full_like(RT, np.nan); INIT, STEP = 60, 30
|
||||
Xc = {d: feats(d) for d in range(7, ND)}
|
||||
for s in range(INIT, ND, STEP):
|
||||
e = min(s + STEP, ND)
|
||||
Xtr = np.vstack([Xc[d] for d in range(7, s)]); ytr = np.concatenate([RT[d] for d in range(7, s)])
|
||||
gb = HistGradientBoostingRegressor(max_depth=4, max_iter=200, learning_rate=0.05, min_samples_leaf=40).fit(Xtr, ytr)
|
||||
for d in range(s, e):
|
||||
ml[d] = gb.predict(Xc[d])
|
||||
clim = np.full_like(RT, np.nan)
|
||||
for d in range(14, ND):
|
||||
wk = dow[d] >= 5
|
||||
past = [k for k in range(max(0, d - 28), d) if (dow[k] >= 5) == wk]
|
||||
if past:
|
||||
clim[d] = RT[past].mean(0)
|
||||
|
||||
valid = np.arange(INIT, ND)
|
||||
fc = {"baseline RT=DA": DA, "climatology": clim, "ML walk-fwd": ml}
|
||||
print(f"\n===== INTRADAY (RT) CAPTURE GATE — {len(valid)} days, perfect-foresight RT ${foreRT[valid].mean()*365/1000:.0f}k/yr/MW =====")
|
||||
print(f"{'RT forecaster':>16} {'capture%':>8} {'EURk/yr/MW':>11}")
|
||||
for nm, yh in fc.items():
|
||||
vv = np.array([dispatch(yh[d], RT[d]) for d in valid]); cap = vv.sum() / foreRT[valid].sum()
|
||||
print(f"{nm:>16} {100*cap:>7.0f}% {vv.mean()*365/1000:>10.0f}")
|
||||
print("\nVERDICT: ML capture >> 'RT=DA' baseline AND > climatology = the engine FINALLY earns its keep")
|
||||
print("(less-seasonal RT is where forecasting/RL beats heuristics). If ML ~ baselines, intraday is noise too.")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
99
scripts/surfer/equity_cohort_challenge.py
Normal file
99
scripts/surfer/equity_cohort_challenge.py
Normal file
@@ -0,0 +1,99 @@
|
||||
#!/usr/bin/env python3
|
||||
"""THE CHALLENGE: does the liquid high-volatility equity cohort fit our edge structure?
|
||||
|
||||
The equity analog of crypto's reflexive cohort: LIQUID (top by dollar-vol -> low cost, avoids the
|
||||
small-cap wall) AND HIGH realized vol (reflexive, retail-driven, less fundamentally-anchored ->
|
||||
where momentum can persist, unlike efficient large-caps). The one slice where reflexivity and
|
||||
liquidity OVERLAP. Test cross-sectional momentum / reversal / low-vol AND long-only momentum
|
||||
(no borrow) within this cohort, gross + net of realistic (liquid ~10bp) cost, OOS, per-year.
|
||||
Reuses the already-downloaded DBEQ data (free).
|
||||
"""
|
||||
import math
|
||||
import os
|
||||
import sys
|
||||
|
||||
import numpy as np
|
||||
|
||||
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
|
||||
from equity_factor_gate import load, roll, trailing # noqa: E402
|
||||
from signal_sweep import xs_weights, validate, sharpe_t # noqa: E402
|
||||
import torch # noqa: E402
|
||||
|
||||
DEV = "cuda" if torch.cuda.is_available() else "cpu"
|
||||
DV_FLOOR = 1e7 # $10M/day: liquid -> low cost (avoid the small-cap wall)
|
||||
VOL_PCTILE = 0.60 # keep names above 60th pctile realized vol (the reflexive cohort)
|
||||
COST_BP = 10.0 # liquid names: ~10bp round-trip (illiquidity wall avoided by construction)
|
||||
|
||||
|
||||
def main():
|
||||
insts, days, close, dvol = load()
|
||||
T, N = close.shape
|
||||
lc = np.log(close)
|
||||
R = np.zeros((T, N)); R[1:] = lc[1:] - lc[:-1]; R = np.where(np.isfinite(R), R, 0.0)
|
||||
dv30 = roll(np.mean, np.nan_to_num(dvol), 30)
|
||||
vol63 = roll(np.std, R, 63)
|
||||
year = (1970 + days / 365.25).astype(int)
|
||||
|
||||
# cohort: liquid AND high-vol (reflexive)
|
||||
univ = np.zeros((T, N), bool)
|
||||
for t in range(T):
|
||||
liq = (dv30[t] > DV_FLOOR) & np.isfinite(close[t]) & np.isfinite(vol63[t])
|
||||
e = np.where(liq)[0]
|
||||
if len(e) > 50:
|
||||
thr = np.quantile(vol63[t, e], VOL_PCTILE)
|
||||
univ[t, e[vol63[t, e] >= thr]] = True
|
||||
print(f"loaded {N} insts {T}d; cohort/day ~{int(univ.sum(1).mean())} (liquid>${DV_FLOOR/1e6:.0f}M & high-vol)")
|
||||
|
||||
def held_weekly(w, K=5):
|
||||
a = 2.0 / (5 + 1)
|
||||
for t in range(1, T):
|
||||
w[t] = a * w[t] + (1 - a) * w[t - 1]
|
||||
wh = w.copy(); last = 0
|
||||
for t in range(T):
|
||||
if t % K == 0:
|
||||
last = t
|
||||
wh[t] = wh[last] = w[last]
|
||||
return wh
|
||||
|
||||
def pnl_ls(sig, net=True): # long-short market-neutral
|
||||
s = sig.copy(); s[~univ] = np.nan
|
||||
w = held_weekly(xs_weights(s))
|
||||
g = np.sum(w[:-1] * R[1:], axis=1)
|
||||
if not net:
|
||||
return g
|
||||
return g - np.sum(np.abs(w[1:] - w[:-1]), axis=1) * COST_BP / 1e4
|
||||
|
||||
def pnl_long(sig, q=0.10, net=True): # long-only top-decile (no borrow)
|
||||
s = sig.copy(); s[~univ] = np.nan
|
||||
w = np.zeros((T, N))
|
||||
for t in range(T):
|
||||
e = np.where(np.isfinite(s[t]) & univ[t])[0]
|
||||
if len(e) > 20:
|
||||
k = max(int(q * len(e)), 5)
|
||||
top = e[np.argsort(-s[t, e])[:k]]; w[t, top] = 1.0 / k
|
||||
w = held_weekly(w)
|
||||
g = np.sum(w[:-1] * R[1:], axis=1)
|
||||
if net:
|
||||
g -= np.sum(np.abs(w[1:] - w[:-1]), axis=1) * COST_BP / 1e4
|
||||
return g
|
||||
|
||||
# cohort equal-weight benchmark (the beta of the cohort)
|
||||
ewb = np.array([R[t][univ[t - 1]].mean() if t > 0 and univ[t - 1].any() else 0.0 for t in range(T)])[1:]
|
||||
|
||||
T_ = lambda x: torch.tensor(x[np.isfinite(x)], device=DEV, dtype=torch.float64)
|
||||
sigs = {"mom_63_skip5": trailing(lc, 63, skip=5), "mom_126_skip5": trailing(lc, 126, skip=5),
|
||||
"reversal_5": -trailing(lc, 5), "lowvol_63": -vol63}
|
||||
print(f"\n===== LIQUID HIGH-VOL EQUITY COHORT — cross-sectional (net {COST_BP}bp) =====")
|
||||
print(f"cohort equal-weight (beta) Sharpe: {sharpe_t(T_(ewb)):+.2f}")
|
||||
print(f"{'factor':>16} {'L/S gross':>9} {'L/S NET':>8} {'OOS':>6} {'CPCVmed':>8} {'DSR':>5} {'LongOnly NET':>12} | per-year(L/S net)")
|
||||
for nm, sg in sigs.items():
|
||||
g = pnl_ls(sg, net=False); p = pnl_ls(sg, net=True); lo = pnl_long(sg, net=True)
|
||||
v = validate(p, days, 20)
|
||||
py = " ".join(f"{y}:{sharpe_t(T_(p[year[1:]==y])):+.1f}" for y in range(2023, 2027) if (year[1:] == y).sum() > 40)
|
||||
print(f"{nm:>16} {sharpe_t(T_(g)):>+9.2f} {v['full']:>+8.2f} {v['oos']:>+6.2f} {v['med']:>+8.2f} {v['dsr']:>5.2f} {sharpe_t(T_(lo)):>+12.2f} | {py}")
|
||||
print("\nVERDICT: a factor with L/S NET full+OOS+CPCVmed>0 & DSR>0.5 in the reflexive-liquid cohort = the fit.")
|
||||
print("If momentum still negative even here, the cohort doesn't rescue it (efficiency/regime, not cost).")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
125
scripts/surfer/equity_factor_gate.py
Normal file
125
scripts/surfer/equity_factor_gate.py
Normal file
@@ -0,0 +1,125 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Small/mid-cap US equity cross-sectional factor gate (DBEQ daily, realistic small-cap costs).
|
||||
|
||||
The less-efficient corner: exclude mega-caps (too efficient) and illiquid micro-caps
|
||||
(untradeable); keep the small/mid liquid band where your small capital is an advantage.
|
||||
Test XS momentum (12-1 style), short-term reversal, low-vol, residual momentum — GROSS and
|
||||
NET of ILLIQUIDITY-SCALED cost (small-cap spreads ~30-150bp, the honest killer). Weekly
|
||||
rebalance + smoothing. Point-in-time universe (incl delisted) -> survivorship-aware.
|
||||
"""
|
||||
import math
|
||||
import os
|
||||
import sys
|
||||
|
||||
import numpy as np
|
||||
|
||||
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
|
||||
from signal_sweep import xs_weights, validate, sharpe_t # noqa: E402
|
||||
import torch # noqa: E402
|
||||
|
||||
DEV = "cuda" if torch.cuda.is_available() else "cpu"
|
||||
DAY_NS = 86_400 * 10**9
|
||||
OUT = "data/surfer/dbeq_ohlcv1d.dbn"
|
||||
EXCLUDE_TOP = 50 # drop mega-caps (efficient)
|
||||
DV_FLOOR = 2e6 # $2M/day min (tradeable)
|
||||
TOPK = 600 # small/mid liquid band size
|
||||
|
||||
|
||||
def roll(fn, X, L):
|
||||
out = np.full_like(X, np.nan)
|
||||
for t in range(L, len(X)):
|
||||
out[t] = fn(X[t - L:t], axis=0)
|
||||
return out
|
||||
|
||||
|
||||
def trailing(lc, L, skip=0):
|
||||
out = np.full_like(lc, np.nan)
|
||||
if skip:
|
||||
out[L + skip:] = lc[L:-skip] - lc[:-(L + skip)]
|
||||
else:
|
||||
out[L:] = lc[L:] - lc[:-L]
|
||||
return out
|
||||
|
||||
|
||||
def load():
|
||||
import databento as db
|
||||
a = db.DBNStore.from_file(OUT).to_ndarray()
|
||||
iid = a["instrument_id"]; ts = a["ts_event"].astype(np.int64)
|
||||
close = a["close"].astype(np.float64) / 1e9
|
||||
vol = a["volume"].astype(np.float64)
|
||||
day = ts // DAY_NS
|
||||
days = np.unique(day); insts = np.unique(iid)
|
||||
dix = np.searchsorted(days, day); iix = np.searchsorted(insts, iid)
|
||||
T, N = len(days), len(insts)
|
||||
C = np.full((T, N), np.nan); DVOL = np.full((T, N), np.nan)
|
||||
C[dix, iix] = np.where(close > 0, close, np.nan)
|
||||
DVOL[dix, iix] = close * vol
|
||||
return insts, days, C, DVOL
|
||||
|
||||
|
||||
def main():
|
||||
insts, days, close, dvol = load()
|
||||
T, N = close.shape
|
||||
lc = np.log(close)
|
||||
R = np.zeros((T, N)); R[1:] = lc[1:] - lc[:-1]; R = np.where(np.isfinite(R), R, 0.0)
|
||||
dv30 = roll(np.mean, np.nan_to_num(dvol), 30)
|
||||
vol63 = roll(np.std, R, 63)
|
||||
year = (1970 + days / 365.25).astype(int)
|
||||
print(f"loaded {N} instruments, {T} days ({days.min()}..{days.max()})")
|
||||
|
||||
# point-in-time small/mid-cap universe: drop top mega-caps, require liquidity floor, take band
|
||||
univ = np.zeros((T, N), bool)
|
||||
for t in range(T):
|
||||
elig = np.where((dv30[t] > DV_FLOOR) & np.isfinite(close[t]))[0]
|
||||
if len(elig) > EXCLUDE_TOP + 20:
|
||||
order = elig[np.argsort(-dv30[t, elig])]
|
||||
band = order[EXCLUDE_TOP:EXCLUDE_TOP + TOPK] # skip mega-caps, take next TOPK
|
||||
univ[t, band] = True
|
||||
|
||||
# illiquidity-scaled round-trip cost (bp): small-caps 30-150bp, mid 5-30bp
|
||||
rt_cost = np.clip(60.0 / np.sqrt(np.maximum(dv30, 1.0) / 1e6), 5.0, 150.0) / 1e4
|
||||
|
||||
def held_weekly(wt, K=5):
|
||||
wh = wt.copy()
|
||||
last = 0
|
||||
for t in range(T):
|
||||
if t % K == 0:
|
||||
last = t
|
||||
wh[t] = wt[last]
|
||||
return wh
|
||||
|
||||
def pnl(sig, net=True):
|
||||
s = sig.copy(); s[~univ] = np.nan
|
||||
w = xs_weights(s)
|
||||
a = 2.0 / (5 + 1) # smooth span 5
|
||||
for t in range(1, T):
|
||||
w[t] = a * w[t] + (1 - a) * w[t - 1]
|
||||
w = held_weekly(w)
|
||||
gross = np.sum(w[:-1] * R[1:], axis=1)
|
||||
if not net:
|
||||
return gross
|
||||
turn = np.abs(w[1:] - w[:-1])
|
||||
cost = np.sum(turn * rt_cost[1:], axis=1)
|
||||
return gross - cost
|
||||
|
||||
sigs = {
|
||||
"mom_63_skip5": trailing(lc, 63, skip=5),
|
||||
"mom_126_skip5": trailing(lc, 126, skip=5),
|
||||
"reversal_5": -trailing(lc, 5),
|
||||
"lowvol_63": -vol63,
|
||||
}
|
||||
print(f"\n===== SMALL/MID-CAP EQUITY FACTOR GATE (band {EXCLUDE_TOP}-{EXCLUDE_TOP+TOPK}, ${DV_FLOOR/1e6:.0f}M floor) =====")
|
||||
print(f"universe/day ~{int(univ.sum(1).mean())}, weekly rebal+smooth, deflate N=20")
|
||||
print(f"{'factor':>16} {'gross':>6} {'NET':>6} {'IS':>6} {'OOS':>6} {'CPCVmed':>8} {'DSR':>5} | per-year")
|
||||
T_ = lambda x: torch.tensor(x[np.isfinite(x)], device=DEV, dtype=torch.float64)
|
||||
for nm, sg in sigs.items():
|
||||
g = pnl(sg, net=False); p = pnl(sg, net=True)
|
||||
gsr = sharpe_t(T_(g)); v = validate(p, days, 20)
|
||||
py = " ".join(f"{y}:{sharpe_t(T_(p[year[1:]==y])):+.1f}" for y in range(2023, 2027) if (year[1:] == y).sum() > 40)
|
||||
print(f"{nm:>16} {gsr:>+6.2f} {v['full']:>+6.2f} {v['is_']:>+6.2f} {v['oos']:>+6.2f} {v['med']:>+8.2f} {v['dsr']:>5.2f} | {py}")
|
||||
print("\nVERDICT: a factor with NET full+OOS+CPCVmed>0 & DSR>0.5 survives small-cap costs = real.")
|
||||
print("gross>>NET means the edge is eaten by illiquidity cost (the usual small-cap fate).")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
103
scripts/surfer/equity_lowvol_detail.py
Normal file
103
scripts/surfer/equity_lowvol_detail.py
Normal file
@@ -0,0 +1,103 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Decompose the equity low-vol +0.74 in detail: alpha vs beta, which leg, turnover, concentration.
|
||||
|
||||
Questions: (1) Is it cross-sectional ALPHA or a structural short-beta directional bet?
|
||||
(2) Which leg carries it — long low-vol (deployable, no borrow) or short high-vol (expensive)?
|
||||
(3) Turnover (is the cost charge fair?). (4) P&L concentration over time. (5) Beta-neutral residual.
|
||||
"""
|
||||
import math
|
||||
import os
|
||||
import sys
|
||||
|
||||
import numpy as np
|
||||
|
||||
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
|
||||
from equity_factor_gate import load, roll # noqa: E402
|
||||
from signal_sweep import validate, sharpe_t # noqa: E402
|
||||
import torch # noqa: E402
|
||||
|
||||
DEV = "cuda" if torch.cuda.is_available() else "cpu"
|
||||
DV_FLOOR, VOL_PCT, COST = 1e7, 0.60, 10.0
|
||||
|
||||
|
||||
def main():
|
||||
insts, days, close, dvol = load()
|
||||
T, N = close.shape
|
||||
lc = np.log(close)
|
||||
R = np.zeros((T, N)); R[1:] = lc[1:] - lc[:-1]; R = np.where(np.isfinite(R), R, 0.0)
|
||||
dv30 = roll(np.mean, np.nan_to_num(dvol), 30)
|
||||
vol63 = roll(np.std, R, 63)
|
||||
year = (1970 + days / 365.25).astype(int)
|
||||
T_ = lambda x: torch.tensor(np.asarray(x)[np.isfinite(np.asarray(x))], device=DEV, dtype=torch.float64)
|
||||
|
||||
univ = np.zeros((T, N), bool)
|
||||
for t in range(T):
|
||||
e = np.where((dv30[t] > DV_FLOOR) & np.isfinite(close[t]) & np.isfinite(vol63[t]))[0]
|
||||
if len(e) > 50:
|
||||
univ[t, e[vol63[t, e] >= np.quantile(vol63[t, e], VOL_PCT)]] = True
|
||||
|
||||
def held(w, K=5):
|
||||
a = 2.0 / 6
|
||||
for t in range(1, T):
|
||||
w[t] = a * w[t] + (1 - a) * w[t - 1]
|
||||
wh = w.copy(); last = 0
|
||||
for t in range(T):
|
||||
if t % K == 0:
|
||||
last = t
|
||||
wh[t] = w[last]
|
||||
return wh
|
||||
|
||||
def leg(which, q=0.10):
|
||||
"""EW long basket of a decile: 'low'=lowest-vol, 'high'=highest-vol, 'all'=whole cohort."""
|
||||
w = np.zeros((T, N))
|
||||
for t in range(T):
|
||||
e = np.where(univ[t])[0]
|
||||
if len(e) < 20:
|
||||
continue
|
||||
if which == "all":
|
||||
w[t, e] = 1.0 / len(e)
|
||||
else:
|
||||
order = e[np.argsort(vol63[t, e])] # ascending vol
|
||||
k = max(int(q * len(e)), 5)
|
||||
sel = order[:k] if which == "low" else order[-k:]
|
||||
w[t, sel] = 1.0 / k
|
||||
w = held(w)
|
||||
pnl = np.sum(w[:-1] * R[1:], axis=1) - np.sum(np.abs(w[1:] - w[:-1]), axis=1) * COST / 1e4
|
||||
turn = float(np.mean(np.sum(np.abs(w[1:] - w[:-1]), axis=1)))
|
||||
return pnl, turn
|
||||
|
||||
lo, lo_turn = leg("low")
|
||||
hi, hi_turn = leg("high")
|
||||
mkt, _ = leg("all")
|
||||
ls = lo - hi # dollar-neutral L/S (gross of the extra cost already in legs)
|
||||
sr = sharpe_t
|
||||
|
||||
print(f"\n===== EQUITY LOW-VOL +0.74 DECOMPOSITION (cohort liquid>${DV_FLOOR/1e6:.0f}M & vol>{int(VOL_PCT*100)}pct) =====")
|
||||
print(f"cohort EW (beta/market) Sharpe {sr(T_(mkt)):+.2f}")
|
||||
print(f"LONG leg (low-vol decile) Sharpe {sr(T_(lo)):+.2f} alpha-vs-cohort {sr(T_(lo-mkt)):+.2f} turn/period {lo_turn:.2f}")
|
||||
print(f"HIGH-vol decile Sharpe {sr(T_(hi)):+.2f} (short it -> +{sr(T_(mkt-hi)):+.2f} short-alpha-vs-cohort)")
|
||||
print(f"L/S (low - high) Sharpe {sr(T_(ls)):+.2f}")
|
||||
|
||||
# beta decomposition of L/S vs cohort market
|
||||
a, b = np.asarray(mkt), np.asarray(ls)
|
||||
m = np.isfinite(a) & np.isfinite(b); x, yv = a[m], b[m]
|
||||
beta = float(np.cov(x, yv)[0, 1] / (np.var(x) + 1e-12))
|
||||
resid = yv - beta * x
|
||||
print(f"\nL/S beta to cohort-market = {beta:+.2f} (negative = structural short-beta tilt)")
|
||||
print(f"L/S beta-NEUTRAL residual alpha Sharpe = {sr(T_(resid)):+.2f} "
|
||||
f"(this is the TRUE cross-sectional alpha; if ~0 it was just short-beta)")
|
||||
|
||||
# long-leg deployable check (no borrow)
|
||||
vlo = validate(lo - mkt, days, 30) # long-leg market-neutralized (long decile vs cohort)
|
||||
print(f"\nLONG-leg alpha (deployable, no borrow): full {vlo['full']:+.2f} OOS {vlo['oos']:+.2f} CPCVmed {vlo['med']:+.2f} DSR {vlo['dsr']:.2f}")
|
||||
|
||||
# per-year + concentration
|
||||
print("per-year L/S: " + " ".join(f"{y}:{sr(T_(ls[year[1:]==y])):+.2f}" for y in range(2023,2027) if (year[1:]==y).sum()>40))
|
||||
p = np.nan_to_num(ls); tot = p.sum(); top = np.sort(p)[-10:].sum()
|
||||
print(f"P&L concentration: top-10 days = {100*top/ (tot+1e-12):.0f}% of total (high => episodic/fragile)")
|
||||
print(f"turnover: low-leg {lo_turn:.2f}/period, high-leg {hi_turn:.2f}/period (low => slow signal, cost charge fair)")
|
||||
print("\nREAD: if beta-neutral residual ~0 => short-beta bet (fragile); if long-leg alpha DSR>0.5 => deployable long-only edge.")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
129
scripts/surfer/equity_lowvol_gauntlet.py
Normal file
129
scripts/surfer/equity_lowvol_gauntlet.py
Normal file
@@ -0,0 +1,129 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Full gauntlet on the equity low-vol/lottery-aversion lead (liquid high-vol cohort).
|
||||
|
||||
The make-or-break for the first non-crypto edge: (1) name-bootstrap (robust to WHICH names? —
|
||||
the test that killed carry), (2) cohort-param robustness (did $10M/60pct manufacture it?),
|
||||
(3) realistic short-borrow + long-only (is the edge trapped in the un-cheap short?),
|
||||
(4) honest deflation + per-year, (5) correlation to the crypto momentum book (diversifier?).
|
||||
Free — DBEQ on disk.
|
||||
"""
|
||||
import math
|
||||
import os
|
||||
import sys
|
||||
|
||||
import numpy as np
|
||||
|
||||
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
|
||||
from equity_factor_gate import load, roll # noqa: E402
|
||||
from signal_sweep import xs_weights, validate, sharpe_t # noqa: E402
|
||||
import pit_sweep # noqa: E402
|
||||
from surfer_poc import compute_weights, CFG # noqa: E402
|
||||
import torch # noqa: E402
|
||||
|
||||
DEV = "cuda" if torch.cuda.is_available() else "cpu"
|
||||
|
||||
|
||||
def main():
|
||||
insts, days, close, dvol = load()
|
||||
T, N = close.shape
|
||||
lc = np.log(close)
|
||||
R = np.zeros((T, N)); R[1:] = lc[1:] - lc[:-1]; R = np.where(np.isfinite(R), R, 0.0)
|
||||
dv30 = roll(np.mean, np.nan_to_num(dvol), 30)
|
||||
vol63 = roll(np.std, R, 63)
|
||||
year = (1970 + days / 365.25).astype(int)
|
||||
T_ = lambda x: torch.tensor(x[np.isfinite(x)], device=DEV, dtype=torch.float64)
|
||||
|
||||
def cohort(dv_floor, vol_pct):
|
||||
u = np.zeros((T, N), bool)
|
||||
for t in range(T):
|
||||
liq = (dv30[t] > dv_floor) & np.isfinite(close[t]) & np.isfinite(vol63[t])
|
||||
e = np.where(liq)[0]
|
||||
if len(e) > 50:
|
||||
u[t, e[vol63[t, e] >= np.quantile(vol63[t, e], vol_pct)]] = True
|
||||
return u
|
||||
|
||||
def smooth_weekly(w, K=5):
|
||||
a = 2.0 / 6
|
||||
for t in range(1, T):
|
||||
w[t] = a * w[t] + (1 - a) * w[t - 1]
|
||||
wh = w.copy(); last = 0
|
||||
for t in range(T):
|
||||
if t % K == 0:
|
||||
last = t
|
||||
wh[t] = w[last]
|
||||
return wh
|
||||
|
||||
def lowvol_pnl(univ, cols=None, long_only=False, borrow_ann=0.0):
|
||||
sig = (-vol63).copy(); sig[~univ] = np.nan
|
||||
if cols is not None:
|
||||
mask = np.zeros(N, bool); mask[cols] = True; sig[:, ~mask] = np.nan
|
||||
if long_only:
|
||||
w = np.zeros((T, N))
|
||||
for t in range(T):
|
||||
e = np.where(np.isfinite(sig[t]))[0]
|
||||
if len(e) > 20:
|
||||
k = max(int(0.10 * len(e)), 5); w[t, e[np.argsort(-sig[t, e])[:k]]] = 1.0 / k
|
||||
else:
|
||||
w = xs_weights(sig)
|
||||
w = smooth_weekly(w)
|
||||
g = np.sum(w[:-1] * R[1:], axis=1)
|
||||
g -= np.sum(np.abs(w[1:] - w[:-1]), axis=1) * 10.0 / 1e4 # 10bp trade cost
|
||||
if borrow_ann > 0: # borrow on short notional
|
||||
short_notional = np.sum(np.clip(-w[:-1], 0, None), axis=1)
|
||||
g -= short_notional * borrow_ann / 252.0
|
||||
return g
|
||||
|
||||
base = cohort(1e7, 0.60)
|
||||
base_pnl = lowvol_pnl(base)
|
||||
v = validate(base_pnl, days, 30)
|
||||
print(f"\n===== EQUITY LOW-VOL GAUNTLET (liquid high-vol cohort, deflate N=30) =====")
|
||||
print(f"BASE L/S: full {v['full']:+.2f} OOS {v['oos']:+.2f} CPCVmed {v['med']:+.2f} DSR {v['dsr']:.2f}")
|
||||
|
||||
# (1) name-bootstrap
|
||||
cohort_cols = np.where(base.any(0))[0]
|
||||
rng = np.random.default_rng(5); sh = []
|
||||
for _ in range(200):
|
||||
c = rng.choice(cohort_cols, size=max(len(cohort_cols) // 2, 20), replace=False)
|
||||
s = sharpe_t(T_(lowvol_pnl(base, cols=c)))
|
||||
if not math.isnan(s):
|
||||
sh.append(s)
|
||||
sh = np.array(sh)
|
||||
print(f"(1) name-bootstrap(200): frac>0 {np.mean(sh>0):.2f} median {np.median(sh):+.2f} 5th {np.percentile(sh,5):+.2f}")
|
||||
|
||||
# (2) cohort-param robustness
|
||||
print("(2) cohort-param grid (L/S full Sharpe):")
|
||||
for fl in [5e6, 1e7, 2e7, 5e7]:
|
||||
row = []
|
||||
for vp in [0.50, 0.60, 0.70]:
|
||||
row.append(sharpe_t(T_(lowvol_pnl(cohort(fl, vp)))))
|
||||
print(f" ${fl/1e6:>3.0f}M floor: " + " ".join(f"vp{int(vp*100)}:{r:+.2f}" for vp, r in zip([0.50,0.60,0.70], row)))
|
||||
|
||||
# (3) short-borrow + long-only
|
||||
print("(3) short-borrow & long-only (net):")
|
||||
for ba in [0.0, 0.10, 0.30]:
|
||||
print(f" L/S borrow {int(ba*100)}%/yr: {sharpe_t(T_(lowvol_pnl(base, borrow_ann=ba))):+.2f}")
|
||||
lo = lowvol_pnl(base, long_only=True)
|
||||
vlo = validate(lo, days, 30)
|
||||
print(f" LONG-ONLY (no borrow): full {vlo['full']:+.2f} OOS {vlo['oos']:+.2f} CPCVmed {vlo['med']:+.2f} DSR {vlo['dsr']:.2f}")
|
||||
|
||||
# (4) per-year (base L/S)
|
||||
print("(4) per-year (base L/S): " + " ".join(f"{y}:{sharpe_t(T_(base_pnl[year[1:]==y])):+.2f}" for y in range(2023,2027) if (year[1:]==y).sum()>40))
|
||||
|
||||
# (5) correlation to crypto momentum book
|
||||
syms, cdays, cc, cqv, cf = pit_sweep.load()
|
||||
cw, _ = compute_weights(cc, cqv, cdays, CFG)
|
||||
cR = np.zeros_like(cc); cR[1:] = np.log(cc)[1:] - np.log(cc)[:-1]; cR = np.where(np.isfinite(cR), cR, 0.0)
|
||||
cff = np.where(np.isfinite(cf), cf, 0.0)
|
||||
cmom = np.sum(cw[:-1] * (cR - cff)[1:], axis=1)
|
||||
cmd = {int(cdays[1:][i]): cmom[i] for i in range(len(cmom))}
|
||||
emd = {int(days[1:][i]): base_pnl[i] for i in range(len(base_pnl))}
|
||||
common = sorted(set(cmd) & set(emd))
|
||||
a = np.array([cmd[d] for d in common]); b = np.array([emd[d] for d in common])
|
||||
m = np.isfinite(a) & np.isfinite(b)
|
||||
corr = float(np.corrcoef(a[m], b[m])[0, 1]) if m.sum() > 50 else float("nan")
|
||||
print(f"(5) corr to crypto-momentum book: {corr:+.2f} (low => diversifier in a preferred market)")
|
||||
print("\nVERDICT: bootstrap frac>0~1 + grid broadly + long-only positive + DSR>0.5 + low crypto-corr = real 2nd sleeve.")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
115
scripts/surfer/equity_vrp.py
Normal file
115
scripts/surfer/equity_vrp.py
Normal file
@@ -0,0 +1,115 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Equity-index VRP on XSP (mini-SPX) options, 2013-2026 — the one real-prior frontier.
|
||||
|
||||
For each day: find the ~30d ATM straddle (strike where call~=put = the forward), back out
|
||||
implied vol from the straddle price, compare to forward realized vol. VRP = IV - RV; a
|
||||
short-vol seller harvests it. 13y spans 2018/2020/2022 tail events. Tests raw VRP Sharpe,
|
||||
per-year (incl. crashes), tail/skew, AND the tail-managed version (don't sell when IV rising
|
||||
— the gate that fixed crypto VRP). XSP = cash-settled European -> no early-exercise distortion.
|
||||
"""
|
||||
import glob
|
||||
import math
|
||||
import os
|
||||
import sys
|
||||
|
||||
import numpy as np
|
||||
|
||||
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
|
||||
from signal_sweep import sharpe_t # noqa: E402
|
||||
import torch # noqa: E402
|
||||
|
||||
DEV = "cuda" if torch.cuda.is_available() else "cpu"
|
||||
DAY_NS = 86_400 * 10**9
|
||||
|
||||
|
||||
def load_xsp():
|
||||
import databento as db
|
||||
rows = []
|
||||
for p in sorted(glob.glob("data/surfer/xsp/*.dbn")):
|
||||
try:
|
||||
df = db.DBNStore.from_file(p).to_df().reset_index()
|
||||
except Exception:
|
||||
continue
|
||||
if df.empty or "symbol" not in df.columns:
|
||||
continue
|
||||
df = df[df["close"] > 0]
|
||||
rows.append(df[["ts_event", "symbol", "close"]])
|
||||
import pandas as pd
|
||||
d = pd.concat(rows, ignore_index=True)
|
||||
s = d["symbol"].astype(str).str.replace(" ", "", regex=False) # "XSP240119C00450000"
|
||||
d["right"] = s.str[-9]
|
||||
d["strike"] = s.str[-8:].astype(float) / 1000.0
|
||||
d["expiry"] = s.str[-15:-9] # YYMMDD
|
||||
d["day"] = (d["ts_event"].astype("int64") // DAY_NS)
|
||||
return d
|
||||
|
||||
|
||||
def to_epoch_day(yymmdd):
|
||||
import datetime
|
||||
y = 2000 + int(yymmdd[:2]); mo = int(yymmdd[2:4]); da = int(yymmdd[4:6])
|
||||
return (datetime.date(y, mo, da) - datetime.date(1970, 1, 1)).days
|
||||
|
||||
|
||||
def main():
|
||||
d = load_xsp()
|
||||
print(f"loaded {len(d)} XSP option-days, {d['day'].nunique()} trading days")
|
||||
d["exp_day"] = d["expiry"].map(to_epoch_day)
|
||||
d["ttm"] = d["exp_day"] - d["day"]
|
||||
d = d[(d["ttm"] >= 20) & (d["ttm"] <= 45)] # ~30d window
|
||||
|
||||
iv_by_day = {}
|
||||
S_by_day = {}
|
||||
for day, g in d.groupby("day"):
|
||||
# pick the expiry closest to 30d
|
||||
exp = g.iloc[(g["ttm"] - 30).abs().argsort()].iloc[0]["exp_day"]
|
||||
ge = g[g["exp_day"] == exp]
|
||||
calls = ge[ge["right"] == "C"].groupby("strike")["close"].mean() # dedup AM/PM series
|
||||
puts = ge[ge["right"] == "P"].groupby("strike")["close"].mean()
|
||||
common = calls.index.intersection(puts.index)
|
||||
if len(common) < 3:
|
||||
continue
|
||||
diff = (calls[common] - puts[common]).abs()
|
||||
katm = diff.idxmin() # ATM: where C~=P
|
||||
S = float(katm + calls[katm] - puts[katm]) # parity forward
|
||||
straddle = float(calls[katm] + puts[katm])
|
||||
T = float(ge["ttm"].iloc[0]) / 365.0
|
||||
if S <= 0 or T <= 0:
|
||||
continue
|
||||
iv = straddle / (0.8 * S * math.sqrt(T)) # ATM straddle -> implied vol
|
||||
iv_by_day[int(day)] = iv; S_by_day[int(day)] = S
|
||||
|
||||
days = np.array(sorted(S_by_day))
|
||||
S = np.array([S_by_day[x] for x in days]); IV = np.array([iv_by_day[x] for x in days])
|
||||
r = np.zeros(len(days)); r[1:] = np.log(S[1:] / S[:-1])
|
||||
# forward 21-day realized vol (annualized) — what the seller faces
|
||||
H = 21
|
||||
rv_fwd = np.full(len(days), np.nan)
|
||||
for t in range(len(days) - H):
|
||||
rv_fwd[t] = np.std(r[t + 1:t + 1 + H]) * math.sqrt(252)
|
||||
vrp = IV - rv_fwd # premium (positive = seller wins)
|
||||
# short-vol daily P&L proxy: collect implied variance, pay realized squared return
|
||||
iv_lag = np.concatenate([[np.nan], IV[:-1]])
|
||||
svol = (iv_lag ** 2) / 252.0 - r ** 2
|
||||
year = np.array([1970 + x / 365.25 for x in days]).astype(int)
|
||||
T_ = lambda x: torch.tensor(np.asarray(x)[np.isfinite(np.asarray(x))], device=DEV, dtype=torch.float64)
|
||||
|
||||
print(f"\n===== EQUITY-INDEX VRP (XSP, {days.min()}..{days.max()}, {len(days)} days) =====")
|
||||
print(f"mean implied vol {np.nanmean(IV):.1%} mean fwd-realized {np.nanmean(rv_fwd):.1%} "
|
||||
f"mean VRP {np.nanmean(vrp):.1%} (positive => premium exists)")
|
||||
print(f"VRP frac>0 (IV>RV): {np.nanmean(vrp[np.isfinite(vrp)]>0):.2f}")
|
||||
print(f"\nshort-vol daily P&L: Sharpe {sharpe_t(T_(svol)):+.2f} skew {float(((svol[np.isfinite(svol)]-np.nanmean(svol))**3).mean()/np.nanstd(svol)**3):+.2f} worst-day {np.nanmin(svol)/np.nanstd(svol):+.1f}σ")
|
||||
print("per-year short-vol: " + " ".join(f"{y}:{sharpe_t(T_(svol[year==y])):+.1f}" for y in range(2013, 2027) if (year == y).sum() > 60))
|
||||
|
||||
# tail-managed: don't sell when implied vol rising over 5d (the crypto-VRP gate)
|
||||
rising = np.zeros(len(days))
|
||||
for t in range(6, len(days)):
|
||||
rising[t] = 0.0 if IV[t - 1] > IV[t - 6] else 1.0
|
||||
svol_tm = rising * svol
|
||||
print(f"\ntail-managed (don't sell into rising IV): Sharpe {sharpe_t(T_(svol_tm)):+.2f} "
|
||||
f"skew {float(((svol_tm[np.isfinite(svol_tm)]-np.nanmean(svol_tm))**3).mean()/np.nanstd(svol_tm)**3):+.2f} worst {np.nanmin(svol_tm)/np.nanstd(svol_tm):+.1f}σ")
|
||||
print("per-year tail-managed: " + " ".join(f"{y}:{sharpe_t(T_(svol_tm[year==y])):+.1f}" for y in range(2013, 2027) if (year == y).sum() > 60))
|
||||
print("\nVERDICT: VRP frac>0 high + short-vol Sharpe>0 + tail-managed improves skew/recent = real, harvestable equity VRP.")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
87
scripts/surfer/equity_vrp2.py
Normal file
87
scripts/surfer/equity_vrp2.py
Normal file
@@ -0,0 +1,87 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Equity-index VRP, CORRECTED: clean SPY underlying (not noisy parity) + XSP option-implied IV.
|
||||
|
||||
Fixes the bug where parity-derived underlying inflated realized vol (spurious -VRP). SPY ~= XSP
|
||||
(~SPX/10), so SPY is the clean underlying for ATM reference + realized vol. Overlap 2023-2026.
|
||||
VRP = implied (XSP ATM straddle) - realized (clean SPY). Short-vol Sharpe, per-year, tail, tail-managed.
|
||||
"""
|
||||
import math
|
||||
import os
|
||||
import sys
|
||||
|
||||
import numpy as np
|
||||
|
||||
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
|
||||
from equity_vrp import load_xsp, to_epoch_day # noqa: E402
|
||||
from signal_sweep import sharpe_t # noqa: E402
|
||||
import torch # noqa: E402
|
||||
|
||||
DEV = "cuda" if torch.cuda.is_available() else "cpu"
|
||||
DAY_NS = 86_400 * 10**9
|
||||
|
||||
|
||||
def load_spy():
|
||||
import databento as db
|
||||
a = db.DBNStore.from_file("data/surfer/spy_ohlcv1d.dbn").to_ndarray()
|
||||
day = a["ts_event"].astype(np.int64) // DAY_NS
|
||||
close = a["close"].astype(np.float64) / 1e9
|
||||
return {int(day[i]): float(close[i]) for i in range(len(day)) if close[i] > 0}
|
||||
|
||||
|
||||
def main():
|
||||
spy = load_spy()
|
||||
print(f"SPY clean: {len(spy)} days ({min(spy)}..{max(spy)}), level {spy[min(spy)]:.0f}->{spy[max(spy)]:.0f}")
|
||||
d = load_xsp()
|
||||
d["exp_day"] = d["expiry"].map(to_epoch_day)
|
||||
d["ttm"] = d["exp_day"] - d["day"]
|
||||
d = d[(d["ttm"] >= 20) & (d["ttm"] <= 45)]
|
||||
d = d[d["day"].isin(spy.keys())] # overlap with clean SPY
|
||||
|
||||
iv = {}
|
||||
for day, g in d.groupby("day"):
|
||||
S = spy[int(day)]
|
||||
exp = g.iloc[(g["ttm"] - 30).abs().argsort()].iloc[0]["exp_day"]
|
||||
ge = g[g["exp_day"] == exp]
|
||||
calls = ge[ge["right"] == "C"].groupby("strike")["close"].mean()
|
||||
puts = ge[ge["right"] == "P"].groupby("strike")["close"].mean()
|
||||
common = calls.index.intersection(puts.index)
|
||||
if len(common) < 3:
|
||||
continue
|
||||
katm = common[np.abs(np.array(common) - S).argmin()] # ATM = strike nearest clean SPY level
|
||||
straddle = float(calls[katm] + puts[katm])
|
||||
T = float(ge["ttm"].iloc[0]) / 365.0
|
||||
if straddle <= 0 or T <= 0:
|
||||
continue
|
||||
iv[int(day)] = straddle / (0.8 * S * math.sqrt(T))
|
||||
|
||||
days = np.array(sorted(set(iv) & set(spy)))
|
||||
S = np.array([spy[x] for x in days]); IV = np.array([iv[x] for x in days])
|
||||
r = np.zeros(len(days)); r[1:] = np.log(S[1:] / S[:-1])
|
||||
H = 21
|
||||
rv_fwd = np.full(len(days), np.nan)
|
||||
for t in range(len(days) - H):
|
||||
rv_fwd[t] = np.std(r[t + 1:t + 1 + H]) * math.sqrt(252)
|
||||
vrp = IV - rv_fwd
|
||||
iv_lag = np.concatenate([[np.nan], IV[:-1]])
|
||||
svol = (iv_lag ** 2) / 252.0 - r ** 2
|
||||
year = np.array([1970 + x / 365.25 for x in days]).astype(int)
|
||||
T_ = lambda x: torch.tensor(np.asarray(x)[np.isfinite(np.asarray(x))], device=DEV, dtype=torch.float64)
|
||||
fin = np.isfinite(vrp)
|
||||
|
||||
print(f"\n===== EQUITY-INDEX VRP (CORRECTED, clean SPY) — {len(days)} days =====")
|
||||
print(f"mean implied {np.nanmean(IV):.1%} mean fwd-realized(clean) {np.nanmean(rv_fwd):.1%} "
|
||||
f"mean VRP {np.nanmean(vrp[fin]):+.1%} frac>0 {np.mean(vrp[fin]>0):.2f}")
|
||||
sk = float(((svol[np.isfinite(svol)] - np.nanmean(svol)) ** 3).mean() / np.nanstd(svol) ** 3)
|
||||
print(f"short-vol daily P&L: Sharpe {sharpe_t(T_(svol)):+.2f} skew {sk:+.2f} worst {np.nanmin(svol)/np.nanstd(svol):+.1f}σ")
|
||||
print("per-year: " + " ".join(f"{y}:{sharpe_t(T_(svol[year==y])):+.1f}" for y in range(2023, 2027) if (year == y).sum() > 40))
|
||||
rising = np.zeros(len(days))
|
||||
for t in range(6, len(days)):
|
||||
rising[t] = 0.0 if IV[t - 1] > IV[t - 6] else 1.0
|
||||
tm = rising * svol
|
||||
print(f"tail-managed (don't sell into rising IV): Sharpe {sharpe_t(T_(tm)):+.2f} worst {np.nanmin(tm)/np.nanstd(tm):+.1f}σ")
|
||||
print("per-year TM: " + " ".join(f"{y}:{sharpe_t(T_(tm[year==y])):+.1f}" for y in range(2023, 2027) if (year == y).sum() > 40))
|
||||
print("\nVERDICT: mean VRP>0 + frac>0~0.8 confirms premium exists; short-vol Sharpe (esp tail-managed) = harvestable.")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
78
scripts/surfer/equity_vrp3.py
Normal file
78
scripts/surfer/equity_vrp3.py
Normal file
@@ -0,0 +1,78 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Equity-index VRP, robust internal-parity version (self-consistent, no external underlying).
|
||||
|
||||
Per day: implied forward F = MEDIAN over the 5 near-ATM strikes of (K + C - P) [put-call parity,
|
||||
median-smoothed to kill per-strike noise]. Use F consistently for ATM, implied vol, AND the
|
||||
realized-vol series. Internally consistent -> avoids both the single-strike-parity noise and the
|
||||
SPY!=XSP divergence that produced spurious negative VRP. XSP 2013-2026.
|
||||
"""
|
||||
import math
|
||||
import os
|
||||
import sys
|
||||
|
||||
import numpy as np
|
||||
|
||||
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
|
||||
from equity_vrp import load_xsp, to_epoch_day # noqa: E402
|
||||
from signal_sweep import sharpe_t # noqa: E402
|
||||
import torch # noqa: E402
|
||||
|
||||
DEV = "cuda" if torch.cuda.is_available() else "cpu"
|
||||
|
||||
|
||||
def main():
|
||||
d = load_xsp()
|
||||
d["exp_day"] = d["expiry"].map(to_epoch_day)
|
||||
d["ttm"] = d["exp_day"] - d["day"]
|
||||
d = d[(d["ttm"] >= 20) & (d["ttm"] <= 45)]
|
||||
F, IVm = {}, {}
|
||||
for day, g in d.groupby("day"):
|
||||
exp = g.iloc[(g["ttm"] - 30).abs().argsort()].iloc[0]["exp_day"]
|
||||
ge = g[g["exp_day"] == exp]
|
||||
calls = ge[ge["right"] == "C"].groupby("strike")["close"].mean()
|
||||
puts = ge[ge["right"] == "P"].groupby("strike")["close"].mean()
|
||||
common = np.array(sorted(set(calls.index) & set(puts.index)))
|
||||
if len(common) < 5:
|
||||
continue
|
||||
rough = common[np.abs(calls[common].values - puts[common].values).argmin()] # min|C-P| ~ ATM
|
||||
near = common[np.argsort(np.abs(common - rough))[:5]] # 5 nearest strikes
|
||||
f = float(np.median([k + float(calls[k]) - float(puts[k]) for k in near])) # parity forward (median)
|
||||
katm = common[np.abs(common - f).argmin()]
|
||||
straddle = float(calls[katm] + puts[katm])
|
||||
T = float(ge["ttm"].iloc[0]) / 365.0
|
||||
if f <= 0 or straddle <= 0 or T <= 0:
|
||||
continue
|
||||
F[int(day)] = f; IVm[int(day)] = straddle / (0.8 * f * math.sqrt(T))
|
||||
|
||||
days = np.array(sorted(F))
|
||||
S = np.array([F[x] for x in days]); IV = np.array([IVm[x] for x in days])
|
||||
r = np.zeros(len(days)); r[1:] = np.log(S[1:] / S[:-1])
|
||||
r = np.clip(r, -0.25, 0.25) # guard residual parity glitches
|
||||
H = 21
|
||||
rv = np.full(len(days), np.nan)
|
||||
for t in range(len(days) - H):
|
||||
rv[t] = np.std(r[t + 1:t + 1 + H]) * math.sqrt(252)
|
||||
vrp = IV - rv; fin = np.isfinite(vrp)
|
||||
iv_lag = np.concatenate([[np.nan], IV[:-1]])
|
||||
svol = (iv_lag ** 2) / 252.0 - r ** 2
|
||||
year = np.array([1970 + x / 365.25 for x in days]).astype(int)
|
||||
T_ = lambda x: torch.tensor(np.asarray(x)[np.isfinite(np.asarray(x))], device=DEV, dtype=torch.float64)
|
||||
|
||||
print(f"\n===== EQUITY-INDEX VRP (robust internal parity) — {len(days)} days ({days.min()}..{days.max()}) =====")
|
||||
print(f"implied-forward level {S[0]:.0f}->{S[-1]:.0f} (should track SPX/10 ~ 400->650)")
|
||||
print(f"mean implied {np.nanmean(IV):.1%} mean fwd-realized {np.nanmean(rv):.1%} "
|
||||
f"mean VRP {np.nanmean(vrp[fin]):+.1%} frac>0 {np.mean(vrp[fin]>0):.2f}")
|
||||
sk = float(((svol[np.isfinite(svol)] - np.nanmean(svol)) ** 3).mean() / np.nanstd(svol) ** 3)
|
||||
print(f"short-vol daily P&L: Sharpe {sharpe_t(T_(svol)):+.2f} skew {sk:+.2f} worst {np.nanmin(svol)/np.nanstd(svol):+.1f}σ")
|
||||
print("per-year: " + " ".join(f"{y}:{sharpe_t(T_(svol[year==y])):+.1f}" for y in range(2013, 2027) if (year == y).sum() > 60))
|
||||
rising = np.zeros(len(days))
|
||||
for t in range(6, len(days)):
|
||||
rising[t] = 0.0 if IV[t - 1] > IV[t - 6] else 1.0
|
||||
tm = rising * svol
|
||||
print(f"tail-managed: Sharpe {sharpe_t(T_(tm)):+.2f} worst {np.nanmin(tm)/np.nanstd(tm):+.1f}σ")
|
||||
print("per-year TM: " + " ".join(f"{y}:{sharpe_t(T_(tm[year==y])):+.1f}" for y in range(2013, 2027) if (year == y).sum() > 60))
|
||||
print("\nVERDICT: implied-forward tracking ~SPX/10 + mean RV ~13-16% + VRP>0 => measurement clean & premium real.")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
83
scripts/surfer/fetch_caiso.py
Normal file
83
scripts/surfer/fetch_caiso.py
Normal file
@@ -0,0 +1,83 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Fetch CAISO day-ahead (DAM) + real-time (RTM) hourly LMP for one hub (free OASIS, no key).
|
||||
|
||||
The RT-vs-DA spread is the forecast-error-driven, less-seasonal intraday opportunity. Monthly
|
||||
chunks with backoff (OASIS rate-limits hard). Aggregates RT to hourly. Caches per chunk.
|
||||
Saves data/surfer/caiso/{dam,rtm}.json as {hour_key: price}.
|
||||
"""
|
||||
import io
|
||||
import json
|
||||
import os
|
||||
import time
|
||||
import urllib.request
|
||||
import zipfile
|
||||
|
||||
NODE = "TH_NP15_GEN-APND"
|
||||
OUT = "data/surfer/caiso"
|
||||
MONTHS = [(2024, m) for m in range(1, 13)]
|
||||
|
||||
|
||||
def oasis(qn, market, y, m):
|
||||
s = f"{y}{m:02d}01T08:00-0000"
|
||||
ny, nm = (y + 1, 1) if m == 12 else (y, m + 1)
|
||||
e = f"{ny}{nm:02d}01T08:00-0000"
|
||||
u = (f"http://oasis.caiso.com/oasisapi/SingleZip?queryname={qn}&startdatetime={s}"
|
||||
f"&enddatetime={e}&version=1&market_run_id={market}&node={NODE}&resultformat=6")
|
||||
for a in range(5):
|
||||
try:
|
||||
raw = urllib.request.urlopen(urllib.request.Request(u, headers={"User-Agent": "curl/8"}), timeout=90).read()
|
||||
z = zipfile.ZipFile(io.BytesIO(raw))
|
||||
txt = z.read(z.namelist()[0]).decode()
|
||||
return txt
|
||||
except Exception as ex:
|
||||
time.sleep(12 * (a + 1))
|
||||
return None
|
||||
|
||||
|
||||
def parse_hourly(txt):
|
||||
lines = txt.strip().split("\n")
|
||||
hdr = lines[0].split(",")
|
||||
if "LMP_TYPE" not in hdr:
|
||||
return {} # error report / wrong format
|
||||
iT = hdr.index("LMP_TYPE")
|
||||
iV = hdr.index("MW") if "MW" in hdr else (hdr.index("PRC") if "PRC" in hdr else -1)
|
||||
if iV < 0:
|
||||
return {}
|
||||
iD = hdr.index("OPR_DT"); iH = hdr.index("OPR_HR")
|
||||
agg = {}
|
||||
for ln in lines[1:]:
|
||||
c = ln.split(",")
|
||||
if len(c) <= max(iT, iV, iD, iH):
|
||||
continue
|
||||
if c[iT] != "LMP":
|
||||
continue
|
||||
k = f"{c[iD]}H{int(c[iH]):02d}"
|
||||
agg.setdefault(k, []).append(float(c[iV]))
|
||||
return {k: sum(v) / len(v) for k, v in agg.items()}
|
||||
|
||||
|
||||
def fetch(market, qn):
|
||||
os.makedirs(OUT, exist_ok=True)
|
||||
out = {}
|
||||
for y, m in MONTHS:
|
||||
cf = f"{OUT}/{market}_{y}{m:02d}.json"
|
||||
if os.path.exists(cf):
|
||||
out.update(json.load(open(cf))); continue
|
||||
txt = oasis(qn, market, y, m)
|
||||
if not txt:
|
||||
print(f" {market} {y}-{m:02d}: FAILED"); continue
|
||||
h = parse_hourly(txt)
|
||||
json.dump(h, open(cf, "w")); out.update(h)
|
||||
print(f" {market} {y}-{m:02d}: {len(h)} hrs")
|
||||
time.sleep(6)
|
||||
return out
|
||||
|
||||
|
||||
def main():
|
||||
dam = fetch("DAM", "PRC_LMP"); rtm = fetch("RTPD", "PRC_RTPD_LMP") # 15-min RT pre-dispatch
|
||||
json.dump(dam, open(f"{OUT}/dam.json", "w")); json.dump(rtm, open(f"{OUT}/rtm.json", "w"))
|
||||
print(f"DONE: DAM {len(dam)} hrs, RTM {len(rtm)} hrs")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
103
scripts/surfer/fetch_crypto.py
Normal file
103
scripts/surfer/fetch_crypto.py
Normal file
@@ -0,0 +1,103 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Fetch daily klines + funding history for major USDT perps (Binance, free, no key).
|
||||
|
||||
Caches per-symbol npz to data/surfer/crypto/ (gitignored): day, open, close, funding_daily
|
||||
(sum of the 8h funding rates that day). Curated long-history majors → reduces (not eliminates)
|
||||
survivorship; v1 caveat documented. Crypto is 24/7 → no roll / no overnight gap.
|
||||
"""
|
||||
import json
|
||||
import os
|
||||
import time
|
||||
import urllib.request
|
||||
|
||||
import numpy as np
|
||||
|
||||
OUT = "data/surfer/crypto"
|
||||
DAY_MS = 86_400_000
|
||||
TOP_N = 80 # programmatic universe: top-N USDT perps by 24h quote-volume (removes hand-selection bias)
|
||||
|
||||
|
||||
def get(url):
|
||||
req = urllib.request.Request(url, headers={"User-Agent": "curl/8"})
|
||||
return json.load(urllib.request.urlopen(req, timeout=30))
|
||||
|
||||
|
||||
def universe(n):
|
||||
info = get("https://fapi.binance.com/fapi/v1/exchangeInfo")
|
||||
perps = {s["symbol"] for s in info["symbols"]
|
||||
if s.get("contractType") == "PERPETUAL" and s.get("quoteAsset") == "USDT"
|
||||
and s.get("status") == "TRADING"}
|
||||
tick = get("https://fapi.binance.com/fapi/v1/ticker/24hr")
|
||||
vol = {t["symbol"]: float(t["quoteVolume"]) for t in tick if t["symbol"] in perps}
|
||||
return sorted(vol, key=lambda s: -vol[s])[:n]
|
||||
|
||||
|
||||
SYMS = universe(TOP_N)
|
||||
|
||||
|
||||
def klines(sym):
|
||||
out = []
|
||||
end = None
|
||||
for _ in range(20):
|
||||
u = f"https://fapi.binance.com/fapi/v1/klines?symbol={sym}&interval=1d&limit=1500"
|
||||
if end:
|
||||
u += f"&endTime={end}"
|
||||
k = get(u)
|
||||
if not k:
|
||||
break
|
||||
out = k + out
|
||||
end = k[0][0] - 1
|
||||
if len(k) < 1500:
|
||||
break
|
||||
time.sleep(0.15)
|
||||
# dedup by openTime
|
||||
d = {int(r[0]): (float(r[1]), float(r[4])) for r in out}
|
||||
days = np.array(sorted(d))
|
||||
op = np.array([d[t][0] for t in days]); cl = np.array([d[t][1] for t in days])
|
||||
return days // DAY_MS, op, cl
|
||||
|
||||
|
||||
def funding(sym, start_ms):
|
||||
out = []
|
||||
st = start_ms
|
||||
for _ in range(80): # forward pagination (startTime works; endTime didn't)
|
||||
u = f"https://fapi.binance.com/fapi/v1/fundingRate?symbol={sym}&startTime={st}&limit=1000"
|
||||
f = get(u)
|
||||
if not f:
|
||||
break
|
||||
out += f
|
||||
st = f[-1]["fundingTime"] + 1
|
||||
if len(f) < 1000:
|
||||
break
|
||||
time.sleep(0.12)
|
||||
daily = {}
|
||||
for r in out:
|
||||
daily.setdefault(int(r["fundingTime"]) // DAY_MS, 0.0)
|
||||
daily[int(r["fundingTime"]) // DAY_MS] += float(r["fundingRate"])
|
||||
return daily
|
||||
|
||||
|
||||
def main():
|
||||
os.makedirs(OUT, exist_ok=True)
|
||||
ok = 0
|
||||
for sym in SYMS:
|
||||
outp = f"{OUT}/{sym}.npz"
|
||||
if os.path.exists(outp):
|
||||
print(f" {sym}: cached"); ok += 1; continue
|
||||
try:
|
||||
kd, op, cl = klines(sym)
|
||||
if len(kd) < 400:
|
||||
print(f" {sym}: too short ({len(kd)}d), skip"); continue
|
||||
fmap = funding(sym, int(kd.min()) * DAY_MS)
|
||||
fund = np.array([fmap.get(int(d), 0.0) for d in kd])
|
||||
np.savez(outp, day=kd, open=op, close=cl, funding=fund)
|
||||
print(f" {sym}: {len(kd)}d ({kd.min()}..{kd.max()}) fundcov={np.mean(fund!=0):.2f}")
|
||||
ok += 1
|
||||
time.sleep(0.2)
|
||||
except Exception as e:
|
||||
print(f" {sym}: FAIL {type(e).__name__} {str(e)[:80]}")
|
||||
print(f"DONE: {ok}/{len(SYMS)} symbols -> {OUT}/")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
103
scripts/surfer/fetch_crypto_pit.py
Normal file
103
scripts/surfer/fetch_crypto_pit.py
Normal file
@@ -0,0 +1,103 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Point-in-time crypto fetch: broad live universe + KNOWN-DEAD perps, WITH volume.
|
||||
|
||||
For a survivorship-free test: delisted Binance perps still return klines up to delisting,
|
||||
so including them lets a date-by-date "top-N by trailing dollar-volume" universe contain
|
||||
dead coins while they traded and drop them when they die. Saves day, close, qvol, funding
|
||||
to data/surfer/crypto_pit/ (gitignored).
|
||||
"""
|
||||
import json
|
||||
import os
|
||||
import time
|
||||
import urllib.request
|
||||
|
||||
import numpy as np
|
||||
|
||||
OUT = "data/surfer/crypto_pit"
|
||||
DAY_MS = 86_400_000
|
||||
TOP_N = 120
|
||||
# famous delisted / dead Binance USDT perps (the survivorship risk); klines available post-delist
|
||||
DEAD = ["LUNAUSDT", "ANCUSDT", "SRMUSDT", "HNTUSDT", "MATICUSDT", "FTTUSDT", "RAYUSDT",
|
||||
"WAVESUSDT", "BNXUSDT", "SCUSDT", "OCEANUSDT", "AGIXUSDT", "CVCUSDT", "TLMUSDT",
|
||||
"DENTUSDT", "KEYUSDT", "CTKUSDT", "TOMOUSDT", "AKROUSDT", "BLZUSDT", "COMBOUSDT",
|
||||
"FTMUSDT", "DGBUSDT", "REEFUSDT", "SLPUSDT", "IDEXUSDT", "LINAUSDT", "NEBLUSDT",
|
||||
"RADUSDT", "BTSUSDT", "STMXUSDT", "MDTUSDT", "AMBUSDT", "GTCUSDT", "DARUSDT"]
|
||||
|
||||
|
||||
def get(url):
|
||||
try:
|
||||
return json.load(urllib.request.urlopen(urllib.request.Request(url, headers={"User-Agent": "curl/8"}), timeout=30))
|
||||
except Exception:
|
||||
return None
|
||||
|
||||
|
||||
def universe(n):
|
||||
info = get("https://fapi.binance.com/fapi/v1/exchangeInfo")
|
||||
perps = {s["symbol"] for s in info["symbols"]
|
||||
if s.get("contractType") == "PERPETUAL" and s.get("quoteAsset") == "USDT" and s.get("status") == "TRADING"}
|
||||
tick = get("https://fapi.binance.com/fapi/v1/ticker/24hr")
|
||||
vol = {t["symbol"]: float(t["quoteVolume"]) for t in tick if t["symbol"] in perps}
|
||||
return sorted(vol, key=lambda s: -vol[s])[:n]
|
||||
|
||||
|
||||
def klines(sym):
|
||||
out, end = [], None
|
||||
for _ in range(20):
|
||||
u = f"https://fapi.binance.com/fapi/v1/klines?symbol={sym}&interval=1d&limit=1500"
|
||||
if end:
|
||||
u += f"&endTime={end}"
|
||||
k = get(u)
|
||||
if not k:
|
||||
break
|
||||
out = k + out
|
||||
end = k[0][0] - 1
|
||||
if len(k) < 1500:
|
||||
break
|
||||
time.sleep(0.15)
|
||||
d = {int(r[0]): (float(r[4]), float(r[7])) for r in out} # close, quote-volume
|
||||
days = np.array(sorted(d))
|
||||
cl = np.array([d[t][0] for t in days]); qv = np.array([d[t][1] for t in days])
|
||||
return days // DAY_MS, cl, qv
|
||||
|
||||
|
||||
def funding(sym, start_ms):
|
||||
out, st = [], start_ms
|
||||
for _ in range(80):
|
||||
f = get(f"https://fapi.binance.com/fapi/v1/fundingRate?symbol={sym}&startTime={st}&limit=1000")
|
||||
if not f:
|
||||
break
|
||||
out += f
|
||||
st = f[-1]["fundingTime"] + 1
|
||||
if len(f) < 1000:
|
||||
break
|
||||
time.sleep(0.12)
|
||||
daily = {}
|
||||
for r in out:
|
||||
daily.setdefault(int(r["fundingTime"]) // DAY_MS, 0.0)
|
||||
daily[int(r["fundingTime"]) // DAY_MS] += float(r["fundingRate"])
|
||||
return daily
|
||||
|
||||
|
||||
def main():
|
||||
os.makedirs(OUT, exist_ok=True)
|
||||
syms = sorted(set(universe(TOP_N)) | set(DEAD))
|
||||
ok = 0
|
||||
for sym in syms:
|
||||
outp = f"{OUT}/{sym}.npz"
|
||||
if os.path.exists(outp):
|
||||
ok += 1; continue
|
||||
kd, cl, qv = klines(sym)
|
||||
if len(kd) < 200:
|
||||
print(f" {sym}: short/none ({len(kd)}d), skip"); continue
|
||||
fmap = funding(sym, int(kd.min()) * DAY_MS)
|
||||
fund = np.array([fmap.get(int(d), 0.0) for d in kd])
|
||||
np.savez(outp, day=kd, close=cl, qvol=qv, funding=fund)
|
||||
alive = "DEAD" if qv[-30:].mean() == 0 else "live"
|
||||
print(f" {sym}: {len(kd)}d ({kd.min()}..{kd.max()}) {alive}")
|
||||
ok += 1
|
||||
time.sleep(0.2)
|
||||
print(f"DONE: {ok}/{len(syms)} -> {OUT}/")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
83
scripts/surfer/fetch_daily.py
Normal file
83
scripts/surfer/fetch_daily.py
Normal file
@@ -0,0 +1,83 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Download daily OHLCV for the surfer universe — continuous front-month, BUDGET-CAPPED.
|
||||
|
||||
Re-confirms get_cost (free, no download) and ABORTS if over the cap before spending.
|
||||
Saves raw DBN to data/surfer/ (gitignored). Reads DATABENTO_API_KEY from env (never printed).
|
||||
"""
|
||||
import datetime
|
||||
import os
|
||||
import sys
|
||||
|
||||
import databento as db
|
||||
|
||||
CAP_USD = 40.00 # hard ceiling on cumulative get_cost; covered by $125 free credits ($0 cash)
|
||||
PER_ROOT_CAP = 10.00 # per-root sanity cap
|
||||
DS = "GLBX.MDP3"
|
||||
SCHEMA = "ohlcv-1d"
|
||||
START = "2010-06-06"
|
||||
END = datetime.date.today().isoformat() # fetch through today (dynamic) — was hardcoded; needed for daily forward-track
|
||||
ROOTS = ["ES", "NQ", "YM", "RTY", "ZN", "ZB", "ZF", "ZT", "6E", "6J", "6B", "6A",
|
||||
"6C", "GC", "SI", "HG", "CL", "NG", "RB", "ZC", "ZS", "ZW"]
|
||||
# parent symbology = all outright expiries per root (fast; continuous chain resolution 504s).
|
||||
# We build the continuous series ourselves (volume-roll + ratio-adjust) from these.
|
||||
STYPE = "parent"
|
||||
OUT_DIR = "data/surfer"
|
||||
|
||||
|
||||
def main():
|
||||
key = os.environ.get("DATABENTO_API_KEY")
|
||||
if not key:
|
||||
print("DATABENTO_API_KEY not set"); return 2
|
||||
client = db.Historical(key)
|
||||
|
||||
global END
|
||||
try: # clamp END to the dataset's available end (GLBX lags ~1-3d)
|
||||
rng = client.metadata.get_dataset_range(dataset=DS)
|
||||
avail = (rng.get("end") or "")[:10]
|
||||
if avail:
|
||||
END = min(END, avail)
|
||||
except Exception:
|
||||
pass
|
||||
|
||||
os.makedirs(OUT_DIR, exist_ok=True)
|
||||
print(f"fetching {SCHEMA} per-root ({len(ROOTS)} roots, {STYPE}) {START}..{END}; "
|
||||
f"per-root get_cost gate (≤${PER_ROOT_CAP}), cumulative cap ${CAP_USD}")
|
||||
import time
|
||||
total_recs = 0
|
||||
spent = 0.0
|
||||
failed = []
|
||||
for r in ROOTS:
|
||||
sym = r + ".FUT"
|
||||
out_r = f"{OUT_DIR}/{r}.dbn"
|
||||
# per-root cost gate (free, no download)
|
||||
try:
|
||||
c_r = client.metadata.get_cost(dataset=DS, symbols=[sym], schema=SCHEMA,
|
||||
start=START, end=END, stype_in=STYPE)
|
||||
except Exception as e:
|
||||
print(f" {r}: get_cost failed ({type(e).__name__}); skipping"); failed.append(r); continue
|
||||
if c_r > PER_ROOT_CAP or spent + c_r > CAP_USD:
|
||||
print(f" {r}: cost ${c_r:.4f} would breach cap (spent ${spent:.2f}) — SKIP"); failed.append(r); continue
|
||||
ok = False
|
||||
for attempt in range(1, 4):
|
||||
try:
|
||||
data = client.timeseries.get_range(dataset=DS, symbols=[sym], schema=SCHEMA,
|
||||
start=START, end=END, stype_in=STYPE)
|
||||
data.to_file(out_r)
|
||||
n = sum(1 for _ in data)
|
||||
total_recs += n; spent += c_r
|
||||
print(f" {r}: ${c_r:.4f} {n:,} recs -> {out_r} (cum ${spent:.2f})")
|
||||
ok = True
|
||||
break
|
||||
except Exception as e:
|
||||
print(f" {r}: attempt {attempt} {type(e).__name__}; retry in 3s")
|
||||
time.sleep(3)
|
||||
if not ok:
|
||||
failed.append(r)
|
||||
print(f"DONE: {total_recs:,} records, get_cost-cum ${spent:.2f}, "
|
||||
f"{len(ROOTS)-len(failed)}/{len(ROOTS)} roots ok"
|
||||
+ (f"; FAILED/SKIPPED: {failed}" if failed else ""))
|
||||
return 1 if failed else 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
sys.exit(main())
|
||||
49
scripts/surfer/fetch_earnings.py
Normal file
49
scripts/surfer/fetch_earnings.py
Normal file
@@ -0,0 +1,49 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Fetch Nasdaq earnings calendar (date + symbol + surprise%) for all weekdays in the DBEQ range,
|
||||
for a proper PEAD test. Free, no key. Cached per date; gentle pacing + backoff.
|
||||
"""
|
||||
import datetime
|
||||
import json
|
||||
import os
|
||||
import time
|
||||
import urllib.request
|
||||
|
||||
OUT = "data/surfer/earnings"
|
||||
START = datetime.date(2023, 3, 28)
|
||||
END = datetime.date(2026, 5, 31)
|
||||
|
||||
|
||||
def get(u):
|
||||
for a in range(4):
|
||||
try:
|
||||
req = urllib.request.Request(u, headers={"User-Agent": "Mozilla/5.0", "Accept": "application/json"})
|
||||
return json.loads(urllib.request.urlopen(req, timeout=25).read())
|
||||
except Exception:
|
||||
if a == 3:
|
||||
return None
|
||||
time.sleep(4 * (a + 1))
|
||||
return None
|
||||
|
||||
|
||||
def main():
|
||||
os.makedirs(OUT, exist_ok=True)
|
||||
d = START
|
||||
n = fail = 0
|
||||
while d <= END:
|
||||
if d.weekday() < 5:
|
||||
f = f"{OUT}/{d}.json"
|
||||
if not os.path.exists(f):
|
||||
r = get(f"https://api.nasdaq.com/api/calendar/earnings?date={d}")
|
||||
if r is None:
|
||||
fail += 1
|
||||
else:
|
||||
rows = (r.get("data") or {}).get("rows") or []
|
||||
json.dump([(x.get("symbol"), x.get("surprise")) for x in rows], open(f, "w"))
|
||||
n += 1
|
||||
time.sleep(1.3)
|
||||
d += datetime.timedelta(days=1)
|
||||
print(f"DONE: fetched {n} new days, {fail} failed; cache dir {OUT}")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
33
scripts/surfer/fetch_equities.py
Normal file
33
scripts/surfer/fetch_equities.py
Normal file
@@ -0,0 +1,33 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Fetch all US equities daily OHLCV (DBEQ.BASIC) for the small/mid-cap factor test.
|
||||
|
||||
Budget-capped, get_cost-gated. Saves DBN to data/surfer/dbeq_ohlcv1d.dbn.
|
||||
"""
|
||||
import os
|
||||
import sys
|
||||
|
||||
import databento as db
|
||||
|
||||
CAP = 70.0
|
||||
DS, SCH = "DBEQ.BASIC", "ohlcv-1d"
|
||||
START, END = "2023-03-28", "2026-06-05"
|
||||
OUT = "data/surfer/dbeq_ohlcv1d.dbn"
|
||||
|
||||
|
||||
def main():
|
||||
c = db.Historical(os.environ["DATABENTO_API_KEY"])
|
||||
if os.path.exists(OUT):
|
||||
print("already downloaded"); return 0
|
||||
cost = c.metadata.get_cost(dataset=DS, symbols=["ALL_SYMBOLS"], schema=SCH, start=START, end=END)
|
||||
print(f"get_cost=${cost:.2f} cap=${CAP:.2f}")
|
||||
if cost > CAP:
|
||||
print("ABORT over cap"); return 1
|
||||
data = c.timeseries.get_range(dataset=DS, symbols=["ALL_SYMBOLS"], schema=SCH, start=START, end=END)
|
||||
os.makedirs("data/surfer", exist_ok=True)
|
||||
data.to_file(OUT)
|
||||
print(f"saved {OUT}")
|
||||
return 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
sys.exit(main())
|
||||
81
scripts/surfer/fetch_es_1m.py
Normal file
81
scripts/surfer/fetch_es_1m.py
Normal file
@@ -0,0 +1,81 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Fetch full-history ES front-month OHLCV-1m (continuous .c.0), year-chunked, BUDGET-CAPPED.
|
||||
|
||||
get_cost-gated (aborts over cap, no download); year chunks with quarter fallback on 504.
|
||||
Saves per-chunk DBN to data/surfer/es1m/ (gitignored). Key from env, never printed.
|
||||
"""
|
||||
import os
|
||||
import sys
|
||||
import time
|
||||
|
||||
import databento as db
|
||||
|
||||
CAP_USD = 25.00
|
||||
DS = "GLBX.MDP3"
|
||||
SCHEMA = "ohlcv-1m"
|
||||
SYM = "ES.c.0"
|
||||
STYPE = "continuous"
|
||||
OUT_DIR = "data/surfer/es1m"
|
||||
|
||||
|
||||
def fetch(client, start, end, out):
|
||||
data = client.timeseries.get_range(dataset=DS, symbols=[SYM], schema=SCHEMA,
|
||||
start=start, end=end, stype_in=STYPE)
|
||||
data.to_file(out)
|
||||
return sum(1 for _ in data)
|
||||
|
||||
|
||||
def main():
|
||||
key = os.environ.get("DATABENTO_API_KEY")
|
||||
if not key:
|
||||
print("DATABENTO_API_KEY not set"); return 2
|
||||
client = db.Historical(key)
|
||||
cost = client.metadata.get_cost(dataset=DS, symbols=[SYM], schema=SCHEMA,
|
||||
start="2010-06-06", end="2026-06-05", stype_in=STYPE)
|
||||
print(f"aggregate get_cost=${cost:.4f} cap=${CAP_USD:.2f}")
|
||||
if cost > CAP_USD:
|
||||
print("ABORT: over cap — nothing downloaded."); return 1
|
||||
os.makedirs(OUT_DIR, exist_ok=True)
|
||||
total = 0
|
||||
for year in range(2010, 2027):
|
||||
y0, y1 = f"{year}-01-01", f"{year+1}-01-01"
|
||||
if year == 2010:
|
||||
y0 = "2010-06-06"
|
||||
if year == 2026:
|
||||
y1 = "2026-06-05"
|
||||
out = f"{OUT_DIR}/ES_{year}.dbn"
|
||||
if os.path.exists(out):
|
||||
print(f" {year}: exists, skip"); continue
|
||||
ok = False
|
||||
for attempt in range(1, 3):
|
||||
try:
|
||||
n = fetch(client, y0, y1, out); total += n
|
||||
print(f" {year}: {n:,} recs -> {out}")
|
||||
ok = True; break
|
||||
except Exception as e:
|
||||
print(f" {year}: attempt {attempt} {type(e).__name__}; retry")
|
||||
time.sleep(3)
|
||||
if not ok: # fall back to quarter chunks
|
||||
print(f" {year}: year failed → quarter fallback")
|
||||
for q, (m0, m1) in enumerate([("01-01", "04-01"), ("04-01", "07-01"),
|
||||
("07-01", "10-01"), ("10-01", "12-31")], 1):
|
||||
qs, qe = f"{year}-{m0}", f"{year}-{m1}"
|
||||
if year == 2010 and q == 1:
|
||||
qs = "2010-06-06"
|
||||
if year == 2026 and q >= 3:
|
||||
continue
|
||||
qout = f"{OUT_DIR}/ES_{year}_q{q}.dbn"
|
||||
if os.path.exists(qout):
|
||||
continue
|
||||
try:
|
||||
n = fetch(client, qs, qe, qout); total += n
|
||||
print(f" {year} q{q}: {n:,} -> {qout}")
|
||||
except Exception as e:
|
||||
print(f" {year} q{q}: FAILED {type(e).__name__}")
|
||||
time.sleep(1)
|
||||
print(f"DONE: {total:,} records into {OUT_DIR}/")
|
||||
return 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
sys.exit(main())
|
||||
35
scripts/surfer/fetch_more_futures.py
Normal file
35
scripts/surfer/fetch_more_futures.py
Normal file
@@ -0,0 +1,35 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Expand the daily futures universe (16y ohlcv-1d, parent) for a diversified multi-asset test.
|
||||
|
||||
Budget-capped, get_cost-gated. Saves to data/surfer/<root>.dbn alongside the existing 22 roots.
|
||||
"""
|
||||
import os
|
||||
import sys
|
||||
|
||||
import databento as db
|
||||
|
||||
CAP = 20.0
|
||||
DS = "GLBX.MDP3"
|
||||
START, END = "2010-06-06", "2026-06-05"
|
||||
ADD = ["6S", "6N", "6M", "HO", "PL", "PA", "UB", "ZL", "ZM", "ZO", "KE", "ZR", "LE", "HE", "GF", "EMD", "NKD"]
|
||||
|
||||
|
||||
def main():
|
||||
c = db.Historical(os.environ["DATABENTO_API_KEY"])
|
||||
todo = [r for r in ADD if not os.path.exists(f"data/surfer/{r}.dbn")]
|
||||
cost = sum(c.metadata.get_cost(dataset=DS, symbols=[r + ".FUT"], schema="ohlcv-1d",
|
||||
start=START, end=END, stype_in="parent") for r in todo)
|
||||
print(f"to fetch: {todo}\naggregate get_cost=${cost:.2f} cap=${CAP:.2f}")
|
||||
if cost > CAP:
|
||||
print("ABORT over cap"); return 1
|
||||
for r in todo:
|
||||
data = c.timeseries.get_range(dataset=DS, symbols=[r + ".FUT"], schema="ohlcv-1d",
|
||||
start=START, end=END, stype_in="parent")
|
||||
data.to_file(f"data/surfer/{r}.dbn")
|
||||
print(f" {r}: saved")
|
||||
print(f"DONE: {len(todo)} roots")
|
||||
return 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
sys.exit(main())
|
||||
49
scripts/surfer/fetch_xsp.py
Normal file
49
scripts/surfer/fetch_xsp.py
Normal file
@@ -0,0 +1,49 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Fetch XSP (mini-SPX) options daily OHLCV for the equity-index VRP test. HARD-CAPPED.
|
||||
|
||||
get_cost-gated: prints the exact cost and ABORTS if above CAP — no surprise charges.
|
||||
Single dataset/schema/symbol/window. ohlcv-1d only (NOT definition).
|
||||
"""
|
||||
import os
|
||||
import sys
|
||||
|
||||
import databento as db
|
||||
|
||||
CAP = 70.00 # hard cap — abort if cost exceeds this
|
||||
DS, SCH, SYM = "OPRA.PILLAR", "ohlcv-1d", "XSP.OPT"
|
||||
START, END = "2013-04-01", "2026-06-05"
|
||||
OUT = "data/surfer/xsp_ohlcv1d.dbn"
|
||||
|
||||
|
||||
def main():
|
||||
c = db.Historical(os.environ["DATABENTO_API_KEY"])
|
||||
if os.path.exists(OUT):
|
||||
print(f"already downloaded: {OUT}"); return 0
|
||||
cost = c.metadata.get_cost(dataset=DS, symbols=[SYM], schema=SCH, start=START, end=END, stype_in="parent")
|
||||
gb = c.metadata.get_billable_size(dataset=DS, symbols=[SYM], schema=SCH, start=START, end=END, stype_in="parent") / 1e9
|
||||
print(f"EXACT get_cost = ${cost:.2f} ({gb:.2f} GB) CAP = ${CAP:.2f}")
|
||||
if cost > CAP:
|
||||
print(f"ABORT: ${cost:.2f} exceeds cap ${CAP:.2f} — NOTHING downloaded.")
|
||||
return 1
|
||||
print(f"under cap by ${CAP - cost:.2f} -> downloading {SYM} {SCH} year-chunked (same total cost)")
|
||||
os.makedirs("data/surfer/xsp", exist_ok=True)
|
||||
for yr in range(2013, 2027):
|
||||
ys = f"{yr}-01-01" if yr > 2013 else "2013-04-01"
|
||||
ye = f"{yr+1}-01-01" if yr < 2026 else "2026-06-05"
|
||||
out = f"data/surfer/xsp/xsp_{yr}.dbn"
|
||||
if os.path.exists(out):
|
||||
print(f" {yr}: exists"); continue
|
||||
for attempt in range(3):
|
||||
try:
|
||||
d = c.timeseries.get_range(dataset=DS, symbols=[SYM], schema=SCH, start=ys, end=ye, stype_in="parent")
|
||||
d.to_file(out)
|
||||
print(f" {yr}: saved ({os.path.getsize(out)/1e6:.1f} MB)")
|
||||
break
|
||||
except Exception as e:
|
||||
print(f" {yr}: attempt {attempt+1} {type(e).__name__}; retry")
|
||||
print(f"DONE -> data/surfer/xsp/")
|
||||
return 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
sys.exit(main())
|
||||
90
scripts/surfer/fetch_xvenue_hist.py
Normal file
90
scripts/surfer/fetch_xvenue_hist.py
Normal file
@@ -0,0 +1,90 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Fetch historical funding (Binance + Bybit) for liquid coins on both, to BACKTEST the cross-venue
|
||||
funding arb instead of waiting for live paper-forward. Aggregates 8h settlements to daily per venue.
|
||||
Cached. Saves data/surfer/xvenue/panel.json = {coin: {date: [binance_daily, bybit_daily]}}.
|
||||
"""
|
||||
import datetime
|
||||
import json
|
||||
import os
|
||||
import time
|
||||
import urllib.request
|
||||
|
||||
OUT = "data/surfer/xvenue"
|
||||
NCOINS = 60
|
||||
DAYS = 330
|
||||
|
||||
|
||||
def get(u, post=None):
|
||||
data = json.dumps(post).encode() if post else None
|
||||
h = {"User-Agent": "Mozilla/5.0", "Accept": "application/json"}
|
||||
for a in range(4):
|
||||
try:
|
||||
return json.loads(urllib.request.urlopen(urllib.request.Request(u, data=data, headers=h), timeout=25).read())
|
||||
except Exception:
|
||||
if a == 3:
|
||||
return None
|
||||
time.sleep(3 * (a + 1))
|
||||
return None
|
||||
|
||||
|
||||
def base(s):
|
||||
return s[:-4] if s.endswith("USDT") else s
|
||||
|
||||
|
||||
def day_of(ms):
|
||||
return datetime.datetime.utcfromtimestamp(int(ms) / 1000).strftime("%Y-%m-%d")
|
||||
|
||||
|
||||
def binance_hist(coin):
|
||||
r = get(f"https://fapi.binance.com/fapi/v1/fundingRate?symbol={coin}USDT&limit=1000")
|
||||
d = {}
|
||||
for x in (r or []):
|
||||
d.setdefault(day_of(x["fundingTime"]), 0.0)
|
||||
d[day_of(x["fundingTime"])] += float(x["fundingRate"])
|
||||
return d
|
||||
|
||||
|
||||
def bybit_hist(coin):
|
||||
out, end = {}, None
|
||||
for _ in range(6):
|
||||
u = f"https://api.bybit.com/v5/market/funding/history?category=linear&symbol={coin}USDT&limit=200"
|
||||
if end:
|
||||
u += f"&endTime={end}"
|
||||
r = get(u)
|
||||
lst = ((r or {}).get("result") or {}).get("list") or []
|
||||
if not lst:
|
||||
break
|
||||
for x in lst:
|
||||
t = x["fundingRateTimestamp"]
|
||||
out.setdefault(day_of(t), 0.0)
|
||||
out[day_of(t)] += float(x["fundingRate"])
|
||||
end = int(lst[-1]["fundingRateTimestamp"]) - 1
|
||||
time.sleep(0.3)
|
||||
return out
|
||||
|
||||
|
||||
def main():
|
||||
os.makedirs(OUT, exist_ok=True)
|
||||
binv = {base(x["symbol"]): float(x["quoteVolume"]) for x in get("https://fapi.binance.com/fapi/v1/ticker/24hr") if x["symbol"].endswith("USDT")}
|
||||
byr = get("https://api.bybit.com/v5/market/tickers?category=linear")["result"]["list"]
|
||||
bybv = {base(x["symbol"]): float(x.get("turnover24h", 0)) for x in byr if x["symbol"].endswith("USDT")}
|
||||
coins = sorted([c for c in binv if c in bybv and binv[c] > 10e6 and bybv[c] > 10e6], key=lambda c: -binv[c])[:NCOINS]
|
||||
print(f"universe: {len(coins)} coins on both venues >$10M/day")
|
||||
panel = {}
|
||||
for i, c in enumerate(coins):
|
||||
cf = f"{OUT}/{c}.json"
|
||||
if os.path.exists(cf):
|
||||
panel[c] = json.load(open(cf)); continue
|
||||
bn = binance_hist(c); by = bybit_hist(c)
|
||||
days = sorted(set(bn) & set(by))
|
||||
panel[c] = {d: [bn[d], by[d]] for d in days}
|
||||
json.dump(panel[c], open(cf, "w"))
|
||||
if i % 10 == 0:
|
||||
print(f" {i+1}/{len(coins)} {c}: {len(panel[c])} common days")
|
||||
time.sleep(0.4)
|
||||
json.dump(panel, open(f"{OUT}/panel.json", "w"))
|
||||
print(f"DONE: {len(panel)} coins -> {OUT}/panel.json")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
120
scripts/surfer/fetch_xvenue_hist2.py
Normal file
120
scripts/surfer/fetch_xvenue_hist2.py
Normal file
@@ -0,0 +1,120 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Deep multi-venue funding history (Binance + OKX + Hyperliquid, ~1yr) for a confident cross-venue
|
||||
backtest. Binance & HL are deep (~330d); OKX ~3mo. Aggregate to daily per venue. Cache per coin.
|
||||
Saves data/surfer/xvenue/panel2.json = {coin: {date: {bn, okx, hl}}}.
|
||||
"""
|
||||
import datetime
|
||||
import json
|
||||
import os
|
||||
import time
|
||||
import urllib.request
|
||||
|
||||
OUT = "data/surfer/xvenue2"
|
||||
NCOINS = 35
|
||||
DAYS = 330
|
||||
|
||||
|
||||
def get(u, post=None):
|
||||
data = json.dumps(post).encode() if post else None
|
||||
h = {"User-Agent": "Mozilla/5.0", "Accept": "application/json"}
|
||||
if post:
|
||||
h["Content-Type"] = "application/json"
|
||||
for a in range(4):
|
||||
try:
|
||||
return json.loads(urllib.request.urlopen(urllib.request.Request(u, data=data, headers=h), timeout=25).read())
|
||||
except Exception:
|
||||
if a == 3:
|
||||
return None
|
||||
time.sleep(3 * (a + 1))
|
||||
return None
|
||||
|
||||
|
||||
def base(s):
|
||||
return s[:-4] if s.endswith("USDT") else s
|
||||
|
||||
|
||||
def day_of(ms):
|
||||
return datetime.datetime.utcfromtimestamp(int(ms) / 1000).strftime("%Y-%m-%d")
|
||||
|
||||
|
||||
def binance_hist(coin, now_ms):
|
||||
out, start = {}, now_ms - DAYS * 86400000
|
||||
for _ in range(8): # paginate back full DAYS window
|
||||
r = get(f"https://fapi.binance.com/fapi/v1/fundingRate?symbol={coin}USDT&startTime={start}&limit=1000")
|
||||
if not r:
|
||||
break
|
||||
for x in r:
|
||||
k = day_of(x["fundingTime"]); out[k] = out.get(k, 0.0) + float(x["fundingRate"])
|
||||
last = int(r[-1]["fundingTime"])
|
||||
if last <= start or len(r) < 2:
|
||||
break
|
||||
start = last + 1
|
||||
return out
|
||||
|
||||
|
||||
def okx_hist(coin):
|
||||
out, after = {}, None
|
||||
for _ in range(6):
|
||||
u = f"https://www.okx.com/api/v5/public/funding-rate-history?instId={coin}-USDT-SWAP&limit=100"
|
||||
if after:
|
||||
u += f"&after={after}"
|
||||
r = get(u); data = (r or {}).get("data") or []
|
||||
if not data:
|
||||
break
|
||||
for x in data:
|
||||
k = day_of(x["fundingTime"]); out[k] = out.get(k, 0.0) + float(x["fundingRate"])
|
||||
after = data[-1]["fundingTime"]; time.sleep(0.2)
|
||||
return out
|
||||
|
||||
|
||||
def hl_hist(coin, now_ms):
|
||||
out, start = {}, now_ms - DAYS * 86400000
|
||||
for _ in range(30):
|
||||
r = get("https://api.hyperliquid.xyz/info", post={"type": "fundingHistory", "coin": coin, "startTime": start})
|
||||
if not r:
|
||||
break
|
||||
for x in r:
|
||||
k = day_of(x["time"]); out[k] = out.get(k, 0.0) + float(x["fundingRate"])
|
||||
last = int(r[-1]["time"])
|
||||
if last <= start or len(r) < 2:
|
||||
break
|
||||
start = last + 1
|
||||
return out
|
||||
|
||||
|
||||
def main():
|
||||
os.makedirs(OUT, exist_ok=True)
|
||||
now_ms = int(time.time() * 1000)
|
||||
binv = {base(x["symbol"]): float(x["quoteVolume"]) for x in get("https://fapi.binance.com/fapi/v1/ticker/24hr") if x["symbol"].endswith("USDT")}
|
||||
hlmeta = get("https://api.hyperliquid.xyz/info", post={"type": "metaAndAssetCtxs"})
|
||||
hlcoins = {u["name"] for u in hlmeta[0]["universe"]}
|
||||
coins = sorted([c for c in binv if c in hlcoins and binv[c] > 10e6], key=lambda c: -binv[c])[:NCOINS]
|
||||
print(f"universe: {len(coins)} coins (Binance>$10M & on Hyperliquid)")
|
||||
panel = {}
|
||||
for i, c in enumerate(coins):
|
||||
cf = f"{OUT}/{c}.json"
|
||||
if os.path.exists(cf):
|
||||
panel[c] = json.load(open(cf)); continue
|
||||
bn = binance_hist(c, now_ms); okx = okx_hist(c); hl = hl_hist(c, now_ms)
|
||||
days = sorted(set(bn) | set(okx) | set(hl))
|
||||
rec = {}
|
||||
for d in days:
|
||||
v = {}
|
||||
if d in bn:
|
||||
v["bn"] = bn[d]
|
||||
if d in okx:
|
||||
v["okx"] = okx[d]
|
||||
if d in hl:
|
||||
v["hl"] = hl[d]
|
||||
if len(v) >= 2:
|
||||
rec[d] = v
|
||||
panel[c] = rec
|
||||
json.dump(rec, open(cf, "w"))
|
||||
print(f" {i+1}/{len(coins)} {c}: {len(rec)} days w/ >=2 venues (bn{len(bn)} okx{len(okx)} hl{len(hl)})")
|
||||
time.sleep(0.3)
|
||||
json.dump(panel, open(f"{OUT}/panel2.json", "w"))
|
||||
print(f"DONE: {len(panel)} coins -> {OUT}/panel2.json")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
215
scripts/surfer/floor_experiment.py
Normal file
215
scripts/surfer/floor_experiment.py
Normal file
@@ -0,0 +1,215 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Surfer Phase 0 — diversified-trend FLOOR + validation, on the GPU (PyTorch).
|
||||
|
||||
Loads per-root parent OHLCV-1d (data/surfer/*.dbn), builds ratio-adjusted continuous
|
||||
front-month series (volume-roll), then runs the floor + the anti-overfit validation
|
||||
ENTIRELY on the GPU (torch.cuda): TSMOM(1/3/12mo) → inverse-vol size → vol-target →
|
||||
net-of-cost daily portfolio returns; CPCV(5th-pct OOS Sharpe), Deflated Sharpe, PBO,
|
||||
IS↔OOS rank-consistency. Prints the PASS/FAIL verdict.
|
||||
|
||||
Data-loading/continuous-assembly is CPU (I/O staging, like mapped-pinned upload); ALL
|
||||
strategy + statistics compute is on the GPU. Verdict gates per the spec §5.
|
||||
"""
|
||||
import glob
|
||||
import math
|
||||
from itertools import combinations
|
||||
|
||||
import numpy as np
|
||||
import torch
|
||||
|
||||
DEV = "cuda" if torch.cuda.is_available() else None
|
||||
TICK_NA = float("nan")
|
||||
|
||||
|
||||
# ---------- data loading + continuous assembly (CPU staging) ----------
|
||||
_MONTHS = "FGHJKMNQUVXZ"
|
||||
|
||||
def load_continuous(path):
|
||||
"""Roll-neutralized continuous daily series from a root's parent OHLCV-1d.
|
||||
Front = max-volume OUTRIGHT per day; daily return only when the front contract is
|
||||
unchanged (roll-day returns ZEROED — no ratio-adjust, no fabricated roll jumps).
|
||||
Returns (day_int[N], adj_close[N]) where adj_close = exp(cumsum(clean_return))."""
|
||||
import re
|
||||
import databento as db
|
||||
root = path.split("/")[-1][:-4]
|
||||
pat = re.compile("^" + re.escape(root) + "[" + _MONTHS + r"]\d{1,2}$") # outrights only (e.g. RBZ5)
|
||||
df = db.DBNStore.from_file(path).to_df().reset_index()
|
||||
df = df[df["close"] > 0].copy()
|
||||
df = df[df["symbol"].astype(str).str.match(pat)]
|
||||
if df.empty:
|
||||
raise ValueError("no outright contracts after filter")
|
||||
df["day"] = (df["ts_event"].astype("int64") // (86_400 * 10**9))
|
||||
# (day, instrument) -> close lookup, for proper roll-day return of the HELD contract
|
||||
cl = {(int(d), int(i)): float(c)
|
||||
for d, i, c in zip(df["day"], df["instrument_id"], df["close"])}
|
||||
idx = df.groupby("day")["volume"].idxmax()
|
||||
f = df.loc[idx, ["day", "instrument_id", "close"]].sort_values("day")
|
||||
days = f["day"].to_numpy(np.int64)
|
||||
inst = f["instrument_id"].to_numpy(np.int64)
|
||||
close = f["close"].to_numpy(np.float64)
|
||||
lr = np.zeros(len(days))
|
||||
for t in range(1, len(days)):
|
||||
a, b = int(inst[t - 1]), int(inst[t])
|
||||
if a == b:
|
||||
lr[t] = np.log(close[t] / close[t - 1])
|
||||
else: # roll: use the HELD (old front 'a') contract's own return through the roll day
|
||||
ca_t = cl.get((int(days[t]), a))
|
||||
if ca_t is not None and ca_t > 0:
|
||||
lr[t] = np.log(ca_t / close[t - 1])
|
||||
# else: 'a' not trading on day t → leave 0 (rare)
|
||||
return days, np.exp(np.cumsum(lr))
|
||||
|
||||
|
||||
def build_panel():
|
||||
paths = sorted(glob.glob("data/surfer/*.dbn"))
|
||||
series = {}
|
||||
for p in paths:
|
||||
root = p.split("/")[-1][:-4]
|
||||
try:
|
||||
d, c = load_continuous(p)
|
||||
if len(d) > 300:
|
||||
series[root] = (d, c)
|
||||
except Exception as e:
|
||||
print(f" skip {root}: {type(e).__name__} {e}")
|
||||
roots = sorted(series)
|
||||
alldays = sorted(set().union(*[set(series[r][0].tolist()) for r in roots]))
|
||||
didx = {d: i for i, d in enumerate(alldays)}
|
||||
C = np.full((len(alldays), len(roots)), np.nan)
|
||||
for j, r in enumerate(roots):
|
||||
d, c = series[r]
|
||||
for dd, cc in zip(d.tolist(), c.tolist()):
|
||||
C[didx[dd], j] = cc
|
||||
return roots, np.array(alldays), C
|
||||
|
||||
|
||||
# ---------- floor + validation (GPU) ----------
|
||||
def sharpe(x): # x: [..., T] torch; Sharpe over last dim, NaN-aware
|
||||
m = torch.nanmean(x, dim=-1)
|
||||
v = torch.nanmean((x - m.unsqueeze(-1)) ** 2, dim=-1).clamp_min(1e-18).sqrt()
|
||||
return m / v * math.sqrt(252)
|
||||
|
||||
|
||||
def run(roots, days, C):
|
||||
Ct = torch.tensor(C, device=DEV, dtype=torch.float64) # [T, N]
|
||||
logC = torch.log(Ct)
|
||||
R = torch.zeros_like(logC); R[1:] = logC[1:] - logC[:-1] # daily log returns
|
||||
R = torch.nan_to_num(R, nan=0.0)
|
||||
T, N = R.shape
|
||||
|
||||
# EWMA vol (per root), halflife 21d
|
||||
a = 1 - math.exp(math.log(0.5) / 21)
|
||||
v = torch.zeros_like(R); v[0] = R[0] ** 2
|
||||
for t in range(1, T):
|
||||
v[t] = a * R[t] ** 2 + (1 - a) * v[t - 1]
|
||||
sig_vol = (v * 252).clamp_min(1e-12).sqrt() # annualized vol [T,N]
|
||||
|
||||
# CONTINUOUS vol-normalized TSMOM (Baz et al / AQR): tanh of the trend t-stat,
|
||||
# averaged over 1/3/12mo lookbacks (252 skips last 21d). Strength in [-1,1] →
|
||||
# multiplied by inverse-vol sizing below (no double vol-counting; sign-floor superseded).
|
||||
import os
|
||||
mode = os.environ.get("FLOOR_SIGNAL", "sign") # "sign" (canonical MOP/AQR) or "cont"
|
||||
s = torch.zeros_like(R)
|
||||
for L, skip in [(21, 0), (63, 0), (252, 21)]:
|
||||
raw = torch.zeros_like(R)
|
||||
for t in range(L + skip, T):
|
||||
e = t - skip
|
||||
raw[t] = logC[e] - logC[e - L] # trend log-return over L
|
||||
if mode == "cont":
|
||||
norm = raw / (sig_vol * math.sqrt(L / 252.0)).clamp_min(1e-6) # ≈ trend Sharpe
|
||||
s = s + torch.tanh(norm)
|
||||
else:
|
||||
s = s + torch.sign(raw)
|
||||
s = (s / 3.0)
|
||||
s = torch.nan_to_num(s, nan=0.0)
|
||||
|
||||
# inverse-vol target weights, vol-targeted to 10% annual
|
||||
# No-trade band on the continuous signal (turnover control; standard CTA practice).
|
||||
# Pre-registered band = 0.10 in [-1,1] signal units: hold position until target moves > band.
|
||||
band = 0.10
|
||||
held = torch.zeros_like(s)
|
||||
held[0] = s[0]
|
||||
for t in range(1, T):
|
||||
move = (s[t] - held[t - 1]).abs() > band
|
||||
held[t] = torch.where(move, s[t], held[t - 1])
|
||||
risk_budget = 0.10 / math.sqrt(N)
|
||||
w = held * (risk_budget / sig_vol.clamp_min(1e-6))
|
||||
w = torch.nan_to_num(w, nan=0.0)
|
||||
# daily portfolio return (lag weights), net of cost
|
||||
port = (w[:-1] * R[1:]).sum(dim=1) # [T-1]
|
||||
turn = (w[1:] - w[:-1]).abs().sum(dim=1)
|
||||
cost1 = torch.nan_to_num(turn * (2.0 / 1e4), nan=0.0) # ~2bp round-trip on weight turnover
|
||||
|
||||
def voltarget(x): # 10% annual, trailing-63d realized
|
||||
sc = torch.ones_like(x)
|
||||
for t in range(63, len(x)):
|
||||
rv = x[t-63:t].std() * math.sqrt(252)
|
||||
sc[t] = (0.10 / rv.clamp_min(1e-6))
|
||||
return x * sc.clamp(0, 5)
|
||||
|
||||
net = voltarget(port - cost1)
|
||||
net2x = voltarget(port - 2 * cost1) # SV cost stress (2x)
|
||||
full_sr = float(sharpe(net))
|
||||
full_sr_2x = float(sharpe(net2x))
|
||||
# CPCV: 10 blocks, leave-2-out test → 45 paths
|
||||
n = len(net); nb = 10; k = 2
|
||||
blocks = [torch.arange(i*n//nb, (i+1)*n//nb, device=DEV) for i in range(nb)]
|
||||
oos_sr = []
|
||||
for combo in combinations(range(nb), k):
|
||||
idx = torch.cat([blocks[b] for b in combo])
|
||||
oos_sr.append(float(sharpe(net[idx])))
|
||||
oos_sr = np.array(oos_sr)
|
||||
oos_5pct = float(np.percentile(oos_sr, 5))
|
||||
|
||||
# IS(<2024)/OOS(>=2024) split by day epoch (days are UTC-day ints)
|
||||
split_day = 19723 # 2024-01-01 in days since epoch
|
||||
is_mask = days[1:] < split_day
|
||||
oos_mask = ~is_mask
|
||||
sr_is = float(sharpe(net[torch.tensor(is_mask, device=DEV)]))
|
||||
sr_oos = float(sharpe(net[torch.tensor(oos_mask, device=DEV)])) if oos_mask.sum() > 30 else float("nan")
|
||||
|
||||
# Deflated Sharpe — HONEST n_trials = trend-floor variants tried (sign + continuous ≈ 3),
|
||||
# NOT the 64 unrelated RL-microstructure commits.
|
||||
n_trials = 3
|
||||
sr_daily = full_sr / math.sqrt(252)
|
||||
sr0 = (1/math.sqrt(n)) * ((1-0.5772)*_ppf(1-1.0/n_trials) + 0.5772*_ppf(1-1.0/(n_trials*math.e)))
|
||||
dsr = _ncdf((sr_daily - sr0) * math.sqrt(n - 1))
|
||||
med = float(np.median(oos_sr))
|
||||
|
||||
print("\n================ SURFER PHASE 0 — FLOOR VERDICT (GPU, continuous TSMOM + no-trade band) ================")
|
||||
print(f"device={DEV} roots={len(roots)} days={n} span≈{n/252:.1f}y")
|
||||
print(f"full-sample Sharpe (net) = {full_sr:+.3f} (2x-cost: {full_sr_2x:+.3f})")
|
||||
print(f"IS(<2024) / OOS(>=2024) Sharpe = {sr_is:+.3f} / {sr_oos:+.3f}")
|
||||
print(f"CPCV 45 paths: median={med:+.3f} 5th-pct={oos_5pct:+.3f}")
|
||||
print(f"Deflated Sharpe (n_trials={n_trials}) = {dsr:.3f}")
|
||||
print("---- DEVELOP-grade (is there a real edge worth building on?) ----")
|
||||
d1, d2, d3 = med > 0, (sr_is > 0 and sr_oos > 0), full_sr_2x > 0
|
||||
print(f" D1 CPCV median > 0 : {'PASS' if d1 else 'FAIL'} ({med:+.3f})")
|
||||
print(f" D2 IS & OOS both positive : {'PASS' if d2 else 'FAIL'} ({sr_is:+.2f}/{sr_oos:+.2f})")
|
||||
print(f" D3 survives 2x cost : {'PASS' if d3 else 'FAIL'} ({full_sr_2x:+.3f})")
|
||||
print("---- DEPLOY-grade (bet real money?) ----")
|
||||
p1, p3 = oos_5pct > 0, dsr > 0.95
|
||||
print(f" P1 CPCV 5th-pct OOS > 0 : {'PASS' if p1 else 'FAIL'} ({oos_5pct:+.3f})")
|
||||
print(f" P3 Deflated Sharpe > 0.95 : {'PASS' if p3 else 'FAIL'} ({dsr:.3f})")
|
||||
print(f"==> DEVELOP: {'PASS → real edge; build the ML regime overlay (must beat this floor OOS)' if (d1 and d2 and d3) else 'FAIL → no developable edge'}")
|
||||
print(f"==> DEPLOY : {'PASS' if (p1 and p3) else 'NOT YET (expected for a bare floor; ML overlay + tuning target this)'}")
|
||||
|
||||
|
||||
def _ppf(p): # inverse normal cdf (Acklam approx)
|
||||
from statistics import NormalDist
|
||||
return NormalDist().inv_cdf(p)
|
||||
def _ncdf(x):
|
||||
from statistics import NormalDist
|
||||
return NormalDist().cdf(x)
|
||||
|
||||
|
||||
def main():
|
||||
if DEV is None:
|
||||
raise SystemExit("CUDA not available to torch")
|
||||
roots, days, C = build_panel()
|
||||
if not roots:
|
||||
raise SystemExit("no data in data/surfer/*.dbn yet")
|
||||
run(roots, days, C)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
80
scripts/surfer/fng_gate.py
Normal file
80
scripts/surfer/fng_gate.py
Normal file
@@ -0,0 +1,80 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Phase A (non-price) — Fear & Greed sentiment gate (market-timing signal).
|
||||
|
||||
Does crypto sentiment (alternative.me Fear & Greed, daily since 2018) predict forward
|
||||
crypto-market returns? Contrarian hypothesis: extreme fear -> buy, extreme greed -> sell.
|
||||
Market = equal-weight return over the crypto_pit universe. Tests contrarian-level,
|
||||
extreme-only, and sentiment-change signals -> forward 1/7/30d returns, IC + timing-strategy
|
||||
Sharpe (net 5bp), per-year, OOS. A market-timing sleeve is directional -> diversifies the
|
||||
market-neutral momentum book if it works.
|
||||
"""
|
||||
import json
|
||||
import math
|
||||
import os
|
||||
import sys
|
||||
import urllib.request
|
||||
|
||||
import numpy as np
|
||||
|
||||
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
|
||||
import pit_sweep # noqa: E402
|
||||
from signal_sweep import sharpe_t # noqa: E402
|
||||
import torch # noqa: E402
|
||||
|
||||
DEV = "cuda" if torch.cuda.is_available() else "cpu"
|
||||
DAY = 86_400
|
||||
|
||||
|
||||
def fng():
|
||||
cache = "data/surfer/fng.json"
|
||||
if os.path.exists(cache):
|
||||
return json.load(open(cache))
|
||||
d = json.load(urllib.request.urlopen(urllib.request.Request(
|
||||
"https://api.alternative.me/fng/?limit=0&format=json", headers={"User-Agent": "curl/8"}), timeout=30))["data"]
|
||||
out = {int(r["timestamp"]) // DAY: float(r["value"]) for r in d}
|
||||
os.makedirs("data/surfer", exist_ok=True); json.dump(out, open(cache, "w"))
|
||||
return {int(k): v for k, v in out.items()}
|
||||
|
||||
|
||||
def main():
|
||||
fg = fng()
|
||||
syms, days, close, qv, fund = pit_sweep.load()
|
||||
R = np.zeros_like(close); R[1:] = np.log(close)[1:] - np.log(close)[:-1]; R = np.where(np.isfinite(R), R, 0.0)
|
||||
mkt = {int(days[t]): float(np.nanmean(R[t][np.isfinite(close[t])])) for t in range(len(days))} # EW market ret day t
|
||||
common = sorted(set(fg) & set(mkt))
|
||||
f = np.array([fg[d] for d in common]); r = np.array([mkt[d] for d in common])
|
||||
cum = np.concatenate([[0], np.cumsum(r)])
|
||||
year = (1970 + np.array(common) / 365.25).astype(int)
|
||||
n = len(common); split = int(0.7 * n)
|
||||
|
||||
def fwd(h):
|
||||
out = np.full(n, np.nan)
|
||||
out[:n - h] = cum[h + 1:n + 1] - cum[1:n - h + 1] # sum of mkt returns t+1..t+h
|
||||
return out
|
||||
|
||||
sigs = {
|
||||
"contrarian_level": -(f - 50),
|
||||
"extreme_only": np.where(f < 25, 1.0, np.where(f > 75, -1.0, 0.0)),
|
||||
"sentiment_chg": np.concatenate([[np.nan] * 7, f[7:] - f[:-7]]), # rising sentiment (momentum)
|
||||
}
|
||||
print(f"\n===== FEAR&GREED GATE — market-timing, {n} days ({common[0]}..{common[-1]}), cost 5bp =====")
|
||||
print(f"{'signal':>16} " + " ".join(f"h={h}:IC" for h in [1, 7, 30]) + " timing(h=7): full/IS/OOS/peryr")
|
||||
for nm, s in sigs.items():
|
||||
ics = []
|
||||
for h in [1, 7, 30]:
|
||||
fw = fwd(h); m = np.isfinite(s) & np.isfinite(fw)
|
||||
ics.append(float(np.corrcoef(s[m], fw[m])[0, 1]) if m.sum() > 100 else float("nan"))
|
||||
# timing strategy at h=7 (weekly hold): position = sign(s) lagged, daily mkt return, ~weekly turnover
|
||||
pos = np.sign(s); pnl = np.full(n, np.nan)
|
||||
pnl[1:] = pos[:-1] * r[1:]
|
||||
turn = np.abs(np.diff(np.nan_to_num(pos)))
|
||||
pnl[1:] -= turn * 5e-4
|
||||
T = lambda x: torch.tensor(x[np.isfinite(x)], device=DEV, dtype=torch.float64)
|
||||
full = sharpe_t(T(pnl)); isr = sharpe_t(T(pnl[:split])); oos = sharpe_t(T(pnl[split:]))
|
||||
py = " ".join(f"{y}:{sharpe_t(T(pnl[year==y])):+.1f}" for y in range(2019, 2027) if (year == y).sum() > 60)
|
||||
print(f"{nm:>16} " + " ".join(f"{ic:+.3f}" for ic in ics) + f" {full:+.2f}/{isr:+.2f}/{oos:+.2f} {py}")
|
||||
print("\nVERDICT: contrarian IC>0 (fear->up) + timing OOS Sharpe>0 net = real sentiment signal (directional sleeve).")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
Some files were not shown because too many files have changed in this diff Show More
Reference in New Issue
Block a user