Files
fxhnt/docs/superpowers/plans/CONSOLIDATION-PROGRESS.md
Jeroen Grusewski 4e7adf42f2 docs(migrate): scheduling-consistency upgrade complete + verified
All-Dagster K8sRunLauncher upgrade done: nightly ETL+precompute runs in
per-run 6-8Gi Jobs (daemon never OOMs), bybit_book_precompute asset folds in
the compare-measured/backtest-refs crons, execution stays in crons. Verified
end-to-end (run SUCCESS, daemon 0 restarts, all data fresh 2026-07-19). VRP +
compare-measured crons deleted; factory + ucits armed. Documented.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-19 10:14:28 +00:00

383 lines
35 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Foxhunt→Bizworx Consolidation — Live Progress (resume anchor)
## 🚦 T8 CUTOVER EXECUTED (2026-07-18 ~20:15 UTC) — fxhnt.ai NOW SERVES FROM BIZWORX ✅
**The fund has cut over. foxhunt is FROZEN but INTACT (rollback available). DELETE NOTHING until T10 observation passes.**
- **C1 (freeze foxhunt):** suspended fund cronjobs (ucits-rebalancer, ucits-volume-ibkr, compare-measured-precompute)
+ scaled foxhunt dagster/fxhnt-dashboard/ib-gateway → 0/0. Infra cronjobs (postgres-backup/pvc-autoscaler/
billing-scraper) LEFT RUNNING (they die at teardown, not now). PASS.
- **C2 (delta DB migrate):** pg_dump -Fc fxhnt (47.8M) from frozen foxhunt → restore into bizworx. GOTCHA:
`pg_restore --clean --if-exists` over the existing rehearsal copy FAILED (dropped _timescaledb_internal schema
mid-restore → 1136 errors, DB stuck restoring:on). FIX (rehearsal-proven path): DROP+CREATE empty fxhnt DB,
then pre_restore → pg_restore (NO --clean) → post_restore = exit 0, 0 errors. EVAL C2 PASS: all counts match
frozen source (forward_nav 106, backtest_summary 6, bybit_sleeve_ret 7023, xsp_option_bars 814373, exec_fills 9,
ibkr_account_nav 4, forward_anchor 20; max_anchor_t0 2026-07-14, max_fwd_nav 2026-07-17; 37 tables/3 hypertables).
NOTE: counts were IDENTICAL to the rehearsal (fund already suspended during parallel work → zero real delta).
LESSON: NEVER pg_restore --clean over a live timescaledb DB; always restore into a fresh empty DB.
- **C3 (bizworx fund live):** scaled dagster/cockpit/ib-gateway→1 + rollout-restart cockpit&dagster (fresh DB
conns). EVAL C3 PASS: cockpit serves real migrated data over the proxy (/ + /strategy/multistrat + /xsfunding
= 200, shows PASS/WAIT/Sharpe/multistrat/xsfunding); ib-gateway 1/1 STABLE (~12min, foxhunt released IBKR
session); dagster 3/3, ingress 3/3, postgres 1/1.
- **C4 (DNS flip):** Scaleway fxhnt.ai records dashboard/cockpit/dagster/git A → **100.93.141.11** (bizworx
fxhnt-ingress), TTL 60 (was 100.95.225.27 foxhunt proxy). SYNTAX: `scw dns record set dns-zone=fxhnt.ai
name=<n> type=A ttl=60 values.0=100.93.141.11` (values.0=, NOT data=). EVAL C4 PASS over REAL public DNS:
dashboard/git/dagster.fxhnt.ai → 100.93.141.11, HTTP 200, tls_verify=0. Content confirmed: git=bizworx Gitea,
dashboard=cockpit "Overview · fxhnt" w/ real data, SSH `git ls-remote ssh://git@git.fxhnt.ai/gitadmin/fxhnt.git`
returns real refs (HEAD d66d1132) via socat path. (SSH hostkey-changed warning = correct, it's bizworx now.)
- **C5 (arm fund):** un-suspended ONLY fxhnt-ucits-rebalancer on bizworx (schedule 0 9 * * 1-5 = 09:00 UTC wkdays;
next run tomorrow AM). US(ibkr)/bybit/vrp/factory/ucits-volume all stay SUSPENDED. EVAL C5 (first armed run
green) is OBSERVABLE at next 09:00 UTC — deps verified (DB+ib-gateway+git-sync all live).
- **ROLLBACK (still trivial until T11):** `scw dns record set ... values.0=100.95.225.27` (back to foxhunt proxy)
+ `kubectl --context $FOX -n foxhunt scale deploy/dagster deploy/fxhnt-dashboard deploy/ib-gateway --replicas=1`
+ un-suspend foxhunt fund cronjobs. foxhunt Postgres still holds pre-cutover state.
- **NEXT:** T9 verify (cockpit LOCAL browser + prod, gate/forward-track intact) → T10 observation window (watch
the first ucits run + a day) → T11 Phase-2 triage → T12 teardown. **ROTATE the pasted tailscale key.** DELETE NOTHING.
Branch `chore/foxhunt-bizworx-consolidation`. Real date 2026-07-18. Plan:
`docs/superpowers/plans/2026-07-18-foxhunt-to-bizworx-consolidation.md`. Contexts:
`FOX=foxhunt-34a1e3c4-…`, `BIZ=bizworx-prod-kapsule-05e1dae3-…` (both in ~/.kube/config).
## Done + verified
- **Access restored:** appended this box's egress `51.158.111.225/32` to the foxhunt Kapsule API ACL
(`scw k8s acl`, entry `bizworx-devbox-v4`) + `scw k8s kubeconfig install` → foxhunt context works.
- **T1** ns `foxhunt` created on bizworx; StorageClass `sbs-default` matches; source = PG16.13/tsdb2.25.1.
- **T2** 6 fund secrets copied FOX→BIZ ns foxhunt (db-credentials, databento, tiingo-api, ibkr-credentials,
argo-git-ssh-key, scw-registry). NOT copied: bybit-testnet (absent/suspended), tailscale-auth (deferred).
- **T3** gitea consolidated → bizworx gitea (ns gitea): `gitadmin/fxhnt` (8 heads/0 tags, master d66d113)
+ `gitadmin/foxhunt` (13 heads/105 tags, main 8213f29a5), refs verified identical. foxhunt (625MB) done
via direct pack-copy (HTTP push + server-side index-pack both failed on size/OOM). `git.fxhnt.ai` STILL
points at foxhunt (DNS re-point is T8). Caveat: foxhunt archive's gitea DB metadata may lag (disk-correct).
- **T4** manifests re-targeted (committed): git-sync `gitea-sshd.foxhunt``gitea-ssh.gitea.svc.cluster.local:22`;
nodeSelector→`large`; ib-gateway pull-secret removed (public ghcr image). Server-side dry-run clean.
- **T5/T6** Postgres on bizworx PINNED to `timescale/timescaledb:2.25.1-pg16` (latest=2.28.3 broke catalog
restore — REHEARSAL caught it). Rehearsal restore of `fxhnt` DB: exit 0, 0 errors, all row counts match
(forward_nav 106, backtest_summary 6, bybit_sleeve_ret 7023, xsp_option_bars 814373, exec_fills 9,
ibkr_account_nav 4), 37/37 tables, 3/3 hypertables. bizworx fxhnt DB currently holds this rehearsal copy.
## T7 IN PROGRESS (parallel verify — foxhunt still fully live)
- **7.1 DONE** — argo-git-ssh-key deploy-key registered read-only on bizworx gitadmin/fxhnt (id 2);
verified `git ls-remote` via the key = 8 refs.
- **cockpit DONE** — `fxhnt-dashboard` deployed on bizworx; git-sync OK (cloned d66d113 from bizworx gitea);
cockpit container Running+ready, renders migrated data (overview + /strategy/xsfunding + /strategy/multistrat
all HTTP OK with real WAIT/PASS gates). Verified from INSIDE the container (port-forward blocked — see below).
- **BLOCKER / decision pending — tailscale:** dagster + cockpit pods have a tailscale sidecar needing
`tailscale-auth` (not copied) → sidecar CreateContainerConfigError → pod phase Pending → port-forward
refuses. Cockpit itself is fine. DECIDE: **A** fresh reusable tailscale auth-key (bizworx pods join as
extra nodes alongside live foxhunt ones; removed at cutover) OR **B** disable sidecar for the parallel
phase, verify via port-forward, enable tailscale at T8 cutover with real hostnames.
## T7 REMAINING
- ib-gateway → apply + verify IBKR paper connect (no 2FA/IP-limit per operator) + whatIf.
- dagster → apply (check its NetworkPolicy allows egress to ns gitea for git-sync — cross-ns) + trigger
nightly asset graph green.
- CronJobs → apply all SUSPENDED; then one `fxhnt-ucits-rebalancer` dry-run (EU hours).
- NP note: fund NPs (dagster, cronjobs) allow egress to gitea in-ns; on bizworx gitea is ns `gitea`
may need namespaceSelector. (cockpit had no NP → git-sync worked.)
## HARD RULE (operator): verify-first, per-step eval, DELETE NOTHING until Phase-1 verified + observation
window (T10) passes. foxhunt stays live/intact through T8, suspended-but-intact until T12 teardown.
Rollback until T11 = re-point DNS + un-suspend foxhunt.
## T7 update (2026-07-18, later)
- tailscale-auth secret created on bizworx (operator gave a fresh reusable key — OPTION A; key was pasted in
transcript, ROTATE after cutover). cockpit pod now 2/2 Running (cockpit + tailscale sidecar joined).
- ib-gateway deployed: IBC logs in to DU9600528 (Simulated Trading) — IBKR paper connect WORKS (no 2FA/IP).
⚠️ pod had 2 restarts during IBC startup churn — WATCH it stabilises to 1/1 ready; verify API port 4003/4004
+ a whatIf before arming any rebalancer.
- dagster was Pending on missing PVC `fxhnt-surfer-data` (dropped from migration; manifest still mounts it) →
created empty 5Gi PVC (sbs-default); dagster now scheduling (git-sync init). unlock_calendar.json rebuilds.
- REMAINING T7: confirm ib-gateway 1/1 + whatIf; confirm dagster git-sync (cross-ns NP to gitea) + trigger
nightly asset graph green; apply CronJobs SUSPENDED; one fxhnt-ucits-rebalancer dry-run (EU hours).
- Then T8 cutover (freeze foxhunt, delta restore, DNS re-point git.fxhnt.ai + dashboard.fxhnt.ai) → T9 verify
→ T10 observe → T11 triage → T12 teardown. foxhunt still fully live; delete NOTHING yet.
## T7 update 2 (2026-07-18)
- dagster FIXED + 3/3 Running: (1) created `dagster` DB on bizworx postgres (fresh dagster, not migrated);
(2) dagster NP gitea-egress needed a cross-ns `namespaceSelector {kubernetes.io/metadata.name: gitea}`
(committed in dagster.yaml) → git-sync from bizworx gitea now OK. webserver+daemon+tailscale all ready.
- ib-gateway: connects to DU9600528 (IBC login OK) but exits 0 after ~6min = IBKR single-session conflict with
the STILL-LIVE foxhunt ib-gateway. Connect verified; stable run only after foxhunt-suspend (T8). NOT a bug.
- T7 CORE VERIFIED: postgres+cockpit+dagster all up on bizworx against migrated data.
- REMAINING T7: (a) apply the SAME cross-ns gitea-NP fix to every cronjob NP (jobs/fxhnt-*.yaml) BEFORE they
run (else git-sync blocked like dagster was); (b) apply CronJobs SUSPENDED; (c) one ucits dry-run (needs
ib-gateway sole-session → do at/after cutover, or briefly scale foxhunt ib-gateway to 0); (d) optional dagster
asset-graph run. Then T8 cutover.
- ALSO for cutover: the tailscale bizworx pods joined as extra nodes (fresh key); at T8 suspend foxhunt so the
real hostnames + IBKR session move cleanly. ROTATE the pasted tailscale key post-cutover.
## P1 EXECUTION (2026-07-18) — wildcard cert + the cross-project DNS-01 IAM fix
- Applied `infra/k8s/services/fxhnt-wildcard-cert.yaml` (Certificate fxhnt.ai + *.fxhnt.ai, ClusterIssuer
letsencrypt-prod, secret fxhnt-wildcard-tls, ns foxhunt) on BIZ.
- **DNS-01 FAILED: `403 Forbidden: domain not found`.** Root cause: `fxhnt.ai` DNS zone lives in the **foxhunt
project c293eb98**; the bizworx cert-manager key (`SCWW9261…`, app `bizworx-prod-runtime` 48418a71) has
`DomainsDNSFullAccess` but SCOPED to the **bizworx project 96d0717d** only → can't touch fxhnt.ai's zone.
- **DEAD END (operator's first pick "move the domain to bizworx project"): Scaleway BLOCKS it.** fxhnt.ai is an
EXTERNAL domain (`scw domain domain list` → IS EXTERNAL=true, REGISTRAR=EXTERNAL). `scw dns zone update
… project-id=…` on the apex → **`Root zone can't be updated`**; on a scratch sub-zone → **`Subdomain not
allowed`**; `scw domain domain` has NO move-project subcommand. Verified via a throwaway `_cutovertest.fxhnt.ai`
sub-zone (created+move-attempt+deleted). So the domain CANNOT be reassigned to the bizworx project via API.
- **ACTUAL FIX (operator's 2nd pick, applied): add an IAM rule.** `scw iam rule create
policy-id=8573ce25-… permission-set-names.0=DomainsDNSFullAccess project-ids.0=c293eb98-…` → policy
`bizworx-prod-runtime-scoped` now has 2 rules (orig 6-perms@96d0717d untouched + new DNS@c293eb98). This lets
the existing bizworx cert-manager key write the DNS-01 TXT on fxhnt.ai. Domain STAYS external in c293eb98;
only IAM changed. NOTE for teardown: this rule references the foxhunt project — harmless to keep, or tighten
later; if the foxhunt PROJECT (not just cluster) is ever deleted, the fxhnt.ai domain/zone must first move
registrar-side or certs break. (Cluster teardown ≠ project deletion; the DNS zone + domain survive T12.)
- Forced cert-manager retry (deleted order+CR) after the IAM change → 403 GONE, challenges progressing normally
(apex fxhnt.ai challenge = `valid`/"Successfully authorized domain"; wildcard `*.fxhnt.ai` challenge finishing
its DNS-01 propagation). TXT is being written to the fxhnt.ai zone by cert-manager (verified via
`scw dns record list dns-zone=fxhnt.ai | grep acme`). P1 is SUCCEEDING — cert will go READY shortly.
If resuming and cert not yet READY: `kubectl --context $BIZ -n foxhunt get certificate fxhnt-wildcard -o wide`;
if stuck >15min, `kubectl delete order,challenge,certificaterequest --all -n foxhunt` to force a clean retry
(IAM rule is permanent so it'll succeed).
## P2 DRAFTED (2026-07-18) — fxhnt-ingress proxy (NOT applied yet; waiting on P1 cert READY)
- `infra/k8s/services/fxhnt-ingress-proxy.yaml` (committed ce1c978, draft): nginx (wildcard TLS via
fxhnt-wildcard-tls) + socat (git SSH→gitea-ssh.gitea:22) + tailscale sidecar (TS_HOSTNAME=fxhnt-ingress).
Fronts 3 fund vhosts BY CLUSTERIP: git→gitea-http.gitea:3000, dashboard/cockpit→fxhnt-dashboard.foxhunt:80,
dagster→dagster.foxhunt:80. Has its own NP (egress to those svcs + apiserver 172.16.0.11:6443 for TS + DERP).
Deps satisfied: tailscale-auth secret present; dagster NP ingress already allows fxhnt-ingress label.
## P1+P2 COMPLETE (2026-07-18) — bizworx fxhnt.ai front verified, NO DNS change yet ✅
- **P1 DONE:** fxhnt-wildcard cert READY (LE, SAN=*.fxhnt.ai + fxhnt.ai). IAM cross-project fix held.
- **P2 DONE:** fxhnt-ingress proxy applied + 3/3 Running. Tailnet identity = **`fxhnt-ingress` @ `100.93.141.11`**
(clean, no -1 suffix). THIS IS THE C4 DNS TARGET.
- **EVAL P2 PASS (from devbox over the real tailnet → 100.93.141.11, the exact browser path post-flip):**
dashboard / cockpit / dagster / git .fxhnt.ai ALL → **HTTP 200, tls=0 (cert chain valid)**. foxhunt still live.
- Fixes needed to get P2 green (all committed):
1. **RBAC:** TS sidecar needs SA+Role+RoleBinding for get/create/update/patch on its ts-state Secret (added
`tailscale-fxhnt-ingress` SA, mirrors tailscale-dagster). Without it: "missing get/update permission".
2. **dagster NP ingress was stale on-cluster:** the 9396f53 edit (foxhunt→fxhnt-ingress label) was committed
but the cluster still had `tailscale-gitlab-proxy` → dagster vhost hung. Re-applied dagster.yaml → fixed.
LESSON: `git commit` ≠ `kubectl apply`; always re-apply after editing a live manifest.
3. **gitea headless-service / CGNAT trap:** gitea-http & gitea-ssh are HEADLESS (ClusterIP:None) → DNS returns
the gitea POD IP in 100.64.0.0/15, which the proxy's KERNEL-MODE tailscale sidecar (TS_USERSPACE=false)
CANNOT reach (CGNAT overlap). FIX: added a non-headless **`gitea-clusterip`** Service (ns gitea, 10.32.x,
selector app.kubernetes.io/instance=gitea+name=gitea, ports 3000+22); nginx+socat now target it. cockpit &
dagster were fine because they have real 10.32.x ClusterIPs. (Kernel-mode TS reaches 10.32.x ClusterIPs but
NOT 100.64.x pod IPs — verified both ways.)
- **Verify artifact:** `curl --resolve <host>:443:100.93.141.11 https://<host>/` from any tailnet node.
- PHASE-1 PREP COMPLETE. Ready to schedule the C1C5 cutover window (see T8-CUTOVER-RUNBOOK.md). C4 DNS target
= 100.93.141.11 for records dashboard/cockpit/dagster/git. foxhunt fully live; delete NOTHING.
## T8 PREP (2026-07-18) — DNS/ingress map + fixes (foxhunt still fully live)
### Gotchas hit this session
- **dagster tailscale sidecar CrashLoop (context deadline exceeded reaching kubernetes.default.svc).** TWO-STAGE
root cause — the FIRST fix (06d12fd, `10.32.0.1/32:443`) was WRONG and did not hold (regressed to 2/3):
* `kubernetes.default.svc` → ClusterIP `10.32.0.1:443`, BUT **kube-proxy DNATs it to the REAL apiserver
endpoint `172.16.0.11:6443` BEFORE the NetworkPolicy is evaluated.** The NP sees dest `172.16.0.11:6443`,
which is inside the NP's OTHER except entry `172.16.0.0/16` → blocked. So allowing the ClusterIP never matched.
* CORRECT FIX: `ipBlock 172.16.0.11/32` port **6443** (the real endpoint, post-DNAT). apiserver endpoint is
`kubectl -n default get endpoints kubernetes` = 172.16.0.11 (Kapsule fixed control-plane).
* SIDE EFFECT this exposed: when the tailscale sidecar is NotReady, the whole dagster pod is NotReady →
**dagster Service has NO endpoints → dagster.fxhnt.ai vhost would 502**. So this MUST be green before C4 DNS flip.
* daemon+webserver themselves were always healthy. Cockpit's sidecar works only because cockpit has NO NP.
* GENERAL LESSON for any pod with a tailscale sidecar + a restrictive egress NP on bizworx: must allow
egress to 172.16.0.11:6443 (apiserver). The cronjob/rebalancer NPs do NOT have tailscale sidecars so they're
unaffected — but if any future fund pod gets a TS sidecar under an NP, add this rule.
- **foxhunt API ACL lapsed AGAIN**: devbox egress IP drifted `51.158.111.225`→`51.158.125.1`. Foxhunt Kapsule
admin API (`*.api.k8s.fr-par.scw.cloud:6443`) is IP-allowlisted; bizworx keeps working (tailscale-cgnat range).
FIX: `scw k8s acl add cluster-id=34a1e3c4… acls.0.ip=<newIP>/32 acls.0.description=bizworx-devbox-vN` (APPENDS;
verified all entries survive). Re-check `curl ifconfig.me` vs `scw k8s acl list` whenever foxhunt 6443 resets.
### THE fxhnt.ai DNS/ingress architecture (drives T8 DNS cutover)
- DNS zone = **Scaleway `fxhnt.ai`** (project c293eb98=foxhunt). ALL 13 A-records → **`100.95.225.27`** = tailnet
node **`foxhunt-gitlab`** = pod **`tailscale-gitlab-proxy` (ns foxhunt, 4/4)** = nginx(vhosts)+2×socat+tailscale.
**This pod DIES at foxhunt teardown (T12) → DNS MUST move off 100.95.225.27 regardless.**
- nginx vhost map (what each fxhnt.ai host fronts):
- `git.fxhnt.ai` → gitea-web.foxhunt:3000 ✅FUND → **bizworx gitea (ns gitea), already mirrored**
- `dashboard/cockpit.fxhnt.ai` → 100.81.150.18:80 (fxhnt-dashboard node→cockpit:8080) ✅FUND → **bizworx cockpit**
- `dagster.fxhnt.ai` → dagster.foxhunt:80 ✅FUND → **bizworx dagster**
- grafana ❌ (operator: "grafana can go too" — drop, bizworx has own obs) · api(50051 legacy Rust gRPC) ❌
· mail(stalwart)/chat(mattermost)/minio ❌legacy → all Phase-2 drop. ONLY git+dashboard/cockpit+dagster move.
- **CRITICAL SAFETY FACT:** the FUND PIPELINE never uses fxhnt.ai — git-sync=`gitea-ssh.gitea.svc.cluster.local:22`
(in-cluster), cockpit/dagster serve over ClusterIP+own tailnet IP. fxhnt.ai is BROWSER/human access only.
→ the fund keeps running through the DNS flip; DNS cutover only affects operator browser access. Low-risk.
- Tailnet identities: live foxhunt `fxhnt-dashboard`=100.81.150.18, `fxhnt-dagster`=100.116.123.29. New bizworx
cockpit joined as **`fxhnt-dashboard-1`=100.107.99.35** (`-1` because live one holds the name). Each bizworx pod
serves TLS DIRECTLY via `tailscale serve` on its own MagicDNS cert (verified: cockpit `tailscale serve status`
shows https://fxhnt-dashboard-1.tailff20e.ts.net → 127.0.0.1:8080). No bizworx proxy pod needed.
### DNS cutover DECISION (operator): "RENAME bizworx pods to take over the identity"
- At cutover: suspend/scale foxhunt fund pods → they RELEASE `fxhnt-dashboard`/`fxhnt-dagster` tailnet names →
restart bizworx pods so they re-register WITHOUT the `-1` suffix (reclaim the names + get their IPs) → then
repoint the Scaleway `fxhnt.ai` A-records (dashboard,cockpit,dagster,git) from the dying 100.95.225.27 to the
reclaimed bizworx pod tailnet IPs. Records to change: `scw dns record` set on zone fxhnt.ai.
- TLS RESOLVED: old proxy served a cert-manager wildcard `*.fxhnt.ai` (ClusterIssuer letsencrypt-prod, Scaleway
DNS-01 webhook). **bizworx already has the SAME stack** (cert-manager + scaleway-certmanager-webhook 1/1 +
scaleway-api creds + letsencrypt-prod ClusterIssuer, already issuing LE certs in ~10 ns incl grafana-tls).
→ issuing `*.fxhnt.ai` on bizworx = just a Certificate resource. NO blocker.
- CHOSEN CUTOVER SHAPE (supersedes the naive pod-IP flip): build a **stripped fund-only nginx+tailscale proxy on
bizworx** (`fxhnt-ingress`) fronting the 3 fund vhosts via ClusterIP (gitea-http.gitea:3000, fxhnt-dashboard
.foxhunt:80, dagster.foxhunt:80) + socat SSH gitea-ssh.gitea:22, terminating the wildcard `*.fxhnt.ai` cert.
Point the 4 fxhnt.ai A-records at THIS proxy's tailnet IP. Because it uses ClusterIP not pod tailnet names,
the "reclaim fxhnt-dashboard identity" rename is NO LONGER on the critical path (optional cosmetic).
- **FULL EXECUTABLE RUNBOOK written: `docs/superpowers/plans/T8-CUTOVER-RUNBOOK.md`** (P1 wildcard cert →
P2 build+verify proxy pre-flip → C1 freeze foxhunt → C2 delta dump/restore → C3 bizworx live+ib-gw sole session
→ C4 DNS flip → C5 arm ucits. Every step has VERIFY + ROLLBACK. Nothing destructive; foxhunt suspended-not-deleted.)
- Confirmed svc names for the proxy: gitea-http.gitea:3000 (headless), gitea-ssh.gitea:22 (headless),
fxhnt-dashboard.foxhunt:80, dagster.foxhunt:80, ib-gateway.foxhunt:4002/4004.
## T7 update 3 (2026-07-18) — cronjobs
- Cross-ns gitea-NP fix applied to ALL 15 jobs/fxhnt-*.yaml (namespaceSelector kubernetes.io/metadata.name=gitea).
- 7 CronJobs applied to bizworx ns foxhunt, ALL suspend=true (verified): bybit-rebalancer,
compare-measured-precompute, factory, ibkr-rebalancer, ucits-rebalancer, ucits-volume-ibkr, vrp-rebalancer.
(real-money safety: none run parallel to the live foxhunt ucits-rebalancer.)
- The 9 one-time Jobs (opra-backfill/warehouse-migrate/paper-backfill/bybit-*-job/deribit) NOT applied (they
run immediately); apply on demand post-cutover with the NP fix already in place.
- REMAINING before T8: one ucits dry-run (needs SOLE IBKR session — briefly scale foxhunt ib-gateway to 0,
or do it as the first act of cutover) + optional dagster asset-graph run. Then T8 cutover.
## T9 VERIFY (2026-07-18) — cockpit + gate intact against MIGRATED bizworx DB ✅
- **Local cockpit** (`./scripts/dev-cockpit-local.sh`, current-context=bizworx → serves the MIGRATED DB read-only
on http://127.0.0.1:8801). ENV GOTCHA fixed: the script's `uv run` needed `--extra web --extra pg` (uvicorn in
`web`, psycopg in `pg`); a bare `uv run` fails ModuleNotFoundError. Patched dev-cockpit-local.sh.
- **NUMBERS CROSS-CHECKED (read the numbers, per feedback_verify_cockpit_locally_not_prod):**
- /paper IBKR card = **$1,029,015** == ibkr_account_nav.nlv latest (2026-07-17 = 1029014.63). "not connected"
ABSENT (the card that silently degraded in a prior session renders the REAL DU account here).
- bybit Sh 2.11 == backtest_summary bybit_4edge 2.1067; unlock Sh 1.03 == 1.0258. All real, DB-backed.
- /, /strategy/multistrat, /strategy/xsfunding, /paper all HTTP 200 with real content.
- CAVEAT: Playwright MCP browser not connected this session → verified by rendering real HTML + reading
numbers (not pixels). Local cockpit left running on :8801 for the operator to click through.
- **T9b forward tracks INTACT (no spurious re-inception from migration):** forward_anchor t0 dates are ORIGINAL
historical inceptions (sixtyforty 2026-06-15, unlock 2026-06-20, positioning 2026-07-06, bybit_4edge/multistrat/
xsfunding 2026-07-07, vrp/multistrat_exec 2026-07-14) — NOT today (2026-07-18). definition_hash/version chains +
archived_at + record_mode all preserved. DB-anchor-SSOT survived byte-for-byte; recon gate will recompute correctly.
- **STILL PENDING:** T9 real-money confirm = watch the first armed fxhnt-ucits-rebalancer run at 09:00 UTC (tomorrow).
Then T10 observe → T11 triage → T12 teardown. ROTATE the pasted tailscale key. DELETE NOTHING.
## T9 VISUAL BROWSER PASS (2026-07-18) — real Chrome via Playwright MCP ✅
- Playwright MCP now available (plugins reloaded); installed the `chrome` channel (Google Chrome 150 at
/opt/google/chrome/chrome — the MCP hardcodes that path, bundled chromium alone isn't enough).
- Drove the LOCAL cockpit (http://127.0.0.1:8801, bizworx migrated DB) in a REAL browser — full render, all data:
- Overview: Bybit 4-edge $832,534 +738% Sh 1.74 maxDD -13%, deploy $75k/75%, gate 11/21 WAIT; 4 edges
(trend Sh0.72 / unlock Sh1.03 / xsfunding Sh3.62 / positioning Sh2.55); forward-tracks table w/ correct
gates: sixtyforty 23/20 GO, unlock 26/14 PASS, multistrat 9/20 WAIT + real IBKR subtrack "real -0.77% ·
9 trades · 5 holdings".
- /paper: IBKR paper account **$1,029,015 · 1 strategy trading · 5 holdings · 0.9% behind plan** (the real
DU account — NOT "not connected"). Bybit book $832,534.
- Full-page screenshot saved (.playwright-mcp/cockpit-overview-bizworx.png). 1 non-material console error.
- Matches DB cross-checks exactly. T9 cockpit verification COMPLETE (visual + numeric).
## FOXHUNT SHUTDOWN (2026-07-18 ~20:55 UTC) — compute stopped, cluster+PVCs INTACT (not destroyed)
Operator decision after verifying migration complete: "shut down everything now, we can spin up if we need to
reboot." Verified first: (1) bizworx has ZERO runtime dep on foxhunt (its postgres/gitea/git-sync/registry are
all bizworx-local — the `postgres.foxhunt.svc` refs are bizworx's OWN ns foxhunt, different cluster); (2) the
pre-cutover state is durably backed up OFF-cluster: `~/Work/fxhnt-cutover-backup/fxhnt_precutover_2026-07-18.dump`
(48M, 805 objects, all tables verified via pg_restore --list in the bizworx pod); (3) fund live+serving on
bizworx (cockpit 2/2, dagster 3/3, ingress 3/3, ib-gateway 1/1; dashboard.fxhnt.ai HTTP 200).
- Scaled ALL foxhunt-ns deploys + statefulsets to 0 (incl postgres/gitea/tailscale-gitlab-proxy = the
instant-rollback path — operator accepted dropping it). Also scaled cert-manager + tailscale-subnet-router to 0.
- Suspended ALL foxhunt cronjobs (fund + infra: postgres-backup/pvc-autoscaler/billing-scraper). Deleted leftover
Completed/Error Job pods so they don't pin nodes.
- Pool `platform` (8b520ebb…): autoscaling=true, **min-size=0** (was 1). Can't force size=0 while autoscaling on;
the Kapsule autoscaler drains nodes to 0 once only DaemonSets+kube-system remain (~10-15min scale-down delay).
Monitoring the 3 dev1_l nodes → 0. `ci-compile-cpu` pool already 0.
- **ROLLBACK now = restore the .dump into a fresh pg (few min), NOT an instant DNS flip.** foxhunt cluster + ALL
PVCs still EXIST (training-data 500Gi, minio 150Gi, postgres 100Gi, gitlab-gitaly 50Gi, gitea valkey, etc.) —
scaling back up reschedules pods on data intact. terraform destroy (T12, irreversible) still DEFERRED.
- REVERSIBILITY: `scw k8s pool update 8b520ebb… min-size=1 size=1` + scale deploys back up → foxhunt returns.
- STILL PENDING: tomorrow 09:00 UTC bizworx ucits-rebalancer run (real-money confirm); then T12 terraform destroy
(with operator confirm). ROTATE the pasted tailscale key.
## DATABENTO / .dbn DATA MIGRATION (2026-07-18 ~21:10 UTC) — operator caught this before teardown
Found databento .dbn data on foxhunt PVCs NOT covered by the Postgres migration. Split cleanly:
- **Fund-relevant (migrated):** `/data/surfer/*.dbn` — 22 futures .dbn (ES/NQ/GC/CL/NG/6E/ZC/ZS... universe,
updated Jul 12-13) + *_state.json. READ BY the fund's tsmom/carry RESEARCH dagster assets
(tsmom_research.py, carry_research.py, assets.py:247 read_front_arrays `$FXHNT_SURFER_DATA_DIR/<sym>.dbn`).
bizworx had only an EMPTY fxhnt-surfer-data PVC → research would've failed. **MIGRATED**: tar (33M gz) foxhunt
inspector → devbox → bizworx writer pod → extracted into fxhnt-surfer-data. Verified from the dagster pod's
OWN view: 22 .dbn at /data/surfer/*.dbn + state JSONs. (Layout note: PVC mounted at /data on dagster,
FXHNT_SURFER_DATA_DIR=/data/surfer → files live at PVC-root/surfer/.)
- **Archived Rust-ML (NOT fund-used): ~171GB .dbn** on training-data-pvc (59 files, 151GB: ES.FUT MBP10/trades
quarterly) + test-data-pvc (21 files, 20GB). Operator decision (option 3): push to PRIVATE Scaleway Object
Storage, NOT a bizworx PVC, and "cannot move 175GB over a tunnel" → **upload INTRA-SCALEWAY** (foxhunt pod →
s3.fr-par.scw.cloud, both fr-par, no tunnel/devbox hop).
- Created private bucket **`fxhnt-dbn-archive`** (fr-par, bizworx project 96d0717d, versioning off, ACL empty).
- S3 creds → foxhunt secret `dbn-archive-s3` (from scw config, never printed).
- `dbn-uploader` pod (rclone/rclone image) mounts training+test PVCs RO; `rclone copy --include "*.dbn*"
--transfers 8` → scw:fxhnt-dbn-archive/{training,test}/. Connectivity tested (rclone lsd rc=0). 171.4GB / 80 files.
- Upload running in background in-pod (logs /tmp/rclone-*.log, /tmp/rclone-done sentinel). Monitored.
- To restore archived data later: `rclone copy scw:fxhnt-dbn-archive/... <dest>` (creds in bizworx-prod-runtime
key or scw config). Kept for the archived Rust codebase per [[project_rust_infra_decommissioned_phase1]].
- NOTE: I bumped platform pool min-size back to 1 (from 0) so the inspector/uploader pods can schedule; after the
upload completes + surfer verified, set min-size=0 again to resume the drain. Cleanup pods: data-inspector, dbn-uploader.
## .dbn ARCHIVE — PIVOTED to block snapshots (rclone upload was a dead end) ✅
- The rclone S3 upload from the STARVED foxhunt cluster STALLED: dev1_l node egress + slow scw-bssd PVC read
(20MB/s) → 76KiB/s, ETA 4 DAYS, byte count frozen. Killed it.
- **RIGHT SOLUTION: Scaleway block-volume SNAPSHOTS** (instant, storage-layer, no tunnel, restorable, cheap).
The .dbn PVCs are scw block volumes → snapshotted directly:
- `fxhnt-training-dbn-archive-2026-07-18` (snap 564a4ccb-ee69-4514-8a64-9932222a0334, 537GB vol = the 151GB
training .dbn) — ParentVol cb327a05 (training-data-pvc).
- `fxhnt-test-dbn-archive-2026-07-18` (snap 5ba0049f-799a-4694-b7a0-0192cf0cda11, 50GB = the 20GB test .dbn)
— ParentVol dc144e62 (test-data-pvc).
- BOTH `available` in ~20s, in the **bizworx project 96d0717d** (survive foxhunt cluster teardown), zone fr-par-2.
- **RESTORE:** `scw block volume create from-snapshot-id=<snap> zone=fr-par-2` → attach to a pod → read .dbn.
- Cleaned up: dbn-uploader + data-inspector pods deleted, dbn-archive-s3 secret deleted.
- CORRECTION (2026-07-18 later): there is NO stray junk. The earlier "132MB / 23 objects" were the FIRST partial
`.dbn` transfers under `training/` — the resumed upload continues them (rclone skips already-present). Verified:
bucket holds ONLY legit `training/*.dbn*` (singletest/ + speedtest/ test files were deleted during testing = 0
objects). **DO NOT wipe/delete `fxhnt-dbn-archive` — it IS the active archive we're building.**
- Node drain RESUMED (platform pool min-size 1->0; no blocking pods). Autoscaler drops the last node to 0.
## STATE SUMMARY (2026-07-18 ~21:15 UTC)
- fxhnt.ai fund LIVE on bizworx (T8+T9 done, cockpit verified in real Chrome). foxhunt compute → 0 (draining).
- Data fully preserved: pg dump (~/Work/fxhnt-cutover-backup), surfer .dbn (migrated to bizworx PVC + verified
from dagster pod), archived 171GB .dbn (2 block snapshots in bizworx project). foxhunt PVCs still exist too.
- PENDING: tomorrow 09:00 UTC ucits run (real-money confirm); then T12 terraform destroy (operator confirm) —
which also removes the source PVCs, but snapshots + dump + bizworx-live cover everything. ROTATE tailscale key.
Cleanup the stray fxhnt-dbn-archive bucket.
## .dbn S3 COPY (2026-07-18 ~21:53 UTC) — operator: also copy compressed .dbn.zst to S3, KEEP snapshots
The .dbn.zst are already compressed → S3 copy is a good browsable archive alongside the snapshots (belt+braces).
- Measured real S3 throughput: ~10-17 MB/s from these small VMs (dev1_l/gp1_xs) — NOT compression-limited, it's
node network egress. Single-file test: 537MiB in 52s = 10MB/s, PVC read 217MB/s (fine). The earlier "0 B/s
stall" was rclone stats not updating until a big multipart completed + too many concurrent 64M-chunk streams
choking the node. FIX: --transfers 2 --s3-chunk-size 32M --s3-upload-concurrency 2 (node-friendly) → moving.
- Restore-snapshot-to-volume path abandoned (Kapsule CSI import of a raw scw volume = rabbit hole). Instead
uploading directly from a foxhunt pod (dbn-uploader) that mounts training+test PVCs RO — data still live there.
- Upload RUNNING in background: rclone copy .dbn* → scw:fxhnt-dbn-archive/{training,test}/. 137GiB, ETA ~6-7h at
~6MB/s. Persistent Monitor watching for completion. Sentinel /tmp/rclone-done in the pod.
- KEEP THE SNAPSHOTS (operator) — validated safety net; only after S3 upload verified (object count + sizes) do
we consider anything. Both archives coexist.
- foxhunt platform pool held at min-size=1 for the uploader; set min-size=0 to resume drain AFTER upload done.
- RESUME/CHECK: `kubectl --context $FOX -n foxhunt exec dbn-uploader -- sh -c 'cat /tmp/rclone-done; tail -1 /tmp/rclone-training.log'`
and `... rclone size scw:fxhnt-dbn-archive`. If pod died: re-create it (manifest in git history commit 2438294-ish) — rclone copy is resumable/idempotent.
## .dbn S3 COPY COMPLETE + VERIFIED (2026-07-19 ~02:41 UTC) ✅
- Upload finished: **80 objects / 171.393 GiB** in `fxhnt-dbn-archive` (private, fr-par, bizworx project).
- VERIFIED byte-for-byte vs source PVCs: `rclone check --size-only --one-way`:
- training: **0 differences, 59 matching files** ✅
- test: **0 differences, 21 matching files** ✅
- source .dbn* count = 80 = bucket object count = exact.
- **TWO independent verified archives now exist** (operator: keep snapshots): (1) S3 object storage
`s3://fxhnt-dbn-archive/{training,test}/*.dbn.zst` (browsable, per-file restore), (2) block snapshots
fxhnt-training/test-dbn-archive-2026-07-18 (564a4ccb / 5ba0049f, bizworx project). Both safe post-teardown.
- Cleaned up: dbn-uploader + s3-speedtest pods, dbn-archive-s3 secrets (both clusters). Node drain RESUMED
(platform pool min-size 1->0; no blocking pods) → autoscaler drops the last dev1_l to 0. foxhunt compute → ~$0.
- REMAINING (next session / operator): tomorrow 09:00 UTC first armed ucits run (real-money confirm); T12
terraform destroy foxhunt Kapsule (removes source PVCs — both .dbn archives + pg dump + bizworx-live cover it);
rotate the pasted tailscale key. Restore .dbn: `rclone copy scw:fxhnt-dbn-archive/... <dest>` OR
`scw block volume create from-snapshot.snapshot-id=<snap> perf-iops=5000 zone=fr-par-2`.
## SCHEDULING CONSISTENCY — ALL-DAGSTER K8sRunLauncher UPGRADE DONE (2026-07-19) ✅
Operator: "go all-Dagster properly + build the image right." Executed the 5-task plan
(docs/superpowers/plans/2026-07-19-production-scheduling-consistency.md). Result = consistent single-orchestrator
ETL, execution isolated in crons, nothing suspended-by-accident, VRP gone, documented.
- DIAGNOSIS (corrected): Dagster ETL was NEVER broken (23:30 ran SUCCESS, forward_nav fresh). The stale
bybit_sleeve_ret was the SUSPENDED compare-measured cron (exists because the compute needs 6-8Gi, OOM-killed
the 4Gi daemon). Fix = K8sRunLauncher, not naive daemon assets.
- BUILT: dagster-k8s dep added; image built via in-cluster **Kaniko** (devbox podman blocked by sandbox mount
perms) → `fxhnt-cockpit:dagster-k8s-25e3d7c`, verified + retagged `:latest` (skopeo Job).
- RBAC: tailscale-dagster SA can create Jobs/pods/events. K8sRunLauncher configured (launched Jobs git-sync
master, no PVC, fund secrets). New asset `bybit_book_precompute` (op_tags 6-8Gi) wraps `_persist_bybit_book`
(moved to application/bybit_book_persist.py). MERGED branch→master (git-sync pulls master).
- VERIFIED (the money shot): triggered combined_book_forward_job → launched Job dagster-run-* → materialized all
14 assets incl the heavy precompute → **daemon 0 restarts** → bybit_sleeve_ret + backtest_summary + forward_nav
ALL fresh 2026-07-19. Run SUCCESS 6m54s.
- CRONS (final): DELETED compare-measured (→ asset) + vrp (shelved). ARMED factory + ucits-rebalancer +
ucits-volume-ibkr. SUSPENDED bybit-rebalancer (no creds) + ibkr-rebalancer (PRIIPs). Manifests updated.
- Memory: [[reference_fxhnt_dagster_k8srunlauncher_scheduling]] created; VRP files + MEMORY.md updated.
- NOTE: the fund now runs off MASTER (28-commit consolidation branch merged 06156bd; +cron changes 1ecf398).