Runs nvidia-smi on GPU node in parallel with fetch-binary, triggering
H100 autoscale during compilation so the node is ready when hyperopt
starts. Exits immediately to free GPU resources.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
The foxhunt-runtime image runs as non-root user 'foxhunt' and uses
/bin/sh (dash), not bash. Two issues:
1. &>/dev/null is bash-only — use >/dev/null 2>&1 for POSIX sh
2. Fallback download to /tmp (writable), not /usr/local/bin (root-only)
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
The build-ci-image template used volumeClaimTemplates (PVC) which
weren't inherited when called via templateRef from ci-pipeline.
Restructured to single pod: init container (alpine/git) clones repo
into emptyDir, main container (kaniko) builds and pushes image from
the same volume. No PVC needed, works correctly via templateRef.
Tested: foxhunt-runtime image rebuilt successfully with kubectl.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
build-ci-image used volumeClaimTemplates (PVC) shared between a
git-clone pod and a kaniko-build pod in a DAG. When called via
templateRef from ci-pipeline, the VCT wasn't inherited, causing
"volume 'workspace' not found" errors.
Fixed by merging into a single pod: init container (git clone) +
main container (kaniko build) sharing an emptyDir volume. No PVC
needed, works correctly when called via templateRef.
Also removed stale compile-training-template.yaml reference from
kustomization (file doesn't exist, template is inline in ci-pipeline).
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Add kubectl v1.31.4 to foxhunt-runtime Dockerfile so deploy step
doesn't need to download it each run. Deploy step falls back to
curl download if kubectl not found (for current image version).
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- Deploy step: use foxhunt-runtime image (cached on platform node) +
curl kubectl, replacing bitnami/kubectl:1.31 which doesn't exist.
- Sensor: add podGC:OnPodCompletion and ttlStrategy to workflow spec
(not inherited from WorkflowTemplate via workflowTemplateRef).
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Three issues prevented services from being deployed after push:
1. No deploy step: pipeline compiled + uploaded binaries to MinIO but
never restarted service deployments. Added deploy-services step
(bitnami/kubectl) that runs rollout restart after compile-services.
2. compile-services couldn't schedule: targeted platform node (4 CPU,
6Gi allocatable) but requested 6Gi memory. Moved to ci-compile-cpu
(POP2-32C-128G) alongside compile-training — both fit simultaneously
(28/32 CPU, 48/128Gi).
3. No podGC: completed Argo pods held resources indefinitely, blocking
subsequent pipeline steps. Added podGC:OnPodCompletion.
Also: detect-changes now triggers service rebuild when CI pipeline
template or k8s service manifests change.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
DaemonSet init container now:
- Detects GPU nodes (99-nvidia.toml in conf.d)
- Creates v3-compatible registry config drop-in
- Restarts containerd via chroot if config files changed
- Idempotent: skips restart if files already exist
Requires hostPID + privileged for chroot /proc/1/root.
On new nodes, first run creates files + restarts containerd
(kills pod), second run finds files unchanged (no restart).
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- Hyperopt: catch early stopping errors and return penalty metrics
instead of aborting entire run. Trials that stop early are scored
as poor (objective=-500) so optimizer avoids those configs.
- DaemonSet: detect GPU nodes (NVIDIA conf.d overlay) and create
v3-compatible registry config drop-in for containerd v2.1+.
Old grpc.v1.cri.registry path is silently ignored by containerd v2.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Kapsule nodes use nameserver 127.0.53.53 which resolves external DNS
but returns NXDOMAIN for .svc.cluster.local. Containerd uses node DNS
for image pulls, causing ErrImagePull for cluster-internal registry refs.
Deploy node-dns-fix DaemonSet (kube-system) that:
1. Points /etc/resolv.conf to kube-dns ClusterIP (10.32.0.10)
2. Creates containerd certs.d config for HTTP GitLab registry
3. Watches for Kapsule node reconciliation resets (every 30s)
4. Restores original resolv.conf on SIGTERM
Also bump compile-services memory: 6Gi→10Gi limit (was OOM-killed at 5.7GB).
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- Add api_uptime_seconds gauge to API service (5s update interval),
matching the pattern used by trading/backtesting/ml-training services
- Remove obsolete Web Gateway Uptime and Active WebSocket Connections
panels from cockpit (web-gateway is not deployed)
- Add API Gateway Uptime (api_uptime_seconds) and API Auth Rate
(api_auth_requests_total) panels in their place
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Join node metrics with node_uname_info to resolve instance IPs to
hostnames (e.g. scw-foxhunt-platform-694db...) in all 10 node-level
panels: CPU, Memory, Disk I/O, Network I/O, Load Average, Disk Usage.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Add __meta_kubernetes_pod_phase=Running filter to training-pods job.
Completed/failed Argo workflow step pods were showing as DOWN targets
because their containers are no longer running to serve /metrics.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- Replace api_gateway_ with api_ prefix to match actual metric exports (cockpit, observability)
- Change grpc_code!="OK" to grpc_code!="0" — gRPC uses numeric status codes (cockpit, trading)
- Update gRPC service filter to actual exported_service labels: trading-agent-service,
broker-gateway, backtesting-service, ml-training-service, data-acquisition-service (trading)
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- Remove duplicate annotated-pods and gitlab-annotated-pods scrape jobs
(all targets already covered by dedicated jobs)
- Drop Argo controller from foxhunt-services job (dedicated HTTPS job exists)
- Resolve datasource variable placeholders to actual UIDs in dashboard JSON
- Add Ingress rule to Prometheus netpol allowing Grafana on port 9090
Targets: 49 → 18 (zero duplicates)
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Rewrite batch_hierarchical_softmax_actions to keep all sampling on GPU:
reshape [N,45]→[N,5,9], max-pool→[N,5] exposure Q, Gumbel-max over
exposures, parallel Gumbel-max over all sub-actions, gather chosen
exposure's sub-action. Only CPU transfer: N flat indices.
Add imagePullPolicy: Always to compile-services and compile-training
containers so rebuilt builder images are never stale-cached.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
grpc_pass only handles native gRPC (HTTP/2 frames). The web dashboard
uses gRPC-web via Connect protocol (HTTP/1.1 POST requests). Switch
api.fxhnt.ai nginx block to proxy_pass so tonic_web::GrpcWebLayer on
the API service can handle the protocol conversion.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
TypeScript 5.9 erasableSyntaxOnly rejects enum declarations in .ts
files (they emit runtime code). Switch protoc-gen-es from target=ts
to target=js+dts which generates .js runtime + .d.ts type declarations.
Also fix type mismatches in useApi.ts (OrderSide/OrderType enum types,
StartBacktestRequest field mapping) and MLDashboard signal defaulting.
Add web-dashboard network policy for MinIO egress + proxy ingress.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
containerd on Kapsule nodes uses host DNS which can't resolve
.svc.cluster.local names. Switch all image references from
gitlab-registry.foxhunt.svc.cluster.local:5000 to localhost:30500
(NodePort on the GitLab registry). Also make training runtime image
configurable via TRAINING_RUNTIME_IMAGE env var.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- nginx pod serves Vite-built static assets from MinIO S3
- initContainer fetches dist/ from s3://foxhunt-binaries/web-dashboard/
- SPA routing via try_files, aggressive caching for hashed assets
- dashboard.fxhnt.ai proxied to web-dashboard:80 via tailscale nginx
- CI pipeline: build-web-dashboard task (npm ci + build + rclone upload)
- detect-changes: needs-dashboard output for web-dashboard/ paths
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Split the combined grafana+dashboard nginx block so dashboard.fxhnt.ai
proxies to web-gateway:3000 (the actual trading dashboard) instead of
Grafana.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Training data PVCs (ReadWriteOnce) don't work with ephemeral GPU nodes
that autoscale across availability zones. Replace with rclone download
from MinIO to /tmp at the start of each GPU step.
- Remove training-data-pvc parameter and volume mount
- Add rclone download from s3:foxhunt-training-data/futures-baseline/
to /tmp/futures-baseline/ in hyperopt, train-best, and evaluate steps
- Add MINIO_ACCESS_KEY/SECRET_KEY env vars to all GPU steps
- Add cuda-compute-cap parameter (default "90")
- Change data-dir default to /tmp/futures-baseline
- 36 .dbn.zst files (53MB total) uploaded to MinIO bucket
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Rename ci-training → ci-training-l40s, add H100/H100x2/H100-SXM/H100-SXM-8
pool resources (all autoscaling 0→1). Remove dev pool (DevPod on platform).
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
build-ci-image templateRef needs registry-auth volume defined in parent
workflow. Secret was also named wrong (scw-registry-credentials →
scw-registry) and needed key projection (.dockerconfigjson → config.json).
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
detect-changes was falling back to Cargo.toml (rebuild-all) because jq
is not installed in foxhunt-runtime:latest. Replaced with grep+sed
extraction that needs no external dependencies. Also added jq to the
Dockerfile for future use.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
T=0.1 was too conservative for degenerate configs — trials with uniform
Q-values still collapsed to 1 unique action. Making temperature a
hyperopt parameter lets TPE learn the optimal exploration-exploitation
balance per config. Range [0.01, 2.0] log-scale.
Also sets training-workflow default gpu-pool to ci-training-h100.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Two bugs fixed:
1. jq parsed webhook as plain array but payload is {body:{commits:[...]}} —
fell back to Cargo.toml catch-all, triggering full recompile every push.
2. Single needs-compile flag gave no granularity — now split into
needs-services and needs-training with path-based detection.
crates/ml/ only → training + services (ml-training-service depends on ml)
services/ only → services only, skip training (saves ~5.5min CUDA compile)
infra/ only → skip both compiles entirely
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
The sensor passes GitLab commit data containing single quotes in
commit messages (e.g. Merge branch 'feature-name'). When embedded
directly in the shell script via Argo template substitution, this
breaks the shell syntax. Pass as a container env var instead.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
When using templateRef, Argo only borrows the template definition — not
spec-level volumes or volumeClaimTemplates from the source WorkflowTemplate.
Add workspace VCT, sccache PVC, and git-ssh-key secret at the ci-pipeline
spec level so compile-training's sub-templates can access them.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- Replace simulated login in fxt CLI with real AuthServiceClient gRPC call
- Add auth proto to fxt build.rs compile list and lib.rs proto module
- Add api.access permission to JWT claims issued by AuthGrpcService
- Remove orphaned ci-sensor.yaml (replaced by cluster-applied version)
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
argo-workflow SA needs workflows/create and workflowtemplates/get
permissions to submit CI pipelines from sensor triggers.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- Remove unnecessary case map — service names already match Cargo package names
- Add sensor pod serviceAccountName: argo-workflow (RBAC fix)
- Add sensor config to events/ for version tracking
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- Replace compile-api-gateway + compile-web-gateway with single compile-api task
- Update detect-changes script: services/api/ instead of services/api_gateway/
- Remove web-gateway from change detection and compile targets
- Fix templateRef parameter resolution (inputs.parameters instead of workflow.parameters)
- Add SSH key setup and cluster-internal GitLab host for compile-service
- Use ci-builder image (has ssh) instead of ci-builder-cpu for service compilation
- Add compile-training and build-ci-image templates to kustomization
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Add Training cockpit sub-tabs (Live/History) with expandable per-epoch
detail panel showing financial metrics, RL diagnostics, and gradient health.
History tab fetches completed/failed runs from ListTrainingJobs RPC.
Add scrollable pod table to Pods cockpit and compact pods summary to
Overview bottom row. Fix K8s API egress in api-gateway NetworkPolicy
(monitoring-service→prometheus, apply existing K8s API CIDR rules).
Add 62 new unit tests across event_loop (51) and data_fetcher (11)
covering all StateUpdate variants, ring buffer eviction, epoch history
accumulation, pod scroll helpers, and navigation edge cases.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Add nginx basic auth (htpasswd) to ci.fxhnt.ai proxy block and
switch Argo to server auth mode. Login: jgrusewski / Welcome01.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Rename DNS record from argo→ci and update nginx proxy server_name
to match. Applied via terragrunt and kubectl.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Switch Argo server from --auth-mode=server (no auth) to
--auth-mode=client (K8s ServiceAccount token). Created
jgrusewski-argo-token secret bound to argo-workflows-server SA
with admin RBAC for full workflow management.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Hyperopt requires trials > n_initial. When running quick smoke tests
with few trials, auto-reduce n_initial to trials-1 so workflows don't
fail with "trials must be greater than n_initial".
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
The GPU runners used runtime_class_name=nvidia for CUDA access but
didn't request nvidia.com/gpu as a K8s resource, allowing multiple
training pods to share a GPU without K8s awareness. Add pod_spec
patches to request GPU resources properly so K8s enforces mutual
exclusion between GitLab CI and Argo training jobs.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
The default-deny-all NetworkPolicy blocks all egress for foxhunt-labeled
pods. Argo workflow pods need:
- Egress to MinIO (9000) for binary fetch and result upload
- Egress to K8s API (443/6443) for executor task result reporting
- Egress to Tempo (4317) for OTLP traces
- DNS already covered by allow-dns policy
Also fix PVC permission denied by setting fsGroup: 1000 to match the
foxhunt user (UID 1000) in the runtime images.
Smoke test fetch-binary step: Succeeded.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- Helm values (controller + server on platform node, MinIO artifact repo)
- WorkflowTemplate: parameterized 5-step DAG (fetch→hyperopt→train→eval→upload)
- Nginx proxy for argo.fxhnt.ai → Argo Server :2746
- DNS A record for argo.fxhnt.ai
- MinIO bucket foxhunt-training-results for Argo artifacts
- Kustomization for kubectl apply -k
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Rename 7 service binaries from snake_case to kebab-case to match
K8s deployment names. Update Cargo.toml package/bin names, K8s
manifest S3 paths and commands, and cross-crate dependency keys.
- api_gateway → api-gateway
- trading_service → trading-service
- broker_gateway_service → broker-gateway
- ml_training_service → ml-training-service
- backtesting_service → backtesting-service
- trading_agent_service → trading-agent-service
- data_acquisition_service → data-acquisition-service
broker-gateway gets an explicit [lib] name = "broker_gateway_service"
since its new package name maps to broker_gateway (not the original
broker_gateway_service used in source code). All other services map
correctly with Rust's automatic hyphen-to-underscore conversion.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
On Ampere+ GPUs the training dtype is BF16. Network forward passes
with BF16 inputs return BF16 tensors, but all loss-path arithmetic
(rewards, dones, gamma, PER weights, Huber constants, atoms, clip
epsilon) was created as F32 from Tensor::from_vec. Candle forbids
mixed-dtype binary ops, causing "dtype mismatch in {mul,sub,add}"
on every training step.
DQN fixes (7 mismatch sites):
- Cast current_q_values to F32 immediately after forward pass
- Cast all target network outputs (standard, dueling, C51, IQN) to F32
- Keep rewards/dones/gamma/weights/atoms/half/delta as F32 (remove
.to_dtype(training_dtype) casts — now use DType::F32 sentinel)
- Entropy regularization and CQL paths now match (F32 + F32)
PPO fixes (3 mismatch sites):
- Use DType::F32 for clip epsilon one_tensor (was training_dtype BF16)
- Cast critic.forward() output to F32 in both MLP and LSTM value loss
- Cast target_returns to F32 for symlog/normalization path
Infra: revert runner-h100 from SXM2 (zero Scaleway quota) back to
ci-training-h100 PCIe pool.
2704 ml lib tests pass, 260 PPO tests pass, 0 clippy warnings.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>