The nodeSelector pointed to non-existent 'ci-training' pool. The actual
Scaleway pool is 'ci-training-h100', matching the CI pipeline template.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Three bugs in the GPU experience collection hot path:
1. gpu_batch_to_experiences() had hardcoded state_dim=43 but the CUDA kernel
outputs states at the ALIGNED dimension (56 with OFI, 48 without). After
sample 0, every replay buffer entry had corrupted state vectors — the
network was learning from garbage data.
2. GPU path never called monitor.track_reward(), so mean_reward was always
reported as 0.0 in epoch logs despite the agent generating real rewards.
3. Action tracking was double-counted (direct array write + track_action_by_exposure),
inflating diversity metrics by 2x. Consolidated into single bounded call.
Also adds missing app.kubernetes.io/component label to job-template.yaml
so Prometheus training-pods scrape job discovers training pods.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
GPU experience collector was falling back to CPU because it required a
curiosity VarMap, but curiosity_weight=0.0 means the module is never
created. Fix: make curiosity optional (CuriosityWeightSet::zeros() for
GPU buffers, curiosity_scale=0.0 in kernel config). Also require
explicit binary-tag SHA in Argo training workflow (no "latest" fallback).
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Add QuestDB ILP sink for training metrics, update Prometheus scrape
configs, and fix network policies for monitoring stack connectivity.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- New compile-and-train-template.yaml: compile + GPU warmup in parallel,
then fetch-binary → hyperopt → train-best → evaluate → upload-results
- Compile step outputs SHA tag via Argo output parameter, fetch-binary
uses it directly (no 'latest' package indirection)
- Remove trading-service from CI deploy step — must be explicitly enabled
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Root cause: preload_data() called loader.ofi_features.take() on the internal
DQN trainer, but load_training_data() only loads OHLCV bars — it never
populates ofi_features. The OFI loading is done by the hyperopt adapter's
own load_ofi_features() method.
Fixes:
- preload_data() now calls self.load_ofi_features() directly
- load_ofi_features() uses self.mbp10_data_dir when set (was hardcoded ../mbp10)
- input_dim pre-computation uses OFI-aware size (51 when enabled, 43 otherwise)
- CI 'latest' package: delete-then-upload to avoid GitLab duplicate file issue
- Warn when mbp10_data_dir is set but no OFI features loaded
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- Rename node-dns-fix → node-bootstrap, fix nvidia gate race condition
(always create conf.d/ and write 50-registry.toml, don't gate on
99-nvidia.toml which doesn't exist on fresh autoscaled GPU nodes)
- Update busybox 1.36 → 1.37
- Fix fetch-binary: use PRIVATE-TOKEN/gitlab-pat (not DEPLOY-TOKEN)
- Fix upload-results: use gitlab-pat (gitlab-ci-token didn't exist)
- Add 'latest' rolling package version in both compile-services and
compile-training (re-upload after CalVer upload)
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Service pods' initContainers need egress to gitlab-webservice-default:8181
to fetch release binaries from the Generic Package Registry. Without this,
curl hangs for 2+ minutes then times out (default-deny-all blocks it).
Covers all app.kubernetes.io/part-of=foxhunt pods in one policy.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Argo Workflows can't resolve {{tasks.X.outputs}} inside template
bodies — only in DAG task arguments. Pass tag via arguments →
inputs.parameters in compile-services, compile-training,
upload-release, and deploy-services templates.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Replace MinIO binary distribution with GitLab Generic Package Registry.
Every code push to main auto-creates a CalVer tag (vYYYY.MM.N),
compiles, uploads binaries to GitLab packages, creates a Release
with auto-generated notes, and deploys via deployment patching.
- New CI templates: create-tag, upload-release
- Modified: compile-services/training upload to GitLab packages
- Modified: deploy-services patches FOXHUNT_RELEASE on deployments
- All 7 service initContainers fetch from GitLab (curl, deploy token)
- Training job-template binary fetch from GitLab (data stays MinIO)
- MinIO retains: sccache, training data, model checkpoints
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
The training-data PVC already has all data (OHLCV, MBP-10, trades).
Mount it read-only at /data in hyperopt/train/evaluate steps instead
of rcloning ~50GB from MinIO every step.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- CI pipeline: add gpu-warmup step parallel to compile-training,
triggers GPU node autoscale so it's ready when training starts
- Training workflow: add mbp10-data-dir and trades-data-dir parameters,
download MBP-10 + trades data in all steps (hyperopt, train, evaluate)
- Use parameterized paths consistently (no hardcoded /tmp paths)
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Add --mbp10-data-dir and --trades-data-dir CLI args to
hyperopt_baseline_rl binary so hyperopt trials can use real
order book and trade data for VPIN/Kyle's Lambda features.
- DQNTrainer: add mbp10_data_dir/trades_data_dir fields + with_ofi_data_dirs() builder
- DQNHyperparameters: pipe through from trainer instead of hardcoded None
- download-trades-job: fix nodeSelector to ci-compile-cpu (platform pool full)
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- Expand training-data-pvc from 10Gi → 200Gi (OHLCV + MBP-10 + headroom)
- Add data-sync-job: rclone sync from MinIO → PVC (delta-only, fast repeats)
- Add sync-training-data init container to training job template
(auto-syncs data from MinIO before each training run)
- Add training-job + data-sync-job NetworkPolicies (MinIO, Tempo, Pushgateway)
- Add app.kubernetes.io/part-of: foxhunt to training pod template
- download_baseline: add --parallel N flag for concurrent Databento downloads
- download_baseline: delete local files after MinIO upload (saves scratch space)
- Bump download job scratch volume to 100Gi
Architecture: Databento → MinIO (source of truth) → PVC (local cache).
First sync pulls everything; subsequent runs only sync deltas (seconds).
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
The download job pod was missing app.kubernetes.io/part-of: foxhunt,
so the default-deny-all egress policy blocked it and MinIO's ingress
policy rejected it. Also removed hostname pinning (was a misdiagnosis).
Added data-download-job NetworkPolicy allowing egress to MinIO:9000
and external HTTPS:443 (Databento API).
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- download_baseline now reads schema from universe TOML config
(was hardcoded to ohlcv-1m, now supports mbp-10 and other schemas)
- Add --rclone-dest flag for direct upload to MinIO after each download
- Add config/universe-es-mbp10.toml: ES.FUT MBP-10, same date range
- Add infra/k8s/training/download-mbp10-job.yaml: K8s Job to download
ES.FUT MBP-10 data in-cluster and upload to MinIO
Usage:
kubectl apply -f infra/k8s/training/download-mbp10-job.yaml
Prereqs: databento-credentials secret, download_baseline binary in MinIO
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- Remove all hardcoded Y-axis min/max bounds (41 removals) — dynamic scaling
- Eliminate Hyperopt section: merge into Training Status + Run Summary
- Deduplicate trial progress (3 panels → 1 gauge + 1 stat)
- Combine ML Jobs + Errors into single panel
- Add Current Trial and Hyperopt Elapsed to status bar
- Add Best Objective Δ (deriv) trend indicator
- Smooth line interpolation + gradient fills on all timeseries
- Stacked area for Action Distribution, gradient fill for Replay Buffer
- Shared crosshair tooltips, consistent threshold colors
- Status bar: stat panels with sparklines, trial before epoch
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
nvidia-smi is driver-mounted by the GPU operator, which may not be
ready when the warmup pod starts on a fresh autoscaled node. The
warmup's purpose is just triggering autoscale, not GPU validation.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Previous fix removed ALL trailing backslashes from --s3-no-check-bucket,
but 3 of 6 occurrences need the continuation for --transfers=8 on the
next line. Restores \ on lines where --transfers follows.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Same bug as service manifests — \\ after --s3-no-check-bucket caused
chmod to be parsed as rclone args. Fixed in 6 places.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Both Dockerfiles already install mold v2.35.1 but the registry images
are stale builds without it. This change triggers rebuild-ci-builder
and rebuild-ci-builder-cpu pipeline steps via detect-changes.
.cargo/config.toml uses -fuse-ld=mold — without mold in the image,
linking falls back to the system default (slower).
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Kubernetes defaults to Always for :latest tags, forcing registry
round-trips that fail on fresh GPU nodes where containerd HTTP-only
registry config has a race condition with HTTPS fallback.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
The Argo Events sensor had no nodeSelector and randomly landed on the
H100 GPU node, preventing autoscale-down. Pin it to DEV1-L to avoid
wasting expensive GPU node hours on a lightweight event listener.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Multi-window backtest: splits validation data into 3 non-overlapping
windows and aggregates with mean(Sharpe) - 0.5*std(Sharpe), penalizing
inconsistency and reducing overfit to a single data segment.
Top-K ensemble: hyperopt now emits top_k_params (top 5 trials) in JSON
output. train_baseline_rl gains --ensemble-top-k flag to train multiple
models per fold from different hyperopt configs, saving checkpoints as
dqn_ensemble_{k}_fold_{n}.safetensors.
Workflow template: adds ensemble-top-k parameter (default 5) and passes
--ensemble-top-k to the train-best step.
2720 tests pass, 0 clippy warnings.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Add clamp_max() to PromQL queries and hard Y-axis limits to prevent
early-epoch degenerate values from blowing up panel scaling (Sharpe
showing 12 instead of 1.3, Profit Factor at 30k).
Run Summary: clamp_max on Best Val Loss (5), Sharpe (5), Profit Factor (20)
Training Quality: clamp_max on Epoch Sharpe (5), Sortino (10), PF (20)
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Replace stat panels with timeseries using smooth line interpolation,
gradient fill, and multi-tooltip. Layout changed from 6x1 to 3x2 grid.
Legends show model/fold only when multiple series exist.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Add 6 stat panels showing overall training run metrics: Best Val Loss,
Best Sharpe, Best Win Rate, Min Max Drawdown, Total Epochs, and Best
Profit Factor. Placed between Training Status and Training Curves rows.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Runs nvidia-smi on GPU node in parallel with fetch-binary, triggering
H100 autoscale during compilation so the node is ready when hyperopt
starts. Exits immediately to free GPU resources.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
The foxhunt-runtime image runs as non-root user 'foxhunt' and uses
/bin/sh (dash), not bash. Two issues:
1. &>/dev/null is bash-only — use >/dev/null 2>&1 for POSIX sh
2. Fallback download to /tmp (writable), not /usr/local/bin (root-only)
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
The build-ci-image template used volumeClaimTemplates (PVC) which
weren't inherited when called via templateRef from ci-pipeline.
Restructured to single pod: init container (alpine/git) clones repo
into emptyDir, main container (kaniko) builds and pushes image from
the same volume. No PVC needed, works correctly via templateRef.
Tested: foxhunt-runtime image rebuilt successfully with kubectl.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
build-ci-image used volumeClaimTemplates (PVC) shared between a
git-clone pod and a kaniko-build pod in a DAG. When called via
templateRef from ci-pipeline, the VCT wasn't inherited, causing
"volume 'workspace' not found" errors.
Fixed by merging into a single pod: init container (git clone) +
main container (kaniko build) sharing an emptyDir volume. No PVC
needed, works correctly when called via templateRef.
Also removed stale compile-training-template.yaml reference from
kustomization (file doesn't exist, template is inline in ci-pipeline).
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Add kubectl v1.31.4 to foxhunt-runtime Dockerfile so deploy step
doesn't need to download it each run. Deploy step falls back to
curl download if kubectl not found (for current image version).
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- Deploy step: use foxhunt-runtime image (cached on platform node) +
curl kubectl, replacing bitnami/kubectl:1.31 which doesn't exist.
- Sensor: add podGC:OnPodCompletion and ttlStrategy to workflow spec
(not inherited from WorkflowTemplate via workflowTemplateRef).
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Three issues prevented services from being deployed after push:
1. No deploy step: pipeline compiled + uploaded binaries to MinIO but
never restarted service deployments. Added deploy-services step
(bitnami/kubectl) that runs rollout restart after compile-services.
2. compile-services couldn't schedule: targeted platform node (4 CPU,
6Gi allocatable) but requested 6Gi memory. Moved to ci-compile-cpu
(POP2-32C-128G) alongside compile-training — both fit simultaneously
(28/32 CPU, 48/128Gi).
3. No podGC: completed Argo pods held resources indefinitely, blocking
subsequent pipeline steps. Added podGC:OnPodCompletion.
Also: detect-changes now triggers service rebuild when CI pipeline
template or k8s service manifests change.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
DaemonSet init container now:
- Detects GPU nodes (99-nvidia.toml in conf.d)
- Creates v3-compatible registry config drop-in
- Restarts containerd via chroot if config files changed
- Idempotent: skips restart if files already exist
Requires hostPID + privileged for chroot /proc/1/root.
On new nodes, first run creates files + restarts containerd
(kills pod), second run finds files unchanged (no restart).
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- Hyperopt: catch early stopping errors and return penalty metrics
instead of aborting entire run. Trials that stop early are scored
as poor (objective=-500) so optimizer avoids those configs.
- DaemonSet: detect GPU nodes (NVIDIA conf.d overlay) and create
v3-compatible registry config drop-in for containerd v2.1+.
Old grpc.v1.cri.registry path is silently ignored by containerd v2.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Kapsule nodes use nameserver 127.0.53.53 which resolves external DNS
but returns NXDOMAIN for .svc.cluster.local. Containerd uses node DNS
for image pulls, causing ErrImagePull for cluster-internal registry refs.
Deploy node-dns-fix DaemonSet (kube-system) that:
1. Points /etc/resolv.conf to kube-dns ClusterIP (10.32.0.10)
2. Creates containerd certs.d config for HTTP GitLab registry
3. Watches for Kapsule node reconciliation resets (every 30s)
4. Restores original resolv.conf on SIGTERM
Also bump compile-services memory: 6Gi→10Gi limit (was OOM-killed at 5.7GB).
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- Add api_uptime_seconds gauge to API service (5s update interval),
matching the pattern used by trading/backtesting/ml-training services
- Remove obsolete Web Gateway Uptime and Active WebSocket Connections
panels from cockpit (web-gateway is not deployed)
- Add API Gateway Uptime (api_uptime_seconds) and API Auth Rate
(api_auth_requests_total) panels in their place
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Join node metrics with node_uname_info to resolve instance IPs to
hostnames (e.g. scw-foxhunt-platform-694db...) in all 10 node-level
panels: CPU, Memory, Disk I/O, Network I/O, Load Average, Disk Usage.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Add __meta_kubernetes_pod_phase=Running filter to training-pods job.
Completed/failed Argo workflow step pods were showing as DOWN targets
because their containers are no longer running to serve /metrics.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- Replace api_gateway_ with api_ prefix to match actual metric exports (cockpit, observability)
- Change grpc_code!="OK" to grpc_code!="0" — gRPC uses numeric status codes (cockpit, trading)
- Update gRPC service filter to actual exported_service labels: trading-agent-service,
broker-gateway, backtesting-service, ml-training-service, data-acquisition-service (trading)
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- Remove duplicate annotated-pods and gitlab-annotated-pods scrape jobs
(all targets already covered by dedicated jobs)
- Drop Argo controller from foxhunt-services job (dedicated HTTPS job exists)
- Resolve datasource variable placeholders to actual UIDs in dashboard JSON
- Add Ingress rule to Prometheus netpol allowing Grafana on port 9090
Targets: 49 → 18 (zero duplicates)
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Rewrite batch_hierarchical_softmax_actions to keep all sampling on GPU:
reshape [N,45]→[N,5,9], max-pool→[N,5] exposure Q, Gumbel-max over
exposures, parallel Gumbel-max over all sub-actions, gather chosen
exposure's sub-action. Only CPU transfer: N flat indices.
Add imagePullPolicy: Always to compile-services and compile-training
containers so rebuilt builder images are never stale-cached.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
grpc_pass only handles native gRPC (HTTP/2 frames). The web dashboard
uses gRPC-web via Connect protocol (HTTP/1.1 POST requests). Switch
api.fxhnt.ai nginx block to proxy_pass so tonic_web::GrpcWebLayer on
the API service can handle the protocol conversion.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>