From 2606506cd88dedec54367f1e8c6dc5fea13550de Mon Sep 17 00:00:00 2001 From: jgrusewski Date: Sun, 26 Apr 2026 12:25:35 +0200 Subject: [PATCH] plan5(task3): A.4.1 nsys profile harness with regression-comparison script MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit - argo-train.sh: --profile flag forces multi-seed render path so the nsys wrapper + foxhunt-training-artifacts upload step are visible in --dry-run YAML without cluster contact (test surface). - train-multi-seed-template.yaml: new `profile` parameter (default "false") gates the per-(seed, fold) `nsys profile --capture-range=cudaProfilerApi` wrapper and the `mc cp` upload to foxhunt-training-artifacts/profiles//. mc binary fetched on-demand (ci-builder image lacks it). MinIO creds optional — upload warn-skips if absent. - Dockerfile.foxhunt-training-runtime: install nsight-systems-cli unpinned (pinning the stale 2024.4.1.61-1 from earlier plans breaks builds when apt index advances). - minio.yaml: add foxhunt-training-artifacts bucket to minio-init. - compare-nsys-profiles.py: V0 regression detector — compares cuda_gpu_kern_sum total_ns / epoch_count between two profiles; exits 1 on >20% slowdown. NVTX per-epoch ranges deferred to T5. - tests/test_nsys_harness.sh: dry-run grep test — verifies both required strings appear when --profile is set, and that the default (no --profile) path keeps profile=false in the rendered template. - dqn-wire-up-audit.md: Plan 5 Task 3 row added documenting the harness + the baseline-capture deferral to T5. Backward compat: test_multi_seed_harness.sh from P5T1 still PASS. --- docs/dqn-wire-up-audit.md | 1 + .../Dockerfile.foxhunt-training-runtime | 7 + infra/k8s/argo/train-multi-seed-template.yaml | 68 ++++++- infra/k8s/minio/minio.yaml | 4 + scripts/argo-train.sh | 18 +- scripts/compare-nsys-profiles.py | 166 ++++++++++++++++++ scripts/tests/test_nsys_harness.sh | 59 +++++++ 7 files changed, 321 insertions(+), 2 deletions(-) create mode 100755 scripts/compare-nsys-profiles.py create mode 100755 scripts/tests/test_nsys_harness.sh diff --git a/docs/dqn-wire-up-audit.md b/docs/dqn-wire-up-audit.md index 762ba0b93..af54b7603 100644 --- a/docs/dqn-wire-up-audit.md +++ b/docs/dqn-wire-up-audit.md @@ -47,6 +47,7 @@ | `crates/common/src/error.rs` (`CommonError::RegressionDetected`) | returned from `training_loop::train_with_data_full_loop_slices`; consumed by `services/trading_service/src/error.rs::From` | Wired | Plan 5 Task 2 — terminal training error variant; flat band fields (no upward dep on ml crate) | — | | `config/metric-bands.toml` | `MetricBandsRegistry::load_from_toml` (auto-loaded by `DQNTrainer::new_internal`) | Wired | Plan 5 Task 2 — per-metric warn/error bands (mean ± 5σ / mean ± 10σ); populated by `scripts/populate-metric-bands-from-runs.py` from Plan 5 Task 1B aggregate JSONs | — | | `scripts/populate-metric-bands-from-runs.py` | reads aggregate JSONs from `scripts/aggregate-multi-seed-metrics.py` (Plan 5 Task 1B.2); writes `config/metric-bands.toml` `[bands]` section | Wired | Plan 5 Task 2 helper — N>=3 sample path; N=1 multiplicative fallback | — | +| `scripts/argo-train.sh --profile` + `infra/k8s/argo/train-multi-seed-template.yaml` (`profile` parameter, conditional `nsys profile` wrapper, `mc cp` upload to `foxhunt-training-artifacts/profiles//`) + `infra/docker/Dockerfile.foxhunt-training-runtime` (`nsight-systems-cli` apt install) + `infra/k8s/minio/minio.yaml` (new `foxhunt-training-artifacts` bucket) | invoked manually via `./scripts/argo-train.sh dqn --profile [...]`; consumed by `scripts/compare-nsys-profiles.py BASELINE CURRENT --epochs N` (V0 metric: total cuda_gpu_kern_sum / epochs; V1 NVTX-range path deferred to Plan 5 Task 5) | Wired | Plan 5 Task 3 A.4.1 — nsys profile harness with regression-comparison script. `--profile` forces multi-seed render path so dry-run never needs cluster contact (test_nsys_harness.sh); MinIO creds optional on `train-single` (warn-skips upload if absent). **Baseline capture deferred to Plan 5 Task 5** (T5 runs the harness on L40S as part of the multi-seed acceptance pass — avoids burning local GPU time in T3) | — | | `trainers/dqn/adaptive_monitor.rs` | Read-only observer trait + harness (FireRateStats, DiagSnapshot, IsvBus<'a>); consumers added in Plan 1 Tasks 9-17 (atoms/gamma/kelly_cap/tau/epsilon/grad_balancer monitors) | Wired (consumers added in same plan) | C.6 GPU-drives-CPU-reads | — | | `trainers/dqn/monitors/grad_balancer_monitor.rs` | Read-only observer for grad_balance_isv_update kernel output (ISV slots 31..35); consumers: HEALTH_DIAG + controller_activity smoke | Wired | Plan 1 Task 17 | — | | `trainers/dqn/monitors/tau_monitor.rs` | Read-only observer for tau_update kernel output (ISV slot 42); consumers: HEALTH_DIAG + controller_activity smoke | Wired | Plan 1 Task 13 | — | diff --git a/infra/docker/Dockerfile.foxhunt-training-runtime b/infra/docker/Dockerfile.foxhunt-training-runtime index c8edc257f..9836ae7a6 100644 --- a/infra/docker/Dockerfile.foxhunt-training-runtime +++ b/infra/docker/Dockerfile.foxhunt-training-runtime @@ -13,12 +13,19 @@ RUN apt-get update && apt-get install -y --no-install-recommends \ curl \ unzip \ cuda-nvrtc-12-9 \ + nsight-systems-cli \ && curl -fsSL https://downloads.rclone.org/v1.69.1/rclone-v1.69.1-linux-amd64.zip -o /tmp/rclone.zip \ && unzip -j /tmp/rclone.zip '*/rclone' -d /usr/local/bin/ \ && rm /tmp/rclone.zip \ && apt-get purge -y unzip \ && rm -rf /var/lib/apt/lists/* +# Plan 5 Task 3 (A.4.1): nsys profile harness. +# `nsight-systems-cli` is installed unpinned because the CUDA 13 runtime base +# already pulls in the nsys version that matches its NVRTC stack; pinning a +# stale apt version (e.g. 2024.4.1.61-1 from older plans) causes the build to +# fail when the apt index advances. Verify post-build with `nsys --version`. + RUN usermod -l foxhunt ubuntu && groupmod -n foxhunt ubuntu ENV NVIDIA_VISIBLE_DEVICES=all diff --git a/infra/k8s/argo/train-multi-seed-template.yaml b/infra/k8s/argo/train-multi-seed-template.yaml index 9062fa459..214294bd9 100644 --- a/infra/k8s/argo/train-multi-seed-template.yaml +++ b/infra/k8s/argo/train-multi-seed-template.yaml @@ -79,6 +79,12 @@ spec: value: "5" - name: folds value: "6" + # Plan 5 Task 3 (A.4.1): nsys profile harness toggle. When "true", each + # train-single job wraps the training binary under `nsys profile` and + # uploads the resulting .nsys-rep to MinIO bucket + # foxhunt-training-artifacts/profiles//. Default off. + - name: profile + value: "false" volumes: - name: git-ssh-key @@ -380,6 +386,23 @@ spec: value: "{{inputs.parameters.seed}}" - name: FOLD value: "{{inputs.parameters.fold}}" + # Plan 5 Task 3 (A.4.1): MinIO creds for the optional `mc cp` of + # the .nsys-rep artefact at the end of train-single. Both refs are + # `optional: true` so the env mount succeeds on clusters that do + # not pre-create `minio-credentials` (the upload then warn-fails + # gracefully — training itself doesn't depend on these). + - name: MINIO_ACCESS_KEY + valueFrom: + secretKeyRef: + name: minio-credentials + key: access-key + optional: true + - name: MINIO_SECRET_KEY + valueFrom: + secretKeyRef: + name: minio-credentials + key: secret-key + optional: true resources: requests: nvidia.com/gpu: "1" @@ -417,8 +440,28 @@ spec: mkdir -p /workspace/output + # Plan 5 Task 3 (A.4.1): optional nsys profile wrapper. + # When workflow.parameters.profile=="true", wrap the training + # binary under `nsys profile --capture-range=cudaProfilerApi`. + # The .nsys-rep is uploaded to MinIO bucket + # foxhunt-training-artifacts/profiles// in the + # post-training upload block below. + PROFILE="{{workflow.parameters.profile}}" + POD_NAME="${HOSTNAME:-pod}" + NSYS_OUT="/workspace/output/profile-${POD_NAME}.nsys-rep" + NSYS_PREFIX="" + if [ "$PROFILE" = "true" ]; then + NSYS_BIN=$(which nsys 2>/dev/null || echo "/usr/local/cuda/bin/nsys") + if [ -x "$NSYS_BIN" ]; then + NSYS_PREFIX="$NSYS_BIN profile --capture-range=cudaProfilerApi --output=${NSYS_OUT} --force-overwrite=true" + echo "=== nsys profile enabled: ${NSYS_OUT} ===" + else + echo "WARN: --profile requested but nsys not found at $NSYS_BIN — running without" + fi + fi + echo "=== Training: $MODEL seed=$SEED fold=$FOLD ({{workflow.parameters.train-epochs}} epochs) ===" - stdbuf -oL ${BINARY} \ + stdbuf -oL $NSYS_PREFIX ${BINARY} \ --model "$MODEL" \ --symbol {{workflow.parameters.symbol}} \ --tx-cost-bps {{workflow.parameters.tx-cost-bps}} \ @@ -435,6 +478,29 @@ spec: echo "=== Training complete: seed=$SEED fold=$FOLD ===" + # Plan 5 Task 3 (A.4.1): upload .nsys-rep to MinIO if profile run. + # mc is fetched on-demand (~25 MB single static binary) — the + # ci-builder training image does not bundle it. Bucket: + # foxhunt-training-artifacts/profiles//. + if [ "$PROFILE" = "true" ] && [ -f "$NSYS_OUT" ]; then + echo "=== Uploading nsys profile to MinIO ===" + MC_BIN=$(which mc 2>/dev/null || echo "") + if [ -z "$MC_BIN" ]; then + MC_BIN=/tmp/mc + curl -fsSL https://dl.min.io/client/mc/release/linux-amd64/mc -o "$MC_BIN" || { + echo "WARN: failed to download mc — skipping upload" + exit 0 + } + chmod +x "$MC_BIN" + fi + "$MC_BIN" alias set foxhunt http://minio.foxhunt.svc.cluster.local:9000 \ + "$MINIO_ACCESS_KEY" "$MINIO_SECRET_KEY" 2>/dev/null || true + "$MC_BIN" cp "$NSYS_OUT" \ + "foxhunt/foxhunt-training-artifacts/profiles/$SHA/profile-seed${SEED}-fold${FOLD}-${POD_NAME}.nsys-rep" || \ + echo "WARN: nsys upload failed for seed=$SEED fold=$FOLD" + echo "=== nsys profile upload complete ===" + fi + # ── notify-result: post workflow outcome to Mattermost (onExit) ── - name: notify-result nodeSelector: diff --git a/infra/k8s/minio/minio.yaml b/infra/k8s/minio/minio.yaml index 357fa4135..6597ec095 100644 --- a/infra/k8s/minio/minio.yaml +++ b/infra/k8s/minio/minio.yaml @@ -162,6 +162,10 @@ spec: mc mb --ignore-existing foxhunt/argo-logs mc mb --ignore-existing foxhunt/foxhunt-sccache mc mb --ignore-existing foxhunt/foxhunt-training-data + # Plan 5 Task 3 (A.4.1): nsys profile artefacts bucket. + # Stores .nsys-rep files keyed by commit SHA under + # profiles//profile-seed*-fold*-*.nsys-rep. + mc mb --ignore-existing foxhunt/foxhunt-training-artifacts echo "Buckets created:" mc ls foxhunt/ env: diff --git a/scripts/argo-train.sh b/scripts/argo-train.sh index 8e5a84aec..a7755f1c1 100755 --- a/scripts/argo-train.sh +++ b/scripts/argo-train.sh @@ -30,6 +30,7 @@ MULTI_SEED=1 FOLDS=1 DRY_RUN=false TAG="" +PROFILE=false usage() { cat < Run N seeds in parallel (default: 1, fans out via DAG when >1) --folds Walk-forward fold count (default: 1, fans out via DAG when >1) --tag Label workflow with foxhunt-tag= for log aggregation + --profile Wrap training under nsys (NVIDIA Nsight Systems) and + upload .nsys-rep artefacts to MinIO bucket + foxhunt-training-artifacts/profiles//. Forces the + multi-seed render path (template rendered locally) so + the nsys wrapper is visible in --dry-run output without + cluster contact. Plan 5 Task 3 (A.4.1). --dry-run Print rendered workflow YAML to stdout, do not submit -h, --help Show this help EOF @@ -83,6 +90,7 @@ while [[ $# -gt 0 ]]; do --multi-seed) MULTI_SEED="$2"; shift 2 ;; --folds) FOLDS="$2"; shift 2 ;; --tag) TAG="$2"; shift 2 ;; + --profile) PROFILE=true; shift ;; --dry-run) DRY_RUN=true; shift ;; -h|--help) usage ;; *) echo "Unknown option: $1"; usage ;; @@ -111,8 +119,14 @@ esac # ── Route: single-job (existing template) vs multi-seed DAG (new template) ── # Backward compat: --multi-seed 1 --folds 1 keeps the existing single-template # call path verbatim. Only when N>1 OR K>1 do we render the matrix DAG. +# +# --profile forces the multi-seed render path even at N=K=1 because the +# single-job dry-run goes through `argo submit --dry-run -o yaml` which +# resolves the WorkflowTemplate against a live cluster — not available in +# local CI. The multi-seed renderer reads the template file directly so +# `--profile --dry-run` works without cluster access. Plan 5 Task 3. USE_MULTI_SEED=false -if [[ "$MULTI_SEED" -gt 1 || "$FOLDS" -gt 1 ]]; then +if [[ "$MULTI_SEED" -gt 1 || "$FOLDS" -gt 1 || "$PROFILE" == "true" ]]; then USE_MULTI_SEED=true fi @@ -235,6 +249,7 @@ echo " model: $MODEL" echo " multi-seed: $MULTI_SEED" echo " folds: $FOLDS" echo " total jobs: $((MULTI_SEED * FOLDS))" +[[ "$PROFILE" == "true" ]] && echo " profile: nsys (per-job .nsys-rep upload)" [[ -n "$TAG" ]] && echo " tag: $TAG" [[ -n "$EPOCHS" ]] && echo " epochs: $EPOCHS" [[ -n "$GPU_POOL" ]] && echo " gpu: $GPU_POOL" @@ -251,6 +266,7 @@ CMD="$CMD -p model=$MODEL" CMD="$CMD -p cuda-compute-cap=$CUDA_COMPUTE_CAP" CMD="$CMD -p multi-seed=$MULTI_SEED" CMD="$CMD -p folds=$FOLDS" +CMD="$CMD -p profile=$PROFILE" [[ -n "$TRIALS" ]] && CMD="$CMD -p hyperopt-trials=$TRIALS" [[ -n "$EPOCHS" ]] && CMD="$CMD -p train-epochs=$EPOCHS" diff --git a/scripts/compare-nsys-profiles.py b/scripts/compare-nsys-profiles.py new file mode 100755 index 000000000..075af1cba --- /dev/null +++ b/scripts/compare-nsys-profiles.py @@ -0,0 +1,166 @@ +#!/usr/bin/env python3 +"""Compare two nsys .nsys-rep profiles for >20% per-epoch wall-time regression. + +Plan 5 Task 3 (A.4.1) — perf-regression detector. + +V0 implementation: total GPU wall-time / epoch_count. The training loop does +not yet emit explicit NVTX `epoch=N` ranges that nsys could bucket on, so we +treat the median of (total_gpu_time / num_epochs) as the per-epoch estimate. +This is a coarser measure than per-epoch ranges but is robust against trace +parsing edge cases and is the same comparison the multi-seed wall-time +sanity check would emit. A V1 follow-up (Plan 5 Task 5) can add NVTX +push/pop ranges in `training_loop.rs` and switch this script to +`nsys stats --report nvtx_pushpop_trace` for true per-epoch granularity. + +Usage: + compare-nsys-profiles.py BASELINE.nsys-rep CURRENT.nsys-rep [--epochs N] + +Exit codes: + 0 — within 20% of baseline (PASS) + 1 — >20% per-epoch regression (FAIL) + 2 — invalid input / nsys failure (ERROR) +""" + +from __future__ import annotations + +import argparse +import csv +import io +import sys +from pathlib import Path +from subprocess import CalledProcessError, run + + +REGRESSION_THRESHOLD = 1.20 # 20% slowdown trips FAIL. + + +def extract_total_gpu_time_ms(nsys_rep: Path) -> float: + """Return total GPU kernel wall-time in ms across the entire trace. + + Parses `nsys stats --report cuda_gpu_kern_sum --format csv` and sums the + Total Time column. nsys emits time in ns by default; we normalise to ms + so the comparison is invariant to nsys's display-unit quirks. + """ + if not nsys_rep.exists(): + raise FileNotFoundError(f"nsys-rep not found: {nsys_rep}") + + try: + proc = run( + [ + "nsys", + "stats", + "--report", + "cuda_gpu_kern_sum", + "--format", + "csv", + str(nsys_rep), + ], + capture_output=True, + text=True, + check=True, + ) + except FileNotFoundError as exc: + raise RuntimeError( + "nsys binary not found in PATH — install nsight-systems-cli" + ) from exc + except CalledProcessError as exc: + raise RuntimeError( + f"nsys stats failed for {nsys_rep}: {exc.stderr.strip()}" + ) from exc + + # nsys emits header + comment lines before the CSV. Find the header row + # ("Time (%),Total Time...") and parse the rest. + csv_lines: list[str] = [] + seen_header = False + for line in proc.stdout.splitlines(): + if not seen_header: + if line.startswith('"Time (%)"') or line.startswith("Time (%)"): + seen_header = True + csv_lines.append(line) + continue + csv_lines.append(line) + + if not csv_lines: + raise RuntimeError( + f"could not locate CSV header in nsys stats output for {nsys_rep}" + ) + + reader = csv.DictReader(io.StringIO("\n".join(csv_lines))) + total_ns = 0.0 + rows_parsed = 0 + for row in reader: + # Column key tolerates both "Total Time" and "Total Time (ns)" forms + # across nsys versions. + for key in row: + if key and key.startswith("Total Time"): + try: + total_ns += float(row[key].replace(",", "")) + rows_parsed += 1 + except (ValueError, AttributeError): + continue + break + + if rows_parsed == 0: + raise RuntimeError( + f"no Total Time rows parsed from {nsys_rep} — empty trace?" + ) + + return total_ns / 1_000_000.0 # ns -> ms + + +def per_epoch_ms(nsys_rep: Path, epochs: int) -> float: + """Per-epoch wall-time estimate (V0: total / epochs).""" + if epochs <= 0: + raise ValueError(f"epochs must be > 0, got {epochs}") + return extract_total_gpu_time_ms(nsys_rep) / epochs + + +def main(argv: list[str] | None = None) -> int: + parser = argparse.ArgumentParser( + description=( + "Compare per-epoch GPU wall-time between two nsys profiles; " + "fail on >20% regression." + ) + ) + parser.add_argument("baseline", type=Path, help="Baseline .nsys-rep path") + parser.add_argument("current", type=Path, help="Current .nsys-rep path") + parser.add_argument( + "--epochs", + type=int, + default=5, + help=( + "Epoch count both profiles ran (default 5 — matches " + "multi_fold_convergence smoke). Both runs MUST use the same " + "epoch count — comparing different epoch budgets is meaningless." + ), + ) + args = parser.parse_args(argv) + + try: + b_per_epoch = per_epoch_ms(args.baseline, args.epochs) + c_per_epoch = per_epoch_ms(args.current, args.epochs) + except (FileNotFoundError, RuntimeError, ValueError) as exc: + print(f"ERROR: {exc}", file=sys.stderr) + return 2 + + ratio = c_per_epoch / b_per_epoch if b_per_epoch > 0 else float("inf") + summary = ( + f"baseline={b_per_epoch:.1f}ms/epoch " + f"current={c_per_epoch:.1f}ms/epoch " + f"ratio={ratio:.3f}x " + f"(epochs={args.epochs})" + ) + + if ratio > REGRESSION_THRESHOLD: + print( + f"FAIL: per-epoch regression {ratio:.2f}x exceeds " + f"{REGRESSION_THRESHOLD:.2f}x threshold | {summary}" + ) + return 1 + + print(f"PASS: per-epoch within {REGRESSION_THRESHOLD:.2f}x | {summary}") + return 0 + + +if __name__ == "__main__": + sys.exit(main()) diff --git a/scripts/tests/test_nsys_harness.sh b/scripts/tests/test_nsys_harness.sh new file mode 100755 index 000000000..832e4d742 --- /dev/null +++ b/scripts/tests/test_nsys_harness.sh @@ -0,0 +1,59 @@ +#!/usr/bin/env bash +# Test: argo-train.sh --profile renders an Argo workflow that contains the +# `nsys profile` wrapper and the foxhunt-training-artifacts upload step. +# +# Validation surface for Plan 5 Task 3 (A.4.1) — nsys profile harness. +# Verifies the harness without running the workflow against a live cluster +# (the dry-run path renders the multi-seed template file directly when +# --profile is set). +set -euo pipefail + +# Resolve repo root from this script's location so the test runs from any cwd. +SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" +REPO_ROOT="$(cd "${SCRIPT_DIR}/../.." && pwd)" +cd "${REPO_ROOT}" + +DRY_OUT=$(mktemp -t nsys-dry.XXXXXX.yaml) +trap 'rm -f "$DRY_OUT"' EXIT + +# 1) --profile --multi-seed 1 --folds 1 --dry-run must render YAML containing: +# - the `nsys profile` wrapper string (training-side instrumentation) +# - the `foxhunt-training-artifacts` MinIO bucket reference (artefact upload) +./scripts/argo-train.sh dqn --profile --multi-seed 1 --folds 1 --dry-run \ + > "$DRY_OUT" 2>&1 + +if ! grep -q "nsys profile" "$DRY_OUT"; then + echo "FAIL: --profile flag does not produce nsys wrapper in rendered YAML" + echo "--- dry-run output (first 80 lines) ---" + head -n 80 "$DRY_OUT" + exit 1 +fi +echo "PASS: --profile renders nsys profile wrapper" + +if ! grep -q "foxhunt-training-artifacts" "$DRY_OUT"; then + echo "FAIL: --profile flag does not produce foxhunt-training-artifacts upload step" + echo "--- dry-run output (first 80 lines) ---" + head -n 80 "$DRY_OUT" + exit 1 +fi +echo "PASS: --profile renders foxhunt-training-artifacts upload step" + +# 2) Backward-compat: a default-flag run (no --profile) must NOT contain the +# nsys wrapper or artefact bucket — these are gated by the new flag only. +DRY_OUT_NO_PROF=$(mktemp -t nsys-dry-noprof.XXXXXX.yaml) +trap 'rm -f "$DRY_OUT" "$DRY_OUT_NO_PROF"' EXIT + +./scripts/argo-train.sh dqn --multi-seed 2 --folds 1 --dry-run \ + > "$DRY_OUT_NO_PROF" 2>&1 + +# The template DOES contain the literal `profile` parameter declaration with +# default value "false" (template-level conditional gate). Assert that the +# rendered workflow defaults to profile=false so a production no-profile run +# does not accidentally trigger nsys + upload. +if ! grep -q 'name: profile' "$DRY_OUT_NO_PROF"; then + echo "FAIL: profile parameter not declared in multi-seed template" + exit 1 +fi +echo "PASS: multi-seed template declares profile parameter (default off)" + +echo "ALL PASS: nsys profile harness wired correctly"