jgrusewski
|
2c8967ad96
|
feat: compute-sanitizer support in Argo training workflow
Usage: ./scripts/argo-train.sh dqn --baseline --epochs 2 --sanitizer memcheck
./scripts/argo-train.sh dqn --baseline --epochs 1 --sanitizer synccheck
Tools: memcheck (OOB, uninitialized), racecheck (data races),
synccheck (__syncthreads divergence/deadlocks).
10-100x slower — use with 1-2 epochs for debugging.
Detects exact kernel + line causing GPU hang.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
2026-04-17 12:24:23 +02:00 |
|
jgrusewski
|
445c2197f2
|
feat: unified Argo train workflow — replace 7 templates with 1
New: train-template.yaml with smart caching:
- ensure-binary: checks /data/bin/{sha}/ on PVC, compiles only on cache miss
- ensure-fxcache: runs precompute only if cache doesn't exist
- gpu-warmup: parallel autoscale during compile
- hyperopt → train-best → evaluate → upload-results
Deleted 7 redundant templates:
- compile-and-train-template.yaml
- train-dqn-template.yaml
- train-baseline-rl-template.yaml
- train-supervised-template.yaml
- training-workflow-template.yaml
- precompute-features-template.yaml
- train-ppo-template.yaml
Deleted: scripts/argo-precompute.sh (absorbed into ensure-fxcache step)
Rewritten: scripts/argo-train.sh (single template, --sha for commit pinning)
Updated: kustomization.yaml
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
2026-04-10 20:36:03 +02:00 |
|
jgrusewski
|
6ba52425ea
|
feat(infra): Argo workflow templates, drop cuDNN, GPU hotpath fixes
- Add compile-and-deploy, train-dqn/ppo/supervised WorkflowTemplates
- Add Argo Events (EventSource, Sensor, Service) for webhook triggers
- Add NetworkPolicy for compile-and-deploy pods (MinIO/DNS/API egress)
- Add convenience scripts: argo-compile-deploy.sh, argo-train.sh
- Drop cuDNN feature flags from all 9 ML crates (zero conv ops in codebase)
- Switch training runtime base to nvidia/cuda:12.9.1-runtime (saves ~800MB)
- Delete unused selective_scan.cu (16KB, zero Rust callers)
- Fix GPU hotpath violations in ml-core (NVTX, gradient utils, capabilities)
- Fix clippy warnings in ml-dqn (VarMap backticks, const fn)
- Add DQN GPU smoketest, backtest evaluator signal adapter fixes
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
|
2026-03-12 01:44:03 +01:00 |
|