Commit Graph

3046 Commits

Author SHA1 Message Date
jgrusewski
3cfa99de3f feat(ppo): pre-upload rollout states to GPU, eliminate per-step tensor allocs
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-28 02:48:33 +01:00
jgrusewski
3883284219 feat(ml): add cuda_pipeline module with GPU data pre-upload for DQN/PPO
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-28 02:45:48 +01:00
jgrusewski
43754bf5d5 docs: numeric type standardization implementation plan
12-worktree phased migration: foundation types → core crates →
ML boundary → services → peripheral → cleanup. Full code for
Phase 1 (Price/Quantity/Money/Ratio in common/financial_types),
migration patterns for Phases 2-6.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-28 02:37:53 +01:00
jgrusewski
de3a807670 docs: numeric type standardization design (i64/6dp financial types)
Consolidate competing Price/Quantity implementations into canonical
i64 fixed-point types in common/. Adds Money and Ratio newtypes.
Phased migration plan across 5 phases using parallel worktrees.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-28 02:32:46 +01:00
jgrusewski
52630a77d3 perf: eliminate heap-alloc Decimal→float casts across 19 files (36 instances)
Replace all `.to_string().parse::<f32/f64>()` patterns with
`num_traits::ToPrimitive` methods (`.to_f32()`, `.to_f64()`).
Each string roundtrip heap-allocated per conversion — fatal in
DQN hot loop (300K+ bars × epochs). Decimal stays as canonical
financial type; conversions happen at GPU/float boundaries only.

Also fixes blocking_read() in async context (risk_integration.rs).

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-28 02:29:04 +01:00
jgrusewski
35ac61b68d chore(ml): downgrade P2-C reversal warns to debug, delete stale patch
P2-C reversal/cash warns fire thousands of times per hyperopt trial
during backtest simulation. Also removes stale Kelly patch file (already
applied).

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-28 01:13:37 +01:00
jgrusewski
5579efe7b1 chore(ml): downgrade portfolio-feature warns to debug
These fire every training step when portfolio isn't populated yet —
extremely noisy during hyperopt with parallel trials.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-28 01:10:36 +01:00
jgrusewski
21a8e3410f fix(hyperopt): require GPU for RL hyperopt, propagate device to trainers
The previous "auto-detect" logic forced CPU which was wrong — GPU is
faster for forward/backward even with parallel trials (DQN/PPO use
<350MB of 24GB VRAM across 7 threads).

Changes:
- Remove --device flag, always require CUDA GPU
- Add DQNTrainer::new_with_device() to share CUDA context across trials
- Propagate hyperopt device to internal DQN trainer (was ignoring it)
- PPO/DQN adapters error on missing GPU instead of silent CPU fallback
- Downgrade batch-size clamping from warn to debug

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-28 01:08:46 +01:00
jgrusewski
623d4b52ae fix(docker): correct web-gateway binary path after CI compile split
The compile-services stage outputs binaries to build-out/services/ but
the web-gateway Dockerfile still referenced the old build-out/ path.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-28 00:59:33 +01:00
jgrusewski
9c503f002a feat(infra): scale CPU compile pool to max 4 nodes
POP2-32C-128G pool now allows up to 4 concurrent nodes for parallel
CI compilation (services + training + builder image builds).

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-28 00:43:08 +01:00
jgrusewski
fc48395ad1 Merge branch 'worktree-training-deploy' (CPU pipeline split) 2026-02-28 00:23:20 +01:00
jgrusewski
829d05e870 feat(ci): split compile pipeline into CPU services and CUDA training
- Add CI_BUILDER_CPU_IMAGE variable and build-ci-builder-cpu prepare job
- Add .rust-base-cpu template (no CUDA_COMPUTE_CAP)
- Split compile-services into compile-services (CPU) and compile-training (CUDA)
- Split .kaniko-base into .kaniko-service-base and .kaniko-training-base
- Service builds use build-out/services/, training uses build-out/training/
- Each compile job has conditional changes: rules for its source paths

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-28 00:21:23 +01:00
jgrusewski
c328e3eb3a fix(ci): make infra-plan dependency optional for API/web pipelines
infra-apply needs infra-plan, but infra-plan only runs when infra/**
changes. API/web-triggered pipelines fail to create because the
dependency doesn't exist. Adding optional: true allows infra-apply
to run without infra-plan for manual triggers.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-28 00:20:08 +01:00
jgrusewski
d4bd3d9465 feat(ci): add CUDA-free CI builder image for service compilation
Based on rust:1.89-slim-bookworm (~2-3GB vs ~8GB CUDA devel).
Same toolchain: mold 2.35, protoc 28.3, sccache 0.10, clang, lld.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-28 00:16:28 +01:00
jgrusewski
590883408a feat(hyperopt): auto-detect CPU for parallel RL hyperopt, use all cores
DQN/PPO networks are tiny (3 layers × 128 neurons). Running parallel
hyperopt on GPU wastes cores because CUDA context serializes across
threads — 5 trials on L4 only used 2000m of 6000m requested CPU.

Changes:
- Add --device flag to hyperopt_baseline_rl (auto/cpu/cuda)
- Auto mode forces CPU for parallel runs (no CUDA contention)
- CPU mode uses all available cores (no 2-core reserve)
- Add with_device() builder to DQN/PPO hyperopt trainers
- Downgrade "portfolio value <= 0" and GPU utilization warnings

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-28 00:11:58 +01:00
jgrusewski
979061c80c fix: downgrade remaining DQN training warn noise to debug/info
Portfolio value <= 0 during early exploration is expected behavior,
not a warning. GPU memory utilization advisory is informational.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-27 23:48:19 +01:00
jgrusewski
e67f828f3c fix(infra): use Kapsule built-in pool-name label for node selection
Kapsule doesn't allow node-labels in kubelet_args. Instead, use the
auto-assigned k8s.scaleway.com/pool-name label that every node gets.

- Revert kubelet_args from Terraform (not supported)
- Switch runner node_selector: pool→k8s.scaleway.com/pool-name
- Update CI pipeline KUBERNETES_NODE_SELECTOR to match
- Helm upgraded both runners (CPU + GPU)

No more manual kubectl label after node scale-up.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-27 23:37:43 +01:00
jgrusewski
a9f70b5145 fix(infra): add kubelet node-labels to all Kapsule pools
Nodes from autoscaler now get pool=<name> labels automatically via
kubelet_args, eliminating the need to manually label nodes after
scale-up events.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-27 23:30:14 +01:00
jgrusewski
8e997be2f7 Merge branch 'worktree-training-deploy' 2026-02-27 23:16:39 +01:00
jgrusewski
290d5b29a0 chore: remove docs/plans/ from .gitignore
Plans are part of the project record and should be tracked.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-27 23:09:35 +01:00
jgrusewski
35769ae72f feat(serving): wire approve/reject promotion to PromotionManager
Replace stub approve_promotion and reject_promotion gRPC handlers with
real implementations that call PromotionManager.approve() and .reject().

- approve_promotion: removes from pending, promotes to active model map
- reject_promotion: removes from pending with operator-supplied reason
- Input validation: empty model_id returns INVALID_ARGUMENT
- Error handling: nonexistent model_id returns success=false with message
- Add 4 tests covering approve, reject, and not-found error paths
- Add #[cfg(test)] active_models_for_test() accessor on PromotionManager

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-27 23:05:01 +01:00
jgrusewski
b762fecff8 fix: downgrade DQN warn noise to debug, fix clippy warnings, default dev-release
- DQN portfolio_tracker: Phase 1/2 complete logs warn→debug (noisy per-epoch)
- risk: lazy_static→LazyLock, .map→.inspect, midpoint overflow fixes
- risk: remove unused POSITION_PROCESSING_LATENCY, lazy_static dep
- CI: default DEV_RELEASE=true for fast iteration (thin LTO, 24% faster)

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-27 23:04:37 +01:00
jgrusewski
1b4bfb9d5b feat(serving): wire list_pending_promotions to PromotionManager
Replace the stub list_pending_promotions gRPC handler with a real
implementation that queries PromotionManager::list_pending() and maps
each PendingModel to the proto PendingPromotion message. Adds a unit
test verifying the field mapping from domain type to proto type.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-27 23:02:02 +01:00
jgrusewski
8324b98f72 feat(serving): wire report_job_completion to PromotionManager
Replace the stub report_job_completion gRPC handler with a real
implementation that:
- Validates job_id as UUID upfront
- On failure: updates DB status to Failed (best-effort), returns "failed"
- On success: updates DB status to Completed, looks up model_type and
  symbol from child_jobs table, registers with PromotionManager, and
  returns the actual promotion status string

Also adds JobSpawner::get_job_by_id() to fetch individual child jobs
from PostgreSQL, and two new tests covering promotion status mapping
and the PromotionManager integration path.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-27 22:56:39 +01:00
jgrusewski
b4a5b8d235 fix(infra): fix training-output-pvc storage class for Kapsule
Changed from non-existent scw-bssd-nfs to scw-bssd (RWO).
Training output is single-pod, doesn't need ReadWriteMany.
PVC deployed to foxhunt namespace (WaitForFirstConsumer).

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-27 22:52:15 +01:00
jgrusewski
034e8c4a91 fix(db): expand child_jobs model_type constraint to all 10 models
The original migration 046 only allowed DQN, PPO, MAMBA-2, TFT, TLOB.
Now includes TGGN, LIQUID, KAN, XLSTM, DIFFUSION.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-27 22:51:58 +01:00
jgrusewski
abc5ee6af1 feat(training): add in-process queue consumer for K8s job dispatch
Background tokio task polls JobSpawner.get_next_pending_job() every 5s,
marks found jobs as Running, builds TrainingJobParams, and dispatches
to K8s via K8sDispatcher. On dispatch failure the job is marked Failed.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-27 22:45:57 +01:00
jgrusewski
f3e485c2a1 feat(ml_training_service): wire start_training handler to JobSpawner
Add JobSpawner as the highest-priority dispatch path in start_training.
When job_spawner is available, training requests are persisted to
PostgreSQL via spawn_batch() and a batch ID is returned immediately.
The queue consumer (Task 3) will later poll for pending jobs and
dispatch them to K8s.

Priority order: JobSpawner (DB) → K8sDispatcher (direct) → orchestrator (in-process).

Also adds test_start_training_model_binary_mapping unit test.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-27 22:37:46 +01:00
jgrusewski
a7d18f4f7c fix(ci): revert RL training back to L4 pool
RL (DQN/PPO) stays on L4-1-24G — CPU-bound, low VRAM (<1GB).
Supervised stays on L40S-1-48G — GPU-bound, high VRAM.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-27 22:32:58 +01:00
jgrusewski
b160afe038 feat(ci): route all training to L40S pool (unify GPU resources)
All training jobs (RL + supervised) now use ci-training pool (L40S-1-48G)
instead of splitting RL→L4, supervised→L40S. With parallel hyperopt,
RL benefits from the L40S's more powerful GPU and same 8 vCPU.

Simplifies infrastructure: can potentially decommission L4 pool.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-27 22:31:33 +01:00
jgrusewski
3817b06f19 feat(ml_training_service): add JobSpawner field to MLTrainingServiceImpl
Wire JobSpawner into the gRPC service struct so that training jobs can
be persisted to PostgreSQL before being dispatched to K8s. This is the
first step toward a durable job queue that survives pod restarts.

- Add `job_spawner: Option<Arc<JobSpawner>>` to MLTrainingServiceImpl
- Extend `new()` constructor to accept the spawner parameter
- Add `DatabaseManager::pg_pool()` accessor for cheap PgPool cloning
- Construct JobSpawner in main.rs and pass `Some(job_spawner)` to service
- Add compile-time test verifying the field exists

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-27 22:29:10 +01:00
jgrusewski
2538d52324 docs: add training pipeline & model deployment implementation plan
11 tasks across 2 independent tracks:
- Track 1 (Tasks 1-5): Wire fxt train → gRPC → JobSpawner → K8sDispatcher → GPU Jobs
- Track 2 (Tasks 6-9): Wire promotion RPCs to real PromotionManager
- Shared (Tasks 10-11): Docker build + end-to-end validation

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-27 22:13:29 +01:00
jgrusewski
539f6c2727 fix(ci): increase RL training CPU request for full L4 utilization
RL hyperopt runs parallel PSO trials across CPU cores. Previous 2000m
request underutilized the L4's 8 vCPUs. Bumped to 6000m/7800m so
hyperopt_baseline_rl --parallel 0 can saturate all cores.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-27 22:11:58 +01:00
jgrusewski
93115c1fc5 docs: add training pipeline & model deployment design
Two parallel tracks:
- Track 1: fxt CLI -> gRPC -> JobSpawner (PostgreSQL) -> K8sDispatcher -> GPU Jobs
- Track 2: S3 checkpoints -> model loader -> promotion manager -> serving

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-27 22:08:26 +01:00
jgrusewski
b351259a6c fix(ci): add low resource requests to kaniko builds for concurrent scheduling
Runner defaults (16 CPU, 32Gi) are tuned for compile jobs but too high
for Kaniko image packaging. Only 2 build pods fit on the 32-core node,
queuing the remaining 7 with FailedScheduling: Insufficient cpu.

Set kaniko builds to 1000m/1Gi — all 9 can now run concurrently.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-27 22:06:09 +01:00
jgrusewski
f336ed01b4 fix(infra): update CPU compile pool to POP2-32C-128G (GP1-L has no quota)
GP1-L and PRO2-L both have 0/0 quota in fr-par-2. POP2-HC-32C-64G had
insufficient root disk (20GB). POP2-32C-128G with 100GB SBS root works.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-27 21:36:13 +01:00
jgrusewski
a5d9cec2cd feat(infra): add POP2-HC-32C-64G CPU compile pool for 4x faster builds
Split runners into CPU (compile/build/deploy) and GPU (training):
- CPU runner: tags kapsule,rust,docker → ci-compile-cpu (32 vCPU, 64GB)
- GPU runner: tags kapsule,gpu → ci-compile (L4) or ci-training (L40S)

Changes:
- Add ci-compile-cpu Terraform pool (POP2-HC-32C-64G, autoscale 0-1)
- Set CARGO_BUILD_JOBS=30 to use all CPUs for Rust compilation
- Increase CPU request to 28000m/31000m (was 7000m/7800m on L4)
- Add docker tag to deploy/infra jobs for explicit CPU runner routing
- Update CI header with new pool routing documentation

Cost: EUR 0.85/h (vs EUR 1.1/h for L4) — cheaper AND 4x faster compile

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-27 21:08:42 +01:00
jgrusewski
bb208b29b2 fix(deps): unify workspace dependency versions and clean up CI
- Upgrade dashmap 6.0→6.1, tokio-tungstenite 0.21→0.24 in workspace
- Upgrade rust-version 1.75→1.85 (CI uses Rust 1.89)
- Remove unused arrayfire from ml crate
- Unify member crates to use workspace = true (nalgebra, dashmap, tokio-tungstenite)
- Fix data crate WebSocket connect calls for tungstenite 0.24 API (Url→str)
- Remove redundant KUBERNETES_RUNTIME_CLASS_NAME from training templates
  (runner rev 34 sets nvidia runtime globally)

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-27 20:52:06 +01:00
jgrusewski
4e0dc97041 fix(ci): decouple compile from nvidia runtime
Compile only needs CUDA stubs (baked into ci-builder) + CUDA_COMPUTE_CAP
for candle-kernels PTX generation. No nvidia-smi or real GPU needed.

- Hardcode CUDA_COMPUTE_CAP=89 (L4) in .rust-base variables
- Remove KUBERNETES_RUNTIME_CLASS_NAME from .rust-base
- compile-services: skip LD_LIBRARY_PATH stripping and nvidia-smi
- test: graceful fallback when nvidia-smi unavailable

This unblocks compile on nodes without nvidia RuntimeClass handler
and opens the door to using CPU-only nodes for faster compilation.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-27 20:42:25 +01:00
jgrusewski
312d113c47 fix(ci): use k=v format for node selector overrides
GitLab Runner 18.9 expects KUBERNETES_NODE_SELECTOR_* values in
"key=value" format. Changed from bare values to "pool=services"
and "pool=ci-training".

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-27 20:13:29 +01:00
jgrusewski
6462e29ed3 fix(ci): use simple pool labels for node selector routing
Runner default nodeSelector changed from k8s.scaleway.com/pool-name to
simple pool=<name> label. CI variables now use correct format:
KUBERNETES_NODE_SELECTOR_pool: <value> (lowercase suffix = label key).

Fixes deploy job scheduling failure (wrong label format caused pods to
land on ci-compile instead of services pool).

Also fixes .train-rl-base tag from kapsule-rl (no matching runner) to
kapsule+gpu (matches existing runner).

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-27 20:11:51 +01:00
jgrusewski
5678d8ec2c fix(ci): move nvidia runtime to per-job override for deploy compatibility
Deploy job was failing with "no runtime for nvidia" because the runner
set runtime_class_name=nvidia globally, but deploy routes to the
services pool which has no GPU/nvidia runtime.

Fix: removed global runtime_class_name from runner config, added
runtime_class_name_overwrite_allowed=".*", and set
KUBERNETES_RUNTIME_CLASS_NAME=nvidia in GPU job templates
(.rust-base, .train-rl-base, .train-validate-base).

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-27 19:57:07 +01:00
jgrusewski
8225950f06 feat(ml): parallel PSO hyperopt with IBKR cost defaults
Enable concurrent trial evaluation for DQN/PPO hyperparameter
optimization via clone-per-particle pattern — each PSO particle
clones the trainer and trains independently, replacing the previous
Arc<Mutex> serialization bottleneck. On L4 (8 vCPU) this yields
~4-5x throughput improvement.

Changes:
- DQNTrainer/PPOTrainer: Clone with Arc<AtomicUsize> trial counter
- DQNTrainer: replace unsafe mutable aliasing with Arc<Mutex> for
  best_trial tracking
- ArgminOptimizer: add optimize_parallel() with ParallelObjectiveFunction
  and scoped-thread LHS evaluation
- CLI: --parallel 0 (auto-detect CPUs-2), --initial-capital 35000,
  --tx-cost-bps 0.1 (IBKR ES all-in)
- CI: both hyperopt jobs use --parallel 0 + IBKR ES cost defaults

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-27 19:42:22 +01:00
jgrusewski
b4cfd88634 refactor(ml): drop dead TFT Parquet loader (-130 lines)
train_from_parquet and load_training_data_from_parquet had zero
callers — TFT hyperopt uses train_from_bars via DBN data.
Also removes unused NormalizationParams struct from this module.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-27 18:56:04 +01:00
jgrusewski
3564efe961 fix(ml): fix TFT OOB in historical features — 51 not 54 elements
Feature vectors are [f64; 51] after WAVE 10 (Proxy OFI removal):
  [0..5] static, [5..15] known, [15..51] unknown (36 features)

Old code used fv[15..54].min(fv.len()) which silently produced 36
features per step but expected 39 in Array2 shape → ShapeError.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-27 18:52:35 +01:00
jgrusewski
20e67d4d64 feat(ci): tune RL resources to actual usage — 2000m CPU, concurrent=2
RL env simulation is single-threaded (~1 core actual). Lowering from
7000m to 2000m request allows DQN + PPO hyperopt to run simultaneously
on one L4 node.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-27 18:46:16 +01:00
jgrusewski
694e10c7a6 feat(ci): route deploy job to services pool, not GPU nodes
Deploy only needs kubectl — no reason to consume GPU node CPU.
Routes to services pool with minimal 200m/500m CPU.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-27 18:39:13 +01:00
jgrusewski
6913506669 feat(ci): fix RL runner SA + full L4 node for sequential RL jobs
- Fix ServiceAccount: use serviceAccount.name (not deprecated serviceAccountName)
- concurrent=1: one RL job gets full L4 node (7000m/7800m CPU, 16Gi/40Gi RAM)
- Faster iteration: ~2.3x more CPU per job vs splitting across 2 concurrent

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-27 18:28:34 +01:00
jgrusewski
0267e4949c feat(ci): add RL runner with L4 PVCs for cross-node training
- New gitlab-runner-rl Helm release: tags kapsule-rl, mounts
  training-data-l4-pvc + sccache-l4-pvc (separate RWO volumes)
- .train-rl-base now uses tags: [kapsule-rl] → picked up by RL runner
- Avoids Multi-Attach errors from RWO PVCs shared across L4/L40S nodes

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-27 18:12:20 +01:00
jgrusewski
3ef88c477d feat(ci): route RL→L4, supervised→L40S with tuned resources
- Add .train-rl-base template: no node selector (→ L4 default), CPU 3000m/3800m
- Update .train-validate-base: CPU 1000m/2000m for GPU-bound supervised models
- Switch train-validate-rl, hyperopt-ppo, hyperopt-dqn to .train-rl-base
- Reduce RL hyperopt epochs 15→8 for faster iteration (~2.4h vs 4.5h/trial)
- Supervised jobs unchanged: 8 models still on L40S with --epochs 15

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-27 16:18:47 +01:00