Adds a parent/child GitLab CI pipeline for ML model training: - Generator script produces per-model hyperopt/train/evaluate jobs - Parent pipeline (.gitlab-ci-training.yml) with manual trigger - NFS-backed ReadWriteMany PVC for shared training outputs - Hyperopt params wired into training binaries (DQN, PPO, TFT, Mamba2) - Shared DBN loader eliminates duplicate code across hyperopt adapters - Supervised hyperopt unified to DBN data (was parquet-only) Pipeline: hyperopt (4 models) → train (10 models) → evaluate ensemble Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
18 KiB
ML Training Pipeline Implementation Plan
For Claude: REQUIRED SUB-SKILL: Use superpowers:executing-plans to implement this plan task-by-task.
Goal: Manually-triggered GitLab CI pipeline that runs hyperopt → train → evaluate for the 10-model ML ensemble on H100 GPUs.
Architecture: A parent .gitlab-ci-training.yml accepts trigger variables, runs a shell script that generates a child pipeline YAML with exactly the needed jobs (models × symbols × phases), then triggers the child pipeline. Jobs run in the pre-built training Docker image with PVC mounts for data and output.
Tech Stack: GitLab CI (parent/child pipelines), Bash (YAML generator), Rust/Clap (--params-file arg), Kubernetes PVCs
Task 1: Output PVC manifest
Files:
- Create:
infra/k8s/training/training-output-pvc.yaml
Step 1: Create the PVC manifest
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: training-output-pvc
namespace: foxhunt
labels:
app.kubernetes.io/name: training-output
app.kubernetes.io/part-of: foxhunt
spec:
accessModes:
- ReadWriteOnce
resources:
requests:
storage: 50Gi
storageClassName: scw-bssd
50Gi because safetensors checkpoints for 10 models × 4 symbols × multiple folds can add up. scw-bssd matches the existing training-data-pvc.
Step 2: Commit
git add infra/k8s/training/training-output-pvc.yaml
git commit -m "infra: add training-output-pvc for ML pipeline checkpoints"
Task 2: Pipeline generator script
Files:
- Create:
scripts/generate-training-pipeline.sh
This script reads environment variables and writes a .training-generated.yml artifact with the exact jobs needed.
Step 1: Write the generator script
The script must:
- Parse
MODELSinto a list (expandall→ 10 models,rl→ dqn,ppo,supervised→ 8 models) - Parse
SYMBOLSinto a list (comma-separated) - Map each model to its binary (
dqn/ppo→*_rl, others →*_supervised) - Map each model to whether it has hyperopt support (
dqn,ppo,tft,mamba2→ yes, others → no) - Write stages, variables, and job definitions to YAML
- Respect
PHASEvariable to skip hyperopt or train stages
Key details:
- All jobs use the training image from
$REGISTRY/training:latest - All jobs mount
training-data-pvcat/data(read-only) andtraining-output-pvcat/output - All jobs use
nodeSelector: gpu-trainingand request 1 GPU - Hyperopt jobs for RL use
--data-dir /data(DBN files with symbol subdirs) - Hyperopt jobs for supervised use
--parquet-file— but the supervised hyperopt binary expects Parquet, NOT DBN. This is a design mismatch. For now, the generator skips supervised hyperopt unless aPARQUET_DIRvariable is set. This is a known limitation documented in the script. - Train jobs pass
--params-file /output/$RUN_ID/hyperopt/$MODEL/$SYMBOL/results.jsonif hyperopt ran - The evaluate job
needs:all train jobs
The script location: scripts/generate-training-pipeline.sh
#!/usr/bin/env bash
set -euo pipefail
# --- Input variables (from GitLab CI trigger) ---
SYMBOLS="${SYMBOLS:-ES.FUT}"
MODELS="${MODELS:-all}"
PHASE="${PHASE:-full}"
MAX_PARALLEL="${MAX_PARALLEL:-10}"
EPOCHS="${EPOCHS:-50}"
HYPEROPT_TRIALS="${HYPEROPT_TRIALS:-20}"
RUN_ID="${RUN_ID:-$(date +%Y%m%d-%H%M%S)}"
REGISTRY="${REGISTRY:-rg.fr-par.scw.cloud/foxhunt-ci}"
OUTPUT=".training-generated.yml"
# --- Expand model groups ---
RL_MODELS="dqn ppo"
SUPERVISED_MODELS="tft mamba2 tggn tlob liquid kan xlstm diffusion"
HYPEROPT_MODELS="dqn ppo tft mamba2"
case "$MODELS" in
all) MODEL_LIST="$RL_MODELS $SUPERVISED_MODELS" ;;
rl) MODEL_LIST="$RL_MODELS" ;;
supervised) MODEL_LIST="$SUPERVISED_MODELS" ;;
*) MODEL_LIST=$(echo "$MODELS" | tr ',' ' ') ;;
esac
IFS=',' read -ra SYMBOL_LIST <<< "$SYMBOLS"
# --- Helper: which binary for a model ---
binary_for() {
case "$1" in
dqn|ppo) echo "train_baseline_rl" ;;
*) echo "train_baseline_supervised" ;;
esac
}
hyperopt_binary_for() {
case "$1" in
dqn|ppo) echo "hyperopt_baseline_rl" ;;
tft|mamba2) echo "hyperopt_baseline_supervised" ;;
*) echo "" ;;
esac
}
has_hyperopt() {
for m in $HYPEROPT_MODELS; do
[ "$m" = "$1" ] && return 0
done
return 1
}
is_rl() {
[ "$1" = "dqn" ] || [ "$1" = "ppo" ]
}
# --- Write YAML header ---
cat > "$OUTPUT" <<YAML
# Auto-generated training pipeline — do not edit
# Run ID: $RUN_ID
# Models: $MODEL_LIST
# Symbols: ${SYMBOL_LIST[*]}
# Phase: $PHASE
stages:
- hyperopt
- train
- evaluate
variables:
RUN_ID: "$RUN_ID"
.training-base:
image: ${REGISTRY}/training:latest
tags:
- kapsule
- rust
variables:
NVIDIA_VISIBLE_DEVICES: all
NVIDIA_DRIVER_CAPABILITIES: compute,utility
RUST_LOG: info
SQLX_OFFLINE: "true"
before_script:
- nvidia-smi || true
- mkdir -p /output/\$RUN_ID
YAML
# --- Generate hyperopt jobs ---
HYPEROPT_JOBS=""
if [ "$PHASE" = "full" ] || [ "$PHASE" = "hyperopt" ]; then
for model in $MODEL_LIST; do
has_hyperopt "$model" || continue
for symbol in "${SYMBOL_LIST[@]}"; do
job_name="hyperopt-${model}-${symbol//\./-}"
HYPEROPT_JOBS="$HYPEROPT_JOBS $job_name"
cat >> "$OUTPUT" <<YAML
${job_name}:
extends: .training-base
stage: hyperopt
resource_group: gpu-${job_name}
script:
YAML
if is_rl "$model"; then
cat >> "$OUTPUT" <<YAML
- mkdir -p /output/\$RUN_ID/hyperopt/${model}/${symbol}
- hyperopt_baseline_rl
--model ${model}
--trials ${HYPEROPT_TRIALS}
--epochs 10
--data-dir /data
--symbol ${symbol}
--output /output/\$RUN_ID/hyperopt/${model}/${symbol}/results.json
--base-dir /tmp/ml_training
YAML
else
# Supervised hyperopt needs parquet — skip if not available
cat >> "$OUTPUT" <<YAML
- echo "Supervised hyperopt for ${model} requires --parquet-file"
- echo "Skipping — train will use default config"
- mkdir -p /output/\$RUN_ID/hyperopt/${model}/${symbol}
YAML
fi
done
done
fi
# --- Generate train jobs ---
TRAIN_JOBS=""
if [ "$PHASE" = "full" ] || [ "$PHASE" = "train" ]; then
for model in $MODEL_LIST; do
for symbol in "${SYMBOL_LIST[@]}"; do
job_name="train-${model}-${symbol//\./-}"
TRAIN_JOBS="$TRAIN_JOBS $job_name"
# Build needs list (depend on hyperopt if it ran)
hyperopt_job="hyperopt-${model}-${symbol//\./-}"
if [ "$PHASE" = "full" ] && has_hyperopt "$model"; then
needs="[${hyperopt_job}]"
else
needs="[]"
fi
# Build params-file arg if hyperopt might have run
params_arg=""
if has_hyperopt "$model"; then
params_arg="--hyperopt-params /output/\$RUN_ID/hyperopt/${model}/${symbol}/results.json"
fi
cat >> "$OUTPUT" <<YAML
${job_name}:
extends: .training-base
stage: train
needs: ${needs}
resource_group: gpu-${job_name}
script:
- mkdir -p /output/\$RUN_ID/models/${model}/${symbol}
YAML
if is_rl "$model"; then
cat >> "$OUTPUT" <<YAML
- train_baseline_rl
--model ${model}
--epochs ${EPOCHS}
--data-dir /data
--output-dir /output/\$RUN_ID/models/${model}/${symbol}
${params_arg}
YAML
else
cat >> "$OUTPUT" <<YAML
- train_baseline_supervised
--model ${model}
--epochs ${EPOCHS}
--data-dir /data
--symbol ${symbol}
--output-dir /output/\$RUN_ID/models/${model}/${symbol}
YAML
fi
done
done
fi
# --- Generate evaluate job ---
if [ "$PHASE" = "full" ] || [ "$PHASE" = "eval" ]; then
# Build needs list from all train jobs
if [ -n "$TRAIN_JOBS" ]; then
needs_list=$(echo "$TRAIN_JOBS" | xargs -n1 | sed 's/.*/"&"/' | paste -sd,)
needs="[${needs_list}]"
else
needs="[]"
fi
cat >> "$OUTPUT" <<YAML
evaluate-ensemble:
extends: .training-base
stage: evaluate
needs: ${needs}
script:
- mkdir -p /output/\$RUN_ID/eval
- evaluate_baseline
--models-dir /output/\$RUN_ID/models
--data-dir /data
--output /output/\$RUN_ID/eval/ensemble_report.json
- cat /output/\$RUN_ID/eval/ensemble_report.json
artifacts:
paths:
- /output/\$RUN_ID/eval/ensemble_report.json
when: always
expire_in: 30 days
YAML
fi
echo "Generated pipeline: $OUTPUT"
echo " Models: $MODEL_LIST"
echo " Symbols: ${SYMBOL_LIST[*]}"
echo " Phase: $PHASE"
echo " Run ID: $RUN_ID"
Make it executable.
Step 2: Verify the script generates valid YAML
Run locally with defaults:
chmod +x scripts/generate-training-pipeline.sh
REGISTRY=rg.fr-par.scw.cloud/foxhunt-ci scripts/generate-training-pipeline.sh
cat .training-generated.yml
Expected: YAML with hyperopt jobs for dqn, ppo (tft/mamba2 skipped — no parquet), 10 train jobs for ES.FUT, and 1 evaluate job.
Step 3: Test with subset
MODELS=dqn SYMBOLS=ES.FUT PHASE=train scripts/generate-training-pipeline.sh
cat .training-generated.yml
Expected: YAML with only train-dqn-ES-FUT and evaluate-ensemble. No hyperopt jobs.
Step 4: Commit
git add scripts/generate-training-pipeline.sh
echo ".training-generated.yml" >> .gitignore
git add .gitignore
git commit -m "feat: add training pipeline generator script"
Task 3: Parent pipeline YAML
Files:
- Create:
.gitlab-ci-training.yml
Step 1: Write the parent pipeline
# ML Training Pipeline — manually triggered
#
# Trigger via GitLab UI: CI/CD → Pipelines → Run pipeline
# Select .gitlab-ci-training.yml as the CI config, set variables.
#
# Or via API:
# curl -X POST --fail \
# -F "token=$TRIGGER_TOKEN" \
# -F "ref=main" \
# -F "variables[MODELS]=all" \
# -F "variables[SYMBOLS]=ES.FUT" \
# -F "variables[PHASE]=full" \
# "$CI_API_V4_URL/projects/$CI_PROJECT_ID/trigger/pipeline"
stages:
- prepare
- trigger
variables:
SYMBOLS: "ES.FUT"
MODELS: "all"
PHASE: "full"
MAX_PARALLEL: "10"
EPOCHS: "50"
HYPEROPT_TRIALS: "20"
RUN_ID: ""
REGISTRY: rg.fr-par.scw.cloud/foxhunt-ci
workflow:
rules:
- if: $CI_PIPELINE_SOURCE == "web"
- if: $CI_PIPELINE_SOURCE == "trigger"
- if: $CI_PIPELINE_SOURCE == "api"
generate-jobs:
stage: prepare
image: alpine:3.19
tags:
- kapsule
script:
- apk add --no-cache bash coreutils
- |
if [ -z "$RUN_ID" ]; then
export RUN_ID=$(date +%Y%m%d-%H%M%S)
fi
- bash scripts/generate-training-pipeline.sh
- cat .training-generated.yml
artifacts:
paths:
- .training-generated.yml
expire_in: 1 day
run-training:
stage: trigger
trigger:
include:
- artifact: .training-generated.yml
job: generate-jobs
strategy: depend
Note: The child pipeline jobs reference the training image and GPU runner tags. The parent runs on a lightweight alpine image (no GPU needed).
Step 2: Verify YAML syntax
python3 -c "import yaml; yaml.safe_load(open('.gitlab-ci-training.yml'))" && echo "Valid YAML"
Step 3: Commit
git add .gitlab-ci-training.yml
git commit -m "feat: add manually-triggered ML training pipeline"
Task 4: Wire --hyperopt-params in train_baseline_rl
Files:
- Modify:
crates/ml/examples/train_baseline_rl.rs
The --hyperopt_params arg already exists (line 71-73) but is unused. Wire it up to override DQN/PPO config fields from the hyperopt JSON.
Step 1: Read the hyperopt output format
The hyperopt JSON looks like:
{
"dqn": {
"best_objective": 0.123,
"best_params": { "learning_rate": 0.001, "hidden_dim": 128, ... },
"trials": 20,
"elapsed_secs": 45.3
}
}
Or for single-model runs:
{
"best_objective": 0.123,
"best_params": { "learning_rate": 0.001, ... },
"trials": 20,
"elapsed_secs": 45.3
}
Step 2: Add a helper function to load and apply hyperopt params
Add after the Args struct:
/// Load hyperopt results JSON and extract best_params for the given model.
/// Returns None if file doesn't exist or model key is missing.
fn load_hyperopt_params(path: &Path, model: &str) -> Option<serde_json::Value> {
let content = std::fs::read_to_string(path).ok()?;
let json: serde_json::Value = serde_json::from_str(&content).ok()?;
// Try model-keyed format first: {"dqn": {"best_params": {...}}}
if let Some(params) = json.get(model).and_then(|m| m.get("best_params")) {
return Some(params.clone());
}
// Try flat format: {"best_params": {...}}
if let Some(params) = json.get("best_params") {
return Some(params.clone());
}
None
}
Step 3: Apply params to DQNConfig and PPOConfig
In the DQN training section, after creating the default config, add:
if let Some(ref params_path) = args.hyperopt_params {
if let Some(params) = load_hyperopt_params(params_path, "dqn") {
info!("Loading hyperopt params from: {}", params_path.display());
if let Some(lr) = params.get("learning_rate").and_then(|v| v.as_f64()) {
dqn_config.learning_rate = lr;
info!(" learning_rate = {}", lr);
}
if let Some(hd) = params.get("hidden_dim").and_then(|v| v.as_u64()) {
dqn_config.hidden_dim = hd as usize;
info!(" hidden_dim = {}", hd);
}
// Add additional param mappings as hyperopt search space grows
} else {
info!("No hyperopt params found at {}; using defaults", params_path.display());
}
}
Same pattern for PPO section, reading "ppo" key and mapping to PPOConfig fields.
Step 4: Build and verify
SQLX_OFFLINE=true cargo check -p ml --example train_baseline_rl
Expected: compiles with 0 errors, 0 warnings.
Step 5: Commit
git add crates/ml/examples/train_baseline_rl.rs
git commit -m "feat(ml): wire --hyperopt-params to override DQN/PPO config"
Task 5: Add --hyperopt-params to train_baseline_supervised
Files:
- Modify:
crates/ml/examples/train_baseline_supervised.rs
Step 1: Add the CLI arg
Add to the Args struct after spread_ticks:
/// Optional path to hyperopt results JSON to override default config
#[arg(long)]
hyperopt_params: Option<PathBuf>,
Step 2: Add the same load_hyperopt_params helper
Same function as Task 4.
Step 3: Apply params to model configs
In each model's config construction (TFT, Mamba2, etc.), add the override block. Only TFT and Mamba2 have hyperopt adapters, so only those need param mapping. Example for TFT:
if let Some(ref params_path) = args.hyperopt_params {
if let Some(params) = load_hyperopt_params(params_path, "tft") {
info!("Loading TFT hyperopt params from: {}", params_path.display());
if let Some(lr) = params.get("learning_rate").and_then(|v| v.as_f64()) {
tft_config.learning_rate = lr as f32;
info!(" learning_rate = {}", lr);
}
if let Some(hd) = params.get("hidden_size").and_then(|v| v.as_u64()) {
tft_config.hidden_size = hd as usize;
info!(" hidden_size = {}", hd);
}
}
}
For models without hyperopt (TGGN, TLOB, etc.), no changes — the param file simply won't have their key.
Step 4: Build and verify
SQLX_OFFLINE=true cargo check -p ml --example train_baseline_supervised
Step 5: Commit
git add crates/ml/examples/train_baseline_supervised.rs
git commit -m "feat(ml): add --hyperopt-params to supervised training binary"
Task 6: Update Dockerfile.training for PVC mounts
Files:
- Modify:
infra/k8s/training/job-template.yaml
Step 1: Update the job template
The current template has training-data-pvc at /data and emptyDir at /output. Replace the output volume with the new PVC and add nodeSelector for H100.
Update volumes section:
volumes:
- name: training-data
persistentVolumeClaim:
claimName: training-data-pvc
- name: output
persistentVolumeClaim:
claimName: training-output-pvc
The nodeSelector is already gpu-training which is the H100 pool. No change needed there.
Step 2: Verify YAML
python3 -c "import yaml; yaml.safe_load(open('infra/k8s/training/job-template.yaml'))" && echo "Valid"
Step 3: Commit
git add infra/k8s/training/job-template.yaml
git commit -m "infra: use training-output-pvc instead of emptyDir in job template"
Task 7: Integration test — dry run the full pipeline locally
Step 1: Generate the pipeline for full ensemble
MODELS=all SYMBOLS=ES.FUT PHASE=full scripts/generate-training-pipeline.sh
Step 2: Verify job count
Expected in .training-generated.yml:
- 2 hyperopt jobs (dqn, ppo — tft/mamba2 skipped without parquet)
- 10 train jobs (all models)
- 1 evaluate job
- Total: 13 jobs
grep -c "^[a-z].*:$" .training-generated.yml
Step 3: Verify multi-symbol expansion
MODELS=dqn,ppo SYMBOLS=ES.FUT,NQ.FUT PHASE=full scripts/generate-training-pipeline.sh
Expected:
- 4 hyperopt jobs (dqn×2 symbols, ppo×2 symbols)
- 4 train jobs
- 1 evaluate job
- Total: 9 jobs
Step 4: Verify phase filtering
MODELS=kan SYMBOLS=ES.FUT PHASE=train scripts/generate-training-pipeline.sh
Expected: 0 hyperopt, 1 train, 1 evaluate. Total: 2 jobs.
Step 5: Final commit with all files
git add -A
git status
git commit -m "feat: ML training pipeline — generator, CI config, PVC, param passing"
Task 8: Update design doc with known limitations
Files:
- Modify:
docs/plans/2026-02-26-training-pipeline-design.md
Add a "Known Limitations" section:
- Supervised hyperopt (
hyperopt_baseline_supervised) expects Parquet input, not DBN. The generator currently skips supervised hyperopt. Future work: add DBN support to supervised hyperopt adapters. evaluate_baselinecurrently only evaluates RL models (DQN/PPO). Extending it to load and evaluate all 10 model types needs work on the evaluation binary.- The output PVC is
ReadWriteOnce— only one node can mount it at a time. If parallel jobs land on different nodes, they'll fail. Mitigation: useReadWriteManywith an NFS-backed storage class, or ensure all training pods schedule on the same node.
Commit
git add docs/plans/2026-02-26-training-pipeline-design.md
git commit -m "docs: add known limitations to training pipeline design"