Files
foxhunt/docs/plans/2026-02-27-gpu-resource-optimization-implementation.md
jgrusewski c20ce9fdb8 docs: GPU resource optimization implementation plan
4 tasks: add .train-rl-base template, switch RL jobs, reduce epochs, validate.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-27 16:16:55 +01:00

6.2 KiB

GPU Resource Optimization Implementation Plan

For Claude: REQUIRED SUB-SKILL: Use superpowers:executing-plans to implement this plan task-by-task.

Goal: Route RL training (DQN, PPO) to cheaper L4 GPUs and supervised models to L40S, with tuned CPU/memory per workload.

Architecture: Add a .train-rl-base CI template (no node selector override → runner default ci-compile/L4). Update .train-validate-base CPU down for GPU-bound supervised models. Switch 3 RL jobs to new template. Reduce RL hyperopt epochs 15→8.

Tech Stack: GitLab CI YAML, Kubernetes node selectors via runner config overwrite variables


Task 1: Add .train-rl-base CI template

Files:

  • Modify: .gitlab-ci.yml:413 (insert new template BEFORE existing .train-validate-base)

Step 1: Insert .train-rl-base template above .train-validate-base

Find the comment block starting at line 410 and insert the new template before it. The new template goes right before the # Runs the training binary... comment block:

# --------------------------------------------------------------------------
# RL training base: DQN + PPO → ci-compile pool (L4)
# RL models are CPU-bound (environment simulation), use <1GB VRAM.
# No KUBERNETES_NODE_SELECTOR_POOL → runner default = ci-compile (L4-1-24G)
# --------------------------------------------------------------------------
.train-rl-base:
  stage: train
  image: ${REGISTRY}/training:${CI_COMMIT_SHA}
  tags:
    - kapsule
    - gpu
  variables:
    # RL is CPU-bound — give it most of the L4's 7800m allocatable CPU.
    # 3000m request allows 2 concurrent RL jobs on one L4 node.
    KUBERNETES_CPU_REQUEST: "3000m"
    KUBERNETES_CPU_LIMIT: "3800m"
    KUBERNETES_MEMORY_REQUEST: "8Gi"
    KUBERNETES_MEMORY_LIMIT: "16Gi"
  rules:
    - if: $CI_COMMIT_BRANCH == "main" && $CI_PIPELINE_SOURCE == "push"
      when: manual
      allow_failure: true
    - if: $CI_PIPELINE_SOURCE == "schedule" && $TRAIN_VALIDATE == "true"
    - if: $CI_PIPELINE_SOURCE == "web" || $CI_PIPELINE_SOURCE == "api"
      when: manual
      allow_failure: true
  before_script:
    - export LD_LIBRARY_PATH=$(echo "$LD_LIBRARY_PATH" | tr ':' '\n' | grep -v stubs | tr '\n' ':' | sed 's/:$//')
    - nvidia-smi
    - mkdir -p ${CI_PROJECT_DIR}/output

Insert this BEFORE the existing line 410 (# Runs the training binary for a few steps...).

Step 2: Update .train-validate-base comment and CPU

In the existing .train-validate-base (now shifted down), update:

  • Comment: change to clarify it's for supervised models on L40S
  • CPU request: "2000m""1000m"
  • CPU limit: "3500m""2000m"

The updated .train-validate-base should read:

# --------------------------------------------------------------------------
# Supervised training base: TFT, Mamba2, TGGN, TLOB, Liquid, KAN, xLSTM, Diffusion → L40S
# Supervised models are GPU-bound — low CPU, high VRAM. Routes to ci-training (L40S-1-48G).
# --------------------------------------------------------------------------
.train-validate-base:
  stage: train
  image: ${REGISTRY}/training:${CI_COMMIT_SHA}
  tags:
    - kapsule
    - gpu
  variables:
    # Route to L40S pool (48GB VRAM for large supervised models)
    KUBERNETES_NODE_SELECTOR_POOL: "k8s.scaleway.com/pool-name=ci-training"
    # Supervised is GPU-bound — minimal CPU. Low request allows 3-4 concurrent jobs.
    KUBERNETES_CPU_REQUEST: "1000m"
    KUBERNETES_CPU_LIMIT: "2000m"
    KUBERNETES_MEMORY_REQUEST: "16Gi"
    KUBERNETES_MEMORY_LIMIT: "40Gi"

Step 3: Commit

git add .gitlab-ci.yml
git commit -m "feat(ci): add .train-rl-base template for L4 routing, tune supervised CPU"

Task 2: Switch RL jobs to .train-rl-base

Files:

  • Modify: .gitlab-ci.yml — 3 jobs: train-validate-rl, hyperopt-ppo, hyperopt-dqn

Step 1: Change train-validate-rl extends

Find train-validate-rl: and change:

# BEFORE:
train-validate-rl:
  extends: .train-validate-base

# AFTER:
train-validate-rl:
  extends: .train-rl-base

Step 2: Change hyperopt-ppo extends

Find hyperopt-ppo: and change:

# BEFORE:
hyperopt-ppo:
  extends: .train-validate-base

# AFTER:
hyperopt-ppo:
  extends: .train-rl-base

Step 3: Change hyperopt-dqn extends

Find hyperopt-dqn: and change:

# BEFORE:
hyperopt-dqn:
  extends: .train-validate-base

# AFTER:
hyperopt-dqn:
  extends: .train-rl-base

Step 4: Commit

git add .gitlab-ci.yml
git commit -m "feat(ci): route RL jobs (DQN, PPO) to L4 via .train-rl-base"

Task 3: Reduce RL hyperopt epochs 15 → 8

Files:

  • Modify: .gitlab-ci.ymlhyperopt-ppo and hyperopt-dqn scripts

Step 1: Change hyperopt-ppo epochs

In the hyperopt-ppo script section, find --epochs 15 and change to --epochs 8.

Step 2: Change hyperopt-dqn epochs

In the hyperopt-dqn script section, find --epochs 15 and change to --epochs 8.

Step 3: Commit

git add .gitlab-ci.yml
git commit -m "feat(ci): reduce RL hyperopt epochs 15→8 for faster iteration"

Task 4: Validate pipeline and push

Step 1: Lint the YAML

python3 -c "import yaml; yaml.safe_load(open('.gitlab-ci.yml'))" && echo "YAML valid"

Expected: YAML valid

Step 2: Verify job routing by checking extends

grep -E '(train-validate-rl|hyperopt-ppo|hyperopt-dqn):' -A1 .gitlab-ci.yml | grep extends

Expected: All 3 show extends: .train-rl-base

grep -E '(train-validate-tft|train-validate-mamba2|hyperopt-tft|hyperopt-mamba2):' -A1 .gitlab-ci.yml | grep extends

Expected: All show extends: .train-validate-base

Step 3: Verify epochs

grep -B5 'epochs 8' .gitlab-ci.yml | grep -E '(hyperopt-ppo|hyperopt-dqn)'

Expected: Both RL hyperopt jobs show epochs 8.

grep -B5 'epochs 15' .gitlab-ci.yml | head -20

Expected: Only supervised hyperopt jobs (tft, mamba2, liquid, tggn, tlob, kan, xlstm, diffusion) still have --epochs 15.

Step 4: Push and verify

git push origin main

Trigger a pipeline via Rails console and verify:

  • RL jobs (train-validate-rl, hyperopt-ppo, hyperopt-dqn) schedule on ci-compile (L4) node
  • Supervised jobs schedule on ci-training (L40S) node
  • RL hyperopt shows --epochs 8 in log output