Files
foxhunt/crates/ml-alpha/cuda/rl_rollout_steps_controller.cu
jgrusewski 083a88f7c3 feat(rl): adaptive controller floors — 12 controllers, all signal-driven
Replaces hardcoded thresholds AND clamp bounds across 12 RL controllers
with observed-signal-driven ISV-slot bounds. Eliminates the architectural
failure mode that surfaced in walk-forward fold 0 (alpha-rl-m9cx5) and
fold 1 (alpha-rl-jgdh6): under Phase 4.5 advantage normalization, eight
controllers (PPO clip, target_tau, rollout_steps, entropy_coef, per_α,
gamma, reward_scale, q_distill_lambda) saturated at extrema within ~50
steps and stayed pinned for the rest of training — same pattern across
two different data slices, confirming the bug is structural rather than
data-size-dependent.

Mechanism: each controller's noise floor was hardcoded as a small
fraction of its target (`_NOISE_FLOOR_FRAC = 0.01f`), calibrated against
a pre-Phase-4.5 signal regime. Phase 4.5 normalization reduces operating
KL by ~50× — observed signal stays below the 1%-of-target floor, the
Schulman widen path fires continuously (asymmetric in the wrong
direction), and ε hits MAX 0.50 within ~50 training steps. ksll2's full-
data run (n_folds=1) happened to escape via a single above-band KL
observation that triggered tighten; both walk-forward folds (3 and 6
files) did not.

Per `feedback_adaptive_not_tuned`, `feedback_isv_for_adaptive_bounds`,
`pearl_controller_anchors_isv_driven`: every threshold now derives from
observed signal statistics (Welford online variance) rather than
constants calibrated against a prior signal regime. User explicitly
expanded scope mid-implementation: "if all clamps ISV bound should be
added to this spec and tasks" — applying the principle consistently
means hardcoded MIN/MAX clamp bounds count too, not just the saturating
noise floors. 12 controllers, single atomic commit per
`feedback_no_partial_refactor`.

Spec: docs/superpowers/specs/2026-05-30-adaptive-controller-floor-design.md
Plan: docs/superpowers/plans/2026-05-30-adaptive-controller-floor-plan.md

# New kernel
- rl_signal_variance_update.cu (80 LOC) — Welford online variance.
  Single-thread single-block. Sentinel-zero skip per pearl. Per-controller
  Welford triple (count, mean, M²) drives every adaptive noise floor.

# Per-controller refactors (12 .cu files)
- ppo_clip + target_tau + rollout_steps + entropy_coef + per_α —
  adaptive noise floor = max(target × 0.5, sqrt(observed_var) × 2);
  asymmetric Schulman (tighten on single observation, widen requires 3
  consecutive below-band).
- gamma (Special G) — hardcoded GAMMA_MIN = 0.995 → adaptive via
  Welford MEAN of trade duration. Per spec Q6 resolution: the Welford
  mean's natural N-smoothing lag breaks the gamma↔trade_duration
  feedback loop without explicit step-period gating. 100-observation
  warmup falls back to EMA before Welford has enough samples.
- reward_scale (Special R) — asymmetric DECREASE rate cap (5% per
  step) on both bootstrap-replace and Wiener-blend paths; bootstrap-
  fraction floor (10% of bootstrap) until 100 trades close.
- q_distill_lambda (Special Q) — hardcoded MIN_LAMBDA = 0.05 → adaptive
  via Welford on q_distill_kl_ema: max(0.001, std × 0.05).
- v_blend_alpha (Phase 4.4) — 5 hardcoded constants → 5 ISV slots
  (DEAD_SIGNAL_FLOOR, TARGET_TRACK_RATIO, SCHULMAN_STEP, EMA_ALPHA,
  BOOTSTRAP_ALPHA) + adaptive dead-signal floor from V_scalar magnitude
  variance.
- ppo_ratio_clamp — adaptive MIN/MAX from observed log-ratio variance.
  Architectural 2.0 absolute floor preserved (don't degenerate to
  vanilla policy gradient).
- reward_clamp — V_BOUND_FLOOR, V_BOUND_EWMA_ALPHA, MIN_WIN, MIN_RATIO,
  MAX_RATIO, MIN/MAX_MARGIN, MARGIN_TOLERANCE/ADJUST_RATE,
  CLIP_RATE_EMA_ALPHA → 10 ISV slots.
- gate_threshold — hardcoded alpha = 0.01 → ISV slot 638.
- Clamp-bound expansion: EPS_MIN/MAX, TAU_MIN, ROLLOUT_MIN/MAX,
  COEF_MIN/MAX, PER_ALPHA_MIN/MAX, GAMMA_MAX, MAX_LAMBDA, KL_TOLERANCE,
  LAMBDA_RAMP_RATE, LAMBDA_DECAY_RATE → all ISV.
- WIENER_ALPHA_FLOOR shared across 9 controllers → single ISV slot 659.

# Trainer integration (integrated.rs)
- rl_signal_variance_update kernel loaded + helper method
  `launch_rl_signal_variance_update`.
- Per-step Welford launches for 9 controller inputs, placed between
  EMA producers and rl_fused_controllers in the per-step pipeline.
- Input-slot lookup array for rl_fused_controllers updated: rollout_steps
  now consumes RL_ADV_VAR_PRE_NORM_INDEX (emitted by
  rl_advantage_normalize before in-place normalize) instead of the
  post-norm advantage_var_ratio (definitionally ~0 under Phase 4.5).
- ~30 new ISV bootstrap entries in with_controllers_bootstrapped.

# ISV slot allocation (isv_slots.rs)
- 72 new slots, RL_SLOTS_END 588 → 660. +288 bytes mapped-pinned.
- 9 Welford variance triples + 5 asymmetric Schulman counters +
  3 new input signals (var_pre_norm, gamma_min_adaptive, reward_mag) +
  v_blend (5 + 3 Welford) + ppo_ratio_clamp (1 + 3 Welford) +
  reward_clamp (5) + gate_threshold (1) + q_distill (1) + 20 clamp bounds +
  shared wiener floor.

# Bootstrap-clamp consistency fix (caught by Task 16 testing)
The original draft bootstrapped RL_EPS_BOOTSTRAP at 0.01 (= KL target),
but EPS_MIN was 0.05 — bootstrap value below clamp range. First post-
bootstrap step always snapped ε to MIN regardless of signal direction
("snap to MIN" behavior the unit tests surfaced). Fixed by bootstrapping
ε at MIN (0.05) so asymmetric Schulman operates from a valid state.

# Validation
- cargo build --release: clean
- cargo build --tests: clean
- 3 GPU Welford kernel tests (G1 constant→0var, G2 sequence→known var,
  G3 sentinel skip): all pass
- 5 GPU adaptive-floor invariant tests (G3 PPO clip holds at bootstrap
  when signal below floor, G4 tightens on single above-band, G5 widens
  only after 3 consecutive below-band, G6 post-warmup uses Welford mean,
  G6 pre-warmup uses EMA): all pass
- integrated_trainer_smoke (full GPU pipeline): passes
- 1k local smoke at b=128: 1000/1000 steps, completed_clean, no NaN,
  controllers genuinely adapting (γ 0.90→0.974, per_α 0.40→0.52,
  q_distill_λ 0.05→0.21, ε held at 0.05 = MIN per Phase 4.5 small-KL
  regime — was previously stuck at MAX 0.50 throughout fold 0/1)
- compute-sanitizer memcheck b=128 5 steps: 0 errors

Cluster validation (G7-G9) submitted as follow-up runs.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-30 15:41:51 +02:00

175 lines
8.8 KiB
Plaintext
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
// rl_rollout_steps_controller.cu — emits PPO rollout buffer flush size
// to ISV[RL_N_ROLLOUT_STEPS_INDEX=404].
//
// Phase E.3b of the integrated RL trainer
// (docs/superpowers/plans/2026-05-22-integrated-rl-trainer.md).
//
// The rollout length sets how many on-policy transitions PPO accumulates
// before flushing into an update batch. Per
// `pearl_controller_anchors_isv_driven` and
// `feedback_isv_for_adaptive_bounds`, this is NOT a hardcoded constant:
// it adapts to the noise level of the advantage estimator.
//
// Controller logic (Schulman-style bounded step):
// * input > ADV_VAR_RATIO_TARGET × TOLERANCE → noisy → widen rollout
// by ADJUST_RATE.
// * input < ADV_VAR_RATIO_TARGET / TOLERANCE → stable → shrink rollout
// by ADJUST_RATE.
// * in-band → hold.
//
// CRITICAL: the prior design used `scale = clamp(input/target, 0.5, 2.0)`
// which saturated to ±2× on the SIGN of (input target), not the
// magnitude. With typical streaming input 110 and target=0.1, every
// step doubled until the controller hit ROLLOUT_MAX in ≤4 steps
// (canonical: alpha-rl-kc2h9 fold0 had n_rollout pegged at MAX 100% of
// 50000 steps despite no actual signal change). Schulman-style bounded
// discrete adjustment converges smoothly to MAX/MIN when signal is
// consistently out-of-band, but doesn't slam there from one
// observation.
//
// Noise-floor gate holds the controller when input is below
// ADV_VAR_RATIO_NOISE_FLOOR = ADV_VAR_RATIO_TARGET × 0.01 — at that
// magnitude the streaming kernel's signal is dominated by numerical
// noise / cold-start, not a real "advantages are clean" indication
// worth shrinking the rollout for.
//
// Bootstrap discipline (per `pearl_first_observation_bootstrap`): the
// ISV slot starts at 0.0 sentinel. First emit writes
// `ROLLOUT_BOOTSTRAP = 2048` directly (the canonical PPO default).
// Subsequent emits Wiener-α blend with floor 0.4 (per
// `pearl_wiener_alpha_floor_for_nonstationary` — the advantage
// distribution drifts as π_new co-adapts, breaking stationarity).
//
// Bounds: rollout ∈ [256, 8192]. Below 256 PPO has too little data per
// update and the surrogate is dominated by sampling noise; above 8192
// updates lag the data so far behind that the importance ratios blow
// past the clip band.
#define RL_N_ROLLOUT_STEPS_INDEX 404
// ROLLOUT MIN/MAX clamp bounds are now ISV-driven per the 2026-05-30
// clamp-bound extension. Runtime-tunable + visible in diag.
#define RL_ROLLOUT_MIN_INDEX 643
#define RL_ROLLOUT_MAX_INDEX 644
// ISV-driven bootstrap + target.
#define RL_ROLLOUT_BOOTSTRAP_INDEX 475
#define RL_ADV_VAR_RATIO_TARGET_INDEX 449
// Schulman params from shared global slots (same as ppo_clip, target_tau).
#define RL_SCHULMAN_TOLERANCE_INDEX 468
#define RL_SCHULMAN_ADJUST_RATE_INDEX 469
// Wiener-α floor — shared across 9 controllers (slot 659).
#define RL_WIENER_ALPHA_FLOOR_INDEX 659
// Adaptive controller floors (spec 2026-05-30-adaptive-controller-floor-design):
// noise floor derives from observed advantage-variance signal via Welford
// triples, asymmetric Schulman widening requires N consecutive below-band
// observations. Replaces hardcoded ADV_VAR_RATIO_NOISE_FLOOR_FRAC=0.01f
// (calibrated against the pre-Phase-4.5 advantage regime; under Phase 4.5
// post-norm advantage variance is definitionally near-zero, so the
// controller is also rewired to consume RL_ADV_VAR_PRE_NORM_INDEX=612 —
// the pre-normalization variance emitted by rl_advantage_normalize.cu).
#define RL_ADV_VAR_VAR_COUNT_INDEX 596
#define RL_ADV_VAR_VAR_M2_INDEX 598
#define RL_ADV_VAR_BELOW_COUNT_INDEX 599
#define NOISE_FLOOR_TARGET_FRAC 0.5f // floor ≥ 50% of target
#define NOISE_FLOOR_STD_MULTIPLIER 2.0f // floor ≥ 2σ of observed signal
#define WIDEN_PATIENCE_CONSECUTIVE 3.0f // widen requires N below-band steps
// ─────────────────────────────────────────────────────────────────────
// rl_rollout_steps_controller:
// Single-thread controller — writes ONE float to
// isv[RL_N_ROLLOUT_STEPS_INDEX].
//
// Inputs:
// isv [≥ RL_N_ROLLOUT_STEPS_INDEX+1] — ISV bus
// alpha — Wiener-α from controller stats;
// floored at WIENER_ALPHA_FLOOR before
// the blend.
// advantage_var_over_abs_mean — caller-computed
// `var(A) / max(|mean A|, ε)` EMA;
// Phase F produces this via a small
// reduce kernel over the rollout
// buffer's advantage column.
//
// Outputs:
// isv[RL_N_ROLLOUT_STEPS_INDEX] — rollout length ∈ [ROLLOUT_MIN, ROLLOUT_MAX]
// ─────────────────────────────────────────────────────────────────────
// Phase R5: scalar input arg replaced with `input_slot` ISV index so the
// EMA producer (Phase R3 ema_update_per_step targeting
// ISV[RL_ADVANTAGE_VAR_RATIO_EMA_INDEX=421]) feeds this controller
// without any host roundtrip per `feedback_cpu_is_read_only`.
extern "C" __global__ void rl_rollout_steps_controller(
float* __restrict__ isv,
float alpha,
int input_slot
) {
if (threadIdx.x != 0 || blockIdx.x != 0) return;
const float prev = isv[RL_N_ROLLOUT_STEPS_INDEX];
// Bootstrap on sentinel 0.0 per pearl_first_observation_bootstrap.
if (prev == 0.0f) {
isv[RL_N_ROLLOUT_STEPS_INDEX] = isv[RL_ROLLOUT_BOOTSTRAP_INDEX];
return;
}
// Cold-start gate: sentinel-zero EMA (streaming kernel hasn't
// fired yet).
const float advantage_var_over_abs_mean = isv[input_slot];
if (advantage_var_over_abs_mean == 0.0f) return;
// ISV-driven regression target — read from
// RL_ADV_VAR_RATIO_TARGET_INDEX (seeded at trainer init).
const float adv_var_target = isv[RL_ADV_VAR_RATIO_TARGET_INDEX];
// Adaptive noise floor — signal-driven per
// `pearl_zscore_normalization_for_magnitude_asymmetric_signals` and
// `feedback_adaptive_not_tuned`. Welford sample variance =
// M² / (count 1) when count > 1.
const float adv_count = isv[RL_ADV_VAR_VAR_COUNT_INDEX];
const float adv_var = (adv_count > 1.0f)
? isv[RL_ADV_VAR_VAR_M2_INDEX] / (adv_count - 1.0f)
: 0.0f;
const float adv_std = sqrtf(adv_var);
const float adv_var_noise_floor = fmaxf(adv_var_target * NOISE_FLOOR_TARGET_FRAC,
adv_std * NOISE_FLOOR_STD_MULTIPLIER);
if (advantage_var_over_abs_mean < adv_var_noise_floor) return;
// Asymmetric Schulman (spec 2026-05-30 Section "Approach C, folded in"):
// widening rollout fires on a single above-band observation — noisy
// advantages are a safety signal we act on fast (collect more samples
// before updating). Shrinking rollout requires N consecutive below-band
// observations so a single quiet step can't push the buffer too small.
// Below-band counter is per-controller in ISV.
const float tolerance = isv[RL_SCHULMAN_TOLERANCE_INDEX];
const float adjust_rate = isv[RL_SCHULMAN_ADJUST_RATE_INDEX];
float scale;
if (advantage_var_over_abs_mean > adv_var_target * tolerance) {
scale = adjust_rate; // widen — single observation
isv[RL_ADV_VAR_BELOW_COUNT_INDEX] = 0.0f; // reset patience
} else if (advantage_var_over_abs_mean < adv_var_target / tolerance) {
const float new_count = isv[RL_ADV_VAR_BELOW_COUNT_INDEX] + 1.0f;
isv[RL_ADV_VAR_BELOW_COUNT_INDEX] = new_count;
scale = (new_count >= WIDEN_PATIENCE_CONSECUTIVE) ? (1.0f / adjust_rate) : 1.0f;
} else {
isv[RL_ADV_VAR_BELOW_COUNT_INDEX] = 0.0f; // in-band → reset
scale = 1.0f;
}
const float rollout_min = isv[RL_ROLLOUT_MIN_INDEX];
const float rollout_max = isv[RL_ROLLOUT_MAX_INDEX];
float target = prev * scale;
target = fmaxf(rollout_min, fminf(target, rollout_max));
// First-observation replace-directly per
// `pearl_first_observation_bootstrap`.
if (prev == isv[RL_ROLLOUT_BOOTSTRAP_INDEX]) {
isv[RL_N_ROLLOUT_STEPS_INDEX] = target;
return;
}
// Wiener-α blend with floor per pearl_wiener_alpha_floor_for_nonstationary.
const float a = fmaxf(alpha, isv[RL_WIENER_ALPHA_FLOOR_INDEX]);
float out = (1.0f - a) * prev + a * target;
out = fmaxf(rollout_min, fminf(out, rollout_max));
isv[RL_N_ROLLOUT_STEPS_INDEX] = out;
}