fix(rl): bounded multiplicative step + noise-floor on rollout_steps + per_α

kc2h9 confirmed: clamping streaming-kernel outputs to [≤100, ≤30]
had ZERO behavioral impact because rl_rollout_steps_controller's
prior design used `scale = clamp(input/target, 0.5, 2.0)` — the
scale saturated to ±2× on the SIGN of (input − target), not the
magnitude. With target=0.1 and typical input=1–10 the controller
slammed to MAX in ≤4 steps regardless of whether input was 4 or
3e5. Bit-identical losses between gxhr8 and kc2h9 confirmed the
saturation.

## Fix 1: rl_rollout_steps_controller — same Schulman pattern as ppo_clip

  * input > TARGET × 1.5     → scale = 1.5      (widen)
  * input < TARGET / 1.5     → scale = 1/1.5    (shrink)
  * in-band                   → scale = 1.0      (hold)
  * input < TARGET × 0.01    → return            (noise floor — hold prev)

Per-step adjustment bounded at 1.5×, so rollout_steps drifts
smoothly toward MIN/MAX rather than slamming there. The noise-floor
gate matches the pattern from
`pearl_multiplicative_controllers_need_bounded_step_and_noise_floor`
applied to the ε and τ controllers earlier in R9.

## Fix 2: rl_per_alpha_controller — noise-floor gate (defensive)

per_α uses a LINEAR lift `0.4 + 0.2·(kurt-3)/7` (not multiplicative),
so it doesn't have the saturation bug. But added a noise-floor gate
at KURT_NOISE_FLOOR = 1.0 so a sub-Gaussian kurtosis reading from
the streaming estimator's startup window (when per-step batch-mean
deviations are small before tails develop) doesn't drag α toward
PER_ALPHA_MIN on cold-start.

## Diag bake-in (per user request "bake in diags")

JSONL gains a `controller_branch` block exposing the
multiplicative-controller inputs alongside their design targets:

  controller_branch: {
    rollout_steps_input:   isv[421],   rollout_steps_target:   0.1,
    ppo_clip_input:        isv[419],   ppo_clip_target:        0.01,
    target_tau_input:      isv[418],   target_tau_target:      0.01,
    per_alpha_input:       isv[422],   per_alpha_target:       0.6,
  }

Post-hoc analysis can compute the branch each step (WIDEN / HOLD /
SHRINK / NOISE) by comparing input/target against the ±33%
tolerance band, revealing whether each controller is being driven
by real signal or sitting in the in-band hold zone. Targets are
reflected from the kernel #defines (synchronised by code review at
the controller-cu file level — there's no ISV slot for these
design constants because they're fundamental to the controller's
behaviour, not adaptive).

## Verified gates (local sm_86)

  G1 isv_bootstrap   
  G3 controllers     
  G4 target_update   
  integrated_smoke   

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
This commit is contained in:
jgrusewski
2026-05-23 22:53:10 +02:00
parent 66115007ab
commit 708c121f20
3 changed files with 120 additions and 28 deletions

View File

@@ -52,6 +52,16 @@
// market data), the lift fully spans 0.2 of the [PER_ALPHA_MIN,
// PER_ALPHA_MAX] band. Beyond ~30 the mapping saturates at α_max.
#define KURT_LIFT_SCALE 7.0f
// Noise-floor gate: if the streaming kurtosis estimator emits a value
// below this magnitude, treat it as "no signal" (sentinel-zero proxy)
// and hold α at the prior value. Without this, on the first few steps
// after the streaming estimator initialises, a sub-Gaussian kurtosis
// reading (e.g. 0.5 — near-uniform per-step batch means before tails
// develop) would yield target = 0.4 = PER_ALPHA_MIN, dragging α toward
// MIN despite zero real signal. Matches the defensive noise-floor
// pattern on the other controllers (per
// pearl_multiplicative_controllers_need_bounded_step_and_noise_floor).
#define KURT_NOISE_FLOOR 1.0f
#define WIENER_ALPHA_FLOOR 0.4f
// ─────────────────────────────────────────────────────────────────────
@@ -87,6 +97,19 @@ extern "C" __global__ void rl_per_alpha_controller(
const float prev = isv[RL_PER_ALPHA_INDEX];
// Noise-floor gate: kurtosis below KURT_NOISE_FLOOR is dominated
// by streaming-estimator startup noise (per-step batch-mean
// differences before tails accumulate). Hold α at prev to avoid
// dragging toward PER_ALPHA_MIN on cold-start.
const float td_kurtosis_ema = isv[input_slot];
if (td_kurtosis_ema > 0.0f && td_kurtosis_ema < KURT_NOISE_FLOOR) {
// Real signal present but below the noise floor — hold prev.
// (Strict sentinel zero handled by the prev==0 bootstrap path
// below, which derives target from the current EMA so the
// first-non-zero observation still seeds correctly.)
if (prev != 0.0f) return;
}
// Compute target from the current input EMA via a piecewise linear
// lift. Shared between bootstrap and per-step paths so the dead-zone
// coincidence with a hardcoded bootstrap value cannot recur.
@@ -97,7 +120,6 @@ extern "C" __global__ void rl_per_alpha_controller(
// The 0.4-0.6 baseline keeps the steady-state output near PER's
// canonical 0.6 (when input EMA stabilises at kurt=10) while
// leaving headroom to lift toward 1.0 under heavy tails.
const float td_kurtosis_ema = isv[input_slot];
const float kurt_excess = fmaxf(0.0f, td_kurtosis_ema - KURT_GAUSSIAN);
float target = 0.4f + 0.2f * (kurt_excess / KURT_LIFT_SCALE);
target = fmaxf(PER_ALPHA_MIN, fminf(target, PER_ALPHA_MAX));

View File

@@ -10,12 +10,28 @@
// `feedback_isv_for_adaptive_bounds`, this is NOT a hardcoded constant:
// it adapts to the noise level of the advantage estimator.
//
// Controller logic:
// * High var(A) / |mean A| → advantage estimates are noisy → grow
// rollout length so PPO has more samples per update and the
// surrogate gradient averages over the noise.
// * Low var(A) / |mean A| → advantages are stable → shrink rollout
// so the data is fresher and policy update lag stays small.
// Controller logic (Schulman-style bounded step):
// * input > ADV_VAR_RATIO_TARGET × TOLERANCE → noisy → widen rollout
// by ADJUST_RATE.
// * input < ADV_VAR_RATIO_TARGET / TOLERANCE → stable → shrink rollout
// by ADJUST_RATE.
// * in-band → hold.
//
// CRITICAL: the prior design used `scale = clamp(input/target, 0.5, 2.0)`
// which saturated to ±2× on the SIGN of (input target), not the
// magnitude. With typical streaming input 110 and target=0.1, every
// step doubled until the controller hit ROLLOUT_MAX in ≤4 steps
// (canonical: alpha-rl-kc2h9 fold0 had n_rollout pegged at MAX 100% of
// 50000 steps despite no actual signal change). Schulman-style bounded
// discrete adjustment converges smoothly to MAX/MIN when signal is
// consistently out-of-band, but doesn't slam there from one
// observation.
//
// Noise-floor gate holds the controller when input is below
// ADV_VAR_RATIO_NOISE_FLOOR = ADV_VAR_RATIO_TARGET × 0.01 — at that
// magnitude the streaming kernel's signal is dominated by numerical
// noise / cold-start, not a real "advantages are clean" indication
// worth shrinking the rollout for.
//
// Bootstrap discipline (per `pearl_first_observation_bootstrap`): the
// ISV slot starts at 0.0 sentinel. First emit writes
@@ -29,12 +45,22 @@
// updates lag the data so far behind that the importance ratios blow
// past the clip band.
#define RL_N_ROLLOUT_STEPS_INDEX 404
#define ROLLOUT_MIN 256.0f
#define ROLLOUT_MAX 8192.0f
#define ROLLOUT_BOOTSTRAP 2048.0f
#define ADV_VAR_RATIO_TARGET 0.1f
#define WIENER_ALPHA_FLOOR 0.4f
#define RL_N_ROLLOUT_STEPS_INDEX 404
#define ROLLOUT_MIN 256.0f
#define ROLLOUT_MAX 8192.0f
#define ROLLOUT_BOOTSTRAP 2048.0f
#define ADV_VAR_RATIO_TARGET 0.1f
// Below this magnitude the streaming kernel's signal is numerical
// noise / cold-start, not a real divergence-from-target — hold.
#define ADV_VAR_RATIO_NOISE_FLOOR (ADV_VAR_RATIO_TARGET * 0.01f) // = 1e-3
// In-band tolerance: input within ±33 % of target → hold.
#define ADV_VAR_RATIO_TOLERANCE 1.5f
// Per-step bounded multiplicative adjustment. 1.5× == 50% growth /
// 33% shrinkage per fire. Prevents one observation from sending
// rollout from BOOTSTRAP (2048) to MAX (8192) in 2 steps as the
// prior 2.0× clamp design did.
#define ADV_VAR_RATIO_ADJUST_RATE 1.5f
#define WIENER_ALPHA_FLOOR 0.4f
// ─────────────────────────────────────────────────────────────────────
// rl_rollout_steps_controller:
@@ -73,23 +99,36 @@ extern "C" __global__ void rl_rollout_steps_controller(
return;
}
// Cold-start gate: when advantage_var_ratio
// EMA is sentinel-zero (no rollout has accumulated advantages
// yet), the multiplicative ratio collapses to 0, which makes
// scale clamp to 0.5 and target halve every step. Within ~6
// cold-start steps the rollout length has migrated from 2048 →
// ~430 despite zero real signal. Hold at bootstrap until first
// observation.
// Cold-start gate: sentinel-zero EMA (streaming kernel hasn't
// fired yet).
const float advantage_var_over_abs_mean = isv[input_slot];
if (advantage_var_over_abs_mean == 0.0f) return;
// Multiplicative adaptation: if the var-ratio exceeds the target
// (noisy advantages), scale up the rollout; if it falls below, scale
// down. Clamped to [0.5, 2.0] per step so we never double-halve in a
// single emit (per pearl_blend_formulas_must_have_permanent_floor:
// step-size bounds even before the Wiener blend).
const float ratio_raw = advantage_var_over_abs_mean / ADV_VAR_RATIO_TARGET;
const float scale = fmaxf(0.5f, fminf(2.0f, ratio_raw));
float target = prev * scale;
// Noise-floor gate: input below ADV_VAR_RATIO_NOISE_FLOOR is
// dominated by numerical noise — hold rollout unchanged.
// Mirrors the ppo_clip / target_tau controllers' design after the
// alpha-rl-mjzfk multiplicative-blow-up incident.
if (advantage_var_over_abs_mean < ADV_VAR_RATIO_NOISE_FLOOR) return;
// Bounded multiplicative adjustment (Schulman-style adaptive).
// At most ADV_VAR_RATIO_ADJUST_RATE × shift per step regardless of
// how far input is from target. After several consecutive
// out-of-band observations the output drifts smoothly toward
// MIN/MAX, but a single observation can't slam it there.
float scale;
if (advantage_var_over_abs_mean
> ADV_VAR_RATIO_TARGET * ADV_VAR_RATIO_TOLERANCE) {
// Noisy → widen.
scale = ADV_VAR_RATIO_ADJUST_RATE;
} else if (advantage_var_over_abs_mean
< ADV_VAR_RATIO_TARGET / ADV_VAR_RATIO_TOLERANCE) {
// Clean → shrink.
scale = 1.0f / ADV_VAR_RATIO_ADJUST_RATE;
} else {
// In-band: hold.
scale = 1.0f;
}
float target = prev * scale;
target = fmaxf(ROLLOUT_MIN, fminf(target, ROLLOUT_MAX));
// First-observation replace-directly per

View File

@@ -603,6 +603,37 @@ fn main() -> Result<()> {
"clamp": isv[RL_TD_KURTOSIS_CLAMP_INDEX],
},
},
// R9 — controller decision diagnostics (which adaptation
// branch the multiplicative controllers fired this step).
//
// ratio = input_ema / TARGET (per controller's design)
// branch = "WIDEN" if ratio > TOLERANCE = 1.5
// "HOLD" if ratio ∈ [1/1.5, 1.5]
// "SHRINK" if ratio < 1/1.5
// "NOISE" if input_ema is below the noise floor
//
// For rl_rollout_steps_controller (TARGET = 0.1) and
// rl_ppo_clip_controller (TARGET = 0.01), this reveals
// whether the controller is being driven by real divergence
// signal or sitting in the in-band hold zone. Use to spot
// controllers stuck saturating at MAX/MIN (which would
// show 100% WIDEN or 100% SHRINK).
//
// Targets and tolerances are reflected from the kernel
// #defines (kept synchronised by code review at the
// controller-cu file level — there's no ISV slot for these
// design constants because they're fundamental to the
// controller's behaviour, not adaptive).
"controller_branch": {
"rollout_steps_input": isv[RL_ADVANTAGE_VAR_RATIO_EMA_INDEX],
"rollout_steps_target": 0.1f32,
"ppo_clip_input": isv[RL_KL_PI_EMA_INDEX],
"ppo_clip_target": 0.01f32,
"target_tau_input": isv[RL_Q_DIVERGENCE_EMA_INDEX],
"target_tau_target": 0.01f32,
"per_alpha_input": isv[RL_TD_KURTOSIS_EMA_INDEX],
"per_alpha_target": 0.6f32, // kurt=10 fixed point
},
"done_count": done_count,
// Action histogram: indices match `Action` enum
// (0=ShortLarge, 1=ShortSmall, 2=Hold, 3=FlatFromLong,