Replaces hardcoded thresholds AND clamp bounds across 12 RL controllers
with observed-signal-driven ISV-slot bounds. Eliminates the architectural
failure mode that surfaced in walk-forward fold 0 (alpha-rl-m9cx5) and
fold 1 (alpha-rl-jgdh6): under Phase 4.5 advantage normalization, eight
controllers (PPO clip, target_tau, rollout_steps, entropy_coef, per_α,
gamma, reward_scale, q_distill_lambda) saturated at extrema within ~50
steps and stayed pinned for the rest of training — same pattern across
two different data slices, confirming the bug is structural rather than
data-size-dependent.
Mechanism: each controller's noise floor was hardcoded as a small
fraction of its target (`_NOISE_FLOOR_FRAC = 0.01f`), calibrated against
a pre-Phase-4.5 signal regime. Phase 4.5 normalization reduces operating
KL by ~50× — observed signal stays below the 1%-of-target floor, the
Schulman widen path fires continuously (asymmetric in the wrong
direction), and ε hits MAX 0.50 within ~50 training steps. ksll2's full-
data run (n_folds=1) happened to escape via a single above-band KL
observation that triggered tighten; both walk-forward folds (3 and 6
files) did not.
Per `feedback_adaptive_not_tuned`, `feedback_isv_for_adaptive_bounds`,
`pearl_controller_anchors_isv_driven`: every threshold now derives from
observed signal statistics (Welford online variance) rather than
constants calibrated against a prior signal regime. User explicitly
expanded scope mid-implementation: "if all clamps ISV bound should be
added to this spec and tasks" — applying the principle consistently
means hardcoded MIN/MAX clamp bounds count too, not just the saturating
noise floors. 12 controllers, single atomic commit per
`feedback_no_partial_refactor`.
Spec: docs/superpowers/specs/2026-05-30-adaptive-controller-floor-design.md
Plan: docs/superpowers/plans/2026-05-30-adaptive-controller-floor-plan.md
# New kernel
- rl_signal_variance_update.cu (80 LOC) — Welford online variance.
Single-thread single-block. Sentinel-zero skip per pearl. Per-controller
Welford triple (count, mean, M²) drives every adaptive noise floor.
# Per-controller refactors (12 .cu files)
- ppo_clip + target_tau + rollout_steps + entropy_coef + per_α —
adaptive noise floor = max(target × 0.5, sqrt(observed_var) × 2);
asymmetric Schulman (tighten on single observation, widen requires 3
consecutive below-band).
- gamma (Special G) — hardcoded GAMMA_MIN = 0.995 → adaptive via
Welford MEAN of trade duration. Per spec Q6 resolution: the Welford
mean's natural N-smoothing lag breaks the gamma↔trade_duration
feedback loop without explicit step-period gating. 100-observation
warmup falls back to EMA before Welford has enough samples.
- reward_scale (Special R) — asymmetric DECREASE rate cap (5% per
step) on both bootstrap-replace and Wiener-blend paths; bootstrap-
fraction floor (10% of bootstrap) until 100 trades close.
- q_distill_lambda (Special Q) — hardcoded MIN_LAMBDA = 0.05 → adaptive
via Welford on q_distill_kl_ema: max(0.001, std × 0.05).
- v_blend_alpha (Phase 4.4) — 5 hardcoded constants → 5 ISV slots
(DEAD_SIGNAL_FLOOR, TARGET_TRACK_RATIO, SCHULMAN_STEP, EMA_ALPHA,
BOOTSTRAP_ALPHA) + adaptive dead-signal floor from V_scalar magnitude
variance.
- ppo_ratio_clamp — adaptive MIN/MAX from observed log-ratio variance.
Architectural 2.0 absolute floor preserved (don't degenerate to
vanilla policy gradient).
- reward_clamp — V_BOUND_FLOOR, V_BOUND_EWMA_ALPHA, MIN_WIN, MIN_RATIO,
MAX_RATIO, MIN/MAX_MARGIN, MARGIN_TOLERANCE/ADJUST_RATE,
CLIP_RATE_EMA_ALPHA → 10 ISV slots.
- gate_threshold — hardcoded alpha = 0.01 → ISV slot 638.
- Clamp-bound expansion: EPS_MIN/MAX, TAU_MIN, ROLLOUT_MIN/MAX,
COEF_MIN/MAX, PER_ALPHA_MIN/MAX, GAMMA_MAX, MAX_LAMBDA, KL_TOLERANCE,
LAMBDA_RAMP_RATE, LAMBDA_DECAY_RATE → all ISV.
- WIENER_ALPHA_FLOOR shared across 9 controllers → single ISV slot 659.
# Trainer integration (integrated.rs)
- rl_signal_variance_update kernel loaded + helper method
`launch_rl_signal_variance_update`.
- Per-step Welford launches for 9 controller inputs, placed between
EMA producers and rl_fused_controllers in the per-step pipeline.
- Input-slot lookup array for rl_fused_controllers updated: rollout_steps
now consumes RL_ADV_VAR_PRE_NORM_INDEX (emitted by
rl_advantage_normalize before in-place normalize) instead of the
post-norm advantage_var_ratio (definitionally ~0 under Phase 4.5).
- ~30 new ISV bootstrap entries in with_controllers_bootstrapped.
# ISV slot allocation (isv_slots.rs)
- 72 new slots, RL_SLOTS_END 588 → 660. +288 bytes mapped-pinned.
- 9 Welford variance triples + 5 asymmetric Schulman counters +
3 new input signals (var_pre_norm, gamma_min_adaptive, reward_mag) +
v_blend (5 + 3 Welford) + ppo_ratio_clamp (1 + 3 Welford) +
reward_clamp (5) + gate_threshold (1) + q_distill (1) + 20 clamp bounds +
shared wiener floor.
# Bootstrap-clamp consistency fix (caught by Task 16 testing)
The original draft bootstrapped RL_EPS_BOOTSTRAP at 0.01 (= KL target),
but EPS_MIN was 0.05 — bootstrap value below clamp range. First post-
bootstrap step always snapped ε to MIN regardless of signal direction
("snap to MIN" behavior the unit tests surfaced). Fixed by bootstrapping
ε at MIN (0.05) so asymmetric Schulman operates from a valid state.
# Validation
- cargo build --release: clean
- cargo build --tests: clean
- 3 GPU Welford kernel tests (G1 constant→0var, G2 sequence→known var,
G3 sentinel skip): all pass
- 5 GPU adaptive-floor invariant tests (G3 PPO clip holds at bootstrap
when signal below floor, G4 tightens on single above-band, G5 widens
only after 3 consecutive below-band, G6 post-warmup uses Welford mean,
G6 pre-warmup uses EMA): all pass
- integrated_trainer_smoke (full GPU pipeline): passes
- 1k local smoke at b=128: 1000/1000 steps, completed_clean, no NaN,
controllers genuinely adapting (γ 0.90→0.974, per_α 0.40→0.52,
q_distill_λ 0.05→0.21, ε held at 0.05 = MIN per Phase 4.5 small-KL
regime — was previously stuck at MAX 0.50 throughout fold 0/1)
- compute-sanitizer memcheck b=128 5 steps: 0 errors
Cluster validation (G7-G9) submitted as follow-up runs.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
163 lines
8.4 KiB
Plaintext
163 lines
8.4 KiB
Plaintext
// rl_gamma_controller.cu — emits γ to ISV[RL_GAMMA_INDEX=400].
|
||
//
|
||
// Phase C of the integrated RL trainer
|
||
// (docs/superpowers/plans/2026-05-22-integrated-rl-trainer.md).
|
||
//
|
||
// γ is the Bellman discount factor used by the categorical Bellman
|
||
// projection (Phase E) and by the PPO advantage estimator (Phase D).
|
||
// Per `pearl_controller_anchors_isv_driven` and
|
||
// `feedback_isv_for_adaptive_bounds`, γ is NOT a hardcoded constant; it
|
||
// adapts so that
|
||
//
|
||
// γ^(mean_trade_duration_events) ≈ 0.5
|
||
//
|
||
// i.e. roughly half of the discounted value remains by the time the
|
||
// average trade closes. This anchors the credit-assignment horizon to
|
||
// the actual trade timescale rather than a tuned constant; if the
|
||
// strategy lengthens its holding period, γ rises automatically.
|
||
//
|
||
// Bootstrap discipline (per `pearl_first_observation_bootstrap`): the
|
||
// ISV slot starts at 0.0 (sentinel "uninitialised"). On the first
|
||
// emit the kernel DERIVES γ from the current input slot via the same
|
||
// target formula the per-step path uses, instead of writing a
|
||
// hardcoded `GAMMA_BOOTSTRAP = 0.99`. At sentinel input
|
||
// (trade_duration_ema = 0 → clamped d = 1) target = 0.5 → clamped to
|
||
// GAMMA_MIN = 0.90, so the cold-start γ is the floor. This eliminates
|
||
// the dead-zone where target(d ≈ 69) = 0.99 = hardcoded bootstrap froze
|
||
// the Wiener blend at canonical long-horizon γ for any realistic
|
||
// trade duration that happened to land near 69 events.
|
||
//
|
||
// Trade-off: cold-start γ = 0.90 (floor) is more myopic than the
|
||
// previous hardcoded 0.99. Once the trade_duration_ema stabilises
|
||
// (within ~10 episodes), the controller drifts γ up toward target
|
||
// (0.986 at d=50, 0.99 at d=69, etc.). The brief myopic warm-up is
|
||
// acceptable cost for guaranteed responsiveness — a frozen controller
|
||
// at canonical γ is worse than a controller that starts low and
|
||
// climbs.
|
||
//
|
||
// Subsequent emits use a Wiener-α blend with floor 0.4 per
|
||
// `pearl_wiener_alpha_floor_for_nonstationary` (the target γ_target
|
||
// drifts as the strategy's holding period co-adapts with the policy,
|
||
// which violates the stationarity precondition of Wiener-optimal α).
|
||
//
|
||
// Bounds: γ ∈ [0.90, 0.999] enforced by clamp to prevent runaway
|
||
// (γ → 1 makes credit assignment infinite-horizon; γ → 0.5 makes
|
||
// the policy myopic).
|
||
|
||
#define RL_GAMMA_INDEX 400
|
||
// γ MAX clamp ceiling is now ISV-driven per the 2026-05-30 clamp-bound
|
||
// extension. γ MIN already adaptive via RL_GAMMA_MIN_ADAPTIVE_INDEX (613).
|
||
#define RL_GAMMA_MAX_INDEX 649
|
||
// Wiener-α floor — shared across 9 controllers (slot 659).
|
||
#define RL_WIENER_ALPHA_FLOOR_INDEX 659
|
||
|
||
// Adaptive controller floors (spec 2026-05-30 Special case G):
|
||
// hardcoded `GAMMA_MIN = 0.995f` replaced with an adaptive bound derived
|
||
// from the Welford mean of trade duration. With short-horizon trade
|
||
// regimes (d_ema = 15.6 in fold 0) the static 0.995 floor pinned gamma
|
||
// for the entire run; adaptive bound `1 - 1/(d_smoothed × 2)` lets the
|
||
// controller drop gamma low enough for the current regime while keeping
|
||
// a hard `GAMMA_MIN_ABSOLUTE = 0.9` to prevent bandit-mode degenerate.
|
||
//
|
||
// Smoothed-duration choice: Welford mean (slot 609) is used after a
|
||
// 100-observation warmup so the gamma↔trade_duration feedback loop is
|
||
// damped by the natural N-smoothing of running mean. Before warmup, the
|
||
// EMA at slot 417 is used directly (less stable but available
|
||
// immediately at cold-start).
|
||
#define RL_GAMMA_MIN_ADAPTIVE_INDEX 613
|
||
#define RL_TRADE_DUR_VAR_COUNT_INDEX 608
|
||
#define RL_TRADE_DUR_VAR_MEAN_INDEX 609
|
||
#define RL_MEAN_TRADE_DURATION_EMA_INDEX 417
|
||
#define GAMMA_MIN_ABSOLUTE 0.9f
|
||
#define HORIZON_MULTIPLIER 2.0f
|
||
#define WELFORD_WARMUP_OBS 100.0f
|
||
// Floor for `d × HORIZON_MULTIPLIER` to prevent the `1/d` term from
|
||
// blowing up when trade duration is very small (cold start, no closes).
|
||
#define HORIZON_FLOOR 10.0f
|
||
|
||
|
||
// ─────────────────────────────────────────────────────────────────────
|
||
// rl_gamma_controller:
|
||
// Single-thread controller — writes ONE float to isv[RL_GAMMA_INDEX].
|
||
//
|
||
// Inputs:
|
||
// isv [≥ RL_GAMMA_INDEX+1] — ISV bus
|
||
// alpha — Wiener-α from the controller's own
|
||
// signal stats (caller computes this
|
||
// upstream from γ-divergence variance).
|
||
// Floored at WIENER_ALPHA_FLOOR before
|
||
// the blend.
|
||
// mean_trade_duration_events — EMA of trade hold time in event count.
|
||
// Caller is responsible for the EMA
|
||
// (Phase E); we just read the scalar.
|
||
//
|
||
// Outputs:
|
||
// isv[RL_GAMMA_INDEX] — γ ∈ [GAMMA_MIN, GAMMA_MAX]
|
||
// ─────────────────────────────────────────────────────────────────────
|
||
// Phase R5: scalar input arg replaced with `input_slot` ISV index so the
|
||
// EMA producer (Phase R3 ema_update_per_step targeting
|
||
// ISV[RL_MEAN_TRADE_DURATION_EMA_INDEX=417]) feeds this controller
|
||
// without any host roundtrip per `feedback_cpu_is_read_only`. At
|
||
// bootstrap (R1) the kernel returns before the input read, so `input_slot`
|
||
// can be the same ISV[417] sentinel-zero slot — the read just doesn't
|
||
// happen on that path.
|
||
extern "C" __global__ void rl_gamma_controller(
|
||
float* __restrict__ isv,
|
||
float alpha,
|
||
int input_slot
|
||
) {
|
||
if (threadIdx.x != 0 || blockIdx.x != 0) return;
|
||
|
||
const float gamma_prev = isv[RL_GAMMA_INDEX];
|
||
|
||
// Adaptive GAMMA_MIN (spec 2026-05-30 Special case G).
|
||
// Replaces hardcoded `GAMMA_MIN = 0.995f`. With d_ema = 15.6 (fold 0)
|
||
// the static 0.995 floor pinned γ for the entire run — the target
|
||
// formula `0.5^(1/d)` produces γ_target = 0.936 at d = 15.6, which
|
||
// clamps to 0.995 = floor and never moves. Adaptive bound
|
||
// `1 - 1/(d × 2)` gives γ_min ≈ 0.968 at d = 15.6 — leaving room
|
||
// for the controller to track the actual trade horizon.
|
||
//
|
||
// After 100 Welford observations the running mean is used (more
|
||
// stable than EMA against per-batch outliers); before that the EMA
|
||
// at slot 417 is used directly so cold-start has a usable signal.
|
||
const float d_welford_count = isv[RL_TRADE_DUR_VAR_COUNT_INDEX];
|
||
const float d_welford_mean = isv[RL_TRADE_DUR_VAR_MEAN_INDEX];
|
||
const float d_smoothed = (d_welford_count >= WELFORD_WARMUP_OBS)
|
||
? d_welford_mean
|
||
: isv[RL_MEAN_TRADE_DURATION_EMA_INDEX];
|
||
const float adaptive_min = fmaxf(GAMMA_MIN_ABSOLUTE,
|
||
1.0f - 1.0f / fmaxf(d_smoothed * HORIZON_MULTIPLIER,
|
||
HORIZON_FLOOR));
|
||
isv[RL_GAMMA_MIN_ADAPTIVE_INDEX] = adaptive_min;
|
||
|
||
// Compute target from the current input EMA. Shared between
|
||
// bootstrap and per-step paths so the dead-zone coincidence with
|
||
// a hardcoded bootstrap cannot recur.
|
||
// Target: γ^d ≈ 0.5 ⇒ γ = 0.5^(1/d). Clamp d ≥ 1 so a
|
||
// single-event trade doesn't push γ to 0.5.
|
||
const float mean_trade_duration_events = isv[input_slot];
|
||
const float d = fmaxf(mean_trade_duration_events, 1.0f);
|
||
const float gamma_max = isv[RL_GAMMA_MAX_INDEX];
|
||
float gamma_target = powf(0.5f, 1.0f / d);
|
||
gamma_target = fmaxf(adaptive_min, fminf(gamma_target, gamma_max));
|
||
|
||
// Bootstrap on sentinel 0.0 per pearl_first_observation_bootstrap:
|
||
// first emit replaces directly with the computed target. At cold
|
||
// start (input EMA also sentinel-zero), clamped d=1, target=0.5,
|
||
// clamped to adaptive_min (≥ GAMMA_MIN_ABSOLUTE = 0.90). Any
|
||
// non-sentinel input produces target ≥ adaptive_min → ≤ 0.999.
|
||
if (gamma_prev == 0.0f) {
|
||
isv[RL_GAMMA_INDEX] = gamma_target;
|
||
return;
|
||
}
|
||
|
||
// Wiener-α blend with floor per pearl_wiener_alpha_floor_for_nonstationary.
|
||
const float a = fmaxf(alpha, isv[RL_WIENER_ALPHA_FLOOR_INDEX]);
|
||
float gamma_new = (1.0f - a) * gamma_prev + a * gamma_target;
|
||
|
||
// Clamp into the bounded range; runaway protection.
|
||
gamma_new = fmaxf(adaptive_min, fminf(gamma_new, gamma_max));
|
||
isv[RL_GAMMA_INDEX] = gamma_new;
|
||
}
|