Files
foxhunt/crates/ml-alpha/cuda/rl_gamma_controller.cu
jgrusewski 083a88f7c3 feat(rl): adaptive controller floors — 12 controllers, all signal-driven
Replaces hardcoded thresholds AND clamp bounds across 12 RL controllers
with observed-signal-driven ISV-slot bounds. Eliminates the architectural
failure mode that surfaced in walk-forward fold 0 (alpha-rl-m9cx5) and
fold 1 (alpha-rl-jgdh6): under Phase 4.5 advantage normalization, eight
controllers (PPO clip, target_tau, rollout_steps, entropy_coef, per_α,
gamma, reward_scale, q_distill_lambda) saturated at extrema within ~50
steps and stayed pinned for the rest of training — same pattern across
two different data slices, confirming the bug is structural rather than
data-size-dependent.

Mechanism: each controller's noise floor was hardcoded as a small
fraction of its target (`_NOISE_FLOOR_FRAC = 0.01f`), calibrated against
a pre-Phase-4.5 signal regime. Phase 4.5 normalization reduces operating
KL by ~50× — observed signal stays below the 1%-of-target floor, the
Schulman widen path fires continuously (asymmetric in the wrong
direction), and ε hits MAX 0.50 within ~50 training steps. ksll2's full-
data run (n_folds=1) happened to escape via a single above-band KL
observation that triggered tighten; both walk-forward folds (3 and 6
files) did not.

Per `feedback_adaptive_not_tuned`, `feedback_isv_for_adaptive_bounds`,
`pearl_controller_anchors_isv_driven`: every threshold now derives from
observed signal statistics (Welford online variance) rather than
constants calibrated against a prior signal regime. User explicitly
expanded scope mid-implementation: "if all clamps ISV bound should be
added to this spec and tasks" — applying the principle consistently
means hardcoded MIN/MAX clamp bounds count too, not just the saturating
noise floors. 12 controllers, single atomic commit per
`feedback_no_partial_refactor`.

Spec: docs/superpowers/specs/2026-05-30-adaptive-controller-floor-design.md
Plan: docs/superpowers/plans/2026-05-30-adaptive-controller-floor-plan.md

# New kernel
- rl_signal_variance_update.cu (80 LOC) — Welford online variance.
  Single-thread single-block. Sentinel-zero skip per pearl. Per-controller
  Welford triple (count, mean, M²) drives every adaptive noise floor.

# Per-controller refactors (12 .cu files)
- ppo_clip + target_tau + rollout_steps + entropy_coef + per_α —
  adaptive noise floor = max(target × 0.5, sqrt(observed_var) × 2);
  asymmetric Schulman (tighten on single observation, widen requires 3
  consecutive below-band).
- gamma (Special G) — hardcoded GAMMA_MIN = 0.995 → adaptive via
  Welford MEAN of trade duration. Per spec Q6 resolution: the Welford
  mean's natural N-smoothing lag breaks the gamma↔trade_duration
  feedback loop without explicit step-period gating. 100-observation
  warmup falls back to EMA before Welford has enough samples.
- reward_scale (Special R) — asymmetric DECREASE rate cap (5% per
  step) on both bootstrap-replace and Wiener-blend paths; bootstrap-
  fraction floor (10% of bootstrap) until 100 trades close.
- q_distill_lambda (Special Q) — hardcoded MIN_LAMBDA = 0.05 → adaptive
  via Welford on q_distill_kl_ema: max(0.001, std × 0.05).
- v_blend_alpha (Phase 4.4) — 5 hardcoded constants → 5 ISV slots
  (DEAD_SIGNAL_FLOOR, TARGET_TRACK_RATIO, SCHULMAN_STEP, EMA_ALPHA,
  BOOTSTRAP_ALPHA) + adaptive dead-signal floor from V_scalar magnitude
  variance.
- ppo_ratio_clamp — adaptive MIN/MAX from observed log-ratio variance.
  Architectural 2.0 absolute floor preserved (don't degenerate to
  vanilla policy gradient).
- reward_clamp — V_BOUND_FLOOR, V_BOUND_EWMA_ALPHA, MIN_WIN, MIN_RATIO,
  MAX_RATIO, MIN/MAX_MARGIN, MARGIN_TOLERANCE/ADJUST_RATE,
  CLIP_RATE_EMA_ALPHA → 10 ISV slots.
- gate_threshold — hardcoded alpha = 0.01 → ISV slot 638.
- Clamp-bound expansion: EPS_MIN/MAX, TAU_MIN, ROLLOUT_MIN/MAX,
  COEF_MIN/MAX, PER_ALPHA_MIN/MAX, GAMMA_MAX, MAX_LAMBDA, KL_TOLERANCE,
  LAMBDA_RAMP_RATE, LAMBDA_DECAY_RATE → all ISV.
- WIENER_ALPHA_FLOOR shared across 9 controllers → single ISV slot 659.

# Trainer integration (integrated.rs)
- rl_signal_variance_update kernel loaded + helper method
  `launch_rl_signal_variance_update`.
- Per-step Welford launches for 9 controller inputs, placed between
  EMA producers and rl_fused_controllers in the per-step pipeline.
- Input-slot lookup array for rl_fused_controllers updated: rollout_steps
  now consumes RL_ADV_VAR_PRE_NORM_INDEX (emitted by
  rl_advantage_normalize before in-place normalize) instead of the
  post-norm advantage_var_ratio (definitionally ~0 under Phase 4.5).
- ~30 new ISV bootstrap entries in with_controllers_bootstrapped.

# ISV slot allocation (isv_slots.rs)
- 72 new slots, RL_SLOTS_END 588 → 660. +288 bytes mapped-pinned.
- 9 Welford variance triples + 5 asymmetric Schulman counters +
  3 new input signals (var_pre_norm, gamma_min_adaptive, reward_mag) +
  v_blend (5 + 3 Welford) + ppo_ratio_clamp (1 + 3 Welford) +
  reward_clamp (5) + gate_threshold (1) + q_distill (1) + 20 clamp bounds +
  shared wiener floor.

# Bootstrap-clamp consistency fix (caught by Task 16 testing)
The original draft bootstrapped RL_EPS_BOOTSTRAP at 0.01 (= KL target),
but EPS_MIN was 0.05 — bootstrap value below clamp range. First post-
bootstrap step always snapped ε to MIN regardless of signal direction
("snap to MIN" behavior the unit tests surfaced). Fixed by bootstrapping
ε at MIN (0.05) so asymmetric Schulman operates from a valid state.

# Validation
- cargo build --release: clean
- cargo build --tests: clean
- 3 GPU Welford kernel tests (G1 constant→0var, G2 sequence→known var,
  G3 sentinel skip): all pass
- 5 GPU adaptive-floor invariant tests (G3 PPO clip holds at bootstrap
  when signal below floor, G4 tightens on single above-band, G5 widens
  only after 3 consecutive below-band, G6 post-warmup uses Welford mean,
  G6 pre-warmup uses EMA): all pass
- integrated_trainer_smoke (full GPU pipeline): passes
- 1k local smoke at b=128: 1000/1000 steps, completed_clean, no NaN,
  controllers genuinely adapting (γ 0.90→0.974, per_α 0.40→0.52,
  q_distill_λ 0.05→0.21, ε held at 0.05 = MIN per Phase 4.5 small-KL
  regime — was previously stuck at MAX 0.50 throughout fold 0/1)
- compute-sanitizer memcheck b=128 5 steps: 0 errors

Cluster validation (G7-G9) submitted as follow-up runs.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-30 15:41:51 +02:00

163 lines
8.4 KiB
Plaintext
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
// rl_gamma_controller.cu — emits γ to ISV[RL_GAMMA_INDEX=400].
//
// Phase C of the integrated RL trainer
// (docs/superpowers/plans/2026-05-22-integrated-rl-trainer.md).
//
// γ is the Bellman discount factor used by the categorical Bellman
// projection (Phase E) and by the PPO advantage estimator (Phase D).
// Per `pearl_controller_anchors_isv_driven` and
// `feedback_isv_for_adaptive_bounds`, γ is NOT a hardcoded constant; it
// adapts so that
//
// γ^(mean_trade_duration_events) ≈ 0.5
//
// i.e. roughly half of the discounted value remains by the time the
// average trade closes. This anchors the credit-assignment horizon to
// the actual trade timescale rather than a tuned constant; if the
// strategy lengthens its holding period, γ rises automatically.
//
// Bootstrap discipline (per `pearl_first_observation_bootstrap`): the
// ISV slot starts at 0.0 (sentinel "uninitialised"). On the first
// emit the kernel DERIVES γ from the current input slot via the same
// target formula the per-step path uses, instead of writing a
// hardcoded `GAMMA_BOOTSTRAP = 0.99`. At sentinel input
// (trade_duration_ema = 0 → clamped d = 1) target = 0.5 → clamped to
// GAMMA_MIN = 0.90, so the cold-start γ is the floor. This eliminates
// the dead-zone where target(d ≈ 69) = 0.99 = hardcoded bootstrap froze
// the Wiener blend at canonical long-horizon γ for any realistic
// trade duration that happened to land near 69 events.
//
// Trade-off: cold-start γ = 0.90 (floor) is more myopic than the
// previous hardcoded 0.99. Once the trade_duration_ema stabilises
// (within ~10 episodes), the controller drifts γ up toward target
// (0.986 at d=50, 0.99 at d=69, etc.). The brief myopic warm-up is
// acceptable cost for guaranteed responsiveness — a frozen controller
// at canonical γ is worse than a controller that starts low and
// climbs.
//
// Subsequent emits use a Wiener-α blend with floor 0.4 per
// `pearl_wiener_alpha_floor_for_nonstationary` (the target γ_target
// drifts as the strategy's holding period co-adapts with the policy,
// which violates the stationarity precondition of Wiener-optimal α).
//
// Bounds: γ ∈ [0.90, 0.999] enforced by clamp to prevent runaway
// (γ → 1 makes credit assignment infinite-horizon; γ → 0.5 makes
// the policy myopic).
#define RL_GAMMA_INDEX 400
// γ MAX clamp ceiling is now ISV-driven per the 2026-05-30 clamp-bound
// extension. γ MIN already adaptive via RL_GAMMA_MIN_ADAPTIVE_INDEX (613).
#define RL_GAMMA_MAX_INDEX 649
// Wiener-α floor — shared across 9 controllers (slot 659).
#define RL_WIENER_ALPHA_FLOOR_INDEX 659
// Adaptive controller floors (spec 2026-05-30 Special case G):
// hardcoded `GAMMA_MIN = 0.995f` replaced with an adaptive bound derived
// from the Welford mean of trade duration. With short-horizon trade
// regimes (d_ema = 15.6 in fold 0) the static 0.995 floor pinned gamma
// for the entire run; adaptive bound `1 - 1/(d_smoothed × 2)` lets the
// controller drop gamma low enough for the current regime while keeping
// a hard `GAMMA_MIN_ABSOLUTE = 0.9` to prevent bandit-mode degenerate.
//
// Smoothed-duration choice: Welford mean (slot 609) is used after a
// 100-observation warmup so the gamma↔trade_duration feedback loop is
// damped by the natural N-smoothing of running mean. Before warmup, the
// EMA at slot 417 is used directly (less stable but available
// immediately at cold-start).
#define RL_GAMMA_MIN_ADAPTIVE_INDEX 613
#define RL_TRADE_DUR_VAR_COUNT_INDEX 608
#define RL_TRADE_DUR_VAR_MEAN_INDEX 609
#define RL_MEAN_TRADE_DURATION_EMA_INDEX 417
#define GAMMA_MIN_ABSOLUTE 0.9f
#define HORIZON_MULTIPLIER 2.0f
#define WELFORD_WARMUP_OBS 100.0f
// Floor for `d × HORIZON_MULTIPLIER` to prevent the `1/d` term from
// blowing up when trade duration is very small (cold start, no closes).
#define HORIZON_FLOOR 10.0f
// ─────────────────────────────────────────────────────────────────────
// rl_gamma_controller:
// Single-thread controller — writes ONE float to isv[RL_GAMMA_INDEX].
//
// Inputs:
// isv [≥ RL_GAMMA_INDEX+1] — ISV bus
// alpha — Wiener-α from the controller's own
// signal stats (caller computes this
// upstream from γ-divergence variance).
// Floored at WIENER_ALPHA_FLOOR before
// the blend.
// mean_trade_duration_events — EMA of trade hold time in event count.
// Caller is responsible for the EMA
// (Phase E); we just read the scalar.
//
// Outputs:
// isv[RL_GAMMA_INDEX] — γ ∈ [GAMMA_MIN, GAMMA_MAX]
// ─────────────────────────────────────────────────────────────────────
// Phase R5: scalar input arg replaced with `input_slot` ISV index so the
// EMA producer (Phase R3 ema_update_per_step targeting
// ISV[RL_MEAN_TRADE_DURATION_EMA_INDEX=417]) feeds this controller
// without any host roundtrip per `feedback_cpu_is_read_only`. At
// bootstrap (R1) the kernel returns before the input read, so `input_slot`
// can be the same ISV[417] sentinel-zero slot — the read just doesn't
// happen on that path.
extern "C" __global__ void rl_gamma_controller(
float* __restrict__ isv,
float alpha,
int input_slot
) {
if (threadIdx.x != 0 || blockIdx.x != 0) return;
const float gamma_prev = isv[RL_GAMMA_INDEX];
// Adaptive GAMMA_MIN (spec 2026-05-30 Special case G).
// Replaces hardcoded `GAMMA_MIN = 0.995f`. With d_ema = 15.6 (fold 0)
// the static 0.995 floor pinned γ for the entire run — the target
// formula `0.5^(1/d)` produces γ_target = 0.936 at d = 15.6, which
// clamps to 0.995 = floor and never moves. Adaptive bound
// `1 - 1/(d × 2)` gives γ_min ≈ 0.968 at d = 15.6 — leaving room
// for the controller to track the actual trade horizon.
//
// After 100 Welford observations the running mean is used (more
// stable than EMA against per-batch outliers); before that the EMA
// at slot 417 is used directly so cold-start has a usable signal.
const float d_welford_count = isv[RL_TRADE_DUR_VAR_COUNT_INDEX];
const float d_welford_mean = isv[RL_TRADE_DUR_VAR_MEAN_INDEX];
const float d_smoothed = (d_welford_count >= WELFORD_WARMUP_OBS)
? d_welford_mean
: isv[RL_MEAN_TRADE_DURATION_EMA_INDEX];
const float adaptive_min = fmaxf(GAMMA_MIN_ABSOLUTE,
1.0f - 1.0f / fmaxf(d_smoothed * HORIZON_MULTIPLIER,
HORIZON_FLOOR));
isv[RL_GAMMA_MIN_ADAPTIVE_INDEX] = adaptive_min;
// Compute target from the current input EMA. Shared between
// bootstrap and per-step paths so the dead-zone coincidence with
// a hardcoded bootstrap cannot recur.
// Target: γ^d ≈ 0.5 ⇒ γ = 0.5^(1/d). Clamp d ≥ 1 so a
// single-event trade doesn't push γ to 0.5.
const float mean_trade_duration_events = isv[input_slot];
const float d = fmaxf(mean_trade_duration_events, 1.0f);
const float gamma_max = isv[RL_GAMMA_MAX_INDEX];
float gamma_target = powf(0.5f, 1.0f / d);
gamma_target = fmaxf(adaptive_min, fminf(gamma_target, gamma_max));
// Bootstrap on sentinel 0.0 per pearl_first_observation_bootstrap:
// first emit replaces directly with the computed target. At cold
// start (input EMA also sentinel-zero), clamped d=1, target=0.5,
// clamped to adaptive_min (≥ GAMMA_MIN_ABSOLUTE = 0.90). Any
// non-sentinel input produces target ≥ adaptive_min → ≤ 0.999.
if (gamma_prev == 0.0f) {
isv[RL_GAMMA_INDEX] = gamma_target;
return;
}
// Wiener-α blend with floor per pearl_wiener_alpha_floor_for_nonstationary.
const float a = fmaxf(alpha, isv[RL_WIENER_ALPHA_FLOOR_INDEX]);
float gamma_new = (1.0f - a) * gamma_prev + a * gamma_target;
// Clamp into the bounded range; runaway protection.
gamma_new = fmaxf(adaptive_min, fminf(gamma_new, gamma_max));
isv[RL_GAMMA_INDEX] = gamma_new;
}