feat(sp11): B1b — structural reward composition refactor (production flip)
Per spec §3.5.3 amended at 7ddaf9c51 on main: experience_env_step
reward composition decomposed from 8+ inline accumulation sites into
explicit per-component locals (r_popart, r_cf, r_trail, r_micro,
r_opp_cost, r_bonus), then composed as Σ w_i × r_i with controller
weights from ISV[340..346).
Trail reward extraction (§3.5.4): trail-fire P&L now flows through
r_trail (forced-exit signal) instead of r_popart (voluntary-exit
signal). REWARD_TRAIL_WEIGHT_INDEX has real signal — controller can
weight forced-exit vs voluntary-exit P&L differently. rc[2] (the
prior structural-placeholder slot) now carries trail magnitude.
Universal post-composition modifiers (§3.4.4): drawdown / capital-
floor / inventory / churn / conviction-scale / cf-flip apply AFTER
the weighted Σ, unweighted. They are risk constraints and structural
operators, NOT learning components — agent cannot weigh them away.
Mean=1 normalization (B0, §3.4.3): weights normalize to mean=1 so per-
bar `w_active × r_active` ≈ pre-SP11 absolute scale on average.
Sentinel-defense: experience_env_step runs at start of epoch, SP11
controller runs at end (training_loop.rs ~3475). At fold 0 epoch 0
step 0 the controller has not yet emitted, so ISV[340..346) hold
sentinel 0. Defense: fmaxf(w_raw, 0.01) — same Invariant-1 hard floor
the controller enforces post-renorm. Cold-start scale = 1% of
pre-SP11; Pearl A bootstrap on first emit replaces sentinel.
cf_reward path: out_rewards[cf_off] now writes controller-weighted
cf reward (w_cf × r_cf with sentinel-defense). Loss-kernel
cf_weight=0.3f at mse:318/c51:789 (structural Q-blend, NOT reward
weight) UNTOUCHED per §3.5 amendment.
Mutual exclusivity preserved (popart / trail / micro / opp_cost):
exactly one path fires per bar; others stay 0. The cascade scalar
`reward` mirrors per-component locals so C.4/D.4b bonus blocks that
read in-progress trade reward (Q-cap pattern from
pearl_one_unbounded_signal_per_reward) keep bit-identical compounding.
After cascade, `reward` is overwritten with r_weighted; post-
composition modifiers operate on r_weighted as before.
This is the production-flip commit. Trainer is now on the SP11
controller end-to-end (modulo replay-time curiosity which lands in
B1c after Layer C audit per §3.5.5).
Verification:
- cargo check + build clean (1m32s release).
- 6/6 SP11 GPU oracle tests pass (none exercise env_step directly).
- 14/14 contract tests pass (sp5_isv_slots=10, state_reset_registry=4).
- Local smoke (RTX 3050 Ti, 20-epoch magnitude_distribution) verifies
HEALTH_DIAG sp11_reward weights drift epoch-over-epoch:
epoch 0: w_pop=1.000 w_cf=1.000 w_tr=1.000 ... (uniform sentinel-defense floor → mean=1)
epoch 3: w_pop=1.991 w_cf=2.036 w_tr=0.493 ... (controller redistributes)
epoch 9: w_pop=1.823 w_cf=1.887 w_tr=0.572 ... (mean ≈ 1.0 preserved, Σ ≈ 6)
EVAL_DIST bit-identical to B1a baseline (eq=0.803 eh=0.197 ef=0.000)
— pre-existing magnitude eval-collapse pathology
(project_magnitude_eval_collapse_kelly_capped) unchanged by B1b.
Audit doc updated (Invariant 7): docs/isv-slots.md SP11 section now
reflects Layer B status with B0/B1a/B1b/B1c rollout timeline.
This commit is contained in:
@@ -2529,7 +2529,7 @@ extern "C" __global__ void experience_env_step(
|
||||
}
|
||||
|
||||
/* ================================================================
|
||||
* REWARD v7: Counterfactual Branch Attribution
|
||||
* REWARD v7 + SP11 B1b: Controller-weighted composition
|
||||
*
|
||||
* Fixes v6 penalty stacking (95% breakeven win rate → ~52%).
|
||||
* Novel components:
|
||||
@@ -2545,8 +2545,65 @@ extern "C" __global__ void experience_env_step(
|
||||
* also removed from reward layer — negative-tail compression
|
||||
* relocated to c51_loss_kernel.cu::block_bellman_project_f;
|
||||
* upper +10 cap retained inline.
|
||||
*
|
||||
* SP11 B1b (2026-05-04): structural decomposition into 6 explicit
|
||||
* per-component reward locals — replaces 8+ inline accumulation sites.
|
||||
* Composed as `Σ w_i × r_i` with controller weights from ISV[340..346)
|
||||
* (mean=1 normalization per spec §3.4.3). Universal post-composition
|
||||
* modifiers (drawdown / capital-floor / inventory / churn /
|
||||
* conviction-scale / cf-flip) apply UNWEIGHTED after the weighted Σ
|
||||
* — risk constraints and structural operators, NOT learning components.
|
||||
*
|
||||
* Site → component mapping:
|
||||
* r_popart ← line ~2612 trade P&L (vol-normalized, capped) on
|
||||
* voluntary-exit segment_complete.
|
||||
* r_trail ← line ~2612 trade P&L on trail-fire segment_complete
|
||||
* (§3.5.4 extraction). Mutually exclusive with r_popart.
|
||||
* r_micro ← line ~2869 OFI dense micro-reward, OR line ~2886
|
||||
* positioned-bar holding-cost fallback (when OFI off).
|
||||
* r_opp_cost ← line ~2987 Flat opportunity-cost (Plan 3 B.1).
|
||||
* r_bonus ← Σ of: line ~2670 C.4 timing, line ~2691 D.4a
|
||||
* persistence, line ~2742 D.4b regime penalty (sub-),
|
||||
* line ~3035 B.2 trade-attempt, line ~3085 D.4c
|
||||
* conviction-consistency.
|
||||
* r_cf ← cf-tuple reward written at out_rewards[cf_off]
|
||||
* (composed in the CF block at line ~3300, separately
|
||||
* weighted by w_cf).
|
||||
*
|
||||
* Mutual exclusivity (popart / trail / micro / opp_cost / fallback):
|
||||
* exactly one of these alternative paths fires per bar; the others
|
||||
* remain at 0. The weighted Σ reduces to `w_active × r_active` for the
|
||||
* active path, plus `w_bonus × r_bonus` (additive intra-trade).
|
||||
*
|
||||
* Cold-start sentinel defense: ISV slots [340..346) are sentinel-0
|
||||
* before the controller's first emit (the SP11 controller runs in
|
||||
* the per-epoch metrics block AFTER experience collection — see
|
||||
* `training_loop.rs` ~3475). Reading raw ISV at fold 0 epoch 0 step
|
||||
* 0 yields w = 0 → r_weighted = 0 (reward magnitude collapse). Defense:
|
||||
* `fmaxf(w_raw, SP11_WEIGHT_HARD_FLOOR=0.01)` per spec §3.4.2 — the
|
||||
* same Invariant-1 floor the controller enforces post-renorm. After
|
||||
* the first controller emit, Pearl A bootstrap replaces the sentinel
|
||||
* with the first observation; weights then evolve via the controller.
|
||||
* ================================================================ */
|
||||
|
||||
/* Per-component reward locals (SP11 B1b). Each population site below
|
||||
* sets exactly one of {r_popart, r_trail, r_micro, r_opp_cost} (mutually
|
||||
* exclusive paths) and may accumulate into r_bonus (intra-trade
|
||||
* additive). r_cf is set inside the CF block separately. */
|
||||
float r_popart = 0.0f;
|
||||
float r_trail = 0.0f;
|
||||
float r_micro = 0.0f;
|
||||
float r_opp_cost = 0.0f;
|
||||
float r_bonus = 0.0f;
|
||||
|
||||
/* Trade-cumulative running scalar — tracks the in-progress reward
|
||||
* across the if/else cascade. C.4/D.4b bonus blocks read this to
|
||||
* Q-cap their `|reward|` multiplicand (the existing Q-bounded
|
||||
* pattern from `pearl_one_unbounded_signal_per_reward.md`).
|
||||
* Updated in lockstep with the per-component locals so the bonus
|
||||
* compounding semantics are bit-identical to pre-SP11. After the
|
||||
* cascade closes, this local is overwritten with `r_weighted`
|
||||
* (`Σ w_i × r_i`) and feeds into the post-composition modifiers. */
|
||||
float reward = 0.0f;
|
||||
|
||||
if (segment_complete && segment_hold_time > 0.0f) {
|
||||
@@ -2607,9 +2664,27 @@ extern "C" __global__ void experience_env_step(
|
||||
* Negative-tail compression moved to c51_loss_kernel.cu::
|
||||
* block_bellman_project_f (Huber-style Q-target smoothing).
|
||||
* Upper +10 cap kept inline for numerical safety against adversarial
|
||||
* reward explosions from rare large wins. */
|
||||
* reward explosions from rare large wins.
|
||||
*
|
||||
* SP11 B1b §3.5.4: split forced-exit (trail-fire) P&L from
|
||||
* voluntary-exit P&L. Both paths produce the same vol-normalized
|
||||
* capped scalar; controller weight `w_trail` vs `w_popart` lets
|
||||
* SP11 weigh them differently. `trail_triggered` flips
|
||||
* `exiting_trade=1` upstream (line ~2371), so this branch fires
|
||||
* for both — the segregation matters only for component attribution. */
|
||||
float base_reward = 2.0f * vol_normalized_return;
|
||||
reward = fminf(base_reward, 10.0f);
|
||||
float capped_pnl = fminf(base_reward, 10.0f);
|
||||
if (trail_triggered) {
|
||||
r_trail = capped_pnl;
|
||||
} else {
|
||||
r_popart = capped_pnl;
|
||||
}
|
||||
reward = capped_pnl;
|
||||
/* C.2 attribution: rc[2] now carries the trail signal when the
|
||||
* trail-fire branch fires (was a structural placeholder pre-B1b).
|
||||
* The reward_component_ema kernel will pick up real trail magnitude
|
||||
* for ISV-driven component-mag-ratio canaries. */
|
||||
reward_components_per_sample[out_off * 6 + 2] = r_trail;
|
||||
|
||||
/* ── Layer 3: CEA counterfactual loop REMOVED (4-branch: direction(4) replaces exposure(9)).
|
||||
* Dense micro-reward (Gem 1) below provides directional gradient signal. ── */
|
||||
@@ -2667,7 +2742,8 @@ extern "C" __global__ void experience_env_step(
|
||||
* (bars_early / segment_hold_time)
|
||||
* pnl_capped
|
||||
* conviction_core;
|
||||
reward += timing_bonus;
|
||||
r_bonus += timing_bonus;
|
||||
reward += timing_bonus; /* keep cascade scalar in lockstep for Q-cap reads below */
|
||||
/* C.2 attribution — accumulate into rc[5] bonus slot (see above). */
|
||||
reward_components_per_sample[out_off * 6 + 5] += timing_bonus;
|
||||
}
|
||||
@@ -2688,7 +2764,8 @@ extern "C" __global__ void experience_env_step(
|
||||
if (drawdown_depth > 1e-6f) { /* skip straight-line profit trades */
|
||||
const float persist_bonus = shaping_scale * conviction_core * drawdown_depth
|
||||
* tanhf(reward / fmaxf(1e-4f, drawdown_depth));
|
||||
reward += persist_bonus;
|
||||
r_bonus += persist_bonus;
|
||||
reward += persist_bonus; /* keep cascade scalar in lockstep for Q-cap reads below */
|
||||
reward_components_per_sample[out_off * 6 + 5] += persist_bonus;
|
||||
}
|
||||
}
|
||||
@@ -2739,7 +2816,8 @@ extern "C" __global__ void experience_env_step(
|
||||
const float reward_capped = fminf(reward_unit, 1.0f) * reward_norm;
|
||||
const float penalty = shaping_scale * conviction_core
|
||||
* bars_late_frac * reward_capped;
|
||||
reward -= penalty;
|
||||
r_bonus -= penalty;
|
||||
reward -= penalty; /* keep cascade scalar in lockstep for Q-cap reads below */
|
||||
reward_components_per_sample[out_off * 6 + 5] -= penalty;
|
||||
}
|
||||
/* Reset shift-bar for the NEXT trade (in addition to Step 3's lifecycle resets,
|
||||
@@ -2866,24 +2944,33 @@ extern "C" __global__ void experience_env_step(
|
||||
+ w_book * sign_pos * book_aggression
|
||||
+ w_hold * hold_quality
|
||||
- spread_penalty;
|
||||
reward = shaping_scale * micro_reward_scale * tanhf(quality / micro_reward_temp);
|
||||
r_micro = shaping_scale * micro_reward_scale * tanhf(quality / micro_reward_temp);
|
||||
reward = r_micro; /* mutually exclusive with segment_complete; reward starts at 0 here */
|
||||
/* Task 0.8: captures the micro-reward magnitude BEFORE downstream
|
||||
* drawdown / conviction / sign-flip transforms. The assignment to
|
||||
* `reward` above is an OVERWRITE (not additive) — the segment_complete
|
||||
* `r_micro` above is an OVERWRITE (not additive) — the segment_complete
|
||||
* branch that precedes this `else if` already returned before reaching
|
||||
* here if it fired, so `reward` at this line IS the micro term in
|
||||
* here if it fired, so `r_micro` at this line IS the micro term in
|
||||
* isolation. Downstream transforms (drawdown penalty, conviction
|
||||
* scaling, counterfactual sign flip) mutate `reward` further before
|
||||
* `total_reward_per_sample` is written, so the ratio
|
||||
* scaling, counterfactual sign flip) mutate the post-composition
|
||||
* `reward` further before `total_reward_per_sample` is written, so
|
||||
* the ratio
|
||||
* Σ|micro_pre_transform| / Σ|total_post_transform|
|
||||
* is a magnitude ratio (diagnostic for relative term scale), not a
|
||||
* true additive decomposition. See accessor docstring. */
|
||||
micro_reward_per_sample[out_off] = reward;
|
||||
micro_reward_per_sample[out_off] = r_micro;
|
||||
/* C.2: capture micro component before downstream transforms. */
|
||||
reward_components_per_sample[out_off * 6 + 3] = reward;
|
||||
reward_components_per_sample[out_off * 6 + 3] = r_micro;
|
||||
} else if (!segment_complete && fabsf(position) > 0.001f) {
|
||||
/* Fallback: original holding cost when micro_reward_scale=0 or no OFI */
|
||||
reward = -shaping_scale * holding_cost_rate * fabsf(position);
|
||||
/* Fallback: original holding cost when micro_reward_scale=0 or no OFI.
|
||||
* SP11 B1b: maps to r_micro (positioned-non-completion slot in the
|
||||
* mutual-exclusivity chain — same semantic role as the OFI branch
|
||||
* above, just degraded). Captured in micro_reward_per_sample +
|
||||
* rc[3] for diagnostic continuity with the OFI path. */
|
||||
r_micro = -shaping_scale * holding_cost_rate * fabsf(position);
|
||||
reward = r_micro;
|
||||
micro_reward_per_sample[out_off] = r_micro;
|
||||
reward_components_per_sample[out_off * 6 + 3] = r_micro;
|
||||
} else if (!segment_complete) {
|
||||
/* Plan 3 D.4c: Welford-style EMA of conviction_core during Flat bars.
|
||||
* Updates running mean and squared-deviation EMAs that the consumer
|
||||
@@ -2984,10 +3071,11 @@ extern "C" __global__ void experience_env_step(
|
||||
: 1.0f;
|
||||
/* conviction_core ∈ [0, 1] — ISV-driven per-sample adaptive
|
||||
* multiplier. Replaces the prior 0.5 hardcoded constant. */
|
||||
reward = -shaping_scale * holding_cost_rate
|
||||
* conviction_core * vol_proxy_flat * q_abs_ref * kl_amp_b1;
|
||||
r_opp_cost = -shaping_scale * holding_cost_rate
|
||||
* conviction_core * vol_proxy_flat * q_abs_ref * kl_amp_b1;
|
||||
reward = r_opp_cost; /* mutually exclusive with segment_complete + micro paths */
|
||||
/* C.2 attribution: record opp_cost component. */
|
||||
reward_components_per_sample[out_off * 6 + 4] = reward;
|
||||
reward_components_per_sample[out_off * 6 + 4] = r_opp_cost;
|
||||
}
|
||||
}
|
||||
|
||||
@@ -3032,7 +3120,8 @@ extern "C" __global__ void experience_env_step(
|
||||
float kl_amp_b2 = fmaxf(1.0f, isv_signals_ptr[ISV_STATE_KL_AMP_IDX]);
|
||||
float bonus = conviction_core * vol_proxy_b2 * novelty;
|
||||
float bonus_scaled = shaping_scale * bonus * kl_amp_b2;
|
||||
reward += bonus_scaled;
|
||||
r_bonus += bonus_scaled;
|
||||
reward += bonus_scaled; /* keep cascade scalar in lockstep */
|
||||
/* C.2 attribution: record bonus component magnitude at rc[5]. */
|
||||
reward_components_per_sample[out_off * 6 + 5] = bonus_scaled;
|
||||
}
|
||||
@@ -3082,7 +3171,8 @@ extern "C" __global__ void experience_env_step(
|
||||
float vol_proxy_c = fmaxf(atr_pct_c, 0.0001f);
|
||||
if (vol_proxy_c > 0.01f) vol_proxy_c = 0.01f;
|
||||
const float bonus = shaping_scale * vol_proxy_c * stability * conviction_core;
|
||||
reward += bonus;
|
||||
r_bonus += bonus;
|
||||
reward += bonus; /* keep cascade scalar in lockstep */
|
||||
reward_components_per_sample[out_off * 6 + 5] += bonus;
|
||||
}
|
||||
}
|
||||
@@ -3094,6 +3184,78 @@ extern "C" __global__ void experience_env_step(
|
||||
ps[PS_PRE_ENTRY_CONVICTION_VAR_EMA] = 0.0f;
|
||||
}
|
||||
|
||||
/* ================================================================
|
||||
* SP11 B1b: Controller-weighted reward composition
|
||||
*
|
||||
* Spec §3.5.3 amended at main commit 7ddaf9c51. Replaces the
|
||||
* pre-SP11 in-place `reward = X / reward += X` cascade with explicit
|
||||
* `Σ w_i × r_i` over the 5 on-policy components populated above.
|
||||
* The cf component (`r_cf`) is composed separately inside the CF
|
||||
* block at the bottom of this kernel (the cf-tuple's reward written
|
||||
* to `out_rewards[cf_off]` is `w_cf × r_cf` per spec §3.5.0/§3.5).
|
||||
*
|
||||
* Mean=1 normalization (§3.4.3): controller writes weights to ISV
|
||||
* with mean(weights) ≈ 1 (Σ ≈ 6 across 6 components). The per-bar
|
||||
* weighted Σ reduces to `w_active × r_active + w_bonus × r_bonus`
|
||||
* which preserves pre-SP11 absolute scale on average (uniform
|
||||
* weights = 1.0 each, identical to pre-SP11's single-path-fires
|
||||
* semantic).
|
||||
*
|
||||
* Sentinel-defense (§3.4.2 + cold-start ordering): the SP11
|
||||
* controller runs in the per-epoch metrics block AFTER experience
|
||||
* collection (training_loop.rs ~3475), so at fold 0 epoch 0 step 0
|
||||
* the controller has NOT yet emitted — ISV slots [340..346) are
|
||||
* sentinel 0. Reading raw 0 weights would collapse `r_weighted` to
|
||||
* 0 (full reward magnitude loss until after epoch 0). Defense:
|
||||
* `fmaxf(w_raw, SP11_WEIGHT_HARD_FLOOR=0.01)` — same Invariant-1
|
||||
* floor the controller enforces post-renorm. Cold-start reward
|
||||
* scale = 0.01 × |r_active| (1% of pre-SP11) on the first emit;
|
||||
* Pearl A bootstrap on the controller's first observation replaces
|
||||
* sentinel with raw → mean-1 weights take over from epoch 1 on.
|
||||
*
|
||||
* NOTE: The 0.01 hard-floor on cold-start IS Invariant-1
|
||||
* (numerical safety). It's the absolute minimum on every weight at
|
||||
* every step — including the first-emit sentinel state. The
|
||||
* controller's `WEIGHT_HARD_FLOOR=0.01` is the same constant
|
||||
* referenced via `SP11_WEIGHT_HARD_FLOOR` in the Rust slot
|
||||
* constants (sp11_isv_slots.rs:56). Hardcoded inline in CUDA
|
||||
* because cudarc kernels can't include Rust headers; magic-number
|
||||
* value sourced from the spec — `feedback_no_quickfixes` carve-out
|
||||
* for Invariant-1 anchors per spec §3.4.2.
|
||||
* ================================================================ */
|
||||
{
|
||||
/* On-policy composition: 5 components (popart / trail / micro /
|
||||
* opp_cost / bonus). The cf-reward path is composed separately in
|
||||
* the CF block below using the same Invariant-1 floor pattern.
|
||||
* r_cf is structurally absent from this Σ because the cf-tuple's
|
||||
* reward lives at out_rewards[cf_off], not out_rewards[out_off]. */
|
||||
const float w_pop_raw = (isv_signals_ptr != NULL) ? isv_signals_ptr[340] : 0.0f;
|
||||
const float w_tr_raw = (isv_signals_ptr != NULL) ? isv_signals_ptr[342] : 0.0f;
|
||||
const float w_mi_raw = (isv_signals_ptr != NULL) ? isv_signals_ptr[343] : 0.0f;
|
||||
const float w_oc_raw = (isv_signals_ptr != NULL) ? isv_signals_ptr[344] : 0.0f;
|
||||
const float w_bn_raw = (isv_signals_ptr != NULL) ? isv_signals_ptr[345] : 0.0f;
|
||||
|
||||
/* Sentinel-defense: Invariant-1 hard floor (= SP11_WEIGHT_HARD_FLOOR).
|
||||
* Before the controller's first emit, raw ISV slots are 0 → defaults
|
||||
* to the 0.01 floor. Post-emit, the controller already enforces this
|
||||
* floor in its renorm; the fmaxf is therefore idempotent on the
|
||||
* post-emit path. NULL isv_signals_ptr (no controller wiring) also
|
||||
* routes through the 0.01 floor — equivalent to "all 5 on-policy
|
||||
* components weighted equally at the minimum scale". */
|
||||
const float w_pop = fmaxf(w_pop_raw, 0.01f);
|
||||
const float w_tr = fmaxf(w_tr_raw, 0.01f);
|
||||
const float w_mi = fmaxf(w_mi_raw, 0.01f);
|
||||
const float w_oc = fmaxf(w_oc_raw, 0.01f);
|
||||
const float w_bn = fmaxf(w_bn_raw, 0.01f);
|
||||
|
||||
const float r_weighted = w_pop * r_popart
|
||||
+ w_tr * r_trail
|
||||
+ w_mi * r_micro
|
||||
+ w_oc * r_opp_cost
|
||||
+ w_bn * r_bonus;
|
||||
reward = r_weighted; /* the post-composition modifiers below operate on r_weighted */
|
||||
}
|
||||
|
||||
/* ---- Drawdown penalty (from trade_physics.cuh) ──────────────────────
|
||||
* Phase 3: gated by shaping_scale. The drawdown penalty is *behavioral*
|
||||
* (steers the policy away from large equity drops); the *physics* of
|
||||
@@ -3414,14 +3576,33 @@ extern "C" __global__ void experience_env_step(
|
||||
}
|
||||
|
||||
if (isnan(cf_reward) || isinf(cf_reward)) cf_reward = 0.0f;
|
||||
|
||||
/* SP11 B1b §3.5.0/§3.5: apply w_cf to the cf-tuple reward written
|
||||
* to the replay buffer. Loss kernels (mse, c51, IQN target, CQL,
|
||||
* Bellman) read this buffer slot and consume the controller-weighted
|
||||
* value directly. Same Invariant-1 hard floor as the on-policy
|
||||
* weights — pre-emit sentinel 0 → 0.01 baseline scale.
|
||||
*
|
||||
* Note (spec §3.5 amendment): the `cf_weight = 0.3f` constants at
|
||||
* `mse_loss_kernel.cu:318` and `c51_loss_kernel.cu:789` are
|
||||
* STRUCTURAL Q-distribution mixing weights for Hold-sample magnitude
|
||||
* branch gradient — a different signal, NOT migrated here. They
|
||||
* stay at 0.3f untouched. The SP11 controller weight applies ONLY
|
||||
* to the cf-reward path in this kernel. */
|
||||
const float w_cf_raw = (isv_signals_ptr != NULL) ? isv_signals_ptr[341] : 0.0f;
|
||||
const float w_cf = fmaxf(w_cf_raw, 0.01f);
|
||||
const float r_cf = cf_reward;
|
||||
const float cf_reward_weighted = w_cf * r_cf;
|
||||
|
||||
out_actions[cf_off] = cf_action;
|
||||
out_rewards[cf_off] = cf_reward;
|
||||
out_rewards[cf_off] = cf_reward_weighted;
|
||||
out_dones[cf_off] = (float)done;
|
||||
/* C.2: component [1] = counterfactual reward magnitude for this sample.
|
||||
* Written at out_off (on-policy slot), not cf_off, so the EMA kernel's
|
||||
* [N*L, 6] reduction over out_off slots captures it alongside the other
|
||||
* on-policy components. */
|
||||
reward_components_per_sample[out_off * 6 + 1] = cf_reward;
|
||||
* on-policy components. Stores the post-weight value so the EMA
|
||||
* tracks the same magnitude the loss kernels actually train on. */
|
||||
reward_components_per_sample[out_off * 6 + 1] = cf_reward_weighted;
|
||||
}
|
||||
|
||||
/* ---- Advance timestep counter ---- */
|
||||
|
||||
@@ -184,7 +184,7 @@ gradient contribution = `iqn_budget` (same as SP5 Layer B). The averaged
|
||||
**Files changed:** `crates/ml/src/cuda_pipeline/gpu_iqn_head.rs`,
|
||||
`crates/ml/src/trainers/dqn/fused_training.rs`.
|
||||
|
||||
## SP11: Reward as controlled subsystem (Tasks A0 + A1 + A2)
|
||||
## SP11: Reward as controlled subsystem (Tasks A0 + A1 + A2 + B0 + B1a + B1b)
|
||||
|
||||
SP11 Task A0 (Fix 39, 2026-05-04) extends `ISV_TOTAL_DIM` from 340 → 360 by
|
||||
allocating 20 contiguous slots at `[340..360)` for the reward-subsystem
|
||||
@@ -217,10 +217,33 @@ gets a fold-reset registry entry (`sp11_novelty_hash`) that closes the
|
||||
A0 deferral; the projection matrix is frozen for the run lifetime and
|
||||
intentionally has NO registry entry.
|
||||
|
||||
**Layer A is still additive — no consumer reads any [340..360) slot in
|
||||
this commit. Training behaviour remains identical to A1/A0; only the
|
||||
HEALTH_DIAG `sp11_reward` line is observably new. Atomic consumer
|
||||
migration (Layer B) wires downstream.**
|
||||
**Layer A was additive (A0/A1/A2). Layer B atomic consumer migration:**
|
||||
|
||||
- **B0 (commit `302992f63`):** controller renormalisation flipped from
|
||||
`Σ=1` to `mean=1` (Σ=N=6) per spec §3.4.3 amendment so per-bar
|
||||
`w_active × r_active` ≈ pre-SP11 absolute scale on average. Per-
|
||||
component cap MAX_WEIGHT=3.0 (Invariant-1, prevents winner-take-all).
|
||||
- **B1a (commit `d5e1214f2`):** saboteur GPU multiplication via
|
||||
`SABOTEUR_INTENSITY_MULT_INDEX` and SimHash novelty-buffer
|
||||
`state_stride` correctness for B1c curiosity wiring.
|
||||
- **B1b (this commit):** structural reward decomposition in
|
||||
`experience_env_step` — 8+ inline accumulation sites replaced with
|
||||
explicit per-component locals (`r_popart`, `r_cf`, `r_trail`,
|
||||
`r_micro`, `r_opp_cost`, `r_bonus`) composed as `Σ w_i × r_i` with
|
||||
controller weights from ISV[340..346). Universal post-composition
|
||||
modifiers (drawdown / capital-floor / inventory / churn /
|
||||
conviction-scale / cf-flip) apply unweighted per spec §3.4.4. Trail
|
||||
P&L extraction (§3.5.4) — forced-exit P&L now flows through
|
||||
`r_trail` (rc[2]); voluntary-exit through `r_popart`. cf-tuple
|
||||
reward at `out_rewards[cf_off]` is now `w_cf × r_cf` (the
|
||||
`cf_weight=0.3f` constants in `mse_loss_kernel.cu:318` /
|
||||
`c51_loss_kernel.cu:789` are STRUCTURAL Q-blend weights, NOT reward
|
||||
weights — left UNTOUCHED per §3.5 amendment). Sentinel-defense:
|
||||
`fmaxf(w_raw, 0.01)` covers the cold-start ordering gap (the
|
||||
controller runs at end-of-epoch, so step 0 of fold 0 reads sentinel
|
||||
0; defense pins minimum scale at the Invariant-1 hard floor).
|
||||
- **B1c (pending):** replay-time curiosity bonus per §3.5.5 (Layer C
|
||||
audit gate — `rewards_buf` single-read-site verification).
|
||||
|
||||
A1 producers:
|
||||
- `val_sharpe_delta_compute_kernel.cu` — two-pass: raw Δ + squared
|
||||
@@ -273,7 +296,13 @@ launcher's seed derivation.
|
||||
|
||||
All 20 slots are FoldReset (sentinel 0; Pearl A bootstraps from first
|
||||
producer launch on each fold per `pearl_first_observation_bootstrap.md`).
|
||||
After A2 all 20 slots have producers; consumers wire in Layer B.
|
||||
After A2 all 20 slots have producers; B1b wires the on-policy reward
|
||||
composition consumer for [340..346) (the 6 component weights).
|
||||
B1c will wire the curiosity consumer for [346] + [349] (replay-time).
|
||||
[347] saboteur multiplier already consumed in B1a's
|
||||
gpu_experience_collector site. [348] `REWARD_WEIGHT_FLOOR_INDEX` is
|
||||
self-consumed by the controller on the next step. [350..360) canaries
|
||||
are read by the controller (intra-cycle).
|
||||
|
||||
The novelty-hash device buffer (`GpuDqnTrainer.novelty_hash_buf`,
|
||||
1M slots × f32 mapped-pinned) lands in A2 alongside its registry entry
|
||||
|
||||
Reference in New Issue
Block a user