286ea26e2a0c02f557237f40225cd16dfdd433b4
202 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
f5632649ca |
spec(ml-alpha): per-horizon attention pool design (C20)
Captures the brainstormed "alternative attention pool variants"
follow-on from the original real-LOB integration brainstorm (Axis 1,
deferred from the LOB workstream as a separate model-side spec).
Design:
Replace shared learned query Q[HIDDEN_DIM] with per-horizon queries
Q_h[N_HORIZONS, HIDDEN_DIM]. Per-horizon softmax + context vectors
feed multi-horizon heads directly (PATH A) — each horizon attends
to a different part of the K=6000 LN_b output sequence. CfC k=0
state is initialised by the MEAN of per-horizon contexts so the
K-loop recurrence + state amplification (per
pearl_state_amplifies_short_horizon_into_long_horizon) survives.
Heads consume per-horizon context concat CfC h_K (residual) with a
default 75/25 weight split.
Falsifiable claim (§0): A/B-tested win means h6000 mean_auc lifts by
≥ +0.01 absolute OR per-horizon distribution shifts toward short
horizons (h1000, h300) with no net h6000 loss. The 3-fold variance
band on the current architecture (mean_auc 0.7749 ± 0.024) means a
+0.01 lift is within noise — a meaningful effect needs ≥ +0.024 or
qualitative distributional shift.
Two new kernels (per_horizon_attention_pool_fwd + _bwd) + signature
extension on multi_horizon_heads_{fwd,bwd}. Variant-toggle config flag
(SharedQuery vs PerHorizonQuery) keeps the existing path fully
functional; new variant is opt-in. CheckpointV1 → V2 with explicit
discriminant + optional q_h field; V1 files load as SharedQuery, new
V2 training writes the discriminant.
Three validation rings:
1. Per-(b,h) numgrad parity at K=16
2. One-epoch smoke (no NaN, loss decreases)
3. 30-epoch × 3-fold A/B (#204) — decision gate per §0 falsifiable claim
Implementation explicitly deferred. The decision to invest depends on
(a) GPU time budget (~3-6 hrs on L40S × 5 GPUs for the A/B), (b)
whether per-horizon cost-frontier sweeps (#202 follow-ups) surface
viable horizons beyond h6000 that would benefit from per-horizon
specialisation, and (c) the 3-fold variance noise floor making the
expected effect size visible.
Next step when ready: invoke superpowers:writing-plans against this
spec for the ~6-8 commit implementation plan.
Closes the "good to have" question from the recent brainstorm with a
concrete decision framework rather than ad-hoc implementation.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
|
||
|
|
bfbaea2661 |
spec(ml-backtesting): real-LOB integration design
Greenfield LOB simulator + execution policy + backtest harness inside existing ml-backtesting crate. Turns the AUC predictor into a P&L producer at IBKR+Scaleway latency profile (100ms baseline). Key design decisions: - Reuse trainer's MultiHorizonLoader, Mbp10RawInput, snap_feature_assemble cubin, CfcTrunk captured graph verbatim (zero train-vs-deploy skew). - Per-horizon ISV-Kelly + adaptive WeightedByRealizedSharpe aggregator (no static horizon mask; policy self-shifts capital). - Block-per-backtest CUDA (32 threads/warp/block), ~1.4 KB shared mem. - Three captured graphs (perception, step-event, decision); host branches only between captures. - Mapped-pinned for all CPU/GPU buffers (hard requirement per feedback_no_htod_htoh_only_mapped_pinned). - Three validation rings: N=1 bit-exact fixtures, trainer-parity check, N>1 invariant fuzz; production-day replay as Ring 3. Single binary fxt-backtest with run + aggregate subcommands. No new crates; extends existing ml-backtesting alongside barrier_backtest. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> |
||
|
|
ea1d8703b7 |
spec(perf): revise K-loop parallelization design — full critical pass
Revises all 12 issues from the self-review pass:
1. (HIGH) Soften bit-equivalence claim — single-sample helper and
batched kernel at B=1 are different CUDA kernels; FP order may
differ. Acceptance is relative_eq ≤ 1e-6, not bit-exact.
2. (HIGH) Explicit asymmetry: only cfc_step has a single-sample GPU
oracle. GRN/VSN/attn bwd rely on smoke + chain-rule preservation.
feedback_no_cpu_test_fallbacks.md forbids a CPU reference oracle.
3. (HIGH) Realistic targets — 3× floor, 5× stretch. Drops 15× claim
which was Amdahl-bounded under any reasonable assumption.
4. (MED) AdamW-after-reducer invariant stated explicitly: final grad
buffer is OVERWRITE by reducer, meaningful only after reducer ran.
5. (MED) New B=32 smoke test (stacked_trainer_loss_shrinks_at_batch_32)
actually exercises the cross-batch reduction path; existing
perception_overfit suite is all B=1.
6. (MED) Rollout commit 1 bundles reduce_axis0 + first consumer
(cfc_step refactor) to avoid feedback_wire_everything_up.md
orphan-kernel anti-pattern.
7. (LOW) Drop "merge to ml-alpha-phase-a" — user already chose direct
commits to that branch; clarify in Rollout.
8. (LOW) Add explicit scratch-sizing formula:
scratch_bytes ≈ B × Σ(param tensor sizes per kernel).
9. (LOW) Remove resolved open question (memset_zeros ordering).
10. (STRUCT) Post-refactor bottleneck analysis section — names the
next L40S floor (Mamba2 scan, launch latency, cuBLAS).
11. (STRUCT) cudaFuncSetAttribute note — refactored cfc_bwd's per-
block shared mem drops to 1 KiB; no attribute change needed.
12. (STRUCT) Explicit Rollback section — atomic-commit-per-kernel +
revert strategy; rollback baseline is the spec commit
(
|
||
|
|
54aa69c108 |
spec(perf): K-loop block-per-batch parallelization design
Documents the design for fixing the single-SM bottleneck in five backward/K-loop kernels: cfc_step (fwd+bwd), GRN bwd, VSN bwd, and attention_pool bwd. All currently use grid=(1,1,1) with an internal n_batch loop — on L40S (142 SMs) with B=32 this puts <1% of the GPU to work in the K-loop critical path. Architecture: block-per-batch (grid=(B,1,1)) for the kernel body, plus per-batch grad scratch buffers reduced via a single parameterised reduce_axis0 kernel (block tree-reduce, no atomicAdd per feedback_no_atomicadd.md). Same pattern as the existing LayerNorm bwd reducer — CUDA-Graph-safe, debuggable, and consistent with foxhunt's no-cooperative-groups discipline. Target: ≥3× epoch wall speedup (stretch 8-15×). Makes 3-fold CV tractable (10.5h → 2-3h) and unblocks decision-stride / state-dim sweeps that compound the gain. Acceptance gates: (a) all 8 perception_overfit smokes still converge, (b) new B=1 bit-equivalence test asserts the refactored batched bwd kernel at B=1 matches the existing single-sample helper byte-for-byte, (c) cluster A/B vs t6z89 baseline shows AUC trajectory within ±0.005 and epoch wall ≥3× faster. Atomic refactor per kernel — one commit per kernel covering the kernel rewrite, scratch buffer alloc, reducer launch wiring, and smoke. No "_legacy" parallel kernels per feedback_no_legacy_aliases.md + feedback_no_partial_refactor.md. Open implementation-plan decisions: exact memset_zeros ordering inside the captured graph, batch-vs-per-tensor reducer launches, optimal block_dim for reduce_axis0 itself. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> |
||
|
|
ef35b10f3e |
refactor(ml-alpha): drop competitive-gate machinery — stacked is default
Per user direction "no gating, this is the new default": the stacked
Mamba2 -> CfC -> heads design is THE production architecture. There's
no competing-baseline comparison to run. Validation reduces to normal
training metrics (per-horizon val AUC, train loss curve, sanity floor
of >0.5 AUC).
Deletions:
- crates/ml-alpha/src/gate/cfc_vs_mamba2.rs (gate verdict logic)
- crates/ml-alpha/src/gate/mod.rs
- crates/ml-alpha/examples/alpha_gate.rs (gate runner binary)
Renames:
- crates/ml-alpha/src/gate/auc.rs -> crates/ml-alpha/src/eval/auc.rs
- lib.rs: pub mod gate -> pub mod eval (gate implied comparison;
eval doesn't)
Spec amendments:
- Drop the "Gate baseline strategy" amendment (committed earlier
this session)
- Reframe the stacked-architecture amendment as a "decision" not a
"gate"; production path is unambiguous
- Reframe Section 4 "Validation gate: CfC must meet Mamba2" -> just
"Validation: per-horizon val AUC" with the >0.5 sanity floor
Doc cleanups: stale "Mamba2 gate baseline" mentions in build.rs and
pinned_mem.rs replaced with neutral wording. The Argo template
comment about "downstream gate consumption" becomes "for monitoring".
Test status: all 26+ ml-alpha tests pass. AUC tests (6/6) still pass
under the eval:: namespace.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
|
||
|
|
6e868fd136 |
spec(ml-alpha): gate-baseline ablation strategy amendment
Defines the Mamba2-only baseline for the stacked-vs-baseline gate verdict as an ablation of the SAME PerceptionTrainer (a --bypass-cfc flag), not a separate model. Apples-to-apples; same data window, same hyperparameters, same code path. The only difference is whether the CfC step is in the loop. Three ablation options evaluated: 1. --bypass-cfc flag (recommended): Mamba2 -> heads directly 2. --mamba2-state-dim 2 (crippled Mamba2, CfC stays) 3. Frozen CfC initialized to identity (no code branch needed) Option 1 wins on clarity: it answers "is CfC additive on top of Mamba2" unambiguously, with the same Mamba2 capacity and same training regime in both arms. Concrete next-session work documented (1-2 hours): - PerceptionTrainerConfig.bypass_cfc: bool + step() branch - alpha_train --bypass-cfc CLI flag - alpha-perception-template.yaml workflow parameter + bash branch - submit both runs, fetch summaries, alpha_gate, commit verdict gate_verdict logic unchanged — the cfc/mamba2 naming in the report becomes stacked/bypass at the binding layer; the verdict math is generic. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> |
||
|
|
3389a59281 |
spec(ml-alpha): stacked Mamba2 -> CfC amendment
Mid-execution architecture revision: Mamba2 stays as a sequence encoder; CfC becomes the layer on top (replacing the Phase 1d.3 MLP stacker). Gate becomes 'stacked AUC >= Mamba2-only stacker AUC at every horizon' — proves the CfC layer is additive, rather than CfC alone beating Mamba2 alone. Plan 1 kernels (cfc_step, heads, projection, BCE, AdamW, Graph A) are unchanged. Only CfcTrunk's forward path gains a Mamba2 prefix that consumes the snapshot stream and emits a 128-dim h_mamba which CfC reads. The Mamba2 kernel (mamba2_alpha_kernel.cubin) is already in the build. Option B (parallel + fused dual-stream) is documented as the Plan 2 fallback if the stacked gate fails. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> |
||
|
|
90d14aaba5 |
spec(ml-alpha): GPU/CPU contract, CUDA Graphs, cold-start, disconnect handling
Critical-review pass on the CfC+PPO design adds: - Public API: mapped-pinned ingress slots, single DtoH per decision - Section 5.5: CUDA Graph capture topology (perception / policy / training), kernel inventory, cuBLAS Lt epilogue fusion, persistent ring layouts, determinism mode, build-time cubin compilation, perf budget - Section 5.6: Databento/IBKR disconnect + failure response, never-do list - Section 6 cold-start: 4-state bootstrap, Phase B walk-forward stats as pre-deployment kill-switch baseline, always-live ISV controllers Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> |
||
|
|
d4ea54d780 |
spec(ml-alpha): CfC + PPO greenfield design
Approved design document for the ml-alpha rebuild. Seven sections plus
appendices, all 7 sections + ISV-driven auto-tuning addition approved
via the brainstorming-skill flow.
Key architectural decisions:
- Full closed-loop scope: training + offline inference + live IBKR
- Three components: ml-alpha library, alpha_train binary, AlphaPpoStrategy
in trading_agent_service (reuses existing data_acquisition,
trading_service, broker_gateway, and the 1177-LOC IBKR adapter that
already exists in trading_engine)
- Snapshot-level CfC perception trunk (Closed-form Continuous-time, the
trainable LTC variant from Hasani 2022) — chosen over Mamba2 for
native multi-time-scale via per-cell learnable τ; hard validation
gate requires CfC AUC ≥ Mamba2 baseline before continuing
- 5 multi-horizon heads at {30, 100, 300, 1000, 6000} snapshots forward
- Decision-stride 50 snapshots (~1 min); empirically validated this
session (per-bar→stride=200 flipped 3-fold mean Sharpe from -4.29 to
+1.78 at quarter-tick cost)
- Action: target_position categorical ∈ {-10, ..., +10}; max_train=10
fixed at architecture level, max_live ≤ 10 runtime-cappable
- Reward in price units (position-size invariant): dense per-segment
PnL − C_trade·|Δcontracts| − C_vol·vol_excess; C_trade=0.034 fixed
(IBKR $1.70 RT ÷ $50/point), C_vol=0 default
- Full online learning with EWC anchoring + 80/20 offline/live replay
buffer + kill-switch on σ-band metric deviations (loss, mean reward,
hit-rate, KL, entropy)
- All hyperparameters that vary during training are ISV-driven
(controller outputs, not magic numbers) per
pearl_controller_anchors_isv_driven.md
- Fully GPU-resident on hot path per feedback_cpu_is_read_only.md
- $35k starting capital; conservative max_live ramp 1→2→3 contracts
with manual config bumps gated on realized track record
- Documented IBKR-specific quirks (PDT, mid-session margin spikes,
rollover, halts, rate limits) with concrete design responses
Supersedes the abandoned 2026-05-16-alpha-ppo-trainer.md (DQN-port plan)
which inherited the wrong architectural shape.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
|
||
|
|
2fe76f2f34 |
docs(phase-e-4-a): execution status + T10 backward_from_h_enriched patch sketch
Two documentation deliverables produced while T14 backtest runs: 1. Plan update (specs/2026-05-15-phase-e-4-a-temporal-foundation.md): adds 'Execution Status' section reflecting actual T1-T14 progression. T5 deferred (real MBP-10 peek), T9 skipped (GRN moved to E.4.B per integration notes), T10 partial (new C51 grad-input kernel landed but Mamba2 backward wiring deferred), T14 in flight. Documents the 4 execution learnings: research-first saved a week of duplicate kernel work; cheap falsification experiments (Path 2, Path 3) avoided expensive investments; C51 borrow was the largest single Sharpe-lift in the session; GpuTensor/CudaSlice interop friction is the real integration cost. 2. T10 patch sketch (specs/2026-05-15-t10-mamba2-backward-from-h-enriched.md): ready-to-apply patch for ml-alpha::Mamba2Block adding a new public method backward_from_h_enriched(cache, d_h_enriched). Bypasses the W_out projection backward, accepts the [B, hidden_dim] gradient from C51's grad-input kernel directly, zero-initialises dw_out/db_out (AdamW step on zero grad is a no-op with correct moment decay — effectively freezes W_out params which is correct semantics since Phase E never uses them). Includes the smoke binary wiring snippet that consumes the new method via launch_alpha_c51_grad_input → Mamba2 backward → AdamW step. Application gated on T14 backtest validation — if frozen Mamba2 already lifts Sharpe, T10 becomes optimisation rather than prerequisite. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> |
||
|
|
eb9047fc30 |
docs(phase-e): E.4 temporal encoder design + E.4.A implementation plan
Design doc (specs/): TFT-style architecture for Phase E execution
policy — sliding window → Mamba2 SSM → GRN trunk → MoE regime gate
→ C51 head → Thompson selector. Two core pillars added per user:
A) Full L1-L10 LOB depth input via hybrid MBP-10 peek
B) ISV-continual-learning: controllers fire at training AND
inference; Q-net weights frozen at inference but effective
policy adapts via ISV modulation
Plan doc (plans/): 14-task implementation plan for E.4.A foundation
(window buffer + L1-L10 depth + Mamba2 forward+backward + ISV-eval
controllers). Falsification gates: smoke R_mean improvement ≥ 50%,
backtest cost=0 Sharpe ≥ +8 (no regression vs C51-flat +10.41),
half-tick Sharpe ≥ -8 (closes 5pt+ of 10pt gap to Phase 1d.4
baseline -4.0).
TGGN (foxhunt Temporal Graph Gated Network) explicitly deferred to
Phase E.5+: existing CPU graph implementation + GPU adapter at
ml-supervised/src/tgnn/ — marginal benefit for single-instrument ES
futures vs the TFT-Mamba2 stack; revisit for multi-instrument
extension or production HFT inference layer.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
|
||
|
|
588a6d38af |
docs(phase-e-4-a): mamba2 + grn integration research — use ml-alpha Mamba2Block
Findings: - Production Mamba2 (gpu_dqn_trainer) is coupled to SH2=256 trunk + ofi_embed + ISV[8] temporal routing — not portable to Phase E. - ml-alpha::mamba2_block::Mamba2Block is from-scratch, fully configurable (in_dim/hidden_dim/state_dim/seq_len), GPU-pure with forward_train/backward/AdamW. Used in Phase 1d.2 to lift AUC 0.50 to 0.66. ml-alpha is already a workspace dep of ml. - GRN skipped for E.4.A — Mamba2 output goes straight to C51 head. Reintroduce GRN in E.4.B if Sharpe gates don't pass. - Controller-at-inference: kernel has no training-mode branches; Wiener state preserved across episodes/cost cells for natural live-deployment simulation. Revises Tasks 8-10 of the plan: use Mamba2Block API instead of writing custom kernels. Only new CUDA needed: alpha_c51_grad_input (C51 gradient w.r.t. input features, for Mamba2 backward chain). Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> |
||
|
|
9e6637c065 |
docs(claude): foxhunt-aware agents & skills design — phase 1 (12 items)
Pattern-cluster auditors + workflow-skill hybrid that internalize accumulated project wisdom (15 feedback_*, 30+ pearl_*, 12+ project_* memory files; 30 SP specs) into foxhunt-namespaced agents and skills under .claude/agents/foxhunt/ and .claude/skills/foxhunt/. Phase 1: 5 auditor agents (isv-discipline, gpu-contract, reward-controller, code-hygiene, sp-critical-reviewer) + 5 workflow skills (sp-spec-writer, isv-slot-scaffolder, pearl-distiller, argo-deploy-helper, smoke-pilot) + 2 maintenance skills (memory-curator, stale-worktree-cleaner). Hook integration is warn-only PostToolUse via a single thin router; agents read memory but only pearl-distiller and memory-curator may write into memory/. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
235e838422 |
feat(sp20): Phase 1.4 (Path C) — kernel-arg refactor + aggregation kernel + wire-up
Atomic Phase 1.4 commit per `feedback_no_partial_refactor`. Wires the
SP20 fused-producer chain (Stats → Aggregate → EMAs → Controllers) into
GpuExperienceCollector's per-rollout-step path with one new aggregation
kernel + Phase 1.2 EMA kernel-signature refactor + production wire-up
+ tests + audit/spec/plan amendments, all in one commit.
## Path C decision rationale
The Phase 1.2 EMA kernel originally took 9 scalar value-args. The Phase
1.4 wire-up site (per-rollout-step in GpuExperienceCollector) needs to
feed per-env GPU-resident trade signals (trade_close_per_sample,
step_ret_per_sample, hold_at_exit_per_sample, packed factored
actions_out) into the kernel — forbidden via host sync
(`feedback_cpu_is_read_only`) and via memcpy_dtoh/htod
(`feedback_no_htod_htoh_only_mapped_pinned`). Path C refactors the EMA
kernel to take a device struct pointer (`SP20EmaInputs* ema_inputs`) +
adds an aggregation kernel that writes the struct on the GPU. Phase 1.2
had zero production consumers yet, so the sig change + Phase 1.4 wire-up
land atomically.
## What this commit contains
1. **EMA kernel signature refactor** (Phase 1.2 → Path C):
`sp20_emas_compute_kernel.cu` swaps 9 scalar args for a single
`const SP20EmaInputs* __restrict__ ema_inputs` device pointer.
Math semantics bit-identical. Rust `EmaInputs` mirrored as
`#[repr(C)]` byte-for-byte; new `pack_inputs_into_f32_view` helper
for tests + collectors that fill the struct via a mapped-pinned
f32-aliased buffer.
2. **New aggregation kernel** (`sp20_aggregate_inputs_kernel.cu`):
Per-env arrays (sliced to current rollout step) + sp20_stats outputs
(p50/std) + aux_dir_acc_reduce_kernel output [0] → SP20EmaInputs
struct. Block tree-reduce (4 stripes × bdim × 4 bytes shmem); no
atomicAdd. Aggregation rules per spec §4.5:
- is_close = OR over envs
- is_win = (≥ 0.5 of closed envs were wins)
- trade_duration = round(mean(hold_at_exit) over closed envs)
- action_is_hold = strict majority (count*2 > n_envs) HOLD
- alpha = 0.0 (Phase 2 forward ref — reward kernel)
- per_bar_hold_reward = 0.0 (Phase 3.2 forward ref — Hold-cost dual)
- aux_logits_p50/std/dir_acc forwarded from upstream
3. **Production wire-up in GpuExperienceCollector**:
4 kernel handles + 4 mapped-pinned buffers (struct fields, alloc,
init); per-rollout-step launch sequence after env_step (step 5c):
`Stats → Aggregate → EMAs → Controllers`. All on the same stream;
stream-implicit producer→consumer ordering. Gated on
`isv_signals_dev_ptr != 0 && trainer_params_ptr != 0` (matches the
existing SP14-β EGF chain pattern).
4. **Fold-boundary reset**:
3 new RegistryEntry records (sp20_ema_inputs_buf,
sp20_emas_internal_buf, sp20_emas_obs_count_buf) + 13 new dispatch
arms in training_loop.rs (10 SP20 ISV slot resets + 3 buffer resets).
Closes the pre-existing `every_fold_and_soft_reset_entry_has_dispatch_arm`
regression (was failing on fresh check after the SP20 ISV slot
registrations landed in commit
|
||
|
|
4e21c38b85 |
fixup(sp20): Phase 1.1 K=3 → K=2 retarget (production aux head is K=2)
The Phase 1.1 sp20_stats_compute kernel (
|
||
|
|
9d9ca3e6e7 |
spec(sp19+20): apply Q1(b) — Hold-reward EMA for Q-scale comparability
Q1 design decision (b): add explicit hold_reward_ema to center the per-bar Hold reward, so Q(Hold) and Q(trade) targets are scale-comparable. The marginal Q(trade) > Q(Hold) preference now comes from data variance in each state (high-aux states pull Q(Hold) more negative), not from structural scale asymmetry that depends on cost_scale magnitude. Q2 decision: keep 4-quadrant fixed (no ramping partials). Already in spec. Changes: - §4.2 Hold opp-cost: dual emission documented — R_per_bar_centered (= R_per_bar - hold_reward_ema) for Q-target/replay tuple, R_per_bar uncentered for hold_baseline_buffer (Component 1 baseline) - §4.5 Kernel 1: hold_reward_ema added (per-step on Hold-state bars only) - §5 data flow: per-bar reward path shows the centered/uncentered split - ISV slots: 9 → 10 (HOLD_REWARD_EMA_INDEX added) - §8 footprint: Component 2 LoC 70 → 90 (+20 for dual emission) - Total LoC estimate: 1620 |
||
|
|
defbd0abe1 |
spec(sp19+20): patch 7 review issues
P1 (must-fix bugs/gaps): 1. ASYM_RATIO_INDEX → LOSS_CAP_INDEX with explicit formula in §4.1 and §4.5 (was double source-of-truth — formula in §4.1, ISV slot orphaned) 2. Replay buffer schema change (per-bar aux_conf) added to §8 implementation footprint (~50 LoC additional, was invisible in original spec) 3. hold_baseline_buffer size specified = LOOKAHEAD_HORIZON_MAX = 30 bars (§4.2) 4. "4-tier gate" typo → "5-tier gate" in §7 P2 (clarifications): 5. alpha_ema centering documented as load-bearing for cold-start learning (§4.1) — not just for Q-target stability. EV is slightly negative at WR=46% with uninformed SP19 label; advantage-style centering rescues cold-start. 6. Q-scale asymmetry between Hold (uncentered) and trade (alpha_ema centered) documented as INTENTIONAL design choice (§4.2) — produces marginal Q(trade) > Q(Hold) preference that counteracts the Q(Hold) attractor. Behavioral test sp20_pure_noise validates the no-signal Hold default still works. P3 (tightening): 7. Behavioral test thresholds tightened (§4.6): - sp20_pure_trend: WR > 70% → > 90%, PF > 2.0 → > 3.0 - sp20_pure_noise: Hold% > 80% → > 95%, trades < 50 → < 20 |
||
|
|
730337375f |
spec(sp19+20): WR-first reward + multi-horizon label utilization
Combines SP19 (multi-horizon labels, already landed) + SP20 (WR-first reward) into one spec per pearl_no_deferrals_for_complementary_fixes. Goal: WR ≥ 55%, textbook PF ≥ 2.0, walk-forward stable, per-regime stable. Six components, atomic ship per feedback_no_partial_refactor: 1. Reward kernel (event-driven, 4-quadrant, asymmetric clamp ramped from wr_ema, multi-horizon directional ground-truth check) 2. Hold opportunity-cost (per-bar, dual emission for real reward + Hold baseline buffer) 3. n-step credit distributor (uniform over trade duration, fixes SP18 B-leg self-bootstrap bug at gpu_experience_collector.rs:4154) 4. Aux→Q confidence gate (sigmoid threshold, mean_a Q baseline avoids Hold-everywhere punishment) 5. 3 fused producer kernels (sp20_emas, sp20_controllers, sp20_stats) driving 9 ISV slots in [510..520) 6. 7 behavioral tests on RTX 3050 Ti, gating L40S deployment ~1,550 LoC total. Spec includes data flow, error handling philosophy, 4-tier test gate, success criteria with explicit failure modes. Cross-references: - pearl_event_driven_reward_density_alignment (design principle) - pearl_audit_unboundedness_for_implicit_asymmetry (asymmetric clamp) - pearl_separate_aux_trunk_when_shared_starves (aux gate leverages SP14-C) - project_metric_pipeline_inflation_audit (WR honest, Sharpe honest, goal grounded) - project_goal_wr_55_pf_2 (the goal this spec implements) |
||
|
|
44e1f58b69 |
docs(sp18): copy spec + plan to feat/sp18-combined branch
Mirrors the SP17 PP.1 plan-copy pattern. Plan and spec source-of-truth live on the sister `feat/sp17-dueling-q-network` worktree; this commit copies them onto the SP18 branch so subsequent commits reference local paths. Spec: docs/superpowers/specs/2026-05-08-sp18-reward-shape-hold-attractor-design.md Plan: docs/superpowers/plans/2026-05-08-sp18-reward-shape-hold-attractor.md Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
5417e27567 |
docs(sp15): amend spec v2 — 10 fixes from second critical review
Second critical review pass found 10 more issues, 3 critical. All addressed: CRITICAL: - §8.2 (3.1): ALPHA_SPLIT cold-start was unspecified — formula produces 0/0=0, not the claimed 0.5 sentinel. Now: ISV slot initialized DIRECTLY to 0.5 in trainer constructor; formula takes over only after both grad-norm EMAs accumulate N_warm non-zero observations - §9.2 (3.5.4): plasticity now performs TWO-STEP recovery: (1) Flat all positions at fire bar (close current losing trade), (2) engage warmup cooldown forcing Hold for M_warm bars. Without step 1, forced Hold preserved the losing position that drove drawdown for the entire 200-bar warmup - §12.2: stale "5-10%" baseline cost estimate updated to "15-25%" matching §6.4 (was contradicting earlier amendment) IMPORTANT: - §9.2 (3.5.5): DD_TRAJECTORY_DECREASING threshold 0.02 hardcoded → ISV-driven via new slot DD_TRAJECTORY_FLOOR (slot 441, 25th percentile of running dd_pct distribution) - §8.2 (3.5): HOLD_FLOOR_ALPHA tracked from rolling 95th percentile of |Q_dir| (NOT running max — was outlier-ratchet vulnerable) - §9.2 (3.5.3): MEDIAN_STREAK_LENGTH formerly undefined in cooldown K formula → ISV-driven via new slot 442, running median of observed loss-streak lengths via two-heap algorithm - §9.2 (3.5.2): asymmetric reward × α split compound interaction explicitly stated as intentional with POS_CAP as binding ceiling NIT: - §7.4: "2C: Group 3 (4 tests)" → "(5 tests)" (was off-by-one after 2.22 added) - §9.2 (3.5.4): Xavier → Kaiming-He init for advantage head reset (architecturally appropriate for ReLU-gated activation chain) ISV_TOTAL_DIM: 441 → 443 post-SP15 (added DD_TRAJECTORY_FLOOR and MEDIAN_STREAK_LENGTH at slots [441..443)). 46 SP15 slots total. File: 799 → 811 lines. Spec is now consistent end-to-end with no contradictions between sections, no hardcoded values violating feedback_isv_for_adaptive_ bounds, and no underspecified load-bearing parameters. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
f3cb0c78c2 |
docs(sp15): amend spec — 14 fixes from critical review
Critical review (combined fresh-eyes subagent + self-critique) flagged 14 issues. All addressed in this amendment. CRITICAL (blocks implementation): - §4.3 NEW: ISV slot allocation map [397..441) per phase, prevents parallel-dispatch ISV-table conflict (was: undefined, would break Approach B) - §4.4 NEW: Phase 2A ↔ Phase 1.2 ABI contract (LobBar struct as canonical contract, Phase 2A defines first; Phase 1.2 conforms). Restores parallelism dependency - §9.2 (3.5.4): plasticity injection now MANDATES cooldown engagement via PLASTICITY_WARM_BARS_REMAINING slot 438. NEW anchor test 2.22 plasticity_cooldown_interlock prevents random-tail-output retrigger loop - §9.2 (3.5.5): recovery curriculum replaced episode-level metadata (didn't exist in transition-level PER) with per-step DD_TRAJECTORY_ DECREASING signal — same goal, no PER restructure IMPORTANT (rework prevention): - §9.2 (3.5.3): cooldown K threshold from running mean of per-trade PnL (NOT variance — variance is LOW during streaks, would delay cooldown when needed) - §8.2 (3.5): Hold floor function form specified as bounded sigmoid with α/k/ε₀ ISV-driven; was "monotonic_inv_func" (undefined family) - §6.2: cost kernel per-side semantics explicit. Half-spread × side_ indicator at entry AND exit; commission_per_rt on close. Resolves round-trip vs per-side ambiguity - §6.2: OFI impact λ initial value 2e-4 + ISV refit methodology; was "empirical fit from MES historical" (hand-wavy) - §6.4: 8 baselines now mandate shared trunk forward; honest 15-25% overhead estimate (was 5-10%, optimistic) - §12.3: Q9 burn policy explicitly honor-system, not enforcement. Mitigations described as deterrents not enforcement - §10.6: Phase 4.5 thresholds CLI-config-driven with rationale, anchored empirically post-Phase-1 (was hardcoded > 1.0) NIT: - §9.2 (3.5.2): multiplier vs SP12 cap interaction explicit (applied BEFORE cap) - §8.2 (3.1): α=0.7 hardcoded → ISV-driven from grad ratio (was feedback_isv_for_adaptive_bounds violation) - §13: pearls reframed as candidates pending validation per discipline rule; pre-naming pressure removed Test count: 21 → 22 (added 2.22). Phase 2 LOC: 2000 → 2100. ISV_TOTAL_DIM: 396 → 441 post-SP15. All 14 issues addressed in spec text. Ready for user review. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
fb5f2394d0 |
docs(sp15): trader discipline and recovery design spec
Design spec for SP15 — addresses the train-dd4xl downward spiral post-mortem (8 epochs, trades 131k→64k, sharpe 79→44 with monotone descent across active_frac, dir_entropy) and the walk-forward audit finding (test_start..test_end slice generated but never consumed, val IS the selection set, no sealed test). Six phases over ~13-21 days: - Phase 0: SP14 EGF ISV-driven retune (gate1=closed forever fix) - Phase 1: honest numbers — unified sharpe kernel, cost-net (commission + spread + OFI-impact, dev/prod parity), drawdown reporting, 8 counterfactual baselines, dd_pct as foundational state input, --holdout-quarters/--dev-quarters CLI, consume abandoned test slice - Phase 2: 21 behavioral tests on dev RTX 3050 Ti, pre-commit hook gates argo-train.sh - Phase 3: 5 trader teachings (r_quality/r_discipline split, explicit cost, quadratic DD, regret, confidence-aware Hold floor) - Phase 3.5: 4 recovery mechanisms (asymmetric reward under DD, cooldown gate, plasticity injection, recovery curriculum in PER) - Phase 4: L40S walk-forward Q1-Q7, Q8 final dev, sealed Q9 OOS Cross-cutting discipline rule: no new pearl/controller/kernel/ISV slot/reward term ships without a Phase 2 behavioral test that proves the intended trader behavior. Hard rule, no exceptions. Decisions captured from 7 clarifying questions answered 2026-05-06. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
40e737a181 |
docs(sp14): spec v3 — K=4 direction + ISV-adaptive β rate limiter
Two further user-flagged corrections from second-critical review:
1. Direction Q-head emits K=4 actions, NOT K=3.
Authoritative source: state_layout.cuh:123-126
#define DIR_SHORT 0 // open/maintain short
#define DIR_HOLD 1 // keep current position (no-op)
#define DIR_LONG 2 // open/maintain long
#define DIR_FLAT 3 // close all to zero
The "default: 3 — Short/Flat/Long" comment at gpu_dqn_trainer.rs:2438
is stale pre-SP13. Production callsites all set branch_0_size: 4.
SP13 added DIR_HOLD as a separate fourth direction action (Hold-pricing).
The config default comment was never updated when DIR_HOLD landed.
This is exactly the feedback_trust_code_not_docs failure mode — a
single stale comment would have silently corrupted the q_disagreement
signal (Hold/Flat being indices 1/3 instead of just Flat=1).
Updates:
- B.1 action-space context: 4 actions with Hold AND Flat both
non-committal (Hold = keep position, Flat = exit to zero)
- B.2.3 q_disagreement mapping: K=4↔K=2 with both Hold and Flat
masked from disagreement signal (no new directional commitment
to evaluate); only Short and Long picks contribute to disagreement
- Edge case handling for all-Hold/all-Flat batches
2. Adaptive β rate limiter (was structural β=0.9).
Per feedback_isv_for_adaptive_bounds, β should be signal-driven
not hardcoded. v3 derives β from variance of α_grad_raw, mirroring
the k_aux/k_q variance-driven steepness pattern in B.2.5.
Formula:
β = clip(β_base + variance_alpha_raw / variance_ref_alpha,
[β_base, β_max])
β_base = 0.5 (light smoothing baseline; ~2-step half-life)
β_max = 0.95 (heavy smoothing; ~20-step half-life)
Stable α_grad_raw → β = β_base (preserves directional intent)
Volatile α_grad_raw → β → β_max (dampens jitter)
Adds 2 ISV slots:
ALPHA_GRAD_RAW_VARIANCE_EMA_INDEX (Welford variance)
BETA_RATE_LIMITER_ADAPTIVE_INDEX (current β value)
ISV slot count: 11 → 13 (net +2 for variance + adaptive β).
LOC estimate: ~1150 → ~1180 (negligible delta; 3 Welford
variances now in alpha_grad_compute_kernel instead of 2).
HEALTH_DIAG pearl_egf_diag emit updated to expose all three
adaptive scalars (α, β, k_aux/k_q) plus all three driving
variances (var_alpha, var_aux, var_q) for full observability.
Verified against current code at HEAD
|
||
|
|
d243a6f080 |
docs(sp14): spec v2 — critical review corrections
8 corrections from code-anchored critical review at HEAD
|
||
|
|
eaf4adcb98 |
docs(sp14): spec — aux→Q wire + Earned Gradient Flow pearl
Designs the SP14 chain on top of SP13 Layer B (HEAD
|
||
|
|
ba83fcd1f5 |
docs(sp13): v3 spec + plan — Hold-pricing replaces Hold-elimination
P0a.T3 v2 implementer's audit revealed `DirectionAction` enum doesn't exist; the codebase uses an 8-variant fused `ExposureLevel` (ShortSmall/Half/Full, Hold, LongSmall/Half/Full, Flat) with cross-crate consumers across 77 files and 32+ test files pinning the 8-variant invariant. Atomic Hold elimination would cascade massively. User insight (2026-05-04): Hold being FREE is the bug, not Hold itself. MFT trading legitimately needs multi-bar holds; we want the model to use them deliberately, not as a CQL-bias lazy default. Holding isn't free in the real world — broker fees, margin interest, opportunity cost. v3 reframes as Hold-pricing: - 4-way action space stays; ExposureLevel::Hold stays; no cross-crate cascade - 3 new ISV slots (380-382): HOLD_COST_INDEX, HOLD_RATE_TARGET_INDEX, HOLD_RATE_OBSERVED_EMA_INDEX - Hold-rate observer: small GPU kernel + Pearls A+D smoothing - Hold-cost controller: 5-line deficit-driven formula (excess > target → cost rises 1×→5× base; observed ≤ target → relax) - Per-bar reward subtraction at action == DIR_HOLD site - 2 GPU oracle tests for the controller P0a.T3 cuts from ~250 LOC + 32-test cascade → ~120 LOC additive. T1+T2 already-staged work unchanged. T4/T5/Layer B/C/D structure preserved. Tension with pearl_event_driven_reward_density_alignment acknowledged in spec — per-bar Hold cost is exposure-NEGATIVE (away from Hold), models real economic carry, ISV-bounded by controller. Inverse of the pearl's failure mode. Faithful reward modeling, not artificial shaping. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
0ca45ef61d |
docs(sp13): v2 spec + plan after critical review
Spec v2 supersedes v1 with five major fixes from the critical review: - Drop alpha-vs-benchmark reward (mathematically tautological — Long trades had alpha = -costs always against always-long benchmark) - Split Phase 0 into 0a (Hold-only, clean test of user hypothesis) + 0b (aux amplification probe, only if 0a partial), avoiding the v1 confounded experiment - Bound direction-skill bonus relative to |alpha| (cap_ratio × |alpha|) to prevent reward gaming on small-alpha correct-direction losers; spec adds 8-quadrant worked-example matrix verifying the no-negative-EV invariant - Replace 5-epoch absolute-threshold gate with 10-epoch trajectory criterion to avoid false-negatives from aux head underconvergence - Aux_w controller adds dual-EMA stagnation detector (decays toward base when no improvement) — prevents permanent destabilization of Q-head in data-limited case Plan v2 mirrors spec changes: - Phase 0a (Hold elimination + dir_acc instrumentation, ~300 LOC, 1 atomic commit) - Phase 0b (aux_w controller replacement, conditional, ~80 LOC) - Layer B (aux head regression -> binary classification, ~120 LOC) - Layer C (skill bonus + luck discount with calibrated bounds, ~150 LOC) - Layer D (30-epoch validation + 3 new pearls) ISV slot allocation [372..380): drops slot 371 (BENCHMARK_PNL_CUMULATIVE), adds 374 (AUX_DIR_ACC_LONG_EMA — stagnation), 379 (SKILL_BONUS_CAP_RATIO). Plan integrates Explore-agent touch-list (50-70 sites) with corrected enum ordering: Short=0, Long=1, Flat=2 (preserves codebase Short-first convention, avoids 30+ stale-comment churn). Adds direction-bias signal handling at experience_kernels.cu:5159 (Hold's 0.5 softening gate vanishes -> [1, 1, 0]). Three new pearls planned for Layer D close-out: - pearl_redefine_success_for_predictive_skill - pearl_skill_bonus_must_be_alpha_bounded (calibration lesson) - pearl_reward_quadrant_audit_required (meta-pearl) Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
ecab584c32 |
spec(sp12): per-trade event-driven reward composition (v3)
Three architectural changes in one unified design to push the model from
HFT noise extraction (62% trade rate, sharpe-gaming) toward MFT alpha-hunting:
1. Asymmetric bounded cap (-10/+5) — restores loss aversion erased by
SP11 symmetric cap. 2:1 ratio matches Kahneman/Tversky prospect theory.
Anchor: pearl_audit_unboundedness_for_implicit_asymmetry.
2. Min-hold soft penalty with temperature curriculum — patience requirement
at exit. Soft factor = deficit/(deficit+T), T anneals 50→5 over 50 epochs.
Forces commitment without paralysis in early training.
3. Zero per-bar shaping (gate micro/opp_cost on events) — eliminates
continuous-reward gradient that pulls toward continuous exposure.
Anchor: pearl_event_driven_reward_density_alignment.
Combined: reward fires only on trade events with prospect-theory loss
aversion + commitment requirement. Pure per-trade event-driven Q-learning
properly aligned with per-trade P&L objective.
~50 LOC across 3-4 files. No new ISV slots in Phase 1 (constants only).
Cost ~€1.30 (€0.30 smoke + €1.00 30-epoch validation).
Empirical motivation: train-multi-seed-pmbwn 50-epoch on commit
|
||
|
|
85069ba75e |
spec(sp11): amend B1b smoke-recovery — z-score normalization for mag-ratio canary
smoke-test-6wd2c on commit |
||
|
|
9395b983cc |
spec(sp11): fix all 13 review issues — straight-up implementation
Critical bugs fixed:
1. Z-score formula: was mean(Δ)/mean(|Δ|), bounded to [-1,+1] — sigmoid
never saturated, controller stuck in [0.27, 0.73]. Now true Z-score:
delta_ema / sqrt(delta_var_ema) with two-pass var. Renamed canary
slot 351 from VAL_SHARPE_STD_EMA to VAL_SHARPE_VAR_EMA.
2. Component weight renormalization added — Σweights = 1 enforced
post-floor so PopArt's reward-magnitude EMA stays uncontaminated.
3. SABOTEUR_MIN bound violation — engagement-floor (0.1) was silently
pulling output below 0.5 stated minimum. Added post-multiplication
clamp to [SABOTEUR_MIN, SABOTEUR_MAX].
4. ratios[] undeclared — now __shared__ float ratios[6] block-loaded
once from ISV.
5. §3.6 slot count off-by-five (15 → 20). All 20 slots get FoldReset
entries; cold-start window behavior specified explicitly (no NaN path).
Architectural fixes:
6. Q-overconfidence concern addressed in scope — controller's
REWARD_CF_WEIGHT_INDEX covers CQL conservatism. No SP11′ deferral;
if T10 fails, extend controller outputs (e.g., target-update τ).
7. Curiosity recompute-at-replay specified (§3.5.1) — replay buffer
stores base reward only; curiosity added per-tuple at replay time
against current visit-count and current ISV. Eliminates stale-signal
replay contamination that would worsen ep1-overfitting.
Smaller fixes:
8. Novelty signal specified — SimHash 42×16 projection → 16-bit code,
1M-slot per-bucket count table, 1/sqrt(1+count). Hash table lives
in state-reset registry (cleared at fold boundary).
9. Saboteur engagement operationally defined — per-bar
|reward_with_saboteur − reward_without_saboteur| > EPS_ENGAGEMENT,
block tree-reduce, EMA'd. EPS_ENGAGEMENT = 0.01 × PnL_EMA.
10. Duplicate sentence in §7 component-starvation paragraph removed.
11. Validation criteria strengthened (§9): aggregation across 9
trajectories specified; primary metric = median peak-epoch ≥ 10
(baseline median peak-epoch = 1); secondary = drop-from-peak ≤ 5%
by ep20; tertiary = ep20 mean ≥ baseline ep1 mean. §9.3 fix-forward
response codified — no rollback.
12. Phases split per project pattern — Layer A additive (3 commits:
slots, canaries, controller), Layer B atomic consumer migration
(1 commit, including replay-time curiosity), Layer C validation +
close-out. No falsification gate; smoke is validation only.
13. Pearls A+D applied to controller outputs (§3.4.1) — chained
apply_pearls_ad_kernel after controller writes scratch, smoothed
values land in ISV. Consumers read smoothed slots.
Constants surviving (Invariant-1 anchors only, all rate-not-regime):
EPS_DIV, WEIGHT_HARD_FLOOR, SABOTEUR_MIN/MAX, CURIOSITY_PERMANENT_FRACTION,
CURIOSITY_BOUND_FRACTION, WEIGHT_FLOOR_FRACTION, ENGAGEMENT_FLOOR.
Spec is now full straight-up implementation; smoke validates Layer A
infrastructure and Layer B consumers, T10 validates SP11 success metric.
~1550 LOC across 3 commits in Layer A + 1 atomic commit in Layer B +
close-out in Layer C.
|
||
|
|
d49fbbe4e2 |
spec(sp11): reframe curiosity as feature + add permanent floor
User correction: curiosity is the *fix* for the ep1-peak overfitting pathology, not a hazard to defend against. Reframed §7 from "Risks" to "Design notes" — curiosity bound is a signal-relative scale, not a defensive cap. Fix the contradiction this exposed in the formula: previous `curiosity_pressure = stagnant_or_worse * curiosity_bound` went to zero when improving, which would cancel the always-on exploration the §7 narrative now relies on. Replace with permanent-floor pattern per pearl_blend_formulas_must_have_permanent_floor: curiosity_floor = 0.2 * curiosity_bound (CURIOSITY_PERMANENT_FRACTION) curiosity_dynamic = stagnant_or_worse * curiosity_bound curiosity_pressure = max(curiosity_dynamic, curiosity_floor) Now curiosity is always ≥ 20% of bound (anti-overfitting baseline) and rises toward the bound when stagnant (stagnation breaker). Updated unit-test guidance to assert pressure > 0 even at z=+10. CURIOSITY_PERMANENT_FRACTION=0.2 added to Invariant-1 fraction list. Saboteur-rising-with-improvement reframed as adversarial-load feature rather than over-stress risk. |
||
|
|
a72a23e6b1 |
spec(sp11): reward as controlled subsystem — fully ISV-driven
Brainstorm spec for SP11. Resolves the policy-stagnation pathology
surfaced in T10 train-multi-seed-xkjkb seed-0 ep0-14: model finds a
stable fixed point at ep1 (peak val sharpe 80.61), then OVERFITS to
it across remaining epochs (decline 80.61 → 70.58). Q-values grow but
val performance declines because reward function has no improvement
pressure.
Architecture (every input ISV-driven):
- Z-score-driven adaptation (no hardcoded "improving" threshold):
improvement_z = val_sharpe_delta_ema / max(val_sharpe_std_ema, EPS)
- 10 ISV outputs: 6 component weights + curiosity_pressure +
saboteur_intensity_mult + adaptive weight_floor + curiosity_bound
- 5 ISV canaries: val_sharpe_delta + val_sharpe_std (Z-score noise
estimate) + 6 per-component grad ratios + saboteur engagement +
PnL magnitude EMA (signal-relative curiosity bound)
- 4 new producer kernels (controller + 3 canary computers)
- Audit + migrate hardcoded cf_weight=0.3 in mse_loss_kernel.cu:318
and c51_loss_kernel.cu:789, plus other shaping multipliers
- NEW reward dimension: curiosity bonus, bounded by PnL magnitude
Per pearl_controller_anchors_isv_driven: every threshold replaced with
sigmoid(z) — no constants encode "what counts as improving". Per
pearl_blend_formulas_must_have_permanent_floor: every weight has
adaptive floor preventing zero-out. Per pearl_engagement_rate_self_
correction: saboteur intensity self-corrects via engagement rate canary.
Per pearl_cold_start_exit_signal_or: improvement signal OR'd from
multiple canaries so single-signal-failure doesn't stall controller.
New pearl authored alongside spec: pearl_reward_as_controlled_subsystem
— meta-principle that every reward path degree of freedom is a unified
controller output. Subsumes controller-anchor pearl at the reward layer.
Scope: 20 ISV slots, 4 producer kernels, audit + migration of
hardcoded reward shaping constants, 1 atomic commit.
~1300-1700 LOC; ~2.5-3 hours subagent work.
Success metric: val_sharpe[ep20] > val_sharpe[ep1] (Fix 33-38 baseline
peaked at ep1; SP11 should shift peak later as model continues
learning).
|
||
|
|
6ae9bc0055 |
spec(sp10): unconditional Thompson selector + ISV-driven temperature
Brainstorm spec for SP10. Resolves the val-Flat-collapse pathology that persisted through Fix 33-37: the eval-time argmax in experience_action_ select picks Hold deterministically every val bar (dir_entropy=0, trade_count=1 in 214k bars) regardless of controller state. Architecture: - Delete `if (eval_mode)` argmax branch in direction-selector kernel - Use temperature-blended Thompson: q_eff = E[Q] + τ × (Thompson - E[Q]) - τ = clamp(intent_eval_divergence / divergence_target, 0.5, 2.0) - τ self-corrects: collapse → high τ; healthy → low τ; permanent 0.5 floor - Reuses SP9's intent_eval_divergence_compute_kernel (extended, not new) Per pearl_controller_anchors_isv_driven: τ is ISV-driven, no hardcoded constants beyond Invariant 1 numerical anchors (clamp range). Per pearl_blend_formulas_must_have_permanent_floor: MIN_TEMP=0.5 ensures eval ALWAYS has stochasticity. The pearl_thompson_for_distributional_action_selection was about Bellman TARGET argmax (selector/target symmetry). It does NOT prohibit Thompson at the rollout selector. SP10 amends the pearl to clarify. Scope: 1 ISV slot, kernel modification (no new kernel — extend SP9's producer), 1 consumer kernel rewrite, test update, pearl amendment, audit doc Fix 38. ~300-500 LOC, single atomic commit per feedback_no_partial_refactor. |
||
|
|
38d2408604 |
spec(sp9): Kelly cold-start warmup floor — fully ISV-driven design
Brainstorm spec for SP9 (Fix 37). Resolves the val-Flat-collapse pathology
surfaced in T10 train-multi-seed-wsnc6 ep1-3: training intent_dist_f=0.47
but eval_dist_f=0.00 because Kelly cap pinned eval mag to Quarter, with
trade_count=1 in 214k bars (chicken-and-egg: cap closed → no trades → no
Kelly samples → cap stays closed).
Architecture per pearl_controller_anchors_isv_driven and
pearl_cold_start_exit_signal_or:
final_kelly_f = max(measured_kelly_f, warmup_floor)
warmup_floor = base_floor × (1 − combined_confidence)
base_floor = clamp(q_var_mag / q_var_mag_ema, MIN, MAX)
combined_confidence = max(statistical, behavioral, temporal)
All thresholds ISV-driven via Pearl D Wiener-α; only Invariant 1 numerical
anchors (clamp ranges, EPS) remain as constants.
Scope: 9 new ISV slots, 6 new GPU producer kernels including a
mandatory eval_dist GPU migration (eliminating DtoH readback at
gpu_backtest_evaluator.rs:1353 per feedback_no_cpu_compute_strict).
~1000-1200 LOC, single atomic commit per feedback_no_partial_refactor.
Pearls authored as part of brainstorm:
- pearl_controller_anchors_isv_driven (regime-encoded constants)
- pearl_cold_start_exit_signal_or (OR'd signals across statistical /
behavioral / temporal axes)
|
||
|
|
8f484f3312 |
spec(sp7): v2 — flatness-gated targets + Wiener α + sentinel-bootstrap
Self-review found 4 violations of feedback_isv_for_adaptive_bounds in v1: 1. TARGET_CQL_OVER_IQN=2.0 hardcoded → flatness-modulated per-branch target_ratio = ANCHOR × (1 − flatness[b]). Realizes the original pearl_q_flatness_gates_conservatism brainstorm idea. 2. TARGET_C51_OVER_IQN=1.0 hardcoded → mirror direction: target_ratio = ANCHOR × flatness[b]. 3. ALPHA_META=0.01 fixed → Wiener-optimal α from per-branch state-tracker EMAs (16 new ISV slots: diff_var/sample_var × 2 heads × 4 branches). 4. Cascaded floor cleanup: compute_adaptive_budgets line 3384 floors reads to 0.02/0.05 always, fighting the controller. Replaced with sentinel-aware bootstrap: cold-start uses BOOTSTRAP value, post-write takes ISV value verbatim. Acceptance criteria rewritten to be ISV-relative where the corresponding signal exists (q-spread vs NOISY_SIGMA, label_scale vs NOISY_SIGMA band). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
0808ebe3e5 |
spec(sp7): loss-balance controller — adapt CQL/C51 budgets to grad-norm ratio
Outcome-driven controller replacing Pearl 2's hardcoded-0 CQL budget and
floor-pinned C51 budget with multiplicative ratio adaptation. Targets
`grad_split_bwd cql/iqn = 2.0` and `c51/iqn = 1.0` per slice (trunk, dir,
mag). Slow EMA (α=0.01) consistent with existing controller pearls; Pearl
A sentinel-bootstrap on fold boundary; Pearl D Wiener-optimal smoothing.
Replaces ghost docstring at fused_training.rs:3409 ("0.10×(1−regime)×health"
formula was never implemented). Pearl 2 keeps owning IQN budget (reference
denominator) and FLATNESS_BASE (consumed by NoisyNet σ pearl); stops
writing CQL/C51/ENS slots so the new controller is the sole driver.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
ca6e0007da |
docs(sp6): brainstorm spec — per-branch consumer infrastructure refactor
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
6e6e0fa114 | docs(sp5): remove duplicate Layer A close-out section (superseded by Layer D) | ||
|
|
af8fe57d1c |
docs(sp5): apply critical review fixes — Layer D, Kelly carve-out, Pearl 4 mitigations, acceptance loosened
Critical self-review surfaced 15 issues; user directed fixes:
- "a no cpu path": KEEP host-EMA close-out (rule compliance) — but split
into separate Layer D atomic commit (was mis-scoped as Layer A)
- Pearl 4 kept: 3 concrete risks documented (constant-β proof break,
β2 memory reset destabilization, ε numerical envelope), structural
envelope bounds added, ALPHA_META halved, ε-only fall-back path defined
- Pearl 6 Kelly cross-fold persistence carve-out: separate slot range
280..286, NOT in SP4 fold-reset registry, Invariant 1 architectural exception
- Commit ordering Pearls 1 → 3 → 2 → 4 → 5 → 6 → 8 → 1-ext (resolves
Pearl 2 circular dependency on 1+3)
- Acceptance criteria: correctness gates (must pass) + performance gates
(loosened to "not catastrophically negative", within 2σ of pre-SP5)
- Pearl 7 timing: explicitly Layer C step 4, post-Layer-B + 3-seed validation
- Pearl 8 enumeration: 4 slots (TRAIL_DIST_PER_DIR per direction)
- Pearl 9 collapsed: 0 slots (Thompson achieved via Pearl 1's atom adaptation)
- Total slot count corrected: 110 (was 120-128 inconsistent)
Layer structure: A (additive, 8 commits) → B (atomic, 11 consumers) → C
(validation + Pearl 7 investigation) → D (host-EMA close-out, separate
atomic commit). Layer D split off from A's "close-out" because PnL
aggregation pipeline migration is its own architectural concern.
User final review pending before invoking writing-plans skill.
|
||
|
|
3361944456 |
docs(sp5): resolve open questions Q1=c, Q2=a, Q3=a
Per user direction:
Q1 (Layer A granularity): per-pearl commits (~9 commits, decision c)
Q2 (Pearl 7 timing): investigate-only in SP5; fix in follow-up
if Bin(2,0.5) persists post-Pearls-1-3 (decision a)
Q3 (Pearl 4 Adam β): include as designed; accept theoretical risk
with Pearls A+D + EPS_CLAMP_FLOOR mitigation;
Layer C smoke monitors for destabilization;
fall-back to ε-only if observed (decision a)
Spec section updated:
- Layer A description: per-pearl commit structure
- Pearl 7 framed as investigation-only with conditional follow-up
- Pearl 4 documents theoretical caveat + mitigation + fallback
User final review pending before invoking writing-plans skill.
|
||
|
|
5f1c1eec51 |
docs(sp5): comprehensive per-branch + per-group adaptation spec — close all hardcoded-multiplier deferrals
SP5 design covers every known adaptive-parameter deferral in the DQN training loop in a single coherent project. After SP5: zero hardcoded multipliers, every adaptive value ISV-driven via Pearls A+D. 9 pearls + 1 sweep close-out + 1 validation milestone: 1-3. Per-branch atom span / loss budget / NoisyNet σ (52 slots) 4. Per-group Adam β1/β2/ε ISV-driven (24 slots) 5. Per-branch IQN τ schedule (20 slots) 6. Kelly cap signal-driven floors (6 slots) 7. dist_q/h/f Bin(2,0.5) audit + action_select fix (0-8 slots) 8. Trail stop signal-driven thresholds (6-8 slots) 9. Thompson direction-branch temperature (4 slots) 1-ext. Per-branch C51 num_atoms (4 slots) Layer A close-out: 5 host-EMA host→GPU migrations Validation: 3-seed × 50-epoch acceptance gate Total: 120-128 new ISV slots, ~5000-7500 LOC, 11-13 producer kernels, ~12 consumer migrations. Layer A (additive infrastructure, ~15 commits) → Layer B (atomic consumer migration, single coordinated commit) → Layer C (validation + cleanup). Mirrors SP4's layer pattern. Triggering data: train-multi-seed-cv2mw 50-epoch L40S baseline (terminated F0 ep10) revealed magnitude head Q-flatness, eval collapse, and frozen action distributions. Plus all SP4 close-out + sweep deferrals folded in per user direction "no deferrals — make a single plan based on ALL findings". Spec at: docs/superpowers/specs/2026-05-01-sp5-magnitude-differentiation-and-eval-collapse-design.md User review pending before invoking writing-plans skill. |
||
|
|
24accea774 |
fix(sp4): migrate fast/slow grad-norm EMA update to GPU per feedback_no_cpu_compute_strict
Layer C close-out C1 redesigned. Original plan claimed
grad_norm_slow_ema_pinned was orphan post-Mech-6 migration; verification
surfaced a SECOND live consumer (fold_warmup_factor_update, commit
|
||
|
|
d8666a232f |
docs(sp4): fix all 11 review findings — spec is now source of truth
Resolved every issue from the critical self-review:
1. Param-group count: 7 → 8 throughout. Slot total: 36 → 40 (8 groups × 3
per-group families = 24 + 16 single = 40). IQL high-tau and low-tau
are 2 distinct param groups (separate buffers + Adam states), not
one. The "(Resolved during implementation)" hand-wave deleted.
2. Reset semantics section rewritten — no longer references Xavier-
derived bootstraps. Sentinel 0 per Pearl A; Wiener-state triples
reset to 0 alongside ISV slots (160 reset entries total).
3-4. Weight decay and L1 lambda hardcoded α/bootstrap removed. Both
now follow universal Pearls A+D contract (sentinel-detect +
Wiener adaptive). For L1, the natural curriculum (λ ramps as
gradient differentiates) emerges from the signal itself, no
"Bootstrap = 0.0" constant needed.
5. ε naming collision resolved: ε_div = 1e-8 (Pearl D division-safety),
ε_clamp_floor = 1.0 (Pearl A consumer cold-start floor). Distinct
names, distinct purposes.
6. Per-signal kernel signature updated: removed stale `ema_alpha` arg,
added `wiener_state` pointer + `wiener_state_offset` + uniform
`alpha_meta` (structural meta-EMA constant).
7. Pearl C engagement counter race fixed. Replaced `clamp_engage_buf
[group] += 1` (race) with proper register-then-tree-reduce pattern
mirroring `dqn_grad_norm_kernel`. Per-thread register counter →
block-shared tree-reduce → single block-leader writes to
`clamp_engage_per_block_buf[engage_buf_offset + blockIdx.x]`.
Host sums across blocks. No atomicAdd, consistent with feedback.
8. Pearl D state buffer allocation spelled out: `wiener_state_buf:
MappedF32Buffer` of size 141 floats (47 slots × 3), per
`feedback_no_htod_htoh_only_mapped_pinned`. Reset-registry entries:
40 + 40×3 (new bound slots + Wiener triples) + 7×3 (retrofit
existing producers' Wiener states) = 181 reset entries.
9. `grad_norm_slow_ema` retirement spelled out in Layer C: removed
entirely (sole consumer Mech 6 migrates to ISV[GRAD_CLIP_BOUND]).
Also documented Q_ABS_REF=16 and H_S2_RMS_EMA=96 transition to
"monitoring-only ISV" — producers stay running, only the orphaned
Mech 1/2/5/6/9/10 consumer reads removed.
10. Histogram bin range fixed: linear-spaced bins from 0 to step_max
(avoids log(0) singularity from earlier log-spaced design).
Linear is also better for p99: top-of-distribution gets ~0.4%
resolution per bin. Degenerate "step_max == 0" branch handles
all-zero signals gracefully (skip ISV update, leave previous bound).
11. Pearl D-subsumes-Pearl-A claim CORRECTED: mathematically wrong.
At t=0, Pearl D's formula yields `x_mean[0] = 0`, not `x[0]`.
Pearl A's first-observation replacement requires an explicit
sentinel-detection branch in the producer. Both pearls are
necessary and complementary; they are not hierarchical.
Slot count math now consistent: 1 target_q + 4 atom_pos +
3×8 per-group + 1 grad_clip + 1 h_s2 + 8 wd_rate + 1 l1_lambda = 40.
Producer count: 15 fused (1 + 4 + 8 + 1 + 1).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
074b4f0865 |
docs(sp4): fold all 4 pearls into spec — A, B, C, D
PEARL A — First-observation bootstrap: eliminates Xavier-derived formulas (2.33 z-score, √2 std, √(2/K_in)) from the bootstrap section. Sentinel ISV[X]=0 at fold reset; producer step 0 detects sentinel, replaces directly with step_observation. Subsequent steps EMA-blend. Consumer cold-start safety via .max(1.0) numerical floor only (Adam-ε category, not magnitude). Truly zero magnitude constants in bootstrap. PEARL B — Fused per-param-group statistics oracle: producer count 36 → 14. Per-group fused kernel reads (params, grads, adam_m, adam_v) once and computes WEIGHT_BOUND, ADAM_M_BOUND, ADAM_V_BOUND, WD_RATE in one multi-pass operation. Trunk's oracle adds Pass E for L1_LAMBDA gradient-direction entropy. 4× memory bandwidth reduction. Cleaner conceptual unit (per-param-group bounds = one oracle). PEARL C — Engagement-rate self-correction: detects post-clamp feedback-loop saturation. For in-kernel clamps, theoretical engagement rate = 1% (top 1% by p99 definition). Producer-side rate-deficit EMA detects sustained deviation; force-bumps bound to step_max when detected. Resolves the "in-kernel feedback loop accepted" limitation from first draft. Per-Adam-kernel block-shared-memory engagement counter (no atomicAdd), block-wide reduce, host-side rate-deficit EMA. PEARL D — Wiener-optimal adaptive α: replaces all hardcoded EMA rates across 14 new producers AND 7 existing pre-SP4 producers. Per-step: α* = diff_var / (diff_var + sample_var + ε_num). Theoretically optimal under Wiener-filter analysis; subsumes Pearl A as t=0 edge case (both vars=0 → α=1 → first-observation replacement). On stationary signals: α→0 (smooth). On non-stationary: α→1 (track). Eliminates the recursion problem (α controlled by signal stats, not another α). ε_num = 1e-8 (Adam-ε numerical category). 3 floats of state per producer. 36+ hardcoded α values eliminated codebase-wide. Limitations section restructured: 6 of 8 first-draft limitations RESOLVED by pearls (in-kernel feedback loop, smoke time-budget, producer plumbing, stale-bound, 8 carved-out items, magnitude bootstrap formulas). Remaining 6 limitations are genuinely irreducible (F0 launch-scheduling variance, Layer B atomic flip risk, novel-pearl field-validation, equilibrium-formula non-stationarity, Pearl C counter-state cost, Pearl D state cost). SP4 now closes 100% of the magnitude/regularization/EMA-rate surface. No hardcoded scalars remain in the entire bound/clamp/EMA chain. Producer count: 14. ISV slot count: 36. Unit tests: 36. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
21b6058e07 |
docs(sp4): close all carve-outs — weight decay + L1 lambda fully signal-driven
Folds two new producer signals into SP4 scope, eliminating the earlier "carved out" exception: WEIGHT DECAY (per param-group, 7 new ISV slots): λ = |w·g| / max(||w||², ε) Derived from equilibrium analysis of d/dt(||w||²) = 2(w·g) - 2λ||w||². The equilibrium gradient-projection-onto-weight-direction divided by weight norm. Same theoretical-derivation category as Adam β values. EMA half-life α=0.005 (~140 steps). Bootstrap 1.0. L1 LAMBDA (trunk only, 1 new ISV slot — NOVEL PEARL): λ = (mean(|g|) / mean(|w|)) × D where D = (log K - H_observed) / log K is gradient-direction entropy deficit and H_observed = -Σ p[i]·log p[i], p[i] = ||g[:,i]|| / Σ ||g[:,j]|| L1 regularization-strength derives from gradient-direction entropy deficit across input features. When gradient is uniform across features (D≈0): network hasn't differentiated, λ=0 (no pruning). When gradient concentrates on few features (D≈1): network has identified what matters, λ ramps up to prune the rest. Self-curriculum — L1 strength tracks the emergence of feature differentiation. This extends pearl_adaptive_moe_lambda (regularization strength = EMA- tracked deficit of regularized quantity) to feature-redundancy domain. Pearl-name candidate (post-validation): pearl_signal_driven_regularisation_strength. Bootstrap λ=0 means cold-start = no L1 pruning; ramp-up only after gradient differentiates. Worst-case behavior is "L1 disabled" — graceful. Total ISV slot count: 28 → 36. Total producer kernels: 28 → 36. Effort estimate: 3000-4000 → 3500-4500 LOC, 1-1.5 → 1.5-2 weeks. Acknowledged limitations updated: removed item #8 (carve-out) since no carve-outs remain. Added items for L1 pearl novelty (untested) and weight decay equilibrium-formula non-stationarity. Both have graceful worst-case behavior and explicit validation criteria (#10 and #11) to detect anomalies. SP4 now closes 100% of the magnitude/regularization surface — no hardcoded scalars remain in the entire chain. AdamW config fields for weight_decay and l1_lambda removed from HyperParams to prevent accidental hardcoding regression. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
506f9443fe |
docs(sp4): fix critical issue J — post-clamp diagnostic dead path
Second-pass self-review found a new critical issue: under "diagnostic = clamp engagement", post-clamp diagnostics (Mech 9 weights, Mech 5 Adam m/v slots 36-43, slots 44-45) NEVER fire — because post-clamp |v| ≤ bound by construction, so producer-side comparison always returns false. Fix: split diagnostic implementation by clamp location. - Buffer-based clamps (Mech 1, 2, 10): diagnostic stays in producer (reads pre-clamp buffer, fires when step_max > bound). - In-kernel clamps (Mech 6, 9): diagnostic lives INSIDE the Adam kernel at the clamp step. Each Adam kernel takes a `diag_slot` arg alongside `weight_clamp_max_abs`; on clamp engagement, writes `nan_flags_buf[diag_slot] = 1`. Idempotent per-thread store (all threads writing 1, race-free per existing convention). No atomicAdd. Without this fix, slots 36-45 would be dead diagnostics under SP4 — detecting nothing, providing no signal. With this fix, every diagnostic slot fires on its corresponding clamp engagement regardless of where the clamp lives. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
cfe57109e3 |
docs(sp4): revise spec — fix 8 critical issues from self-review
Self-review found multiple problems with the first draft. This revision addresses all 8 critical issues: 1. P² parallelization claim was WRONG — P² is sequential. Replaced with dynamic-range histogram (3-pass single-block kernel: max-reduce → log-spaced bin → cumulative-from-top to find p99). 256 bins → ~0.4% quantile precision; numerical-precision derivation in same theoretical-constant category as floating-point precision. 2. Bootstrap values had wrong magnitudes — used Xavier σ instead of p99 of max-element. Recomputed: WEIGHT_BOUND[group] = 2.33 × √(2/K_in[group]); H_S2_BOUND = 2.33 × √2 ≈ 3.3. All bootstraps now from theoretical p99 under Xavier-init, not std. 3. EMA half-life of 700 steps (α=0.001) didn't converge in 5-epoch smoke. Revised α=0.005 (~140 steps) for weight/Adam producers — reaches ~99% convergence within one fold's training. Smoke now validates steady-state behavior, not just bootstrap. 4. Weight decay, L1 lambda, CLIP_MULTIPLIER were listed in scope but undesignable in producer-consumer pattern. CARVED OUT explicitly: weight decay + L1 λ → separate research-spec; CLIP_MULTIPLIER and MIN_CLIP subsumed by SP4's GRAD_CLIP_BOUND slot. 5. F0 ≥ 45 acceptance criterion was uncertain. Revised to F0 ≥ 37.5 (matches the 1e30-effectively-unclamped diagnostic smoke). The ~8-point F0 variance from launch-scheduling-shift is independent of clamp value; SP4 cannot guarantee F0=45 even with ideal design. 6. Per-param-group p99 plumbing concretized: each producer takes (offset, length) launch args; main DQN params buffer sliced into trunk/value/branch using existing param_sizes layout knowledge. 7. Pre/post-clamp feedback loop EXPLICIT: producer runs BEFORE consumer for buffer-based clamps (h_s2, target_q, atom_pos — in captured graph immediately before clamp). Producer runs AFTER for in-kernel clamps (weights, Adam m/v — feedback loop accepted with documented soft-anchor dynamics). No more hand-waving. 8. Unit test strategy: per-producer kernel test with synthetic Gaussian input → known p99 ≈ 2.33 → assert |computed - 2.33|<5%. 28 tests total. Catches algorithm bugs before L40S smoke. Effort estimate revised UP: 3500-4500 LOC (from 2-3000), 1.5-2.5 weeks (from 1-2). 28 producer kernels + 28 unit tests is more plumbing than first draft assumed. Self-review limitations section now explicit about: in-kernel feedback loop accepted, smoke time-budget marginally validates Adam EMAs, F0 may not return to 45, Layer B atomic-flip risk, plumbing density, stale-by-one-step bound, histogram-precision is design choice not tuning, carved-out items remain hardcoded post-SP4. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
e00736ce9b |
docs(sp4): signal-driven magnitude control design spec
Comprehensive design replacing every hardcoded magnitude multiplier in the SP3 mechanism stack (Mechs 1, 2, 5, 6, 9, 10) plus pre-SP3 mechs in the same magnitude-control surface, per feedback_isv_for_adaptive_bounds and feedback_adaptive_not_tuned. Core principle: the BOUND lives in an ISV slot, computed by a producer kernel as p99 EMA of observed signal magnitude. Consumer reads the slot and clamps directly — no multiplier between ISV read and clamp. Cold- start ε from theoretical-init bootstrap (Xavier, etc.) — same theoretical- constant category as Adam β values, not tuning knobs. Architecture: - 28 new ISV slots (7 base bounds × per-param-group split where appropriate) - Per-signal P² (Jain-Chlamtac) quantile producer kernels - Diagnostic = clamp engagement (sticky flag from producer's max-comparison) - Migration in 3 layers: additive infra → atomic consumer flip → smoke Out of scope: theoretical/structural constants (Adam β, Xavier formula, attention 1/√d, hidden_dim, num_atoms). EMA rates stay as documented statistical-design parameters (half-life ≈ observation time-window). Estimated: ~2-3000 LOC across 3 layers, 1-2 weeks, 1 L40S smoke. Awaiting user spec review before writing implementation plan. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
ed727b51c0 |
spec(dqn): SP2 + SP3 combined — F0 regression fix + Q-learning structural stability
Designs the fixes for the two SP1 leftover issues identified at SP1 closure
(commit
|
||
|
|
792812baa1 |
spec(dqn): SP1 numerical stability — revise per critical review
9 substantive issues addressed inline:
1. ISV-driven design elevated from 'if applicable' to MANDATORY for
all dynamic bounds in SP1 fixes. Numerical-stability ε bounds are
the only carve-out (Invariant 1). Hardcoded tuning constants for
dynamic ranges explicitly rejected.
2. F0 Sharpe regression criterion changed from absolute (≥55) to
ratio-based (≥95% of latest baseline; floor 53.08 currently).
Prevents iterative erosion across multiple fix commits.
3. 'F1 trending positive' replaced with concrete monotone-improvement
test: Best Sharpe at last epoch ≥ Best Sharpe at first epoch of
the same fold.
4. Pass criterion distinguishes NaN-CLAMPED-TO-ZERO (failure) from
'Genuine grad collapse' (legitimate observation, permitted) per
the existing infrastructure from commit
|