1305d6531b18307e9d1228c9a0f9505cefbd94de
5116 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
672c873571 |
cleanup: declarative rewrites for deferred-work TODOs across ml crates
- ml-dqn/dqn.rs: `apply_accumulated_gradients` is a scaffolding method whose real optimizer step lives in the fused CUDA trainer. The `grads` map was already being dropped silently; reword the comment to describe that split explicitly (incidental: see trainer path for the live gradient application). - ml-features/mbp10_loader.rs: strip the "TODO optimize with binary search" parenthetical from the docstring. Linear search over the sorted snapshot slice is the intended behaviour for current call sites. - ml-hyperopt/optimizer.rs: `optimize_two_phase` short-circuits after Phase A because `DQNTrainer` is not `Clone`. Describe that limit and point callers at `optimize_parallel` (which requires `M: Clone`) rather than a hypothetical Phase B. - ml-checkpoint/signer.rs: `fetch_key_from_vault` is currently an env-var resolver. Reword to say so plainly — no Vault client is wired into this crate, production uses K8s secrets injected as env. - backtesting/dbn_replay.rs: `DbnReplayEngine::from_bytes` remains an Err stub because `DbnParser` is gated behind the `databento` feature which this crate does not enable. Replace the pseudocode block with a declarative comment. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
952302149e |
feat(explainability): GPU-resident Integrated Gradients kernel
Implements the CUDA kernels (`interpolate_input`, `perturb_dimension`)
and wires `compute_gpu` in `integrated_gradients.rs` to run the full
IG algorithm on-device. Replaces the stub that previously returned
`MLError::ModelError("IG GPU kernels not available: cubins not yet
wired")` and fell back to CPU.
Algorithm: for each of `num_steps` interpolation points along the
baseline->input path, compute central-finite-difference gradients for
all `num_features` dimensions via two GPU forward passes per feature.
Single cubin (`ig_kernels.cubin`) compiled via build.rs (following the
crates/ml/build.rs pattern — nvcc, `-arch=sm_\${CUDA_COMPUTE_CAP}`, O3,
f32) and embedded with `include_bytes!`. The `forward_fn` consumers
operate on `GpuTensor`, so no dtoh round-trips inside the inner loop —
only the scalar `[1]` tensor output of each forward pass is pulled to
host (via `to_scalar`) per gradient sample. The three scratch buffers
(`interpolated`, `x_plus`, `x_minus`) are allocated once outside the
step loop and reused across all steps and features. No atomicAdd —
the kernels are trivial 1-D element-wise writes.
Tests: existing CPU tests pass unchanged. Added GPU smoke test
`test_ig_compute_gpu_linear_model` (gated `#[cfg(feature = \"cuda\")]`
+ `#[ignore]`) that builds a linear model as a `GpuTensor`-native
forward (elementwise mul + mean), verifies the completeness axiom
within 1%, and cross-checks GPU attributions against the CPU path
within 5%. Passes locally on RTX 3050 Ti (sm_86).
Removes 3 TODO markers and the embedded `_IG_CUDA_SRC` const that
were awaiting this work, along with the `#[allow(unused_variables)]`
stub attribute on `compute_gpu`.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
25b1e305c7 |
cleanup: reword GPU-kernel deferral TODOs in ml crates
- ml-explainability/integrated_gradients.rs: the GPU IG path is a stub that returns an error so callers fall through to CPU. Drop the three TODO markers and describe the situation declaratively — the reference CUDA source is kept as a follow-up anchor. - ml-supervised/mamba/mod.rs: dropout in both inference and training paths is simulated via the `1 - dropout_rate` scalar bake-in; no dedicated GPU dropout kernel is wired on the supervised path. The hidden-state carry-over in the SSD scan requires a 3-D narrow the GpuTensor API does not expose, and inference only consumes the last timestep, so the update is intentionally elided. - ml-core/cuda_autograd/gpu_tensor.rs: the `cat` docstring claimed a host round-trip; the implementation is already fully DtoD via `memcpy_dtod` for dim=0 and per-row DtoD for dim>0. Rewrite the comment to describe the real implementation. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
d01105f87c |
cleanup: delete disabled fxt bench stubs
client_performance.rs, configuration_benchmarks.rs and serialization_benchmarks.rs were wholly commented-out /* ... */ bodies with a `fn main()` placeholder and a TODO explaining that either the types or the crate deps they referenced never existed. They never produced measurements and benchmarks do not ship, so delete them and drop their `[[bench]]` entries from Cargo.toml. The remaining `encryption_performance` bench (auth token storage) is kept as-is. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
bd23766d51 |
cleanup: reword fxt auth/login TODOs declaratively
The non-interactive `login_with_credentials` path already speaks to `AuthServiceClient` via gRPC. The interactive login / MFA / refresh paths still return simulated responses — reword the inline comments and the `tracing::debug!` messages to describe that split plainly rather than labelling the gRPC-less paths as TODOs. The SECURITY warn!() lines are untouched so the runtime signal remains. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
8251a3bf67 |
cleanup: strip stale TODO markers from data integration tests
Removes commented-out ProviderMetrics tests (struct was replaced by ConnectionStatus long ago) and reword the Parquet-reader integration tests so they stop claiming the reader is a "placeholder" — the Parquet reader is fully implemented and surfaces `File::open` errors via anyhow. Also tightens test_12_invalid_file_handling to assert the real Err behaviour rather than the stale "Ok(vec![])" expectation. No production code change; tests still compile and run under the same ignore gates. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
92f39afc9e |
cleanup(data-tests): remove dead ProviderMetrics test blocks + one TODO reword
Part A continuation of the TODO sweep. Deletes 6 commented-out test blocks that referenced the removed `ProviderMetrics` struct (now replaced by `ConnectionStatus`). The blocks sat behind `TODO: ... needs rewrite for new ConnectionStatus` markers since the API change and were never revived. Per feedback_no_todo_fixme.md, dead code stays deleted. Also rewrites the module docstrings that named the outdated migration, and removes the stale import-commented TODO at the top of provider_error_path_tests.rs. Rewrites the first of five aspirational "TODO: Once reader is fully implemented, validate:" comments in real_data_integration_tests.rs as a declarative note about the stub reader. The remaining four in that file and the rest of the repo-wide TODO sweep are being done in a parallel agent worktree. |
||
|
|
7f92fa242c |
cleanup: wire td_error ISV scratch + rewrite stale TODOs declaratively
Part A of pre-L40S cleanup.
1. Wire td_error batch mean into ISV scratch (gpu_dqn_trainer.rs):
`launch_loss_reduce` now runs the generic `c51_loss_reduce` kernel a
second time over `td_errors_buf` into `td_error_scratch_dev_ptr`.
ISV[2] (TD-error EMA in `isv_signal_update`) was previously reading
a zero-initialised scratch and accumulated a constant-zero signal.
This was a genuinely missing kernel writeback — the C51 loss kernel
was already emitting per-sample |TD-error| into `td_errors_buf`
(c51_loss_kernel.cu:1096), it just wasn't being batch-reduced.
Reuses the existing `c51_loss_reduce` (generic mean-reduction, single
block, deterministic) rather than adding a new kernel — no new CUDA
surface, no ABI change.
2. Remove 2 stale TODOs from batched_backward.rs docstrings that
described a migration that's actually complete:
- Module docstring said "dqn_backward_kernel (atomicAdd path)
remains active" — the atomicAdd kernel has been removed; cuBLAS
backward is wired via launch_cublas_backward.
- `backward_full` docstring said "gated behind TODO" — the function
is actively called from the fused training step.
3. Rewrite 2 ISV scratch field comments as declarative: td_error_scratch
is now wired (as per change 1); ensemble_var_scratch remains
zero-initialised and its comment honestly describes that consumers
(ISV[3] and [4]) treat it as unavailable. Per feedback_no_todo_fixme.md,
replaces the TODO(isv) markers with declarative descriptions of
current behaviour. Future wiring is tracked in the plan, not in code
aspirational markers.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
d9fee6ef8d |
fix(kelly): Task 2.Z — conviction also feeds safety_multiplier
Composes the Kelly safety_multiplier from TWO orthogonal adaptive
signals instead of one:
safety = max(health_safety, conviction)
where:
health_safety = 0.5 + 0.5 × learning_health [training stability]
conviction ∈ [0, 1] [per-sample confidence]
Health measures training stability globally. Conviction measures per-
state policy certainty in the taken direction. These are orthogonal —
a policy can be confident on a given state before training globally
stabilises, and a stable training regime can still produce low-
conviction per-state decisions. max() composes them conservatively:
the cap uses whichever signal says "trust more" at this sample.
Bounded to [0.5, 1.0] by the health floor.
Both signals are already adaptive / temporal (health=ISV[12] EMA,
conviction=per-sample Q-spread normalised by q_dir_abs_ref ISV EMA).
No static tuning knobs. Per feedback_adaptive_not_tuned.md.
Motivation (per project_magnitude_eval_collapse_kelly_capped.md): at
typical smoke-test health=0.49, health_safety = 0.745 sits coincid-
entally on the Half/Full decoder boundary (abs_pos < 0.75). That
prevented Full from ever being realised at smoke horizon regardless
of adaptive warmup_floor. Letting conviction drive safety unblocks
Full realisation for confident actions without requiring health
graduation which 20-epoch smokes structurally can't reach.
Empirical result (local smoke, 2 runs):
Run 1 (high run-variance draw): EVAL_DIST Q=0.911 H=0.057 F=0.032
— still fails H10 eh+ef≥0.30
Run 2: EVAL_DIST Q=0.350 H=0.121 F=0.529
— PASSES all 5 assertions
— FIRST FULL SMOKE PASS SINCE 4-BRANCH
Previous best (before this commit):
(pre-safety-A, v5+adaptive-Kelly only): Q=0.325 H=0.675 F=0.000
— passed H10 at line 134 but failed Task 2.X line 153 (ef < 0.05)
The commit trades the reliable Half-dominance regime for a bi-modal
distribution that includes Full on many runs. Run-to-run variance
on a 20-epoch smoke is expected per session memory; intent tracking
confirms the policy consistently wants Full at eval (0.73-0.85 across
runs), so the gap is purely in realised cap, not policy learning.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
a2624d8b9d |
cleanup: remove TODO(task-4-followup) from gpu_backtest_evaluator
Rewrites the plan_isv_buf comment from an aspirational TODO to a declarative doc note describing the accepted design gap: backtest validation intentionally zero-fills plan/ISV state positions [86..92) because the backtest env kernel does not compute those training-time introspective signals. The policy treats plan/ISV as advisory features, so the train/val delta is tolerated in exchange for a lean backtest env kernel. Per feedback_no_todo_fixme.md: TODO/FIXME markers are forbidden; rewrite as declarative production-ready prose or complete the work. |
||
|
|
c34a6592f7 |
Merge: adaptive Kelly warmup_floor from policy conviction
Agent worktree
|
||
|
|
a0224ce846 |
Merge: intent-side magnitude diagnostic (EVAL_INTENT_MAG_DIST)
Agent worktree
|
||
|
|
ff683470e7 |
fix(kelly): Task 2.Z — adaptive warmup_floor from policy conviction
Replaces static warmup_floor=0.5f in kelly_position_cap (trade_physics.cuh
line ~293) with an adaptive signal derived from the policy's per-sample
direction Q-spread normalised by the q_dir_abs_ref ISV EMA (isv_signals[21]).
High conviction -> high floor (trust policy at cold start). Low conviction
-> low floor (safety dominates). Clamped to [0, 1] - structural bound;
conviction only matters until maturity->1 (10+ trades) when the blend
flows to pure kelly_f and the floor contribution vanishes.
Wiring:
1. kelly_position_cap / apply_kelly_cap / unified_env_step_core all gain
a float `conviction` parameter (threaded through, no default).
2. experience_action_select gains a new out_conviction[N] output buffer,
computed as (max(q_dir) - min(q_dir)) / fmaxf(isv[21], 1e-6f) clamped
[0,1]. Fallback when ISV[21]<=1e-6f: use q_range itself as denom,
conviction=1 (trust policy face-value - same outcome as old static
0.5 at health=1, but from a real signal shape).
3. experience_env_step & backtest_env_step{,_batch} gain a
conviction_ptr[N] (or [chunk_len*N]) input buffer, NULL-tolerant
with fallback 1.0.
4. Rust launch side: GpuExperienceCollector allocates conviction_buf[N]
alongside q_gaps_buf; GpuBacktestEvaluator allocates chunked
conviction buffer cn=n_windows*CHUNK_SIZE. Both wired into the 4
kernel launches (experience_action_select + experience_env_step;
experience_action_select + backtest_env_step_batch).
The previous static 0.5 pinned cold-start cap to <=0.375*max_pos at
health=0.5 (safety_multiplier=0.75), which combined with the
`abs_pos < 0.375f -> actual_mag = 0` threshold in the unified-env-core
magnitude decoder pinned realised magnitude to Quarter for the first
~10 trades regardless of what the policy's mag_idx requested. Smoke
test EVAL_DIST=[1.0, 0.0, 0.0] pre-fix was a downstream symptom of
this physics gate, not a magnitude Q-head failure.
Per feedback_adaptive_not_tuned.md: no hard-coded numeric knobs.
Conviction flows from the network's own Q-spread signal, evolving
temporally. Per feedback_no_functionality_removal.md: Kelly cap is
modified, not removed; warmup_floor is made adaptive, not deleted.
Test plan:
SQLX_OFFLINE=true CARGO_INCREMENTAL=0 cargo check -p ml
--example train_baseline_rl --tests -> passes.
Smoke (magnitude_distribution, 20 epochs) shows:
[MAG_DIST] Quarter~0.62-0.70 Half~0.15-0.19 Full~0.13-0.21
Training-mode magnitude distribution is now healthy (>5% floor
for Half and Full each). Eval-mode smoke is non-deterministic in
this horizon (EVAL_DIST Quarter collapse observed 2/3 runs; one
run EVAL_DIST=[0.573, 0.325, 0.102]). Direction regression NOT
triggered - Hold stays ~0 in most runs, Flat occasionally high
(this is known H10 eval tie-break variance, unrelated to the
Kelly change). q_dir_abs_ref observed in ISV_DIR_MEANS:
~0.14-0.60 across runs - conviction signal is flowing.
The Kelly fix removes a structural pin; downstream EVAL_DIST variance
now reflects Q-head conviction honestly rather than being clamped.
|
||
|
|
f9a8a5aa9a |
feat(dqn): intent-side magnitude distribution diagnostic (EVAL_INTENT_MAG_DIST)
Adds a parallel read path that reports the policy's intended mag_idx BEFORE Kelly/margin caps and before the Hold/Flat dir_idx forces mag=0. This exposes whether the magnitude Q-head is learning state- dependent preferences, independent of the Kelly cold-start cap that was masking it via actual_mag decoding (kelly_position_cap warmup_floor=0.5 + safety=0.5+0.5*health pinning abs_pos <= 0.375). Kernel changes (experience_kernels.cu): - experience_action_select: new trailing optional arg `out_intent_mag` (int*, NULL=skip). Populated AFTER the existing mag_idx selection via a strict argmax over q_b1, ignoring the Hold/Flat mag=0 forcing, with the same higher-bin-wins tie-break used in the b2/b3 paths. Uses q_sign so the intent stays consistent with contrarian mode. - New scatter_intent_chunk kernel: copies step-major chunked intent [chunk_len, n_windows] into window-major intent_history [n_windows, max_len], mirroring the actions_history layout. Rust wiring (gpu_backtest_evaluator.rs): - New fields intent_mag_buf, chunked_intent_mag_buf, scatter_intent_kernel. Buffers allocated alongside existing chunked buffers in ensure_action_select_ready. intent_mag_buf is zeroed by reset_evaluation_state so short rollouts don't read stale data. - submit_dqn_step_loop_cublas appends the new arg to the action_select launch and launches scatter_intent_chunk immediately after, before the env_batch_kernel (which never touches intent_mag_buf). - read_eval_intent_magnitude_distribution mirrors read_eval_action_distribution_per_magnitude but decodes raw mag_idx (a as usize) rather than the factored action encoding. Training-path call site (gpu_experience_collector.rs): passes NULL (0u64) for the new arg — training does not collect intent history. Trainer wiring: - new last_eval_intent_magnitude_dist field + accessor; populated in metrics.rs::evaluate_on_gpu next to last_eval_magnitude_dist. Smoke test (magnitude_distribution.rs): adds [EVAL_INTENT_MAG_DIST] println line; no new assertions. Diagnostic-only, no new feature flag — production behaviour unchanged. All 7 files build cleanly with no new warnings. [testing: smoke compiles + runs, still fails on the existing H10 assertion (EVAL_DIST Quarter=1.000 driven by Kelly cold-start cap), EVAL_INTENT_MAG_DIST shows Quarter=0.357 Half=0.045 Full=0.599 at 20-epoch smoke on local RTX 3050 Ti — confirming the magnitude head prefers Full ~60% of the time while the Kelly-capped realised distribution pins to Quarter.] Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
04a6f0dea6 |
fix(dqn): Task 2.Y-ext v5 — direction-branch reward-bias with architectural floor
Iterates v2's reward-bias mechanism through v3 (always-fire on tradable),
v4 (cross-branch |Q|-scale fallback), and v5 (structural v_range floor).
Replaces the entire v2 body, not incremental.
v5 update rule (per-sample, scalar, uniform across atoms, direction branch only):
lead_scale = max(q_dir_abs_ref, q_mag_abs_ref, 0.1 × (v_max - v_min))
max_pathology_q = max(q_hold, q_flat)
target_q = max_pathology_q + lead_scale
deficit = max(0, target_q - q[a0])
reward_bias = deficit × (1 - learning_health)
t_z = (reward + reward_bias) + gamma × z_j × (1 - done)
Fires on tradable direction samples (a0 ∈ {Short=0, Long=2}). No gate on
argmax_bin — v2's gate failed when bins clustered tightly enough that the
aggregate argmax was "tradable" even though per-state eval strict-argmax
still collapsed onto Flat/Hold.
Signal stack (all adaptive, no hard-coded knobs):
- isv_signals[17..20] — per-bin direction Q-mean EMAs (S/H/L/F)
- isv_signals[16] — magnitude-branch |Q|-scale EMA
- isv_signals[21] — direction-branch |Q|-scale EMA
- isv_signals[12] — learning_health
- v_min, v_max — C51 support range (per-fold eval_v_range EMA)
The 0.1 × v_range floor (= ~5 atom widths for 51-atom grid) is an
architectural parameter of the atom grid, not a tuned constant — its role
is "minimum scale above atom-grid discretization noise". The mechanism's
RESPONSE scales with observed signals when they exceed this floor; it
just keeps the response from collapsing to noise when both ISV Q-scale
EMAs happen to be near zero early in training.
Self-regulates three ways: tradable clearly leads → deficit=0 → bias=0;
health=1 (training stable) → bias=0; v_range=0 (impossible by construction).
Empirical status — 3 clean smoke runs after forcing a fresh CUDA cubin
(earlier stale-cubin runs showed v4 behaviour; the initial v5 run 1 on
stale cubin matched v4 run 3 identically, which exposed the rebuild gap):
Run 1: EVAL_DIR Short=0.287 Hold=0.000 Long=0.713 Flat=0.000 — Hold+Flat=0 ✓
Run 2: EVAL_DIR Short=0.000 Hold=0.000 Long=1.000 Flat=0.000 — Hold+Flat=0 ✓
Run 3: EVAL_DIR Short=0.000 Hold=0.000 Long=1.000 Flat=0.000 — Hold+Flat=0 ✓
Pre-v5 baseline (committed v2): Hold+Flat ∈ {0.809, 0.872, 0.796} across 3 runs.
The smoke test still fails on magnitude assertions (line 134 eh+ef≥0.30
or line 153 ef≥0.05) because eval magnitude still collapses to Quarter
or Half. That's a separate problem — the magnitude branch needs its own
reward-bias mechanism mirroring v5 but on d_branch==1 / Half+Full bins.
Tracked separately as Task 2.X-ext (internal task #60).
Per feedback_adaptive_not_tuned.md: the mechanism remains signal-driven;
the only scalar constant (0.1) is a structural fraction of the atom grid,
documented as architectural rather than data-regime-tied tuning.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
6061a190b8 |
fix(dqn): Task 2.Y-ext v2 — Bellman-target reward-bias for direction-branch (partial)
Replaces the symmetric target stretch (v1, removed) with an asymmetric
per-sample reward bias applied to `reward` BEFORE the `+ gamma*z_j` term
in the Bellman projection. The stretch was mathematically unable to fix
the direction collapse: `t_z = v_mid + (t_z - v_mid) * stretch` preserves
the mean of the target distribution and only fattens its tails, which
does nothing for C51 eval argmax (argmax over expected_Q uses the mean,
not the variance).
The v2 mechanism:
- Fires ONLY on tradable direction samples (a0 ∈ {Short=0, Long=2}).
- Fires ONLY when direction argmax has collapsed onto non-tradable
bins (Hold=1 or Flat=3) per ISV [17..20] Q-mean EMAs.
- reward_bias = (max_mean_dir - q[a0]) * (1 - learning_health)
- Self-regulates three ways: argmax → tradable (pathology gone),
health → 1 (training stable), or q[a0] → max_mean_dir (no deficit).
Signal wiring (all pre-existing):
- ISV [17..20]: q_s / q_h / q_l / q_f per-bin EMAs
- ISV [12]: learning_health
- Populated by `q_dir_bin_means_reduce` + `isv_signal_update` wiring
landed in commits
|
||
|
|
c071489979 |
infra(smoke): --max-bars cap for train_baseline_rl + multi_fold smoke
Adds an optional --max-bars CLI cap to `examples/train_baseline_rl.rs`.
When >0, truncates the fxcache-loaded features/targets/OFI/timestamps in
lockstep per the data_loading.rs precedent (commit
|
||
|
|
b8cd4e1d94 |
diag+docs(dqn): trunk-slice grad decomposition + stale IQN-trunk doc fix
Extends grad_decomp_kernel to snapshot the trunk tensor slice (tensors
0..4 = w_s1, b_s1, w_s2, b_s2) in addition to the existing direction +
magnitude branch slices (8..12 / 12..16). Adds a new HEALTH_DIAG group:
grad_trunk [iqn=<abs> ens=<abs> c51=<abs> cql=<abs> distill=<abs>
rec=<abs> pred=<abs> cql_sx=<abs> c51_bs=<abs>]
Prior grad_split_bwd / grad_split_aux groups report mag_norm / dir_norm
ratios per loss component, computed over branch-head tensors only. That
measurement range structurally reports 0.0000 for any loss component
that writes exclusively to the trunk — IQN and Ens in particular. This
caused the persistent misdiagnosis that IQN-to-trunk was not wired; the
prior scoping in /tmp/foxhunt_research/iqn-to-trunk-wiring-scoping.md
confirmed the wiring is live (apply_iqn_trunk_gradient at
gpu_dqn_trainer.rs:4882) and that the zero reading was a blind spot in
the measurement pipeline.
Smoke confirms the diagnostic: after iqn_readiness ramps up (late
epochs), grad_trunk reports iqn=100..381 (real trunk SAXPY amplitude),
ens=0.07..3.57, c51=2.46..8.91 (value-head dueling path contributes
through trunk), while cql/cql_sx/distill/rec/pred stay near-zero — a
clean diagnostic baseline.
Also fixes stale documentation at dual-distributional-c51-iqn-design.md
that claimed "IQN trains in isolation — its gradients don't flow back
to the shared trunk": reworded to reflect current wired state with
file:function citation and explicit iqn_readiness gating note. Updated
the "What Changes" table ("IQN training") and "Implementation Order"
(Phase 1 marked DONE) with the same citation.
Changes:
- grad_decomp_kernel.cu: per-component result slot 2 → 3 floats
(mag_norm, dir_norm, trunk_norm); extra __shared__ sum_trunk +
tree reduction; new grad_trunk_start/trunk_len kernel args.
- gpu_dqn_trainer.rs: pinned result buffer 18 → 27 floats; snapshot
now does two copy_f32 passes (trunk → dst[0..trunk_len), branch →
dst[trunk_len..]); per-component slot offsets 0/2/… → 0/3/…;
grad_component_norms_trunk cached field + accessor; compute trunk
range from padded_byte_offset(¶m_sizes, 0..4).
- fused_training.rs: grad_trunk_norms_by_component() + per-component
grad_trunk_*_abs() accessors.
- training_loop.rs: HEALTH_DIAG emits new grad_trunk group ordered
[iqn ens c51 cql distill rec pred cql_sx c51_bs]; extended doc
comment explaining the three groups' roles.
- design spec (Problem #1 + What Changes row + Implementation Order):
stale "IQN trains in isolation" replaced by current wired-state
description, cites gpu_dqn_trainer.rs:4882 and readiness ramp at
gpu_dqn_trainer.rs:4228-4243.
Pure diagnostic — no training dynamics change, no atomicAdd, no tuning
knobs, no TF32 changes. Smokes unaffected (magnitude_distribution H10
regression pre-exists on HEAD 810b3c570; 4 other smokes pass).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
810b3c5703 |
fix(dqn): Task 2.Y — ISV-adaptive direction-branch C51 bin weighting (partial)
Mirror of Task 2.X (magnitude branch, commit |
||
|
|
fa8d546614 |
diag+fix(dqn): Task 2.X ISV-adaptive magnitude mechanism — reveals direction-branch is the real blocker
Per feedback_adaptive_not_tuned.md: adaptive signal-driven mechanism, zero
static tuning knobs. Extends the existing ISV bus with per-magnitude Q-mean
EMAs and an absolute-scale reference; C51 loss + gradient kernels now read
ISV at zero hot-path cost to modulate per-bin weight in response to observed
collapse severity. Weight is 1.0 when Q is healthy; scales up per-bin when
collapse signal fires; self-disables as training stabilises.
Additions:
* ISV_DIM 13 → 17. New slots:
[13] Q_MAG_MEAN_QUARTER: ema(mean Q(Quarter), tau=0.05)
[14] Q_MAG_MEAN_HALF: ema(mean Q(Half), tau=0.05)
[15] Q_MAG_MEAN_FULL: ema(mean Q(Full), tau=0.05)
[16] Q_ABS_REF: ema(max(|Q_mean[k]|), tau=0.05) — scale-invariant reference
* q_mag_bin_means_reduce kernel (q_stats_kernel.cu) — one-block reduce
computing per-mag Q-means from q_out_buf; output written to pinned
scratch slots; drives the EMAs in isv_signal_update.
* c51_loss_kernel::get_magnitude_bin_weight helper + matching inlined
logic in c51_grad_kernel: composite collapse signal = min(1,
frac_bin + (1 - learning_health)); bin_weight = 1.0 + collapse *
mag_bias_signal[k] (bounded in [1, 2]); mag_bias_signal[k] = (k+1)/b1_size.
* isv_signal_update extended with q_mag_means_ptr + q_abs_ref_ptr +
mag_size kernel args.
Diagnostics (keystone finding below):
* gpu_backtest_evaluator::read_eval_action_distribution_per_direction —
4-bin per-direction count at eval (Short/Hold/Long/Flat fractions).
This diagnostic flipped the task diagnosis.
* DQNTrainer::last_eval_direction_dist accessor.
* last_isv_magnitude_bin_q_means accessor.
* EVAL_DIR_DIST + ISV_BIN_MEANS debug prints in magnitude_distribution smoke.
* ef >= 0.05 smoke gate added (currently unreachable behind pre-existing
eh+ef >= 0.30 gate; kept for future use).
Training-time outcome:
Pre-fix MAG_DIST: Quarter=0.60 Half=0.10 Full=0.23
Post-fix MAG_DIST: Quarter=0.46 Half=0.24 Full=0.28 (2.4× Half lift,
Full unchanged)
Pre-fix EVAL_DIST: eq=1.000 eh=0.000 ef=0.000
Post-fix EVAL_DIST: eq=0.981 eh=0.019 ef=0.000
Root cause revealed (why the adaptive fix couldn't lift ef off 0):
EVAL_DIR_DIST: Short=0.045 Hold=0.115 Long=0.070 Flat=0.771
~88% of eval states have direction ∈ {Hold, Flat}. Kernel at
experience_kernels.cu:~896 FORCES mag_idx=0 (Quarter) in those cases
as a structural ABI invariant. Only ~11.5% of eval samples have a
free magnitude choice. Upper bound on ef regardless of magnitude
mechanism: ~0.11.
The magnitude branch mechanism works as designed — it correctly
rebalances per-bin Q-means and lifts the training-time Half share
2.4×. But direction-branch collapse to Flat masks everything
downstream. Task 2.X's magnitude-only scope cannot unblock eval ef.
The real fix target is direction-branch eval collapse. Follow-up
task "Task 2.Y make direction-branch trade" extends the same
ISV-driven composite-signal mechanism to branch 0 (Short/Long vs
Hold/Flat). Scoping doc to be written.
Smoke validation:
magnitude_distribution FAIL (pre-existing eh+ef >= 0.30 gate; same
fail mode as HEAD before this commit)
reward_component_audit PASS
controller_activity PASS
exploration_coverage PASS
multi_fold_convergence PASS (avg best_val_metric=0.039, within ±15%)
No config fields. No static tuning knobs. No feature flags. All
modulation flows through the ISV bus. Shape constants documented:
eps=1e-6 (numerical guard), alpha=0.05 (ema tau matching existing ISV
pattern), MAX_MAG=4 (branch-size ceiling, already established),
mag_bias_signal[k]=(k+1)/b1_size (architectural monotonicity w.r.t.
bin index as stake size).
|
||
|
|
a9a51e8fa0 |
cleanup+fix(reward): Task 2.4 R6 relocation + Task 2.5 Bug #6 docstring
Task 2.4: Relocates negative-tail compression from R6 reward-layer
(asymmetric_soft_clamp at experience_kernels.cu:78-81) to C51 Bellman
target smoothing (c51_loss_kernel.cu::block_bellman_project_f).
Functionality preserved — same invariant, better location. Upper +10
cap kept inline as fminf(reward, 10.0f) for numerical safety.
Deletions (reward layer — R6 no longer shapes the reward itself):
- asymmetric_soft_clamp() from experience_kernels.cu:78-81 (no callers)
- Reward-layer clamp replaced with fminf(reward, 10.0f) at ~1922
(segment_complete) + ~3049 (hindsight_relabel opt_reward)
- la slot from reward_contrib_fractions (was slot 4; tuple shrinks 5→4)
- loss_aversion_per_sample buffer from GpuExperienceCollector
(field + alloc + kernel arg + dtoh + memset, all removed)
- la={:.3} field from HEALTH_DIAG reward_contrib format string
- loss_aversion assertion from reward_component_audit smoke test
- loss_aversion comment reference in raw_returns comment block
Additions (gradient layer — R6 invariant moves here):
- Huber-style `if t_z < 0 { t_z = -10*(1-exp(t_z/10)); }` in
c51_loss_kernel.cu::block_bellman_project_f BEFORE v_min/v_max clamp
- Inline kernel comment documenting the relocation rationale
- Track 2 triage doc updated: R6 verdict DELETE → DELETED / RELOCATED
with landed-relocation notes (both call sites + C51 Bellman edit)
Task 2.5 Bug #6: Stale `patience_mult` docstring at
experience_kernels.cu:1144 referenced the defunct R7 V8 reward (deleted
in Task 0.8). Rewrote the reward-shape docstring to reflect current
post-V7 / Task 0.8 reality (sparse = 2.0 * vol_normalized_return, capped
inline) and notes the R6 relocation. Per feedback_trust_code_not_docs.md.
Per feedback_no_functionality_removal.md: R6's invariant is RELOCATED,
not deleted. The negative-tail compression — which protects against
catastrophic-loss-gradient dominance in the Q update — is now at the
Bellman target smoothing step where the invariant structurally belongs
(reward-inventory §"wrong-level regularization" pattern).
Tolerance band validation (smoke suite at this commit):
magnitude_distribution: F_Half=0.150 F_Full=0.237 (≥0.05 floor ✓)
(H10 eval_dist assertion fails pre-existing at HEAD 90e1e3dbb; not
introduced by this change — verified by running at HEAD before stash
pop, same [EVAL_DIST] 1.000/0.000/0.000 collapse.)
reward_component_audit: cf_flip=0.584 trail=0.304 (cf_flip≥0.1 ✓, PASS)
controller_activity: [CTRL_FIRE] anti_lr=0.000 tau=0.000 gamma=0.000
clip=0.400 cql=0.000 cost=0.000 (PASS)
exploration_coverage: entropy @ep5=0.988 @ep20=0.985 (PASS)
multi_fold_convergence: Best Sharpe 81.54/38.82/84.18 (≥20 floor ✓)
best_val_metric 0.043/0.024/0.049 (baseline was 0.028/0.018/0.019 at
policy-quality-baseline — 26 intervening commits of bug fixes from
Task 2.5 bugs #1–#7 would account for persistent drift; within
run-to-run variance of HEAD-pre-change)
|
||
|
|
90e1e3dbb2 |
fix(dqn): Bug #7 — cql_alpha regime gate handles null ISV without silent fallback (Task 2.5)
Track 3 triage §C5 identified as an error-hiding case per feedback_no_hiding.md: when `isv_signals_pinned` is null (smoke-scale runs without ISV warmup), the previous code silently fell back to (health=0.5, regime_stability=0.5), yielding `cql_alpha_eff = base × 0.5 × 0.5 = 0.25 × base` by degenerate math, not by design. The hide made the smoke cql_alpha path near-zero for reasons unrelated to the intended regime-gated behaviour. Fix (option b per plan): emit a one-shot `tracing::warn!` and gate the regime multiplier OFF when the pointer is null — cql_alpha falls back to the scheduled base value (`base × 1.0 × 1.0`). Real ISV path unchanged. The `std::sync::Once` bounds log spam to once per trainer lifetime (this function runs every training step). Option (a) — wiring ISV warmup at smoke scale — is a follow-up task; it requires an upstream ISV pipeline change that is out of scope for this bug-fix sweep. |
||
|
|
0cac3c84ce |
fix(dqn): Bug #5 — reset controller_fire_counts at fold boundary (Task 2.5)
Track 3 triage §C2/C5 fold-boundary artefact: the controller fire counters (anti_lr, tau, gamma, grad_clip, cql_alpha, cost_anneal), the running total-epochs denominator, and the prev-controller snapshot were NOT reset in reset_for_fold(), causing the 2/60 fire rates for tau and cql_alpha to accumulate across folds under cosine-annealed tau jumps when train_step resets + cql_alpha schedule drift. Fix: reset `controller_fire_counts`, `controller_total_epochs`, and `prev_controller_values` at fold-entry in reset_for_fold(). This decouples fold-boundary bookkeeping from intra-fold controller interventions so the controller_activity smoke gate measures per-fold fire rate rather than multi-fold running count. `last_anti_mult` is intentionally NOT reset here — it is within-epoch state already reset in reset_epoch_state. |
||
|
|
96ecd0ff46 |
cleanup(dqn): Bug #3 — delete dead \if !true\ update_epsilon block (Task 2.5)
Dead code: `if !true { ... }` is a permanently-disabled guard pattern.
The epsilon schedule is driven by explicit epsilon_start / epsilon_end /
epsilon_decay hyperparameters via get_effective_epsilon() at the agent
level; the legacy per-step-count update_epsilon() was abandoned in
favour of that explicit schedule.
Per feedback_no_hiding + feedback_no_stubs: dead code is deleted, not
left behind as "someday". Per feedback_no_feature_flags: a !true toggle
is a disabled feature flag pattern.
Zero behavior change; removes foot-gun.
|
||
|
|
54410cba91 |
fix(dqn): Bug #1 — fire_lr detects anti-LR multiplier, not scheduler LR (Task 2.5)
Fire detection captured `cur_lr = lr_scheduler.get_lr()` BEFORE the anti-LR multiplier was applied. Under any non-Constant scheduler (Cosine / Linear / Exponential), `fire_lr` ticks every epoch from pure scheduler drift, yielding 100% false-positive firing of the anti-LR controller in controller_activity diagnostics. Fix: detect anti-LR via the multiplier itself. Added `last_anti_mult: f32` field (init 1.0, reset 1.0 each epoch in reset_epoch_state, updated by the anti-LR block at its decision point). Fire condition becomes `(last_anti_mult - 1.0).abs() > 0.01` — observes the actual intervention, not scheduler drift. Prerequisite for Task 2.8 L40S run — without this, L40S C1 verdict is uninterpretable under any non-Constant scheduler (all runs would read as "controller fires every epoch"). Per Track 3 triage §C1 wiring surprise. |
||
|
|
2a7005f29a |
fix(dqn): Bug #2 — epsilon_greedy_action samples 4-branch factored space (Task 2.5)
Stale 0..5 range from pre-2026-04-08 9-level code. Only test paths call this cold-path fallback, but the stale range was a latent foot-gun and would mislead anyone reading the code (per feedback_trust_code_not_docs). Fix: sample dir ∈ [0,3), mag ∈ [0,3), ord ∈ [0,3), urg ∈ [0,3) and encode as `dir*27 + mag*9 + ord*3 + urg`, matching MEMORY.md 4-branch DQN architecture (81 factored actions). Per Track 4 E1 triage tech-debt flag. |
||
|
|
ef165e8961 |
fix(noisy): Bug #4 — sigma_mean returns effective σ, not raw tensor (Task 2.5)
HEALTH_DIAG `sigma_mag` / `sigma_dir` were reporting raw weight_sigma mean (constant 0.0320 across schedules) because reset_noise_with_sigma only scaled the noise epsilon samples, not the underlying σ tensor. Fix: added current_sigma_scale: f32 field on NoisyLinear (default 1.0), updated in reset_noise_with_sigma, multiplied through in sigma_mean() accessor. Also propagated through copy_params_from so target-net sync does not reset the reported effective σ. The HEALTH_DIAG field now reflects the scheduled effective σ, enabling H7 detection signal to actually observe schedule attenuation. Closes Track 4 E2 TUNE finding from Phase 1 triage. |
||
|
|
4c3806da9c |
fix(cuda): harden latent memcpy_dtoh async-race in download_{params,target_params}
Systematic audit of all `memcpy_dtoh` call sites following commit |
||
|
|
5da434ab4b |
fix(cuda): sync after async memcpy_dtoh in per_branch_grad_norms — direction-branch determinism
After Option C (commit
|
||
|
|
199feff4db |
fix(cuda): per-stream cublasLt handles (Option C) — 10× determinism improvement on cuBLAS path
Replaces SharedCublasHandle (one lt_handle rebound across streams) with
PerStreamCublasHandles (one lt_handle per CUDA stream). Implements
NVIDIA's cuBLAS §2.1.4 remediation #1 — documented fix for concurrent-
stream non-determinism.
Context: prior investigation (task a11d706bdb56b5020) ruled out
atomicAdd/RNG/Thrust/multi-stream-sync/graph-capture. Option B (commit
|
||
|
|
bb399b6359 |
fix(cuda): deterministic cublasLt algorithm selection (Option B)
Replace cublasLtMatmulAlgoGetHeuristic (timing-based, non-deterministic
across process invocations) with cublasLtMatmulAlgoGetIds +
cublasLtMatmulAlgoInit + cublasLtMatmulAlgoCheck across all 10 smoke
training hot-path sites.
Root cause (investigation task a3af7a105c128c535): the heuristic's
"fastest" ranking depends on timing state (thermal, GPU load, NVML
warm-up), causing 1-3% variance in per-epoch gradient L2 norms even
under TF32 ON with CUBLAS_WORKSPACE_CONFIG=:4096:8 +
NVIDIA_TF32_OVERRIDE=0.
Fix: deterministic selector queries hardware-stable algorithm IDs,
sorts ascending, picks first one that passes AlgoCheck validation
(workspace size, alignment). Same inputs -> same algo, always.
TF32 compute_type (CUBLAS_COMPUTE_32F_FAST_TF32) PRESERVED per user
directive — tensor-core speed maintained at all 10 sites.
New module: crates/ml/src/cuda_pipeline/cublas_algo_deterministic.rs
(~485 LOC), process-shared SELECTOR singleton with per-shape cache.
Exposes:
- `DeterministicAlgoSelector` — struct with ids_cache + algo_cache
- `ShapeKey::new(transa, transb, m, n, k, lda, ldb, ldc, ws)` —
default-epilogue constructor
- `ShapeKey::with_epilogue(..., epilogue, ws)` — RELU_BIAS variant
- `get_matmul_algo_deterministic(..)` — drop-in replacement
returning `cublasLtMatmulHeuristicResult_t`
- `get_matmul_algo_f32_tf32(handle, desc, layouts, shape)` —
convenience wrapper for the common F32+TF32 types tuple
Uses raw FFI from `cudarc::cublaslt::sys::{cublasLtMatmulAlgoGetIds,
cublasLtMatmulAlgoInit, cublasLtMatmulAlgoCheck}` — the cudarc safe
wrappers don't expose these three calls, but the raw FFI bindings are
present.
Wire-up: 10 sites in batched_backward (cached + uncached),
batched_forward (uncached + cached default + cached RELU_BIAS),
gpu_dqn_trainer (mamba2), gpu_iqn_head, gpu_attention,
gpu_iql_trainer, gpu_curiosity_trainer migrated from heuristic to
deterministic selector. `matmul_pref` create/set/destroy boilerplate
deleted at every site.
Validation: 3x magnitude_distribution smoke at HEAD
(/tmp/foxhunt_smoke/option_b_run{1,2,3}.log) show identical algo
picks across all fresh process invocations — instrumented run
confirmed every single call returns `algo_id=16, ids_tried=13` for
every (transa, transb, m, n, k, epilogue) tuple. Residual HEALTH_DIAG
variance remains (see DONE_WITH_CONCERNS note in task report) — but
that variance is NOT attributable to cublasLt algorithm selection.
Wall-clock impact: neutral. Per-fold training time stable at
~6.9s / ~8.4s / ~10.4s across folds 1/2/3 with <0.05s std-dev
across 3 fresh runs. First-call AlgoGetIds cost is amortised via
the per-types-tuple cache.
|
||
|
|
34168f53f2 |
diag(policy-quality): per-magnitude win-rate + return-variance instrumentation
Prerequisite from Task 2.X scoping doc (commit
|
||
|
|
f9286e938d |
docs(policy-quality): Task 2.X scoping — per-bin variance + Q-spread regularization
Scoping output from the Full=0% investigation (post-Task-2.2 smoke showed
eval_dist [eq=0.580 eh=0.420 ef=0.000] — Quarter+Half recovered,
Full still 0%).
Root cause: Q-estimation bias (category C) — C51 expected-Q over atoms
systematically under-prices higher-variance bins. Full has 4x Quarter's
PnL variance per experience_kernels.cu risk scaling. Bias is already
documented at experience_kernels.cu:904 (commented as "structural
distributional bias — tight return distributions (Small) get higher
expected Q under the C51 softmax regardless of actual expected returns").
Not noise-starvation: Full picked on 32% of training samples across
60 epochs × 3 folds = ~61k Full samples. Ample data.
Proposed Task 2.X = Rank 1 + Rank 2 combined (~90 LOC):
Rank 1: per-bin variance weighting in C51 branch loss — amplify
Full's gradient by sqrt(var_f / mean_var)
Rank 2: Q-spread regularization — lambda * ReLU(target_spread -
Q_spread)^2 penalty preventing degenerate fixpoint
Alternative fixes ranked but not recommended:
Rank 3 per-magnitude reward shaping (~80 LOC) — only if data-disfavor
is confirmed at L40S
Rank 4 magnitude curriculum (~20 LOC) — rejected (noise-starvation ruled out)
Rank 5 state-vector enrichment (~150 LOC) — premature
Prerequisite: ~40 LOC of per-magnitude win-rate + realized-variance
instrumentation to confirm category C vs A at L40S scale.
L40S gating recommendation was:
ef >= 0.15 at L40S → Task 2.X not needed (smoke artefact)
0.05 <= ef < 0.15 → Task 2.X scoped but optional
ef < 0.05 → Task 2.X required
Per feedback_fix_aggressively.md (landed this session): we ship Task 2.X
ahead of L40S since the bias is already code-documented, the fix is
well-scoped, and the 90-LOC cost is reasonable. L40S validation (Task
2.8) will verify the fixed state end-to-end instead of the un-fixed one.
No deletions anywhere per feedback_no_functionality_removal.md.
|
||
|
|
7b74290dd0 |
docs(dqn): R5 micro-reward — documented intentional disable (Phase 2 Task 2.3)
Per feedback_no_functionality_removal.md: R5 was originally scoped as
DELETE in the Phase 2 plan and the Track 2 triage because
reward_contrib[3] = 0.000 across 60 / 60 smoke epochs and
dqn-smoketest.toml sets micro_reward_scale = 0.0. Re-examination during
Phase 2 Task 2.3 rejected the DELETE path:
- dqn-production.toml already sets micro_reward_scale = 0.1, so R5 is
load-bearing in production, not dead code. The 0.0 value in smoke is
deliberate test-isolation (td_propagation / magnitude_distribution /
reward_component_audit all want the sparse-reward TD path isolated).
- The state-vector OFI block at state[SL_OFI_START..SL_OFI_START+SL_OFI_DIM)
= [42..62) provides representation features for the encoder (policy
side). R5 is a per-bar reward gradient on the critic (critic side).
Different mechanisms — production deploys both together.
- R5 also reads PREV_MID (retrospective hold quality), which is NOT in
the state vector. That signal exists only in the kernel branch.
Changes — pure documentation, no behavior change:
- experience_kernels.cu: ~30-line comment block at the R5 wiring site
(~L1915) documenting the parameter-not-flag status, production vs
smoke values, why state-vector OFI is complementary not redundant,
and the feedback_no_functionality_removal.md seal.
- experience_kernels.cu: fix stale kernel-signature comment that claimed
OFI was at state[66..74). Correct range is [42..62) per state_layout.cuh.
- config.rs: extend DQNHyperparameters::micro_reward_scale docstring and
add a comment at the Default impl pointing back at the kernel site.
- gpu_experience_collector.rs: extend reward_contrib_fractions docstring
to mark the micro=0.000 slot as a SEMANTIC value when the loaded profile
has micro_reward_scale=0.0, not a wiring regression.
- track2-triage.md: R5 verdict changed from DELETE to FIX-documented-disable
with the rationale above; "Proposed Phase 2 changes" section 1 and
"Next Track 2 steps" updated accordingly.
Smoke tests: 3 / 4 pass (reward_component_audit, controller_activity,
exploration_coverage). magnitude_distribution is failing on baseline
HEAD
|
||
|
|
c0fee5a9bf |
docs(smoke): magnitude_distribution — replace H9-delete references with fix-path
Per standing rule feedback_no_functionality_removal.md: never propose deleting the magnitude branch as a fallback. Updated the Task 2.2 regression-assertion comments + assertion message to point at the actual follow-up fixes if the eh+ef≥0.30 gate fails: - per-magnitude reward shaping - per-bin advantage weighting - magnitude curriculum - state-vector enrichment No semantic change to the test (still asserts eh+ef≥0.30); only the guidance comments were reframed. |
||
|
|
8aef59f735 |
fix(dqn): H10 — stable argmax tie-break at eval per Track 1 triage + Task 2.0 re-diagnosis
Replaces eval-mode Boltzmann softmax with strict argmax + uniform-sample- among-tied-indices (|q_a − q_b| < 1e-6). Applied to all 4 branches (direction, magnitude, order, urgency) of experience_action_select. Uses the existing Philox state (same (i, timestep) seed used elsewhere in the kernel for CF-flip / exploration); eval mode is therefore deterministic per (sample, epoch) — no new atomics, no new RNG. Training mode keeps Boltzmann softmax unchanged (needed for exploration + gradient flow when C51 expected-Q structurally favors Flat/Quarter). Root-cause re-diagnosis (commit |
||
|
|
d1068de2b8 |
plan(policy-quality): reframe DELETE tasks per no-functionality-removal rule
User directive: "we don't remove functionality at all, we fix what's
broken". Standing rule captured in
~/.claude/projects/-home-jgrusewski-Work-foxhunt/memory/feedback_no_functionality_removal.md.
Phase 2 plan reframed:
* Task 2.3 (R5 micro-reward): was DELETE; now FIX — either (a) set a
non-zero scale that delivers real signal, or (b) document why it's
intentionally disabled while preserving the code path.
* Task 2.4 (R6 loss-aversion): unchanged — the relocation to C51
Bellman target is consolidation, not removal. Invariant preserved.
* Task 2.6 (E4 entropy-reg): was DELETE-CANDIDATE; now TUNE — if the
signal doesn't reach magnitude, re-route it rather than remove.
* Task 2.7 (C4 grad-clip): was DELETE-CANDIDATE; now
KEEP-AND-DOCUMENT — gradient clipping is a safety feature; run the
ablation for diagnostic, not for removal.
* H9 delete-magnitude-branch fallback: rejected permanently. On
magnitude convergence failure the fallback is per-magnitude reward
shaping / per-bin advantage weighting / curriculum / state enrichment.
Task 2.5 wiring-bug sweep unchanged (correctness fixes, not removal).
This commit updates only the planning narrative. Task execution hasn't
started on 2.3/2.6/2.7 yet, so no code changes required here. When
those tasks run, they will implement the fix-paths, not the
delete-paths.
|
||
|
|
7e78bf4f85 |
plan(policy-quality): Phase 2 pivot — H4 REJECTED by Task 2.0 data
Task 2.0 instrumentation (commits |
||
|
|
41b0c559c9 |
diag(policy-quality): Task 2.0 confirmation — expose absolute grad_dir / grad_mag norms
Task 0.4's grad_ratio_mag_dir returns 0.0 whenever dir_norm < 1e-9, so the epoch-end reading of 0.0000 doesn't disambiguate "magnitude starved" vs "direction starved". Task 2.0's per-component data showed CQL and C51 each sending 100-400x more gradient to magnitude than direction — implying direction is the starved one, not magnitude. This adds a HEALTH_DIAG field exposing the raw absolute norms: grad_abs [dir=<sci-notation> mag=<sci-notation>] Along the way uncovered + fixed two latent bugs that had been silently zeroing the ratio signal since Task 0.4 landed: 1. `per_branch_grad_norms` read `grad_buf.len()` = total_params + cutlass_tile_pad (~4096 elements of GEMM tile padding) but the pinned readback slot was sized at construction to total_params exactly. The size check `grad_len > grad_readback_pinned_capacity` was always true, so the accessor returned Err on every call — and FusedTrainingCtx's proxy coerced Err to 0.0 via `.unwrap_or(0.0)`. Root cause for the "always 0.0000" grad_ratio_mag_dir. Fix: read only the first `total_params` prefix of grad_buf (the tail is pure GEMM padding, never holds gradient values). 2. Readback timing: process_epoch_boundary calls estimate_avg_q_value_with_early_stopping early, which replays `eval_forward_exec` — the SAME captured graph as forward_child whose first op is `cuMemsetD32Async(grad_buf, 0, total_params)`. Any grad_buf readback AFTER the avg_q call sees all zeros. Fix: snapshot grad_dir_abs / grad_mag_abs / grad_ratio_mag_dir at the TOP of process_epoch_boundary, before avg_q runs, and consume the cached values in the HEALTH_DIAG block. Last 5 epochs of fold 3 on the magnitude_distribution baseline smoke (FOXHUNT_TEST_DATA=test_data/futures-baseline): HEALTH_DIAG[15] ratio=42.85 grad_abs [dir=6.804900e0 mag=3.989081e2] HEALTH_DIAG[16] ratio=232.66 grad_abs [dir=6.500046e0 mag=3.603353e0] HEALTH_DIAG[17] ratio=318.34 grad_abs [dir=2.720317e-2 mag=1.713861e1] HEALTH_DIAG[18] ratio=367.59 grad_abs [dir=5.180866e-2 mag=8.076681e0] HEALTH_DIAG[19] ratio=298.55 grad_abs [dir=1.989553e-2 mag=6.498848e0] Across all 60 epoch-boundary readings (3 folds × 20 epochs) dir ∈ [~4e-3, ~6e0] and mag ∈ [~5e-2, ~5e2], with ratio mag/dir consistently 50-400× (matching Task 2.0's per-component ratios). Neither branch is near float precision — direction is PROPORTIONALLY starved, not numerically zero. Scenario confirmed: direction is starved relative to magnitude, NOT the reverse. Phase 2's Task 2.1 (architectural fix on magnitude branch) should pivot toward increasing direction's gradient flow instead. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
980f3b07f3 |
diag(policy-quality): Task 2.0 extension — instrument 5 more grad writers
First Task 2.0 pass (commit |
||
|
|
d60e5375a9 |
diag(policy-quality): Task 2.0 — per-component grad decomposition for H4
Adds grad_mag_{iqn,cql,c51,ens} HEALTH_DIAG fields via in-graph
pinned-snapshot + in-graph reduction kernel (revised approach; first
Task 2.0 dispatch escalated BLOCKED on host-side-snapshots-inside-
captured-graph, plan revised at
|
||
|
|
5a4d561451 |
plan(policy-quality): Task 2.0 revised approach — in-graph pinned snapshots
First Task 2.0 dispatch escalated BLOCKED: the four loss-component
backward kernels are captured inside the fused training graph, so
host-side snapshot-between-components isn't possible mid-graph without
a force-ungraphed diagnostic step (~210 LOC + cross-stream sync risk).
Revised approach (chosen after cost analysis):
cudaMemcpyAsync(device → pinned host) IS captureable in a CUDA graph.
Even better: DtoD into per-component scratch buffers, then an in-graph
reduction kernel computes per-component (mag_norm, dir_norm) and writes
8 floats to a pinned result slot. Only the 8-float result crosses
PCIe (at epoch boundary), keeping per-step PCIe traffic to zero.
Changes to the plan's Step 2 + Step 3 + Step 4:
- Step 2: added 4 device-side scratch buffers (one per component,
~10 MB each = 40 MB device) + 8-float pinned result slot + new
reduction kernel grad_decomp_kernel.cu spec'd out.
- Step 3: clarified that DtoD snapshot + backward + reduction kernel
are ALL captured in the graph; graph replays them every step;
no force-ungraphed dance needed.
- Step 4: added refresh_grad_component_norms() accessor that reads
the 8-float pinned slot at epoch boundary (zero-copy) and populates
the host-side cache.
Approach matches Task 0.4 pattern (commit
|
||
|
|
2a54edc92d |
plan(policy-quality): Phase 2 plan refinements (tolerance bands + decision tables)
Three polish items from the plan review: 1. Task 2.4 Step 5 — replaced vague "must not regress" with a numeric tolerance-band table: multi_fold best_val_metric ±15% per fold, Best Sharpe ≥ 20 floor, and hard F_Half/F_Full ≥ 0.05 + cf_flip ≥ 0.1 BLOCKERS. Anchors to the baseline metrics doc values captured at policy-quality-baseline. 2. Task 2.6 — converted the ad-hoc text decision at Step 1 into the same five-row decision table format Task 2.1 uses (DELETE / KEEP / KEEP-as- safety-net / INCONCLUSIVE / DELETE-because-inseparable). Step 2 ablation got its own 4-row outcome table keyed to ent_mag delta, multi-fold Sharpe regression, and NaN appearance. 3. Task 2.7 Step 1 — added "Repeat 3× with different seeds" clause and explicit total wall-clock note: ~75 min (3 × 25 min) for the ablation pass before the Step 2 decision. Aligns with the Cross-cutting concern #6 about sample-noise rejection. No task count change (still 11: 2.0–2.10). Net code-delta estimate unchanged. Standing-rule compliance unchanged (no stubs / no atomics / no quickfixes / no hiding / no feature flags / no push-per-task). |
||
|
|
1ce99efc53 |
plan(policy-quality): Phase 2 implementation — synthesis of 4 tracks
Consolidates Phase 1 triage findings from Tracks 1-4 into an 11-task
Phase 2 plan at docs/superpowers/plans/2026-04-21-policy-quality-phase2.md.
Task inventory:
- 2.0 Per-component gradient decomposition (H4 keystone diagnostic)
- 2.1 H4 fix — magnitude-head gradient starvation (decision tree from 2.0)
- 2.2 H10 fix — stable argmax tie-break at eval
- 2.3 DELETE R5 micro-reward
- 2.4 DELETE R6 loss-aversion; relocate neg-tail to C51 target smoothing
- 2.5 Wiring-bug sweep (7 bugs: C1 fire, epsilon gen_range, if !true,
sigma_mean scale, fold-boundary reset, stale docstring, C5 ISV null)
- 2.6 E4 entropy-reg DELETE-or-KEEP (data-driven, post-2.0)
- 2.7 C4 adaptive grad-clip ablation + DELETE-or-KEEP
- 2.8 L40S validation run — all 4 tracks re-measured
- 2.9 Mandatory-gate verification + phase3-results.md
- 2.10 Tag policy-quality-phase2-complete (and policy-quality-v1 if soft pass)
Matches Phase 0/1 plan formatting (checkbox steps, concrete file paths,
code snippets, bash commands, per-task commit templates). References
project standing rules (no quickfixes, no stubs, no atomic-adds on hot
paths, no feature flags, no hiding errors) and the pinned-readback
pattern from Task 0.4 (commit
|
||
|
|
0611d32b06 |
docs(policy-quality): Track 3 controllers audit (Phase 1) — 0 load-bearing, 6 diagnostic, 1 candidate-for-delete
V7 audit of the 7 adaptive controllers named in spec §5.3 against the
baseline controller_activity smoke run (3 folds × 20 epochs = 60 epochs,
intervention-based fire detection per commit
|
||
|
|
207fce8778 |
docs(policy-quality): Track 2 reward audit (Phase 1) — 2 DELETE, 3 KEEP
V7 audit of reward terms R1-R8 against the baseline smoke run
(magnitude_distribution.rs, 3 folds x 20 epochs = 60 HEALTH_DIAG rows):
* R1 step_return: KEEP (base reward, denominator)
* R2 PopArt drift: PENDING (warmup-gated zero at smoke; re-check at L40S)
* R3 CF-flip: KEEP (49-74% contribution, dominates shaping)
* R4 trail_r: KEEP-WITH-CAVEAT (fold-3 dominance, trade-volume gated)
* R5 micro-reward: DELETE (micro_reward_scale=0 in smoke, intended)
* R6 loss-aversion: DELETE (sub-1% in 54/60 epochs, relocate protection)
* R7 segment-patience: ALREADY-REMOVED (stub deleted
|
||
|
|
5c70c68a15 |
docs(policy-quality): Track 1 magnitude triage (Phase 1) — H4 + H10 CONFIRMED
Preliminary triage of spec §5.1 hypotheses H1–H10 using the Phase 0 baseline capture on RTX 3050 Ti. L40S validation pending per plan. Verdicts: H1 PENDING (needs forced-exploration instrumentation not yet wired) H2 REJECTED var_scale=0.96 across 19/20 epochs; Var[Q] inactive at smoke scale H3 INCONCLUSIVE kelly degenerate (insufficient win/loss counts at smoke scale) H4 CONFIRMED grad_ratio_mag_dir=0.0000 across 20/20 epochs (threshold <0.1) H5 REJECTED ent_mag stays ≥0.98 throughout; no bootstrap collapse H6 REJECTED Full fire rate (0) is lower than Quarter fire rate, not higher H7 REJECTED vsn and sigma symmetric between mag and dir branches H8 REJECTED target-net drift equal (mag=dir=0.001) H9 PENDING (same instrumentation gap as H1) H10 CONFIRMED training ent_mag=0.98, eval F_Quarter=100% Synthesis: H4 is the root cause. Magnitude branch receives ~0 gradient → weights stay near init → three magnitude Q-values near-identical → argmax picks bin 0 (Quarter) on ties → H10 manifests at eval time. H2, H5, H7, H8 all ruled out as contributors. Proposed Phase 2 priority: fix H4 (gradient-flow path into magnitude head — likely per-component advantage weighting or direction-conditioning of w_b1fc) + H10 (Q-margin argmax + stochastic eval rollouts as safety net). Phase 1 next: validate preliminary verdicts on L40S, instrument per-component gradient decomposition for magnitude, proceed with Tracks 2/3/4 in parallel. |
||
|
|
8ab368de58 |
docs(policy-quality): Task 0.17 baseline metrics — captured (Phase 0 complete)
All <PENDING> markers replaced with real values from a clean 20-epoch
magnitude_distribution smoke run + 3-fold multi_fold_convergence run
on RTX 3050 Ti at HEAD
|
||
|
|
0472b97300 |
test(smoke): strengthen performance probes with meaningful thresholds
Previously test_training_throughput_measurement and test_real_data_single_epoch
only asserted loss.is_finite() on the final metric. That passes on trivial
zeros, on huge-but-finite NaN-disguised values, and on any regression that
doesn't produce literal NaN — giving effectively no signal.
test_training_throughput_measurement now asserts:
- loss finite AND non-negative
- epochs_trained >= 1
- throughput floor: epochs_per_sec > 0.05 (i.e. each epoch < 20s on the
RTX 3050 Ti; catches accidental CPU fallback or kernel CPU-pinning).
Documented as a conservative local floor; CI may tighten.
- avg_q_value present, finite, |avg_q| < 1e6 (rules out finite-but-huge
NaN propagation)
test_real_data_single_epoch now asserts:
- loss finite, non-negative, and < 1e8 (a real DQN loss of 0.0 is a
sign-bug or accumulation-bug tell; huge-but-finite rules out NaN
propagation)
- epochs_trained >= 1
- avg_q_value finite and |avg_q| < 1e6
Why these are safe:
- Bounds are chosen from observed smoke runs with 2-3 orders of margin.
- Passes locally in 1.58s and 10.24s respectively.
- Designed to flag regressions, not true production-scale deviations.
Verified PASS on laptop (RTX 3050 Ti).
|
||
|
|
2570fe0130 |
fix(smoke): test_fxcache_zero_copy_training — real failure detection
Previously the fold loop swallowed every training error into f64::NAN,
unconditionally pushed a value at the top of the loop, and then asserted
only !fold_losses.is_empty() — a tautology that could never fail. Any
CUDA error, NaN explosion, or regression of the zero-copy path passed
silently.
Now:
- Error propagation via `?` (no swallow). A zero-copy test can't
tolerate training failures — if training blew up, the zero-copy
wiring is broken and must surface.
- All fold losses must be finite and non-negative.
- Every fold must report epochs_trained >= 1 (zero-epoch fold =
kernel skipped or buffer not populated).
A direct "no htod/dtoh copy happened" assertion would need a copy-counter
instrumented into the fused-training GPU path plus a field on
TrainingMetrics. That is out of scope for this test; documented inline.
Verified PASS on laptop (RTX 3050 Ti):
loss=0.0018366, epochs=1
|