Commit Graph

2782 Commits

Author SHA1 Message Date
jgrusewski
a400288ae0 feat: comprehensive TOML profiles + config consistency test
Task 7: Update all three DQN TOML profiles with every configurable parameter.
- dqn-smoketest: add [distributional], [advanced], [exploration] (noisy_sigma,
  entropy_coefficient, count_bonus), [risk] (max_position, loss_aversion),
  fix learning_rate 0.0003 -> 0.00003, add reward_scale + gamma
- dqn-production: add reward_scale, exploration params (noisy/entropy/count_bonus),
  advanced params (n_steps/tau/c51_warmup/her/iqn_lambda/spectral_norm),
  remove hardcoded v_min/v_max (now computed), remove duplicate tau from [training]
- dqn-hyperopt: add search space bounds for spectral_norm_sigma_max,
  c51_warmup_epochs, her_ratio

Profile system additions (training_profile.rs):
- TrainingSection: add reward_scale (recomputes v_min/v_max on apply)
- ExplorationSection: add noisy_sigma_init, entropy_coefficient,
  count_bonus_coefficient, q_gap_threshold
- AdvancedSection: add n_steps, tau, c51_warmup_epochs, her_ratio,
  iqn_lambda, spectral_norm_sigma_max, gradient_clip_norm
- RiskSection: add max_position (alias for max_position_absolute), loss_aversion
- RewardSection: add reward_scale
- SearchSpaceSection: add spectral_norm_sigma_max, c51_warmup_epochs, her_ratio
- apply_to: gamma change now recomputes v_min/v_max automatically

Task 8: Config consistency integration tests (5 tests):
- test_config_consistency_across_structs: v_range computed not hardcoded,
  fill simulation bounds, q_clip symmetry
- test_toml_profile_applies_all_fields: smoketest profile applies every
  new section field correctly
- test_production_profile_applies_all_sections: production profile end-to-end
- test_hyperopt_search_space_has_new_bounds: new search space fields parse
- test_reward_scale_recomputes_v_range: gamma override triggers v_range recomputation

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-25 10:27:45 +01:00
jgrusewski
8777288880 feat: v_range computed from reward_scale + gamma — config flows HP → DQNConfig → GpuConfig
Task 4: Add `reward_scale` field (default 10.0) to DQNHyperparameters with
`computed_v_min()`/`computed_v_max()` methods. Formula:
v_range = (reward_scale / (1 - gamma) * 1.2).clamp(20, 300).
conservative() now computes v_min/v_max = +-240 (was hardcoded +-50).
Hyperopt adapter uses same formula instead of hardcoded max_abs_reward.

Task 5: Verified DQNConfig receives v_min/v_max from DQNHyperparameters
in constructor.rs (lines 285-286). Chain intact.

Task 6: Verified GpuDqnTrainConfig receives v_min/v_max from DQNConfig
in fused_training.rs (lines 169-170). Chain intact.

All Default impls updated: DQNConfig, GpuDqnTrainConfig,
ExperienceCollectorConfig, DqnBacktestConfig — zero hardcoded v_min/v_max.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-25 10:16:26 +01:00
jgrusewski
e8f54c37c1 feat: expose experience collector params in TOML training profiles
Add 7 new fields to DQNHyperparameters (fill_ioc_fill_prob,
fill_limit_fill_min, fill_limit_fill_max, fill_spread_cost_frac,
fill_spread_capture_frac, q_clip_min, q_clip_max) and wire them
through training_profile.rs into ExperienceCollectorConfig construction
in training_loop.rs. Previously these 7 values were hardcoded at the
construction site; now they flow from TOML [experience.fill_simulation]
and [risk] sections. Default values match the prior hardcoded constants
so existing behavior is unchanged.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-25 10:05:01 +01:00
jgrusewski
1a51e67ba0 fix: unify Default impls — all config structs match conservative() values
DQNConfig::default(): learning_rate 1e-4→3e-5, gamma 0.99→0.95,
batch_size 64→1024, v_min/v_max -25/25→-50/50, q_clip -100/100→-200/200.

DQNConfig::conservative(): learning_rate 1e-4→3e-5, gamma 0.99→0.95,
batch_size 32→1024, v_min/v_max -25/25→-50/50, q_clip -500/500→-200/200.

DQNConfig::emergency_safe_defaults(): v_min/v_max -25/25→-50/50,
q_clip -500/500→-200/200.

GpuDqnTrainConfig::default(): state_dim 72→48, v_min/v_max -2/2→-50/50,
lr 3e-4→3e-5, weight_decay 1e-5→1e-4, batch_size 256→64.

ExperienceCollectorConfig::default(): use_noisy_nets false→true,
use_distributional false→true, num_atoms 1→51,
v_min/v_max -2/2→-50/50, q_clip -500/500→-200/200.

DqnBacktestConfig::from_network_dims(): v_min/v_max -2/2→-50/50.
GpuExperienceCollector constructor: v_min/v_max -2/2→-50/50.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-25 09:57:04 +01:00
jgrusewski
04d8802c94 refactor: remove 8 always-on use_ booleans — features are mandatory
Remove use_double_dqn, use_dueling, use_per, use_branching,
use_distributional, use_noisy_nets, use_huber_loss, and use_cql
from DQNConfig, DQNHyperparameters, and DqnParams structs.

These features are always enabled (Rainbow DQN standard). The boolean
flags were dead code — every constructor set them to true, and the
only code paths that set them to false were in tests that disabled
features for simplicity. With the fields removed, the features are
unconditionally active, eliminating ~490 lines of dead configuration.

Key changes:
- Struct field declarations removed from 3 core config structs
- Conditional branches (if use_X { ... } else { ... }) simplified:
  dueling/branching/PER network creation is now unconditional
- Checkpoint metadata hardcodes "true" for backward compatibility
- Hyperopt search space index 11 (use_branching) fixed at 1.0
- TOML/YAML config files cleaned of removed fields
- Tests that toggled these flags updated or rewritten

45 files changed, -487 net lines. Zero new test failures.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-25 09:45:54 +01:00
jgrusewski
70a9499f33 docs: config unification plan — single source of truth, 8 tasks 2026-03-25 09:01:12 +01:00
jgrusewski
eceaf45b11 feat: --full mode for 3-phase hyperopt pipeline (BC → RL → Refinement)
- CampaignMode enum: Quick, Standard, Full
- dqn_full() constructor: 20 trials × 100 epochs (covers all 3 phases)
- fxt tune start --full flag: auto-sets 20 trials, 100 epochs
- Phase 1 (BC): MSE warmup + expert demos + DT pretrain
- Phase 2 (RL): all 25 features, C51 ramp, HER
- Phase 3 (Refine): pure C51, shrink-and-perturb

Usage: fxt tune start --model dqn --full --gpu

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-25 08:44:41 +01:00
jgrusewski
704ee72412 Revert "fix: stop clearing feature cache on every Argo run — auto-invalidates via content hash"
This reverts commit eab24bd288.
2026-03-25 08:40:10 +01:00
jgrusewski
eab24bd288 fix: stop clearing feature cache on every Argo run — auto-invalidates via content hash 2026-03-25 08:37:59 +01:00
jgrusewski
5e7cb5d9ff fix: backtest metrics — multiplicative compounding, chunked drawdown, per-trade win rate
6 bugs in backtest_metrics_kernel.cu:
1. Additive return accumulation → multiplicative (equity *= 1+r)
2. Absolute drawdown → fractional ((peak-current)/peak)
3. Total return from additive sum → from compounded equity
4. Strided bar processing → consecutive chunks (correct drawdown)
5. (Sharpe unchanged — arithmetic mean is correct)
6. Per-bar win count → per-trade win tracking

Before: Return=17975%, MaxDD=100% (impossible). After: honest metrics.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-25 08:09:56 +01:00
jgrusewski
480b07864e fix: backtest metrics — multiplicative compounding, chunked drawdown, per-trade win rate
Five bugs in the GPU backtest metrics kernel produced impossible results
(17,975% return, 100% MaxDD on every trial):

1. Additive return accumulation (local_cum += r) replaced with
   multiplicative compounding (local_cum *= 1+r, init 1.0)
2. Absolute drawdown (peak - current) replaced with fractional
   drawdown ((peak - current) / peak)
3. Strided bar processing (thread sees every Nth bar) replaced with
   consecutive chunked processing for correct drawdown tracking
4. Per-bar win counting replaced with per-trade win/loss tracking
5. Total return now computed via multiplicative equity product
   reduction across threads (s_sorted[0] - 1.0)

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-25 08:07:56 +01:00
jgrusewski
07545a3df3 docs: reward v6 fixes plan 2026-03-25 01:49:53 +01:00
jgrusewski
7ed4e5ca90 fix: reward v6 — ATR vol proxy, tanh squash, loss aversion ordering, remove double penalty
Five reward computation fixes in experience_env_step CUDA kernel:

1. Replace CUSUM vol proxy with ATR(14): CUSUM at feature[41] is a binary
   direction indicator [-1,1,0], NOT volatility. When CUSUM≈0, vol_proxy
   became 0.0001 causing 10000x reward amplification. ATR(14) at feature[9]
   is actual realized volatility — reverse the safe_normalize encoding
   (ln(atr)+7)/16 to recover atr_pct = exp(norm*16-7) / price.

2. Move loss aversion BEFORE squash: previously applied after hard clamp,
   creating asymmetric [-15, +10] range making expected reward negative
   even for fair strategies. Now applied pre-squash for smooth asymmetry.

3. Replace hard clamp with tanh soft squash: fmaxf(-10, fminf(10, reward))
   destroyed tail information (1% and 5% wins both → 10.0). tanh preserves
   that larger wins produce proportionally larger rewards.

4. Remove turnover penalty: the 0.05*|delta|/max_position penalty double-
   counted transaction costs already deducted from cash via Almgren-Chriss
   impact model at line ~679, over-penalizing necessary rebalancing.

5. Clarify CUSUM spread_scale usage: CUSUM at feature[41] is correctly used
   as market-stress proxy for spread widening in tx cost computation — this
   is distinct from the (now-fixed) vol proxy for reward normalization.

Also: annotate min_hold_bars=5 as hyperopt candidate.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-25 01:42:41 +01:00
jgrusewski
852ee87c38 feat: reward v6 — sparse trade-completion only (ETDQN validated)
Eliminates the dense per-bar reward component entirely. Per-bar ES returns
have SNR of 0.001 — mathematically unlearnable. The model trained on noise.

Reward v6 design (Takara et al. 2023, ETDQN):
- During trade: reward = 0.0 (ZERO — no noise)
- At trade exit: reward = 10.0 × vol_normalized(trade_return)
- Turnover penalty: -0.05 × |delta_position| / max_position
- Loss aversion: 1.5× on negative rewards

Vol normalization (Zhang 2020): CUSUM proxy for realized volatility.
Makes rewards comparable across trending vs ranging regimes.

Results (50-epoch local smoke test):
- v5: Sharpe -0.82 to -0.37, Return -25% to -10%, 0 profitable epochs
- v6: Sharpe -0.37 to +0.25, Return -9% to +6.6%, 7 profitable epochs
- PF >1.0 in 7 epochs (was 0). MaxDD 6-17% (was 14-32%).

Also: C51 v_range widened to ±50 (default), hyperopt computes ±(10/(1-γ)×1.2).

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-25 01:22:53 +01:00
jgrusewski
d7fcdae711 feat: raw portfolio returns buffer for accurate Sharpe/MaxDD + multiplicative equity curve
- Added raw_returns_out GPU buffer alongside rewards_out in experience kernel
- Portfolio return = (equity_t - equity_{t-1}) / equity_{t-1} per bar (no shaping)
- collect_trade_stats() downloads raw returns (not RL rewards) for financials
- MaxDD now uses multiplicative compounding: equity *= (1 + r_t)
- total_return computed from compounded equity curve

Before: MaxDD 94-2213% (using shaped rewards). After: MaxDD 17-32% (honest).
The model shows -15% return per epoch with PF 0.7-0.9 — no edge yet, needs H100 hyperopt.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-25 01:05:03 +01:00
jgrusewski
c0352bb051 fix: DBN data loader produces proper 4-element targets [preproc, preproc, raw, raw]
The DBN loader was producing 2-element targets [close, next_close] which got
zero-padded to 4 elements. Now produces full 4-element layout matching the
kernel's expected format: [0:1]=network input, [2:3]=raw prices for portfolio sim.

Kernel reads tgt[2:3] for raw_close/raw_next (restored to original design).
All other kernels (DT, PPO, expert demos) also read tgt[2:3] correctly.

Stale feature cache invalidated by this data format change.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-25 00:41:09 +01:00
jgrusewski
33f840d588 fix: target index mismatch — kernel read tgt[2:3] but DBN loader puts prices at tgt[0:1]
ROOT CAUSE of 0% win rate: The experience kernel read raw_close from tgt[2] and
raw_next from tgt[3], but the DBN data loader produces 2-element target vectors
[current_close, next_close] which get zero-padded to 4 elements. So tgt[2:3]=0,
triggering the degenerate price guard (raw_close=1.0, raw_next=1.0), making
every trade's P&L exactly zero → classified as loss.

Fix: read tgt[0] and tgt[1] which contain the actual raw prices.

Result: wins=264/594 (44.4% win rate), PF improving 0.44→0.88→0.98 over 3 epochs.

Also: segment-based trade detection correctly counts reversals,
old_pos_pnl uses saved pre-update position for correct P&L computation.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-25 00:37:32 +01:00
jgrusewski
a69174e99f fix: segment-based trade lifecycle + reversal P&L computation
- Trade detection now counts reversals (S100→L50) as completed segments
- old_pos_pnl saved before position update for correct reversal P&L
- realized_pnl writeback uses old position PnL on reversal bars
- 0% win rate persists — needs deeper investigation (likely tx cost interaction)

WIP: The trade_return formula produces correct sign for raw market moves,
but every trade still shows as a loss. Suspect tx costs on both entry AND
exit of each reversal segment exceed the 1-bar price movement.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-24 23:57:04 +01:00
jgrusewski
70e38508fc feat: three-phase pipeline + hyperopt search space expansion
Task 9: Document ensemble consensus limitation — ensemble heads live on
FusedTrainingCtx (training-only), not accessible at inference time.
The ensemble already provides value via diversity gradient during training.

Task 10: Add three-phase training pipeline to the epoch loop:
- Phase 1 (Behavioral Cloning): C51 warmup + expert demos (already existed)
- Phase 2 (Full-Stack RL): all features, expert decay (already existed)
- Phase 3 (Refinement, last 20%): force expert_ratio=0, shrink-and-perturb
  at phase boundary for plasticity consolidation

Task 11: Expand hyperopt search space from 41D to 45D with 4 new dims:
- her_ratio [0.0, 0.5]: HER relabeling ratio (was hardcoded 0.0)
- curiosity_weight [0.0, 0.2]: intrinsic reward weight (was fixed 0.0)
- use_cvar_action_selection [0.0, 1.0]: risk-aware IQN action scoring
- cvar_alpha [0.01, 0.2]: CVaR confidence level
All wired through build_hyperparams to DQNHyperparameters.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-24 23:19:12 +01:00
jgrusewski
bff4b2807c feat: ensemble value backward + Decision Transformer Phase 1 integration
Task 7: Wire ensemble diversity gradient through value head into shared trunk.
The KL gradient kernel was computing d_logits but gradient flow stopped there.
Now: d_logits → cuBLAS backward W_v2 → ReLU mask → backward W_v1 → d_h_s2 →
trunk backward (layers 2,1) → SAXPY(diversity_weight) into grad_buf.
Uses launch_dx_only for value head layers (skip dW/db — only d_h_s2 needed).
graph_adam sees combined C51 + IQN + ensemble diversity in single update.

Task 8: Decision Transformer Phase 1 already fully implemented. DT runs before
main training loop when dt_pretrain_epochs > 0: builds trajectories from GPU
data, runs pretrain_step per epoch/batch, logs loss. Verified and confirmed.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-24 23:00:13 +01:00
jgrusewski
679a483de5 feat: wire CVaR action selection + implement CQL conservative loss GPU kernel
Task 4 — CVaR Action Selection:
- Add use_cvar_action_selection and cvar_alpha fields to DQNHyperparameters
- Wire from hyperparams into DQNConfig constructor (was hardcoded to false)
- Default: enabled (true) with alpha=0.05 (worst 5% quantile tail)
- Unblocks risk-aware position scaling via IQN head's compute_cvar_q()

Task 5 — Curiosity Wiring (verified active):
- GpuCuriosityTrainer trains forward model on GPU experience data
- train_curiosity_gpu() called from training_loop after experience collection
- curiosity_weight=0.05 (Task 2) gates trainer creation — active when >0
- Intrinsic reward injection into DQN kernel deferred (Phase 4+, per kernel docs)

Task 6 — CQL Conservative Loss GPU Kernel:
- Add use_cql/cql_alpha to GpuDqnTrainConfig (wired from DQNHyperparameters)
- Implement cql_logit_grad_kernel: computes dCQL/d_logits for Branching Dueling C51
  - Per-branch logsumexp penalty with softmax gradient through expectation chain
  - One thread per sample, iterates 3 branches (exposure, order, urgency)
- Add apply_cql_gradient() method: launches CQL kernel + cuBLAS backward_full
  - Accumulates CQL parameter gradients into grad_buf (beta=1.0)
  - Same injection pattern as IQN trunk gradient
- Wire into FusedTrainingCtx::run_full_step() between graph_forward and graph_adam

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-24 22:45:49 +01:00
jgrusewski
cfcaa68572 feat: activate all dormant feature defaults + verify Kelly sizing
- DQNHyperparameters::conservative(): q_gap_threshold 0.0→0.05 (Tier 2 conviction gating active)
- DQNHyperparameters::conservative(): her_ratio 0.0→0.2 (20% HER relabeling enabled)
- All other Tier 2 defaults already active: count_bonus_coefficient=Some(0.1),
  curiosity_weight=0.1, use_cql=true, cql_alpha=0.1, enable_kelly_sizing=true,
  kelly_fractional=0.5, kelly_max_fraction=0.25
- Kelly sizing in experience_kernels.cu confirmed active: ps[14:17] wired,
  f*=(b*p-q)/b formula correct, half-Kelly safety applied, total_trades>=20 gate present

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-03-24 22:27:39 +01:00
jgrusewski
263997ad31 feat: reward v5 — trade-aware hybrid with dynamic trailing stop
Replace reward v4 (pure mark-to-market return) with reward v5, a two-component
trade-aware hybrid that separates dense per-bar signal from sparse trade-exit signal:

Dense (every bar, weight 0.1): raw_pnl / equity when in a trade, zero when flat.
Keeps gradients flowing without overwhelming the sparse trade completion signal.

Sparse (at trade exit, weight 2.0): trade_return * patience_multiplier where
patience = sqrt(hold_time / expected_hold). Regime-adaptive expected hold via
ADX: trending (ADX>30) = 20 bars, ranging (ADX<20) = 8 bars, default = 12 bars.

Dynamic trailing stop: regime-adaptive trail distance (0.5% base, widens with
volatility via CUSUM and trend via ADX). Activates when trade is profitable and
held > 2 bars. Locks in profits by forcing exit when unrealized P&L drops below
the trailing floor.

Also fixes kernel signature mismatch: removes 7 old reward v2 parameters
(w_dsr, w_pnl, w_dd, w_idle, dd_threshold, time_decay_rate, eta) that were
already removed from the Rust launcher in a prior commit.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-24 22:16:31 +01:00
jgrusewski
be2672596d docs: master plan — maximize all 25 DQN features for profitable ES trading
12 tasks covering:
- Reward v5: trade-aware hybrid (dense + sparse) with patience multiplier
- Dynamic trailing stop as environment physics (regime-adaptive)
- Activate 11 dormant features: IQN CVaR, HER, ensemble, DT, CQL, curiosity,
  count bonus, entropy, Kelly sizing, Q-gap conviction, CVaR action selection
- Three-phase training: behavioral cloning → online RL → DT refinement
- Ensemble consensus action selection with uncertainty-based sizing
- Full hyperopt search space (50+ dimensions)
- Smoke test validation: model must produce winning trades

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-24 22:00:45 +01:00
jgrusewski
825db90f23 feat: comprehensive DQN training pipeline overhaul — 16 bug fixes, MSE warmup, financial metrics
Major fixes:
- C51 v_range calibrated for reward v4 (±2.0, was ±25/±0.5)
- Wrong Flat index in Q-gap filter (qe[4]→qe[2] in branching_action_select)
- hold_time tracks total position duration (was only losing bars)
- Entropy coefficient wired to C51 backward kernel (0.001, was unwired)
- Count bonus wired to GPU action selection (per-branch UCB)
- Q-gap warmup ramp (0→threshold over 5 epochs, was static)
- IQN lambda gradient scaling (max_grad_norm × (1+lambda))
- PER beta annealing 4x faster (500 steps, was 2000)
- Reward normalization disabled (scrambled per-bar returns)
- Capital floor uses natural return (was hardcoded -1.0)
- Financial metrics pipeline: real per-trade GPU stats (was Trades=1)

New features:
- MSE loss CUDA kernel for C51 warmup phase
- Blended MSE→C51 loss with linear alpha ramp
- GPU trade_stats_reduce kernel for per-trade financial metrics
- TradeStats struct with real win/loss/PF from portfolio states
- Behavioral smoke test (Q-values, action entropy, trades)
- 50-epoch convergence test with anomaly detection
- c51_warmup_epochs in hyperopt search space (41D)

Dead code removed:
- portfolio_sim_kernel (150 lines CUDA)
- DSR/PnL/drawdown reward v2 computations
- 7 dead kernel params from env_step signature
- GpuPortfolioSimulator (never called)
- Reward normalization block + state fields

0 warnings, 0 errors, 1241 unit tests + 8 smoke tests pass.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-24 21:48:38 +01:00
jgrusewski
3b25382864 feat: reward v4 — pure mark-to-market portfolio return per bar
Replaces the 8-component dense shaping + sparse trade completion with:
  reward_t = (equity_t - equity_{t-1}) / equity_{t-1}

- Losing bars get NEGATIVE reward (every bar, not just exit)
- Flat bars get ZERO reward
- Trade entry: tx_cost hits cash → immediate negative reward
- Loss aversion: losses weighted 1.5x (prospect theory)
- No more DSR, no dense shaping, no hold penalty

NOTE: v_range needs recalibration for percentage returns (~0.001/bar)
instead of the old reward scale (~0.01-2.0). Current v_range=42 is
1000x too wide for the new reward magnitude.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-24 16:41:15 +01:00
jgrusewski
02169e16e2 fix: risk management re-enabled + backtest tx_cost consistency
1. Risk management (CVaR, conviction, Kelly) re-enabled as ENVIRONMENT PHYSICS:
   - Agent observes scaling via portfolio state features
   - Learns to account for risk limits in its policy
   - No longer destroys credit assignment (scaling is physics, not action override)
   - Kelly uses half-Kelly (0.5x) for safety, activates after 20 trades

2. Backtest tx_cost now uses training's transaction_cost_multiplier from hyperopt
   (was hardcoded 0.1 bps — 17x lower than training). Training and eval see same costs.

3. Backtest env tx_cost formula expanded to match training:
   - Square-root market impact (Almgren-Chriss)
   - Order-type premiums (Market=0, IoC=+2bps, LimitMaker=-5bps)

Result: first POSITIVE Sharpe (+0.0838) in project history. 134K trades.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-24 16:15:18 +01:00
jgrusewski
511a502ea4 fix: 3 more audit findings — monitoring defines, backtest tx_cost consistency
1. Monitoring kernel: prepend common_device_functions.cuh for DQN_ORDER_ACTIONS
2. Backtest metrics kernel: prepend common_device_functions.cuh
3. Backtest gather kernel: prepend common_device_functions.cuh
4. Backtest tx_cost: use training's transaction_cost_multiplier from hyperopt
   (was hardcoded 0.1 bps — 17x lower than training's ~1.7 bps multiplier)

All standalone kernels now consistently include common_device_functions.cuh
for DQN action space defines. No more hardcoded constants.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-24 15:48:40 +01:00
jgrusewski
432fcdc7ee fix: 8 critical bugs — tx cost 10000x, action decode, state mismatch, monitoring
CATASTROPHIC fixes:
1. TX cost missing *0.0001f bps conversion — trades cost $17K instead of $1
2. Backtest action decode: raw factored int→800% exposure (should decode exposure_idx)
3. State mismatch: training 66 features, backtest 45 — Q-values at eval were garbage
4. Position scaling disabled: CVaR/conviction/Kelly destroyed credit assignment

Monitoring fixes:
5. Q-value labels: 5-element array→9-element for 9-action exposure space
6. Monitoring kernel: add common_device_functions.cuh for DQN_ORDER_ACTIONS defines
7. Backtest metrics: factored action decode for buy/sell/hold classification
8. Backtest gather: add common_device_functions.cuh for MARKET_DIM/PORTFOLIO_DIM

Result: agent now trades 71K times (was 1) with all 9 exposure levels explored.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-24 15:32:39 +01:00
jgrusewski
ddac49509b fix: 5 critical training fixes — IQN sequencing, grad clip, v_range, reward, normalization
1. IQN gradient sequencing: train_step_gpu replays forward ONLY, caller injects
   IQN/attention/ensemble gradients into grad_buf, then calls replay_adam_and_readback().
   Single Adam sees combined C51+IQN gradient. Previously IQN was a NO-OP (SAXPY
   happened after both graphs completed — Adam already consumed gradients).

2. Gradient clip 1.0 → 10.0: C51 with 101 atoms × 3 branches produces 300x larger
   gradients than standard DQN. Clip at 1.0 made effective LR ~7e-12. Result:
   grad_norm 137K → 490 (280x reduction, network actually learns now).

3. max_abs_reward 3.0 → 1.5: tighter C51 support [-42, +42] instead of [-84, +84].
   Q-values at 24 (58% of v_max) instead of 81 (96%). 2x atom resolution.

4. choppy_bonus removed: Flat reward was 0.02 on 60-70% of bars, dominating
   normalized reward distribution. Now Flat gets exactly 0.0.

5. Reward normalization: Welford EMA was broken (alpha=0.01 over 150K samples →
   variance converges to zero → divides by 1e-8 → Q-value explosion). Fixed with
   batch-level mean/std + EMA blending + variance floor 0.01.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-24 12:38:37 +01:00
jgrusewski
b0b8c94d77 feat: complete GPU training pipeline — all features wired, zero CPU hot path
Split CUDA Graph (forward + adam phases with gradient injection point):
- IQN trunk gradient flows through single Adam (no dual optimizer conflict)
- Spectral norm runs BEFORE forward (not after Adam — no tug-of-war)
- σ_max in 40D hyperopt search space [1.0, 10.0]

Attention Phase B backward:
- Full gradient flow through 4-head self-attention weights
- Separate Adam optimizer for attention params
- Backward kernel recomputes forward from saved_input (memory-efficient)

Ensemble multi-head:
- Real cuBLAS value head forward per ensemble head (was copying head 0 logits)
- KL diversity gradient kernel with hierarchical reduction
- forward_value_head() on CublasForward for per-head SGEMM

Regime PER scaling:
- Kernel reads target ADX/CUSUM from states_buf directly (zero CPU readback)
- Removed 2x memcpy_dtoh per training step

Decision Transformer:
- 14 CUDA kernels (embed, causal attention, FFN, CE loss + backward + trajectory building)
- GPU-native trajectory builder (return-to-go reverse cumsum, momentum expert actions)
- Wired into training loop with dt_pretrain_epochs config

HER Future/Final:
- episode_ids flow through PER buffer (GpuBatch, GpuReplayBuffer, GpuBatchSlices)
- GPU-native donor sampling (binary search on episode boundaries)
- Strategy dispatch in fused_training.rs

Backtest SEGV fix:
- Missing q_gaps_buf argument in action_select kernel launch
- Dynamic branch_sizes from agent (not hardcoded)

Local test: objective=10.48, Sharpe=0.0419, 175K trades, zero errors

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-24 11:28:36 +01:00
jgrusewski
8e1508ab7a feat: add DT pre-training config fields and loop hook to DQN trainer
- Add 5 DT fields to DQNHyperparameters (dt_pretrain_epochs, dt_context_len,
  dt_embed_dim, dt_num_layers, dt_target_return) with conservative() defaults
  (dt_pretrain_epochs=0 = disabled)
- Wire pre-training phase before the main DQN epoch loop: logs config when
  dt_pretrain_epochs > 0, noting that pretrain_step() kernels are ready in
  decision_transformer.rs and full integration awaits trajectory data pipeline
- Add dt_pretrain_epochs to DQNParams struct, Default impl, from_continuous,
  and the hyperparams struct literal in the hyperopt adapter (fixed to 0 for
  all hyperopt trials, not in search space)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-03-24 02:22:43 +01:00
jgrusewski
28e2bcbf22 feat: HER Future/Final strategies via GPU episode boundary tracking
Add GPU-native episode boundary tracking to enable HER Future and Final
strategies, which previously could not be used (only Random worked).

- Create her_episode_kernel.cu with 3 kernels:
  - fill_episode_ids_kernel: fills episode_ids[i] = i/L, zero CPU
  - her_sample_future_donors: samples random donor LATER in same episode via LCG
  - her_find_episode_end: finds last transition index of same episode
- Add episode_ids CudaSlice<i32> field to GpuExperienceBatch; filled by
  fill_episode_ids_gpu() on every collect_experiences_gpu() call
- Wire fill_episode_ids_kernel compilation into GpuExperienceCollector::new()
- Add future_donors_func, episode_end_func, and rng_states fields to GpuHer
- Add relabel_batch_with_strategy() dispatching to GPU episode kernels based
  on HerGpuStrategy (Future or Final); errors on Random (use existing path)
- All episode ID computation uses direct integer arithmetic (i/L), no scan

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-03-24 02:21:21 +01:00
jgrusewski
28475ff3ec feat: Ensemble multi-head Q-network with KL diversity loss (Task 5)
Adds K independent value/advantage head weight sets sharing a common DQN
trunk. Provides uncertainty estimation (Q-value variance across heads) and
diversity regularization (KL divergence between head distributions).

Architecture:
- Head 0 stays inside CUDA Graph (zero overhead for ensemble_count=1 default)
- Heads 1..K-1 run outside CUDA Graph using post-graph save_h_s2 activations
- DtoD clone of head weights at init with stream-sync; diversity grows over training
- Pairwise KL uses symmetrized Jensen–Shannon divergence for numerical stability

New files:
- ensemble_kernels.cu: two NVRTC kernels — ensemble_aggregate_kernel (mean/var
  Q-values across K heads) and ensemble_diversity_kernel (hierarchical warp→block
  reduction matching dqn_grad_norm_kernel pattern, no flat atomicAdd)
- compile_ensemble_kernels() function in gpu_dqn_trainer.rs
- New GpuDqnTrainer accessors: on_v_logits_buf(), tg_h_v_scratch_ptr()

FusedTrainingCtx changes:
- ensemble_extra_heads: Vec<(DuelingWeightSet, BranchingWeightSet)>
- Pre-allocated GPU buffers (logits, mean_q, var_q, diversity_loss)
- run_ensemble_step() method runs after CUDA Graph replay (EventTrackingGuard)
- Wired in run_full_step() between IQN PER step and spectral norm step

Config: ensemble_count=1 (default, zero overhead), ensemble_diversity_weight=0.01

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-24 02:07:39 +01:00
jgrusewski
617b9fb718 feat: attention backward pass Phase A (residual passthrough)
Add backward_residual() method to GpuAttention that implements
gradient flow through the residual connection for frozen attention
weights. This is Phase A of the attention backward implementation.

The residual connection (output = input + attention(input)) ensures
that gradients flow through unchanged: d_input = d_output. With frozen
weights, we don't compute attention weight gradients, making this a
literal no-op when the input and output buffers alias (as they do in
apply_iqn_trunk_gradient).

Phase B will add attention path gradients when weights are unfrozen.
2026-03-24 01:52:22 +01:00
jgrusewski
a008d1671b feat: wire GpuAttention forward pass into DQN fused training pipeline
Implements Task 1 from docs/superpowers/plans/2026-03-24-remaining-gpu-features.md.
The existing attention CUDA kernel (attention_kernel.cu) is now called as a
post-graph operation that modifies save_h_s2 for the next graph replay (1-step lag).

Changes:
- config.rs: add `use_attention: bool` to DQNHyperparameters (default: false)
- fused_training.rs: add `gpu_attention: Option<GpuAttention>` to FusedTrainingCtx;
  initialize from hyperparams.use_attention in new() using shared_h2 as state_dim;
  call apply_attention_forward() in run_full_step() after EMA (Step 3b), before IQL
- gpu_dqn_trainer.rs: add GpuAttention import and apply_attention_forward() method
  that stream-syncs, wraps in EventTrackingGuard, calls attention.forward(), then
  DtoD-copies the attended output back into save_h_s2

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-03-24 01:42:03 +01:00
jgrusewski
df398b51d0 feat: IQN backward flows gradient to shared trunk (dual gradient source)
The IQN backward kernel now computes dL/d(h_s2) = dL/d(combined) ⊙ embed
and outputs it to d_h_s2_buf [B, hidden_dim]. Previously this was
explicitly NOT computed (comment: "trunk trained by C51").

Now the shared trunk receives BOTH gradient signals:
  C51: dense cross-entropy gradient (can be noisy/steep)
  IQN: bounded Huber quantile gradient (always stable)

The IQN gradient stabilizes trunk training when C51's gradient is steep.
With both signals, the trunk learns from C51's distributional knowledge
AND IQN's risk-aware quantile knowledge simultaneously.

New: d_h_s2_buf allocated in GpuIqnHead, zeroed before backward,
accumulated via atomicAdd across all quantiles per sample.

Accessors added: GpuDqnTrainer::bw_d_h_s2_buf(), shared_h2()

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-23 21:13:01 +01:00
jgrusewski
54b3382f8c config: lower default LR 1e-4→3e-5 for stability with 72-dim state
With the larger state space (72 dims: market + portfolio + multi-TF)
and complex reward structure, the old 1e-4 default produces steep
loss surfaces and doubling grad_norm per epoch. 3e-5 gives the
optimizer a gentler starting point. Hyperopt still searches [1e-5, 3e-4].

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-23 21:06:12 +01:00
jgrusewski
23fb59db2b feat: Layer 3 — IQN primary with fixed τ (QR-DQN, CUDA Graph compatible)
Layer 3a: Fixed τ midpoints replace random sampling in IQN.
τ_i = (2i-1)/(2N) for i=1..N, pre-computed in constructor.
Eliminates sample_taus_kernel launch → fully deterministic →
CUDA Graph compatible. IQN becomes equivalent to QR-DQN.

Layer 3b: IQN per_sample_loss replaces C51 td_errors for PER.
GPU DtoD copy — zero CPU. IQN's Huber quantile loss is bounded
by construction (max = kappa²/2 per quantile). This eliminates
the PER feedback explosion that caused C51 loss to reach 242K.

No conditional flag — IQN is always primary when active.
Three-layer defense complete:
  L1: Per-sample CE clamp at 50 (breaks PER loop)
  L2: Label smoothing ε=0.01 on Bellman target (prevents log(0))
  L3: IQN Huber loss for PER priorities (bounded by construction)

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-23 20:52:20 +01:00
jgrusewski
a51345a141 fix: Layer 1+2 — C51 loss clamp + label smoothing (prevents 242K explosion)
Layer 1: Per-sample CE clamped to MAX_PER_SAMPLE_CE=50 before IS weight.
Breaks PER feedback loop: high loss → high priority → high IS weight → repeat.

Layer 2: Bellman target smoothed with ε=0.01 uniform mix after projection.
Prevents any atom from having zero probability → no log(near-zero) in CE.

Root cause: C51 cross-entropy is unbounded when target and predicted
distributions are maximally misaligned. PER amplifies pathological samples.
These two layers cap the maximum possible CE and prevent the extreme
misalignment from occurring.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-23 20:45:11 +01:00
jgrusewski
909650f2a1 plan: three-layer stable distributional loss architecture
Layer 1: Per-sample C51 loss clamp (MAX_CE=50) breaks PER feedback loop
Layer 2: Label smoothing (ε=0.01) on Bellman target prevents log(0)
Layer 3: IQN primary with fixed τ (QR-DQN style, CUDA Graph compatible)

5 tasks, each independently testable and committable.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-23 20:42:08 +01:00
jgrusewski
f5cb953082 fix: gradient clip 10.0→1.0 — prevents grad accumulation with complex reward
grad_norm was growing 5x per epoch (523→2673→13K→67K→NaN). The old
clip at 10.0 allowed gradients to accumulate. With clip=1.0, the
Adam optimizer receives bounded updates.

Trial 2 (TPE-guided params) trains cleanly for 4 epochs:
  train_loss: 4.5→3.3 (decreasing!)
  Q-value: 12-19 (stable)
  grad_norm: 162-567 (bounded)

Trial 1 (random initial params) still NaN's — this is expected and
handled by the hyperopt penalty (1M objective). The optimizer learns
to avoid unstable parameter regions.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-23 14:40:30 +01:00
jgrusewski
1936c160b5 fix: NaN guards + multi-timeframe feature clamping
Gradient explosion from unbounded multi-timeframe features (240-bar
return could be ±50%). All multi-timeframe outputs now clamped:
- Return: ±10%
- Volatility: 0-10%
- Volume ratio: 0-5x
- Momentum: [0, 1] (already bounded)

NaN guards added:
- Kelly: NaN check on continuous/discrete Kelly before blend
- target_position: final NaN→0 after all scaling
- reward: final NaN→0 before writing to replay buffer

Local test shows Q-gap=18.7 (epoch 5) — 18x better than start-of-session 0.000.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-23 14:07:21 +01:00
jgrusewski
b7ddf812cb feat: multi-head attention + Decision Transformer architecture (GPU-only)
P10: Multi-Head Feature Attention (attention_kernel.cu + gpu_attention.rs)
- 4-head self-attention over 72-dim state features
- Learned W_Q, W_K, W_V, W_O projections (21K params, ~82KB)
- Softmax + residual connection + layer normalization
- Xavier initialization, one warp per sample (32 threads)
- Zero CPU: all weights on GPU, kernel launch only

P11: Decision Transformer (decision_transformer.rs)
- Sequence model conditioned on return-to-go (Chen et al., 2021)
- Token embedding: (state + return-to-go + action) → embed_dim
- N transformer layers with causal multi-head attention + FFN
- Context window: 20 timesteps (attend to recent trading history)
- Pre-training on offline walk-forward data, fine-tuning with DQN
- Architecture defined, CUDA transformer kernels to follow
- Config: embed_dim=128, num_layers=3, num_heads=4

Both modules are GPU-only — CudaSlice allocations, kernel launches,
zero memcpy_dtoh in hot path.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-23 13:01:13 +01:00
jgrusewski
d789bac2a2 feat: add expert demonstration generator for DQN imitation learning warmup
Add ExpertDemoGenerator using MA crossover strategy filtered by ADX trend
strength to produce (bar_index, exposure_action_index) pairs. Fast/slow EMA
crossover with ADX > threshold triggers Long100/Short100 signals, otherwise Flat.
Includes effective_ratio() for linear decay of expert_demo_ratio over training
epochs (controlled by expert_demo_ratio and expert_demo_decay_epochs fields
in DQNHyperparameters). Comprehensive unit tests for EMA, ADX, decay, and
edge cases included.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-23 12:54:30 +01:00
jgrusewski
b370b93cc9 feat: add curriculum learning with ADX-based difficulty scoring to walk-forward windows
Add difficulty_score field to WalkForwardWindow computed via compute_difficulty()
which calculates mean ADX(14) over training bars. Add DifficultyPhase enum (Easy/Mixed/Full)
and filter_by_difficulty() for curriculum-based window filtering — Easy phase trains
only on trending markets (ADX>30), Mixed duplicates trending windows for 2x weight,
Full uses all windows equally. The curriculum_phase field in DQNHyperparameters
(committed in previous ensemble commit) controls the active phase.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-23 12:54:01 +01:00
jgrusewski
0718dabd4c feat: GPU-native multi-timeframe features — 4 windows × 4 features
16 new features computed directly in the state_gather CUDA kernel
from market data already on GPU. No CPU feature engineering.

Lookback windows: 5, 15, 60, 240 bars (≈5min, 15min, 1hr, 4hr)
Features per window:
  - Return over N bars (directional bias)
  - Volatility (high-low range / close)
  - Volume trend (current / N-bar average)
  - Momentum (position within N-bar range [0=bottom, 1=top])

State dim: 66 raw (72 aligned) without OFI, 74 raw (80 aligned) with.
The model now sees price action at 4 timescales simultaneously.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-23 12:52:36 +01:00
jgrusewski
140b9426c3 feat: add ensemble_count and ensemble_diversity_weight to DQNHyperparameters
Wire ensemble agent configuration into DQN hyperparameters:
- ensemble_count: number of agents in ensemble (default 1 = no ensemble)
- ensemble_diversity_weight: KL-divergence diversity loss weight (default 0.01)

Both fields default to backward-compatible values (single agent, no
diversity loss) and are plumbed through the conservative() constructor.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-23 12:52:29 +01:00
jgrusewski
d1b183c6dc feat: add Phase Risk (P4) to four-phase hyperopt — 8D risk parameter search
Add HyperoptPhase::Risk variant that fixes dynamics + architecture from
Phase 2 best params and searches only risk parameters (8D):
kelly_fractional, dd_threshold, loss_aversion, time_decay_rate,
q_gap_threshold, w_dsr, w_pnl, w_dd.

CLI: --phase risk (alongside fast, full, reward)

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-23 12:51:49 +01:00
jgrusewski
33d91cce5b spec: profitability roadmap — 11 findings from deep research
4 critical root causes fixed this session (stop-loss, dense shaping,
action aliasing, HFT objective weight). 7 remaining improvements
documented with priority order: three-phase hyperopt, behavioral
cloning warm start, curriculum learning, ensemble, multi-timeframe,
attention, Decision Transformer.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-23 12:45:06 +01:00