apply_position_mask() in factored_q_network.rs looped 0..45 against a
5-wide Q-values tensor — would panic at runtime. Changed to 0..5 using
ExposureLevel::from_index() directly. Updated stale "45 actions" comments
in 5 files (dqn.rs, reward.rs, hyperopt/adapters/dqn.rs, curriculum.rs).
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
A: Switch backtest eval from Gumbel softmax to greedy argmax (batch_greedy_actions)
so hyperopt Sharpe reflects the agent's actual learned policy, not noisy sampling.
C: Disable reward normalization (enable_normalization=false). EMA normalizer with
±3.0 clipping was flattening the reward landscape, preventing the agent from
distinguishing large winners from scratch trades.
D: Wire tx_cost_bps (0.1 bps for IBKR ES) through to EvaluationEngine via
new_with_fee_rate(). Previously hardcoded at 15 bps (150x mismatch with actual
commission costs), massively penalizing every trade in backtest.
E: Scale PnL reward by agent's target exposure in calculate_pnl_reward().
Previously, a Short100 action received POSITIVE reward when market went up
(pct_return ignored position direction). Now: reward = pct_return × exposure.
2735 tests pass, 0 clippy warnings.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Early stopping with 8-epoch trials returns penalty metrics (no backtest),
causing ALL trials to hit FALLBACK OBJECTIVE = 44.6 regardless of actual
model quality. The Sharpe was 1.36-1.61 but natural fluctuation triggered
"Sharpe worsening" at epoch 5, killing the backtest evaluation.
Fix: disable early stopping when epochs ≤ 12. With ~90s per trial, running
all 8 epochs is cheap. Long training runs (50 epochs) still use it.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
The C2 fix from the previous session was too aggressive — it eliminated
ALL exploration mechanisms simultaneously:
1. epsilon forced to 0.0 when noisy nets active (select_action)
2. count bonus removed from Q-value computation (metrics only)
3. noisy_epsilon_floor config field declared but never read
This left noisy nets as the sole exploration mechanism, which produces
perturbations too small to overcome Q-value gaps (A4=0.12 vs others≈0.02).
Result: 20/20 hyperopt trials hit fallback objective with 1/5 diversity.
Fixes:
- select_action: use noisy_epsilon_floor (not 0.0) as effective_epsilon
when noisy nets active — guarantees minimum random action rate
- select_action: re-enable UCB count bonus on Q-values before argmax
(both IQN and standard paths) for directed exploration
- select_action_with_confidence: same fixes for consistency
- trainer: set epsilon to noisy_epsilon_floor (not 0.0) at init
- hyperopt: widen noisy_epsilon_floor range from [0.0, 0.05] to
[0.03, 0.15] with default 0.05
Exploration now has two complementary mechanisms:
- noisy_epsilon_floor: random actions feed diverse replay buffer
- count bonus: UCB term biases greedy selection toward under-visited actions
- noisy nets: weight perturbation adds stochasticity to Q-values
select_action_inference (production) is unchanged — pure exploitation.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Critical bug: all 3 DQN action selection methods (select_action,
select_action_with_confidence, select_action_inference) used
FactoredAction::from_index() which maps indices 0-4 to exposure_idx=0
(Short100) via division by 9. This is the root cause of action
diversity collapse during both training and production inference.
Fix: ExposureLevel::from_index() + OrderRouter::route_default() in all
DQN paths. Also fixes hyperopt objective thresholds (<10 → <3 for
5-action degenerate detection), stale defaults/comments, integration
test configs.
Files: dqn.rs (3 methods), trainer.rs (validation + select_action),
hyperopt/adapters/dqn.rs (thresholds), dqn_model.rs (comments),
train_baseline_rl.rs (default), reward.rs (comment),
dqn_integration.rs + ensemble_integration.rs (num_actions).
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- agent.rs: select_action_factored() now uses ExposureLevel::from_index()
+ OrderRouter::route_default() instead of FactoredAction::from_index().
Previously, indices 0-4 mapped to all-Short100 variants in the 45-action
space — now correctly maps to 5 distinct exposure levels.
- hyperopt: plateau_window .max(3) → .max(5) to prevent premature early
stopping with short trial epochs.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Root cause: 45 factored actions (5 exposure × 3 order × 3 urgency) caused
reward degeneracy — 9 actions per exposure level produced nearly identical
rewards since order type/urgency had 1000-4000x weaker signal than PnL.
This collapsed action diversity as DQN couldn't differentiate actions.
Changes:
- DQN now outputs 5 Q-values (Short100, Short50, Flat, Long50, Long100)
- New OrderRouter deterministically maps exposure → (order_type, urgency)
based on spread and volatility microstructure signals
- PPO retains full 45-action space (separate CUDA constants DQN_NUM_ACTIONS
vs PPO_NUM_ACTIONS)
- CUDA kernels: DQN diversity entropy uses 5 categories, PPO keeps 45
- Phase B: pnl_history cleared per epoch so Sharpe reflects current epoch
(was accumulating across all epochs, causing frozen Sharpe metric)
24 files, 2728 tests pass, 0 clippy warnings
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Sharpe-based early stopping kills every hyperopt trial at epoch 4
because compute_epoch_financials() is deterministic (greedy argmax on
fixed validation data) — the model doesn't change enough in 8 short
epochs to shift any argmax decisions, making Sharpe bit-identical
across epochs and triggering plateau detection immediately.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
The previous C4 fix re-enabled early stopping with adaptive plateau
window but still used val-loss as the stopping metric. Val-loss (TD
Bellman residual) can plateau while trading strategy still improves.
Now:
- Best-checkpoint saved when epoch Sharpe improves (not val-loss)
- Plateau detection checks sharpe_history (not val_loss_history)
- Patience-based EarlyStopping receives -Sharpe (negate for lower=better API)
- Per-epoch Sharpe extracted from compute_epoch_financials() (already computed)
This ensures early stopping and best-model selection track the metric
that actually matters for hyperopt: trading performance.
2720 tests pass, 0 clippy warnings.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Multi-window backtest: splits validation data into 3 non-overlapping
windows and aggregates with mean(Sharpe) - 0.5*std(Sharpe), penalizing
inconsistency and reducing overfit to a single data segment.
Top-K ensemble: hyperopt now emits top_k_params (top 5 trials) in JSON
output. train_baseline_rl gains --ensemble-top-k flag to train multiple
models per fold from different hyperopt configs, saving checkpoints as
dqn_ensemble_{k}_fold_{n}.safetensors.
Workflow template: adds ensemble-top-k parameter (default 5) and passes
--ensemble-top-k to the train-best step.
2720 tests pass, 0 clippy warnings.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
4 root causes of 45-action DQN collapsing to 1-6 actions:
1. Batch epsilon ignoring noisy_epsilon_floor: select_actions_batch()
and select_actions_batch_gpu() used get_epsilon() which returns 0.0
with noisy nets — zero random exploration in the training path.
Added get_effective_epsilon() that respects noisy_epsilon_floor.
2. Entropy coefficient too weak: default 0.05 with bounds (0.01, 0.2)
produced max ~0.19 bonus vs TD loss of 4+. Bumped default to 0.1,
widened bounds to (0.05, 0.5) for effective anti-collapse.
3. count_bonus_coefficient not in search space: was hardcoded at 0.1
in from_continuous(). Promoted to 31st search dimension with bounds
(0.05, 1.0) so PSO/TPE can optimize exploration strength.
4. Diversity penalty too coarse: objective had <10 unique actions
short-circuit but nothing for 10-20. Added graduated penalty that
linearly ramps from 3.0 (10 actions) to 0.0 (20 actions).
Also fixes pre-existing clippy impl_trait_in_params in optimizer.rs.
2720 tests pass, 0 clippy warnings.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Wire record_hyperopt_trial_duration, set_hyperopt_best_objective,
set_hyperopt_trial_best_loss, and set_hyperopt_elapsed into all three
optimizer paths (PSO sequential, PSO parallel, TPE). These metrics were
registered but never called during the optimization loop, causing
"Best Objective Over Time", "Trial Duration", and "Elapsed Time"
Grafana panels to show "No data".
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Ephemeral Argo workflow pods terminate after training completes, causing
Prometheus to lose all scraped metrics. Add push_to_gateway() to POST
final metrics to the existing pushgateway service so they persist on the
Grafana training dashboard after pod completion.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
The batch training path (select_actions_batch → forward()) does NOT
increment DQN::total_steps. Only select_action() does. When
warmup_steps > 0, train_step() checks total_steps < warmup_steps
and returns (0.0, 0.0) — zero loss, zero gradients. The model
never trained; "results" were random initialization Q-values.
Fix: force warmup_steps=0 in hyperopt adapter and remove
warmup_ratio from the 31D→30D search space (saves a dimension
for TPE/PSO effectiveness).
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- Replace Silverman's bandwidth (h = 1.06σn^(-1/5)) with Scott's rule
(h = 0.7σn^(-1/(d+4))) for tighter kernels in high-D parameter spaces
- Add best-trial injection: always evaluate EI at best known point plus
5 small perturbations (±5%), preventing optimizer from forgetting peaks
- Scale n_candidates dynamically: max(256, 8*n_dims) instead of fixed 100
- Reduce gamma from 0.25 to 0.15 when trials < 50 for tighter exploitation
- Wire model_name through PSO/TPE paths for per-trial Prometheus metrics
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Re-export HardwareTimestamp through data crate instead of ml/risk
depending directly on trading_engine. Reduces coupling between
the ML pipeline and the trading engine.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Two issues causing all 20 hyperopt trials to have identical f64::MAX
objective (TPE optimizer blind):
1. Val-loss plateau early stopping fired at epoch 5-6 of every 8-epoch
trial (plateau_window=5 too aggressive for short runs). Disabled
early_stopping_enabled for hyperopt; gradient-collapse patience
still active as safety net.
2. Penalty metrics used f64::MAX for gradient_norm/q_value_std which
produced ~3.6e+308 objective. Changed to 100.0 so TPE can still
differentiate between early-stopped trials by other metric fields.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- Hyperopt: catch early stopping errors and return penalty metrics
instead of aborting entire run. Trials that stop early are scored
as poor (objective=-500) so optimizer avoids those configs.
- DaemonSet: detect GPU nodes (NVIDIA conf.d overlay) and create
v3-compatible registry config drop-in for containerd v2.1+.
Old grpc.v1.cri.registry path is silently ignored by containerd v2.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Fixes from deep investigation audit (LOW/MEDIUM priority):
1. CVaR penalty: hard cliff (0 or 10) → smooth ramp with gradient signal
for PSO. Formula: min(10, max(0, -cvar-0.05)*200).
2. Clip outliers leakage: data_loading.rs now computes clip bounds from
training portion only (first 80%), then applies to full series.
Log returns and windowed normalize are causal (no leakage).
3. Noisy sigma scheduler: hyperopt now matches conservative() defaults
(enabled, initial=0.8, final=0.4) so hyperopt-found params
generalize to train_best without scheduler mismatch.
4. evaluate_supervised.rs: NormStats fallback from test data (leakage)
replaced with bail! matching evaluate_baseline.rs behavior.
5. Doc comments: stale 27D references updated to 31D (4 locations).
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
4 bugs found by deep investigation agents:
1. HIGH: minimum_profit_factor (search dim 30) was never forwarded from
DQNHyperparameters to DQNConfig — trainer hardcoded 1.5, making the
entire dimension wasted. Added field to DQNHyperparameters, wired
through trainer.rs.
2. HIGH: Backtest EvaluationEngine used hardcoded $10K initial capital
while training used $35K (self.initial_capital). Returns/Sharpe were
3.5x distorted. Now uses self.initial_capital.
3. MEDIUM: calculate_hft_activity_score_wave10 multiplied already-100x
buy_pct/sell_pct by 100 again, making the diversity penalty threshold
(15%) unreachable (values were ~2700). Removed double multiplication.
4. MEDIUM: Sortino ratio returned 0.0 for all-positive returns (no
downside deviation), penalizing perfect strategies in the 40%-weighted
composite score. Now returns 100.0 (capped) when mean return > 0.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Degenerate trials with <10/45 unique actions now short-circuit the objective
to a graduated penalty (8.0 for 1 action → 0 for 10+), ignoring composite
score entirely. Previously, phantom Sharpe=2317 drove objective to -369k,
trapping TPE in the low-temp region.
Combined with temp floor raise (0.1→0.5 from prior commit), this eliminates
both the supply (no low temps) and demand (no reward) for degenerate trials.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Hyperopt TPE was stuck chasing phantom Sharpe=2317 from degenerate trials
with 2/45 unique actions (temp 0.15-0.20). Two fixes:
1. Raise eval_softmax_temp bounds from [0.1, 2.0] to [0.5, 2.0] — temps
below 0.3 consistently produce 1-3/45 actions regardless of model quality
2. Add diversity_penalty to objective: 50*(1 - unique/10) for <10/45 actions,
so degenerate trials score badly even if they accidentally reach the TPE
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Rewrite batch_hierarchical_softmax_actions to keep all sampling on GPU:
reshape [N,45]→[N,5,9], max-pool→[N,5] exposure Q, Gumbel-max over
exposures, parallel Gumbel-max over all sub-actions, gather chosen
exposure's sub-action. Only CPU transfer: N flat indices.
Add imagePullPolicy: Always to compile-services and compile-training
containers so rebuilt builder images are never stale-cached.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Flat Gumbel-max over 45 actions can collapse to a single exposure
bucket (e.g., 2/45 unique actions all in Long100) even with adequate
temperature, because order/urgency diversity doesn't produce trades.
Two-stage sampling: first Gumbel-max over 5 exposure levels (max Q
per level), then Gumbel-max over 9 order×urgency combos within the
chosen exposure. Guarantees exposure-level diversity proportional to
actual Q-value differences.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- Raise eval_softmax_temp lower bound from 0.01 to 0.1 (log scale) to
prevent near-greedy collapse in hyperopt walk-forward eval. Temps below
0.05 produced degenerate 1-trade trials even with softmax sampling.
- Mount MinIO CA cert in CI compile pods and append to system trust store
so sccache S3 backend can verify MinIO's self-signed TLS certificate.
- Add openssh-client to ci-builder-cpu Dockerfile (missing, broke git
clone over SSH).
- Bump MinIO memory limits from 512Mi to 2Gi (OOMKilled under load).
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Add device_pool to DQN/PPO hyperopt trainers for round-robin GPU assignment
per trial. Binary detects all CUDA devices, scales VRAM budget by GPU count.
Single-GPU: no behavior change (pool of 1).
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
T=0.1 was too conservative for degenerate configs — trials with uniform
Q-values still collapsed to 1 unique action. Making temperature a
hyperopt parameter lets TPE learn the optimal exploration-exploitation
balance per config. Range [0.01, 2.0] log-scale.
Also sets training-workflow default gpu-pool to ci-training-h100.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
DQN hyperopt walk-forward eval collapsed to 1 trade when Q-values were
nearly uniform (greedy argmax always picked the same action). Replace
batch_greedy_actions with batch_softmax_actions using the Gumbel-max
trick (argmax(Q/T + Gumbel(0,1))) — fully GPU-resident, no CPU
softmax/sampling. Temperature 0.1: nearly greedy when Q-values are
separated, diverse when uniform. Logs unique_actions/45 per backtest.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Replace legacy Buy/Sell/Hold collapse with exposure-aware evaluation.
Long100→Long50 now generates a partial-close trade instead of being
a no-op. Combined with graduated trade penalty, PSO now gets gradient
signal across the entire 26D search space.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Add calculate_trade_insufficiency_penalty() that produces monotonically
decreasing penalties for low trade counts (10.0 at 0 trades, 0.0 at 100+),
giving PSO gradient signal across the degenerate plateau where all trials
produce the same 1.25 objective.
Wire the penalty into extract_objective() with short-circuit for <10 trades
(skip composite score since it's all zeros anyway).
Include 4 unit tests verifying zero/one/graduated/no-plateau properties.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Add factored evaluation support to EvaluationEngine with continuous
exposure tracking (-1.0 to +1.0), partial position changes, and
order-type-aware transaction costs for the 45-action factored space.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- JWT issuer now foxhunt-api across all 16 files (services, tests, config, docker-compose)
- Remove serde alias api_gateway_url from FxtConfig (no backwards compat)
- Remove api_gateway CLI alias from e2e orchestrator
- All services must deploy simultaneously for JWT validation to match
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
When use_noisy_nets=true (the conservative() default), epsilon never
decayed from 1.0 because (1) the trainer skipped update_epsilon() and
(2) DQNAgentType::set_epsilon() was a no-op for RegimeConditional agents.
This caused ALL training actions to be random — Q-values were learned but
never used for action selection.
Fix: set stored epsilon to 0.0 at training start when noisy nets are on.
Exploration is provided by NoisyLinear weight perturbation + the separate
noisy_epsilon_floor (5% safety floor for 45-action spaces).
Also fixes:
- Zstd-compressed .dbn file detection via magic bytes (0x28B52FFD)
- Test data path resolution using ancestors().find() for workspace root
- Test assertions updated for epsilon < 0.01 with noisy nets
2698 lib tests pass, 0 clippy warnings.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Add live training metrics monitor CLI command (streaming & one-shot) using
the monitoring gRPC service. Update DQN tests to match post-fix defaults:
IQN disabled, CQL alpha=0.1, v_min/v_max widened, 26D search space.
- train.rs: `fxt train monitor [--once] [--model X] [--interval N]`
- Rewrite gradient collapse test for BF16 mixed precision awareness
- Update inference test config to match trainer defaults (IQN off, CQL on)
- Update production smoke test for 26D parameter space
- Add dqn_action_collapse_fix_test.rs verifying all 6 root cause fixes
- Add planning docs for monitoring service and epoch financial metrics
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- Expand search space from 25D to 26D with cql_alpha (0.0-0.5)
- v_min bounds: (-3,-1) → (-15,-3), v_max: (1,3) → (3,15)
- Default use_qr_dqn: true → false (IQN disabled)
- Wire cql_alpha into DQNHyperparameters construction
- Update all 18 adapter tests for new dimensions and ranges
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Hold penalty was divided by 1000 (producing ~0.00001), negligible vs
transaction costs (0.05-0.15%). Now uses hold_penalty_weight directly
from hyperopt (0.01-2.0 range).
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Delete stub rainbow_agent.rs (fake select_action, hardcoded train loss)
and promote rainbow_agent_impl.rs to rainbow_agent.rs. Unify on the
real 8-field RainbowAgentMetrics from rainbow_config.rs, replacing the
4-field stub version (epsilon→exploration_rate, average_loss→current_loss).
5 files changed: -667/+447 lines, 2698+32 tests pass, 0 clippy.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>