Commit Graph

1908 Commits

Author SHA1 Message Date
jgrusewski
1d87fa41bc fix(infra): remove trailing backslash in training workflow rclone commands
Same bug as service manifests — \\ after --s3-no-check-bucket caused
chmod to be parsed as rclone args. Fixed in 6 places.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-07 00:35:27 +01:00
jgrusewski
5dbe529ee6 fix(ci): rebuild CI builder images with mold linker
Both Dockerfiles already install mold v2.35.1 but the registry images
are stale builds without it. This change triggers rebuild-ci-builder
and rebuild-ci-builder-cpu pipeline steps via detect-changes.

.cargo/config.toml uses -fuse-ld=mold — without mold in the image,
linking falls back to the system default (slower).

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-07 00:20:09 +01:00
jgrusewski
b31329931f fix(infra): remove MinIO TLS, fix sccache 0% cache hits, update pool selectors
- Remove all HTTPS/TLS from MinIO (plain HTTP for internal cluster traffic)
- Fix sccache 0% cache hit rate (rustls rejected self-signed MinIO cert)
- Remove hardcoded URLs from k8s_dispatcher.rs (S3_ENDPOINT, TRAINING_RUNTIME_IMAGE,
  CALLBACK_ENDPOINT now required env vars)
- Update GitLab registry S3 credentials to HTTP endpoint
- Fix PVC manifest (20Gi → 100Gi to match cluster)
- Fix nodeSelector: infra/foxhunt → platform (match actual node pool)
- Fix rclone trailing backslash causing chmod to be parsed as rclone args
- Remove minio-ca-cert ConfigMap references from all manifests
- Update trading-service GPU overlay to l40s pool

20 files changed, -118 lines net

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-06 23:53:05 +01:00
jgrusewski
8f35c4f480 fix(ml): fix stale C2 defaults in trainer + clarify comments
- noisy_epsilon_floor fallback: 0.05 → 0.0 (C2: NoisyNet only)
- count_bonus_coefficient fallback: 0.1 → 0.0 (C2: disabled)
- use_count_bonus: true → false (C2: conflicts with NoisyNet)
- Fix 4 stale comments saying "REMOVED" → "FIXED to 0.0"

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-06 23:02:31 +01:00
jgrusewski
9fc1b629c4 feat(ml): M2 PER priority staleness tracking for N-step returns
Replace stub StalenessTracking trait with real StalenessTracker:
- Per-slot timestamp tracking (last_updated_step per experience)
- get_stale_indices() returns oldest-first for refresh
- Wired into PrioritizedReplayBuffer (mark_updated on priority update)
- Epoch-boundary refresh in DQN trainer: recomputes priorities for
  stale entries (>500 steps old) via Q-value forward pass
- 7 staleness tests + 23 PER tests pass, 0 clippy

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-06 22:55:58 +01:00
jgrusewski
b2610d24b2 fix(ml): implement 8 DQN hyperopt research findings (C1-C3, H1-H3, M1, M3)
- C1: Narrow V_min/V_max from [-15,+15] to [-2,+2] for C51 atom resolution
- C2: Remove exploration stacking (28D→26D), keep NoisyNet only
- C3: Fix entropy regularizer for 5-action space, remove 2x/3x multipliers
- H1: Narrow gamma to [0.88, 0.96] (avoid overnight gap discounting)
- H2: Raise tau lower bound to 0.005 (target net tracks in short trials)
- H3: Cap kelly_fractional at 0.75 (no full-Kelly ruin)
- M1: Enable QR-DQN when num_atoms > 100 (hyperopt-toggled)
- M3: Add Q-value gap logging (Q_best - Q_second_best) per epoch

2735 tests pass, 0 clippy warnings

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-06 22:40:46 +01:00
jgrusewski
14c3ca2507 fix(ml): use backtest action distribution in objective, not training counts
The objective function was using buy/sell/hold percentages from training
metrics instead of the backtest. This caused the optimizer to receive stale
action distribution signals that diverged from actual backtest behavior.

Added buy_action_pct/sell_action_pct/hold_action_pct to BacktestMetrics,
counted during the multi-window backtest loop, and used in extract_objective
when backtest metrics are available.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-06 22:10:59 +01:00
jgrusewski
b1f377babd fix(ml): gate patience early stopping behind early_stopping_enabled, reduce search to 28D
Patience-based early stopping (WAVE 24) was not gated by early_stopping_enabled,
causing ALL hyperopt trials to terminate early and return penalty metrics with
backtest_metrics: None — no backtest ever ran, producing FALLBACK OBJECTIVE 44.6.

Also removes eval_softmax_temp from search space (29D→28D) since backtest uses
greedy argmax, widens gamma range (0.95-0.99→0.90-0.999) for better TPE signal,
and updates all tests to match.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-06 21:56:53 +01:00
jgrusewski
aa7cc44c31 Merge feature/action-diversity-fix: remove diversity hard gate
Resolves early_stopping comment conflict (keep main's wording).

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-06 21:26:30 +01:00
jgrusewski
05df216f05 fix(ml): remove diversity hard gate, accept greedy mono-action policies
Greedy argmax legitimately converges to 1-2 actions when the model
is confident. The hard gate (unique_actions < 3 → 8.0 penalty) was
blocking ALL trials from using real backtest metrics, forcing FALLBACK.

Changes:
- Remove hard diversity short-circuit gate (unique_actions < 3)
- Reduce soft diversity penalty from 3.0 to 0.8 max (mild TPE signal)
- Fix early_stopping_enabled: false (was self.epochs > 12)
- Update test to validate mono-action policies produce good objectives

Multi-window backtest (3 windows, mean - 0.5*std) already handles
phantom Sharpe from single-bucket flukes.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-06 21:24:42 +01:00
jgrusewski
faeae5ef01 fix(ml): fix position mask panic + stale 45-action comments after DQN 5-action refactor
apply_position_mask() in factored_q_network.rs looped 0..45 against a
5-wide Q-values tensor — would panic at runtime. Changed to 0..5 using
ExposureLevel::from_index() directly. Updated stale "45 actions" comments
in 5 files (dqn.rs, reward.rs, hyperopt/adapters/dqn.rs, curriculum.rs).

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-06 20:52:55 +01:00
jgrusewski
e0bee7640b Merge remote-tracking branch 'origin/feature/action-diversity-fix'
# Conflicts:
#	crates/ml/src/hyperopt/adapters/dqn.rs
2026-03-06 20:47:08 +01:00
jgrusewski
ba269d7ce7 fix(ml): improve DQN backtest Sharpe — greedy eval, exposure-scaled reward, cost alignment
A: Switch backtest eval from Gumbel softmax to greedy argmax (batch_greedy_actions)
   so hyperopt Sharpe reflects the agent's actual learned policy, not noisy sampling.

C: Disable reward normalization (enable_normalization=false). EMA normalizer with
   ±3.0 clipping was flattening the reward landscape, preventing the agent from
   distinguishing large winners from scratch trades.

D: Wire tx_cost_bps (0.1 bps for IBKR ES) through to EvaluationEngine via
   new_with_fee_rate(). Previously hardcoded at 15 bps (150x mismatch with actual
   commission costs), massively penalizing every trade in backtest.

E: Scale PnL reward by agent's target exposure in calculate_pnl_reward().
   Previously, a Short100 action received POSITIVE reward when market went up
   (pct_return ignored position direction). Now: reward = pct_return × exposure.

2735 tests pass, 0 clippy warnings.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-06 20:25:42 +01:00
jgrusewski
1950ba009d fix(ml): disable early stopping for short hyperopt trials
Early stopping with 8-epoch trials returns penalty metrics (no backtest),
causing ALL trials to hit FALLBACK OBJECTIVE = 44.6 regardless of actual
model quality. The Sharpe was 1.36-1.61 but natural fluctuation triggered
"Sharpe worsening" at epoch 5, killing the backtest evaluation.

Fix: disable early stopping when epochs ≤ 12. With ~90s per trial, running
all 8 epochs is cheap. Long training runs (50 epochs) still use it.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-06 18:49:01 +01:00
jgrusewski
e412f12e33 fix(ml): restore exploration mechanisms for DQN action diversity
The C2 fix from the previous session was too aggressive — it eliminated
ALL exploration mechanisms simultaneously:

1. epsilon forced to 0.0 when noisy nets active (select_action)
2. count bonus removed from Q-value computation (metrics only)
3. noisy_epsilon_floor config field declared but never read

This left noisy nets as the sole exploration mechanism, which produces
perturbations too small to overcome Q-value gaps (A4=0.12 vs others≈0.02).
Result: 20/20 hyperopt trials hit fallback objective with 1/5 diversity.

Fixes:
- select_action: use noisy_epsilon_floor (not 0.0) as effective_epsilon
  when noisy nets active — guarantees minimum random action rate
- select_action: re-enable UCB count bonus on Q-values before argmax
  (both IQN and standard paths) for directed exploration
- select_action_with_confidence: same fixes for consistency
- trainer: set epsilon to noisy_epsilon_floor (not 0.0) at init
- hyperopt: widen noisy_epsilon_floor range from [0.0, 0.05] to
  [0.03, 0.15] with default 0.05

Exploration now has two complementary mechanisms:
- noisy_epsilon_floor: random actions feed diverse replay buffer
- count bonus: UCB term biases greedy selection toward under-visited actions
- noisy nets: weight perturbation adds stochasticity to Q-values

select_action_inference (production) is unchanged — pure exploitation.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-06 18:27:04 +01:00
jgrusewski
03062f2401 fix(ml): use exposure index (0-4) in DQN experience storage and tracking
FactoredAction::to_index() returns 0-44 (exposure×9 + order×3 + urgency),
but DQN Q-network has 5 outputs. Storing the factored index in Experience
caused train_step() to reject valid actions (e.g. Flat→index 19 > num_actions=5).

Fix: use action.exposure as usize (0-4) for:
- Experience storage (training loop + validation adapter)
- Count bonus tracking (diversity metrics)
- Safety action counts (diversity monitor)
- Debug logging (action indices)
- Test assertions (< 45 → < 5)

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-06 16:27:17 +01:00
jgrusewski
09710e590f fix(ml): audit — purge all FactoredAction::from_index from DQN paths
Critical bug: all 3 DQN action selection methods (select_action,
select_action_with_confidence, select_action_inference) used
FactoredAction::from_index() which maps indices 0-4 to exposure_idx=0
(Short100) via division by 9. This is the root cause of action
diversity collapse during both training and production inference.

Fix: ExposureLevel::from_index() + OrderRouter::route_default() in all
DQN paths. Also fixes hyperopt objective thresholds (<10 → <3 for
5-action degenerate detection), stale defaults/comments, integration
test configs.

Files: dqn.rs (3 methods), trainer.rs (validation + select_action),
hyperopt/adapters/dqn.rs (thresholds), dqn_model.rs (comments),
train_baseline_rl.rs (default), reward.rs (comment),
dqn_integration.rs + ensemble_integration.rs (num_actions).

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-06 16:20:58 +01:00
jgrusewski
5ace5dd24f fix(ml): DQN hyperopt overhaul — C2 triple exploration + eval 5-action compat
Critical fixes:
- hyperopt backtest: ExposureLevel::from_index() replaces FactoredAction::from_index()
  which mapped all DQN indices 0-4 to Short100 (every trial ran all-short)
- evaluate_baseline: same fix + DQN num_actions default 5, PPO hardcoded 45
- simulate_chunk_trades: is_dqn dispatch for DQN vs PPO action decoding

C2 triple exploration stacking:
- Removed count bonus UCB from Q-value computation in both batch paths
  (noisy nets are sole exploration mechanism)
- Narrowed noisy_epsilon_floor from [0.02, 0.10] to [0.0, 0.05], default 0.0
- Removed count_bonus_coefficient from search space (30D → 29D)
- Count bonus module kept for diversity metrics tracking only

Smoke tests: 6 new tests verifying 5-action space, 29D search space,
epsilon floor defaults, count bonus coefficient fixed at 0.1

2735 tests passed, 0 clippy warnings.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-06 15:29:02 +01:00
jgrusewski
e3c3cf0fa3 docs: add DQN hyperopt overhaul implementation plan
7 tasks covering: 5-action compatibility in eval/backtest paths,
C2 triple exploration stacking fix (remove count bonus from Q-values,
narrow epsilon floor, reduce search space 30D→29D), PPO isolation
verification.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-06 14:32:37 +01:00
jgrusewski
d1601b3720 fix(ml): complete Phase B3 + agent single-action path
- agent.rs: select_action_factored() now uses ExposureLevel::from_index()
  + OrderRouter::route_default() instead of FactoredAction::from_index().
  Previously, indices 0-4 mapped to all-Short100 variants in the 45-action
  space — now correctly maps to 5 distinct exposure levels.
- hyperopt: plateau_window .max(3) → .max(5) to prevent premature early
  stopping with short trial epochs.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-06 14:15:14 +01:00
jgrusewski
f6727f1103 fix(ml): reduce DQN action space from 45 to 5 exposure levels + OrderRouter
Root cause: 45 factored actions (5 exposure × 3 order × 3 urgency) caused
reward degeneracy — 9 actions per exposure level produced nearly identical
rewards since order type/urgency had 1000-4000x weaker signal than PnL.
This collapsed action diversity as DQN couldn't differentiate actions.

Changes:
- DQN now outputs 5 Q-values (Short100, Short50, Flat, Long50, Long100)
- New OrderRouter deterministically maps exposure → (order_type, urgency)
  based on spread and volatility microstructure signals
- PPO retains full 45-action space (separate CUDA constants DQN_NUM_ACTIONS
  vs PPO_NUM_ACTIONS)
- CUDA kernels: DQN diversity entropy uses 5 categories, PPO keeps 45
- Phase B: pnl_history cleared per epoch so Sharpe reflects current epoch
  (was accumulating across all epochs, causing frozen Sharpe metric)

24 files, 2728 tests pass, 0 clippy warnings

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-06 14:07:42 +01:00
jgrusewski
9a10e82fa7 fix(ml): disable val-loss early stopping in hyperopt, fix penalty metrics
Sharpe-based early stopping kills every hyperopt trial at epoch 4
because compute_epoch_financials() is deterministic (greedy argmax on
fixed validation data) — the model doesn't change enough in 8 short
epochs to shift any argmax decisions, making Sharpe bit-identical
across epochs and triggering plateau detection immediately.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-06 13:10:27 +01:00
jgrusewski
4e0d1fcbe6 fix(ci): add imagePullPolicy: IfNotPresent to training workflow
Kubernetes defaults to Always for :latest tags, forcing registry
round-trips that fail on fresh GPU nodes where containerd HTTP-only
registry config has a race condition with HTTPS fallback.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-06 12:49:34 +01:00
jgrusewski
2f8fa1ab19 Merge feature/action-diversity-fix: DQN hyperopt overhaul (B1-B3, C1-C4)
7 root cause fixes for DQN hyperopt train/eval mismatch and reward corruption:
- B1: Eval mode (noisy noise disabled, softmax action selection)
- B3: Per-bar portfolio state sync in eval
- C1: Extrinsic-only replay buffer (curiosity removed from rewards)
- C2: Single exploration (noisy nets only, no epsilon/count bonus on Q-values)
- C3: Neutral hold reward, search space 31D to 30D
- C4: Sharpe-based early stopping (replaces val-loss plateau)

2720 tests, 0 failures, 0 clippy warnings.
2026-03-06 11:44:39 +01:00
jgrusewski
ce847fd0d4 docs: DQN hyperopt overhaul design — 7 remaining root cause fixes
Addresses eval/training mismatch (B1-B3), reward architecture (C1-C3),
and early stopping (C4). See design doc for full analysis.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-06 11:44:02 +01:00
jgrusewski
3ae5a295f1 fix(ml): C4 complete — Sharpe-based early stopping replaces val-loss
The previous C4 fix re-enabled early stopping with adaptive plateau
window but still used val-loss as the stopping metric. Val-loss (TD
Bellman residual) can plateau while trading strategy still improves.

Now:
- Best-checkpoint saved when epoch Sharpe improves (not val-loss)
- Plateau detection checks sharpe_history (not val_loss_history)
- Patience-based EarlyStopping receives -Sharpe (negate for lower=better API)
- Per-epoch Sharpe extracted from compute_epoch_financials() (already computed)

This ensures early stopping and best-model selection track the metric
that actually matters for hyperopt: trading performance.

2720 tests pass, 0 clippy warnings.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-06 11:17:44 +01:00
jgrusewski
86b950df10 fix(ml): DQN hyperopt overhaul — 7 root cause fixes (B1-B3, C1-C4)
B1: Eval mode — disable noisy layer noise, use softmax action selection
    (was greedy argmax, causing train/eval policy mismatch)
B3: Per-bar portfolio state sync in eval (was frozen within 1024-bar chunks)
C1: Extrinsic-only replay buffer — curiosity reward no longer stored
    (was corrupting Q-values to learn novelty instead of trading P&L)
C2: Single exploration mechanism — noisy nets only. Removed count bonus
    from Q-values and epsilon floor (was triple-stacking exploration)
C3: Neutral hold reward (0.0) — removed hold_penalty_weight from 31D→30D
    search space (was biasing Q-values toward excessive trading)
C4: Re-enabled early stopping with adaptive plateau_window = epochs/2

2720 tests pass, 0 failures, 0 clippy warnings.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-06 10:23:51 +01:00
Administrator
4658d7e914 Merge branch 'feature/action-diversity-fix' into 'main'
feat(ml): multi-window backtest + top-K ensemble training

See merge request root/foxhunt!2
2026-03-06 07:48:06 +00:00
jgrusewski
84366d8dd8 fix(infra): pin CI sensor pod to DEV1-L platform nodes
The Argo Events sensor had no nodeSelector and randomly landed on the
H100 GPU node, preventing autoscale-down. Pin it to DEV1-L to avoid
wasting expensive GPU node hours on a lightweight event listener.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-06 08:45:59 +01:00
jgrusewski
5a9fa1534b feat(ml): multi-window backtest objective + top-K ensemble training
Multi-window backtest: splits validation data into 3 non-overlapping
windows and aggregates with mean(Sharpe) - 0.5*std(Sharpe), penalizing
inconsistency and reducing overfit to a single data segment.

Top-K ensemble: hyperopt now emits top_k_params (top 5 trials) in JSON
output. train_baseline_rl gains --ensemble-top-k flag to train multiple
models per fold from different hyperopt configs, saving checkpoints as
dqn_ensemble_{k}_fold_{n}.safetensors.

Workflow template: adds ensemble-top-k parameter (default 5) and passes
--ensemble-top-k to the train-best step.

2720 tests pass, 0 clippy warnings.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-06 08:42:03 +01:00
jgrusewski
2b77a7fbe1 fix(dashboard): clamp degenerate outliers in training metrics panels
Add clamp_max() to PromQL queries and hard Y-axis limits to prevent
early-epoch degenerate values from blowing up panel scaling (Sharpe
showing 12 instead of 1.3, Profit Factor at 30k).

Run Summary: clamp_max on Best Val Loss (5), Sharpe (5), Profit Factor (20)
Training Quality: clamp_max on Epoch Sharpe (5), Sortino (10), PF (20)

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-06 02:22:21 +01:00
jgrusewski
5323cf11e7 fix(dashboard): convert Run Summary to smooth timeseries, 3x2 layout
Replace stat panels with timeseries using smooth line interpolation,
gradient fill, and multi-tooltip. Layout changed from 6x1 to 3x2 grid.
Legends show model/fold only when multiple series exist.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-06 01:37:31 +01:00
Administrator
3cca2494f3 Merge branch 'feature/action-diversity-fix' into 'main'
fix(ml): resolve DQN action diversity collapse in hyperopt

See merge request root/foxhunt!1
2026-03-06 00:29:30 +00:00
jgrusewski
9067402284 feat(dashboard): add Run Summary row to training dashboard
Add 6 stat panels showing overall training run metrics: Best Val Loss,
Best Sharpe, Best Win Rate, Min Max Drawdown, Total Epochs, and Best
Profit Factor. Placed between Training Status and Training Curves rows.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-06 01:28:29 +01:00
jgrusewski
98694a0ca2 fix(ml): resolve DQN action diversity collapse in hyperopt
4 root causes of 45-action DQN collapsing to 1-6 actions:

1. Batch epsilon ignoring noisy_epsilon_floor: select_actions_batch()
   and select_actions_batch_gpu() used get_epsilon() which returns 0.0
   with noisy nets — zero random exploration in the training path.
   Added get_effective_epsilon() that respects noisy_epsilon_floor.

2. Entropy coefficient too weak: default 0.05 with bounds (0.01, 0.2)
   produced max ~0.19 bonus vs TD loss of 4+. Bumped default to 0.1,
   widened bounds to (0.05, 0.5) for effective anti-collapse.

3. count_bonus_coefficient not in search space: was hardcoded at 0.1
   in from_continuous(). Promoted to 31st search dimension with bounds
   (0.05, 1.0) so PSO/TPE can optimize exploration strength.

4. Diversity penalty too coarse: objective had <10 unique actions
   short-circuit but nothing for 10-20. Added graduated penalty that
   linearly ramps from 3.0 (10 actions) to 0.0 (20 actions).

Also fixes pre-existing clippy impl_trait_in_params in optimizer.rs.

2720 tests pass, 0 clippy warnings.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-06 01:27:54 +01:00
jgrusewski
371598ecb5 fix(metrics): emit per-trial hyperopt metrics for Grafana panels
Wire record_hyperopt_trial_duration, set_hyperopt_best_objective,
set_hyperopt_trial_best_loss, and set_hyperopt_elapsed into all three
optimizer paths (PSO sequential, PSO parallel, TPE). These metrics were
registered but never called during the optimization loop, causing
"Best Objective Over Time", "Trial Duration", and "Elapsed Time"
Grafana panels to show "No data".

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-06 00:57:48 +01:00
jgrusewski
06a875e6fc fix(metrics): push training metrics to pushgateway before pod exit
Ephemeral Argo workflow pods terminate after training completes, causing
Prometheus to lose all scraped metrics. Add push_to_gateway() to POST
final metrics to the existing pushgateway service so they persist on the
Grafana training dashboard after pod completion.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-06 00:32:09 +01:00
jgrusewski
b6f6d9c7c9 fix(ml): remove warmup_steps bug that prevented DQN training
The batch training path (select_actions_batch → forward()) does NOT
increment DQN::total_steps. Only select_action() does. When
warmup_steps > 0, train_step() checks total_steps < warmup_steps
and returns (0.0, 0.0) — zero loss, zero gradients. The model
never trained; "results" were random initialization Q-values.

Fix: force warmup_steps=0 in hyperopt adapter and remove
warmup_ratio from the 31D→30D search space (saves a dimension
for TPE/PSO effectiveness).

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-06 00:29:42 +01:00
jgrusewski
351bdaf8f1 feat(infra): add GPU warmup step to training pipeline
Runs nvidia-smi on GPU node in parallel with fetch-binary, triggering
H100 autoscale during compilation so the node is ready when hyperopt
starts. Exits immediately to free GPU resources.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-06 00:07:41 +01:00
jgrusewski
f2938b19e8 fix(ml): improve TPE exploitation with Scott bandwidth, best-trial injection
- Replace Silverman's bandwidth (h = 1.06σn^(-1/5)) with Scott's rule
  (h = 0.7σn^(-1/(d+4))) for tighter kernels in high-D parameter spaces
- Add best-trial injection: always evaluate EI at best known point plus
  5 small perturbations (±5%), preventing optimizer from forgetting peaks
- Scale n_candidates dynamically: max(256, 8*n_dims) instead of fixed 100
- Reduce gamma from 0.25 to 0.15 when trials < 50 for tighter exploitation
- Wire model_name through PSO/TPE paths for per-trial Prometheus metrics

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-05 23:54:30 +01:00
jgrusewski
08a927639e fix(ci): POSIX-compatible kubectl detection in deploy step
The foxhunt-runtime image runs as non-root user 'foxhunt' and uses
/bin/sh (dash), not bash. Two issues:
1. &>/dev/null is bash-only — use >/dev/null 2>&1 for POSIX sh
2. Fallback download to /tmp (writable), not /usr/local/bin (root-only)

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-05 23:26:35 +01:00
jgrusewski
cc4e0c5a2d refactor: remove trading_engine dep from ml and risk crates
Re-export HardwareTimestamp through data crate instead of ml/risk
depending directly on trading_engine. Reduces coupling between
the ML pipeline and the trading engine.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-05 23:19:38 +01:00
jgrusewski
7b60fd5f86 fix(ci): Docker build uses init container + emptyDir instead of PVC
The build-ci-image template used volumeClaimTemplates (PVC) which
weren't inherited when called via templateRef from ci-pipeline.

Restructured to single pod: init container (alpine/git) clones repo
into emptyDir, main container (kaniko) builds and pushes image from
the same volume. No PVC needed, works correctly via templateRef.

Tested: foxhunt-runtime image rebuilt successfully with kubectl.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-05 23:11:55 +01:00
jgrusewski
9cd4da9acf fix(ci): restructure Docker build to single pod with emptyDir
build-ci-image used volumeClaimTemplates (PVC) shared between a
git-clone pod and a kaniko-build pod in a DAG. When called via
templateRef from ci-pipeline, the VCT wasn't inherited, causing
"volume 'workspace' not found" errors.

Fixed by merging into a single pod: init container (git clone) +
main container (kaniko build) sharing an emptyDir volume. No PVC
needed, works correctly when called via templateRef.

Also removed stale compile-training-template.yaml reference from
kustomization (file doesn't exist, template is inline in ci-pipeline).

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-05 22:58:42 +01:00
jgrusewski
171fe86194 fix(ci): bake kubectl into runtime image for faster deployments
Add kubectl v1.31.4 to foxhunt-runtime Dockerfile so deploy step
doesn't need to download it each run. Deploy step falls back to
curl download if kubectl not found (for current image version).

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-05 22:49:51 +01:00
jgrusewski
4e71cbf549 fix(ci): use runtime image for deploy step, add podGC to sensor
- Deploy step: use foxhunt-runtime image (cached on platform node) +
  curl kubectl, replacing bitnami/kubectl:1.31 which doesn't exist.
- Sensor: add podGC:OnPodCompletion and ttlStrategy to workflow spec
  (not inherited from WorkflowTemplate via workflowTemplateRef).

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-05 22:46:11 +01:00
jgrusewski
e9481c6107 fix(ml): disable val-loss early stopping in hyperopt, fix penalty metrics
Two issues causing all 20 hyperopt trials to have identical f64::MAX
objective (TPE optimizer blind):

1. Val-loss plateau early stopping fired at epoch 5-6 of every 8-epoch
   trial (plateau_window=5 too aggressive for short runs). Disabled
   early_stopping_enabled for hyperopt; gradient-collapse patience
   still active as safety net.

2. Penalty metrics used f64::MAX for gradient_norm/q_value_std which
   produced ~3.6e+308 objective. Changed to 100.0 so TPE can still
   differentiate between early-stopped trials by other metric fields.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-05 22:45:03 +01:00
jgrusewski
a77c873da4 fix(ci): add deploy step, podGC, move compile-services to ci-compile-cpu
Three issues prevented services from being deployed after push:

1. No deploy step: pipeline compiled + uploaded binaries to MinIO but
   never restarted service deployments. Added deploy-services step
   (bitnami/kubectl) that runs rollout restart after compile-services.

2. compile-services couldn't schedule: targeted platform node (4 CPU,
   6Gi allocatable) but requested 6Gi memory. Moved to ci-compile-cpu
   (POP2-32C-128G) alongside compile-training — both fit simultaneously
   (28/32 CPU, 48/128Gi).

3. No podGC: completed Argo pods held resources indefinitely, blocking
   subsequent pipeline steps. Added podGC:OnPodCompletion.

Also: detect-changes now triggers service rebuild when CI pipeline
template or k8s service manifests change.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-05 22:34:44 +01:00
jgrusewski
5d5cce72d2 fix(infra): auto-restart containerd on GPU nodes for registry config
DaemonSet init container now:
- Detects GPU nodes (99-nvidia.toml in conf.d)
- Creates v3-compatible registry config drop-in
- Restarts containerd via chroot if config files changed
- Idempotent: skips restart if files already exist

Requires hostPID + privileged for chroot /proc/1/root.
On new nodes, first run creates files + restarts containerd
(kills pod), second run finds files unchanged (no restart).

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-05 22:12:17 +01:00
jgrusewski
248b2fec70 fix(ml,infra): handle early stopping in hyperopt, GPU containerd registry
- Hyperopt: catch early stopping errors and return penalty metrics
  instead of aborting entire run. Trials that stop early are scored
  as poor (objective=-500) so optimizer avoids those configs.
- DaemonSet: detect GPU nodes (NVIDIA conf.d overlay) and create
  v3-compatible registry config drop-in for containerd v2.1+.
  Old grpc.v1.cri.registry path is silently ignored by containerd v2.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-05 22:07:09 +01:00