Eliminates the f64→f32 cudarc ABI trap (feedback_cudarc_f64_f32_abi.md,
task #82) at the type level: hyperparameters consumed by CUDA kernels
now live as f32 in Rust, cast once at the TOML/PSO ingest boundary
instead of at every kernel call site.
Structs changed:
- DQNHyperparameters (crates/ml/src/trainers/dqn/config.rs) —
~85 scalar fields migrated from f64 → f32. Covers all
kernel-facing scalars: reward weights (w_pnl/w_dd/w_idle,
dd_threshold, cea_weight, micro_reward_*, price_confirm_weight,
book_aggression_weight, hold_quality_weight), exploration
(epsilon_* and the 4 branch mults, noisy_sigma_*, count_bonus,
noise_sigma, q_gap_threshold), distributional RL (v_min, v_max,
reward_scale, iqn_lambda, qr_kappa, spectral_*,
gradient_collapse_multiplier), fill simulation (5 fill_*
fields), risk/Kelly (kelly_fractional, kelly_max_fraction,
max_leverage, max_position_absolute, minimum_profit_factor),
ensemble/curiosity (curiosity_weight,
curiosity_q_penalty_lambda, ensemble_*, beta_*, variance_cap),
anti-LR (anti_lr_*, adversarial_dd_threshold,
beta_penalty_strength), walk-forward (wf_*),
experience (avg_spread, transaction_cost_multiplier,
holding_cost_rate, churn_penalty_scale, contract_multiplier,
margin_pct, tick_size, bars_per_day, cash_reserve_percent),
misc kernel scalars (gamma, tau, huber_delta, q_clip_*,
shrink_perturb_*, regime_replay_decay, per_alpha,
per_beta_start, dt_target_return, etc.). Also
`noisy_epsilon_floor: Option<f32>` and
`count_bonus_coefficient: Option<f32>`.
- `computed_v_min` / `computed_v_max` now return f32.
- `compute_max_position` returns f32 (f64 internally for the
notional division).
Fields preserved as f64 (precision-sensitive, NOT kernel-facing
scalars — per task spec and feedback_cudarc_f64_f32_abi.md):
- `learning_rate` — tested at 1e-10 tolerance; f32 rounds to
2e-12 for a 1e-5 LR.
- `entropy_coefficient` — tested at 1e-9 tolerance.
- `weight_decay` — tiny 1e-5..1e-3 range.
- `adam_epsilon` — 1e-8 default; f32 preserves denorms here but
paired with weight_decay/learning_rate for symmetry.
- `gradient_clip_norm: Option<f64>` — grad norms are f64
accumulators by project convention.
- `min_loss_improvement_pct`, `q_value_floor` — early-stopping
long-horizon stats.
- `cql_alpha` — flows into DQNConfig (ml-dqn) still f64.
- `min_learning_rate`, `lr_min` — paired with learning_rate.
- Family intensity scalars (6 `*_intensity` fields) — f64 PSO
search space; the intensity applies via `as f32` at each call
site in `apply_family_scaling`.
No checkpoint format change: DQNHyperparameters has
`#[derive(Debug, Clone)]` only (not Serialize/Deserialize), so the
TOML-ingest path is the only serde boundary and already casts
explicitly via `hp.field = v as f32;` in
`DqnTrainingProfile::apply_to`. PSO hyperopt bounds stay f64 in
`SearchSpaceSection` and cast at the adapter boundary.
Call-site impact:
- ~30 `as f32` casts removed from hot paths (fused_training
FusedConfig builder, training_loop kernel launches,
constructor DQNConfig builder, action.rs GPU action selector,
trainer/mod.rs WF config). Kernels now receive the hp field
directly via `&hp.x`.
- ~75 `as f32` casts added at the `apply_to` / hyperopt-adapter
ingest boundary — the single conversion point.
- Cross-crate contracts (DQNConfig in ml-dqn, PortfolioTracker
in ml-core, GAECalculator, DropoutScheduler, NoisySigmaScheduler,
KellyOptimizerConfig, RewardConfig) retain their f64 signatures;
ml calls cast at the boundary with `f64::from(hp.x)` so the
contract is explicit and greppable.
Two test-side adjustments:
- `test_kelly_fields_are_public` now asserts `f32` for
kelly_fractional / kelly_max_fraction (these migrated).
- `test_early_stopping_termination` uses `f64::from(...)` to
preserve the f64 threshold computation against the now-f32
`gradient_collapse_multiplier`.
Verified:
- `SQLX_OFFLINE=true CARGO_INCREMENTAL=0 RUSTC_WRAPPER=sccache
cargo check --workspace` — clean, zero new warnings.
- `cargo check --workspace --tests` — clean.
- `cargo test -p ml --lib training_profile::tests` — all 18
migrated tests pass. The one pre-existing failure
(`test_production_profile_applies_all_sections` n_steps
mismatch, 1 vs 5) reproduces on stashed baseline, so it is
unrelated to this change.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>