Files
foxhunt/docs/superpowers/plans/2026-03-25-config-unification.md

13 KiB
Raw Blame History

Config Unification Plan — Single Source of Truth

For agentic workers: REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (- [ ]) syntax for tracking.

Goal: Eliminate all hardcoded config values and use_ boolean bloat. Every training parameter flows from TOML → DQNHyperparameters → downstream configs. No Default impl should diverge from the TOML-driven path.

Architecture: Single config flow: TOML file → DqnTrainingProfile::load() → DQNHyperparameters → GpuDqnTrainConfig + GpuExperienceCollectorConfig + DQNConfig. Default impls match conservative() exactly. All use_ booleans that are always true are removed — the features are always on.

Tech Stack: Rust config structs, TOML deserialization, training_profile.rs.


Problem Statement

Three config structs have independent Default impls with conflicting values:

Parameter DQNConfig GpuDqnTrainConfig ExperienceCollectorConfig
state_dim 48 72
v_min/v_max -25/25 -2/2 -2/2
learning_rate 1e-4 3e-4
weight_decay 1e-4 1e-5
batch_size 64 256
num_atoms 51 51 1
use_noisy_nets true false
use_distributional true false

Plus 26 ExperienceCollectorConfig parameters are hardcoded with no TOML exposure.

Design Principles

  1. TOML is the source of truth — every tunable value has a TOML field
  2. conservative() reads TOML — no hardcoded literals in the function
  3. Default impls match conservative() — tests behave like production
  4. No use_ booleans for always-on features — if it's always true, remove the flag
  5. One config flow — TOML → DQNHyperparameters → everything else

Task 1: Clean up use_ booleans (config.rs + dqn.rs)

Files:

  • Modify: crates/ml/src/trainers/dqn/config.rs — DQNHyperparameters
  • Modify: crates/ml-dqn/src/dqn.rs — DQNConfig

Always-on features (remove the use_ field, hardwire to true internally):

  • use_double_dqn — always true since Wave 6.4
  • use_dueling — always true (Rainbow DQN standard)
  • use_per — always true (memory says "PER always enabled, non-PER paths are dead code")
  • use_branching — always true (9×3×3 action space)
  • use_distributional — always true (C51 is the primary loss)
  • use_noisy_nets — always true (replaces epsilon-greedy)
  • use_huber_loss — always true (more robust than MSE)
  • use_cql — always true (activated this session)

For each:

  • Step 1: Remove the field from DQNHyperparameters
  • Step 2: Remove the field from DQNConfig
  • Step 3: In constructors that read these fields, replace if self.use_X { ... } with just { ... }
  • Step 4: In training_profile.rs, remove the TOML field if it exists (or ignore it)
  • Step 5: Search for ALL references with grep -rn "use_double_dqn\|use_dueling\|use_per\b\|use_branching\|use_distributional\|use_noisy_nets\|use_huber_loss\|use_cql\b" and update each site
  • Step 6: Build + test

Keep as legitimate config flags:

  • use_cvar_action_selection — experimental, toggleable

  • enable_kelly_sizing — can be disabled for ablation

  • enable_circuit_breaker — safety feature, must be toggleable

  • enable_action_masking — experimental

  • enable_regime_qnetwork — experimental

  • enable_early_stopping — must be toggleable

  • Step 7: Commit: "refactor: remove always-on use_ booleans — 8 features are mandatory"


Task 2: Unify Default impls to match conservative()

Files:

  • Modify: crates/ml-dqn/src/dqn.rs — DQNConfig::default()
  • Modify: crates/ml/src/cuda_pipeline/gpu_dqn_trainer.rs — GpuDqnTrainConfig::default()
  • Modify: crates/ml/src/cuda_pipeline/gpu_experience_collector.rs — GpuExperienceCollectorConfig default

Every Default impl must produce values identical to what conservative() → constructor flow would produce. This ensures tests behave like production.

  • Step 1: In DQNConfig::default():

    • v_min: -50.0, v_max: 50.0 (match conservative/reward v6)
    • learning_rate: 3e-5 (match conservative)
    • weight_decay: 1e-4 (match conservative)
    • batch_size: 64 (match smoketest, conservative)
    • num_atoms: 51
    • state_dim: 48 (match actual feature dim with padding)
    • num_actions: 9 (match 9-action exposure space)
  • Step 2: In GpuDqnTrainConfig::default():

    • state_dim: 72state_dim: 48 (or remove default entirely — always set by constructor)
    • v_min: -50.0, v_max: 50.0
    • lr: 3e-5 → match conservative
    • weight_decay: 1e-4
    • batch_size: 64
    • branch_0_size: 9 (correct — 9 exposure levels)
  • Step 3: In GpuExperienceCollectorConfig default:

    • use_noisy_nets: true (was false — mismatch)
    • use_distributional: true (was false — mismatch)
    • num_atoms: 51 (was 1 — mismatch)
    • v_min: -50.0, v_max: 50.0
    • q_clip_min: -200.0, q_clip_max: 200.0 (match reward v6 scale)
  • Step 4: Build + all tests pass

  • Step 5: Commit: "fix: unify Default impls to match conservative() — no more config mismatches"


Task 3: Expose ExperienceCollectorConfig params in TOML

Files:

  • Modify: crates/ml/src/training_profile.rs — add ExperienceSection
  • Modify: config/training/dqn-smoketest.toml — add [experience] section
  • Create: config/training/dqn-production.toml — add [experience] section (if not exists)
  • Modify: crates/ml/src/trainers/dqn/trainer/training_loop.rs — read from profile

26 parameters need TOML exposure. Group them:

[experience]
gpu_n_episodes = 32              # episodes per experience collection round
gpu_timesteps_per_episode = 100  # bars per episode
initial_capital = 100000.0       # starting equity

[experience.fill_simulation]
ioc_fill_prob = 0.85
limit_fill_min = 0.30
limit_fill_max = 0.80
spread_cost_frac = 0.50
spread_capture_frac = 0.50

[experience.barriers]
profit_mult = 1.02
loss_mult = 0.98
max_bars = 500.0
scale = 0.5

[risk]
max_position = 2.0
loss_aversion = 1.5
q_clip_min = -200.0
q_clip_max = 200.0
  • Step 1: Add ExperienceSection, FillSimulationSection, BarrierSection, RiskSection to training_profile.rs
  • Step 2: Implement apply_to() for each section
  • Step 3: Update dqn-smoketest.toml with [experience] section
  • Step 4: In training_loop.rs, pass TOML values to ExperienceCollectorConfig instead of hardcoded defaults
  • Step 5: Build + test
  • Step 6: Commit: "feat: expose 26 experience collector params in TOML training profiles"

Task 4: Wire v_range computation from reward scale

Files:

  • Modify: crates/ml/src/trainers/dqn/config.rs — compute v_min/v_max from reward_scale + gamma
  • Modify: crates/ml/src/hyperopt/adapters/dqn.rs — same computation in from_continuous()

v_min/v_max should NEVER be hardcoded. They should be computed:

let reward_scale = 10.0; // reward v6 REWARD_SCALE
let v_range = (reward_scale / (1.0 - gamma) * 1.2).clamp(20.0, 300.0);
let v_min = -v_range;
let v_max = v_range;
  • Step 1: Add reward_scale: f64 to DQNHyperparameters (default 10.0)
  • Step 2: In conservative(), compute v_min/v_max from reward_scale + gamma
  • Step 3: Remove hardcoded v_min/v_max from conservative()
  • Step 4: In hyperopt adapter from_continuous(), use same formula
  • Step 5: Add reward_scale to TOML and hyperopt search space
  • Step 6: Build + test
  • Step 7: Commit: "feat: v_range computed from reward_scale + gamma — no more hardcoded C51 bounds"

Task 5: Consolidate DQNConfig → derive from DQNHyperparameters

Files:

  • Modify: crates/ml-dqn/src/dqn.rs — DQNConfig construction
  • Modify: crates/ml/src/trainers/dqn/trainer/constructor.rs — DQNConfig from hyperparams

DQNConfig should NOT have its own defaults. It should be constructed FROM DQNHyperparameters:

impl From<&DQNHyperparameters> for DQNConfig {
    fn from(hp: &DQNHyperparameters) -> Self {
        Self {
            state_dim: hp.state_dim(),
            num_actions: hp.num_actions(),
            learning_rate: hp.learning_rate,
            gamma: hp.gamma,
            batch_size: hp.batch_size,
            v_min: hp.v_min(),
            v_max: hp.v_max(),
            num_atoms: hp.num_atoms,
            // ... all fields derived
        }
    }
}
  • Step 1: Add From<&DQNHyperparameters> for DQNConfig
  • Step 2: Use this conversion in constructor.rs instead of field-by-field copy
  • Step 3: Remove DQNConfig::default() (or make it delegate to From<conservative()>)
  • Step 4: Build + test
  • Step 5: Commit: "refactor: DQNConfig derived from DQNHyperparameters — single source of truth"

Task 6: Consolidate GpuDqnTrainConfig → derive from DQNHyperparameters

Files:

  • Modify: crates/ml/src/cuda_pipeline/gpu_dqn_trainer.rs
  • Modify: crates/ml/src/trainers/dqn/fused_training.rs

Same pattern: GpuDqnTrainConfig should be constructed FROM DQNHyperparameters, not hardcoded:

impl GpuDqnTrainConfig {
    pub fn from_hyperparams(hp: &DQNHyperparameters, dqn_config: &DQNConfig) -> Self {
        Self {
            state_dim: dqn_config.state_dim,
            num_atoms: hp.num_atoms,
            v_min: hp.v_min() as f32,
            v_max: hp.v_max() as f32,
            lr: hp.learning_rate as f32,
            weight_decay: hp.weight_decay as f32,
            // ... all derived
        }
    }
}
  • Step 1: Add from_hyperparams() constructor
  • Step 2: Use it in fused_training.rs instead of field-by-field construction
  • Step 3: Remove or deprecate GpuDqnTrainConfig::default()
  • Step 4: Build + test
  • Step 5: Commit: "refactor: GpuDqnTrainConfig derived from hyperparams — no independent defaults"

Task 7: Update TOML files with all parameters

Files:

  • Modify: config/training/dqn-smoketest.toml
  • Modify: config/training/dqn-hyperopt.toml
  • Create or modify: config/training/dqn-production.toml

Each TOML file should have ALL sections:

[training]
epochs = 50
batch_size = 64
learning_rate = 0.00003
gamma = 0.95
weight_decay = 0.0001
reward_scale = 10.0

[distributional]
num_atoms = 51
# v_min/v_max computed from reward_scale + gamma

[exploration]
epsilon_start = 1.0
epsilon_end = 0.02
epsilon_decay = 0.95
noisy_sigma_init = 0.5
entropy_coefficient = 0.001
count_bonus_coefficient = 0.1
q_gap_threshold = 0.05

[experience]
gpu_n_episodes = 32
gpu_timesteps_per_episode = 100
initial_capital = 100000.0

[risk]
max_position = 2.0
loss_aversion = 1.5
kelly_fractional = 0.5
kelly_max_fraction = 0.25

[advanced]
n_steps = 3
tau = 0.005
c51_warmup_epochs = 5
her_ratio = 0.2
cql_alpha = 0.1
iqn_lambda = 0.25
spectral_norm_sigma_max = 3.0
gradient_clip_norm = 10.0

[early_stopping]
enabled = false
patience = 10
  • Step 1: Create comprehensive dqn-smoketest.toml
  • Step 2: Create comprehensive dqn-production.toml
  • Step 3: Update dqn-hyperopt.toml with new sections
  • Step 4: Verify all TOML fields are read by training_profile.rs
  • Step 5: Build + test
  • Step 6: Commit: "feat: comprehensive TOML training profiles — every parameter configurable"

Task 8: Integration test — config consistency

Files:

  • Create: crates/ml/src/trainers/dqn/smoke_tests/config_consistency.rs

Write a test that verifies all three config structs produce identical values when constructed through the proper flow:

#[test]
fn test_config_consistency() {
    let hp = DQNHyperparameters::conservative();
    let dqn_config = DQNConfig::from(&hp);
    let gpu_config = GpuDqnTrainConfig::from_hyperparams(&hp, &dqn_config);

    // All must agree
    assert_eq!(dqn_config.num_atoms as usize, hp.num_atoms);
    assert!((dqn_config.v_min - hp.v_min()).abs() < 1e-6);
    assert!((gpu_config.v_min as f64 - hp.v_min()).abs() < 0.01);
    assert!((gpu_config.lr as f64 - hp.learning_rate).abs() < 1e-8);
    // ... all shared fields
}
  • Step 1: Write config consistency test
  • Step 2: Test passes
  • Step 3: Commit: "test: config consistency — all 3 structs agree on shared parameters"

Execution Order

Task 1 (remove use_ booleans) → Task 2 (unify defaults) → Task 3 (TOML exposure)
    → Task 4 (v_range computation) → Task 5 (DQNConfig from HP)
    → Task 6 (GpuConfig from HP) → Task 7 (TOML files) → Task 8 (consistency test)

Each task produces a working, testable state. Total: ~8 commits.