Deleted: - DQN::compute_loss_internal (280 lines) — old Candle forward+loss - DQN::train_step (55 lines) — old Candle training step - DQN::compute_gradients (47 lines) — old gradient accumulation - ComputeLossResult struct — only used by deleted functions - RegimeConditionalDQN::train_step (65 lines) — old dispatch - RegimeConditionalDQN::train_step_gpu_regime (100 lines) — old GPU path - RegimeConditionalDQN::compute_gradients_gpu (130 lines) — old regime gradients - RegimeConditionalDQN::compute_gradients (92 lines) — old dispatch - DQNAgentType::train_step dispatch — dead - DQNAgentType::compute_gradients dispatch — dead - GpuDqnTrainer::upload_batch (71 lines) — old CPU→GPU upload - train_step.rs (500 lines) — entire module including ensure_fused_ctx - dqn_benchmark.rs — used old train_step - examples.rs — used old train_step - validation/adapters.rs (289 lines) — used old train_step - dqn/trainable_adapter.rs — used old train_step - gpu_smoketest.rs — tested old train_step - Gradient accumulation path in training_loop.rs (144 lines) - IQN d_h_s2().clone() → raw pointer (zero alloc) - Causal intervention format! string alloc removed - Dead HER relabel functions (320 lines) Kept: - ensure_fused_ctx logic inlined into training_loop.rs - set_noise_sigma_scale re-added to RegimeConditionalDQN Fixed: - GpuReplayBuffer max_batch_size wired from batch_size parameter (was hardcoded 1024, blocking batch_size=8192) Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
55 lines
2.1 KiB
Rust
55 lines
2.1 KiB
Rust
//! GPU Training Benchmark System for Production ML Models
|
|
//!
|
|
//! This module provides comprehensive benchmarking for GPU-based model training
|
|
//! with real market data from Databento (DBN format). Measures actual training
|
|
//! performance, memory usage, and stability metrics across different models.
|
|
//!
|
|
//! # Model Benchmarks
|
|
//!
|
|
//! - `dqn_benchmark` - Deep Q-Network (Module 6a) - 50-150MB VRAM
|
|
//! - `ppo_benchmark` - Proximal Policy Optimization (Module 6b) - 80-200MB VRAM
|
|
//! - `mamba2_benchmark` - MAMBA-2 State Space Model (Module 6c) - 150-500MB VRAM
|
|
//! - `tft_benchmark` - Temporal Fusion Transformer (Module 6d) - 1.5-2.5GB VRAM
|
|
//!
|
|
//! # Core Components
|
|
//!
|
|
//! - `stability_validator` - Training stability detection (Module 5)
|
|
//! - `memory_profiler` - GPU memory tracking
|
|
//! - `gpu_hardware` - GPU hardware detection
|
|
//! - `statistical_sampler` - Performance statistics
|
|
//! - `batch_size_finder` - Optimal batch size selection
|
|
//! - `data_loader` - Real market data loading
|
|
|
|
pub mod batch_size_finder;
|
|
pub mod data_loader;
|
|
pub mod gpu_hardware;
|
|
pub mod mamba2_benchmark;
|
|
pub mod memory_profiler;
|
|
pub mod performance_tracker;
|
|
pub mod ppo_benchmark;
|
|
pub mod stability_validator;
|
|
pub mod statistical_sampler;
|
|
pub mod tft_benchmark;
|
|
|
|
// Re-export core types
|
|
pub use batch_size_finder::{BatchSizeConfig, BatchSizeFinder};
|
|
pub use data_loader::{DataStatistics, DbnDataLoader, MarketDataPoint};
|
|
pub use gpu_hardware::GpuHardwareManager;
|
|
pub use memory_profiler::{MemoryProfiler, MemorySnapshot};
|
|
pub use performance_tracker::{
|
|
PerformanceBaseline, PerformanceMetrics, PerformanceTracker, RegressionItem, RegressionResult,
|
|
};
|
|
pub use stability_validator::{GradientHealth, LossTrend, StabilityMetrics, StabilityValidator};
|
|
pub use statistical_sampler::{BenchmarkStatistics, StatisticalSampler};
|
|
|
|
// Re-export DQN benchmark types
|
|
|
|
// Re-export PPO benchmark types
|
|
pub use ppo_benchmark::{PpoBenchmarkResult, PpoBenchmarkRunner};
|
|
|
|
// Re-export MAMBA-2 benchmark types
|
|
pub use mamba2_benchmark::{Mamba2BenchmarkResult, Mamba2BenchmarkRunner};
|
|
|
|
// Re-export TFT benchmark types
|
|
pub use tft_benchmark::{TftBenchmarkResult, TftBenchmarkRunner};
|