Root cause: shmem_max_in_dim only included trunk dims (state_dim, shared_h1, shared_h2) but not head dims (value_h, adv_h). When hidden_dim_base=32 made the trunk narrow while heads stayed at 128, the BF16 weight tile for branch output (255×128=32640 BF16 elements) overflowed the shared memory region (12288 BF16 elements). On H100 the overflow landed in unused-but-mapped hardware shmem (silent corruption). On RTX 3050 (48KB physical shmem) it hit unmapped memory → CUDA_ERROR_ILLEGAL_ADDRESS. Changes: - gpu_dqn_trainer.rs: shmem_max_in_dim includes value_h/adv_h - Remove all #[ignore] from smoke tests (feature_coverage, training_stability, gpu_residency) - Smoke tests use real .dbn data from test_data/ (hard error if missing) - Remove synthetic_data() fallback — no fake data in tests - GPU-direct DtoD training path (train_step_gpu, FusedTrainScalars) - GPU-native PER priority update kernel (zero CPU readback) - IQN dual-head integration (gpu_iqn_head.rs) - BF16 dtype fixes across 6 model adapters - Hyperopt 30D→31D (iqn_lambda) - portfolio_transformer: unconditional BF16 (remove dead CPU branches) - liquid/adapter: all tests use Cuda(0) directly - Fix pre-existing gpu_kernel_parity_test.rs (stale args) - Fix pre-existing evaluate_baseline.rs (removed fields) Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
11 KiB
Executable File
11 KiB
Executable File