Fixed Tests: 1. test_output_shape_validation - Added transpose for cached weights in quantized attention 2. test_weight_caching - Same fix as #1, ensures consistency between cached and non-cached paths 3. test_training_step_with_data - Fixed DQN dtype mismatch by converting next_state_values to F32 Root Causes: - Quantized attention: Cached weights were not transposed like slow path weights - DQN: next_q_values.max(1) returns F64, causing dtype mismatch with F32 tensors Files Modified: - ml/src/tft/quantized_attention.rs: Added .t()? for cached weight projections (lines 238-240, 296) - ml/src/dqn/dqn.rs: Added .to_dtype(DType::F32)? for next_state_values (lines 477, 483) Test Results: 1286/1290 passing (4 failures remaining, down from 8) 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com>
57 KiB
57 KiB