Wave 1 (Architecture & Design - 5 agents): - Multi-model training orchestration (DQN, PPO, MAMBA-2, TFT-INT8) - Sequential training strategy (95.9% GPU headroom, 6.3min total) - Hybrid multi-asset strategy (2x parallel, 22% GPU usage, 12-18min) - Backward compatible gRPC API design with oneof pattern - TDD test pyramid (67 tests: 24 unit + 28 integration + 15 E2E) - Implementation roadmap (20 agents, 2.5 weeks, 13,280 LOC) Wave 2 (Core TLI Commands - 5 agents): - tli train start: Multi-model, multi-asset job submission (14 tests ✅) - tli train watch: Real-time streaming with weighted progress (10 tests ✅) - tli train status: Color-coded formatted status display (10 tests ✅) - tli train list: Filtering, sorting, pagination support (12 tests ✅) - tli train stop: Graceful cancellation with checkpoints (11 tests ✅) Status: - 57/57 tests passing (100% TDD compliance) - ~4,095 LOC (tests + implementation + docs) - 3.5 hours actual vs 15-20 hours estimated (78% faster) - Zero compilation errors, production-ready code - Full documentation: WAVE_2_TLI_COMMANDS_COMPLETE.md Next: Wave 3 (Multi-Asset Multi-Model Backend Logic - 5 agents) 🤖 Generated with Claude Code Co-Authored-By: Claude <noreply@anthropic.com>
57 lines
1.5 KiB
Bash
Executable File
57 lines
1.5 KiB
Bash
Executable File
#!/bin/bash
|
|
# Memory Profiling Script for ML Training
|
|
# Usage: ./profile_memory.sh <MODEL> <PARQUET_FILE> [EPOCHS]
|
|
#
|
|
# Example: ./profile_memory.sh dqn test_data/ES_FUT_small.parquet 3
|
|
|
|
set -e
|
|
|
|
MODEL=$1
|
|
PARQUET=$2
|
|
EPOCHS=${3:-3}
|
|
|
|
if [ -z "$MODEL" ] || [ -z "$PARQUET" ]; then
|
|
echo "Usage: $0 <MODEL> <PARQUET_FILE> [EPOCHS]"
|
|
echo " MODEL: dqn, ppo, mamba2, tft"
|
|
echo " PARQUET_FILE: path to training data"
|
|
echo " EPOCHS: number of epochs (default: 3)"
|
|
exit 1
|
|
fi
|
|
|
|
echo "=========================================="
|
|
echo "Memory Profiling: train_${MODEL}"
|
|
echo "Data: $PARQUET"
|
|
echo "Epochs: $EPOCHS"
|
|
echo "=========================================="
|
|
echo
|
|
|
|
# Get base directory (foxhunt root)
|
|
SCRIPT_DIR="$( cd "$( dirname "${BASH_SOURCE[0]}" )" && pwd )"
|
|
FOXHUNT_ROOT="$( cd "$SCRIPT_DIR/../.." && pwd )"
|
|
|
|
cd "$FOXHUNT_ROOT"
|
|
|
|
# Check if parquet file exists
|
|
if [ ! -f "$PARQUET" ]; then
|
|
echo "Error: Parquet file not found: $PARQUET"
|
|
exit 1
|
|
fi
|
|
|
|
# Run training with memory profiling
|
|
echo "Starting profiling run..."
|
|
echo
|
|
|
|
/usr/bin/time -v cargo run -p ml --example train_${MODEL} --release -- \
|
|
--parquet-file "$PARQUET" \
|
|
--epochs "$EPOCHS" \
|
|
2>&1 | tee /tmp/profile_${MODEL}_$$.log
|
|
|
|
echo
|
|
echo "=========================================="
|
|
echo "Memory Profile Summary"
|
|
echo "=========================================="
|
|
grep -E "(Maximum resident|User time|System time|Percent of CPU|Elapsed|Minor.*page faults|Major.*page faults)" /tmp/profile_${MODEL}_$$.log || true
|
|
|
|
echo
|
|
echo "Full log saved to: /tmp/profile_${MODEL}_$$.log"
|