- Created data/examples/download_ml_training_data.rs using reqwest + Databento HTTP API - Downloaded 90 days × 4 symbols (ES.FUT, NQ.FUT, ZN.FUT, 6E.FUT) - Files saved to test_data/real/databento/ml_training/ - Total: 360 files, 15 MB compressed DBN format - Used existing Rust pattern from download_nq_fut.rs - API key loaded from .env file - 100% success rate (360/360 files) - Ready for ML training benchmarks Next: Create simplified training benchmark for RTX 3050 Ti GPU measurements
412 lines
12 KiB
Markdown
412 lines
12 KiB
Markdown
# Kaggle Dataset Technical Specifications
|
|
## Wave 153 Data Source Bake-Off
|
|
|
|
**Last Updated**: 2025-10-12
|
|
|
|
---
|
|
|
|
## 📊 Dataset Specifications
|
|
|
|
### BTC/USD Dataset
|
|
|
|
| Attribute | Value |
|
|
|-----------|-------|
|
|
| **Dataset Name** | Bitcoin BTC, 7 Exchanges, 1m Full Historical Data |
|
|
| **Kaggle ID** | `imranbukhari/comprehensive-btcusd-1m-data` |
|
|
| **URL** | https://www.kaggle.com/datasets/imranbukhari/comprehensive-btcusd-1m-data |
|
|
| **File Size** | 262 MB |
|
|
| **Format** | CSV (8 files) |
|
|
| **Last Update** | 2025-10-10 (4 days ago) |
|
|
| **Update Frequency** | Daily |
|
|
| **Estimated Rows** | ~3.8 million |
|
|
| **Date Range** | Full historical (earliest available to Oct 2025) |
|
|
| **Exchanges** | 7 (Binance, Coinbase, Bitfinex, Bitstamp, BitMEX, KuCoin, OKX) |
|
|
| **Usability Score** | 10.0/10 |
|
|
| **Downloads** | 3,073 |
|
|
| **Views** | 11,900 |
|
|
| **License** | CC BY-SA 4.0 |
|
|
|
|
### ETH/USD Dataset
|
|
|
|
| Attribute | Value |
|
|
|-----------|-------|
|
|
| **Dataset Name** | Ethereum ETH, 7 Exchanges, 1m Full Historical Data |
|
|
| **Kaggle ID** | `imranbukhari/comprehensive-ethusd-1m-data` |
|
|
| **URL** | https://www.kaggle.com/datasets/imranbukhari/comprehensive-ethusd-1m-data |
|
|
| **File Size** | 22 MB |
|
|
| **Format** | CSV (8 files) |
|
|
| **Last Update** | 2025-09-15 (1 month ago) |
|
|
| **Update Frequency** | Monthly |
|
|
| **Estimated Rows** | ~320,000 |
|
|
| **Date Range** | Full historical (earliest available to Sep 2025) |
|
|
| **Exchanges** | 7 (Binance, Coinbase, Bitfinex, Bitstamp, Kraken, KuCoin, OKX) |
|
|
| **Usability Score** | 10.0/10 |
|
|
| **Downloads** | Not specified |
|
|
| **Views** | Not specified |
|
|
| **License** | CC BY-SA 4.0 |
|
|
|
|
---
|
|
|
|
## 📁 File Structure
|
|
|
|
### Files Included (Per Dataset)
|
|
|
|
Both BTC and ETH datasets contain 8 CSV files:
|
|
|
|
```
|
|
comprehensive-btcusd-1m-data/
|
|
├── BTCUSD_1m_Binance.csv # Binance exchange data
|
|
├── BTCUSD_1m_Coinbase.csv # Coinbase exchange data
|
|
├── BTCUSD_1m_Bitfinex.csv # Bitfinex exchange data
|
|
├── BTCUSD_1m_Bitstamp.csv # Bitstamp exchange data
|
|
├── BTCUSD_1m_BitMEX.csv # BitMEX exchange data
|
|
├── BTCUSD_1m_KuCoin.csv # KuCoin exchange data
|
|
├── BTCUSD_1m_OKX.csv # OKX exchange data
|
|
└── BTCUSD_1m_Combined_Index.csv # ⭐ RECOMMENDED: Multi-exchange aggregate
|
|
```
|
|
|
|
**Recommended File**: Use `*_Combined_Index.csv` for longest continuous series.
|
|
|
|
---
|
|
|
|
## 🔧 Data Schema
|
|
|
|
### CSV Column Structure
|
|
|
|
| Column | Type | Description | Example | Notes |
|
|
|--------|------|-------------|---------|-------|
|
|
| `timestamp` | DateTime | UTC timestamp (start of 1-min interval) | `2024-10-01 00:00:00` | ISO 8601 format |
|
|
| `open` | Float | Price at interval start | `42000.00` | USD |
|
|
| `high` | Float | Highest price in interval | `42010.00` | USD, >= all others |
|
|
| `low` | Float | Lowest price in interval | `41990.00` | USD, <= all others |
|
|
| `close` | Float | Price at interval end | `42005.00` | USD |
|
|
| `volume` | Float | Total traded volume | `1.234` | Base currency (BTC/ETH) |
|
|
| `exchange` | String | Source exchange (optional) | `Binance` | May be absent in some files |
|
|
|
|
### OHLCV Constraints
|
|
|
|
Must satisfy:
|
|
```
|
|
high >= open
|
|
high >= close
|
|
high >= low
|
|
low <= open
|
|
low <= close
|
|
volume >= 0
|
|
```
|
|
|
|
---
|
|
|
|
## 🧮 Size Calculations
|
|
|
|
### BTC Dataset
|
|
- **File Size**: 262 MB
|
|
- **Estimated Rows**: ~3.8 million
|
|
- **Bytes per Row**: ~69 bytes (compressed CSV)
|
|
- **Expected Days**: ~2,639 days (~7.2 years)
|
|
- **Start Date Estimate**: ~2018-01-01 (if updated to 2025-10-10)
|
|
|
|
### ETH Dataset
|
|
- **File Size**: 22 MB
|
|
- **Estimated Rows**: ~320,000
|
|
- **Bytes per Row**: ~69 bytes (compressed CSV)
|
|
- **Expected Days**: ~222 days (~7.3 months)
|
|
- **Start Date Estimate**: ~2025-02-01 (if updated to 2025-09-15)
|
|
|
|
**Note**: Actual date ranges unknown without download.
|
|
|
|
---
|
|
|
|
## 🔄 Preprocessing Pipeline
|
|
|
|
### Applied Transformations
|
|
|
|
1. **Multi-Exchange Aggregation**
|
|
- Combines data from 7 exchanges
|
|
- Averaging methodology (exact algorithm unclear)
|
|
- Creates `*_Combined_Index.csv`
|
|
|
|
2. **Deduplication**
|
|
- Removes duplicate timestamps
|
|
- Ensures unique 1-minute intervals
|
|
|
|
3. **Chronological Sorting**
|
|
- Data sorted by timestamp ascending
|
|
|
|
4. **UTC Standardization**
|
|
- All timestamps converted to UTC
|
|
- Aligned to 1-minute boundaries
|
|
|
|
5. **Gap Filling**
|
|
- Missing intervals filled (methodology unclear)
|
|
- Options: interpolation, forward-fill, nulls?
|
|
|
|
6. **Outlier Filtering**
|
|
- Anomalous prices/volumes removed
|
|
- Threshold criteria not documented
|
|
|
|
7. **Validation**
|
|
- OHLCV constraints enforced
|
|
- Continuity checks applied
|
|
|
|
### Preprocessing Impact
|
|
|
|
⚠️ **RED FLAGS**:
|
|
- NOT raw data - may hide exchange-specific anomalies
|
|
- Gap-filling may introduce synthetic data
|
|
- Averaging may smooth volatility spikes
|
|
- Outlier removal may delete valid extreme events
|
|
|
|
✅ **BENEFITS**:
|
|
- Cleaner data for ML training
|
|
- No missing intervals (continuous series)
|
|
- Multi-exchange view reduces bias
|
|
- Author claims superior MAE performance
|
|
|
|
---
|
|
|
|
## 📈 Data Quality Metrics
|
|
|
|
### Claimed Attributes
|
|
- ✅ **Completeness**: 100% (no gaps)
|
|
- ✅ **Continuity**: Unbroken time series
|
|
- ✅ **Accuracy**: Outliers filtered
|
|
- ✅ **Consistency**: Standardized format
|
|
- ✅ **Timeliness**: Daily updates (BTC)
|
|
|
|
### Validation Required
|
|
- ⚠️ **OHLCV Constraints**: Must verify High >= Low, etc.
|
|
- ⚠️ **Timestamp Continuity**: Should be 1440 rows/day
|
|
- ⚠️ **Volume Validation**: Check zeros, negatives, outliers
|
|
- ⚠️ **Completeness**: Calculate actual % of expected intervals
|
|
- ⚠️ **Preprocessing Impact**: Compare individual vs combined files
|
|
|
|
---
|
|
|
|
## 🧪 Sample Data Validation Script
|
|
|
|
### Python Validation Code
|
|
|
|
```python
|
|
import pandas as pd
|
|
import numpy as np
|
|
|
|
def validate_kaggle_ohlcv(filepath):
|
|
"""Validate Kaggle OHLCV dataset quality."""
|
|
|
|
# Load data
|
|
df = pd.read_csv(filepath, parse_dates=['timestamp'])
|
|
|
|
# Basic stats
|
|
print(f"Rows: {len(df):,}")
|
|
print(f"Date Range: {df['timestamp'].min()} to {df['timestamp'].max()}")
|
|
print(f"Days: {(df['timestamp'].max() - df['timestamp'].min()).days}")
|
|
|
|
# OHLCV Constraints
|
|
high_violations = ~(
|
|
(df['high'] >= df['open']) &
|
|
(df['high'] >= df['close']) &
|
|
(df['high'] >= df['low'])
|
|
)
|
|
low_violations = ~(
|
|
(df['low'] <= df['open']) &
|
|
(df['low'] <= df['close'])
|
|
)
|
|
print(f"High Violations: {high_violations.sum()}")
|
|
print(f"Low Violations: {low_violations.sum()}")
|
|
|
|
# Timestamp Continuity
|
|
df['date'] = df['timestamp'].dt.date
|
|
daily_counts = df.groupby('date').size()
|
|
expected_per_day = 1440 # 1-minute intervals
|
|
completeness = (daily_counts == expected_per_day).mean() * 100
|
|
print(f"Completeness: {completeness:.2f}% (days with all 1440 rows)")
|
|
print(f"Missing Intervals: {(expected_per_day - daily_counts).sum()}")
|
|
|
|
# Volume Validation
|
|
zero_volume_pct = (df['volume'] == 0).mean() * 100
|
|
negative_volume = (df['volume'] < 0).any()
|
|
volume_outliers = (df['volume'] > df['volume'].quantile(0.999)).sum()
|
|
print(f"Zero Volume: {zero_volume_pct:.2f}%")
|
|
print(f"Negative Volume: {negative_volume}")
|
|
print(f"Volume Outliers (>99.9%ile): {volume_outliers}")
|
|
|
|
# Price Outliers
|
|
for col in ['open', 'high', 'low', 'close']:
|
|
outliers = (df[col] > df[col].quantile(0.999)).sum()
|
|
print(f"{col.capitalize()} Outliers: {outliers}")
|
|
|
|
return {
|
|
'row_count': len(df),
|
|
'completeness_pct': completeness,
|
|
'ohlcv_violations': high_violations.sum() + low_violations.sum(),
|
|
'zero_volume_pct': zero_volume_pct,
|
|
'volume_outliers': volume_outliers
|
|
}
|
|
|
|
# Example usage
|
|
btc_stats = validate_kaggle_ohlcv('BTCUSD_1m_Combined_Index.csv')
|
|
eth_stats = validate_kaggle_ohlcv('ETHUSD_1m_Combined_Index.csv')
|
|
```
|
|
|
|
---
|
|
|
|
## 🔌 Integration with Foxhunt
|
|
|
|
### Recommended Usage Pattern
|
|
|
|
```rust
|
|
// Rust integration pseudocode
|
|
|
|
use chrono::{DateTime, Utc};
|
|
use serde::{Deserialize, Serialize};
|
|
|
|
#[derive(Debug, Deserialize, Serialize)]
|
|
struct KaggleOHLCV {
|
|
timestamp: DateTime<Utc>,
|
|
open: f64,
|
|
high: f64,
|
|
low: f64,
|
|
close: f64,
|
|
volume: f64,
|
|
}
|
|
|
|
impl KaggleOHLCV {
|
|
/// Validate OHLCV constraints
|
|
fn is_valid(&self) -> bool {
|
|
self.high >= self.open
|
|
&& self.high >= self.close
|
|
&& self.high >= self.low
|
|
&& self.low <= self.open
|
|
&& self.low <= self.close
|
|
&& self.volume >= 0.0
|
|
}
|
|
|
|
/// Convert to Foxhunt MarketDataEvent
|
|
fn to_market_event(&self) -> MarketDataEvent {
|
|
MarketDataEvent {
|
|
symbol: "BTC/USD".to_string(),
|
|
timestamp: self.timestamp,
|
|
price: self.close,
|
|
volume: self.volume,
|
|
bid: None, // Not available in OHLCV
|
|
ask: None, // Not available in OHLCV
|
|
high: Some(self.high),
|
|
low: Some(self.low),
|
|
open: Some(self.open),
|
|
}
|
|
}
|
|
}
|
|
|
|
/// Load Kaggle CSV for backtesting
|
|
async fn load_kaggle_data(path: &str) -> Result<Vec<KaggleOHLCV>> {
|
|
let mut rdr = csv::Reader::from_path(path)?;
|
|
let mut events = Vec::new();
|
|
|
|
for result in rdr.deserialize() {
|
|
let record: KaggleOHLCV = result?;
|
|
if !record.is_valid() {
|
|
warn!("Invalid OHLCV record: {:?}", record);
|
|
continue;
|
|
}
|
|
events.push(record);
|
|
}
|
|
|
|
Ok(events)
|
|
}
|
|
```
|
|
|
|
---
|
|
|
|
## 📊 Performance Characteristics
|
|
|
|
### File I/O
|
|
- **CSV Read**: ~1-2 seconds for 262MB BTC file (SSD)
|
|
- **Memory Usage**: ~500MB for full BTC dataset in memory
|
|
- **Parsing**: Standard CSV, highly optimized by pandas/polars
|
|
|
|
### Recommended Tools
|
|
- **Python**: pandas, polars (10x faster for large files)
|
|
- **Rust**: csv crate, serde
|
|
- **Compression**: gzip/zstd for storage (3-5x reduction)
|
|
|
|
---
|
|
|
|
## 🔐 License Compliance
|
|
|
|
### CC BY-SA 4.0 Requirements
|
|
|
|
**Attribution Required**:
|
|
```
|
|
Data source: "Bitcoin BTC, 7 Exchanges, 1m Full Historical Data"
|
|
by Imran Bukhari, Kaggle. Licensed under CC BY-SA 4.0.
|
|
https://www.kaggle.com/datasets/imranbukhari/comprehensive-btcusd-1m-data
|
|
```
|
|
|
|
**ShareAlike**: Derivatives must use same license.
|
|
|
|
**Allowed**:
|
|
- Commercial use ✅
|
|
- Distribution ✅
|
|
- Modification ✅
|
|
- Private use ✅
|
|
|
|
**Conditions**:
|
|
- Attribution required
|
|
- ShareAlike for derivatives
|
|
- No warranty provided
|
|
|
|
---
|
|
|
|
## 🚨 Known Issues & Limitations
|
|
|
|
1. **Preprocessing Opacity**
|
|
- Gap-filling methodology unclear
|
|
- Averaging algorithm not documented
|
|
- Outlier thresholds unknown
|
|
|
|
2. **Data Lag**
|
|
- BTC: 4-day lag
|
|
- ETH: 1-month lag
|
|
- Not suitable for live trading
|
|
|
|
3. **Missing Metadata**
|
|
- Exact start dates unknown
|
|
- Exchange weights in combined index unclear
|
|
- Volume units assumed (not explicit)
|
|
|
|
4. **Potential Biases**
|
|
- Multi-exchange averaging may smooth volatility
|
|
- Outlier removal may delete valid extreme events
|
|
- Gap filling may introduce synthetic data
|
|
|
|
5. **Download Barriers**
|
|
- Kaggle account required
|
|
- Large files (262MB BTC)
|
|
- Cannot sample without full download
|
|
|
|
---
|
|
|
|
## 📞 Support & Resources
|
|
|
|
### Dataset Author
|
|
- **Name**: Imran Bukhari
|
|
- **Kaggle Profile**: https://www.kaggle.com/imranbukhari
|
|
- **Contact**: Via Kaggle discussion tab
|
|
|
|
### Documentation
|
|
- **Usage Notebook**: Included with download
|
|
- **Diagnostics**: Data quality functions provided
|
|
- **Aggregation Methodology**: Coming soon (per author)
|
|
|
|
### Community
|
|
- **Discussions**: https://www.kaggle.com/datasets/imranbukhari/comprehensive-btcusd-1m-data/discussion
|
|
- **Code Examples**: https://www.kaggle.com/datasets/imranbukhari/comprehensive-btcusd-1m-data/code
|
|
|
|
---
|
|
|
|
**Document Version**: 1.0
|
|
**Last Updated**: 2025-10-12
|
|
**Wave 153 Agent**: Kaggle Technical Analysis
|
|
**Status**: ✅ COMPLETE
|