bart
Datasets
All datasets matching “bart”variouscryptodata
variouscryptodata
Crypto market datasets collected as a by-product of our own research and
published so they are not lost. One sub-folder per dataset; each appended
nightly where collection is still running.
folder
what
coverage
cadence
polymarket_updown_orderbook/
Polymarket Up/Down (5m/15m) order books, 10 levels, BTC/ETH/SOL/XRP/DOGE/HYPE/BNB, with Binance spot reference
2026-05-24 → present
appended nightly (previous UTC day)
hyperliquid_trades/
Hyperliquid perp… See the full description on the dataset page: https://huggingface.co/datasets/Barthel/variouscryptodata.bartholomew-dataset-v3
BART Dataset v3
The final version of the BART pretraining corpus, focused on removing anything that betrays a
post-1930 origin. This is our best vintage dataset yet.
Documents
146,031 (97.52% of v2)
Characters
102,798,688,961 (96.73% of v2)
Tokens
~23B (estimated)
Shards
473 (one per v2 shard, same basename)
Source
BART Dataset v2
Cutoff
1930
Schema
single string column text
Lineage — three cumulative filtering stages over the same corpus:… See the full description on the dataset page: https://huggingface.co/datasets/jbduran/bartholomew-dataset-v3.rayst3rEventHubDatasetbartholomew-dataset-v2
BART Dataset v2
The second version of the BART pretraining corpus, focused on stripping low-quality text —
boilerplate and OCR corruption — out of
v1.
Documents
149,745 (93.44% of v1)
Characters
106,274,384,672 (89.50% of v1)
Tokens
~24B (estimated)
Shards
473 (one per v1 shard, same basename)
Source
BART Dataset v1
Schema
single string column text
Lineage — three cumulative filtering stages over the same corpus:
Institutional Books 1.0 →
v1 →
v2 →
v3… See the full description on the dataset page: https://huggingface.co/datasets/jbduran/bartholomew-dataset-v2.vcc2026-broad-development-20260908
VCC 2026 broad development artifacts
The B7 RunPod science-output recovery is complete: 256 files, 5,106,104,214 bytes.
Contents
Files
Checkpoints
64
Fit receipts
64
Prediction arrays
64
Prediction receipts
64
Use archive-index.json for current locations, SHA-256 values, sizes, and immutable per-file revisions. STATUS.md records completed work and remaining scientific comparisons. The recovery receipt proves all 109 files missing from the previous archive… See the full description on the dataset page: https://huggingface.co/datasets/barthazian/vcc2026-broad-development-20260908.
