Barthel/variouscryptodata
variouscryptodata Crypto market datasets collected as a by-product of our own research and published so they are not lost. One sub-folder per dataset; each appended nightly where collection is still running. folder what coverage cadence polymarket_updown_orderbook/ Polymarket Up/Down (5m/15m) order books, 10 levels, BTC/ETH/SOL/XRP/DOGE/HYPE/BNB, with Binance spot reference 2026-05-24 → present appended nightly (previous UTC day) hyperliquid_trades/ Hyperliquid perp… See the full description on the dataset page: https://huggingface.co/datasets/Barthel/variouscryptodata.
variouscryptodata
Crypto market datasets collected as a by-product of our own research and published so they are not lost. One sub-folder per dataset; each appended nightly where collection is still running.
polymarketupdownorderbook
Continuous top-10-level order book snapshots for Polymarket's short-dated crypto Up/Down markets (5-minute and 15-minute, BTC/ETH/SOL/XRP/DOGE/HYPE/BNB), captured every ~6 seconds per market (≈2–3 s in the last minute before a market closes), with the Binance spot price, the market's strike and the time left to resolution on every row.
- Granularity: one row per (market, snapshot); ~250 000 rows/day.
- Both outcomes: the
UpandDownbooks are recorded side by side. - Layout:
polymarket_updown_orderbook/data/date=YYYY-MM-DD/book_depth.parquet(Hive-style day partitions).
Polymarket's CLOB API is live-only; to our knowledge no public archive of the order books of these short-dated markets exists.
Schema
Prices in Up/Down are outcome-token prices in USDC (0–1); sizes are in shares. Parse with json.loads (Python) or json_extract (DuckDB). Schema v1 (19 days, 2026-05-24 → 2026-06-11) has 11 columns; v2 (from 2026-06-12) has 15. Reading with union_by_name/diagonal concat handles both.
Quick start
import polars as pl, json
df = pl.read_parquet(
"hf://datasets/Barthel/variouscryptodata/polymarket_updown_orderbook/data/date=2026-08-21/book_depth.parquet")
row = df.filter(pl.col("asset") == "btc").row(0, named=True)
book = json.loads(row["Up"])
print(book["bids"][0], book["asks"][0], row["spot"], row["secs_left"])-- DuckDB
SELECT asset, period_min, count(*)
FROM read_parquet('hf://datasets/Barthel/variouscryptodata/polymarket_updown_orderbook/data/*/*.parquet', union_by_name=true)
GROUP BY 1,2 ORDER BY 1,2;Collection notes (read before modelling)
- Best-effort single-host capture: short gaps (seconds to minutes) occur around reconnects and host maintenance; treat
tsspacing as irregular. - Snapshot cadence is adaptive: ~6 s per market normally, ~2–3 s during the last minute before a market closes; all live markets (7 assets × 2 periods) are polled in the same cycle.
spotis the Binance spot price fetched by the collector at snapshot time, not Polymarket's resolution oracle; use it for analysis, not as ground truth for settlement.- Nothing here is investment advice; no trading strategy is included.
- Every upload is scanned automatically for credentials before publishing.
hyperliquid_trades
Every public trade print streamed from Hyperliquid's WebSocket trades channel for 30 markets: the perps BTC, ETH, SOL, HYPE, AVAX and 25 HIP-3 markets (tokenised equities, indices, commodities and FX such as xyz:NVDA, xyz:TSLA, xyz:GOLD, xyz:SP500, xyz:EUR, cash:USA500, km:US500 — the coin column uses Hyperliquid's dex:NAME form). One Parquet per UTC day, all markets in one file. Roughly 100 rows per day carry exchange timestamps far outside the file's day: on every (re)subscription Hyperliquid replays a market's most recent trades, and for the three dormant markets cash:SILVER, cash:USA500, km:US500 (no trades at all during the collection period so far) those ~30 replayed trades date from June/July 2026 and recur in every daily file; filter on time_ms if that matters to you.
- Layout:
hyperliquid_trades/data/date=YYYY-MM-DD/trades.parquet - Rows: ~0.6–3.6 million/day (median ≈2 M; weekends lowest); sorted by
time_ms; de-duplicated on (coin,tid).
Deliberately not included: counterparty wallet addresses (users) and transaction hashes — this dataset is about prices and flow, not about who traded. Collection is best-effort from a single WebSocket client; brief gaps (reconnects) can occur. HIP-3 markets follow their own trading hours, so zero-trade stretches there are normal, not gaps.
import polars as pl
t = pl.read_parquet("hf://datasets/Barthel/variouscryptodata/hyperliquid_trades/data/date=2026-08-21/trades.parquet")
print(t.group_by("coin").agg(pl.len(), (pl.col("px")*pl.col("sz")).sum().alias("notional")).sort("notional", descending=True).head(10))hyperliquidmainnetarchive (part 1 static 2026-07-18 → 2026-08-04; part 2 from 2026-08-22, appended nightly)
An 18-day capture (2026-07-18 ~14:03 UTC → 2026-08-04 ~13:25 UTC; first and last day partial) of Hyperliquid's public WebSocket feed by a shadow-trading research bot (no orders were sent from this data). Coins were the bot's watch-list at the time — 16 liquid perps: BTC, ETH, SOL, HYPE, XRP, AVAX, NEAR, ONDO, UNI, WLD, ZEC, PUMP, TRUMP, FARTCOIN, LIT, VVV (15 coins on 2026-07-18 → 07-20; AVAX was added 2026-07-21, 16 from then on). One Parquet per table per UTC day: hyperliquid_mainnet_archive/<table>/date=YYYY-MM-DD/<table>.parquet.
Part 2 (from 2026-08-22 ~19:49 UTC, first day partial; appended nightly): a dedicated read-only collector subscribes l2Book + trades for every perp on the main exchange and every HIP-3 market (≈320 markets, re-discovered every 6 h, so newly listed markets appear automatically), and bbo for the ~40 highest-volume markets (ranked at collector start; new high-volume entrants are added at discovery, none removed). Same table layout and column names as part 1, with these differences:
asset_ctx,markandfundingcome from the RESTmetaAndAssetCtxsendpoint once per minute (part 1: streamed, ~5 s).bbois derived from eachl2booksnapshot (identical definition to part 1).l2book: up to 20 levels per side — thin HIP-3 markets can have fewer (seebid_levels/ask_levels); the ~5.4 s cadence is the exchange's push rate.candlesare 5-minute bars built fromtradescaptured live (receive latency ≤ 60 s), only buckets with ≥ 1 trade (part 1: exchange candle stream incl. empty buckets); buckets around collector (re)starts may be partial.trades: on every (re)subscription Hyperliquid replays the last ~30 trades of a market, so each daily file can contain a few older trades per market — filter ontime_msif that matters.- Part-2 volumes:
l2book/bbo≈ 5 M rows/day,trades≈ 5–8 M,bbo_stream≈ 9–11 M,asset_ctx/mark/funding≈ 0.46 M,candles≤ 92 k. - Receive latency (recvms − timems): l2book ≈ 0.6 s median / ≈ 1–2 s p99; trades/bbo ≈ 0.35 s median; bursts up to ~10 s at (re)subscribe.
Gap between part 1 and part 2: 2026-08-04 13:25 → 2026-08-22 19:49 UTC.
Notes: time_ms/open_ms are exchange timestamps; recv_ms is the collector's receive time (single host, best-effort; latency typically ≈0.45 s median and ≈1 s p99, with occasional bursts up to ~10–30 s). Hyperliquid's own complete history is available from the exchange's requester-pays S3 archive; this is a free, partial mirror for convenience.
import polars as pl, json
b = pl.read_parquet("hf://datasets/Barthel/variouscryptodata/hyperliquid_mainnet_archive/l2book/date=2026-07-27/l2book.parquet")
snap = b.filter(pl.col("coin") == "ETH").row(0, named=True)
bids, asks = json.loads(snap["bids"]), json.loads(snap["asks"])
print(bids[0], asks[0]) # [px, sz, n_orders]License & citation
Data © the collector, released under CC-BY-4.0. Underlying quotes originate from Polymarket's public CLOB API, Binance's public REST API and Hyperliquid's public WebSocket API. If you use this data, please cite “Barthel/variouscryptodata (Hugging Face dataset)”.
