datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dclm-baseline-1.0-parquet
DCLM-baseline
Note: this is an identical copy of https://huggingface.co/datasets/mlfoundations/dclm-baseline-1.0, where all the files have been mapped to a parquet format.
DCLM-baseline is a 4T token / 3B document pretraining dataset that achieves strong performance on language model benchmarks.
Below are comparisions of model trained on DCLM-baseline with other models in the 7B regime.
Model
Params
Tokens
Open dataset?
CORE
MMLU
EXTENDED
Open weights, closed datasets… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/dclm-baseline-1.0-parquet.ettin-parquetTCGA-12K-parquet
TCGA-12K Parquet
Attribution
This dataset contains 224 x 224 JPEG patches from whole-slide images originally downloaded from The Cancer Genome Atlas (TCGA) that are available in the NCI Genomic Data Commons (GDC) Open Access tier. We mirror and repackage a commonly used ~12k WSI subset in parquet format for ease of training. We exclude patches that did not pass HSV thresholding, following the procedure in Kaiko.AI's Midnight paper. Patches were randomly sampled across… See the full description on the dataset page: https://huggingface.co/datasets/medarc/TCGA-12K-parquet.nfpc-parquet-dataset
AML Mule Account Detection Challenge
Data Format: All files are in Apache Parquet format (Snappy compression). Use pandas.read_parquet(), pyarrow.parquet, or any Parquet-compatible reader. Transaction files are split across batch-N/ subdirectories.
Primary Objective/Problem Statement
Identify mule accounts used for money laundering from banking transaction and account data. Given labelled training data and unlabelled test accounts, predict which test accounts are mules.… See the full description on the dataset page: https://huggingface.co/datasets/preetisheoran/nfpc-parquet-dataset.NQ-F_1min_OHLCV_Parquetblbooks-parquet
Dataset Card for British Library Books
This dataset is the same as https://huggingface.co/datasets/TheBritishLibrary/blbooks, however, this version is stored as parquet to avoid needing to run a datasets script. This also makes loading this dataset much quicker.
Dataset Summary
This dataset consists of books digitised by the British Library in partnership with Microsoft. The dataset includes ~25 million pages of out of copyright texts. The majority of the texts were… See the full description on the dataset page: https://huggingface.co/datasets/biglam/blbooks-parquet.usc-x-24-us-election-parquetThis is a version of the USC X 24 US Election Twitter/X Dataset from USC, cleaned and converted to parquet.
The repository contains multiple directories named part_{part_number}, where each directory consists of chunk files prefixed with a timeline. Each chunk file contains 50,000 tweets related to the US elections 2024. Specifically, each subdirectory labeled with the prefix "part" contains 20 chunk files, resulting in a total of 1,000,000 tweets per part... Check out our memo that provides… See the full description on the dataset page: https://huggingface.co/datasets/deadbirds/usc-x-24-us-election-parquet.StockChina-Minute-Parquet
China Stock Market 1-Minute Bar Dataset (Parquet)
This dataset provides high-frequency 1-minute historical candlestick and trading data for Chinese A-Share stocks (SSE / SZSE: .XSHE, .XSHG).
Attribution & Source Credit
This dataset is an optimized Parquet conversion of the original CSV dataset created by jobs-git:
Original Dataset: jobs-git/StockChina-Minute
Original uncompressed CSV size: ~121.5 GB across 1,450 stock symbols.
Optimizations in this… See the full description on the dataset page: https://huggingface.co/datasets/q1232990/StockChina-Minute-Parquet.hypersim-episodes-v3-parquet
hypersim-episodes-v3-parquet
Per-frame Parquet dataset for ReCAST tracker training.
Schema
One row per frame, grouped by episode_id. Arrow memory-mapped access
enables reading specific frames without loading entire episodes.
Column
Type
Description
episode_id
int32
Episode identifier
frame_idx
int32
Frame index within episode
jpeg
binary
JPEG-encoded RGB frame
depth
list<float32>
Flat H×W depth map
seg
list<uint16>
Semantic segmentation (empty if… See the full description on the dataset page: https://huggingface.co/datasets/OSResight/hypersim-episodes-v3-parquet.TCGA-12K-parquet-shuffled
TCGA-12K Parquet (Shuffled)
Attribution
This dataset contains 224 x 224 JPEG patches from whole-slide images originally downloaded from The Cancer Genome Atlas (TCGA) that are available in the NCI Genomic Data Commons (GDC) Open Access tier. We mirror and repackage a commonly used ~12k WSI subset in parquet format for ease of training. We exclude patches that did not pass HSV thresholding, following the procedure in Kaiko.AI's Midnight paper. Patches were randomly sampled… See the full description on the dataset page: https://huggingface.co/datasets/medarc/TCGA-12K-parquet-shuffled.tw-broker-parquetCESNET-TLS-YEAR22-PARQUET
CESNET-TLS-Year22 — canonical flow parquet
CESNET-TLS-Year22 (507,739,073 TLS
flows over the full year 2022 from the CESNET2 backbone, 180 service labels)
converted from the cesnet-datazoo ORIG HDF5 database into a canonical
flow-record parquet schema: 357 daily parquet files, exactly 507,739,073
rows, 39.5 GB zstd.
label_service carries the authoritative APP label decoded from the
PyTables enum embedded in the source database; servicemap.csv (included)
documents the services.… See the full description on the dataset page: https://huggingface.co/datasets/Lystea/CESNET-TLS-YEAR22-PARQUET.congressional-record-parquet
Congressional Record 43rd–114th Congresses — Parquet Edition
This repository contains a Parquet-formatted derivative of the Stanford
Congressional Record for the 43rd–114th Congresses: Parsed Speeches and Phrase Counts
dataset.
The files were prepared as a compact, query-friendly research corpus for historical
text search with tools such as DuckDB and the Congressional Record Explorer.
Coverage
Congresses: 43rd–114th
One Parquet file per Congress
Files:… See the full description on the dataset page: https://huggingface.co/datasets/yeeder/congressional-record-parquet.CICIOT2023-PARQUET
CICIoT2023 — ipfixprobe flow records (Parquet)
1,479,074,715 bidirectional network flows re-exported from the raw PCAPs of
CICIoT2023 (Canadian
Institute for Cybersecurity, University of New Brunswick) with
ipfixprobe 5.7.0, stored as 308 Parquet
files (~19.4 GB) covering 33 attack classes + benign traffic from the
105-device IoT testbed.
The original dataset ships ~587 GB of PCAPs and CSV features computed with a
closed pipeline. This conversion provides an alternative… See the full description on the dataset page: https://huggingface.co/datasets/Lystea/CICIOT2023-PARQUET.tartanair-episodes-v3-parquet
tartanair-episodes-v3-parquet
Per-frame Parquet dataset for ReCAST tracker training.
Schema
One row per frame, grouped by episode_id. Arrow memory-mapped access
enables reading specific frames without loading entire episodes.
Column
Type
Description
episode_id
int32
Episode identifier
frame_idx
int32
Frame index within episode
jpeg
binary
JPEG-encoded RGB frame
depth
list<float32>
Flat H×W depth map
seg
list<uint16>
Semantic segmentation (empty if… See the full description on the dataset page: https://huggingface.co/datasets/OSResight/tartanair-episodes-v3-parquet.hover-parquet
Dataset Card for HoVer (Parquet Format)
Note: This is a scriptless, Parquet-based version of the HoVer dataset for seamless integration with HuggingFace datasets library. No trust_remote_code required!
Quick Start
from datasets import load_dataset
# Load the dataset (no trust_remote_code needed!)
dataset = load_dataset("vincentkoc/hover-parquet")
# Access splits
train = dataset["train"]
validation = dataset["validation"]
test = dataset["test"]
# Example usage… See the full description on the dataset page: https://huggingface.co/datasets/vincentkoc/hover-parquet.setica-tts-parquets-vpsDanbooru-2026-parquet-metadatanucl-parquet-data
Licensing. These shards are derived from ENDF/B-VIII.0, a US Government
work — public domain in the US under 17 U.S.C. §105 — in the NJOY-processed
pointwise form published by the OpenMC project. Neither the evaluation nor its
processed form is ours to relicense, so the previous license: mit tag on this
dataset was incorrect and has been removed. MIT covers the nucl-parquet code
and conversion, not the bundled evaluated data. Per-library terms are recorded
in data/licenses.toml.
Cite: D.A.… See the full description on the dataset page: https://huggingface.co/datasets/gerchowl/nucl-parquet-data.institutional_investors_parquet_by_stockgui_parquet_hfx-community-notes-parquet-20250222All Twitter/X Community Notes data converted to Parquet.
https://communitynotes.x.com/guide/en/about/introduction
Pulled Feb 22, 2025
full-fold-the-rag-parquet-merged0222IOT23-PARQUET
IoT-23 — canonical flow parquet (light path)
IoT-23 (Stratosphere
Laboratory, CTU University: real IoT malware infections + benign IoT device
captures) converted from the light distribution's labeled Zeek conn logs
into a canonical flow-record parquet schema: 23 captures, 325.3M rows,
7.8 GB zstd. One parquet per capture — leave-one-capture-out splits rebuild
from filenames.
Fidelity caveats, by construction (Tier A only):
Source is conn.log.labeled, not pcap: TCP flag… See the full description on the dataset page: https://huggingface.co/datasets/Lystea/IOT23-PARQUET.analytics-parquetsfuss-parquet
FUSS Parquet Dataset
This dataset provides the Free Universal Sound Separation (FUSS) Dataset as a set of parquet files.
The Free Universal Sound Separation (FUSS) Dataset is a database of arbitrary sound mixtures and source-level references, for use in experiments on arbitrary sound separation.
This is the official sound separation data for the DCASE2020 Challenge Task 4: Sound Event Detection and Separation in Domestic Environments.
Overview: FUSS audio data is sourced from a… See the full description on the dataset page: https://huggingface.co/datasets/scaleinvariant/fuss-parquet.btc_20260818T1100.v2.1s.parquet
BTC Up/Down 5m — one-second bars, market and underlying on one clock
Polymarket runs a Bitcoin Up or Down market every five minutes, resolved off a Chainlink
TWAP of BTC/USD. This is one hour of that market on a one-second grid — and, on the same
grid, the spot and perpetual books from the four exchanges the price comes from, plus the
oracle stream the market resolves against.
Every grid is complete and gapless, and the winning outcome is already joined onto the… See the full description on the dataset page: https://huggingface.co/datasets/RabbitGo777/btc_20260818T1100.v2.1s.parquet.Danbooru-2026-parquet-metadatamunch-1-latent-NEW-parquet
🎙️ Urdu TTS Latent Dataset — munch-1-latent-NEW-parquet
Pre-computed DACVAE latent representations for 51,021 Urdu utterances, ready for TTS model training. No audio decoding required at training time — load the dataset, reshape the binary blob, and train.
Source
Field
Value
Source audio
Humair332/Urdu-munch-1
Codec
Aratako/Semantic-DACVAE-Japanese-32dim
Codec sample rate
48,000 Hz
Encoder hop size
1,920 samples
Latent frame rate
25.0 Hz
Latent dim… See the full description on the dataset page: https://huggingface.co/datasets/zuhri025/munch-1-latent-NEW-parquet.IndicVoice-latent-NEW-parquet
