datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DatasetWithCapitalLettersLiveHouse-TS-test
LiveHouse-TS
A prospective benchmark for univariate time-series forecasting.
Real streams. Predictions frozen in advance. Results as the future unfolds.
Explore the leaderboard ↗
·
Submit a model
·
Public data
·
Evaluation
·
Cite
Forecast first. Evaluate later.
LiveHouse-TS asks a simple question: how well does a model predict data that
has not arrived yet? Models receive… See the full description on the dataset page: https://huggingface.co/datasets/Saxon0520/LiveHouse-TS-test.cyp-challenge-train-test
CYP Challenge Train/Test Dataset
A high-quality experimental dataset for predicting inhibition of the major drug-metabolizing Cytochrome P450 enzymes (CYP1A2, CYP2C9, CYP2D6, CYP3A4), released as part of the OpenADMET CYP Inhibition Blind Challenge.
Blog post: Announcing OpenADMET’s CYP inhibition blind challenge
Challenge Space: OpenADMET CYP Inhibition Blind Challenge
Challenge period: August 17, 2026 - November 3, 2026
Produced by: OpenADMET
CHANGELOG
Updated… See the full description on the dataset page: https://huggingface.co/datasets/openadmet/cyp-challenge-train-test.protein_data_testsplit 1, 2 -> for sequences
split 3, 4 -> for residues
SpatialLM-Testset
SpatialLM Testset
Project page | Paper | Code
We provide a test set of 107 preprocessed point clouds and their corresponding GT layouts, point clouds are reconstructed from RGB videos using MASt3R-SLAM. SpatialLM-Testset is quite challenging compared to prior clean RGBD scan datasets due to the noises and occlusions in the point clouds reconstructed from monocular RGB videos.
Folder Structure
Outlines of the dataset files:… See the full description on the dataset page: https://huggingface.co/datasets/manycore-research/SpatialLM-Testset.pxr-challenge-train-test
PXR Challenge Train/Test Dataset
A high-quality experimental dataset for predicting human Pregnane-X Receptor (PXR) induction, comprising over 11,000 compounds screened using a high-fidelity in-house assay. This is the largest publicly available PXR activity dataset, released as part of the OpenADMET PXR Induction Blind Challenge.
Blog post: Announcing the Next OpenADMET Blind Challenge: Predicting PXR Induction
Challenge Space: openadmet/pxr-challenge
Challenge period: April 1… See the full description on the dataset page: https://huggingface.co/datasets/openadmet/pxr-challenge-train-test.protein_data_test_2vindr-cxr-testsetcontextual_testCheck out the paper.
factur-x-test-data
Factur-X / ZUGFeRD Test Data
Well-formed UN/CEFACT Cross Industry Invoice XML with synthetic seller, buyer, and totals for e-invoicing tests.
Free to use under CC0-1.0 — public domain dedication, no attribution required.
Files
factur-x-small.json / factur-x-small.csv — 100 rows (documents: 25)
factur-x-medium.json / factur-x-medium.csv — 2,000 rows (documents: 250)
factur-x-large.json / factur-x-large.csv — 20,000 rows (documents: 2,500)
Deterministic: the same… See the full description on the dataset page: https://huggingface.co/datasets/StanzaAPI/factur-x-test-data.iso20022-test-data
ISO 20022 Test Data
Well-formed pain.001.001.09 Customer Credit Transfer Initiation XML with synthetic debtors, creditors, and amounts.
Free to use under CC0-1.0 — public domain dedication, no attribution required.
Files
iso20022-small.json / iso20022-small.csv — 100 rows (documents: 25)
iso20022-medium.json / iso20022-medium.csv — 2,000 rows (documents: 250)
iso20022-large.json / iso20022-large.csv — 20,000 rows (documents: 2,500)
Deterministic: the same size… See the full description on the dataset page: https://huggingface.co/datasets/StanzaAPI/iso20022-test-data.dialogsum-test
Dataset Card for DIALOGSum Corpus
Dataset Description
Links
Homepage: https://aclanthology.org/2021.findings-acl.449
Repository: https://github.com/cylnlp/dialogsum
Paper: https://aclanthology.org/2021.findings-acl.449
Point of Contact: https://huggingface.co/knkarthick
Dataset Summary
DialogSum is a large-scale dialogue summarization dataset, consisting of 13,460 (Plus 100 holdout data for topic generation) dialogues with corresponding… See the full description on the dataset page: https://huggingface.co/datasets/neil-code/dialogsum-test.epcis-test-data
EPCIS Test Data
ObjectEvent records in EPCIS 2.0 JSON-LD for track-and-trace tests, with synthetic SGTIN EPCs.
Free to use under CC0-1.0 — public domain dedication, no attribution required.
Files
epcis-small.json / epcis-small.csv — 100 rows (documents: 25)
epcis-medium.json / epcis-medium.csv — 2,000 rows (documents: 250)
epcis-large.json / epcis-large.csv — 20,000 rows (documents: 2,500)
Deterministic: the same size always produces the same rows.… See the full description on the dataset page: https://huggingface.co/datasets/StanzaAPI/epcis-test-data.peppol-test-data
Peppol BIS Test Data
Well-formed UBL Invoice documents with the Peppol BIS Billing 3.0 customization ID and synthetic parties.
Free to use under CC0-1.0 — public domain dedication, no attribution required.
Files
peppol-small.json / peppol-small.csv — 100 rows (documents: 25)
peppol-medium.json / peppol-medium.csv — 2,000 rows (documents: 250)
peppol-large.json / peppol-large.csv — 20,000 rows (documents: 2,500)
Deterministic: the same size always produces the… See the full description on the dataset page: https://huggingface.co/datasets/StanzaAPI/peppol-test-data.x12-test-data
ANSI X12 Test Data
Structurally complete X12 837P claim documents (ISA, GS, ST, BHT, NM1, CLM, SE, GE, IEA) with varying control numbers. Synthetic claims; not real patients or payers.
Free to use under CC0-1.0 — public domain dedication, no attribution required.
Files
x12-small.json / x12-small.csv — 100 rows (documents: 25)
x12-medium.json / x12-medium.csv — 2,000 rows (documents: 250)
x12-large.json / x12-large.csv — 20,000 rows (documents: 2,500)… See the full description on the dataset page: https://huggingface.co/datasets/StanzaAPI/x12-test-data.cbam-test-data
CBAM Test Data
Rows for CBAM reporting tests: declaration id, CN code, quantity, unit, embedded emissions, and country of origin. Synthetic values.
Free to use under CC0-1.0 — public domain dedication, no attribution required.
Files
cbam-small.json / cbam-small.csv — 100 rows (documents: 25)
cbam-medium.json / cbam-medium.csv — 2,000 rows (documents: 250)
cbam-large.json / cbam-large.csv — 20,000 rows (documents: 2,500)
Deterministic: the same size always… See the full description on the dataset page: https://huggingface.co/datasets/StanzaAPI/cbam-test-data.vat-test-data
VAT Test Data
National-format VAT numbers for every country in the VAT registry, with the 2026 standard rate. Format-shaped synthetic values for fixtures; not verified registrations.
Free to use under CC0-1.0 — public domain dedication, no attribution required.
Files
vat-small.json / vat-small.csv — 100 rows (documents: 25)
vat-medium.json / vat-medium.csv — 2,000 rows (documents: 250)
vat-large.json / vat-large.csv — 20,000 rows (documents: 2,500)… See the full description on the dataset page: https://huggingface.co/datasets/StanzaAPI/vat-test-data.map-testbias-test-gpt-sentences
Dataset Card for "BiasTestGPT: Generated Test Sentences"
Dataset of sentences for bias testing in open-sourced Pretrained Language Models generated using ChatGPT and other generative Language Models.
This dataset is used and actively populated by the BiasTestGPT HuggingFace Tool.
BiasTestGPT HuggingFace Tool
Dataset with Bias Specifications
Project Landing Page
Dataset Structure
The dataset is structured as a set of CSV files with names corresponding to the social… See the full description on the dataset page: https://huggingface.co/datasets/AnimaLab/bias-test-gpt-sentences.videophy2_testProject: https://github.com/Hritikbansal/videophy/tree/main/VIDEOPHY2
caption: original prompt in the dataset
video_url: generated video (using original prompt or upsampled caption, depending on the video model)
sa: semantic adherence score (1-5) from human evaluation
pc: physical commonsense score (1-5) from human evaluation
joint: computed as sa >= 4, pc >= 4
physics_rules_followed: list of physics rules followed in the video as judged by human annotators (1)
physics_rules_unfollowed: list… See the full description on the dataset page: https://huggingface.co/datasets/videophysics/videophy2_test.2025_Virtual_Cell_Challenge_Test_Dataiban-test-data
IBAN Test Data
One valid test IBAN per country, generated deterministically and checked with the ISO 7064 MOD-97 algorithm. None is a real account.
Free to use under CC0-1.0 — public domain dedication, no attribution required.
Files
iban-small.json / iban-small.csv — 100 rows (documents: 25)
iban-medium.json / iban-medium.csv — 2,000 rows (documents: 250)
iban-large.json / iban-large.csv — 20,000 rows (documents: 2,500)
Deterministic: the same size always… See the full description on the dataset page: https://huggingface.co/datasets/StanzaAPI/iban-test-data.gs1-gtin-test-data
GS1 GTIN Test Data
GS1 GTIN-14 identifiers with correctly computed MOD-10 check digits, plus a GS1 element string with batch and expiry (AI 01/10/17).
Free to use under CC0-1.0 — public domain dedication, no attribution required.
Files
gs1-gtin-small.json / gs1-gtin-small.csv — 100 rows (documents: 25)
gs1-gtin-medium.json / gs1-gtin-medium.csv — 2,000 rows (documents: 250)
gs1-gtin-large.json / gs1-gtin-large.csv — 20,000 rows (documents: 2,500)
Deterministic:… See the full description on the dataset page: https://huggingface.co/datasets/StanzaAPI/gs1-gtin-test-data.udi-test-data
UDI Test Data
Device identifier, lot, and serial strings in the GS1 element-string shape used by UDI carriers. Synthetic devices; not real products.
Free to use under CC0-1.0 — public domain dedication, no attribution required.
Files
udi-small.json / udi-small.csv — 100 rows (documents: 25)
udi-medium.json / udi-medium.csv — 2,000 rows (documents: 250)
udi-large.json / udi-large.csv — 20,000 rows (documents: 2,500)
Deterministic: the same size always produces… See the full description on the dataset page: https://huggingface.co/datasets/StanzaAPI/udi-test-data.test-dataset-v1container-test-data
Container Number Test Data
Shipping container numbers with a correctly computed ISO 6346 check digit. Owner codes are synthetic.
Free to use under CC0-1.0 — public domain dedication, no attribution required.
Files
container-small.json / container-small.csv — 100 rows (documents: 25)
container-medium.json / container-medium.csv — 2,000 rows (documents: 250)
container-large.json / container-large.csv — 20,000 rows (documents: 2,500)
Deterministic: the same size… See the full description on the dataset page: https://huggingface.co/datasets/StanzaAPI/container-test-data.casp14-casp15-cameo-test-proteinslei-test-data
LEI Test Data
Twenty-character LEIs with a correctly computed ISO 7064 MOD-97-10 check-digit pair. Synthetic entities; not registered organisations.
Free to use under CC0-1.0 — public domain dedication, no attribution required.
Files
lei-small.json / lei-small.csv — 100 rows (documents: 25)
lei-medium.json / lei-medium.csv — 2,000 rows (documents: 250)
lei-large.json / lei-large.csv — 20,000 rows (documents: 2,500)
Deterministic: the same size always produces… See the full description on the dataset page: https://huggingface.co/datasets/StanzaAPI/lei-test-data.test2upscale_board_data
