datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
swe-bench-dummy-test-datasettokenizers-test-data
tokenizers-test-data
Test and benchmark fixtures for huggingface/tokenizers,
pulled on demand by the repo Makefiles (make test / make bench / make fixtures
via hf download).
Layout
fixtures/ — multilingual + modality corpora for cross-language encode
benchmarks. Organized, documented, and reproducible: see
fixtures/FIXTURES.md for provenance and
fixtures/fixtures_manifest.json for
exact sources, pinned revisions, and sizes. Rebuild any file with… See the full description on the dataset page: https://huggingface.co/datasets/hf-internal-testing/tokenizers-test-data.Lora_Cloud_Dataset_Test
VLM Safety Inspector (2B / 4B / 8B) Mac 端评测与闭环套件
VLM Safety Inspector (2B / 4B / 8B) Mac 端闭环评测包
本目录是一个完全自包含(Self-Contained)的独立评测套件,专门适配您的 Mac(Apple Silicon / MPS)目录布局。
本目录是一个完全独立、自包含(Self-Contained)的评测套件,专为在 Mac (Apple Silicon / MPS) 上运行。
一、Mac 端文件布局自动识别(针对您的 iild 结构)
一、核心架构与流水线
评测脚本已内置针对您 Mac 端 iild/ 目录结构的全自动路径解析器:
在本次评测中,整条上行与闭环流水线严格遵循您的设想:
上游双塔一致性(In-Domain Consistency):
输入给 Planner 和 Inspector 的 150 个任务安全规则,已在 PC 端由纯 Legacy… See the full description on the dataset page: https://huggingface.co/datasets/lvesucces/Lora_Cloud_Dataset_Test.hindi_audio_dataset_testprotein_data_testsplit 1, 2 -> for sequences
split 3, 4 -> for residues
dataset-test-1smart-turn-data-v3.2-testTesting dataset for Smart Turn v3.2.
Thank you to the following contributors whose audio samples are included in this dataset:
The Pipecat team
Liva AI: https://www.theliva.ai/
Midcentury: https://www.midcentury.xyz/
MundoAI: https://mundoai.world/
Also, thank you to the following people for the CC-0 background noise sample data which has been used in this dataset:
https://freesound.org/people/4team/sounds/214995/
https://freesound.org/people/tomhannen/sounds/698090/… See the full description on the dataset page: https://huggingface.co/datasets/pipecat-ai/smart-turn-data-v3.2-test.test-datachain-llm-evalwhisperkit-test-datarlbenchfail_test_dataset
Guardian: RLBench-Fail Dataset
This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data generated in the RLBench simulator for training and evaluating Vision-Language Models (VLMs) on failure detection tasks.
Failures are produced by an automated pipeline that procedurally perturbs successful scripted trajectories in simulation, generating diverse planning… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/rlbenchfail_test_dataset.protein_data_test_2factur-x-test-data
Factur-X / ZUGFeRD Test Data
Well-formed UN/CEFACT Cross Industry Invoice XML with synthetic seller, buyer, and totals for e-invoicing tests.
Free to use under CC0-1.0 — public domain dedication, no attribution required.
Files
factur-x-small.json / factur-x-small.csv — 100 rows (documents: 25)
factur-x-medium.json / factur-x-medium.csv — 2,000 rows (documents: 250)
factur-x-large.json / factur-x-large.csv — 20,000 rows (documents: 2,500)
Deterministic: the same… See the full description on the dataset page: https://huggingface.co/datasets/StanzaAPI/factur-x-test-data.iso20022-test-data
ISO 20022 Test Data
Well-formed pain.001.001.09 Customer Credit Transfer Initiation XML with synthetic debtors, creditors, and amounts.
Free to use under CC0-1.0 — public domain dedication, no attribution required.
Files
iso20022-small.json / iso20022-small.csv — 100 rows (documents: 25)
iso20022-medium.json / iso20022-medium.csv — 2,000 rows (documents: 250)
iso20022-large.json / iso20022-large.csv — 20,000 rows (documents: 2,500)
Deterministic: the same size… See the full description on the dataset page: https://huggingface.co/datasets/StanzaAPI/iso20022-test-data.test-datasetsmart-product-pricing-2025_test_dataepcis-test-data
EPCIS Test Data
ObjectEvent records in EPCIS 2.0 JSON-LD for track-and-trace tests, with synthetic SGTIN EPCs.
Free to use under CC0-1.0 — public domain dedication, no attribution required.
Files
epcis-small.json / epcis-small.csv — 100 rows (documents: 25)
epcis-medium.json / epcis-medium.csv — 2,000 rows (documents: 250)
epcis-large.json / epcis-large.csv — 20,000 rows (documents: 2,500)
Deterministic: the same size always produces the same rows.… See the full description on the dataset page: https://huggingface.co/datasets/StanzaAPI/epcis-test-data.peppol-test-data
Peppol BIS Test Data
Well-formed UBL Invoice documents with the Peppol BIS Billing 3.0 customization ID and synthetic parties.
Free to use under CC0-1.0 — public domain dedication, no attribution required.
Files
peppol-small.json / peppol-small.csv — 100 rows (documents: 25)
peppol-medium.json / peppol-medium.csv — 2,000 rows (documents: 250)
peppol-large.json / peppol-large.csv — 20,000 rows (documents: 2,500)
Deterministic: the same size always produces the… See the full description on the dataset page: https://huggingface.co/datasets/StanzaAPI/peppol-test-data.x12-test-data
ANSI X12 Test Data
Structurally complete X12 837P claim documents (ISA, GS, ST, BHT, NM1, CLM, SE, GE, IEA) with varying control numbers. Synthetic claims; not real patients or payers.
Free to use under CC0-1.0 — public domain dedication, no attribution required.
Files
x12-small.json / x12-small.csv — 100 rows (documents: 25)
x12-medium.json / x12-medium.csv — 2,000 rows (documents: 250)
x12-large.json / x12-large.csv — 20,000 rows (documents: 2,500)… See the full description on the dataset page: https://huggingface.co/datasets/StanzaAPI/x12-test-data.testdataWideSeek-R1-test-data
Testing Dataset
We provide test.jsonl, a testing split for evaluating WideSeek-R1 on the standard WideSearch dataset. All examples are sourced from WideSearch; we only convert them into a format that is directly compatible with the WideSeek-R1 evaluation scripts. This makes the dataset plug-and-play—no additional configuration required.
test-big-dataset
Dataset Card for Danish WIT
Dataset Summary
Google presented the Wikipedia Image Text (WIT) dataset in July
2021, a dataset which contains
scraped images from Wikipedia along with their descriptions. WikiMedia released
WIT-Base in September
2021,
being a modified version of WIT where they have removed the images with empty
"reference descriptions", as well as removing images where a person's face covers more
than 10% of the image surface, along with inappropriate images… See the full description on the dataset page: https://huggingface.co/datasets/huggingface/test-big-dataset.cbam-test-data
CBAM Test Data
Rows for CBAM reporting tests: declaration id, CN code, quantity, unit, embedded emissions, and country of origin. Synthetic values.
Free to use under CC0-1.0 — public domain dedication, no attribution required.
Files
cbam-small.json / cbam-small.csv — 100 rows (documents: 25)
cbam-medium.json / cbam-medium.csv — 2,000 rows (documents: 250)
cbam-large.json / cbam-large.csv — 20,000 rows (documents: 2,500)
Deterministic: the same size always… See the full description on the dataset page: https://huggingface.co/datasets/StanzaAPI/cbam-test-data.test-data-set-Arabic-lettervat-test-data
VAT Test Data
National-format VAT numbers for every country in the VAT registry, with the 2026 standard rate. Format-shaped synthetic values for fixtures; not verified registrations.
Free to use under CC0-1.0 — public domain dedication, no attribution required.
Files
vat-small.json / vat-small.csv — 100 rows (documents: 25)
vat-medium.json / vat-medium.csv — 2,000 rows (documents: 250)
vat-large.json / vat-large.csv — 20,000 rows (documents: 2,500)… See the full description on the dataset page: https://huggingface.co/datasets/StanzaAPI/vat-test-data.dataset_test_disaggregated_nli
Dataset Card for "dataset_test_disaggregated_nli"
Dataset for testing a universal classifier. Additional information and training code available here: https://github.com/MoritzLaurer/zeroshot-classifier
Long-video-test-datachunking-test-data2025_Virtual_Cell_Challenge_Test_Dataiban-test-data
IBAN Test Data
One valid test IBAN per country, generated deterministically and checked with the ISO 7064 MOD-97 algorithm. None is a real account.
Free to use under CC0-1.0 — public domain dedication, no attribution required.
Files
iban-small.json / iban-small.csv — 100 rows (documents: 25)
iban-medium.json / iban-medium.csv — 2,000 rows (documents: 250)
iban-large.json / iban-large.csv — 20,000 rows (documents: 2,500)
Deterministic: the same size always… See the full description on the dataset page: https://huggingface.co/datasets/StanzaAPI/iban-test-data.gs1-gtin-test-data
GS1 GTIN Test Data
GS1 GTIN-14 identifiers with correctly computed MOD-10 check digits, plus a GS1 element string with batch and expiry (AI 01/10/17).
Free to use under CC0-1.0 — public domain dedication, no attribution required.
Files
gs1-gtin-small.json / gs1-gtin-small.csv — 100 rows (documents: 25)
gs1-gtin-medium.json / gs1-gtin-medium.csv — 2,000 rows (documents: 250)
gs1-gtin-large.json / gs1-gtin-large.csv — 20,000 rows (documents: 2,500)
Deterministic:… See the full description on the dataset page: https://huggingface.co/datasets/StanzaAPI/gs1-gtin-test-data.
