Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01klieret /swe-bench-dummy-test-datasettextn<1K0 likes50k downloads1y agoHugging Face02hf-internal-testing /tokenizers-test-data tokenizers-test-data Test and benchmark fixtures for huggingface/tokenizers, pulled on demand by the repo Makefiles (make test / make bench / make fixtures via hf download). Layout fixtures/ — multilingual + modality corpora for cross-language encode benchmarks. Organized, documented, and reproducible: see fixtures/FIXTURES.md for provenance and fixtures/fixtures_manifest.json for exact sources, pinned revisions, and sizes. Rebuild any file with… See the full description on the dataset page: https://huggingface.co/datasets/hf-internal-testing/tokenizers-test-data.textn<1K0 likes45k downloads12d agoHugging Face03lvesucces /Lora_Cloud_Dataset_Test VLM Safety Inspector (2B / 4B / 8B) Mac 端评测与闭环套件 VLM Safety Inspector (2B / 4B / 8B) Mac 端闭环评测包 本目录是一个完全自包含(Self-Contained)的独立评测套件,专门适配您的 Mac(Apple Silicon / MPS)目录布局。 本目录是一个完全独立、自包含(Self-Contained)的评测套件,专为在 Mac (Apple Silicon / MPS) 上运行。 一、Mac 端文件布局自动识别(针对您的 iild 结构) 一、核心架构与流水线 评测脚本已内置针对您 Mac 端 iild/ 目录结构的全自动路径解析器: 在本次评测中,整条上行与闭环流水线严格遵循您的设想: 上游双塔一致性(In-Domain Consistency): 输入给 Planner 和 Inspector 的 150 个任务安全规则,已在 PC 端由纯 Legacy… See the full description on the dataset page: https://huggingface.co/datasets/lvesucces/Lora_Cloud_Dataset_Test.imagen<1K0 likes2.8k downloads24d agoHugging Face04manavtabbly /hindi_audio_dataset_testaudion<1K0 likes2.6k downloads1y agoHugging Face05heispv /protein_data_testsplit 1, 2 -> for sequences split 3, 4 -> for residues textn<1K0 likes2.2k downloads2y agoHugging Face06lighteval-tests-datasets /dataset-test-1textn<1K0 likes1.8k downloads2y agoHugging Face07pipecat-ai /smart-turn-data-v3.2-testTesting dataset for Smart Turn v3.2. Thank you to the following contributors whose audio samples are included in this dataset: The Pipecat team Liva AI: https://www.theliva.ai/ Midcentury: https://www.midcentury.xyz/ MundoAI: https://mundoai.world/ Also, thank you to the following people for the CC-0 background noise sample data which has been used in this dataset: https://freesound.org/people/4team/sounds/214995/ https://freesound.org/people/tomhannen/sounds/698090/… See the full description on the dataset page: https://huggingface.co/datasets/pipecat-ai/smart-turn-data-v3.2-test.audio10K<n<100K3 likes1.2k downloads11d agoHugging Face08dvcorg /test-datachain-llm-evaltextn<1K1 likes1.1k downloads6h agoHugging Face09argmaxinc /whisperkit-test-dataaudion<1K0 likes893 downloads5mo agoHugging Face10paulpacaud /rlbenchfail_test_dataset Guardian: RLBench-Fail Dataset This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data generated in the RLBench simulator for training and evaluating Vision-Language Models (VLMs) on failure detection tasks. Failures are produced by an automated pipeline that procedurally perturbs successful scripted trajectories in simulation, generating diverse planning… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/rlbenchfail_test_dataset.tabularvisual-question-answering1K<n<10K0 likes892 downloads8mo agoHugging Face11heispv /protein_data_test_2textn<1K0 likes876 downloads2y agoHugging Face12StanzaAPI /factur-x-test-data Factur-X / ZUGFeRD Test Data Well-formed UN/CEFACT Cross Industry Invoice XML with synthetic seller, buyer, and totals for e-invoicing tests. Free to use under CC0-1.0 — public domain dedication, no attribution required. Files factur-x-small.json / factur-x-small.csv — 100 rows (documents: 25) factur-x-medium.json / factur-x-medium.csv — 2,000 rows (documents: 250) factur-x-large.json / factur-x-large.csv — 20,000 rows (documents: 2,500) Deterministic: the same… See the full description on the dataset page: https://huggingface.co/datasets/StanzaAPI/factur-x-test-data.textn<1K0 likes797 downloads8d agoHugging Face13StanzaAPI /iso20022-test-data ISO 20022 Test Data Well-formed pain.001.001.09 Customer Credit Transfer Initiation XML with synthetic debtors, creditors, and amounts. Free to use under CC0-1.0 — public domain dedication, no attribution required. Files iso20022-small.json / iso20022-small.csv — 100 rows (documents: 25) iso20022-medium.json / iso20022-medium.csv — 2,000 rows (documents: 250) iso20022-large.json / iso20022-large.csv — 20,000 rows (documents: 2,500) Deterministic: the same size… See the full description on the dataset page: https://huggingface.co/datasets/StanzaAPI/iso20022-test-data.textn<1K0 likes794 downloads8d agoHugging Face14plantcad /test-datasettextn<1K0 likes778 downloads1y agoHugging Face15kpkom /smart-product-pricing-2025_test_dataimage10K<n<100K0 likes777 downloads1y agoHugging Face16StanzaAPI /epcis-test-data EPCIS Test Data ObjectEvent records in EPCIS 2.0 JSON-LD for track-and-trace tests, with synthetic SGTIN EPCs. Free to use under CC0-1.0 — public domain dedication, no attribution required. Files epcis-small.json / epcis-small.csv — 100 rows (documents: 25) epcis-medium.json / epcis-medium.csv — 2,000 rows (documents: 250) epcis-large.json / epcis-large.csv — 20,000 rows (documents: 2,500) Deterministic: the same size always produces the same rows.… See the full description on the dataset page: https://huggingface.co/datasets/StanzaAPI/epcis-test-data.textn<1K0 likes760 downloads8d agoHugging Face17StanzaAPI /peppol-test-data Peppol BIS Test Data Well-formed UBL Invoice documents with the Peppol BIS Billing 3.0 customization ID and synthetic parties. Free to use under CC0-1.0 — public domain dedication, no attribution required. Files peppol-small.json / peppol-small.csv — 100 rows (documents: 25) peppol-medium.json / peppol-medium.csv — 2,000 rows (documents: 250) peppol-large.json / peppol-large.csv — 20,000 rows (documents: 2,500) Deterministic: the same size always produces the… See the full description on the dataset page: https://huggingface.co/datasets/StanzaAPI/peppol-test-data.textn<1K0 likes725 downloads8d agoHugging Face18StanzaAPI /x12-test-data ANSI X12 Test Data Structurally complete X12 837P claim documents (ISA, GS, ST, BHT, NM1, CLM, SE, GE, IEA) with varying control numbers. Synthetic claims; not real patients or payers. Free to use under CC0-1.0 — public domain dedication, no attribution required. Files x12-small.json / x12-small.csv — 100 rows (documents: 25) x12-medium.json / x12-medium.csv — 2,000 rows (documents: 250) x12-large.json / x12-large.csv — 20,000 rows (documents: 2,500)… See the full description on the dataset page: https://huggingface.co/datasets/StanzaAPI/x12-test-data.textn<1K0 likes690 downloads8d agoHugging Face19roitberg-group /testdataimagen<1K0 likes640 downloads1y agoHugging Face20WideSeek-R1 /WideSeek-R1-test-data Testing Dataset We provide test.jsonl, a testing split for evaluating WideSeek-R1 on the standard WideSearch dataset. All examples are sourced from WideSearch; we only convert them into a format that is directly compatible with the WideSeek-R1 evaluation scripts. This makes the dataset plug-and-play—no additional configuration required. texttext-generationn<1K0 likes640 downloads5mo agoHugging Face21huggingface /test-big-dataset Dataset Card for Danish WIT Dataset Summary Google presented the Wikipedia Image Text (WIT) dataset in July 2021, a dataset which contains scraped images from Wikipedia along with their descriptions. WikiMedia released WIT-Base in September 2021, being a modified version of WIT where they have removed the images with empty "reference descriptions", as well as removing images where a person's face covers more than 10% of the image surface, along with inappropriate images… See the full description on the dataset page: https://huggingface.co/datasets/huggingface/test-big-dataset.imageimage-to-text100K<n<1M0 likes607 downloads2y agoHugging Face22StanzaAPI /cbam-test-data CBAM Test Data Rows for CBAM reporting tests: declaration id, CN code, quantity, unit, embedded emissions, and country of origin. Synthetic values. Free to use under CC0-1.0 — public domain dedication, no attribution required. Files cbam-small.json / cbam-small.csv — 100 rows (documents: 25) cbam-medium.json / cbam-medium.csv — 2,000 rows (documents: 250) cbam-large.json / cbam-large.csv — 20,000 rows (documents: 2,500) Deterministic: the same size always… See the full description on the dataset page: https://huggingface.co/datasets/StanzaAPI/cbam-test-data.tabularn<1K0 likes597 downloads8d agoHugging Face23masumtechnonext /test-data-set-Arabic-letteraudio10K<n<100K0 likes585 downloads2mo agoHugging Face24StanzaAPI /vat-test-data VAT Test Data National-format VAT numbers for every country in the VAT registry, with the 2026 standard rate. Format-shaped synthetic values for fixtures; not verified registrations. Free to use under CC0-1.0 — public domain dedication, no attribution required. Files vat-small.json / vat-small.csv — 100 rows (documents: 25) vat-medium.json / vat-medium.csv — 2,000 rows (documents: 250) vat-large.json / vat-large.csv — 20,000 rows (documents: 2,500)… See the full description on the dataset page: https://huggingface.co/datasets/StanzaAPI/vat-test-data.textn<1K0 likes570 downloads8d agoHugging Face25MoritzLaurer /dataset_test_disaggregated_nli Dataset Card for "dataset_test_disaggregated_nli" Dataset for testing a universal classifier. Additional information and training code available here: https://github.com/MoritzLaurer/zeroshot-classifier text1M<n<10M1 likes544 downloads3y agoHugging Face26sfsdfsafsddsfsdafsa /Long-video-test-datatextn<1K2 likes523 downloads3y agoHugging Face27chunking-ai /chunking-test-dataaudion<1K0 likes507 downloads1y agoHugging Face28Arcticbun /2025_Virtual_Cell_Challenge_Test_Datatabularn<1K0 likes449 downloads5mo agoHugging Face29StanzaAPI /iban-test-data IBAN Test Data One valid test IBAN per country, generated deterministically and checked with the ISO 7064 MOD-97 algorithm. None is a real account. Free to use under CC0-1.0 — public domain dedication, no attribution required. Files iban-small.json / iban-small.csv — 100 rows (documents: 25) iban-medium.json / iban-medium.csv — 2,000 rows (documents: 250) iban-large.json / iban-large.csv — 20,000 rows (documents: 2,500) Deterministic: the same size always… See the full description on the dataset page: https://huggingface.co/datasets/StanzaAPI/iban-test-data.textn<1K0 likes449 downloads8d agoHugging Face30StanzaAPI /gs1-gtin-test-data GS1 GTIN Test Data GS1 GTIN-14 identifiers with correctly computed MOD-10 check digits, plus a GS1 element string with batch and expiry (AI 01/10/17). Free to use under CC0-1.0 — public domain dedication, no attribution required. Files gs1-gtin-small.json / gs1-gtin-small.csv — 100 rows (documents: 25) gs1-gtin-medium.json / gs1-gtin-medium.csv — 2,000 rows (documents: 250) gs1-gtin-large.json / gs1-gtin-large.csv — 20,000 rows (documents: 2,500) Deterministic:… See the full description on the dataset page: https://huggingface.co/datasets/StanzaAPI/gs1-gtin-test-data.textn<1K0 likes425 downloads8d agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.