Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01heispv /protein_data_testsplit 1, 2 -> for sequences split 3, 4 -> for residues textn<1K0 likes2.2k downloads2y agoHugging Face02heispv /protein_data_test_2textn<1K0 likes876 downloads2y agoHugging Face03StanzaAPI /factur-x-test-data Factur-X / ZUGFeRD Test Data Well-formed UN/CEFACT Cross Industry Invoice XML with synthetic seller, buyer, and totals for e-invoicing tests. Free to use under CC0-1.0 — public domain dedication, no attribution required. Files factur-x-small.json / factur-x-small.csv — 100 rows (documents: 25) factur-x-medium.json / factur-x-medium.csv — 2,000 rows (documents: 250) factur-x-large.json / factur-x-large.csv — 20,000 rows (documents: 2,500) Deterministic: the same… See the full description on the dataset page: https://huggingface.co/datasets/StanzaAPI/factur-x-test-data.textn<1K0 likes797 downloads9d agoHugging Face04StanzaAPI /iso20022-test-data ISO 20022 Test Data Well-formed pain.001.001.09 Customer Credit Transfer Initiation XML with synthetic debtors, creditors, and amounts. Free to use under CC0-1.0 — public domain dedication, no attribution required. Files iso20022-small.json / iso20022-small.csv — 100 rows (documents: 25) iso20022-medium.json / iso20022-medium.csv — 2,000 rows (documents: 250) iso20022-large.json / iso20022-large.csv — 20,000 rows (documents: 2,500) Deterministic: the same size… See the full description on the dataset page: https://huggingface.co/datasets/StanzaAPI/iso20022-test-data.textn<1K0 likes794 downloads9d agoHugging Face05StanzaAPI /epcis-test-data EPCIS Test Data ObjectEvent records in EPCIS 2.0 JSON-LD for track-and-trace tests, with synthetic SGTIN EPCs. Free to use under CC0-1.0 — public domain dedication, no attribution required. Files epcis-small.json / epcis-small.csv — 100 rows (documents: 25) epcis-medium.json / epcis-medium.csv — 2,000 rows (documents: 250) epcis-large.json / epcis-large.csv — 20,000 rows (documents: 2,500) Deterministic: the same size always produces the same rows.… See the full description on the dataset page: https://huggingface.co/datasets/StanzaAPI/epcis-test-data.textn<1K0 likes760 downloads9d agoHugging Face06StanzaAPI /peppol-test-data Peppol BIS Test Data Well-formed UBL Invoice documents with the Peppol BIS Billing 3.0 customization ID and synthetic parties. Free to use under CC0-1.0 — public domain dedication, no attribution required. Files peppol-small.json / peppol-small.csv — 100 rows (documents: 25) peppol-medium.json / peppol-medium.csv — 2,000 rows (documents: 250) peppol-large.json / peppol-large.csv — 20,000 rows (documents: 2,500) Deterministic: the same size always produces the… See the full description on the dataset page: https://huggingface.co/datasets/StanzaAPI/peppol-test-data.textn<1K0 likes725 downloads9d agoHugging Face07StanzaAPI /x12-test-data ANSI X12 Test Data Structurally complete X12 837P claim documents (ISA, GS, ST, BHT, NM1, CLM, SE, GE, IEA) with varying control numbers. Synthetic claims; not real patients or payers. Free to use under CC0-1.0 — public domain dedication, no attribution required. Files x12-small.json / x12-small.csv — 100 rows (documents: 25) x12-medium.json / x12-medium.csv — 2,000 rows (documents: 250) x12-large.json / x12-large.csv — 20,000 rows (documents: 2,500)… See the full description on the dataset page: https://huggingface.co/datasets/StanzaAPI/x12-test-data.textn<1K0 likes690 downloads9d agoHugging Face08StanzaAPI /cbam-test-data CBAM Test Data Rows for CBAM reporting tests: declaration id, CN code, quantity, unit, embedded emissions, and country of origin. Synthetic values. Free to use under CC0-1.0 — public domain dedication, no attribution required. Files cbam-small.json / cbam-small.csv — 100 rows (documents: 25) cbam-medium.json / cbam-medium.csv — 2,000 rows (documents: 250) cbam-large.json / cbam-large.csv — 20,000 rows (documents: 2,500) Deterministic: the same size always… See the full description on the dataset page: https://huggingface.co/datasets/StanzaAPI/cbam-test-data.tabularn<1K0 likes597 downloads9d agoHugging Face09StanzaAPI /vat-test-data VAT Test Data National-format VAT numbers for every country in the VAT registry, with the 2026 standard rate. Format-shaped synthetic values for fixtures; not verified registrations. Free to use under CC0-1.0 — public domain dedication, no attribution required. Files vat-small.json / vat-small.csv — 100 rows (documents: 25) vat-medium.json / vat-medium.csv — 2,000 rows (documents: 250) vat-large.json / vat-large.csv — 20,000 rows (documents: 2,500)… See the full description on the dataset page: https://huggingface.co/datasets/StanzaAPI/vat-test-data.textn<1K0 likes570 downloads9d agoHugging Face10Arcticbun /2025_Virtual_Cell_Challenge_Test_Datatabularn<1K0 likes449 downloads5mo agoHugging Face11StanzaAPI /iban-test-data IBAN Test Data One valid test IBAN per country, generated deterministically and checked with the ISO 7064 MOD-97 algorithm. None is a real account. Free to use under CC0-1.0 — public domain dedication, no attribution required. Files iban-small.json / iban-small.csv — 100 rows (documents: 25) iban-medium.json / iban-medium.csv — 2,000 rows (documents: 250) iban-large.json / iban-large.csv — 20,000 rows (documents: 2,500) Deterministic: the same size always… See the full description on the dataset page: https://huggingface.co/datasets/StanzaAPI/iban-test-data.textn<1K0 likes449 downloads9d agoHugging Face12StanzaAPI /gs1-gtin-test-data GS1 GTIN Test Data GS1 GTIN-14 identifiers with correctly computed MOD-10 check digits, plus a GS1 element string with batch and expiry (AI 01/10/17). Free to use under CC0-1.0 — public domain dedication, no attribution required. Files gs1-gtin-small.json / gs1-gtin-small.csv — 100 rows (documents: 25) gs1-gtin-medium.json / gs1-gtin-medium.csv — 2,000 rows (documents: 250) gs1-gtin-large.json / gs1-gtin-large.csv — 20,000 rows (documents: 2,500) Deterministic:… See the full description on the dataset page: https://huggingface.co/datasets/StanzaAPI/gs1-gtin-test-data.textn<1K0 likes425 downloads9d agoHugging Face13StanzaAPI /udi-test-data UDI Test Data Device identifier, lot, and serial strings in the GS1 element-string shape used by UDI carriers. Synthetic devices; not real products. Free to use under CC0-1.0 — public domain dedication, no attribution required. Files udi-small.json / udi-small.csv — 100 rows (documents: 25) udi-medium.json / udi-medium.csv — 2,000 rows (documents: 250) udi-large.json / udi-large.csv — 20,000 rows (documents: 2,500) Deterministic: the same size always produces… See the full description on the dataset page: https://huggingface.co/datasets/StanzaAPI/udi-test-data.textn<1K0 likes386 downloads9d agoHugging Face14codezerro /test-dataset-v1image100K<n<1M0 likes381 downloads1y agoHugging Face15StanzaAPI /container-test-data Container Number Test Data Shipping container numbers with a correctly computed ISO 6346 check digit. Owner codes are synthetic. Free to use under CC0-1.0 — public domain dedication, no attribution required. Files container-small.json / container-small.csv — 100 rows (documents: 25) container-medium.json / container-medium.csv — 2,000 rows (documents: 250) container-large.json / container-large.csv — 20,000 rows (documents: 2,500) Deterministic: the same size… See the full description on the dataset page: https://huggingface.co/datasets/StanzaAPI/container-test-data.tabularn<1K0 likes345 downloads9d agoHugging Face16StanzaAPI /lei-test-data LEI Test Data Twenty-character LEIs with a correctly computed ISO 7064 MOD-97-10 check-digit pair. Synthetic entities; not registered organisations. Free to use under CC0-1.0 — public domain dedication, no attribution required. Files lei-small.json / lei-small.csv — 100 rows (documents: 25) lei-medium.json / lei-medium.csv — 2,000 rows (documents: 250) lei-large.json / lei-large.csv — 20,000 rows (documents: 2,500) Deterministic: the same size always produces… See the full description on the dataset page: https://huggingface.co/datasets/StanzaAPI/lei-test-data.textn<1K0 likes319 downloads9d agoHugging Face17openadmet /openadmet-expansionrx-challenge-test-data-blinded OpenADMET-ExpansionRx Challenge blinded test dataset This dataset contains real-work ADMET data from a recently prosecuted series of drug discovery campaigns by Expansion Therapeutics on RNA mediated diseases. While optimising candidate molecules for their preclinical programs Expansion collected a variety of ADMET data for off-targets and properties of interest in the traditional game of “whack-a-mole” familiar to all drug hunters. Now, they’ve made the bold and generous decision… See the full description on the dataset page: https://huggingface.co/datasets/openadmet/openadmet-expansionrx-challenge-test-data-blinded.text1K<n<10K5 likes229 downloads1y agoHugging Face18abidlabs /test-translation-datasettextn<1K0 likes178 downloads5y agoHugging Face19BigDocs /test_tsv_datasettext10K<n<100K0 likes147 downloads2y agoHugging Face20iulusoy /test-datatexttext-classificationn<1K0 likes67 downloads3y agoHugging Face21roberto-armas /ml_data_test_detection_bank_transaction_frauds_unbalanced ML Data Test Detection Bank Transaction Frauds Unbalanced The project provides a quick and accessible dataset designed for learning and experimenting with machine learning algorithms, specifically in the context of detecting fraudulent bank transactions. It is intended for practicing and applying concepts such as Random Forest, Support Vector Machines (SVM), and Synthetic Minority Over-sampling Technique (SMOTE) to address unbalanced classification problems. Note: This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/roberto-armas/ml_data_test_detection_bank_transaction_frauds_unbalanced.tabular1K<n<10K2 likes64 downloads1y agoHugging Face22luckeciano /pku-llama3.1-8b-dataset-test-generationstext1M<n<10M0 likes52 downloads2y agoHugging Face23Youlln /Dataset-Achats-Test Dataset d'Achats Départementaux 🏛️ Un dataset de classification d'intitulés d'achats pour les départements français, généré automatiquement via l'API OpenAI. 📊 Description Ce dataset contient 10 000 exemples d'intitulés d'achats répartis en 200 catégories typiques des marchés publics départementaux français. Structure des données Format : CSV avec colonnes text et label Langue : Français Domaine : Commande publique départementale Exemples par catégorie : 50… See the full description on the dataset page: https://huggingface.co/datasets/Youlln/Dataset-Achats-Test.text10K<n<100K0 likes52 downloads1y agoHugging Face24open-neo /open-neo-public-test-data Open-Neo HCC1395/HCC1395BL Public WGS Test Data This dataset provides a public paired whole-genome test case for validating the Open-Neo installation, input-QC, DNA evidence, purity/CNV, HLA, candidate generation, presentation, ranking, and reporting paths. Samples Role Sample BioSample Source material Tumor HCC1395 (WGS_EA_T_1) SAMN10102573 Breast carcinoma cell line, ATCC CRL-2324 Matched normal HCC1395BL (WGS_EA_N_1) SAMN10102574 B-lymphoblast cell… See the full description on the dataset page: https://huggingface.co/datasets/open-neo/open-neo-public-test-data.textothern<1K0 likes49 downloads1mo agoHugging Face25landogayatri /hf-free-maxxing-test-datatabularn<1K0 likes39 downloads17d agoHugging Face260xmoose0xmoose0xmoose /xss-data-viewer-test XSS Test Dataset texttext-classificationn<1K0 likes36 downloads26d agoHugging Face27ApertureQA /testdataset-hdupteststst textn<1K0 likes34 downloads8mo agoHugging Face28hf-sec-research-1 /sec-data-url-test Image URL Test imagen<1K0 likes32 downloads3d agoHugging Face29parnoux /hate_speech_open_data_original_class_test_settabulartext-classification1K<n<10K1 likes30 downloads4y agoHugging Face30krishan-CSE /Sinhala_dataset_Testtext1K<n<10K0 likes28 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.