datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
protein_data_testsplit 1, 2 -> for sequences
split 3, 4 -> for residues
protein_data_test_2factur-x-test-data
Factur-X / ZUGFeRD Test Data
Well-formed UN/CEFACT Cross Industry Invoice XML with synthetic seller, buyer, and totals for e-invoicing tests.
Free to use under CC0-1.0 — public domain dedication, no attribution required.
Files
factur-x-small.json / factur-x-small.csv — 100 rows (documents: 25)
factur-x-medium.json / factur-x-medium.csv — 2,000 rows (documents: 250)
factur-x-large.json / factur-x-large.csv — 20,000 rows (documents: 2,500)
Deterministic: the same… See the full description on the dataset page: https://huggingface.co/datasets/StanzaAPI/factur-x-test-data.iso20022-test-data
ISO 20022 Test Data
Well-formed pain.001.001.09 Customer Credit Transfer Initiation XML with synthetic debtors, creditors, and amounts.
Free to use under CC0-1.0 — public domain dedication, no attribution required.
Files
iso20022-small.json / iso20022-small.csv — 100 rows (documents: 25)
iso20022-medium.json / iso20022-medium.csv — 2,000 rows (documents: 250)
iso20022-large.json / iso20022-large.csv — 20,000 rows (documents: 2,500)
Deterministic: the same size… See the full description on the dataset page: https://huggingface.co/datasets/StanzaAPI/iso20022-test-data.epcis-test-data
EPCIS Test Data
ObjectEvent records in EPCIS 2.0 JSON-LD for track-and-trace tests, with synthetic SGTIN EPCs.
Free to use under CC0-1.0 — public domain dedication, no attribution required.
Files
epcis-small.json / epcis-small.csv — 100 rows (documents: 25)
epcis-medium.json / epcis-medium.csv — 2,000 rows (documents: 250)
epcis-large.json / epcis-large.csv — 20,000 rows (documents: 2,500)
Deterministic: the same size always produces the same rows.… See the full description on the dataset page: https://huggingface.co/datasets/StanzaAPI/epcis-test-data.peppol-test-data
Peppol BIS Test Data
Well-formed UBL Invoice documents with the Peppol BIS Billing 3.0 customization ID and synthetic parties.
Free to use under CC0-1.0 — public domain dedication, no attribution required.
Files
peppol-small.json / peppol-small.csv — 100 rows (documents: 25)
peppol-medium.json / peppol-medium.csv — 2,000 rows (documents: 250)
peppol-large.json / peppol-large.csv — 20,000 rows (documents: 2,500)
Deterministic: the same size always produces the… See the full description on the dataset page: https://huggingface.co/datasets/StanzaAPI/peppol-test-data.x12-test-data
ANSI X12 Test Data
Structurally complete X12 837P claim documents (ISA, GS, ST, BHT, NM1, CLM, SE, GE, IEA) with varying control numbers. Synthetic claims; not real patients or payers.
Free to use under CC0-1.0 — public domain dedication, no attribution required.
Files
x12-small.json / x12-small.csv — 100 rows (documents: 25)
x12-medium.json / x12-medium.csv — 2,000 rows (documents: 250)
x12-large.json / x12-large.csv — 20,000 rows (documents: 2,500)… See the full description on the dataset page: https://huggingface.co/datasets/StanzaAPI/x12-test-data.cbam-test-data
CBAM Test Data
Rows for CBAM reporting tests: declaration id, CN code, quantity, unit, embedded emissions, and country of origin. Synthetic values.
Free to use under CC0-1.0 — public domain dedication, no attribution required.
Files
cbam-small.json / cbam-small.csv — 100 rows (documents: 25)
cbam-medium.json / cbam-medium.csv — 2,000 rows (documents: 250)
cbam-large.json / cbam-large.csv — 20,000 rows (documents: 2,500)
Deterministic: the same size always… See the full description on the dataset page: https://huggingface.co/datasets/StanzaAPI/cbam-test-data.vat-test-data
VAT Test Data
National-format VAT numbers for every country in the VAT registry, with the 2026 standard rate. Format-shaped synthetic values for fixtures; not verified registrations.
Free to use under CC0-1.0 — public domain dedication, no attribution required.
Files
vat-small.json / vat-small.csv — 100 rows (documents: 25)
vat-medium.json / vat-medium.csv — 2,000 rows (documents: 250)
vat-large.json / vat-large.csv — 20,000 rows (documents: 2,500)… See the full description on the dataset page: https://huggingface.co/datasets/StanzaAPI/vat-test-data.2025_Virtual_Cell_Challenge_Test_Dataiban-test-data
IBAN Test Data
One valid test IBAN per country, generated deterministically and checked with the ISO 7064 MOD-97 algorithm. None is a real account.
Free to use under CC0-1.0 — public domain dedication, no attribution required.
Files
iban-small.json / iban-small.csv — 100 rows (documents: 25)
iban-medium.json / iban-medium.csv — 2,000 rows (documents: 250)
iban-large.json / iban-large.csv — 20,000 rows (documents: 2,500)
Deterministic: the same size always… See the full description on the dataset page: https://huggingface.co/datasets/StanzaAPI/iban-test-data.gs1-gtin-test-data
GS1 GTIN Test Data
GS1 GTIN-14 identifiers with correctly computed MOD-10 check digits, plus a GS1 element string with batch and expiry (AI 01/10/17).
Free to use under CC0-1.0 — public domain dedication, no attribution required.
Files
gs1-gtin-small.json / gs1-gtin-small.csv — 100 rows (documents: 25)
gs1-gtin-medium.json / gs1-gtin-medium.csv — 2,000 rows (documents: 250)
gs1-gtin-large.json / gs1-gtin-large.csv — 20,000 rows (documents: 2,500)
Deterministic:… See the full description on the dataset page: https://huggingface.co/datasets/StanzaAPI/gs1-gtin-test-data.test-dataset-v1udi-test-data
UDI Test Data
Device identifier, lot, and serial strings in the GS1 element-string shape used by UDI carriers. Synthetic devices; not real products.
Free to use under CC0-1.0 — public domain dedication, no attribution required.
Files
udi-small.json / udi-small.csv — 100 rows (documents: 25)
udi-medium.json / udi-medium.csv — 2,000 rows (documents: 250)
udi-large.json / udi-large.csv — 20,000 rows (documents: 2,500)
Deterministic: the same size always produces… See the full description on the dataset page: https://huggingface.co/datasets/StanzaAPI/udi-test-data.iata-awb-test-data
IATA Air Waybill Test Data
Air waybill numbers combining a 3-digit airline prefix, a 7-digit serial, and a MOD-7 check digit (IATA Resolution 600a).
Free to use under CC0-1.0 — public domain dedication, no attribution required.
Files
iata-awb-small.json / iata-awb-small.csv — 100 rows (documents: 25)
iata-awb-medium.json / iata-awb-medium.csv — 2,000 rows (documents: 250)
iata-awb-large.json / iata-awb-large.csv — 20,000 rows (documents: 2,500)
Deterministic:… See the full description on the dataset page: https://huggingface.co/datasets/StanzaAPI/iata-awb-test-data.container-test-data
Container Number Test Data
Shipping container numbers with a correctly computed ISO 6346 check digit. Owner codes are synthetic.
Free to use under CC0-1.0 — public domain dedication, no attribution required.
Files
container-small.json / container-small.csv — 100 rows (documents: 25)
container-medium.json / container-medium.csv — 2,000 rows (documents: 250)
container-large.json / container-large.csv — 20,000 rows (documents: 2,500)
Deterministic: the same size… See the full description on the dataset page: https://huggingface.co/datasets/StanzaAPI/container-test-data.lei-test-data
LEI Test Data
Twenty-character LEIs with a correctly computed ISO 7064 MOD-97-10 check-digit pair. Synthetic entities; not registered organisations.
Free to use under CC0-1.0 — public domain dedication, no attribution required.
Files
lei-small.json / lei-small.csv — 100 rows (documents: 25)
lei-medium.json / lei-medium.csv — 2,000 rows (documents: 250)
lei-large.json / lei-large.csv — 20,000 rows (documents: 2,500)
Deterministic: the same size always produces… See the full description on the dataset page: https://huggingface.co/datasets/StanzaAPI/lei-test-data.test-translation-datasetopenadmet-expansionrx-challenge-test-data-blinded
OpenADMET-ExpansionRx Challenge blinded test dataset
This dataset contains real-work ADMET data from a recently prosecuted series of drug discovery campaigns by Expansion Therapeutics on RNA mediated diseases. While optimising candidate molecules for their preclinical programs Expansion collected a variety of ADMET data for off-targets and properties of interest in the traditional game of “whack-a-mole” familiar to all drug hunters. Now, they’ve made the bold and generous decision… See the full description on the dataset page: https://huggingface.co/datasets/openadmet/openadmet-expansionrx-challenge-test-data-blinded.test_tsv_datasettest-dataopen-neo-public-test-data
Open-Neo HCC1395/HCC1395BL Public WGS Test Data
This dataset provides a public paired whole-genome test case for validating the
Open-Neo installation, input-QC, DNA evidence, purity/CNV, HLA, candidate
generation, presentation, ranking, and reporting paths.
Samples
Role
Sample
BioSample
Source material
Tumor
HCC1395 (WGS_EA_T_1)
SAMN10102573
Breast carcinoma cell line, ATCC CRL-2324
Matched normal
HCC1395BL (WGS_EA_N_1)
SAMN10102574
B-lymphoblast cell… See the full description on the dataset page: https://huggingface.co/datasets/open-neo/open-neo-public-test-data.ml_data_test_detection_bank_transaction_frauds_unbalanced
ML Data Test Detection Bank Transaction Frauds Unbalanced
The project provides a quick and accessible dataset designed for learning and experimenting with machine learning algorithms, specifically in the context of detecting fraudulent bank transactions. It is intended for practicing and applying concepts such as Random Forest, Support Vector Machines (SVM), and Synthetic Minority Over-sampling Technique (SMOTE) to address unbalanced classification problems.
Note: This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/roberto-armas/ml_data_test_detection_bank_transaction_frauds_unbalanced.southern-cross-ai-test-datasetstest-datasetpku-llama3.1-8b-dataset-test-generationsDataset-Achats-Test
Dataset d'Achats Départementaux 🏛️
Un dataset de classification d'intitulés d'achats pour les départements français, généré automatiquement via l'API OpenAI.
📊 Description
Ce dataset contient 10 000 exemples d'intitulés d'achats répartis en 200 catégories typiques des marchés publics départementaux français.
Structure des données
Format : CSV avec colonnes text et label
Langue : Français
Domaine : Commande publique départementale
Exemples par catégorie : 50… See the full description on the dataset page: https://huggingface.co/datasets/Youlln/Dataset-Achats-Test.hf-free-maxxing-test-dataSinhala_dataset_Testxss-data-viewer-test
XSS Test Dataset
