datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cyp-challenge-train-test
CYP Challenge Train/Test Dataset
A high-quality experimental dataset for predicting inhibition of the major drug-metabolizing Cytochrome P450 enzymes (CYP1A2, CYP2C9, CYP2D6, CYP3A4), released as part of the OpenADMET CYP Inhibition Blind Challenge.
Blog post: Announcing OpenADMET’s CYP inhibition blind challenge
Challenge Space: OpenADMET CYP Inhibition Blind Challenge
Challenge period: August 17, 2026 - November 3, 2026
Produced by: OpenADMET
CHANGELOG
Updated… See the full description on the dataset page: https://huggingface.co/datasets/openadmet/cyp-challenge-train-test.pxr-challenge-train-test
PXR Challenge Train/Test Dataset
A high-quality experimental dataset for predicting human Pregnane-X Receptor (PXR) induction, comprising over 11,000 compounds screened using a high-fidelity in-house assay. This is the largest publicly available PXR activity dataset, released as part of the OpenADMET PXR Induction Blind Challenge.
Blog post: Announcing the Next OpenADMET Blind Challenge: Predicting PXR Induction
Challenge Space: openadmet/pxr-challenge
Challenge period: April 1… See the full description on the dataset page: https://huggingface.co/datasets/openadmet/pxr-challenge-train-test.vindr-cxr-testsetcbam-test-data
CBAM Test Data
Rows for CBAM reporting tests: declaration id, CN code, quantity, unit, embedded emissions, and country of origin. Synthetic values.
Free to use under CC0-1.0 — public domain dedication, no attribution required.
Files
cbam-small.json / cbam-small.csv — 100 rows (documents: 25)
cbam-medium.json / cbam-medium.csv — 2,000 rows (documents: 250)
cbam-large.json / cbam-large.csv — 20,000 rows (documents: 2,500)
Deterministic: the same size always… See the full description on the dataset page: https://huggingface.co/datasets/StanzaAPI/cbam-test-data.videophy2_testProject: https://github.com/Hritikbansal/videophy/tree/main/VIDEOPHY2
caption: original prompt in the dataset
video_url: generated video (using original prompt or upsampled caption, depending on the video model)
sa: semantic adherence score (1-5) from human evaluation
pc: physical commonsense score (1-5) from human evaluation
joint: computed as sa >= 4, pc >= 4
physics_rules_followed: list of physics rules followed in the video as judged by human annotators (1)
physics_rules_unfollowed: list… See the full description on the dataset page: https://huggingface.co/datasets/videophysics/videophy2_test.2025_Virtual_Cell_Challenge_Test_Datatest-dataset-v1iata-awb-test-data
IATA Air Waybill Test Data
Air waybill numbers combining a 3-digit airline prefix, a 7-digit serial, and a MOD-7 check digit (IATA Resolution 600a).
Free to use under CC0-1.0 — public domain dedication, no attribution required.
Files
iata-awb-small.json / iata-awb-small.csv — 100 rows (documents: 25)
iata-awb-medium.json / iata-awb-medium.csv — 2,000 rows (documents: 250)
iata-awb-large.json / iata-awb-large.csv — 20,000 rows (documents: 2,500)
Deterministic:… See the full description on the dataset page: https://huggingface.co/datasets/StanzaAPI/iata-awb-test-data.container-test-data
Container Number Test Data
Shipping container numbers with a correctly computed ISO 6346 check digit. Owner codes are synthetic.
Free to use under CC0-1.0 — public domain dedication, no attribution required.
Files
container-small.json / container-small.csv — 100 rows (documents: 25)
container-medium.json / container-medium.csv — 2,000 rows (documents: 250)
container-large.json / container-large.csv — 20,000 rows (documents: 2,500)
Deterministic: the same size… See the full description on the dataset page: https://huggingface.co/datasets/StanzaAPI/container-test-data.FARM_training_test
FARM Aerial Radio Map (ARM) Dataset
Paper:
FARM: Foundational Aerial Radio Map for Intelligent Low-Altitude Networking (https://arxiv.org/abs/2604.17362)
Overview
This repository releases the constructed ARM datasets based on ARM-Omni for FARM training, in-domain evaluation (D1-D10), and zero-shot evaluation (P1, F1, and A1). The dataset coverage is summarized below:
Dataset
Frequencies (GHz)
Max Rx Height (m)
Beamwidths
Map Grid Size
Volume
D1
2.1… See the full description on the dataset page: https://huggingface.co/datasets/jliang097/FARM_training_test.upscale_board_datavideophy_test_publicWe have uploaded the videos at: https://huggingface.co/videophysics/videophy-test-videos/tree/main
For more details, please visit:
project github: https://github.com/Hritikbansal/videophy
project website: https://videophy.github.io/
testing-goldstandard-cuthill
Dataset Card for Curated Gold Standard Hoyal Cuthill Dataset
Dataset Description
Dorsal full body images of subspecies of Heliconius erato and Heliconius melpomene (18 subspecies total).
There are 960 images with 320 specimens (3 images of each specimen: Original/ Bird transformed/ Butterfly transformed)
The original images are low-resolution RGB photographs (photographs were "cropped and resized to a height of 64 pixels (maintaining the original image aspect ratio and… See the full description on the dataset page: https://huggingface.co/datasets/jrw2989/testing-goldstandard-cuthill.juliet_test_suite_c_1_3
Dataset Card for the Juliet Test Suite 1.3
Dataset Summary
This Datasets contains all test cases from the NIST's Juliet test suite for the C and C++ programming languages. The dataset contains a benign and a defective implementation of each sample, which have been extracting by means of the OMITGOOD and OMITBAD preprocessor macros of the Juliet test suite.
Supported Tasks and Leaderboards
Software defect prediction, code clone detection.
Languages… See the full description on the dataset page: https://huggingface.co/datasets/LorenzH/juliet_test_suite_c_1_3.Dense-Evolution-Ising-Tests
🔬 Quantum Phase Transitions, Variational Gradients, and Error Mitigation
This repository contains a rigorous empirical study, raw datasets, and quantum error mitigation protocols executed on Dense Evolution—a high-performance Statevector quantum simulator. Utilizing 64-bit double precision (complex128) and hardware-accelerated static compilation via the JAX XLA engine, this project maps the non-linear physics of the Transverse Field Ising Model (TFIM) and Tight-Binding… See the full description on the dataset page: https://huggingface.co/datasets/Tatopenn/Dense-Evolution-Ising-Tests.patents_claims_1.5m_traim_testTest_Semantic_Searcharc_agi_2_human_testing
ARC-AGI-2 Human testing data
This file contains data from human testing sessions on ARC-AGI tasks.
Each row represents a single test attempt by a human participant on a specific task-test pair in the "Public Train" or "Public Eval" ARC-AGI-2 datasets. Not all tasks in the released "Public Train"
sets were tested, so these results are not comprehensive. This data does not include tasks from "Semi Private Evaluation" or "Private Evaluation"
Column Descriptions… See the full description on the dataset page: https://huggingface.co/datasets/arcprize/arc_agi_2_human_testing.Obj_testargilla-invalid-rowsafrimgsm-translate-test
Dataset Card for afrimgsm-translate-test
Dataset Summary
AFRIMGSM-TT is an evaluation dataset comprising translations of the GSM8k dataset from 16 African languages and 1 high resource language into English using NLLB.
It includes test sets across all 17 languages.
Languages
There are 17 languages available :
Dataset Structure
Data Instances
The examples look like this for English:
from datasets import load_dataset
data =… See the full description on the dataset page: https://huggingface.co/datasets/masakhane/afrimgsm-translate-test.Dr.Sparse-OTF-test-set
Dr.Sparse OTF Test Set
100 sparse matrices from the SuiteSparse Matrix Collection,
converted to the flat binary format the Dr.Sparse
benchmark harness reads. This is the held-out evaluation set for LLM-generated
CUDA sparse kernels (SpMV / SpMM / SpGEMM), kept separate from the matrices the
models were developed against.
Layout
Matrices are grouped into size tiers by row count, the convention Dr.Sparse task
discovery scans for:
tier
rows
matrices
size… See the full description on the dataset page: https://huggingface.co/datasets/KinGeorge/Dr.Sparse-OTF-test-set.dimensional-weight-calculator-test-cases
Dimensional Weight Calculator Test Cases
This small tabular dataset is designed for testing dimensional-weight calculator implementations. It covers ordinary calculations, comparison and rounding boundaries, unit labels, exact ties, alternative divisors, dimension-order permutations, and invalid-input handling.
The data is a software-test resource, not a collection of observed shipments. It contains no customer, order, seller, product, price, inventory, account, or carrier-rate… See the full description on the dataset page: https://huggingface.co/datasets/hang008613950785007/dimensional-weight-calculator-test-cases.ml_data_test_detection_bank_transaction_frauds_unbalanced
ML Data Test Detection Bank Transaction Frauds Unbalanced
The project provides a quick and accessible dataset designed for learning and experimenting with machine learning algorithms, specifically in the context of detecting fraudulent bank transactions. It is intended for practicing and applying concepts such as Random Forest, Support Vector Machines (SVM), and Synthetic Minority Over-sampling Technique (SMOTE) to address unbalanced classification problems.
Note: This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/roberto-armas/ml_data_test_detection_bank_transaction_frauds_unbalanced.FinCast-Paper-testtest_merra_pm25cyp-challenge-train-test
CYP Challenge Train/Test Dataset
A high-quality experimental dataset for predicting inhibition of the major drug-metabolizing Cytochrome P450 enzymes (CYP1A2, CYP2C9, CYP2D6, CYP3A4), released as part of the OpenADMET CYP Inhibition Blind Challenge.
Blog post: Announcing OpenADMET’s CYP inhibition blind challenge
Challenge Space: OpenADMET CYP Inhibition Blind Challenge
Challenge period: August 17, 2026 - November 3, 2026
Produced by: OpenADMET
CHANGELOG
Updated… See the full description on the dataset page: https://huggingface.co/datasets/ks121/cyp-challenge-train-test.test-datasetSC-train-valid-test_SDG-Descriptionsnli-label:
(0) entailment
(2) contradiction
human_methylation_bench_ver1_test
Human DNA Methylation Dataset ver1
This dataset is a benchmark dataset for predicting the aging clock, curated from publicly available DNA methylation data. The original benchmark dataset was published by Dmitrii Kriukov et al. (2024) by integrating data from 65 individual studies.
To improve usability, we ensured unique sample IDs (excluding duplicate data, GSE118468 and GSE118469) and randomly split the data into training and testing subsets (train : test = 7 : 3) to… See the full description on the dataset page: https://huggingface.co/datasets/openaging/human_methylation_bench_ver1_test.
