datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Enlatics_benchmarking
GAIA-style Evaluation Results (Public)
This dataset contains GAIA-inspired benchmark question results for LLM evaluation.
What is inside
grok_answers.csv: question-by-question outputs, model used, response time (seconds), and run status.
Notes
These tasks are designed in a GAIA-style (multi-hop, web-grounded questions).
Official GAIA leaderboard submission access was restricted for our account, so results are published here for transparency and… See the full description on the dataset page: https://huggingface.co/datasets/enlatics/Enlatics_benchmarking.URSA-benchmarking-sets
URSA benchmarking sets
Benchmark for reaction plausibility, single-step and multistep retrosynthesis from Zagribelnyy et al. (2026).
It bundles two task families: reaction-level plausibility judgment and retrosynthesis target sets.
Reaction plausibility
URSA-reaction-plausibility-bench-2026.csv — 1,000 reactions predicted by different models hand-labeled by expert oragnic chemists (500 plausible / 500 implausible), for benchmarking reaction-level… See the full description on the dataset page: https://huggingface.co/datasets/insilicomedicine/URSA-benchmarking-sets.UM-DLP-Public-Benchmarking-Dataset
UM DLP Public Benchmarking Dataset
Description
The UM DLP Public Benchmarking Dataset is a publicly available collection designed specifically to stress test Data Loss Prevention (DLP) systems, helping identify detection gaps, false positives, and false negatives for ongoing improvement.
This benchmark dataset contains 1,343 manually validated records across six major categories relevant to financial and sensitive data risks:
Financial Data (Account information about… See the full description on the dataset page: https://huggingface.co/datasets/alibustami/UM-DLP-Public-Benchmarking-Dataset.benchmarking-cultures-25
Benchmarking-Cultures-25 Dataset
This dataset accompanies the Unsteady Metrics and Benchmarking Cultures of AI Model Builders paper submitted to FAccT 2026 by Stefan Baack, Christo Buschek and Maty Bohacek.
The dataset contains the following parts:
core: The curated Benchmarking-Cultures-25 dataset.
derived: Datasets that were derived from the core dataset and informed the FAccT submission.
figures: Figures generated from derived data and used in the paper.
docs: Data dictionaries… See the full description on the dataset page: https://huggingface.co/datasets/matybohacek/benchmarking-cultures-25.cross_species_benchmarking**Repository: https://d-script.readthedocs.io/en/stable/data.html
**Reference: Sledzieski, S., Singh, R., Cowen, L. & Berger, B. D-SCRIPT translates genome to phenome with sequence-based, structure-aware, genome-scale predictions of protein-protein interactions. Cell Systems 12, 969-982.e6 (2021).
Bernett_benchmarking
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Dataset Sources [optional]
**Repository: https://doi.org/10.6084/m9.figshare.21591618.v3
**Reference:
Bernett, J., Blumenthal, D. B. & List, M. Cracking the black box of deep sequence-based protein–protein interaction prediction. Briefings in Bioinformatics 25, bbae076… See the full description on the dataset page: https://huggingface.co/datasets/danliu1226/Bernett_benchmarking.URSA-benchmarking-sets
URSA benchmarking sets
Benchmark for reaction plausibility, single-step and multistep retrosynthesis from Zagribelnyy et al. (2026).
It bundles two task families: reaction-level plausibility judgment and retrosynthesis target sets.
Reaction plausibility
URSA-reaction-plausibility-bench-2026.csv — 1,000 reactions predicted by different models hand-labeled by expert oragnic chemists (500 plausible / 500 implausible), for benchmarking reaction-level… See the full description on the dataset page: https://huggingface.co/datasets/introvoyz041/URSA-benchmarking-sets.africa-synth-energy-pv-performance-benchmarking-africa-benin
Africa Synth Energy Pv Performance Benchmarking Africa Benin | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: csv - Sector: energy - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-energy-pv-performance-benchmarking-africa-benin.UM-DLP-Public-Benchmarking-Dataset
UM DLP Public Benchmarking Dataset
Description
The UM DLP Public Benchmarking Dataset is a publicly available collection designed specifically to stress test Data Loss Prevention (DLP) systems, helping identify detection gaps, false positives, and false negatives for ongoing improvement.
This benchmark dataset contains 1,343 manually validated records across six major categories relevant to financial and sensitive data risks:
Financial Data (Account information… See the full description on the dataset page: https://huggingface.co/datasets/xorushi/UM-DLP-Public-Benchmarking-Dataset.virus_human_benchmarking**Repository: http://kurata35.bio.kyutech.ac.jp/LSTM-PHV/download_page
**Reference: Tsukiyama, S., Hasan, M. M., Fujii, S. & Kurata, H. LSTM-PHV: prediction of human-virus protein–protein interactions by LSTM with word2vec. Briefings in Bioinformatics 22, bbab228 (2021).
Agri_STT_Benchmarking_DatasetThis is a domain-specific, multilingual agricultural speech dataset with a primary focus on Hindi, Telugu, and Odia, designed for speech-to-text and automatic speech recognition (ASR) tasks. It features human-annotated transcriptions and is intended for benchmarking ASR model performance in real-world agricultural scenarios.
This paper presents a comprehensive benchmark of 10 ASR models for agricultural advisory use across Hindi, Telugu, and Odia, using 10,934 real-world Farmer.Chat audio… See the full description on the dataset page: https://huggingface.co/datasets/bullseye-4/Agri_STT_Benchmarking_Dataset.virus_human_benchmarking**Repository: http://kurata35.bio.kyutech.ac.jp/LSTM-PHV/download_page
**Reference: Tsukiyama, S., Hasan, M. M., Fujii, S. & Kurata, H. LSTM-PHV: prediction of human-virus protein–protein interactions by LSTM with word2vec. Briefings in Bioinformatics 22, bbab228 (2021).
