datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
external-benchmarking
Vector Search Benchmarks
This repo contains datasets for benchmarking vector search performance, to help Superlinked prioritize integration partners.
For performing actual benchmarking on this dataset, see the github repository README.
Overview
We reviewed a number of publicly available datasets and noted 3 core problems + here is how this dataset fixes them:
Problems of other vector search benchmarks
How this dataset solves it
Not enough metadata of… See the full description on the dataset page: https://huggingface.co/datasets/superlinked/external-benchmarking.groundtruth-dynamic-benchmarking
Groundtruth Dynamic Benchmarking — Geology
Question sets and grading rubrics for evaluating LLMs on real-world geological
reasoning. Every question is authored from a real source corpus, and every
claim in the grading key carries an evidence locator back to that corpus —
nothing is synthetic. Licensing/redistribution status varies by corpus — see
License.
This dataset holds the questions, grading rubrics, and source corpora.
Running an evaluation (generating answers from a model… See the full description on the dataset page: https://huggingface.co/datasets/EigenformAI/groundtruth-dynamic-benchmarking.cdsm_benchmarking_data
CDSM Collagen Structure Benchmark — Data
Structures and scores for a benchmark comparing a deterministic collagen
triple-helix builder (CDSM) against four co-folding models — Boltz-2, Chai-1,
Protenix and AlphaFold3, the last in both with-MSA (af3_msa) and no-MSA
(af3_nomsa) conditions — on 80 experimentally resolved collagen triple
helices from the RCSB PDB.
Code: https://github.com/bm-howard/cdsm_benchmarking
Layout
Prefix
Contents
Size
experimental/… See the full description on the dataset page: https://huggingface.co/datasets/CollagenHelixLabs/cdsm_benchmarking_data.Enlatics_benchmarking
GAIA-style Evaluation Results (Public)
This dataset contains GAIA-inspired benchmark question results for LLM evaluation.
What is inside
grok_answers.csv: question-by-question outputs, model used, response time (seconds), and run status.
Notes
These tasks are designed in a GAIA-style (multi-hop, web-grounded questions).
Official GAIA leaderboard submission access was restricted for our account, so results are published here for transparency and… See the full description on the dataset page: https://huggingface.co/datasets/enlatics/Enlatics_benchmarking.Agri_STT_Benchmarking_Dataset
Agri STT Benchmarking Dataset
10,808 farmer voice queries in Hindi, Telugu and Odia, with reference transcripts, for benchmarking automatic speech recognition in agricultural contexts. The audio is included in this repository.
Every recording is a smallholder farmer speaking a question to Farmer.Chat, an AI advisory service run by Digital Green. Reference transcripts were produced by human annotators. Nothing here is read from a script or recorded in a studio, so the audio… See the full description on the dataset page: https://huggingface.co/datasets/DigiGreen/Agri_STT_Benchmarking_Dataset.URSA-benchmarking-sets
URSA benchmarking sets
Benchmark for reaction plausibility, single-step and multistep retrosynthesis from Zagribelnyy et al. (2026).
It bundles two task families: reaction-level plausibility judgment and retrosynthesis target sets.
Reaction plausibility
URSA-reaction-plausibility-bench-2026.csv — 1,000 reactions predicted by different models hand-labeled by expert oragnic chemists (500 plausible / 500 implausible), for benchmarking reaction-level… See the full description on the dataset page: https://huggingface.co/datasets/insilicomedicine/URSA-benchmarking-sets.malayalam_msc_benchmarkingminipile_benchmarkingUM-DLP-Public-Benchmarking-Dataset
UM DLP Public Benchmarking Dataset
Description
The UM DLP Public Benchmarking Dataset is a publicly available collection designed specifically to stress test Data Loss Prevention (DLP) systems, helping identify detection gaps, false positives, and false negatives for ongoing improvement.
This benchmark dataset contains 1,343 manually validated records across six major categories relevant to financial and sensitive data risks:
Financial Data (Account information about… See the full description on the dataset page: https://huggingface.co/datasets/alibustami/UM-DLP-Public-Benchmarking-Dataset.PDB-Single
PDB-Single: Precise Debugging Benchmarking — single-line bug set
📄 Paper ·
💻 Code ·
🌐 Project page ·
🏆 Leaderboard
PDB-Single is the single-line bug set of the PDB (Precise Debugging Benchmarking) evaluation suite. Every example pairs a ground-truth program with a synthesized buggy version plus a line-level edit script (gt_diff) that encodes the minimal correct fix.
Source datasets: BigCodeBench + LiveCodeBench
Sibling datasets:… See the full description on the dataset page: https://huggingface.co/datasets/Precise-Debugging-Benchmarking/PDB-Single.benchmarking-cultures-25
Benchmarking-Cultures-25 Dataset
This dataset accompanies the Unsteady Metrics and Benchmarking Cultures of AI Model Builders paper submitted to FAccT 2026 by Stefan Baack, Christo Buschek and Maty Bohacek.
The dataset contains the following parts:
core: The curated Benchmarking-Cultures-25 dataset.
derived: Datasets that were derived from the core dataset and informed the FAccT submission.
figures: Figures generated from derived data and used in the paper.
docs: Data dictionaries… See the full description on the dataset page: https://huggingface.co/datasets/matybohacek/benchmarking-cultures-25.malayalam_common_voice_benchmarkingcross_species_benchmarking**Repository: https://d-script.readthedocs.io/en/stable/data.html
**Reference: Sledzieski, S., Singh, R., Cowen, L. & Berger, B. D-SCRIPT translates genome to phenome with sequence-based, structure-aware, genome-scale predictions of protein-protein interactions. Cell Systems 12, 969-982.e6 (2021).
PDB-Single-Full
PDB-Single-Full: Precise Debugging Benchmarking — unfiltered single-line bug pool
📄 Paper ·
💻 Code ·
🌐 Project page ·
🏆 Leaderboard
PDB-Single-Full is the unfiltered single-line bug pool of the PDB (Precise Debugging Benchmarking) evaluation suite. Every example pairs a ground-truth program with a synthesized buggy version plus a line-level edit script (gt_diff) that encodes the minimal correct fix.
Source datasets: BigCodeBench +… See the full description on the dataset page: https://huggingface.co/datasets/Precise-Debugging-Benchmarking/PDB-Single-Full.PDB-Multi
PDB-Multi: Precise Debugging Benchmarking — multi-line bug subset (2–4 line blocks)
📄 Paper ·
💻 Code ·
🌐 Project page ·
🏆 Leaderboard
PDB-Multi is the multi-line bug subset (2–4 line blocks) of the PDB (Precise Debugging Benchmarking) evaluation suite. Every example pairs a ground-truth program with a synthesized buggy version plus a line-level edit script (gt_diff) that encodes the minimal correct fix.
Source datasets: BigCodeBench +… See the full description on the dataset page: https://huggingface.co/datasets/Precise-Debugging-Benchmarking/PDB-Multi.Bernett_benchmarking
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Dataset Sources [optional]
**Repository: https://doi.org/10.6084/m9.figshare.21591618.v3
**Reference:
Bernett, J., Blumenthal, D. B. & List, M. Cracking the black box of deep sequence-based protein–protein interaction prediction. Briefings in Bioinformatics 25, bbae076… See the full description on the dataset page: https://huggingface.co/datasets/danliu1226/Bernett_benchmarking.quantum-error-mitigation-and-benchmarking
Neura Parse — Quantum Error Mitigation, Characterization & Benchmarking
A pre-fault-tolerance, code-backed vertical on getting trustworthy answers from noisy hardware and rigorously measuring device quality: error-mitigation techniques, characterization/tomography protocols, and benchmarking suites. Runnable Mitiq, pyGSTi, and Qiskit Experiments pipelines with honest sampling-overhead and bias/variance accounting — the practitioner and research toolkit the general dataset… See the full description on the dataset page: https://huggingface.co/datasets/Neura-parse/quantum-error-mitigation-and-benchmarking.ASR-Benchmarking-Dataset
Hindi STT Benchmarking Eval
Overview
This dataset packages the Hindi eval split used for STT benchmarking across six Vistaar-derived parts: IndicTTS, FLEURS, CommonVoice, Kathbath, Kathbath noisy, and MUCS. Each row contains the audio, original reference transcript, and raw plus normalized transcripts from Ringg, ElevenLabs, Deepgram, and Sarvam.
The dataset contains 10,000 utterances and about 15.5 hours of 16 kHz mono WAV audio.
The dataset is published as part-specific… See the full description on the dataset page: https://huggingface.co/datasets/RinggAI/ASR-Benchmarking-Dataset.llama-south-africa-benchmarking
Dataset Card for Evaluation run of chad-brouze/llama-8b-south-africa
Dataset automatically created during the evaluation run of model chad-brouze/llama-8b-south-africa
The dataset is composed of 17 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 14 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/africa-intelligence/llama-south-africa-benchmarking.PDB-Wild
PDB-Wild: Precise Debugging Benchmarking — multi-line and repository-level bugs
📄 Paper ·
💻 Code ·
🌐 Project page ·
🏆 Leaderboard
PDB-Wild is the multi-line and repository-level bug set of the PDB (Precise Debugging Benchmarking) evaluation suite. Every example pairs a ground-truth program with a synthesized buggy version plus a line-level edit script (gt_diff) that encodes the minimal correct fix.
Source datasets: BigCodeBench +… See the full description on the dataset page: https://huggingface.co/datasets/Precise-Debugging-Benchmarking/PDB-Wild.wikidata_benchmarkingThis dataset was compiled for the purpose of finetuning models in the context of benchmarking for art historical research.
Images scraped from Wikimedia Commons via Wikidata; metadata scraped from
Wikidata (CC0). Image licenses vary per file (predominantly public domain,
some CC-BY-SA) see the license_short_name / license_url columns in the
parquet files for the exact terms of each individual image, and the
commons file page for full details.
Benchmarking_tpLMs_dataaya101-benchmarking
Dataset Card for Evaluation run of CohereForAI/aya-101
Dataset automatically created during the evaluation run of model CohereForAI/aya-101
The dataset is composed of 5 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/africa-intelligence/aya101-benchmarking.Treebank-Benchmarking
Turkish Treebank Benchmarking
This is the repo for Turkish treebank benchmarking, namely evaluating Tranformer models on POS-Dep-Morph task.
For the data, we used two treebank, IMST and BOUN. We converted conllu format to json lines for being compatible to HF dataset formats.
Here are treebank sizes at a glance:
Dataset
train lines
dev lines
test lines
BOUN
7803
979
979
IMST
3435
1100
1100
A typical instance from the dataset looks like:
{
"id": "ins_1267"… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/Treebank-Benchmarking.building-energy-benchmarking-requirements
Building Energy Benchmarking and Performance Standard Requirements by Jurisdiction
Canonical, always-current version: https://referencesource.org/building-energy-benchmarking-requirements/
Machine-readable: https://referencesource.org/building-energy-benchmarking-requirements/data.json — this mirror is a point-in-time copy.
Last verified: 2026-08-15
Stale after: 2027-02-11 (past this date, prefer the canonical copy —
it re-verifies on a cadence this snapshot does not)
Records:… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/building-energy-benchmarking-requirements.URSA-benchmarking-sets
URSA benchmarking sets
Benchmark for reaction plausibility, single-step and multistep retrosynthesis from Zagribelnyy et al. (2026).
It bundles two task families: reaction-level plausibility judgment and retrosynthesis target sets.
Reaction plausibility
URSA-reaction-plausibility-bench-2026.csv — 1,000 reactions predicted by different models hand-labeled by expert oragnic chemists (500 plausible / 500 implausible), for benchmarking reaction-level… See the full description on the dataset page: https://huggingface.co/datasets/introvoyz041/URSA-benchmarking-sets.africa-synth-energy-pv-performance-benchmarking-africa-benin
Africa Synth Energy Pv Performance Benchmarking Africa Benin | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: csv - Sector: energy - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-energy-pv-performance-benchmarking-africa-benin.UM-DLP-Public-Benchmarking-Dataset
UM DLP Public Benchmarking Dataset
Description
The UM DLP Public Benchmarking Dataset is a publicly available collection designed specifically to stress test Data Loss Prevention (DLP) systems, helping identify detection gaps, false positives, and false negatives for ongoing improvement.
This benchmark dataset contains 1,343 manually validated records across six major categories relevant to financial and sensitive data risks:
Financial Data (Account information… See the full description on the dataset page: https://huggingface.co/datasets/xorushi/UM-DLP-Public-Benchmarking-Dataset.faithfulness-precision-spl-context-gpt-benchmarkingtopical-guardrailing-nemo-benchmarking
