Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01superlinked /external-benchmarking Vector Search Benchmarks This repo contains datasets for benchmarking vector search performance, to help Superlinked prioritize integration partners. For performing actual benchmarking on this dataset, see the github repository README. Overview We reviewed a number of publicly available datasets and noted 3 core problems + here is how this dataset fixes them: Problems of other vector search benchmarks How this dataset solves it Not enough metadata of… See the full description on the dataset page: https://huggingface.co/datasets/superlinked/external-benchmarking.image10M<n<100M0 likes716 downloads1y agoHugging Face02EigenformAI /groundtruth-dynamic-benchmarking Groundtruth Dynamic Benchmarking — Geology Question sets and grading rubrics for evaluating LLMs on real-world geological reasoning. Every question is authored from a real source corpus, and every claim in the grading key carries an evidence locator back to that corpus — nothing is synthetic. Licensing/redistribution status varies by corpus — see License. This dataset holds the questions, grading rubrics, and source corpora. Running an evaluation (generating answers from a model… See the full description on the dataset page: https://huggingface.co/datasets/EigenformAI/groundtruth-dynamic-benchmarking.textquestion-answeringn<1K0 likes376 downloads1mo agoHugging Face03CollagenHelixLabs /cdsm_benchmarking_data CDSM Collagen Structure Benchmark — Data Structures and scores for a benchmark comparing a deterministic collagen triple-helix builder (CDSM) against four co-folding models — Boltz-2, Chai-1, Protenix and AlphaFold3, the last in both with-MSA (af3_msa) and no-MSA (af3_nomsa) conditions — on 80 experimentally resolved collagen triple helices from the RCSB PDB. Code: https://github.com/bm-howard/cdsm_benchmarking Layout Prefix Contents Size experimental/… See the full description on the dataset page: https://huggingface.co/datasets/CollagenHelixLabs/cdsm_benchmarking_data.tabular10K<n<100K0 likes297 downloads1mo agoHugging Face04enlatics /Enlatics_benchmarking GAIA-style Evaluation Results (Public) This dataset contains GAIA-inspired benchmark question results for LLM evaluation. What is inside grok_answers.csv: question-by-question outputs, model used, response time (seconds), and run status. Notes These tasks are designed in a GAIA-style (multi-hop, web-grounded questions). Official GAIA leaderboard submission access was restricted for our account, so results are published here for transparency and… See the full description on the dataset page: https://huggingface.co/datasets/enlatics/Enlatics_benchmarking.textn<1K0 likes264 downloads8mo agoHugging Face05DigiGreen /Agri_STT_Benchmarking_Dataset Agri STT Benchmarking Dataset 10,808 farmer voice queries in Hindi, Telugu and Odia, with reference transcripts, for benchmarking automatic speech recognition in agricultural contexts. The audio is included in this repository. Every recording is a smallholder farmer speaking a question to Farmer.Chat, an AI advisory service run by Digital Green. Reference transcripts were produced by human annotators. Nothing here is read from a script or recorded in a studio, so the audio… See the full description on the dataset page: https://huggingface.co/datasets/DigiGreen/Agri_STT_Benchmarking_Dataset.audioautomatic-speech-recognition10K<n<100K3 likes260 downloads2mo agoHugging Face06insilicomedicine /URSA-benchmarking-sets URSA benchmarking sets Benchmark for reaction plausibility, single-step and multistep retrosynthesis from Zagribelnyy et al. (2026). It bundles two task families: reaction-level plausibility judgment and retrosynthesis target sets. Reaction plausibility URSA-reaction-plausibility-bench-2026.csv — 1,000 reactions predicted by different models hand-labeled by expert oragnic chemists (500 plausible / 500 implausible), for benchmarking reaction-level… See the full description on the dataset page: https://huggingface.co/datasets/insilicomedicine/URSA-benchmarking-sets.text1K<n<10K2 likes206 downloads8d agoHugging Face07kurianbenoy /malayalam_msc_benchmarkingtabular10K<n<100K1 likes179 downloads3y agoHugging Face08kenhktsui /minipile_benchmarkingtabular1M<n<10M0 likes175 downloads2y agoHugging Face09alibustami /UM-DLP-Public-Benchmarking-Dataset UM DLP Public Benchmarking Dataset Description The UM DLP Public Benchmarking Dataset is a publicly available collection designed specifically to stress test Data Loss Prevention (DLP) systems, helping identify detection gaps, false positives, and false negatives for ongoing improvement. This benchmark dataset contains 1,343 manually validated records across six major categories relevant to financial and sensitive data risks: Financial Data (Account information about… See the full description on the dataset page: https://huggingface.co/datasets/alibustami/UM-DLP-Public-Benchmarking-Dataset.tabulartext-classification1K<n<10K4 likes163 downloads1y agoHugging Face10Precise-Debugging-Benchmarking /PDB-Single PDB-Single: Precise Debugging Benchmarking — single-line bug set 📄 Paper &nbsp;·&nbsp; 💻 Code &nbsp;·&nbsp; 🌐 Project page &nbsp;·&nbsp; 🏆 Leaderboard PDB-Single is the single-line bug set of the PDB (Precise Debugging Benchmarking) evaluation suite. Every example pairs a ground-truth program with a synthesized buggy version plus a line-level edit script (gt_diff) that encodes the minimal correct fix. Source datasets: BigCodeBench + LiveCodeBench Sibling datasets:… See the full description on the dataset page: https://huggingface.co/datasets/Precise-Debugging-Benchmarking/PDB-Single.tabulartext-generation1K<n<10K1 likes157 downloads5d agoHugging Face11matybohacek /benchmarking-cultures-25 Benchmarking-Cultures-25 Dataset This dataset accompanies the Unsteady Metrics and Benchmarking Cultures of AI Model Builders paper submitted to FAccT 2026 by Stefan Baack, Christo Buschek and Maty Bohacek. The dataset contains the following parts: core: The curated Benchmarking-Cultures-25 dataset. derived: Datasets that were derived from the core dataset and informed the FAccT submission. figures: Figures generated from derived data and used in the paper. docs: Data dictionaries… See the full description on the dataset page: https://huggingface.co/datasets/matybohacek/benchmarking-cultures-25.text1K<n<10K4 likes150 downloads5mo agoHugging Face12kurianbenoy /malayalam_common_voice_benchmarkingtabular1K<n<10K1 likes128 downloads3y agoHugging Face13danliu1226 /cross_species_benchmarking**Repository: https://d-script.readthedocs.io/en/stable/data.html **Reference: Sledzieski, S., Singh, R., Cowen, L. & Berger, B. D-SCRIPT translates genome to phenome with sequence-based, structure-aware, genome-scale predictions of protein-protein interactions. Cell Systems 12, 969-982.e6 (2021). text100K<n<1M2 likes124 downloads1y agoHugging Face14Precise-Debugging-Benchmarking /PDB-Single-Full PDB-Single-Full: Precise Debugging Benchmarking — unfiltered single-line bug pool 📄 Paper &nbsp;·&nbsp; 💻 Code &nbsp;·&nbsp; 🌐 Project page &nbsp;·&nbsp; 🏆 Leaderboard PDB-Single-Full is the unfiltered single-line bug pool of the PDB (Precise Debugging Benchmarking) evaluation suite. Every example pairs a ground-truth program with a synthesized buggy version plus a line-level edit script (gt_diff) that encodes the minimal correct fix. Source datasets: BigCodeBench +… See the full description on the dataset page: https://huggingface.co/datasets/Precise-Debugging-Benchmarking/PDB-Single-Full.tabulartext-generation1K<n<10K0 likes113 downloads5d agoHugging Face15Precise-Debugging-Benchmarking /PDB-Multi PDB-Multi: Precise Debugging Benchmarking — multi-line bug subset (2–4 line blocks) 📄 Paper &nbsp;·&nbsp; 💻 Code &nbsp;·&nbsp; 🌐 Project page &nbsp;·&nbsp; 🏆 Leaderboard PDB-Multi is the multi-line bug subset (2–4 line blocks) of the PDB (Precise Debugging Benchmarking) evaluation suite. Every example pairs a ground-truth program with a synthesized buggy version plus a line-level edit script (gt_diff) that encodes the minimal correct fix. Source datasets: BigCodeBench +… See the full description on the dataset page: https://huggingface.co/datasets/Precise-Debugging-Benchmarking/PDB-Multi.tabulartext-generationn<1K0 likes81 downloads5d agoHugging Face16danliu1226 /Bernett_benchmarking Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Dataset Sources [optional] **Repository: https://doi.org/10.6084/m9.figshare.21591618.v3 **Reference: Bernett, J., Blumenthal, D. B. & List, M. Cracking the black box of deep sequence-based protein–protein interaction prediction. Briefings in Bioinformatics 25, bbae076… See the full description on the dataset page: https://huggingface.co/datasets/danliu1226/Bernett_benchmarking.text100K<n<1M0 likes62 downloads1y agoHugging Face17Neura-parse /quantum-error-mitigation-and-benchmarking Neura Parse — Quantum Error Mitigation, Characterization & Benchmarking A pre-fault-tolerance, code-backed vertical on getting trustworthy answers from noisy hardware and rigorously measuring device quality: error-mitigation techniques, characterization/tomography protocols, and benchmarking suites. Runnable Mitiq, pyGSTi, and Qiskit Experiments pipelines with honest sampling-overhead and bias/variance accounting — the practitioner and research toolkit the general dataset… See the full description on the dataset page: https://huggingface.co/datasets/Neura-parse/quantum-error-mitigation-and-benchmarking.tabulartext-generation100K<n<1M0 likes60 downloads3mo agoHugging Face18RinggAI /ASR-Benchmarking-Dataset Hindi STT Benchmarking Eval Overview This dataset packages the Hindi eval split used for STT benchmarking across six Vistaar-derived parts: IndicTTS, FLEURS, CommonVoice, Kathbath, Kathbath noisy, and MUCS. Each row contains the audio, original reference transcript, and raw plus normalized transcripts from Ringg, ElevenLabs, Deepgram, and Sarvam. The dataset contains 10,000 utterances and about 15.5 hours of 16 kHz mono WAV audio. The dataset is published as part-specific… See the full description on the dataset page: https://huggingface.co/datasets/RinggAI/ASR-Benchmarking-Dataset.audioautomatic-speech-recognition10K<n<100K1 likes58 downloads5mo agoHugging Face19africa-intelligence /llama-south-africa-benchmarking Dataset Card for Evaluation run of chad-brouze/llama-8b-south-africa Dataset automatically created during the evaluation run of model chad-brouze/llama-8b-south-africa The dataset is composed of 17 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 14 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/africa-intelligence/llama-south-africa-benchmarking.tabular1K<n<10K0 likes46 downloads2y agoHugging Face20Precise-Debugging-Benchmarking /PDB-Wild PDB-Wild: Precise Debugging Benchmarking — multi-line and repository-level bugs 📄 Paper &nbsp;·&nbsp; 💻 Code &nbsp;·&nbsp; 🌐 Project page &nbsp;·&nbsp; 🏆 Leaderboard PDB-Wild is the multi-line and repository-level bug set of the PDB (Precise Debugging Benchmarking) evaluation suite. Every example pairs a ground-truth program with a synthesized buggy version plus a line-level edit script (gt_diff) that encodes the minimal correct fix. Source datasets: BigCodeBench +… See the full description on the dataset page: https://huggingface.co/datasets/Precise-Debugging-Benchmarking/PDB-Wild.tabulartext-generationn<1K0 likes46 downloads5d agoHugging Face21chcaa /wikidata_benchmarkingThis dataset was compiled for the purpose of finetuning models in the context of benchmarking for art historical research. Images scraped from Wikimedia Commons via Wikidata; metadata scraped from Wikidata (CC0). Image licenses vary per file (predominantly public domain, some CC-BY-SA) see the license_short_name / license_url columns in the parquet files for the exact terms of each individual image, and the commons file page for full details. image1K<n<10K0 likes40 downloads1mo agoHugging Face22yk0 /Benchmarking_tpLMs_datatext0 likes39 downloads1y agoHugging Face23africa-intelligence /aya101-benchmarking Dataset Card for Evaluation run of CohereForAI/aya-101 Dataset automatically created during the evaluation run of model CohereForAI/aya-101 The dataset is composed of 5 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/africa-intelligence/aya101-benchmarking.tabular1K<n<10K0 likes36 downloads2y agoHugging Face24turkish-nlp-suite /Treebank-Benchmarking Turkish Treebank Benchmarking This is the repo for Turkish treebank benchmarking, namely evaluating Tranformer models on POS-Dep-Morph task. For the data, we used two treebank, IMST and BOUN. We converted conllu format to json lines for being compatible to HF dataset formats. Here are treebank sizes at a glance: Dataset train lines dev lines test lines BOUN 7803 979 979 IMST 3435 1100 1100 A typical instance from the dataset looks like: { "id": "ins_1267"… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/Treebank-Benchmarking.text10K<n<100K0 likes34 downloads8mo agoHugging Face25referencesource /building-energy-benchmarking-requirements Building Energy Benchmarking and Performance Standard Requirements by Jurisdiction Canonical, always-current version: https://referencesource.org/building-energy-benchmarking-requirements/ Machine-readable: https://referencesource.org/building-energy-benchmarking-requirements/data.json — this mirror is a point-in-time copy. Last verified: 2026-08-15 Stale after: 2027-02-11 (past this date, prefer the canonical copy — it re-verifies on a cadence this snapshot does not) Records:… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/building-energy-benchmarking-requirements.textn<1K0 likes34 downloads12d agoHugging Face26introvoyz041 /URSA-benchmarking-sets URSA benchmarking sets Benchmark for reaction plausibility, single-step and multistep retrosynthesis from Zagribelnyy et al. (2026). It bundles two task families: reaction-level plausibility judgment and retrosynthesis target sets. Reaction plausibility URSA-reaction-plausibility-bench-2026.csv — 1,000 reactions predicted by different models hand-labeled by expert oragnic chemists (500 plausible / 500 implausible), for benchmarking reaction-level… See the full description on the dataset page: https://huggingface.co/datasets/introvoyz041/URSA-benchmarking-sets.text1K<n<10K0 likes29 downloads2mo agoHugging Face27electricsheepafrica /africa-synth-energy-pv-performance-benchmarking-africa-benin Africa Synth Energy Pv Performance Benchmarking Africa Benin | Africa (Electric Sheep Africa metadata inventory) Size category: 10K<n<100K - Formats: csv - Sector: energy - Engineered by Electric Sheep Africa TL;DR This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context. What This Dataset Covers… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-energy-pv-performance-benchmarking-africa-benin.tabulartabular-classification10K<n<100K0 likes28 downloads2mo agoHugging Face28xorushi /UM-DLP-Public-Benchmarking-Dataset UM DLP Public Benchmarking Dataset Description The UM DLP Public Benchmarking Dataset is a publicly available collection designed specifically to stress test Data Loss Prevention (DLP) systems, helping identify detection gaps, false positives, and false negatives for ongoing improvement. This benchmark dataset contains 1,343 manually validated records across six major categories relevant to financial and sensitive data risks: Financial Data (Account information… See the full description on the dataset page: https://huggingface.co/datasets/xorushi/UM-DLP-Public-Benchmarking-Dataset.tabulartext-classification1K<n<10K0 likes25 downloads2mo agoHugging Face29onepaneai /faithfulness-precision-spl-context-gpt-benchmarkingtextn<1K0 likes24 downloads2y agoHugging Face30onepaneai /topical-guardrailing-nemo-benchmarkingtextn<1K0 likes23 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.