Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mackelab /benchmarking_sbi_runs Benchmarking SBI Runs This dataset contains the raw, per-run results underlying the manuscript "Benchmarking Simulation-Based Inference" (Lueckmann, Boelts, Greenberg, Goncalves & Macke, AISTATS 2021). It is a direct migration of the Git LFS data from mackelab/benchmarking_sbi_runs on GitHub. For compiled, ready-to-use dataframes built from these raw results (and the code that produced them), see the companion repository:… See the full description on the dataset page: https://huggingface.co/datasets/mackelab/benchmarking_sbi_runs.100K<n<1M0 likes14k downloads2mo agoHugging Face02DabbyOWL /PDE_Inverse_Problem_Benchmarking PDEInvBench: A Comprehensive Dataset and Design Space Exploration of Neural Networks for PDE Inverse Problems This is the official dataset for the paper PDEInvBench: A Comprehensive Dataset and Design Space Exploration of Neural Networks for PDE Inverse Problems. Code: GitHub - ASK-Berkeley/PDEInvBench Sample Usage You can use the provided script from the codebase to batch download the data: pip install huggingface_hub python3 huggingface_pdeinv_download.py --dataset… See the full description on the dataset page: https://huggingface.co/datasets/DabbyOWL/PDE_Inverse_Problem_Benchmarking.imageother100M<n<1B3 likes8.6k downloads5mo agoHugging Face03facebook /map-anything-benchmarking MapAnything Benchmarking Dataset Dataset Description This dataset contains the WAI format data used for benchmarking feed-forward 3D reconstruction models in the MapAnything codebase. Please see our Data Processing README for more details. Citation If you use this dataset in your research, please cite our paper: @inproceedings{keetha2026mapanything, title={{MapAnything}: Universal Feed-Forward Metric {3D} Reconstruction}, author={Nikhil Keetha and Norman… See the full description on the dataset page: https://huggingface.co/datasets/facebook/map-anything-benchmarking.image-to-3d100B<n<1T8 likes3.8k downloads9mo agoHugging Face04Precise-Debugging-Benchmarking /PDB-Results PDB-Results: model outputs and scores 📄 Paper &nbsp;·&nbsp; 💻 Code &nbsp;·&nbsp; 🌐 Project page &nbsp;·&nbsp; 🏆 Leaderboard Raw debugging outputs and evaluator scores for every model evaluated on the PDB (Precise Debugging Benchmarking) suite, so that every reported number can be inspected and recomputed. Evaluation sets Filename tag Set Tasks Models pdb_single_hard PDB-Single 5,751 (BigCodeBench 2,525 + LiveCodeBench 3,226) 4 pdb_single… See the full description on the dataset page: https://huggingface.co/datasets/Precise-Debugging-Benchmarking/PDB-Results.0 likes741 downloads5d agoHugging Face05superlinked /external-benchmarking Vector Search Benchmarks This repo contains datasets for benchmarking vector search performance, to help Superlinked prioritize integration partners. For performing actual benchmarking on this dataset, see the github repository README. Overview We reviewed a number of publicly available datasets and noted 3 core problems + here is how this dataset fixes them: Problems of other vector search benchmarks How this dataset solves it Not enough metadata of… See the full description on the dataset page: https://huggingface.co/datasets/superlinked/external-benchmarking.image10M<n<100M0 likes716 downloads1y agoHugging Face06EigenformAI /groundtruth-dynamic-benchmarking-submissions Groundtruth Dynamic Benchmarking — Geology — Submissions Community-submitted evaluation runs against the groundtruth-dynamic-benchmarking geology rubrics, feeding the leaderboard. We are currently running two tracks: model benchmarking (comparing different models with no special harness) and harness benchmarking (comparing different harnesses using a single standard model - GLM 4.7). Each submission is a pointwise rubric score: one model, scored 0–10 per question against a… See the full description on the dataset page: https://huggingface.co/datasets/EigenformAI/groundtruth-dynamic-benchmarking-submissions.0 likes448 downloads1mo agoHugging Face07EigenformAI /groundtruth-dynamic-benchmarking Groundtruth Dynamic Benchmarking — Geology Question sets and grading rubrics for evaluating LLMs on real-world geological reasoning. Every question is authored from a real source corpus, and every claim in the grading key carries an evidence locator back to that corpus — nothing is synthetic. Licensing/redistribution status varies by corpus — see License. This dataset holds the questions, grading rubrics, and source corpora. Running an evaluation (generating answers from a model… See the full description on the dataset page: https://huggingface.co/datasets/EigenformAI/groundtruth-dynamic-benchmarking.textquestion-answeringn<1K0 likes376 downloads1mo agoHugging Face08CollagenHelixLabs /cdsm_benchmarking_data CDSM Collagen Structure Benchmark — Data Structures and scores for a benchmark comparing a deterministic collagen triple-helix builder (CDSM) against four co-folding models — Boltz-2, Chai-1, Protenix and AlphaFold3, the last in both with-MSA (af3_msa) and no-MSA (af3_nomsa) conditions — on 80 experimentally resolved collagen triple helices from the RCSB PDB. Code: https://github.com/bm-howard/cdsm_benchmarking Layout Prefix Contents Size experimental/… See the full description on the dataset page: https://huggingface.co/datasets/CollagenHelixLabs/cdsm_benchmarking_data.tabular10K<n<100K0 likes297 downloads1mo agoHugging Face09enlatics /Enlatics_benchmarking GAIA-style Evaluation Results (Public) This dataset contains GAIA-inspired benchmark question results for LLM evaluation. What is inside grok_answers.csv: question-by-question outputs, model used, response time (seconds), and run status. Notes These tasks are designed in a GAIA-style (multi-hop, web-grounded questions). Official GAIA leaderboard submission access was restricted for our account, so results are published here for transparency and… See the full description on the dataset page: https://huggingface.co/datasets/enlatics/Enlatics_benchmarking.textn<1K0 likes264 downloads8mo agoHugging Face10DigiGreen /Agri_STT_Benchmarking_Dataset Agri STT Benchmarking Dataset 10,808 farmer voice queries in Hindi, Telugu and Odia, with reference transcripts, for benchmarking automatic speech recognition in agricultural contexts. The audio is included in this repository. Every recording is a smallholder farmer speaking a question to Farmer.Chat, an AI advisory service run by Digital Green. Reference transcripts were produced by human annotators. Nothing here is read from a script or recorded in a studio, so the audio… See the full description on the dataset page: https://huggingface.co/datasets/DigiGreen/Agri_STT_Benchmarking_Dataset.audioautomatic-speech-recognition10K<n<100K3 likes260 downloads2mo agoHugging Face11Skoleatlas /benchmark-dk-interaktivt-benchmarkingunivers Benchmark.dk - interaktivt benchmarkingunivers (komplet data-høst) Komplet høst af datamaterialet bag Indenrigs- og Sundhedsministeriets Benchmarkingenheds interaktive benchmarkingunivers: https://www.benchmark.dk/interaktivt-benchmarkingunivers. Høstet 1. juni 2026. Del af Silkeborg Skoleatlas - 4 af de deri indeholdte skole/dagtilbud-tabeller indgår også kurateret i atlassets egen silkeborg-benchmark-national-datasæt, men dette repo er den fulde, ukuraterede kilde: alle 6… See the full description on the dataset page: https://huggingface.co/datasets/Skoleatlas/benchmark-dk-interaktivt-benchmarkingunivers.tabular-classification1K<n<10K0 likes260 downloads3mo agoHugging Face12aviadcohz /Detecture_Benchmarking Detecture Benchmarking Suite A five-dataset benchmark suite for VLM-guided multi-texture segmentation, released alongside the Detecture architecture (VLM-guided multi-texture segmentation via multiplexed grounding). This bundle contains the training set, one in-domain test set, and three out-of-domain evaluation benchmarks used to score Detecture against baseline model families (SAM 3 vanilla, Grounded-SAM 3, Sa2VA, and Qwen2SAM zero-shot) in the Detecture paper. Layout… See the full description on the dataset page: https://huggingface.co/datasets/aviadcohz/Detecture_Benchmarking.imageimage-segmentation1K<n<10K0 likes258 downloads6mo agoHugging Face13allenai /fluid-benchmarking Fluid Language Model Benchmarking This dataset provides IRT models for ARC Challenge, GSM8K, HellaSwag, MMLU, TruthfulQA, and WinoGrande. Furthermore, it contains results for pretraining checkpoints of Amber-6.7B, K2-65B, OLMo1-7B, OLMo2-7B, Pythia-2.8B, and Pythia-6.9B, evaluated on these six benchmarks. 🚀 Usage For utilities to use the dataset and to replicate the results from the paper, please see the corresponding GitHub… See the full description on the dataset page: https://huggingface.co/datasets/allenai/fluid-benchmarking.3 likes225 downloads1y agoHugging Face14insilicomedicine /URSA-benchmarking-sets URSA benchmarking sets Benchmark for reaction plausibility, single-step and multistep retrosynthesis from Zagribelnyy et al. (2026). It bundles two task families: reaction-level plausibility judgment and retrosynthesis target sets. Reaction plausibility URSA-reaction-plausibility-bench-2026.csv — 1,000 reactions predicted by different models hand-labeled by expert oragnic chemists (500 plausible / 500 implausible), for benchmarking reaction-level… See the full description on the dataset page: https://huggingface.co/datasets/insilicomedicine/URSA-benchmarking-sets.text1K<n<10K2 likes206 downloads8d agoHugging Face15RomeroLab-Duke /protein-fitness-datasets-for-benchmarking-ft-esm2-strategies Protein Fitness Datasets for Benchmarking ESM-2 Fine-Tuning Strategies Dataset description This repository contains processed protein sequence–function datasets for CreiLOV, avGFP, and Ube4b. The variants and experimental measurements were obtained in previously published deep mutational scanning studies: Chen, Y. et al. Deep Mutational Scanning of an Oxygen-Independent Fluorescent Protein CreiLOV for Comprehensive Profiling of Mutational and Epistatic Effects.… See the full description on the dataset page: https://huggingface.co/datasets/RomeroLab-Duke/protein-fitness-datasets-for-benchmarking-ft-esm2-strategies.0 likes191 downloads2mo agoHugging Face16kurianbenoy /malayalam_msc_benchmarkingtabular10K<n<100K1 likes179 downloads3y agoHugging Face17broadinstitute /Celldega_Visualization_Benchmarkingvideon<1K0 likes179 downloads2mo agoHugging Face18open-source-benchmarking /os-world-modified0 likes176 downloads1y agoHugging Face19kenhktsui /minipile_benchmarkingtabular1M<n<10M0 likes175 downloads2y agoHugging Face20alibustami /UM-DLP-Public-Benchmarking-Dataset UM DLP Public Benchmarking Dataset Description The UM DLP Public Benchmarking Dataset is a publicly available collection designed specifically to stress test Data Loss Prevention (DLP) systems, helping identify detection gaps, false positives, and false negatives for ongoing improvement. This benchmark dataset contains 1,343 manually validated records across six major categories relevant to financial and sensitive data risks: Financial Data (Account information about… See the full description on the dataset page: https://huggingface.co/datasets/alibustami/UM-DLP-Public-Benchmarking-Dataset.tabulartext-classification1K<n<10K4 likes163 downloads1y agoHugging Face21mwmathis /DLCspeed_benchmarking Dataset Card for DLC Speed Benchmarking ZIP This supports the dlc-live benchmarking zip formally hosted on our Harvard Rowland server. All information can be found in our publication: Real-time, low-latency closed-loop feedback using markerless posture tracking Gary A Kane, Gonçalo Lopes, Jonny L Saunders, Alexander Mathis, Mackenzie W Mathis https://elifesciences.org/articles/61909 Direct Use """ DeepLabCut Toolbox (deeplabcut.org) © A. & M. Mathis Labs Licensed… See the full description on the dataset page: https://huggingface.co/datasets/mwmathis/DLCspeed_benchmarking.0 likes157 downloads1y agoHugging Face22Precise-Debugging-Benchmarking /PDB-Single PDB-Single: Precise Debugging Benchmarking — single-line bug set 📄 Paper &nbsp;·&nbsp; 💻 Code &nbsp;·&nbsp; 🌐 Project page &nbsp;·&nbsp; 🏆 Leaderboard PDB-Single is the single-line bug set of the PDB (Precise Debugging Benchmarking) evaluation suite. Every example pairs a ground-truth program with a synthesized buggy version plus a line-level edit script (gt_diff) that encodes the minimal correct fix. Source datasets: BigCodeBench + LiveCodeBench Sibling datasets:… See the full description on the dataset page: https://huggingface.co/datasets/Precise-Debugging-Benchmarking/PDB-Single.tabulartext-generation1K<n<10K1 likes157 downloads5d agoHugging Face23matybohacek /benchmarking-cultures-25 Benchmarking-Cultures-25 Dataset This dataset accompanies the Unsteady Metrics and Benchmarking Cultures of AI Model Builders paper submitted to FAccT 2026 by Stefan Baack, Christo Buschek and Maty Bohacek. The dataset contains the following parts: core: The curated Benchmarking-Cultures-25 dataset. derived: Datasets that were derived from the core dataset and informed the FAccT submission. figures: Figures generated from derived data and used in the paper. docs: Data dictionaries… See the full description on the dataset page: https://huggingface.co/datasets/matybohacek/benchmarking-cultures-25.text1K<n<10K4 likes150 downloads5mo agoHugging Face24nyamtulla /benchmarking-the-benchmarks-data Benchmarking the Benchmarks — Raw SLM Safety Evaluation Runs Raw evaluation data for the ESORICS 2026 paper: Nyamtulla Shaik, Fengjun Li, Bo Luo — University of Kansas Benchmarking the Benchmarks: Evaluating Automated Safety Benchmarks for Small Language Models. ESORICS 2026. arXiv:2608.17183 Code and processed data: https://github.com/nyamtulla/benchmarking-the-benchmarks ⚠️ Content warning This dataset contains adversarial safety prompts and model responses… See the full description on the dataset page: https://huggingface.co/datasets/nyamtulla/benchmarking-the-benchmarks-data.text-generation100K<n<1M2 likes144 downloads1mo agoHugging Face25aviadcohz /Detecture_ICLR_Benchmarking Detecture ICLR Benchmarking The four evaluation routes reported in the ICLR 2027 submission on sub-semantic image segmentation, bundled so the published numbers can be reproduced from a single download. This is a smaller, paper-aligned release. An earlier bundle, aviadcohz/Detecture_Benchmarking, carried five datasets at 4.3 GB for a previous version of this work. That one included routes the current paper does not report. This release carries only what the paper evaluates on.… See the full description on the dataset page: https://huggingface.co/datasets/aviadcohz/Detecture_ICLR_Benchmarking.imageimage-segmentation1K<n<10K0 likes141 downloads1mo agoHugging Face26kurianbenoy /malayalam_common_voice_benchmarkingtabular1K<n<10K1 likes128 downloads3y agoHugging Face27danliu1226 /cross_species_benchmarking**Repository: https://d-script.readthedocs.io/en/stable/data.html **Reference: Sledzieski, S., Singh, R., Cowen, L. & Berger, B. D-SCRIPT translates genome to phenome with sequence-based, structure-aware, genome-scale predictions of protein-protein interactions. Cell Systems 12, 969-982.e6 (2021). text100K<n<1M2 likes124 downloads1y agoHugging Face28Precise-Debugging-Benchmarking /PDB-Single-Full PDB-Single-Full: Precise Debugging Benchmarking — unfiltered single-line bug pool 📄 Paper &nbsp;·&nbsp; 💻 Code &nbsp;·&nbsp; 🌐 Project page &nbsp;·&nbsp; 🏆 Leaderboard PDB-Single-Full is the unfiltered single-line bug pool of the PDB (Precise Debugging Benchmarking) evaluation suite. Every example pairs a ground-truth program with a synthesized buggy version plus a line-level edit script (gt_diff) that encodes the minimal correct fix. Source datasets: BigCodeBench +… See the full description on the dataset page: https://huggingface.co/datasets/Precise-Debugging-Benchmarking/PDB-Single-Full.tabulartext-generation1K<n<10K0 likes113 downloads5d agoHugging Face29kishor-turing /osworld-benchmarking-gold-files0 likes82 downloads1y agoHugging Face30Precise-Debugging-Benchmarking /PDB-Multi PDB-Multi: Precise Debugging Benchmarking — multi-line bug subset (2–4 line blocks) 📄 Paper &nbsp;·&nbsp; 💻 Code &nbsp;·&nbsp; 🌐 Project page &nbsp;·&nbsp; 🏆 Leaderboard PDB-Multi is the multi-line bug subset (2–4 line blocks) of the PDB (Precise Debugging Benchmarking) evaluation suite. Every example pairs a ground-truth program with a synthesized buggy version plus a line-level edit script (gt_diff) that encodes the minimal correct fix. Source datasets: BigCodeBench +… See the full description on the dataset page: https://huggingface.co/datasets/Precise-Debugging-Benchmarking/PDB-Multi.tabulartext-generationn<1K0 likes81 downloads5d agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.