Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01superlinked /external-benchmarking Vector Search Benchmarks This repo contains datasets for benchmarking vector search performance, to help Superlinked prioritize integration partners. For performing actual benchmarking on this dataset, see the github repository README. Overview We reviewed a number of publicly available datasets and noted 3 core problems + here is how this dataset fixes them: Problems of other vector search benchmarks How this dataset solves it Not enough metadata of… See the full description on the dataset page: https://huggingface.co/datasets/superlinked/external-benchmarking.image10M<n<100M0 likes716 downloads1y agoHugging Face02CollagenHelixLabs /cdsm_benchmarking_data CDSM Collagen Structure Benchmark — Data Structures and scores for a benchmark comparing a deterministic collagen triple-helix builder (CDSM) against four co-folding models — Boltz-2, Chai-1, Protenix and AlphaFold3, the last in both with-MSA (af3_msa) and no-MSA (af3_nomsa) conditions — on 80 experimentally resolved collagen triple helices from the RCSB PDB. Code: https://github.com/bm-howard/cdsm_benchmarking Layout Prefix Contents Size experimental/… See the full description on the dataset page: https://huggingface.co/datasets/CollagenHelixLabs/cdsm_benchmarking_data.tabular10K<n<100K0 likes297 downloads1mo agoHugging Face03kurianbenoy /malayalam_msc_benchmarkingtabular10K<n<100K1 likes179 downloads3y agoHugging Face04kenhktsui /minipile_benchmarkingtabular1M<n<10M0 likes175 downloads2y agoHugging Face05alibustami /UM-DLP-Public-Benchmarking-Dataset UM DLP Public Benchmarking Dataset Description The UM DLP Public Benchmarking Dataset is a publicly available collection designed specifically to stress test Data Loss Prevention (DLP) systems, helping identify detection gaps, false positives, and false negatives for ongoing improvement. This benchmark dataset contains 1,343 manually validated records across six major categories relevant to financial and sensitive data risks: Financial Data (Account information about… See the full description on the dataset page: https://huggingface.co/datasets/alibustami/UM-DLP-Public-Benchmarking-Dataset.tabulartext-classification1K<n<10K4 likes163 downloads1y agoHugging Face06Precise-Debugging-Benchmarking /PDB-Single PDB-Single: Precise Debugging Benchmarking — single-line bug set 📄 Paper &nbsp;·&nbsp; 💻 Code &nbsp;·&nbsp; 🌐 Project page &nbsp;·&nbsp; 🏆 Leaderboard PDB-Single is the single-line bug set of the PDB (Precise Debugging Benchmarking) evaluation suite. Every example pairs a ground-truth program with a synthesized buggy version plus a line-level edit script (gt_diff) that encodes the minimal correct fix. Source datasets: BigCodeBench + LiveCodeBench Sibling datasets:… See the full description on the dataset page: https://huggingface.co/datasets/Precise-Debugging-Benchmarking/PDB-Single.tabulartext-generation1K<n<10K1 likes157 downloads5d agoHugging Face07kurianbenoy /malayalam_common_voice_benchmarkingtabular1K<n<10K1 likes128 downloads3y agoHugging Face08Precise-Debugging-Benchmarking /PDB-Single-Full PDB-Single-Full: Precise Debugging Benchmarking — unfiltered single-line bug pool 📄 Paper &nbsp;·&nbsp; 💻 Code &nbsp;·&nbsp; 🌐 Project page &nbsp;·&nbsp; 🏆 Leaderboard PDB-Single-Full is the unfiltered single-line bug pool of the PDB (Precise Debugging Benchmarking) evaluation suite. Every example pairs a ground-truth program with a synthesized buggy version plus a line-level edit script (gt_diff) that encodes the minimal correct fix. Source datasets: BigCodeBench +… See the full description on the dataset page: https://huggingface.co/datasets/Precise-Debugging-Benchmarking/PDB-Single-Full.tabulartext-generation1K<n<10K0 likes113 downloads5d agoHugging Face09Precise-Debugging-Benchmarking /PDB-Multi PDB-Multi: Precise Debugging Benchmarking — multi-line bug subset (2–4 line blocks) 📄 Paper &nbsp;·&nbsp; 💻 Code &nbsp;·&nbsp; 🌐 Project page &nbsp;·&nbsp; 🏆 Leaderboard PDB-Multi is the multi-line bug subset (2–4 line blocks) of the PDB (Precise Debugging Benchmarking) evaluation suite. Every example pairs a ground-truth program with a synthesized buggy version plus a line-level edit script (gt_diff) that encodes the minimal correct fix. Source datasets: BigCodeBench +… See the full description on the dataset page: https://huggingface.co/datasets/Precise-Debugging-Benchmarking/PDB-Multi.tabulartext-generationn<1K0 likes81 downloads5d agoHugging Face10Neura-parse /quantum-error-mitigation-and-benchmarking Neura Parse — Quantum Error Mitigation, Characterization & Benchmarking A pre-fault-tolerance, code-backed vertical on getting trustworthy answers from noisy hardware and rigorously measuring device quality: error-mitigation techniques, characterization/tomography protocols, and benchmarking suites. Runnable Mitiq, pyGSTi, and Qiskit Experiments pipelines with honest sampling-overhead and bias/variance accounting — the practitioner and research toolkit the general dataset… See the full description on the dataset page: https://huggingface.co/datasets/Neura-parse/quantum-error-mitigation-and-benchmarking.tabulartext-generation100K<n<1M0 likes60 downloads3mo agoHugging Face11africa-intelligence /llama-south-africa-benchmarking Dataset Card for Evaluation run of chad-brouze/llama-8b-south-africa Dataset automatically created during the evaluation run of model chad-brouze/llama-8b-south-africa The dataset is composed of 17 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 14 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/africa-intelligence/llama-south-africa-benchmarking.tabular1K<n<10K0 likes46 downloads2y agoHugging Face12Precise-Debugging-Benchmarking /PDB-Wild PDB-Wild: Precise Debugging Benchmarking — multi-line and repository-level bugs 📄 Paper &nbsp;·&nbsp; 💻 Code &nbsp;·&nbsp; 🌐 Project page &nbsp;·&nbsp; 🏆 Leaderboard PDB-Wild is the multi-line and repository-level bug set of the PDB (Precise Debugging Benchmarking) evaluation suite. Every example pairs a ground-truth program with a synthesized buggy version plus a line-level edit script (gt_diff) that encodes the minimal correct fix. Source datasets: BigCodeBench +… See the full description on the dataset page: https://huggingface.co/datasets/Precise-Debugging-Benchmarking/PDB-Wild.tabulartext-generationn<1K0 likes46 downloads5d agoHugging Face13africa-intelligence /aya101-benchmarking Dataset Card for Evaluation run of CohereForAI/aya-101 Dataset automatically created during the evaluation run of model CohereForAI/aya-101 The dataset is composed of 5 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/africa-intelligence/aya101-benchmarking.tabular1K<n<10K0 likes36 downloads2y agoHugging Face14electricsheepafrica /africa-synth-energy-pv-performance-benchmarking-africa-benin Africa Synth Energy Pv Performance Benchmarking Africa Benin | Africa (Electric Sheep Africa metadata inventory) Size category: 10K<n<100K - Formats: csv - Sector: energy - Engineered by Electric Sheep Africa TL;DR This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context. What This Dataset Covers… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-energy-pv-performance-benchmarking-africa-benin.tabulartabular-classification10K<n<100K0 likes28 downloads2mo agoHugging Face15xorushi /UM-DLP-Public-Benchmarking-Dataset UM DLP Public Benchmarking Dataset Description The UM DLP Public Benchmarking Dataset is a publicly available collection designed specifically to stress test Data Loss Prevention (DLP) systems, helping identify detection gaps, false positives, and false negatives for ongoing improvement. This benchmark dataset contains 1,343 manually validated records across six major categories relevant to financial and sensitive data risks: Financial Data (Account information… See the full description on the dataset page: https://huggingface.co/datasets/xorushi/UM-DLP-Public-Benchmarking-Dataset.tabulartext-classification1K<n<10K0 likes25 downloads2mo agoHugging Face16africa-intelligence /llama-benchmarking Dataset Card for Evaluation run of meta-llama/Llama-3.1-8B-Instruct Dataset automatically created during the evaluation run of model meta-llama/Llama-3.1-8B-Instruct The dataset is composed of 5 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/africa-intelligence/llama-benchmarking.tabular1K<n<10K1 likes19 downloads2y agoHugging Face17AI4BD /Translation_Benchmarking_datasets_alltabular10K<n<100K0 likes19 downloads2y agoHugging Face18onepaneai /hallucination-invalid-questions-mysql-explanation-benchmarkingtabularn<1K1 likes17 downloads2y agoHugging Face19africa-intelligence /InkubaLM-benchmarking Dataset Card for Evaluation run of lelapa/InkubaLM-0.4B Dataset automatically created during the evaluation run of model lelapa/InkubaLM-0.4B The dataset is composed of 5 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/africa-intelligence/InkubaLM-benchmarking.tabular1K<n<10K1 likes16 downloads2y agoHugging Face20lamm-mit /collagen-cdsm_benchmarking_datagated CDSM Collagen Structure Benchmark — Data Structures and scores for a benchmark comparing a deterministic collagen triple-helix builder (CDSM) against four co-folding models — Boltz-2, Chai-1, Protenix and AlphaFold3, the last in both with-MSA (af3_msa) and no-MSA (af3_nomsa) conditions — on 80 experimentally resolved collagen triple helices from the RCSB PDB. Code: https://github.com/bm-howard/cdsm_benchmarking Layout Prefix Contents Size experimental/… See the full description on the dataset page: https://huggingface.co/datasets/lamm-mit/collagen-cdsm_benchmarking_data.tabular10K<n<100K0 likes15 downloads1mo agoHugging Face21onepaneai /hallucination-valid-questions-mysql-explanation-benchmarkingtabularn<1K0 likes14 downloads2y agoHugging Face22onepaneai /profanity-gpt-spl-benchmarkingtabularn<1K0 likes13 downloads2y agoHugging Face23onepaneai /hallucination-invalid-questions-mysql-falcon-explanation-benchmarkingtabularn<1K0 likes12 downloads2y agoHugging Face24MKipke /benchmarking-wikidataimage1K<n<10K0 likes12 downloads7mo agoHugging Face25onepaneai /faithfulness-f1score-spl-prompt-falcon-benchmarkingtabularn<1K0 likes11 downloads2y agoHugging Face26AI4BD /Translation_Benchmarking_datasets_all_florestabular1K<n<10K0 likes11 downloads2y agoHugging Face27onepaneai /faithfulness-f1score-spl-prompt-gpt-benchmarkingtabularn<1K0 likes9 downloads2y agoHugging Face28onepaneai /polarity-gpt-spl-benchmarkingtabularn<1K0 likes9 downloads2y agoHugging Face29africa-intelligence /aya23-benchmarking Dataset Card for Evaluation run of CohereForAI/aya-23-8B Dataset automatically created during the evaluation run of model CohereForAI/aya-23-8B The dataset is composed of 5 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/africa-intelligence/aya23-benchmarking.tabular1K<n<10K1 likes9 downloads2y agoHugging Face30africa-intelligence /aya-benchmarking Dataset Card for Evaluation run of CohereForAI/aya-23-8B Dataset automatically created during the evaluation run of model CohereForAI/aya-23-8B The dataset is composed of 5 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/africa-intelligence/aya-benchmarking.tabularn<1K0 likes6 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.