Team Ai
7 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01EigenformAI /groundtruth-dynamic-benchmarking Groundtruth Dynamic Benchmarking — Geology Question sets and grading rubrics for evaluating LLMs on real-world geological reasoning. Every question is authored from a real source corpus, and every claim in the grading key carries an evidence locator back to that corpus — nothing is synthetic. Licensing/redistribution status varies by corpus — see License. This dataset holds the questions, grading rubrics, and source corpora. Running an evaluation (generating answers from a model… See the full description on the dataset page: https://huggingface.co/datasets/EigenformAI/groundtruth-dynamic-benchmarking.textquestion-answeringn<1K0 likes376 downloads2mo agoHugging Face02Precise-Debugging-Benchmarking /PDB-Single PDB-Single: Precise Debugging Benchmarking — single-line bug set 📄 Paper &nbsp;·&nbsp; 💻 Code &nbsp;·&nbsp; 🌐 Project page &nbsp;·&nbsp; 🏆 Leaderboard PDB-Single is the single-line bug set of the PDB (Precise Debugging Benchmarking) evaluation suite. Every example pairs a ground-truth program with a synthesized buggy version plus a line-level edit script (gt_diff) that encodes the minimal correct fix. Source datasets: BigCodeBench + LiveCodeBench Sibling datasets:… See the full description on the dataset page: https://huggingface.co/datasets/Precise-Debugging-Benchmarking/PDB-Single.tabulartext-generation1K<n<10K1 likes157 downloads5d agoHugging Face03nyamtulla /benchmarking-the-benchmarks-data Benchmarking the Benchmarks — Raw SLM Safety Evaluation Runs Raw evaluation data for the ESORICS 2026 paper: Nyamtulla Shaik, Fengjun Li, Bo Luo — University of Kansas Benchmarking the Benchmarks: Evaluating Automated Safety Benchmarks for Small Language Models. ESORICS 2026. arXiv:2608.17183 Code and processed data: https://github.com/nyamtulla/benchmarking-the-benchmarks ⚠️ Content warning This dataset contains adversarial safety prompts and model responses… See the full description on the dataset page: https://huggingface.co/datasets/nyamtulla/benchmarking-the-benchmarks-data.text-generation100K<n<1M2 likes144 downloads1mo agoHugging Face04Precise-Debugging-Benchmarking /PDB-Single-Full PDB-Single-Full: Precise Debugging Benchmarking — unfiltered single-line bug pool 📄 Paper &nbsp;·&nbsp; 💻 Code &nbsp;·&nbsp; 🌐 Project page &nbsp;·&nbsp; 🏆 Leaderboard PDB-Single-Full is the unfiltered single-line bug pool of the PDB (Precise Debugging Benchmarking) evaluation suite. Every example pairs a ground-truth program with a synthesized buggy version plus a line-level edit script (gt_diff) that encodes the minimal correct fix. Source datasets: BigCodeBench +… See the full description on the dataset page: https://huggingface.co/datasets/Precise-Debugging-Benchmarking/PDB-Single-Full.tabulartext-generation1K<n<10K0 likes113 downloads5d agoHugging Face05Precise-Debugging-Benchmarking /PDB-Multi PDB-Multi: Precise Debugging Benchmarking — multi-line bug subset (2–4 line blocks) 📄 Paper &nbsp;·&nbsp; 💻 Code &nbsp;·&nbsp; 🌐 Project page &nbsp;·&nbsp; 🏆 Leaderboard PDB-Multi is the multi-line bug subset (2–4 line blocks) of the PDB (Precise Debugging Benchmarking) evaluation suite. Every example pairs a ground-truth program with a synthesized buggy version plus a line-level edit script (gt_diff) that encodes the minimal correct fix. Source datasets: BigCodeBench +… See the full description on the dataset page: https://huggingface.co/datasets/Precise-Debugging-Benchmarking/PDB-Multi.tabulartext-generationn<1K0 likes81 downloads5d agoHugging Face06Neura-parse /quantum-error-mitigation-and-benchmarking Neura Parse — Quantum Error Mitigation, Characterization & Benchmarking A pre-fault-tolerance, code-backed vertical on getting trustworthy answers from noisy hardware and rigorously measuring device quality: error-mitigation techniques, characterization/tomography protocols, and benchmarking suites. Runnable Mitiq, pyGSTi, and Qiskit Experiments pipelines with honest sampling-overhead and bias/variance accounting — the practitioner and research toolkit the general dataset… See the full description on the dataset page: https://huggingface.co/datasets/Neura-parse/quantum-error-mitigation-and-benchmarking.tabulartext-generation100K<n<1M0 likes60 downloads3mo agoHugging Face07Precise-Debugging-Benchmarking /PDB-Wild PDB-Wild: Precise Debugging Benchmarking — multi-line and repository-level bugs 📄 Paper &nbsp;·&nbsp; 💻 Code &nbsp;·&nbsp; 🌐 Project page &nbsp;·&nbsp; 🏆 Leaderboard PDB-Wild is the multi-line and repository-level bug set of the PDB (Precise Debugging Benchmarking) evaluation suite. Every example pairs a ground-truth program with a synthesized buggy version plus a line-level edit script (gt_diff) that encodes the minimal correct fix. Source datasets: BigCodeBench +… See the full description on the dataset page: https://huggingface.co/datasets/Precise-Debugging-Benchmarking/PDB-Wild.tabulartext-generationn<1K0 likes46 downloads5d agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.