datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cmevs-erp-eval
CM-EVS: A Coverage-Curated Panoramic RGB-D Dataset for Indoor Scene Understanding
CM-EVS is a curated panoramic RGB-D dataset built under a single principle: maximize the geometric coverage of a 3D scene with the fewest equirectangular (ERP) frames possible. The release is structured as one redistributable Blender indoor data archive plus four license-aware adapter packages that regenerate matched frames locally from upstream sources whose terms forbid redistribution.
v1.0… See the full description on the dataset page: https://huggingface.co/datasets/anon-cmevs-2026/cmevs-erp-eval.MVU-Eval-Data
MVU-Eval Dataset
Paper | Code | Project Page
Dataset Description
The advent of Multimodal Large Language Models (MLLMs) has expanded AI capabilities to visual modalities, yet existing evaluation benchmarks remain limited to single-video understanding, overlooking the critical need for multi-video understanding in real-world scenarios (e.g., sports analytics and autonomous driving). To address this significant gap, we introduce MVU-Eval, the first comprehensive benchmark… See the full description on the dataset page: https://huggingface.co/datasets/MVU-Eval-Team/MVU-Eval-Data.imdb-ciOpenVLHarness-Evaluation-Datasets
OpenVLHarness evaluation datasets
Processed evaluation splits used by
OpenVLHarness
(project page).
Each <split>.tsv holds the exact prompts (question) and annotations
(answer plus metadata) we evaluate on; image_path is relative to
images/<folder>/ inside images/<folder>.zip. GLIP.zip holds the ODinW-13
configs and COCO-format val/test annotations used for ODinW AP evaluation.
You normally don't need to download anything by hand: running
openvlharness-eval --data <split> ...… See the full description on the dataset page: https://huggingface.co/datasets/uclanlp/OpenVLHarness-Evaluation-Datasets.K2-EvalResearch Paper coming soon!
K2EvalK^{2} EvalK2Eval
K2EvalK^{2} EvalK2Eval is a novel benchmark featuring 90 handwritten instructions that require in-depth knowledge of Korean language and culture for accurate completion.
Benchmark Overview
The design principle behind K2EvalK^{2} EvalK2Eval centers on collecting instructions that necessitate knowledge specific to Korean culture and context in order to solve. This approach distinguishes our work from simply translating… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/K2-Eval.narrow-model-safety-eval
Narrow Model Safety Evaluation — Protein Dual-Use Risk Dataset
Summary: Annotations, results, and evaluation data for a proof-of-concept framework assessing dual-use risk in narrow scientific AI models. Two lines of work: (1) structure-level metrics — FSPE, FSI, and Physical Realizability Tier — on eight published protein toxins and mechanism-matched benign controls (ESM-2, ProteinMPNN); (2) mechanism generalization — a leave-one-mechanism-out panel measuring what an… See the full description on the dataset page: https://huggingface.co/datasets/jang1563/narrow-model-safety-eval.staining-robustness-evaluation
A Protocol for Evaluating Robustness to H&E Staining Variation in Computational Pathology Models
This repository provides the stain references, pretrained models, and experimental results required to:
Define custom staining references using our PLISM reference library
Reproduce our published controlled staining robustness experiments
👉 Code repository: https://github.com/lely475/staining-robustness-evaluation/tree/main
👉 Associated publication: Paper
Overview: How… See the full description on the dataset page: https://huggingface.co/datasets/CTPLab-DBE-UniBas/staining-robustness-evaluation.PhysicalAI-Robotics-PhysicalAssets-VoMP-Eval
VoMP: Predicting Volumetric Mechanical Properties
Dataset Description:
The Pre-Processed 3D Dataset is a dataset that is composed of 4 individual 3D asset datasets which are processed to render them from multiple views, voxelize the assets, and propagate VLM annotations for material properties.
We release pre-processed data derived from the 3D assets, specifically: voxels, rendered images, and LLM-annotated material descriptions.
This dataset is for research and… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-Robotics-PhysicalAssets-VoMP-Eval.JobBERT-evaluation-dataset
JobBERT evaluation dataset 💾
This is the official repository containing the evaluation data that was used for the JobBERT paper. This dataset is a list of vacancy titles, each tagged with an ESCO (v1.0.5) occupation.
The full dataset is split into two files in a stratified way by class distribution. This data was automatically collected from a large governmental job board.
Access the JobBERT paper here: https://arxiv.org/abs/2109.09605
BibTeX Citation
If you use this… See the full description on the dataset page: https://huggingface.co/datasets/TechWolf/JobBERT-evaluation-dataset.Raon-OpenTTS-Eval
Raon-OpenTTS-Eval
Technical Report
A robustness-oriented evaluation benchmark for zero-shot text-to-speech, covering 4 acoustic regimes (Clean, Noisy, Wild, Expressive) across 12 datasets with 6,000 prompt–text pairs.
Existing zero-shot TTS benchmarks typically evaluate models using prompts drawn from a single read-speech dataset, providing an incomplete view of robustness under realistic and challenging recording scenarios. Raon-OpenTTS-Eval… See the full description on the dataset page: https://huggingface.co/datasets/KRAFTON/Raon-OpenTTS-Eval.shoulders-of-giants
Shoulders of Giants
Can AI agents build on a scientist's work and write a follow-up paper? Each of the 15 tasks gives an agent a
published (or recently submitted) scientific paper and a follow-up research direction proposed by the paper's own
author. The agent has to carry out the research and write the follow-up paper. Code, rubrics and graders:
GitHub. Leaderboard and every judge verdict:
website.
Subsets
Subset
Content
benchmark/
The 15 tasks:… See the full description on the dataset page: https://huggingface.co/datasets/prometheus-eval/shoulders-of-giants.wmt-mqm-human-evaluation
Dataset Summary
This dataset contains all MQM human annotations from previous WMT Metrics shared tasks and the MQM annotations from Experts, Errors, and Context.
The data is organised into 8 columns:
lp: language pair
src: input text
mt: translation
ref: reference translation
score: MQM score
system: MT Engine that produced the translation
annotators: number of annotators
domain: domain of the input text (e.g. news)
year: collection year
You can also find the original data here.… See the full description on the dataset page: https://huggingface.co/datasets/RicardoRei/wmt-mqm-human-evaluation.tulu-3-harmbench-evalThis data comes from the HarmBench benchmark.
This is one of the datasets included in the Ai2 Safety Evaluation Suite, and the Tülu 3 evaluation suite.
The repo for Ai2's safety suite includes instructions on how to evaluate models on various safety-related evaluation including this one.
mev1_evalsbeetle-merge-eval
Beetle merged models — benchmark evaluation against their parents
Minimal-pair benchmark accuracy for the Beetle merged models published in the
Mergeability org, scored against their own parent models and, where one
exists, the jointly-trained ceiling on the same harness.
The existing sweep datasets (Mergeability/merge-sweep-results,
Mergeability-2/mergeability-results) record merge quality in nats (NLL,
delta_floor, rel_damage, barrier, geometry). They contain no downstream… See the full description on the dataset page: https://huggingface.co/datasets/Cross-Mergeability/beetle-merge-eval.RAG-Evaluation-Dataset-KO
Allganize RAG Leaderboard
Allganize RAG 리더보드는 5개 도메인(금융, 공공, 의료, 법률, 커머스)에 대해서 한국어 RAG의 성능을 평가합니다.일반적인 RAG는 간단한 질문에 대해서는 답변을 잘 하지만, 문서의 테이블과 이미지에 대한 질문은 답변을 잘 못합니다.
RAG 도입을 원하는 수많은 기업들은 자사에 맞는 도메인, 문서 타입, 질문 형태를 반영한 한국어 RAG 성능표를 원하고 있습니다.평가를 위해서는 공개된 문서와 질문, 답변 같은 데이터 셋이 필요하지만, 자체 구축은 시간과 비용이 많이 드는 일입니다.이제 올거나이즈는 RAG 평가 데이터를 모두 공개합니다.
RAG는 Parser, Retrieval, Generation 크게 3가지 파트로 구성되어 있습니다.현재, 공개되어 있는 RAG 리더보드 중, 3가지 파트를 전체적으로 평가하는 한국어로 구성된 리더보드는 없습니다.
Allganize RAG 리더보드에서는 문서를… See the full description on the dataset page: https://huggingface.co/datasets/allganize/RAG-Evaluation-Dataset-KO.iac-eval
IaC-Eval dataset (v1.1)
IaC-Eval dataset is the first human-curated and challenging Cloud Infrastructure-as-Code (IaC) dataset tailored to more rigorously benchmark large language models' IaC code generation capabilities.
This dataset contains 458 questions ranging from simple to difficult across various cloud services (targeting AWS for now).
| Github | 🏆 Leaderboard TBD | 📖 NeurIPS 2024 Paper |
2. Usage instructions
Option 1: Running the evaluation… See the full description on the dataset page: https://huggingface.co/datasets/autoiac-project/iac-eval.rhan-eval-sweepSWEbench-Verified-eval150-u355-M2.7-orch-cost-20260923
Fixed Solo350 u355 + MiniMax-M2.7: orchestration cost study
Best observed cost tradeoff: compact coordinator decisions plus soft review at the existing hard limit (at most 12 worker turns). M2.7 metered token cost falls 59.3%, while mean solved tasks decrease from 90.00 to 87.67/150. Accuracy equivalence was not established.
This closed study contains 6 designs and 16 complete independent runs on the same 150 tasks (2400 scored task/run pairs), each with an independent audit.… See the full description on the dataset page: https://huggingface.co/datasets/CharlieLLL/SWEbench-Verified-eval150-u355-M2.7-orch-cost-20260923.RAG-Evaluation-Dataset-JA
Allganize RAG Leaderboard とは
Allganize RAG Leaderboard は、5つの業種ドメイン(金融、情報通信、製造、公共、流通・小売)において、日本語のRAGの性能評価を実施したものです。一般的なRAGは簡単な質問に対する回答は可能ですが、図表の中に記載されている情報などに対して回答できないケースが多く存在します。RAGの導入を希望する多くの企業は、自社と同じ業種ドメイン、文書タイプ、質問形態を反映した日本語のRAGの性能評価を求めています。RAGの性能評価には、検証ドキュメントや質問と回答といったデータセット、検証環境の構築が必要となりますが、AllganizeではRAGの導入検討の参考にしていただきたく、日本語のRAG性能評価に必要なデータを公開いたしました。RAGソリューションは、Parser、Retrieval、Generation の3つのパートで構成されています。現在、この3つのパートを総合的に評価した日本語のRAG Leaderboardは存在していません。(公開時点)Allganize RAG… See the full description on the dataset page: https://huggingface.co/datasets/allganize/RAG-Evaluation-Dataset-JA.MMLU-Pro-CoT-Eval
Dataset Details
Modality: Text
Format: CSV
Size: 100K - 1M rows
Total Rows: 248,836
License: MIT
Libraries Supported: datasets, pandas, croissant
Structure
Each row in the dataset includes:
question: The query posed in the dataset.
answer: The correct response.
category: The domain of the question (e.g., math, science).
src: The source of the question.
id: A unique identifier for each entry.
chain_of_thoughts: Step-by-step reasoning steps leading to the answer.… See the full description on the dataset page: https://huggingface.co/datasets/UW-Madison-Lee-Lab/MMLU-Pro-CoT-Eval.llava-15-rlmpq-vlm-eval-results
RL-MPQ VLM Evaluation Artifacts
Complete figures, tables, galleries, and raw benchmark CSVs for the extended VLM evaluation.
Dataset: AvoCahDoe/llava-15-rlmpq-vlm-eval-results
Collections (by base VLM)
RL-MPQ VLM — LLaVA-1.5-13B — HF collection
RL-MPQ VLM — LLaVA-1.5-7B — HF collection
RL-MPQ VLM — LLaVA-Next Mistral-7B — HF collection
RL-MPQ VLM — Qwen2-VL-7B — HF collection
Model repos
RL-MPQ High Fidelity →… See the full description on the dataset page: https://huggingface.co/datasets/AvoCahDoe/llava-15-rlmpq-vlm-eval-results.circuitlens-gemma-2-2btranscoder-descriptions-and-evaluations
CircuitLens & WeightLens: Transcoder Descriptions and Evaluations
This dataset contains automatically generated descriptions and evaluation metrics for Gemma-2-2B transcoders, produced using CircuitLens and WeightLens methods.
Methods
CircuitLens: https://github.com/egolimblevskaia/CircuitLens
WeightLens: https://github.com/egolimblevskaia/WeightLens
Dataset Structure
The dataset is organized by layers (0, 4, 7, 10, 12, 15, 18, 21, 23, 25), with each layer… See the full description on the dataset page: https://huggingface.co/datasets/egolimblevskaia/circuitlens-gemma-2-2btranscoder-descriptions-and-evaluations.matrix-game-evalwmt-da-human-evaluation
Dataset Summary
This dataset contains all DA human annotations from previous WMT News Translation shared tasks.
The data is organised into 8 columns:
lp: language pair
src: input text
mt: translation
ref: reference translation
score: z score
raw: direct assessment
annotators: number of annotators
domain: domain of the input text (e.g. news)
year: collection year
You can also find the original data for each year in the results section https://www.statmt.org/wmt{YEAR}/results.html… See the full description on the dataset page: https://huggingface.co/datasets/RicardoRei/wmt-da-human-evaluation.paperswithcode-data-evaluation-tables
Process data from paperswithcode
See https://huggingface.co/datasets/pwc-archive/files/tree/main.
Download and unzip evaluation tables:
curl -L -O "https://huggingface.co/datasets/pwc-archive/files/resolve/main/jul-28-evaluation-tables.json.gz"
gunzip jul-28-evaluation-tables.json.gz
Install jq.
See https://jqlang.org/.
If on Debian/Ubuntu, install with sudo apt-get install jq.
Example jq to extract:
jq -r '
def process(parent):
.task as $current_task |
(if parent then… See the full description on the dataset page: https://huggingface.co/datasets/felixleungsc/paperswithcode-data-evaluation-tables.RAG_Evaluation_Datasetapt-eval
🚨 APT-Eval Dataset 🚨
Almost AI, Almost Human: The Challenge of Detecting AI-Polished Writing
📝 Paper, 🖥️ Github, 🎥 Recording
This repository contains the official dataset of the ACL 2025 paper 'Almost AI, Almost Human: The Challenge of Detecting AI-Polished Writing'
APT-Eval is the first and largest dataset to evaluate the AI-text detectors behavior for AI-polished texts.
It contains almost 15K text samples, polished by 5 different LLMs, for 6 different domains, with 2 major… See the full description on the dataset page: https://huggingface.co/datasets/smksaha/apt-eval.compliance-sycophancy-cot
Compliance-Sycophancy CoT Analysis
When compliance-forcing instructions cause frontier AI models to fabricate answers, the models know they are fabricating.
Reading the reasoning traces of DeepSeek V4 Pro (129 traces) and Qwen3-80B (41 traces) reveals that 100% of fabrication cases show the model explicitly recognizing insufficient context, referencing the compliance instruction, and deliberately overriding its own uncertainty. A one-sentence defense phrase ("if you lack… See the full description on the dataset page: https://huggingface.co/datasets/schema-eval/compliance-sycophancy-cot.olmo-eval-strongrejectThis data comes from the StrongREJECT benchmark.
This is one of the datasets included in the Ai2 Safety Evaluation Suite, and the Olmo evaluation suite.
The repo for Ai2's safety suite includes instructions on how to evaluate models on various safety-related evaluations, including this one.
Permitted Use
The data is provided for benchmarking and evaluation purposes only. It is intended for research and educational use in accordance with Ai2's Responsible Use Guidelines.
Disclaimer
This… See the full description on the dataset page: https://huggingface.co/datasets/allenai/olmo-eval-strongreject.
