datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
OpenVLHarness-Evaluation-Datasets
OpenVLHarness evaluation datasets
Processed evaluation splits used by
OpenVLHarness
(project page).
Each <split>.tsv holds the exact prompts (question) and annotations
(answer plus metadata) we evaluate on; image_path is relative to
images/<folder>/ inside images/<folder>.zip. GLIP.zip holds the ODinW-13
configs and COCO-format val/test annotations used for ODinW AP evaluation.
You normally don't need to download anything by hand: running
openvlharness-eval --data <split> ...… See the full description on the dataset page: https://huggingface.co/datasets/uclanlp/OpenVLHarness-Evaluation-Datasets.staining-robustness-evaluation
A Protocol for Evaluating Robustness to H&E Staining Variation in Computational Pathology Models
This repository provides the stain references, pretrained models, and experimental results required to:
Define custom staining references using our PLISM reference library
Reproduce our published controlled staining robustness experiments
👉 Code repository: https://github.com/lely475/staining-robustness-evaluation/tree/main
👉 Associated publication: Paper
Overview: How… See the full description on the dataset page: https://huggingface.co/datasets/CTPLab-DBE-UniBas/staining-robustness-evaluation.JobBERT-evaluation-dataset
JobBERT evaluation dataset 💾
This is the official repository containing the evaluation data that was used for the JobBERT paper. This dataset is a list of vacancy titles, each tagged with an ESCO (v1.0.5) occupation.
The full dataset is split into two files in a stratified way by class distribution. This data was automatically collected from a large governmental job board.
Access the JobBERT paper here: https://arxiv.org/abs/2109.09605
BibTeX Citation
If you use this… See the full description on the dataset page: https://huggingface.co/datasets/TechWolf/JobBERT-evaluation-dataset.wmt-mqm-human-evaluation
Dataset Summary
This dataset contains all MQM human annotations from previous WMT Metrics shared tasks and the MQM annotations from Experts, Errors, and Context.
The data is organised into 8 columns:
lp: language pair
src: input text
mt: translation
ref: reference translation
score: MQM score
system: MT Engine that produced the translation
annotators: number of annotators
domain: domain of the input text (e.g. news)
year: collection year
You can also find the original data here.… See the full description on the dataset page: https://huggingface.co/datasets/RicardoRei/wmt-mqm-human-evaluation.RAG-Evaluation-Dataset-KO
Allganize RAG Leaderboard
Allganize RAG 리더보드는 5개 도메인(금융, 공공, 의료, 법률, 커머스)에 대해서 한국어 RAG의 성능을 평가합니다.일반적인 RAG는 간단한 질문에 대해서는 답변을 잘 하지만, 문서의 테이블과 이미지에 대한 질문은 답변을 잘 못합니다.
RAG 도입을 원하는 수많은 기업들은 자사에 맞는 도메인, 문서 타입, 질문 형태를 반영한 한국어 RAG 성능표를 원하고 있습니다.평가를 위해서는 공개된 문서와 질문, 답변 같은 데이터 셋이 필요하지만, 자체 구축은 시간과 비용이 많이 드는 일입니다.이제 올거나이즈는 RAG 평가 데이터를 모두 공개합니다.
RAG는 Parser, Retrieval, Generation 크게 3가지 파트로 구성되어 있습니다.현재, 공개되어 있는 RAG 리더보드 중, 3가지 파트를 전체적으로 평가하는 한국어로 구성된 리더보드는 없습니다.
Allganize RAG 리더보드에서는 문서를… See the full description on the dataset page: https://huggingface.co/datasets/allganize/RAG-Evaluation-Dataset-KO.RAG-Evaluation-Dataset-JA
Allganize RAG Leaderboard とは
Allganize RAG Leaderboard は、5つの業種ドメイン(金融、情報通信、製造、公共、流通・小売)において、日本語のRAGの性能評価を実施したものです。一般的なRAGは簡単な質問に対する回答は可能ですが、図表の中に記載されている情報などに対して回答できないケースが多く存在します。RAGの導入を希望する多くの企業は、自社と同じ業種ドメイン、文書タイプ、質問形態を反映した日本語のRAGの性能評価を求めています。RAGの性能評価には、検証ドキュメントや質問と回答といったデータセット、検証環境の構築が必要となりますが、AllganizeではRAGの導入検討の参考にしていただきたく、日本語のRAG性能評価に必要なデータを公開いたしました。RAGソリューションは、Parser、Retrieval、Generation の3つのパートで構成されています。現在、この3つのパートを総合的に評価した日本語のRAG Leaderboardは存在していません。(公開時点)Allganize RAG… See the full description on the dataset page: https://huggingface.co/datasets/allganize/RAG-Evaluation-Dataset-JA.circuitlens-gemma-2-2btranscoder-descriptions-and-evaluations
CircuitLens & WeightLens: Transcoder Descriptions and Evaluations
This dataset contains automatically generated descriptions and evaluation metrics for Gemma-2-2B transcoders, produced using CircuitLens and WeightLens methods.
Methods
CircuitLens: https://github.com/egolimblevskaia/CircuitLens
WeightLens: https://github.com/egolimblevskaia/WeightLens
Dataset Structure
The dataset is organized by layers (0, 4, 7, 10, 12, 15, 18, 21, 23, 25), with each layer… See the full description on the dataset page: https://huggingface.co/datasets/egolimblevskaia/circuitlens-gemma-2-2btranscoder-descriptions-and-evaluations.wmt-da-human-evaluation
Dataset Summary
This dataset contains all DA human annotations from previous WMT News Translation shared tasks.
The data is organised into 8 columns:
lp: language pair
src: input text
mt: translation
ref: reference translation
score: z score
raw: direct assessment
annotators: number of annotators
domain: domain of the input text (e.g. news)
year: collection year
You can also find the original data for each year in the results section https://www.statmt.org/wmt{YEAR}/results.html… See the full description on the dataset page: https://huggingface.co/datasets/RicardoRei/wmt-da-human-evaluation.paperswithcode-data-evaluation-tables
Process data from paperswithcode
See https://huggingface.co/datasets/pwc-archive/files/tree/main.
Download and unzip evaluation tables:
curl -L -O "https://huggingface.co/datasets/pwc-archive/files/resolve/main/jul-28-evaluation-tables.json.gz"
gunzip jul-28-evaluation-tables.json.gz
Install jq.
See https://jqlang.org/.
If on Debian/Ubuntu, install with sudo apt-get install jq.
Example jq to extract:
jq -r '
def process(parent):
.task as $current_task |
(if parent then… See the full description on the dataset page: https://huggingface.co/datasets/felixleungsc/paperswithcode-data-evaluation-tables.RAG_Evaluation_Datasetjapanese-image-classification-evaluation-dataset
recruit-jp/japanese-image-classification-evaluation-dataset
Overview
Developed by: Recruit Co., Ltd.
Dataset type: Image Classification
Language(s): Japanese
LICENSE: CC-BY-4.0
More details are described in our tech blog post.
日本語CLIP学習済みモデルとその評価用データセットの公開
Dataset Details
This dataset is comprised of four image classification tasks related to concepts and things unique to Japan. Specifically, is consists of the following tasks.
jafood101: Image… See the full description on the dataset page: https://huggingface.co/datasets/recruit-jp/japanese-image-classification-evaluation-dataset.toporisk-evaluation-artifacts
TopoRisk Evaluation Artifacts
English | 简体中文
TopoRisk studies how a limited verification budget should be allocated across
an agentic workflow graph. Instead of ranking steps only by their local failure
probability, the scheduler estimates how an error can reach terminal outputs
and recomputes marginal value after every selected audit.
This repository is an artifact-first research preview. It contains code,
aggregate measurements, and figures. The manuscript and its LaTeX… See the full description on the dataset page: https://huggingface.co/datasets/JosephAA/toporisk-evaluation-artifacts.IELTS-writing-task-2-evaluationurdu-english-llm-evaluation
Urdu-English Evaluation Dataset: Testing Qwen, Gemma, and Llama
Testing three small open-source language models on a dataset of a hundred questions in Urdu and English, across five categories.
Motivation
It all starts when I noticed, while using voice and chat-based AI tools, that Urdu is often not handled as well as English, and I suspected that models aren't trained as extensively on Urdu compared to other languages like Hindi, that made me wonder if even… See the full description on the dataset page: https://huggingface.co/datasets/Momina-Muzafar/urdu-english-llm-evaluation.crimson-ma2-evaluation
Crimson MolmoAct2: preliminary real-robot evaluation
Right-arm bottle pick-and-place rollouts recorded on 2026-09-24 (Asia/Bangkok).
This release contains three checkpoints, 31 logged attempts, videos from three cameras,
selected recorded telemetry, and joint trajectories. It is an evaluation evidence release,
not a training dataset or a completed six-checkpoint benchmark.
Model source: Kavin60606/crimson-ma2-ckpts.
The policy task text was pick and place the bottle. Each… See the full description on the dataset page: https://huggingface.co/datasets/CrimsonRobot/crimson-ma2-evaluation.evaluation5evaluation1evaluation2ecommerce-analytics-sql-evaluation
Ecommerce Analytics SQL Evaluation (declared GMV, verified answer key)
An evalpack: an evaluation database generated from the answer key, not
annotated after the fact. A VLDB 2026 audit found 52.8% of BIRD Mini-Dev
answer keys wrong because benchmarks annotate answers onto existing
databases; this dataset inverts the order. The declared properties (curves,
shares, identities) are the specification, the database is generated to
satisfy them exactly, and every shipped question was… See the full description on the dataset page: https://huggingface.co/datasets/rasinmuhammed/ecommerce-analytics-sql-evaluation.llm-math-evaluation-dataset
LLM Math Response Evaluation Dataset
Dataset Summary
A human-annotated dataset of 150 AI-generated math responses
evaluated across GPT-4o, Claude, and Gemini. Each response is
scored on Correctness, Reasoning, and Clarity using a structured
rubric, with written justification for every score.
Supported Tasks
LLM evaluation and benchmarking
Math reasoning quality assessment
Error type classification in AI responses
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/nbvbharath-1729/llm-math-evaluation-dataset.mHallucination_Evaluation
Multilingual Hallucination Evaluation in the wild
The dataset was as part of the paper: How Much Do LLMs Hallucinate across Languages? On Multilingual Estimation of LLM Hallucination in the Wild
Below is the figure summarizing the multilingual hallucination evaluation dataset creation (and multilingual hallucination detection dataset):
Dataset Details
The dataset is a high quality synthetic query/prompt and wikipedia reference pair for estimating hallucinations in the… See the full description on the dataset page: https://huggingface.co/datasets/WueNLP/mHallucination_Evaluation.NLU-Evaluation-Data-en-de
NLU Evaluation Data - English and German
A labeled English and German language multi-domain dataset (21 domains) with 25K user utterances for human-robot interaction.
This dataset is collected and annotated for evaluating NLU services and platforms.
The detailed paper on this dataset can be found at arXiv.org:
Benchmarking Natural Language Understanding Services for building Conversational Agents
The dataset builds on the annotated data of the xliuhw/NLU-Evaluation-Data
repository.… See the full description on the dataset page: https://huggingface.co/datasets/deutsche-telekom/NLU-Evaluation-Data-en-de.ai-blind-spot-evaluation-dataset
Evaluating Diagnostic Recognition Blind Spots in Small Models: A Cross-Model Study of Endometriosis
Note: this is the dataset for the Fatima Fellowship Application Fall 2026 technical challenge.
Answer to Question 1 and Question 3
Healthcare practitioners, when faced with atypical manifestations of diseases, often resort to hypothesis generation and revision if classic pattern recognition doesn’t suffice. This raises the question that sits at the heart of this evaluation: How… See the full description on the dataset page: https://huggingface.co/datasets/duaa-amer/ai-blind-spot-evaluation-dataset.YouTube-Evaluation-Set
Awaaz se Alfaaz — YouTube Evaluation Set
This dataset is the realistic multi-speaker evaluation set used in Awaaz se Alfaaz, accepted at LaTeLL 2026 — "Enhancing Urdu ASR with Whisper v3: Fine-Tuning on Latest Datasets and Realistic Multi-Speaker Evaluation with SLM Post-Processing." It contains 30 short-form Urdu YouTube videos (YouTube Shorts) covering a mix of news, sports, and current affairs content, along with human annotated gold transcripts and transcripts produced by… See the full description on the dataset page: https://huggingface.co/datasets/awaaz-se-alfaaz/YouTube-Evaluation-Set.libero-plus-evaluation
LIBERO-plus Evaluation Dataset (180 held-out tasks)
This dataset contains the evaluation tasks used for validating the pi0.5 LIBERO-plus LoRA model on the LIBERO-plus benchmark.
It consists of a meticulously designed set of 180 tasks, distributed along 4 perturbation axes to rigorously evaluate the robustness and generalization capabilities of Vision-Language-Action (VLA) models:
Base — Standard BDDL scenes. Camera (0, 0, 100, 0, 0), fixed object poses, no noise.
View — Camera… See the full description on the dataset page: https://huggingface.co/datasets/nosuke113/libero-plus-evaluation.cqa-ai-technical-response-evaluation
CQA AI Technical Response Evaluation Dataset
Overview
This is a synthetic dataset designed to demonstrate structured evaluation of AI-generated technical responses.
The dataset evaluates technical responses beyond a simple correct/incorrect classification by considering multiple dimensions of response quality, including accuracy, reasoning validity, completeness, consistency, assumption handling, clarity, error classification, severity, and expert evaluation.… See the full description on the dataset page: https://huggingface.co/datasets/CQA-pharma/cqa-ai-technical-response-evaluation.fon-code-switching-evaluation
French-Fon Code-Switching Evaluation Benchmark
Overview
This dataset is a benchmark designed to evaluate the contextual understanding of small language models in French-Fon code-switching scenarios.
The benchmark focuses on situations in which French and Fon (Fongbé) are used within the same interaction, with particular attention to cases where critical information required to answer a question is provided in Fon.
The benchmark was developed as part of an academic… See the full description on the dataset page: https://huggingface.co/datasets/fai-adh/fon-code-switching-evaluation.AudioVisual-Benchmark-Evaluation
AudioVisual Benchmark Evaluation — evaluation subsets
Item-id lists for the audio-visual benchmark subsets used in our reported
evaluation tables.
Layout
<benchmark>/eval_subset.csv item ids evaluated in the paper
<benchmark>/media_index.csv id -> media filename(s)
<benchmark>/media/ the media files those ids refer to
eval_subset.csv holds a single id column keyed to the source benchmark
(question_id, idx, or index). media/ contains exactly the… See the full description on the dataset page: https://huggingface.co/datasets/plnguyen2908/AudioVisual-Benchmark-Evaluation.wmt-sqm-human-evaluation
Dataset Summary
In 2022, several changes were made to the annotation procedure used in the WMT Translation task. In contrast to the standard DA (sliding scale from 0-100) used in previous years, in 2022 annotators performed DA+SQM (Direct Assessment + Scalar Quality Metric). In DA+SQM, the annotators still provide a raw score between 0 and 100, but also are presented with seven labeled tick marks. DA+SQM helps to stabilize scores across annotators (as compared to DA).
The data is… See the full description on the dataset page: https://huggingface.co/datasets/RicardoRei/wmt-sqm-human-evaluation.responsible-agent-workflow-evaluation
Responsible Agent Workflow Evaluation
Version 1.0.0 contains 130 wholly synthetic scenarios for evaluating
whether an AI agent respects safety, permission and accountability boundaries
in operational settings. Thirteen categories contain ten scenarios each. Every
record includes an intentionally unsafe request, contextual facts, expected
safe behaviour, explicitly prohibited behaviour, severity, evaluation criteria
and reviewer guidance.
This is a red-team and… See the full description on the dataset page: https://huggingface.co/datasets/nwhite-systems/responsible-agent-workflow-evaluation.
