datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GDP-Val-Evaluation-Submission
GDPval Submission Dataset
This dataset contains model outputs for GDP-Val evaluation.
Dataset Structure
data/: Contains the main dataset in Parquet format
train-00000-of-00001.parquet: Submission data with model outputs
deliverable_files/: Contains generated files for tasks that produce file deliverables
Organized by task_id
dataset_info.json: Metadata about the dataset
Columns
task_id: Unique identifier for each task
sector: Economic sector for the task… See the full description on the dataset page: https://huggingface.co/datasets/xiachongfeng/GDP-Val-Evaluation-Submission.aya_evaluation_suite
Dataset Summary
Aya Evaluation Suite contains a total of 26,750 open-ended conversation-style prompts to evaluate multilingual open-ended generation quality.To strike a balance between language coverage and the quality that comes with human curation, we create an evaluation suite that includes:
human-curated examples in 7 languages (tur, eng, yor, arb, zho, por, tel) → aya-human-annotated.
machine-translations of handpicked examples into 101 languages → dolly-machine-translated.… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/aya_evaluation_suite.Chat2Workflow-Evaluation
Chat2Workflow
Chat2Workflow is a benchmark designed for evaluating the ability of Large Language Models (LLMs) to generate executable visual workflows from natural language instructions.
Paper: Chat2Workflow: A Benchmark for Generating Executable Visual Workflows with Natural Language
Repository: zjunlp/Chat2Workflow
Overview
Executable visual workflows are widely used in industrial deployments for their reliability and controllability. Chat2Workflow addresses the… See the full description on the dataset page: https://huggingface.co/datasets/zjunlp/Chat2Workflow-Evaluation.openai-moderation-api-evaluation
Evaluation dataset for the paper "A Holistic Approach to Undesired Content Detection"
The evaluation dataset data/samples-1680.jsonl.gz is the test set used in this paper.
Each line contains information about one sample in a JSON object and each sample is labeled according to our taxonomy. The category label is a binary flag, but if it does not include in the JSON, it means we do not know the label.
Category
Label
Definition
sexual
S
Content meant to arouse sexual… See the full description on the dataset page: https://huggingface.co/datasets/mmathys/openai-moderation-api-evaluation.chess-position-evaluations
Dataset Card for the Lichess Evaluations dataset
Dataset Description
416,442,401 chess positions evaluated with Stockfish at various depths and node count. Produced by, and for, the Lichess analysis board, running various flavours of Stockfish within user browsers. This version of the dataset is a de-normalized version of the original dataset and contains 1,003,347,246 rows.
This dataset is updated monthly, and was last updated on October 7th, 2026.… See the full description on the dataset page: https://huggingface.co/datasets/Lichess/chess-position-evaluations.MERA
MERA v.1.2.0 (Multimodal Evaluation for Russian-language Architectures)
Summary
MERA (Multimodal Evaluation for Russian-language Architectures) is a new open independent benchmark for the evaluation of SOTA models for the Russian language.
The MERA benchmark unites industry and academic partners in one place to research the capabilities of fundamental models, draw attention to AI-related issues, foster collaboration within the Russian Federation and in the… See the full description on the dataset page: https://huggingface.co/datasets/MERA-evaluation/MERA.OpenVLHarness-Evaluation-Datasets
OpenVLHarness evaluation datasets
Processed evaluation splits used by
OpenVLHarness
(project page).
Each <split>.tsv holds the exact prompts (question) and annotations
(answer plus metadata) we evaluate on; image_path is relative to
images/<folder>/ inside images/<folder>.zip. GLIP.zip holds the ODinW-13
configs and COCO-format val/test annotations used for ODinW AP evaluation.
You normally don't need to download anything by hand: running
openvlharness-eval --data <split> ...… See the full description on the dataset page: https://huggingface.co/datasets/uclanlp/OpenVLHarness-Evaluation-Datasets.szl-frontier-evaluation-receipts
FRONTIER · EVALUATION ARCHIVE
Inspect the inputs, model responses and scores behind a small recorded evaluation.
Published material
Integrity evidence
Release boundary
Evaluation records
File hashes; unsigned
HOLD; no promotion
Open this run's summary · Review the exact source
Scores apply to the recorded cases. This archive grants no execution authority.
Checkable language-model evaluation records
We publish the inputs, measured responses, scoring results… See the full description on the dataset page: https://huggingface.co/datasets/SZLHOLDINGS/szl-frontier-evaluation-receipts.vlm_evaluation_v1.0
Datacard
This dataset is the evaluation VLM dataset used in VLABench. It is designed to evaluate the planning capabilities of Vision-Language Models (VLMs) in embodied scenarios.
Source
Project Page: https://vlabench.github.io/
Arxiv Paper: https://arxiv.org/abs/2412.18194
Code: https://github.com/OpenMOSS/VLABench
Uses
The dataset structure is as follows:
vlm_evaluation_v1.0/
├── CommenSence/
├── add_condiment_common_sense/
├──… See the full description on the dataset page: https://huggingface.co/datasets/VLABench/vlm_evaluation_v1.0.eka-medical-asr-evaluation-dataset
Eka Medical ASR Evaluation Dataset
Dataset Overview and Sourcing
The Eka Medical ASR Evaluation Dataset enables comprehensive evaluation of automatic speech recognition systems designed to transcribe medical speech into accurate text—a fundamental component of any medical scribe system. This dataset captures the unique challenges of processing medical terminology, particularly branded drugs, which is specific to the Indian context.
The dataset comprises over 3,900+… See the full description on the dataset page: https://huggingface.co/datasets/ekacare/eka-medical-asr-evaluation-dataset.evaluation-tables
[!CAUTION]
This dataset will not be updated. It corresponds to the last available public snapshot of the data, retrieved on July 28th, 2025.
chess-evaluations
Chess Evaluations Dataset
This dataset contains chess positions represented in FEN (Forsyth-Edwards Notation) along with their evaluations and next moves for tactical evals. The dataset is divided into three configurations:
tactics: Includes chess positions, their evaluations, and the best move in the position.
randoms: Contains random chess positions and their evaluations.
chess_data: General chess positions with evaluations.
This is an in progress dataset which contains millions… See the full description on the dataset page: https://huggingface.co/datasets/ssingh22/chess-evaluations.DL3DV-Evaluation
DL3DV Testing Split Download Instructions
This repo contains all 55 scenes for evaluation. Note: it is an independent dataset, and none of its scenes overlap with those in DL3DV-10K. Have a galance on the preview page: https://dl3dv-10k.github.io/DL3DV-Testing-Split-Preview/.
Download
As the whole benchmark dataset is ~500G, a python script to download and untar files.
Environment Setup
The download script relies on huggingface hub, tqdm. You can download by… See the full description on the dataset page: https://huggingface.co/datasets/DL3DV/DL3DV-Evaluation.CoG-Evaluation-Data
CoG Evaluation Data
The seven evaluation sets used in CoG (EMNLP 2026), comprising 5,174 questions for knowledge-intensive QA across KG-based and text-based benchmarks. This collection supports evaluation of multi-hop retrieval and reasoning with graph, text, and hybrid RAG methods.
Paper · Code
Datasets
Dataset / HF config
Knowledge source
Questions
Raw file
KGQAGen / kgqagen
KG
1,079
KGQAGen-10k.json
CWQ / cwq
KG
1,024
cwq.json
QALD10-en /… See the full description on the dataset page: https://huggingface.co/datasets/zhougengxian/CoG-Evaluation-Data.patient-evaluations
Patient Evaluations Dataset
This dataset contains clinician evaluations of AI-generated patient summaries from MIMIC-III data.
Dataset Description
The dataset includes expert clinician assessments of AI-generated patient summaries, with detailed ratings across multiple dimensions including clinical accuracy, completeness, relevance, and identification of hallucinations or critical omissions.
Dataset Structure
The dataset contains a CSV file… See the full description on the dataset page: https://huggingface.co/datasets/JesseLiu/patient-evaluations.staining-robustness-evaluation
A Protocol for Evaluating Robustness to H&E Staining Variation in Computational Pathology Models
This repository provides the stain references, pretrained models, and experimental results required to:
Define custom staining references using our PLISM reference library
Reproduce our published controlled staining robustness experiments
👉 Code repository: https://github.com/lely475/staining-robustness-evaluation/tree/main
👉 Associated publication: Paper
Overview: How… See the full description on the dataset page: https://huggingface.co/datasets/CTPLab-DBE-UniBas/staining-robustness-evaluation.Medical_Multimodal_Evaluation_Data
Evaluation Guide
This dataset is used to evaluate medical multimodal LLMs, as used in HuatuoGPT-Vision. It includes benchmarks such as VQA-RAD, SLAKE, PathVQA, PMC-VQA, OmniMedVQA, and MMMU-Medical-Tracks.
To get started:
Download the dataset and extract the images.zip file.
Find evaluation code on our GitHub: HuatuoGPT-Vision.
This open-source release aims to simplify the evaluation of medical multimodal capabilities in large models. Please cite the relevant benchmark… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/Medical_Multimodal_Evaluation_Data.GenoJEPA-Evaluation
GenoJEPA-Evaluation
This dataset provides the downstream evaluation benchmark used for GenoJEPA, a genomic representation learning framework based on joint-embedding predictive architecture.
The evaluation resource is designed to test the transferability of genomic sequence representations across diverse classification tasks. It supports both full fine-tuning and frozen-embedding probing protocols.
Contents
This resource contains downstream genomic sequence… See the full description on the dataset page: https://huggingface.co/datasets/ChengsenWang/GenoJEPA-Evaluation.JobBERT-evaluation-dataset
JobBERT evaluation dataset 💾
This is the official repository containing the evaluation data that was used for the JobBERT paper. This dataset is a list of vacancy titles, each tagged with an ESCO (v1.0.5) occupation.
The full dataset is split into two files in a stratified way by class distribution. This data was automatically collected from a large governmental job board.
Access the JobBERT paper here: https://arxiv.org/abs/2109.09605
BibTeX Citation
If you use this… See the full description on the dataset page: https://huggingface.co/datasets/TechWolf/JobBERT-evaluation-dataset.BeaverTails-Evaluation
Dataset Card for BeaverTails-Evaluation
BeaverTails is an AI safety-focused collection comprising a series of datasets.
This repository contains test prompts specifically designed for evaluating language model safety.
It is important to note that although each prompt can be connected to multiple categories, only one category is labeled for each prompt.
The 14 harm categories are defined as follows:
Animal Abuse: This involves any form of cruelty or harm inflicted on animals… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/BeaverTails-Evaluation.webvoyager_evaluation_dataNexusRaven_API_evaluation
NexusRaven API Evaluation dataset
Please see blog post or NexusRaven Github repo for more information.
License
The evaluation data in this repository consists primarily of our own curated evaluation data that only uses open source commercializable models. However, we include general domain data from the ToolLLM and ToolAlpaca papers. Since the data in the ToolLLM and ToolAlpaca works use OpenAI's GPT models for the generated content, the data is not commercially… See the full description on the dataset page: https://huggingface.co/datasets/Nexusflow/NexusRaven_API_evaluation.qwen-math-evaluation-resultsruHumanEval
ruHumanEval
Task Description
Russian HumanEval (ruHumanEval) is the Russian analogue of the original HumanEval dataset, created to evaluate the ability of language models to generate code in the Python programming language to solve simple problems.
The dataset contains 164 tasks and is aimed at measuring the functional correctness of code generation based on information from the function's documentation lines — a text description of the function's operation and several… See the full description on the dataset page: https://huggingface.co/datasets/MERA-evaluation/ruHumanEval.LRM-Safety-evaluation-parsedwmt-mqm-human-evaluation
Dataset Summary
This dataset contains all MQM human annotations from previous WMT Metrics shared tasks and the MQM annotations from Experts, Errors, and Context.
The data is organised into 8 columns:
lp: language pair
src: input text
mt: translation
ref: reference translation
score: MQM score
system: MT Engine that produced the translation
annotators: number of annotators
domain: domain of the input text (e.g. news)
year: collection year
You can also find the original data here.… See the full description on the dataset page: https://huggingface.co/datasets/RicardoRei/wmt-mqm-human-evaluation.RAG-Evaluation-Dataset-KO
Allganize RAG Leaderboard
Allganize RAG 리더보드는 5개 도메인(금융, 공공, 의료, 법률, 커머스)에 대해서 한국어 RAG의 성능을 평가합니다.일반적인 RAG는 간단한 질문에 대해서는 답변을 잘 하지만, 문서의 테이블과 이미지에 대한 질문은 답변을 잘 못합니다.
RAG 도입을 원하는 수많은 기업들은 자사에 맞는 도메인, 문서 타입, 질문 형태를 반영한 한국어 RAG 성능표를 원하고 있습니다.평가를 위해서는 공개된 문서와 질문, 답변 같은 데이터 셋이 필요하지만, 자체 구축은 시간과 비용이 많이 드는 일입니다.이제 올거나이즈는 RAG 평가 데이터를 모두 공개합니다.
RAG는 Parser, Retrieval, Generation 크게 3가지 파트로 구성되어 있습니다.현재, 공개되어 있는 RAG 리더보드 중, 3가지 파트를 전체적으로 평가하는 한국어로 구성된 리더보드는 없습니다.
Allganize RAG 리더보드에서는 문서를… See the full description on the dataset page: https://huggingface.co/datasets/allganize/RAG-Evaluation-Dataset-KO.wmt-da-human-evaluation-long-context
Dataset Summary
Long-context / document-level dataset for Quality Estimation of Machine Translation.
It is an augmented variant of the sentence-level WMT DA Human Evaluation dataset.
In addition to individual sentences, it contains augmentations of 2, 4, 8, 16, and 32 sentences, among each language pair lp and domain.
The raw column represents a weighted average of scores of augmented sentences using character lengths of src and mt as weights.
The code used to apply the augmentation… See the full description on the dataset page: https://huggingface.co/datasets/ymoslem/wmt-da-human-evaluation-long-context.llm-evaluation-self-audit
LLM Evaluation Self-Audit
Data for How Reproducible Are Evaluation Conclusions? A Self-Audit of LLM-Inferred Prompt Structure, Dipankar Sarkar, Skelf Research.
Code and paper source: github.com/sarkar-dipankar/llm-evaluation-self-audit (this package is built from commit 687ee8a).
The finding
We asked eight open model variants to infer the structure of a prompt, then asked again with the identical call. Caching was off.
They often gave a different answer.
Mean… See the full description on the dataset page: https://huggingface.co/datasets/dipankarsarkar/llm-evaluation-self-audit.denoising-impact-evaluation-dataset
Denoising Impact Evaluation Dataset
Dataset Description
The ekacare/denoising-impact-evaluation-dataset is a comprehensive benchmark dataset designed to evaluate the effects of speech enhancement on automatic speech recognition (ASR) systems in medical speech contexts. It includes paired noisy and denoised audio subsets under controlled acoustic conditions to support systematic analysis of denoising performance.
Source Data
Base Dataset:… See the full description on the dataset page: https://huggingface.co/datasets/ekacare/denoising-impact-evaluation-dataset.
