datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
evaluation-results@misc{muennighoff2022crosslingual,
title={Crosslingual Generalization through Multitask Finetuning},
author={Niklas Muennighoff and Thomas Wang and Lintang Sutawika and Adam Roberts and Stella Biderman and Teven Le Scao and M Saiful Bari and Sheng Shen and Zheng-Xin Yong and Hailey Schoelkopf and Xiangru Tang and Dragomir Radev and Alham Fikri Aji and Khalid Almubarak and Samuel Albanie and Zaid Alyafeai and Albert Webson and Edward Raff and Colin Raffel},
year={2022},
eprint={2211.01786},
archivePrefix={arXiv},
primaryClass={cs.CL}
}GDP-Val-Evaluation-Submission
GDPval Submission Dataset
This dataset contains model outputs for GDP-Val evaluation.
Dataset Structure
data/: Contains the main dataset in Parquet format
train-00000-of-00001.parquet: Submission data with model outputs
deliverable_files/: Contains generated files for tasks that produce file deliverables
Organized by task_id
dataset_info.json: Metadata about the dataset
Columns
task_id: Unique identifier for each task
sector: Economic sector for the task… See the full description on the dataset page: https://huggingface.co/datasets/xiachongfeng/GDP-Val-Evaluation-Submission.vam-cross-evaluation-artifactsgrobid-evaluation
GROBID End-to-End Evaluation Dataset
Reference corpora used for GROBID end-to-end
benchmarking of scientific-article structuring.
Documentation: https://grobid.readthedocs.io/en/latest/End-to-end-evaluation/
Latest benchmarking scores: https://grobid.readthedocs.io/en/latest/Benchmarking/
Official archive (Zenodo): https://zenodo.org/record/7708580
Dataset summary
These are the datasets used for GROBID end-to-end benchmarking, covering:
metadata extraction… See the full description on the dataset page: https://huggingface.co/datasets/sciencialab/grobid-evaluation.aya_evaluation_suite
Dataset Summary
Aya Evaluation Suite contains a total of 26,750 open-ended conversation-style prompts to evaluate multilingual open-ended generation quality.To strike a balance between language coverage and the quality that comes with human curation, we create an evaluation suite that includes:
human-curated examples in 7 languages (tur, eng, yor, arb, zho, por, tel) → aya-human-annotated.
machine-translations of handpicked examples into 101 languages → dolly-machine-translated.… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/aya_evaluation_suite.evaluation_logs
Evaluation logs from "Auditing Games for Sandbagging"
This dataset provides evaluation transcripts produced for the paper "Auditing Games for Sandbagging". Transcripts are provided in Inspect .eval format, see https://github.com/AI-Safety-Institute/sabotage_games for a guide to viewing them.
Dataset Details
evaluation_transcripts/handover_evals contains the transcripts provided by the red team to the blue team at the beginning of the main round of the game, showing… See the full description on the dataset page: https://huggingface.co/datasets/sandbagging-games/evaluation_logs.evaluationChat2Workflow-Evaluation
Chat2Workflow
Chat2Workflow is a benchmark designed for evaluating the ability of Large Language Models (LLMs) to generate executable visual workflows from natural language instructions.
Paper: Chat2Workflow: A Benchmark for Generating Executable Visual Workflows with Natural Language
Repository: zjunlp/Chat2Workflow
Overview
Executable visual workflows are widely used in industrial deployments for their reliability and controllability. Chat2Workflow addresses the… See the full description on the dataset page: https://huggingface.co/datasets/zjunlp/Chat2Workflow-Evaluation.Egocentric_10K_Evaluation
Dataset Card for Egocentric_10K_Evaluation
This is a FiftyOne dataset with 30000 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("Voxel51/Egocentric_10K_Evaluation")
# Launch the App
session = fo.launch_app(dataset)
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/Egocentric_10K_Evaluation.openai-moderation-api-evaluation
Evaluation dataset for the paper "A Holistic Approach to Undesired Content Detection"
The evaluation dataset data/samples-1680.jsonl.gz is the test set used in this paper.
Each line contains information about one sample in a JSON object and each sample is labeled according to our taxonomy. The category label is a binary flag, but if it does not include in the JSON, it means we do not know the label.
Category
Label
Definition
sexual
S
Content meant to arouse sexual… See the full description on the dataset page: https://huggingface.co/datasets/mmathys/openai-moderation-api-evaluation.chess-position-evaluations
Dataset Card for the Lichess Evaluations dataset
Dataset Description
409,710,113 chess positions evaluated with Stockfish at various depths and node count. Produced by, and for, the Lichess analysis board, running various flavours of Stockfish within user browsers. This version of the dataset is a de-normalized version of the original dataset and contains 988,851,570 rows.
This dataset is updated monthly, and was last updated on September 27th, 2026.… See the full description on the dataset page: https://huggingface.co/datasets/Lichess/chess-position-evaluations.MERA
MERA (Multimodal Evaluation for Russian-language Architectures)
Summary
MERA (Multimodal Evaluation for Russian-language Architectures) is a new open independent benchmark for the evaluation of SOTA models for the Russian language.
The MERA benchmark unites industry and academic partners in one place to research the capabilities of fundamental models, draw attention to AI-related issues, foster collaboration within the Russian Federation and in the international arena… See the full description on the dataset page: https://huggingface.co/datasets/MERA-evaluation/MERA.staining-robustness-evaluation
A Protocol for Evaluating Robustness to H&E Staining Variation in Computational Pathology Models
This repository provides the stain references, pretrained models, and experimental results required to:
Define custom staining references using our PLISM reference library
Reproduce our published controlled staining robustness experiments
👉 Code repository: https://github.com/lely475/staining-robustness-evaluation/tree/main
👉 Associated publication: Paper
Overview: How… See the full description on the dataset page: https://huggingface.co/datasets/CTPLab-DBE-UniBas/staining-robustness-evaluation.VeriLoop-Coder-E1-Evaluation-Evidence
VeriLoop Coder-E1 Evaluation Evidence
This repository contains the public evaluation-evidence packages
referenced by the official VeriLoop Coder-E1 benchmark result files.
Model repository:
tsinghua-sigs-robot-lab/veriloop-coder-e1
Evidence packages
Benchmark
Evidence directory
DeepSWE
veriloop-coder-e1-deepswe-evaluation-evidence-v1.0.0
SWE-bench Pro
veriloop-coder-e1-swe-bench-pro-evaluation-evidence-v1.0.0
SWE-bench Verified… See the full description on the dataset page: https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-Coder-E1-Evaluation-Evidence.user-evaluationsGPT-4o-evaluation-biases
A database to support the evaluation of gender biases in GPT-4o output
The database and its construction process are described in the paper "A database to support the evaluation of gender biases in GPT-4o output" by Mehner et al., presented at the 1st ISCA/ITG Workshop on Diversity in Large Speech and Language Models (Berlin, Februar 20, 2025).
Introduction
This is a database of prompts and answers generated with GPT-4o-mini and GPT-4o in a pretest and a main test… See the full description on the dataset page: https://huggingface.co/datasets/mtec-TUB/GPT-4o-evaluation-biases.hallucination-evaluation-results
Hallucination evaluation artifacts
Canonical COCO and AMBER detection packages are under detection/<dataset>/models/<model>/probes/<probe>/. See detection/index.json for active and pending packages. Each published package has a manifest with exact checkpoint, cache and test-result paths. Seeds are 0, 42 and 1337. TruthPrInt retains its native validation TPR-at-1%-FPR checkpoint; other detection probes use best pooled validation AUROC. Detection thresholds are validation-best F1.… See the full description on the dataset page: https://huggingface.co/datasets/ToiTenBao/hallucination-evaluation-results.ShapeR-Evaluation
ShapeR Evaluation Dataset
We introduce a new dataset of in-the-wild sequences with paired posed multi-view images, SLAM
point clouds, and individually complete 3D shape annotations for 178 objects across 7 diverse scenes. In contrast to existing real-world 3D reconstruction datasets which are either captured in controlled setups or have merged object and background geometries or incomplete shapes, this dataset is designed to capture real-world challenges like occlusions, clutter… See the full description on the dataset page: https://huggingface.co/datasets/facebook/ShapeR-Evaluation.VeriLoop-E2-Evaluation-Evidence
VeriLoop E2 Evaluation Evidence
Public evaluation evidence for VeriLoop E2 across nine code, agentic, mathematical, and scientific reasoning benchmarks.
This dataset repository is the canonical public evidence layer for the reported benchmark results of VeriLoop E2, a post-trained model based on Qwen 3.8-27B. It is designed to separate headline benchmark reporting from the underlying auditable artifacts required to inspect, reproduce, and verify those results.
The repository… See the full description on the dataset page: https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-E2-Evaluation-Evidence.haotian_data-GPS-lm-evaluation-harness
Language Model Evaluation Harness
Latest News 📣
[2025/03] Added support for steering HF models!
[2025/02] Added SGLang support!
[2024/09] We are prototyping allowing users of LM Evaluation Harness to create and evaluate on text+image multimodal input, text output tasks, and have just added the hf-multimodal and vllm-vlm model types and mmmu task as a prototype feature. We welcome users to try out this in-progress feature and stress-test it for themselves, and suggest… See the full description on the dataset page: https://huggingface.co/datasets/happynew111/haotian_data-GPS-lm-evaluation-harness.music-off-policy-evaluation-benchmark
Music Off-Policy Evaluation Dataset
Music Off-Policy Evaluation Dataset is a dataset designed for Off-Policy Evaluation (OPE) research. It contains logged interactions from the home page of Amazon Music.
Use cases:
Benchmarking OPE estimators
Evaluating counterfactual ranking policies offline
License
Music Off-Policy Evaluation Benchmark © 2026 by Amazon is licensed under Creative Commons Attribution-NonCommercial 4.0 International.… See the full description on the dataset page: https://huggingface.co/datasets/amazon/music-off-policy-evaluation-benchmark.vlm_evaluation_v1.0
Datacard
This dataset is the evaluation VLM dataset used in VLABench. It is designed to evaluate the planning capabilities of Vision-Language Models (VLMs) in embodied scenarios.
Source
Project Page: https://vlabench.github.io/
Arxiv Paper: https://arxiv.org/abs/2412.18194
Code: https://github.com/OpenMOSS/VLABench
Uses
The dataset structure is as follows:
vlm_evaluation_v1.0/
├── CommenSence/
├── add_condiment_common_sense/
├──… See the full description on the dataset page: https://huggingface.co/datasets/VLABench/vlm_evaluation_v1.0.nyush-galaxea-a1-lingbot-va-real-world-evaluations
LingBot-VA on Galaxea A1 — Real-World Evaluations
Fruit-placement rollouts and open-loop diagnostics of object grounding,
layout generalization, and predicted robot motion.
Fruit step-1000: lemon-to-plate rollout in the Official layout.
Evidence
Scale
Real closed-loop rollouts
61 archived; 60 scored
Matched base-model controls
9 predictions
Post-trained diagnostics
48 full-horizon predictions; 1,211 rolling futures
Controlled OOD studies
558 predictions… See the full description on the dataset page: https://huggingface.co/datasets/pengyue-polaron/nyush-galaxea-a1-lingbot-va-real-world-evaluations.szl-frontier-evaluation-receipts
FRONTIER · EVALUATION ARCHIVE
Inspect the inputs, model responses and scores behind a small recorded evaluation.
Published material
Integrity evidence
Release boundary
Evaluation records
File hashes; unsigned
HOLD; no promotion
Open this run's summary · Review the exact source
Scores apply to the recorded cases. This archive grants no execution authority.
Checkable language-model evaluation records
We publish the inputs, measured responses, scoring results… See the full description on the dataset page: https://huggingface.co/datasets/SZLHOLDINGS/szl-frontier-evaluation-receipts.eka-medical-asr-evaluation-dataset
Eka Medical ASR Evaluation Dataset
Dataset Overview and Sourcing
The Eka Medical ASR Evaluation Dataset enables comprehensive evaluation of automatic speech recognition systems designed to transcribe medical speech into accurate text—a fundamental component of any medical scribe system. This dataset captures the unique challenges of processing medical terminology, particularly branded drugs, which is specific to the Indian context.
The dataset comprises over 3,900+… See the full description on the dataset page: https://huggingface.co/datasets/ekacare/eka-medical-asr-evaluation-dataset.chess-evaluations
Chess Evaluations Dataset
This dataset contains chess positions represented in FEN (Forsyth-Edwards Notation) along with their evaluations and next moves for tactical evals. The dataset is divided into three configurations:
tactics: Includes chess positions, their evaluations, and the best move in the position.
randoms: Contains random chess positions and their evaluations.
chess_data: General chess positions with evaluations.
This is an in progress dataset which contains millions… See the full description on the dataset page: https://huggingface.co/datasets/ssingh22/chess-evaluations.evaluation-tables
[!CAUTION]
This dataset will not be updated. It corresponds to the last available public snapshot of the data, retrieved on July 28th, 2025.
evaluation_dataset
Data manifest and release plan
Smoke-Eval/ contains the held-out evaluation sequences copied from the
Windows evaluation workspace. All released inference and metric paths resolve
to this directory. The source configurations also refer to two additional
datasets that are not required for the selected Smoke-Eval inference check.
Dataset
Role
Expected release contents
Current status
Rice / Ti-Zed-Dji
GRT and refinement training/validation; RGB, radar, and depth inputs… See the full description on the dataset page: https://huggingface.co/datasets/mypersonalsharingspot11/evaluation_dataset.AndroidFlux_Evaluation_Outputtoken_evaluation
