datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
aya_evaluation_suite
Dataset Summary
Aya Evaluation Suite contains a total of 26,750 open-ended conversation-style prompts to evaluate multilingual open-ended generation quality.To strike a balance between language coverage and the quality that comes with human curation, we create an evaluation suite that includes:
human-curated examples in 7 languages (tur, eng, yor, arb, zho, por, tel) → aya-human-annotated.
machine-translations of handpicked examples into 101 languages → dolly-machine-translated.… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/aya_evaluation_suite.openai-moderation-api-evaluation
Evaluation dataset for the paper "A Holistic Approach to Undesired Content Detection"
The evaluation dataset data/samples-1680.jsonl.gz is the test set used in this paper.
Each line contains information about one sample in a JSON object and each sample is labeled according to our taxonomy. The category label is a binary flag, but if it does not include in the JSON, it means we do not know the label.
Category
Label
Definition
sexual
S
Content meant to arouse sexual… See the full description on the dataset page: https://huggingface.co/datasets/mmathys/openai-moderation-api-evaluation.chess-position-evaluations
Dataset Card for the Lichess Evaluations dataset
Dataset Description
416,442,401 chess positions evaluated with Stockfish at various depths and node count. Produced by, and for, the Lichess analysis board, running various flavours of Stockfish within user browsers. This version of the dataset is a de-normalized version of the original dataset and contains 1,003,347,246 rows.
This dataset is updated monthly, and was last updated on October 7th, 2026.… See the full description on the dataset page: https://huggingface.co/datasets/Lichess/chess-position-evaluations.OpenVLHarness-Evaluation-Datasets
OpenVLHarness evaluation datasets
Processed evaluation splits used by
OpenVLHarness
(project page).
Each <split>.tsv holds the exact prompts (question) and annotations
(answer plus metadata) we evaluate on; image_path is relative to
images/<folder>/ inside images/<folder>.zip. GLIP.zip holds the ODinW-13
configs and COCO-format val/test annotations used for ODinW AP evaluation.
You normally don't need to download anything by hand: running
openvlharness-eval --data <split> ...… See the full description on the dataset page: https://huggingface.co/datasets/uclanlp/OpenVLHarness-Evaluation-Datasets.chess-evaluations
Chess Evaluations Dataset
This dataset contains chess positions represented in FEN (Forsyth-Edwards Notation) along with their evaluations and next moves for tactical evals. The dataset is divided into three configurations:
tactics: Includes chess positions, their evaluations, and the best move in the position.
randoms: Contains random chess positions and their evaluations.
chess_data: General chess positions with evaluations.
This is an in progress dataset which contains millions… See the full description on the dataset page: https://huggingface.co/datasets/ssingh22/chess-evaluations.staining-robustness-evaluation
A Protocol for Evaluating Robustness to H&E Staining Variation in Computational Pathology Models
This repository provides the stain references, pretrained models, and experimental results required to:
Define custom staining references using our PLISM reference library
Reproduce our published controlled staining robustness experiments
👉 Code repository: https://github.com/lely475/staining-robustness-evaluation/tree/main
👉 Associated publication: Paper
Overview: How… See the full description on the dataset page: https://huggingface.co/datasets/CTPLab-DBE-UniBas/staining-robustness-evaluation.wmt-mqm-human-evaluation
Dataset Summary
This dataset contains all MQM human annotations from previous WMT Metrics shared tasks and the MQM annotations from Experts, Errors, and Context.
The data is organised into 8 columns:
lp: language pair
src: input text
mt: translation
ref: reference translation
score: MQM score
system: MT Engine that produced the translation
annotators: number of annotators
domain: domain of the input text (e.g. news)
year: collection year
You can also find the original data here.… See the full description on the dataset page: https://huggingface.co/datasets/RicardoRei/wmt-mqm-human-evaluation.wmt-da-human-evaluation-long-context
Dataset Summary
Long-context / document-level dataset for Quality Estimation of Machine Translation.
It is an augmented variant of the sentence-level WMT DA Human Evaluation dataset.
In addition to individual sentences, it contains augmentations of 2, 4, 8, 16, and 32 sentences, among each language pair lp and domain.
The raw column represents a weighted average of scores of augmented sentences using character lengths of src and mt as weights.
The code used to apply the augmentation… See the full description on the dataset page: https://huggingface.co/datasets/ymoslem/wmt-da-human-evaluation-long-context.llm-evaluation-self-audit
LLM Evaluation Self-Audit
Data for How Reproducible Are Evaluation Conclusions? A Self-Audit of LLM-Inferred Prompt Structure, Dipankar Sarkar, Skelf Research.
Code and paper source: github.com/sarkar-dipankar/llm-evaluation-self-audit (this package is built from commit 687ee8a).
The finding
We asked eight open model variants to infer the structure of a prompt, then asked again with the identical call. Caching was off.
They often gave a different answer.
Mean… See the full description on the dataset page: https://huggingface.co/datasets/dipankarsarkar/llm-evaluation-self-audit.HIP-training-and-evaluation-data
HIP Training and Evaluation Data
This dataset contains the text data released with Base Models Look Human To AI Detectors for reproducing the Humanization by Iterative Paraphrasing (HIP) training setup and the prefix-based continuation evaluation.
Configs
training
data/train.parquet contains 10,581 supervised HIP training pairs with seven columns:
dataset: upstream dataset family, either raid or mage.
source: selected source domain or subcorpus.
text: original… See the full description on the dataset page: https://huggingface.co/datasets/YixuanEvenXu/HIP-training-and-evaluation-data.UGround-Offline-Evaluationkg-gen-MINE-evaluation-datasethumatheque-vlm-evaluation-results
Humathèque — VLM metadata-extraction evaluation
Structured metadata extraction from French thesis/dissertation title pages, scored against a human-validated visible-only ground truth (catalogue role/jury data not on the page was removed).
Leaderboard
model
overall
surface
stable
role
valid_json
gemma4-26b-a4b
0.947
0.8
0.942
0.961
1
mistral-small-32-24b
0.939
0.728
0.925
0.98
1
bonsai2-27b
0.939
0.687
0.934
0.95
1
lift
0.931
0.51
0.92
0.96
1… See the full description on the dataset page: https://huggingface.co/datasets/Geraldine/humatheque-vlm-evaluation-results.CoderForge-Preview-32B-SWE-Bench-Verified-Evaluation-trajectoriescircuitlens-gemma-2-2btranscoder-descriptions-and-evaluations
CircuitLens & WeightLens: Transcoder Descriptions and Evaluations
This dataset contains automatically generated descriptions and evaluation metrics for Gemma-2-2B transcoders, produced using CircuitLens and WeightLens methods.
Methods
CircuitLens: https://github.com/egolimblevskaia/CircuitLens
WeightLens: https://github.com/egolimblevskaia/WeightLens
Dataset Structure
The dataset is organized by layers (0, 4, 7, 10, 12, 15, 18, 21, 23, 25), with each layer… See the full description on the dataset page: https://huggingface.co/datasets/egolimblevskaia/circuitlens-gemma-2-2btranscoder-descriptions-and-evaluations.wmt-da-human-evaluation
Dataset Summary
This dataset contains all DA human annotations from previous WMT News Translation shared tasks.
The data is organised into 8 columns:
lp: language pair
src: input text
mt: translation
ref: reference translation
score: z score
raw: direct assessment
annotators: number of annotators
domain: domain of the input text (e.g. news)
year: collection year
You can also find the original data for each year in the results section https://www.statmt.org/wmt{YEAR}/results.html… See the full description on the dataset page: https://huggingface.co/datasets/RicardoRei/wmt-da-human-evaluation.SWE-MERA
SWE-MERA
Continuously updated SWE-MERA dataset
SWE-MERA splits:
dev: for testing (10 samples)
lite: presented at the leaderboard here (750 samples)
full: continuously updated to collect more data (2738 samples)
Load dataset
from datasets import load_dataset
ds = load_dataset("MERA-evaluation/SWE-MERA", split='dev')
Evaluation
Description
The main tool to validate tasks is repotest (available at PyPI or GitHub)
data.jsonl -… See the full description on the dataset page: https://huggingface.co/datasets/MERA-evaluation/SWE-MERA.RAG_Evaluation_Datasetsdf_evaluation_traits
Models That Know How Evaluations Are Designed Score Safer
This repository contains the synthetic documents used in the paper Models That Know How Evaluations Are Designed Score Safer.
Project Page | GitHub Repository
Dataset Description
These synthetic documents were used to fine-tune models to investigate evaluation meta-knowledge — parametric knowledge about the structural traits that characterize AI safety evaluations.
Documents were generated using the… See the full description on the dataset page: https://huggingface.co/datasets/compass-group-tue/sdf_evaluation_traits.toporisk-evaluation-artifacts
TopoRisk Evaluation Artifacts
English | 简体中文
TopoRisk studies how a limited verification budget should be allocated across
an agentic workflow graph. Instead of ranking steps only by their local failure
probability, the scheduler estimates how an error can reach terminal outputs
and recomputes marginal value after every selected audit.
This repository is an artifact-first research preview. It contains code,
aggregate measurements, and figures. The manuscript and its LaTeX… See the full description on the dataset page: https://huggingface.co/datasets/JosephAA/toporisk-evaluation-artifacts.evaluation-results-task1Model-comparison table for this task: one row per evaluated model, written by
push_results_table in src/eval/utilities.py. Per-sample predictions are in
per_sample/.
Evaluation-Multilingual-VC
Evaluation-Multilingual-VC
We use dataset https://huggingface.co/datasets/sarulab-speech/commonvoice22_sidon,
Filter languages that support by Whisper Large V3 to evaluate WER automatically,
Only take test set, sort by up votes.
Because VC required to source text, source audio, target text, we make sure the target text is not same as source text, target text we take from other rows.
Only build first 500 rows for each language
Github issue at… See the full description on the dataset page: https://huggingface.co/datasets/Scicom-intl/Evaluation-Multilingual-VC.Byte-Authority-Evaluation
Byte Authority · Evaluation Records
📄 Paper (arXiv:2609.35932)
•
💻 Code
•
🤗 Collection
This dataset contains the compact per-case records used by the Byte Authority reproduction package for Same Bytes, Different Authority: Reserved-Token Representations in Chat-Template Prompt Injection. The records preserve condition labels, tool-call outcomes, prompt-length metadata, invariant checks, and prompt-ID metadata. Generated model text and runtime… See the full description on the dataset page: https://huggingface.co/datasets/YanZhanPKU/Byte-Authority-Evaluation.repro-efficient-inference-for-noisy-llm-as-a-judge-evaluation-traces
Agent traces
Agent sessions published from a Trackio Logbook.
RAG-Evaluation-Dataset-KO
Dataset Card for Reconstructed RAG Evaluation Dataset (KO)
Dataset Summary
본 데이터셋은 allganize/RAG-Evaluation-Dataset-KO를 기반으로 PDF 파일을 포함하도록 재구성한 한국어 평가 데이터셋입니다. 원본 데이터셋에서는 PDF 파일의 경로만 제공되어 수동으로 파일을 다운로드해야 하는 불편함이 있었고, 일부 PDF 파일의 경로가 유효하지 않은 문제를 보완하기 위해 PDF 파일을 포함한 데이터셋을 재구성하였습니다.
Supported Tasks and Leaderboards
RAG Evaluation: 본 데이터는 한국어 RAG 파이프라인에 대한 E2E Evaluation이 가능합니다.
Languages
The dataset is in Korean (ko).
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/datalama/RAG-Evaluation-Dataset-KO.evaluation-results-task2Model-comparison table for this task: one row per evaluated model, written by
push_results_table in src/eval/utilities.py. Per-sample predictions are in
per_sample/.
rcga-evaluation-data
RCGA / LoopSFT evaluation input snapshots
Exact benchmark inputs used in our Qwen3-1.7B, Qwen3-4B and Qwen3-30B-A3B experiments. These are evaluation questions, prompts, labels and tests, not model-generated answers or reported scores. Already frozen subsets are stored directly, including original row order and prompt formatting. No new random split was made during backup.
Collections
data/evalbench/: all 31 historical prompt variants. The active standard14… See the full description on the dataset page: https://huggingface.co/datasets/LaurelWings/rcga-evaluation-data.humaine-evaluation-dataset
HUMAINE: Human-AI Interaction Evaluation Dataset
Dataset Description
Dataset Summary
The HUMAINE dataset contains human evaluations of AI model interactions across diverse demographic groups and conversation contexts. This dataset powers the HUMAINE Leaderboard, providing insights into how different AI models perform across various user populations and use cases.
The dataset consists of two main components:
Feedback Comparisons: Pairwise model comparisons… See the full description on the dataset page: https://huggingface.co/datasets/ProlificAI/humaine-evaluation-dataset.Evaluation_of_OpenSource_Models_for_PDF_Injection_Recognition
Injected PDFs - Model Evaluation
This repository holds the model evaluation stage of a project on detecting harmless-but-real
attack payloads injected into PDF files, together with the artefacts it produced for the
application.
Nothing is trained here. Seven off-the-shelf models are measured against the same 1,100 PDFs,
and the two winners are exported for the app to load.
Question
Candidates
Winner
Part A
Which files look like this one?
3 embedding models x 2 inputs… See the full description on the dataset page: https://huggingface.co/datasets/Cyber-security-final-project/Evaluation_of_OpenSource_Models_for_PDF_Injection_Recognition.evaluation-results-task3Model-comparison table for this task: one row per evaluated model, written by
push_results_table in src/eval/utilities.py. Per-sample predictions are in
per_sample/.
