datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
OpenVLHarness-Evaluation-Datasets
OpenVLHarness evaluation datasets
Processed evaluation splits used by
OpenVLHarness
(project page).
Each <split>.tsv holds the exact prompts (question) and annotations
(answer plus metadata) we evaluate on; image_path is relative to
images/<folder>/ inside images/<folder>.zip. GLIP.zip holds the ODinW-13
configs and COCO-format val/test annotations used for ODinW AP evaluation.
You normally don't need to download anything by hand: running
openvlharness-eval --data <split> ...… See the full description on the dataset page: https://huggingface.co/datasets/uclanlp/OpenVLHarness-Evaluation-Datasets.staining-robustness-evaluation
A Protocol for Evaluating Robustness to H&E Staining Variation in Computational Pathology Models
This repository provides the stain references, pretrained models, and experimental results required to:
Define custom staining references using our PLISM reference library
Reproduce our published controlled staining robustness experiments
👉 Code repository: https://github.com/lely475/staining-robustness-evaluation/tree/main
👉 Associated publication: Paper
Overview: How… See the full description on the dataset page: https://huggingface.co/datasets/CTPLab-DBE-UniBas/staining-robustness-evaluation.wmt-mqm-human-evaluation
Dataset Summary
This dataset contains all MQM human annotations from previous WMT Metrics shared tasks and the MQM annotations from Experts, Errors, and Context.
The data is organised into 8 columns:
lp: language pair
src: input text
mt: translation
ref: reference translation
score: MQM score
system: MT Engine that produced the translation
annotators: number of annotators
domain: domain of the input text (e.g. news)
year: collection year
You can also find the original data here.… See the full description on the dataset page: https://huggingface.co/datasets/RicardoRei/wmt-mqm-human-evaluation.circuitlens-gemma-2-2btranscoder-descriptions-and-evaluations
CircuitLens & WeightLens: Transcoder Descriptions and Evaluations
This dataset contains automatically generated descriptions and evaluation metrics for Gemma-2-2B transcoders, produced using CircuitLens and WeightLens methods.
Methods
CircuitLens: https://github.com/egolimblevskaia/CircuitLens
WeightLens: https://github.com/egolimblevskaia/WeightLens
Dataset Structure
The dataset is organized by layers (0, 4, 7, 10, 12, 15, 18, 21, 23, 25), with each layer… See the full description on the dataset page: https://huggingface.co/datasets/egolimblevskaia/circuitlens-gemma-2-2btranscoder-descriptions-and-evaluations.wmt-da-human-evaluation
Dataset Summary
This dataset contains all DA human annotations from previous WMT News Translation shared tasks.
The data is organised into 8 columns:
lp: language pair
src: input text
mt: translation
ref: reference translation
score: z score
raw: direct assessment
annotators: number of annotators
domain: domain of the input text (e.g. news)
year: collection year
You can also find the original data for each year in the results section https://www.statmt.org/wmt{YEAR}/results.html… See the full description on the dataset page: https://huggingface.co/datasets/RicardoRei/wmt-da-human-evaluation.RAG_Evaluation_Datasettoporisk-evaluation-artifacts
TopoRisk Evaluation Artifacts
English | 简体中文
TopoRisk studies how a limited verification budget should be allocated across
an agentic workflow graph. Instead of ranking steps only by their local failure
probability, the scheduler estimates how an error can reach terminal outputs
and recomputes marginal value after every selected audit.
This repository is an artifact-first research preview. It contains code,
aggregate measurements, and figures. The manuscript and its LaTeX… See the full description on the dataset page: https://huggingface.co/datasets/JosephAA/toporisk-evaluation-artifacts.crimson-ma2-evaluation
Crimson MolmoAct2: preliminary real-robot evaluation
Right-arm bottle pick-and-place rollouts recorded on 2026-09-24 (Asia/Bangkok).
This release contains three checkpoints, 31 logged attempts, videos from three cameras,
selected recorded telemetry, and joint trajectories. It is an evaluation evidence release,
not a training dataset or a completed six-checkpoint benchmark.
Model source: Kavin60606/crimson-ma2-ckpts.
The policy task text was pick and place the bottle. Each… See the full description on the dataset page: https://huggingface.co/datasets/CrimsonRobot/crimson-ma2-evaluation.evaluation5evaluation1evaluation2ecommerce-analytics-sql-evaluation
Ecommerce Analytics SQL Evaluation (declared GMV, verified answer key)
An evalpack: an evaluation database generated from the answer key, not
annotated after the fact. A VLDB 2026 audit found 52.8% of BIRD Mini-Dev
answer keys wrong because benchmarks annotate answers onto existing
databases; this dataset inverts the order. The declared properties (curves,
shares, identities) are the specification, the database is generated to
satisfy them exactly, and every shipped question was… See the full description on the dataset page: https://huggingface.co/datasets/rasinmuhammed/ecommerce-analytics-sql-evaluation.llm-math-evaluation-dataset
LLM Math Response Evaluation Dataset
Dataset Summary
A human-annotated dataset of 150 AI-generated math responses
evaluated across GPT-4o, Claude, and Gemini. Each response is
scored on Correctness, Reasoning, and Clarity using a structured
rubric, with written justification for every score.
Supported Tasks
LLM evaluation and benchmarking
Math reasoning quality assessment
Error type classification in AI responses
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/nbvbharath-1729/llm-math-evaluation-dataset.NLU-Evaluation-Data-en-de
NLU Evaluation Data - English and German
A labeled English and German language multi-domain dataset (21 domains) with 25K user utterances for human-robot interaction.
This dataset is collected and annotated for evaluating NLU services and platforms.
The detailed paper on this dataset can be found at arXiv.org:
Benchmarking Natural Language Understanding Services for building Conversational Agents
The dataset builds on the annotated data of the xliuhw/NLU-Evaluation-Data
repository.… See the full description on the dataset page: https://huggingface.co/datasets/deutsche-telekom/NLU-Evaluation-Data-en-de.cqa-ai-technical-response-evaluation
CQA AI Technical Response Evaluation Dataset
Overview
This is a synthetic dataset designed to demonstrate structured evaluation of AI-generated technical responses.
The dataset evaluates technical responses beyond a simple correct/incorrect classification by considering multiple dimensions of response quality, including accuracy, reasoning validity, completeness, consistency, assumption handling, clarity, error classification, severity, and expert evaluation.… See the full description on the dataset page: https://huggingface.co/datasets/CQA-pharma/cqa-ai-technical-response-evaluation.fon-code-switching-evaluation
French-Fon Code-Switching Evaluation Benchmark
Overview
This dataset is a benchmark designed to evaluate the contextual understanding of small language models in French-Fon code-switching scenarios.
The benchmark focuses on situations in which French and Fon (Fongbé) are used within the same interaction, with particular attention to cases where critical information required to answer a question is provided in Fon.
The benchmark was developed as part of an academic… See the full description on the dataset page: https://huggingface.co/datasets/fai-adh/fon-code-switching-evaluation.wmt-sqm-human-evaluation
Dataset Summary
In 2022, several changes were made to the annotation procedure used in the WMT Translation task. In contrast to the standard DA (sliding scale from 0-100) used in previous years, in 2022 annotators performed DA+SQM (Direct Assessment + Scalar Quality Metric). In DA+SQM, the annotators still provide a raw score between 0 and 100, but also are presented with seven labeled tick marks. DA+SQM helps to stabilize scores across annotators (as compared to DA).
The data is… See the full description on the dataset page: https://huggingface.co/datasets/RicardoRei/wmt-sqm-human-evaluation.extrinsic-evaluations
Extrinsic evaluations — the union view
One tidy long-format table of every extrinsic (downstream, task-level) evaluation
produced across the 2026-08-26 mergeability workstreams, so that a single file answers
"how did model X score on benchmark Y" regardless of which experiment produced it.
The per-experiment datasets remain the authoritative record of their own methods,
figures and caveats. This is the union view, not a replacement, and it deliberately
carries no analysis of its… See the full description on the dataset page: https://huggingface.co/datasets/Cross-Mergeability/extrinsic-evaluations.saas-finance-sql-evaluation
SaaS Finance SQL Evaluation (MRR waterfalls that reconcile exactly)
An evalpack: an evaluation database generated from the answer key, not
annotated after the fact. A VLDB 2026 audit found 52.8% of BIRD Mini-Dev
answer keys wrong because benchmarks annotate answers onto existing
databases; this dataset inverts the order. The declared properties (curves,
shares, identities) are the specification, the database is generated to
satisfy them exactly, and every shipped question was… See the full description on the dataset page: https://huggingface.co/datasets/rasinmuhammed/saas-finance-sql-evaluation.normative_evaluation_llms_everyday_dilemmascontinuous-evaluation-in-programming-course-logs
Continuous Evaluation in Programming Course - Faculty and Student Logs
Two tables from a first-year programming course run at National Institute in India, July to December 2023. Students worked through a ladder of 60 programming tasks (O01 to O60), marked each task complete themselves on a course dashboard, and a teaching assistant then checked every claimed task in a short viva.
Files
self_reported_claims.csv — 203 rows, one per student.
column
meaning… See the full description on the dataset page: https://huggingface.co/datasets/vicharanashala-org/continuous-evaluation-in-programming-course-logs.edtech-sql-evaluation
EdTech SQL Evaluation (declared learning curve, verified answer key)
An evalpack: an evaluation database generated from the answer key, not
annotated after the fact. A VLDB 2026 audit found 52.8% of BIRD Mini-Dev
answer keys wrong because benchmarks annotate answers onto existing
databases; this dataset inverts the order. The declared properties (curves,
shares, identities) are the specification, the database is generated to
satisfy them exactly, and every shipped question was… See the full description on the dataset page: https://huggingface.co/datasets/rasinmuhammed/edtech-sql-evaluation.memefact-llm-evaluations
MemeFact LLM Evaluations Dataset
This dataset contains 7,680 evaluation records where state-of-the-art Large Language Models (LLMs) assessed fact-checking memes according to specific quality criteria. The dataset provides comprehensive insights into how different AI models evaluate visual-textual content and how these evaluations compare to human judgments.
Dataset Description
Overview
The "MemeFact LLM Evaluations" dataset documents a systematic… See the full description on the dataset page: https://huggingface.co/datasets/sergiogpinto/memefact-llm-evaluations.Misal-Evaluation-v0.1chess_position_evaluationsremote-ai-evaluation-training-market-snapshot
Dataset Description
This is an aggregate August 22, 2026 research snapshot from Specialist AI Work, an independent PatchMedia tracker of reviewed remote AI evaluation, AI training, data annotation-adjacent, and expert-review opportunities.
The live Specialist AI Work inventory has advanced since this snapshot. The counts in this repository describe the immutable August 22 research object; they are not a claim about today's inventory.
Reporting date: 2026-08-22
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/patchmedia-org/remote-ai-evaluation-training-market-snapshot.llm-commit-message-evaluation
Dataset Card for LLM Commit Message Evaluation
The LLM Commit Message Evaluation dataset is designed to evaluate and compare the performance of Large Language Models (LLMs) in generating high-quality git commit messages. It contains real-world code diffs, issue descriptions, and issue titles extracted from open-source repositories (such as OWASP/Nest).
For each code change, the dataset provides the original human-written commit message alongside commit messages generated by… See the full description on the dataset page: https://huggingface.co/datasets/g-for-gour/llm-commit-message-evaluation.Quantum_Gate_Performance_Evaluation
🧪 Quantum Gate Performance Dataset
📘 Title:
Comprehensive Quantum Gate Performance Analysis: A Comparative Study of Noise and No-Noise Effects
📂 Dataset Description:
This repository contains benchmarking results for 13 quantum gates (e.g., H, CNOT, Toffoli) tested under noisy and noise-free conditions, based on 1000 simulation runs per gate configuration. Total 26000 rows and 13 columns.
📊 Features include:
Gate Type
Execution Time
Error Rate
Fidelity… See the full description on the dataset page: https://huggingface.co/datasets/ismielabir/Quantum_Gate_Performance_Evaluation.algozee_mas-feature-evaluation-dataset
MAS Feature Evaluation Dataset
A Structured Dataset for Multi-Agent Capability Analysis
Dataset Info
Source: Kaggle
Original Size: 0.01 MB
Kaggle Downloads: 14
Files: 1
Files
mas_dataset.csv
Mirrored from Kaggle
perplexity_evaluation
SaulLM-7B: Pioneering the first Legal Large Language Model
Perplexity Analysis
This dataset presents the data used in the paper "SaulLM-7B: Pioneering the first Legal Large Language Model" in "6.3 Perplexity Analysis" section.
The dataset contains the perplexity scores of SaulLM-7B, Llama2-7B and Mistral-7B across a corpora of recent text.
Cleaning
We proceeded to standardize the data by removing any special characters using unicodedata normalization.
We also… See the full description on the dataset page: https://huggingface.co/datasets/Equall/perplexity_evaluation.
