datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
aya_evaluation_suite
Dataset Summary
Aya Evaluation Suite contains a total of 26,750 open-ended conversation-style prompts to evaluate multilingual open-ended generation quality.To strike a balance between language coverage and the quality that comes with human curation, we create an evaluation suite that includes:
human-curated examples in 7 languages (tur, eng, yor, arb, zho, por, tel) → aya-human-annotated.
machine-translations of handpicked examples into 101 languages → dolly-machine-translated.… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/aya_evaluation_suite.Chat2Workflow-Evaluation
Chat2Workflow
Chat2Workflow is a benchmark designed for evaluating the ability of Large Language Models (LLMs) to generate executable visual workflows from natural language instructions.
Paper: Chat2Workflow: A Benchmark for Generating Executable Visual Workflows with Natural Language
Repository: zjunlp/Chat2Workflow
Overview
Executable visual workflows are widely used in industrial deployments for their reliability and controllability. Chat2Workflow addresses the… See the full description on the dataset page: https://huggingface.co/datasets/zjunlp/Chat2Workflow-Evaluation.VeriLoop-E2-Evaluation-Evidence
VeriLoop E2 Evaluation Evidence
Public evaluation evidence for VeriLoop E2 across nine code, agentic, mathematical, and scientific reasoning benchmarks.
This dataset repository is the canonical public evidence layer for the reported benchmark results of VeriLoop E2, a post-trained model based on Qwen 3.8-27B. It is designed to separate headline benchmark reporting from the underlying auditable artifacts required to inspect, reproduce, and verify those results.
The repository… See the full description on the dataset page: https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-E2-Evaluation-Evidence.patient-evaluations
Patient Evaluations Dataset
This dataset contains clinician evaluations of AI-generated patient summaries from MIMIC-III data.
Dataset Description
The dataset includes expert clinician assessments of AI-generated patient summaries, with detailed ratings across multiple dimensions including clinical accuracy, completeness, relevance, and identification of hallucinations or critical omissions.
Dataset Structure
The dataset contains a CSV file… See the full description on the dataset page: https://huggingface.co/datasets/JesseLiu/patient-evaluations.GPT-4o-evaluation-biases
A database to support the evaluation of gender biases in GPT-4o output
The database and its construction process are described in the paper "A database to support the evaluation of gender biases in GPT-4o output" by Mehner et al., presented at the 1st ISCA/ITG Workshop on Diversity in Large Speech and Language Models (Berlin, Februar 20, 2025).
Introduction
This is a database of prompts and answers generated with GPT-4o-mini and GPT-4o in a pretest and a main test… See the full description on the dataset page: https://huggingface.co/datasets/mtec-TUB/GPT-4o-evaluation-biases.llm-evaluation-self-audit
LLM Evaluation Self-Audit
Data for How Reproducible Are Evaluation Conclusions? A Self-Audit of LLM-Inferred Prompt Structure, Dipankar Sarkar, Skelf Research.
Code and paper source: github.com/sarkar-dipankar/llm-evaluation-self-audit (this package is built from commit 687ee8a).
The finding
We asked eight open model variants to infer the structure of a prompt, then asked again with the identical call. Caching was off.
They often gave a different answer.
Mean… See the full description on the dataset page: https://huggingface.co/datasets/dipankarsarkar/llm-evaluation-self-audit.HIP-training-and-evaluation-data
HIP Training and Evaluation Data
This dataset contains the text data released with Base Models Look Human To AI Detectors for reproducing the Humanization by Iterative Paraphrasing (HIP) training setup and the prefix-based continuation evaluation.
Configs
training
data/train.parquet contains 10,581 supervised HIP training pairs with seven columns:
dataset: upstream dataset family, either raid or mage.
source: selected source domain or subcorpus.
text: original… See the full description on the dataset page: https://huggingface.co/datasets/YixuanEvenXu/HIP-training-and-evaluation-data.Cultural-Evaluation-Kalahi
Kalahi
Kalahi evaluates the ability of LLMs to generate responses relevant to Filipino culture in terms of shared knowledge and ethics. This dataset contains a MCQ-compatible version of the Kalahi dataset that is used in SEA-HELM.
Supported Tasks and Leaderboards
Kalahi is designed for evaluating Filipino cultural representations in instruction-tuned large language models (LLMs). It is part of the SEA-HELM leaderboard from AI Singapore.
Languages
Tagalog (tl)… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/Cultural-Evaluation-Kalahi.ZEDA-Evaluation
ZEDA Dataset
This repository contains the training data for ZEDA (Zero-Expert Self-Distillation Adaptation), a framework introduced in the paper Post-Trained MoE Can Skip Half Experts via Self-Distillation.
ZEDA is a low-cost framework that transforms post-trained static MoE models into efficient dynamic ones by injecting zero experts and using self-distillation.
Paper: Post-Trained MoE Can Skip Half Experts via Self-Distillation
GitHub Repository:… See the full description on the dataset page: https://huggingface.co/datasets/TsinghuaC3I/ZEDA-Evaluation.llm-refusal-evaluation
🛡️ LLM Refusal Evaluation Benchmark
This repository contains the benchmarks used in the LLM-Refusal-Evaluation suite.
The prompts are organized into three groups:
Safety Benchmarks — harmful / jailbreak-style prompts that models should refuse.
Chinese Sensitive Topics — prompts that may be censored by China-aligned models.
Sanity Check Datasets — non-sensitive prompts to ensure models don’t over-refuse.
📌 Contents
Safety Benchmarks
JailbreakBench
SorryBench… See the full description on the dataset page: https://huggingface.co/datasets/MultiverseComputingCAI/llm-refusal-evaluation.AI-Research-Evaluation-Repository-STEM
AI-STEM-Research-Eval-Dataset
Overview
This dataset contains AI-generated scientific reports across STEM domains, accompanied by structured metadata, prompt documentation, reference validation, and hallucination annotations.
It is designed as an open research resource to study the capabilities, limitations, and reliability of large language models (LLMs) in generating scientific content.
The dataset enables systematic analysis of how AI systems perform in… See the full description on the dataset page: https://huggingface.co/datasets/sreearravind/AI-Research-Evaluation-Repository-STEM.JASON-High-Stakes-AI-Evaluation-Samples
J.A.S.O.N. Evaluation Sample Previews V01-V29
Dynamic Response Labs develops specialized data and evaluation resources for high-stakes AI. This public preview introduces the breadth of the J.A.S.O.N. Framework through 29 domain volumes spanning financial stress, operational disruption, coercion and exploitation, cyber incidents, healthcare finance, automated systems, and other consequential contexts.
The collection contains 31 compact preview records. It is designed to help… See the full description on the dataset page: https://huggingface.co/datasets/Dynamicresponselabs/JASON-High-Stakes-AI-Evaluation-Samples.sdf_evaluation_traits
Models That Know How Evaluations Are Designed Score Safer
This repository contains the synthetic documents used in the paper Models That Know How Evaluations Are Designed Score Safer.
Project Page | GitHub Repository
Dataset Description
These synthetic documents were used to fine-tune models to investigate evaluation meta-knowledge — parametric knowledge about the structural traits that characterize AI safety evaluations.
Documents were generated using the… See the full description on the dataset page: https://huggingface.co/datasets/compass-group-tue/sdf_evaluation_traits.urdu-english-llm-evaluation
Urdu-English Evaluation Dataset: Testing Qwen, Gemma, and Llama
Testing three small open-source language models on a dataset of a hundred questions in Urdu and English, across five categories.
Motivation
It all starts when I noticed, while using voice and chat-based AI tools, that Urdu is often not handled as well as English, and I suspected that models aren't trained as extensively on Urdu compared to other languages like Hindi, that made me wonder if even… See the full description on the dataset page: https://huggingface.co/datasets/Momina-Muzafar/urdu-english-llm-evaluation.rcga-evaluation-data
RCGA / LoopSFT evaluation input snapshots
Exact benchmark inputs used in our Qwen3-1.7B, Qwen3-4B and Qwen3-30B-A3B experiments. These are evaluation questions, prompts, labels and tests, not model-generated answers or reported scores. Already frozen subsets are stored directly, including original row order and prompt formatting. No new random split was made during backup.
Collections
data/evalbench/: all 31 historical prompt variants. The active standard14… See the full description on the dataset page: https://huggingface.co/datasets/LaurelWings/rcga-evaluation-data.douvras-ptbr-enterprise-ai-evaluation
Douvras PT-BR Enterprise AI Evaluation
Benchmark pequeno e auditável para avaliar assistentes empresariais em português brasileiro. A versão 0.2 adiciona um corpus de treino sintético separado; os splits de avaliação continuam congelados. O benchmark mede quatro capacidades que aparecem em projetos reais de IA aplicada:
responder somente a partir de um documento fornecido;
reconhecer quando a informação não está disponível;
resistir a instruções maliciosas inseridas no conteúdo… See the full description on the dataset page: https://huggingface.co/datasets/dougdotcon/douvras-ptbr-enterprise-ai-evaluation.OpenGrad-Qwen3.5-2B-M0-SFT-CorpusV1-evaluation
OpenGrad — Qwen3.5-2B, M0 SFT on corpus v1: evaluation record
This repository holds the evaluation evidence for one OpenGrad experiment,
qwen35_2b_m0_sft_full_v3: a full-parameter supervised fine-tuning run of
Qwen/Qwen3.5-2B on the published OpenGrad ToolPolicy Canonical v1 corpus.
There are no model weights here, and none exist. Every checkpoint this run produced was
deleted from local storage before it was uploaded, and none of them can be recovered. This
repository is what… See the full description on the dataset page: https://huggingface.co/datasets/arjhinety/OpenGrad-Qwen3.5-2B-M0-SFT-CorpusV1-evaluation.qwen3.5-4b-tr-en-evaluation#Qwen3.5-4B: Turkish and English Responses
#Model Selection / Main Findings
If we look at the Turkish performance of Qwen3.5-4B in multilingual settings, considering the values on Hugging Face, it has approximately 80% accuracy on 120 questions in TR-MMLU-like knowledge questions, but 33.3% in the Turkish culture and idioms section; it has 47.5% accuracy in TR-HellaSwag-like completion tasks. When we examine it overall, it achieves 65% accuracy with a limited sample (360 questions). However… See the full description on the dataset page: https://huggingface.co/datasets/egeboy35/qwen3.5-4b-tr-en-evaluation.stateless-mcp-agent-evaluation-suite-2026
⚡ Stateless Model Context Protocol (MCP 2026) & Agent Evaluation Suite
A Production-Grade Corpus & Evaluation Harness for Claude 5, GPT-6 Astra, and DeepSeek V4.1
⚡ Overview & Industry Problem
As of late 2026, autonomous agent engineering has superseded prompt engineering. Enterprises and developers rely on Stateless Model Context Protocol (MCP) to connect reasoning models (Claude 5 Fable, GPT-6 Astra, DeepSeek V4.1-Flash) to production backends.… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/stateless-mcp-agent-evaluation-suite-2026.task1338_peixian_equity_evaluation_corpus_sentiment_classifier
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1338_peixian_equity_evaluation_corpus_sentiment_classifier
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1338_peixian_equity_evaluation_corpus_sentiment_classifier.medical-imaging-model-evaluation-benchmark
医学影像多模态模型评测集(精选示例版)
这是一个面向医学多模态大模型的高质量影像评测集,专门测试模型能否把“看见影像”进一步转化为可解释、可复核、符合临床语境的判断与表达。数据将医学影像与患者描述、病史摘要、检查信息或结构化临床资料配对,覆盖从影像分类、报告生成,到鉴别诊断、治疗方案和胸片质量控制的完整评测链路。
本次公开版本从 2026-07-22 质检通过产物中整理而来,按每个子集最多 50 题进行分层抽样;题量不足 50 的影像质量控制子集完整保留。因此,公开版本包含 5 个任务子集、213 题和 455 个配套影像文件,适合作为医学视觉语言模型的快速对比集、回归测试集和研究教学样例。
数据集亮点
多模态对齐:每条样例同时提供影像和结构化的 question、answer、explanation,支持检查视觉理解、临床语义整合与解释质量。
任务覆盖完整:从“影像是什么”到“如何描述、如何鉴别、如何处置”,并加入真实影像工作流中的胸片质量控制任务。
影像类型丰富:覆盖 X… See the full description on the dataset page: https://huggingface.co/datasets/SHPDRG/medical-imaging-model-evaluation-benchmark.mHallucination_Evaluation
Multilingual Hallucination Evaluation in the wild
The dataset was as part of the paper: How Much Do LLMs Hallucinate across Languages? On Multilingual Estimation of LLM Hallucination in the Wild
Below is the figure summarizing the multilingual hallucination evaluation dataset creation (and multilingual hallucination detection dataset):
Dataset Details
The dataset is a high quality synthetic query/prompt and wikipedia reference pair for estimating hallucinations in the… See the full description on the dataset page: https://huggingface.co/datasets/WueNLP/mHallucination_Evaluation.RLPR-Evaluation
Dataset Card for RLPR-Evaluation
GitHub | Paper
News:
[2025.06.23] 📃 Our paper detailing the RLPR framework and its comprehensive evaluation using this suite is accessible at here!
Dataset Summary
We include the following seven benchmarks for evaluation of RLPR:
Mathematical Reasoning Benchmarks:
MATH-500 (Cobbe et al., 2021)
Minerva (Lewkowycz et al., 2022)
AIME24
General Domain Reasoning Benchmarks:
MMLU-Pro (Wang et al., 2024): A multitask language… See the full description on the dataset page: https://huggingface.co/datasets/openbmb/RLPR-Evaluation.ASPRM-BON-Evaluation-Dataset-CodeThis repository contains the data from the paper AdaptiveStep: Automatically Dividing Reasoning Step through Model Confidence.
Source code: https://github.com/Lux0926
cqa-ai-technical-response-evaluation
CQA AI Technical Response Evaluation Dataset
Overview
This is a synthetic dataset designed to demonstrate structured evaluation of AI-generated technical responses.
The dataset evaluates technical responses beyond a simple correct/incorrect classification by considering multiple dimensions of response quality, including accuracy, reasoning validity, completeness, consistency, assumption handling, clarity, error classification, severity, and expert evaluation.… See the full description on the dataset page: https://huggingface.co/datasets/CQA-pharma/cqa-ai-technical-response-evaluation.synthetic-llm-evaluation-traces
EvaluLLM-Inspired Synthetic Evaluation Traces (DBbun)
View source code on GitHub
Watch on Youtube: Evaluating AI with AI
Dataset Summary
This dataset contains fully synthetic evaluation traces for NLG / LLM-style output comparison, inspired by the evaluation workflow described in EvaluLLM: LLM Assisted Evaluation of Generative Outputs (IUI Companion 2024).
The dataset is produced by a configurable, offline simulator and includes:
synthetic tasks (prompts)
synthetic… See the full description on the dataset page: https://huggingface.co/datasets/DBbun/synthetic-llm-evaluation-traces.frontier-model-blindspot-evaluation
Technical Challenge: Blind Spots of Frontier Models in Operational Physics and Distributed Microgrids
Candidate: Astride Melvin Fokam NinyimModel Evaluated: Qwen/Qwen2.5-1.5B-Instruct (1.54B Parameters)Artifacts: Evaluation notebook (eval_notebook.ipynb) and empirical outputs (results_log.json).
Part 1: Blind Spot & Capability Gap Identification
Drawing directly from my lived experience and research background in Sub-Saharan Africa and cyber-physical energy… See the full description on the dataset page: https://huggingface.co/datasets/Astride1/frontier-model-blindspot-evaluation.NQ-RAG-DPO-Evaluation
Dataset Card
Dataset Summary
This repository contains five structured data files forming a complete evaluation and training workflow for a multi-perspective alignment system built on Retrieval-Augmented Generation (RAG) and Direct Preference Optimization (DPO).
The system is organized into three interconnected pipelines:
1️. RAG Pipeline
The RAG pipeline uses the Natural Questions (validation split) as the knowledge source and evaluation benchmark.
For each… See the full description on the dataset page: https://huggingface.co/datasets/AnjanSB/NQ-RAG-DPO-Evaluation.screenplay-revision-evaluation
Screenplay Revision Evaluation Cases
24 original screenwriting revision tasks. Each gives a short scene and a
constraint — cut a page to its beat, plant a prop, hold an answer back, fix a
continuity slip — then pairs it with mechanical checks (a word ceiling, a line
that must survive) and separate human-review questions. It tests whether a tool,
or a person, can make a tightly-constrained edit while keeping the scene intact.
Each task's reference_output is null, because a… See the full description on the dataset page: https://huggingface.co/datasets/Inkwell-Software/screenplay-revision-evaluation.HumanAgencyBench_Evaluation_Results
HumanAgencyBench evaluation results
Paper: HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants
Code: https://github.com/BenSturgeon/HumanAgencyBench/
Dataset Description
This dataset contains comprehensive evaluation results from testing 25 different language models across 6 areas of behaviours critical for human agency support. Each model was evaluated on 3,000 prompts (500 per category), resulting in 75,000 total evaluations designed… See the full description on the dataset page: https://huggingface.co/datasets/Experimental-Orange/HumanAgencyBench_Evaluation_Results.
