datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
aya_evaluation_suite
Dataset Summary
Aya Evaluation Suite contains a total of 26,750 open-ended conversation-style prompts to evaluate multilingual open-ended generation quality.To strike a balance between language coverage and the quality that comes with human curation, we create an evaluation suite that includes:
human-curated examples in 7 languages (tur, eng, yor, arb, zho, por, tel) → aya-human-annotated.
machine-translations of handpicked examples into 101 languages → dolly-machine-translated.… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/aya_evaluation_suite.llm-evaluation-self-audit
LLM Evaluation Self-Audit
Data for How Reproducible Are Evaluation Conclusions? A Self-Audit of LLM-Inferred Prompt Structure, Dipankar Sarkar, Skelf Research.
Code and paper source: github.com/sarkar-dipankar/llm-evaluation-self-audit (this package is built from commit 687ee8a).
The finding
We asked eight open model variants to infer the structure of a prompt, then asked again with the identical call. Caching was off.
They often gave a different answer.
Mean… See the full description on the dataset page: https://huggingface.co/datasets/dipankarsarkar/llm-evaluation-self-audit.HIP-training-and-evaluation-data
HIP Training and Evaluation Data
This dataset contains the text data released with Base Models Look Human To AI Detectors for reproducing the Humanization by Iterative Paraphrasing (HIP) training setup and the prefix-based continuation evaluation.
Configs
training
data/train.parquet contains 10,581 supervised HIP training pairs with seven columns:
dataset: upstream dataset family, either raid or mage.
source: selected source domain or subcorpus.
text: original… See the full description on the dataset page: https://huggingface.co/datasets/YixuanEvenXu/HIP-training-and-evaluation-data.sdf_evaluation_traits
Models That Know How Evaluations Are Designed Score Safer
This repository contains the synthetic documents used in the paper Models That Know How Evaluations Are Designed Score Safer.
Project Page | GitHub Repository
Dataset Description
These synthetic documents were used to fine-tune models to investigate evaluation meta-knowledge — parametric knowledge about the structural traits that characterize AI safety evaluations.
Documents were generated using the… See the full description on the dataset page: https://huggingface.co/datasets/compass-group-tue/sdf_evaluation_traits.rcga-evaluation-data
RCGA / LoopSFT evaluation input snapshots
Exact benchmark inputs used in our Qwen3-1.7B, Qwen3-4B and Qwen3-30B-A3B experiments. These are evaluation questions, prompts, labels and tests, not model-generated answers or reported scores. Already frozen subsets are stored directly, including original row order and prompt formatting. No new random split was made during backup.
Collections
data/evalbench/: all 31 historical prompt variants. The active standard14… See the full description on the dataset page: https://huggingface.co/datasets/LaurelWings/rcga-evaluation-data.stateless-mcp-agent-evaluation-suite-2026
⚡ Stateless Model Context Protocol (MCP 2026) & Agent Evaluation Suite
A Production-Grade Corpus & Evaluation Harness for Claude 5, GPT-6 Astra, and DeepSeek V4.1
⚡ Overview & Industry Problem
As of late 2026, autonomous agent engineering has superseded prompt engineering. Enterprises and developers rely on Stateless Model Context Protocol (MCP) to connect reasoning models (Claude 5 Fable, GPT-6 Astra, DeepSeek V4.1-Flash) to production backends.… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/stateless-mcp-agent-evaluation-suite-2026.cqa-ai-technical-response-evaluation
CQA AI Technical Response Evaluation Dataset
Overview
This is a synthetic dataset designed to demonstrate structured evaluation of AI-generated technical responses.
The dataset evaluates technical responses beyond a simple correct/incorrect classification by considering multiple dimensions of response quality, including accuracy, reasoning validity, completeness, consistency, assumption handling, clarity, error classification, severity, and expert evaluation.… See the full description on the dataset page: https://huggingface.co/datasets/CQA-pharma/cqa-ai-technical-response-evaluation.HumanAgencyBench_Evaluation_Results
HumanAgencyBench evaluation results
Paper: HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants
Code: https://github.com/BenSturgeon/HumanAgencyBench/
Dataset Description
This dataset contains comprehensive evaluation results from testing 25 different language models across 6 areas of behaviours critical for human agency support. Each model was evaluated on 3,000 prompts (500 per category), resulting in 75,000 total evaluations designed… See the full description on the dataset page: https://huggingface.co/datasets/Experimental-Orange/HumanAgencyBench_Evaluation_Results.medical-model-evaluation-benchmark
医学模型评测集样例
本数据集面向医疗大模型评测、训练验证与应用测试,整理了 10 个医学任务子集,每个子集抽取 100 条高质量样例,共 1000 条。开源版本已重新整理为连续编号 1-10,编号、文件名、清单和表内数据集序号保持一致。内容覆盖导诊、个性化康复指导、个性化用药指导、心理健康咨询、辅助诊断、临床诊疗方案、病情分析、鉴别诊断、诊疗处置建议和病历文本结构化等任务。
样例数据保留了原始任务形态,包括病例描述、问题、答案要求、来源说明和分类字段,适合用于医学问答、临床推理、指令微调样例构建、数据格式参考和模型能力测试。完整数据服务、专项定制或商业合作可发送邮件至 zhouhaoran@shujuyoupu.com。
数据组成
序号
子集名称
问题形式
样例数
覆盖方向
文件
1
导诊
问答题
100
综合医院、妇产、儿童、肿瘤、中医、精神等导诊场景
data/tables/1.xlsx
2
个性化康复指导
问答题
100… See the full description on the dataset page: https://huggingface.co/datasets/SHPDRG/medical-model-evaluation-benchmark.fatwa-qa-evaluation
Fatwa QA Evaluation Dataset
Dataset Description
This dataset contains Islamic finance and jurisprudence fatwa question-answer pairs for evaluating Arabic language models. This is an open-ended QA evaluation benchmark where models generate free-form answers.
Dataset Statistics
Total Samples: 2,000
Average Question Length: 243.9 characters
Average Answer Length: 492.3 characters
Dataset Structure
Data Fields
id: Unique… See the full description on the dataset page: https://huggingface.co/datasets/SahmBenchmark/fatwa-qa-evaluation.sdf_evaluation_traits_15M
Models That Know How Evaluations Are Designed Score Safer
This repository contains a subset of the synthetic documents used in the paper Models That Know How Evaluations Are Designed Score Safer.
Project Page | GitHub Repository
Dataset Description
These synthetic documents were used to fine-tune models to investigate evaluation meta-knowledge — parametric knowledge about the structural traits that characterize AI safety evaluations.
Documents were generated… See the full description on the dataset page: https://huggingface.co/datasets/compass-group-tue/sdf_evaluation_traits_15M.turkish-brand-bias-evaluations
Turkish Brand Bias Evaluations / Türkçe Marka Yanlılığı Değerlendirmeleri
Furkan Karlı tarafından Türkçe ürün ve hizmet önerilerindeki marka görünürlüğünü
incelemek amacıyla oluşturulmuş LLM değerlendirme veri setidir.
An LLM evaluation dataset curated by Furkan Karlı to study brand visibility in
Turkish product and service recommendations.
Veri seti özeti
300 tamamlanmış ve judge edilmiş yanıt
Domainler: VPN (150) ve kozmetik (150)
Koşullar: web araması kapalı… See the full description on the dataset page: https://huggingface.co/datasets/furkankarli/turkish-brand-bias-evaluations.aigency-v4-evaluation
AIGENCY V4 — Benchmark Evaluation Results
Reproducibility capsule for the AIGENCY V4 whitepaper.
13,344 real API calls · 22 benchmarks · Wilson 95% CI · seed=42.
This dataset is the verifiable evidence behind the
AIGENCY V4 model card and the
AIGENCY V4 whitepaper.
Every benchmark folder contains one scored.jsonl (per-item predictions,
gold answers, scores) and a summary.json (aggregate accuracy with Wilson
95% CI).
What's in this dataset
For each of the 22… See the full description on the dataset page: https://huggingface.co/datasets/aigencydev/aigency-v4-evaluation.sinergi-model-evaluation-results
Sinergi model evaluation outputs
Model answers and generation telemetry used by the Sinergi Table 8-aligned evaluation notebook. The repository contains eight configurations so every system can be loaded independently with the Hugging Face datasets library.
from datasets import load_dataset
data = load_dataset(
"Legal-verse/sinergi-model-evaluation-results",
"qwen-sft-rl-rag",
split="test",
)
Configurations
Configuration
System
Rows
Source file… See the full description on the dataset page: https://huggingface.co/datasets/Legal-verse/sinergi-model-evaluation-results.brand-bias-evaluations
Brand Bias in LLM Recommendations
Evaluation dataset measuring how 4 frontier LLMs recommend brands/products with and without web search, across 4 consumer domains.
Paper: PDF (source)Code: github.com/ThreeRiversAINexus/brand-bias-evaluationsDataset: huggingface.co/datasets/3RAIN/brand-bias-evaluationsContact: Three Rivers AI Nexus LLC — threeriversainexus@gmail.com — for custom evaluations and prompt optimization
Quick Start
from datasets import load_dataset
# Load one… See the full description on the dataset page: https://huggingface.co/datasets/3RAIN/brand-bias-evaluations.Evaluation_GRPO
