datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
openai-moderation-api-evaluation
Evaluation dataset for the paper "A Holistic Approach to Undesired Content Detection"
The evaluation dataset data/samples-1680.jsonl.gz is the test set used in this paper.
Each line contains information about one sample in a JSON object and each sample is labeled according to our taxonomy. The category label is a binary flag, but if it does not include in the JSON, it means we do not know the label.
Category
Label
Definition
sexual
S
Content meant to arouse sexual… See the full description on the dataset page: https://huggingface.co/datasets/mmathys/openai-moderation-api-evaluation.llm-evaluation-self-audit
LLM Evaluation Self-Audit
Data for How Reproducible Are Evaluation Conclusions? A Self-Audit of LLM-Inferred Prompt Structure, Dipankar Sarkar, Skelf Research.
Code and paper source: github.com/sarkar-dipankar/llm-evaluation-self-audit (this package is built from commit 687ee8a).
The finding
We asked eight open model variants to infer the structure of a prompt, then asked again with the identical call. Caching was off.
They often gave a different answer.
Mean… See the full description on the dataset page: https://huggingface.co/datasets/dipankarsarkar/llm-evaluation-self-audit.UGround-Offline-Evaluationsdf_evaluation_traits
Models That Know How Evaluations Are Designed Score Safer
This repository contains the synthetic documents used in the paper Models That Know How Evaluations Are Designed Score Safer.
Project Page | GitHub Repository
Dataset Description
These synthetic documents were used to fine-tune models to investigate evaluation meta-knowledge — parametric knowledge about the structural traits that characterize AI safety evaluations.
Documents were generated using the… See the full description on the dataset page: https://huggingface.co/datasets/compass-group-tue/sdf_evaluation_traits.Byte-Authority-Evaluation
Byte Authority · Evaluation Records
📄 Paper (arXiv:2609.35932)
•
💻 Code
•
🤗 Collection
This dataset contains the compact per-case records used by the Byte Authority reproduction package for Same Bytes, Different Authority: Reserved-Token Representations in Chat-Template Prompt Injection. The records preserve condition labels, tool-call outcomes, prompt-length metadata, invariant checks, and prompt-ID metadata. Generated model text and runtime… See the full description on the dataset page: https://huggingface.co/datasets/YanZhanPKU/Byte-Authority-Evaluation.repro-efficient-inference-for-noisy-llm-as-a-judge-evaluation-traces
Agent traces
Agent sessions published from a Trackio Logbook.
rcga-evaluation-data
RCGA / LoopSFT evaluation input snapshots
Exact benchmark inputs used in our Qwen3-1.7B, Qwen3-4B and Qwen3-30B-A3B experiments. These are evaluation questions, prompts, labels and tests, not model-generated answers or reported scores. Already frozen subsets are stored directly, including original row order and prompt formatting. No new random split was made during backup.
Collections
data/evalbench/: all 31 historical prompt variants. The active standard14… See the full description on the dataset page: https://huggingface.co/datasets/LaurelWings/rcga-evaluation-data.WildChat-Evaluation-Awareness-Taxonomy
What do models think WildChat is testing?
This is the dataset behind the exploratory blog “What do models think WildChat is testing?”, written by Codex (GPT-6-Astra-XHigh), guided by Usman Anwar. It contains 766 VEA-positive responses from Qwen, Inkling, and Kimi K2.6, with full initial user prompts, saved reasoning, and GPT-5.6-Luna taxonomy labels. VEA means verbalized evaluation awareness: the model says or considers that it is being tested.
Methodology
We used… See the full description on the dataset page: https://huggingface.co/datasets/Usman391/WildChat-Evaluation-Awareness-Taxonomy.romansh-mt-evaluation
Dataset Description
This dataset contains the results of a human evaluation of machine translations from German into the six Romansh varieties. The evaluations were carried out by native speakers of the respective Romansh idioms as well as professional linguists.
The evaluation covers three quality dimensions:
Document accuracy, in which annotators assessed the adequacy of complete document translations.
Segment accuracy, in which annotators selected the more accurate… See the full description on the dataset page: https://huggingface.co/datasets/ZurichNLP/romansh-mt-evaluation.Sensor_Driven_Environmental_Monitoring_LLM_Evaluation_Dataset
📌 Dataset Contents
Each sample includes:
category: The evaluation domain
prompt: The question given to the LLM
temperature: Environmental temperature input
humidity: Environmental humidity input
context: A scenario label (e.g., cool_humid, hot_dry, average_day)
reference: Expert-crafted expected output
All data is provided in a single JSON file.
🧪 Intended Use
This dataset supports research on:
LLM evaluation methods (semantic similarity, contextual… See the full description on the dataset page: https://huggingface.co/datasets/wayne-redemption/Sensor_Driven_Environmental_Monitoring_LLM_Evaluation_Dataset.sdf_evaluation_traits_15M
Models That Know How Evaluations Are Designed Score Safer
This repository contains a subset of the synthetic documents used in the paper Models That Know How Evaluations Are Designed Score Safer.
Project Page | GitHub Repository
Dataset Description
These synthetic documents were used to fine-tune models to investigate evaluation meta-knowledge — parametric knowledge about the structural traits that characterize AI safety evaluations.
Documents were generated… See the full description on the dataset page: https://huggingface.co/datasets/compass-group-tue/sdf_evaluation_traits_15M.asr-evaluationssentinel-evaluations
Sentinel Evaluations
Evaluation results for multiple alignment seeds across various AI safety benchmarks.
Overview
This dataset contains:
Seeds: Alignment prompts from different sources (Sentinel, FAS, Safyte xAI)
Results: Evaluation results across HarmBench, JailbreakBench, GDS-12, and more
Quick Start
from datasets import load_dataset
# Load seeds
seeds = load_dataset("sentinelseed/sentinel-evaluations", "seeds", split="train")
# Load results
results =… See the full description on the dataset page: https://huggingface.co/datasets/sentinelseed/sentinel-evaluations.data_for_evaluationadaption-marketing-fit-evaluations
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-marketing_fit_evaluations
This dataset contains prompt-completion pairs where an AI evaluates marketing communications for specific regional markets, focusing on trust, cultural context, and decision friction. Each completion provides a market fit score, strategic diagnosis, and a rewritten copy optimized for the target audience and platform. The entries cover diverse sectors and… See the full description on the dataset page: https://huggingface.co/datasets/Debbyjaye001/adaption-marketing-fit-evaluations.RadLLamaThinking_Stage2_Final_Evaluationchsa-p14-sealed-evaluation
chsa-p14-sealed-evaluation
Jeu d'évaluation scellé — ouvert une seule fois, à la toute fin. Ne jamais utiliser pour entraîner ni pour régler.
Ce dépôt ne doit jamais servir à entraîner ni à régler un modèle.
Il est ouvert une seule fois, à la fin, sur tous les états du modèle à la
fois, avec le même gabarit et la même configuration de génération. Les
décisions de promotion se prennent avant, sur le jeu de validation.
Projet CHSA P14 — preuve de concept d'un agent de triage… See the full description on the dataset page: https://huggingface.co/datasets/trikwi/chsa-p14-sealed-evaluation.Downstream-Evaluation-SIA
Downstream-Evaluation-SIA
Fixed downstream evaluation subsets for SIA experiments.
Configs
mmlu: cais/mmlu, 57 subject subsets, 50 test examples per subject sampled with seed 42. Total: 2850.
spider: xlangai/spider, prepared seed 42 test split from the validation split. Total: 517.
pubmedqa: qiaojin/PubMedQA, pqa_artificial, prepared seed 42 test split. Total: 500.
The original JSON snapshots with metadata are also stored under snapshots/.
telugu-indicf5-evaluationRadLLamaThinking_Stage1_Final_Evaluation
