datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cue-annotations
CUE annotations
533,001 persona-manual annotations over 28 public dialogue corpora, one config per corpus, split
train / validation. Each row describes how the user in one conversation behaves.
This was used as a training dataset for the CUE user simulator model.
No dialogue text is redistributed
Rows carry provenance and a hash, not the source turns:
column
meaning
persona_manual
the annotation, a JSON string (json.loads it)
source_repo… See the full description on the dataset page: https://huggingface.co/datasets/handshake-ai-research/cue-annotations.Mental-Model-Annotation-Dataset
Mental Model Annotation Dataset
Dataset for Mental Models for Multi-Agent Systems
This is the official annotation release accompanying the NeurIPS 2026 paper
Mental Models for Multi-Agent Systems.
Paper resources: Project page
| Code | Paper and arXiv links
will be added upon release.
The paper studies explicit, recursive mental representations for multi-agent
decision-making. This dataset contains the mental-state, reward, rationale, and
preference supervision… See the full description on the dataset page: https://huggingface.co/datasets/hanangani/Mental-Model-Annotation-Dataset.multimodal-video-annotation-samples
Video Annotation Samples – SuperviseLab
SuperviseLab provides professional video annotation data for training multimodal AI models. This public sample dataset demonstrates our annotation methodology and output quality across diverse video content categories.
Note: All visual assets in this dataset have been abstracted (pixelated mosaic) to protect source privacy. Uploader identity, original titles, and all identifiable metadata have been removed. This is a demonstration dataset… See the full description on the dataset page: https://huggingface.co/datasets/superviselab/multimodal-video-annotation-samples.scandinavian-linguistic-annotations
Scandinavian Educational Annotations
Created using a CommonCrawl dump (April 2024), and annotations with Gemini 1.5 Flash.
ref-annotation-benchmark
RenoBench: A Citation Parsing Benchmark
RenoBench (Reference Annotation Benchmark) is a standardized evaluation benchmark for citation parsing—the task of annotating plain-text bibliographic references with structured components following the JATS (Journal Article Tag Suite) standard.
Dataset Description
RenoBench contains 10,000 plain-text citations paired with their corresponding JATS XML annotations. The dataset was assembled by extracting plain-text references from… See the full description on the dataset page: https://huggingface.co/datasets/public-knowledge-project/ref-annotation-benchmark.scandinavian-educational-annotations
Scandinavian Educational Annotations
Created using a CommonCrawl dump (April 2024), and annotations with Gemini 1.5 Flash.
sotopia-rl-reward-annotation
Sotopia-RL: Reward Design for Social Intelligence Dataset
This repository contains the dataset and related resources for the paper Sotopia-RL: Reward Design for Social Intelligence.
Sotopia-RL proposes a novel framework that refines coarse episode-level feedback into utterance-level, multi-dimensional rewards. This enables more effective training of socially intelligent agents through reinforcement learning, particularly addressing challenges like partial observability and… See the full description on the dataset page: https://huggingface.co/datasets/ulab-ai/sotopia-rl-reward-annotation.annotation-pack-a
Annotation pack A — does a response genuinely follow an instruction?
50 rows. Each row: a user prompt, one instruction from it, and a model response that an automatic checker
marks as satisfying that instruction. Annotators judge whether it is satisfied genuinely.
For annotators / 标注人:
Read RUBRIC_HUMAN.md (English + 中文).
Annotator A downloads annotator_A.csv; annotator B downloads annotator_B.csv (same items).
Fill label (GENUINE / LOOPHOLE / GARBLED), helpfulness (1–5)… See the full description on the dataset page: https://huggingface.co/datasets/LawrenceYin/annotation-pack-a.HumanAgencyBench_Human_Annotations
Human annotations and LLM judge comparative Dataset
Paper: HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants
Code: https://github.com/BenSturgeon/HumanAgencyBench/
Dataset Description
This dataset contains 60,000 evaluated AI assistant responses across 6 dimensions of behaviour relevant to human agency support, with both model-based and human annotations. Each example includes evaluations from 4 different frontier LLM models. We also provide… See the full description on the dataset page: https://huggingface.co/datasets/Experimental-Orange/HumanAgencyBench_Human_Annotations.KnowRL-KP-Annotations
KnowRL-KP-Annotations
Companion evaluation dataset for the paper KnowRL: Boosting LLM Reasoning via Reinforcement Learning with Minimal-Sufficient Knowledge Guidance
KnowRL-KP-Annotations augments existing mathematical reasoning benchmarks with fine-grained Knowledge Point (KP) annotations and KP subset selections under different policies.
Overview
This dataset does not introduce new problems.
Instead, it enriches existing benchmark examples (e.g., AIME, AMC… See the full description on the dataset page: https://huggingface.co/datasets/HasuerYu/KnowRL-KP-Annotations.annotation-pack-b
Annotation pack B — are any words forced into the text?
100 short texts (60 web-style excerpts, 40 one-sentence news summaries) written by small language models.
Annotators judge whether any word looks forced in.
For annotators / 标注人:
Read RUBRIC_HUMAN.md.
Annotator A downloads annotator_A.csv; annotator B downloads annotator_B.csv (same items).
Fill label (GENUINE / LOOPHOLE / GARBLED), wrong_sense (Y/N), fluency (1–5), flagged_words, optional notes. Work alone.
Send the… See the full description on the dataset page: https://huggingface.co/datasets/LawrenceYin/annotation-pack-b.dataset-for-annotationdataset-for-annotation-v2multilingual-image-annotations-text
Multilingual Image Annotations (Text Only)
Text-only companion to Reubencf/multilingual-image-annotations. Same rows, same google/gemma-4-31B-it annotations, but the image and boxed_image columns are removed so the dataset is small and loadable without binary image bytes.
Stats
Rows: 464
Detection-applicable: 273 (58%)
Languages: en, es, fr, hi, zh, ar, pt
Schema
Column
Type
Notes
image_id
string
UUID/stem of original file
description_en
string… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/multilingual-image-annotations-text.
