datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Mental-Model-Annotation-Dataset
Mental Model Annotation Dataset
Dataset for Mental Models for Multi-Agent Systems
This is the official annotation release accompanying the NeurIPS 2026 paper
Mental Models for Multi-Agent Systems.
Paper resources: Project page
| Code | Paper and arXiv links
will be added upon release.
The paper studies explicit, recursive mental representations for multi-agent
decision-making. This dataset contains the mental-state, reward, rationale, and
preference supervision… See the full description on the dataset page: https://huggingface.co/datasets/hanangani/Mental-Model-Annotation-Dataset.multimodal-video-annotation-samples
Video Annotation Samples – SuperviseLab
SuperviseLab provides professional video annotation data for training multimodal AI models. This public sample dataset demonstrates our annotation methodology and output quality across diverse video content categories.
Note: All visual assets in this dataset have been abstracted (pixelated mosaic) to protect source privacy. Uploader identity, original titles, and all identifiable metadata have been removed. This is a demonstration dataset… See the full description on the dataset page: https://huggingface.co/datasets/superviselab/multimodal-video-annotation-samples.sotopia-rl-reward-annotation
Sotopia-RL: Reward Design for Social Intelligence Dataset
This repository contains the dataset and related resources for the paper Sotopia-RL: Reward Design for Social Intelligence.
Sotopia-RL proposes a novel framework that refines coarse episode-level feedback into utterance-level, multi-dimensional rewards. This enables more effective training of socially intelligent agents through reinforcement learning, particularly addressing challenges like partial observability and… See the full description on the dataset page: https://huggingface.co/datasets/ulab-ai/sotopia-rl-reward-annotation.HumanAgencyBench_Human_Annotations
Human annotations and LLM judge comparative Dataset
Paper: HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants
Code: https://github.com/BenSturgeon/HumanAgencyBench/
Dataset Description
This dataset contains 60,000 evaluated AI assistant responses across 6 dimensions of behaviour relevant to human agency support, with both model-based and human annotations. Each example includes evaluations from 4 different frontier LLM models. We also provide… See the full description on the dataset page: https://huggingface.co/datasets/Experimental-Orange/HumanAgencyBench_Human_Annotations.
