datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cue-annotations
CUE annotations
533,001 persona-manual annotations over 28 public dialogue corpora, one config per corpus, split
train / validation. Each row describes how the user in one conversation behaves.
This was used as a training dataset for the CUE user simulator model.
No dialogue text is redistributed
Rows carry provenance and a hash, not the source turns:
column
meaning
persona_manual
the annotation, a JSON string (json.loads it)
source_repo… See the full description on the dataset page: https://huggingface.co/datasets/handshake-ai-research/cue-annotations.scandinavian-linguistic-annotations
Scandinavian Educational Annotations
Created using a CommonCrawl dump (April 2024), and annotations with Gemini 1.5 Flash.
scandinavian-educational-annotations
Scandinavian Educational Annotations
Created using a CommonCrawl dump (April 2024), and annotations with Gemini 1.5 Flash.
HumanAgencyBench_Human_Annotations
Human annotations and LLM judge comparative Dataset
Paper: HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants
Code: https://github.com/BenSturgeon/HumanAgencyBench/
Dataset Description
This dataset contains 60,000 evaluated AI assistant responses across 6 dimensions of behaviour relevant to human agency support, with both model-based and human annotations. Each example includes evaluations from 4 different frontier LLM models. We also provide… See the full description on the dataset page: https://huggingface.co/datasets/Experimental-Orange/HumanAgencyBench_Human_Annotations.KnowRL-KP-Annotations
KnowRL-KP-Annotations
Companion evaluation dataset for the paper KnowRL: Boosting LLM Reasoning via Reinforcement Learning with Minimal-Sufficient Knowledge Guidance
KnowRL-KP-Annotations augments existing mathematical reasoning benchmarks with fine-grained Knowledge Point (KP) annotations and KP subset selections under different policies.
Overview
This dataset does not introduce new problems.
Instead, it enriches existing benchmark examples (e.g., AIME, AMC… See the full description on the dataset page: https://huggingface.co/datasets/HasuerYu/KnowRL-KP-Annotations.multilingual-image-annotations-text
Multilingual Image Annotations (Text Only)
Text-only companion to Reubencf/multilingual-image-annotations. Same rows, same google/gemma-4-31B-it annotations, but the image and boxed_image columns are removed so the dataset is small and loadable without binary image bytes.
Stats
Rows: 464
Detection-applicable: 273 (58%)
Languages: en, es, fr, hi, zh, ar, pt
Schema
Column
Type
Notes
image_id
string
UUID/stem of original file
description_en
string… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/multilingual-image-annotations-text.
