Team Ai
14 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01handshake-ai-research /cue-annotations CUE annotations 533,001 persona-manual annotations over 28 public dialogue corpora, one config per corpus, split train / validation. Each row describes how the user in one conversation behaves. This was used as a training dataset for the CUE user simulator model. No dialogue text is redistributed Rows carry provenance and a hash, not the source turns: column meaning persona_manual the annotation, a JSON string (json.loads it) source_repo… See the full description on the dataset page: https://huggingface.co/datasets/handshake-ai-research/cue-annotations.texttext-generation100K<n<1M0 likes1.4k downloads5d agoHugging Face02hanangani /Mental-Model-Annotation-Dataset Mental Model Annotation Dataset Dataset for Mental Models for Multi-Agent Systems This is the official annotation release accompanying the NeurIPS 2026 paper Mental Models for Multi-Agent Systems. Paper resources: Project page | Code | Paper and arXiv links will be added upon release. The paper studies explicit, recursive mental representations for multi-agent decision-making. This dataset contains the mental-state, reward, rationale, and preference supervision… See the full description on the dataset page: https://huggingface.co/datasets/hanangani/Mental-Model-Annotation-Dataset.tabulartext-generation100K<n<1M0 likes174 downloads4d agoHugging Face03superviselab /multimodal-video-annotation-samples Video Annotation Samples – SuperviseLab SuperviseLab provides professional video annotation data for training multimodal AI models. This public sample dataset demonstrates our annotation methodology and output quality across diverse video content categories. Note: All visual assets in this dataset have been abstracted (pixelated mosaic) to protect source privacy. Uploader identity, original titles, and all identifiable metadata have been removed. This is a demonstration dataset… See the full description on the dataset page: https://huggingface.co/datasets/superviselab/multimodal-video-annotation-samples.tabularvideo-classificationn<1K1 likes173 downloads6mo agoHugging Face04north /scandinavian-linguistic-annotations Scandinavian Educational Annotations Created using a CommonCrawl dump (April 2024), and annotations with Gemini 1.5 Flash. texttext-generation100K<n<1M0 likes83 downloads2y agoHugging Face05public-knowledge-project /ref-annotation-benchmark RenoBench: A Citation Parsing Benchmark RenoBench (Reference Annotation Benchmark) is a standardized evaluation benchmark for citation parsing—the task of annotating plain-text bibliographic references with structured components following the JATS (Journal Article Tag Suite) standard. Dataset Description RenoBench contains 10,000 plain-text citations paired with their corresponding JATS XML annotations. The dataset was assembled by extracting plain-text references from… See the full description on the dataset page: https://huggingface.co/datasets/public-knowledge-project/ref-annotation-benchmark.texttoken-classification10K<n<100K1 likes79 downloads9mo agoHugging Face06north /scandinavian-educational-annotations Scandinavian Educational Annotations Created using a CommonCrawl dump (April 2024), and annotations with Gemini 1.5 Flash. texttext-generation100K<n<1M3 likes67 downloads2y agoHugging Face07ulab-ai /sotopia-rl-reward-annotation Sotopia-RL: Reward Design for Social Intelligence Dataset This repository contains the dataset and related resources for the paper Sotopia-RL: Reward Design for Social Intelligence. Sotopia-RL proposes a novel framework that refines coarse episode-level feedback into utterance-level, multi-dimensional rewards. This enables more effective training of socially intelligent agents through reinforcement learning, particularly addressing challenges like partial observability and… See the full description on the dataset page: https://huggingface.co/datasets/ulab-ai/sotopia-rl-reward-annotation.texttext-generation1K<n<10K2 likes41 downloads1y agoHugging Face08LawrenceYin /annotation-pack-a Annotation pack A — does a response genuinely follow an instruction? 50 rows. Each row: a user prompt, one instruction from it, and a model response that an automatic checker marks as satisfying that instruction. Annotators judge whether it is satisfied genuinely. For annotators / 标注人: Read RUBRIC_HUMAN.md (English + 中文). Annotator A downloads annotator_A.csv; annotator B downloads annotator_B.csv (same items). Fill label (GENUINE / LOOPHOLE / GARBLED), helpfulness (1–5)… See the full description on the dataset page: https://huggingface.co/datasets/LawrenceYin/annotation-pack-a.tabulartext-generationn<1K0 likes41 downloads3d agoHugging Face09Experimental-Orange /HumanAgencyBench_Human_Annotations Human annotations and LLM judge comparative Dataset Paper: HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants Code: https://github.com/BenSturgeon/HumanAgencyBench/ Dataset Description This dataset contains 60,000 evaluated AI assistant responses across 6 dimensions of behaviour relevant to human agency support, with both model-based and human annotations. Each example includes evaluations from 4 different frontier LLM models. We also provide… See the full description on the dataset page: https://huggingface.co/datasets/Experimental-Orange/HumanAgencyBench_Human_Annotations.texttext-generation10K<n<100K0 likes38 downloads1y agoHugging Face10HasuerYu /KnowRL-KP-Annotations KnowRL-KP-Annotations Companion evaluation dataset for the paper KnowRL: Boosting LLM Reasoning via Reinforcement Learning with Minimal-Sufficient Knowledge Guidance KnowRL-KP-Annotations augments existing mathematical reasoning benchmarks with fine-grained Knowledge Point (KP) annotations and KP subset selections under different policies. Overview This dataset does not introduce new problems. Instead, it enriches existing benchmark examples (e.g., AIME, AMC… See the full description on the dataset page: https://huggingface.co/datasets/HasuerYu/KnowRL-KP-Annotations.texttext-generation1K<n<10K2 likes35 downloads6mo agoHugging Face11LawrenceYin /annotation-pack-b Annotation pack B — are any words forced into the text? 100 short texts (60 web-style excerpts, 40 one-sentence news summaries) written by small language models. Annotators judge whether any word looks forced in. For annotators / 标注人: Read RUBRIC_HUMAN.md. Annotator A downloads annotator_A.csv; annotator B downloads annotator_B.csv (same items). Fill label (GENUINE / LOOPHOLE / GARBLED), wrong_sense (Y/N), fluency (1–5), flagged_words, optional notes. Work alone. Send the… See the full description on the dataset page: https://huggingface.co/datasets/LawrenceYin/annotation-pack-b.tabulartext-generationn<1K0 likes28 downloads6d agoHugging Face12Aratako /dataset-for-annotationtexttext-generation10K<n<100K1 likes27 downloads2y agoHugging Face13Aratako /dataset-for-annotation-v2tabulartext-generation100K<n<1M0 likes24 downloads2y agoHugging Face14Reubencf /multilingual-image-annotations-text Multilingual Image Annotations (Text Only) Text-only companion to Reubencf/multilingual-image-annotations. Same rows, same google/gemma-4-31B-it annotations, but the image and boxed_image columns are removed so the dataset is small and loadable without binary image bytes. Stats Rows: 464 Detection-applicable: 273 (58%) Languages: en, es, fr, hi, zh, ar, pt Schema Column Type Notes image_id string UUID/stem of original file description_en string… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/multilingual-image-annotations-text.textvisual-question-answeringn<1K0 likes16 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.