Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ljvmiranda921 /gsd-humaneval-annotationstext1K<n<10K1 likes3.3k downloads15d agoHugging Face02LumiOpen /hpltv2-llama33-edu-annotation HPLT version 2.0 educational annotations This dataset contains annotations derived from HPLT v2 cleaned samples. There are 500,000 annotations for each language if the source contains at least 500,000 samples. We prompt Llama-3.3-70B-Instruct to score web pages based on their educational value following FineWeb-Edu classifier. Note 1: The dataset contains the prompt (using the first 1500 characters of the text sample), the scores, and the full Llama 3 generation. The column "idx"… See the full description on the dataset page: https://huggingface.co/datasets/LumiOpen/hpltv2-llama33-edu-annotation.tabular10M<n<100M3 likes1.8k downloads1y agoHugging Face03JQL-AI /JQL-LLM-Edu-Annotations 📚 JQL Educational Quality Annotations from LLMs This dataset provides 17,186,606 documents with high-quality LLM annotations for evaluating the educational value of web documents, and serves as a benchmark for training and evaluating multilingual LLM annotators as described in the JQL paper. 📝 Dataset Summary Multilingual document-level quality annotations scored on a 0–5 educational value scale by three state-of-the-art LLMs: Gemma-3-27B-it, Mistral-3.1-24B-it… See the full description on the dataset page: https://huggingface.co/datasets/JQL-AI/JQL-LLM-Edu-Annotations.tabular10M<n<100M2 likes1.3k downloads1y agoHugging Face04ce-amtic /ProcVQA-20M-annotationsgated ProcVQA-20M Annotations Project Page | arXiv | Code | Model | Media This repository contains the text annotations for the ProcVQA-20M dataset. The full image files are hosted separately on ProcVQA-20M-media. Overview This dataset is constructed from over 26 embodied datasets, comprising: 20M QA pairs for training 330K original trajectories 50M annotated frames from ~5,000 hours of manipulation data 200+ different tasks Dataset Structure The… See the full description on the dataset page: https://huggingface.co/datasets/ce-amtic/ProcVQA-20M-annotations.image10K<n<100K1 likes764 downloads5mo agoHugging Face05Bofeee5675 /GUI-Net-1M-relative-annotationstext100K<n<1M0 likes335 downloads1y agoHugging Face06nvidia /SEED-Timeline-Annotations Timeline Annotations for BONES-SEED Humanoid Motion Dataset Dataset Description: This dataset provides additional text description annotations from the BONES-SEED humanoid motion dataset. For each motion, this dataset provides an overview text description of the entire motion at a high level, along with a “timeline” of annotated segments within the motion. Each segment generally contains a single atomic action and is defined by a start time, end time, and text… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/SEED-Timeline-Annotations.text100K<n<1M7 likes291 downloads6mo agoHugging Face07Bofeee5675 /GUI-Net-1M-absolute-annotationstext100K<n<1M2 likes268 downloads1y agoHugging Face08ambient-intelligence-labs /egolongqa-synth-annotations EgoLongQA synthetic MCQs, teacher traces and annotation outputs Everything produced by the annotation and synthesis pipelines for the AI Wearables Challenge 2026 EgoLongQA ≤2B track, other than the distillation set (which lives in infinitylogesh/egolongqa-junior-distill). ⚠️ Read this before counting rows The synthetic set is 943 questions over 408 videos, and it is stored two ways: file rows shape training_sets/train_synth_v3.jsonl 943 flat — one row… See the full description on the dataset page: https://huggingface.co/datasets/ambient-intelligence-labs/egolongqa-synth-annotations.tabularvisual-question-answering1K<n<10K0 likes213 downloads1mo agoHugging Face09Himpq /kwext-bilibili-video-title-annotations KwExt Bilibili Video Title Annotations This dataset is a model-assisted annotation set for the KwExt keyword extraction project. The current snapshot contains 5,000 Chinese Bilibili video titles from annotation stages video_title_zh_001 through video_title_zh_005, with 1,000 records in each stage. The release is intended for early experiments with: extracting title-grounded keywords and ranking their importance; broad semantic tags for retrieval and RAG metadata; dense tag… See the full description on the dataset page: https://huggingface.co/datasets/Himpq/kwext-bilibili-video-title-annotations.tabulartoken-classification1K<n<10K0 likes183 downloads23d agoHugging Face10hanangani /Mental-Model-Annotation-Dataset Mental Model Annotation Dataset Dataset for Mental Models for Multi-Agent Systems This is the official annotation release accompanying the NeurIPS 2026 paper Mental Models for Multi-Agent Systems. Paper resources: Project page | Code | Paper and arXiv links will be added upon release. The paper studies explicit, recursive mental representations for multi-agent decision-making. This dataset contains the mental-state, reward, rationale, and preference supervision… See the full description on the dataset page: https://huggingface.co/datasets/hanangani/Mental-Model-Annotation-Dataset.tabulartext-generation100K<n<1M0 likes174 downloads4d agoHugging Face11superviselab /multimodal-video-annotation-samples Video Annotation Samples – SuperviseLab SuperviseLab provides professional video annotation data for training multimodal AI models. This public sample dataset demonstrates our annotation methodology and output quality across diverse video content categories. Note: All visual assets in this dataset have been abstracted (pixelated mosaic) to protect source privacy. Uploader identity, original titles, and all identifiable metadata have been removed. This is a demonstration dataset… See the full description on the dataset page: https://huggingface.co/datasets/superviselab/multimodal-video-annotation-samples.tabularvideo-classificationn<1K1 likes173 downloads6mo agoHugging Face12LianeMarilin /4k-video-annotations 4K Video Annotations — Shot Segmentation and Camera Motion This dataset contains 12 frame-accurate shot clips segmented from five short cinematic video sequences. Every clip is paired with a detailed, manually reviewed annotation covering visible content, subject actions, shot scale, camera angle, camera movement, movement direction, stabilization, composition, lighting, color, pacing, transitions, timecodes, and technical properties. The footage depicts a tense nighttime… See the full description on the dataset page: https://huggingface.co/datasets/LianeMarilin/4k-video-annotations.imagen<1K0 likes171 downloads26d agoHugging Face13MBZUAI /video_annotation_pipeline 👁️ Semi-Automatic Video Annotation Pipeline 📝 Description Video-ChatGPT introduces the VideoInstruct100K dataset, which employs a semi-automatic annotation pipeline to generate 75K instruction-tuning QA pairs. To address the limitations of this annotation process, we present VCG+112K dataset developed through an improved annotation pipeline. Our approach improves the accuracy and quality of instruction tuning pairs by improving keyframe extraction, leveraging SoTA… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI/video_annotation_pipeline.textn<1K2 likes150 downloads2y agoHugging Face14JQL-AI /JQL-Human-Edu-Annotations 📚 JQL Multilingual Educational Quality Annotations This dataset provides high-quality human annotations for evaluating the educational value of web documents, and serves as a benchmark for training and evaluating multilingual LLM annotators as described in the JQL paper. 📝 Dataset Summary Documents: 511 English texts Annotations: 3 human ratings per document (0–5 scale) Translations: Into 35 European languages using DeepL and GPT-4o Purpose: For training and… See the full description on the dataset page: https://huggingface.co/datasets/JQL-AI/JQL-Human-Edu-Annotations.texttext-classification10K<n<100K5 likes144 downloads1y agoHugging Face15ltg /normistral-fluency-annotationManual fluency annotations for Fluent Alignment with Disfluent Judges: Post-training for Lower-resource Languages Citation @misc{samuel2025fluentalignmentdisfluentjudges, title={Fluent Alignment with Disfluent Judges: Post-training for Lower-resource Languages}, author={David Samuel and Lilja Øvrelid and Erik Velldal and Andrey Kutuzov}, year={2025}, eprint={2512.08777}, archivePrefix={arXiv}, primaryClass={cs.CL}… See the full description on the dataset page: https://huggingface.co/datasets/ltg/normistral-fluency-annotation.textn<1K0 likes128 downloads10mo agoHugging Face16aiobservatory /annotations The AI Observatory A public measurement platform aggregating real-world AI conversations from seven sources under a unified 145-feature taxonomy. This dataset card accompanies paper The AI Observatory: A Public Measure of Real-World AI Use. 📄 Paper: [anonymous OpenReview link] 📊 Dashboard: https://project-ai-observatory.vercel.app/ 💾 Anonymous code: https://anonymous.4open.science/r/ai-observatory/README.md TL;DR 23,158 conversations, 85,633 turns, ~5,000… See the full description on the dataset page: https://huggingface.co/datasets/aiobservatory/annotations.tabulartext-classification100K<n<1M3 likes119 downloads2mo agoHugging Face17nikhilchandak /gpqa-diamond-annotations GPQA Diamond Dataset This dataset contains filtered JSONL files of human annotations on question specificity, answer uniqueness, answer matching to the ground truth for different models for the GPQA Diamond dataset. The dataset was annotated by two human graders. It contains 198 (original size) * 2 = 396 rows as each rows is repeated twice (one for each human). A human grader given the question, actual answer and model response, has to answer whether the response matches the… See the full description on the dataset page: https://huggingface.co/datasets/nikhilchandak/gpqa-diamond-annotations.tabularn<1K1 likes112 downloads1y agoHugging Face18netprtony /pokemon-cards-image-and-annotationsimage1K<n<10K0 likes112 downloads1y agoHugging Face19BAAI /CCI3-HQ-Annotation-Benchmark CCI3-HQ-Annotation-Benchmark These 14k samples were randomly extracted from a large corpus of Chinese texts, containing both the original text and corresponding labels. They can be used to evaluate the quality of Chinese corpora. Citation If you use this benchmark or the CCI3-HQ dataset, please cite: @misc{wang2024cci30hqlargescalechinesedataset, title={CCI3.0-HQ: a large-scale Chinese dataset of high quality designed for pre-training large language models}… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/CCI3-HQ-Annotation-Benchmark.document10K<n<100K4 likes101 downloads2mo agoHugging Face20bluolightning /manga109s-line-annotations Manga109-s Text Line Annotations High-precision, line-level bounding box and polygon annotations for the Manga109-s Dataset, supporting both full manga pages and speech bubble crops. Furigana is not labeled and is almost entirely excluded from line labels. Includes 8-point oriented polygons for slanted/rotated text lines. The annotation process is documented in METHODOLOGY.md (WIP). Notice: This dataset contains zero dialogue text and zero images. It requires your own local… See the full description on the dataset page: https://huggingface.co/datasets/bluolightning/manga109s-line-annotations.textimage-text-to-textn<1K1 likes94 downloads22d agoHugging Face21PrentisAI /ScreenRef-Annotations ScreenRef — Annotations Annotations only. This repository contains no screenshots. It holds every annotation of ScreenRef (484 MB); the screenshots are in the gated repository PrentisAI/ScreenRef. The annotations are licensed under the same LICENSE as the screenshots: Academic Research only, no redistribution. Contents Path What it is Size anno/ 58 JSONL files — all tasks (the full recipe), 633,369 rows 240 MB anno_point/ 46 JSONL files — same screens… See the full description on the dataset page: https://huggingface.co/datasets/PrentisAI/ScreenRef-Annotations.tabularimage-text-to-text1M<n<10M0 likes90 downloads11d agoHugging Face22BUT-FIT /orca-audio-qa-annotations ORCA Audio QA Annotations Annotation data for training and evaluating ORCA (Open-ended Response Correctness Assessment), a scoring model for audio question-answering tasks. Paper: ORCA: Open-ended Response Correctness Assessment for Audio Question Answering — accepted to TACL 2026 Code & usage: github.com/BUTSpeechFIT/ORCA Pretrained Models: orca-olmo-2-1b-multinomial orca-gemma-3-4b-it-multinomial orca-llama-3.2-3b-it-multinomial Dataset overview ORCA is… See the full description on the dataset page: https://huggingface.co/datasets/BUT-FIT/orca-audio-qa-annotations.texttext-classification100K<n<1M0 likes88 downloads3mo agoHugging Face23MongoDB /wikipedia-22-12-en-annotationtabular10K<n<100K0 likes78 downloads2y agoHugging Face24jamesding0302 /memgen-annotations MemGen Annotations This is the annotation dataset for the paper How Well Does Generative Recommendation Generalize?. The annotations categorize evaluation instances under the leave-one-out protocol: test split uses the last item in the user history sequence as target, val split uses the second-to-last item as target. Columns sample_id: row index within the split in the original dataset. user_id: raw user identifier (join key). master: one of memorization… See the full description on the dataset page: https://huggingface.co/datasets/jamesding0302/memgen-annotations.textother100K<n<1M1 likes76 downloads7mo agoHugging Face25jkminder /apertus-annotation-feedbacktextn<1K0 likes74 downloads4mo agoHugging Face26rafmacalaba /data-use-annotations Data-use annotations Public store of keep/drop rulings from the annotation review app (human_labeling/review.html). Files rulings/<annotator>.jsonl — one file per annotator, one JSON object per ruling: key (span UID), ruling (DATA_MENTION keep / NON_MENTION drop), queue (gold / sample), annotator (required, set in the UI), ts. Last write per (queue, key, annotator) wins. from datasets import load_dataset ds = load_dataset("rafmacalaba/data-use-annotations") #… See the full description on the dataset page: https://huggingface.co/datasets/rafmacalaba/data-use-annotations.texttext-classificationn<1K0 likes65 downloads1mo agoHugging Face27Hannibal52Barca /icl-sarm-annotations ICL SARM Subtask Annotations Per-episode subtask decomposition (name + start/end frame) generated with a VLM-based annotation pipeline (Qwen3-VL-8B-Instruct, ecot-style plan generation + bidirectional grounding), for the two adityx23 ICL robot manipulation datasets: File Source dataset Episodes Tasks icl-dataset_subtasks.jsonl adityx23/icl-dataset (reference set) 3149 36 icl-demo-dataset_subtasks.jsonl adityx23/icl-demo-dataset (query/test set) 285 27… See the full description on the dataset page: https://huggingface.co/datasets/Hannibal52Barca/icl-sarm-annotations.textrobotics1K<n<10K0 likes65 downloads27d agoHugging Face28valira-ai /objaverse-xl-shape-annotations objaverse-xl-shape-annotations Shape-based textual annotations for 537,841 objects from Objaverse-XL. Each object gets a class label and a short, geometry-focused description. Why does this exist? Objaverse-XL is a large benchmark, but it does not contain any textual descriptions. This dataset was built to fix that. Every description focuses strictly on shape and structure, making it suitable for text-to-3D retrieval and contrastive representation learning tasks… See the full description on the dataset page: https://huggingface.co/datasets/valira-ai/objaverse-xl-shape-annotations.texttext-to-3d100K<n<1M1 likes48 downloads4mo agoHugging Face29gatilin /SenseNova-Vision-Corpus-50M-annotationtext10M<n<100M0 likes48 downloads3mo agoHugging Face30toroe /Dolci-Think-SFT-7B-Propella-Annotationstabular1M<n<10M0 likes44 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.