Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01openeurollm /propella-annotations This dataset contains document annotations produced with propella-1-4b, a small multilingual LLM that annotates text documents across six categories: core content, classification, quality & value, audience & purpose, safety & compliance, and geographic relevance. The annotations can be used to filter, select, and curate LLM training data at scale. Properties Each document is annotated across 18 properties organized into six categories: Category Property Description… See the full description on the dataset page: https://huggingface.co/datasets/openeurollm/propella-annotations.text1B<n<10B21 likes7.3k downloads2mo agoHugging Face02ljvmiranda921 /gsd-humaneval-annotationstext1K<n<10K1 likes3.3k downloads15d agoHugging Face03tanish434 /Truebones-ZOO-Annotations Truebones ZOO Annotations Text prompts, per-clip metadata, rest-pose renders and the exact build pipeline for Truebones ZOO — 1,097 animal motion clips across 74 skeletons: mammals, birds, reptiles, insects, marine and prehistoric creatures. 1.02 hours, 111,158 frames, uniformly 30 fps. Rigs range from 9 to 143 joints; clips from 0.3 to 18.5 seconds. The motion files themselves are not in this repository. Truebones ZOO is a commercial library by Truebones Motions Animation… See the full description on the dataset page: https://huggingface.co/datasets/tanish434/Truebones-ZOO-Annotations.tabular1K<n<10K0 likes3.2k downloads27d agoHugging Face04m-hamza-mughal /beat2-additional-annotations BEAT2 Official Release + Additional Annotations This is a fork of H-Liu1997/BEAT2 that adds annotations contributed by the RAG-Gesture (CVPR 2025) and MIBURI (CVPR 2026) projects. The base BEAT2-English data (motion, audio, TextGrids, semantic labels, pretrained motion-autoencoder weights) is inherited verbatim from upstream; the additional annotations from RAG-Gesture and MIBURI are pushed on top. Citations If you use only the original BEAT2 dataset, please cite… See the full description on the dataset page: https://huggingface.co/datasets/m-hamza-mughal/beat2-additional-annotations.audio1K<n<10K0 likes2.8k downloads4mo agoHugging Face05handshake-ai-research /cue-annotations CUE annotations 533,001 persona-manual annotations over 28 public dialogue corpora, one config per corpus, split train / validation. Each row describes how the user in one conversation behaves. This was used as a training dataset for the CUE user simulator model. No dialogue text is redistributed Rows carry provenance and a hash, not the source turns: column meaning persona_manual the annotation, a JSON string (json.loads it) source_repo… See the full description on the dataset page: https://huggingface.co/datasets/handshake-ai-research/cue-annotations.texttext-generation100K<n<1M0 likes1.4k downloads5d agoHugging Face06JQL-AI /JQL-LLM-Edu-Annotations 📚 JQL Educational Quality Annotations from LLMs This dataset provides 17,186,606 documents with high-quality LLM annotations for evaluating the educational value of web documents, and serves as a benchmark for training and evaluating multilingual LLM annotators as described in the JQL paper. 📝 Dataset Summary Multilingual document-level quality annotations scored on a 0–5 educational value scale by three state-of-the-art LLMs: Gemma-3-27B-it, Mistral-3.1-24B-it… See the full description on the dataset page: https://huggingface.co/datasets/JQL-AI/JQL-LLM-Edu-Annotations.tabular10M<n<100M2 likes1.3k downloads1y agoHugging Face07laion /Emilia-with-Emotion-Annotations2audio10M<n<100M1 likes1.2k downloads1y agoHugging Face08laion /Emilia-with-Emotion-Annotations4audio10M<n<100M1 likes841 downloads1y agoHugging Face09HuggingFaceTB /python-edu-annotations Annotations for 📚 Python-Edu classifier This dataset contains the annotations used for training Python-Edu educational quality classifier. We prompt Llama-3-70B-Instruct to score python programs from StarCoderData based on their educational value. Note: the dataset contains the Python program, the prompt (using the first 1000 characters of the program) and the scores but it doesn't contain the full Llama 3 generation. text100K<n<1M2 likes774 downloads2y agoHugging Face10ce-amtic /ProcVQA-20M-annotationsgated ProcVQA-20M Annotations Project Page | arXiv | Code | Model | Media This repository contains the text annotations for the ProcVQA-20M dataset. The full image files are hosted separately on ProcVQA-20M-media. Overview This dataset is constructed from over 26 embodied datasets, comprising: 20M QA pairs for training 330K original trajectories 50M annotated frames from ~5,000 hours of manipulation data 200+ different tasks Dataset Structure The… See the full description on the dataset page: https://huggingface.co/datasets/ce-amtic/ProcVQA-20M-annotations.image10K<n<100K1 likes764 downloads5mo agoHugging Face11it-just-works /vast27m_annotations VAST-27M Annotations Dataset This dataset contains annotations from the VAST-27M dataset, originally created for the paper "VAST: A Vision-Audio-Subtitle-Text Omni-Modality Foundation Model and Dataset" by Chen et al. (2024). Original Source This dataset is derived from the VAST-27M dataset, which was created by researchers at the University of Chinese Academy of Sciences and the Institute of Automation, Chinese Academy of Science. The original dataset and more… See the full description on the dataset page: https://huggingface.co/datasets/it-just-works/vast27m_annotations.tabular10M<n<100M1 likes697 downloads2y agoHugging Face12HuggingFaceFW /fineweb-edu-llama3-annotations Annotations for 📚 FineWeb-Edu classifier This dataset contains the annotations used for training 📚 FineWeb-Edu educational quality classifier. We prompt Llama-3-70B-Instruct to score web pages from 🍷 FineWeb based on their educational value. Note: the dataset contains the FineWeb text sample, the prompt (using the first 1000 characters of the text sample) and the scores but it doesn't contain the full Llama 3 generation. text100K<n<1M50 likes690 downloads2y agoHugging Face13laion /Emilia-with-Emotion-Annotations3audio10M<n<100M1 likes588 downloads1y agoHugging Face14Linzhan /Truebones-ZOO-Annotations Truebones ZOO Annotations Text prompts, per-clip metadata, rest-pose renders and the exact build pipeline for Truebones ZOO — 1,097 animal motion clips across 74 skeletons: mammals, birds, reptiles, insects, marine and prehistoric creatures. 1.02 hours, 111,158 frames, uniformly 30 fps. Rigs range from 9 to 143 joints; clips from 0.3 to 18.5 seconds. The motion files themselves are not in this repository. Truebones ZOO is a commercial library by Truebones Motions Animation… See the full description on the dataset page: https://huggingface.co/datasets/Linzhan/Truebones-ZOO-Annotations.tabular1K<n<10K3 likes429 downloads1mo agoHugging Face15laion /Emilia-with-Emotion-Annotations5audio10M<n<100M3 likes404 downloads1y agoHugging Face16ncbi /TrialGPT-Criterion-Annotationstext1K<n<10K7 likes346 downloads1mo agoHugging Face17Bofeee5675 /GUI-Net-1M-relative-annotationstext100K<n<1M0 likes335 downloads1y agoHugging Face18YsK-dev /TurkWeb-Edu-AnnotationsV3 TurkWeb-Edu V3 Model: Qwen/Qwen3-30B-A3B-Instruct-2507 Format: Structured JSON (vLLM 0.15.0) tabular100K<n<1M0 likes322 downloads8mo agoHugging Face19nvidia /SEED-Timeline-Annotations Timeline Annotations for BONES-SEED Humanoid Motion Dataset Dataset Description: This dataset provides additional text description annotations from the BONES-SEED humanoid motion dataset. For each motion, this dataset provides an overview text description of the entire motion at a high level, along with a “timeline” of annotated segments within the motion. Each segment generally contains a single atomic action and is defined by a start time, end time, and text… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/SEED-Timeline-Annotations.text100K<n<1M7 likes291 downloads6mo agoHugging Face20ambient-intelligence-labs /egoproactive-synth-annotations EgoProactive synthetic proactive annotations Everything produced by the annotation and synthesis pipelines for the AI Wearables Challenge 2026 EgoProactive Dense timestamped proactive walkthroughs generated with the ambient agent (orchestrator deepseek/deepseek-v4-flash-0731 + vision Qwen3.6-27B), using the held-out-validated dense policy (setup-phase coverage, repetition-collapse, fire-at-onset) and a -0.5s onset correction at chunk-binning. set clips median events/clip… See the full description on the dataset page: https://huggingface.co/datasets/ambient-intelligence-labs/egoproactive-synth-annotations.textvideo-text-to-textn<1K0 likes277 downloads1mo agoHugging Face21obalcells /longfact-annotationstext1K<n<10K2 likes273 downloads1y agoHugging Face22Bofeee5675 /GUI-Net-1M-absolute-annotationstext100K<n<1M2 likes268 downloads1y agoHugging Face23edesaras /CEFR-Sentence-Level-Annotations Dataset Card for Dataset Name 17k english sentences annotated by english education professionals. Original repo for CEFR-SP is located at this repo This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP):… See the full description on the dataset page: https://huggingface.co/datasets/edesaras/CEFR-Sentence-Level-Annotations.tabulartext-classification10K<n<100K6 likes260 downloads2y agoHugging Face24houlab /cossmos-annotations-db ARSMA-web motif annotations (cossmos-annotations-db) Per-motif annotation: base pairs, stacking, sugar puckers, glycosidic conformations and the deposition metadata of the parent entry. It backs the motif browser of ARSMA-web. This is not the occurrence count. One row of instances.parquet is one annotated motif site, and the table has 283,164 of them against 285,165 clips in houlab/motif-db. The two differ because the annotation rows of the 25 CoSSMos classes are CoSSMos's own… See the full description on the dataset page: https://huggingface.co/datasets/houlab/cossmos-annotations-db.tabular1M<n<10M0 likes257 downloads8d agoHugging Face25laion /emolia-voicenet-gemini-annotations Emolia VoiceNet Gemini Annotations 468,180 dimension-level annotations over 236,613 Emolia speech clips, each scored 0-6 (0-2 for the content-safety dimension) on one of 57 perceptual voice / speech dimensions - arousal, valence, brightness, resonance placement, speaking styles, genuineness, recording quality, and more - by Gemini 3.5 Flash (non-thinking, temperature 0). This repository ships the annotations, audio provenance, per-dimension statistics, and the full scoring… See the full description on the dataset page: https://huggingface.co/datasets/laion/emolia-voicenet-gemini-annotations.tabularaudio-classification100K<n<1M0 likes230 downloads3mo agoHugging Face26ambient-intelligence-labs /egolongqa-synth-annotations EgoLongQA synthetic MCQs, teacher traces and annotation outputs Everything produced by the annotation and synthesis pipelines for the AI Wearables Challenge 2026 EgoLongQA ≤2B track, other than the distillation set (which lives in infinitylogesh/egolongqa-junior-distill). ⚠️ Read this before counting rows The synthetic set is 943 questions over 408 videos, and it is stored two ways: file rows shape training_sets/train_synth_v3.jsonl 943 flat — one row… See the full description on the dataset page: https://huggingface.co/datasets/ambient-intelligence-labs/egolongqa-synth-annotations.tabularvisual-question-answering1K<n<10K0 likes213 downloads1mo agoHugging Face27tvonarx /emboss-roof-annotations Emboss 3D Roof Reference Annotations Manual 3D reference meshes and editable annotations for Swiss and Brazilian buildings, prepared for the evaluation and parameter tuning of Emboss. The annotations describe building and roof geometry, including roof superstructures. Emboss source code 3dlabel annotation tool Emboss segmentation model Example reference annotation in 3dlabel: annotated mesh and LiDAR points (Figure D.1(a) in the paper). 3dn<1K0 likes213 downloads27d agoHugging Face28obalcells /longfact-augmented-annotationstext10K<n<100K0 likes203 downloads1y agoHugging Face29touati-kamel /forest-fire-annotations Forest Fire Detection Dataset — Auto-Annotated Bounding-box annotated version of touati-kamel/forest-fire-dataset, built for training forest-fire / smoke / fog object detection models. Overview This dataset contains video frames auto-labeled with bounding boxes for fire and smoke-related visual phenomena, using a zero-shot open-vocabulary object detector (Grounding DINO). It is derived from the original touati-kamel/forest-fire-dataset image classification dataset… See the full description on the dataset page: https://huggingface.co/datasets/touati-kamel/forest-fire-annotations.image10K<n<100K0 likes191 downloads2mo agoHugging Face30Himpq /kwext-bilibili-video-title-annotations KwExt Bilibili Video Title Annotations This dataset is a model-assisted annotation set for the KwExt keyword extraction project. The current snapshot contains 5,000 Chinese Bilibili video titles from annotation stages video_title_zh_001 through video_title_zh_005, with 1,000 records in each stage. The release is intended for early experiments with: extracting title-grounded keywords and ranking their importance; broad semantic tags for retrieval and RAG metadata; dense tag… See the full description on the dataset page: https://huggingface.co/datasets/Himpq/kwext-bilibili-video-title-annotations.tabulartoken-classification1K<n<10K0 likes183 downloads23d agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.