Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01happy8825 /co3d-annotations CO3D Annotations (Derived) This dataset is derived from the CO3D Dataset. For Test Set ONLY, Preprocessed for VGGT input License Original License: Creative Commons Attribution-NonCommercial 4.0 (CC BY-NC 4.0) Non-commercial use only. Attribution required: This dataset is derived from the CO3D dataset © Meta, licensed under CC BY-NC 4.0. 3dother0 likes44k downloads1y agoHugging Face02commoncrawl /gneissweb-annotation-url-testing-v1 GneissWeb Annotations GneissWeb Annotations, powered by IBM Research's GneissWeb methodology, is a dataset of quality and category annotations applied to the Common Crawl corpus. This dataset enables precise filtering of web content across medical, educational, technology, and scientific domains, making it easier to build high-quality corpora for research projects, language models, and specialized applications. Learn more about the annotation process and methodology in our… See the full description on the dataset page: https://huggingface.co/datasets/commoncrawl/gneissweb-annotation-url-testing-v1.tabular10B<n<100B0 likes8k downloads10mo agoHugging Face03openeurollm /propella-annotations This dataset contains document annotations produced with propella-1-4b, a small multilingual LLM that annotates text documents across six categories: core content, classification, quality & value, audience & purpose, safety & compliance, and geographic relevance. The annotations can be used to filter, select, and curate LLM training data at scale. Properties Each document is annotated across 18 properties organized into six categories: Category Property Description… See the full description on the dataset page: https://huggingface.co/datasets/openeurollm/propella-annotations.text1B<n<10B21 likes7.3k downloads2mo agoHugging Face04ljvmiranda921 /gsd-humaneval-annotationstext1K<n<10K1 likes3.6k downloads12d agoHugging Face05m-hamza-mughal /beat2-additional-annotations BEAT2 Official Release + Additional Annotations This is a fork of H-Liu1997/BEAT2 that adds annotations contributed by the RAG-Gesture (CVPR 2025) and MIBURI (CVPR 2026) projects. The base BEAT2-English data (motion, audio, TextGrids, semantic labels, pretrained motion-autoencoder weights) is inherited verbatim from upstream; the additional annotations from RAG-Gesture and MIBURI are pushed on top. Citations If you use only the original BEAT2 dataset, please cite… See the full description on the dataset page: https://huggingface.co/datasets/m-hamza-mughal/beat2-additional-annotations.audio1K<n<10K0 likes3.5k downloads4mo agoHugging Face06tanish434 /Truebones-ZOO-Annotations Truebones ZOO Annotations Text prompts, per-clip metadata, rest-pose renders and the exact build pipeline for Truebones ZOO — 1,097 animal motion clips across 74 skeletons: mammals, birds, reptiles, insects, marine and prehistoric creatures. 1.02 hours, 111,158 frames, uniformly 30 fps. Rigs range from 9 to 143 joints; clips from 0.3 to 18.5 seconds. The motion files themselves are not in this repository. Truebones ZOO is a commercial library by Truebones Motions Animation… See the full description on the dataset page: https://huggingface.co/datasets/tanish434/Truebones-ZOO-Annotations.tabular1K<n<10K0 likes3.2k downloads24d agoHugging Face07laion /laions_got_talent_enhanced_just_flash_annotations0 likes2.4k downloads2y agoHugging Face08mlfoundations /dcvlm_pool_small_annotations DCVLM-Pool (small) — per-sample annotations Every filtering annotation we computed for the small data pool of our DataComp-VLM paper: image quality, image–text alignment, language ID, text-quality classifiers, multimodal perplexity, decontamination scores and more — up to 167 fields per sample (180 distinct fields overall), for all 120,940,134 samples across 166 source datasets. These are the raw annotations, not a filtered dataset. They are the inputs our curation pipeline… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/dcvlm_pool_small_annotations.image-text-to-text100M<n<1B0 likes2.3k downloads2mo agoHugging Face09laion /Emilia-with-Emotion-Annotations Dataset Card for Emilia with Emotion Annotations Dataset Description This dataset is an enhanced version of the Emilia dataset, enriched with detailed emotion annotations. The annotations were generated using models from the EmoNet suite to provide deeper insight into the emotional content of speech. This work is based on the research and models described in the blog post "Do They See What We See?". The annotations include 54 scores for each sample, covering a wide range… See the full description on the dataset page: https://huggingface.co/datasets/laion/Emilia-with-Emotion-Annotations.29 likes2.3k downloads1y agoHugging Face10LumiOpen /hpltv2-llama33-edu-annotation HPLT version 2.0 educational annotations This dataset contains annotations derived from HPLT v2 cleaned samples. There are 500,000 annotations for each language if the source contains at least 500,000 samples. We prompt Llama-3.3-70B-Instruct to score web pages based on their educational value following FineWeb-Edu classifier. Note 1: The dataset contains the prompt (using the first 1500 characters of the text sample), the scores, and the full Llama 3 generation. The column "idx"… See the full description on the dataset page: https://huggingface.co/datasets/LumiOpen/hpltv2-llama33-edu-annotation.tabular10M<n<100M3 likes1.9k downloads1y agoHugging Face11mitermix /audiosnippets_small_with_detailed_annotationaudio100K<n<1M1 likes1.8k downloads2y agoHugging Face12ciscoriordan /open-greek-corpus-annotations Open Greek Corpus Annotations Token-level linguistic annotations for the Open Greek Corpus: lemma, part of speech (UD UPOS), and morphology (UD features) for every served token. Three provenance classes, never confused thanks to per-token provenance and confidence tiers: gold treebank annotations where an openly licensed MANUAL treebank covers a work (GLAUx's treebank layers, MACULA Greek for the NT), GLAUx's own automatic annotation as the middle auto: class, and model… See the full description on the dataset page: https://huggingface.co/datasets/ciscoriordan/open-greek-corpus-annotations.token-classification0 likes1.7k downloads3mo agoHugging Face13mitermix /audiosnippets_small_with_detailed_annotation2audio1M<n<10M1 likes1.6k downloads2y agoHugging Face14yuanbopang /urap26-annotation-videos URAP26 Annotation Videos Public 1080p RGB proxies paired with YOLOMG compensated-difference videos for browser annotation. videoobject-detectionn<1K0 likes1.6k downloads2mo agoHugging Face15JQL-AI /JQL-LLM-Edu-Annotations 📚 JQL Educational Quality Annotations from LLMs This dataset provides 17,186,606 documents with high-quality LLM annotations for evaluating the educational value of web documents, and serves as a benchmark for training and evaluating multilingual LLM annotators as described in the JQL paper. 📝 Dataset Summary Multilingual document-level quality annotations scored on a 0–5 educational value scale by three state-of-the-art LLMs: Gemma-3-27B-it, Mistral-3.1-24B-it… See the full description on the dataset page: https://huggingface.co/datasets/JQL-AI/JQL-LLM-Edu-Annotations.tabular10M<n<100M2 likes1.5k downloads1y agoHugging Face16laion /laions_got_talent_enhanced_flash_annotations_and_long_captions18 likes1.4k downloads2y agoHugging Face17MaxwellmuF /Soofi-sft-annotation-filteredtext1M<n<10M0 likes1.1k downloads6mo agoHugging Face18VR-VLA /VR-egodex-annotation-converted-v6.0 VR-egodex-annotation-converted-v6.0 EgoDex converted from LeRobot v2.1 into the Layer-1 v0.6.0 annotation schema, with per-clip narration included as language sidecars. 314,839 clips · 78,282,306 frames · 724.8 hours @ 30 fps · 129 tasks 100% narration coverage (1 sidecar per clip) 71 GB annotations + 2.3 GB narratives Videos are NOT included. This release contains annotations and narration only. Source video lives in griffinlabs/EgoDex-LeRobot-v3.0; orig_id in the manifest… See the full description on the dataset page: https://huggingface.co/datasets/VR-VLA/VR-egodex-annotation-converted-v6.0.tabularrobotics10M<n<100M0 likes925 downloads25d agoHugging Face19laion /Emilia-with-Emotion-Annotations4audio10M<n<100M1 likes841 downloads1y agoHugging Face20contralabs /video-detail-annotation Video Detail Annotation — Contra Labs A human-annotated evaluation dataset comparing AI-generated product videos across three leading video generation models: Google Veo 3.1, Adobe Firefly Video, and Grok Imagine (xAI). Annotations were collected by professional video editors sourced from the Contra network, a platform connecting top independent professionals across creative and technical fields. Annotators reviewed each video and left timestamped, dimension-specific comments on… See the full description on the dataset page: https://huggingface.co/datasets/contralabs/video-detail-annotation.videovideo-classificationn<1K3 likes802 downloads3mo agoHugging Face21it-just-works /vast27m_annotations VAST-27M Annotations Dataset This dataset contains annotations from the VAST-27M dataset, originally created for the paper "VAST: A Vision-Audio-Subtitle-Text Omni-Modality Foundation Model and Dataset" by Chen et al. (2024). Original Source This dataset is derived from the VAST-27M dataset, which was created by researchers at the University of Chinese Academy of Sciences and the Institute of Automation, Chinese Academy of Science. The original dataset and more… See the full description on the dataset page: https://huggingface.co/datasets/it-just-works/vast27m_annotations.tabular10M<n<100M1 likes798 downloads2y agoHugging Face22HuggingFaceTB /python-edu-annotations Annotations for 📚 Python-Edu classifier This dataset contains the annotations used for training Python-Edu educational quality classifier. We prompt Llama-3-70B-Instruct to score python programs from StarCoderData based on their educational value. Note: the dataset contains the Python program, the prompt (using the first 1000 characters of the program) and the scores but it doesn't contain the full Llama 3 generation. text100K<n<1M2 likes781 downloads2y agoHugging Face23ce-amtic /ProcVQA-20M-annotationsgated ProcVQA-20M Annotations Project Page | arXiv | Code | Model | Media This repository contains the text annotations for the ProcVQA-20M dataset. The full image files are hosted separately on ProcVQA-20M-media. Overview This dataset is constructed from over 26 embodied datasets, comprising: 20M QA pairs for training 330K original trajectories 50M annotated frames from ~5,000 hours of manipulation data 200+ different tasks Dataset Structure The… See the full description on the dataset page: https://huggingface.co/datasets/ce-amtic/ProcVQA-20M-annotations.image10K<n<100K1 likes764 downloads5mo agoHugging Face24P-Kevin /hssd-annotations hssd-annotations English | 中文 A standalone, locally-stored, zero-dependency Python API to search and retrieve HSSD assets and their full annotation set — for downstream scene generation (e.g. SceneSmith) and articulation/clearance research. Every annotation family is merged into one per-asset record keyed by the HSSD asset id. The library ships in post-replacement form: each asset is linked to its articulated realization (official HSSD articulated, or a PartNet-Mobility… See the full description on the dataset page: https://huggingface.co/datasets/P-Kevin/hssd-annotations.1 likes752 downloads1mo agoHugging Face25HuggingFaceFW /fineweb-edu-llama3-annotations Annotations for 📚 FineWeb-Edu classifier This dataset contains the annotations used for training 📚 FineWeb-Edu educational quality classifier. We prompt Llama-3-70B-Instruct to score web pages from 🍷 FineWeb based on their educational value. Note: the dataset contains the FineWeb text sample, the prompt (using the first 1000 characters of the text sample) and the scores but it doesn't contain the full Llama 3 generation. text100K<n<1M50 likes645 downloads2y agoHugging Face26chfeng /probe-tip-annotations-data0 likes608 downloads6mo agoHugging Face27MCG-NJU /VideoChat3-Training-Data-Annotations VideoChat3-Stage3-Training-Data This repository includes all annotation files used across the four training stages of VideoChat3, from Stage 0 to Stage 3. You can refer to the provided source-data links to download videos, images, and other multimedia data for training. In videochat3_data_annotations, we also provide a source field to indicate the source dataset for each entry. To facilitate Stage 3 training reproduction using the high-quality open-source datasets we collected… See the full description on the dataset page: https://huggingface.co/datasets/MCG-NJU/VideoChat3-Training-Data-Annotations.video-text-to-text2 likes591 downloads5d agoHugging Face28laion /Emilia-with-Emotion-Annotations2audio10M<n<100M1 likes586 downloads1y agoHugging Face29handshake-ai-research /cue-annotations CUE annotations 533,001 persona-manual annotations over 28 public dialogue corpora, one config per corpus, split train / validation. Each row describes how the user in one conversation behaves. This was used as a training dataset for the CUE user simulator model. No dialogue text is redistributed Rows carry provenance and a hash, not the source turns: column meaning persona_manual the annotation, a JSON string (json.loads it) source_repo… See the full description on the dataset page: https://huggingface.co/datasets/handshake-ai-research/cue-annotations.texttext-generation100K<n<1M0 likes580 downloads2d agoHugging Face30titoruizh /Drone-Orthomosaic-Vehicles-Yolo-annotation Dataset Tailings Mining Vehicles & Instruments (High-Res Drone Imagery) Dataset Summary This dataset contains high-resolution aerial imagery focused on vehicle detection and geotechnical monitoring instruments within active mining environments (tailings dams). The data was acquired using a DJI Zenmuse P1 sensor at 120m altitude. Photogrammetric Context The images originate from large-scale georeferenced orthomosaics generated from bi-daily… See the full description on the dataset page: https://huggingface.co/datasets/titoruizh/Drone-Orthomosaic-Vehicles-Yolo-annotation.imageobject-detection1K<n<10K3 likes551 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.