datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gneissweb-annotation-url-testing-v1
GneissWeb Annotations
GneissWeb Annotations, powered by IBM Research's GneissWeb methodology, is a dataset of quality and category annotations applied to the Common Crawl corpus.
This dataset enables precise filtering of web content across medical, educational, technology, and scientific domains, making it easier to build high-quality corpora for research projects, language models, and specialized applications.
Learn more about the annotation process and methodology in our… See the full description on the dataset page: https://huggingface.co/datasets/commoncrawl/gneissweb-annotation-url-testing-v1.Truebones-ZOO-Annotations
Truebones ZOO Annotations
Text prompts, per-clip metadata, rest-pose renders and the exact build pipeline for
Truebones ZOO — 1,097 animal motion clips across 74 skeletons: mammals, birds,
reptiles, insects, marine and prehistoric creatures. 1.02 hours, 111,158 frames, uniformly
30 fps. Rigs range from 9 to 143 joints; clips from 0.3 to 18.5 seconds.
The motion files themselves are not in this repository. Truebones ZOO is a commercial
library by Truebones Motions Animation… See the full description on the dataset page: https://huggingface.co/datasets/tanish434/Truebones-ZOO-Annotations.hpltv2-llama33-edu-annotation
HPLT version 2.0 educational annotations
This dataset contains annotations derived from HPLT v2 cleaned samples.
There are 500,000 annotations for each language if the source contains at least 500,000 samples.
We prompt Llama-3.3-70B-Instruct to score web pages based on their educational value following FineWeb-Edu classifier.
Note 1: The dataset contains the prompt (using the first 1500 characters of the text sample), the scores, and the full Llama 3 generation. The column "idx"… See the full description on the dataset page: https://huggingface.co/datasets/LumiOpen/hpltv2-llama33-edu-annotation.JQL-LLM-Edu-Annotations
📚 JQL Educational Quality Annotations from LLMs
This dataset provides 17,186,606 documents with high-quality LLM annotations for evaluating the educational value of web documents, and serves as a benchmark for training and evaluating multilingual LLM annotators as described in the JQL paper.
📝 Dataset Summary
Multilingual document-level quality annotations scored on a 0–5 educational value scale by three state-of-the-art LLMs:
Gemma-3-27B-it, Mistral-3.1-24B-it… See the full description on the dataset page: https://huggingface.co/datasets/JQL-AI/JQL-LLM-Edu-Annotations.VR-egodex-annotation-converted-v6.0
VR-egodex-annotation-converted-v6.0
EgoDex converted from LeRobot v2.1 into the Layer-1 v0.6.0 annotation schema, with
per-clip narration included as language sidecars.
314,839 clips · 78,282,306 frames · 724.8 hours @ 30 fps · 129 tasks
100% narration coverage (1 sidecar per clip)
71 GB annotations + 2.3 GB narratives
Videos are NOT included. This release contains annotations and narration only. Source
video lives in griffinlabs/EgoDex-LeRobot-v3.0;
orig_id in the manifest… See the full description on the dataset page: https://huggingface.co/datasets/VR-VLA/VR-egodex-annotation-converted-v6.0.vast27m_annotations
VAST-27M Annotations Dataset
This dataset contains annotations from the VAST-27M dataset, originally created for the paper "VAST: A Vision-Audio-Subtitle-Text Omni-Modality Foundation Model and Dataset" by Chen et al. (2024).
Original Source
This dataset is derived from the VAST-27M dataset, which was created by researchers at the University of Chinese Academy of Sciences and the Institute of Automation, Chinese Academy of Science. The original dataset and more… See the full description on the dataset page: https://huggingface.co/datasets/it-just-works/vast27m_annotations.Truebones-ZOO-Annotations
Truebones ZOO Annotations
Text prompts, per-clip metadata, rest-pose renders and the exact build pipeline for
Truebones ZOO — 1,097 animal motion clips across 74 skeletons: mammals, birds,
reptiles, insects, marine and prehistoric creatures. 1.02 hours, 111,158 frames, uniformly
30 fps. Rigs range from 9 to 143 joints; clips from 0.3 to 18.5 seconds.
The motion files themselves are not in this repository. Truebones ZOO is a commercial
library by Truebones Motions Animation… See the full description on the dataset page: https://huggingface.co/datasets/Linzhan/Truebones-ZOO-Annotations.gneissweb-annotation-host-testing-v1
GneissWeb Annotations
GneissWeb Annotations, powered by IBM Research's GneissWeb methodology, is a dataset of quality and category annotations applied to the Common Crawl corpus.
This dataset enables precise filtering of web content across medical, educational, technology, and scientific domains, making it easier to build high-quality corpora for research projects, language models, and specialized applications.
Learn more about the annotation process and methodology in our… See the full description on the dataset page: https://huggingface.co/datasets/commoncrawl/gneissweb-annotation-host-testing-v1.highlevel_thinking_with_grounding_annotation_split1000_v3_merged_promptsTurkWeb-Edu-AnnotationsV3
TurkWeb-Edu V3
Model: Qwen/Qwen3-30B-A3B-Instruct-2507
Format: Structured JSON (vLLM 0.15.0)
cossmos-annotations-db
ARSMA-web motif annotations (cossmos-annotations-db)
Per-motif annotation: base pairs, stacking, sugar puckers, glycosidic
conformations and the deposition metadata of the parent entry. It backs the
motif browser of ARSMA-web.
This is not the occurrence count. One row of instances.parquet is one
annotated motif site, and the table has 283,164 of them against 285,165 clips in
houlab/motif-db. The two differ because the annotation rows of the 25 CoSSMos
classes are CoSSMos's own… See the full description on the dataset page: https://huggingface.co/datasets/houlab/cossmos-annotations-db.CEFR-Sentence-Level-Annotations
Dataset Card for Dataset Name
17k english sentences annotated by english education professionals. Original repo for CEFR-SP is located at this repo
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP):… See the full description on the dataset page: https://huggingface.co/datasets/edesaras/CEFR-Sentence-Level-Annotations.pick_and_place_annotationThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "Unitree_G1_Inspire",
"total_episodes": 1035,
"total_frames": 322070,
"total_tasks": 14,
"total_videos": 2070,
"total_chunks": 2,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1035"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/KGB0/pick_and_place_annotation.emolia-voicenet-gemini-annotations
Emolia VoiceNet Gemini Annotations
468,180 dimension-level annotations over 236,613 Emolia speech clips,
each scored 0-6 (0-2 for the content-safety dimension) on one of 57 perceptual
voice / speech dimensions - arousal, valence, brightness, resonance placement, speaking
styles, genuineness, recording quality, and more - by Gemini 3.5 Flash (non-thinking,
temperature 0). This repository ships the annotations, audio provenance, per-dimension
statistics, and the full scoring… See the full description on the dataset page: https://huggingface.co/datasets/laion/emolia-voicenet-gemini-annotations.annotation_app_data
Dataset Card for Systematic Review of Acceptability Judgments data
A curated dataset of research articles used in a systematic review of judgment tasks in linguistics. Each entry records article-level metadata and experiment-level methodological features, supporting structured comparison and analysis across studies.
Dataset Description
This annotation dataset comprises systematically coded observations from a corpus of published studies employing judgment tasks in… See the full description on the dataset page: https://huggingface.co/datasets/jasongraf1/annotation_app_data.highlevel_thinking_with_grounding_annotation_split1000_v2_merged_promptsemboss-roof-annotations
Emboss 3D Roof Reference Annotations
Manual 3D reference meshes and editable annotations for Swiss and Brazilian buildings, prepared for the evaluation and parameter tuning of Emboss. The annotations describe building and roof geometry, including roof superstructures.
Emboss source code
3dlabel annotation tool
Emboss segmentation model
Example reference annotation in 3dlabel: annotated mesh and LiDAR points (Figure D.1(a) in the paper).
egolongqa-synth-annotations
EgoLongQA synthetic MCQs, teacher traces and annotation outputs
Everything produced by the annotation and synthesis pipelines for the AI Wearables Challenge 2026
EgoLongQA ≤2B track, other than the distillation set (which lives in
infinitylogesh/egolongqa-junior-distill).
⚠️ Read this before counting rows
The synthetic set is 943 questions over 408 videos, and it is stored two ways:
file
rows
shape
training_sets/train_synth_v3.jsonl
943
flat — one row… See the full description on the dataset page: https://huggingface.co/datasets/ambient-intelligence-labs/egolongqa-synth-annotations.seamless-interaction-jefferson-annotations
Seamless Interaction Jefferson-Style Annotations
An automatic, turn-oriented annotation layer for the
Meta Seamless Interaction Dataset.
It compares the dataset's traditional transcript with an ASR-derived
Jefferson-style condition and supplies speech-act, communicative-purpose,
interactional-signal, alignment, and quality fields.
This is a derived noncommercial research dataset. It does not redistribute
the source audio. Every record retains the original interaction ID, split… See the full description on the dataset page: https://huggingface.co/datasets/kennethli319/seamless-interaction-jefferson-annotations.kwext-bilibili-video-title-annotations
KwExt Bilibili Video Title Annotations
This dataset is a model-assisted annotation set for the KwExt keyword
extraction project. The current snapshot contains 5,000 Chinese Bilibili
video titles from annotation stages video_title_zh_001 through
video_title_zh_005, with 1,000 records in each stage.
The release is intended for early experiments with:
extracting title-grounded keywords and ranking their importance;
broad semantic tags for retrieval and RAG metadata;
dense tag… See the full description on the dataset page: https://huggingface.co/datasets/Himpq/kwext-bilibili-video-title-annotations.highlevel_thinking_with_grounding_annotation_split1000_merged_promptsmultimodal-video-annotation-samples
Video Annotation Samples – SuperviseLab
SuperviseLab provides professional video annotation data for training multimodal AI models. This public sample dataset demonstrates our annotation methodology and output quality across diverse video content categories.
Note: All visual assets in this dataset have been abstracted (pixelated mosaic) to protect source privacy. Uploader identity, original titles, and all identifiable metadata have been removed. This is a demonstration dataset… See the full description on the dataset page: https://huggingface.co/datasets/superviselab/multimodal-video-annotation-samples.4k-video-annotations
4K Video Annotations — Shot Segmentation and Camera Motion
This dataset contains 12 frame-accurate shot clips segmented from five short cinematic video sequences. Every clip is paired with a detailed, manually reviewed annotation covering visible content, subject actions, shot scale, camera angle, camera movement, movement direction, stabilization, composition, lighting, color, pacing, transitions, timecodes, and technical properties.
The footage depicts a tense nighttime… See the full description on the dataset page: https://huggingface.co/datasets/LianeMarilin/4k-video-annotations.example_10kbp_human_annotationsConceptARC_Rule_Annotations
ConceptARC Rule Annotations
Model outputs, natural-language rules and human judgements of those rules on the 480 tasks of the ConceptARC benchmark. This is the data behind the paper Distinguishing Performance From Competence in Evaluations of Humanlike Abstract Reasoning (NeurIPS 2026, Evaluations and Datasets Track).
Paper: arXiv:2510.02125
Project page and interactive viewer: claasbeger.github.io/performance-competence-gap
Authors: Claas Beger, Ryan Yi, Shuhao Fu, Kaleda… See the full description on the dataset page: https://huggingface.co/datasets/ClaasBeger/ConceptARC_Rule_Annotations.lumiopen-hpltv2-llama33-edu-annotation-etscripted_atomic_train_frac_0.3_large_goal_annotationThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "ur5_wsg50_lego_atomic_step",
"total_episodes": 664,
"total_frames": 116214,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:664"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path": null… See the full description on the dataset page: https://huggingface.co/datasets/windfromthenorth/scripted_atomic_train_frac_0.3_large_goal_annotation.SWE-bench_Verified_With_Annotationsexample_eval_only_10kb_human_annotationsTurkWeb-Edu-AnnotationsV3
TurkWeb-Edu V3
Model: Qwen/Qwen3-30B-A3B-Instruct-2507
Format: Structured JSON (vLLM 0.15.0)
