Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01tanish434 /Truebones-ZOO-Annotations Truebones ZOO Annotations Text prompts, per-clip metadata, rest-pose renders and the exact build pipeline for Truebones ZOO — 1,097 animal motion clips across 74 skeletons: mammals, birds, reptiles, insects, marine and prehistoric creatures. 1.02 hours, 111,158 frames, uniformly 30 fps. Rigs range from 9 to 143 joints; clips from 0.3 to 18.5 seconds. The motion files themselves are not in this repository. Truebones ZOO is a commercial library by Truebones Motions Animation… See the full description on the dataset page: https://huggingface.co/datasets/tanish434/Truebones-ZOO-Annotations.tabular1K<n<10K0 likes3.2k downloads27d agoHugging Face02JQL-AI /JQL-LLM-Edu-Annotations 📚 JQL Educational Quality Annotations from LLMs This dataset provides 17,186,606 documents with high-quality LLM annotations for evaluating the educational value of web documents, and serves as a benchmark for training and evaluating multilingual LLM annotators as described in the JQL paper. 📝 Dataset Summary Multilingual document-level quality annotations scored on a 0–5 educational value scale by three state-of-the-art LLMs: Gemma-3-27B-it, Mistral-3.1-24B-it… See the full description on the dataset page: https://huggingface.co/datasets/JQL-AI/JQL-LLM-Edu-Annotations.tabular10M<n<100M2 likes1.3k downloads1y agoHugging Face03it-just-works /vast27m_annotations VAST-27M Annotations Dataset This dataset contains annotations from the VAST-27M dataset, originally created for the paper "VAST: A Vision-Audio-Subtitle-Text Omni-Modality Foundation Model and Dataset" by Chen et al. (2024). Original Source This dataset is derived from the VAST-27M dataset, which was created by researchers at the University of Chinese Academy of Sciences and the Institute of Automation, Chinese Academy of Science. The original dataset and more… See the full description on the dataset page: https://huggingface.co/datasets/it-just-works/vast27m_annotations.tabular10M<n<100M1 likes697 downloads2y agoHugging Face04Linzhan /Truebones-ZOO-Annotations Truebones ZOO Annotations Text prompts, per-clip metadata, rest-pose renders and the exact build pipeline for Truebones ZOO — 1,097 animal motion clips across 74 skeletons: mammals, birds, reptiles, insects, marine and prehistoric creatures. 1.02 hours, 111,158 frames, uniformly 30 fps. Rigs range from 9 to 143 joints; clips from 0.3 to 18.5 seconds. The motion files themselves are not in this repository. Truebones ZOO is a commercial library by Truebones Motions Animation… See the full description on the dataset page: https://huggingface.co/datasets/Linzhan/Truebones-ZOO-Annotations.tabular1K<n<10K3 likes429 downloads1mo agoHugging Face05YsK-dev /TurkWeb-Edu-AnnotationsV3 TurkWeb-Edu V3 Model: Qwen/Qwen3-30B-A3B-Instruct-2507 Format: Structured JSON (vLLM 0.15.0) tabular100K<n<1M0 likes322 downloads8mo agoHugging Face06edesaras /CEFR-Sentence-Level-Annotations Dataset Card for Dataset Name 17k english sentences annotated by english education professionals. Original repo for CEFR-SP is located at this repo This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP):… See the full description on the dataset page: https://huggingface.co/datasets/edesaras/CEFR-Sentence-Level-Annotations.tabulartext-classification10K<n<100K6 likes260 downloads2y agoHugging Face07houlab /cossmos-annotations-db ARSMA-web motif annotations (cossmos-annotations-db) Per-motif annotation: base pairs, stacking, sugar puckers, glycosidic conformations and the deposition metadata of the parent entry. It backs the motif browser of ARSMA-web. This is not the occurrence count. One row of instances.parquet is one annotated motif site, and the table has 283,164 of them against 285,165 clips in houlab/motif-db. The two differ because the annotation rows of the 25 CoSSMos classes are CoSSMos's own… See the full description on the dataset page: https://huggingface.co/datasets/houlab/cossmos-annotations-db.tabular1M<n<10M0 likes257 downloads8d agoHugging Face08laion /emolia-voicenet-gemini-annotations Emolia VoiceNet Gemini Annotations 468,180 dimension-level annotations over 236,613 Emolia speech clips, each scored 0-6 (0-2 for the content-safety dimension) on one of 57 perceptual voice / speech dimensions - arousal, valence, brightness, resonance placement, speaking styles, genuineness, recording quality, and more - by Gemini 3.5 Flash (non-thinking, temperature 0). This repository ships the annotations, audio provenance, per-dimension statistics, and the full scoring… See the full description on the dataset page: https://huggingface.co/datasets/laion/emolia-voicenet-gemini-annotations.tabularaudio-classification100K<n<1M0 likes230 downloads3mo agoHugging Face09ambient-intelligence-labs /egolongqa-synth-annotations EgoLongQA synthetic MCQs, teacher traces and annotation outputs Everything produced by the annotation and synthesis pipelines for the AI Wearables Challenge 2026 EgoLongQA ≤2B track, other than the distillation set (which lives in infinitylogesh/egolongqa-junior-distill). ⚠️ Read this before counting rows The synthetic set is 943 questions over 408 videos, and it is stored two ways: file rows shape training_sets/train_synth_v3.jsonl 943 flat — one row… See the full description on the dataset page: https://huggingface.co/datasets/ambient-intelligence-labs/egolongqa-synth-annotations.tabularvisual-question-answering1K<n<10K0 likes213 downloads1mo agoHugging Face10tvonarx /emboss-roof-annotations Emboss 3D Roof Reference Annotations Manual 3D reference meshes and editable annotations for Swiss and Brazilian buildings, prepared for the evaluation and parameter tuning of Emboss. The annotations describe building and roof geometry, including roof superstructures. Emboss source code 3dlabel annotation tool Emboss segmentation model Example reference annotation in 3dlabel: annotated mesh and LiDAR points (Figure D.1(a) in the paper). 3dn<1K0 likes213 downloads27d agoHugging Face11Himpq /kwext-bilibili-video-title-annotations KwExt Bilibili Video Title Annotations This dataset is a model-assisted annotation set for the KwExt keyword extraction project. The current snapshot contains 5,000 Chinese Bilibili video titles from annotation stages video_title_zh_001 through video_title_zh_005, with 1,000 records in each stage. The release is intended for early experiments with: extracting title-grounded keywords and ranking their importance; broad semantic tags for retrieval and RAG metadata; dense tag… See the full description on the dataset page: https://huggingface.co/datasets/Himpq/kwext-bilibili-video-title-annotations.tabulartoken-classification1K<n<10K0 likes183 downloads23d agoHugging Face12kennethli319 /seamless-interaction-jefferson-annotations Seamless Interaction Jefferson-Style Annotations An automatic, turn-oriented annotation layer for the Meta Seamless Interaction Dataset. It compares the dataset's traditional transcript with an ASR-derived Jefferson-style condition and supplies speech-act, communicative-purpose, interactional-signal, alignment, and quality fields. This is a derived noncommercial research dataset. It does not redistribute the source audio. Every record retains the original interaction ID, split… See the full description on the dataset page: https://huggingface.co/datasets/kennethli319/seamless-interaction-jefferson-annotations.tabularautomatic-speech-recognition100K<n<1M0 likes172 downloads2mo agoHugging Face13LianeMarilin /4k-video-annotations 4K Video Annotations — Shot Segmentation and Camera Motion This dataset contains 12 frame-accurate shot clips segmented from five short cinematic video sequences. Every clip is paired with a detailed, manually reviewed annotation covering visible content, subject actions, shot scale, camera angle, camera movement, movement direction, stabilization, composition, lighting, color, pacing, transitions, timecodes, and technical properties. The footage depicts a tense nighttime… See the full description on the dataset page: https://huggingface.co/datasets/LianeMarilin/4k-video-annotations.imagen<1K0 likes171 downloads26d agoHugging Face14ClaasBeger /ConceptARC_Rule_Annotations ConceptARC Rule Annotations Model outputs, natural-language rules and human judgements of those rules on the 480 tasks of the ConceptARC benchmark. This is the data behind the paper Distinguishing Performance From Competence in Evaluations of Humanlike Abstract Reasoning (NeurIPS 2026, Evaluations and Datasets Track). Paper: arXiv:2510.02125 Project page and interactive viewer: claasbeger.github.io/performance-competence-gap Authors: Claas Beger, Ryan Yi, Shuhao Fu, Kaleda… See the full description on the dataset page: https://huggingface.co/datasets/ClaasBeger/ConceptARC_Rule_Annotations.tabular10K<n<100K0 likes161 downloads8d agoHugging Face15CLS-Lab /narrative-gold-annotations Narrative annotation dataset Human annotations for three narrative-analysis tasks — setting, agency, and event relation — over passages sampled from the Dolma corpus. Annotators & anonymization Annotator identities are anonymized. Each task has a single gold adjudicator plus one or more secondary annotators used for double annotation / agreement. Role Meaning gold The adjudicated / primary label for every released instance. annotator_1 Second… See the full description on the dataset page: https://huggingface.co/datasets/CLS-Lab/narrative-gold-annotations.tabulartext-classification1K<n<10K1 likes153 downloads3mo agoHugging Face16emarro /example_10kbp_human_annotationstabular100K<n<1M0 likes138 downloads1y agoHugging Face17abullard1 /steam-reviews-constructiveness-binary-label-annotations-1.5k 1.5K Steam Reviews Binary Labeled for Constructiveness Dataset Summary This dataset contains 1,461 Steam reviews from 10 of the most reviewed games. Each game has about the same amount of reviews. Each review is annotated with a binary label indicating whether the review is constructive or not. The dataset is designed to support tasks related to text classification, particularly constructiveness detection tasks in the gaming domain. Also available as… See the full description on the dataset page: https://huggingface.co/datasets/abullard1/steam-reviews-constructiveness-binary-label-annotations-1.5k.tabulartext-classification1K<n<10K2 likes121 downloads2y agoHugging Face18aiobservatory /annotations The AI Observatory A public measurement platform aggregating real-world AI conversations from seven sources under a unified 145-feature taxonomy. This dataset card accompanies paper The AI Observatory: A Public Measure of Real-World AI Use. 📄 Paper: [anonymous OpenReview link] 📊 Dashboard: https://project-ai-observatory.vercel.app/ 💾 Anonymous code: https://anonymous.4open.science/r/ai-observatory/README.md TL;DR 23,158 conversations, 85,633 turns, ~5,000… See the full description on the dataset page: https://huggingface.co/datasets/aiobservatory/annotations.tabulartext-classification100K<n<1M3 likes119 downloads2mo agoHugging Face19emarro /example_eval_only_10kb_human_annotationstabular10K<n<100K0 likes115 downloads1y agoHugging Face20Alptekinege /TurkWeb-Edu-AnnotationsV3 TurkWeb-Edu V3 Model: Qwen/Qwen3-30B-A3B-Instruct-2507 Format: Structured JSON (vLLM 0.15.0) tabular100K<n<1M0 likes113 downloads7mo agoHugging Face21nikhilchandak /gpqa-diamond-annotations GPQA Diamond Dataset This dataset contains filtered JSONL files of human annotations on question specificity, answer uniqueness, answer matching to the ground truth for different models for the GPQA Diamond dataset. The dataset was annotated by two human graders. It contains 198 (original size) * 2 = 396 rows as each rows is repeated twice (one for each human). A human grader given the question, actual answer and model response, has to answer whether the response matches the… See the full description on the dataset page: https://huggingface.co/datasets/nikhilchandak/gpqa-diamond-annotations.tabularn<1K1 likes112 downloads1y agoHugging Face22WRBench /wrbench-human-annotations WRBench Human Annotations This dataset contains the human comparison labels used to validate WRBench's automatic evaluation metrics. Version Update: 2026-07-07 We updated the release after rechecking videos that changed during benchmark maintenance. The release now includes: 1,741 clean comparison rows. 4,302 individual human judgments. 585 newly rechecked current-benchmark comparisons, each reviewed by three annotators. Majority-label summaries for the newly… See the full description on the dataset page: https://huggingface.co/datasets/WRBench/wrbench-human-annotations.tabularimage-to-video1K<n<10K0 likes110 downloads3mo agoHugging Face23hirotakahiraki /candor-turntaking-annotations CANDOR - Turn-Taking Annotations Speech transcription and turn-taking annotation dataset built from the CANDOR corpus using NVIDIA Canary-Qwen2.5B ASR. Dataset Description This dataset contains 172,591 transcribed speech segments from the CANDOR conversational speech corpus (1,656 conversations). Each segment is a per-speaker utterance with Canary ASR transcript, designed for turn-taking prediction research. Source Audio corpus: CANDOR (English… See the full description on the dataset page: https://huggingface.co/datasets/hirotakahiraki/candor-turntaking-annotations.tabularautomatic-speech-recognition100K<n<1M1 likes107 downloads7mo agoHugging Face24CLS-Lab /narrative-llm-annotations NarraDolma LLM-Labeled — Distillation Set The intermediate, LLM-labeled dataset that bridges the small human gold set and the full NarraDolma corpus. It contains 5,000 passages sampled from Dolma and labeled by Gemma across all 11 narrative dimensions, stratified by source and topic to preserve the original distribution. These labels are the knowledge-distillation training set used to train NarraBert. Paper: arXiv:2606.19468 Collection: Narratives in LLM Pretraining Data… See the full description on the dataset page: https://huggingface.co/datasets/CLS-Lab/narrative-llm-annotations.tabular10K<n<100K1 likes101 downloads1mo agoHugging Face25sed-i /mania-pattern-annotations osu!mania pattern annotations Snapshot v4 uses publication schema beatmap-lens-annotations version 4. This snapshot contains 600 human judgments, 16402 agent judgments, and 2969 source identities (annotation and required calibration sources). Export implementation: GitHub commit d6b4210550ba. v4: 2,860 complete sections This release retains all 2,860 requested complete sections using explicit metadata-only source access. Each source retains its original SHA-256… See the full description on the dataset page: https://huggingface.co/datasets/sed-i/mania-pattern-annotations.tabular10K<n<100K2 likes97 downloads1d agoHugging Face26huyouare /SWE-bench_Verified_With_Annotationstabularn<1K1 likes94 downloads2y agoHugging Face27PrentisAI /ScreenRef-Annotations ScreenRef — Annotations Annotations only. This repository contains no screenshots. It holds every annotation of ScreenRef (484 MB); the screenshots are in the gated repository PrentisAI/ScreenRef. The annotations are licensed under the same LICENSE as the screenshots: Academic Research only, no redistribution. Contents Path What it is Size anno/ 58 JSONL files — all tasks (the full recipe), 633,369 rows 240 MB anno_point/ 46 JSONL files — same screens… See the full description on the dataset page: https://huggingface.co/datasets/PrentisAI/ScreenRef-Annotations.tabularimage-text-to-text1M<n<10M0 likes90 downloads11d agoHugging Face28toroe /llama-nemotron-science-propella-annotationstabular100K<n<1M0 likes85 downloads8mo agoHugging Face29jablonkagroup /corral-reasoning-annotations Corral – Reasoning Annotations LLM epistemic annotations over Corral traces where the annotator judged that the agents do not reason scientifically 📋 Dataset Summary This dataset is part of the Corral collection accompanying the paper AI scientists produce results without reasoning scientifically. It contains annotated evaluation traces with LLM-generated epistemic annotations across the Corral benchmark. The dataset is exposed as three model-specific… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/corral-reasoning-annotations.tabulartext-classificationn<1K0 likes83 downloads4mo agoHugging Face30Abdu07 /Agentglass-swerebench-annotations Dataset Summary SWE-rebench-OpenHands-Trajectories is a dataset of multi-turn agent trajectories for software engineering tasks, collected using Qwen/Qwen3-Coder-480B-A35B-Instruct with OpenHands (v0.54.0) agent scaffolding. This dataset captures complete agent execution traces as they attempt to resolve real GitHub issues from nebius/SWE-rebench. Each trajectory contains the agent's step-by-step reasoning, actions, and environmental observations. Metric… See the full description on the dataset page: https://huggingface.co/datasets/Abdu07/Agentglass-swerebench-annotations.tabular10K<n<100K0 likes67 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.