Team Ai
28 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01tanish434 /Truebones-ZOO-Annotations Truebones ZOO Annotations Text prompts, per-clip metadata, rest-pose renders and the exact build pipeline for Truebones ZOO — 1,097 animal motion clips across 74 skeletons: mammals, birds, reptiles, insects, marine and prehistoric creatures. 1.02 hours, 111,158 frames, uniformly 30 fps. Rigs range from 9 to 143 joints; clips from 0.3 to 18.5 seconds. The motion files themselves are not in this repository. Truebones ZOO is a commercial library by Truebones Motions Animation… See the full description on the dataset page: https://huggingface.co/datasets/tanish434/Truebones-ZOO-Annotations.tabular1K<n<10K0 likes3.2k downloads27d agoHugging Face02Linzhan /Truebones-ZOO-Annotations Truebones ZOO Annotations Text prompts, per-clip metadata, rest-pose renders and the exact build pipeline for Truebones ZOO — 1,097 animal motion clips across 74 skeletons: mammals, birds, reptiles, insects, marine and prehistoric creatures. 1.02 hours, 111,158 frames, uniformly 30 fps. Rigs range from 9 to 143 joints; clips from 0.3 to 18.5 seconds. The motion files themselves are not in this repository. Truebones ZOO is a commercial library by Truebones Motions Animation… See the full description on the dataset page: https://huggingface.co/datasets/Linzhan/Truebones-ZOO-Annotations.tabular1K<n<10K3 likes429 downloads1mo agoHugging Face03jasongraf1 /annotation_app_data Dataset Card for Systematic Review of Acceptability Judgments data A curated dataset of research articles used in a systematic review of judgment tasks in linguistics. Each entry records article-level metadata and experiment-level methodological features, supporting structured comparison and analysis across studies. Dataset Description This annotation dataset comprises systematically coded observations from a corpus of published studies employing judgment tasks in… See the full description on the dataset page: https://huggingface.co/datasets/jasongraf1/annotation_app_data.tabularn<1K0 likes279 downloads3d agoHugging Face04tvonarx /emboss-roof-annotations Emboss 3D Roof Reference Annotations Manual 3D reference meshes and editable annotations for Swiss and Brazilian buildings, prepared for the evaluation and parameter tuning of Emboss. The annotations describe building and roof geometry, including roof superstructures. Emboss source code 3dlabel annotation tool Emboss segmentation model Example reference annotation in 3dlabel: annotated mesh and LiDAR points (Figure D.1(a) in the paper). 3dn<1K0 likes213 downloads27d agoHugging Face05abullard1 /steam-reviews-constructiveness-binary-label-annotations-1.5k 1.5K Steam Reviews Binary Labeled for Constructiveness Dataset Summary This dataset contains 1,461 Steam reviews from 10 of the most reviewed games. Each game has about the same amount of reviews. Each review is annotated with a binary label indicating whether the review is constructive or not. The dataset is designed to support tasks related to text classification, particularly constructiveness detection tasks in the gaming domain. Also available as… See the full description on the dataset page: https://huggingface.co/datasets/abullard1/steam-reviews-constructiveness-binary-label-annotations-1.5k.tabulartext-classification1K<n<10K2 likes121 downloads2y agoHugging Face06huyouare /SWE-bench_Verified_With_Annotationstabularn<1K1 likes94 downloads2y agoHugging Face07processvenue /INVOICE_ANNOTATION_V2tabularimage-classification1K<n<10K0 likes67 downloads10mo agoHugging Face08soda-lmu /tweet-annotation-sensitivity-2 Tweet Annotation Sensitivity Experiment 2: Annotations in Five Experimental Conditions Attention: This repository contains cases that might be offensive or upsetting. We do not support the views expressed in these hateful posts. Description The dataset contains tweet data annotations of hate speech (HS) and offensive language (OL) in five experimental conditions. The tweet data was sampled from the corpus created by Davidson et al. (2017). We selected 3,000 Tweets for our… See the full description on the dataset page: https://huggingface.co/datasets/soda-lmu/tweet-annotation-sensitivity-2.tabulartext-classification10K<n<100K3 likes48 downloads2y agoHugging Face09vennu95 /llm-delusion-response-annotations LLM Delusion-Like Belief Reinforcement Annotations This dataset contains human annotations of responses generated by conversational large language models (LLMs) to prompts expressing potentially delusion-like or reality-distorted beliefs. The purpose of the dataset is to support evaluation of whether conversational LLM responses may unintentionally reinforce or strengthen delusion-like beliefs. Dataset Files Consensus Dataset… See the full description on the dataset page: https://huggingface.co/datasets/vennu95/llm-delusion-response-annotations.documenttext-classification1K<n<10K0 likes46 downloads4mo agoHugging Face10LawrenceYin /annotation-pack-a Annotation pack A — does a response genuinely follow an instruction? 50 rows. Each row: a user prompt, one instruction from it, and a model response that an automatic checker marks as satisfying that instruction. Annotators judge whether it is satisfied genuinely. For annotators / 标注人: Read RUBRIC_HUMAN.md (English + 中文). Annotator A downloads annotator_A.csv; annotator B downloads annotator_B.csv (same items). Fill label (GENUINE / LOOPHOLE / GARBLED), helpfulness (1–5)… See the full description on the dataset page: https://huggingface.co/datasets/LawrenceYin/annotation-pack-a.tabulartext-generationn<1K0 likes41 downloads2d agoHugging Face11wasanx /gemba_annotation GEMBA Annotation Dataset Description This dataset contains human and machine-generated Multidimensional Quality Metrics (MQM) annotations for machine-translated Thai text. It is intended for evaluating and comparing MT system outputs and error annotation quality across different automated models. The dataset features annotations from three large language models (LLMs): Claude 3.7 Sonnet, Gemini 2.0 Flash, and 4o Mini, alongside human annotations, providing a comprehensive… See the full description on the dataset page: https://huggingface.co/datasets/wasanx/gemba_annotation.tabulartranslation1K<n<10K0 likes33 downloads1y agoHugging Face12LawrenceYin /annotation-pack-b Annotation pack B — are any words forced into the text? 100 short texts (60 web-style excerpts, 40 one-sentence news summaries) written by small language models. Annotators judge whether any word looks forced in. For annotators / 标注人: Read RUBRIC_HUMAN.md. Annotator A downloads annotator_A.csv; annotator B downloads annotator_B.csv (same items). Fill label (GENUINE / LOOPHOLE / GARBLED), wrong_sense (Y/N), fluency (1–5), flagged_words, optional notes. Work alone. Send the… See the full description on the dataset page: https://huggingface.co/datasets/LawrenceYin/annotation-pack-b.tabulartext-generationn<1K0 likes28 downloads6d agoHugging Face13processvenue /INVOICE_ANNOTATION_V1tabularimage-classification1K<n<10K0 likes27 downloads10mo agoHugging Face14ManjuKrish /llm-delusion-response-annotations LLM Delusion-Like Belief Reinforcement Annotations This dataset contains human annotations of responses generated by conversational large language models (LLMs) to prompts expressing potentially delusion-like or reality-distorted beliefs. The purpose of the dataset is to support evaluation of whether conversational LLM responses may unintentionally reinforce or strengthen delusion-like beliefs. Dataset Files Consensus Dataset… See the full description on the dataset page: https://huggingface.co/datasets/ManjuKrish/llm-delusion-response-annotations.documenttext-classification1K<n<10K0 likes24 downloads4mo agoHugging Face15prateek-0-gupta /allaimovies-annotations allaimovies overview annotations The raw per-film codings behind the allaimovies dataset: 3,092 science-fiction films (1911-2026) whose TMDB plot overview was read by gpt-5.4-mini against a fixed rubric (below) with schema-enforced JSON output. 2,069 were coded as having an AI present. This table is the model's output as collected, one row per film, before it was joined with reception, credits and character data; use it if you want to re-check, re-code or compare against another… See the full description on the dataset page: https://huggingface.co/datasets/prateek-0-gupta/allaimovies-annotations.tabulartext-classification1K<n<10K0 likes21 downloads1mo agoHugging Face16bryanchrist /EDUMATH_annotations EDUMATH Annotation Dataset The EDUMATH Annotation Dataset contains 3,012 math word problems annotated by teachers and Gemma 3 27B IT as reported in EDUMATH: Generating Standards-aligned Educational Math Word Problems. The dataset contains the final labels from human annotators in the solvability, accuracy, appropriateness, and standards_alignment columns along with the final label for Meets all Criteria (MaC), which was assigned as described in the paper. The model_labels and… See the full description on the dataset page: https://huggingface.co/datasets/bryanchrist/EDUMATH_annotations.tabular1K<n<10K0 likes20 downloads6mo agoHugging Face17EmmaYee /FER2013-VAD-annotationThis dataset involves train-20240123-14902.csv, publictest-20240508.csv and privatetest-20240506-yh.csv, which could be used for public and private test respectively. 14902, 1298 and 3589 samples are for train, public and private dataset at present. tabular10K<n<100K0 likes19 downloads2mo agoHugging Face18pytorch-survival /gene_annotationstabular10K<n<100K0 likes17 downloads3y agoHugging Face19apbrault /music_annotationstabular10K<n<100K0 likes17 downloads2y agoHugging Face20huggingworld /clinvar-annotationspart of 🧬 Genomic Reasoning Agent LLM-driven agentic system for personal genomic variant interpretation Overview This project builds a multi-step reasoning agent that interprets personal genomic data from 23andMe against biomedical knowledge databases (ClinVar, GWAS Catalog, gnomAD). The agent is trained with GRPO (Group Relative Policy Optimization) using fully verifiable reward signals — no human labelers needed. The core insight mirrors DeepSeek-R1's… See the full description on the dataset page: https://huggingface.co/datasets/huggingworld/clinvar-annotations.tabularn<1K0 likes16 downloads6mo agoHugging Face21iwylin /imbue_dearman_expert_annotations Dataset Overview This dataset was collected as part of the project "IMBUE: Improving Interpersonal Effectiveness through Simulation and Just-in-time Feedback with Human-Language Model Interaction." We are releasing this dataset with the hope that it provides valuable opportunities for researchers to develop and evaluate new LLM-based tools for interpersonal skill training across a range of fields, including natural language processing, conversational AI, and computational… See the full description on the dataset page: https://huggingface.co/datasets/iwylin/imbue_dearman_expert_annotations.tabularn<1K0 likes13 downloads2y agoHugging Face22cemig-ceia /fineweb-edu-gemini-annotations-portuguese-regressiontabular1K<n<10K0 likes12 downloads1y agoHugging Face23bryanchrist /annotations MATHWELL Human Annotation Dataset The MATHWELL Human Annotation Dataset contains 5,084 synthetic word problems and answers generated by MATHWELL, a reference-free educational grade school math word problem generator released in MATHWELL: Generating Educational Math Word Problems Using Teacher Annotations, and comparison models (GPT-4, GPT-3.5, Llama-2, MAmmoTH, and LLEMMA) with expert human annotations for solvability, accuracy, appropriateness, and meets all criteria (MaC).… See the full description on the dataset page: https://huggingface.co/datasets/bryanchrist/annotations.tabular1K<n<10K2 likes11 downloads2y agoHugging Face24TurCOMET /TurkDoc-MT-MQM-Annotationstabular1K<n<10K0 likes11 downloads5mo agoHugging Face25jrfish /Approximation_Political_Neutrality_Annotation_Dataset Dataset Card for Approximation of Political Neutrality Annotation Dataset This dataset card accompanies the paper Political Neturality in AI is Impossible- But Here is How to Approximate it. This the dataset of annotatd generations of political questions. Dataset Details Dataset Description It includes the input propmts, output generations, as well as the annotated approxiamtions of political neutrality techniques. Curated by: Jillian Fisher Language(s)… See the full description on the dataset page: https://huggingface.co/datasets/jrfish/Approximation_Political_Neutrality_Annotation_Dataset.tabular10K<n<100K0 likes10 downloads2y agoHugging Face26actdisease /historical-magazines-genre-annotationgatedtabulartext-classification10K<n<100K0 likes6 downloads7mo agoHugging Face27soda-lmu /tweet-annotation-sensitivity-1 Tweet Annotation Sensitivity Experiment 1: Annotation in Six Experimental Conditions Attention: This repository contains cases that might be offensive or upsetting. We do not support the views expressed in these hateful posts. Description We drew a stratified sample of 20 tweets, that were pre-annotated in a study by Davidson et al. (2017) for Hate Speech / Offensive Language / Neither. The stratification was done with respect to majority-voted class and level of… See the full description on the dataset page: https://huggingface.co/datasets/soda-lmu/tweet-annotation-sensitivity-1.tabulartext-classification1K<n<10K0 likes5 downloads3y agoHugging Face28hbXNov /distill_r1_qwen_math_1.5b_128_solns_math_train_with_correctness_gpt_annotationtabular1K<n<10K0 likes4 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.