datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Truebones-ZOO-Annotations
Truebones ZOO Annotations
Text prompts, per-clip metadata, rest-pose renders and the exact build pipeline for
Truebones ZOO — 1,097 animal motion clips across 74 skeletons: mammals, birds,
reptiles, insects, marine and prehistoric creatures. 1.02 hours, 111,158 frames, uniformly
30 fps. Rigs range from 9 to 143 joints; clips from 0.3 to 18.5 seconds.
The motion files themselves are not in this repository. Truebones ZOO is a commercial
library by Truebones Motions Animation… See the full description on the dataset page: https://huggingface.co/datasets/tanish434/Truebones-ZOO-Annotations.Truebones-ZOO-Annotations
Truebones ZOO Annotations
Text prompts, per-clip metadata, rest-pose renders and the exact build pipeline for
Truebones ZOO — 1,097 animal motion clips across 74 skeletons: mammals, birds,
reptiles, insects, marine and prehistoric creatures. 1.02 hours, 111,158 frames, uniformly
30 fps. Rigs range from 9 to 143 joints; clips from 0.3 to 18.5 seconds.
The motion files themselves are not in this repository. Truebones ZOO is a commercial
library by Truebones Motions Animation… See the full description on the dataset page: https://huggingface.co/datasets/Linzhan/Truebones-ZOO-Annotations.emboss-roof-annotations
Emboss 3D Roof Reference Annotations
Manual 3D reference meshes and editable annotations for Swiss and Brazilian buildings, prepared for the evaluation and parameter tuning of Emboss. The annotations describe building and roof geometry, including roof superstructures.
Emboss source code
3dlabel annotation tool
Emboss segmentation model
Example reference annotation in 3dlabel: annotated mesh and LiDAR points (Figure D.1(a) in the paper).
steam-reviews-constructiveness-binary-label-annotations-1.5k
1.5K Steam Reviews Binary Labeled for Constructiveness
Dataset Summary
This dataset contains 1,461 Steam reviews from 10 of the most reviewed games. Each game has about the same amount of reviews. Each review is annotated with a binary label indicating whether the review is constructive or not. The dataset is designed to support tasks related to text classification, particularly constructiveness detection tasks in the gaming domain.
Also available as… See the full description on the dataset page: https://huggingface.co/datasets/abullard1/steam-reviews-constructiveness-binary-label-annotations-1.5k.SWE-bench_Verified_With_Annotationsllm-delusion-response-annotations
LLM Delusion-Like Belief Reinforcement Annotations
This dataset contains human annotations of responses generated by conversational large language models (LLMs) to prompts expressing potentially delusion-like or reality-distorted beliefs.
The purpose of the dataset is to support evaluation of whether conversational LLM responses may unintentionally reinforce or strengthen delusion-like beliefs.
Dataset Files
Consensus Dataset… See the full description on the dataset page: https://huggingface.co/datasets/vennu95/llm-delusion-response-annotations.llm-delusion-response-annotations
LLM Delusion-Like Belief Reinforcement Annotations
This dataset contains human annotations of responses generated by conversational large language models (LLMs) to prompts expressing potentially delusion-like or reality-distorted beliefs.
The purpose of the dataset is to support evaluation of whether conversational LLM responses may unintentionally reinforce or strengthen delusion-like beliefs.
Dataset Files
Consensus Dataset… See the full description on the dataset page: https://huggingface.co/datasets/ManjuKrish/llm-delusion-response-annotations.allaimovies-annotations
allaimovies overview annotations
The raw per-film codings behind the allaimovies dataset: 3,092 science-fiction films
(1911-2026) whose TMDB plot overview was read by gpt-5.4-mini against a fixed rubric (below) with
schema-enforced JSON output. 2,069 were coded as having an AI present. This table is the
model's output as collected, one row per film, before it was joined with reception, credits and
character data; use it if you want to re-check, re-code or compare against another… See the full description on the dataset page: https://huggingface.co/datasets/prateek-0-gupta/allaimovies-annotations.EDUMATH_annotations
EDUMATH Annotation Dataset
The EDUMATH Annotation Dataset contains 3,012 math word problems annotated by teachers and Gemma 3 27B IT as reported in EDUMATH: Generating Standards-aligned Educational Math Word Problems. The dataset contains the final labels from human annotators in the solvability, accuracy, appropriateness, and standards_alignment columns along with the final label for Meets all Criteria (MaC), which was assigned as described in the paper. The model_labels and… See the full description on the dataset page: https://huggingface.co/datasets/bryanchrist/EDUMATH_annotations.gene_annotationsmusic_annotationsclinvar-annotationspart of
🧬 Genomic Reasoning Agent
LLM-driven agentic system for personal genomic variant interpretation
Overview
This project builds a multi-step reasoning agent that interprets personal genomic data from 23andMe against biomedical knowledge databases (ClinVar, GWAS Catalog, gnomAD). The agent is trained with GRPO (Group Relative Policy Optimization) using fully verifiable reward signals — no human labelers needed.
The core insight mirrors DeepSeek-R1's… See the full description on the dataset page: https://huggingface.co/datasets/huggingworld/clinvar-annotations.imbue_dearman_expert_annotations
Dataset Overview
This dataset was collected as part of the project "IMBUE: Improving Interpersonal Effectiveness through Simulation and Just-in-time Feedback with Human-Language Model Interaction."
We are releasing this dataset with the hope that it provides valuable opportunities for researchers to develop and evaluate new LLM-based tools for interpersonal skill training across a range of fields, including natural language processing, conversational AI, and computational… See the full description on the dataset page: https://huggingface.co/datasets/iwylin/imbue_dearman_expert_annotations.fineweb-edu-gemini-annotations-portuguese-regressionannotations
MATHWELL Human Annotation Dataset
The MATHWELL Human Annotation Dataset contains 5,084 synthetic word problems and answers generated by MATHWELL, a reference-free educational grade school math word problem generator released in MATHWELL: Generating Educational Math Word Problems Using Teacher Annotations, and comparison models (GPT-4, GPT-3.5, Llama-2, MAmmoTH, and LLEMMA) with expert human annotations for solvability, accuracy, appropriateness, and meets all criteria (MaC).… See the full description on the dataset page: https://huggingface.co/datasets/bryanchrist/annotations.TurkDoc-MT-MQM-Annotations
