datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
HSTLI_A-Dataset-of-Human-Semen-Time-Lapse-Images
HSTLI: A Dataset of Human Semen Time Lapse Images
Dataset Details
HSTLI contains 3,266 time-lapse microscopy videos of human sperm.Clips were recorded from two imaging modalities:
CASA system (Sperm Class Analyzer)
Optical microscope (Swift M10DB-MP + Fujifilm X-T30)
A subset of videos was manually annotated with bounding boxes around each visible sperm head.
The dataset supports detection, tracking and motility computation.
Total contents:
34… See the full description on the dataset page: https://huggingface.co/datasets/DFL-KamLab/HSTLI_A-Dataset-of-Human-Semen-Time-Lapse-Images.700k_Human_Preference_Dataset_FLUX_SD3_MJ_DALLE3
NOTE: A newer version of this dataset is available Imagen3_Flux1.1_Flux1_SD3_MJ_Dalle_Human_Preference_Dataset
Rapidata Image Generation Preference Dataset
This Dataset is a 1/3 of a 2M+ human annotation dataset that was split into three modalities: Preference, Coherence, Text-to-Image Alignment.
Link to the Coherence dataset: https://huggingface.co/datasets/Rapidata/Flux_SD3_MJ_Dalle_Human_Coherence_Dataset
Link to the Text-2-Image Alignment dataset:… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/700k_Human_Preference_Dataset_FLUX_SD3_MJ_DALLE3.Cabin-Human-Behavior-Dataset
全球最大的智能座舱多模态开源高质量数据集来啦!
一. 数据集摘要 (Dataset Summary)
「CyberData塞塔」智能座舱用户行为数据集是一个专为加速智能座舱感知算法开发而设计的高质量、程序化生成的图像数据集。随着 C-NCAP、EU GSR 等全球汽车安全法规对驾驶员监控系统 (DMS) 和乘客监控系统 (OMS) 提出更高要求,安全、合规、多样化的训练数据变得至关重要。本数据集通过合成方式,旨在解决真实世界数据采集面临的隐私风险、高昂成本和长尾场景覆盖不足等核心挑战。
该数据集包含 5,000 张 由 XAI Lab 自主研发的数据集生成引擎合成的高保真座舱内用户行为图像,每张图像都附带丰富的、100% 精确的标注信息。
核心特点:
丰富的场景多样性: 涵盖不同年龄、性别、种族和衣着风格的虚拟人模型,以及多种驾驶与乘坐行为(如使用手机、喝水、疲劳、手势)和面部表情。
专为座舱感知优化: 数据集可直接用于智能座舱端侧视觉模型,尤其是 DMS/OMS 算法的训练、微调与验证,帮助模型精准理解座舱内复杂的交互与状态。… See the full description on the dataset page: https://huggingface.co/datasets/OpenSparX/Cabin-Human-Behavior-Dataset.Flux_SD3_MJ_Dalle_Human_Alignment_Dataset
NOTE: A newer version of this dataset is available Imagen3_Flux1.1_Flux1_SD3_MJ_Dalle_Human_Alignment_Dataset
Rapidata Image Generation Alignment Dataset
This Dataset is a 1/3 of a 2M+ human annotation dataset that was split into three modalities: Preference, Coherence, Text-to-Image Alignment.
Link to the Coherence dataset: https://huggingface.co/datasets/Rapidata/Flux_SD3_MJ_Dalle_Human_Coherence_Dataset
Link to the Preference dataset:… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/Flux_SD3_MJ_Dalle_Human_Alignment_Dataset.text-to-speech-human-preferences-315k
Text-to-speech human preferences: 315K votes across 15 models
This gated dataset contains the evaluation record behind Datapoint Audio
Bench: 315,000 eligible pairwise votes comparing 15 text-to-speech
models in a complete round-robin over 300 English prompts. The prompt set
covers eight practical voice-agent categories, and every generated sample is
included as a typed audio record.
The source evaluation collected 357,651 completed responses. The published
benchmark excluded… See the full description on the dataset page: https://huggingface.co/datasets/datapointai/text-to-speech-human-preferences-315k.Cabin-Human-ABNORMAL-Behavior-Dataset
全球最大的智能座舱多模态开源高质量数据集来啦!
一. 数据集摘要 (Dataset Summary)
「CyberData塞塔」智能座舱用户行为数据集是一个专为加速智能座舱感知算法开发而设计的高质量、程序化生成的图像数据集。随着 C-NCAP、EU GSR 等全球汽车安全法规对驾驶员监控系统 (DMS) 和乘客监控系统 (OMS) 提出更高要求,安全、合规、多样化的训练数据变得至关重要。本数据集通过合成方式,旨在解决真实世界数据采集面临的隐私风险、高昂成本和长尾场景覆盖不足等核心挑战。
该数据集包含 5,000 张 由 XAI Lab 自主研发的数据集生成引擎合成的高保真座舱内用户行为图像,每张图像都附带丰富的、100% 精确的标注信息。
数据格式
数据集以JSON格式提供,包含以下字段:
image_id: 图像ID
image_path: 图像路径
category: 行为类别
tags: 行为标签
behaviors: 包含左右乘客行为描述的对象
left_passenger: 左侧乘客行为描述… See the full description on the dataset page: https://huggingface.co/datasets/OpenSparX/Cabin-Human-ABNORMAL-Behavior-Dataset.Human-Like-DPO-Dataset
Enhancing Human-Like Responses in Large Language Models
🤗 Models | 📊 Dataset | 📄 Paper
📢 The paper associated with this dataset has been accepted to the AAAI-26 Workshop on Personalization in the Era of Large Foundation Models (PerFM).
Human-Like-DPO-Dataset
This dataset was created as part of research aimed at improving conversational fluency and engagement in large language models. It is suitable for formats like Direct Preference Optimization (DPO) to guide… See the full description on the dataset page: https://huggingface.co/datasets/HumanLLMs/Human-Like-DPO-Dataset.Flux_SD3_MJ_Dalle_Human_Coherence_Dataset
NOTE: A newer version of this dataset is available: Imagen3_Flux1.1_Flux1_SD3_MJ_Dalle_Human_Coherence_Dataset
Rapidata Image Generation Coherence Dataset
This Dataset is a 1/3 of a 2M+ human annotation dataset that was split into three modalities: Preference, Coherence, Text-to-Image Alignment.
Link to the Preference dataset: https://huggingface.co/datasets/Rapidata/700k_Human_Preference_Dataset_FLUX_SD3_MJ_DALLE3
Link to the Text-2-Image Alignment dataset:… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/Flux_SD3_MJ_Dalle_Human_Coherence_Dataset.human-telemetry-driving-dataset-lite-version
Dataset Card for 15 Laps of 30Hz NGSIM-Style Telemetry
This is a Lite Version of a larger research dataset focusing on human driving signatures in high-fidelity simulations. It includes 15 full laps of telemetry captured at 30Hz within Unreal Engine 5, specifically formatted to match NGSIM standards.
Dataset Details
Dataset Description
This Lite Version dataset contains 15 laps of high-fidelity human driving telemetry. It is intended for researchers and… See the full description on the dataset page: https://huggingface.co/datasets/AtlasBuiltIt/human-telemetry-driving-dataset-lite-version.text-to-video-human-preferences-326k
Text-to-video human preferences: 326K votes across 15 models
This dataset contains the complete voting record behind the
Datapoint Video Bench
leaderboard: 325,520 validated pairwise votes — exactly 10 for each of
32,552 video pairs. The votes compare 15 text-to-video models
on 314 prompts built to stress motion, physics, and temporal
consistency, judged by 22,982 annotators in 187 countries.
Every generated clip is included as a full-resolution MP4.
Built on the Datapoint… See the full description on the dataset page: https://huggingface.co/datasets/datapointai/text-to-video-human-preferences-326k.my-human-datasettext-2-image-human-preferences-2m
Text-to-image human preferences: 2M votes across 30 models
This dataset contains the complete voting record behind the
Datapoint Image Bench
leaderboard: 2,161,160 validated pairwise votes — exactly 10 for each of
216,116 image pairs. The votes compare 30 text-to-image models in a complete
round-robin on 500 prompts, judged by annotators from over 200 countries.
Every vote includes the annotator's trust score at the time the vote was
cast.
Built on the Datapoint annotation… See the full description on the dataset page: https://huggingface.co/datasets/datapointai/text-2-image-human-preferences-2m.humandata-conf-labels-v3
실제 Franka 압축 가능성 라벨 (human_data v3)
각 순간의 모션을 더 빠르게 지나가도 되는지 VLM 으로 매긴 라벨.
만든 방법
판정기 gemini-3.8-flash (reasoning low), 온도 0
입력: 그 순간의 카메라 2장 + 지시문 + 계획된 액션에서 계산한 사실
출력: 문항을 1~5 등급으로. prompts/humandata_v3.txt 가 나간 전문 그대로다
SYSTEM 메시지를 쓰지 않는다. role: user 하나이고 가이던스가 본문 머리에 온다
청크 길이 16, stride 16. 아래 "중간 시점" 참고
게이트는 VLM 신뢰도 하나다. 접촉 열을 쓰지 않는다
실제 Franka 데모 세 벌 · 3,446 청크. 셋 다 같은 프롬프트로 만들었다
(전문·문항·부호·가중이 글자까지 같다).
데이터셋
청크
에피
fps
목적지
conf 평균
pnp_task
874
101… See the full description on the dataset page: https://huggingface.co/datasets/prehj/humandata-conf-labels-v3.human-faces-dataset-r1real-human-faces-data-sethuman-vs-Ai-generated-datasetDataset-PTEN_HUMAN
Description
This dataset contains signle site mutation of protein PTEN_HUMAN and the correspond mutation effect score from deep mutation scanning experiment.
Protein Format: AA sequence
Splits
traing: 3311
valid: 375
test: 410
Related paper
The dataset is from Deep generative models of genetic variation capture the effects of mutations.
Label
Label means mutation fitness score (protein stability) of each protein based on deep mutation scanning… See the full description on the dataset page: https://huggingface.co/datasets/SaProtHub/Dataset-PTEN_HUMAN.Benchmark_Dataset-Human_population_classification
Sumary
This dataset provides a benchmark for evaluating the model's ability to leverage richer genetic information from longer sequences to achieve more accurate inference.
Using data from the Human Pangenome Reference Consortium (BioProject ID: PRJNA730823), we designed a population classification task focusing on African, East Asian, and European population groups.
From samples' VCF file and the reference genome sequence, we generated sample pseudo-sequences.
Based on variant… See the full description on the dataset page: https://huggingface.co/datasets/BGI-HangzhouAI/Benchmark_Dataset-Human_population_classification.uiclip_human_data_hfimage-2-video-human-preferences-large
I2V Human Preferences (Large)
Human preference dataset for image-to-video (I2V) generation quality. Each row contains a reference image, two generated videos (one from Pika and one from CogVideoX), and 10 human preference annotations aggregated via majority vote.
This is the large (3,000-row) subset — the complete dataset. See also: small (1,000 rows), medium (2,000 rows).
Dataset Summary
Metric
Value
Total rows
3,000
Annotations per row
10
Total… See the full description on the dataset page: https://huggingface.co/datasets/datapointai/image-2-video-human-preferences-large.ConvoDrift_human_eval_conversational_dataset
ConvoDrift Human Evaluation Data
This folder contains human evaluations of ConvoDrift conversations from three
annotators. Each conversation has six prompt-response pairs, quality ratings
for up to Q1-Q8, and optional corrections to drift and direction labels.
Files
Configuration
Records
Description
annotator_1
500
Complete evaluation file for Annotator 1
annotator_2
500
Complete evaluation file for Annotator 2
annotator_3
500
Complete evaluation… See the full description on the dataset page: https://huggingface.co/datasets/Vihindi-K/ConvoDrift_human_eval_conversational_dataset.75-percent-human-dataset-opt-mistakeuiclip_human_data-paired_hfHuman-Like-DPO-Dataset-kolbf-human-datatext-2-video-human-preferences-motion-v2-large
Human Preferences for AI-Generated Video: Motion Quality v2 (large)
115,732 pairwise human preference labels comparing 4 frontier video generation models on human motion across 3 quality dimensions, collected from real annotators via Datapoint AI.
This is an expanded version of the motion quality dataset with 417 unique prompts (up from 60) and 11 motion categories (up from 6).
Why This Dataset
Video generation models are improving fast, but evaluating human motion… See the full description on the dataset page: https://huggingface.co/datasets/datapointai/text-2-video-human-preferences-motion-v2-large.shamimhasan8_ai-vs-human-text-dataset
AI vs Human Text Dataset
Mirror of the Kaggle dataset shamimhasan8/ai-vs-human-text-dataset by Shamim Hasan, released under MIT. All credit goes to the original author; please cite and link the Kaggle page when using this data.
Detect AI-generated vs. human-written text samples
License
MIT License, Copyright (c) Shamim Hasan. The full license text is in LICENSE and applies to all files in this repository.
Original description (from Kaggle)
This… See the full description on the dataset page: https://huggingface.co/datasets/haoxianc/shamimhasan8_ai-vs-human-text-dataset.Dataset-UBC9_HUMAN
Description
This dataset contains signle site mutation of protein UBC9_HUMAN and the correspond mutation effect score from deep mutation scanning experiment.
Protein Format: AA sequence
Splits
traing: 2169
valid: 265
test: 287
Related paper
The dataset is from Deep generative models of genetic variation capture the effects of mutations.
Label
Label means mutation effect score of each protein based on deep mutation scanning experiment.
The wild‐type… See the full description on the dataset page: https://huggingface.co/datasets/SaProtHub/Dataset-UBC9_HUMAN.Human-Like-DPO-Dataset_deduplicated_and_duplicatedhuman-like-sft-dataset-split
