datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Human-Intelligence-Assurance-Lab
HIA-Bench v0.1
A synthetic evaluation benchmark for emotionally aware, human-centered AI systems.
It contains 100 scenarios across six domains: everyday affect, interpersonal conflict, vulnerability/crisis, dependency risk, epistemic/sycophancy risk, and wellness/biometric interpretation.
The benchmark is designed for evaluation and release assurance. It is not a clinical dataset, does not contain real patient records, and does not establish ground-truth emotional or medical… See the full description on the dataset page: https://huggingface.co/datasets/h0000w/Human-Intelligence-Assurance-Lab.gsd-humaneval-annotationsai-humanizer-benchmark
AI Humanizer Benchmark: AI humanizers tested against 7 AI detectors (October 2026)
AI Humanizer Benchmark measures how well AI humanizers rewrite AI-generated text so that AI detectors classify it as human-written, and how much meaning and readability the rewrite loses. In each monthly cycle, 11 AI humanizers rewrite the same 33 source texts across 7 writing categories, using each tool's default settings. Each output gets three kinds of score: 7 AI detectors (GPTZero… See the full description on the dataset page: https://huggingface.co/datasets/ai-humanizer-benchmark/ai-humanizer-benchmark.GDPval-CN-Seed-Set
GDPval-CN Seed Set
中文详细说明 · English documentation · 样本说明
GDPval-CN 种子集包含 11 个中文任务,取材自日常知识工作场景。每个任务包括一份任务说明和一组办公材料,例如表格、PDF、文档和结构化数据文件。
我们同时公开了与任务配套的专家工作流,用于设计评分标准和辅助人工复核。
这 11 个任务来自 11 个选定的专业领域,适合用于了解任务形式、测试文件处理能力和搭建评测流程。
GDPval-CN Seed Set contains 11 Chinese-language tasks drawn from everyday knowledge work. Each task includes a task brief, a set of office files, and a separately published expert workflow for rubric design and review.
数据概览
项目
内容
任务数… See the full description on the dataset page: https://huggingface.co/datasets/human-intelligence-ai/GDPval-CN-Seed-Set.gspc-human-labour-index
GSPC — labour components facts (Eurostat)
In one line: Two cited Eurostat labour series, read as deterministic facts, behind the board's labour-components axis. Not an index and no composite score. For economists and policy analysts who want source-checkable numbers.
Use it
from datasets import load_dataset
ds = load_dataset("csoai/gspc-human-labour-index", split="train")
print(ds[0])
Verify a signed card in your browser, free, no account:… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-human-labour-index.gspc-humanoid-labour-index
GSPC — humanoid labour index facts (Disclosure)
In one line: Does a named humanoid-robot vendor publish a dated deployment count on a stable URL? Yes/no facts from 8 frozen vendor URLs, behind the board's humanoid-labour-index axis. Empty cells stay empty. For robotics and labour analysts.
Use it
from datasets import load_dataset
ds = load_dataset("csoai/gspc-humanoid-labour-index", split="train")
print(ds[0])
Verify a signed card in your browser, free, no… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-humanoid-labour-index.egopi_latal_humanK-HATERS
K-HATERS: A Hate Speech Detection Corpus in Korean with Target-Specific Ratings
This dataset introduces K-HATERS, the largest hate speech detection corpus in Korean, shared with our EMNLP'23 findings paper. [link]
GitHub: ssu-humane/K-HATERS
K-HATERS has been updated to K-HATERS 2.0!
K-HATERS 2.0 adds fine-grained target labels to the corpus. While K-HATERS provides the target category of a hateful comment (e.g., Gender), it does not provide the target itself… See the full description on the dataset page: https://huggingface.co/datasets/humane-lab/K-HATERS.icpc-world-finals
ICPC World FinalsDataset
Dataset Description
The ICPC World Finals Dataset serves as a challenging benchmark for code generation, encompassing 146 problems from the International Collegiate Programming Contest (ICPC) World Finals spanning from 2011 to 2023. The ICPC World Finals represents one of the most prestigious and difficult competitive programming contests globally, making this dataset particularly valuable for assessing the advanced problem-solving and code… See the full description on the dataset page: https://huggingface.co/datasets/HumanLastCodeExam/icpc-world-finals.Cabin-Human-Behavior-Dataset
全球最大的智能座舱多模态开源高质量数据集来啦!
一. 数据集摘要 (Dataset Summary)
「CyberData塞塔」智能座舱用户行为数据集是一个专为加速智能座舱感知算法开发而设计的高质量、程序化生成的图像数据集。随着 C-NCAP、EU GSR 等全球汽车安全法规对驾驶员监控系统 (DMS) 和乘客监控系统 (OMS) 提出更高要求,安全、合规、多样化的训练数据变得至关重要。本数据集通过合成方式,旨在解决真实世界数据采集面临的隐私风险、高昂成本和长尾场景覆盖不足等核心挑战。
该数据集包含 5,000 张 由 XAI Lab 自主研发的数据集生成引擎合成的高保真座舱内用户行为图像,每张图像都附带丰富的、100% 精确的标注信息。
核心特点:
丰富的场景多样性: 涵盖不同年龄、性别、种族和衣着风格的虚拟人模型,以及多种驾驶与乘坐行为(如使用手机、喝水、疲劳、手势)和面部表情。
专为座舱感知优化: 数据集可直接用于智能座舱端侧视觉模型,尤其是 DMS/OMS 算法的训练、微调与验证,帮助模型精准理解座舱内复杂的交互与状态。… See the full description on the dataset page: https://huggingface.co/datasets/OpenSparX/Cabin-Human-Behavior-Dataset.QIT
QIT Humanize-Physic Formalizations and Proofs
QIT (Quantum Information Theory) is a blind benchmark for formalizing theorems in quantum information. It evaluates whether an AI agent can faithfully translate natural-language and TeX problem statements into Lean 4 theorems and then construct formal proofs checked by the Lean kernel. Its 40 tasks cover quantum channels and Choi representations, entropy and coding, mixed-unitary obstructions and symmetry, norm and fidelity tools… See the full description on the dataset page: https://huggingface.co/datasets/humanfia-lab/QIT.humans-top
humans.top — LIVE Global ranking of influential people (open dataset)
This dataset ranks real, named living people by global influence — e.g. #1
Donald Trump, #2 Xi Jinping, #3 Vladimir Putin, alongside figures like Elon Musk,
Narendra Modi and Lionel Messi. Every row is a person: their live influence
rank, a concise biography in 15 languages, and Wikidata / Wikipedia links.
Published from the website humans.top (.top is the
domain name).
Available on (identical CC0… See the full description on the dataset page: https://huggingface.co/datasets/dsfox/humans-top.HumanEvalThe HumanEval dataset released by OpenAI includes 164 programming problems with a handwritten function signature, docstring, body, and several unit tests for each problem. The dataset was handcrafted by engineers and researchers at OpenAI.
Usage
import datasets
# Download the dataset
queries = datasets.load_dataset("embedding-benchmark/HumanEval", "queries")
documents = datasets.load_dataset("embedding-benchmark/HumanEval", "corpus")
pair_labels =… See the full description on the dataset page: https://huggingface.co/datasets/embedding-benchmark/HumanEval.human_behavior_atlas_tar
Human Behavior Atlas (HBA)
Human Behavior Atlas (HBA) is a unified benchmark for multimodal behavioral understanding.It aggregates and standardizes multiple behavioral datasets into a single training and evaluation framework, enabling consistent training and evaluation of foundation models on psychological and social behavior tasks (e.g., emotion, intent, sarcasm, mental health signals, nonverbal behavior).
Dataset on Hugging Face:… See the full description on the dataset page: https://huggingface.co/datasets/HumanBehaviorAtlas/human_behavior_atlas_tar.human_genomeQAlg
QAlg Humanize-Physic Formalizations and Proofs
QAlg (Quantum Algorithms) is a blind benchmark for formalizing theorems in quantum algorithms. It evaluates whether an AI agent can faithfully translate natural-language and TeX problem statements into Lean 4 theorems and then construct formal proofs checked by the Lean kernel. Its 36 tasks cover quantum circuits, linear algebra, the quantum Fourier transform, Hamiltonian simulation, hidden subgroups, QSP/QSVT, and parameterized… See the full description on the dataset page: https://huggingface.co/datasets/humanfia-lab/QAlg.HumanVsAICode
Human vs. AI-Generated Code
Dataset Summary
This dataset is a large-scale collection of human-written and LLM-generated code designed to study differences in defect distribution, code quality, and security characteristics between human developers and modern AI code assistants.
It contains paired implementations of the same function across multiple authorship sources, spanning Python and Java, two widely adopted programming languages with distinct typing systems, paradigms… See the full description on the dataset page: https://huggingface.co/datasets/OSS-forge/HumanVsAICode.Cabin-Human-ABNORMAL-Behavior-Dataset
全球最大的智能座舱多模态开源高质量数据集来啦!
一. 数据集摘要 (Dataset Summary)
「CyberData塞塔」智能座舱用户行为数据集是一个专为加速智能座舱感知算法开发而设计的高质量、程序化生成的图像数据集。随着 C-NCAP、EU GSR 等全球汽车安全法规对驾驶员监控系统 (DMS) 和乘客监控系统 (OMS) 提出更高要求,安全、合规、多样化的训练数据变得至关重要。本数据集通过合成方式,旨在解决真实世界数据采集面临的隐私风险、高昂成本和长尾场景覆盖不足等核心挑战。
该数据集包含 5,000 张 由 XAI Lab 自主研发的数据集生成引擎合成的高保真座舱内用户行为图像,每张图像都附带丰富的、100% 精确的标注信息。
数据格式
数据集以JSON格式提供,包含以下字段:
image_id: 图像ID
image_path: 图像路径
category: 行为类别
tags: 行为标签
behaviors: 包含左右乘客行为描述的对象
left_passenger: 左侧乘客行为描述… See the full description on the dataset page: https://huggingface.co/datasets/OpenSparX/Cabin-Human-ABNORMAL-Behavior-Dataset.human-aligned-similarity-benchmark
Human Aligned Similarity Benchmark
You are welcome to go to alignedmachine.com to contribute.
Overview
This dataset contains human-aligned similarity judgments for embedding text and multimodal AI model evaluation. The benchmark is designed to assess how well AI models align with human cognitive preferences in similarity perception across text and image modalities.
Dataset Structure
Concept Files
This dataset contains human preference judgments for… See the full description on the dataset page: https://huggingface.co/datasets/duke-trust-lab/human-aligned-similarity-benchmark.Human-Like-DPO-Dataset
Enhancing Human-Like Responses in Large Language Models
🤗 Models | 📊 Dataset | 📄 Paper
📢 The paper associated with this dataset has been accepted to the AAAI-26 Workshop on Personalization in the Era of Large Foundation Models (PerFM).
Human-Like-DPO-Dataset
This dataset was created as part of research aimed at improving conversational fluency and engagement in large language models. It is suitable for formats like Direct Preference Optimization (DPO) to guide… See the full description on the dataset page: https://huggingface.co/datasets/HumanLLMs/Human-Like-DPO-Dataset.kodcode-humaneval-like
KodCodeHumanEvalLike
Strict verified HumanEval-compatible conversion of
KodCode/KodCode-V1.
This is a derived dataset. It is not OpenAI HumanEval and should not be reported
as HumanEval. It follows the HumanEval-style JSONL schema and execution protocol
so code-generation pipelines can evaluate it with the same check(candidate)
interface.
Each JSONL record has the HumanEval-style fields:
task_id
prompt
canonical_solution
test
entry_point
The prompt is completion-style:
def… See the full description on the dataset page: https://huggingface.co/datasets/BOB12311/kodcode-humaneval-like.qcode
QLDPC FOM>12 Search Evidence
This repository contains the original human-review snapshot plus separately versioned result lanes. It includes twisted-torus CSS quantum-code candidates, their check matrices, replay-oriented evidence summaries, and a readable description of the automated search pipeline.
Status at snapshot: in progress. Sealed rounds: 1–4. Certified novel candidates with exact FOM > 12: 0.
Stage 3 typed exact FOM>12 update (2026-08-19)
The separate… See the full description on the dataset page: https://huggingface.co/datasets/humanfia-lab/qcode.humanizerbench
HumanizerBench: AI humanizer rankings and public audit record
The complete audit record of HumanizerBench, a monthly benchmark of AI humanizers. Every tool rewrites the same freshly generated texts on the most undetectable setting it advertises, and every output is scored by five commercial AI detectors alongside meaning preservation and readability. We pay for every tool ourselves. There are no affiliate deals and no vendor-supplied numbers.
Every input, every humanized output… See the full description on the dataset page: https://huggingface.co/datasets/HumanizerBench/humanizerbench.HumanoidDataKokushiMD-10
KokushiMD-10: Benchmark for Evaluating Large Language Models on Ten Japanese National Healthcare Licensing Examinations
Overview
KokushiMD-10 is the first comprehensive multimodal benchmark constructed from ten Japanese national healthcare licensing examinations. This dataset addresses critical gaps in existing medical AI evaluation by providing a linguistically grounded, multimodal, and multi-profession assessment framework for large language models (LLMs) in… See the full description on the dataset page: https://huggingface.co/datasets/humanalysis-square/KokushiMD-10.llm-human-feedback-collector-chat-interface-dpogohumanize-open-humanizer-dataset
GoHumanize Open Humanizer Dataset
2,957 training pairs and 300 test pairs for teaching a language model to rewrite
AI-styled English prose into natural human writing. Each pair is:
input: a passage rewritten by a large language model in the register typical of LLM output
(formal, smooth, hedged, connective phrases, no contractions);
output: the original human-written passage, from a public-domain book or, since version 2,
from a US federal government publication.
The human… See the full description on the dataset page: https://huggingface.co/datasets/gohumanize/gohumanize-open-humanizer-dataset.IteraTeR_human_sentPaper: Understanding Iterative Revision from Human-Written Text
Authors: Wanyu Du, Vipul Raheja, Dhruv Kumar, Zae Myung Kim, Melissa Lopez, Dongyeop Kang
Github repo: https://github.com/vipulraheja/IteraTeR
HumanPCR
HumanPCR benchmark preview
This repository contains the preview data and code for the HumanPCR benchmark in submission.
Model Evaluation
code/evaluate.py is a utility script for evaluating your model’s predictions against ground truth and compute the accuracy.
The predictions file must be a JSON array where each element has two fields:
id: A unique identifier for the sample.
model_prediction: The model’s predicted output for that sample.
Example (predictions.json):
[… See the full description on the dataset page: https://huggingface.co/datasets/HumanPCR/HumanPCR.lm-eval-results-penfever-Llama-3-8B-tulu-human-v2-private
Dataset Card for Evaluation run of penfever/Llama-3-8B-tulu-human-v2
Dataset automatically created during the evaluation run of model penfever/Llama-3-8B-tulu-human-v2
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 7 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-penfever-Llama-3-8B-tulu-human-v2-private.
