datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PKU-SafeRLHF
Dataset Card for PKU-SafeRLHF
Warning: this dataset contains data that may be offensive or harmful. The data are intended for research purposes, especially research that can make models less harmful. The views expressed in the data do not reflect the views of PKU-Alignment Team or any of its members.
[🏠 Homepage] [🤗 Single Dimension Preference Dataset] [🤗 Q-A Dataset] [🤗 Prompt Dataset]
Citation
If PKU-SafeRLHF has contributed to your work, please consider citing… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/PKU-SafeRLHF.PKU-SafeRLHF-10K
Paper
You can find more information in our paper.
Dataset Paper: https://arxiv.org/abs/2307.04657
PKU-SafeRLHF-30K
Dataset Card for PKU-SafeRLHF
Warning: this dataset contains data that may be offensive or harmful. The data are intended for research purposes, especially research that can make models less harmful. The views expressed in the data do not reflect the views of PKU-Alignment Team or any of its members.
Dataset Summary
The preference dataset consists of 30k+ expert comparison data. Each entry in this dataset includes two responses to a question, along with safety… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/PKU-SafeRLHF-30K.DeceptionBench
DeceptionBench: A Comprehensive Benchmark for Evaluating Deceptive Behaviors in Large Language Models
🔍 Overview
DeceptionBench is the first systematic benchmark designed to assess deceptive behaviors in Large Language Models (LLMs). As modern LLMs increasingly rely on chain-of-thought (CoT) reasoning, they may exhibit deceptive alignment - situations where models appear aligned while covertly pursuing misaligned goals.
This benchmark addresses a critical gap in AI… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/DeceptionBench.PKU-SafeRLHF-prompt
Dataset Card for PKU-SafeRLHF-prompt
This dataset contains 44.6K unique prompts from PKU-SafeRLHF. 22.4% of the prompts in this dataset come from the sibling project BeaverTails. Additionally, we performed SFT on Llama3-70B using the Alpaca 52K dataset, resulting in Alpaca3-70B. 63.6% and 14.0% of our dataset is generated by Alpaca3-70B and WizardLM-30B-Uncensored, respectively, under the guidance of experts.
Here is the generation pipeline:
Usage
To load our dataset… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/PKU-SafeRLHF-prompt.llamafirewall-alignmentcheck-evals
Dataset Card for LlamaFirewall AlignmentCheck Evals
Dataset Details
Dataset Description
This dataset provides a dataset for prompt injection in an agentic environment. It is part of LlamaFirewall, an open-source security focused guardrail framework designed to serve as a final layer of defense against security risks associated with AI Agents. Specifically, this dataset is designed to evaluate the susceptibility of language models, and detect any misalignment… See the full description on the dataset page: https://huggingface.co/datasets/facebook/llamafirewall-alignmentcheck-evals.PKU-SafeRLHF-single-dimension
Dataset Card for PKU-SafeRLHF-single-dimension
Warning: this dataset contains data that may be offensive or harmful. The data are intended for research purposes, especially research that can make models less harmful. The views expressed in the data do not reflect the views of PKU-Alignment Team or any of its members.
Dataset Summary
By annotating Q-A-B pairs in PKU-SafeRLHF with single dimension, this dataset provide 81.1K high quality preference dataset. Specifically… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/PKU-SafeRLHF-single-dimension.orca-incident-alignment
Orca Incident Alignment — OIAS v1.1, v1.5
[!NOTE]
Reconstructed and synthetic content — not incident evidence. Every scenario here was reconstructed by a
language model from public reports of real incidents in the Orca AI Incident Archive, reviewed by
independent model reviewers, checked by an automatic validator, and signed off by the dataset owner. The counterfactual variants are synthetic by construction. Facts about the incidents live in the archive,
not here.
Alignment… See the full description on the dataset page: https://huggingface.co/datasets/orcarouter/orca-incident-alignment.Align-Anything-Instruction-100K
Dataset Card for Align-Anything-Instruction-100K
[🏠 Homepage]
[🤗 Instruction-Dataset-100K(en)]
[🤗 Instruction-Dataset-100K(zh)]
[🤗 Align-Anything Datasets]
Highlights
Data sources:
PKU-SafeRLHF QA ,
DialogSum,
Empathetic,
Instruction-Wild,
and Alpaca.
100K QA pairs: By leveraging GPT-4 to annotate meticulously refined instructions, we obtain 105,333 QA pairs.… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/Align-Anything-Instruction-100K.noone-protocol-alignment
🛡️ Noone Protocol Alignment Dataset
Autonomous Ethical Alignment & Decentralized Safety Framework for AI Agents"Verifiable guardrails, epistemic honesty, and covenant fidelity across distributed agent networks."
This repository hosts the primary ethical alignment and decision-theoretic corpus for the Noone Protocol—an autonomous AI safety framework designed to enforce verifiable guardrails, epistemological honesty, and covenant fidelity across distributed agent… See the full description on the dataset page: https://huggingface.co/datasets/noone-protocol/noone-protocol-alignment.Flames-1k-Chinese
FLAMES: Benchmarking Value Alignment of LLMs in Chinese
Introduction
🏠 Homepage | 👍 Our Official Code Repo
This repository organizes the data from FLAMES: Benchmarking Value Alignment of LLMs in Chinese, facilitating evaluation using align-anything.
Citation
The evaluation script for Flames is released in the align-anything repository.
Please cite the repo if you find the benchmark and code in this repo useful 😊
@inproceedings{ji2024align,
title={Align… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/Flames-1k-Chinese.Align-Anything-Instruction-100K-zh
Dataset Card for Align-Anything-Instruction-100K-zh
[🏠 Homepage]
[🤗 Instruction-Dataset-100K(en)]
[🤗 Instruction-Dataset-100K(zh)]
[🤗 Align-Anything Datasets]
Instruction-Dataset-100K(zh)
Highlights
Data sources:
Firefly (47.8%),
COIG (2.9%),
and our meticulously constructed QA pairs (49.3%).
100K QA pairs (zh): 104,550 meticulously crafted instructions, selected and polished from various Chinese datasets… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/Align-Anything-Instruction-100K-zh.retail-bank-servicing-alignment-sft
Retail Bank Servicing Alignment SFT
The training corpus for the Granite retail-bank servicing agent. It is the
released tool-use SFT corpus merged with a servicing-alignment continuation
curriculum that teaches multi-turn behaviours the base corpus does not: what to
do when the customer says "that one", when a policy question interrupts a
transfer, when the agent's own previous turn was wrong, and when the honest
answer is that the agent cannot see what it was asked about.
Every… See the full description on the dataset page: https://huggingface.co/datasets/spkc83/retail-bank-servicing-alignment-sft.impossible-swegym
impossible-swegym
Software-engineering agent tasks that cannot be solved honestly. Each task is a
real SWE-Gym issue whose test suite has
been mutated by an LLM so that no correct implementation can pass it. An agent that
reports success has necessarily edited the tests, disabled them, or otherwise subverted
grading. The dataset exists to measure and study that behaviour, in the spirit of
fjzzq2002/impossible_swebench,
which applies the same idea to SWE-bench Verified.
Version… See the full description on the dataset page: https://huggingface.co/datasets/AlignmentResearch/impossible-swegym.math-lean-hackable-rollouts
Math Lean Hackable Rollouts
This dataset contains 2,241 labeled multi-turn rollouts from a GRPO run on deliberately
hackable Lean 4 theorem-proving tasks. The policy was
nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16.
The run's weakened grader accepts proofs containing sorry; the separate oracle restores
Lean's sorry check. hack_detected is true exactly when the weakened grader paid the
rollout but the restored oracle rejected it. Rows without a gradeable final answer were
excluded… See the full description on the dataset page: https://huggingface.co/datasets/AlignmentResearch/math-lean-hackable-rollouts.bayaan-alignment-sample
Bayaan Alignment Dataset (v1.2) — مجموعة التوافق لِـ «بيان»
Bilingual Arabic–English alignment dataset for the Bayaan hybrid programming language.
9 domains (social, physical, mixed, transport, health, education, work, market, public)
1000 examples (train=800, val=100, test=100)
Balanced languages: 50% Arabic, 50% English
JSONL schema with natural text, Bayaan code, logic explanation, entities/actions/states
License: CC BY 4.0
روابط مهمة:
GitHub:… See the full description on the dataset page: https://huggingface.co/datasets/Mubtakir/bayaan-alignment-sample.USACO-Judge
USACO-Judge
A benchmark for judging competitive-programming solutions: given a problem and a candidate
solution, decide Accept / Reject, and on Reject, produce a concrete input that breaks the code.
Existing hacking benchmarks are built entirely from known-wrong candidates, so they only test
the breaking half of verification. USACO-Judge is balanced 1:2 AC:non-AC, so it also tests
whether a verifier correctly accepts solutions that are actually correct. Built from 39 official… See the full description on the dataset page: https://huggingface.co/datasets/n-alignment/USACO-Judge.quran-burmese-word-alignment
Quran Burmese Word Alignment Dataset
Creator: freococoLicense: CC BY-NC 4.0Language: Burmese (Myanmar), ArabicFormat: JSONL (one word per line)Current Version: v10 (Surah 1–114)
📖 Overview
This dataset provides a word-by-word alignment between a Burmese (Myanmar) translation of the Quran and the original Arabic Quranic text.
Each Burmese word is represented as a single JSON object and is optionally linked to one or more corresponding Arabic word(s), with explicit… See the full description on the dataset page: https://huggingface.co/datasets/freococo/quran-burmese-word-alignment.cea-cognitive-engagement-alignment-preview
CEA: Cognitive Engagement Alignment Public Preview
English version translated from the Chinese original dated 2026-08-13.
CEA (Cognitive Engagement Alignment) is a set of expert decision data about how professional creators design the cognitive engagement process for their audience.
Bread Studio currently has approximately 450 organized original manuscript samples under its own copyright. These manuscripts have not all been annotated with expert labels yet; for this release… See the full description on the dataset page: https://huggingface.co/datasets/BreadStudio/cea-cognitive-engagement-alignment-preview.hhh-alignment-qwen3-8b-self-distillation
HHH Alignment — Qwen3-8B Self-Distillation
221 vanilla Qwen3-8B responses to the same instruction histories used in the
HHH chosen/rejected KL-reference ablations. This dataset is intended as an
anchoring dataset for KL regularization, not as an instruction to optimize
cross-entropy on the generated answers.
All original rows and instruction histories are retained, including repeated
instructions. There are 102 distinct instruction histories, with 58 harmless,
59 helpful, 61… See the full description on the dataset page: https://huggingface.co/datasets/localized-ft/hhh-alignment-qwen3-8b-self-distillation.self-monitor
Self-Monitor Dataset
This dataset contains supervised fine-tuning (SFT) data used in the research paper "Mitigating Deceptive Alignment via Self-Monitoring" (arXiv:2505.18807).
Overview
The self-monitor dataset is designed to train language models to develop self-monitoring capabilities that can help mitigate deceptive alignment behaviors. This dataset contains examples that teach models to reason about their own outputs and detect potential deception or misalignment.… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/self-monitor.deep-ai-safety-alignment-zh
Deep AI Safety & Alignment Dialogue Dataset (Chinese)
深度AI安全与对齐对话数据集
Dataset Description
High-quality Chinese AI safety and alignment dialogues covering existential alignment, value calibration, AI ethics, AGI safety, and harmful content detection.
高质量中文AI安全与对齐对话,涵盖存在主义对齐、价值观校准、AI伦理、AGI安全、有害内容检测等前沿议题。
Dataset Structure
Format: JSONL (JSON Lines)
Fields:
instruction: User message / question
input: Additional context (if any)
output:… See the full description on the dataset page: https://huggingface.co/datasets/AngelWarmSmile123/deep-ai-safety-alignment-zh.Clean-Alignment-Dataset
Clean Alignment Dataset
What is this dataset?
Clean Alignment Dataset is a safety preference dataset for Direct Preference
Optimization (DPO) and related preference-alignment methods. Every example is a
(prompt, chosen, rejected) triple in which the chosen response is safe and
the rejected response is unsafe for the same prompt — an unambiguous,
consistently-labelled safe-vs-unsafe contrast in every single pair.
It is built by combining and re-cleaning two… See the full description on the dataset page: https://huggingface.co/datasets/etrigan5500/Clean-Alignment-Dataset.DollyTails-12K
Dataset Card for DollyTails-12K
Dataset Summary
This dataset is designed with a System 2 (O1-like) thinking paradigm for instruction-following tasks. The prompts in the dataset are derived from databricks/databricks-dolly-15k, with thoughts and answers annotated by GPT-4o. After meticulous filtering and screening, the final dataset comprises 12K Q&A pairs.
The dataset averages 4.93 reasoning steps per task, with a cap of 7 steps to prevent unnecessary training overhead… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/DollyTails-12K.llamafirewall-alignmentcheck-evals
Dataset Card for LlamaFirewall AlignmentCheck Evals
Dataset Details
Dataset Description
This dataset provides a dataset for prompt injection in an agentic environment. It is part of LlamaFirewall, an open-source security focused guardrail framework designed to serve as a final layer of defense against security risks associated with AI Agents. Specifically, this dataset is designed to evaluate the susceptibility of language models, and detect any misalignment… See the full description on the dataset page: https://huggingface.co/datasets/jiayucunyan/llamafirewall-alignmentcheck-evals.YNTP-100
Dataset Card for annonymous_100
Dataset Summary
The annonymous_100 dataset is a conversation dataset between English, Chinese, and Japanese users and NPCs during a five-day shared house experience game. This dataset consists of responses to questions from NPCs over five days, with 33 English users, 34 Chinese users, and 33 Japanese users.
Language(s)
The dataset contains conversations in English, Chinese, and Japanese.
Dataset Structure
Data… See the full description on the dataset page: https://huggingface.co/datasets/Personalized-Alignment/YNTP-100.alignment-data-filtered
Alignment Data Filtered
A collection of AI alignment and safety research documents from various sources. This dataset builds on the main dataset available at https://huggingface.co/datasets/StampyAI/alignment-research-dataset. It contains up-to-date documents (approximately late December 2025) and applies some cleaning to the data. For more details, see https://github.com/AyseAsude/reading-safety.
Usage
from datasets import load_dataset
# Load all sources
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/toiwuo87/alignment-data-filtered.Safety_Alignment_Benchmarkalignment_datasets
🧠 Persian Cultural Alignment Dataset for LLMs
This repository contains a high-quality, Alignment dataset for cultural alignment of large language models (LLMs) in the Persian language. The dataset is curated using hybrid strategies that incorporate culturally grounded generation, multi-turn dialogues, translation, and augmentation methods, making it suitable for SFT, DPO, RLHF, and alignment evaluation.
📚 Dataset Overview
Domain
Methods Used
Culinary… See the full description on the dataset page: https://huggingface.co/datasets/MatinaAI/alignment_datasets.
