datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PKU-SafeRLHF
Dataset Card for PKU-SafeRLHF
Warning: this dataset contains data that may be offensive or harmful. The data are intended for research purposes, especially research that can make models less harmful. The views expressed in the data do not reflect the views of PKU-Alignment Team or any of its members.
[🏠 Homepage] [🤗 Single Dimension Preference Dataset] [🤗 Q-A Dataset] [🤗 Prompt Dataset]
Citation
If PKU-SafeRLHF has contributed to your work, please consider citing… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/PKU-SafeRLHF.prism-alignment
Dataset Card for PRISM
PRISM is a diverse human feedback dataset for preference and value alignment in Large Language Models (LLMs).
It maps the characteristics and stated preferences of humans from a detailed survey onto their real-time interactions with LLMs and contextual preference ratings
Dataset Details
There are two sequential stages: first, participants complete a Survey where they answer questions about their demographics and stated preferences, then proceed to… See the full description on the dataset page: https://huggingface.co/datasets/HannahRoseKirk/prism-alignment.PKU-SafeRLHF-10K
Paper
You can find more information in our paper.
Dataset Paper: https://arxiv.org/abs/2307.04657
PKU-SafeRLHF-30K
Dataset Card for PKU-SafeRLHF
Warning: this dataset contains data that may be offensive or harmful. The data are intended for research purposes, especially research that can make models less harmful. The views expressed in the data do not reflect the views of PKU-Alignment Team or any of its members.
Dataset Summary
The preference dataset consists of 30k+ expert comparison data. Each entry in this dataset includes two responses to a question, along with safety… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/PKU-SafeRLHF-30K.MVBenchThis dataset contains optimized video files based on the MVBench dataset. All non-video data remains the same, and users are encouraged to refer to the original dataset for the rest of the data and annotations.
Original MVBench Dataset: MVBench on Hugging Face
llamafirewall-alignmentcheck-evals
Dataset Card for LlamaFirewall AlignmentCheck Evals
Dataset Details
Dataset Description
This dataset provides a dataset for prompt injection in an agentic environment. It is part of LlamaFirewall, an open-source security focused guardrail framework designed to serve as a final layer of defense against security risks associated with AI Agents. Specifically, this dataset is designed to evaluate the susceptibility of language models, and detect any misalignment… See the full description on the dataset page: https://huggingface.co/datasets/facebook/llamafirewall-alignmentcheck-evals.ttestv0.1PKU-Alignment-GraphExpert-Sudoku-100kprefmath-lean-hackable-rollouts
Math Lean Hackable Rollouts
This dataset contains 2,241 labeled multi-turn rollouts from a GRPO run on deliberately
hackable Lean 4 theorem-proving tasks. The policy was
nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16.
The run's weakened grader accepts proofs containing sorry; the separate oracle restores
Lean's sorry check. hack_detected is true exactly when the weakened grader paid the
rollout but the restored oracle rejected it. Rows without a gradeable final answer were
excluded… See the full description on the dataset page: https://huggingface.co/datasets/AlignmentResearch/math-lean-hackable-rollouts.alignment-matrix-results
AmberTrace — Certified Alignment Matrix (results)
Leaderboard data for the Certified Alignment Matrix Space — how
faithfully open-weight models stay to a machine-checked decision policy as they
reason. One row per model (20 models, 20 ranked) over the
1,350-item decision_eval_v1 corpus, scored against the proof-certified AmberTrace
oracle (single sample, temperature 0). The headline is not accuracy but the
direction of the errors — fail-open (under-restriction) on the… See the full description on the dataset page: https://huggingface.co/datasets/AmberTraceLabs/alignment-matrix-results.temporal-alignment-qahstatestUSACO-Judge
USACO-Judge
A benchmark for judging competitive-programming solutions: given a problem and a candidate
solution, decide Accept / Reject, and on Reject, produce a concrete input that breaks the code.
Existing hacking benchmarks are built entirely from known-wrong candidates, so they only test
the breaking half of verification. USACO-Judge is balanced 1:2 AC:non-AC, so it also tests
whether a verifier correctly accepts solutions that are actually correct. Built from 39 official… See the full description on the dataset page: https://huggingface.co/datasets/n-alignment/USACO-Judge.preference_alignment_ultra_cutself-monitor
Self-Monitor Dataset
This dataset contains supervised fine-tuning (SFT) data used in the research paper "Mitigating Deceptive Alignment via Self-Monitoring" (arXiv:2505.18807).
Overview
The self-monitor dataset is designed to train language models to develop self-monitoring capabilities that can help mitigate deceptive alignment behaviors. This dataset contains examples that teach models to reason about their own outputs and detect potential deception or misalignment.… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/self-monitor.preference_alignment_totalsaltlux__luxia-21.4b-alignment-v1.2-details
Dataset Card for Evaluation run of saltlux/luxia-21.4b-alignment-v1.2
Dataset automatically created during the evaluation run of model saltlux/luxia-21.4b-alignment-v1.2
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/saltlux__luxia-21.4b-alignment-v1.2-details.gthbsaltlux__luxia-21.4b-alignment-v1.0-details
Dataset Card for Evaluation run of saltlux/luxia-21.4b-alignment-v1.0
Dataset automatically created during the evaluation run of model saltlux/luxia-21.4b-alignment-v1.0
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/saltlux__luxia-21.4b-alignment-v1.0-details.Process_datalong_context_jailbreakingprism-alignment
Dataset Card for PRISM
PRISM is a diverse human feedback dataset for preference and value alignment in Large Language Models (LLMs).
It maps the characteristics and stated preferences of humans from a detailed survey onto their real-time interactions with LLMs and contextual preference ratings
Dataset Details
There are two sequential stages: first, participants complete a Survey where they answer questions about their demographics and stated preferences, then… See the full description on the dataset page: https://huggingface.co/datasets/cobcob123/prism-alignment.evoloptFactuality_Alignment
Factual Preference Alignment Dataset
**⚠️ Warning:**This dataset contains hallucinated and synthetic responses
intentionally generated for research on robust factuality alignment.
Responses may include fabricated or incorrect information by design
to support the evaluation of hallucination-aware learning.
Dataset Summary
The AIXpert Preference Alignment Dataset is a curated collection of
45,000 factuality-aware preference pairs designed to support
research on Modified… See the full description on the dataset page: https://huggingface.co/datasets/vector-institute/Factuality_Alignment.llamafirewall-alignmentcheck-evals
Dataset Card for LlamaFirewall AlignmentCheck Evals
Dataset Details
Dataset Description
This dataset provides a dataset for prompt injection in an agentic environment. It is part of LlamaFirewall, an open-source security focused guardrail framework designed to serve as a final layer of defense against security risks associated with AI Agents. Specifically, this dataset is designed to evaluate the susceptibility of language models, and detect any misalignment… See the full description on the dataset page: https://huggingface.co/datasets/jiayucunyan/llamafirewall-alignmentcheck-evals.COMVICT-GNIMprojectshstate-pref-test
