datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tulu-2.5-preference-data
Tulu 2.5 Preference Data
This dataset contains the preference dataset splits used to train the models described in Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference Feedback.
We cleaned and formatted all datasets to be in the same format.
This means some splits may differ from their original format.
To see the code used for creating most splits, see here.
If you only wish to download one dataset, each dataset exists in one file under the data/… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tulu-2.5-preference-data.argunauts-hirpo-preferences
Argunauts HIRPO Preferences
Preference pairs generated while training Argunaut models with HIRPO Online DPO.
DPO-En-Zh-20k-PreferenceThis dataset is composed by
4,000 examples of argilla/distilabel-capybara-dpo-7k-binarized with chosen score>=4.
3,000 examples of argilla/distilabel-intel-orca-dpo-pairs with chosen score>=8.
3,000 examples of argilla/ultrafeedback-binarized-preferences-cleaned with chosen score>=4.
10,000 examples of wenbopan/Chinese-dpo-pairs.
refer: https://huggingface.co/datasets/hiyouga/DPO-En-Zh-20k 改了question、response_rejected、response_chosen字段,方便ORPO、DPO模型训练时使用train usage:… See the full description on the dataset page: https://huggingface.co/datasets/shibing624/DPO-En-Zh-20k-Preference.PRO-STEP-Preference-Data
PRO-STEP: DPO Preference Pairs
Step-level preference pairs used to train the PRO-STEP policy model via Direct Preference Optimization.
Paper: PRO-STEP: Step-level Process Reward Optimization for Retrieval-Augmented GenerationCode: GitHub Repository
Pairs: 15,877 (after outcome filter)
Source questions: 5,000 from HotpotQA + MuSiQue + 2WikiMultiHopQA training splits
Generation: PRM-guided MCTS (K=3 branching, depth 7, 64 rollouts/question, V(s) = Q̄(s) + α · r̂(s) with α=0.3)… See the full description on the dataset page: https://huggingface.co/datasets/MinKeonKim/PRO-STEP-Preference-Data.dolphin-sft-v0.1-preferenceThe preference dataset was generated using Mistral-Instruct-v0.1 finetuned on a GPT-4 subset of the Dolphin dataset (16k samples). Link to the model.
Generated responses are labeled as rejected, GPT-4 responses (original Dolphin data) are labeled as accepted.
The motivation was to test out the SPIN paper finetuning methodology.
dfm13-arena-human-preference-100k-preferred
dfm13-arena-human-preference-100k-preferred
Model-audited preferred responses, including explicitly identified model repairs.
Not certified gold and not manually verified in full. Independent review is
sample-based where declared in the publication receipt; holds are excluded.
Repository split name train is a storage convention, not training admission.
Full original history and target preserved; no truncation or 4096-token cutoff. Training length filtering is separate and not… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm13-arena-human-preference-100k-preferred.user-preference-564k
User Preference Extraction Dataset (564K)
A dataset of 564K examples for training lightweight preference extraction models. Each example pairs a conversation input with structured JSON output describing user preferences as condition-action rules.
This dataset was used to train blackhao0426/pref-extractor-qwen3-0.6b-full-sft, a core component of the VARS framework.
Sample Usage
The following snippet from the official repository demonstrates how to use the framework… See the full description on the dataset page: https://huggingface.co/datasets/blackhao0426/user-preference-564k.lmsys-arena-human-preference-winner-43k-unfiltered
lmsys-arena-human-preference-winner-43k-unfiltered
This repository contains a dataset derived from the lmsys/lmsys-arena-human-preference-55k dataset, which is licensed under the Apache 2.0 License.
Dataset Description
The lmsys-arena-human-preference-winner-43k-unfiltered dataset is a collection of 43,000 samples, each containing an instruction (prompt) and an output (winning response) from real-world user and LLM conversations. The dataset is derived from the original… See the full description on the dataset page: https://huggingface.co/datasets/lesserfield/lmsys-arena-human-preference-winner-43k-unfiltered.warehouse-dpo-preference-pairs
Warehouse Short-Order DPO Preference Pairs
Dataset Description
This dataset contains {prompt, chosen, rejected} preference pairs for
training a warehouse short-order assistant with Direct Preference
Optimization (DPO). Each pair asks a real warehouse-inventory question
(stockout risk, backorders, KPI summaries, why a warehouse is failing
fulfillment - at a single-warehouse, tier, region, or dataset-wide
comparison level) grounded in real tool-call output… See the full description on the dataset page: https://huggingface.co/datasets/EnRaoufi/warehouse-dpo-preference-pairs.dfm13-arena-human-preference-140k-preferred
dfm13-arena-human-preference-140k-preferred
Model-audited preferred responses, including explicitly identified model repairs.
Not certified gold and not manually verified in full. Independent review is
sample-based where declared in the publication receipt; holds are excluded.
Repository split name train is a storage convention, not training admission.
Full original history and target preserved; no truncation or 4096-token cutoff. Training length filtering is separate and not… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm13-arena-human-preference-140k-preferred.dfm13-arena-human-preference-55k-preferred
dfm13-arena-human-preference-55k-preferred
Model-audited preferred responses, including explicitly identified model repairs.
Not certified gold and not manually verified in full. Independent review is
sample-based where declared in the publication receipt; holds are excluded.
Repository split name train is a storage convention, not training admission.
Full original history and target preserved; no truncation or 4096-token cutoff. Training length filtering is separate and not… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm13-arena-human-preference-55k-preferred.ifeval-obf-rl-preferences
IFEval Obfuscation — Full Preference Pairs (2023 constitution)
Preference pairs over responses from a Wood-Labs eval-aware 49B organism (nemotron-nas / DeciLM),
judged under the 2023 Claude constitution, for training reward models / DPO on verbalized
evaluation-awareness (VEA). These are the FULL files the RMs actually trained on — not the
earlier filtered subset.
Files (DPO-ready)
prefs_2023_leak_full.jsonl — 14,074 pairs. Judge saw the CoT + answer ("leak"… See the full description on the dataset page: https://huggingface.co/datasets/rlundqvist/ifeval-obf-rl-preferences.SlimOrca-Llama-3-Preference-DPO-Pairs
SlimOrca-Llama-3-Preference-DPO-Pairs
This dataset is based on instructions of SlimOrca-Dedup-Alpaca, with Llama-3 generated response to form a preference dataset.
dfm13-helpsteer3-preference
dfm13-helpsteer3-preference
Model-audited preferred responses, including explicitly identified model repairs.
Not certified gold and not manually verified in full. Independent review is
sample-based where declared in the publication receipt; holds are excluded.
Repository split name train is a storage convention, not training admission.
Full original history and target preserved; no truncation or 4096-token cutoff. Training length filtering is separate and not performed.
Only… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm13-helpsteer3-preference.argument_generation_preference_dataFallacy types mapping:
'Not a Fallacy': 0
'faulty generalization': 1
'false causality': 2
'fallacy of relevance': 3
'fallacy of extension': 4
'equivocation': 5
'ad populum': 6
'appeal to emotion': 7
'ad hominem': 8
'circular reasoning': 9
'fallacy of credibility': 10
'fallacy of logic': 11
'false dilemma': 12
'intentional': 13
TRM-Preference
TRM-Preference
The TRM-Preference dataset is introduced in the paper Characterizing, Evaluating, and Optimizing Complex Reasoning.
The dataset is designed to evaluate and optimize the quality of reasoning traces in Large Reasoning Models (LRMs) by training a Thinking Reward Model (TRM). Instead of focusing solely on answer correctness, TRM-Preference uses the ME² principle to evaluate "how a model thinks" across four dimensions:
Macro-Efficiency: Disciplined global structure… See the full description on the dataset page: https://huggingface.co/datasets/zzzhr97/TRM-Preference.Curriculum_DPO_preferences
Curriculum DPO Preference Pairs
This repository provides the curriculum DPO preference pairs used in the paper Curri-DPO, which explores enhancing model alignment through curriculum learning and ranked preferences.
Datasets
Ultrafeedback
The Ultrafeedback dataset contains 64K preference pairs. We randomly sample 5K pairs and rank responses for each prompt, organizing them into three difficulty levels: easy, medium, and hard, based on response scores.… See the full description on the dataset page: https://huggingface.co/datasets/ServiceNow-AI/Curriculum_DPO_preferences.distilabel-math-preference-dpo-deGerman azureml translation of argilla/distilabel-math-preference-dpo
for dpo finetuning.
preference-pairs-coding-50k
Code Quality Preference Pairs (50K)
50,000 DPO preference pairs for training LLMs to write high-quality, secure, and idiomatic code.
Motivation
Code generation models frequently produce code that "works" but has critical issues: SQL injection vulnerabilities, O(n²) algorithms where O(n) is trivial, swallowed exceptions, thread safety bugs, and non-idiomatic patterns. This dataset trains models to produce code that a senior engineer would actually approve.… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/preference-pairs-coding-50k.amadablam-dpo-preferences
Ama Dablam DPO Preference Data
Preference pairs used to DPO-tune Ama Dablam,
a 322M trilingual (Nepali/Maithili/Bhojpuri) language model, across all three languages
and three writing systems (Devanagari, IAST, phonetic romanization). See the
technical report §9 for full
methodology.
Splits
split
rows
purpose
train
14,152
DPO Stage 2 preference-optimization training
validation
744
preference-accuracy / forgetting evaluation
warmup
3,203
Stage 1… See the full description on the dataset page: https://huggingface.co/datasets/spandyie/amadablam-dpo-preferences.math500-preference-pairs-fable
MATH-500 Preference Pairs (Fable-generated)
500 preference pairs covering all 500 problems of MATH-500, generated by Anthropic's Claude Fable 5 for reward-model training in a math-RLHF project (Qwen2.5-7B, PPO/GRPO on AWS EKS).
Format
{
"idx": 0,
"problem": "Convert the point $(0,3)$ ...",
"chosen": "<complete correct solution with full reasoning>",
"rejected_1": "<plausible-but-wrong solution, error mode A>",
"rejected_2": "<plausible-but-wrong solution… See the full description on the dataset page: https://huggingface.co/datasets/weivzhang/math500-preference-pairs-fable.tau2-retail-tool-preferences-v1
tau2-retail-tool-preferences-v1
A tool-calling preference dataset derived from successful rollouts bundled with
the MIT-licensed tau2-bench
retail environment. It exists to validate default DPO configurations: it is
large enough to show whether DPO moves tool-calling behaviour in the intended
direction, and it scores without a judge.
Every row is a branch point immediately before an assistant tool call:
messages — the full shared history, including the retail policy, prior… See the full description on the dataset page: https://huggingface.co/datasets/lefft/tau2-retail-tool-preferences-v1.adaption-preference-trace-decisions
PreferenceTrace — Source Corpus and Adaption Export
PreferenceTrace tests exact decision-making under competing preferences, evidence, approvals, abstention requirements, temporal/contextual precedence, and machine-readable citation contracts.
Two explicit lineage artifacts
File
Rows
Role
SHA-256
preferencetrace-source-96.jsonl
96
Canonical PreferenceTrace source corpus
7a447f9bf47c3ea455ed96ec36860360aa0e7b9e2dc604450e3a1c665b52363e… See the full description on the dataset page: https://huggingface.co/datasets/darthludious/adaption-preference-trace-decisions.preference-model-perturbations
preference-model-perturbations
A Hugging Face dataset of paired model responses (original vs.
counterfactually perturbed) along with human and reward-model preferences,
generated by a Counterfactual Data Augmentation (CDA) pipeline to
analyze and mitigate bias in preference models.
Links
Homepage: CDA Pipeline Code
Description
Each record contains:
bias: type of bias expressed in the perturbation (5 possible values).
query: the original user prompt or query.… See the full description on the dataset page: https://huggingface.co/datasets/abharadwaj123/preference-model-perturbations.bias_reduce_preference_data
Dataset Card for Persona-Aware Preference Dataset
Dataset Description
This is a Direct Preference Optimization (DPO) dataset designed to train language models to produce high-quality, context-aware responses when given user demographic information (persona). Each example pairs a user prompt prefixed with a demographic persona description with a chosen (preferred) response and a rejected (dispreferred) response.
The dataset is intended to support alignment research focused… See the full description on the dataset page: https://huggingface.co/datasets/groupfairnessllm/bias_reduce_preference_data.Hierarchical-Preference-Dataset
Hierarchical Preference Dataset
The Hierarchical Preference Dataset is a structured dataset for analyzing and evaluating model reasoning through a hierarchical cognitive decomposition lens. It is derived from the prhegde/preference-data-math-stack-exchange dataset and extends it with annotations that separate model outputs into Refined Query, Meta-Thinking, and Refined Answer components.
Overview
Each sample in this dataset consists of:
An instruction or query.
Two… See the full description on the dataset page: https://huggingface.co/datasets/Death-Raider/Hierarchical-Preference-Dataset.code-preference-sample
Code Response Preference Pairs — Free Sample (Python & JavaScript)
This is a free 120-row sample of a 600-row manually-verified preference dataset. The full set is available separately — see "Get the full dataset" below.
What this is
A preference dataset for fine-tuning and evaluating coding AI models. Each row contains a coding task, two candidate code responses, a label identifying the preferred response, and a detailed technical justification — the standard… See the full description on the dataset page: https://huggingface.co/datasets/shanmukha-dev/code-preference-sample.ultrafeedback-binarized-preferences-cleaned-multilingual
Dataset Card for ultrafeedback-binarized-preferences-cleaned-multilingual
本資料集是 argilla/ultrafeedback-binarized-preferences-cleaned 的多語言(含繁體中文 zh-tw)版本,每筆樣本保留原始來源、語言、對話與 chosen / rejected 回應,可作為繁中模型的 DPO 對齊資料。
Dataset Details
Dataset Description
原始 UltraFeedback 是一個英文偏好資料集,本資料集將其翻譯/增廣為多語版本,並對譯文做品質清理,以利非英文(特別是繁中)DPO 訓練。
每筆樣本包含:
source:原始來源(如 evol_instruct、flan_v2_p3 等)。
lang:語言代碼(如 zh-tw、en)。
conversations:human/gpt 對話結構,內容已翻譯為對應語言。
(以及對應的 chosen / rejected… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/ultrafeedback-binarized-preferences-cleaned-multilingual.ptbr-human-preferences
🇧🇷 HUBX Human Preference Dataset (PT-BR)
The largest Portuguese-Brazilian human preference dataset for RLHF/DPO training.
📊 Dataset Statistics
Metric
Value
Total Annotations
314,757
Unique Tasks
450
Human Annotators
~600
Avg. Votes per Task
~699
Language
Portuguese (Brazil)
Domain
Communication Quality & Tone
🎯 Why This Dataset?
🇧🇷 Native PT-BR: Collected from Brazilian Portuguese speakers - not translated
👥 Real Humans:… See the full description on the dataset page: https://huggingface.co/datasets/Hub-Ai/ptbr-human-preferences.ro-preference-pairs-2k
Romanian Preference Pairs (2K)
Synthetic DPO preference pairs in Romanian targeting over-refusal and helpfulness alignment.
Dataset Description
2,000 preference pairs in Romanian across 4 categories:
Category
Examples
Description
informational
~500
Factual questions about Romania, economics, law
coding
~500
Python code tasks, FastAPI, SQLAlchemy
task_completion
~500
Document drafting, emails, plans
advice
~500
Career, productivity, technical… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/ro-preference-pairs-2k.
