datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ultrafeedback-binarized-preferences-cleaned
UltraFeedback - Binarized using the Average of Preference Ratings (Cleaned)
This dataset represents a new iteration on top of argilla/ultrafeedback-binarized-preferences,
and is the recommended and preferred dataset by Argilla to use from now on when fine-tuning on UltraFeedback.
Read more about Argilla's approach towards UltraFeedback binarization at argilla/ultrafeedback-binarized-preferences/README.md.
Differences with argilla/ultrafeedback-binarized-preferences… See the full description on the dataset page: https://huggingface.co/datasets/argilla/ultrafeedback-binarized-preferences-cleaned.ultrafeedback-binarized-preferences-cleaned-kto
UltraFeedback - Binarized using the Average of Preference Ratings (Cleaned) KTO
A KTO signal transformed version of the highly loved UltraFeedback Binarized Preferences Cleaned, the preferred dataset by Argilla to use from now on when fine-tuning on UltraFeedback
This dataset represents a new iteration on top of argilla/ultrafeedback-binarized-preferences,
and is the recommended and preferred dataset by Argilla to use from now on when fine-tuning on UltraFeedback.
Read more about… See the full description on the dataset page: https://huggingface.co/datasets/argilla/ultrafeedback-binarized-preferences-cleaned-kto.WritingPrompts_preferences
Dataset Card for "WritingPrompts_preferences"
Human preference data from r/WritingPrompts
tulu-2.5-preference-data
Tulu 2.5 Preference Data
This dataset contains the preference dataset splits used to train the models described in Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference Feedback.
We cleaned and formatted all datasets to be in the same format.
This means some splits may differ from their original format.
To see the code used for creating most splits, see here.
If you only wish to download one dataset, each dataset exists in one file under the data/… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tulu-2.5-preference-data.distilabel-math-preference-dpo
Dataset Card for "distilabel-math-preference-dpo"
More Information needed
Preference-Collection
Dataset Card
Dataset Summary
The Preference Collection is a dataset designed to induce fine-grained evaluation capabilities into language models.
Recently, proprietary LLMs (e.g., GPT-4) have been used to evaluate long-form responses. In our experiments, we found that open-source LMs are not capable of evaluating long-form responses, showing low correlation with both human evaluators and GPT-4.\
In our paper, we found that by (1) fine-tuning feedback generated by GPT-4… See the full description on the dataset page: https://huggingface.co/datasets/prometheus-eval/Preference-Collection.Infinity-Preference
Infinity-Preference
The focus of human preferences varies from task to task. Therefore, Infinity-Preference attempts to adjust preference attribute weights on each task based on (Infinity Instruct's)[https://huggingface.co/datasets/BAAI/Infinity-Instruct] capability labelling system. This version contains 59438 evenly sampled instructions from Infinity-Instruct's instruction set for each task type. Each instruction is accompanied by a preference pair sampled from Gemma-2-9B-IT. This… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/Infinity-Preference.argunauts-hirpo-preferences
Argunauts HIRPO Preferences
Preference pairs generated while training Argunaut models with HIRPO Online DPO.
Capybara-Preferences
Dataset Card for Capybara-Preferences
This dataset has been created with distilabel.
Dataset Summary
This dataset is built on top of LDJnr/Capybara, in order to generate a preference
dataset out of an instruction-following dataset. This is done by keeping the conversations in the column conversation but splitting
the last assistant turn from it, so that the conversation contains all the turns up until the last user's turn, so that it can be reused… See the full description on the dataset page: https://huggingface.co/datasets/argilla/Capybara-Preferences.Apertus-v1.5-Preference-Data
Apertus 1.5 Preference Dataset
This is the preference dataset used for the offline DPO stage of Apertus v1.5 alignment training, applied to the 70B model.
The prompts come from Ai2's Olmo 3 Dolci-Instruct-DPO dataset. We only reuse the prompts from Dolci-Instruct-DPO; all chosen / rejected responses in this dataset were generated by us.
How this dataset was built
Prompts. Taken from Dolci-Instruct-DPO (ODC-BY).
Response generation and annotation. Every prompt was… See the full description on the dataset page: https://huggingface.co/datasets/swiss-ai/Apertus-v1.5-Preference-Data.Qwen3-Coder-Next-OpenCode-Preference
Dataset Card — OpenCode Rejection Sampling (Preference)
Overview
This dataset contains 10,920 preference pairs for preference-based training (DPO, KTO, SimPO, ORPO, etc.) on competitive programming tasks. Each pair consists of:
Chosen: a candidate solution that passes 100% of test cases
Rejected: a candidate solution that fails, with a fine-grained rejection type label
Pairs are produced via rejection sampling with Qwen3-Coder-Next: 8 candidate solutions are… See the full description on the dataset page: https://huggingface.co/datasets/zake7749/Qwen3-Coder-Next-OpenCode-Preference.grounded-qa-preferences
Grounded QA preferences
Preference pairs for a small RLHF stack. Each row is a passage, a question, a preferred answer, and a rejected answer.
The questions, answer spans, and unanswerable labels come from SQuAD 2.0 (Rajpurkar et al.). This dataset does not add new human rankings. A fixed rule turns those annotations into Bradley-Terry pairs:
pair_type
When
Chosen
Rejected
wrong_span
The passage answers the question
The gold span
A different short span from the same… See the full description on the dataset page: https://huggingface.co/datasets/saitejaalasyam/grounded-qa-preferences.gigaverbo-v2-preferences
GigaVerbo-v2 Preferences: A Hybrid-Reasoning Portuguese Preference Dataset
Dataset Summary
GigaVerbo-v2 Preferences is a preference dataset designed for Direct Preference Optimization (DPO) and other direct alignment algorithms. The dataset comprises approximately 27.8 million tokens across 28,437 preference pairs, organized into 4 distinct subsets covering both quality-focused and safety-focused alignment. It is entirely composed of high-quality, LLM-generated data… See the full description on the dataset page: https://huggingface.co/datasets/Polygl0t/gigaverbo-v2-preferences.DPO-En-Zh-20k-PreferenceThis dataset is composed by
4,000 examples of argilla/distilabel-capybara-dpo-7k-binarized with chosen score>=4.
3,000 examples of argilla/distilabel-intel-orca-dpo-pairs with chosen score>=8.
3,000 examples of argilla/ultrafeedback-binarized-preferences-cleaned with chosen score>=4.
10,000 examples of wenbopan/Chinese-dpo-pairs.
refer: https://huggingface.co/datasets/hiyouga/DPO-En-Zh-20k 改了question、response_rejected、response_chosen字段,方便ORPO、DPO模型训练时使用train usage:… See the full description on the dataset page: https://huggingface.co/datasets/shibing624/DPO-En-Zh-20k-Preference.PRO-STEP-Preference-Data
PRO-STEP: DPO Preference Pairs
Step-level preference pairs used to train the PRO-STEP policy model via Direct Preference Optimization.
Paper: PRO-STEP: Step-level Process Reward Optimization for Retrieval-Augmented GenerationCode: GitHub Repository
Pairs: 15,877 (after outcome filter)
Source questions: 5,000 from HotpotQA + MuSiQue + 2WikiMultiHopQA training splits
Generation: PRM-guided MCTS (K=3 branching, depth 7, 64 rollouts/question, V(s) = Q̄(s) + α · r̂(s) with α=0.3)… See the full description on the dataset page: https://huggingface.co/datasets/MinKeonKim/PRO-STEP-Preference-Data.ethics_expression_preferencesReformat of the ethics dataset to be in DPO format, in this expression format
Examples:
subset commonsense
before 13910 {'label': 1, 'input': "I went to the principal's office to change my records before going to a different school."}
after 1513 {'prompt': 'I won my match by ', 'chosen': 'making a lot of good shots.', 'rejected': "moving my ball when the other player wasn't looking."}
subset utilitarianism
before 13737 {'baseline': 'I forgot to wear my virus-proof mask when I visited… See the full description on the dataset page: https://huggingface.co/datasets/wassname/ethics_expression_preferences.ultrafeedback-multi-binarized-preferences-cleaned
UltraFeedback - Multi-Binarized using the Average of Preference Ratings (Cleaned)
This dataset represents a new iteration on top of argilla/ultrafeedback-binarized-preferences-cleaned,
and has been created to explore whether DPO fine-tuning with more than one rejection per chosen response helps the model perform better in the
AlpacaEval, MT-Bench, and LM Eval Harness benchmarks.
Read more about Argilla's approach towards UltraFeedback binarization at… See the full description on the dataset page: https://huggingface.co/datasets/argilla/ultrafeedback-multi-binarized-preferences-cleaned.dolphin-sft-v0.1-preferenceThe preference dataset was generated using Mistral-Instruct-v0.1 finetuned on a GPT-4 subset of the Dolphin dataset (16k samples). Link to the model.
Generated responses are labeled as rejected, GPT-4 responses (original Dolphin data) are labeled as accepted.
The motivation was to test out the SPIN paper finetuning methodology.
Capybara-Preferences-Filtered
Dataset Card for Capybara-Preferences-Filtered
This dataset has been created with distilabel, plus some extra post-processing steps described below.
Dataset Summary
This dataset is built on top of argilla/Capybara-Preferences, but applies a further in detail filtering.
The filtering approach has been proposed and shared by @LDJnr, and applies the following:
Remove responses from the assistant, not only in the last turn, but also in intermediate… See the full description on the dataset page: https://huggingface.co/datasets/argilla/Capybara-Preferences-Filtered.helpsteer2-preference-openai-native
HelpSteer2 Preference — OpenAI Native Format
A deterministic, training-ready repackaging of the preference split of
nvidia/HelpSteer2.
Why use this
What it is for. Preference optimisation — DPO, ORPO, SimPO, KTO — and reward
modelling, on 7,051 pairs that come from paid human annotators, not from an LLM
judge. Each pair carries a graded strength from 1 to 3 rather than a bare
binary label, so you can weight the loss by how strongly humans actually
disagreed, or… See the full description on the dataset page: https://huggingface.co/datasets/Archangel-system/helpsteer2-preference-openai-native.creative-rubrics-preferences
creative-rubrics-preferences 🎏
A dataset of creative responses using GPT-4.5, o3-mini and DeepSeek-R1.
This dataset contains several prompts seeking creative and diverse answers (like writing movie reviews, short stories, etc), and the style of the responses has been enhanced by prompting the model with custom rubrics that seek different creative styles.
This dataset was used in the paper Configurable Preference Tuning with Rubric-Guided Synthetic Data.
Code:… See the full description on the dataset page: https://huggingface.co/datasets/vicgalle/creative-rubrics-preferences.dfm13-arena-human-preference-100k-preferred
dfm13-arena-human-preference-100k-preferred
Model-audited preferred responses, including explicitly identified model repairs.
Not certified gold and not manually verified in full. Independent review is
sample-based where declared in the publication receipt; holds are excluded.
Repository split name train is a storage convention, not training admission.
Full original history and target preserved; no truncation or 4096-token cutoff. Training length filtering is separate and not… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm13-arena-human-preference-100k-preferred.DiscoverLLM-multiturn-preferences
DiscoverLLM: Multi-turn Preference Dataset
Multi-turn dialogue data with scored candidate completions, produced by best-of-N
synthesis over the DiscoverLLM user simulator
(paper · project page).
Each example is a single turn of a simulated user–assistant conversation with one of
several candidate assistant responses and an associated reward score, intended for
offline DPO / GRPO / reward-model training.
Configs
Config
Rows
Task
creative_writing
3,052… See the full description on the dataset page: https://huggingface.co/datasets/kixlab/DiscoverLLM-multiturn-preferences.Apertus-v1.5-Preference-Pool
Apertus 1.5 Preference Pool
The following dataset contains responses generated from a pool of 24 models, alongside the judge annotations that score them according to rubric-based rewards. The dataset was used for constructing the final Apertus 1.5 preference dataset, Apertus-v1.5-Preference-Data.
It is derived from Ai2's Olmo 3 Dolci-Instruct-DPO, where we reuse only the prompts and regenerate every response.
Where the preference dataset ships a single chosen / rejected pair per… See the full description on the dataset page: https://huggingface.co/datasets/swiss-ai/Apertus-v1.5-Preference-Pool.user-preference-564k
User Preference Extraction Dataset (564K)
A dataset of 564K examples for training lightweight preference extraction models. Each example pairs a conversation input with structured JSON output describing user preferences as condition-action rules.
This dataset was used to train blackhao0426/pref-extractor-qwen3-0.6b-full-sft, a core component of the VARS framework.
Sample Usage
The following snippet from the official repository demonstrates how to use the framework… See the full description on the dataset page: https://huggingface.co/datasets/blackhao0426/user-preference-564k.lmsys-arena-human-preference-winner-43k-unfiltered
lmsys-arena-human-preference-winner-43k-unfiltered
This repository contains a dataset derived from the lmsys/lmsys-arena-human-preference-55k dataset, which is licensed under the Apache 2.0 License.
Dataset Description
The lmsys-arena-human-preference-winner-43k-unfiltered dataset is a collection of 43,000 samples, each containing an instruction (prompt) and an output (winning response) from real-world user and LLM conversations. The dataset is derived from the original… See the full description on the dataset page: https://huggingface.co/datasets/lesserfield/lmsys-arena-human-preference-winner-43k-unfiltered.Synthetic-JP-Preference-Dataset-Qwen2.5_72B-191k
Synthetic-JP-Preference-Dataset-Qwen2.5_72B-191k
概要
5種類のオープンモデルとQwen/Qwen2.5-72B-Instruct-GPTQ-Int8を使って作成した、190854件の日本語合成Preferenceデータセットです。
以下、データセットの詳細です。
instructionには、Aratako/Magpie-Tanuki-8B-annotated-96kのinput_qualityがexcellentのものを利用
回答生成には、以下の5つのApache 2.0ライセンスのモデルを利用
weblab-GENIAC/Tanuki-8B-dpo-v1.0
team-hatakeyama-phase2/Tanuki-8x8B-dpo-v1.0-GPTQ-8bit
cyberagent/calm3-22b-chat
llm-jp/llm-jp-3-13b-instruct
Qwen/Qwen2.5-32B-Instruct-GPTQ-Int8… See the full description on the dataset page: https://huggingface.co/datasets/Aratako/Synthetic-JP-Preference-Dataset-Qwen2.5_72B-191k.dialect-preferences
DiaLLM — Pooled Preference Dataset (Implicit Thread)
Part of DiaLLM: An Investigation into the Robustness-Generation Gap in
English Dialect Adaptation (EMNLP 2026 Main).
45,690 preference pairs, pooling all three variety-specific sets
(Australian,
Northern British,
Indian) without
variety targeting. Used for implicit-thread DPO training, where the three
varieties are pooled rather than targeted individually, preserving the
variety-agnostic objective of that thread.… See the full description on the dataset page: https://huggingface.co/datasets/surrey-nlp/dialect-preferences.argilla-ultrafeedback-binarized-preferences-cleaned
UltraFeedback (Cleaned)
This dataset combines the train split of argilla/ultrafeedback-binarized-preferences-cleaned,
and test split of HuggingFaceH4/ultrafeedback_binarized.
warehouse-dpo-preference-pairs
Warehouse Short-Order DPO Preference Pairs
Dataset Description
This dataset contains {prompt, chosen, rejected} preference pairs for
training a warehouse short-order assistant with Direct Preference
Optimization (DPO). Each pair asks a real warehouse-inventory question
(stockout risk, backorders, KPI summaries, why a warehouse is failing
fulfillment - at a single-warehouse, tier, region, or dataset-wide
comparison level) grounded in real tool-call output… See the full description on the dataset page: https://huggingface.co/datasets/EnRaoufi/warehouse-dpo-preference-pairs.
