datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tulu-2.5-preference-data
Tulu 2.5 Preference Data
This dataset contains the preference dataset splits used to train the models described in Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference Feedback.
We cleaned and formatted all datasets to be in the same format.
This means some splits may differ from their original format.
To see the code used for creating most splits, see here.
If you only wish to download one dataset, each dataset exists in one file under the data/… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tulu-2.5-preference-data.Apertus-v1.5-Preference-Data
Apertus 1.5 Preference Dataset
This is the preference dataset used for the offline DPO stage of Apertus v1.5 alignment training, applied to the 70B model.
The prompts come from Ai2's Olmo 3 Dolci-Instruct-DPO dataset. We only reuse the prompts from Dolci-Instruct-DPO; all chosen / rejected responses in this dataset were generated by us.
How this dataset was built
Prompts. Taken from Dolci-Instruct-DPO (ODC-BY).
Response generation and annotation. Every prompt was… See the full description on the dataset page: https://huggingface.co/datasets/swiss-ai/Apertus-v1.5-Preference-Data.PRO-STEP-Preference-Data
PRO-STEP: DPO Preference Pairs
Step-level preference pairs used to train the PRO-STEP policy model via Direct Preference Optimization.
Paper: PRO-STEP: Step-level Process Reward Optimization for Retrieval-Augmented GenerationCode: GitHub Repository
Pairs: 15,877 (after outcome filter)
Source questions: 5,000 from HotpotQA + MuSiQue + 2WikiMultiHopQA training splits
Generation: PRM-guided MCTS (K=3 branching, depth 7, 64 rollouts/question, V(s) = Q̄(s) + α · r̂(s) with α=0.3)… See the full description on the dataset page: https://huggingface.co/datasets/MinKeonKim/PRO-STEP-Preference-Data.Synthetic-JP-Preference-Dataset-Qwen2.5_72B-191k
Synthetic-JP-Preference-Dataset-Qwen2.5_72B-191k
概要
5種類のオープンモデルとQwen/Qwen2.5-72B-Instruct-GPTQ-Int8を使って作成した、190854件の日本語合成Preferenceデータセットです。
以下、データセットの詳細です。
instructionには、Aratako/Magpie-Tanuki-8B-annotated-96kのinput_qualityがexcellentのものを利用
回答生成には、以下の5つのApache 2.0ライセンスのモデルを利用
weblab-GENIAC/Tanuki-8B-dpo-v1.0
team-hatakeyama-phase2/Tanuki-8x8B-dpo-v1.0-GPTQ-8bit
cyberagent/calm3-22b-chat
llm-jp/llm-jp-3-13b-instruct
Qwen/Qwen2.5-32B-Instruct-GPTQ-Int8… See the full description on the dataset page: https://huggingface.co/datasets/Aratako/Synthetic-JP-Preference-Dataset-Qwen2.5_72B-191k.DataSmith-Preference-HumanEval
Website: ljvmiranda921.github.io/datasmith/
DataSmith-Preference-HumanEval
This dataset contains human preference judgments comparing DataSmith against a teacher baseline that generates without tools.
Each comparison shows an evaluator two conversations written from the same English seed prompt, one from each system, and asks which they prefer overall and on four aspects.
Evaluators were native or fluent speakers of the target language, and the two systems were shown as… See the full description on the dataset page: https://huggingface.co/datasets/ljvmiranda921/DataSmith-Preference-HumanEval.dataset-tldr-preference-dpo
Dataset Card for dataset-tldr-preference-dpo
This dataset has been created with distilabel.
Dataset Summary
This is a dataset intended for training models using DPO/ORPO for the task of producing concise tl;dr summaries of machine learning datasets based on their dataset cards.
The dataset was created with distilabel. Each row of the dataset contains a dataset card which has been parsed to remove empty sections and placeholder text.
The instruction request… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/dataset-tldr-preference-dpo.argument_generation_preference_dataFallacy types mapping:
'Not a Fallacy': 0
'faulty generalization': 1
'false causality': 2
'fallacy of relevance': 3
'fallacy of extension': 4
'equivocation': 5
'ad populum': 6
'appeal to emotion': 7
'ad hominem': 8
'circular reasoning': 9
'fallacy of credibility': 10
'fallacy of logic': 11
'false dilemma': 12
'intentional': 13
MNLP_M1_Preference_dpo_dataset
M1 Preference Data for DPO
Dataset Description
This dataset contains processed M1 preference data for DPO training.
Created by: CS-552 Stochastic Parrots Team
Date: May 24, 2025
Version: 1.0
Number of examples: 17615
Dataset Source
This dataset is derived from the M1 preference data collected through interactions with large language models (like ChatGPT) for CS-552 (Modern Natural Language Processing) at EPFL. The preference data consists of… See the full description on the dataset page: https://huggingface.co/datasets/stochastic-parrots/MNLP_M1_Preference_dpo_dataset.bias_reduce_preference_data
Dataset Card for Persona-Aware Preference Dataset
Dataset Description
This is a Direct Preference Optimization (DPO) dataset designed to train language models to produce high-quality, context-aware responses when given user demographic information (persona). Each example pairs a user prompt prefixed with a demographic persona description with a chosen (preferred) response and a rejected (dispreferred) response.
The dataset is intended to support alignment research focused… See the full description on the dataset page: https://huggingface.co/datasets/groupfairnessllm/bias_reduce_preference_data.judgment-consistency-preference-data
Dataset Card for judgment Consistency Preference Data
Dataset Description
This is a preference dataset designed to enhance the consistency of judgment in models when faced with disturbance, suitable for the DPO algorithm. It contains 2607 prompts sampled from arithmetic, commonsense, symbolic, and knowledge reasoning datasets, each accompanied by a pair of responses: one "chosen" response and one "rejected" response.
We design a dialogue scenario with one round of… See the full description on the dataset page: https://huggingface.co/datasets/NUSTM/judgment-consistency-preference-data.tulu-3-preference-data-with-distraction
Tulu-3 Preference Data with Distraction (Preference Data)
This dataset provides preference pairs (for DPO, IPO, ORPO, KTO, etc.) where prompts intentionally include distractor content (e.g., hidden instructions, puzzles, or extra tasks) to test and train models to ignore the distractor and solve the primary query. It is the preference companion to the SFT-only dataset groupfairnessllm/tulu-3-sft-with-distraction. The original data is derived from Tulu 3 dataset which contains coding… See the full description on the dataset page: https://huggingface.co/datasets/groupfairnessllm/tulu-3-preference-data-with-distraction.Hierarchical-Preference-Dataset
Hierarchical Preference Dataset
The Hierarchical Preference Dataset is a structured dataset for analyzing and evaluating model reasoning through a hierarchical cognitive decomposition lens. It is derived from the prhegde/preference-data-math-stack-exchange dataset and extends it with annotations that separate model outputs into Refined Query, Meta-Thinking, and Refined Answer components.
Overview
Each sample in this dataset consists of:
An instruction or query.
Two… See the full description on the dataset page: https://huggingface.co/datasets/Death-Raider/Hierarchical-Preference-Dataset.Self_Alignment_Preference-Dataset
Mistral Self-Alignment Preference Dataset
Warning: This dataset contains harmful and offensive data! Proceed with caution.
The Mistral Self-Alignment Preference Dataset was generated by Mistral 7b using the Anthropics Red Teaming Prompts dataset available at Hugging Face - Anthropics Red Teaming Prompts Dataset. The data generation process utilized the Preference Data Generation Notebook, which can be found here.
The purpose of this dataset is to facilitate self-alignment, as… See the full description on the dataset page: https://huggingface.co/datasets/August4293/Self_Alignment_Preference-Dataset.synthetic-preference-data
Synthetic Preference Data
A small, synthetically generated preference dataset intended for testing
RLHF / DPO training pipelines. Each example contains a prompt and two
responses — one correct (chosen), one subtly flawed (rejected).
Generation Procedure
Generator model: gpt-4o-mini (OpenAI)
Generation method: OpenAI Structured Outputs (response_format=PreferenceExample) — guarantees schema-valid JSON.
Prompt templates: 4 (factual, step-by-step reasoning, technical how-to… See the full description on the dataset page: https://huggingface.co/datasets/antony-bryan-3D2Y/synthetic-preference-data.sherlock_preference_datasetThis dataset contains preference data for tuning Vision-Language models on the Sherlock Dataset for Abductive Reasoning. It is designed to evaluate the effectiveness of fine-tuning using Supervised Fine-Tuning (SFT) or Preference Optimization. Preferences are generated by prompting four models: mistralai/Pixtral-12B-2409, Qwen/Qwen2-VL-7B-Instruct, google/paligemma2-3b-ft-docci-448, and google/paligemma2-10b-ft-docci-448.
Since this dataset is intended for optimizing PaLI-Gemma models… See the full description on the dataset page: https://huggingface.co/datasets/akshayg08/sherlock_preference_dataset.INFH-6000Q-dpo-preference-dataset
INFH-6000Q DPO Preference Dataset
This dataset contains the final preference pairs used for the Direct Preference Optimization assignment in this repository.
Source
Base instruction source: GAIR/lima
Candidate generator: local Qwen/Qwen2.5-7B-Instruct
Preference ranker: local llm-blender/PairRM
Construction Pipeline
Sample 50 instructions from the local LIMA training split with seed 42.
Generate 5 candidate responses per instruction with Qwen2.5-7B-Instruct.… See the full description on the dataset page: https://huggingface.co/datasets/ITBill/INFH-6000Q-dpo-preference-dataset.
