datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ultrafeedback-binarized-preferences-cleaned
UltraFeedback - Binarized using the Average of Preference Ratings (Cleaned)
This dataset represents a new iteration on top of argilla/ultrafeedback-binarized-preferences,
and is the recommended and preferred dataset by Argilla to use from now on when fine-tuning on UltraFeedback.
Read more about Argilla's approach towards UltraFeedback binarization at argilla/ultrafeedback-binarized-preferences/README.md.
Differences with argilla/ultrafeedback-binarized-preferences… See the full description on the dataset page: https://huggingface.co/datasets/argilla/ultrafeedback-binarized-preferences-cleaned.distilabel-math-preference-dpo
Dataset Card for "distilabel-math-preference-dpo"
More Information needed
Capybara-Preferences
Dataset Card for Capybara-Preferences
This dataset has been created with distilabel.
Dataset Summary
This dataset is built on top of LDJnr/Capybara, in order to generate a preference
dataset out of an instruction-following dataset. This is done by keeping the conversations in the column conversation but splitting
the last assistant turn from it, so that the conversation contains all the turns up until the last user's turn, so that it can be reused… See the full description on the dataset page: https://huggingface.co/datasets/argilla/Capybara-Preferences.Apertus-v1.5-Preference-Data
Apertus 1.5 Preference Dataset
This is the preference dataset used for the offline DPO stage of Apertus v1.5 alignment training, applied to the 70B model.
The prompts come from Ai2's Olmo 3 Dolci-Instruct-DPO dataset. We only reuse the prompts from Dolci-Instruct-DPO; all chosen / rejected responses in this dataset were generated by us.
How this dataset was built
Prompts. Taken from Dolci-Instruct-DPO (ODC-BY).
Response generation and annotation. Every prompt was… See the full description on the dataset page: https://huggingface.co/datasets/swiss-ai/Apertus-v1.5-Preference-Data.Qwen3-Coder-Next-OpenCode-Preference
Dataset Card — OpenCode Rejection Sampling (Preference)
Overview
This dataset contains 10,920 preference pairs for preference-based training (DPO, KTO, SimPO, ORPO, etc.) on competitive programming tasks. Each pair consists of:
Chosen: a candidate solution that passes 100% of test cases
Rejected: a candidate solution that fails, with a fine-grained rejection type label
Pairs are produced via rejection sampling with Qwen3-Coder-Next: 8 candidate solutions are… See the full description on the dataset page: https://huggingface.co/datasets/zake7749/Qwen3-Coder-Next-OpenCode-Preference.gigaverbo-v2-preferences
GigaVerbo-v2 Preferences: A Hybrid-Reasoning Portuguese Preference Dataset
Dataset Summary
GigaVerbo-v2 Preferences is a preference dataset designed for Direct Preference Optimization (DPO) and other direct alignment algorithms. The dataset comprises approximately 27.8 million tokens across 28,437 preference pairs, organized into 4 distinct subsets covering both quality-focused and safety-focused alignment. It is entirely composed of high-quality, LLM-generated data… See the full description on the dataset page: https://huggingface.co/datasets/Polygl0t/gigaverbo-v2-preferences.PRO-STEP-Preference-Data
PRO-STEP: DPO Preference Pairs
Step-level preference pairs used to train the PRO-STEP policy model via Direct Preference Optimization.
Paper: PRO-STEP: Step-level Process Reward Optimization for Retrieval-Augmented GenerationCode: GitHub Repository
Pairs: 15,877 (after outcome filter)
Source questions: 5,000 from HotpotQA + MuSiQue + 2WikiMultiHopQA training splits
Generation: PRM-guided MCTS (K=3 branching, depth 7, 64 rollouts/question, V(s) = Q̄(s) + α · r̂(s) with α=0.3)… See the full description on the dataset page: https://huggingface.co/datasets/MinKeonKim/PRO-STEP-Preference-Data.ultrafeedback-multi-binarized-preferences-cleaned
UltraFeedback - Multi-Binarized using the Average of Preference Ratings (Cleaned)
This dataset represents a new iteration on top of argilla/ultrafeedback-binarized-preferences-cleaned,
and has been created to explore whether DPO fine-tuning with more than one rejection per chosen response helps the model perform better in the
AlpacaEval, MT-Bench, and LM Eval Harness benchmarks.
Read more about Argilla's approach towards UltraFeedback binarization at… See the full description on the dataset page: https://huggingface.co/datasets/argilla/ultrafeedback-multi-binarized-preferences-cleaned.Capybara-Preferences-Filtered
Dataset Card for Capybara-Preferences-Filtered
This dataset has been created with distilabel, plus some extra post-processing steps described below.
Dataset Summary
This dataset is built on top of argilla/Capybara-Preferences, but applies a further in detail filtering.
The filtering approach has been proposed and shared by @LDJnr, and applies the following:
Remove responses from the assistant, not only in the last turn, but also in intermediate… See the full description on the dataset page: https://huggingface.co/datasets/argilla/Capybara-Preferences-Filtered.DiscoverLLM-multiturn-preferences
DiscoverLLM: Multi-turn Preference Dataset
Multi-turn dialogue data with scored candidate completions, produced by best-of-N
synthesis over the DiscoverLLM user simulator
(paper · project page).
Each example is a single turn of a simulated user–assistant conversation with one of
several candidate assistant responses and an associated reward score, intended for
offline DPO / GRPO / reward-model training.
Configs
Config
Rows
Task
creative_writing
3,052… See the full description on the dataset page: https://huggingface.co/datasets/kixlab/DiscoverLLM-multiturn-preferences.dialect-preferences
DiaLLM — Pooled Preference Dataset (Implicit Thread)
Part of DiaLLM: An Investigation into the Robustness-Generation Gap in
English Dialect Adaptation (EMNLP 2026 Main).
45,690 preference pairs, pooling all three variety-specific sets
(Australian,
Northern British,
Indian) without
variety targeting. Used for implicit-thread DPO training, where the three
varieties are pooled rather than targeted individually, preserving the
variety-agnostic objective of that thread.… See the full description on the dataset page: https://huggingface.co/datasets/surrey-nlp/dialect-preferences.warehouse-dpo-preference-pairs
Warehouse Short-Order DPO Preference Pairs
Dataset Description
This dataset contains {prompt, chosen, rejected} preference pairs for
training a warehouse short-order assistant with Direct Preference
Optimization (DPO). Each pair asks a real warehouse-inventory question
(stockout risk, backorders, KPI summaries, why a warehouse is failing
fulfillment - at a single-warehouse, tier, region, or dataset-wide
comparison level) grounded in real tool-call output… See the full description on the dataset page: https://huggingface.co/datasets/EnRaoufi/warehouse-dpo-preference-pairs.MNLP_M1_Preference_dpo_dataset
M1 Preference Data for DPO
Dataset Description
This dataset contains processed M1 preference data for DPO training.
Created by: CS-552 Stochastic Parrots Team
Date: May 24, 2025
Version: 1.0
Number of examples: 17615
Dataset Source
This dataset is derived from the M1 preference data collected through interactions with large language models (like ChatGPT) for CS-552 (Modern Natural Language Processing) at EPFL. The preference data consists of… See the full description on the dataset page: https://huggingface.co/datasets/stochastic-parrots/MNLP_M1_Preference_dpo_dataset.innoduel-rlhf-real-world-human-preferences-sample
Real-World Human Pairwise Preferences — Public Sample
📦 This is a free, public sample of a commercial dataset.
It contains 1,350 rows curated for inspection. The full dataset has 1.5 million
human pairwise-preference decisions.
Full dataset: https://huggingface.co/datasets/NordosoftOy/innoduel-rlhf
Request access / licensing: see § Access to the full dataset — contact kari.nieminen@nordo.fi.
Use this sample to evaluate the data's quality, structure and… See the full description on the dataset page: https://huggingface.co/datasets/NordosoftOy/innoduel-rlhf-real-world-human-preferences-sample.paraphrasing-preferences-orpo-dpo
Paraphrasing Preference Dataset
A preference dataset for training paraphrase models via DPO, RLHF, or ORPO. Each example contains a source text, a task-specific prompt, and a chosen/rejected paraphrase pair ranked by a composite quality score.
Dataset Summary
Train
Val
Total
Examples
852
95
947
Sources: Quora questions (571), SQuAD 2.0 sentences (218), CNN News sentences (158). The val split is stratified by excellent_in, category, and binned total_delta.… See the full description on the dataset page: https://huggingface.co/datasets/alecccdd/paraphrasing-preferences-orpo-dpo.distilabel-math-preference-dpo-argilla
Dataset Card for "distilabel-math-preference-dpo"
More Information needed
yue-math-preference
Cantonese Math Preference
This dataset is a Cantonese and Simplified Chinese translation of argilla/distilabel-math-preference-dpo. For more detailed information about the original dataset, please refer to the provided link.
This dataset is translated by Gemini Pro and has not undergone any manual verification. The content may be inaccurate or misleading. please keep this in mind when using this dataset.
License
This dataset is provided under the same license as the… See the full description on the dataset page: https://huggingface.co/datasets/hon9kon9ize/yue-math-preference.PreferenceTravelPlanner
PreferenceTravelPlanner Dataset
PreferenceTravelPlanner is a benchmark crafted for evaluating language agents in tool-use and complex planning within multiple constraints and preferences. For more details, see our paper. It is created by augmenting TravelPlanner (See paper for more details) with several common type of preferences under various preference paradigms.
Introduction
In PreferenceTravelPlanner, for a given query, language agents are expected to formulate a… See the full description on the dataset page: https://huggingface.co/datasets/pensieves/PreferenceTravelPlanner.ultrafeedback-binarized-preferences-cleaned
UltraFeedback - Binarized using the Average of Preference Ratings (Cleaned)
This dataset represents a new iteration on top of argilla/ultrafeedback-binarized-preferences,
and is the recommended and preferred dataset by Argilla to use from now on when fine-tuning on UltraFeedback.
Read more about Argilla's approach towards UltraFeedback binarization at argilla/ultrafeedback-binarized-preferences/README.md.
Differences with argilla/ultrafeedback-binarized-preferences… See the full description on the dataset page: https://huggingface.co/datasets/pharaouk/ultrafeedback-binarized-preferences-cleaned.helpsteer2_preference
Introduction
This is a binarized preference datasets from nvidia/HelpSteer2. HelpSteer2 is an open-source Helpfulness Dataset (CC-BY-4.0) that supports aligning models to become more helpful, factually correct and coherent, while being adjustable in terms of the complexity and verbosity of its responses. This dataset has been created in partnership with Scale AI.
I processed the raw data by prioritizing helpfulness, correctness, and coherence to determine which responses were chosen… See the full description on the dataset page: https://huggingface.co/datasets/AIR-hl/helpsteer2_preference.ultrafeedback-binarized-preferences-cleaned-no-refusals
UltraFeedback Binarized Preferences Cleaned No Refusals
A Minos-cleaned version of argilla/ultrafeedback-binarized-preferences-cleaned for use as a neutral helpfulness DPO anchor. Rows are removed when either the chosen or rejected assistant response is classified as a refusal by NousResearch/Minos-v1.
Cleaning version: minos-only-v1-2026-06-23
See manifest.json in the repository files for counts and endpoint metadata.
assignment4-pairrm-preferences-submitAD4Edu-Preferences
Dataset Card for AD4Edu Preferences
Preference pairs over audio descriptions (AD) of slide-based lecture videos, for blind and
low-vision (BLV) students. Each pair is two candidate ADs for the same lecture moment; the task is to
say which better serves a BLV listener under the project's 45-rule lecture-AD standard
(rules_for_slides.yaml; six categories: style, terminology, length, deixis, faithfulness,
non-redundancy).
PRIVATE, derived from copyrighted lecture video. Do not… See the full description on the dataset page: https://huggingface.co/datasets/Hermeneia/AD4Edu-Preferences.lima-qwen2.5-7b-pairrm-preferences
LIMA × Qwen2.5-7B-Instruct × PairRM preference dataset
Preference dataset built for Assignment 4 of the alignment course.
How it was built
Source instructions: 50 instructions sampled with seed=42 from the GAIR/lima training split.
Candidate generation: For each instruction we sampled 5 responses from Qwen/Qwen2.5-7B-Instruct using the official chat template (temperature=0.9, top_p=0.95, max_new_tokens=512).
Ranking: All 5 candidates per instruction were ranked with… See the full description on the dataset page: https://huggingface.co/datasets/Barryzbr12/lima-qwen2.5-7b-pairrm-preferences.INFH-6000Q-dpo-preference-dataset
INFH-6000Q DPO Preference Dataset
This dataset contains the final preference pairs used for the Direct Preference Optimization assignment in this repository.
Source
Base instruction source: GAIR/lima
Candidate generator: local Qwen/Qwen2.5-7B-Instruct
Preference ranker: local llm-blender/PairRM
Construction Pipeline
Sample 50 instructions from the local LIMA training split with seed 42.
Generate 5 candidate responses per instruction with Qwen2.5-7B-Instruct.… See the full description on the dataset page: https://huggingface.co/datasets/ITBill/INFH-6000Q-dpo-preference-dataset.
