datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gdpval_preference_rubricstulu-2.5-preference-data
Tulu 2.5 Preference Data
This dataset contains the preference dataset splits used to train the models described in Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference Feedback.
We cleaned and formatted all datasets to be in the same format.
This means some splits may differ from their original format.
To see the code used for creating most splits, see here.
If you only wish to download one dataset, each dataset exists in one file under the data/… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tulu-2.5-preference-data.pairs_Movies_and_TVwebdev-arena-preference-10k
WebDev Arena Preference Dataset
This dataset contains 10K real-world Webdev Arena battle with 10 state-of-the-art LLMs. More details in the blog post.
Dataset License Agreement
This Agreement contains the terms and conditions that govern your access and use of the WebDev Arena Dataset (Arena Dataset). You may not use the Arena Dataset if you do not accept this Agreement. By clicking to accept, accessing the Arena Dataset, or both, you hereby agree to the terms of the… See the full description on the dataset page: https://huggingface.co/datasets/lmarena-ai/webdev-arena-preference-10k.pairs_Grocery_and_Gourmet_Foodargunauts-hirpo-preferences
Argunauts HIRPO Preferences
Preference pairs generated while training Argunaut models with HIRPO Online DPO.
tool-use-preference-pairs
Tool Use Preference Pairs
Rights & intended use: legacy public research corpus / portfolio
artifact. Hosted frontier-model outputs are research-only inputs under
project policy (synthetic-factory#161):
intended_use: research_only, project_training_policy: blocked. Not
training data for any model-weight update. Machine-readable record:
rights.json.
Release status: The raw, uncurated payload is now published under
data/raw/. It is available for inspection and reproducibility… See the full description on the dataset page: https://huggingface.co/datasets/rmems/tool-use-preference-pairs.FinQA-DPO-Statement-PreferenceCode-Preference-PairsCreator Nicolas Mejia-Petit
My Kofi
Code-Preference-Pairs Dataset
Overview
This dataset was created while created Open-critic-GPT. Here is a little Overview:
The Open-Critic-GPT dataset is a synthetic dataset created to train models in both identifying and fixing bugs in code. The dataset is generated using a unique synthetic data pipeline which involves:
Prompting a local model with an existing code example.
Introducing bugs into the code. While also having the model… See the full description on the dataset page: https://huggingface.co/datasets/Vezora/Code-Preference-Pairs.vLLM-SR-Preference-V1The files in this repo is the LLM-labeled samples that are used as the training dataset for vLLM-SR Preference model V1.
The training file (sharegpt_preference_labeld_with_negative.jsonl) contains 25k records that have sample_id, golden label for the preference-based routing policy, and a set of negative labels that are plausible but do not match the conversation context.
The validation file has the same structure, but only 1% of the training file size. The validation file and the training… See the full description on the dataset page: https://huggingface.co/datasets/ppppqp/vLLM-SR-Preference-V1.1.4b-policy_preference_data_gold_labelledPreference dataset using labels from the AlpacaFarm dataset, generated answers from a 1.4b fine-tuned Pythia policy model, and labelled using the AlpacaFarm 'reward-model-human' as a gold reward model.
Used to train reward models in 'Reward Model Ensembles Mitigate Overoptimization'
magpie-preference
Dataset Card for Magpie Preference Dataset
Dataset Description
The Magpie Preference Dataset is a crowdsourced collection of human preferences on synthetic instruction-response pairs generated using the Magpie approach.
This dataset is continuously updated through user interactions with the Magpie Preference Gradio Space.
What is Magpie?
Magpie is a very interesting new approach to creating synthetic data which doesn't require any seed data:… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/magpie-preference.DPO-En-Zh-20k-PreferenceThis dataset is composed by
4,000 examples of argilla/distilabel-capybara-dpo-7k-binarized with chosen score>=4.
3,000 examples of argilla/distilabel-intel-orca-dpo-pairs with chosen score>=8.
3,000 examples of argilla/ultrafeedback-binarized-preferences-cleaned with chosen score>=4.
10,000 examples of wenbopan/Chinese-dpo-pairs.
refer: https://huggingface.co/datasets/hiyouga/DPO-En-Zh-20k 改了question、response_rejected、response_chosen字段,方便ORPO、DPO模型训练时使用train usage:… See the full description on the dataset page: https://huggingface.co/datasets/shibing624/DPO-En-Zh-20k-Preference.openai-tldr-summarisation-preferences
Human feedback data
This is the version of the dataset used in https://arxiv.org/abs/2310.06452.
If starting a new project we would recommend using https://huggingface.co/datasets/openai/summarize_from_feedback.
See https://github.com/openai/summarize-from-feedback for original details of the dataset.
Here the data is formatted to enable huggingface transformers sequence classification models to be trained as reward functions.
dataset-viber-image-generation-preference-inference-endpoints-battle-flux
Dataset Card for Dataset Name
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/davidberenstein1957/dataset-viber-image-generation-preference-inference-endpoints-battle-flux.proactivity_preference_dataset
ProVoice study 1 — driver state, vehicle context and preferred Level of Autonomy
Driving-simulator data from the population data collection of the ProVoice /
ProActivity project (CARLA 0.10): 12 drivers × 2 sessions, ~20 Hz
multimodal driver-state and vehicle frames, and 1,446 driver-assigned
Level-of-Autonomy (LoA) labels stating how autonomously an in-vehicle
assistant should act on a given task. Drivers were prompted every 20 s about
two randomly drawn in-vehicle tasks and… See the full description on the dataset page: https://huggingface.co/datasets/ProVoice-proactivity/proactivity_preference_dataset.MediQ_AskDocs_preference
This dataset is the preference data subset of MediQ_AskDocs, for the SFT subset, see MediQ_AskDocs.
To cite:
@misc{li2025aligningllmsaskgood,
title={Aligning LLMs to Ask Good Questions A Case Study in Clinical Reasoning},
author={Shuyue Stella Li and Jimin Mun and Faeze Brahman and Jonathan S. Ilgen and Yulia Tsvetkov and Maarten Sap},
year={2025},
eprint={2502.14860},
archivePrefix={arXiv},
primaryClass={cs.CL}… See the full description on the dataset page: https://huggingface.co/datasets/stellalisy/MediQ_AskDocs_preference.repochat-arena-preference-4k
Overview
This dataset contains leaderboard vote data on RepoChat collected from 2024/11/30 to 2025/02/03
For reproducing the leaderboards from this data, refer to the notebook.
License
User prompts are licensed under CC-BY-4.0, and model outputs are governed by the terms of use set by the respective model providers.
PRO-STEP-Preference-Data
PRO-STEP: DPO Preference Pairs
Step-level preference pairs used to train the PRO-STEP policy model via Direct Preference Optimization.
Paper: PRO-STEP: Step-level Process Reward Optimization for Retrieval-Augmented GenerationCode: GitHub Repository
Pairs: 15,877 (after outcome filter)
Source questions: 5,000 from HotpotQA + MuSiQue + 2WikiMultiHopQA training splits
Generation: PRM-guided MCTS (K=3 branching, depth 7, 64 rollouts/question, V(s) = Q̄(s) + α · r̂(s) with α=0.3)… See the full description on the dataset page: https://huggingface.co/datasets/MinKeonKim/PRO-STEP-Preference-Data.personal_preference_eval
Dataset Card for personal_preference_eval
Dataset Description
Dataset for personal preference eval in paper "Linear Alignment: A Closed-form Solution for Aligning Human Preferences without Tuning and Feedback"
Field Description
Field Name
Field Description
index
Index of data point.
domain
Domain of question.
question
User query.
preference_a
Description of user_a.
preference_b
Description of user_b.
preference_c
Description of user_c.… See the full description on the dataset page: https://huggingface.co/datasets/kkuusou/personal_preference_eval.tr-dpo-preferences
tr-dpo-preferences
Türkçe Direct Preference Optimization (DPO) eğitimi için
hazırlanmış sentetik tercih dataseti.
Bence Gayet İyi Bir İş Yapıyorum, demi?
dolphin-sft-v0.1-preferenceThe preference dataset was generated using Mistral-Instruct-v0.1 finetuned on a GPT-4 subset of the Dolphin dataset (16k samples). Link to the model.
Generated responses are labeled as rejected, GPT-4 responses (original Dolphin data) are labeled as accepted.
The motivation was to test out the SPIN paper finetuning methodology.
orm-pairwise-preference-pairs
Pairwise Outcome Reward Model (ORM)
A Robust Preference Learning Model for Agentic Reasoning Systems
📋 Model Description
This is a Pairwise Outcome Reward Model (ORM) designed for agentic reasoning systems. The model learns to rank reasoning traces through relative preference judgments rather than absolute quality scores, achieving superior stability and reproducibility compared to traditional pointwise approaches.
Key Achievements:
✅ 96.3% pairwise accuracy with… See the full description on the dataset page: https://huggingface.co/datasets/LossFunctionLover/orm-pairwise-preference-pairs.LLaVA-Human-Preference-10Kdfm13-arena-human-preference-100k-preferred
dfm13-arena-human-preference-100k-preferred
Model-audited preferred responses, including explicitly identified model repairs.
Not certified gold and not manually verified in full. Independent review is
sample-based where declared in the publication receipt; holds are excluded.
Repository split name train is a storage convention, not training admission.
Full original history and target preserved; no truncation or 4096-token cutoff. Training length filtering is separate and not… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm13-arena-human-preference-100k-preferred.user-preference-564k
User Preference Extraction Dataset (564K)
A dataset of 564K examples for training lightweight preference extraction models. Each example pairs a conversation input with structured JSON output describing user preferences as condition-action rules.
This dataset was used to train blackhao0426/pref-extractor-qwen3-0.6b-full-sft, a core component of the VARS framework.
Sample Usage
The following snippet from the official repository demonstrates how to use the framework… See the full description on the dataset page: https://huggingface.co/datasets/blackhao0426/user-preference-564k.lmsys-arena-human-preference-winner-43k-unfiltered
lmsys-arena-human-preference-winner-43k-unfiltered
This repository contains a dataset derived from the lmsys/lmsys-arena-human-preference-55k dataset, which is licensed under the Apache 2.0 License.
Dataset Description
The lmsys-arena-human-preference-winner-43k-unfiltered dataset is a collection of 43,000 samples, each containing an instruction (prompt) and an output (winning response) from real-world user and LLM conversations. The dataset is derived from the original… See the full description on the dataset page: https://huggingface.co/datasets/lesserfield/lmsys-arena-human-preference-winner-43k-unfiltered.Preference-Conditioned-Heterogeneous-MARL-Microgrid
SEGAN OPSD-Derived Microgrid Multiyear Benchmark
This repository contains the processed multiyear microgrid benchmark used for the study
“Preference-Conditioned Heterogeneous Multi-Agent Reinforcement Learning for Safe Microgrid Energy Management.”
Files
microgrid_opsd_multiyear.csv — processed hourly benchmark data.
opsd_multiyear_metadata.json — provenance, selected OPSD nodes, source-column mapping, scaling notes, and processing metadata.… See the full description on the dataset page: https://huggingface.co/datasets/Tristanchou/Preference-Conditioned-Heterogeneous-MARL-Microgrid.kuza_dpo_preferencemed-llm-triage-es-preference
