datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
arena-human-preference-55kDataset for Kaggle competition on predicting human preference on Chatbot Arena battles.
The training dataset includes over 55,000 real-world user and LLM conversations and user preferences across over 70 state-of-the-art LLMs, such as GPT-4, Claude 2, Llama 2, Gemini, and Mistral models.
Each sample represents a battle consisting of 2 LLMs which answer the same question, with a user label of either prefer model A, prefer model B, tie, or tie (both bad).
Citation
Please cite the… See the full description on the dataset page: https://huggingface.co/datasets/lmarena-ai/arena-human-preference-55k.image-preference-demo
Image dataset for preference aquisition demo
This dataset provides the files used to run the example that we use in this blog post to illustrate how easily
you can set up and run the annotation process to collect a huge preference dataset using Rapidata's API.
The goal is to collect human preferences based on pairwise image matchups.
The dataset contains:
Generated images: A selection of example images generated using Flux.1 and Stable Diffusion. The images are provided in a .zip… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/image-preference-demo.leetcode_preference
Dataset Card for LeetCode Preference
Dataset Summary
This dataset facilitates experiments utilizing Direct Preference Optimization (DPO) as outlined in the paper titled Direct Preference Optimization: Your Language Model is Secretly a Reward Model. This repository provides code pairings crafted by CodeLLaMA-7b. For every LeetCode question posed, CodeLLaMA-7b produces two unique solutions. These are subsequently evaluated and ranked by human experts based on their accuracy… See the full description on the dataset page: https://huggingface.co/datasets/minfeng-ai/leetcode_preference.preference_datasetpreference-forecast
HorizonBench Human Longitudinal Preference Dataset
Release Status
Version 1.0.0 is available through gated research access. The structured tables and sanitized participant-authored text pass the documented release checks. Manual review covered every high-priority passage, a fixed random sample of 600 medium-priority passages, and 155 additional medium-priority passages. Text outside the manual sample received the same automated sanitization applied to every row.… See the full description on the dataset page: https://huggingface.co/datasets/stellalisy/preference-forecast.human_preference_eval_dataset
Human Preference Evaluation Dataset
DigiGreen/human_preference_eval_dataset is a human-annotated preference dataset that supports research and development in evaluating model responses based on human preferences and comparative QA assessment.
This dataset contains real agricultural questions paired with two candidate responses and expert judgments about which response is preferable. It can be used to benchmark preference learning systems, train reward models, and improve… See the full description on the dataset page: https://huggingface.co/datasets/DigiGreen/human_preference_eval_dataset.cleaned-lmsys-arena-human-preference-55k
original dataset
https://huggingface.co/datasets/lmsys/lmsys-arena-human-preference-55k
Use the following code to process the original data to obtain the cleaned data.
import csv
import random
input_file = R'C:\Users\Downloads\train.csv'
output_file = 'cleaned-lmsys-arena-human-preference-55k.csv'
def clean_text(text):
if text.startswith('["') and text.endswith('"]'):
return text[2:-2]
return text
with open(input_file, mode='r', encoding='utf-8') as… See the full description on the dataset page: https://huggingface.co/datasets/REILX/cleaned-lmsys-arena-human-preference-55k.innoduel-rlhf-real-world-human-preferences-sample
Real-World Human Pairwise Preferences — Public Sample
📦 This is a free, public sample of a commercial dataset.
It contains 1,350 rows curated for inspection. The full dataset has 1.5 million
human pairwise-preference decisions.
Full dataset: https://huggingface.co/datasets/NordosoftOy/innoduel-rlhf
Request access / licensing: see § Access to the full dataset — contact kari.nieminen@nordo.fi.
Use this sample to evaluate the data's quality, structure and… See the full description on the dataset page: https://huggingface.co/datasets/NordosoftOy/innoduel-rlhf-real-world-human-preferences-sample.arena-human-preference-55kDataset for Kaggle competition on predicting human preference on Chatbot Arena battles.
The training dataset includes over 55,000 real-world user and LLM conversations and user preferences across over 70 state-of-the-art LLMs, such as GPT-4, Claude 2, Llama 2, Gemini, and Mistral models.
Each sample represents a battle consisting of 2 LLMs which answer the same question, with a user label of either prefer model A, prefer model B, tie, or tie (both bad).
Citation
Please cite the… See the full description on the dataset page: https://huggingface.co/datasets/haiguga/arena-human-preference-55k.gen_image_wordnet_preferencesThis dataset contains generated images. See the associated Hugging Face Collection for examples and additional details: Generated Image Wordnet
JSON_PreferenceLifestyle_Preferences_Dataset
🏠 Lifestyle Preferences Dataset
A tabular dataset for roommate compatibility, preference modeling, and recommendation systems
cive202/Lifestyle_Preferences_Dataset · Hugging Face Hub
Overview
The Lifestyle Preferences Dataset is a tabular dataset containing lifestyle and preference information for machine learning, data analysis, and recommendation-system applications.
It is well suited to studying relationships between individual lifestyle preferences… See the full description on the dataset page: https://huggingface.co/datasets/cive202/Lifestyle_Preferences_Dataset.PreferenceTravelPlanner
PreferenceTravelPlanner Dataset
PreferenceTravelPlanner is a benchmark crafted for evaluating language agents in tool-use and complex planning within multiple constraints and preferences. For more details, see our paper. It is created by augmenting TravelPlanner (See paper for more details) with several common type of preferences under various preference paradigms.
Introduction
In PreferenceTravelPlanner, for a given query, language agents are expected to formulate a… See the full description on the dataset page: https://huggingface.co/datasets/pensieves/PreferenceTravelPlanner.2026.RA.ToM-Hidden-Preference-Probe
2026.RA.ToM-Hidden-Preference-Probe
Baseline and frozen-probe result tables for a negative interpretability result: in a multi-issue negotiation game family, an opponent's hidden preferences are not decodable from the listening model's frozen residual-stream representations of the dialogue above honest baselines.
What the experiment was
Two or six copies of Qwen/Qwen3-8B negotiate several issues, each holding a private score sheet (its own point values per option… See the full description on the dataset page: https://huggingface.co/datasets/siddharthmb/2026.RA.ToM-Hidden-Preference-Probe.JSON_Preference_decomposed
JSON_Preference_decomposed
A length / syntax / semantic decomposition of the original
nruia/JSON_Preference
dataset. For each preference pair (y1, y2), two intermediate responses
y2'' (double prime) and y2' (prime) are added so that the total
alignment gap
G(y1, y2) = log P(y1 | x) - log P(y2 | x)
can be decomposed along a path of intermediate latent representations:
Step
Quantity
Interpretation
1
`log P(y2''
x) - log P(y2
2
`log P(y2'
x) - log P(y2''
3
`log P(y1
x)… See the full description on the dataset page: https://huggingface.co/datasets/Bojian92/JSON_Preference_decomposed.imdb_prefix20_forDPO_gpt2-large-imdb_multi-preferencedahl-preference-kor-med
DAHL: Domain-specific Automated Hallucination Evaluation of Long-Form Text through a Benchmark Dataset in Biomedicine
This repository provides the Korean medical QA preference dataset automatically labeled with DAHL Score.
Link to Paper
Citation
@misc{seo2024dahldomainspecificautomatedhallucination,
title={DAHL: Domain-specific Automated Hallucination Evaluation of Long-Form Text through a Benchmark Dataset in Biomedicine},
author={Jean Seo and Jongwon Lim… See the full description on the dataset page: https://huggingface.co/datasets/seemdog/dahl-preference-kor-med.preference-datasetcompositional-preference-modeling
Dataset Featurization: Compositional Preference Modeling
This repository contains the datasets used in our case study on compositional preference modeling from Dataset Featurization, demonstrating how our unsupervised featurization pipeline can produce features describing human preferences and match expert-level produced features. This case study is built on top of Compositional Preference Modeling (CPM).
HH-RLHF - Featurization
Utilizing HH-RLHF dataset, we provide… See the full description on the dataset page: https://huggingface.co/datasets/Bravansky/compositional-preference-modeling.sujet-finance-instruct-human-preference-13kCulturalKaleidoscope_Preference
🎉 This paper has been accepted for presentation in the Main Track at NAACL 2025.
Project Page: https://neuralsentinel.github.io/KaleidoCulture/
📖 Citation
If you find this useful in your research, please consider citing:
@misc{banerjee2024navigatingculturalkaleidoscopehitchhikers,
title={Navigating the Cultural Kaleidoscope: A Hitchhiker's Guide to Sensitivity in Large Language Models},
author={Somnath Banerjee and Sayan Layek and Hari Shrawgi and Rajarshi… See the full description on the dataset page: https://huggingface.co/datasets/SoftMINER-Group/CulturalKaleidoscope_Preference.PreferenceEvalGPTpairrm-llama-preferences-1744946336
PairRM Preference Dataset
Dataset Description
This dataset contains preference pairs created using the PairRM reward model to evaluate responses generated by the Llama-3.2 model.
Dataset Creation Process
50 instructions were extracted from the Lima dataset
5 responses were generated per instruction using the llama-3.2 chat template
PairRM was applied to create preference pairs
Dataset Statistics
Number of instructions: 50
Number of preference… See the full description on the dataset page: https://huggingface.co/datasets/Likhith003/pairrm-llama-preferences-1744946336.ask_science_preference_datasetcoding-assistance-preferences
Coding Assitance Preferences
Coding Assistance Preferences is a dataset designed to study how programmers evaluate human and AI-generated responses to Python questions.
Each example presents a StackOverflow question, alongside two answers:
one written by a human on the forum and upvoted by users,
and one generated by an LLM (Gemini-2.0-Flash).
Annotators rated which response they preferred, the type of question (theoretical or practical), whether the responses suggest the same… See the full description on the dataset page: https://huggingface.co/datasets/NaomiDerel/coding-assistance-preferences.arena-human-preference-55kDataset for Kaggle competition on predicting human preference on Chatbot Arena battles.
The training dataset includes over 55,000 real-world user and LLM conversations and user preferences across over 70 state-of-the-art LLMs, such as GPT-4, Claude 2, Llama 2, Gemini, and Mistral models.
Each sample represents a battle consisting of 2 LLMs which answer the same question, with a user label of either prefer model A, prefer model B, tie, or tie (both bad).
Citation
Please cite the… See the full description on the dataset page: https://huggingface.co/datasets/snartkun/arena-human-preference-55k.personalized-preferenceimdb_prefix20_forDPO_multi-preference_smallimdb_prefix20_forDPO_multi-preference_v21000 preferences per person
Non-overlapping prompt-responses for different sub-population
2 sub-populations: one that prefers short answers and one that prefers grammatically correct answers
imdb_prefix20_forDPO_multi-preference_v2_with_clustering
