datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
arena-human-preference-55kDataset for Kaggle competition on predicting human preference on Chatbot Arena battles.
The training dataset includes over 55,000 real-world user and LLM conversations and user preferences across over 70 state-of-the-art LLMs, such as GPT-4, Claude 2, Llama 2, Gemini, and Mistral models.
Each sample represents a battle consisting of 2 LLMs which answer the same question, with a user label of either prefer model A, prefer model B, tie, or tie (both bad).
Citation
Please cite the… See the full description on the dataset page: https://huggingface.co/datasets/lmarena-ai/arena-human-preference-55k.leetcode_preference
Dataset Card for LeetCode Preference
Dataset Summary
This dataset facilitates experiments utilizing Direct Preference Optimization (DPO) as outlined in the paper titled Direct Preference Optimization: Your Language Model is Secretly a Reward Model. This repository provides code pairings crafted by CodeLLaMA-7b. For every LeetCode question posed, CodeLLaMA-7b produces two unique solutions. These are subsequently evaluated and ranked by human experts based on their accuracy… See the full description on the dataset page: https://huggingface.co/datasets/minfeng-ai/leetcode_preference.preference-forecast
HorizonBench Human Longitudinal Preference Dataset
Release Status
Version 1.0.0 is available through gated research access. The structured tables and sanitized participant-authored text pass the documented release checks. Manual review covered every high-priority passage, a fixed random sample of 600 medium-priority passages, and 155 additional medium-priority passages. Text outside the manual sample received the same automated sanitization applied to every row.… See the full description on the dataset page: https://huggingface.co/datasets/stellalisy/preference-forecast.innoduel-rlhf-real-world-human-preferences-sample
Real-World Human Pairwise Preferences — Public Sample
📦 This is a free, public sample of a commercial dataset.
It contains 1,350 rows curated for inspection. The full dataset has 1.5 million
human pairwise-preference decisions.
Full dataset: https://huggingface.co/datasets/NordosoftOy/innoduel-rlhf
Request access / licensing: see § Access to the full dataset — contact kari.nieminen@nordo.fi.
Use this sample to evaluate the data's quality, structure and… See the full description on the dataset page: https://huggingface.co/datasets/NordosoftOy/innoduel-rlhf-real-world-human-preferences-sample.arena-human-preference-55kDataset for Kaggle competition on predicting human preference on Chatbot Arena battles.
The training dataset includes over 55,000 real-world user and LLM conversations and user preferences across over 70 state-of-the-art LLMs, such as GPT-4, Claude 2, Llama 2, Gemini, and Mistral models.
Each sample represents a battle consisting of 2 LLMs which answer the same question, with a user label of either prefer model A, prefer model B, tie, or tie (both bad).
Citation
Please cite the… See the full description on the dataset page: https://huggingface.co/datasets/haiguga/arena-human-preference-55k.gen_image_wordnet_preferencesThis dataset contains generated images. See the associated Hugging Face Collection for examples and additional details: Generated Image Wordnet
Lifestyle_Preferences_Dataset
🏠 Lifestyle Preferences Dataset
A tabular dataset for roommate compatibility, preference modeling, and recommendation systems
cive202/Lifestyle_Preferences_Dataset · Hugging Face Hub
Overview
The Lifestyle Preferences Dataset is a tabular dataset containing lifestyle and preference information for machine learning, data analysis, and recommendation-system applications.
It is well suited to studying relationships between individual lifestyle preferences… See the full description on the dataset page: https://huggingface.co/datasets/cive202/Lifestyle_Preferences_Dataset.PreferenceTravelPlanner
PreferenceTravelPlanner Dataset
PreferenceTravelPlanner is a benchmark crafted for evaluating language agents in tool-use and complex planning within multiple constraints and preferences. For more details, see our paper. It is created by augmenting TravelPlanner (See paper for more details) with several common type of preferences under various preference paradigms.
Introduction
In PreferenceTravelPlanner, for a given query, language agents are expected to formulate a… See the full description on the dataset page: https://huggingface.co/datasets/pensieves/PreferenceTravelPlanner.2026.RA.ToM-Hidden-Preference-Probe
2026.RA.ToM-Hidden-Preference-Probe
Baseline and frozen-probe result tables for a negative interpretability result: in a multi-issue negotiation game family, an opponent's hidden preferences are not decodable from the listening model's frozen residual-stream representations of the dialogue above honest baselines.
What the experiment was
Two or six copies of Qwen/Qwen3-8B negotiate several issues, each holding a private score sheet (its own point values per option… See the full description on the dataset page: https://huggingface.co/datasets/siddharthmb/2026.RA.ToM-Hidden-Preference-Probe.compositional-preference-modeling
Dataset Featurization: Compositional Preference Modeling
This repository contains the datasets used in our case study on compositional preference modeling from Dataset Featurization, demonstrating how our unsupervised featurization pipeline can produce features describing human preferences and match expert-level produced features. This case study is built on top of Compositional Preference Modeling (CPM).
HH-RLHF - Featurization
Utilizing HH-RLHF dataset, we provide… See the full description on the dataset page: https://huggingface.co/datasets/Bravansky/compositional-preference-modeling.sujet-finance-instruct-human-preference-13kask_science_preference_datasetcoding-assistance-preferences
Coding Assitance Preferences
Coding Assistance Preferences is a dataset designed to study how programmers evaluate human and AI-generated responses to Python questions.
Each example presents a StackOverflow question, alongside two answers:
one written by a human on the forum and upvoted by users,
and one generated by an LLM (Gemini-2.0-Flash).
Annotators rated which response they preferred, the type of question (theoretical or practical), whether the responses suggest the same… See the full description on the dataset page: https://huggingface.co/datasets/NaomiDerel/coding-assistance-preferences.arena-human-preference-55kDataset for Kaggle competition on predicting human preference on Chatbot Arena battles.
The training dataset includes over 55,000 real-world user and LLM conversations and user preferences across over 70 state-of-the-art LLMs, such as GPT-4, Claude 2, Llama 2, Gemini, and Mistral models.
Each sample represents a battle consisting of 2 LLMs which answer the same question, with a user label of either prefer model A, prefer model B, tie, or tie (both bad).
Citation
Please cite the… See the full description on the dataset page: https://huggingface.co/datasets/snartkun/arena-human-preference-55k.personalized-preferenceimdb_prefix20_forDPO_multi-preference_smallimdb_prefix20_forDPO_multi-preference_v21000 preferences per person
Non-overlapping prompt-responses for different sub-population
2 sub-populations: one that prefers short answers and one that prefers grammatically correct answers
imdb_prefix20_forDPO_multi-preference_v2_with_clusteringimdb_prefix20_forDPO_multi-preferencepreference-study-data
