datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
proactivity_preference_dataset
ProVoice study 1 — driver state, vehicle context and preferred Level of Autonomy
Driving-simulator data from the population data collection of the ProVoice /
ProActivity project (CARLA 0.10): 12 drivers × 2 sessions, ~20 Hz
multimodal driver-state and vehicle frames, and 1,446 driver-assigned
Level-of-Autonomy (LoA) labels stating how autonomously an in-vehicle
assistant should act on a given task. Drivers were prompted every 20 s about
two randomly drawn in-vehicle tasks and… See the full description on the dataset page: https://huggingface.co/datasets/ProVoice-proactivity/proactivity_preference_dataset.PRO-STEP-Preference-Data
PRO-STEP: DPO Preference Pairs
Step-level preference pairs used to train the PRO-STEP policy model via Direct Preference Optimization.
Paper: PRO-STEP: Step-level Process Reward Optimization for Retrieval-Augmented GenerationCode: GitHub Repository
Pairs: 15,877 (after outcome filter)
Source questions: 5,000 from HotpotQA + MuSiQue + 2WikiMultiHopQA training splits
Generation: PRM-guided MCTS (K=3 branching, depth 7, 64 rollouts/question, V(s) = Q̄(s) + α · r̂(s) with α=0.3)… See the full description on the dataset page: https://huggingface.co/datasets/MinKeonKim/PRO-STEP-Preference-Data.LLaVA-Human-Preference-10KPreference-Conditioned-Heterogeneous-MARL-Microgrid
SEGAN OPSD-Derived Microgrid Multiyear Benchmark
This repository contains the processed multiyear microgrid benchmark used for the study
“Preference-Conditioned Heterogeneous Multi-Agent Reinforcement Learning for Safe Microgrid Energy Management.”
Files
microgrid_opsd_multiyear.csv — processed hourly benchmark data.
opsd_multiyear_metadata.json — provenance, selected OPSD nodes, source-column mapping, scaling notes, and processing metadata.… See the full description on the dataset page: https://huggingface.co/datasets/Tristanchou/Preference-Conditioned-Heterogeneous-MARL-Microgrid.warehouse-dpo-preference-pairs
Warehouse Short-Order DPO Preference Pairs
Dataset Description
This dataset contains {prompt, chosen, rejected} preference pairs for
training a warehouse short-order assistant with Direct Preference
Optimization (DPO). Each pair asks a real warehouse-inventory question
(stockout risk, backorders, KPI summaries, why a warehouse is failing
fulfillment - at a single-warehouse, tier, region, or dataset-wide
comparison level) grounded in real tool-call output… See the full description on the dataset page: https://huggingface.co/datasets/EnRaoufi/warehouse-dpo-preference-pairs.preference_alignment_ultra_cutpreference_alignment_totalstudents-subject-preferences
Students' Subject Preferences
A small survey-style dataset recording which school subjects five students like and dislike.
Each row is one student: their ID, the subjects they named as favorites, and the subjects they
named as least favorites. Subject names are in Mongolian Cyrillic.
Files
File
Rows
Description
data/train.jsonl
5
One JSON object per student
Schema
Column
Type
Description
student_id
int
Student identifier… See the full description on the dataset page: https://huggingface.co/datasets/sumya123/students-subject-preferences.preference_tuning_hh_ultraBAAI__Gemma2-9B-IT-Simpo-Infinity-Preference-details
Dataset Card for Evaluation run of BAAI/Gemma2-9B-IT-Simpo-Infinity-Preference
Dataset automatically created during the evaluation run of model BAAI/Gemma2-9B-IT-Simpo-Infinity-Preference
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/BAAI__Gemma2-9B-IT-Simpo-Infinity-Preference-details.ultrafeedback-binarized-preferences-cleaned-deGerman translation from Mixtral (not the best one, and might contain comments etc, despite it prompted not to, but this is mostly for testing purposes atm) of a first part of the dataset as provided by argilla.
argilla_distilabel-math-preference-dpo-PreferenceShareGPTopen_preference_v0.4
This dataset is reformatted version of following datasets
ryota39/synthetic-instruct-gptj-pairwise-ja
ryota39/webgpt_comparisons-ja
label 1 stands for chosen sentence
label 0 stands for rejected sentence
Format
train sample: 199628
validation sample: 1000
test sample: 1000
skip sample: 417
data points which have same chosen and rejected responses were eliminated
{
"index": 33045,
"input": "user: 今年知っておくべき税法の変更点にはどのようなものがありますか。\nassistant:… See the full description on the dataset page: https://huggingface.co/datasets/ryota39/open_preference_v0.4.argilla_ultrafeedback-binarized-preferences-cleaned-PreferenceShareGPTSafety_preferencecx-preference-pairsjudgelm-preference-distributions
JudgeLM Persona Preference Distributions
Multi-annotator pairwise preference labels for training and evaluating probabilistic LLM autoraters, from the paper Judging with Confidence: Calibrating Autoraters to Preference Distributions (EMNLP 2026 Findings).
Each item is a response pair (A, B) from the JudgeLM corpus, judged by a teacher LLM under multiple sampled personas. Aggregating the persona votes yields an empirical preference distribution — a soft target p_b_over_a = Pr[B ≻… See the full description on the dataset page: https://huggingface.co/datasets/CalibratingAutorater/judgelm-preference-distributions.argilla_Capybara-Preferences-PreferenceShareGPTuser_study-preference-personalized_0423_base_filtered
Filtered user study dataset
Source repo: ehejin/user_study-preference-personalized_0423_base
Each row is ONE item review (pre-rating, conversation, post-rating). Submission-level
fields (prolific_pid, demographics, background) are duplicated across rows that share
a submission.
The 25-50 rows here are the FIRST review for each unique pool index, selected the same
way the analysis plot uses — see scripts/plot_vote_shift_3way.py.
Total rows: 50
Magpie-Align_Magpie-Pro-DPO-200K-PreferenceShareGPTSeppeV__SmolLM_pretrained_with_sft_trained_with_1pc_data_on_a_preference_dpo-details
Dataset Card for Evaluation run of SeppeV/SmolLM_pretrained_with_sft_trained_with_1pc_data_on_a_preference_dpo
Dataset automatically created during the evaluation run of model SeppeV/SmolLM_pretrained_with_sft_trained_with_1pc_data_on_a_preference_dpo
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/SeppeV__SmolLM_pretrained_with_sft_trained_with_1pc_data_on_a_preference_dpo-details.assignment4-pairrm-preferences
Assignment 4 Preference Dataset
Generated from GAIR/lima instructions with Qwen2.5-7B-Instruct and ranked with PairRM.
user_study-preference-personalized_0505_NP2_filtered
Filtered user study dataset
Source repo: ehejin/user_study-preference-personalized_0505_NP2
Each row is ONE item review (pre-rating, conversation, post-rating). Submission-level
fields (prolific_pid, demographics, background) are duplicated across rows that share
a submission.
The 25-50 rows here are the FIRST review for each unique pool index, selected the same
way the analysis plot uses — see scripts/plot_vote_shift_3way.py.
Total rows: 50
distilabel-math-preference-dpo-koultrafeedback-binarized-preferences-cleaned-de-2preference_tuningmitigate_preference_dpo
Mitigate toxic self preference with DPO
directory structure
'quality_response/': contains the response for QuALITY dataset.
user_study-preference-personalized_0505_base_filtered
Filtered user study dataset
Source repo: ehejin/user_study-preference-personalized_0505_base
Each row is ONE item review (pre-rating, conversation, post-rating). Submission-level
fields (prolific_pid, demographics, background) are duplicated across rows that share
a submission.
The 25-50 rows here are the FIRST review for each unique pool index, selected the same
way the analysis plot uses — see scripts/plot_vote_shift_3way.py.
Total rows: 50
user_study-preference-281_all_filtered
Combined user study dataset (0505)
Merged from 5 filtered sub-studies. Each row carries a condition
and source_repo field.
Sub-studies:
0505 NP2 → ehejin/user_study-preference-personalized_BASE_filtered
0505 NP1 → ehejin/user_study-preference-personalized_0423_base_filtered
0505 NP3 → ehejin/user_study-preference-personalized_0423_base_personalized_filtered
0505 base → ehejin/user_study-preference-personalized_0505_base_filtered
0505 base personalized →… See the full description on the dataset page: https://huggingface.co/datasets/ehejin/user_study-preference-281_all_filtered.BeaverTails-single-dimension-preference
