datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tulu-2.5-preference-data
Tulu 2.5 Preference Data
This dataset contains the preference dataset splits used to train the models described in Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference Feedback.
We cleaned and formatted all datasets to be in the same format.
This means some splits may differ from their original format.
To see the code used for creating most splits, see here.
If you only wish to download one dataset, each dataset exists in one file under the data/… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tulu-2.5-preference-data.1.4b-policy_preference_data_gold_labelledPreference dataset using labels from the AlpacaFarm dataset, generated answers from a 1.4b fine-tuned Pythia policy model, and labelled using the AlpacaFarm 'reward-model-human' as a gold reward model.
Used to train reward models in 'Reward Model Ensembles Mitigate Overoptimization'
dataset-viber-image-generation-preference-inference-endpoints-battle-flux
Dataset Card for Dataset Name
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/davidberenstein1957/dataset-viber-image-generation-preference-inference-endpoints-battle-flux.proactivity_preference_dataset
ProVoice study 1 — driver state, vehicle context and preferred Level of Autonomy
Driving-simulator data from the population data collection of the ProVoice /
ProActivity project (CARLA 0.10): 12 drivers × 2 sessions, ~20 Hz
multimodal driver-state and vehicle frames, and 1,446 driver-assigned
Level-of-Autonomy (LoA) labels stating how autonomously an in-vehicle
assistant should act on a given task. Drivers were prompted every 20 s about
two randomly drawn in-vehicle tasks and… See the full description on the dataset page: https://huggingface.co/datasets/ProVoice-proactivity/proactivity_preference_dataset.PRO-STEP-Preference-Data
PRO-STEP: DPO Preference Pairs
Step-level preference pairs used to train the PRO-STEP policy model via Direct Preference Optimization.
Paper: PRO-STEP: Step-level Process Reward Optimization for Retrieval-Augmented GenerationCode: GitHub Repository
Pairs: 15,877 (after outcome filter)
Source questions: 5,000 from HotpotQA + MuSiQue + 2WikiMultiHopQA training splits
Generation: PRM-guided MCTS (K=3 branching, depth 7, 64 rollouts/question, V(s) = Q̄(s) + α · r̂(s) with α=0.3)… See the full description on the dataset page: https://huggingface.co/datasets/MinKeonKim/PRO-STEP-Preference-Data.POVID_preference_data_for_VLLMspreference-data-math-stack-exchangeThe preference dataset is derived from the stack exchange dataset which contains questions and answers from the Stack Overflow Data Dump. This contains questions and answers for various topics. For this work, we used only question and answers from math.stackexchange.com sub-folder.
The questions are grouped with answers that are assigned a score corresponding to the Anthropic paper:
score = log2 (1 + upvotes) rounded to the nearest integer, plus 1 if the answer was accepted by the questioner… See the full description on the dataset page: https://huggingface.co/datasets/prhegde/preference-data-math-stack-exchange.argument_generation_preference_dataFallacy types mapping:
'Not a Fallacy': 0
'faulty generalization': 1
'false causality': 2
'fallacy of relevance': 3
'fallacy of extension': 4
'equivocation': 5
'ad populum': 6
'appeal to emotion': 7
'ad hominem': 8
'circular reasoning': 9
'fallacy of credibility': 10
'fallacy of logic': 11
'false dilemma': 12
'intentional': 13
Arabic-preference-data-RLHFmDPO-preference-dataDataset derived from VLFeedback
Images can be found in this zip file
tiny-preference-datasetSky-T1_preference_data_10kpharma-preference-dataset
Pharma DPO Preference Dataset
Pharmaceutical domain preference dataset used for
Direct Preference Optimization (DPO) — Stage 3 of the pharma TinyLlama
fine-tuning pipeline.
Format
Each JSONL record contains 3 fields:
{
"prompt": "### Instruction:\nExplain the mechanism of metformin.\n\n### Response:\n",
"chosen": "Metformin primarily works by ...",
"rejected": "Metformin is a drug that ..."
}
prompt — Alpaca-style instruction prompt (same format as… See the full description on the dataset page: https://huggingface.co/datasets/ThakrePranjal/pharma-preference-dataset.bias_reduce_preference_data
Dataset Card for Persona-Aware Preference Dataset
Dataset Description
This is a Direct Preference Optimization (DPO) dataset designed to train language models to produce high-quality, context-aware responses when given user demographic information (persona). Each example pairs a user prompt prefixed with a demographic persona description with a chosen (preferred) response and a rejected (dispreferred) response.
The dataset is intended to support alignment research focused… See the full description on the dataset page: https://huggingface.co/datasets/groupfairnessllm/bias_reduce_preference_data.Prompt_Preference_DatasetA preference dataset for end2end prompt optimization. Check our usage here.
han-human-preference-assist-dataset-v1
Human Preference Assist Dataset
Overview
A dataset capturing user-specific
preferences during humanoid assistance tasks.
Supports personalization and adaptive interaction.
Data Fields
user_id
preferred_task_style
preferred_speed
interaction_tone
confirmation_required
Intended Use
Personalized robotics systems
Adaptive assistance research
Human-robot interaction modeling
License
MIT
pairrm-preference-datasetPairwise preference dataset generated using Mistral + PairRM.
dataset-viber-chat-generation-preference-inference-endpoints-battle
Dataset Card for Dataset Name
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/davidberenstein1957/dataset-viber-chat-generation-preference-inference-endpoints-battle.SeppeV__SmolLM_pretrained_with_sft_trained_with_1pc_data_on_a_preference_dpo-details
Dataset Card for Evaluation run of SeppeV/SmolLM_pretrained_with_sft_trained_with_1pc_data_on_a_preference_dpo
Dataset automatically created during the evaluation run of model SeppeV/SmolLM_pretrained_with_sft_trained_with_1pc_data_on_a_preference_dpo
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/SeppeV__SmolLM_pretrained_with_sft_trained_with_1pc_data_on_a_preference_dpo-details.pharma-preference-dataset-unsloth
Pharma DPO Preference Dataset — Unsloth Pipeline
Preference dataset in DPO format (prompt / chosen / rejected) used for
Stage 3 DPO training in the Unsloth 3-stage pharma fine-tuning pipeline.
Format
{
"prompt": "### Instruction:\nExplain the mechanism of metformin.\n\n### Response:",
"chosen": "Metformin primarily acts by activating AMPK...",
"rejected": "Metformin mainly works by increasing insulin secretion..."
}
Stats
Total rows: 48… See the full description on the dataset page: https://huggingface.co/datasets/ThakrePranjal/pharma-preference-dataset-unsloth.Hierarchical-Preference-Dataset
Hierarchical Preference Dataset
The Hierarchical Preference Dataset is a structured dataset for analyzing and evaluating model reasoning through a hierarchical cognitive decomposition lens. It is derived from the prhegde/preference-data-math-stack-exchange dataset and extends it with annotations that separate model outputs into Refined Query, Meta-Thinking, and Refined Answer components.
Overview
Each sample in this dataset consists of:
An instruction or query.
Two… See the full description on the dataset page: https://huggingface.co/datasets/Death-Raider/Hierarchical-Preference-Dataset.sherlock_preference_datasetThis dataset contains preference data for tuning Vision-Language models on the Sherlock Dataset for Abductive Reasoning. It is designed to evaluate the effectiveness of fine-tuning using Supervised Fine-Tuning (SFT) or Preference Optimization. Preferences are generated by prompting four models: mistralai/Pixtral-12B-2409, Qwen/Qwen2-VL-7B-Instruct, google/paligemma2-3b-ft-docci-448, and google/paligemma2-10b-ft-docci-448.
Since this dataset is intended for optimizing PaLI-Gemma models… See the full description on the dataset page: https://huggingface.co/datasets/akshayg08/sherlock_preference_dataset.synthetic_preference_dataset_multi_1746753473
allenai/open_instruct: Rejection Sampling Dataset
See https://github.com/allenai/open-instruct/blob/main/docs/algorithms/rejection_sampling.md for more detail
Configs
args:
{'add_timestamp': True,
'hf_entity': 'VGraf',
'hf_repo_id': 'synthetic_preference_dataset_multi',
'hf_repo_id_scores': 'synthetic_preference_dataset_multi_scores',
'input_filename':… See the full description on the dataset page: https://huggingface.co/datasets/VGraf/synthetic_preference_dataset_multi_1746753473.assn2-preference-datasynthetic_preference_dataset_multi_1741068828
allenai/open_instruct: Rejection Sampling Dataset
See https://github.com/allenai/open-instruct/blob/main/docs/algorithms/rejection_sampling.md for more detail
Configs
args:
{'add_timestamp': True,
'hf_entity': 'VGraf',
'hf_repo_id': 'synthetic_preference_dataset_multi',
'hf_repo_id_scores': 'synthetic_preference_dataset_multi_scores',
'input_filename': '/weka/oe-adapt-default/victoriag/synth_data/completions.jsonl',
'max_parallel_requests': 100,
'model':… See the full description on the dataset page: https://huggingface.co/datasets/VGraf/synthetic_preference_dataset_multi_1741068828.synthetic_preference_dataset_multi_1746753469
allenai/open_instruct: Rejection Sampling Dataset
See https://github.com/allenai/open-instruct/blob/main/docs/algorithms/rejection_sampling.md for more detail
Configs
args:
{'add_timestamp': True,
'hf_entity': 'VGraf',
'hf_repo_id': 'synthetic_preference_dataset_multi',
'hf_repo_id_scores': 'synthetic_preference_dataset_multi_scores',
'input_filename':… See the full description on the dataset page: https://huggingface.co/datasets/VGraf/synthetic_preference_dataset_multi_1746753469.bpo_preference_hh_dataPrompt_Preference_DatasetA preference dataset for end2end prompt optimization. Check our usage here.
assignment4-preference-datasetassignment4_preference_dataset
assignment4_preference_dataset
This dataset contains pairwise preference data for Assignment 4.
Files
assignment4_preference_pairs.jsonl: Main preference dataset in JSONL format.
assignment4_preference_pairs.csv: CSV version for quick inspection.
Schema (JSONL)
Each line stores one preference sample with:
instruction/prompt text
chosen response
rejected response
optional metadata fields
Usage
Use this dataset for reward modeling, preference… See the full description on the dataset page: https://huggingface.co/datasets/SuperSteel/assignment4_preference_dataset.
