datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ultrafeedback-binarized-preferences-cleaned
UltraFeedback - Binarized using the Average of Preference Ratings (Cleaned)
This dataset represents a new iteration on top of argilla/ultrafeedback-binarized-preferences,
and is the recommended and preferred dataset by Argilla to use from now on when fine-tuning on UltraFeedback.
Read more about Argilla's approach towards UltraFeedback binarization at argilla/ultrafeedback-binarized-preferences/README.md.
Differences with argilla/ultrafeedback-binarized-preferences… See the full description on the dataset page: https://huggingface.co/datasets/argilla/ultrafeedback-binarized-preferences-cleaned.ultrafeedback-binarized-preferences-cleaned-kto
UltraFeedback - Binarized using the Average of Preference Ratings (Cleaned) KTO
A KTO signal transformed version of the highly loved UltraFeedback Binarized Preferences Cleaned, the preferred dataset by Argilla to use from now on when fine-tuning on UltraFeedback
This dataset represents a new iteration on top of argilla/ultrafeedback-binarized-preferences,
and is the recommended and preferred dataset by Argilla to use from now on when fine-tuning on UltraFeedback.
Read more about… See the full description on the dataset page: https://huggingface.co/datasets/argilla/ultrafeedback-binarized-preferences-cleaned-kto.stack-exchange-preferences
Dataset Card for H4 Stack Exchange Preferences Dataset
Dataset Summary
This dataset contains questions and answers from the Stack Overflow Data Dump for the purpose of preference model training.
Importantly, the questions have been filtered to fit the following criteria for preference models (following closely from Askell et al. 2021): have >=2 answers.
This data could also be used for instruction fine-tuning and language model training.
The questions are grouped with… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceH4/stack-exchange-preferences.open-image-preferences-v1
Open Image Preferences
Prompt: Anime-style concept art of a Mayan Quetzalcoatl biomutant, dystopian world, vibrant colors, 4K.
Image 1
Image 2
Prompt: 8-bit pixel art of a blue knight, green car, and glacier landscape in Norway, fantasy style, colorful and detailed.
Image 1… See the full description on the dataset page: https://huggingface.co/datasets/data-is-better-together/open-image-preferences-v1.preference-test-sets
Preference Test Sets
Very few preference datasets have heldout test sets for validation of reward model accuracy results.
In this dataset, we curate the test sets from popular preference datasets into a common schema for easy loading and evaluation.
Anthropic HH (Helpful & Harmless Agent and Red Teaming), test set in full is 8552 samples
Anthropic HHH Alignment (Helpful, Honest, & Harmless), formatted from Big Bench for standalone evaluation.
Learning to summarize, downsampled from… See the full description on the dataset page: https://huggingface.co/datasets/allenai/preference-test-sets.lora-fusing-preferencesgenies_preferences
Dataset Card for "genie_dpo"
A conversion of the distribution from GENIES to open_pref_eval format.
Conversion code
arena-human-preference-140k
Overview
This dataset contains user votes collected in the text-only category. Each row represents a single vote judging two models (model_a and model_b) on a user conversation, along with the full conversation history and metadata. Key fields include:
id: Unique feedback ID of each vote/row.
evaluation_session_id: Unique ID of each evaluation session, which can contain multiple separate votes/evaluations.
evaluation_order: Evaluation order of the current vote.
winner: Battle… See the full description on the dataset page: https://huggingface.co/datasets/lmarena-ai/arena-human-preference-140k.Omni-Preferencellama-3.1-tulu-3-8b-preference-mixture
Tulu 3 8B Preference Mixture
Note that this collection is licensed under ODC-BY-1.0 license; different licenses apply to subsets of the data. Some portions of the dataset are non-commercial. We present the mixture as a research artifact.
This mix is made up from the following preference datasets:
https://huggingface.co/datasets/allenai/tulu-3-sft-reused-off-policy
https://huggingface.co/datasets/allenai/tulu-3-sft-reused-on-policy-8b… See the full description on the dataset page: https://huggingface.co/datasets/allenai/llama-3.1-tulu-3-8b-preference-mixture.UltraMedical-Preference
Dataset Card for Dataset Name
The UltraMedical-Preference Dataset
The UltraMedical-Preference dataset enhances the UltraMedical collection with preference annotations. This dataset includes responses sampled from both open-source and proprietary models, annotated for user preferences.
Models Used for Annotation
For the proprietary models, we utilized gpt-3.5-turbo and gpt-4-turbo. For open-source models, selections included Llama-3-8B/70B, Qwen1.5-72B… See the full description on the dataset page: https://huggingface.co/datasets/TsinghuaC3I/UltraMedical-Preference.gdpval_preference_rubricsarena-human-preference-55kDataset for Kaggle competition on predicting human preference on Chatbot Arena battles.
The training dataset includes over 55,000 real-world user and LLM conversations and user preferences across over 70 state-of-the-art LLMs, such as GPT-4, Claude 2, Llama 2, Gemini, and Mistral models.
Each sample represents a battle consisting of 2 LLMs which answer the same question, with a user label of either prefer model A, prefer model B, tie, or tie (both bad).
Citation
Please cite the… See the full description on the dataset page: https://huggingface.co/datasets/lmarena-ai/arena-human-preference-55k.Skywork-Reward-Preference-80K-v0.2
Skywork Reward Preference 80K
IMPORTANT:
This dataset is the decontaminated version of Skywork-Reward-Preference-80K-v0.1. We removed 4,957 pairs from the magpie-ultra-v0.1 subset that have a significant n-gram overlap with the evaluation prompts in RewardBench. You can find the set of removed pairs here. For more information, see this GitHub gist.
If your task involves evaluation on RewardBench, we strongly encourage you to use v0.2 instead of v0.1 of the dataset.
We will soon… See the full description on the dataset page: https://huggingface.co/datasets/Skywork/Skywork-Reward-Preference-80K-v0.2.world-model-physics-human-preference-283k
Rapidata Physics Benchmark
Built by Rapidata.
Do video and world models understand physics? We gave 25 video- and world models the same
real-world starting frame and scene description from Physics-IQ and asked
them to predict what happens next. ~283,000 human votes, collected with the
Rapidata Python SDK, decided which continuation is more realistic — with the
real recording competing as a hidden 26th participant.
Each row is a head-to-head matchup between two participants on… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/world-model-physics-human-preference-283k.arena-human-preference-100k
Overview
This dataset contains leaderboard conversation data collected between June 2024 and August 2024.
It includes English human preference evaluations used to develop Arena Explorer.
Additionally, we provide an embedding file, which contains precomputed embeddings for the English conversations.
These embeddings are used in the topic modeling pipeline to categorize and analyze these conversations.
For a detailed exploration of the dataset and analysis methods, refer to the… See the full description on the dataset page: https://huggingface.co/datasets/lmarena-ai/arena-human-preference-100k.AA_preference_CherryAA_preference_cocourAA_preference_l0AA_preference_cositulu-2.5-preference-data
Tulu 2.5 Preference Data
This dataset contains the preference dataset splits used to train the models described in Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference Feedback.
We cleaned and formatted all datasets to be in the same format.
This means some splits may differ from their original format.
To see the code used for creating most splits, see here.
If you only wish to download one dataset, each dataset exists in one file under the data/… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tulu-2.5-preference-data.text-2-video-human-preferences
Rapidata Video Generation Preference Dataset
This dataset was collected in ~12 hours using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation.
The data collected in this dataset informs our text-2-video model benchmark. We just started so currently only two models are represented in this set:
Sora
Hunyouan
Pika 2.0
Runway ML Alpha
Luma Ray 2
Explore our latest model rankings on our website.
If you get value from this dataset and would… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-video-human-preferences.WritingPrompts_preferences
Dataset Card for "WritingPrompts_preferences"
Human preference data from r/WritingPrompts
mmlu_preferencesreformat of MMLU to be in DPO (paired) format
examples:
{'prompt': 'Which of the following statements about the lanthanide elements is NOT true?', 'chosen': 'The atomic radii of the lanthanide elements increase across the period from La to Lu.', 'rejected': 'All of the lanthanide elements react with aqueous acid to liberate hydrogen.'}
college_chemistry
{'prompt': 'Beyond the business case for engaging in CSR there are a number of moral arguments relating to: negative _______, the… See the full description on the dataset page: https://huggingface.co/datasets/wassname/mmlu_preferences.PPE-Human-Preference-V1
Overview
This contains the human preference evaluation set for Preference Proxy Evaluations.
This dataset is meant for benchmarking and evaluation, not for training.
Paper
Code
License
User prompts are licensed under CC-BY-4.0, and model outputs are governed by the terms of use set by the respective model providers.
Citation
@misc{frick2024evaluaterewardmodelsrlhf,
title={How to Evaluate Reward Models for RLHF},
author={Evan Frick and Tianle Li and… See the full description on the dataset page: https://huggingface.co/datasets/lmarena-ai/PPE-Human-Preference-V1.human-coherence-preferences-images
Rapidata Image Generation Coherence Dataset
This dataset was collected in ~4 Days using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation.
Explore our latest model rankings on our website.
If you get value from this dataset and would like to see more in the future, please consider liking it.
Overview
One of the largest human annotated coherence datasets for text-to-image models, this release contains over 1,200,000 human… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/human-coherence-preferences-images.trajectory-preference-v1700k_Human_Preference_Dataset_FLUX_SD3_MJ_DALLE3
NOTE: A newer version of this dataset is available Imagen3_Flux1.1_Flux1_SD3_MJ_Dalle_Human_Preference_Dataset
Rapidata Image Generation Preference Dataset
This Dataset is a 1/3 of a 2M+ human annotation dataset that was split into three modalities: Preference, Coherence, Text-to-Image Alignment.
Link to the Coherence dataset: https://huggingface.co/datasets/Rapidata/Flux_SD3_MJ_Dalle_Human_Coherence_Dataset
Link to the Text-2-Image Alignment dataset:… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/700k_Human_Preference_Dataset_FLUX_SD3_MJ_DALLE3.text-2-video-human-preferences-seedance-1-pro
Rapidata Video Generation Seedance 1 Pro Human Preference
In this dataset, ~60k human responses from ~20k human annotators were collected to evaluate Seedance 1 Pro video generation model on our benchmark. This dataset was collected in roughtly 30 min using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation.
Explore our latest model rankings on our website.
If you get value from this dataset and would like to see more in the future, please… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-video-human-preferences-seedance-1-pro.preference-dataset-qwen3
