Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01cm2435-new /gdpval_preference_rubricsaudion<1K0 likes2.7k downloads6mo agoHugging Face02allenai /tulu-2.5-preference-data Tulu 2.5 Preference Data This dataset contains the preference dataset splits used to train the models described in Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference Feedback. We cleaned and formatted all datasets to be in the same format. This means some splits may differ from their original format. To see the code used for creating most splits, see here. If you only wish to download one dataset, each dataset exists in one file under the data/… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tulu-2.5-preference-data.texttext-generation1M<n<10M18 likes874 downloads2y agoHugging Face03lms-shape-preferences /pairs_Movies_and_TVtextn<1K0 likes401 downloads6mo agoHugging Face04lmarena-ai /webdev-arena-preference-10k WebDev Arena Preference Dataset This dataset contains 10K real-world Webdev Arena battle with 10 state-of-the-art LLMs. More details in the blog post. Dataset License Agreement This Agreement contains the terms and conditions that govern your access and use of the WebDev Arena Dataset (Arena Dataset). You may not use the Arena Dataset if you do not accept this Agreement. By clicking to accept, accessing the Arena Dataset, or both, you hereby agree to the terms of the… See the full description on the dataset page: https://huggingface.co/datasets/lmarena-ai/webdev-arena-preference-10k.text10K<n<100K20 likes395 downloads2y agoHugging Face05lms-shape-preferences /pairs_Grocery_and_Gourmet_Foodtextn<1K0 likes314 downloads6mo agoHugging Face06DebateLabKIT /argunauts-hirpo-preferences Argunauts HIRPO Preferences Preference pairs generated while training Argunaut models with HIRPO Online DPO. texttext-generation100K<n<1M0 likes313 downloads10mo agoHugging Face07rmems /tool-use-preference-pairs Tool Use Preference Pairs Rights & intended use: legacy public research corpus / portfolio artifact. Hosted frontier-model outputs are research-only inputs under project policy (synthetic-factory#161): intended_use: research_only, project_training_policy: blocked. Not training data for any model-weight update. Machine-readable record: rights.json. Release status: The raw, uncurated payload is now published under data/raw/. It is available for inspection and reproducibility… See the full description on the dataset page: https://huggingface.co/datasets/rmems/tool-use-preference-pairs.text1K<n<10K0 likes255 downloads18d agoHugging Face08MSc-Thesis /FinQA-DPO-Statement-Preferencetextn<1K0 likes213 downloads7mo agoHugging Face09Vezora /Code-Preference-PairsCreator Nicolas Mejia-Petit My Kofi Code-Preference-Pairs Dataset Overview This dataset was created while created Open-critic-GPT. Here is a little Overview: The Open-Critic-GPT dataset is a synthetic dataset created to train models in both identifying and fixing bugs in code. The dataset is generated using a unique synthetic data pipeline which involves: Prompting a local model with an existing code example. Introducing bugs into the code. While also having the model… See the full description on the dataset page: https://huggingface.co/datasets/Vezora/Code-Preference-Pairs.text10K<n<100K32 likes211 downloads2y agoHugging Face10ppppqp /vLLM-SR-Preference-V1The files in this repo is the LLM-labeled samples that are used as the training dataset for vLLM-SR Preference model V1. The training file (sharegpt_preference_labeld_with_negative.jsonl) contains 25k records that have sample_id, golden label for the preference-based routing policy, and a set of negative labels that are plausible but do not match the conversation context. The validation file has the same structure, but only 1% of the training file size. The validation file and the training… See the full description on the dataset page: https://huggingface.co/datasets/ppppqp/vLLM-SR-Preference-V1.textn<1K0 likes180 downloads9mo agoHugging Face11tlc4418 /1.4b-policy_preference_data_gold_labelledPreference dataset using labels from the AlpacaFarm dataset, generated answers from a 1.4b fine-tuned Pythia policy model, and labelled using the AlpacaFarm 'reward-model-human' as a gold reward model. Used to train reward models in 'Reward Model Ensembles Mitigate Overoptimization' text10K<n<100K0 likes174 downloads2y agoHugging Face12davanstrien /magpie-preference Dataset Card for Magpie Preference Dataset Dataset Description The Magpie Preference Dataset is a crowdsourced collection of human preferences on synthetic instruction-response pairs generated using the Magpie approach. This dataset is continuously updated through user interactions with the Magpie Preference Gradio Space. What is Magpie? Magpie is a very interesting new approach to creating synthetic data which doesn't require any seed data:… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/magpie-preference.textn<1K15 likes155 downloads17d agoHugging Face13shibing624 /DPO-En-Zh-20k-PreferenceThis dataset is composed by 4,000 examples of argilla/distilabel-capybara-dpo-7k-binarized with chosen score>=4. 3,000 examples of argilla/distilabel-intel-orca-dpo-pairs with chosen score>=8. 3,000 examples of argilla/ultrafeedback-binarized-preferences-cleaned with chosen score>=4. 10,000 examples of wenbopan/Chinese-dpo-pairs. refer: https://huggingface.co/datasets/hiyouga/DPO-En-Zh-20k 改了question、response_rejected、response_chosen字段,方便ORPO、DPO模型训练时使用train usage:… See the full description on the dataset page: https://huggingface.co/datasets/shibing624/DPO-En-Zh-20k-Preference.texttext-generation10K<n<100K18 likes144 downloads2y agoHugging Face14UCL-DARK /openai-tldr-summarisation-preferences Human feedback data This is the version of the dataset used in https://arxiv.org/abs/2310.06452. If starting a new project we would recommend using https://huggingface.co/datasets/openai/summarize_from_feedback. See https://github.com/openai/summarize-from-feedback for original details of the dataset. Here the data is formatted to enable huggingface transformers sequence classification models to be trained as reward functions. texttext-classification100K<n<1M2 likes143 downloads3y agoHugging Face15davidberenstein1957 /dataset-viber-image-generation-preference-inference-endpoints-battle-flux Dataset Card for Dataset Name Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More Information Needed] Paper [optional]: [More Information Needed] Demo [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/davidberenstein1957/dataset-viber-image-generation-preference-inference-endpoints-battle-flux.imagen<1K0 likes134 downloads2y agoHugging Face16ProVoice-proactivity /proactivity_preference_dataset ProVoice study 1 — driver state, vehicle context and preferred Level of Autonomy Driving-simulator data from the population data collection of the ProVoice / ProActivity project (CARLA 0.10): 12 drivers × 2 sessions, ~20 Hz multimodal driver-state and vehicle frames, and 1,446 driver-assigned Level-of-Autonomy (LoA) labels stating how autonomously an in-vehicle assistant should act on a given task. Drivers were prompted every 20 s about two randomly drawn in-vehicle tasks and… See the full description on the dataset page: https://huggingface.co/datasets/ProVoice-proactivity/proactivity_preference_dataset.tabular1M<n<10M0 likes124 downloads25d agoHugging Face17stellalisy /MediQ_AskDocs_preferencegated This dataset is the preference data subset of MediQ_AskDocs, for the SFT subset, see MediQ_AskDocs. To cite: @misc{li2025aligningllmsaskgood, title={Aligning LLMs to Ask Good Questions A Case Study in Clinical Reasoning}, author={Shuyue Stella Li and Jimin Mun and Faeze Brahman and Jonathan S. Ilgen and Yulia Tsvetkov and Maarten Sap}, year={2025}, eprint={2502.14860}, archivePrefix={arXiv}, primaryClass={cs.CL}… See the full description on the dataset page: https://huggingface.co/datasets/stellalisy/MediQ_AskDocs_preference.text100K<n<1M2 likes116 downloads2y agoHugging Face18lmarena-ai /repochat-arena-preference-4k Overview This dataset contains leaderboard vote data on RepoChat collected from 2024/11/30 to 2025/02/03 For reproducing the leaderboards from this data, refer to the notebook. License User prompts are licensed under CC-BY-4.0, and model outputs are governed by the terms of use set by the respective model providers. text1K<n<10K4 likes116 downloads2y agoHugging Face19MinKeonKim /PRO-STEP-Preference-Data PRO-STEP: DPO Preference Pairs Step-level preference pairs used to train the PRO-STEP policy model via Direct Preference Optimization. Paper: PRO-STEP: Step-level Process Reward Optimization for Retrieval-Augmented GenerationCode: GitHub Repository Pairs: 15,877 (after outcome filter) Source questions: 5,000 from HotpotQA + MuSiQue + 2WikiMultiHopQA training splits Generation: PRM-guided MCTS (K=3 branching, depth 7, 64 rollouts/question, V(s) = Q̄(s) + α · r̂(s) with α=0.3)… See the full description on the dataset page: https://huggingface.co/datasets/MinKeonKim/PRO-STEP-Preference-Data.tabulartext-generation10K<n<100K0 likes115 downloads1mo agoHugging Face20kkuusou /personal_preference_eval Dataset Card for personal_preference_eval Dataset Description Dataset for personal preference eval in paper "Linear Alignment: A Closed-form Solution for Aligning Human Preferences without Tuning and Feedback" Field Description Field Name Field Description index Index of data point. domain Domain of question. question User query. preference_a Description of user_a. preference_b Description of user_b. preference_c Description of user_c.… See the full description on the dataset page: https://huggingface.co/datasets/kkuusou/personal_preference_eval.textn<1K5 likes113 downloads3y agoHugging Face21Swagvictoria /tr-dpo-preferences tr-dpo-preferences Türkçe Direct Preference Optimization (DPO) eğitimi için hazırlanmış sentetik tercih dataseti. Bence Gayet İyi Bir İş Yapıyorum, demi? text1K<n<10K1 likes108 downloads2d agoHugging Face22reciperesearch /dolphin-sft-v0.1-preferenceThe preference dataset was generated using Mistral-Instruct-v0.1 finetuned on a GPT-4 subset of the Dolphin dataset (16k samples). Link to the model. Generated responses are labeled as rejected, GPT-4 responses (original Dolphin data) are labeled as accepted. The motivation was to test out the SPIN paper finetuning methodology. texttext-generation10K<n<100K11 likes103 downloads2y agoHugging Face23LossFunctionLover /orm-pairwise-preference-pairs Pairwise Outcome Reward Model (ORM) A Robust Preference Learning Model for Agentic Reasoning Systems 📋 Model Description This is a Pairwise Outcome Reward Model (ORM) designed for agentic reasoning systems. The model learns to rank reasoning traces through relative preference judgments rather than absolute quality scores, achieving superior stability and reproducibility compared to traditional pointwise approaches. Key Achievements: ✅ 96.3% pairwise accuracy with… See the full description on the dataset page: https://huggingface.co/datasets/LossFunctionLover/orm-pairwise-preference-pairs.text10K<n<100K0 likes86 downloads9mo agoHugging Face24zhiqings /LLaVA-Human-Preference-10Ktabular1K<n<10K34 likes83 downloads3y agoHugging Face25schneiderkamplab /dfm13-arena-human-preference-100k-preferred dfm13-arena-human-preference-100k-preferred Model-audited preferred responses, including explicitly identified model repairs. Not certified gold and not manually verified in full. Independent review is sample-based where declared in the publication receipt; holds are excluded. Repository split name train is a storage convention, not training admission. Full original history and target preserved; no truncation or 4096-token cutoff. Training length filtering is separate and not… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm13-arena-human-preference-100k-preferred.texttext-generation10K<n<100K0 likes83 downloads9d agoHugging Face26blackhao0426 /user-preference-564k User Preference Extraction Dataset (564K) A dataset of 564K examples for training lightweight preference extraction models. Each example pairs a conversation input with structured JSON output describing user preferences as condition-action rules. This dataset was used to train blackhao0426/pref-extractor-qwen3-0.6b-full-sft, a core component of the VARS framework. Sample Usage The following snippet from the official repository demonstrates how to use the framework… See the full description on the dataset page: https://huggingface.co/datasets/blackhao0426/user-preference-564k.texttext-generation100K<n<1M2 likes76 downloads7mo agoHugging Face27lesserfield /lmsys-arena-human-preference-winner-43k-unfiltered lmsys-arena-human-preference-winner-43k-unfiltered This repository contains a dataset derived from the lmsys/lmsys-arena-human-preference-55k dataset, which is licensed under the Apache 2.0 License. Dataset Description The lmsys-arena-human-preference-winner-43k-unfiltered dataset is a collection of 43,000 samples, each containing an instruction (prompt) and an output (winning response) from real-world user and LLM conversations. The dataset is derived from the original… See the full description on the dataset page: https://huggingface.co/datasets/lesserfield/lmsys-arena-human-preference-winner-43k-unfiltered.texttext-generation10K<n<100K2 likes74 downloads2y agoHugging Face28Tristanchou /Preference-Conditioned-Heterogeneous-MARL-Microgrid SEGAN OPSD-Derived Microgrid Multiyear Benchmark This repository contains the processed multiyear microgrid benchmark used for the study “Preference-Conditioned Heterogeneous Multi-Agent Reinforcement Learning for Safe Microgrid Energy Management.” Files microgrid_opsd_multiyear.csv — processed hourly benchmark data. opsd_multiyear_metadata.json — provenance, selected OPSD nodes, source-column mapping, scaling notes, and processing metadata.… See the full description on the dataset page: https://huggingface.co/datasets/Tristanchou/Preference-Conditioned-Heterogeneous-MARL-Microgrid.tabulartime-series-forecastingn<1K0 likes74 downloads28d agoHugging Face29kuzaai /kuza_dpo_preferencetext1K<n<10K0 likes74 downloads21d agoHugging Face30apol /med-llm-triage-es-preferencetext1K<n<10K0 likes72 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.