Team Ai
20 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01lmarena-ai /arena-human-preference-55kDataset for Kaggle competition on predicting human preference on Chatbot Arena battles. The training dataset includes over 55,000 real-world user and LLM conversations and user preferences across over 70 state-of-the-art LLMs, such as GPT-4, Claude 2, Llama 2, Gemini, and Mistral models. Each sample represents a battle consisting of 2 LLMs which answer the same question, with a user label of either prefer model A, prefer model B, tie, or tie (both bad). Citation Please cite the… See the full description on the dataset page: https://huggingface.co/datasets/lmarena-ai/arena-human-preference-55k.tabulartext-classification10K<n<100K159 likes2.4k downloads2y agoHugging Face02minfeng-ai /leetcode_preference Dataset Card for LeetCode Preference Dataset Summary This dataset facilitates experiments utilizing Direct Preference Optimization (DPO) as outlined in the paper titled Direct Preference Optimization: Your Language Model is Secretly a Reward Model. This repository provides code pairings crafted by CodeLLaMA-7b. For every LeetCode question posed, CodeLLaMA-7b produces two unique solutions. These are subsequently evaluated and ranked by human experts based on their accuracy… See the full description on the dataset page: https://huggingface.co/datasets/minfeng-ai/leetcode_preference.tabularn<1K7 likes68 downloads3y agoHugging Face03stellalisy /preference-forecastgated HorizonBench Human Longitudinal Preference Dataset Release Status Version 1.0.0 is available through gated research access. The structured tables and sanitized participant-authored text pass the documented release checks. Manual review covered every high-priority passage, a fixed random sample of 600 medium-priority passages, and 155 additional medium-priority passages. Text outside the manual sample received the same automated sanitization applied to every row.… See the full description on the dataset page: https://huggingface.co/datasets/stellalisy/preference-forecast.tabular100K<n<1M0 likes40 downloads24d agoHugging Face04NordosoftOy /innoduel-rlhf-real-world-human-preferences-sample Real-World Human Pairwise Preferences — Public Sample 📦 This is a free, public sample of a commercial dataset. It contains 1,350 rows curated for inspection. The full dataset has 1.5 million human pairwise-preference decisions. Full dataset: https://huggingface.co/datasets/NordosoftOy/innoduel-rlhf Request access / licensing: see § Access to the full dataset — contact kari.nieminen@nordo.fi. Use this sample to evaluate the data's quality, structure and… See the full description on the dataset page: https://huggingface.co/datasets/NordosoftOy/innoduel-rlhf-real-world-human-preferences-sample.tabulartext-generation1K<n<10K0 likes28 downloads2mo agoHugging Face05haiguga /arena-human-preference-55kDataset for Kaggle competition on predicting human preference on Chatbot Arena battles. The training dataset includes over 55,000 real-world user and LLM conversations and user preferences across over 70 state-of-the-art LLMs, such as GPT-4, Claude 2, Llama 2, Gemini, and Mistral models. Each sample represents a battle consisting of 2 LLMs which answer the same question, with a user label of either prefer model A, prefer model B, tie, or tie (both bad). Citation Please cite the… See the full description on the dataset page: https://huggingface.co/datasets/haiguga/arena-human-preference-55k.tabulartext-classification10K<n<100K0 likes27 downloads9mo agoHugging Face06VityaVitalich /gen_image_wordnet_preferencesThis dataset contains generated images. See the associated Hugging Face Collection for examples and additional details: Generated Image Wordnet tabulartext-to-image1K<n<10K0 likes25 downloads2y agoHugging Face07cive202 /Lifestyle_Preferences_Dataset 🏠 Lifestyle Preferences Dataset A tabular dataset for roommate compatibility, preference modeling, and recommendation systems cive202/Lifestyle_Preferences_Dataset · Hugging Face Hub Overview The Lifestyle Preferences Dataset is a tabular dataset containing lifestyle and preference information for machine learning, data analysis, and recommendation-system applications. It is well suited to studying relationships between individual lifestyle preferences… See the full description on the dataset page: https://huggingface.co/datasets/cive202/Lifestyle_Preferences_Dataset.tabular10K<n<100K1 likes21 downloads2mo agoHugging Face08pensieves /PreferenceTravelPlanner PreferenceTravelPlanner Dataset PreferenceTravelPlanner is a benchmark crafted for evaluating language agents in tool-use and complex planning within multiple constraints and preferences. For more details, see our paper. It is created by augmenting TravelPlanner (See paper for more details) with several common type of preferences under various preference paradigms. Introduction In PreferenceTravelPlanner, for a given query, language agents are expected to formulate a… See the full description on the dataset page: https://huggingface.co/datasets/pensieves/PreferenceTravelPlanner.tabulartext-generation1K<n<10K0 likes19 downloads7mo agoHugging Face09siddharthmb /2026.RA.ToM-Hidden-Preference-Probe 2026.RA.ToM-Hidden-Preference-Probe Baseline and frozen-probe result tables for a negative interpretability result: in a multi-issue negotiation game family, an opponent's hidden preferences are not decodable from the listening model's frozen residual-stream representations of the dialogue above honest baselines. What the experiment was Two or six copies of Qwen/Qwen3-8B negotiate several issues, each holding a private score sheet (its own point values per option… See the full description on the dataset page: https://huggingface.co/datasets/siddharthmb/2026.RA.ToM-Hidden-Preference-Probe.tabular10K<n<100K0 likes19 downloads2mo agoHugging Face10Bravansky /compositional-preference-modeling Dataset Featurization: Compositional Preference Modeling This repository contains the datasets used in our case study on compositional preference modeling from Dataset Featurization, demonstrating how our unsupervised featurization pipeline can produce features describing human preferences and match expert-level produced features. This case study is built on top of Compositional Preference Modeling (CPM). HH-RLHF - Featurization Utilizing HH-RLHF dataset, we provide… See the full description on the dataset page: https://huggingface.co/datasets/Bravansky/compositional-preference-modeling.tabularfeature-extraction10K<n<100K0 likes10 downloads2y agoHugging Face11seeyssimon /sujet-finance-instruct-human-preference-13ktabular10K<n<100K0 likes7 downloads2y agoHugging Face12mattany /ask_science_preference_datasettabular1K<n<10K0 likes6 downloads1y agoHugging Face13NaomiDerel /coding-assistance-preferences Coding Assitance Preferences Coding Assistance Preferences is a dataset designed to study how programmers evaluate human and AI-generated responses to Python questions. Each example presents a StackOverflow question, alongside two answers: one written by a human on the forum and upvoted by users, and one generated by an LLM (Gemini-2.0-Flash). Annotators rated which response they preferred, the type of question (theoretical or practical), whether the responses suggest the same… See the full description on the dataset page: https://huggingface.co/datasets/NaomiDerel/coding-assistance-preferences.tabulartext-classificationn<1K0 likes6 downloads1y agoHugging Face14snartkun /arena-human-preference-55kDataset for Kaggle competition on predicting human preference on Chatbot Arena battles. The training dataset includes over 55,000 real-world user and LLM conversations and user preferences across over 70 state-of-the-art LLMs, such as GPT-4, Claude 2, Llama 2, Gemini, and Mistral models. Each sample represents a battle consisting of 2 LLMs which answer the same question, with a user label of either prefer model A, prefer model B, tie, or tie (both bad). Citation Please cite the… See the full description on the dataset page: https://huggingface.co/datasets/snartkun/arena-human-preference-55k.tabulartext-classification10K<n<100K0 likes6 downloads3mo agoHugging Face15tools-o /personalized-preferencetabularn<1K0 likes5 downloads2y agoHugging Face16keertanavc /imdb_prefix20_forDPO_multi-preference_smalltabular10K<n<100K0 likes4 downloads2y agoHugging Face17keertanavc /imdb_prefix20_forDPO_multi-preference_v21000 preferences per person Non-overlapping prompt-responses for different sub-population 2 sub-populations: one that prefers short answers and one that prefers grammatically correct answers tabular10K<n<100K0 likes4 downloads2y agoHugging Face18kvseet17 /imdb_prefix20_forDPO_multi-preference_v2_with_clusteringtabular10K<n<100K0 likes4 downloads2y agoHugging Face19keertanavc /imdb_prefix20_forDPO_multi-preferencetabular10K<n<100K0 likes3 downloads2y agoHugging Face20andrew-gordon /preference-study-datatabularn<1K0 likes3 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.