Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01allenai /tulu-2.5-preference-data Tulu 2.5 Preference Data This dataset contains the preference dataset splits used to train the models described in Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference Feedback. We cleaned and formatted all datasets to be in the same format. This means some splits may differ from their original format. To see the code used for creating most splits, see here. If you only wish to download one dataset, each dataset exists in one file under the data/… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tulu-2.5-preference-data.texttext-generation1M<n<10M18 likes874 downloads2y agoHugging Face02tlc4418 /1.4b-policy_preference_data_gold_labelledPreference dataset using labels from the AlpacaFarm dataset, generated answers from a 1.4b fine-tuned Pythia policy model, and labelled using the AlpacaFarm 'reward-model-human' as a gold reward model. Used to train reward models in 'Reward Model Ensembles Mitigate Overoptimization' text10K<n<100K0 likes174 downloads2y agoHugging Face03davidberenstein1957 /dataset-viber-image-generation-preference-inference-endpoints-battle-flux Dataset Card for Dataset Name Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More Information Needed] Paper [optional]: [More Information Needed] Demo [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/davidberenstein1957/dataset-viber-image-generation-preference-inference-endpoints-battle-flux.imagen<1K0 likes134 downloads2y agoHugging Face04ProVoice-proactivity /proactivity_preference_dataset ProVoice study 1 — driver state, vehicle context and preferred Level of Autonomy Driving-simulator data from the population data collection of the ProVoice / ProActivity project (CARLA 0.10): 12 drivers × 2 sessions, ~20 Hz multimodal driver-state and vehicle frames, and 1,446 driver-assigned Level-of-Autonomy (LoA) labels stating how autonomously an in-vehicle assistant should act on a given task. Drivers were prompted every 20 s about two randomly drawn in-vehicle tasks and… See the full description on the dataset page: https://huggingface.co/datasets/ProVoice-proactivity/proactivity_preference_dataset.tabular1M<n<10M0 likes124 downloads25d agoHugging Face05MinKeonKim /PRO-STEP-Preference-Data PRO-STEP: DPO Preference Pairs Step-level preference pairs used to train the PRO-STEP policy model via Direct Preference Optimization. Paper: PRO-STEP: Step-level Process Reward Optimization for Retrieval-Augmented GenerationCode: GitHub Repository Pairs: 15,877 (after outcome filter) Source questions: 5,000 from HotpotQA + MuSiQue + 2WikiMultiHopQA training splits Generation: PRM-guided MCTS (K=3 branching, depth 7, 64 rollouts/question, V(s) = Q̄(s) + α · r̂(s) with α=0.3)… See the full description on the dataset page: https://huggingface.co/datasets/MinKeonKim/PRO-STEP-Preference-Data.tabulartext-generation10K<n<100K0 likes115 downloads1mo agoHugging Face06YiyangAiLab /POVID_preference_data_for_VLLMstext10K<n<100K8 likes68 downloads3y agoHugging Face07prhegde /preference-data-math-stack-exchangeThe preference dataset is derived from the stack exchange dataset which contains questions and answers from the Stack Overflow Data Dump. This contains questions and answers for various topics. For this work, we used only question and answers from math.stackexchange.com sub-folder. The questions are grouped with answers that are assigned a score corresponding to the Anthropic paper: score = log2 (1 + upvotes) rounded to the nearest integer, plus 1 if the answer was accepted by the questioner… See the full description on the dataset page: https://huggingface.co/datasets/prhegde/preference-data-math-stack-exchange.text10K<n<100K6 likes49 downloads3y agoHugging Face08lucamouchel /argument_generation_preference_dataFallacy types mapping: 'Not a Fallacy': 0 'faulty generalization': 1 'false causality': 2 'fallacy of relevance': 3 'fallacy of extension': 4 'equivocation': 5 'ad populum': 6 'appeal to emotion': 7 'ad hominem': 8 'circular reasoning': 9 'fallacy of credibility': 10 'fallacy of logic': 11 'false dilemma': 12 'intentional': 13 texttext-generation1K<n<10K0 likes45 downloads2y agoHugging Face09FreedomIntelligence /Arabic-preference-data-RLHFtext10K<n<100K4 likes40 downloads3y agoHugging Face10fwnlp /mDPO-preference-dataDataset derived from VLFeedback Images can be found in this zip file text1K<n<10K8 likes38 downloads2y agoHugging Face11llamafactory /tiny-preference-datasettextn<1K0 likes34 downloads10mo agoHugging Face12NovaSky-AI /Sky-T1_preference_data_10ktext1K<n<10K15 likes29 downloads2y agoHugging Face13ThakrePranjal /pharma-preference-dataset Pharma DPO Preference Dataset Pharmaceutical domain preference dataset used for Direct Preference Optimization (DPO) — Stage 3 of the pharma TinyLlama fine-tuning pipeline. Format Each JSONL record contains 3 fields: { "prompt": "### Instruction:\nExplain the mechanism of metformin.\n\n### Response:\n", "chosen": "Metformin primarily works by ...", "rejected": "Metformin is a drug that ..." } prompt — Alpaca-style instruction prompt (same format as… See the full description on the dataset page: https://huggingface.co/datasets/ThakrePranjal/pharma-preference-dataset.textn<1K1 likes27 downloads3mo agoHugging Face14groupfairnessllm /bias_reduce_preference_data Dataset Card for Persona-Aware Preference Dataset Dataset Description This is a Direct Preference Optimization (DPO) dataset designed to train language models to produce high-quality, context-aware responses when given user demographic information (persona). Each example pairs a user prompt prefixed with a demographic persona description with a chosen (preferred) response and a rejected (dispreferred) response. The dataset is intended to support alignment research focused… See the full description on the dataset page: https://huggingface.co/datasets/groupfairnessllm/bias_reduce_preference_data.texttext-generationn<1K0 likes21 downloads6mo agoHugging Face15Junrulu /Prompt_Preference_DatasetA preference dataset for end2end prompt optimization. Check our usage here. text10K<n<100K1 likes19 downloads3y agoHugging Face16ariefansclub /han-human-preference-assist-dataset-v1 Human Preference Assist Dataset Overview A dataset capturing user-specific preferences during humanoid assistance tasks. Supports personalization and adaptive interaction. Data Fields user_id preferred_task_style preferred_speed interaction_tone confirmation_required Intended Use Personalized robotics systems Adaptive assistance research Human-robot interaction modeling License MIT textn<1K0 likes19 downloads8mo agoHugging Face17jasperyeoh2 /pairrm-preference-datasetPairwise preference dataset generated using Mistral + PairRM. textn<1K0 likes18 downloads1mo agoHugging Face18davidberenstein1957 /dataset-viber-chat-generation-preference-inference-endpoints-battle Dataset Card for Dataset Name Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More Information Needed] Paper [optional]: [More Information Needed] Demo [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/davidberenstein1957/dataset-viber-chat-generation-preference-inference-endpoints-battle.textn<1K0 likes17 downloads2y agoHugging Face19open-llm-leaderboard /SeppeV__SmolLM_pretrained_with_sft_trained_with_1pc_data_on_a_preference_dpo-detailsgated Dataset Card for Evaluation run of SeppeV/SmolLM_pretrained_with_sft_trained_with_1pc_data_on_a_preference_dpo Dataset automatically created during the evaluation run of model SeppeV/SmolLM_pretrained_with_sft_trained_with_1pc_data_on_a_preference_dpo The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/SeppeV__SmolLM_pretrained_with_sft_trained_with_1pc_data_on_a_preference_dpo-details.tabular10K<n<100K0 likes15 downloads2y agoHugging Face20ThakrePranjal /pharma-preference-dataset-unsloth Pharma DPO Preference Dataset — Unsloth Pipeline Preference dataset in DPO format (prompt / chosen / rejected) used for Stage 3 DPO training in the Unsloth 3-stage pharma fine-tuning pipeline. Format { "prompt": "### Instruction:\nExplain the mechanism of metformin.\n\n### Response:", "chosen": "Metformin primarily acts by activating AMPK...", "rejected": "Metformin mainly works by increasing insulin secretion..." } Stats Total rows: 48… See the full description on the dataset page: https://huggingface.co/datasets/ThakrePranjal/pharma-preference-dataset-unsloth.textn<1K0 likes15 downloads3mo agoHugging Face21Death-Raider /Hierarchical-Preference-Dataset Hierarchical Preference Dataset The Hierarchical Preference Dataset is a structured dataset for analyzing and evaluating model reasoning through a hierarchical cognitive decomposition lens. It is derived from the prhegde/preference-data-math-stack-exchange dataset and extends it with annotations that separate model outputs into Refined Query, Meta-Thinking, and Refined Answer components. Overview Each sample in this dataset consists of: An instruction or query. Two… See the full description on the dataset page: https://huggingface.co/datasets/Death-Raider/Hierarchical-Preference-Dataset.texttext-generation1K<n<10K0 likes13 downloads1y agoHugging Face22akshayg08 /sherlock_preference_datasetThis dataset contains preference data for tuning Vision-Language models on the Sherlock Dataset for Abductive Reasoning. It is designed to evaluate the effectiveness of fine-tuning using Supervised Fine-Tuning (SFT) or Preference Optimization. Preferences are generated by prompting four models: mistralai/Pixtral-12B-2409, Qwen/Qwen2-VL-7B-Instruct, google/paligemma2-3b-ft-docci-448, and google/paligemma2-10b-ft-docci-448. Since this dataset is intended for optimizing PaLI-Gemma models… See the full description on the dataset page: https://huggingface.co/datasets/akshayg08/sherlock_preference_dataset.texttext-generation100K<n<1M0 likes10 downloads2y agoHugging Face23VGraf /synthetic_preference_dataset_multi_1746753473 allenai/open_instruct: Rejection Sampling Dataset See https://github.com/allenai/open-instruct/blob/main/docs/algorithms/rejection_sampling.md for more detail Configs args: {'add_timestamp': True, 'hf_entity': 'VGraf', 'hf_repo_id': 'synthetic_preference_dataset_multi', 'hf_repo_id_scores': 'synthetic_preference_dataset_multi_scores', 'input_filename':… See the full description on the dataset page: https://huggingface.co/datasets/VGraf/synthetic_preference_dataset_multi_1746753473.textn<1K0 likes9 downloads1y agoHugging Face24rivmttt /assn2-preference-datatextn<1K0 likes9 downloads5mo agoHugging Face25VGraf /synthetic_preference_dataset_multi_1741068828 allenai/open_instruct: Rejection Sampling Dataset See https://github.com/allenai/open-instruct/blob/main/docs/algorithms/rejection_sampling.md for more detail Configs args: {'add_timestamp': True, 'hf_entity': 'VGraf', 'hf_repo_id': 'synthetic_preference_dataset_multi', 'hf_repo_id_scores': 'synthetic_preference_dataset_multi_scores', 'input_filename': '/weka/oe-adapt-default/victoriag/synth_data/completions.jsonl', 'max_parallel_requests': 100, 'model':… See the full description on the dataset page: https://huggingface.co/datasets/VGraf/synthetic_preference_dataset_multi_1741068828.textn<1K0 likes8 downloads2y agoHugging Face26VGraf /synthetic_preference_dataset_multi_1746753469 allenai/open_instruct: Rejection Sampling Dataset See https://github.com/allenai/open-instruct/blob/main/docs/algorithms/rejection_sampling.md for more detail Configs args: {'add_timestamp': True, 'hf_entity': 'VGraf', 'hf_repo_id': 'synthetic_preference_dataset_multi', 'hf_repo_id_scores': 'synthetic_preference_dataset_multi_scores', 'input_filename':… See the full description on the dataset page: https://huggingface.co/datasets/VGraf/synthetic_preference_dataset_multi_1746753469.textn<1K0 likes8 downloads1y agoHugging Face27won-bae /bpo_preference_hh_datatext10K<n<100K0 likes8 downloads1y agoHugging Face28Acamal1 /Prompt_Preference_DatasetA preference dataset for end2end prompt optimization. Check our usage here. text10K<n<100K0 likes7 downloads10mo agoHugging Face29hhhappyshow /assignment4-preference-datasettextn<1K0 likes7 downloads6mo agoHugging Face30SuperSteel /assignment4_preference_dataset assignment4_preference_dataset This dataset contains pairwise preference data for Assignment 4. Files assignment4_preference_pairs.jsonl: Main preference dataset in JSONL format. assignment4_preference_pairs.csv: CSV version for quick inspection. Schema (JSONL) Each line stores one preference sample with: instruction/prompt text chosen response rejected response optional metadata fields Usage Use this dataset for reward modeling, preference… See the full description on the dataset page: https://huggingface.co/datasets/SuperSteel/assignment4_preference_dataset.tabularn<1K0 likes7 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.