Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01allenai /tulu-2.5-preference-data Tulu 2.5 Preference Data This dataset contains the preference dataset splits used to train the models described in Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference Feedback. We cleaned and formatted all datasets to be in the same format. This means some splits may differ from their original format. To see the code used for creating most splits, see here. If you only wish to download one dataset, each dataset exists in one file under the data/… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tulu-2.5-preference-data.texttext-generation1M<n<10M18 likes1.2k downloads2y agoHugging Face02Rapidata /700k_Human_Preference_Dataset_FLUX_SD3_MJ_DALLE3 NOTE: A newer version of this dataset is available Imagen3_Flux1.1_Flux1_SD3_MJ_Dalle_Human_Preference_Dataset Rapidata Image Generation Preference Dataset This Dataset is a 1/3 of a 2M+ human annotation dataset that was split into three modalities: Preference, Coherence, Text-to-Image Alignment. Link to the Coherence dataset: https://huggingface.co/datasets/Rapidata/Flux_SD3_MJ_Dalle_Human_Coherence_Dataset Link to the Text-2-Image Alignment dataset:… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/700k_Human_Preference_Dataset_FLUX_SD3_MJ_DALLE3.imagetext-to-image10K<n<100K20 likes912 downloads2y agoHugging Face03KORMo-Team /preference-dataset-qwen30 likes877 downloads1y agoHugging Face04allenai /preference-datasets-tulutext1M<n<10M8 likes711 downloads3y agoHugging Face05WPRM /preference_data_llama_factory_wo_checklist Dataset Card for "preference_data_llama_factory_wo_checklist" More Information needed image10K<n<100K0 likes432 downloads1y agoHugging Face06swiss-ai /Apertus-v1.5-Preference-Data Apertus 1.5 Preference Dataset This is the preference dataset used for the offline DPO stage of Apertus v1.5 alignment training, applied to the 70B model. The prompts come from Ai2's Olmo 3 Dolci-Instruct-DPO dataset. We only reuse the prompts from Dolci-Instruct-DPO; all chosen / rejected responses in this dataset were generated by us. How this dataset was built Prompts. Taken from Dolci-Instruct-DPO (ODC-BY). Response generation and annotation. Every prompt was… See the full description on the dataset page: https://huggingface.co/datasets/swiss-ai/Apertus-v1.5-Preference-Data.tabulartext-generation100K<n<1M4 likes271 downloads1mo agoHugging Face07OpenRLHF /preference_dataset_mixture2_and_safe_pku Copy from https://huggingface.co/datasets/weqweasdas/preference_dataset_mixture2_and_safe_pku Reward Model Overview This is the data mixture used for the reward model weqweasdas/RM-Mistral-7B, trained with the script https://github.com/WeiXiongUST/RLHF-Reward-Modeling . Also see a short blog for the training details (data mixture, parameters...): https://www.notion.so/Reward-Modeling-for-RLHF-abe03f9afdac42b9a5bee746844518d0 Model Details If you have any question… See the full description on the dataset page: https://huggingface.co/datasets/OpenRLHF/preference_dataset_mixture2_and_safe_pku.tabular100K<n<1M8 likes235 downloads2y agoHugging Face08jessierenjie /pair_preference_model_dataset_add_emoji_to_win_rate0.1_rrm_newtext1M<n<10M0 likes227 downloads2y agoHugging Face09TianqiLiuAI /pair_preference_model_dataset_add_prefix_to_win_rate0.1_rrm_0p2text1M<n<10M0 likes194 downloads2y agoHugging Face10TianqiLiuAI /pair_preference_model_dataset_gemma2_2b_rrm_0p2text1M<n<10M0 likes174 downloads2y agoHugging Face11WPRM /preference_data_llama_factory_corrected_format_text_onlyimage10K<n<100K0 likes164 downloads1y agoHugging Face12TianqiLiuAI /pair_preference_model_dataset_add_prefix_to_win_rate0.1_rrmtext1M<n<10M0 likes152 downloads2y agoHugging Face13WPRM /preference_data_llama_factory_corrected_formatimage10K<n<100K0 likes148 downloads1y agoHugging Face14cornfieldrm /pair-preference-dataset-700K_standardtabular100K<n<1M0 likes138 downloads2y agoHugging Face15davidberenstein1957 /dataset-viber-image-generation-preference-inference-endpoints-battle-flux Dataset Card for Dataset Name Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More Information Needed] Paper [optional]: [More Information Needed] Demo [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/davidberenstein1957/dataset-viber-image-generation-preference-inference-endpoints-battle-flux.imagen<1K0 likes134 downloads2y agoHugging Face16cornfieldrm /preference_dataset-standard_format-v2.2text100K<n<1M0 likes131 downloads2y agoHugging Face17cornfieldrm /pair-preference-dataset-700K_subset-15-out-of-16_standardtabular100K<n<1M0 likes120 downloads2y agoHugging Face18weqweasdas /preference_dataset_mixture2_and_safe_pku Reward Model Overview This is the data mixture used for the reward model weqweasdas/RM-Mistral-7B, trained with the script https://github.com/WeiXiongUST/RLHF-Reward-Modeling . Also see a short blog for the training details (data mixture, parameters...): https://www.notion.so/Reward-Modeling-for-RLHF-abe03f9afdac42b9a5bee746844518d0 Model Details If you have any question with this reward model and also any question about reward modeling, feel free to drop me an… See the full description on the dataset page: https://huggingface.co/datasets/weqweasdas/preference_dataset_mixture2_and_safe_pku.tabular100K<n<1M12 likes119 downloads2y agoHugging Face19ProVoice-proactivity /proactivity_preference_dataset ProVoice study 1 — driver state, vehicle context and preferred Level of Autonomy Driving-simulator data from the population data collection of the ProVoice / ProActivity project (CARLA 0.10): 12 drivers × 2 sessions, ~20 Hz multimodal driver-state and vehicle frames, and 1,446 driver-assigned Level-of-Autonomy (LoA) labels stating how autonomously an in-vehicle assistant should act on a given task. Drivers were prompted every 20 s about two randomly drawn in-vehicle tasks and… See the full description on the dataset page: https://huggingface.co/datasets/ProVoice-proactivity/proactivity_preference_dataset.tabular1M<n<10M0 likes118 downloads21d agoHugging Face20MinKeonKim /PRO-STEP-Preference-Data PRO-STEP: DPO Preference Pairs Step-level preference pairs used to train the PRO-STEP policy model via Direct Preference Optimization. Paper: PRO-STEP: Step-level Process Reward Optimization for Retrieval-Augmented GenerationCode: GitHub Repository Pairs: 15,877 (after outcome filter) Source questions: 5,000 from HotpotQA + MuSiQue + 2WikiMultiHopQA training splits Generation: PRM-guided MCTS (K=3 branching, depth 7, 64 rollouts/question, V(s) = Q̄(s) + α · r̂(s) with α=0.3)… See the full description on the dataset page: https://huggingface.co/datasets/MinKeonKim/PRO-STEP-Preference-Data.tabulartext-generation10K<n<100K0 likes115 downloads1mo agoHugging Face21sam-paech /gemma-3-27b-it-antislop-ftpo-preference-datasettext10K<n<100K0 likes109 downloads1y agoHugging Face22yufan /Preference_Dataset_Merged Dataset Overview This Dataset consists of the following open-sourced preference dataset Arena Human Preference Anthropic HH MT-Bench Human Judgement Ultra Feedback Tulu3 Preference Dataset Skywork-Reward-Preference-80K-v0.2 Cleaning Cleaning Method 1: Only keep the following Language using FastText language detection(EN/DE/ES/ZH/IT/JA/FR) Cleaning Method 2: Remove duplicates to ensure each prompt appears only once Cleaning Method 3: Remove datasets where… See the full description on the dataset page: https://huggingface.co/datasets/yufan/Preference_Dataset_Merged.text100K<n<1M0 likes104 downloads2y agoHugging Face23ringos /tulu-2_5-preference-data-sharegpttext1M<n<10M1 likes101 downloads2y agoHugging Face24sdiazlor /math-preference-dataset Dataset Card for math-preference-dataset This dataset has been created with distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/sdiazlor/math-preference-dataset/raw/main/pipeline.yaml" or explore the configuration: distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/sdiazlor/math-preference-dataset.tabularn<1K0 likes99 downloads2y agoHugging Face25Aratako /Synthetic-JP-Preference-Dataset-Qwen2.5_72B-191k Synthetic-JP-Preference-Dataset-Qwen2.5_72B-191k 概要 5種類のオープンモデルとQwen/Qwen2.5-72B-Instruct-GPTQ-Int8を使って作成した、190854件の日本語合成Preferenceデータセットです。 以下、データセットの詳細です。 instructionには、Aratako/Magpie-Tanuki-8B-annotated-96kのinput_qualityがexcellentのものを利用 回答生成には、以下の5つのApache 2.0ライセンスのモデルを利用 weblab-GENIAC/Tanuki-8B-dpo-v1.0 team-hatakeyama-phase2/Tanuki-8x8B-dpo-v1.0-GPTQ-8bit cyberagent/calm3-22b-chat llm-jp/llm-jp-3-13b-instruct Qwen/Qwen2.5-32B-Instruct-GPTQ-Int8… See the full description on the dataset page: https://huggingface.co/datasets/Aratako/Synthetic-JP-Preference-Dataset-Qwen2.5_72B-191k.texttext-generation100K<n<1M6 likes93 downloads2y agoHugging Face26alexnik /1.4b-policy_preference_data_armorm_gold_labelledtext10K<n<100K0 likes92 downloads2y agoHugging Face27weqweasdas /preference_dataset_mix2 Dataset Card for "preference_dataset_mix2" More Information needed tabular100K<n<1M3 likes87 downloads3y agoHugging Face28prhegde /preference-data-math-stack-exchangeThe preference dataset is derived from the stack exchange dataset which contains questions and answers from the Stack Overflow Data Dump. This contains questions and answers for various topics. For this work, we used only question and answers from math.stackexchange.com sub-folder. The questions are grouped with answers that are assigned a score corresponding to the Anthropic paper: score = log2 (1 + upvotes) rounded to the nearest integer, plus 1 if the answer was accepted by the questioner… See the full description on the dataset page: https://huggingface.co/datasets/prhegde/preference-data-math-stack-exchange.text10K<n<100K6 likes83 downloads3y agoHugging Face29yflantmy /universal-preference-hijacking-datasets Phi: Preference Hijacking in Multi-modal Large Language Models at Inference Time Figure 1: Examples of Phi, which can hijack MLLM's preference toward the image. Figure 2: Example of a universal hijacking perturbation, which can be transferred across different images. This dataset is used to train and evaluate the universal hijacking perturbations in the paper "Phi: Preference Hijacking in Multi-modal Large Language Models at Inference Time", accepted at EMNLP… See the full description on the dataset page: https://huggingface.co/datasets/yflantmy/universal-preference-hijacking-datasets.imagequestion-answering1K<n<10K0 likes82 downloads1y agoHugging Face30TaylorAI /RLCD-generated-preference-data-split Dataset Card for "RLCD-generated-preference-data-split" More Information needed tabular100K<n<1M0 likes75 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.