Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01swiss-ai /Apertus-v1.5-Preference-Data Apertus 1.5 Preference Dataset This is the preference dataset used for the offline DPO stage of Apertus v1.5 alignment training, applied to the 70B model. The prompts come from Ai2's Olmo 3 Dolci-Instruct-DPO dataset. We only reuse the prompts from Dolci-Instruct-DPO; all chosen / rejected responses in this dataset were generated by us. How this dataset was built Prompts. Taken from Dolci-Instruct-DPO (ODC-BY). Response generation and annotation. Every prompt was… See the full description on the dataset page: https://huggingface.co/datasets/swiss-ai/Apertus-v1.5-Preference-Data.tabulartext-generation100K<n<1M4 likes277 downloads2mo agoHugging Face02OpenRLHF /preference_dataset_mixture2_and_safe_pku Copy from https://huggingface.co/datasets/weqweasdas/preference_dataset_mixture2_and_safe_pku Reward Model Overview This is the data mixture used for the reward model weqweasdas/RM-Mistral-7B, trained with the script https://github.com/WeiXiongUST/RLHF-Reward-Modeling . Also see a short blog for the training details (data mixture, parameters...): https://www.notion.so/Reward-Modeling-for-RLHF-abe03f9afdac42b9a5bee746844518d0 Model Details If you have any question… See the full description on the dataset page: https://huggingface.co/datasets/OpenRLHF/preference_dataset_mixture2_and_safe_pku.tabular100K<n<1M8 likes219 downloads2y agoHugging Face03cornfieldrm /pair-preference-dataset-700K_standardtabular100K<n<1M0 likes124 downloads2y agoHugging Face04ProVoice-proactivity /proactivity_preference_dataset ProVoice study 1 — driver state, vehicle context and preferred Level of Autonomy Driving-simulator data from the population data collection of the ProVoice / ProActivity project (CARLA 0.10): 12 drivers × 2 sessions, ~20 Hz multimodal driver-state and vehicle frames, and 1,446 driver-assigned Level-of-Autonomy (LoA) labels stating how autonomously an in-vehicle assistant should act on a given task. Drivers were prompted every 20 s about two randomly drawn in-vehicle tasks and… See the full description on the dataset page: https://huggingface.co/datasets/ProVoice-proactivity/proactivity_preference_dataset.tabular1M<n<10M0 likes124 downloads25d agoHugging Face05MinKeonKim /PRO-STEP-Preference-Data PRO-STEP: DPO Preference Pairs Step-level preference pairs used to train the PRO-STEP policy model via Direct Preference Optimization. Paper: PRO-STEP: Step-level Process Reward Optimization for Retrieval-Augmented GenerationCode: GitHub Repository Pairs: 15,877 (after outcome filter) Source questions: 5,000 from HotpotQA + MuSiQue + 2WikiMultiHopQA training splits Generation: PRM-guided MCTS (K=3 branching, depth 7, 64 rollouts/question, V(s) = Q̄(s) + α · r̂(s) with α=0.3)… See the full description on the dataset page: https://huggingface.co/datasets/MinKeonKim/PRO-STEP-Preference-Data.tabulartext-generation10K<n<100K0 likes115 downloads1mo agoHugging Face06weqweasdas /preference_dataset_mixture2_and_safe_pku Reward Model Overview This is the data mixture used for the reward model weqweasdas/RM-Mistral-7B, trained with the script https://github.com/WeiXiongUST/RLHF-Reward-Modeling . Also see a short blog for the training details (data mixture, parameters...): https://www.notion.so/Reward-Modeling-for-RLHF-abe03f9afdac42b9a5bee746844518d0 Model Details If you have any question with this reward model and also any question about reward modeling, feel free to drop me an… See the full description on the dataset page: https://huggingface.co/datasets/weqweasdas/preference_dataset_mixture2_and_safe_pku.tabular100K<n<1M12 likes104 downloads2y agoHugging Face07cornfieldrm /pair-preference-dataset-700K_subset-15-out-of-16_standardtabular100K<n<1M0 likes75 downloads2y agoHugging Face08weqweasdas /preference_dataset_mix2 Dataset Card for "preference_dataset_mix2" More Information needed tabular100K<n<1M3 likes74 downloads3y agoHugging Face09monology /stackexchange-preference-datatabular100K<n<1M0 likes68 downloads1y agoHugging Face10sdiazlor /math-preference-dataset Dataset Card for math-preference-dataset This dataset has been created with distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/sdiazlor/math-preference-dataset/raw/main/pipeline.yaml" or explore the configuration: distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/sdiazlor/math-preference-dataset.tabularn<1K0 likes65 downloads2y agoHugging Face11TaylorAI /RLCD-generated-preference-data-split Dataset Card for "RLCD-generated-preference-data-split" More Information needed tabular100K<n<1M0 likes61 downloads3y agoHugging Face12cornfieldrm /pair-preference-dataset-700K_subset-7-out-of-8_standardtabular100K<n<1M0 likes60 downloads2y agoHugging Face13WPRM /preference_data_sh_wo_checklisttabular10K<n<100K0 likes57 downloads1y agoHugging Face14cornfieldrm /pair-preference-dataset-700K_subset-9-out-of-10_standardtabular100K<n<1M0 likes54 downloads2y agoHugging Face15weqweasdas /preference_dataset_mixture2_and_safe_pku150k Dataset Card for "preference_dataset_mixture2_and_safe_pku150k" More Information needed tabular100K<n<1M0 likes51 downloads3y agoHugging Face16dayone3nder /1004_ti2t_preference_dataset_supple_30ktabular10K<n<100K0 likes50 downloads2y agoHugging Face17PrimeIntellect /SYNTHETIC-1-Preference-Data SYNTHETIC-1: Two Million Crowdsourced Reasoning Traces from Deepseek-R1 SYNTHETIC-1 is a reasoning dataset obtained from Deepseek-R1, generated with crowdsourced compute and annotated with diverse verifiers such as LLM judges or symbolic mathematics verifiers. This is the SFT version of the dataset - the raw data and SFT dataset can be found in our 🤗 SYNTHETIC-1 Collection. The dataset consists of the following tasks and verifiers that were implemented in our library genesys:… See the full description on the dataset page: https://huggingface.co/datasets/PrimeIntellect/SYNTHETIC-1-Preference-Data.tabular10K<n<100K6 likes47 downloads2y agoHugging Face18cornfieldrm /pair-preference-dataset-700K_subset-2-of-3_standardtabular100K<n<1M0 likes46 downloads2y agoHugging Face19Yuriy81 /example-preference-dataset Dataset Card for example-preference-dataset This dataset has been created with distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/Yuriy81/example-preference-dataset/raw/main/pipeline.yaml" or explore the configuration: distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/Yuriy81/example-preference-dataset.tabularn<1K0 likes46 downloads2y agoHugging Face20cornfieldrm /pair-preference-dataset-700K_subset-2-of-4_gemma-2b_1of4_iter3_conf-0.8_bs128_lr1e-5_conf-0.8tabular10K<n<100K0 likes45 downloads2y agoHugging Face21distilabel-internal-testing /example-generate-preference-dataset Dataset Card for example-preference-dataset This dataset has been created with distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/sdiazlor/example-preference-dataset/raw/main/pipeline.yaml" or explore the configuration: distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/distilabel-internal-testing/example-generate-preference-dataset.tabularn<1K0 likes44 downloads2y agoHugging Face22transitionGap /gst-india-preference-dataset-prep-smalltabularn<1K0 likes40 downloads2y agoHugging Face23preference-team /dataset-for-annotation-v2-annotated 合成された質問と2つの応答文のペアに対して、日本語を母語とするチームメンバーが、好ましい応答文を人手でアノテーションしました HelpSteer2-preferenceに習い、選好だけでなく選好の強度も[-3, 3]の範囲で付与しました アノテーション時のメモ ジャンルは以下 簡単な一般知識(wikipediaを読まずに回答できる系) 難しめの一般知識(wikipediaを読んだら回答できる系) 歴史上の出来事の論述 医療知識(応急処置系) 機械学習の課題と解決方法 化学式の解説 架空の物語生成 ロールプレイ 詩の創作 素因数分解や偶数奇数判定などの簡単な数学タスク コーディングタスク アルゴリズムやシステムのメリデメの解説 美術や思想についての論述 日本語の文法の解説 その他 LLMの定型文として登場するフレーズは、 「もちろんです」 「~も見逃せません」 「~も見過ごせません」 「総じて、」 「まず始めに、~さらに、~次に、~まとめると、」 データセットを目視で読み込んだ印象… See the full description on the dataset page: https://huggingface.co/datasets/preference-team/dataset-for-annotation-v2-annotated.tabular1K<n<10K3 likes40 downloads1y agoHugging Face24cornfieldrm /pair-preference-dataset-700K_subset-4-of-4_gemma-2b_1of4_iter1_conf-0.8_bs128_lr1e-5_conf-0.8tabular10K<n<100K0 likes39 downloads2y agoHugging Face25living-box /preference_dataset_mixture2_and_safe_pku Reward Model Overview This is the data mixture used for the reward model weqweasdas/RM-Mistral-7B, trained with the script https://github.com/WeiXiongUST/RLHF-Reward-Modeling . Also see a short blog for the training details (data mixture, parameters...): https://www.notion.so/Reward-Modeling-for-RLHF-abe03f9afdac42b9a5bee746844518d0 Model Details If you have any question with this reward model and also any question about reward modeling, feel free to drop me an… See the full description on the dataset page: https://huggingface.co/datasets/living-box/preference_dataset_mixture2_and_safe_pku.tabular100K<n<1M0 likes39 downloads9mo agoHugging Face26cornfieldrm /pair-preference-dataset-700K_subset-4-of-4_gemma-2b_1of4_iter1_bs128_lr1e-5_conf-0.9tabular10K<n<100K0 likes37 downloads2y agoHugging Face27dayone3nder /0930_ta2t_preference_dataset_20Ktabular10K<n<100K0 likes37 downloads2y agoHugging Face28cornfieldrm /pair-preference-dataset-700K_subset-3-of-4_gemma-2b_1of4_iter1_conf-0.8_bs128_lr1e-5_conf-0.8tabular10K<n<100K0 likes35 downloads2y agoHugging Face29cornfieldrm /pair-preference-dataset-700K_subset-2-of-2_standardtabular100K<n<1M0 likes34 downloads2y agoHugging Face30stochastic-parrots /MNLP_M1_Preference_dpo_dataset M1 Preference Data for DPO Dataset Description This dataset contains processed M1 preference data for DPO training. Created by: CS-552 Stochastic Parrots Team Date: May 24, 2025 Version: 1.0 Number of examples: 17615 Dataset Source This dataset is derived from the M1 preference data collected through interactions with large language models (like ChatGPT) for CS-552 (Modern Natural Language Processing) at EPFL. The preference data consists of… See the full description on the dataset page: https://huggingface.co/datasets/stochastic-parrots/MNLP_M1_Preference_dpo_dataset.tabulartext-generation10K<n<100K0 likes33 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.