Team Ai
4 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01swiss-ai /Apertus-v1.5-Preference-Data Apertus 1.5 Preference Dataset This is the preference dataset used for the offline DPO stage of Apertus v1.5 alignment training, applied to the 70B model. The prompts come from Ai2's Olmo 3 Dolci-Instruct-DPO dataset. We only reuse the prompts from Dolci-Instruct-DPO; all chosen / rejected responses in this dataset were generated by us. How this dataset was built Prompts. Taken from Dolci-Instruct-DPO (ODC-BY). Response generation and annotation. Every prompt was… See the full description on the dataset page: https://huggingface.co/datasets/swiss-ai/Apertus-v1.5-Preference-Data.tabulartext-generation100K<n<1M4 likes277 downloads2mo agoHugging Face02MinKeonKim /PRO-STEP-Preference-Data PRO-STEP: DPO Preference Pairs Step-level preference pairs used to train the PRO-STEP policy model via Direct Preference Optimization. Paper: PRO-STEP: Step-level Process Reward Optimization for Retrieval-Augmented GenerationCode: GitHub Repository Pairs: 15,877 (after outcome filter) Source questions: 5,000 from HotpotQA + MuSiQue + 2WikiMultiHopQA training splits Generation: PRM-guided MCTS (K=3 branching, depth 7, 64 rollouts/question, V(s) = Q̄(s) + α · r̂(s) with α=0.3)… See the full description on the dataset page: https://huggingface.co/datasets/MinKeonKim/PRO-STEP-Preference-Data.tabulartext-generation10K<n<100K0 likes115 downloads1mo agoHugging Face03stochastic-parrots /MNLP_M1_Preference_dpo_dataset M1 Preference Data for DPO Dataset Description This dataset contains processed M1 preference data for DPO training. Created by: CS-552 Stochastic Parrots Team Date: May 24, 2025 Version: 1.0 Number of examples: 17615 Dataset Source This dataset is derived from the M1 preference data collected through interactions with large language models (like ChatGPT) for CS-552 (Modern Natural Language Processing) at EPFL. The preference data consists of… See the full description on the dataset page: https://huggingface.co/datasets/stochastic-parrots/MNLP_M1_Preference_dpo_dataset.tabulartext-generation10K<n<100K0 likes33 downloads1y agoHugging Face04ITBill /INFH-6000Q-dpo-preference-dataset INFH-6000Q DPO Preference Dataset This dataset contains the final preference pairs used for the Direct Preference Optimization assignment in this repository. Source Base instruction source: GAIR/lima Candidate generator: local Qwen/Qwen2.5-7B-Instruct Preference ranker: local llm-blender/PairRM Construction Pipeline Sample 50 instructions from the local LIMA training split with seed 42. Generate 5 candidate responses per instruction with Qwen2.5-7B-Instruct.… See the full description on the dataset page: https://huggingface.co/datasets/ITBill/INFH-6000Q-dpo-preference-dataset.tabulartext-generationn<1K0 likes5 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.