Team Ai
4 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01HowieHwong /PPOpt-data PersonaAtlas Dataset Dataset Summary PersonaAtlas is a synthetic multi-turn conversational dataset designed for persona-aware alignment of large language models. It contains 10,462 conversation samples derived from 2,055 unique personas, enabling research on personalized response generation and preference modeling. Each example includes: A structured persona profile (persona) with user preference features The source prompt (original_query) The full dialog… See the full description on the dataset page: https://huggingface.co/datasets/HowieHwong/PPOpt-data.texttext-generation10K<n<100K2 likes60 downloads8mo agoHugging Face02christinakopi /big-math-ppo-mix-40k-qwq-solutions Big Math PPO Mix 40k QwQ Solutions Teacher solutions generated by Qwen/QwQ-32B-Preview for the Big Math PPO Mix prompts, intended for knowledge distillation into a student policy. This is the combined dataset, merging: big-math-ppo-mix-30k prompts → 22,584 correct solutions (75.3% of 30,000) big-math-ppo-mix-extra-10k prompts → 7,986 correct solutions (79.9% of 10,000) Total: 30,570 verified-correct solutions (40,000 prompts attempted). Generation vLLM sampling… See the full description on the dataset page: https://huggingface.co/datasets/christinakopi/big-math-ppo-mix-40k-qwq-solutions.tabulartext-generation10K<n<100K0 likes21 downloads4mo agoHugging Face030x7o /oasst2-ru-ppo OASST-RU-PPO Dataset Description The oasst-ru-ppo dataset is designed for optimizing language models using Proximal Policy Optimization (PPO). It is specifically tailored for Russian language models and is created from a collection of dialogues with associated rewards. Dataset Creation The dataset is created from the original oasst2 dataset, which contains a series of dialogs. Each dialog is a sequence of responses, where each response is a text… See the full description on the dataset page: https://huggingface.co/datasets/0x7o/oasst2-ru-ppo.texttext-generation1K<n<10K3 likes12 downloads3y agoHugging Face04widebluesky /wbs-llm-ppo-demo widebluesky/wbs-llm-ppo-demo Independent synthetic prompt/reference dataset used to train the PPO demo model. This repository contains only this task's data and can be loaded independently. Contents and split policy Eight English fixtures: train 4, validation 2, test 2. Fields are id, group_id, split, messages (role/content pairs ending in a user turn), and reference. The reference is used to compute reward and evaluation; it is never appended to policy inputs.… See the full description on the dataset page: https://huggingface.co/datasets/widebluesky/wbs-llm-ppo-demo.texttext-generationn<1K0 likes1 downloads2d agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.