datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PPOpt-data
PersonaAtlas Dataset
Dataset Summary
PersonaAtlas is a synthetic multi-turn conversational dataset designed for persona-aware alignment of large language models. It contains 10,462 conversation samples derived from 2,055 unique personas, enabling research on personalized response generation and preference modeling.
Each example includes:
A structured persona profile (persona) with user preference features
The source prompt (original_query)
The full dialog… See the full description on the dataset page: https://huggingface.co/datasets/HowieHwong/PPOpt-data.big-math-ppo-mix-40k-qwq-solutions
Big Math PPO Mix 40k QwQ Solutions
Teacher solutions generated by Qwen/QwQ-32B-Preview for the Big Math PPO Mix
prompts, intended for knowledge distillation into a student policy.
This is the combined dataset, merging:
big-math-ppo-mix-30k prompts → 22,584 correct solutions (75.3% of 30,000)
big-math-ppo-mix-extra-10k prompts → 7,986 correct solutions (79.9% of 10,000)
Total: 30,570 verified-correct solutions (40,000 prompts attempted).
Generation
vLLM sampling… See the full description on the dataset page: https://huggingface.co/datasets/christinakopi/big-math-ppo-mix-40k-qwq-solutions.oasst2-ru-ppo
OASST-RU-PPO Dataset
Description
The oasst-ru-ppo dataset is designed for optimizing language models using Proximal Policy Optimization (PPO). It is specifically tailored for Russian language models and is created from a collection of dialogues with associated rewards.
Dataset Creation
The dataset is created from the original oasst2 dataset, which contains a series of dialogs. Each dialog is a sequence of responses, where each response is a text… See the full description on the dataset page: https://huggingface.co/datasets/0x7o/oasst2-ru-ppo.wbs-llm-ppo-demo
widebluesky/wbs-llm-ppo-demo
Independent synthetic prompt/reference dataset used to train the PPO demo model.
This repository contains only this task's data and can be loaded independently.
Contents and split policy
Eight English fixtures: train 4, validation 2, test 2. Fields are id, group_id,
split, messages (role/content pairs ending in a user turn), and reference.
The reference is used to compute reward and evaluation; it is never appended to policy inputs.… See the full description on the dataset page: https://huggingface.co/datasets/widebluesky/wbs-llm-ppo-demo.
