Team Ai
Datasetpublic

LemonTea03/BinaryOPD-Data

Do We Really Need KL Divergence for On-Policy Distillation of Large Language Models? This dataset supports the paper Do We Really Need KL Divergence for On-Policy Distillation of Large Language Models?. It contains training and evaluation data for on-policy distillation (OPD) experiments, including BinaryOPD, positive/negative reward experiments, and consensus multi-teacher on-policy distillation (C-MOPD). Code and additional details are available in the GitHub repository.… See the full description on the dataset page: https://huggingface.co/datasets/LemonTea03/BinaryOPD-Data.

sourceHugging Faceupdated 5d agoView on Hugging Face
0likes198downloads
Dataset Card

Do We Really Need KL Divergence for On-Policy Distillation of Large Language Models?

This dataset supports the paper Do We Really Need KL Divergence for On-Policy Distillation of Large Language Models?. It contains training and evaluation data for on-policy distillation (OPD) experiments, including BinaryOPD, positive/negative reward experiments, and consensus multi-teacher on-policy distillation (C-MOPD).

Code and additional details are available in the GitHub repository.

Dataset Structure

SplitFileContent
train/DAPO-Math-17k.parquetDAPO math training set
train/DeepMath-103K-filtered-level6.parquetDeepMath, difficulty ≥ 6
train/Eurus-RL-Data-code.parquetEurus code training set
train/math_and_code.parquetBalanced DeepMath + Eurus mixture for MOPD/C-MOPD
eval/math_eval.parquetAIME24/25, AMC, MATH500, Minerva, OlympiadBench, IMO-Bench
eval/livecodebench_v6.parquetLiveCodeBench v6 evaluation slice
eval/HumanEval.parquetHumanEval
eval/MBPP.parquetMBPP

Usage

Please refer to the GitHub README for training scripts, reward variants, multi-teacher distillation experiments, and evaluation instructions.