Osye/openhermes2.5-Perplexity_filtered_top30
OpenHermes 2.5 - Perlexity Filtered (Top 30%) A filtered subset of OpenHermes 2.5 containing the top 30% highest perplexity samples scored by Qwen2.5-3B-Instruct Dataset Summary Source teknium/OpenHermes-2.5 Size 300466 samples (from ~1M original) Filter method Cross-entropy loss scored by Qwen2.5-3B-Instruct (4-bit NF4) Kept samples above the 70th percentile loss threshold Why this dataset? High perplexity samples are the examples a model finds… See the full description on the dataset page: https://huggingface.co/datasets/Osye/openhermes2.5-Perplexity_filtered_top30.
OpenHermes 2.5 - Perlexity Filtered (Top 30%)
A filtered subset of OpenHermes 2.5 containing the top 30% highest perplexity samples scored by Qwen2.5-3B-Instruct
Dataset Summary
- Source teknium/OpenHermes-2.5
- Size 300466 samples (from ~1M original)
- Filter method Cross-entropy loss scored by Qwen2.5-3B-Instruct (4-bit NF4)
- Kept samples above the 70th percentile loss threshold
Why this dataset?
High perplexity samples are the examples a model finds most surprising, they tend to be more diverse, complex, and informative for finetuning compared to "easy" repetitive samples. This filtered subset is useful for efficient SFT.
License
Inherits the license of OpenHermes 2.5 - Apache 2.0.
github training code: https://github.com/Malum0x/selective-qlora
