Team Ai
Datasetpublic

Osye/openhermes2.5-Perplexity_filtered_top30

OpenHermes 2.5 - Perlexity Filtered (Top 30%) A filtered subset of OpenHermes 2.5 containing the top 30% highest perplexity samples scored by Qwen2.5-3B-Instruct Dataset Summary Source teknium/OpenHermes-2.5 Size 300466 samples (from ~1M original) Filter method Cross-entropy loss scored by Qwen2.5-3B-Instruct (4-bit NF4) Kept samples above the 70th percentile loss threshold Why this dataset? High perplexity samples are the examples a model finds… See the full description on the dataset page: https://huggingface.co/datasets/Osye/openhermes2.5-Perplexity_filtered_top30.

sourceHugging Faceapache-2.0updated 7mo agoView on Hugging Face
1likes25downloads
Dataset Card

OpenHermes 2.5 - Perlexity Filtered (Top 30%)

A filtered subset of OpenHermes 2.5 containing the top 30% highest perplexity samples scored by Qwen2.5-3B-Instruct

Dataset Summary

  • —Source teknium/OpenHermes-2.5
  • —Size 300466 samples (from ~1M original)
  • —Filter method Cross-entropy loss scored by Qwen2.5-3B-Instruct (4-bit NF4)
  • —Kept samples above the 70th percentile loss threshold

Why this dataset?

High perplexity samples are the examples a model finds most surprising, they tend to be more diverse, complex, and informative for finetuning compared to "easy" repetitive samples. This filtered subset is useful for efficient SFT.

License

Inherits the license of OpenHermes 2.5 - Apache 2.0.

github training code: https://github.com/Malum0x/selective-qlora