Team Ai
Datasetpublic

sail/regmix-data-sample

RegMix Data Sample Dataset Description The RegMix Data Sample is a curated dataset derived from the Pile-Uncopyrighted, specifically designed for the RegMix paper (https://huggingface.co/papers/2407.01492). This dataset aims to facilitate the automatic identification of high-performing data mixtures for language model pre-training by formulating it as a regression task. Key Features: Size: Approximately 20GB disk space, 5B tokens Distribution:… See the full description on the dataset page: https://huggingface.co/datasets/sail/regmix-data-sample.

sourceHugging Facemitupdated 2y agoView on Hugging Face
2likes764downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

Team Ai shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face