Team Ai
Datasetpublic

timaeus/pile-dm_mathematics

Dataset Creation Process These subsets were created by streaming over the rows from monology/pile-uncopyrighted and filtering by the meta column. Each subset is generally limited to the first 100,000 qualifying rows encountered. Citations If you use this dataset, please cite the original Pile papers: @article{gao2020pile, title={The Pile: An 800GB dataset of diverse text for language modeling}, author={Gao, Leo and Biderman, Stella and Black, Sid and Golding… See the full description on the dataset page: https://huggingface.co/datasets/timaeus/pile-dm_mathematics.

sourceHugging Faceupdated 2y agoView on Hugging Face
1likes119downloads
settings

This repository belongs to timaeus on Hugging Face.

Team Ai never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namepile-dm_mathematics
visibilitypublic
licencenot set
gatedno
ownertimaeus
Account settings
timaeus/pile-dm_mathematics · Team Ai