Team Ai
Datasetpublic

timaeus/pile-dm_mathematics

Dataset Creation Process These subsets were created by streaming over the rows from monology/pile-uncopyrighted and filtering by the meta column. Each subset is generally limited to the first 100,000 qualifying rows encountered. Citations If you use this dataset, please cite the original Pile papers: @article{gao2020pile, title={The Pile: An 800GB dataset of diverse text for language modeling}, author={Gao, Leo and Biderman, Stella and Black, Sid and Golding… See the full description on the dataset page: https://huggingface.co/datasets/timaeus/pile-dm_mathematics.

sourceHugging Faceupdated 2y agoView on Hugging Face
1likes112downloads
3 commits on main
17347912y ago

Update README.md

jqhoogland
bd87a382y ago

Upload dataset

jqhoogland
3f555d42y ago

initial commit

jqhoogland