Team Ai
Datasetpublic

timaeus/pile-dm_mathematics

Dataset Creation Process These subsets were created by streaming over the rows from monology/pile-uncopyrighted and filtering by the meta column. Each subset is generally limited to the first 100,000 qualifying rows encountered. Citations If you use this dataset, please cite the original Pile papers: @article{gao2020pile, title={The Pile: An 800GB dataset of diverse text for language modeling}, author={Gao, Leo and Biderman, Stella and Black, Sid and Golding… See the full description on the dataset page: https://huggingface.co/datasets/timaeus/pile-dm_mathematics.

sourceHugging Faceupdated 2y agoView on Hugging Face
1likes117downloads
../
filetrain-00000-of-00002.parquet197.8 MBdownload
filetrain-00001-of-00002.parquet197.9 MBdownload

timaeus/pile-dm_mathematics · main · files are served by the source, never re-hosted here