Team Ai
Datasetpublic

HuggingFaceFW/finepdfs

Liberating 3T of the finest tokens from PDFs What is this? As we run out of web pages to process, the natural question has always been: what to do next? Only a few knew about a data source that everyone avoided for ages, due to its incredible extraction cost and complexity: PDFs. 📄 FinePDFs is exactly that. It is the largest publicly available corpus sourced exclusively from PDFs, containing about 3 trillion tokens across 475 million documents in 1733 languages. Compared to… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/finepdfs.

sourceHugging Faceodc-byupdated 6mo agoView on Hugging Face
946likes56kdownloads
3 commits on main
220bac36mo ago

Update README.md

thomwolf
89f54119mo ago

Update README.md

hynky
d8e854410mo ago

Super-squash branch 'main' using huggingface_hub

hynky