Team Ai
Datasetpublic

HuggingFaceFW/fineweb

๐Ÿท FineWeb 15 trillion tokens of the finest data the ๐ŸŒ web has to offer What is it? The ๐Ÿท FineWeb dataset consists of more than 18.5T tokens (originally 15T tokens) of cleaned and deduplicated english web data from CommonCrawl. The data processing pipeline is optimized for LLM performance and ran on the ๐Ÿญ datatrove library, our large scale data processing library. ๐Ÿท FineWeb was originally meant to be a fully open replication of ๐Ÿฆ… RefinedWeb, with aโ€ฆ See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/fineweb.

sourceHugging Faceodc-byupdated 1y agoView on Hugging Face
3.6klikes314kdownloads
../
file000_00000.parquet2.00 GBdownload
file001_00000.parquet2.00 GBdownload
file002_00000.parquet2.00 GBdownload
file003_00000.parquet2.00 GBdownload
file004_00000.parquet2.00 GBdownload
file005_00000.parquet2.00 GBdownload
file006_00000.parquet2.00 GBdownload
file007_00000.parquet2.00 GBdownload
file008_00000.parquet2.00 GBdownload
file009_00000.parquet2.00 GBdownload
file010_00000.parquet2.00 GBdownload
file011_00000.parquet2.00 GBdownload
file012_00000.parquet2.00 GBdownload
file013_00000.parquet2.00 GBdownload
file014_00000.parquet548.3 MBdownload

HuggingFaceFW/fineweb ยท main ยท files are served by the source, never re-hosted here