HuggingFaceFW/fineweb
๐ท FineWeb 15 trillion tokens of the finest data the ๐ web has to offer What is it? The ๐ท FineWeb dataset consists of more than 18.5T tokens (originally 15T tokens) of cleaned and deduplicated english web data from CommonCrawl. The data processing pipeline is optimized for LLM performance and ran on the ๐ญ datatrove library, our large scale data processing library. ๐ท FineWeb was originally meant to be a fully open replication of ๐ฆ RefinedWeb, with aโฆ See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/fineweb.
3.6k314k
../
000_00000.parquetdownload
001_00000.parquetdownload
002_00000.parquetdownload
003_00000.parquetdownload
004_00000.parquetdownload
005_00000.parquetdownload
006_00000.parquetdownload
007_00000.parquetdownload
008_00000.parquetdownload
009_00000.parquetdownload
010_00000.parquetdownload
011_00000.parquetdownload
012_00000.parquetdownload
013_00000.parquetdownload
014_00000.parquetdownload
