HuggingFaceFW/fineweb-2
š„ FineWeb2 A sparkling update with 1000s of languages What is it? This is the second iteration of the popular š· FineWeb dataset, bringing high quality pretraining data to over 1000 š£ļø languages. The š„ FineWeb2 dataset is fully reproducible, available under the permissive ODC-By 1.0 license and extensively validated through hundreds of ablation experiments. In particular, on the set of 9 diverse languages we used to guide our processing decisions, š„⦠See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/fineweb-2.
898114k
