Team Ai
Datasetpublic

HuggingFaceFW/fineweb-2

šŸ„‚ FineWeb2 A sparkling update with 1000s of languages What is it? This is the second iteration of the popular šŸ· FineWeb dataset, bringing high quality pretraining data to over 1000 šŸ—£ļø languages. The šŸ„‚ FineWeb2 dataset is fully reproducible, available under the permissive ODC-By 1.0 license and extensively validated through hundreds of ablation experiments. In particular, on the set of 9 diverse languages we used to guide our processing decisions, šŸ„‚ā€¦ See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/fineweb-2.

sourceHugging Faceodc-byupdated 1y agoView on Hugging Face
898likes114kdownloads

HuggingFaceFW/fineweb-2 Ā· main Ā· files are served by the source, never re-hosted here