Team Ai
20 results

Dutch

BramVanroy /wikipedia_culturax_dutch Filtered CulturaX + Wikipedia for Dutch This is a combined and filtered version of CulturaX and Wikipedia, only including Dutch. It is intended for the training of LLMs. Different configs are available based on the number of tokens (see a section below with an overview). This can be useful if you want to know exactly how many tokens you have. Great for using as a streaming dataset, too. Tokens are counted as white-space tokens, so depending on your tokenizer, you'll likely end up… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/wikipedia_culturax_dutch.texttext-generation1B<n<10B6 likes5.8k downloads2y agoHugging Facessmits /fineweb-2-dutchtabulartext-generation10M<n<100M3 likes2.9k downloads2y agoHugging FacevGassen /Dutch-Basisbestandwetten-Legislation-Laws-XML-Cleantext10K<n<100K0 likes2.7k downloads1y agoHugging Facessmits /tokenized-falcon2-dutch-20481M<n<10M0 likes2.3k downloads2y agoHugging Facessmits /processed-falcon-dutch-dataset100K<n<1M0 likes2k downloads2y agoHugging Facedanish-foundation-models /dutch-dynaword 🧨 Dutch Dynaword Version 1.0.1 (Changelog) Language nld, Nederlands, Dutch License Openly Licensed, See the respective dataset Models For model trained used this data see danish-foundation-models Contact If you have question about this project please create an issue here Dataset Description Number of samples: 14.45M Number of tokens (Llama 3): 37.89B Average document length in tokens (min, max): 2.62K (2, 5.45M) Dataset… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/dutch-dynaword.imagetext-generation10M<n<100M3 likes1k downloads1mo agoHugging Face