Team Ai
Datasetpublicgated

mrinaldi/TestiMole

Dataset Card for TestiMole -- A multi-billion tokens Italian text corpus Testimole is a large linguistic resource for Italian obtained through a massive web scraping effort. As of June 2024, it is one of the largest datasets for the Italian language, if not the largest, publicly available, consisting of almost 100B tokens counted with the Tiktoken cl100k BPE tokenizer. It consists mainly of conversational data (Italian Usenet hierarchies, Italian message boards, Italian… See the full description on the dataset page: https://huggingface.co/datasets/mrinaldi/TestiMole.

sourceHugging Faceupdated 4mo agoView on Hugging Face
12likes60downloads
Dataset Card

No card is published for this repository, or it could not be fetched from Hugging Face right now.