mrinaldi/TestiMole
Dataset Card for TestiMole -- A multi-billion tokens Italian text corpus Testimole is a large linguistic resource for Italian obtained through a massive web scraping effort. As of June 2024, it is one of the largest datasets for the Italian language, if not the largest, publicly available, consisting of almost 100B tokens counted with the Tiktoken cl100k BPE tokenizer. It consists mainly of conversational data (Italian Usenet hierarchies, Italian message boards, Italian… See the full description on the dataset page: https://huggingface.co/datasets/mrinaldi/TestiMole.
1260
No card is published for this repository, or it could not be fetched from Hugging Face right now.
