Team Ai
Datasetpublic

agentlans/multilingual-text

Multilingual Text Dataset This dataset contains a curated selection of rows from multiple input datasets, where each row includes a text chunk of approximately 2000 tokens (as measured by Llama 3.1 tokenizer) verified to be written in the correct language. Only rows with properly classified language chunks are retained, ensuring high-quality multilingual data for analysis or model training. Preprocessing Steps Normalized whitespace, punctuation, Unicode… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/multilingual-text.

sourceHugging Faceodc-byupdated 1y agoView on Hugging Face
5likes354downloads
settings

This repository belongs to agentlans on Hugging Face.

Team Ai never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namemultilingual-text
visibilitypublic
licenceodc-by
gatedno
owneragentlans
Account settings
agentlans/multilingual-text · Team Ai