agentlans/multilingual-text
Multilingual Text Dataset This dataset contains a curated selection of rows from multiple input datasets, where each row includes a text chunk of approximately 2000 tokens (as measured by Llama 3.1 tokenizer) verified to be written in the correct language. Only rows with properly classified language chunks are retained, ensuring high-quality multilingual data for analysis or model training. Preprocessing Steps Normalized whitespace, punctuation, Unicode… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/multilingual-text.
Update README.md
Update README.md
Update README.md
Upload 10 files
Update README.md
Update README.md
Upload 6 files
Update README.md
Update README.md
Upload 6 files
Update README.md
Upload 4 files
Update README.md
Upload 10 files
Update README.md
initial commit
