Team Ai
Datasetpublic

agentlans/multilingual-text

Multilingual Text Dataset This dataset contains a curated selection of rows from multiple input datasets, where each row includes a text chunk of approximately 2000 tokens (as measured by Llama 3.1 tokenizer) verified to be written in the correct language. Only rows with properly classified language chunks are retained, ensuring high-quality multilingual data for analysis or model training. Preprocessing Steps Normalized whitespace, punctuation, Unicode… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/multilingual-text.

sourceHugging Faceodc-byupdated 1y agoView on Hugging Face
5likes354downloads
16 commits on main
19aa0e31y ago

Update README.md

agentlans
b26e1151y ago

Update README.md

agentlans
af362231y ago

Update README.md

agentlans
59b0adc1y ago

Upload 10 files

agentlans
4df47461y ago

Update README.md

agentlans
575dd511y ago

Update README.md

agentlans
47e13d21y ago

Upload 6 files

agentlans
76aed131y ago

Update README.md

agentlans
8503dee1y ago

Update README.md

agentlans
05554f11y ago

Upload 6 files

agentlans
f41442a1y ago

Update README.md

agentlans
511231c1y ago

Upload 4 files

agentlans
584021b1y ago

Update README.md

agentlans
716c1421y ago

Upload 10 files

agentlans
c7d801c1y ago

Update README.md

agentlans
93876701y ago

initial commit

agentlans