Team Ai
Datasetpublic

procesaur/sr-tokenizer-test

Sr Tokenizer test This dataset provides a large Serbian text corpus designed for training and evaluating of tokenizers for Serbian language models. It combines multiple sources of Serbian text in both Cyrillic and Latin scripts, unified into a consistent JSONL format with id and text fields. Dataset Structure Metadata has been stripped; Each record is a JSON object with: id: unique identifier text: raw Serbian text Source coprora Znanje(sr)… See the full description on the dataset page: https://huggingface.co/datasets/procesaur/sr-tokenizer-test.

sourceHugging Facecc-by-nc-sa-4.0updated 5mo agoView on Hugging Face
0likes76downloads

procesaur/sr-tokenizer-test · main · files are served by the source, never re-hosted here