procesaur/sr-tokenizer-test
Sr Tokenizer test This dataset provides a large Serbian text corpus designed for training and evaluating of tokenizers for Serbian language models. It combines multiple sources of Serbian text in both Cyrillic and Latin scripts, unified into a consistent JSONL format with id and text fields. Dataset Structure Metadata has been stripped; Each record is a JSON object with: id: unique identifier text: raw Serbian text Source coprora Znanje(sr)… See the full description on the dataset page: https://huggingface.co/datasets/procesaur/sr-tokenizer-test.
076
