Team Ai
Datasetpublic

procesaur/sr-tokenizer-test

Sr Tokenizer test This dataset provides a large Serbian text corpus designed for training and evaluating of tokenizers for Serbian language models. It combines multiple sources of Serbian text in both Cyrillic and Latin scripts, unified into a consistent JSONL format with id and text fields. Dataset Structure Metadata has been stripped; Each record is a JSON object with: id: unique identifier text: raw Serbian text Source coprora Znanje(sr)… See the full description on the dataset page: https://huggingface.co/datasets/procesaur/sr-tokenizer-test.

sourceHugging Facecc-by-nc-sa-4.0updated 5mo agoView on Hugging Face
0likes76downloads
settings

This repository belongs to procesaur on Hugging Face.

Team Ai never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namesr-tokenizer-test
visibilitypublic
licencecc-by-nc-sa-4.0
gatedno
ownerprocesaur
Account settings
procesaur/sr-tokenizer-test · Team Ai