PaxiAI/Vietnamese-Tokenizer-Corpus
Vietnamese Tokenizer Training Corpus v1 This dataset is the training corpus used to build PaxiAI/Vietnamese-Tokenizer. It was created by sampling and combining Vietnamese, English, and source-code datasets into an approximately 8 GiB corpus intended specifically for tokenizer training. Its purpose is to provide a diverse and representative sample from which a Vietnamese-focused Byte-level BPE vocabulary can be learned. Dataset Summary The corpus contains… See the full description on the dataset page: https://huggingface.co/datasets/PaxiAI/Vietnamese-Tokenizer-Corpus.
028
No card is published for this repository, or it could not be fetched from Hugging Face right now.
