Minuri/sinhala-test-set-50k
Sinhala Test Set - 50K Sentences A held-out Sinhala test set of 50,000 sentences drawn from the Minuri/diverse_sinhala_dataset corpus. Used for perplexity evaluation of three continually pretrained LLaMA 3.2 1B variants (Models A, B, C) as part of a diversity-driven Sinhala language model adaptation study. Dataset Description This test set was held out strictly from all three pretraining corpora (A, B, C) to enable unbiased perplexity measurement. It covers… See the full description on the dataset page: https://huggingface.co/datasets/Minuri/sinhala-test-set-50k.
Sinhala Test Set - 50K Sentences
A held-out Sinhala test set of 50,000 sentences drawn from the Minuri/diverse_sinhala_dataset corpus. Used for perplexity evaluation of three continually pretrained LLaMA 3.2 1B variants (Models A, B, C) as part of a diversity-driven Sinhala language model adaptation study.
Dataset Description
This test set was held out strictly from all three pretraining corpora (A, B, C) to enable unbiased perplexity measurement. It covers multiple domains as classified by a fine-tuned XLM-RoBERTa domain classifier (94% macro-F1).
Perplexity Results (on this test set)
Source Datasets (via parent corpus)
Dataset Structure
Splits
Format
Available in both JSONL and CSV formats.
Intended Uses
- Perplexity evaluation of Sinhala language models
- Held-out benchmark for continual pretraining experiments
Related Repositories
Sources & Licenses
This dataset contains sentences derived from the following source datasets. Users must comply with the license terms of each:
This dataset is released under CC BY-SA 4.0 in compliance with the ShareAlike terms of Wikipedia and NSINA.
