BDRC/danyig-pedri-binary-script-classifier
Danyig vs Pedri Binary Script Classification Dataset Stage-2 binary classifier for distinguishing Danyig (5 subscripts: DraDring, DraRing, Drathung, Gongshabma, Tsegdrig) from Pedri (2 subscripts: Peri, Petsuk). Real-only, all images human-reviewed. Images per class Class train val test All Danyig 480 60 60 600 Pedri 480 60 60 600 Total 960 120 120 1,200 Splits Manuscript-stratified split — each manuscript work appears in exactly… See the full description on the dataset page: https://huggingface.co/datasets/BDRC/danyig-pedri-binary-script-classifier.
0103
1---2license: mit3task_categories:4 - image-classification5tags:6 - tibetan7 - manuscript8 - script-classification9 - bdrc10 - danyig11 - pedri12 - binary13pretty_name: Danyig vs Pedri Binary Script Classification14size_categories:15 - 1K<n<10K16dataset_info:17 features:18 - name: id19 dtype: string20 - name: image_bytes21 dtype: image22 - name: script23 dtype:24 class_label:25 names:26 '0': Danyig27 '1': Pedri28 - name: script_type29 dtype: string30 splits:31 - name: train32 num_bytes: 53090000033 num_examples: 96034 - name: validation35 num_bytes: 5950000036 num_examples: 12037 - name: test38 num_bytes: 7990000039 num_examples: 12040 download_size: 67030000041 dataset_size: 67030000042configs:43 - config_name: default44 data_files:45 - split: train46 path: "train-*-of-*.parquet"47 - split: validation48 path: "val-*-of-*.parquet"49 - split: test50 path: "test-*-of-*.parquet"51---52 53# Danyig vs Pedri Binary Script Classification Dataset54 55Stage-2 binary classifier for distinguishing **Danyig** (5 subscripts: DraDring, DraRing, Drathung, Gongshabma, Tsegdrig) from **Pedri** (2 subscripts: Peri, Petsuk). Real-only, all images human-reviewed.56 57## Images per class58 59| Class | train | val | test | **All** |60|-------|------:|----:|-----:|--------:|61| Danyig | 480 | 60 | 60 | 600 |62| Pedri | 480 | 60 | 60 | 600 |63| **Total** | **960** | **120** | **120** | **1,200** |64 65## Splits66 67Manuscript-stratified split — each manuscript work appears in exactly one of train / val / test (no data leakage across splits).68 69| Split | Images | Works |70|-------|-------:|------:|71| train | 960 | 555 |72| validation | 120 | 12 |73| test | 120 | 116 |74| **Total** | **1,200** | |75 76Page-level split manifest: [`splits/pedri-danyig_combined.json`](splits/pedri-danyig_combined.json).77 78## Parquet schema79 80| Column | Type | Description |81|--------|------|-------------|82| `id` | string | BDRC page id (e.g. `W3CN502-I3CN212840005`) |83| `image_bytes` | binary | JPEG/PNG/TIF page image |84| `script` | string | `Danyig` or `Pedri` |85| `script_type` | string | Subscript name (e.g. `Tsegdrig`, `Petsuk`) |86 87See [`split_stats.json`](split_stats.json) and [`split_stats.md`](split_stats.md) for row-level counts.88 89## Load in Python90 91```python92from datasets import load_dataset93 94ds = load_dataset("BDRC/danyig-pedri-binary-script-classifier")95train = ds["train"] # 96096val = ds["validation"] # 12097test = ds["test"] # 12098```99 100```python101from io import BytesIO102from PIL import Image103 104row = train[0]105img = Image.open(BytesIO(row["image_bytes"])).convert("RGB")106print(row["id"], row["script"])107```108 109## Citation110 111```bibtex112@misc{bdrc_danyig_pedri_binary,113 title = {Danyig vs Pedri Binary Script Classification Dataset},114 author = {Buddhist Digital Resource Center and OpenPecha},115 year = {2026},116 url = {https://huggingface.co/datasets/BDRC/danyig-pedri-binary-script-classifier},117 note = {Images from BDRC}118}119```120 121## License122 123Images taken from the open access collection of the Buddhist Digital Resource Center. Not all images are in the public domain, some are from recent publications possibly under copyright. We provide the images under the Fair Use copyright exception, but any reuse of this dataset will have to be based on a copyright analysis. We provide the classification data under the CC0 1.0 Universal (Public Domain Dedication).124 125## Acknowledgements126 127All images are provided by the Buddhist Digital Resource Center (BDRC). This dataset was developed by Dharmaduta from specifications provided by BDRC for the project "The BDRC Etext Corpus", with funding from the Khyentse Foundation. **[Buddhist Digital Resource Center](https://www.bdrc.io)** (BDRC). Developed by Dharmaduta / OpenPecha.128 