Team Ai
Datasetpublic

BDRC/danyig-pedri-binary-script-classifier

Danyig vs Pedri Binary Script Classification Dataset Stage-2 binary classifier for distinguishing Danyig (5 subscripts: DraDring, DraRing, Drathung, Gongshabma, Tsegdrig) from Pedri (2 subscripts: Peri, Petsuk). Real-only, all images human-reviewed. Images per class Class train val test All Danyig 480 60 60 600 Pedri 480 60 60 600 Total 960 120 120 1,200 Splits Manuscript-stratified split — each manuscript work appears in exactly… See the full description on the dataset page: https://huggingface.co/datasets/BDRC/danyig-pedri-binary-script-classifier.

sourceHugging Facemitupdated 4mo agoView on Hugging Face
0likes103downloads
README.md128 linesDownload Raw Back to root
1---2license: mit3task_categories:4  - image-classification5tags:6  - tibetan7  - manuscript8  - script-classification9  - bdrc10  - danyig11  - pedri12  - binary13pretty_name: Danyig vs Pedri Binary Script Classification14size_categories:15  - 1K<n<10K16dataset_info:17  features:18    - name: id19      dtype: string20    - name: image_bytes21      dtype: image22    - name: script23      dtype:24        class_label:25          names:26            '0': Danyig27            '1': Pedri28    - name: script_type29      dtype: string30  splits:31    - name: train32      num_bytes: 53090000033      num_examples: 96034    - name: validation35      num_bytes: 5950000036      num_examples: 12037    - name: test38      num_bytes: 7990000039      num_examples: 12040  download_size: 67030000041  dataset_size: 67030000042configs:43  - config_name: default44    data_files:45      - split: train46        path: "train-*-of-*.parquet"47      - split: validation48        path: "val-*-of-*.parquet"49      - split: test50        path: "test-*-of-*.parquet"51---52 53# Danyig vs Pedri Binary Script Classification Dataset54 55Stage-2 binary classifier for distinguishing **Danyig** (5 subscripts: DraDring, DraRing, Drathung, Gongshabma, Tsegdrig) from **Pedri** (2 subscripts: Peri, Petsuk). Real-only, all images human-reviewed.56 57## Images per class58 59| Class | train | val | test | **All** |60|-------|------:|----:|-----:|--------:|61| Danyig | 480 | 60 | 60 | 600 |62| Pedri | 480 | 60 | 60 | 600 |63| **Total** | **960** | **120** | **120** | **1,200** |64 65## Splits66 67Manuscript-stratified split — each manuscript work appears in exactly one of train / val / test (no data leakage across splits).68 69| Split | Images | Works |70|-------|-------:|------:|71| train | 960 | 555 |72| validation | 120 | 12 |73| test | 120 | 116 |74| **Total** | **1,200** | |75 76Page-level split manifest: [`splits/pedri-danyig_combined.json`](splits/pedri-danyig_combined.json).77 78## Parquet schema79 80| Column | Type | Description |81|--------|------|-------------|82| `id` | string | BDRC page id (e.g. `W3CN502-I3CN212840005`) |83| `image_bytes` | binary | JPEG/PNG/TIF page image |84| `script` | string | `Danyig` or `Pedri` |85| `script_type` | string | Subscript name (e.g. `Tsegdrig`, `Petsuk`) |86 87See [`split_stats.json`](split_stats.json) and [`split_stats.md`](split_stats.md) for row-level counts.88 89## Load in Python90 91```python92from datasets import load_dataset93 94ds = load_dataset("BDRC/danyig-pedri-binary-script-classifier")95train = ds["train"]       # 96096val   = ds["validation"]  # 12097test  = ds["test"]        # 12098```99 100```python101from io import BytesIO102from PIL import Image103 104row = train[0]105img = Image.open(BytesIO(row["image_bytes"])).convert("RGB")106print(row["id"], row["script"])107```108 109## Citation110 111```bibtex112@misc{bdrc_danyig_pedri_binary,113  title  = {Danyig vs Pedri Binary Script Classification Dataset},114  author = {Buddhist Digital Resource Center and OpenPecha},115  year   = {2026},116  url    = {https://huggingface.co/datasets/BDRC/danyig-pedri-binary-script-classifier},117  note   = {Images from BDRC}118}119```120 121## License122 123Images taken from the open access collection of the Buddhist Digital Resource Center. Not all images are in the public domain, some are from recent publications possibly under copyright. We provide the images under the Fair Use copyright exception, but any reuse of this dataset will have to be based on a copyright analysis. We provide the classification data under the CC0 1.0 Universal (Public Domain Dedication).124 125## Acknowledgements126 127All images are provided by the Buddhist Digital Resource Center (BDRC). This dataset was developed by Dharmaduta from specifications provided by BDRC for the project "The BDRC Etext Corpus", with funding from the Khyentse Foundation. **[Buddhist Digital Resource Center](https://www.bdrc.io)** (BDRC). Developed by Dharmaduta / OpenPecha.128