BDRC/ALL-BDRC-alignments
Tibetan OCR — ALL-BDRC-alignments 79,572 page images of Tibetan woodblock prints (uchen) aligned page-by-page with hand-verified Unicode transcriptions, line breaks preserved. Transcriptions come from the Asian Classics Input Project (ACIP) Sungbum corpus via the Asian Legacy Library (ALL), normalized to Unicode and manually matched to BDRC scans. This is the largest clean uchen woodblock set in the BDRC Tibetan OCR release — released jointly by the Asian Legacy Library (ALL)… See the full description on the dataset page: https://huggingface.co/datasets/BDRC/ALL-BDRC-alignments.
Tibetan OCR — ALL-BDRC-alignments
<!-- DRAFT. Flagship training set (CC0-ready). Numbers measured 2026-09-01. Public release name: ALL-BDRC-alignments (Asian Legacy Library + BDRC). -->
79,572 page images of Tibetan woodblock prints (uchen) aligned page-by-page with hand-verified Unicode transcriptions, line breaks preserved. Transcriptions come from the Asian Classics Input Project (ACIP) Sungbum corpus via the Asian Legacy Library (ALL), normalized to Unicode and manually matched to BDRC scans. This is the largest clean uchen woodblock set in the BDRC Tibetan OCR release — released jointly by the Asian Legacy Library (ALL) and BDRC, hence the name ALL-BDRC.
Part of the BDRC Tibetan OCR release; see the model `BDRC/tibetan-ocr` and the evaluation benchmark.
At a glance
How it was made
The transcriptions were made by ACIP mostly from blockprints and manuscripts that also exist in the BDRC collection, but no link existed between the two databases. This dataset is the result of three manual steps, carried out by multiple annotators over roughly three months in a collaboration between BDRC and **Dharmaduta**:
- Transcription normalization — the original ACIP files were converted to TEI/XML and then to plain Unicode strings representing the original transcription, with no emendations or markup.
- Identification — each ACIP transcription was matched to a BDRC scan, mostly by manual search. Only matches of the exact version (block print / manuscript) were kept.
- Alignment — transcription and images were aligned page by page, by hand. No changes were made to the transcriptions, and only exact transcriptions of the full image were kept.
Sources & collections
The scans come from the following BDRC image collections:
Contents & schema
One row per page. Columns:
from datasets import load_dataset
ds = load_dataset("BDRC/ALL-BDRC-alignments", split="train")
print(ds[0]["transcription"]); ds[0]["image"]Note: some source scans are very high resolution. For training and inference we recommend capping the long side (e.g. ≤ 5000 × 2500 px).
Splits
Single train split (this is a training resource). For held-out evaluation use the separate BDRC Tibetan OCR benchmark.
How the images are packaged
Embedded image bytes in sharded Parquet (HF datasets Image feature); the viewer renders them and load_dataset streams them.
License, attribution & image rights
- License: CC0-1.0 — use, modify, redistribute freely.
- Attribution requested (not required): please cite the Buddhist Digital Resource Center (BDRC) and the Asian Legacy Library (ALL).
- ⚠ Image-rights disclaimer: BDRC claims no copyright over the scans and ALL claims none over the transcriptions, but verifying the copyright status of each underlying original book is the user's responsibility. Neither BDRC nor ALL can be held liable for misuse of the data by third parties.
Credits
Transcriptions by the Asian Classics Input Project (ACIP) via the Asian Legacy Library. Alignments produced for BDRC by Dharmaduta. Manuscript identification: Tenzin Rabten, Sonam Gyal, Lhujam Gyal, Jampa Gonpo, Tsethar Dolma. Page-by-page alignment: Gade, Kunsang, Lekshey Gyatso, Kaldhen, Choying Tenpa. Funded by the Khyentse Foundation ("The BDRC Etext Corpus").
Citation
@misc{bdrc_all_bdrc_alignments_2026,
title = {Tibetan OCR --- ALL-BDRC-alignments (ACIP Sungbum, woodblock uchen)},
author = {Buddhist Digital Resource Center and Asian Legacy Library},
year = {2026},
howpublished = {Hugging Face},
note = {https://huggingface.co/datasets/BDRC/ALL-BDRC-alignments}
}