Team Ai
Datasetpublic

BDRC/berkeley-transcriptions

Tibetan OCR — Berkeley 8,866 page images of Tibetan text with page-level Unicode transcriptions, mostly dbu-med (u-med) manuscripts with a small woodblock (uchen) portion. These works were transcribed by Geshe Dangsong Namgyal, from 2016 to 2026, in his role as a Data System Analyst at University of California, Berkeley. This work was initiated by Prof. Kurt Keutzer with the longstanding hope of producing training data for Handwritten Text Recognition for dbu med and OCR of… See the full description on the dataset page: https://huggingface.co/datasets/BDRC/berkeley-transcriptions.

sourceHugging Facecc0-1.0updated 2mo agoView on Hugging Face
0likes65downloads
settings

This repository belongs to BDRC on Hugging Face.

Team Ai never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

nameberkeley-transcriptions
visibilitypublic
licencecc0-1.0
gatedno
ownerBDRC
Account settings
BDRC/berkeley-transcriptions · Team Ai