Team Ai
Datasetpublic

BDRC/berkeley-transcriptions

Tibetan OCR — Berkeley 8,866 page images of Tibetan text with page-level Unicode transcriptions, mostly dbu-med (u-med) manuscripts with a small woodblock (uchen) portion. These works were transcribed by Geshe Dangsong Namgyal, from 2016 to 2026, in his role as a Data System Analyst at University of California, Berkeley. This work was initiated by Prof. Kurt Keutzer with the longstanding hope of producing training data for Handwritten Text Recognition for dbu med and OCR of… See the full description on the dataset page: https://huggingface.co/datasets/BDRC/berkeley-transcriptions.

sourceHugging Facecc0-1.0updated 2mo agoView on Hugging Face
0likes65downloads
4 commits on main
ed7b6312mo ago

Update README.md

Eroux
748d65e2mo ago

Update README.md

Eroux
6fcd8a32mo ago

Add files using upload-large-folder tool

Eroux
3ddb5842mo ago

initial commit

Eroux