deepcopy/handwritten-text-recognition-bongabdo
Dataset Card for Bongabdo Dataset Summary Bongabdo is a curated dataset of full-page Bangla (Bengali) handwritten text, intended for use in offline handwriting recognition tasks using modern neural architectures. It includes high-resolution scanned images of handwritten Bangla scripts, transcriptions, and rich per-document metadata. The data has been contributed by people of diverse age groups, occupations, and genders, making it well-suited for training robust… See the full description on the dataset page: https://huggingface.co/datasets/deepcopy/handwritten-text-recognition-bongabdo.
Dataset Card for Bongabdo
Dataset Summary
Bongabdo is a curated dataset of full-page Bangla (Bengali) handwritten text, intended for use in offline handwriting recognition tasks using modern neural architectures. It includes high-resolution scanned images of handwritten Bangla scripts, transcriptions, and rich per-document metadata. The data has been contributed by people of diverse age groups, occupations, and genders, making it well-suited for training robust handwriting recognition models that generalize well across handwriting styles.
The dataset is especially valuable given the low-resource nature of the Bengali language in handwriting datasets. It was introduced in the research paper:
Towards Full-page Offline Bangla Handwritten Text Recognition using Image-to-Sequence Architecture Ayanabha Ghosh, 2023 – IEEE Silchar Subsection Conference, Silchar, Assam, India.
This dataset was ported from https://www.kaggle.com/datasets/joebeachcapital/handwritten-text-recognition-bongabdo
Supported Tasks
- Offline Handwritten Text Recognition (HTR)
- Document Layout Understanding (potentially, using zone-level metadata)
- Bangla NLP Preprocessing (from human handwriting)
Dataset Structure
Each example in the dataset contains the following:
Features
Data Splits
The current release does not include predefined training/validation/test splits. Users are encouraged to split the dataset as needed using train_test_split() from the datasets library or other strategies.
Usage Example
from datasets import load_dataset
ds = load_dataset("deepcopy/handwritten-text-recognition-bongabdo")
sample = ds["train"][0]
sample["image"].show()
print(sample["text"])Citation
If you use this dataset in your research or development, please cite the following paper:
@inproceedings{ghosh2023bangla,
title={Towards Full-page Offline Bangla Handwritten Text Recognition using Image-to-Sequence Architecture},
author={Ghosh, Ayanabha},
booktitle={IEEE Silchar Subsection Conference},
year={2023},
address={Silchar, Assam, India}
}Licensing
Attribution 4.0 International (CC BY 4.0)
