Team Ai
Datasetpublic

deepcopy/handwritten-text-recognition-bongabdo

Dataset Card for Bongabdo Dataset Summary Bongabdo is a curated dataset of full-page Bangla (Bengali) handwritten text, intended for use in offline handwriting recognition tasks using modern neural architectures. It includes high-resolution scanned images of handwritten Bangla scripts, transcriptions, and rich per-document metadata. The data has been contributed by people of diverse age groups, occupations, and genders, making it well-suited for training robust… See the full description on the dataset page: https://huggingface.co/datasets/deepcopy/handwritten-text-recognition-bongabdo.

sourceHugging Faceupdated 1y agoView on Hugging Face
0likes28downloads
Dataset Card

Dataset Card for Bongabdo

Dataset Summary

Bongabdo is a curated dataset of full-page Bangla (Bengali) handwritten text, intended for use in offline handwriting recognition tasks using modern neural architectures. It includes high-resolution scanned images of handwritten Bangla scripts, transcriptions, and rich per-document metadata. The data has been contributed by people of diverse age groups, occupations, and genders, making it well-suited for training robust handwriting recognition models that generalize well across handwriting styles.

The dataset is especially valuable given the low-resource nature of the Bengali language in handwriting datasets. It was introduced in the research paper:

Towards Full-page Offline Bangla Handwritten Text Recognition using Image-to-Sequence Architecture Ayanabha Ghosh, 2023 – IEEE Silchar Subsection Conference, Silchar, Assam, India.

This dataset was ported from https://www.kaggle.com/datasets/joebeachcapital/handwritten-text-recognition-bongabdo

Supported Tasks

  • —Offline Handwritten Text Recognition (HTR)
  • —Document Layout Understanding (potentially, using zone-level metadata)
  • —Bangla NLP Preprocessing (from human handwriting)

Dataset Structure

Each example in the dataset contains the following:

Features

Column NameTypeDescription
imageImageFull-page scanned handwritten Bangla document.
textstringTranscription of the handwritten text from the document (Unicode Bangla).
SNint64Serial number of the example.
FilenamestringName of the image and annotation files.
UsernamestringAn anonymized contributor ID.
Ageint64Age of the writer.
GenderstringGender of the writer (M/F).
OccupationstringWriter’s occupation (e.g., Student).
CategorystringTopical category of the text (e.g., News - Travel).
Char Countfloat64Number of characters in the transcription.
Article linkstringOriginal source URL (for typed text, not the handwriting).
StrikeboolWhether the handwriting includes strike-throughs.
Bangla - EnglishboolWhether the text includes code-switching between Bangla and English.
Multi - ParagraphboolWhether the document contains multiple paragraphs.

Data Splits

The current release does not include predefined training/validation/test splits. Users are encouraged to split the dataset as needed using train_test_split() from the datasets library or other strategies.

Usage Example

python
from datasets import load_dataset

ds = load_dataset("deepcopy/handwritten-text-recognition-bongabdo")
sample = ds["train"][0]
sample["image"].show()
print(sample["text"])

Citation

If you use this dataset in your research or development, please cite the following paper:

bibtex
@inproceedings{ghosh2023bangla,
  title={Towards Full-page Offline Bangla Handwritten Text Recognition using Image-to-Sequence Architecture},
  author={Ghosh, Ayanabha},
  booktitle={IEEE Silchar Subsection Conference},
  year={2023},
  address={Silchar, Assam, India}
}

Licensing

Attribution 4.0 International (CC BY 4.0)