Team Ai
Datasetpublicgated

shiima/Preprocessing-khalidchawtany-ckb_simple_ocr_dataset

CKB Simple OCR Dataset (Processed) Dataset Description This dataset contains 15,000 Central Kurdish (Sorani) OCR samples from the khalidchawtany/ckb_simple_ocr_dataset, preprocessed using the asosoft library for Kurdish text normalization. Languages Central Kurdish (ckb) / Sorani Kurdish Dataset Structure The dataset maintains the same structure as the original CKB Simple OCR dataset with the following columns: image text… See the full description on the dataset page: https://huggingface.co/datasets/shiima/Preprocessing-khalidchawtany-ckb_simple_ocr_dataset.

sourceHugging Facemitupdated 9mo agoView on Hugging Face
0likes5downloads
Dataset Card

No card is published for this repository, or it could not be fetched from Hugging Face right now.