Team Ai
Datasetpublic

HumynLabs/Chinese_Documents_Dataset_PDF

Chinese Documents Dataset (PDF) This dataset consists of a curated collection of Chinese-language documents in PDF format. It includes textbooks, research papers, articles, public-domain books, and official documents written in Simplified and Traditional Chinese. The dataset supports AI research in OCR, document understanding, and multilingual text extraction. Contact For queries or collaborations related to this dataset, contact: anoushka@kgen.io… See the full description on the dataset page: https://huggingface.co/datasets/HumynLabs/Chinese_Documents_Dataset_PDF.

sourceHugging Facecc-by-4.0updated 11mo agoView on Hugging Face
0likes585downloads
4 commits on main
184ad5311mo ago

Update README.md

KAI
91bb34e11mo ago

Upload 30 files

KAI
25a371211mo ago

Update README.md

KAI
5520aea11mo ago

initial commit

KAI