Team Ai
Datasetpublic

HumynLabs/Japanese_Documents_Dataset_PDF

Japanese Documents Dataset (PDF) This dataset contains a curated collection of Japanese-language documents in PDF format. The corpus includes textbooks, research papers, news articles, public-domain books, and government publications written in Japanese. It is intended to support AI research in OCR, document understanding, and multilingual text recognition. Contact For queries or collaborations related to this dataset, contact: anoushka@kgen.io… See the full description on the dataset page: https://huggingface.co/datasets/HumynLabs/Japanese_Documents_Dataset_PDF.

sourceHugging Facecc-by-4.0updated 11mo agoView on Hugging Face
2likes912downloads
4 commits on main
18e534611mo ago

Update README.md

KAI
8f093ae11mo ago

Upload 63 files

KAI
bd68df211mo ago

Update README.md

KAI
01510f811mo ago

initial commit

KAI