leonarb/olmocr-demo
0
olmOCR Markdown Converter
This Space uses the olmOCR model pipeline to convert PDFs (including scientific papers) into markdown .txt files that retain document structure, headers, and basic math formatting — ready for Calibre/Kindle or downstream parsing.
- ✅ Vision + text anchor OCR pipeline (via
olmOCR) - ✅ Extracts semantic structure via PDF TOC
- ✅ Outputs clean
.txtin markdown format - ✅ Hugging Face Gradio Space with GPU support
Example Use
Upload a scientific paper in PDF and download a markdown .txt version with preserved headers and inline structure.
Built by @BenedictRichardLeonardi using olmOCR
