Favicon of MonkeyOCR

MonkeyOCR

Local OCR and document parsing software that converts PDFs and images to Markdown, recognizes tables and formulas, and runs on NVIDIA GPUs.

MonkeyOCR is a local AI document parser for developers and researchers working with English and Chinese PDFs or images. It extracts text, formulas and tables while identifying page structure and relationships between blocks. That makes it useful for documents where plain text extraction loses reading order or separates content from its layout.

Its approach combines layout detection, content recognition and relationship prediction rather than asking a large multimodal model to interpret an entire page at once. The output includes formatted Markdown, a PDF with detected layouts marked, and JSON containing block positions, types, content and relationships. It can process batches of files or recognize just text, tables or formulas.

You can run it locally, deploy it through Docker, or use its Windows support. A Gradio browser interface and FastAPI service provide interactive access and application integration; a separate hosted demo accepts document uploads. Local deployments process documents on your own hardware. Inference backends include LMDeploy and vLLM. The smaller MonkeyOCR-pro-1.2B model favors speed and lower hardware demands, while MonkeyOCR-pro-3B favors parsing accuracy. AWQ quantization is supported, and testing includes an NVIDIA RTX 4060 with 8GB VRAM for the smaller or quantized 3B model.

The source code is open source under Apache 2.0, but the model weights have separate restrictions: academic research and non-commercial evaluation are allowed, while production use and commercial deployment require a written commercial license. Support remains limited for handwriting, photographed text, Traditional Chinese and multilingual documents.

Similar to MonkeyOCR