MonkeyOCR is a local AI document parser for developers and researchers working with English and Chinese PDFs or images. It extracts text, formulas and tables while identifying page structure and relationships between blocks. That makes it useful for documents where plain text extraction loses reading order or separates content from its layout.
Its approach combines layout detection, content recognition and relationship prediction rather than asking a large multimodal model to interpret an entire page at once. The output includes formatted Markdown, a PDF with detected layouts marked, and JSON containing block positions, types, content and relationships. It can process batches of files or recognize just text, tables or formulas.
You can run it locally, deploy it through Docker, or use its Windows support. A Gradio browser interface and FastAPI service provide interactive access and application integration; a separate hosted demo accepts document uploads. Local deployments process documents on your own hardware. Inference backends include LMDeploy and vLLM. The smaller MonkeyOCR-pro-1.2B model favors speed and lower hardware demands, while MonkeyOCR-pro-3B favors parsing accuracy. AWQ quantization is supported, and testing includes an NVIDIA RTX 4060 with 8GB VRAM for the smaller or quantized 3B model.
The source code is open source under Apache 2.0, but the model weights have separate restrictions: academic research and non-commercial evaluation are allowed, while production use and commercial deployment require a written commercial license. Support remains limited for handwriting, photographed text, Traditional Chinese and multilingual documents.
Claim this page and we'll verify you by hand. MonkeyOCR gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find MonkeyOCR?Promote it
Something wrong or outdated on this page?
9.4KUpdated 2 days agoMIT
macOS · Windows · Linux · Android · Docker · Web#Batch processing#LM Studio integration#MCP
xberg, formerly Kreuzberg, is a local document extraction engine for developers building AI search, document processing, and retrieval-augmented generation applications. It reads PDFs, Office files, scanned images, email, and nested archives, extracting text, tables, images, and metadata through one shared engine. It's open source under MIT.
19.7KUpdated 6 months agoApache-2.0
Linux · Docker · Web#Batch processing#Distributed execution#OpenAI-compatible API
15.5KUpdated 3 days agoApache-2.0
macOS · Windows · Linux · Docker#Multilingual
Unstructured is a local document processing library for developers building LLM applications and document ingestion pipelines. It turns PDFs, Word documents, HTML, emails and images into document elements that applications can use. The Python library is open source under Apache 2.0 and runs on your own hardware, including through Docker images for x86_64 and Apple Silicon.
6.4KUpdated 1 day agoApache-2.0
Docker · Web
docTR is an open-source Python OCR library for developers building document processing tools and researchers comparing text recognition models. It reads PDFs and images on your own hardware, locating words and recognizing their text. The library uses PyTorch and carries the Apache 2.0 license.
9.2KUpdated 6 months agoMIT
Docker#Hugging Face integration#Multilingual#Multimodal input
40.1KUpdated 2 weeks agoApache-2.0
macOS · Linux · Web#Batch processing#llama.cpp backend#Multilingual
olmOCR is an open-source OCR toolkit for turning PDFs and image documents into text for LLM datasets and training. It suits researchers and developers who need readable document content, including pages where columns, figures, or mathematical notation make text extraction difficult. You can run it on your own GPU, including through Docker, or use a remote inference server.
dots.ocr is a self-hosted document parser that combines multilingual text recognition and page layout analysis in one vision-language model. It's for developers and teams converting PDFs or document images into structured text while running inference on their own hardware. The Python project is open source under the MIT license.
Marker is a local document converter for developers and teams turning PDFs, scans and Office files into structured text. It preserves tables, equations and page structure for document processing and AI workflows. Its pipeline reads embedded PDF text and uses Surya OCR where text is missing or damaged, rather than reading every page through a vision model.