dots.ocr is a self-hosted document parser that combines multilingual text recognition and page layout analysis in one vision-language model. It's for developers and teams converting PDFs or document images into structured text while running inference on their own hardware. The Python project is open source under the MIT license.
The parser identifies page elements such as titles, captions, footnotes and tables, then places them in reading order. It preserves the original language rather than translating the text. Output includes JSON with element positions, categories and text, plus Markdown for the document content. Tables use HTML, and mathematical formulas use LaTeX. A layout image shows where the detected elements sit on the page.
You can extract text alone, detect layout without transcription, or combine both. The parser also produces Markdown that omits page headers and footers. These choices make it useful when you need page structure as well as searchable text.
Inference supports vLLM and Hugging Face Transformers, with Docker deployment available. It can use NVIDIA GPUs through CUDA, and CPU inference is also supported. A hosted live demo provides a separate way to try the model.
The related dots.mocr model extends the document work to webpage screenshots and text in photographs. It can also convert structured graphics into SVG code, with dots.mocr-svg focused on that task. Complex tables and formulas remain difficult, SVG results can be imperfect, and document parsing can occasionally fail.
Claim this page and we'll verify you by hand. dots.ocr gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find dots.ocr?Promote it
Something wrong or outdated on this page?
23.9KUpdated 8 months agoMIT
Linux#Batch processing#Hugging Face integration#Multimodal input
DeepSeek-OCR is an open-source OCR model for developers building document processing tools and researchers studying how AI reads text through images. It runs on your own hardware with NVIDIA CUDA GPUs. Its distinctive focus is visual text compression: representing document images with compact sets of vision tokens for a language model to read.
90.4KUpdated 2 weeks agoApache-2.0
Web#Multilingual#ONNX#Structured output
21.4KUpdated 3 weeks agoApache-2.0
macOS · Web#Batch processing#llama.cpp backend#Multilingual
6.7KUpdated 2 months agoApache-2.0
Windows · Docker · Web#Batch processing#Hugging Face integration#Multilingual
15.5KUpdated 3 days agoApache-2.0
macOS · Windows · Linux · Docker#Multilingual
Unstructured is a local document processing library for developers building LLM applications and document ingestion pipelines. It turns PDFs, Word documents, HTML, emails and images into document elements that applications can use. The Python library is open source under Apache 2.0 and runs on your own hardware, including through Docker images for x86_64 and Apple Silicon.
9.4KUpdated 2 days agoMIT
macOS · Windows · Linux · Android · Docker · Web#Batch processing#LM Studio integration#MCP
PaddleOCR is an open source OCR and document parsing toolkit for developers building document search, RAG systems and AI agents. It runs on your own hardware or a self-hosted server and turns PDFs and images into structured Markdown or JSON. The Python toolkit uses PaddlePaddle and carries the Apache 2.0 license.
Surya is a local OCR toolkit for developers extracting text and structure from PDFs and document images. It combines text recognition, layout analysis and table recognition in one vision-language model, so results retain page structure and reading order rather than just the words.
MonkeyOCR is a local AI document parser for developers and researchers working with English and Chinese PDFs or images. It extracts text, formulas and tables while identifying page structure and relationships between blocks. That makes it useful for documents where plain text extraction loses reading order or separates content from its layout.
xberg, formerly Kreuzberg, is a local document extraction engine for developers building AI search, document processing, and retrieval-augmented generation applications. It reads PDFs, Office files, scanned images, email, and nested archives, extracting text, tables, images, and metadata through one shared engine. It's open source under MIT.