Favicon of dots.ocr

dots.ocr

A self-hosted OCR model that extracts text and page structure from PDFs and images, with vLLM, Hugging Face Transformers and CPU inference support.

dots.ocr is a self-hosted document parser that combines multilingual text recognition and page layout analysis in one vision-language model. It's for developers and teams converting PDFs or document images into structured text while running inference on their own hardware. The Python project is open source under the MIT license.

The parser identifies page elements such as titles, captions, footnotes and tables, then places them in reading order. It preserves the original language rather than translating the text. Output includes JSON with element positions, categories and text, plus Markdown for the document content. Tables use HTML, and mathematical formulas use LaTeX. A layout image shows where the detected elements sit on the page.

You can extract text alone, detect layout without transcription, or combine both. The parser also produces Markdown that omits page headers and footers. These choices make it useful when you need page structure as well as searchable text.

Inference supports vLLM and Hugging Face Transformers, with Docker deployment available. It can use NVIDIA GPUs through CUDA, and CPU inference is also supported. A hosted live demo provides a separate way to try the model.

The related dots.mocr model extends the document work to webpage screenshots and text in photographs. It can also convert structured graphics into SVG code, with dots.mocr-svg focused on that task. Complex tables and formulas remain difficult, SVG results can be imperfect, and document parsing can occasionally fail.

Similar to dots.ocr