Favicon of olmOCR

olmOCR

An open-source OCR toolkit that converts PDFs and images into Markdown using a local GPU or an OpenAI-compatible inference server. Apache 2.0 licensed.

Screenshot of olmOCR website

olmOCR is an open-source OCR toolkit for turning PDFs and image documents into text for LLM datasets and training. It suits researchers and developers who need readable document content, including pages where columns, figures, or mathematical notation make text extraction difficult. You can run it on your own GPU, including through Docker, or use a remote inference server.

It accepts PDF, PNG, and JPEG documents and produces Markdown. Its document parsing handles equations, tables, handwriting, and complex formatting, preserves a natural reading order across multi-column pages and insets, and removes headers and footers automatically. These capabilities matter when the extracted text needs to follow the document's content rather than its visual placement on the page.

The local model requires a GPU. Remote processing works with vLLM or another inference platform that implements the OpenAI API, and Cirrascale is a verified external provider. Choosing remote inference sends document content to that server; the project also offers an online demo.

For larger collections, olmOCR can distribute conversion across multiple worker nodes and use AWS S3 for document input and coordination. It exports Dolma files as well as Markdown, and includes olmOCR-Bench for evaluating OCR performance on difficult document layouts. The Python toolkit uses the Apache 2.0 license.

Similar to olmOCR