
olmOCR is an open-source OCR toolkit for turning PDFs and image documents into text for LLM datasets and training. It suits researchers and developers who need readable document content, including pages where columns, figures, or mathematical notation make text extraction difficult. You can run it on your own GPU, including through Docker, or use a remote inference server.
It accepts PDF, PNG, and JPEG documents and produces Markdown. Its document parsing handles equations, tables, handwriting, and complex formatting, preserves a natural reading order across multi-column pages and insets, and removes headers and footers automatically. These capabilities matter when the extracted text needs to follow the document's content rather than its visual placement on the page.
The local model requires a GPU. Remote processing works with vLLM or another inference platform that implements the OpenAI API, and Cirrascale is a verified external provider. Choosing remote inference sends document content to that server; the project also offers an online demo.
For larger collections, olmOCR can distribute conversion across multiple worker nodes and use AWS S3 for document input and coordination. It exports Dolma files as well as Markdown, and includes olmOCR-Bench for evaluating OCR performance on difficult document layouts. The Python toolkit uses the Apache 2.0 license.
Claim this page with an email at olmocr.allenai.org. olmOCR gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find olmOCR?Promote it
Something wrong or outdated on this page?
9.4KUpdated 2 days agoMIT
macOS · Windows · Linux · Android · Docker · Web#Batch processing#LM Studio integration#MCP
xberg, formerly Kreuzberg, is a local document extraction engine for developers building AI search, document processing, and retrieval-augmented generation applications. It reads PDFs, Office files, scanned images, email, and nested archives, extracting text, tables, images, and metadata through one shared engine. It's open source under MIT.
40.1KUpdated 2 weeks agoApache-2.0
macOS · Linux · Web#Batch processing#llama.cpp backend#Multilingual
6.7KUpdated 2 months agoApache-2.0
Windows · Docker · Web#Batch processing#Hugging Face integration#Multilingual
6.4KUpdated 1 day agoApache-2.0
Docker · Web
docTR is an open-source Python OCR library for developers building document processing tools and researchers comparing text recognition models. It reads PDFs and images on your own hardware, locating words and recognizing their text. The library uses PyTorch and carries the Apache 2.0 license.
15.5KUpdated 3 days agoApache-2.0
macOS · Windows · Linux · Docker#Multilingual
Unstructured is a local document processing library for developers building LLM applications and document ingestion pipelines. It turns PDFs, Word documents, HTML, emails and images into document elements that applications can use. The Python library is open source under Apache 2.0 and runs on your own hardware, including through Docker images for x86_64 and Apple Silicon.
21.4KUpdated 3 weeks agoApache-2.0
macOS · Web#Batch processing#llama.cpp backend#Multilingual
Marker is a local document converter for developers and teams turning PDFs, scans and Office files into structured text. It preserves tables, equations and page structure for document processing and AI workflows. Its pipeline reads embedded PDF text and uses Surya OCR where text is missing or damaged, rather than reading every page through a vision model.
MonkeyOCR is a local AI document parser for developers and researchers working with English and Chinese PDFs or images. It extracts text, formulas and tables while identifying page structure and relationships between blocks. That makes it useful for documents where plain text extraction loses reading order or separates content from its layout.
Surya is a local OCR toolkit for developers extracting text and structure from PDFs and document images. It combines text recognition, layout analysis and table recognition in one vision-language model, so results retain page structure and reading order rather than just the words.