Zerox is an MIT-licensed OCR library for developers preparing documents for AI applications. Its Node.js and Python packages run on your own machine or server, while cloud vision models read the document pages and produce Markdown. Document conversion happens locally, but page images go to the selected model provider, so this workflow needs internet access and provider credentials.
The library accepts PDFs, images and office documents such as DOCX, from local files or URLs. Its vision-based approach addresses documents with unusual layouts, tables and charts, where the page's visual structure matters alongside the text. You can keep the Markdown for downstream processing or, in Node.js, extract specific fields as structured data using JSON Schema.
Model choices include GPT-4o and GPT-4.1 through OpenAI or Azure OpenAI, Claude through AWS Bedrock, and Google Gemini. Python also supports Vertex AI and uses LiteLLM for model-provider access. Omni offers a hosted OCR demo separately from the packages you run yourself.
Both packages support selected-page processing and concurrent requests. They also offer a formatting mode that helps keep tables consistent across pages, at the cost of slower processing. Node.js adds orientation correction and edge trimming, plus structured extraction for individual pages or whole documents. It can use a different model for extraction than for OCR. Python supports custom system prompts for controlling the Markdown conversion.
Claim this page and we'll verify you by hand. Zerox gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Zerox?Promote it
Something wrong or outdated on this page?
23.9KUpdated 8 months agoMIT
Linux#Batch processing#Hugging Face integration#Multimodal input
DeepSeek-OCR is an open-source OCR model for developers building document processing tools and researchers studying how AI reads text through images. It runs on your own hardware with NVIDIA CUDA GPUs. Its distinctive focus is visual text compression: representing document images with compact sets of vision tokens for a language model to read.
80.8KUpdated 1 day ago
macOS · Windows · Linux#llama.cpp backend#MCP#MLX
9.4KUpdated 2 days agoMIT
macOS · Windows · Linux · Android · Docker · Web#Batch processing#LM Studio integration#MCP
68.2KUpdated 1 day agoMIT
macOS · Windows · Linux#MCP#Works offline
40.1KUpdated 2 weeks agoApache-2.0
macOS · Linux · Web#Batch processing#llama.cpp backend#Multilingual
19.7KUpdated 6 months agoApache-2.0
Linux · Docker · Web#Batch processing#Distributed execution#OpenAI-compatible API
MinerU parses documents locally into structured text for AI agents, RAG systems and knowledge bases. It's for people working with scanned PDFs, academic papers and Office files whose tables, formulas or page layouts need more care than plain text extraction.
xberg, formerly Kreuzberg, is a local document extraction engine for developers building AI search, document processing, and retrieval-augmented generation applications. It reads PDFs, Office files, scanned images, email, and nested archives, extracting text, tables, images, and metadata through one shared engine. It's open source under MIT.
Docling is an MIT-licensed, open source document parser for developers turning files into structured content for search and AI applications. It runs locally on macOS, Linux, and Windows, including in air-gapped environments. Its PDF processing identifies page layout and reading order, extracts tables, code, and formulas, and classifies images.
Marker is a local document converter for developers and teams turning PDFs, scans and Office files into structured text. It preserves tables, equations and page structure for document processing and AI workflows. Its pipeline reads embedded PDF text and uses Surya OCR where text is missing or damaged, rather than reading every page through a vision model.
olmOCR is an open-source OCR toolkit for turning PDFs and image documents into text for LLM datasets and training. It suits researchers and developers who need readable document content, including pages where columns, figures, or mathematical notation make text extraction difficult. You can run it on your own GPU, including through Docker, or use a remote inference server.