Surya is a local OCR toolkit for developers extracting text and structure from PDFs and document images. It combines text recognition, layout analysis and table recognition in one vision-language model, so results retain page structure and reading order rather than just the words.
It identifies headings, footnotes, figures and other page elements, with coordinates and confidence scores in structured output. OCR can cover a full page or individual blocks. HTML output preserves tables and includes equations as KaTeX-compatible LaTeX alongside surrounding text. Table recognition detects rows, columns and cells, with support for spanning cells and header rows. Separate models handle text-line detection and OCR error detection.
Surya supports multilingual documents, including Arabic, Chinese, Hindi and Japanese. It's built for documents. Examples include newspapers, textbooks, tax forms and handwritten notes; text in photographs and natural scenes isn't its focus. A Python interface suits document-processing applications, while a Streamlit app lets users inspect PDFs and images interactively.
Inference runs on your hardware through vllm on NVIDIA GPUs or llama.cpp on CPU and Apple Silicon. It can also connect to an existing OpenAI-compatible inference server. Datalab offers a separate hosted platform that runs Surya and variants of Chandra.
The code is open source under Apache 2.0. Model weights use a modified AI Pubs Open Rail-M license, with permissions for research, personal projects and qualifying startups; broader commercial use requires separate licensing.
Claim this page and we'll verify you by hand. Surya gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Surya?Promote it
Something wrong or outdated on this page?
90.4KUpdated 2 weeks agoApache-2.0
Web#Multilingual#ONNX#Structured output
PaddleOCR is an open source OCR and document parsing toolkit for developers building document search, RAG systems and AI agents. It runs on your own hardware or a self-hosted server and turns PDFs and images into structured Markdown or JSON. The Python toolkit uses PaddlePaddle and carries the Apache 2.0 license.
23.9KUpdated 8 months agoMIT
Linux#Batch processing#Hugging Face integration#Multimodal input
DeepSeek-OCR is an open-source OCR model for developers building document processing tools and researchers studying how AI reads text through images. It runs on your own hardware with NVIDIA CUDA GPUs. Its distinctive focus is visual text compression: representing document images with compact sets of vision tokens for a language model to read.
9.2KUpdated 6 months agoMIT
Docker#Hugging Face integration#Multilingual#Multimodal input
dots.ocr is a self-hosted document parser that combines multilingual text recognition and page layout analysis in one vision-language model. It's for developers and teams converting PDFs or document images into structured text while running inference on their own hardware. The Python project is open source under the MIT license.
40.1KUpdated 2 weeks agoApache-2.0
macOS · Linux · Web#Batch processing#llama.cpp backend#Multilingual
Marker is a local document converter for developers and teams turning PDFs, scans and Office files into structured text. It preserves tables, equations and page structure for document processing and AI workflows. Its pipeline reads embedded PDF text and uses Surya OCR where text is missing or damaged, rather than reading every page through a vision model.
9.4KUpdated 2 days agoMIT
macOS · Windows · Linux · Android · Docker · Web#Batch processing#LM Studio integration#MCP
xberg, formerly Kreuzberg, is a local document extraction engine for developers building AI search, document processing, and retrieval-augmented generation applications. It reads PDFs, Office files, scanned images, email, and nested archives, extracting text, tables, images, and metadata through one shared engine. It's open source under MIT.
6.7KUpdated 2 months agoApache-2.0
Windows · Docker · Web#Batch processing#Hugging Face integration#Multilingual
MonkeyOCR is a local AI document parser for developers and researchers working with English and Chinese PDFs or images. It extracts text, formulas and tables while identifying page structure and relationships between blocks. That makes it useful for documents where plain text extraction loses reading order or separates content from its layout.