Favicon of Surya

Surya

Local OCR software for PDFs and images, with reading order, tables and math. Runs on CPU, Apple Silicon or NVIDIA GPUs; code uses Apache 2.0.

Surya is a local OCR toolkit for developers extracting text and structure from PDFs and document images. It combines text recognition, layout analysis and table recognition in one vision-language model, so results retain page structure and reading order rather than just the words.

It identifies headings, footnotes, figures and other page elements, with coordinates and confidence scores in structured output. OCR can cover a full page or individual blocks. HTML output preserves tables and includes equations as KaTeX-compatible LaTeX alongside surrounding text. Table recognition detects rows, columns and cells, with support for spanning cells and header rows. Separate models handle text-line detection and OCR error detection.

Surya supports multilingual documents, including Arabic, Chinese, Hindi and Japanese. It's built for documents. Examples include newspapers, textbooks, tax forms and handwritten notes; text in photographs and natural scenes isn't its focus. A Python interface suits document-processing applications, while a Streamlit app lets users inspect PDFs and images interactively.

Inference runs on your hardware through vllm on NVIDIA GPUs or llama.cpp on CPU and Apple Silicon. It can also connect to an existing OpenAI-compatible inference server. Datalab offers a separate hosted platform that runs Surya and variants of Chandra.

The code is open source under Apache 2.0. Model weights use a modified AI Pubs Open Rail-M license, with permissions for research, personal projects and qualifying startups; broader commercial use requires separate licensing.

Similar to Surya