Favicon of PaddleOCR

PaddleOCR

An open source OCR toolkit that runs locally, converts PDFs and images to Markdown or JSON, and supports multilingual text under Apache 2.0.

Screenshot of PaddleOCR website

PaddleOCR is an open source OCR and document parsing toolkit for developers building document search, RAG systems and AI agents. It runs on your own hardware or a self-hosted server and turns PDFs and images into structured Markdown or JSON. The Python toolkit uses PaddlePaddle and carries the Apache 2.0 license.

Its models cover different document tasks. PaddleOCR-VL uses a compact vision-language model to read text, tables, formulas and charts, with support for difficult material such as ancient documents, rare characters and seals. PP-Structure preserves document layout and hierarchy while providing detailed coordinates for text and individual table cells.

For text recognition, PP-OCR handles multilingual documents with Chinese, English, Japanese and Latin-script languages in a single model. It also reads text outside conventional pages, including street scenes, digital displays, dot-matrix characters and markings on industrial components. Model sizes range from small options for constrained devices to larger ones for server workloads.

Local deployment supports CPU inference, with OpenVINO, ONNX Runtime and TensorRT available for acceleration. The toolkit also supports multiple GPUs and model training. Developers can integrate it into C++, C# and Java applications.

The website offers a separate hosted interface and API for online processing. PP-ChatOCR integrates ERNIE for extracting information from documents, and official MCP and Agent Skills services expose OCR and parsing capabilities to AI applications.

Similar to PaddleOCR