
PaddleOCR is an open source OCR and document parsing toolkit for developers building document search, RAG systems and AI agents. It runs on your own hardware or a self-hosted server and turns PDFs and images into structured Markdown or JSON. The Python toolkit uses PaddlePaddle and carries the Apache 2.0 license.
Its models cover different document tasks. PaddleOCR-VL uses a compact vision-language model to read text, tables, formulas and charts, with support for difficult material such as ancient documents, rare characters and seals. PP-Structure preserves document layout and hierarchy while providing detailed coordinates for text and individual table cells.
For text recognition, PP-OCR handles multilingual documents with Chinese, English, Japanese and Latin-script languages in a single model. It also reads text outside conventional pages, including street scenes, digital displays, dot-matrix characters and markings on industrial components. Model sizes range from small options for constrained devices to larger ones for server workloads.
Local deployment supports CPU inference, with OpenVINO, ONNX Runtime and TensorRT available for acceleration. The toolkit also supports multiple GPUs and model training. Developers can integrate it into C++, C# and Java applications.
The website offers a separate hosted interface and API for online processing. PP-ChatOCR integrates ERNIE for extracting information from documents, and official MCP and Agent Skills services expose OCR and parsing capabilities to AI applications.
Claim this page with an email at paddleocr.ai. PaddleOCR gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find PaddleOCR?Promote it
Something wrong or outdated on this page?
23.9KUpdated 8 months agoMIT
Linux#Batch processing#Hugging Face integration#Multimodal input
DeepSeek-OCR is an open-source OCR model for developers building document processing tools and researchers studying how AI reads text through images. It runs on your own hardware with NVIDIA CUDA GPUs. Its distinctive focus is visual text compression: representing document images with compact sets of vision tokens for a language model to read.
68.2KUpdated 1 day agoMIT
macOS · Windows · Linux#MCP#Works offline
21.4KUpdated 3 weeks agoApache-2.0
macOS · Web#Batch processing#llama.cpp backend#Multilingual
9.2KUpdated 6 months agoMIT
Docker#Hugging Face integration#Multilingual#Multimodal input
huggingface.coComputer Vision Models
#Hugging Face integration#Multimodal input#Structured output
76.8KUpdated 2 days agoApache-2.0
Windows#Multilingual
Docling is an MIT-licensed, open source document parser for developers turning files into structured content for search and AI applications. It runs locally on macOS, Linux, and Windows, including in air-gapped environments. Its PDF processing identifies page layout and reading order, extracts tables, code, and formulas, and classifies images.
Surya is a local OCR toolkit for developers extracting text and structure from PDFs and document images. It combines text recognition, layout analysis and table recognition in one vision-language model, so results retain page structure and reading order rather than just the words.
dots.ocr is a self-hosted document parser that combines multilingual text recognition and page layout analysis in one vision-language model. It's for developers and teams converting PDFs or document images into structured text while running inference on their own hardware. The Python project is open source under the MIT license.
Florence-2 is Microsoft's open-source vision model for developers who want to process images on their own hardware. It handles several image tasks through text prompts, so one model can generate descriptions, read text and locate objects. It runs locally with PyTorch and Hugging Face Transformers on a CPU or CUDA GPU, and uses the MIT license.
Tesseract is an open source OCR engine for extracting text from images, with a command line program and a library developers can embed in their own applications. It's suited to document processing workflows and software that needs text recognition. The project uses the Apache 2.0 license and doesn't include a graphical app.