
Tesseract is an open source OCR engine for extracting text from images, with a command line program and a library developers can embed in their own applications. It's suited to document processing workflows and software that needs text recognition. The project uses the Apache 2.0 license and doesn't include a graphical app.
Its neural network engine uses LSTM models to recognize lines of text. A legacy engine recognizes character patterns and requires compatible trained language data. Tesseract supports UTF-8 text and recognition across many languages, and you can train it to recognize additional languages.
It accepts PNG, JPEG and TIFF images. Output can be plain text or a searchable PDF, including a PDF containing only an invisible text layer. For applications that need structured OCR results, it also produces hOCR HTML, TSV, ALTO and PAGE formats.
Developers can use the libtesseract C or C++ API to add recognition to an application, while the command line program suits scripted processing. This gives Tesseract a role both as a standalone OCR tool and as a component inside other software.
Image quality affects recognition results, so poor source images may need cleanup before processing. Tesseract uses Leptonica to read input images; it doesn't accept PDF documents as input.
Claim this page and we'll verify you by hand. Tesseract gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Tesseract?Promote it
Something wrong or outdated on this page?
68.2KUpdated 1 day agoMIT
macOS · Windows · Linux#MCP#Works offline
Docling is an MIT-licensed, open source document parser for developers turning files into structured content for search and AI applications. It runs locally on macOS, Linux, and Windows, including in air-gapped environments. Its PDF processing identifies page layout and reading order, extracts tables, code, and formulas, and classifies images.
90.4KUpdated 2 weeks agoApache-2.0
Web#Multilingual#ONNX#Structured output
PaddleOCR is an open source OCR and document parsing toolkit for developers building document search, RAG systems and AI agents. It runs on your own hardware or a self-hosted server and turns PDFs and images into structured Markdown or JSON. The Python toolkit uses PaddlePaddle and carries the Apache 2.0 license.
30KUpdated 10 months agoApache-2.0
Windows#Multilingual
EasyOCR is a Python OCR library for developers who want to extract text from images on their own hardware. It reads text in photographs and dense documents, so it can serve both scene-text recognition and document processing. It's open source under Apache 2.0.
23.9KUpdated 8 months agoMIT
Linux#Batch processing#Hugging Face integration#Multimodal input
DeepSeek-OCR is an open-source OCR model for developers building document processing tools and researchers studying how AI reads text through images. It runs on your own hardware with NVIDIA CUDA GPUs. Its distinctive focus is visual text compression: representing document images with compact sets of vision tokens for a language model to read.
20.3KUpdated 4 days agoApache-2.0
macOS · Windows · Linux · Docker · Web#Batch processing#Multilingual#Voice cloning
ebook2audiobook turns non-DRM ebooks into narrated audio with chapters and metadata, for readers who want audio editions of their own books. It runs locally on Windows, macOS and Linux, with Docker support and a browser interface built with Gradio. It's open source under Apache 2.0.
796Updated 2 weeks agoMIT
macOS · Windows · Linux#llama.cpp backend#LM Studio integration#Multilingual
Interpreter reads Japanese text from game windows and displays English translations over your screen. It's built for people playing Japanese retro games, with MeikiOCR tuned for game text and pixel fonts. The default OCR and Sugoi V4 translation models run locally, so text stays on your computer and no internet connection is needed after setup. It's open source under the MIT license, with support for Windows, macOS and Linux.