Favicon of Tesseract

Tesseract

An open source OCR engine with a command line tool and C/C++ APIs. Recognizes text in images and creates searchable PDFs under the Apache 2.0 license.

Screenshot of Tesseract website

Tesseract is an open source OCR engine for extracting text from images, with a command line program and a library developers can embed in their own applications. It's suited to document processing workflows and software that needs text recognition. The project uses the Apache 2.0 license and doesn't include a graphical app.

Its neural network engine uses LSTM models to recognize lines of text. A legacy engine recognizes character patterns and requires compatible trained language data. Tesseract supports UTF-8 text and recognition across many languages, and you can train it to recognize additional languages.

It accepts PNG, JPEG and TIFF images. Output can be plain text or a searchable PDF, including a PDF containing only an invisible text layer. For applications that need structured OCR results, it also produces hOCR HTML, TSV, ALTO and PAGE formats.

Developers can use the libtesseract C or C++ API to add recognition to an application, while the command line program suits scripted processing. This gives Tesseract a role both as a standalone OCR tool and as a component inside other software.

Image quality affects recognition results, so poor source images may need cleanup before processing. Tesseract uses Leptonica to read input images; it doesn't accept PDF documents as input.

Similar to Tesseract