Favicon of OCRmyPDF

OCRmyPDF

An open-source command-line OCR tool for searchable PDFs. Runs locally on Linux, macOS, Windows and FreeBSD, using Tesseract for text recognition.

Screenshot of OCRmyPDF website

OCRmyPDF turns scanned PDFs into documents you can search and copy text from while preserving the resolution of their original images. It's a local command-line tool for people digitizing paper records and developers building document processing systems. Your documents stay on your machine.

It places recognized text beneath the scanned image, with attention to accurate alignment for copying and pasting. Where possible, it adds that text without altering other PDF content. It can produce PDF/A files for long-term storage, optimize images to reduce file size, and straighten or clean scans before recognition. It also validates input and output PDFs.

Tesseract handles text recognition through its language packs, including documents with multiple languages. OCRmyPDF uses available CPU cores to process work in parallel and can handle PDFs with thousands of pages. It runs on Linux, macOS, Windows and FreeBSD, with Docker images for x64 and ARM. The software is open source under the Mozilla Public License 2.0, and requires Tesseract and Ghostscript.

A plugin interface lets other OCR engines replace Tesseract. OCRmyPDF-AppleOCR uses Apple Vision on macOS; OCRmyPDF-EasyOCR uses the PyTorch-based EasyOCR engine, for which a GPU is strongly recommended. OCRmyPDF-PaddleOCR provides GPU-accelerated recognition through PaddleOCR. For document collections that need a searchable management interface, paperless-ngx integrates OCRmyPDF into its processing workflow.

Similar to OCRmyPDF