
OCRmyPDF turns scanned PDFs into documents you can search and copy text from while preserving the resolution of their original images. It's a local command-line tool for people digitizing paper records and developers building document processing systems. Your documents stay on your machine.
It places recognized text beneath the scanned image, with attention to accurate alignment for copying and pasting. Where possible, it adds that text without altering other PDF content. It can produce PDF/A files for long-term storage, optimize images to reduce file size, and straighten or clean scans before recognition. It also validates input and output PDFs.
Tesseract handles text recognition through its language packs, including documents with multiple languages. OCRmyPDF uses available CPU cores to process work in parallel and can handle PDFs with thousands of pages. It runs on Linux, macOS, Windows and FreeBSD, with Docker images for x64 and ARM. The software is open source under the Mozilla Public License 2.0, and requires Tesseract and Ghostscript.
A plugin interface lets other OCR engines replace Tesseract. OCRmyPDF-AppleOCR uses Apple Vision on macOS; OCRmyPDF-EasyOCR uses the PyTorch-based EasyOCR engine, for which a GPU is strongly recommended. OCRmyPDF-PaddleOCR provides GPU-accelerated recognition through PaddleOCR. For document collections that need a searchable management interface, paperless-ngx integrates OCRmyPDF into its processing workflow.
Claim this page and we'll verify you by hand. OCRmyPDF gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find OCRmyPDF?Promote it
Something wrong or outdated on this page?
20.3KUpdated 4 days agoApache-2.0
macOS · Windows · Linux · Docker · Web#Batch processing#Multilingual#Voice cloning
ebook2audiobook turns non-DRM ebooks into narrated audio with chapters and metadata, for readers who want audio editions of their own books. It runs locally on Windows, macOS and Linux, with Docker support and a browser interface built with Gradio. It's open source under Apache 2.0.
796Updated 2 weeks agoMIT
macOS · Windows · Linux#llama.cpp backend#LM Studio integration#Multilingual
37.3KUpdated 24 hours agoAGPL-3.0
macOS · Windows · Docker · Web#Batch processing#Hugging Face integration#MCP
14.4KUpdated 24 hours agoMIT
macOS · Windows · Linux#Batch processing#LM Studio integration#Multilingual
15.5KUpdated 3 days agoApache-2.0
macOS · Windows · Linux · Docker#Multilingual
Unstructured is a local document processing library for developers building LLM applications and document ingestion pipelines. It turns PDFs, Word documents, HTML, emails and images into document elements that applications can use. The Python library is open source under Apache 2.0 and runs on your own hardware, including through Docker images for x86_64 and Apple Silicon.
9.4KUpdated 2 days agoMIT
macOS · Windows · Linux · Android · Docker · Web#Batch processing#LM Studio integration#MCP
Interpreter reads Japanese text from game windows and displays English translations over your screen. It's built for people playing Japanese retro games, with MeikiOCR tuned for game text and pixel fonts. The default OCR and Sugoi V4 translation models run locally, so text stays on your computer and no internet connection is needed after setup. It's open source under the MIT license, with support for Windows, macOS and Linux.
PDFMathTranslate translates scientific PDFs while keeping their page layout, formulas, charts, contents pages and annotations. It's for researchers, students and others who need to read papers in another language without losing the relationship between the text and its figures. It produces both translated PDFs and bilingual documents for comparison with the original.
Subtitle Edit is an MIT-licensed subtitle editor for Windows, macOS and Linux. It's for people creating captions, translating dialogue or fixing subtitles that don't match the video. Its core editing, conversion and video playback work offline on your device, with optional AI tools for transcription and translation.
xberg, formerly Kreuzberg, is a local document extraction engine for developers building AI search, document processing, and retrieval-augmented generation applications. It reads PDFs, Office files, scanned images, email, and nested archives, extracting text, tables, images, and metadata through one shared engine. It's open source under MIT.