Umi-OCR turns screenshots, images, and scanned documents into text on Windows and Linux. It works offline. The software is free under the MIT License, and recognition runs locally with built-in language libraries. It suits people who need to copy text from a screen, process a folder of images, or make scans searchable without sending them to an online OCR service.
Screenshot recognition captures text from a selected area or an image pasted into the app. For larger jobs, batch recognition handles local images and saves the results as plain text, Markdown, JSONL, or CSV. Its text processing can put columns and paragraphs into reading order, retain spacing in code screenshots, and handle vertical layouts when the OCR engine supports them. You can also exclude areas that contain recurring watermarks or logos, so they don't appear in batch results.
Document recognition extracts text from PDF, XPS, EPUB, MOBI, FB2, and CBZ files. For scanned PDFs, Umi-OCR can produce a searchable PDF with the recognized text behind the page image. Ignore areas can keep headers and footers out of the extracted text. The app also reads QR codes and barcodes from images and can generate code images from text.
Umi-OCR supports the offline RapidOCR-json and PaddleOCR-json engines through plugins. It has a multilingual interface, including English, Japanese, and Traditional Chinese, and provides command-line and HTTP interfaces for people who want to connect OCR to other software.
Claim this page and we'll verify you by hand. Umi-OCR gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Umi-OCR?Promote it
Something wrong or outdated on this page?
14.4KUpdated 24 hours agoMIT
macOS · Windows · Linux#Batch processing#LM Studio integration#Multilingual
Subtitle Edit is an MIT-licensed subtitle editor for Windows, macOS and Linux. It's for people creating captions, translating dialogue or fixing subtitles that don't match the video. Its core editing, conversion and video playback work offline on your device, with optional AI tools for transcription and translation.
20.3KUpdated 4 days agoApache-2.0
macOS · Windows · Linux · Docker · Web#Batch processing#Multilingual#Voice cloning
796Updated 2 weeks agoMIT
macOS · Windows · Linux#llama.cpp backend#LM Studio integration#Multilingual
9.4KUpdated 2 days agoMIT
macOS · Windows · Linux · Android · Docker · Web#Batch processing#LM Studio integration#MCP
982Updated 1 year ago
macOS · Windows · Linux · Docker#Home Assistant integration#Image-to-image#Multimodal input
68.2KUpdated 1 day agoMIT
macOS · Windows · Linux#MCP#Works offline
ebook2audiobook turns non-DRM ebooks into narrated audio with chapters and metadata, for readers who want audio editions of their own books. It runs locally on Windows, macOS and Linux, with Docker support and a browser interface built with Gradio. It's open source under Apache 2.0.
Interpreter reads Japanese text from game windows and displays English translations over your screen. It's built for people playing Japanese retro games, with MeikiOCR tuned for game text and pixel fonts. The default OCR and Sugoi V4 translation models run locally, so text stays on your computer and no internet connection is needed after setup. It's open source under the MIT license, with support for Windows, macOS and Linux.
xberg, formerly Kreuzberg, is a local document extraction engine for developers building AI search, document processing, and retrieval-augmented generation applications. It reads PDFs, Office files, scanned images, email, and nested archives, extracting text, tables, images, and metadata through one shared engine. It's open source under MIT.
CodeProject.AI Server gives developers a shared API for AI tasks that run on their own hardware. It's a self-hosted service for adding image analysis, text processing and generation to applications. Processing stays on the machine running the server, without cloud calls or sending data outside your device or network.
Docling is an MIT-licensed, open source document parser for developers turning files into structured content for search and AI applications. It runs locally on macOS, Linux, and Windows, including in air-gapped environments. Its PDF processing identifies page layout and reading order, extracts tables, code, and formulas, and classifies images.