Offline OCR and Document Scanning Tools

Turn scans, photos and PDFs into searchable text with engines like Tesseract and PaddleOCR, or run OCRmyPDF over a whole archive.

34 tools
Favicon of Docling

Docling

2 videos
Open source document parser that runs locally on macOS, Linux, and Windows, with OCR and structured exports for AI applications.

68.2KUpdated 1 day agoMIT

macOS · Windows · Linux#MCP#Works offline

Favicon of MinerU

MinerU

1 video
Local document parser converts PDFs, Office files and images into Markdown, JSON and LaTeX on macOS, Windows and Linux, with optional remote parsing.

80.8KUpdated 1 day ago

macOS · Windows · Linux#llama.cpp backend#MCP#MLX

A self-hosted document management system that makes scanned paperwork searchable. Runs on your own server with Docker and uses the GPL-3.0 license.

46.2KUpdated 22 hours agoGPL-3.0

Docker · Web#Role-based access#Single sign-on

A self-hosted AI server that gives applications a REST API for local image and text processing on Windows, macOS, Linux and Docker.

982Updated 1 year ago

macOS · Windows · Linux · Docker#Home Assistant integration#Image-to-image#Multimodal input

Local AI PDF parser that converts scientific papers to Markdown with LaTeX math and tables. Runs on CPU or GPU, with MIT code and CC-BY-NC weights.

10.1KUpdated 2 years agoMIT

Windows#Batch processing

An open-source Raycast extension for local LLM chat and writing on macOS and Windows. Connects to local or remote Ollama servers under the MIT license.

253Updated 4 weeks agoMIT

macOS · Windows#MCP#Multilingual#Multimodal input

An open-source screen history tool that runs offline on Windows, macOS and Linux, with local AI search and screenshots stored on your device.

2.9KUpdated 1 year agoAGPL-3.0

macOS · Windows · Linux · Web#Semantic search#Works offline

A local vision-language model for image analysis, document extraction and video understanding, with Apache 2.0 weights and Hugging Face Transformers support.

huggingface.coOCR and Document Scanning

Linux#Batch processing#Hugging Face integration#Multimodal input

An open-source Python document parser for LLM applications. Run it locally or in Docker to process PDFs, Word files and images under Apache 2.0.

15.5KUpdated 3 days agoApache-2.0

macOS · Windows · Linux · Docker#Multilingual

An open-source Python OCR library that reads PDFs and images locally with PyTorch, runs on CPU or GPU, and supports Docker deployment.

6.4KUpdated 1 day agoApache-2.0

Docker · Web

An open-source document extraction engine with a Rust core, CPU-only processing, Docker deployment, and support for Ollama, LM Studio, and vLLM.

9.4KUpdated 2 days agoMIT

macOS · Windows · Linux · Android · Docker · Web#Batch processing#LM Studio integration#MCP

A self-hosted AI add-on for paperless-ngx that extracts text and organizes documents using local Ollama models or cloud APIs. Open source under MIT.

2.7KUpdated 6 days agoMIT

Docker · Web#Batch processing#Human approval#Multimodal input

Open-source OCR library for Node.js and Python that converts documents to Markdown using cloud vision models. MIT licensed; page images leave your machine.

12.3KUpdated 1 year agoMIT

Linux#Multimodal input#Structured output

Favicon of EasyOCR

EasyOCR

2 videos
An open-source Python OCR library that extracts text from images locally, supports multilingual reading, and uses the Apache 2.0 license.

30KUpdated 10 months agoApache-2.0

Windows#Multilingual

Bilingual translation for web pages, formatted PDFs and video subtitles, with local Ollama models or cloud translation engines.

immersivetranslate.comOCR and Document Scanning

macOS · iOS · Android · Browser Extension#Inpainting#Multilingual

Local OCR model that extracts plain or formatted text from images, supports multi-page recognition, and works with Hugging Face Transformers.

8.2KUpdated 2 years ago

#Batch processing#GGUF#Hugging Face integration

Local OCR model that converts images and PDFs to text or Markdown. Runs on NVIDIA GPUs with vLLM or Transformers under the MIT license.

23.9KUpdated 8 months agoMIT

Linux#Batch processing#Hugging Face integration#Multimodal input

A local PDF extraction toolkit in Python with OCR, layout detection and table recognition. Runs on CPU or GPU and uses the AGPL-3.0 license.

10KUpdated 2 years agoAGPL-3.0

#Hugging Face integration#Multilingual

Local OCR and document parsing software that converts PDFs and images to Markdown, recognizes tables and formulas, and runs on NVIDIA GPUs.

6.7KUpdated 2 months agoApache-2.0

Windows · Docker · Web#Batch processing#Hugging Face integration#Multilingual

Open-source Python toolkit for document layout detection and OCR workflows, with pretrained deep learning models and an Apache 2.0 license.

5.8KUpdated 4 years agoApache-2.0

Favicon of Tesseract

Tesseract

4 videos
An open source OCR engine with a command line tool and C/C++ APIs. Recognizes text in images and creates searchable PDFs under the Apache 2.0 license.

76.8KUpdated 2 days agoApache-2.0

Windows#Multilingual

An open-source vision model that runs locally with PyTorch and Hugging Face Transformers, supports CPU or CUDA GPUs, and uses the MIT license.

huggingface.coComputer Vision Models

#Hugging Face integration#Multimodal input#Structured output

An open-source PDF translator that preserves formulas and layouts, runs locally or in Docker, and supports Ollama, Google, DeepL and OpenAI.

37.3KUpdated 1 day agoAGPL-3.0

macOS · Windows · Docker · Web#Batch processing#Hugging Face integration#MCP

Self-hosted computer vision server for images and video, with Docker support, NVIDIA GPU acceleration, and optional Roboflow hosted compute.

2.5KUpdated 1 day ago

macOS · Windows · Linux · Docker#Batch processing#Code execution#Multimodal input

Local audiobook converter turns ebooks into narrated audio with chapters and voice cloning. Runs on Windows, macOS and Linux under Apache 2.0.

20.3KUpdated 4 days agoApache-2.0

macOS · Windows · Linux · Docker · Web#Batch processing#Multilingual#Voice cloning

Local document converter turns PDFs and Office files into Markdown, JSON or HTML, with OCR on CPU, NVIDIA GPUs or Apple Silicon and optional LLM support.

40.1KUpdated 2 weeks agoApache-2.0

macOS · Linux · Web#Batch processing#llama.cpp backend#Multilingual

An open-source OCR toolkit that converts PDFs and images into Markdown using a local GPU or an OpenAI-compatible inference server. Apache 2.0 licensed.

19.7KUpdated 6 months agoApache-2.0

Linux · Docker · Web#Batch processing#Distributed execution#OpenAI-compatible API

A self-hosted OCR model that extracts text and page structure from PDFs and images, with vLLM, Hugging Face Transformers and CPU inference support.

9.2KUpdated 6 months agoMIT

Docker#Hugging Face integration#Multilingual#Multimodal input

An open-source subtitle editor for Windows, macOS and Linux, with Whisper transcription and translation through Ollama or LM Studio.

14.4KUpdated 24 hours agoMIT

macOS · Windows · Linux#Batch processing#LM Studio integration#Multilingual

Local OCR software for PDFs and images, with reading order, tables and math. Runs on CPU, Apple Silicon or NVIDIA GPUs; code uses Apache 2.0.

21.4KUpdated 3 weeks agoApache-2.0

macOS · Web#Batch processing#llama.cpp backend#Multilingual

Favicon of PaddleOCR

PaddleOCR

2 videos
An open source OCR toolkit that runs locally, converts PDFs and images to Markdown or JSON, and supports multilingual text under Apache 2.0.

90.4KUpdated 2 weeks agoApache-2.0

Web#Multilingual#ONNX#Structured output

An open-source command-line OCR tool for searchable PDFs. Runs locally on Linux, macOS, Windows and FreeBSD, using Tesseract for text recognition.

34.9KUpdated 2 days agoMPL-2.0

macOS · Windows · Linux · Docker#Multilingual

Free, MIT-licensed OCR software for Windows and Linux that reads screenshots, image batches, and scanned documents offline.

47.6KUpdated 10 months agoMIT

Windows · Linux#Batch processing#Multilingual#Works offline

An offline screen translator for Windows, macOS and Linux that reads Japanese game text and displays English overlays. Open source under MIT.

796Updated 2 weeks agoMIT

macOS · Windows · Linux#llama.cpp backend#LM Studio integration#Multilingual

More in Documents, Writing and Research