68.2KUpdated 1 day agoMIT
macOS · Windows · Linux#MCP#Works offline
Docling is an MIT-licensed, open source document parser for developers turning files into structured content for search and AI applications. It runs locally on macOS, Linux, and Windows, including in air-gapped environments. Its PDF processing identifies page layout and reading order, extracts tables, code, and formulas, and classifies images.
80.8KUpdated 1 day ago
macOS · Windows · Linux#llama.cpp backend#MCP#MLX
MinerU parses documents locally into structured text for AI agents, RAG systems and knowledge bases. It's for people working with scanned PDFs, academic papers and Office files whose tables, formulas or page layouts need more care than plain text extraction.
46.2KUpdated 22 hours agoGPL-3.0
Docker · Web#Role-based access#Single sign-on
Paperless-ngx is a self-hosted document management system for people who want to keep scanned paperwork in a searchable digital archive. It brings document scanning, indexing and storage together, so you can find records by their contents rather than sift through paper folders. You can run it on a local server at home, and it supports deployment with Docker.
982Updated 1 year ago
macOS · Windows · Linux · Docker#Home Assistant integration#Image-to-image#Multimodal input
CodeProject.AI Server gives developers a shared API for AI tasks that run on their own hardware. It's a self-hosted service for adding image analysis, text processing and generation to applications. Processing stays on the machine running the server, without cloud calls or sending data outside your device or network.
10.1KUpdated 2 years agoMIT
Windows#Batch processing
Nougat is a local AI PDF parser for researchers and developers who need scientific papers as usable text, including their equations and tables. It converts academic PDFs into Markdown-style documents with LaTeX notation, so the output retains structure that plain text extraction can lose.
253Updated 4 weeks agoMIT
macOS · Windows#MCP#Multilingual#Multimodal input
Raycast Ollama brings Ollama models into Raycast on macOS and Windows, for people who want AI chat and text assistance within their desktop launcher. It connects to Ollama on your own machine or a remote server you choose. The extension is open source under the MIT license, and local inference doesn't require an Ollama API key.
2.9KUpdated 1 year agoAGPL-3.0
macOS · Windows · Linux · Web#Semantic search#Works offline
OpenRecall records your screen at regular intervals and makes that history searchable with local AI. It's a free, open-source alternative to Microsoft's Windows Recall and Rewind.ai for people who want to find something they previously saw on their computer. It runs on Windows, macOS and Linux, with a browser interface served from your own machine.
huggingface.coOCR and Document Scanning
Linux#Batch processing#Hugging Face integration#Multimodal input
Qwen2.5-VL is a vision-language model you can run on your own hardware to answer questions about images and video. It's aimed at developers building document processing tools, visual assistants and agents that interact with computer or phone screens. The instruction-tuned 7B model has Apache 2.0 licensing and works with Hugging Face Transformers, with weights available in Safetensors format.
15.5KUpdated 3 days agoApache-2.0
macOS · Windows · Linux · Docker#Multilingual
Unstructured is a local document processing library for developers building LLM applications and document ingestion pipelines. It turns PDFs, Word documents, HTML, emails and images into document elements that applications can use. The Python library is open source under Apache 2.0 and runs on your own hardware, including through Docker images for x86_64 and Apple Silicon.
6.4KUpdated 1 day agoApache-2.0
Docker · Web
docTR is an open-source Python OCR library for developers building document processing tools and researchers comparing text recognition models. It reads PDFs and images on your own hardware, locating words and recognizing their text. The library uses PyTorch and carries the Apache 2.0 license.
9.4KUpdated 2 days agoMIT
macOS · Windows · Linux · Android · Docker · Web#Batch processing#LM Studio integration#MCP
xberg, formerly Kreuzberg, is a local document extraction engine for developers building AI search, document processing, and retrieval-augmented generation applications. It reads PDFs, Office files, scanned images, email, and nested archives, extracting text, tables, images, and metadata through one shared engine. It's open source under MIT.
2.7KUpdated 6 days agoMIT
Docker · Web#Batch processing#Human approval#Multimodal input
Paperless-GPT adds AI text extraction and document organization to an existing paperless-ngx library. It runs in Docker on your own server and suits people who want less manual sorting of scanned paperwork. The project is open source under the MIT license.
12.3KUpdated 1 year agoMIT
Linux#Multimodal input#Structured output
Zerox is an MIT-licensed OCR library for developers preparing documents for AI applications. Its Node.js and Python packages run on your own machine or server, while cloud vision models read the document pages and produce Markdown. Document conversion happens locally, but page images go to the selected model provider, so this workflow needs internet access and provider credentials.
30KUpdated 10 months agoApache-2.0
Windows#Multilingual
EasyOCR is a Python OCR library for developers who want to extract text from images on their own hardware. It reads text in photographs and dense documents, so it can serve both scene-text recognition and document processing. It's open source under Apache 2.0.
immersivetranslate.comOCR and Document Scanning
macOS · iOS · Android · Browser Extension#Inpainting#Multilingual
Immersive Translate is an AI translation extension and mobile app that keeps original text alongside its translation. It's aimed at students, researchers and people who read foreign-language material for work. The bilingual page layout lets readers compare passages without replacing the source text.
8.2KUpdated 2 years ago
#Batch processing#GGUF#Hugging Face integration
GOT-OCR2.0 is an OCR model for developers and researchers who want to extract text from images on their own hardware. It handles both plain text and formatted output through a single model, with recognition modes for selected regions and documents spanning multiple pages. The Python codebase builds on Vary.
23.9KUpdated 8 months agoMIT
Linux#Batch processing#Hugging Face integration#Multimodal input
DeepSeek-OCR is an open-source OCR model for developers building document processing tools and researchers studying how AI reads text through images. It runs on your own hardware with NVIDIA CUDA GPUs. Its distinctive focus is visual text compression: representing document images with compact sets of vision tokens for a language model to read.
10KUpdated 2 years agoAGPL-3.0
#Hugging Face integration#Multilingual
PDF-Extract-Kit is a local AI model toolbox for developers and researchers building document processing applications. It extracts text, tables and mathematical formulas from PDFs, with separate models for identifying page elements and recognizing their contents. It's open source under AGPL-3.0, written in Python, and supports CPU or GPU execution on your own hardware.
6.7KUpdated 2 months agoApache-2.0
Windows · Docker · Web#Batch processing#Hugging Face integration#Multilingual
MonkeyOCR is a local AI document parser for developers and researchers working with English and Chinese PDFs or images. It extracts text, formulas and tables while identifying page structure and relationships between blocks. That makes it useful for documents where plain text extraction loses reading order or separates content from its layout.
5.8KUpdated 4 years agoApache-2.0
LayoutParser is an open-source Python library for developers and researchers who need to detect page structure in document images and turn OCR output into structured data. Its pretrained deep learning models share a common interface, so you can work with models trained on different document datasets without rewriting the surrounding pipeline.
76.8KUpdated 2 days agoApache-2.0
Windows#Multilingual
Tesseract is an open source OCR engine for extracting text from images, with a command line program and a library developers can embed in their own applications. It's suited to document processing workflows and software that needs text recognition. The project uses the Apache 2.0 license and doesn't include a graphical app.
huggingface.coComputer Vision Models
#Hugging Face integration#Multimodal input#Structured output
Florence-2 is Microsoft's open-source vision model for developers who want to process images on their own hardware. It handles several image tasks through text prompts, so one model can generate descriptions, read text and locate objects. It runs locally with PyTorch and Hugging Face Transformers on a CPU or CUDA GPU, and uses the MIT license.
37.3KUpdated 1 day agoAGPL-3.0
macOS · Windows · Docker · Web#Batch processing#Hugging Face integration#MCP
PDFMathTranslate translates scientific PDFs while keeping their page layout, formulas, charts, contents pages and annotations. It's for researchers, students and others who need to read papers in another language without losing the relationship between the text and its figures. It produces both translated PDFs and bilingual documents for comparison with the original.
2.5KUpdated 1 day ago
macOS · Windows · Linux · Docker#Batch processing#Code execution#Multimodal input
Roboflow Inference is a self-hosted computer vision server for teams building camera and image analysis systems. It runs on your own computer, server, or edge device and combines model predictions with workflows for tracking, counting, measuring, and responding to events. Roboflow also offers hosted servers and a Serverless Cloud API, where processing runs on its infrastructure.
20.3KUpdated 4 days agoApache-2.0
macOS · Windows · Linux · Docker · Web#Batch processing#Multilingual#Voice cloning
ebook2audiobook turns non-DRM ebooks into narrated audio with chapters and metadata, for readers who want audio editions of their own books. It runs locally on Windows, macOS and Linux, with Docker support and a browser interface built with Gradio. It's open source under Apache 2.0.
40.1KUpdated 2 weeks agoApache-2.0
macOS · Linux · Web#Batch processing#llama.cpp backend#Multilingual
Marker is a local document converter for developers and teams turning PDFs, scans and Office files into structured text. It preserves tables, equations and page structure for document processing and AI workflows. Its pipeline reads embedded PDF text and uses Surya OCR where text is missing or damaged, rather than reading every page through a vision model.
19.7KUpdated 6 months agoApache-2.0
Linux · Docker · Web#Batch processing#Distributed execution#OpenAI-compatible API
olmOCR is an open-source OCR toolkit for turning PDFs and image documents into text for LLM datasets and training. It suits researchers and developers who need readable document content, including pages where columns, figures, or mathematical notation make text extraction difficult. You can run it on your own GPU, including through Docker, or use a remote inference server.
9.2KUpdated 6 months agoMIT
Docker#Hugging Face integration#Multilingual#Multimodal input
dots.ocr is a self-hosted document parser that combines multilingual text recognition and page layout analysis in one vision-language model. It's for developers and teams converting PDFs or document images into structured text while running inference on their own hardware. The Python project is open source under the MIT license.
14.4KUpdated 24 hours agoMIT
macOS · Windows · Linux#Batch processing#LM Studio integration#Multilingual
Subtitle Edit is an MIT-licensed subtitle editor for Windows, macOS and Linux. It's for people creating captions, translating dialogue or fixing subtitles that don't match the video. Its core editing, conversion and video playback work offline on your device, with optional AI tools for transcription and translation.
21.4KUpdated 3 weeks agoApache-2.0
macOS · Web#Batch processing#llama.cpp backend#Multilingual
Surya is a local OCR toolkit for developers extracting text and structure from PDFs and document images. It combines text recognition, layout analysis and table recognition in one vision-language model, so results retain page structure and reading order rather than just the words.
90.4KUpdated 2 weeks agoApache-2.0
Web#Multilingual#ONNX#Structured output
PaddleOCR is an open source OCR and document parsing toolkit for developers building document search, RAG systems and AI agents. It runs on your own hardware or a self-hosted server and turns PDFs and images into structured Markdown or JSON. The Python toolkit uses PaddlePaddle and carries the Apache 2.0 license.
34.9KUpdated 2 days agoMPL-2.0
macOS · Windows · Linux · Docker#Multilingual
OCRmyPDF turns scanned PDFs into documents you can search and copy text from while preserving the resolution of their original images. It's a local command-line tool for people digitizing paper records and developers building document processing systems. Your documents stay on your machine.
47.6KUpdated 10 months agoMIT
Windows · Linux#Batch processing#Multilingual#Works offline
Umi-OCR turns screenshots, images, and scanned documents into text on Windows and Linux. It works offline. The software is free under the MIT License, and recognition runs locally with built-in language libraries. It suits people who need to copy text from a screen, process a folder of images, or make scans searchable without sending them to an online OCR service.
796Updated 2 weeks agoMIT
macOS · Windows · Linux#llama.cpp backend#LM Studio integration#Multilingual
Interpreter reads Japanese text from game windows and displays English translations over your screen. It's built for people playing Japanese retro games, with MeikiOCR tuned for game text and pixel fonts. The default OCR and Sugoi V4 translation models run locally, so text stays on your computer and no internet connection is needed after setup. It's open source under the MIT license, with support for Windows, macOS and Linux.