Paperless-ngx is a self-hosted document management system for people who want to keep scanned paperwork in a searchable digital archive. It brings document scanning, indexing and storage together, so you can find records by their contents rather than sift through paper folders. You can run it on a local server at home, and it supports deployment with Docker.
Its focus is document archiving: keeping records such as invoices and tax paperwork available for later searches. OCR is part of the project, and the interface supports multiple languages. Paperless-ngx is the official successor to Paperless and Paperless-ng, with a community team responsible for development and support. It's open source under GPL-3.0.
Hosting it yourself lets you choose where the archive lives, but the storage has a concrete privacy limitation: Paperless-ngx stores document information in clear text without encryption. The server therefore needs to be a machine you trust, particularly for financial records or identity documents. A hosted demo is available for trying the interface; it resets its contents frequently and isn't suitable for confidential documents.
Optional AI features can suggest document titles, dates, tags and correspondents, or answer questions about one or several documents. They support local Ollama models and OpenAI-compatible endpoints, with an optional RAG index. These features require configuration; choosing a remote provider sends document content to that provider.
Claim this page and we'll verify you by hand. Paperless-ngx gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Paperless-ngx?Promote it
Something wrong or outdated on this page?
6.4KUpdated 1 day agoApache-2.0
Docker · Web
docTR is an open-source Python OCR library for developers building document processing tools and researchers comparing text recognition models. It reads PDFs and images on your own hardware, locating words and recognizing their text. The library uses PyTorch and carries the Apache 2.0 license.
20.3KUpdated 4 days agoApache-2.0
macOS · Windows · Linux · Docker · Web#Batch processing#Multilingual#Voice cloning
6.7KUpdated 2 months agoApache-2.0
Windows · Docker · Web#Batch processing#Hugging Face integration#Multilingual
19.7KUpdated 6 months agoApache-2.0
Linux · Docker · Web#Batch processing#Distributed execution#OpenAI-compatible API
2.7KUpdated 6 days agoMIT
Docker · Web#Batch processing#Human approval#Multimodal input
37.3KUpdated 24 hours agoAGPL-3.0
macOS · Windows · Docker · Web#Batch processing#Hugging Face integration#MCP
ebook2audiobook turns non-DRM ebooks into narrated audio with chapters and metadata, for readers who want audio editions of their own books. It runs locally on Windows, macOS and Linux, with Docker support and a browser interface built with Gradio. It's open source under Apache 2.0.
MonkeyOCR is a local AI document parser for developers and researchers working with English and Chinese PDFs or images. It extracts text, formulas and tables while identifying page structure and relationships between blocks. That makes it useful for documents where plain text extraction loses reading order or separates content from its layout.
olmOCR is an open-source OCR toolkit for turning PDFs and image documents into text for LLM datasets and training. It suits researchers and developers who need readable document content, including pages where columns, figures, or mathematical notation make text extraction difficult. You can run it on your own GPU, including through Docker, or use a remote inference server.
Paperless-GPT adds AI text extraction and document organization to an existing paperless-ngx library. It runs in Docker on your own server and suits people who want less manual sorting of scanned paperwork. The project is open source under the MIT license.
PDFMathTranslate translates scientific PDFs while keeping their page layout, formulas, charts, contents pages and annotations. It's for researchers, students and others who need to read papers in another language without losing the relationship between the text and its figures. It produces both translated PDFs and bilingual documents for comparison with the original.