Favicon of Paperless-AI

Paperless-AI

A self-hosted document AI extension for Paperless-ngx that classifies files and searches your archive with Ollama or cloud APIs. MIT licensed.

Paperless-AI is a self-hosted extension for Paperless-ngx users who want automatic document sorting and chat with their archive. It requires an existing Paperless-ngx instance and runs in Docker, with a browser interface for reviewing and processing documents. The project is no longer maintained.

It detects newly added documents, analyzes their contents, and assigns titles, tags, document types and correspondents. Rules let you restrict which documents it processes and control the tags it applies. You can also request AI tagging manually, which is useful when you want to review sensitive files before processing them.

Document chat uses retrieval-augmented generation (RAG) to answer questions using your archive. You can search by meaning and ask about details inside documents, such as a contract date or a bill amount, without relying on exact keyword matches. This combines archive search with the same document classification workflow used for incoming files.

The AI backend determines where analysis happens. Ollama supports local LLM use with Mistral, Llama, Phi-3 and Gemma-2, while OpenAI, Azure, DeepSeek and other cloud services process requests remotely. It also works with OpenAI-compatible backends such as LiteLLM, vLLM and FastChat. Choosing a local backend lets you keep model inference on your own hardware; choosing a cloud API sends document content to that provider for analysis. The extension is open source under the MIT license.

After initial configuration, the documented setup requires restarting the container to build the RAG index.

Similar to Paperless-AI