Favicon of PrivateGPT

PrivateGPT

An open-source local AI API under Apache 2.0 that connects to Ollama, llama.cpp and other OpenAI-compatible servers for document retrieval and agent workflows.

Screenshot of PrivateGPT website

PrivateGPT is a self-hosted API layer for developers building AI applications around local models. It adds document retrieval, database access and agent tools to an existing model server. Local workflows can work offline and keep data within your environment; web search and connections to online providers need internet access.

It doesn't run models itself. It connects to Ollama, llama.cpp, vLLM, LocalAI or another OpenAI-compatible inference server. Its application API follows Claude-style patterns, with streaming responses, asynchronous processing and token counting. The Python project is open source under Apache 2.0, with Docker deployment available.

Document and artifact ingestion lets applications use uploaded files as context, while retrieval supplies source citations with answers. PrivateGPT also supports database queries, text-to-SQL and CSV analysis. Built-in web search, web fetching and code execution sit alongside custom tools and MCP connectors for agent workflows. It coordinates embeddings, retrieval and model calls through the same API.

PrivateGPT can act as a local backend for Claude Code, OpenCode, VS Code and Cline, or support automation through n8n. Its built-in workbench UI serves testing and demos, so the main audience is developers who need a backend for their own products. The same team builds Zylon, a separate on-premise enterprise platform that adds infrastructure management, governance, user management and operational support.

Similar to PrivateGPT