
PrivateGPT is a self-hosted API layer for developers building AI applications around local models. It adds document retrieval, database access and agent tools to an existing model server. Local workflows can work offline and keep data within your environment; web search and connections to online providers need internet access.
It doesn't run models itself. It connects to Ollama, llama.cpp, vLLM, LocalAI or another OpenAI-compatible inference server. Its application API follows Claude-style patterns, with streaming responses, asynchronous processing and token counting. The Python project is open source under Apache 2.0, with Docker deployment available.
Document and artifact ingestion lets applications use uploaded files as context, while retrieval supplies source citations with answers. PrivateGPT also supports database queries, text-to-SQL and CSV analysis. Built-in web search, web fetching and code execution sit alongside custom tools and MCP connectors for agent workflows. It coordinates embeddings, retrieval and model calls through the same API.
PrivateGPT can act as a local backend for Claude Code, OpenCode, VS Code and Cline, or support automation through n8n. Its built-in workbench UI serves testing and demos, so the main audience is developers who need a backend for their own products. The same team builds Zylon, a separate on-premise enterprise platform that adds infrastructure management, governance, user management and operational support.
Claim this page with an email at privategpt.dev. PrivateGPT gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find PrivateGPT?Promote it
Something wrong or outdated on this page?
20.1KUpdated 2 days agoMIT
macOS · Linux · Docker · Web#Code execution#llama.cpp backend#OpenAI-compatible API
DB-GPT is a self-hosted AI data assistant for teams analyzing business data and developers building data applications. It turns plain-language requests into SQL queries and Python analysis, then produces charts, dashboards, or HTML reports. You can run it on macOS or Linux, with Docker deployment also supported.
31.2KUpdated 21 hours agoApache-2.0
Docker · Web#Knowledge graphs#MCP#Multi-user access
4.4KUpdated 7 months agoApache-2.0
Docker · Web#Batch processing#Multimodal input#Ollama integration
157.6KUpdated 3 hours ago
Docker · Web#Code execution#MCP#OpenAI-compatible API
18.3KUpdated 24 hours agoMIT
macOS · Windows · Linux · Docker · Web#Human approval#Hybrid search#llama.cpp backend
29.8KUpdated 1 day ago
Docker · Web#LLM tracing#MCP#Multi-user access
FastGPT is a self-hosted AI agent builder for teams that want assistants to answer questions using company documents and carry out business workflows. Its visual editor connects model calls, knowledge retrieval and tools into applications for customer support, internal knowledge search and document review. You can run the platform on your own server through Docker or use the vendor's hosted service.
Cognee gives AI agents persistent memory across sessions, connecting documents, code, and conversations in a searchable knowledge graph. It's for developers who want agents to retain project context and teams whose knowledge sits across tickets, discussions, and repositories. The Python package is open source under Apache 2.0.
Cognita is a self-hosted RAG framework for developers building applications that answer questions using their own documents. The project is archived and no longer maintained. It combines a browser interface for document Q&A with reusable components built on LangChain and LlamaIndex, under the Apache 2.0 open-source license.
Dify is a source-available platform for teams building AI agents and apps on a visual canvas. Its Community Edition runs on your own server with Docker. Dify also offers a hosted cloud service, while Enterprise deployments can run in a VPC or on a self-hosted server. The Community Edition uses a custom Apache 2.0 derivative license.
DocsGPT is an MIT-licensed, open-source platform for teams that want AI search, assistants and agents over their own documents. It can run on your servers with local models, including fully air-gapped deployments where documents and questions stay inside your network. Answers include the source title and page number so readers can check the evidence.