
Morphik Core is a self-hosted multimodal retrieval engine for developers building AI applications around visually rich documents. It searches diagrams, schematics, charts, and datasheets alongside text, so applications can retrieve information that text extraction alone can miss. You can run it on your own server, including through Docker, or use Morphik's hosted service.
Its visual search uses techniques such as ColPali and represents whole document pages to retain visual context. It supports search across PDFs, images, and videos. Developers can combine document ingestion and retrieval with structured information extraction, metadata labeling, classification, and bounding boxes. Knowledge graphs provide another way to explore the stored information.
A Python SDK and REST API let developers use the engine in their applications. The browser console supports file uploads, connections to data sources, search, and chat with documents. Its web interface can also be embedded in another application with custom branding. MCP support connects it to Claude, Open Web UI, and other MCP clients; data integrations include Google Suite, Slack, and Confluence.
For shared deployments, Morphik Core supports multiple tenants, role-based access control, and access rules expressed in natural language. Deployment choices include on-premises infrastructure and cloud environments such as AWS, GCP, and Azure. Support for self-hosted deployments is limited to installation guidance and community help.
The code uses the Business Source License 1.1, which places conditions on production use. Its additional use grant and commercial terms govern self-hosted deployments.
Claim this page with an email at dev.morphik.ai. Morphik Core gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Morphik Core?Promote it
Something wrong or outdated on this page?
39.9KUpdated 4 days agoMIT
macOS · Windows · Linux · Docker · Web#Knowledge graphs#LLM tracing#Multimodal input
LightRAG combines knowledge graphs with vector search to answer questions across a document collection. It's a self-hosted Python framework for developers building document assistants, particularly where answers depend on relationships between facts in different files, such as legal or financial material.
91.5KUpdated 7 hours agoApache-2.0
macOS · Windows · Linux · Docker#Hybrid search#MCP#Multi-agent workflows
12KUpdated 1 week agoApache-2.0
Docker · Web#Batch processing#Human approval#Multi-agent workflows
31.2KUpdated 21 hours agoApache-2.0
Docker · Web#Knowledge graphs#MCP#Multi-user access
4.4KUpdated 7 months agoApache-2.0
Docker · Web#Batch processing#Multimodal input#Ollama integration
18.3KUpdated 24 hours agoMIT
macOS · Windows · Linux · Docker · Web#Human approval#Hybrid search#llama.cpp backend
RAGFlow is an Apache 2.0 licensed RAG engine for teams building AI agents that need to answer questions from their own documents. It can run on a self-hosted server through Docker on Windows, macOS or Linux. A separate hosted cloud service is available.
Bisheng is an open source, self-hosted platform for teams building AI applications around business documents and processes. Its visual workflow editor combines automated tasks with human feedback, including intervention during multi-turn conversations. It's suited to document review, support ticket assistance and report generation that need more control than a single chatbot exchange.
Cognee gives AI agents persistent memory across sessions, connecting documents, code, and conversations in a searchable knowledge graph. It's for developers who want agents to retain project context and teams whose knowledge sits across tickets, discussions, and repositories. The Python package is open source under Apache 2.0.
Cognita is a self-hosted RAG framework for developers building applications that answer questions using their own documents. The project is archived and no longer maintained. It combines a browser interface for document Q&A with reusable components built on LangChain and LlamaIndex, under the Apache 2.0 open-source license.
DocsGPT is an MIT-licensed, open-source platform for teams that want AI search, assistants and agents over their own documents. It can run on your servers with local models, including fully air-gapped deployments where documents and questions stay inside your network. Answers include the source title and page number so readers can check the evidence.