Favicon of Langchain-Chatchat

Langchain-Chatchat

Self-hosted AI document chat and agents with Ollama and Xinference support. Runs offline with local models on Windows, macOS and Linux.

Langchain-Chatchat is a self-hosted application for asking questions about your own documents and using AI agents. It focuses on Chinese-language use and open models, with a fully offline setup that can keep documents and model processing on your hardware. Its code is open source under Apache 2.0.

Document answers use retrieval-augmented generation (RAG): the app finds relevant passages in your files and gives them to the model as context. You don't need to train a model on those files. It supports local knowledge-base management and combines keyword and vector search through BM25 and KNN.

The app connects to Ollama, Xinference, LocalAI and FastChat, with models including GLM-4-Chat, Qwen2-Instruct and Llama3. Your choice of backend determines the hardware it can use, including CPUs, GPUs, NPUs and Apple's Metal acceleration. It supports Python 3.8 through 3.11 on Windows, macOS and Linux, with Docker deployment available.

Agents can choose tools automatically, and you can also select tools yourself when a model struggles with tool selection. Beyond document chat, the app supports database questions, web search, arXiv papers, Wolfram queries and image generation. Image conversations work with vision models such as qwen-vl-chat.

A Streamlit browser interface supports multiple conversations and custom system prompts; a FastAPI API lets developers connect other applications. Optional cloud connections through One API include OpenAI, Azure OpenAI and Anthropic Claude. Those requests go to external services rather than staying within the offline setup.

Similar to Langchain-Chatchat