Favicon of Serge

Serge

A self-hosted LLM chat interface that runs models through llama.cpp in Docker without remote API keys. Source is licensed under MIT and Apache 2.0.

Screenshot of Serge website

Serge is a self-hosted chat interface for people who want to run language models on their own hardware and talk to them in a browser. The project is archived and no longer maintained. It uses llama.cpp to run models locally, with Alpaca as a named chat model and LLaMA also referenced in its memory requirements.

The application runs in Docker, while the browser provides the chat interface. It doesn't rely on a remote AI API or require API keys. The website states that Serge collects no data and carries no ads, so its core offering is a personal assistant hosted under your control rather than a connection to an external model service.

Serge stores chat history and model parameters locally. It also exposes an API with documentation, which matters if you want to use the local model beyond the supplied chat screen. Its SvelteKit frontend sits alongside a Python API built with FastAPI and LangChain, backed by llama.cpp.

Its source is available under MIT and Apache 2.0 licenses. Windows use requires Docker Desktop and WSL2. Available memory determines which models you can run, and LLaMA can crash if the machine doesn't have enough RAM for the selected model.

Similar to Serge