
Serge is a self-hosted chat interface for people who want to run language models on their own hardware and talk to them in a browser. The project is archived and no longer maintained. It uses llama.cpp to run models locally, with Alpaca as a named chat model and LLaMA also referenced in its memory requirements.
The application runs in Docker, while the browser provides the chat interface. It doesn't rely on a remote AI API or require API keys. The website states that Serge collects no data and carries no ads, so its core offering is a personal assistant hosted under your control rather than a connection to an external model service.
Serge stores chat history and model parameters locally. It also exposes an API with documentation, which matters if you want to use the local model beyond the supplied chat screen. Its SvelteKit frontend sits alongside a Python API built with FastAPI and LangChain, backed by llama.cpp.
Its source is available under MIT and Apache 2.0 licenses. Windows use requires Docker Desktop and WSL2. Available memory determines which models you can run, and LLaMA can crash if the machine doesn't have enough RAM for the selected model.
Claim this page with an email at serge.chat. Serge gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Serge?Promote it
Something wrong or outdated on this page?
1.6KUpdated 1 year agoMIT
Windows · Docker · Web#llama.cpp backend#LM Studio integration#Multimodal input
Amica is a locally runnable interface for talking with customizable 3D AI characters. It's for people who want an animated, voiced character as the face of their AI assistant, with a choice of local LLM backends or cloud services. The project builds on Pixiv's ChatVRM.
12KUpdated 12 months agoApache-2.0
macOS · Windows · Linux · Docker · Web#Code execution#llama.cpp backend#Multi-user access
11.9KUpdated 4 days agoAGPL-3.0
macOS · Windows · Linux · Android · Docker · Web#GGUF#Hugging Face integration#llama.cpp backend
25.8KUpdated 4 months agoApache-2.0
macOS · Windows · Linux · Docker · Web#Hybrid search#llama.cpp backend#Multi-user access
4.8KUpdated 3 weeks agoApache-2.0
macOS · Windows · Linux · Docker · Web#GGUF#Hugging Face integration#llama.cpp backend
47.7KUpdated 1 month agoAGPL-3.0
macOS · Windows · Linux · Docker · Web#GGUF#llama.cpp backend#LoRA
h2oGPT is a self-hosted ChatGPT alternative for people who want to chat with local models and ask questions about their own documents. The project is archived and no longer maintained. It's open source under Apache 2.0, with support for Linux, macOS, Windows and Docker.
KoboldCpp pairs local model inference with a browser interface built for chat, creative writing and roleplay. A fork of llama.cpp, it bundles KoboldAI Lite with tools for keeping character details and story context alongside your conversations. It's open source under AGPL-3.0.
kotaemon is a self-hosted document chat app for people who want to ask questions across their files and check where the answers came from. It runs in a browser on Windows, macOS or Linux, with Docker also supported. The project uses the Apache 2.0 license.
Lollms WebUI is a local, single-user AI interface for people who want text chat and media generation in one place. It runs on Windows, macOS and Linux, with Docker support, and lets writers, developers and other users choose models and task-specific personalities. It's free and open source under Apache 2.0. The project receives minimal maintenance.
text-generation-webui, also called TextGen, runs language models on your own hardware through a desktop app or a self-hosted browser interface. It's for people who want private chat and writing tools, and developers who need a local model API. It works offline without telemetry; web search and page fetching use the internet.