LlamaGPT is a self-hosted ChatGPT alternative for people who want general chat or coding help on their own computer or home server. It runs models locally and keeps conversation data on your device. After the initial model download, it works offline.
The supported models include Nous Hermes Llama 2 for general conversation, Code Llama for coding, and Phind Code Llama. The selection covers smaller models for machines with less memory and larger models for more capable hardware. It uses quantized GGML and GGUF models to reduce memory requirements. The smallest listed chat model needs about 6.3 GB of memory; the largest needs about 41.4 GB.
You can use its browser-based chat interface on an umbrelOS home server, an M1 or M2 Mac, or an x86 or arm64 machine running Docker. Kubernetes deployment is also supported. CPU operation is supported, and NVIDIA GPUs can accelerate inference through CUDA. Hardware affects response speed: the published tests include both an M1 Max MacBook Pro and a Raspberry Pi 4, with substantially slower generation on the Pi.
For developers, llama-cpp-python provides an OpenAI-compatible API, so applications that use that API format can connect to the local model server. The chat interface builds on Chatbot UI, while llama.cpp handles model inference. LlamaGPT is open source under the MIT license. The supported model weights have their own license terms.
Claim this page and we'll verify you by hand. LlamaGPT gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find LlamaGPT?Promote it
Something wrong or outdated on this page?
12KUpdated 12 months agoApache-2.0
macOS · Windows · Linux · Docker · Web#Code execution#llama.cpp backend#Multi-user access
h2oGPT is a self-hosted ChatGPT alternative for people who want to chat with local models and ask questions about their own documents. The project is archived and no longer maintained. It's open source under Apache 2.0, with support for Linux, macOS, Windows and Docker.
11.9KUpdated 4 days agoAGPL-3.0
macOS · Windows · Linux · Android · Docker · Web#GGUF#Hugging Face integration#llama.cpp backend
4.8KUpdated 3 weeks agoApache-2.0
macOS · Windows · Linux · Docker · Web#GGUF#Hugging Face integration#llama.cpp backend
47.7KUpdated 1 month agoAGPL-3.0
macOS · Windows · Linux · Docker · Web#GGUF#llama.cpp backend#LoRA
544Updated 14 hours agoMIT
macOS · iOS · Web#Code execution#Distributed execution#Hugging Face integration
12.5KUpdated 4 months agoApache-2.0
Docker · Web#Hugging Face integration#OpenAI-compatible API
KoboldCpp pairs local model inference with a browser interface built for chat, creative writing and roleplay. A fork of llama.cpp, it bundles KoboldAI Lite with tools for keeping character details and story context alongside your conversations. It's open source under AGPL-3.0.
Lollms WebUI is a local, single-user AI interface for people who want text chat and media generation in one place. It runs on Windows, macOS and Linux, with Docker support, and lets writers, developers and other users choose models and task-specific personalities. It's free and open source under Apache 2.0. The project receives minimal maintenance.
text-generation-webui, also called TextGen, runs language models on your own hardware through a desktop app or a self-hosted browser interface. It's for people who want private chat and writing tools, and developers who need a local model API. It works offline without telemetry; web search and page fetching use the internet.
Pooled runs a single open model across browser tabs on laptops, desktops and phones, combining their memory when the model won't fit on one device. It's for people who want local AI chat or a coding assistant using hardware they already have. It's open source under the MIT license and requires no account or per-device installation.
OpenLLM is a self-hosted LLM server for developers who want to connect their applications to models running on their own hardware or servers. Its OpenAI-compatible API works with clients built for that interface, including the OpenAI Python client and LlamaIndex. The project is open source under the Apache License 2.0.