Favicon of LlamaGPT

LlamaGPT

A self-hosted ChatGPT alternative that runs Llama 2 and Code Llama locally, with an MIT license and an OpenAI-compatible API.

LlamaGPT is a self-hosted ChatGPT alternative for people who want general chat or coding help on their own computer or home server. It runs models locally and keeps conversation data on your device. After the initial model download, it works offline.

The supported models include Nous Hermes Llama 2 for general conversation, Code Llama for coding, and Phind Code Llama. The selection covers smaller models for machines with less memory and larger models for more capable hardware. It uses quantized GGML and GGUF models to reduce memory requirements. The smallest listed chat model needs about 6.3 GB of memory; the largest needs about 41.4 GB.

You can use its browser-based chat interface on an umbrelOS home server, an M1 or M2 Mac, or an x86 or arm64 machine running Docker. Kubernetes deployment is also supported. CPU operation is supported, and NVIDIA GPUs can accelerate inference through CUDA. Hardware affects response speed: the published tests include both an M1 Max MacBook Pro and a Raspberry Pi 4, with substantially slower generation on the Pi.

For developers, llama-cpp-python provides an OpenAI-compatible API, so applications that use that API format can connect to the local model server. The chat interface builds on Chatbot UI, while llama.cpp handles model inference. LlamaGPT is open source under the MIT license. The supported model weights have their own license terms.

Similar to LlamaGPT