Favicon of text-generation-webui

text-generation-webui

Local LLM desktop app for Windows, macOS and Linux. Run GGUF models offline or connect other apps through OpenAI- and Anthropic-compatible APIs.

text-generation-webui, also called TextGen, runs language models on your own hardware through a desktop app or a self-hosted browser interface. It's for people who want private chat and writing tools, and developers who need a local model API. It works offline without telemetry; web search and page fetching use the internet.

Chat supports instruction-following models and custom characters. You can discuss images with vision models, attach PDF, DOCX or text documents, edit messages and branch conversations. A notebook provides free-form text generation outside the chat format.

The portable desktop app runs GGUF models through llama.cpp on Windows, macOS and Linux. CPU-only operation is supported, alongside NVIDIA and AMD GPUs; the full installation also supports Apple Silicon. Docker deployment is available. Backend choices include ik_llama.cpp, Transformers, ExLlamaV3 and TensorRT-LLM, and you can switch models or backends without restarting.

Its OpenAI- and Anthropic-compatible APIs let other applications use local models for chat and text generation. Tool calling supports custom functions and MCP servers, including tools for web search and calculations.

The full installation includes LoRA fine-tuning on chat or raw text datasets, plus image generation with diffusers models such as Z-Image-Turbo. Extensions add text-to-speech, voice input and translation. The project is open source under AGPL-3.0.

Similar to text-generation-webui