Favicon of koboldcpp

koboldcpp

Local LLM runner for GGUF and GGML models on Windows, macOS and Linux, with CPU or GPU support, a browser UI and an AGPL-3.0 license.

KoboldCpp pairs local model inference with a browser interface built for chat, creative writing and roleplay. A fork of llama.cpp, it bundles KoboldAI Lite with tools for keeping character details and story context alongside your conversations. It's open source under AGPL-3.0.

It runs GGUF and older GGML models on Windows, macOS and Linux as a self-contained executable. Models are separate downloads. You can use a CPU alone or move some or all model processing to a GPU, with CUDA for NVIDIA, Vulkan for NVIDIA and AMD, and Metal support on macOS. Docker is also supported, while Android through Termux and Raspberry Pi require builds from source.

The interface has chat, adventure and storywriting modes, plus memory, world information and author's notes. It accepts Tavern Character Cards and can save or export ongoing stories. TextDB adds retrieval from text collections, and tool calling supports MCP servers.

Its capabilities extend beyond text. It generates and edits images with models such as SDXL, Flux and Qwen Image, supports video and music generation, and recognizes images. Whisper handles speech recognition; Qwen3TTS and Kokoro are among its speech generation backends. OpenAI-compatible and Ollama-compatible APIs let other applications use its inference server.

Local models run on your hardware. Colab and RunPod deployments run on cloud GPUs, while the frontend can also send requests to external providers such as Anthropic, OpenRouter and Gemini.

Similar to koboldcpp