KoboldCpp pairs local model inference with a browser interface built for chat, creative writing and roleplay. A fork of llama.cpp, it bundles KoboldAI Lite with tools for keeping character details and story context alongside your conversations. It's open source under AGPL-3.0.
It runs GGUF and older GGML models on Windows, macOS and Linux as a self-contained executable. Models are separate downloads. You can use a CPU alone or move some or all model processing to a GPU, with CUDA for NVIDIA, Vulkan for NVIDIA and AMD, and Metal support on macOS. Docker is also supported, while Android through Termux and Raspberry Pi require builds from source.
The interface has chat, adventure and storywriting modes, plus memory, world information and author's notes. It accepts Tavern Character Cards and can save or export ongoing stories. TextDB adds retrieval from text collections, and tool calling supports MCP servers.
Its capabilities extend beyond text. It generates and edits images with models such as SDXL, Flux and Qwen Image, supports video and music generation, and recognizes images. Whisper handles speech recognition; Qwen3TTS and Kokoro are among its speech generation backends. OpenAI-compatible and Ollama-compatible APIs let other applications use its inference server.
Local models run on your hardware. Colab and RunPod deployments run on cloud GPUs, while the frontend can also send requests to external providers such as Anthropic, OpenRouter and Gemini.
Claim this page and we'll verify you by hand. koboldcpp gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find koboldcpp?Promote it
Something wrong or outdated on this page?
130KUpdated 38 minutes agoMIT
Web#Code execution#GGUF#Hugging Face integration
llama.cpp runs language models on your own hardware and can serve them from a machine you control. It’s an MIT-licensed, open source inference engine for people building local AI apps, running a private model server, or using a model directly from the command line. It supports vision-language models too.
544Updated 14 hours agoMIT
macOS · iOS · Web#Code execution#Distributed execution#Hugging Face integration
12.5KUpdated 4 months agoApache-2.0
Docker · Web#Hugging Face integration#OpenAI-compatible API
182KUpdated 16 hours agoMIT
macOS · Windows · Linux · Docker#GGUF#llama.cpp backend#Multimodal input
1.6KUpdated 1 year agoMIT
Windows · Docker · Web#llama.cpp backend#LM Studio integration#Multimodal input
655Updated 2 days agoApache-2.0
macOS · Windows · Linux#GGUF#Hugging Face integration#llama.cpp backend
Pooled runs a single open model across browser tabs on laptops, desktops and phones, combining their memory when the model won't fit on one device. It's for people who want local AI chat or a coding assistant using hardware they already have. It's open source under the MIT license and requires no account or per-device installation.
OpenLLM is a self-hosted LLM server for developers who want to connect their applications to models running on their own hardware or servers. Its OpenAI-compatible API works with clients built for that interface, including the OpenAI Python client and LlamaIndex. The project is open source under the Apache License 2.0.
Ollama runs language models on your own computer or server. It provides a command-line runner and a local API for people building AI applications or connecting existing tools to models they host themselves. The software is distributed under the MIT license.
Amica is a locally runnable interface for talking with customizable 3D AI characters. It's for people who want an animated, voiced character as the face of their AI assistant, with a choice of local LLM backends or cloud services. The project builds on Pixiv's ChatVRM.
Docker Model Runner lets developers run and serve AI models on their own computer or server using Docker Desktop, Docker Engine or the standalone dmr binary. It pulls models from Docker Hub, OCI registries, and Hugging Face, then stores them locally. Inference runs locally too.