Harbor is a CLI and companion app for people experimenting with AI on their own hardware. It manages a local LLM development environment, connecting model backends to chat interfaces and supporting services so you don't have to configure each connection yourself. It's open source under Apache 2.0.
Docker Compose handles the container services. Backends include Ollama, llama.cpp and vLLM, with chat interfaces such as Open WebUI and LibreChat. On macOS, Docker Model Runner, MLX and oMLX can run inference directly on the host with Metal acceleration. GGUF models work through llama.cpp; the PrismML backend supports Ternary Bonsai 2 GGUFs on CPU, NVIDIA and AMD ROCm hardware.
The connections extend beyond chat. SearXNG supplies web search to Open WebUI, Perplexica and Local Deep Research. Speaches adds speech recognition and text-to-speech, while ComfyUI connects FLUX image generation to Open WebUI. MCP integrations let chat interfaces use external tools, and Dify or n8n can support larger AI workflows.
For coding, Harbor connects local OpenAI-compatible backends to installed tools including Codex, Claude Code and OpenCode. Harbor Boost adds workflows for web research, checking deliverables and reviewing scope or style before an agent answers.
Configuration profiles let you keep different model setups, and built-in benchmarks can evaluate models against your own tasks. Services are accessible over your LAN, including from a phone. Harbor can also export your selected stack as a standalone Docker Compose file.
Claim this page and we'll verify you by hand. Harbor gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Harbor?Promote it
Something wrong or outdated on this page?
3.1KUpdated 1 day agoMIT
macOS · Windows · Linux · Docker#GGUF#Hugging Face integration#llama.cpp backend
RamaLama runs and serves AI models on your own hardware using OCI containers. It's aimed at developers who want local chat or a self-hosted inference API with a container workflow they can also use in production. The project uses the MIT license.
77.4KUpdated 1 year agoMIT
macOS · Windows · Linux · Docker#GGUF#llama.cpp backend#OpenAI-compatible API
12KUpdated 12 months agoApache-2.0
macOS · Windows · Linux · Docker · Web#Code execution#llama.cpp backend#Multi-user access
11.9KUpdated 4 days agoAGPL-3.0
macOS · Windows · Linux · Android · Docker · Web#GGUF#Hugging Face integration#llama.cpp backend
5.8KUpdated 1 day agoApache-2.0
macOS · Windows · Linux · iOS · Android · Docker#GGUF#Hugging Face integration#llama.cpp backend
5.8KUpdated 2 days agoMIT
macOS · Windows · Linux · Docker · Web#GGUF#Image-to-image#llama.cpp backend
GPT4All is a local AI chatbot for people who want to run language models on their own desktop or laptop and keep conversations on their machine. Its LocalDocs feature lets you ask questions about your own documents without sending them to a cloud service. It suits developers, teams and individuals who want control over their models and data.
h2oGPT is a self-hosted ChatGPT alternative for people who want to chat with local models and ask questions about their own documents. The project is archived and no longer maintained. It's open source under Apache 2.0, with support for Linux, macOS, Windows and Docker.
KoboldCpp pairs local model inference with a browser interface built for chat, creative writing and roleplay. A fork of llama.cpp, it bundles KoboldAI Lite with tools for keeping character details and story context alongside your conversations. It's open source under AGPL-3.0.
Lemonade is an open source local AI server for people who want to use models on their own hardware or connect them to apps and agents. It handles chat, coding, image generation, speech, transcription, and embeddings. A built-in interface lets you use those capabilities directly, while its server makes them available to other software.
llama-swap is a self-hosted proxy for people running several AI models on their own hardware. It starts the model server a request needs and swaps out another when necessary, so you don't have to keep every model loaded or manage separate API connections in your apps.