
Lemonade is an open source local AI server for people who want to use models on their own hardware or connect them to apps and agents. It handles chat, coding, image generation, speech, transcription, and embeddings. A built-in interface lets you use those capabilities directly, while its server makes them available to other software.
It runs on Windows, macOS, and Linux, with Docker also available. Local inference can use a CPU or supported NVIDIA, AMD, and Apple Silicon GPUs; some configurations use AMD Ryzen AI NPUs. That range of hardware support is useful if you want one local AI tool across different computers. Lemonade supports GGUF, FLM, and ONNX models, including custom GGUF and ONNX models from Hugging Face or ModelScope.
Apps can connect through OpenAI, Anthropic, and Ollama compatible APIs. Named integrations include Open WebUI, AnythingLLM, and Claude Code. Developers can also bundle an embeddable version into their own applications, giving users local model access without a separate Lemonade installation.
Lemonade is licensed under Apache 2.0 and sends no telemetry. Local use keeps data on your hardware; routing to a server or cloud API is available when you choose it. Model downloads and registry lookups can contact Hugging Face or ModelScope at your request.
Claim this page with an email at lemonade-server.ai. Lemonade gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Lemonade?Promote it
Something wrong or outdated on this page?
9.6KUpdated 1 day agoApache-2.0
macOS · Windows · Linux · Docker · Web#Batch processing#llama.cpp backend#Multimodal input
Xinference serves language, speech and multimodal models through a shared API on your own computer or servers. It's an open source platform under Apache 2.0 for developers and researchers who want to build applications around models they host. You can also deploy it on cloud infrastructure.
3.7KUpdated 5 months agoMIT
Docker#OpenAI-compatible API#Streaming inference
lmstudio.aiComputer and Browser Agents
macOS · Windows · Linux#llama.cpp backend#MCP#MLX
77KUpdated 21 hours agoApache-2.0
macOS · Windows · Linux · Docker · Web#Code execution#GGUF#Image-to-image
655Updated 2 days agoApache-2.0
macOS · Windows · Linux#GGUF#Hugging Face integration#llama.cpp backend
77.4KUpdated 1 year agoMIT
macOS · Windows · Linux · Docker#GGUF#llama.cpp backend#OpenAI-compatible API
Speaches is a self-hosted speech server for developers who want transcription, translation and speech generation on their own hardware. Its OpenAI-compatible API lets applications use local speech models through tools and SDKs built for OpenAI's API. The project is open source under the MIT license.
LM Studio is a desktop application for downloading and running language models on macOS, Windows and Linux. You can search for models, manage downloads and chat with them through the app. Downloaded models can run offline, including document chat that uses files on your computer.
Unsloth brings model training and everyday AI use into a desktop app for people who want to run models on their own hardware. Its no-code interface covers chat, fine-tuning and media generation on macOS, Windows and Linux. The Unsloth software is open source under Apache 2.0.
Docker Model Runner lets developers run and serve AI models on their own computer or server using Docker Desktop, Docker Engine or the standalone dmr binary. It pulls models from Docker Hub, OCI registries, and Hugging Face, then stores them locally. Inference runs locally too.
GPT4All is a local AI chatbot for people who want to run language models on their own desktop or laptop and keep conversations on their machine. Its LocalDocs feature lets you ask questions about your own documents without sending them to a cloud service. It suits developers, teams and individuals who want control over their models and data.