Favicon of Lemonade

Lemonade

An open source local AI server for chat, image generation, and speech on Windows, macOS, and Linux, with APIs for apps and agents.

Screenshot of Lemonade website

Lemonade is an open source local AI server for people who want to use models on their own hardware or connect them to apps and agents. It handles chat, coding, image generation, speech, transcription, and embeddings. A built-in interface lets you use those capabilities directly, while its server makes them available to other software.

It runs on Windows, macOS, and Linux, with Docker also available. Local inference can use a CPU or supported NVIDIA, AMD, and Apple Silicon GPUs; some configurations use AMD Ryzen AI NPUs. That range of hardware support is useful if you want one local AI tool across different computers. Lemonade supports GGUF, FLM, and ONNX models, including custom GGUF and ONNX models from Hugging Face or ModelScope.

Apps can connect through OpenAI, Anthropic, and Ollama compatible APIs. Named integrations include Open WebUI, AnythingLLM, and Claude Code. Developers can also bundle an embeddable version into their own applications, giving users local model access without a separate Lemonade installation.

Lemonade is licensed under Apache 2.0 and sends no telemetry. Local use keeps data on your hardware; routing to a server or cloud API is available when you choose it. Model downloads and registry lookups can contact Hugging Face or ModelScope at your request.

Similar to Lemonade