Favicon of OpenLLM

OpenLLM

An open-source LLM server that runs locally or in the cloud, with OpenAI-compatible APIs, a chat UI, and an Apache 2.0 license.

OpenLLM is a self-hosted LLM server for developers who want to connect their applications to models running on their own hardware or servers. Its OpenAI-compatible API works with clients built for that interface, including the OpenAI Python client and LlamaIndex. The project is open source under the Apache License 2.0.

A built-in browser chat interface and terminal chat let you test models directly before connecting an application. Supported models include DeepSeek R1, Llama, Gemma, Mistral, and Qwen2.5, alongside Qwen2.5-Coder for coding tasks. You can also serve custom models through added model repositories; those repositories must be public.

OpenLLM uses BentoML for model serving and vLLM as an inference backend. It supports Docker and Kubernetes deployments, so the same tool covers local testing and server deployment. GPU requirements depend on the model: listed smaller Gemma models need 12 GB of GPU memory, while larger models require multiple GPUs. Gated models require Hugging Face access and a token.

With local hosting, model inference runs on your machine; on a self-hosted server, it runs on that server. BentoCloud is a separate managed cloud deployment option where inference runs in the cloud. It provides autoscaling, model orchestration, and observability for hosted deployments.

Similar to OpenLLM