OpenLLM is a self-hosted LLM server for developers who want to connect their applications to models running on their own hardware or servers. Its OpenAI-compatible API works with clients built for that interface, including the OpenAI Python client and LlamaIndex. The project is open source under the Apache License 2.0.
A built-in browser chat interface and terminal chat let you test models directly before connecting an application. Supported models include DeepSeek R1, Llama, Gemma, Mistral, and Qwen2.5, alongside Qwen2.5-Coder for coding tasks. You can also serve custom models through added model repositories; those repositories must be public.
OpenLLM uses BentoML for model serving and vLLM as an inference backend. It supports Docker and Kubernetes deployments, so the same tool covers local testing and server deployment. GPU requirements depend on the model: listed smaller Gemma models need 12 GB of GPU memory, while larger models require multiple GPUs. Gated models require Hugging Face access and a token.
With local hosting, model inference runs on your machine; on a self-hosted server, it runs on that server. BentoCloud is a separate managed cloud deployment option where inference runs in the cloud. It provides autoscaling, model orchestration, and observability for hosted deployments.
Claim this page and we'll verify you by hand. OpenLLM gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find OpenLLM?Promote it
Something wrong or outdated on this page?
11.9KUpdated 4 days agoAGPL-3.0
macOS · Windows · Linux · Android · Docker · Web#GGUF#Hugging Face integration#llama.cpp backend
KoboldCpp pairs local model inference with a browser interface built for chat, creative writing and roleplay. A fork of llama.cpp, it bundles KoboldAI Lite with tools for keeping character details and story context alongside your conversations. It's open source under AGPL-3.0.
49.3KUpdated 2 hours agoMIT
macOS · Linux · Docker · Web#Code execution#Human approval#llama.cpp backend
130KUpdated 36 minutes agoMIT
Web#Code execution#GGUF#Hugging Face integration
544Updated 14 hours agoMIT
macOS · iOS · Web#Code execution#Distributed execution#Hugging Face integration
3.1KUpdated 3 months agoMIT
macOS · Windows · Linux#Distributed execution#Hugging Face integration#Quantization
7.7KUpdated 5 days agoMIT
macOS · Windows · Linux · Docker · Web#Code execution#GGUF#Hugging Face integration
LocalAI runs language models, speech, vision and image generation on hardware you control. It's for developers and teams that want a self-hosted AI server for their apps without sending model requests to a cloud service. Its OpenAI-compatible API works with existing clients, and it also accepts Anthropic, Ollama and ElevenLabs API calls.
llama.cpp runs language models on your own hardware and can serve them from a machine you control. It’s an MIT-licensed, open source inference engine for people building local AI apps, running a private model server, or using a model directly from the command line. It supports vision-language models too.
Pooled runs a single open model across browser tabs on laptops, desktops and phones, combining their memory when the model won't fit on one device. It's for people who want local AI chat or a coding assistant using hardware they already have. It's open source under the MIT license and requires no account or per-device installation.
Distributed Llama runs a local LLM across several computers, sharing both the computation and the model's memory use. It's for people who want to use their own networked hardware for inference rather than keep the entire workload on one machine. The C++ project is open source under the MIT license.
mistral.rs is an open source inference engine for running models on your own computer or self-hosted server. It's for developers building AI applications and people who want local chat, multimodal models and agent tools in the same runtime. The Rust project uses the MIT license.