
Docker Model Runner lets developers run and serve AI models on their own computer or server using Docker Desktop, Docker Engine or the standalone dmr binary. It pulls models from Docker Hub, OCI registries, and Hugging Face, then stores them locally. Inference runs locally too.
It runs on Windows and macOS through Docker Desktop, and on Linux through Docker Engine. The project is open source under Apache 2.0. Its llama.cpp engine runs GGUF models across supported platforms. For Safetensors models, vLLM serves high throughput workloads on NVIDIA GPUs under Linux or Windows with WSL2; Diffusers generates images with Stable Diffusion on NVIDIA GPUs under Linux. Docker Engine also supports CPU, AMD, and Vulkan backends, while macOS supports Apple Silicon.
OpenAI and Ollama compatible APIs let applications use the models, including coding assistants such as Cline, Continue, Cursor, and Aider. Docker Compose and Testcontainers can work with it as well. Developers can manage local models from the Docker Desktop interface or command line, inspect logs, and view prompt and response details. It can package GGUF and Safetensors files as OCI artifacts for distribution through container registries.
On Linux, inference engines run inside containers. On macOS and Windows, they run in system sandboxes. The Model Runner API has no built-in authentication: any client that can reach it can send inference requests, pull models, or load them.
Claim this page with an email at docs.docker.com. Docker Model Runner gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Docker Model Runner?Promote it
Something wrong or outdated on this page?
182KUpdated 16 hours agoMIT
macOS · Windows · Linux · Docker#GGUF#llama.cpp backend#Multimodal input
Ollama runs language models on your own computer or server. It provides a command-line runner and a local API for people building AI applications or connecting existing tools to models they host themselves. The software is distributed under the MIT license.
130KUpdated 36 minutes agoMIT
Web#Code execution#GGUF#Hugging Face integration
77.4KUpdated 1 year agoMIT
macOS · Windows · Linux · Docker#GGUF#llama.cpp backend#OpenAI-compatible API
11.9KUpdated 4 days agoAGPL-3.0
macOS · Windows · Linux · Android · Docker · Web#GGUF#Hugging Face integration#llama.cpp backend
5.8KUpdated 1 day agoApache-2.0
macOS · Windows · Linux · iOS · Android · Docker#GGUF#Hugging Face integration#llama.cpp backend
7.7KUpdated 5 days agoMIT
macOS · Windows · Linux · Docker · Web#Code execution#GGUF#Hugging Face integration
llama.cpp runs language models on your own hardware and can serve them from a machine you control. It’s an MIT-licensed, open source inference engine for people building local AI apps, running a private model server, or using a model directly from the command line. It supports vision-language models too.
GPT4All is a local AI chatbot for people who want to run language models on their own desktop or laptop and keep conversations on their machine. Its LocalDocs feature lets you ask questions about your own documents without sending them to a cloud service. It suits developers, teams and individuals who want control over their models and data.
KoboldCpp pairs local model inference with a browser interface built for chat, creative writing and roleplay. A fork of llama.cpp, it bundles KoboldAI Lite with tools for keeping character details and story context alongside your conversations. It's open source under AGPL-3.0.
Lemonade is an open source local AI server for people who want to use models on their own hardware or connect them to apps and agents. It handles chat, coding, image generation, speech, transcription, and embeddings. A built-in interface lets you use those capabilities directly, while its server makes them available to other software.
mistral.rs is an open source inference engine for running models on your own computer or self-hosted server. It's for developers building AI applications and people who want local chat, multimodal models and agent tools in the same runtime. The Rust project uses the MIT license.