Favicon of Docker Model Runner

Docker Model Runner

Run AI models through Docker Desktop, Docker Engine or a standalone binary, with local inference and OpenAI and Ollama compatible APIs.

Screenshot of Docker Model Runner website

Docker Model Runner lets developers run and serve AI models on their own computer or server using Docker Desktop, Docker Engine or the standalone dmr binary. It pulls models from Docker Hub, OCI registries, and Hugging Face, then stores them locally. Inference runs locally too.

It runs on Windows and macOS through Docker Desktop, and on Linux through Docker Engine. The project is open source under Apache 2.0. Its llama.cpp engine runs GGUF models across supported platforms. For Safetensors models, vLLM serves high throughput workloads on NVIDIA GPUs under Linux or Windows with WSL2; Diffusers generates images with Stable Diffusion on NVIDIA GPUs under Linux. Docker Engine also supports CPU, AMD, and Vulkan backends, while macOS supports Apple Silicon.

OpenAI and Ollama compatible APIs let applications use the models, including coding assistants such as Cline, Continue, Cursor, and Aider. Docker Compose and Testcontainers can work with it as well. Developers can manage local models from the Docker Desktop interface or command line, inspect logs, and view prompt and response details. It can package GGUF and Safetensors files as OCI artifacts for distribution through container registries.

On Linux, inference engines run inside containers. On macOS and Windows, they run in system sandboxes. The Model Runner API has no built-in authentication: any client that can reach it can send inference requests, pull models, or load them.

Similar to Docker Model Runner