Favicon of RamaLama

RamaLama

An open-source local LLM runner for Linux, macOS and Windows via WSL2, with Podman or Docker isolation and llama.cpp or vLLM inference.

Screenshot of RamaLama website

RamaLama runs and serves AI models on your own hardware using OCI containers. It's aimed at developers who want local chat or a self-hosted inference API with a container workflow they can also use in production. The project uses the MIT license.

It works on Linux and macOS, plus Windows through WSL2 with Docker or Podman. RamaLama detects available GPUs and selects a matching runtime image, falling back to the CPU when needed. Hardware support includes NVIDIA, AMD and Intel GPUs, Apple Silicon, and Ascend NPUs. It supports llama.cpp and vLLM; Apple Silicon Macs can also use MLX directly on the host, outside containers.

Models can come from Hugging Face, ModelScope, Ollama or OCI registries such as Docker Hub, Quay, Pulp and Artifactory. RamaLama can package models as OCI images and convert Safetensors into quantized GGUF models. It also includes model inspection, benchmarking and perplexity measurement for comparing inference performance and model behavior.

Container isolation is a central part of its approach. By default, models run in rootless Podman or Docker containers with read-only model mounts. Local chat containers have networking disabled, so the inference process can't send data out over the network. Model and runtime downloads use remote registries; inference runs on your machine. Temporary data written inside the container is removed when the session ends.

Similar to RamaLama