
RamaLama runs and serves AI models on your own hardware using OCI containers. It's aimed at developers who want local chat or a self-hosted inference API with a container workflow they can also use in production. The project uses the MIT license.
It works on Linux and macOS, plus Windows through WSL2 with Docker or Podman. RamaLama detects available GPUs and selects a matching runtime image, falling back to the CPU when needed. Hardware support includes NVIDIA, AMD and Intel GPUs, Apple Silicon, and Ascend NPUs. It supports llama.cpp and vLLM; Apple Silicon Macs can also use MLX directly on the host, outside containers.
Models can come from Hugging Face, ModelScope, Ollama or OCI registries such as Docker Hub, Quay, Pulp and Artifactory. RamaLama can package models as OCI images and convert Safetensors into quantized GGUF models. It also includes model inspection, benchmarking and perplexity measurement for comparing inference performance and model behavior.
Container isolation is a central part of its approach. By default, models run in rootless Podman or Docker containers with read-only model mounts. Local chat containers have networking disabled, so the inference process can't send data out over the network. Model and runtime downloads use remote registries; inference runs on your machine. Temporary data written inside the container is removed when the session ends.
Claim this page with an email at ramalama.ai. RamaLama gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find RamaLama?Promote it
Something wrong or outdated on this page?
7.7KUpdated 5 days agoMIT
macOS · Windows · Linux · Docker · Web#Code execution#GGUF#Hugging Face integration
mistral.rs is an open source inference engine for running models on your own computer or self-hosted server. It's for developers building AI applications and people who want local chat, multimodal models and agent tools in the same runtime. The Rust project uses the MIT license.
15.8KUpdated 2 days agoApache-2.0
Web#Distributed execution#Hugging Face integration#LoRA
655Updated 2 days agoApache-2.0
macOS · Windows · Linux#GGUF#Hugging Face integration#llama.cpp backend
77.4KUpdated 1 year agoMIT
macOS · Windows · Linux · Docker#GGUF#llama.cpp backend#OpenAI-compatible API
3.2KUpdated 5 days agoApache-2.0
macOS · Linux · Docker#GGUF#llama.cpp backend#MCP
11.9KUpdated 4 days agoAGPL-3.0
macOS · Windows · Linux · Android · Docker · Web#GGUF#Hugging Face integration#llama.cpp backend
ms-swift is a Python framework for developers and researchers who want to train and deploy language or multimodal models on their own hardware. It brings fine-tuning, evaluation and model serving into one project, with support for Qwen3, DeepSeek-R1, Llama4 and Mistral, plus multimodal models such as Qwen3-VL and InternVL3.5. It's open source under Apache 2.0.
Docker Model Runner lets developers run and serve AI models on their own computer or server using Docker Desktop, Docker Engine or the standalone dmr binary. It pulls models from Docker Hub, OCI registries, and Hugging Face, then stores them locally. Inference runs locally too.
GPT4All is a local AI chatbot for people who want to run language models on their own desktop or laptop and keep conversations on their machine. Its LocalDocs feature lets you ask questions about your own documents without sending them to a cloud service. It suits developers, teams and individuals who want control over their models and data.
Harbor is a CLI and companion app for people experimenting with AI on their own hardware. It manages a local LLM development environment, connecting model backends to chat interfaces and supporting services so you don't have to configure each connection yourself. It's open source under Apache 2.0.
KoboldCpp pairs local model inference with a browser interface built for chat, creative writing and roleplay. A fork of llama.cpp, it bundles KoboldAI Lite with tools for keeping character details and story context alongside your conversations. It's open source under AGPL-3.0.