Favicon of LocalAI

LocalAI

A self-hosted AI runtime with an OpenAI-compatible API. It runs models on CPUs or GPUs and keeps inference on your own hardware.

Screenshot of LocalAI website

LocalAI runs language models, speech, vision and image generation on hardware you control. It's for developers and teams that want a self-hosted AI server for their apps without sending model requests to a cloud service. Its OpenAI-compatible API works with existing clients, and it also accepts Anthropic, Ollama and ElevenLabs API calls.

The same runtime handles chat, transcription, speech synthesis, object detection, video and 3D workloads. A GPU is optional. LocalAI supports CPU-only machines as well as NVIDIA, AMD, Intel and Apple Silicon acceleration. It runs on macOS and Linux, in Docker or Podman containers, and on Kubernetes.

LocalAI can use different engines for different models, including llama.cpp, vLLM, SGLang and MLX. Its small core obtains backends as models need them. That approach gives teams a way to serve different kinds of AI work through one API without bundling every engine into the base runtime. It supports GGUF models, including the project's APEX quantizations, which are designed to fit larger models into less GPU memory.

The integrated web interface and model gallery provide another way to work with models. For heavier use, distributed mode can place requests across machines according to available hardware and where a model is loaded. LocalAI is open source under the MIT License. Its built-in agent can use tools, MCP servers and skills, and asks for approval before running commands that change state.

Similar to LocalAI