
LocalAI runs language models, speech, vision and image generation on hardware you control. It's for developers and teams that want a self-hosted AI server for their apps without sending model requests to a cloud service. Its OpenAI-compatible API works with existing clients, and it also accepts Anthropic, Ollama and ElevenLabs API calls.
The same runtime handles chat, transcription, speech synthesis, object detection, video and 3D workloads. A GPU is optional. LocalAI supports CPU-only machines as well as NVIDIA, AMD, Intel and Apple Silicon acceleration. It runs on macOS and Linux, in Docker or Podman containers, and on Kubernetes.
LocalAI can use different engines for different models, including llama.cpp, vLLM, SGLang and MLX. Its small core obtains backends as models need them. That approach gives teams a way to serve different kinds of AI work through one API without bundling every engine into the base runtime. It supports GGUF models, including the project's APEX quantizations, which are designed to fit larger models into less GPU memory.
The integrated web interface and model gallery provide another way to work with models. For heavier use, distributed mode can place requests across machines according to available hardware and where a model is loaded. LocalAI is open source under the MIT License. Its built-in agent can use tools, MCP servers and skills, and asks for approval before running commands that change state.
Claim this page with an email at localai.io. LocalAI gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find LocalAI?Promote it
Something wrong or outdated on this page?
47.7KUpdated 1 month agoApache-2.0
macOS · Linux · Web#Distributed execution#Hugging Face integration#MLX
exo is a local LLM runner that combines your devices into a cluster, letting you use models too large for one machine's memory. It's for people who want to run large models on their own hardware and developers connecting existing AI clients to local inference. It runs on macOS and Linux under the Apache 2.0 license.
223Updated 16 hours agoApache-2.0
macOS · Linux#GGUF#Git integration#Guardrails
LLMKube is a free, open-source Kubernetes operator for teams and homelab owners running local LLM inference across their own hardware. It manages Linux GPU servers and Apple Silicon Macs together, so a mixed fleet can serve models through the same platform. It uses the Apache 2.0 license.
12.5KUpdated 4 months agoApache-2.0
Docker · Web#Hugging Face integration#OpenAI-compatible API
OpenLLM is a self-hosted LLM server for developers who want to connect their applications to models running on their own hardware or servers. Its OpenAI-compatible API works with clients built for that interface, including the OpenAI Python client and LlamaIndex. The project is open source under the Apache License 2.0.
3.1KUpdated 3 months agoMIT
macOS · Windows · Linux#Distributed execution#Hugging Face integration#Quantization
Distributed Llama runs a local LLM across several computers, sharing both the computation and the model's memory use. It's for people who want to use their own networked hardware for inference rather than keep the entire workload on one machine. The C++ project is open source under the MIT license.
1.3KUpdated 1 day agoApache-2.0
Web#LoRA#Multimodal input#Ollama integration
KubeAI is an open source Kubernetes operator for teams serving AI models on their own infrastructure or cloud clusters. It manages model servers and scales them with demand, including starting from zero running replicas. It uses the Apache 2.0 license and can run on CPUs, GPUs or TPUs, including in a local Kubernetes cluster.
2.6KUpdated 20 hours agoApache-2.0
Web#OpenAI-compatible API#Prompt caching
vLLM Production Stack is an open source inference stack for teams serving LLMs on their own Kubernetes GPU clusters. It brings request routing and monitoring around vLLM, so applications can move from one serving instance to a distributed deployment without changing their code. It requires a GPU-enabled Kubernetes environment.