Self-Hosted LLM Inference Servers

Model servers built for throughput, such as vLLM and SGLang, and exo for splitting one model across several machines.

62 tools
Favicon of Slotstream

Slotstream

1 video
Local LLM runner for Apple Silicon Macs that runs Qwen3.8-Flash-Next from SSD. Works offline after download and connects to coding agents and chat apps.

407Updated 2 days agoMIT

macOS#MLX#Multimodal input#OpenAI-compatible API

Favicon of Reef

Reef

1 video
Self-hosted AI agent infrastructure that learns from feedback, trains weights with Slime and SGLang, or improves prompts and skills without local training GPUs.

7.4KUpdated 12 hours agoApache-2.0

Linux#Agent Skills#OpenAI-compatible API#Prompt versioning

A browser-based local LLM tool that pools laptop, desktop and phone GPUs for chat and coding. Open source under MIT, with no account required.

544Updated 14 hours agoMIT

macOS · iOS · Web#Code execution#Distributed execution#Hugging Face integration

Favicon of exo

exo

1 video
An open-source local LLM runner for macOS and Linux that splits models across devices and works offline with downloaded models. Apache 2.0 licensed.

47.7KUpdated 1 month agoApache-2.0

macOS · Linux · Web#Distributed execution#Hugging Face integration#MLX

Favicon of vLLM

vLLM

9 videos
An open source LLM serving engine that runs on your hardware, supports NVIDIA and AMD GPUs, and provides an OpenAI-compatible API.

93KUpdated 1 hour agoApache-2.0

macOS · Docker#Batch processing#Distributed execution#GGUF

Favicon of llama.cpp

llama.cpp

11 videos
An open source local LLM engine for GGUF models, with CPU and GPU support, a built-in web UI, and an OpenAI-compatible server.

130KUpdated 40 minutes agoMIT

Web#Code execution#GGUF#Hugging Face integration

Favicon of Ollama

Ollama

31 videos
Open-source local LLM runner for macOS, Windows, Linux and Docker, with optional cloud models and coding agent integrations.

182KUpdated 16 hours agoMIT

macOS · Windows · Linux · Docker#GGUF#llama.cpp backend#Multimodal input

Favicon of SGLang

SGLang

3 videos
An open-source inference framework for serving language and multimodal models on your own hardware, with an OpenAI-compatible API.

36.7KUpdated 1 hour agoApache-2.0

#Batch processing#Distributed execution#LoRA

A self-hosted AI runtime with an OpenAI-compatible API. It runs models on CPUs or GPUs and keeps inference on your own hardware.

49.3KUpdated 2 hours agoMIT

macOS · Linux · Docker · Web#Code execution#Human approval#llama.cpp backend

Self-hosted LLM inference for Kubernetes with NVIDIA, AMD and Apple Silicon support, OpenAI-compatible APIs, and an Apache 2.0 license.

223Updated 16 hours agoApache-2.0

macOS · Linux#GGUF#Git integration#Guardrails

A community Helm chart that deploys Ollama on Kubernetes, with CPU or NVIDIA and AMD GPU support. Open source under the MIT license.

592Updated 5 days agoMIT

Docker#Ollama integration

An open-source computer vision API that runs offline on your hardware, with face recognition, object detection and support for custom models.

818Updated 4 years agoApache-2.0

macOS · Windows · Linux · Docker#Home Assistant integration#Works offline

A self-hosted AI server that gives applications a REST API for local image and text processing on Windows, macOS, Linux and Docker.

982Updated 1 year ago

macOS · Windows · Linux · Docker#Home Assistant integration#Image-to-image#Multimodal input

A self-hosted text-to-speech server using Piper and Coqui XTTS v2, with voice cloning and an OpenAI-compatible API. Archived and no longer maintained.

857Updated 2 years agoAGPL-3.0

macOS · Windows · Linux · Docker#Multilingual#ONNX#OpenAI-compatible API

Self-hosted AI application server with an OpenAI-compatible API, local Ollama and vLLM backends, document search and agent tool calling. MIT licensed.

8.4KUpdated 20 hours agoMIT

#Agent Skills#Batch processing#Guardrails

Self-hosted LLM inference engine for Hugging Face models, with OpenAI-compatible APIs, multimodal support, and CPU or GPU execution under AGPL-3.0.

1.9KUpdated 3 weeks agoAGPL-3.0

macOS · Windows · Linux · Docker#Batch processing#Distributed execution#Hugging Face integration

An on-device AI SDK that runs text, image and audio models on macOS, Windows and Linux, with GGUF, MLX and an OpenAI-compatible API.

qualcomm/GenieXInference Libraries and Bindings

macOS · Windows · Linux#GGUF#Hugging Face integration#llama.cpp backend

A self-hosted Kubernetes operator that deploys Hugging Face models with vLLM, manages GPU capacity, and runs fine-tuning and document retrieval services.

1KUpdated 7 days ago

#Distributed execution#Hugging Face integration#LoRA

A self-hosted AI serving framework that deploys models and pipelines on Kubernetes, on-premises or in the cloud, under the Business Source License.

4.8KUpdated 8 months ago

A self-hosted LLM inference server under Apache 2.0, with Docker deployment, multi-GPU support and an OpenAI-compatible chat API. The project is archived.

10.9KUpdated 6 months agoApache-2.0

Linux · Docker#Batch processing#Distributed execution#Hugging Face integration

A self-hosted LLM serving stack built on vLLM, with an OpenAI-compatible API, request routing and GPU cluster monitoring. Licensed under Apache 2.0.

2.6KUpdated 20 hours agoApache-2.0

Web#OpenAI-compatible API#Prompt caching

Open-source LLM serving infrastructure for Kubernetes with multi-node inference, demand-based autoscaling, LoRA management and vLLM integration.

5.1KUpdated 20 hours agoApache-2.0

#Batch processing#Distributed execution#LoRA

A free, open-source face recognition server that runs in Docker, uses FaceNet and InsightFace, and supports CPU and GPU processing.

8.3KUpdated 3 years agoApache-2.0

macOS · Windows · Linux · Docker · Web#Multi-user access#Role-based access

An open-source LLM server that runs locally or in the cloud, with OpenAI-compatible APIs, a chat UI, and an Apache 2.0 license.

12.5KUpdated 4 months agoApache-2.0

Docker · Web#Hugging Face integration#OpenAI-compatible API

A self-hosted model serving library for Python and LLM APIs. Run it on a laptop or cluster with request batching, streaming, and CPU or GPU resources.

44KUpdated 18 hours agoApache-2.0

#Batch processing#Hugging Face integration#ONNX

An open-source LLM training toolkit that turns documents into specialist datasets, with offline generation on macOS and Linux and optional cloud compute.

1.9KUpdated 3 months agoMIT

macOS · Windows · Linux#Distributed execution#llama.cpp backend#Quantization

Self-hostable MLOps software for experiment tracking, pipelines, dataset versioning and model serving, with an Apache 2.0 Python SDK.

6.9KUpdated 2 days agoApache-2.0

Docker · Web#Code execution#Git integration#Multi-user access

Face analysis toolkit for self-hosted recognition on CPU or NVIDIA GPU, local video face redaction, and commercially licensed models.

29.9KUpdated 3 weeks ago

macOS · Linux · iOS · Android · Web#Image-to-image#ONNX#Quantization

Self-hosted AI compute orchestration under MPL-2.0 for training and inference on GPU clouds, Kubernetes, VMs and bare-metal servers.

2.3KUpdated 1 day agoMPL-2.0

macOS · Windows · Linux · Docker#Agent Skills#Batch processing#Multi-user access

A local LLM inference engine for sparse models, with CPU and GPU support on Linux and Windows. Open source under MIT, with CPU-only support on Apple Silicon.

9.8KUpdated 5 months agoMIT

macOS · Windows · Linux#Batch processing#GGUF#Hugging Face integration

Open-source distributed LLM software runs inference and fine-tuning across shared GPUs, with public or private networks and support for Llama 3.1.

10.6KUpdated 2 years agoMIT

macOS · Windows · Linux · Docker#Hugging Face integration

An open-source voice AI framework that processes speech directly, with local inference on Mac and iPhone through MLX and self-hosted server backends.

11.2KUpdated 5 months agoApache-2.0

macOS · iOS · Web#Hugging Face integration#MLX#Quantization

A self-hosted LLM API server that runs ExLlamaV3 models on your hardware, with OpenAI-compatible endpoints and an AGPL-3.0 license.

1.4KUpdated 2 days agoAGPL-3.0

Windows · Linux · Docker#Batch processing#Distributed execution#Hugging Face integration

An open source local LLM runner that splits inference and memory across your computers. Runs on Linux, macOS and Windows under the MIT license.

3.1KUpdated 3 months agoMIT

macOS · Windows · Linux#Distributed execution#Hugging Face integration#Quantization

Favicon of llm-d

llm-d

1 video
An open-source LLM inference stack for self-hosted Kubernetes clusters, with vLLM and SGLang backends and support for GPUs, TPUs, XPUs and CPUs.

4.7KUpdated 23 hours agoApache-2.0

#Batch processing#Distributed execution#OpenAI-compatible API

A self-hosted text-to-speech API for Kokoro-82M. Generate speech locally on CPU, NVIDIA GPU or Apple Silicon, with multi-speaker audio and captions.

5.5KUpdated 3 weeks agoApache-2.0

macOS · Windows · Linux · Docker · Web#Home Assistant integration#Multilingual#OpenAI-compatible API

More in Run Models Locally