Self-Hosted LLM Inference Servers

Model servers built for throughput, such as vLLM and SGLang, and exo for splitting one model across several machines.

63 tools
A self-hosted text-to-speech API for Kokoro-82M. Generate speech locally on CPU, NVIDIA GPU or Apple Silicon, with multi-speaker audio and captions.

5.5KUpdated 3 weeks agoApache-2.0

macOS · Windows · Linux · Docker · Web#Home Assistant integration#Multilingual#OpenAI-compatible API

Self-hosted AI inference server runs TensorRT, PyTorch and ONNX models on GPUs or CPUs, with dynamic batching and a BSD-3-Clause license.

11KUpdated 1 week agoBSD-3-Clause

Windows · Linux · Docker#Batch processing#ONNX

An open-source local LLM training framework with LoRA, multimodal support, and deployment through vLLM, SGLang or LMDeploy. Apache 2.0 licensed.

15.8KUpdated 11 hours agoApache-2.0

Web#Distributed execution#Hugging Face integration#LoRA

Python library for running GGUF models locally through llama.cpp, with a self-hosted OpenAI-compatible server and CPU or GPU support.

10.6KUpdated 1 week agoMIT

macOS · Windows · Linux#GGUF#Hugging Face integration#llama.cpp backend

An open-source local LLM runner for Linux, macOS and Windows via WSL2, with Podman or Docker isolation and llama.cpp or vLLM inference.

3.1KUpdated 1 day agoMIT

macOS · Windows · Linux · Docker#GGUF#Hugging Face integration#llama.cpp backend

Favicon of koboldcpp

koboldcpp

1 video
Local LLM runner for GGUF and GGML models on Windows, macOS and Linux, with CPU or GPU support, a browser UI and an AGPL-3.0 license.

11.9KUpdated 4 days agoAGPL-3.0

macOS · Windows · Linux · Android · Docker · Web#GGUF#Hugging Face integration#llama.cpp backend

An open-source Python framework for AI model serving. Build inference APIs and multi-model pipelines locally or deploy with Docker under Apache 2.0.

8.9KUpdated 3 weeks agoApache-2.0

Docker#Batch processing#ControlNet#Distributed execution

Open-source Python toolkit for training and serving LLMs on your own hardware, with Apache 2.0 licensing, quantization and multi-GPU support.

13.7KUpdated 3 weeks agoApache-2.0

#Hugging Face integration#LoRA#Quantization

An open-source local LLM compiler and deployment engine with GPU support across desktop, browser and mobile platforms, plus an OpenAI-compatible API.

23.2KUpdated 1 day agoApache-2.0

macOS · Windows · Linux · iOS · Android · Web#OpenAI-compatible API

A local LLM inference engine with OpenAI and Anthropic-compatible APIs. Runs on macOS, Linux and Windows with CPU, CUDA or Apple Silicon support.

7.7KUpdated 5 days agoMIT

macOS · Windows · Linux · Docker · Web#Code execution#GGUF#Hugging Face integration

A self-hosted text embedding server with a REST API, CPU and GPU support, and offline operation with downloaded model weights. Apache 2.0 licensed.

5.1KUpdated 1 week agoApache-2.0

macOS · Linux · Docker#Batch processing#Hugging Face integration#LLM tracing

Self-hosted computer vision server for images and video, with Docker support, NVIDIA GPU acceleration, and optional Roboflow hosted compute.

2.5KUpdated 1 day ago

macOS · Windows · Linux · Docker#Batch processing#Code execution#Multimodal input

Favicon of Lemonade

Lemonade

1 video
An open source local AI server for chat, image generation, and speech on Windows, macOS, and Linux, with APIs for apps and agents.

5.8KUpdated 6 hours agoApache-2.0

macOS · Windows · Linux · iOS · Android · Docker#GGUF#Hugging Face integration#llama.cpp backend

Open-source LLM serving toolkit for your own GPU servers, with quantization, text and vision models, and OpenAI-compatible APIs. Apache 2.0 licensed.

8.1KUpdated 3 days agoApache-2.0

#Batch processing#Distributed execution#Hugging Face integration

A self-hosted LLM inference library built on PyTorch for NVIDIA GPUs, with a Python API, OpenAI-compatible serving, and multi-node support.

14.8KUpdated 3 hours ago

Docker#Batch processing#Distributed execution#LoRA

Favicon of Speaches

Speaches

1 video
Self-hosted speech API for transcription, translation and speech generation. Runs via Docker on CPU or GPU with faster-whisper, Kokoro and Piper.

3.7KUpdated 5 months agoMIT

Docker#OpenAI-compatible API#Streaming inference

Self-hosted AI model serving platform for Linux, Windows and macOS. Run language, speech and image models through an OpenAI-compatible API under Apache 2.0.

9.6KUpdated 7 hours agoApache-2.0

macOS · Windows · Linux · Docker · Web#Batch processing#llama.cpp backend#Multimodal input

An open-source local LLM framework that splits work across CPUs and GPUs, with SGLang serving and LlamaFactory fine-tuning under Apache 2.0.

19.5KUpdated 11 hours agoApache-2.0

Docker#LoRA#Multimodal input#Prompt caching

Favicon of GPT4All

GPT4All

1 video
An open-source local AI chatbot for Windows, macOS and Linux. Run models without a GPU or cloud API, and chat privately with your documents.

77.4KUpdated 1 year agoMIT

macOS · Windows · Linux · Docker#GGUF#llama.cpp backend#OpenAI-compatible API

Self-hosted AI inference operator for Kubernetes with vLLM, Ollama and an OpenAI-compatible API. Runs on CPUs, GPUs or TPUs under Apache 2.0.

1.3KUpdated 2 days agoApache-2.0

Web#LoRA#Multimodal input#Ollama integration

Self-hosted text-to-speech server runs Chatterbox models on CPU or GPU, with voice cloning, audiobook generation and an OpenAI-compatible API.

1.5KUpdated 4 months agoMIT

macOS · Windows · Linux · Docker · Web#Hugging Face integration#Multilingual#OpenAI-compatible API

Self-hosted speech-to-text API that runs in Docker on CPU or CUDA GPUs, with Whisper, Faster Whisper and WhisperX. Open source under MIT.

3.3KUpdated 2 months agoMIT

Docker · Web#Multilingual#Speaker diarization#Voice activity detection

A self-hosted inference framework that coordinates NVIDIA GPU clusters with vLLM, SGLang or TensorRT-LLM and exposes an OpenAI-compatible API.

8.2KUpdated 1 hour ago

#Distributed execution#Multimodal input#OpenAI-compatible API

Favicon of OpenVINO

OpenVINO

2 videos
Open-source AI inference toolkit for local or self-hosted deployment on Linux, Windows and macOS, with CPU, Intel GPU and NPU support.

10.9KUpdated 1 day agoApache-2.0

macOS · Windows · Linux#Hugging Face integration#Multimodal input#ONNX

Favicon of KServe

KServe

3 videos
Self-hosted AI model serving platform for Kubernetes. Serve LLMs and predictive models with vLLM, Hugging Face support and an OpenAI-compatible API.

6.1KUpdated 6 hours agoApache-2.0

#Hugging Face integration#ONNX#OpenAI-compatible API

Self-hosted embedding and reranking API with MIT licensing, Hugging Face models, and CPU, NVIDIA, AMD and Apple MPS support.

2.9KUpdated 6 months agoMIT

macOS · Docker#Batch processing#Hugging Face integration#Multimodal input

Run AI models through Docker Desktop, Docker Engine or a standalone binary, with local inference and OpenAI and Ollama compatible APIs.

655Updated 3 days agoApache-2.0

macOS · Windows · Linux#GGUF#Hugging Face integration#llama.cpp backend

More in Run Models Locally