TabbyAPI is a self-hosted LLM API server built around ExLlamaV3, for people who want local model inference behind an OpenAI-compatible API. It's the official server for that backend. The project targets personal use and small groups, and its maintainers explicitly advise against using it for production workloads.
It supports Exl3 models and FP16/BF16 weights. A published Docker image runs with CUDA on NVIDIA GPUs, while parallel batching with paged attention supports NVIDIA Ampere GPUs and newer. The code is open source under AGPL-3.0.
The API covers chat completions and tool/function calling, so applications can use locally served models through the same API style they use for OpenAI. Embedding support is an optional stack, included in the latest-extras Docker image. For applications that need predictable output formats, it can constrain responses with JSON schemas, regular expressions, or EBNF grammars.
TabbyAPI can load and unload models and download them from Hugging Face. It supports draft models for speculative decoding and continuous batching for concurrent requests. Chat templates use Jinja2 and follow Hugging Face conventions, while an internal proxy lets the server override client sampling parameters. It also supports AI Horde.
Claim this page and we'll verify you by hand. TabbyAPI gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find TabbyAPI?Promote it
Something wrong or outdated on this page?
9.6KUpdated 1 day agoApache-2.0
macOS · Windows · Linux · Docker · Web#Batch processing#llama.cpp backend#Multimodal input
Xinference serves language, speech and multimodal models through a shared API on your own computer or servers. It's an open source platform under Apache 2.0 for developers and researchers who want to build applications around models they host. You can also deploy it on cloud infrastructure.
5.1KUpdated 1 week agoApache-2.0
macOS · Linux · Docker#Batch processing#Hugging Face integration#LLM tracing
2.9KUpdated 6 months agoMIT
macOS · Docker#Batch processing#Hugging Face integration#Multimodal input
8.4KUpdated 20 hours agoMIT
#Agent Skills#Batch processing#Guardrails
1.3KUpdated 1 day agoApache-2.0
Web#LoRA#Multimodal input#Ollama integration
1.9KUpdated 3 weeks agoAGPL-3.0
macOS · Windows · Linux · Docker#Batch processing#Distributed execution#Hugging Face integration
Text Embeddings Inference is a self-hosted server for developers who need text embeddings for search and retrieval applications. It serves models through a REST API on your own hardware and can run offline once model weights are downloaded. The Rust project is open source under Apache 2.0.
Infinity Embeddings is a self-hosted server for developers building semantic search and retrieval-augmented generation applications. It runs embedding and reranking models on your own hardware, with support for image and audio search alongside text. It's open source under MIT.
OGX, formerly Llama Stack, is a self-hosted AI application server for developers building chat apps, document search or AI agents. It brings model inference, file storage, vector search and agent orchestration into one process. You can run it on a laptop, in a datacenter or in the cloud. It's open source under MIT.
KubeAI is an open source Kubernetes operator for teams serving AI models on their own infrastructure or cloud clusters. It manages model servers and scales them with demand, including starting from zero running replicas. It uses the Apache 2.0 license and can run on CPUs, GPUs or TPUs, including in a local Kubernetes cluster.
Sonar is a self-hosted inference engine for developers and teams serving Hugging Face-compatible language and multimodal models on their own hardware. Based on vLLM, it adds model and quantization formats, sampling methods, and deployment features. It's open source under AGPL-3.0.