Favicon of TabbyAPI

TabbyAPI

A self-hosted LLM API server that runs ExLlamaV3 models on your hardware, with OpenAI-compatible endpoints and an AGPL-3.0 license.

TabbyAPI is a self-hosted LLM API server built around ExLlamaV3, for people who want local model inference behind an OpenAI-compatible API. It's the official server for that backend. The project targets personal use and small groups, and its maintainers explicitly advise against using it for production workloads.

It supports Exl3 models and FP16/BF16 weights. A published Docker image runs with CUDA on NVIDIA GPUs, while parallel batching with paged attention supports NVIDIA Ampere GPUs and newer. The code is open source under AGPL-3.0.

The API covers chat completions and tool/function calling, so applications can use locally served models through the same API style they use for OpenAI. Embedding support is an optional stack, included in the latest-extras Docker image. For applications that need predictable output formats, it can constrain responses with JSON schemas, regular expressions, or EBNF grammars.

TabbyAPI can load and unload models and download them from Hugging Face. It supports draft models for speculative decoding and continuous batching for concurrent requests. Chat templates use Jinja2 and follow Hugging Face conventions, while an internal proxy lets the server override client sampling parameters. It also supports AI Horde.

Similar to TabbyAPI