Favicon of Text Generation Inference

Text Generation Inference

A self-hosted LLM inference server under Apache 2.0, with Docker deployment, multi-GPU support and an OpenAI-compatible chat API. The project is archived.

Screenshot of Text Generation Inference website

Text Generation Inference (TGI) is a self-hosted LLM server for developers and teams serving models through an API on their own hardware. The repository is archived; its README describes maintenance mode and recommends other inference engines for new deployments. Its focus is handling concurrent generation requests and making efficient use of GPU memory.

TGI runs locally or on your own server, with an official Docker container available. It's open source under Apache 2.0. Supported models include Llama, Falcon, StarCoder, BLOOM, GPT-NeoX and T5, along with fine-tuned models. Its Messages API provides OpenAI Chat Completion API compatible responses, so applications built around that interface can use a TGI backend.

Continuous batching combines incoming requests to improve throughput, while tensor parallelism spreads inference across multiple GPUs. Flash Attention and Paged Attention optimize supported model architectures. Quantization reduces VRAM needs through formats and methods including AWQ, GPTQ and bitsandbytes; TGI also loads Safetensors weights.

For chat applications, it streams tokens as they're generated. Structured output guidance constrains responses to predefined schemas for function calling and tool use. Generation controls include stop sequences and repetition penalties, and the API can return log probabilities.

TGI includes Prometheus metrics and OpenTelemetry tracing for teams monitoring a deployed service. Hugging Face uses it to power Hugging Chat, its Inference API and hosted Inference Endpoints; those hosted services run on Hugging Face infrastructure, while a self-hosted deployment runs inference on your server.

Similar to Text Generation Inference