
Text Generation Inference (TGI) is a self-hosted LLM server for developers and teams serving models through an API on their own hardware. The repository is archived; its README describes maintenance mode and recommends other inference engines for new deployments. Its focus is handling concurrent generation requests and making efficient use of GPU memory.
TGI runs locally or on your own server, with an official Docker container available. It's open source under Apache 2.0. Supported models include Llama, Falcon, StarCoder, BLOOM, GPT-NeoX and T5, along with fine-tuned models. Its Messages API provides OpenAI Chat Completion API compatible responses, so applications built around that interface can use a TGI backend.
Continuous batching combines incoming requests to improve throughput, while tensor parallelism spreads inference across multiple GPUs. Flash Attention and Paged Attention optimize supported model architectures. Quantization reduces VRAM needs through formats and methods including AWQ, GPTQ and bitsandbytes; TGI also loads Safetensors weights.
For chat applications, it streams tokens as they're generated. Structured output guidance constrains responses to predefined schemas for function calling and tool use. Generation controls include stop sequences and repetition penalties, and the API can return log probabilities.
TGI includes Prometheus metrics and OpenTelemetry tracing for teams monitoring a deployed service. Hugging Face uses it to power Hugging Chat, its Inference API and hosted Inference Endpoints; those hosted services run on Hugging Face infrastructure, while a self-hosted deployment runs inference on your server.
Claim this page and we'll verify you by hand. Text Generation Inference gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Text Generation Inference?Promote it
Something wrong or outdated on this page?
1.9KUpdated 3 weeks agoAGPL-3.0
macOS · Windows · Linux · Docker#Batch processing#Distributed execution#Hugging Face integration
Sonar is a self-hosted inference engine for developers and teams serving Hugging Face-compatible language and multimodal models on their own hardware. Based on vLLM, it adds model and quantization formats, sampling methods, and deployment features. It's open source under AGPL-3.0.
11.9KUpdated 4 days agoAGPL-3.0
macOS · Windows · Linux · Android · Docker · Web#GGUF#Hugging Face integration#llama.cpp backend
7.7KUpdated 5 days agoMIT
macOS · Windows · Linux · Docker · Web#Code execution#GGUF#Hugging Face integration
182KUpdated 16 hours agoMIT
macOS · Windows · Linux · Docker#GGUF#llama.cpp backend#Multimodal input
1.4KUpdated 2 days agoAGPL-3.0
Windows · Linux · Docker#Batch processing#Distributed execution#Hugging Face integration
5.1KUpdated 1 week agoApache-2.0
macOS · Linux · Docker#Batch processing#Hugging Face integration#LLM tracing
KoboldCpp pairs local model inference with a browser interface built for chat, creative writing and roleplay. A fork of llama.cpp, it bundles KoboldAI Lite with tools for keeping character details and story context alongside your conversations. It's open source under AGPL-3.0.
mistral.rs is an open source inference engine for running models on your own computer or self-hosted server. It's for developers building AI applications and people who want local chat, multimodal models and agent tools in the same runtime. The Rust project uses the MIT license.
Ollama runs language models on your own computer or server. It provides a command-line runner and a local API for people building AI applications or connecting existing tools to models they host themselves. The software is distributed under the MIT license.
TabbyAPI is a self-hosted LLM API server built around ExLlamaV3, for people who want local model inference behind an OpenAI-compatible API. It's the official server for that backend. The project targets personal use and small groups, and its maintainers explicitly advise against using it for production workloads.
Text Embeddings Inference is a self-hosted server for developers who need text embeddings for search and retrieval applications. It serves models through a REST API on your own hardware and can run offline once model weights are downloaded. The Rust project is open source under Apache 2.0.