Favicon of Xinference

Xinference

Self-hosted AI model serving platform for Linux, Windows and macOS. Run language, speech and image models through an OpenAI-compatible API under Apache 2.0.

Screenshot of Xinference website

Xinference serves language, speech and multimodal models through a shared API on your own computer or servers. It's an open source platform under Apache 2.0 for developers and researchers who want to build applications around models they host. You can also deploy it on cloud infrastructure.

Its OpenAI-compatible API lets applications use self-hosted models while retaining a familiar interface, including function calling. A web interface provides model management and interaction, and you can register your own model weights alongside the built-in choices. Integrations include LangChain, LlamaIndex, Dify and Chatbox; Xagent connects model serving with agent planning and tool use.

The scope extends beyond chat. Xinference supports text embeddings and reranking for document search, image generation, speech transcription and text-to-speech. Multimodal models can process images and audio, and it also serves video generation models and traditional machine learning models.

Xinference runs on Linux, Windows and macOS, with CPU and GPU support, including Metal. Its inference engines include vLLM, llama.cpp and TensorRT, with GGML support as well. Docker deployment supports NVIDIA GPUs, and Kubernetes deployment supports GPU clusters. For workloads that outgrow one machine, distributed inference spans multiple workers. Automatic batching groups concurrent requests, and the vLLM integration can share its model context cache across replicas.

Similar to Xinference