
Xinference serves language, speech and multimodal models through a shared API on your own computer or servers. It's an open source platform under Apache 2.0 for developers and researchers who want to build applications around models they host. You can also deploy it on cloud infrastructure.
Its OpenAI-compatible API lets applications use self-hosted models while retaining a familiar interface, including function calling. A web interface provides model management and interaction, and you can register your own model weights alongside the built-in choices. Integrations include LangChain, LlamaIndex, Dify and Chatbox; Xagent connects model serving with agent planning and tool use.
The scope extends beyond chat. Xinference supports text embeddings and reranking for document search, image generation, speech transcription and text-to-speech. Multimodal models can process images and audio, and it also serves video generation models and traditional machine learning models.
Xinference runs on Linux, Windows and macOS, with CPU and GPU support, including Metal. Its inference engines include vLLM, llama.cpp and TensorRT, with GGML support as well. Docker deployment supports NVIDIA GPUs, and Kubernetes deployment supports GPU clusters. For workloads that outgrow one machine, distributed inference spans multiple workers. Automatic batching groups concurrent requests, and the vLLM integration can share its model context cache across replicas.
Claim this page and we'll verify you by hand. Xinference gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Xinference?Promote it
Something wrong or outdated on this page?
1.3KUpdated 1 day agoApache-2.0
Web#LoRA#Multimodal input#Ollama integration
KubeAI is an open source Kubernetes operator for teams serving AI models on their own infrastructure or cloud clusters. It manages model servers and scales them with demand, including starting from zero running replicas. It uses the Apache 2.0 license and can run on CPUs, GPUs or TPUs, including in a local Kubernetes cluster.
5.8KUpdated 1 day agoApache-2.0
macOS · Windows · Linux · iOS · Android · Docker#GGUF#Hugging Face integration#llama.cpp backend
3.7KUpdated 5 months agoMIT
Docker#OpenAI-compatible API#Streaming inference
1.4KUpdated 2 days agoAGPL-3.0
Windows · Linux · Docker#Batch processing#Distributed execution#Hugging Face integration
49.3KUpdated 2 hours agoMIT
macOS · Linux · Docker · Web#Code execution#Human approval#llama.cpp backend
18.5KUpdated 3 days agoApache-2.0
macOS · Windows · Linux · Docker · Web#Batch processing#MLX#Multilingual
Lemonade is an open source local AI server for people who want to use models on their own hardware or connect them to apps and agents. It handles chat, coding, image generation, speech, transcription, and embeddings. A built-in interface lets you use those capabilities directly, while its server makes them available to other software.
Speaches is a self-hosted speech server for developers who want transcription, translation and speech generation on their own hardware. Its OpenAI-compatible API lets applications use local speech models through tools and SDKs built for OpenAI's API. The project is open source under the MIT license.
TabbyAPI is a self-hosted LLM API server built around ExLlamaV3, for people who want local model inference behind an OpenAI-compatible API. It's the official server for that backend. The project targets personal use and small groups, and its maintainers explicitly advise against using it for production workloads.
LocalAI runs language models, speech, vision and image generation on hardware you control. It's for developers and teams that want a self-hosted AI server for their apps without sending model requests to a cloud service. Its OpenAI-compatible API works with existing clients, and it also accepts Anthropic, Ollama and ElevenLabs API calls.
VideoLingo is a self-hosted video translation app for creators and educators who need bilingual subtitles or dubbed versions of their videos. It brings transcription, translation and subtitle timing into one browser interface, with dubbing as an optional output. The project is open source under Apache 2.0; a separate hosted service offers subtitle translation and dubbing.