
Ray Serve is a self-hosted Python library for developers building inference APIs that combine models with application logic. It runs on a laptop, on-premise servers, Kubernetes, or cloud infrastructure you choose. It's open source under Apache 2.0.
Its main distinction is model composition: one application can connect several models, preprocessing steps, database queries, and response validation in Python. Those parts can use different machine learning frameworks and run on separate machines, with each stage scaling independently as traffic changes. This suits ML engineers and LLM developers whose services need several steps to produce an answer.
Ray Serve works with PyTorch, TensorFlow, Keras, Scikit-learn, and Hugging Face Transformers, as well as arbitrary Python code. FastAPI integration provides HTTP request parsing and validation. You're free to use models optimized with tools such as ONNXRuntime; Serve itself doesn't perform model-specific optimization.
For LLM applications, it supports streamed responses, dynamic request batching, and serving across multiple GPUs or machines. Built on Ray's distributed runtime, it can allocate CPU and GPU resources per model and share a GPU among deployments through fractional allocation.
The same serving application can move from local testing to a cluster with little or no code change. Its scope is serving: model lifecycle management and performance visualization require other tools.
Claim this page with an email at docs.ray.io. Ray Serve gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Ray Serve?Promote it
Something wrong or outdated on this page?
8.9KUpdated 3 weeks agoApache-2.0
Docker#Batch processing#ControlNet#Distributed execution
BentoML is a Python framework for developers turning AI models into services on their own hardware or servers. It supports self-hosted inference APIs and multi-model applications, with Apache 2.0 licensing. You can develop and debug locally, then deploy the services in Docker containers, on Kubernetes, or in your own cloud.
6KUpdated 22 hours agoApache-2.0
#Hugging Face integration#ONNX#OpenAI-compatible API
11KUpdated 1 week agoBSD-3-Clause
Windows · Linux · Docker#Batch processing#ONNX
5.1KUpdated 20 hours agoApache-2.0
#Batch processing#Distributed execution#LoRA
2.3KUpdated 1 day agoMPL-2.0
macOS · Windows · Linux · Docker#Agent Skills#Batch processing#Multi-user access
1KUpdated 7 days ago
#Distributed execution#Hugging Face integration#LoRA
Kaito manages self-hosted LLM inference, fine-tuning, and document retrieval services in a Kubernetes cluster. It's for teams that want to run models on infrastructure they control while reducing the work of sizing GPU resources and managing model deployments. The project is open source under Apache 2.0.
KServe is an open source platform for teams serving LLMs and predictive machine learning models on their own Kubernetes infrastructure. It puts both kinds of workloads under a common serving API, so teams can manage different model frameworks through the same platform. It uses the Apache 2.0 license.
Triton Inference Server, offered by NVIDIA as Dynamo-Triton, is a self-hosted AI inference server for teams deploying models in applications. It serves models from different frameworks through one server, with support for on-premises hardware, cloud infrastructure and edge devices. It's open source under the BSD-3-Clause license.
AIBrix is open-source infrastructure for teams serving large language models on their own Kubernetes clusters. It focuses on the work around inference: directing requests, scaling capacity and managing models across servers. Enterprise infrastructure teams can use its components to build a self-hosted model service. It's licensed under Apache 2.0.
dstack is a self-hosted orchestration tool for AI teams managing compute across GPU clouds and their own servers. It puts cluster management, training jobs and model inference behind one interface, so teams can use different providers and accelerators without maintaining a separate workflow for each environment. It's open source under the Mozilla Public License 2.0.