Favicon of Ray Serve

Ray Serve

A self-hosted model serving library for Python and LLM APIs. Run it on a laptop or cluster with request batching, streaming, and CPU or GPU resources.

Screenshot of Ray Serve website

Ray Serve is a self-hosted Python library for developers building inference APIs that combine models with application logic. It runs on a laptop, on-premise servers, Kubernetes, or cloud infrastructure you choose. It's open source under Apache 2.0.

Its main distinction is model composition: one application can connect several models, preprocessing steps, database queries, and response validation in Python. Those parts can use different machine learning frameworks and run on separate machines, with each stage scaling independently as traffic changes. This suits ML engineers and LLM developers whose services need several steps to produce an answer.

Ray Serve works with PyTorch, TensorFlow, Keras, Scikit-learn, and Hugging Face Transformers, as well as arbitrary Python code. FastAPI integration provides HTTP request parsing and validation. You're free to use models optimized with tools such as ONNXRuntime; Serve itself doesn't perform model-specific optimization.

For LLM applications, it supports streamed responses, dynamic request batching, and serving across multiple GPUs or machines. Built on Ray's distributed runtime, it can allocate CPU and GPU resources per model and share a GPU among deployments through fractional allocation.

The same serving application can move from local testing to a cluster with little or no code change. Its scope is serving: model lifecycle management and performance visualization require other tools.

Similar to Ray Serve