Favicon of BentoML

BentoML

An open-source Python framework for AI model serving. Build inference APIs and multi-model pipelines locally or deploy with Docker under Apache 2.0.

Screenshot of BentoML website

BentoML is a Python framework for developers turning AI models into services on their own hardware or servers. It supports self-hosted inference APIs and multi-model applications, with Apache 2.0 licensing. You can develop and debug locally, then deploy the services in Docker containers, on Kubernetes, or in your own cloud.

Its scope extends beyond LLMs. BentoML serves custom and fine-tuned models across text, image, audio, and embedding workloads. Supported frameworks and runtimes include PyTorch, Transformers, JAX, vLLM, SGLang, and TRT-LLM. Models named in its catalog include Llama, DeepSeek, Qwen, GPT-OSS, and Flux.

For applications that need several models, BentoML can compose them into pipelines for RAG and other workflows. It supports interactive requests, asynchronous tasks, and batch inference. Adaptive batching, model parallelism, and distributed serving help use CPU and GPU resources efficiently, including running large models across multiple GPUs. Its packaging brings code, models, and dependencies together for reproducible deployment.

The broader Bento platform adds deployment automation, performance monitoring, access controls, and autoscaling down to zero. It also supports rollbacks and canary, shadow, and A/B testing. The deployment choice matters: services can run on infrastructure you control, while the hosted BentoCloud offering supplies cloud compute with NVIDIA and AMD GPUs.

Similar to BentoML