
BentoML is a Python framework for developers turning AI models into services on their own hardware or servers. It supports self-hosted inference APIs and multi-model applications, with Apache 2.0 licensing. You can develop and debug locally, then deploy the services in Docker containers, on Kubernetes, or in your own cloud.
Its scope extends beyond LLMs. BentoML serves custom and fine-tuned models across text, image, audio, and embedding workloads. Supported frameworks and runtimes include PyTorch, Transformers, JAX, vLLM, SGLang, and TRT-LLM. Models named in its catalog include Llama, DeepSeek, Qwen, GPT-OSS, and Flux.
For applications that need several models, BentoML can compose them into pipelines for RAG and other workflows. It supports interactive requests, asynchronous tasks, and batch inference. Adaptive batching, model parallelism, and distributed serving help use CPU and GPU resources efficiently, including running large models across multiple GPUs. Its packaging brings code, models, and dependencies together for reproducible deployment.
The broader Bento platform adds deployment automation, performance monitoring, access controls, and autoscaling down to zero. It also supports rollbacks and canary, shadow, and A/B testing. The deployment choice matters: services can run on infrastructure you control, while the hosted BentoCloud offering supplies cloud compute with NVIDIA and AMD GPUs.
Claim this page with an email at bentoml.com. BentoML gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find BentoML?Promote it
Something wrong or outdated on this page?
2.3KUpdated 1 day agoMPL-2.0
macOS · Windows · Linux · Docker#Agent Skills#Batch processing#Multi-user access
dstack is a self-hosted orchestration tool for AI teams managing compute across GPU clouds and their own servers. It puts cluster management, training jobs and model inference behind one interface, so teams can use different providers and accelerators without maintaining a separate workflow for each environment. It's open source under the Mozilla Public License 2.0.
49.3KUpdated 2 hours agoMIT
macOS · Linux · Docker · Web#Code execution#Human approval#llama.cpp backend
12.5KUpdated 4 months agoApache-2.0
Docker · Web#Hugging Face integration#OpenAI-compatible API
11KUpdated 1 week agoBSD-3-Clause
Windows · Linux · Docker#Batch processing#ONNX
9.6KUpdated 1 day agoApache-2.0
macOS · Windows · Linux · Docker · Web#Batch processing#llama.cpp backend#Multimodal input
6.9KUpdated 2 days agoApache-2.0
Docker · Web#Code execution#Git integration#Multi-user access
LocalAI runs language models, speech, vision and image generation on hardware you control. It's for developers and teams that want a self-hosted AI server for their apps without sending model requests to a cloud service. Its OpenAI-compatible API works with existing clients, and it also accepts Anthropic, Ollama and ElevenLabs API calls.
OpenLLM is a self-hosted LLM server for developers who want to connect their applications to models running on their own hardware or servers. Its OpenAI-compatible API works with clients built for that interface, including the OpenAI Python client and LlamaIndex. The project is open source under the Apache License 2.0.
Triton Inference Server, offered by NVIDIA as Dynamo-Triton, is a self-hosted AI inference server for teams deploying models in applications. It serves models from different frameworks through one server, with support for on-premises hardware, cloud infrastructure and edge devices. It's open source under the BSD-3-Clause license.
Xinference serves language, speech and multimodal models through a shared API on your own computer or servers. It's an open source platform under Apache 2.0 for developers and researchers who want to build applications around models they host. You can also deploy it on cloud infrastructure.
ClearML is an MLOps suite for recording experiments, managing datasets and running ML workloads. Its Apache 2.0 Python SDK connects to a ClearML Server, available as a hosted service or open-source software you deploy yourself. ClearML Agent handles job orchestration and reproducibility.