Seldon Core 2 is an AI model serving framework for teams running production machine learning and LLM applications on Kubernetes. It can run on your own infrastructure or in a cloud environment. Its focus is managing individual models and connected applications within the same deployment system.
Models can share inference servers, reducing the need for separate infrastructure for each model. Memory overcommit lets teams deploy a collection of models whose combined memory requirements exceed available memory, rather than reserve capacity for every unused model. Autoscaling applies to both models and application components, with built-in or custom scaling logic.
For applications with multiple stages, Seldon Core supports composable pipelines that use Kafka to stream data between components. Custom components can add application logic, LLMs, drift detection or outlier detection to those pipelines. This makes it relevant to teams whose serving needs extend beyond a single prediction endpoint.
It also supports model comparisons. A/B tests route traffic between candidate models or pipelines, while shadow deployments let teams evaluate candidates alongside an existing deployment.
Seldon Core 2 uses the Business Source License 1.1. It permits non-production use and a limited production grant for non-profit educational institutions; other production use needs commercial terms.
Claim this page and we'll verify you by hand. Seldon Core gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Seldon Core?Promote it
Something wrong or outdated on this page?
5.1KUpdated 20 hours agoApache-2.0
#Batch processing#Distributed execution#LoRA
AIBrix is open-source infrastructure for teams serving large language models on their own Kubernetes clusters. It focuses on the work around inference: directing requests, scaling capacity and managing models across servers. Enterprise infrastructure teams can use its components to build a self-hosted model service. It's licensed under Apache 2.0.
8.9KUpdated 3 weeks agoApache-2.0
Docker#Batch processing#ControlNet#Distributed execution
6.9KUpdated 2 days agoApache-2.0
Docker · Web#Code execution#Git integration#Multi-user access
2.3KUpdated 1 day agoMPL-2.0
macOS · Windows · Linux · Docker#Agent Skills#Batch processing#Multi-user access
1KUpdated 7 days ago
#Distributed execution#Hugging Face integration#LoRA
Kaito manages self-hosted LLM inference, fine-tuning, and document retrieval services in a Kubernetes cluster. It's for teams that want to run models on infrastructure they control while reducing the work of sizing GPU resources and managing model deployments. The project is open source under Apache 2.0.
6KUpdated 22 hours agoApache-2.0
#Hugging Face integration#ONNX#OpenAI-compatible API
BentoML is a Python framework for developers turning AI models into services on their own hardware or servers. It supports self-hosted inference APIs and multi-model applications, with Apache 2.0 licensing. You can develop and debug locally, then deploy the services in Docker containers, on Kubernetes, or in your own cloud.
ClearML is an MLOps suite for recording experiments, managing datasets and running ML workloads. Its Apache 2.0 Python SDK connects to a ClearML Server, available as a hosted service or open-source software you deploy yourself. ClearML Agent handles job orchestration and reproducibility.
dstack is a self-hosted orchestration tool for AI teams managing compute across GPU clouds and their own servers. It puts cluster management, training jobs and model inference behind one interface, so teams can use different providers and accelerators without maintaining a separate workflow for each environment. It's open source under the Mozilla Public License 2.0.
KServe is an open source platform for teams serving LLMs and predictive machine learning models on their own Kubernetes infrastructure. It puts both kinds of workloads under a common serving API, so teams can manage different model frameworks through the same platform. It uses the Apache 2.0 license.