RouteLLM is a self-hosted Python framework for developers who want to split requests between a stronger LLM and a cheaper model. It judges which prompts need the stronger model, so an application doesn't have to send every request to its most expensive provider. You control the cost-quality tradeoff through a routing threshold.
It includes a Python client compatible with OpenAI's client interface and an OpenAI-compatible server for existing applications. LiteLLM provides connections to local models through Ollama, cloud providers such as Anthropic and Gemini, and other OpenAI-compatible endpoints. The framework is open source under the Apache 2.0 license.
Local routing doesn't guarantee offline use. The mf and sw_ranking routers require an OpenAI API key and send embedding requests to OpenAI even when the models generating answers run locally. Requests routed to cloud models also go to their providers.
The supplied routers use preference data to choose between models, with approaches including matrix factorization and BERT classification. They were trained on GPT-4 and Mixtral 8x7B, but can route between other model pairs without retraining. Threshold calibration helps match routing decisions to the kinds of prompts an application receives.
For teams comparing routing strategies, the evaluation framework measures performance on MMLU, GSM8K and MT Bench and produces comparison plots. Developers can add their own routers and benchmarks.
Claim this page and we'll verify you by hand. RouteLLM gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find RouteLLM?Promote it
Something wrong or outdated on this page?
8.5KUpdated 23 hours agoApache-2.0
Web#LLM tracing#MCP#Multimodal input
Bifrost is a self-hosted AI gateway for developers and teams whose applications use multiple model providers. It puts Ollama, custom model deployments, and cloud services behind one OpenAI-compatible API, so applications can switch models without maintaining a separate integration for each provider.
59.9KUpdated 55 minutes ago
#Guardrails#MCP#Multi-user access
28.2KUpdated 1 day agoApache-2.0
Docker · Web#Batch processing#LLM tracing#MCP
11.7KUpdated 4 months agoApache-2.0
Docker · Web#Batch processing#LLM tracing#Multimodal input
20.3KUpdated 1 hour agoMIT
#Human approval#LLM tracing#MCP
9.4KUpdated 1 day agoApache-2.0
macOS · Windows · Linux · iOS · Android · Web#Batch processing#LLM tracing#Structured output
LiteLLM gives platform teams one place to manage access to LLMs across providers. Its self-hosted AI gateway puts cloud services and internal or locally hosted models behind an OpenAI-compatible API, so applications can change models without changing their integration. Developers can also use its Python SDK directly.
MLflow brings agent tracing, LLM evaluation, and model experiment tracking into a platform you can run locally or on your own servers. It's for developers and teams who need to understand failures, compare changes, and monitor AI applications in production. It's open source under Apache 2.0.
TensorZero is a self-hosted platform for developers building LLM applications. The project is archived and no longer maintained. It combines a model gateway with tools for inspecting responses, evaluating workflows, and improving prompts using production data and human feedback.
Pydantic AI is a Python SDK for developers building AI agents into their own applications. Its main draw is Pydantic validation across agent tools and results, so an agent can return structured data that application code can check and use. The SDK is MIT licensed.
BAML is a programming language for developers building AI agents, with typed model calls and local tracing built into the language. It runs standalone on macOS, Linux and Windows, or alongside an existing application. The language is open source under Apache 2.0, and works offline.