Favicon of RouteLLM

RouteLLM

A self-hosted LLM routing framework in Python that connects to Ollama and cloud models through LiteLLM. Open source under Apache 2.0.

RouteLLM is a self-hosted Python framework for developers who want to split requests between a stronger LLM and a cheaper model. It judges which prompts need the stronger model, so an application doesn't have to send every request to its most expensive provider. You control the cost-quality tradeoff through a routing threshold.

It includes a Python client compatible with OpenAI's client interface and an OpenAI-compatible server for existing applications. LiteLLM provides connections to local models through Ollama, cloud providers such as Anthropic and Gemini, and other OpenAI-compatible endpoints. The framework is open source under the Apache 2.0 license.

Local routing doesn't guarantee offline use. The mf and sw_ranking routers require an OpenAI API key and send embedding requests to OpenAI even when the models generating answers run locally. Requests routed to cloud models also go to their providers.

The supplied routers use preference data to choose between models, with approaches including matrix factorization and BERT classification. They were trained on GPT-4 and Mixtral 8x7B, but can route between other model pairs without retraining. Threshold calibration helps match routing decisions to the kinds of prompts an application receives.

For teams comparing routing strategies, the evaluation framework measures performance on MMLU, GSM8K and MT Bench and produces comparison plots. Developers can add their own routers and benchmarks.

Similar to RouteLLM