Favicon of LMDeploy

LMDeploy

Open-source LLM serving toolkit for your own GPU servers, with quantization, text and vision models, and OpenAI-compatible APIs. Apache 2.0 licensed.

Screenshot of LMDeploy website

LMDeploy is an open-source toolkit for developers serving language and vision-language models on their own hardware. It combines model compression with inference and self-hosted APIs, so teams can use it for batch processing or as the model backend for an application. It uses the Apache 2.0 license.

Its two inference engines suit different development needs. TurboMind focuses on inference performance with optimized CUDA kernels; the Python-based PyTorch engine makes extending model support more accessible. NVIDIA GPU support includes GeForce RTX 50 series, and a separate LMDeploy-Jetson project supports offline deployment on NVIDIA Jetson hardware.

Quantization reduces the memory needed to serve models. LMDeploy supports AWQ, GPTQ and quantization of the attention cache, including combinations of AWQ, cache quantization and automatic prefix caching. Continuous batching lets the server process incoming requests together, while multi-GPU and multi-node serving support larger deployments. A request distributor routes work across models and machines.

Supported models include Llama, Qwen, InternLM, DeepSeek and gpt-oss. Vision-language support includes LLaVA, Qwen-VL and InternVL. The serving layer provides OpenAI-compatible APIs and Anthropic-compatible endpoints, with tool calling, reasoning outputs, structured output and LoRA serving. Batch inference runs locally; model downloads use Hugging Face by default, with ModelScope and openMind Hub as alternatives.

Similar to LMDeploy