
LMDeploy is an open-source toolkit for developers serving language and vision-language models on their own hardware. It combines model compression with inference and self-hosted APIs, so teams can use it for batch processing or as the model backend for an application. It uses the Apache 2.0 license.
Its two inference engines suit different development needs. TurboMind focuses on inference performance with optimized CUDA kernels; the Python-based PyTorch engine makes extending model support more accessible. NVIDIA GPU support includes GeForce RTX 50 series, and a separate LMDeploy-Jetson project supports offline deployment on NVIDIA Jetson hardware.
Quantization reduces the memory needed to serve models. LMDeploy supports AWQ, GPTQ and quantization of the attention cache, including combinations of AWQ, cache quantization and automatic prefix caching. Continuous batching lets the server process incoming requests together, while multi-GPU and multi-node serving support larger deployments. A request distributor routes work across models and machines.
Supported models include Llama, Qwen, InternLM, DeepSeek and gpt-oss. Vision-language support includes LLaVA, Qwen-VL and InternVL. The serving layer provides OpenAI-compatible APIs and Anthropic-compatible endpoints, with tool calling, reasoning outputs, structured output and LoRA serving. Batch inference runs locally; model downloads use Hugging Face by default, with ModelScope and openMind Hub as alternatives.
Claim this page and we'll verify you by hand. LMDeploy gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find LMDeploy?Promote it
Something wrong or outdated on this page?
7.7KUpdated 5 days agoMIT
macOS · Windows · Linux · Docker · Web#Code execution#GGUF#Hugging Face integration
mistral.rs is an open source inference engine for running models on your own computer or self-hosted server. It's for developers building AI applications and people who want local chat, multimodal models and agent tools in the same runtime. The Rust project uses the MIT license.
10.9KUpdated 23 hours agoApache-2.0
macOS · Windows · Linux#Hugging Face integration#Multimodal input#ONNX
1.9KUpdated 3 weeks agoAGPL-3.0
macOS · Windows · Linux · Docker#Batch processing#Distributed execution#Hugging Face integration
2.3KUpdated 1 year agoMIT
Linux#Batch processing#GGUF#Hugging Face integration
6.1KUpdated 5 days ago
macOS · iOS · Android#Hugging Face integration#Multimodal input#Quantization
Cactus is an on-device AI engine for developers building automation into mobile apps, wearables and embedded devices. Its Needle model handles tool calling locally, so a device can turn a request into an action without an internet connection. The focus is small devices, including smart home hardware, robots and microcontrollers.
5.1KUpdated 20 hours ago
macOS · Windows · Linux · iOS · Android · Web#MLX#Multimodal input#OpenAI-compatible API
ExecuTorch is PyTorch's runtime for developers building AI into mobile apps, desktop software and embedded devices. It runs models on the user's hardware, with support for Android, iOS, Linux, macOS and Windows, as well as microcontrollers. Developers can reuse a PyTorch model across targets, though hardware-specific deployments need their own exported model files.
OpenVINO is an Apache 2.0 licensed toolkit for developers who want to run AI models locally or serve them on their own infrastructure. It converts and optimizes models for inference, with support for x86 and ARM CPUs, Intel integrated and discrete GPUs, and Intel NPUs. Its runtime works on Linux, Windows and macOS.
Sonar is a self-hosted inference engine for developers and teams serving Hugging Face-compatible language and multimodal models on their own hardware. Based on vLLM, it adds model and quantization formats, sampling methods, and deployment features. It's open source under AGPL-3.0.
AutoAWQ is a Python library for developers who want to compress and run LLMs on their own hardware using 4-bit Activation-aware Weight Quantization (AWQ). The project is archived and no longer maintained. It's open source under the MIT license. It installs as a Python package, with optional kernel or Intel CPU dependencies.