mistral.rs is an open source inference engine for running models on your own computer or self-hosted server. It's for developers building AI applications and people who want local chat, multimodal models and agent tools in the same runtime. The Rust project uses the MIT license.
It runs on macOS, Linux and Windows, with Metal acceleration on Apple Silicon, CUDA on NVIDIA GPUs, and CPU execution. Docker images support CPU and CUDA deployments. It can also spread inference across multiple GPUs or machines.
Supported models include Qwen3, Gemma 4 and Muse Glimmer 30B. The engine handles text, image, video and audio inputs, along with speech generation, image generation and embeddings. It loads Hugging Face checkpoints, GGUF files and UQFF quantizations, and can quantize models locally. Hardware-aware recommendations help match model memory use to the available devices.
A built-in browser chat interface displays reasoning, generated plots and files. OpenAI-compatible and Anthropic-compatible APIs let applications connect to the same server, while Python and Rust SDKs let developers embed inference directly in their apps.
The agent runtime can execute local Python and shell commands, call custom tools and connect to MCP servers. It also supports web search and reusable skill bundles. Python sessions retain state, and shell sessions have sandboxing and approval controls. LoRA adapter selection, continuous batching and prefix caching support customized models and shared serving workloads.
Claim this page and we'll verify you by hand. mistral.rs gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find mistral.rs?Promote it
Something wrong or outdated on this page?
77.4KUpdated 1 year agoMIT
macOS · Windows · Linux · Docker#GGUF#llama.cpp backend#OpenAI-compatible API
GPT4All is a local AI chatbot for people who want to run language models on their own desktop or laptop and keep conversations on their machine. Its LocalDocs feature lets you ask questions about your own documents without sending them to a cloud service. It suits developers, teams and individuals who want control over their models and data.
qualcomm/GenieXInference Libraries and Bindings
macOS · Windows · Linux#GGUF#Hugging Face integration#llama.cpp backend
9.8KUpdated 5 months agoMIT
macOS · Windows · Linux#Batch processing#GGUF#Hugging Face integration
3.1KUpdated 1 day agoMIT
macOS · Windows · Linux · Docker#GGUF#Hugging Face integration#llama.cpp backend
10.9KUpdated 23 hours agoApache-2.0
macOS · Windows · Linux#Hugging Face integration#Multimodal input#ONNX
23.2KUpdated 23 hours agoApache-2.0
macOS · Windows · Linux · iOS · Android · Web#OpenAI-compatible API
MLC LLM is an open-source compiler and deployment engine for developers who want to run language models on their own hardware or inside apps. Its main distinction is the range of devices it targets: the same underlying engine, MLCEngine, serves desktop, browser and mobile deployments. The project uses the Apache 2.0 license.
Nexa SDK is an on-device AI inference framework for developers building applications that process text, images or audio on users' hardware. It runs models locally across CPUs, GPUs and NPUs, with a shared interface for different backends. Its scope includes language and vision models, speech recognition, speech synthesis and image generation.
PowerInfer is a local LLM inference engine for developers and researchers who want to run large models on a PC with a consumer GPU. It splits work between the CPU and GPU to reduce GPU memory demands and data transfers. The code is open source under the MIT license.
RamaLama runs and serves AI models on your own hardware using OCI containers. It's aimed at developers who want local chat or a self-hosted inference API with a container workflow they can also use in production. The project uses the MIT license.
OpenVINO is an Apache 2.0 licensed toolkit for developers who want to run AI models locally or serve them on their own infrastructure. It converts and optimizes models for inference, with support for x86 and ARM CPUs, Intel integrated and discrete GPUs, and Intel NPUs. Its runtime works on Linux, Windows and macOS.