Favicon of mistral.rs

mistral.rs

A local LLM inference engine with OpenAI and Anthropic-compatible APIs. Runs on macOS, Linux and Windows with CPU, CUDA or Apple Silicon support.

mistral.rs is an open source inference engine for running models on your own computer or self-hosted server. It's for developers building AI applications and people who want local chat, multimodal models and agent tools in the same runtime. The Rust project uses the MIT license.

It runs on macOS, Linux and Windows, with Metal acceleration on Apple Silicon, CUDA on NVIDIA GPUs, and CPU execution. Docker images support CPU and CUDA deployments. It can also spread inference across multiple GPUs or machines.

Supported models include Qwen3, Gemma 4 and Muse Glimmer 30B. The engine handles text, image, video and audio inputs, along with speech generation, image generation and embeddings. It loads Hugging Face checkpoints, GGUF files and UQFF quantizations, and can quantize models locally. Hardware-aware recommendations help match model memory use to the available devices.

A built-in browser chat interface displays reasoning, generated plots and files. OpenAI-compatible and Anthropic-compatible APIs let applications connect to the same server, while Python and Rust SDKs let developers embed inference directly in their apps.

The agent runtime can execute local Python and shell commands, call custom tools and connect to MCP servers. It also supports web search and reusable skill bundles. Python sessions retain state, and shell sessions have sandboxing and approval controls. LoRA adapter selection, continuous batching and prefix caching support customized models and shared serving workloads.

Similar to mistral.rs