Favicon of Mistral Small and Large

Mistral Small and Large

LLM models for local or cloud inference, with downloadable weights and function calling. The Apache 2.0 inference library is archived.

Screenshot of Mistral Small and Large website

Mistral Small and Large are downloadable language models for developers building chat, reasoning and tool-using applications on their own infrastructure. Capabilities and hardware requirements depend on the release. Mistral Small 3.1 adds image understanding to text generation, while Mistral Large 2 is a larger text model.

The official mistral-inference repository documents downloading weights, using the Python interface and running interactive chat from the command line. Its examples cover function calling and multimodal input for Mistral Small 3.1. A deployment image uses vLLM to serve the models. The inference library is archived and no longer maintained, so check a current serving framework before starting a new deployment.

Local inference needs enough GPU memory for the chosen weights and context length. Larger releases require considerably more resources than Small models; a shared model family name does not imply the same hardware requirements.

Licenses also differ by release. Mistral Small 3.1 weights and the inference code are Apache 2.0. The Mistral Large 2 weights listed in the repository use the non-commercial Mistral AI Research License. Check the exact model card before commercial deployment. Mistral also offers hosted inference through its API and cloud providers.

Similar to Mistral Small and Large