
Mistral Small and Large are downloadable language models for developers building chat, reasoning and tool-using applications on their own infrastructure. Capabilities and hardware requirements depend on the release. Mistral Small 3.1 adds image understanding to text generation, while Mistral Large 2 is a larger text model.
The official mistral-inference repository documents downloading weights, using the Python interface and running interactive chat from the command line. Its examples cover function calling and multimodal input for Mistral Small 3.1. A deployment image uses vLLM to serve the models. The inference library is archived and no longer maintained, so check a current serving framework before starting a new deployment.
Local inference needs enough GPU memory for the chosen weights and context length. Larger releases require considerably more resources than Small models; a shared model family name does not imply the same hardware requirements.
Licenses also differ by release. Mistral Small 3.1 weights and the inference code are Apache 2.0. The Mistral Large 2 weights listed in the repository use the non-commercial Mistral AI Research License. Check the exact model card before commercial deployment. Mistral also offers hosted inference through its API and cloud providers.
Claim this page with an email at mistral.ai. Mistral Small and Large gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Mistral Small and Large?Promote it
Something wrong or outdated on this page?
273Updated 2 years agoApache-2.0
Linux#Guardrails#Hugging Face integration#LM Studio integration
Granite is IBM's family of open-source AI models for developers and businesses that want to run and customize AI on their own hardware or servers. The language-model repository listed here is archived and no longer maintained. The broader family includes models for language, speech, document understanding and forecasting, released under Apache 2.0 for research and commercial use.
2.1KUpdated 3 weeks agoApache-2.0
Linux#GGUF#Guardrails#Hugging Face integration
7.7KUpdated 12 months ago
#Hugging Face integration#Multimodal input#Quantization
Llama is Meta's family of large language models for developers, researchers and businesses that want to run models on their own hardware or servers. Its downloadable weights let you build generative AI applications with local inference. Access requires license acceptance and approval, and the weights use custom licensing for research and commercial use.
5.8KUpdated 2 days agoApache-2.0
Android#LM Studio integration#LoRA#Multilingual
10.8KUpdated 3 months agoApache-2.0
Docker#Hugging Face integration#Multimodal input#Tool calling
9.2KUpdated 6 months agoMIT
Docker#Hugging Face integration#Multilingual#Multimodal input
Nemotron is NVIDIA's family of AI models for developers building agents that reason, write code and call tools. You can run models locally for private, offline work or deploy them on your own servers. NVIDIA publishes model weights, training data and recipes so teams can inspect and adapt the models for their applications.
Gemma is Google DeepMind’s family of open-weight AI models for developers building applications that can run on their own hardware. Its range covers compact models for phones and IoT devices alongside larger Gemma 4 models for reasoning on personal computers and servers. Some applications can work offline, keeping model inference on the device. Google AI Studio and Google Cloud are also available for hosted use.
Pixtral is a family of Mistral models for developers who want to run multimodal AI on their own hardware. The associated mistral-inference project is archived and no longer maintained. That status applies to the inference library.
dots.ocr is a self-hosted document parser that combines multilingual text recognition and page layout analysis in one vision-language model. It's for developers and teams converting PDFs or document images into structured text while running inference on their own hardware. The Python project is open source under the MIT license.