
Pixtral is a family of Mistral models for developers who want to run multimodal AI on their own hardware. The associated mistral-inference project is archived and no longer maintained. That status applies to the inference library.
The available models include Pixtral 12B, Pixtral 12B Base and Pixtral Large Instruct, with weights distributed through Hugging Face. The base and instruction variants give developers a choice of model within the same family. Mistral's library supports multimodal instruction following and lets applications run models through Python or a command-line interface.
Local inference uses downloaded model weights. The library requires a GPU, so this route is for people with GPU hardware rather than those looking for a CPU-only application. It's a developer library rather than a ready-made chat interface.
Pixtral 12B weights and the inference code use Apache 2.0. Pixtral Large Instruct uses the Mistral AI Research License, which restricts use to the purposes it authorizes; check its terms before commercial deployment. For server deployment, the repository also includes code for a vLLM serving image that uses Transformers in place of Mistral's reference implementation.
Claim this page with an email at mistral.ai. Pixtral gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Pixtral?Promote it
Something wrong or outdated on this page?
10.8KUpdated 3 months agoApache-2.0
Docker#Hugging Face integration#Multimodal input#Tool calling
Mistral Small and Large are downloadable language models for developers building chat, reasoning and tool-using applications on their own infrastructure. Capabilities and hardware requirements depend on the release. Mistral Small 3.1 adds image understanding to text generation, while Mistral Large 2 is a larger text model.
9.2KUpdated 6 months agoMIT
Docker#Hugging Face integration#Multilingual#Multimodal input
10.6KUpdated 2 years agoApache-2.0
Docker · Web#Hugging Face integration#Multimodal input
1.3KUpdated 7 months agoApache-2.0
Windows · Docker#Hugging Face integration#Multimodal input#OpenAI-compatible API
273Updated 2 years agoApache-2.0
Linux#Guardrails#Hugging Face integration#LM Studio integration
2.1KUpdated 3 weeks agoApache-2.0
Linux#GGUF#Guardrails#Hugging Face integration
dots.ocr is a self-hosted document parser that combines multilingual text recognition and page layout analysis in one vision-language model. It's for developers and teams converting PDFs or document images into structured text while running inference on their own hardware. The Python project is open source under the MIT license.
Grounding DINO finds objects in images using category names or descriptive phrases you supply. It's a local AI model for developers and computer vision researchers who need detection beyond a fixed set of labels, including people building dataset annotation tools.
JoyCaption is an open-weight image captioning model for people preparing datasets to train or fine-tune diffusion models. It runs on your own GPU and covers both SFW and NSFW images, including photography, anime, digital art and furry artwork. Automated captions reduce the need to write descriptions by hand or find images that already have usable text.
Granite is IBM's family of open-source AI models for developers and businesses that want to run and customize AI on their own hardware or servers. The language-model repository listed here is archived and no longer maintained. The broader family includes models for language, speech, document understanding and forecasting, released under Apache 2.0 for research and commercial use.
Nemotron is NVIDIA's family of AI models for developers building agents that reason, write code and call tools. You can run models locally for private, offline work or deploy them on your own servers. NVIDIA publishes model weights, training data and recipes so teams can inspect and adapt the models for their applications.