
Codestral and Devstral are Mistral coding models with downloadable weights for local or self-hosted development tools. Codestral focuses on code generation and fill-in-the-middle completion, where the model fills a gap between existing code. Devstral is designed for software engineering agents that explore a repository, use tools and edit multiple files.
Codestral releases include the 22B model and the 7B Codestral-Mamba model. The official mistral-inference library documents Python and command-line use, but that library is archived and no longer maintained. Its examples also cover instruction following and function calling.
Devstral Small 1.0 is a text-only 24B model developed with All Hands AI. Its model card recommends the OpenHands agent scaffold and documents local serving through vLLM, Transformers, LM Studio, llama.cpp and Ollama. Hardware needs depend on precision and context length; the card describes local deployment on an RTX 4090 or a Mac with 32GB RAM.
Licenses differ between releases. Codestral 22B uses the Mistral AI Non-Production License, which restricts production use. Devstral Small 1.0 uses Apache 2.0. The Apache license on the inference library does not determine the terms of every model. Check the exact weights you plan to deploy. Mistral also offers hosted API access; downloaded weights let you run inference on your own hardware.
Claim this page with an email at mistral.ai. Codestral and Devstral gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Codestral and Devstral?Promote it
Something wrong or outdated on this page?
huggingface.coCoding Models
#Hugging Face integration#LoRA#Multilingual
GLM-4.5 is an open-source language model for developers building AI agents and coding tools on their own servers. It combines reasoning with tool calling and offers a choice between thinking mode for complex tasks and non-thinking mode for direct responses. The MIT license permits commercial use and modification.
273Updated 2 years agoApache-2.0
Linux#Guardrails#Hugging Face integration#LM Studio integration
11.1KUpdated 11 months ago
#Hugging Face integration#Quantization#Tool calling
Kimi K2 is Moonshot AI's language model series for developers building coding assistants and AI agents, and researchers who want a foundation model to customize. You can run its checkpoints on your own infrastructure or use Moonshot's hosted API. Local inference runs on your hardware; the hosted API sends requests to Moonshot's service.
3.2KUpdated 1 year agoApache-2.0
#Hugging Face integration#Multilingual#Tool calling
7KUpdated 11 months agoMIT
#Hugging Face integration#Multilingual
DeepSeek-Coder is a family of code language models for developers who want downloadable models for their own coding tools. It generates code, completes existing code, and fills gaps within a file using the surrounding context. DeepSeek also offers a hosted chat interface; the downloadable weights are separate from that service.
huggingface.coCoding Models
#Hugging Face integration
Qwen2.5-Coder is a family of code-focused language models for developers who want to run a coding assistant on their own hardware or build a self-hosted code agent. It covers code generation, reasoning about code and fixing errors. The 32B instruction-tuned model is open source under Apache 2.0 and accepts requests in a chat format.
Granite is IBM's family of open-source AI models for developers and businesses that want to run and customize AI on their own hardware or servers. The language-model repository listed here is archived and no longer maintained. The broader family includes models for language, speech, document understanding and forecasting, released under Apache 2.0 for research and commercial use.
MiniMax-M1 is an open-source reasoning model for developers building agents or working on complex software and mathematical problems. Its million-token context window makes it a candidate for tasks with long inputs that also need extended reasoning. You can serve the model on your own infrastructure through vLLM or use it through Transformers.