
MiniMax-M1 is an open-source reasoning model for developers building agents or working on complex software and mathematical problems. Its million-token context window makes it a candidate for tasks with long inputs that also need extended reasoning. You can serve the model on your own infrastructure through vLLM or use it through Transformers.
The model builds on MiniMax-Text-01 and combines a mixture-of-experts architecture with lightning attention. That attention design reduces computation during long generations compared with the original DeepSeek-R1 in the reported comparison. This matters for workloads where the model spends substantial time reasoning before it produces an answer.
M1 supports function calling: it can identify when a task needs an external function and produce structured arguments for that call. Developers can use this capability as part of an AI agent that interacts with tools. Its reinforcement-learning training covers mathematical problems and software engineering tasks in sandbox environments; general uses include summarization, translation, question answering and creative writing.
The weights use the Apache 2.0 license. Hugging Face hosts the MiniMax-M1-40k and MiniMax-M1-80k variants, which differ in their thinking budgets. vLLM is the recommended backend for production serving, while Transformers provides another deployment path.
Claim this page with an email at minimax.io. MiniMax-M1 gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find MiniMax-M1?Promote it
Something wrong or outdated on this page?
huggingface.coCoding Models
#Hugging Face integration#LoRA#Multilingual
GLM-4.5 is an open-source language model for developers building AI agents and coding tools on their own servers. It combines reasoning with tool calling and offers a choice between thinking mode for complex tasks and non-thinking mode for direct responses. The MIT license permits commercial use and modification.
273Updated 2 years agoApache-2.0
Linux#Guardrails#Hugging Face integration#LM Studio integration
11.1KUpdated 11 months ago
#Hugging Face integration#Quantization#Tool calling
Kimi K2 is Moonshot AI's language model series for developers building coding assistants and AI agents, and researchers who want a foundation model to customize. You can run its checkpoints on your own infrastructure or use Moonshot's hosted API. Local inference runs on your hardware; the hosted API sends requests to Moonshot's service.
huggingface.coOpen-Weight LLMs
#Guardrails#Hugging Face integration#Multilingual
107Updated 1 year ago
#GGUF#Hugging Face integration#Multilingual
EXAONE 4.0 is a family of language models from LG AI Research that combines general language tasks and complex problem solving in the same model. It's aimed at developers building multilingual AI applications, including on-device apps and agents that use tools. It supports English, Korean and Spanish.
3.9KUpdated 1 week agoApache-2.0
#Hugging Face integration#Multilingual#Tool calling
Granite is IBM's family of open-source AI models for developers and businesses that want to run and customize AI on their own hardware or servers. The language-model repository listed here is archived and no longer maintained. The broader family includes models for language, speech, document understanding and forecasting, released under Apache 2.0 for research and commercial use.
Command A is an open weights language model from Cohere and Cohere Labs for researchers and developers building self-hosted chatbots, document assistants and AI agents. Its focus is business tasks that combine multilingual text, supplied documents and external tools. You can run the model on your own hardware; Cohere also offers hosted chat through a playground and Hugging Face Space.
SmolLM3 is a 3B parameter language model from Hugging Face for developers and researchers who want to run an LLM on their own hardware. It comes as a base model and an instruction-tuned model for chat, reasoning and tool calling. Both run locally.