
Kimi K2 is Moonshot AI's language model series for developers building coding assistants and AI agents, and researchers who want a foundation model to customize. You can run its checkpoints on your own infrastructure or use Moonshot's hosted API. Local inference runs on your hardware; the hosted API sends requests to Moonshot's service.
The two variants serve different needs. Kimi-K2-Base is intended for fine-tuning and custom applications. Kimi-K2-Instruct handles general chat, reasoning and tool use without a long thinking mode. It can decide when to call tools supplied by an application, so an agent can act on requests as well as generate text.
K2 uses a mixture-of-experts architecture: it activates part of the model for each token rather than using all its parameters at once. Its 128K-token context gives applications room for lengthy inputs. The model's evaluations cover repository fixes, multilingual coding and terminal tasks, alongside tool calling, mathematics and general knowledge.
For self-hosted use, the block-fp8 checkpoints are available on Hugging Face. Supported inference engines include vLLM, SGLang, KTransformers and TensorRT-LLM. Moonshot's hosted API supports OpenAI and Anthropic API conventions, which matters for applications already built around those interfaces.
K2 has one trillion total parameters with 32 billion activated per token, so self-hosting still requires substantial GPU infrastructure. Its weights use a modified MIT license with additional conditions.
Claim this page with an email at kimi.com. Kimi K2 gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Kimi K2?Promote it
Something wrong or outdated on this page?
huggingface.coCoding Models
#Hugging Face integration#LoRA#Multilingual
GLM-4.5 is an open-source language model for developers building AI agents and coding tools on their own servers. It combines reasoning with tool calling and offers a choice between thinking mode for complex tasks and non-thinking mode for direct responses. The MIT license permits commercial use and modification.
273Updated 2 years agoApache-2.0
Linux#Guardrails#Hugging Face integration#LM Studio integration
3.2KUpdated 1 year agoApache-2.0
#Hugging Face integration#Multilingual#Tool calling
107Updated 1 year ago
#GGUF#Hugging Face integration#Multilingual
EXAONE 4.0 is a family of language models from LG AI Research that combines general language tasks and complex problem solving in the same model. It's aimed at developers building multilingual AI applications, including on-device apps and agents that use tools. It supports English, Korean and Spanish.
20.4KUpdated 2 months agoApache-2.0
macOS · Linux#Hugging Face integration#LM Studio integration#Ollama integration
2.1KUpdated 3 weeks agoApache-2.0
Linux#GGUF#Guardrails#Hugging Face integration
Granite is IBM's family of open-source AI models for developers and businesses that want to run and customize AI on their own hardware or servers. The language-model repository listed here is archived and no longer maintained. The broader family includes models for language, speech, document understanding and forecasting, released under Apache 2.0 for research and commercial use.
MiniMax-M1 is an open-source reasoning model for developers building agents or working on complex software and mathematical problems. Its million-token context window makes it a candidate for tasks with long inputs that also need extended reasoning. You can serve the model on your own infrastructure through vLLM or use it through Transformers.
gpt-oss is a pair of OpenAI reasoning models for developers who want to run a local LLM or host one on their own server. The models are open weight and licensed under Apache 2.0. OpenAI also has a hosted browser demo, separate from running the models on your hardware.
Nemotron is NVIDIA's family of AI models for developers building agents that reason, write code and call tools. You can run models locally for private, offline work or deploy them on your own servers. NVIDIA publishes model weights, training data and recipes so teams can inspect and adapt the models for their applications.