Favicon of Kimi K2

Kimi K2

A self-hosted language model series for coding and tool use, with Base and Instruct variants and a hosted OpenAI/Anthropic-compatible API.

Screenshot of Kimi K2 website

Kimi K2 is Moonshot AI's language model series for developers building coding assistants and AI agents, and researchers who want a foundation model to customize. You can run its checkpoints on your own infrastructure or use Moonshot's hosted API. Local inference runs on your hardware; the hosted API sends requests to Moonshot's service.

The two variants serve different needs. Kimi-K2-Base is intended for fine-tuning and custom applications. Kimi-K2-Instruct handles general chat, reasoning and tool use without a long thinking mode. It can decide when to call tools supplied by an application, so an agent can act on requests as well as generate text.

K2 uses a mixture-of-experts architecture: it activates part of the model for each token rather than using all its parameters at once. Its 128K-token context gives applications room for lengthy inputs. The model's evaluations cover repository fixes, multilingual coding and terminal tasks, alongside tool calling, mathematics and general knowledge.

For self-hosted use, the block-fp8 checkpoints are available on Hugging Face. Supported inference engines include vLLM, SGLang, KTransformers and TensorRT-LLM. Moonshot's hosted API supports OpenAI and Anthropic API conventions, which matters for applications already built around those interfaces.

K2 has one trillion total parameters with 32 billion activated per token, so self-hosting still requires substantial GPU infrastructure. Its weights use a modified MIT license with additional conditions.

Similar to Kimi K2