Favicon of GLM-4.5

GLM-4.5

Open-source LLM for self-hosted AI agents, with thinking and direct-response modes, MIT licensing, and support for vLLM and SGLang.

Screenshot of GLM-4.5 website

GLM-4.5 is an open-source language model for developers building AI agents and coding tools on their own servers. It combines reasoning with tool calling and offers a choice between thinking mode for complex tasks and non-thinking mode for direct responses. The MIT license permits commercial use and modification.

The family includes the full GLM-4.5 model and the smaller GLM-4.5-Air. Both use a mixture-of-experts design, which activates only part of the model for each token. Air has lower hardware requirements, though the documented configurations still use server-class NVIDIA GPUs.

You can self-host the models through Transformers, vLLM or SGLang. The latter two support tool calling with OpenAI-style tool descriptions and can separate reasoning from the response. Z.ai and Zhipu AI also provide hosted API access; that route runs inference on their services rather than your hardware.

Hardware is a major consideration. The documented FP8 configuration for the full model uses eight H100 or four H200 GPUs, while Air uses two H100s or one H200. The documented server configurations also have substantial system memory requirements. Using the full 128K context needs additional GPU capacity.

The downloadable weights include base models and reasoning models in BF16, plus FP8 versions of the reasoning models. The models support English and Chinese, and developers can fine-tune them with Llama Factory or Swift.

Similar to GLM-4.5