
GLM-4.5 is an open-source language model for developers building AI agents and coding tools on their own servers. It combines reasoning with tool calling and offers a choice between thinking mode for complex tasks and non-thinking mode for direct responses. The MIT license permits commercial use and modification.
The family includes the full GLM-4.5 model and the smaller GLM-4.5-Air. Both use a mixture-of-experts design, which activates only part of the model for each token. Air has lower hardware requirements, though the documented configurations still use server-class NVIDIA GPUs.
You can self-host the models through Transformers, vLLM or SGLang. The latter two support tool calling with OpenAI-style tool descriptions and can separate reasoning from the response. Z.ai and Zhipu AI also provide hosted API access; that route runs inference on their services rather than your hardware.
Hardware is a major consideration. The documented FP8 configuration for the full model uses eight H100 or four H200 GPUs, while Air uses two H100s or one H200. The documented server configurations also have substantial system memory requirements. Using the full 128K context needs additional GPU capacity.
The downloadable weights include base models and reasoning models in BF16, plus FP8 versions of the reasoning models. The models support English and Chinese, and developers can fine-tune them with Llama Factory or Swift.
Claim this page and we'll verify you by hand. GLM-4.5 gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find GLM-4.5?Promote it
Something wrong or outdated on this page?
273Updated 2 years agoApache-2.0
Linux#Guardrails#Hugging Face integration#LM Studio integration
Granite is IBM's family of open-source AI models for developers and businesses that want to run and customize AI on their own hardware or servers. The language-model repository listed here is archived and no longer maintained. The broader family includes models for language, speech, document understanding and forecasting, released under Apache 2.0 for research and commercial use.
11.1KUpdated 11 months ago
#Hugging Face integration#Quantization#Tool calling
Kimi K2 is Moonshot AI's language model series for developers building coding assistants and AI agents, and researchers who want a foundation model to customize. You can run its checkpoints on your own infrastructure or use Moonshot's hosted API. Local inference runs on your hardware; the hosted API sends requests to Moonshot's service.
3.2KUpdated 1 year agoApache-2.0
#Hugging Face integration#Multilingual#Tool calling
huggingface.coOpen-Weight LLMs
#Guardrails#Hugging Face integration#Multilingual
107Updated 1 year ago
#GGUF#Hugging Face integration#Multilingual
EXAONE 4.0 is a family of language models from LG AI Research that combines general language tasks and complex problem solving in the same model. It's aimed at developers building multilingual AI applications, including on-device apps and agents that use tools. It supports English, Korean and Spanish.
127Updated 12 months ago
macOS#GGUF#Hugging Face integration#llama.cpp backend
Falcon-H1 is a family of language models from the Technology Innovation Institute for developers building local AI applications or serving models on their own hardware. Its hybrid design combines attention with Mamba2 state space models to reduce memory use and speed up inference. Model sizes cover smaller devices and larger server deployments.
MiniMax-M1 is an open-source reasoning model for developers building agents or working on complex software and mathematical problems. Its million-token context window makes it a candidate for tasks with long inputs that also need extended reasoning. You can serve the model on your own infrastructure through vLLM or use it through Transformers.
Command A is an open weights language model from Cohere and Cohere Labs for researchers and developers building self-hosted chatbots, document assistants and AI agents. Its focus is business tasks that combine multilingual text, supplied documents and external tools. You can run the model on your own hardware; Cohere also offers hosted chat through a playground and Hugging Face Space.