
EXAONE 4.0 is a family of language models from LG AI Research that combines general language tasks and complex problem solving in the same model. It's aimed at developers building multilingual AI applications, including on-device apps and agents that use tools. It supports English, Korean and Spanish.
Both models have reasoning and non-reasoning modes. The non-reasoning mode handles general use, while reasoning mode lets the model work through harder problems. Tool use supports applications where an AI agent needs to call external functions as part of a task.
The family includes a 32B model aimed at higher performance and a smaller 1.2B model designed for on-device applications. Developers considering local AI can choose quantized models in GGUF, GPTQ or AWQ formats. The models also work with Hugging Face Transformers, and the quantized variants are available through Hugging Face.
Evaluations cover math, coding, instruction following and general knowledge. Korean evaluations also test practical and professional knowledge, a relevant distinction for applications that need Korean expertise alongside English or Spanish. The 32B model combines local and full attention to handle context.
The weights use the EXAONE AI Model License Agreement 1.2 - NC, which allows research and educational use and restricts commercial use and competing-model development.
Claim this page with an email at lgresearch.ai. EXAONE gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find EXAONE?Promote it
Something wrong or outdated on this page?
huggingface.coOpen-Weight LLMs
#Guardrails#Hugging Face integration#Multilingual
Command A is an open weights language model from Cohere and Cohere Labs for researchers and developers building self-hosted chatbots, document assistants and AI agents. Its focus is business tasks that combine multilingual text, supplied documents and external tools. You can run the model on your own hardware; Cohere also offers hosted chat through a playground and Hugging Face Space.
127Updated 12 months ago
macOS#GGUF#Hugging Face integration#llama.cpp backend
huggingface.coCoding Models
#Hugging Face integration#LoRA#Multilingual
20.4KUpdated 2 months agoApache-2.0
macOS · Linux#Hugging Face integration#LM Studio integration#Ollama integration
273Updated 2 years agoApache-2.0
Linux#Guardrails#Hugging Face integration#LM Studio integration
huggingface.coOpen-Weight LLMs
#Hugging Face integration#LoRA#Multilingual
Falcon-H1 is a family of language models from the Technology Innovation Institute for developers building local AI applications or serving models on their own hardware. Its hybrid design combines attention with Mamba2 state space models to reduce memory use and speed up inference. Model sizes cover smaller devices and larger server deployments.
GLM-4.5 is an open-source language model for developers building AI agents and coding tools on their own servers. It combines reasoning with tool calling and offers a choice between thinking mode for complex tasks and non-thinking mode for direct responses. The MIT license permits commercial use and modification.
gpt-oss is a pair of OpenAI reasoning models for developers who want to run a local LLM or host one on their own server. The models are open weight and licensed under Apache 2.0. OpenAI also has a hosted browser demo, separate from running the models on your hardware.
Granite is IBM's family of open-source AI models for developers and businesses that want to run and customize AI on their own hardware or servers. The language-model repository listed here is archived and no longer maintained. The broader family includes models for language, speech, document understanding and forecasting, released under Apache 2.0 for research and commercial use.
Jamba is AI21's language model family for teams building AI applications on their own servers. The documented Large 1.7 model combines Mamba state-space models with Transformer attention to process long context efficiently. A 256K-token context window makes it relevant for work that depends on lengthy documents, such as investment research, due diligence and reviewing procurement responses.