Best Open-Weight LLMs to Run Locally

Language models you can download and run, such as Qwen3 and gpt-oss, for chat, reasoning and tool calling.

22 tools
An open-weight local LLM family from Google DeepMind for phones, PCs and servers, with Ollama and LM Studio support and an Apache 2.0 JAX library.

5.8KUpdated 2 days agoApache-2.0

Android#LM Studio integration#LoRA#Multilingual

Favicon of Qwen3

Qwen3

1 video
A language model family with public weights for local CPU or GPU use through Ollama, llama.cpp, and LM Studio, plus server deployment.

27.7KUpdated 9 months ago

#Batch processing#GGUF#Hugging Face integration

An open-source audio AI model for local transcription, translation and Q&A. Runs offline with vLLM or Transformers under Apache 2.0.

10.8KUpdated 3 months agoApache-2.0

#Hugging Face integration#Multilingual#Multimodal input

A self-hosted LLM with a 256K context window, vLLM and Transformers support, and research and commercial use under the Jamba Open Model License.

huggingface.coOpen-Weight LLMs

#Hugging Face integration#LoRA#Multilingual

An open weights LLM you can self-host with Transformers, with document citations, tool calling and a CC-BY-NC-4.0 noncommercial license.

huggingface.coOpen-Weight LLMs

#Guardrails#Hugging Face integration#Multilingual

A family of AI models you can run offline with Ollama, llama.cpp or LM Studio, with open weights and training data for building specialized agents.

2.1KUpdated 3 weeks agoApache-2.0

Linux#GGUF#Guardrails#Hugging Face integration

LLM models for local or cloud inference, with downloadable weights and function calling. The Apache 2.0 inference library is archived.

10.8KUpdated 3 months agoApache-2.0

Docker#Hugging Face integration#Multimodal input#Tool calling

Open-source LLM for self-hosted AI agents, with thinking and direct-response modes, MIT licensing, and support for vLLM and SGLang.

huggingface.coCoding Models

#Hugging Face integration#LoRA#Multilingual

A local LLM family for chat, coding and multilingual tasks, with GGUF and Hugging Face formats for CPU or GPU use and support for llama.cpp and MLX.

127Updated 12 months ago

macOS#GGUF#Hugging Face integration#llama.cpp backend

A self-hosted language model series for coding and tool use, with Base and Instruct variants and a hosted OpenAI/Anthropic-compatible API.

11.1KUpdated 11 months ago

#Hugging Face integration#Quantization#Tool calling

A self-hosted language model for reasoning and agent tasks, with fast and slow thinking modes, a 256K context window, and Transformers and vLLM support.

820Updated 1 year ago

#Quantization#Tool calling

Language models for English, Korean and Spanish, with reasoning and tool use. Includes an on-device model and GGUF, GPTQ and AWQ formats.

107Updated 1 year ago

#GGUF#Hugging Face integration#Multilingual

An open-source local LLM for reasoning and math. Run it on your own hardware without cloud connectivity, or use Microsoft's hosted inference APIs.

3.9KUpdated 1 month agoMIT

#Hugging Face integration#Ollama integration#Works offline

An open-source AI model family for on-premises deployment, licensed under Apache 2.0, with language, speech, vision and guardrail models.

273Updated 2 years agoApache-2.0

Linux#Guardrails#Hugging Face integration#LM Studio integration

A bilingual local LLM family for English and Chinese, with chat and base models, version-specific licensing, and quantized variants for consumer GPUs.

7.8KUpdated 2 years agoApache-2.0

Docker#Hugging Face integration#llama.cpp backend#Multilingual

Open-source reasoning LLM for self-hosted use through vLLM or Transformers, with million-token context and function calling under Apache 2.0.

3.2KUpdated 1 year agoApache-2.0

#Hugging Face integration#Multilingual#Tool calling

An open-source LLM family with reasoning and chat models, Transformers support, and a long-context variant that accepts up to one million tokens.

7.3KUpdated 1 year agoApache-2.0

#Hugging Face integration#Tool calling#Web search

A language model family with downloadable weights, training data and code for people who want to run models locally or study how they're built.

6.7KUpdated 10 months agoApache-2.0

#Tool calling

A local LLM family for developers and researchers, with downloadable weights, text and vision models, and custom licensing for research and commercial use.

7.7KUpdated 12 months ago

#Hugging Face integration#Multimodal input#Quantization

DeepSeek-V3 and R1 are downloadable language-model families for local text generation and reasoning.

92KUpdated 1 year agoMIT

#Hugging Face integration

An open source 3B local LLM that runs on CPU or CUDA GPU and supports direct or reasoning responses, six languages and a 128k context window.

3.9KUpdated 1 week agoApache-2.0

#Hugging Face integration#Multilingual#Tool calling

Open-weight local LLMs under Apache 2.0, with 20B and 120B models that work with Ollama, LM Studio and vLLM.

20.4KUpdated 2 months agoApache-2.0

macOS · Linux#Hugging Face integration#LM Studio integration#Ollama integration

More in Open Models