Open-Weight AI Models to Run Locally

Open LLMs such as Qwen3 and DeepSeek, plus models for code, images, video, speech and search, with weights you can download and run yourself.

Subcategories

100+ tools
An open-source text-to-video model you can run locally through ComfyUI or Python, with Apache 2.0 licensing and LoRA fine-tuning.

3.7KUpdated 11 months agoApache-2.0

Web#Hugging Face integration#LoRA

Local AI video generator with downloadable models, NVIDIA GPU support and LoRA fine-tuning. Code and the 2B model use Apache 2.0.

13KUpdated 11 months agoApache-2.0

Windows · Web#Hugging Face integration#LoRA#Multimodal input

A local text-to-image model for Chinese and English prompts, with Diffusers support, a community ComfyUI wrapper, and Apache 2.0 code.

1.1KUpdated 2 years agoApache-2.0

Web#Batch processing#Hugging Face integration#LoRA

An Apache 2.0 text-to-image model for local ComfyUI workflows. Its original repository is deprecated in favor of newer Chroma1 checkpoints.

huggingface.coImage Generation Models

An image generation model you can run locally, with text-to-image and image-to-image support, MIT-licensed Python code and separately licensed weights.

27.3KUpdated 9 months agoMIT

Web#Hugging Face integration#Image-to-image

Favicon of Qwen-Image

Qwen-Image

1 video
An open-source image generation and editing model for local deployment, with Chinese text rendering, multi-image edits and an Apache 2.0 license.

8.4KUpdated 8 months agoApache-2.0

Web#Image-to-image#LoRA#Multimodal input

A local AI image generation model with text-guided image editing, Diffusers support, and a reference implementation requiring at least 10GB of GPU VRAM.

73.5KUpdated 4 years ago

#Guardrails#Hugging Face integration#Image-to-image

A local vision-language model for image analysis, document extraction and video understanding, with Apache 2.0 weights and Hugging Face Transformers support.

huggingface.coOCR and Document Scanning

Linux#Batch processing#Hugging Face integration#Multimodal input

Multimodal AI models you can run locally with Mistral's GPU inference library, which is open source under Apache 2.0 and archived.

10.8KUpdated 3 months agoApache-2.0

Docker#Hugging Face integration#Multimodal input#Tool calling

Open-source vision language model for local image and text tasks, with Apache 2.0 licensing, Transformers support, and a small GPU memory footprint.

3.9KUpdated 1 week agoApache-2.0

#Hugging Face integration#LoRA#Multimodal input

Open-source coding LLMs for code generation, reasoning and fixes. Run the 32B instruction model locally with Transformers or self-host it with vLLM.

huggingface.coCoding Models

#Hugging Face integration

A self-hosted LLM with a 256K context window, vLLM and Transformers support, and research and commercial use under the Jamba Open Model License.

huggingface.coOpen-Weight LLMs

#Hugging Face integration#LoRA#Multilingual

Downloadable coding models for local completion and tool-using agents. Codestral and Devstral releases have different model licenses.

10.8KUpdated 3 months agoApache-2.0

Docker#Hugging Face integration#Tool calling

An open weights LLM you can self-host with Transformers, with document citations, tool calling and a CC-BY-NC-4.0 noncommercial license.

huggingface.coOpen-Weight LLMs

#Guardrails#Hugging Face integration#Multilingual

A family of AI models you can run offline with Ollama, llama.cpp or LM Studio, with open weights and training data for building specialized agents.

2.1KUpdated 3 weeks agoApache-2.0

Linux#GGUF#Guardrails#Hugging Face integration

LLM models for local or cloud inference, with downloadable weights and function calling. The Apache 2.0 inference library is archived.

10.8KUpdated 3 months agoApache-2.0

Docker#Hugging Face integration#Multimodal input#Tool calling

Open-source LLM for self-hosted AI agents, with thinking and direct-response modes, MIT licensing, and support for vLLM and SGLang.

huggingface.coCoding Models

#Hugging Face integration#LoRA#Multilingual

An open-source portrait animation tool that runs on Ubuntu with an NVIDIA GPU, turning a still image and English speech into a talking video.

8.7KUpdated 2 years agoMIT

Linux#Hugging Face integration#Multimodal input#ONNX

A local LLM family for chat, coding and multilingual tasks, with GGUF and Hugging Face formats for CPU or GPU use and support for llama.cpp and MLX.

127Updated 12 months ago

macOS#GGUF#Hugging Face integration#llama.cpp backend

Open-source text-to-speech software for local voice cloning and streaming speech generation, with Apache 2.0 licensing and NVIDIA GPU deployment.

23.8KUpdated 4 months agoApache-2.0

Linux · Docker · Web#Hugging Face integration#Multilingual#Streaming inference

Local vision-language models for image and video understanding, with Apache 2.0 code, mobile deployment and support for Ollama and llama.cpp.

26.5KUpdated 3 weeks agoApache-2.0

macOS · iOS · Android · Web#GGUF#Hugging Face integration#llama.cpp backend

A local image model with Chinese and English prompts, text rendering, ComfyUI support, Apache 2.0 code, and separate model-weight terms.

4.6KUpdated 2 years agoApache-2.0

Web#ControlNet#Hugging Face integration#Image-to-image

A self-hosted language model series for coding and tool use, with Base and Instruct variants and a hosted OpenAI/Anthropic-compatible API.

11.1KUpdated 11 months ago

#Hugging Face integration#Quantization#Tool calling

Vision-language models that answer image questions, read text in images and detect objects, with downloadable weights and GPU research code.

3.5KUpdated 1 year agoApache-2.0

#Hugging Face integration#Multimodal input

An open-source text-to-image model for local or self-hosted generation, with 4K output, PyTorch training code and Hugging Face Diffusers support.

1.9KUpdated 2 years agoApache-2.0

Web#Hugging Face integration

A self-hosted vision-language model project for image chat, with a local Gradio interface, GPU inference and Apache 2.0 code.

25KUpdated 2 years agoApache-2.0

macOS · Web#LoRA#Multimodal input#Quantization

A self-hosted language model for reasoning and agent tasks, with fast and slow thinking modes, a 256K context window, and Transformers and vLLM support.

820Updated 1 year ago

#Quantization#Tool calling

Language models for English, Korean and Spanish, with reasoning and tool use. Includes an on-device model and GGUF, GPTQ and AWQ formats.

107Updated 1 year ago

#GGUF#Hugging Face integration#Multilingual

Code completion models that run locally on CPU or GPUs, with Hugging Face Transformers support and LoRA fine-tuning.

2.1KUpdated 3 years agoApache-2.0

#Hugging Face integration#LoRA#Quantization

Open-weight coding models with downloadable weights, base and instruction-tuned variants, and support for project-level code completion.

7KUpdated 11 months agoMIT

#Hugging Face integration#Multilingual

An open-source local LLM for reasoning and math. Run it on your own hardware without cloud connectivity, or use Microsoft's hosted inference APIs.

3.9KUpdated 1 month agoMIT

#Hugging Face integration#Ollama integration#Works offline

An open-source AI model family for on-premises deployment, licensed under Apache 2.0, with language, speech, vision and guardrail models.

273Updated 2 years agoApache-2.0

Linux#Guardrails#Hugging Face integration#LM Studio integration

A bilingual local LLM family for English and Chinese, with chat and base models, version-specific licensing, and quantized variants for consumer GPUs.

7.8KUpdated 2 years agoApache-2.0

Docker#Hugging Face integration#llama.cpp backend#Multilingual

Open-source reasoning LLM for self-hosted use through vLLM or Transformers, with million-token context and function calling under Apache 2.0.

3.2KUpdated 1 year agoApache-2.0

#Hugging Face integration#Multilingual#Tool calling

An open-source LLM family with reasoning and chat models, Transformers support, and a long-context variant that accepts up to one million tokens.

7.3KUpdated 1 year agoApache-2.0

#Hugging Face integration#Tool calling#Web search

An open-source image and video generation framework with 4K text-to-image models, laptop GPU support, and ComfyUI and Diffusers integrations.

9.2KUpdated 2 weeks agoApache-2.0

#ControlNet#LoRA#Multimodal input