Open-Weight AI Models to Run Locally

Open LLMs such as Qwen3 and DeepSeek, plus models for code, images, video, speech and search, with weights you can download and run yourself.

Subcategories

100+ tools
An MIT-licensed image generation model you can run locally with CUDA, with full and distilled variants, Diffusers support, and a Gradio interface.

2.5KUpdated 1 year agoMIT

Web#Hugging Face integration

An open-source image and text model for local image classification without task-specific training. Runs through PyTorch on CPU or CUDA GPUs under MIT.

34.4KUpdated 6 months agoMIT

#Batch processing#Multimodal input

A local image captioning model with open weights, Apache 2.0 code, SFW and NSFW coverage, and support for ComfyUI and vLLM.

1.3KUpdated 7 months agoApache-2.0

Windows · Docker#Hugging Face integration#Multimodal input#OpenAI-compatible API

An open-source Python text-to-speech library that runs locally on CPU or CUDA GPUs and controls voice style through text descriptions.

5.6KUpdated 2 years agoApache-2.0

macOS#Hugging Face integration

Open-source local text-to-speech built on Qwen2.5, with Chinese and English voice cloning, adjustable voices, and an Apache 2.0 license.

11KUpdated 1 year agoApache-2.0

macOS · Windows · Linux · Web#Hugging Face integration#Multilingual#Voice cloning

An open-source object detection model that finds objects using text prompts. Run it locally with PyTorch on a CPU or NVIDIA GPU. Apache 2.0 licensed.

10.6KUpdated 2 years agoApache-2.0

Docker · Web#Hugging Face integration#Multimodal input

A Python library for running and training CLIP image-text models on your own hardware, with local checkpoints and Hugging Face model support.

14.2KUpdated 5 days ago

#Hugging Face integration#Multimodal input

Open-source text-to-speech software generates custom voices locally on NVIDIA GPUs or Apple Silicon, with Docker support and an Apache 2.0 license.

14.9KUpdated 2 years agoApache-2.0

macOS · Windows · Docker#Streaming inference#Voice cloning

An open-source image generation model that runs locally, edits images, and uses multiple references. Supports Diffusers and carries an MIT license.

4.3KUpdated 10 months agoMIT

Web#Hugging Face integration#Image-to-image#LoRA

A local multimodal AI model for answering image questions and generating pictures, with downloadable weights and a Gradio interface.

17.8KUpdated 2 years agoMIT

Web#Batch processing#Hugging Face integration#Multimodal input

An open-source image generation framework for local use, with ComfyUI, Diffusers and fine-tuning support. Code uses the Apache 2.0 license.

1KUpdated 4 months agoApache-2.0

Linux · Web#Batch processing#Hugging Face integration#LoRA

Open-source text-to-speech built on Llama, with local inference, voice cloning and streaming audio. Uses Apache 2.0; Baseten offers cloud hosting.

6.3KUpdated 10 months agoApache-2.0

#Hugging Face integration#llama.cpp backend#LoRA

An open-source text-to-audio model that runs locally on CPU or NVIDIA GPU, with multilingual speech, voice presets and an MIT license.

39.3KUpdated 2 years agoMIT

#Hugging Face integration#Multilingual

Open-source text-to-speech model for local English dialogue generation, with voice cloning, NVIDIA GPU inference and an Apache 2.0 license.

19.4KUpdated 10 months agoApache-2.0

Docker · Web#Hugging Face integration#Multimodal input#Voice cloning

An open-source Python package for local text embeddings on CPU, with MIT licensing and integrations for Sentence Transformers and LangChain.

2.2KUpdated 1 day agoMIT

#Hugging Face integration#Multilingual

Local AI portrait animation software turns images and audio into talking-head videos, with editable facial landmarks and Apache 2.0 source code.

4.3KUpdated 6 months agoApache-2.0

Linux · Web#Hugging Face integration#Multimodal input

Local text-to-speech software generates speech from a reference voice, supports English and Chinese, and runs with NVIDIA GPUs. Code uses the MIT license.

15.3KUpdated 1 week agoMIT

Docker · Web#Multilingual#Voice cloning

A text-to-music model with melody conditioning and local GPU inference. AudioCraft code is MIT licensed; pretrained weights have a noncommercial license.

23.7KUpdated 2 years agoMIT

#Multimodal input

Favicon of Moonshine

Moonshine

3 videos
An on-device AI voice toolkit for speech recognition, intent recognition and text to speech, with support for desktop, mobile, browsers and Raspberry Pi.

11.2KUpdated 1 month ago

macOS · Windows · Linux · iOS · Android · Web#Multilingual#Streaming inference

Open-source speech recognition model for Mandarin, Cantonese, English, Japanese and Korean. Runs locally on CPU or GPU under the MIT license.

9.4KUpdated 3 weeks agoMIT

Docker#Batch processing#GGUF#Hugging Face integration

Open-source Python library for object detection and segmentation, with pretrained models and deployment exports. Licensed under Apache 2.0.

34.7KUpdated 2 days agoApache-2.0

An open-source AI music model and synthesis engine that runs locally on Apple Silicon, with a macOS app and AUv3 plugin for DAWs. Apache 2.0 licensed.

1.8KUpdated 2 months agoApache-2.0

macOS#MLX#Streaming inference

Face analysis toolkit for self-hosted recognition on CPU or NVIDIA GPU, local video face redaction, and commercially licensed models.

29.9KUpdated 3 weeks ago

macOS · Linux · iOS · Android · Web#Image-to-image#ONNX#Quantization

Open-source voice cloning uses a reference recording to generate multilingual speech with style control. The Python project uses the MIT license.

37.7KUpdated 1 year agoMIT

#Multilingual#Voice cloning

Local OCR model that extracts plain or formatted text from images, supports multi-page recognition, and works with Hugging Face Transformers.

8.2KUpdated 2 years ago

#Batch processing#GGUF#Hugging Face integration

Favicon of Parakeet

Parakeet

1 video
English speech-to-text model that runs locally through NVIDIA NeMo on Linux, with punctuation, word timestamps and CC BY 4.0 licensed weights.

18.5KUpdated 1 day agoApache-2.0

Linux · Docker#Batch processing#Hugging Face integration

Local OCR model that converts images and PDFs to text or Markdown. Runs on NVIDIA GPUs with vLLM or Transformers under the MIT license.

23.9KUpdated 8 months agoMIT

Linux#Batch processing#Hugging Face integration#Multimodal input

Local speech recognition models for English transcription, built on Whisper. Run on CPU or CUDA GPUs with Hugging Face Transformers under the MIT license.

4.1KUpdated 2 years agoMIT

#Batch processing#Hugging Face integration

An open-source voice AI framework that processes speech directly, with local inference on Mac and iPhone through MLX and self-hosted server backends.

11.2KUpdated 5 months agoApache-2.0

macOS · iOS · Web#Hugging Face integration#MLX#Quantization

Local AI lip-sync model for Windows and Linux that matches faces to supplied audio, with NVIDIA GPU support and an MIT-licensed codebase.

6.6KUpdated 1 year ago

Windows · Linux · Web#Batch processing#Inpainting#Multilingual

Favicon of Kokoro

Kokoro

6 videos
Open-source text-to-speech model and library for local speech generation, with multilingual voices, Apache 2.0 licensing and Apple Silicon GPU support.

9.1KUpdated 1 year agoApache-2.0

macOS · Windows#Batch processing#Multilingual#ONNX

A text-to-image model family you can deploy on your own infrastructure, with a consumer-hardware variant and separate API and web access.

1.5KUpdated 2 years agoMIT

#ControlNet#Hugging Face integration#Image-to-image

A vision-language model family with local Python code for training and evaluation, Apache 2.0 licensing, and variants built on OLMo and Qwen2.

937Updated 2 years agoApache-2.0

#Hugging Face integration#Multimodal input#Works offline

A language model family with downloadable weights, training data and code for people who want to run models locally or study how they're built.

6.7KUpdated 10 months agoApache-2.0

#Tool calling

An open-source vision model that runs locally with PyTorch and Hugging Face Transformers, supports CPU or CUDA GPUs, and uses the MIT license.

huggingface.coComputer Vision Models

#Hugging Face integration#Multimodal input#Structured output

Open source desktop app that turns images and motion prompts into video on Windows or Linux with an NVIDIA RTX GPU.

17.3KUpdated 11 months agoApache-2.0

Windows · Linux · Web#Hugging Face integration#Multimodal input