Local & Self-Hosted AI Tools

Browse local and self-hosted AI software, from chat apps and model servers to coding, image, video and voice tools.

800+ tools
A local LLM family for chat, coding and multilingual tasks, with GGUF and Hugging Face formats for CPU or GPU use and support for llama.cpp and MLX.

127Updated 12 months ago

macOS#GGUF#Hugging Face integration#llama.cpp backend

Open-source text-to-speech software for local voice cloning and streaming speech generation, with Apache 2.0 licensing and NVIDIA GPU deployment.

23.8KUpdated 4 months agoApache-2.0

Linux · Docker · Web#Hugging Face integration#Multilingual#Streaming inference

Local vision-language models for image and video understanding, with Apache 2.0 code, mobile deployment and support for Ollama and llama.cpp.

26.5KUpdated 3 weeks agoApache-2.0

macOS · iOS · Android · Web#GGUF#Hugging Face integration#llama.cpp backend

A local image model with Chinese and English prompts, text rendering, ComfyUI support, Apache 2.0 code, and separate model-weight terms.

4.6KUpdated 2 years agoApache-2.0

Web#ControlNet#Hugging Face integration#Image-to-image

A self-hosted language model series for coding and tool use, with Base and Instruct variants and a hosted OpenAI/Anthropic-compatible API.

11.1KUpdated 11 months ago

#Hugging Face integration#Quantization#Tool calling

Vision-language models that answer image questions, read text in images and detect objects, with downloadable weights and GPU research code.

3.5KUpdated 1 year agoApache-2.0

#Hugging Face integration#Multimodal input

An open-source text-to-image model for local or self-hosted generation, with 4K output, PyTorch training code and Hugging Face Diffusers support.

1.9KUpdated 2 years agoApache-2.0

Web#Hugging Face integration

A self-hosted vision-language model project for image chat, with a local Gradio interface, GPU inference and Apache 2.0 code.

25KUpdated 2 years agoApache-2.0

macOS · Web#LoRA#Multimodal input#Quantization

A self-hosted language model for reasoning and agent tasks, with fast and slow thinking modes, a 256K context window, and Transformers and vLLM support.

820Updated 1 year ago

#Quantization#Tool calling

Language models for English, Korean and Spanish, with reasoning and tool use. Includes an on-device model and GGUF, GPTQ and AWQ formats.

107Updated 1 year ago

#GGUF#Hugging Face integration#Multilingual

Code completion models that run locally on CPU or GPUs, with Hugging Face Transformers support and LoRA fine-tuning.

2.1KUpdated 3 years agoApache-2.0

#Hugging Face integration#LoRA#Quantization

Open-weight coding models with downloadable weights, base and instruction-tuned variants, and support for project-level code completion.

7KUpdated 11 months agoMIT

#Hugging Face integration#Multilingual

An open-source local LLM for reasoning and math. Run it on your own hardware without cloud connectivity, or use Microsoft's hosted inference APIs.

3.9KUpdated 1 month agoMIT

#Hugging Face integration#Ollama integration#Works offline

An open-source AI model family for on-premises deployment, licensed under Apache 2.0, with language, speech, vision and guardrail models.

273Updated 2 years agoApache-2.0

Linux#Guardrails#Hugging Face integration#LM Studio integration

A bilingual local LLM family for English and Chinese, with chat and base models, version-specific licensing, and quantized variants for consumer GPUs.

7.8KUpdated 2 years agoApache-2.0

Docker#Hugging Face integration#llama.cpp backend#Multilingual

Open-source reasoning LLM for self-hosted use through vLLM or Transformers, with million-token context and function calling under Apache 2.0.

3.2KUpdated 1 year agoApache-2.0

#Hugging Face integration#Multilingual#Tool calling

An open-source LLM family with reasoning and chat models, Transformers support, and a long-context variant that accepts up to one million tokens.

7.3KUpdated 1 year agoApache-2.0

#Hugging Face integration#Tool calling#Web search

An open-source image and video generation framework with 4K text-to-image models, laptop GPU support, and ComfyUI and Diffusers integrations.

9.2KUpdated 2 weeks agoApache-2.0

#ControlNet#LoRA#Multimodal input

An MIT-licensed image generation model you can run locally with CUDA, with full and distilled variants, Diffusers support, and a Gradio interface.

2.5KUpdated 1 year agoMIT

Web#Hugging Face integration

An open-source image and text model for local image classification without task-specific training. Runs through PyTorch on CPU or CUDA GPUs under MIT.

34.4KUpdated 6 months agoMIT

#Batch processing#Multimodal input

A local image captioning model with open weights, Apache 2.0 code, SFW and NSFW coverage, and support for ComfyUI and vLLM.

1.3KUpdated 7 months agoApache-2.0

Windows · Docker#Hugging Face integration#Multimodal input#OpenAI-compatible API

An open-source Python text-to-speech library that runs locally on CPU or CUDA GPUs and controls voice style through text descriptions.

5.6KUpdated 2 years agoApache-2.0

macOS#Hugging Face integration

Open-source local text-to-speech built on Qwen2.5, with Chinese and English voice cloning, adjustable voices, and an Apache 2.0 license.

11KUpdated 1 year agoApache-2.0

macOS · Windows · Linux · Web#Hugging Face integration#Multilingual#Voice cloning

An open-source object detection model that finds objects using text prompts. Run it locally with PyTorch on a CPU or NVIDIA GPU. Apache 2.0 licensed.

10.6KUpdated 2 years agoApache-2.0

Docker · Web#Hugging Face integration#Multimodal input

A free, self-hosted AI image editor for erasing objects and extending images. Runs on CPU, GPU or Apple Silicon. Apache 2.0; archived and unmaintained.

23.3KUpdated 1 year agoApache-2.0

macOS · Windows · Web#Batch processing#Image-to-image#Inpainting

A Python library for running and training CLIP image-text models on your own hardware, with local checkpoints and Hugging Face model support.

14.2KUpdated 5 days ago

#Hugging Face integration#Multimodal input

Open-source text-to-speech software generates custom voices locally on NVIDIA GPUs or Apple Silicon, with Docker support and an Apache 2.0 license.

14.9KUpdated 2 years agoApache-2.0

macOS · Windows · Docker#Streaming inference#Voice cloning

Open-source machine learning library in C/C++ with CPU, GPU, NPU and browser backends, quantization support, and an MIT license.

15.4KUpdated 6 days agoMIT

Web#Quantization

Local LLM acceleration library for Intel CPUs, GPUs and NPUs. Runs on Windows and Linux, integrates with Ollama and llama.cpp, and is archived.

8.9KUpdated 8 months agoApache-2.0

Windows · Linux · Docker#Distributed execution#GGUF#Hugging Face integration

A local LLM inference library for Windows and Linux with NVIDIA GPUs, MIT licensing, GPTQ and EXL2 support, and an OpenAI-compatible API through TabbyAPI.

4.6KUpdated 7 months agoMIT

Windows · Linux#Batch processing#Quantization#Speculative decoding

An open-source toolkit for building and evaluating LLM apps locally, with Python workflows, a VS Code designer and optional Azure AI collaboration.

11.2KUpdated 5 months agoMIT

VS Code#Code execution#LLM tracing#Visual workflows

A self-hosted LLM inference server under Apache 2.0, with Docker deployment, multi-GPU support and an OpenAI-compatible chat API. The project is archived.

10.9KUpdated 6 months agoApache-2.0

Linux · Docker#Batch processing#Distributed execution#Hugging Face integration

An open-source image generation model that runs locally, edits images, and uses multiple references. Supports Diffusers and carries an MIT license.

4.3KUpdated 10 months agoMIT

Web#Hugging Face integration#Image-to-image#LoRA

A local multimodal AI model for answering image questions and generating pictures, with downloadable weights and a Gradio interface.

17.8KUpdated 2 years agoMIT

Web#Batch processing#Hugging Face integration#Multimodal input

An open-source image generation framework for local use, with ComfyUI, Diffusers and fine-tuning support. Code uses the Apache 2.0 license.

1KUpdated 4 months agoApache-2.0

Linux · Web#Batch processing#Hugging Face integration#LoRA

Open-source text-to-speech built on Llama, with local inference, voice cloning and streaming audio. Uses Apache 2.0; Baseten offers cloud hosting.

6.3KUpdated 10 months agoApache-2.0

#Hugging Face integration#llama.cpp backend#LoRA

An open-source text-to-audio model that runs locally on CPU or NVIDIA GPU, with multilingual speech, voice presets and an MIT license.

39.3KUpdated 2 years agoMIT

#Hugging Face integration#Multilingual

An open-source ChatGPT alternative for Chrome that uses Ollama models locally to chat across tabs, analyze files, and assist with writing.

1.1KUpdated 7 months agoAGPL-3.0

Web · Browser Extension#Multilingual#Multimodal input#Ollama integration

Open-source text-to-speech model for local English dialogue generation, with voice cloning, NVIDIA GPU inference and an Apache 2.0 license.

19.4KUpdated 10 months agoApache-2.0

Docker · Web#Hugging Face integration#Multimodal input#Voice cloning

An open-source NotebookLM alternative with a local browser interface, MIT license, and cloud services through LlamaCloud, OpenAI and ElevenLabs.

2KUpdated 7 months agoMIT

macOS · Windows · Linux · Web#MCP#RAG

An open-source on-device AI framework for Android, iOS, desktop and web, with model conversion from PyTorch, TensorFlow and JAX.

3.5KUpdated 22 hours agoApache-2.0

macOS · Windows · Linux · iOS · Android · Web#Agent Skills#Hugging Face integration#Multimodal input

An AI inference framework for Android, iOS, desktop and browsers, with CPU and Vulkan GPU support and PyTorch and ONNX model conversion.

23.9KUpdated 6 days ago

macOS · Windows · Linux · iOS · Android · Web#ONNX#Quantization

An open-source model quantization library for local LLMs and vision models, with Apache 2.0 licensing and Hugging Face Transformers integration.

960Updated 7 months agoApache-2.0

#Hugging Face integration#LoRA#Quantization

A local LLM app for iPhone, iPad, and Mac that works offline after model download, keeps chats on-device, and connects to Siri and Apple Shortcuts.

privatellm.appAutomation and No-Code AI

macOS · iOS#Multilingual#Quantization#Works offline

Open-source PyTorch training framework for pretraining and fine-tuning models on your own CPUs or GPUs, with Apache 2.0 licensing.

31.4KUpdated 1 week agoApache-2.0

macOS#Distributed execution#ONNX

Local AI voice assistant for Linux and Windows with camera vision, persistent memory and MCP tools. Uses Ollama or cloud APIs. MIT licensed.

5.7KUpdated 2 weeks agoMIT

macOS · Windows · Linux · Docker#Home Assistant integration#MCP#Multi-agent workflows

Open-source LLM pretraining library built on PyTorch for custom datasets, with distributed NVIDIA GPU training. Licensed under Apache 2.0.

2.8KUpdated 1 week agoApache-2.0

Linux#Distributed execution#Hugging Face integration

A browser-based LLM chat app that connects to Ollama and OpenAI servers, stores chat data in your browser, and supports text, vision and reasoning models.

1.2KUpdated 12 months agoMIT

macOS · Windows · Linux · Docker · Web#Multimodal input#Ollama integration#OpenAI-compatible API