Tools to Run AI Models Locally

The software that loads a model and runs it on your machine: one-step apps like Ollama for a laptop, and servers like vLLM for a GPU box.

Subcategories

100+ tools
Favicon of Slotstream

Slotstream

1 video
Local LLM runner for Apple Silicon Macs that runs Qwen3.8-Flash-Next from SSD. Works offline after download and connects to coding agents and chat apps.

407Updated 2 days agoMIT

macOS#MLX#Multimodal input#OpenAI-compatible API

Favicon of Reef

Reef

1 video
Self-hosted AI agent infrastructure that learns from feedback, trains weights with Slime and SGLang, or improves prompts and skills without local training GPUs.

7.4KUpdated 12 hours agoApache-2.0

Linux#Agent Skills#OpenAI-compatible API#Prompt versioning

A browser-based local LLM tool that pools laptop, desktop and phone GPUs for chat and coding. Open source under MIT, with no account required.

544Updated 14 hours agoMIT

macOS · iOS · Web#Code execution#Distributed execution#Hugging Face integration

Favicon of Meetily

Meetily

6 videos
A local AI meeting assistant for macOS and Windows that records and transcribes calls offline, with Ollama or your own API key for summaries.

31.3KUpdated 3 weeks agoMIT

macOS · Windows · Linux#Ollama integration#OpenAI-compatible API#Streaming inference

Open-source diffusion engine in Python for local image, video and music generation, with low-VRAM inference and model training under Apache 2.0.

13.2KUpdated 2 days agoApache-2.0

#ControlNet#Image-to-image#Inpainting

Favicon of exo

exo

1 video
An open-source local LLM runner for macOS and Linux that splits models across devices and works offline with downloaded models. Apache 2.0 licensed.

47.7KUpdated 1 month agoApache-2.0

macOS · Linux · Web#Distributed execution#Hugging Face integration#MLX

Favicon of Jan

Jan

3 videos
A free desktop AI chat app for Windows, macOS and Linux that runs models locally or connects to GPT, Claude and other cloud models.

44.7KUpdated 5 hours ago

macOS · Windows · Linux#Hugging Face integration#MCP#OpenAI-compatible API

A self-hostable photo storage app with on-device AI, encrypted backups and sharing, plus apps for mobile, desktop and the web. Open source under AGPL-3.0.

29.2KUpdated 1 day agoAGPL-3.0

macOS · Windows · Linux · iOS · Android · Web#Multi-user access#Role-based access#Semantic search

Favicon of vLLM

vLLM

9 videos
An open source LLM serving engine that runs on your hardware, supports NVIDIA and AMD GPUs, and provides an OpenAI-compatible API.

93KUpdated 1 hour agoApache-2.0

macOS · Docker#Batch processing#Distributed execution#GGUF

A Python library for running diffusion models locally with PyTorch, including Stable Diffusion, LoRA adapters and Apple Silicon support. Apache 2.0 licensed.

34.6KUpdated 1 day agoApache-2.0

macOS#ControlNet#Hugging Face integration#Image-to-image

A local AI assistant for iPhone, iPad, Mac, and Vision Pro that runs models on device, searches your files, and can block external web connections.

noemaai.comChat With Your Documents

macOS · iOS#GGUF#MCP#MLX

Favicon of LM Studio

LM Studio

16 videos
Download local language models, chat with documents and connect apps to a local model API on macOS, Windows or Linux.

lmstudio.aiComputer and Browser Agents

macOS · Windows · Linux#llama.cpp backend#MCP#MLX

Favicon of llama.cpp

llama.cpp

11 videos
An open source local LLM engine for GGUF models, with CPU and GPU support, a built-in web UI, and an OpenAI-compatible server.

130KUpdated 40 minutes agoMIT

Web#Code execution#GGUF#Hugging Face integration

Favicon of Chatterbox

Chatterbox

3 videos
An open-source text-to-speech model family that runs on your own hardware, clones voices from short clips, and supports offline deployment.

26.6KUpdated 2 months agoMIT

Linux#Multilingual#Voice cloning#Voice conversion

Favicon of MLX

MLX

2 videos
An open-source machine learning array framework with NumPy-style APIs, shared CPU and GPU memory on Apple silicon, and Linux CPU and CUDA backends.

28.6KUpdated 24 hours agoMIT

macOS · Linux#Distributed execution#LoRA

An open-source speech-to-text engine that runs Whisper models locally on desktop and mobile, with CPU-only inference and GPU acceleration. MIT licensed.

54KUpdated 2 days agoMIT

macOS · Windows · Linux · iOS · Android · Docker#Hugging Face integration#Quantization#Streaming inference

Favicon of Ollama

Ollama

31 videos
Open-source local LLM runner for macOS, Windows, Linux and Docker, with optional cloud models and coding agent integrations.

182KUpdated 16 hours agoMIT

macOS · Windows · Linux · Docker#GGUF#llama.cpp backend#Multimodal input

Favicon of SGLang

SGLang

3 videos
An open-source inference framework for serving language and multimodal models on your own hardware, with an OpenAI-compatible API.

36.7KUpdated 1 hour agoApache-2.0

#Batch processing#Distributed execution#LoRA

An open-source Python library for running and training text, vision, audio and multimodal models locally, with Apache 2.0 licensing and PyTorch support.

166.8KUpdated 1 day agoApache-2.0

#Hugging Face integration#Multimodal input

A self-hosted AI runtime with an OpenAI-compatible API. It runs models on CPUs or GPUs and keeps inference on your own hardware.

49.3KUpdated 2 hours agoMIT

macOS · Linux · Docker · Web#Code execution#Human approval#llama.cpp backend

Favicon of Unsloth

Unsloth

6 videos
An open-source local LLM app for macOS, Windows and Linux. Run and train models, generate media, and connect coding agents to your hardware.

77KUpdated 21 hours agoApache-2.0

macOS · Windows · Linux · Docker · Web#Code execution#GGUF#Image-to-image

Favicon of Piper

Piper

1 video
A local text-to-speech engine with a web server, developer APIs and trainable voices. It's open source under GPL-3.0.

5.7KUpdated 2 days agoGPL-3.0

#ONNX

Run local models on your phone or computer for chat, document analysis, image generation and speech transcription. Connect to your own model server.

3.2KUpdated 4 days agoMIT

macOS · Windows · Linux · iOS · Android#GGUF#Human approval#LM Studio integration

Favicon of Locally AI

Locally AI

2 videos
Local AI assistant that runs language and vision models on Apple devices. Works offline after model download, with on-device voice and no account required.

locallyai.appDesktop Chat Apps

macOS · iOS#MLX#Multilingual#Multimodal input

Local AI chat app for macOS that runs GGUF models offline, searches files on-device, and also connects to Claude, ChatGPT, and OpenAI-compatible endpoints.

recurse.chatChat With Your Documents

macOS#GGUF#Hugging Face integration#OpenAI-compatible API

A paid audio transcription app that runs Whisper locally on macOS, iOS and visionOS, with multilingual transcription and subtitle exports.

sindresorhus.comOn-Device and In-Browser AI

macOS · iOS#Batch processing#Multilingual

A free, open-source browser translator for Chrome and Firefox with on-device translation, an offline dictionary, and custom backends such as Ollama.

1.1KUpdated 4 weeks agoBSD-3-Clause

Android · Browser Extension#Multilingual#Ollama integration#Works offline

A local LLM inference SDK for desktop, mobile and browsers. Runs on CPU or GPU, with offline inference and online Picovoice account validation.

318Updated 3 weeks agoApache-2.0

macOS · Windows · Linux · iOS · Android · Web#Quantization

Self-hosted LLM inference for Kubernetes with NVIDIA, AMD and Apple Silicon support, OpenAI-compatible APIs, and an Apache 2.0 license.

223Updated 16 hours agoApache-2.0

macOS · Linux#GGUF#Git integration#Guardrails

A community Helm chart that deploys Ollama on Kubernetes, with CPU or NVIDIA and AMD GPU support. Open source under the MIT license.

592Updated 5 days agoMIT

Docker#Ollama integration

An open-source computer vision API that runs offline on your hardware, with face recognition, object detection and support for custom models.

818Updated 4 years agoApache-2.0

macOS · Windows · Linux · Docker#Home Assistant integration#Works offline

A self-hosted AI server that gives applications a REST API for local image and text processing on Windows, macOS, Linux and Docker.

982Updated 1 year ago

macOS · Windows · Linux · Docker#Home Assistant integration#Image-to-image#Multimodal input

Self-hosted AI chat and document Q&A runs on Linux, macOS and Windows with local or cloud models. Apache 2.0 licensed; archived and no longer maintained.

12KUpdated 12 months agoApache-2.0

macOS · Windows · Linux · Docker · Web#Code execution#llama.cpp backend#Multi-user access

Python toolkit for semantic search and RAG with BGE embedding models, multilingual rerankers, fine-tuning and evaluation. MIT licensed.

12.2KUpdated 1 month agoMIT

#Multilingual#Multimodal input#Semantic search

A self-hosted text-to-speech server using Piper and Coqui XTTS v2, with voice cloning and an OpenAI-compatible API. Archived and no longer maintained.

857Updated 2 years agoAGPL-3.0

macOS · Windows · Linux · Docker#Multilingual#ONNX#OpenAI-compatible API

Local AI music generation library built on Stable Diffusion. Run it on your own hardware with CUDA, Apple Silicon or CPU. MIT licensed and no longer maintained.

3.9KUpdated 2 years agoMIT

macOS · Windows · Linux · Web#Hugging Face integration#Image-to-image#Multimodal input