Apps to Run LLMs Locally

Apps like Ollama, Jan and GPT4All download an open model and run it in one step. Most of them also serve an OpenAI-compatible API.

35 tools
Favicon of Slotstream

Slotstream

1 video
Local LLM runner for Apple Silicon Macs that runs Qwen3.8-Flash-Next from SSD. Works offline after download and connects to coding agents and chat apps.

407Updated 2 days agoMIT

macOS#MLX#Multimodal input#OpenAI-compatible API

A browser-based local LLM tool that pools laptop, desktop and phone GPUs for chat and coding. Open source under MIT, with no account required.

544Updated 14 hours agoMIT

macOS · iOS · Web#Code execution#Distributed execution#Hugging Face integration

Favicon of exo

exo

1 video
An open-source local LLM runner for macOS and Linux that splits models across devices and works offline with downloaded models. Apache 2.0 licensed.

47.7KUpdated 1 month agoApache-2.0

macOS · Linux · Web#Distributed execution#Hugging Face integration#MLX

Favicon of Jan

Jan

3 videos
A free desktop AI chat app for Windows, macOS and Linux that runs models locally or connects to GPT, Claude and other cloud models.

44.7KUpdated 5 hours ago

macOS · Windows · Linux#Hugging Face integration#MCP#OpenAI-compatible API

Favicon of LM Studio

LM Studio

16 videos
Download local language models, chat with documents and connect apps to a local model API on macOS, Windows or Linux.

lmstudio.aiComputer and Browser Agents

macOS · Windows · Linux#llama.cpp backend#MCP#MLX

Favicon of llama.cpp

llama.cpp

11 videos
An open source local LLM engine for GGUF models, with CPU and GPU support, a built-in web UI, and an OpenAI-compatible server.

130KUpdated 40 minutes agoMIT

Web#Code execution#GGUF#Hugging Face integration

Favicon of Ollama

Ollama

31 videos
Open-source local LLM runner for macOS, Windows, Linux and Docker, with optional cloud models and coding agent integrations.

182KUpdated 16 hours agoMIT

macOS · Windows · Linux · Docker#GGUF#llama.cpp backend#Multimodal input

A self-hosted AI runtime with an OpenAI-compatible API. It runs models on CPUs or GPUs and keeps inference on your own hardware.

49.3KUpdated 2 hours agoMIT

macOS · Linux · Docker · Web#Code execution#Human approval#llama.cpp backend

Favicon of Unsloth

Unsloth

6 videos
An open-source local LLM app for macOS, Windows and Linux. Run and train models, generate media, and connect coding agents to your hardware.

77KUpdated 21 hours agoApache-2.0

macOS · Windows · Linux · Docker · Web#Code execution#GGUF#Image-to-image

Favicon of Locally AI

Locally AI

2 videos
Local AI assistant that runs language and vision models on Apple devices. Works offline after model download, with on-device voice and no account required.

locallyai.appDesktop Chat Apps

macOS · iOS#MLX#Multilingual#Multimodal input

Self-hosted AI chat and document Q&A runs on Linux, macOS and Windows with local or cloud models. Apache 2.0 licensed; archived and no longer maintained.

12KUpdated 12 months agoApache-2.0

macOS · Windows · Linux · Docker · Web#Code execution#llama.cpp backend#Multi-user access

Desktop AI app powered by llama.cpp, with downloads for Windows and Intel or Apple Silicon Macs. The desktop app is deprecated and unsupported.

desktop.backyard.aiDesktop Chat Apps

macOS · Windows#llama.cpp backend

A self-hosted ChatGPT alternative that runs Llama 2 and Code Llama locally, with an MIT license and an OpenAI-compatible API.

10.9KUpdated 3 years agoMIT

macOS · Docker · Web#GGUF#llama.cpp backend#OpenAI-compatible API

Local LLM web interface for Windows, macOS and Linux. Use GGUF models, Ollama or cloud APIs, with local chat storage and an Apache 2.0 license.

4.8KUpdated 3 weeks agoApache-2.0

macOS · Windows · Linux · Docker · Web#GGUF#Hugging Face integration#llama.cpp backend

An on-device AI SDK that runs text, image and audio models on macOS, Windows and Linux, with GGUF, MLX and an OpenAI-compatible API.

qualcomm/GenieXInference Libraries and Bindings

macOS · Windows · Linux#GGUF#Hugging Face integration#llama.cpp backend

A local LLM app that runs models offline on iOS and macOS, supports text and vision models, and uses ggml and llama.cpp under the MIT license.

2.1KUpdated 8 months agoMIT

macOS · iOS#llama.cpp backend#Multimodal input#RAG

An open-source LLM server that runs locally or in the cloud, with OpenAI-compatible APIs, a chat UI, and an Apache 2.0 license.

12.5KUpdated 4 months agoApache-2.0

Docker · Web#Hugging Face integration#OpenAI-compatible API

An open-source local LLM stack manager that connects Ollama, llama.cpp and AI apps through Docker Compose. Includes a CLI and companion app.

3.2KUpdated 5 days agoApache-2.0

macOS · Linux · Docker#GGUF#llama.cpp backend#MCP

A local LLM inference engine for sparse models, with CPU and GPU support on Linux and Windows. Open source under MIT, with CPU-only support on Apple Silicon.

9.8KUpdated 5 months agoMIT

macOS · Windows · Linux#Batch processing#GGUF#Hugging Face integration

An on-device AI engine runs automation models locally on phones and tiny devices, with speech, vision and optional cloud routing.

6.1KUpdated 5 days ago

macOS · iOS · Android#Hugging Face integration#Multimodal input#Quantization

An open source local LLM runner that splits inference and memory across your computers. Runs on Linux, macOS and Windows under the MIT license.

3.1KUpdated 3 months agoMIT

macOS · Windows · Linux#Distributed execution#Hugging Face integration#Quantization

A CLI and Python library for local LLMs and remote APIs, with SQLite prompt history, embeddings, and support for OpenAI-compatible endpoints.

12.6KUpdated 1 week agoApache-2.0

#LM Studio integration#Multimodal input#OpenAI-compatible API

Local LLM desktop app for Windows, macOS and Linux. Run GGUF models offline or connect other apps through OpenAI- and Anthropic-compatible APIs.

47.7KUpdated 1 month agoAGPL-3.0

macOS · Windows · Linux · Docker · Web#GGUF#llama.cpp backend#LoRA

A local LLM runner that packages a model and runtime in one file for macOS, Linux, BSD and Windows. Open source under Apache 2.0.

26.1KUpdated 2 weeks ago

macOS · Windows · Linux#GGUF#Hugging Face integration#llama.cpp backend

An open-source local LLM runner for Linux, macOS and Windows via WSL2, with Podman or Docker isolation and llama.cpp or vLLM inference.

3.1KUpdated 1 day agoMIT

macOS · Windows · Linux · Docker#GGUF#Hugging Face integration#llama.cpp backend

Favicon of koboldcpp

koboldcpp

1 video
Local LLM runner for GGUF and GGML models on Windows, macOS and Linux, with CPU or GPU support, a browser UI and an AGPL-3.0 license.

11.9KUpdated 4 days agoAGPL-3.0

macOS · Windows · Linux · Android · Docker · Web#GGUF#Hugging Face integration#llama.cpp backend

A mobile AI assistant that runs GGUF models on iOS and Android. Core chat works offline after a model download and needs no account.

8.4KUpdated 2 days agoMIT

iOS · Android#GGUF#Hugging Face integration#llama.cpp backend

An open-source local LLM compiler and deployment engine with GPU support across desktop, browser and mobile platforms, plus an OpenAI-compatible API.

23.2KUpdated 23 hours agoApache-2.0

macOS · Windows · Linux · iOS · Android · Web#OpenAI-compatible API

A local LLM inference engine with OpenAI and Anthropic-compatible APIs. Runs on macOS, Linux and Windows with CPU, CUDA or Apple Silicon support.

7.7KUpdated 5 days agoMIT

macOS · Windows · Linux · Docker · Web#Code execution#GGUF#Hugging Face integration

Favicon of Lemonade

Lemonade

1 video
An open source local AI server for chat, image generation, and speech on Windows, macOS, and Linux, with APIs for apps and agents.

5.8KUpdated 1 day agoApache-2.0

macOS · Windows · Linux · iOS · Android · Docker#GGUF#Hugging Face integration#llama.cpp backend

Favicon of llama-swap

llama-swap

1 video
A local AI proxy that switches models on demand through OpenAI and Anthropic compatible APIs. Runs on macOS, Windows, Linux and FreeBSD under MIT.

5.8KUpdated 2 days agoMIT

macOS · Windows · Linux · Docker · Web#GGUF#Image-to-image#llama.cpp backend

Favicon of GPT4All

GPT4All

1 video
An open-source local AI chatbot for Windows, macOS and Linux. Run models without a GPU or cloud API, and chat privately with your documents.

77.4KUpdated 1 year agoMIT

macOS · Windows · Linux · Docker#GGUF#llama.cpp backend#OpenAI-compatible API

Favicon of MLX LM

MLX LM

2 videos
A Python package for local LLM inference and fine-tuning on Apple Silicon, built on MLX with Hugging Face model support and an MIT license.

7.2KUpdated 1 day agoMIT

macOS#Batch processing#Distributed execution#Hugging Face integration

Run AI models through Docker Desktop, Docker Engine or a standalone binary, with local inference and OpenAI and Ollama compatible APIs.

655Updated 2 days agoApache-2.0

macOS · Windows · Linux#GGUF#Hugging Face integration#llama.cpp backend

A local AI desktop workspace for Apple Silicon Macs, Windows and Linux x64, with GGUF and MLX models, document search and optional cloud providers.

46Updated 5 days agoAGPL-3.0

macOS · Windows · Linux · Browser Extension#Code execution#GGUF#llama.cpp backend

More in Run Models Locally