Run AI Models in the Browser and on Phones

Run models inside a web page or a mobile app on the CPU, GPU or NPU with toolkits like WebLLM, Transformers.js and ExecuTorch.

51 tools
Favicon of Willow

Willow

1 video
An open-source voice assistant for ESP32-S3-BOX hardware, with offline commands, self-hosted speech recognition, and Home Assistant integration.

3.1KUpdated 2 months agoApache-2.0

#Home Assistant integration#Voice activity detection#Wake word detection

An open source Android AI chat app that runs GGUF models offline through llama.cpp and connects to remote providers, including Ollama and OpenAI.

2.7KUpdated 2 weeks agoMIT

Android#GGUF#Hugging Face integration#llama.cpp backend

An open source engine for running ONNX models on Windows, macOS, Linux, mobile devices and the web, with CPU, GPU and NPU support.

22KUpdated 27 minutes agoMIT

macOS · Windows · Linux · iOS · Android · Web#Distributed execution#ONNX

An open-source JavaScript library that runs AI models in your browser using ONNX Runtime, with CPU execution and optional WebGPU acceleration.

16.3KUpdated 1 week agoApache-2.0

Web#Hugging Face integration#Image-to-image#Multilingual

A mobile AI assistant that runs GGUF models on iOS and Android. Core chat works offline after a model download and needs no account.

8.5KUpdated 2 days agoMIT

iOS · Android#GGUF#Hugging Face integration#llama.cpp backend

Favicon of Draw Things

Draw Things

4 videos
An AI image generator that runs offline on iPhone, iPad and Mac, with on-device LoRA training and optional self-hosted or managed cloud compute.

575Updated 1 day agoGPL-3.0

macOS · Linux · iOS · Docker#Image-to-image#Inpainting#LoRA

An open-source local LLM compiler and deployment engine with GPU support across desktop, browser and mobile platforms, plus an OpenAI-compatible API.

23.2KUpdated 24 hours agoApache-2.0

macOS · Windows · Linux · iOS · Android · Web#OpenAI-compatible API

An open-source Android LLM chat app that runs GGUF models on-device through llama.cpp or connects to Ollama, OpenAI and Claude. Licensed under AGPL-3.0.

2.8KUpdated 1 week agoAGPL-3.0

Android#GGUF#llama.cpp backend#Ollama integration

An MIT-licensed translation app for Windows, macOS and Linux that runs Marian and Bergamot models locally and connects to Chrome and Firefox.

631Updated 2 years agoMIT

macOS · Windows · Linux · Web · Browser Extension#Batch processing#Multilingual#Quantization

Favicon of OpenVINO

OpenVINO

2 videos
Open-source AI inference toolkit for local or self-hosted deployment on Linux, Windows and macOS, with CPU, Intel GPU and NPU support.

10.9KUpdated 24 hours agoApache-2.0

macOS · Windows · Linux#Hugging Face integration#Multimodal input#ONNX

Favicon of WebLLM

WebLLM

2 videos
A local LLM engine that runs models in the browser with WebGPU acceleration, OpenAI API compatibility, and an Apache 2.0 license.

19.2KUpdated 2 weeks agoApache-2.0

Web · Browser Extension#OpenAI-compatible API#Streaming inference#Structured output

Open-source Rust ML framework for running models locally on CPUs, NVIDIA GPUs or in browsers, with Apache 2.0 licensing and quantized LLM support.

21.1KUpdated 2 days agoApache-2.0

macOS · Web#GGUF#Hugging Face integration#Multilingual

Local AI app for Android and iOS. Run Gemma 4 on-device, ask questions about photos, transcribe audio and compare model performance.

24.8KUpdated 21 hours agoApache-2.0

iOS · Android#Agent Skills#Hugging Face integration#Multilingual

Chat with models on your iPhone or iPad through Liquid LEAP, or connect to your own Ollama or LM Studio server. Cloud providers are optional.

liquid.aiMobile Chat Apps

macOS · iOS

An AI dictation app for iPhone and iPad that uses Gemma models locally. It supports offline transcription and optional cloud text features.

apps.apple.comDictation and Voice Typing

iOS#Multilingual#Works offline

More in Run Models Locally