Run AI Models in the Browser and on Phones

Run models inside a web page or a mobile app on the CPU, GPU or NPU with toolkits like WebLLM, Transformers.js and ExecuTorch.

51 tools
Favicon of Slotstream

Slotstream

1 video
Local LLM runner for Apple Silicon Macs that runs Qwen3.8-Flash-Next from SSD. Works offline after download and connects to coding agents and chat apps.

407Updated 2 days agoMIT

macOS#MLX#Multimodal input#OpenAI-compatible API

A browser-based local LLM tool that pools laptop, desktop and phone GPUs for chat and coding. Open source under MIT, with no account required.

544Updated 14 hours agoMIT

macOS · iOS · Web#Code execution#Distributed execution#Hugging Face integration

Favicon of Meetily

Meetily

6 videos
A local AI meeting assistant for macOS and Windows that records and transcribes calls offline, with Ollama or your own API key for summaries.

31.3KUpdated 3 weeks agoMIT

macOS · Windows · Linux#Ollama integration#OpenAI-compatible API#Streaming inference

Favicon of Jan

Jan

3 videos
A free desktop AI chat app for Windows, macOS and Linux that runs models locally or connects to GPT, Claude and other cloud models.

44.7KUpdated 5 hours ago

macOS · Windows · Linux#Hugging Face integration#MCP#OpenAI-compatible API

A self-hostable photo storage app with on-device AI, encrypted backups and sharing, plus apps for mobile, desktop and the web. Open source under AGPL-3.0.

29.2KUpdated 1 day agoAGPL-3.0

macOS · Windows · Linux · iOS · Android · Web#Multi-user access#Role-based access#Semantic search

A local AI assistant for iPhone, iPad, Mac, and Vision Pro that runs models on device, searches your files, and can block external web connections.

noemaai.comChat With Your Documents

macOS · iOS#GGUF#MCP#MLX

Favicon of Chatterbox

Chatterbox

3 videos
An open-source text-to-speech model family that runs on your own hardware, clones voices from short clips, and supports offline deployment.

26.6KUpdated 2 months agoMIT

Linux#Multilingual#Voice cloning#Voice conversion

An open-source speech-to-text engine that runs Whisper models locally on desktop and mobile, with CPU-only inference and GPU acceleration. MIT licensed.

54KUpdated 2 days agoMIT

macOS · Windows · Linux · iOS · Android · Docker#Hugging Face integration#Quantization#Streaming inference

Run local models on your phone or computer for chat, document analysis, image generation and speech transcription. Connect to your own model server.

3.2KUpdated 4 days agoMIT

macOS · Windows · Linux · iOS · Android#GGUF#Human approval#LM Studio integration

Favicon of Locally AI

Locally AI

2 videos
Local AI assistant that runs language and vision models on Apple devices. Works offline after model download, with on-device voice and no account required.

locallyai.appDesktop Chat Apps

macOS · iOS#MLX#Multilingual#Multimodal input

Local AI chat app for macOS that runs GGUF models offline, searches files on-device, and also connects to Claude, ChatGPT, and OpenAI-compatible endpoints.

recurse.chatChat With Your Documents

macOS#GGUF#Hugging Face integration#OpenAI-compatible API

A paid audio transcription app that runs Whisper locally on macOS, iOS and visionOS, with multilingual transcription and subtitle exports.

sindresorhus.comOn-Device and In-Browser AI

macOS · iOS#Batch processing#Multilingual

A free, open-source browser translator for Chrome and Firefox with on-device translation, an offline dictionary, and custom backends such as Ollama.

1.1KUpdated 4 weeks agoBSD-3-Clause

Android · Browser Extension#Multilingual#Ollama integration#Works offline

A local LLM inference SDK for desktop, mobile and browsers. Runs on CPU or GPU, with offline inference and online Picovoice account validation.

318Updated 3 weeks agoApache-2.0

macOS · Windows · Linux · iOS · Android · Web#Quantization

A C++ neural translation library based on Marian NMT, with native and WebAssembly builds for local processing. Open source under MPL 2.0.

550Updated 2 years agoMPL-2.0

Web#Multilingual

An offline AI assistant for Android and iOS that runs models on your phone, keeps chats encrypted on-device, and offers optional hosted models.

layla-network.aiAI Characters and Roleplay

iOS · Android#Code execution#GGUF#llama.cpp backend

An on-device AI SDK that runs text, image and audio models on macOS, Windows and Linux, with GGUF, MLX and an OpenAI-compatible API.

qualcomm/GenieXInference Libraries and Bindings

macOS · Windows · Linux#GGUF#Hugging Face integration#llama.cpp backend

Open-source vision language model for local image and text tasks, with Apache 2.0 licensing, Transformers support, and a small GPU memory footprint.

3.9KUpdated 1 week agoApache-2.0

#Hugging Face integration#LoRA#Multimodal input

An open source on-device AI engine for local LLMs and image models, with iOS, Android, CPU and GPU support under Apache 2.0.

16.2KUpdated 1 day agoApache-2.0

Windows · iOS · Android#Image-to-image#Multimodal input#ONNX

A local LLM app that runs models offline on iOS and macOS, supports text and vision models, and uses ggml and llama.cpp under the MIT license.

2.1KUpdated 8 months agoMIT

macOS · iOS#llama.cpp backend#Multimodal input#RAG

Local vision-language models for image and video understanding, with Apache 2.0 code, mobile deployment and support for Ollama and llama.cpp.

26.5KUpdated 3 weeks agoApache-2.0

macOS · iOS · Android · Web#GGUF#Hugging Face integration#llama.cpp backend

An open-source ChatGPT alternative for Chrome that uses Ollama models locally to chat across tabs, analyze files, and assist with writing.

1.1KUpdated 7 months agoAGPL-3.0

Web · Browser Extension#Multilingual#Multimodal input#Ollama integration

An open-source on-device AI framework for Android, iOS, desktop and web, with model conversion from PyTorch, TensorFlow and JAX.

3.5KUpdated 19 hours agoApache-2.0

macOS · Windows · Linux · iOS · Android · Web#Agent Skills#Hugging Face integration#Multimodal input

An AI inference framework for Android, iOS, desktop and browsers, with CPU and Vulkan GPU support and PyTorch and ONNX model conversion.

23.9KUpdated 6 days ago

macOS · Windows · Linux · iOS · Android · Web#ONNX#Quantization

A local LLM app for iPhone, iPad, and Mac that works offline after model download, keeps chats on-device, and connects to Siri and Apple Shortcuts.

privatellm.appAutomation and No-Code AI

macOS · iOS#Multilingual#Quantization#Works offline

An on-device AI runtime that runs PyTorch models locally on Android, iOS, desktops and embedded hardware, with CPU, GPU, NPU and DSP acceleration.

5.1KUpdated 20 hours ago

macOS · Windows · Linux · iOS · Android · Web#MLX#Multimodal input#OpenAI-compatible API

Favicon of Moonshine

Moonshine

3 videos
An on-device AI voice toolkit for speech recognition, intent recognition and text to speech, with support for desktop, mobile, browsers and Raspberry Pi.

11.2KUpdated 1 month ago

macOS · Windows · Linux · iOS · Android · Web#Multilingual#Streaming inference

A local LLM chat app for iOS, iPadOS, macOS and visionOS. Runs on Apple silicon, works offline and saves chat history locally. Licensed under MIT.

2.3KUpdated 1 year agoMIT

macOS · iOS#MLX#Quantization#Works offline

An open-source toolkit for on-device AI across mobile, web and desktop. Apache 2.0 licensed; input stays local, while API metrics go to Google.

37.1KUpdated 19 hours agoApache-2.0

iOS · Android · Web

Open-source Python tools optimize Hugging Face models for local, edge and cloud hardware, with ONNX Runtime, OpenVINO and TensorRT-LLM integrations.

3.5KUpdated 6 days agoApache-2.0

#Hugging Face integration#ONNX#Quantization

A local LLM chat app for Android that runs GGUF models on-device through llama.cpp. Open source under the Apache 2.0 license.

894Updated 3 months agoApache-2.0

Android#GGUF#llama.cpp backend

An open-source React Native library that runs GGUF models on iOS and Android through llama.cpp, with GPU acceleration and image and audio understanding.

1KUpdated 3 days agoMIT

iOS · Android#GGUF#llama.cpp backend#Multilingual

An on-device AI engine runs automation models locally on phones and tiny devices, with speech, vision and optional cloud routing.

6.1KUpdated 5 days ago

macOS · iOS · Android#Hugging Face integration#Multimodal input#Quantization

An open-source voice AI framework that processes speech directly, with local inference on Mac and iPhone through MLX and self-hosted server backends.

11.2KUpdated 5 months agoApache-2.0

macOS · iOS · Web#Hugging Face integration#MLX#Quantization

An open-source macOS dictation app that runs Whisper and Parakeet models locally on Apple Silicon and transcribes microphone recordings or audio files.

3KUpdated 3 weeks agoMIT

macOS#Batch processing#Hugging Face integration#Multilingual

Local LLM toolkit that runs language and multimodal models on Rockchip NPUs, with model conversion, quantization and C/C++ interfaces.

1.7KUpdated 2 days ago

Linux#Multimodal input#Quantization

More in Run Models Locally