407Updated 2 days agoMIT
macOS#MLX#Multimodal input#OpenAI-compatible API
Slotstream runs Qwen3.8-Flash-Next on Apple Silicon Macs that don't have enough RAM to hold the whole model. It's aimed at people with 16 to 64 GB of memory who want local chat, image questions or a model backend for coding agents. Most model weights stay on the SSD, while frequently used expert networks stay in memory. The full model remains available.
544Updated 14 hours agoMIT
macOS · iOS · Web#Code execution#Distributed execution#Hugging Face integration
Pooled runs a single open model across browser tabs on laptops, desktops and phones, combining their memory when the model won't fit on one device. It's for people who want local AI chat or a coding assistant using hardware they already have. It's open source under the MIT license and requires no account or per-device installation.
31.3KUpdated 3 weeks agoMIT
macOS · Windows · Linux#Ollama integration#OpenAI-compatible API#Streaming inference
Meetily is a local AI meeting assistant for people who want meeting notes while keeping recordings on their own device. It captures calls from Zoom, Google Meet, Microsoft Teams and other meeting software without placing a bot in the meeting. You can watch the transcript appear during the call, then generate a summary.
44.7KUpdated 5 hours ago
macOS · Windows · Linux#Hugging Face integration#MCP#OpenAI-compatible API
Jan gives people a ChatGPT-style chat interface for AI models running on their own computer. It's free and available for Windows, macOS and Linux. Local chats work offline, with the model and conversation kept on your machine. You can also use cloud models when you want access to a hosted provider.
29.2KUpdated 1 day agoAGPL-3.0
macOS · Windows · Linux · iOS · Android · Web#Multi-user access#Role-based access#Semantic search
Ente Photos is a Google Photos alternative for people who want encrypted photo backups and AI search without giving a storage provider access to their library. You can use its hosted service or self-host the server. Face detection, grouping and natural-language search run on your device; the hosted service stores your encrypted backups.
noemaai.comChat With Your Documents
macOS · iOS#GGUF#MCP#MLX
Noema is a private local AI assistant for iPhone, iPad, Mac, and Vision Pro. It's for people who want to chat with models and work with their own files on Apple hardware without depending on a cloud service. Local chats can stay on your device, and file retrieval runs there too.
26.6KUpdated 2 months agoMIT
Linux#Multilingual#Voice cloning#Voice conversion
Chatterbox is an MIT-licensed text-to-speech model family for developers and creators who want to generate speech on their own hardware. You can self-host it on a GPU, including in an air-gapped environment, without an account or API key. Resemble AI also offers separate managed hosting.
54KUpdated 2 days agoMIT
macOS · Windows · Linux · iOS · Android · Docker#Hugging Face integration#Quantization#Streaming inference
whisper.cpp runs OpenAI's Whisper speech recognition models on your own hardware, with fully offline transcription once you've downloaded a model. It's for developers building speech-to-text into applications and people who want to transcribe audio locally. Audio can stay on-device rather than going to a cloud transcription service. The project is open source under the MIT license.
3.2KUpdated 4 days agoMIT
macOS · Windows · Linux · iOS · Android#GGUF#Human approval#LM Studio integration
Off Grid AI runs language models on iOS, Android, macOS, Windows and Linux. You can chat, analyze documents and generate images on your own hardware. The mobile app uses the MIT license, while the desktop app uses AGPL.
locallyai.appDesktop Chat Apps
macOS · iOS#MLX#Multilingual#Multimodal input
Locally AI is a native app for running language and vision models on recent iPhones, iPads, and Macs. It's for people who want a private AI assistant on their own device, with text, image processing, and voice conversations available without cloud processing. Once a model is downloaded, it works offline and doesn't require an account.
recurse.chatChat With Your Documents
macOS#GGUF#Hugging Face integration#OpenAI-compatible API
RecurseChat is a paid Mac app for people who want to chat with AI and ask questions about their files on their own computer. It runs local LLMs without an internet connection or a server, keeping local conversations on the device. The same app also connects to Claude and ChatGPT, so you can choose local processing or a cloud provider.
sindresorhus.comOn-Device and In-Browser AI
macOS · iOS#Batch processing#Multilingual
Aiko is a paid, native transcription app for macOS, iOS and visionOS that processes speech on your device with OpenAI's Whisper model. It's for people turning meetings, lectures or other recordings into text while keeping the audio local, including sensitive recordings.
1.1KUpdated 4 weeks agoBSD-3-Clause
Android · Browser Extension#Multilingual#Ollama integration#Works offline
Linguist is a free browser translation extension for people who read across languages and want control over where their text goes. Its built-in Bergamot translator processes text on your device without an internet connection. You can also choose an external translation provider or connect a local backend such as Ollama or LibreTranslate.
318Updated 3 weeks agoApache-2.0
macOS · Windows · Linux · iOS · Android · Web#Quantization
picoLLM is an on-device inference SDK for developers building apps that run compressed language models on users' hardware. It generates text locally, so prompts don't need to go to a cloud inference service. Its main distinction is Picovoice's compression method, which learns how to allocate precision across model weights rather than applying a fixed allocation.
550Updated 2 years agoMPL-2.0
Web#Multilingual
Bergamot Translator is a C++ library for running neural machine translation on your own device. Developers can integrate it into native applications or build it for WebAssembly. Separate browser integrations, such as the TranslateLocally extension, make the engine accessible to end users. The software is free and open source under the Mozilla Public License 2.0.
layla-network.aiAI Characters and Roleplay
iOS · Android#Code execution#GGUF#llama.cpp backend
Layla is an offline AI assistant for Android and iOS that runs language models on your phone. It's for people who want private everyday chat or an AI companion with custom characters, memory and roleplay. Local conversations stay encrypted on your device; optional cloud mode sends requests to your chosen hosted provider.
qualcomm/GenieXInference Libraries and Bindings
macOS · Windows · Linux#GGUF#Hugging Face integration#llama.cpp backend
Nexa SDK is an on-device AI inference framework for developers building applications that process text, images or audio on users' hardware. It runs models locally across CPUs, GPUs and NPUs, with a shared interface for different backends. Its scope includes language and vision models, speech recognition, speech synthesis and image generation.
3.9KUpdated 1 week agoApache-2.0
#Hugging Face integration#LoRA#Multimodal input
SmolVLM is a compact vision language model from Hugging Face for developers building local AI applications that work with images and text. It can describe pictures, answer questions about diagrams, and read information from documents such as invoices. Its small memory footprint makes on-device use practical on laptops and smaller local setups.
16.2KUpdated 1 day agoApache-2.0
Windows · iOS · Android#Image-to-image#Multimodal input#ONNX
MNN is a lightweight C++ framework for developers who want AI models to run on phones, PCs and embedded devices. It handles inference and training on the device, with a focus on small application footprints and hardware acceleration. The project is open source under Apache 2.0, and Alibaba uses it in apps including Taobao, Youku and DingTalk.
2.1KUpdated 8 months agoMIT
macOS · iOS#llama.cpp backend#Multimodal input#RAG
LLM Farm runs large language models offline on iOS and macOS. It's for people who want on-device AI chat or need to compare how different models perform on Apple hardware before choosing one for a project. The app is open source under the MIT license, with ggml and llama.cpp handling local inference.
26.5KUpdated 3 weeks agoApache-2.0
macOS · iOS · Android · Web#GGUF#Hugging Face integration#llama.cpp backend
MiniCPM-V is a family of local vision-language models for developers building apps that interpret images and video on their own hardware. It supports iOS, Android and HarmonyOS, as well as Mac deployment and server inference. The current repository states that MiniCPM-o/V code and model weights use Apache 2.0.
1.1KUpdated 7 months agoAGPL-3.0
Web · Browser Extension#Multilingual#Multimodal input#Ollama integration
NativeMind brings a local AI assistant into Chrome for people who want help with webpages, documents, and writing without sending that content to a cloud model. The browser extension connects to Ollama on your machine, keeping prompts and AI processing on-device. It requires no account and uses the AGPL-3.0 open-source license.
3.5KUpdated 19 hours agoApache-2.0
macOS · Windows · Linux · iOS · Android · Web#Agent Skills#Hugging Face integration#Multimodal input
LiteRT is Google's open-source framework for developers building AI into apps that run on users' own devices. It succeeds TensorFlow Lite and covers model conversion, optimization and local inference. It's licensed under Apache 2.0.
23.9KUpdated 6 days ago
macOS · Windows · Linux · iOS · Android · Web#ONNX#Quantization
ncnn is a C++ framework for developers building on-device AI into mobile, desktop and embedded applications. Its focus is running neural networks with a small memory footprint and no third-party runtime dependencies. Models run on the target device's CPU or a supported Vulkan GPU.
privatellm.appAutomation and No-Code AI
macOS · iOS#Multilingual#Quantization#Works offline
Private LLM runs AI chat entirely on your iPhone, iPad, or Mac. It's an app for people who want to use language models without sending their conversations to a cloud service. After the first model download, it works offline and requires no account. Conversations stay on-device, with no tracking or logs.
5.1KUpdated 20 hours ago
macOS · Windows · Linux · iOS · Android · Web#MLX#Multimodal input#OpenAI-compatible API
ExecuTorch is PyTorch's runtime for developers building AI into mobile apps, desktop software and embedded devices. It runs models on the user's hardware, with support for Android, iOS, Linux, macOS and Windows, as well as microcontrollers. Developers can reuse a PyTorch model across targets, though hardware-specific deployments need their own exported model files.
11.2KUpdated 1 month ago
macOS · Windows · Linux · iOS · Android · Web#Multilingual#Streaming inference
Moonshine is an on-device AI toolkit for developers building voice agents and applications that listen and speak. It combines speech to text, intent recognition and text to speech in one library. Voice processing stays on the device, and you don't need an account or API keys.
2.3KUpdated 1 year agoMIT
macOS · iOS#MLX#Quantization#Works offline
Fullmoon is an open source chat app for people who want to run language models on their Apple devices and keep conversations local. It supports iOS, iPadOS, macOS and visionOS, with on-device inference optimized for Apple silicon. You can chat fully offline, and the app saves your chat history locally.
37.1KUpdated 19 hours agoApache-2.0
iOS · Android · Web
MediaPipe is an open-source toolkit for developers adding on-device AI to applications on Android, iOS, the web, desktop and edge devices. It pairs pretrained models with APIs for specific tasks, so developers can use existing solutions or customize them for their applications. The project uses the Apache 2.0 license.
3.5KUpdated 6 days agoApache-2.0
#Hugging Face integration#ONNX#Quantization
Optimum is a collection of Python packages for developers who want to train or run Hugging Face models more efficiently on specific hardware. It extends Transformers, Diffusers, TIMM and Sentence Transformers, with integrations for local machines, mobile and edge devices, and cloud accelerators. It's open source under Apache 2.0.
894Updated 3 months agoApache-2.0
Android#GGUF#llama.cpp backend
SmolChat is an Android app for people who want to chat with language models running on their own phone. It uses llama.cpp to run GGUF models on-device, including small language models. Inference stays on your device.
1KUpdated 3 days agoMIT
iOS · Android#GGUF#llama.cpp backend#Multilingual
llama.rn brings llama.cpp into React Native apps so developers can run local LLM inference on iOS and Android. It's an MIT-licensed library for building AI features into a mobile app, with model processing on the device. It uses GGUF models and requires React Native's New Architecture.
6.1KUpdated 5 days ago
macOS · iOS · Android#Hugging Face integration#Multimodal input#Quantization
Cactus is an on-device AI engine for developers building automation into mobile apps, wearables and embedded devices. Its Needle model handles tool calling locally, so a device can turn a request into an action without an internet connection. The focus is small devices, including smart home hardware, robots and microcontrollers.
11.2KUpdated 5 months agoApache-2.0
macOS · iOS · Web#Hugging Face integration#MLX#Quantization
Moshi is a voice AI model and dialogue framework that can listen while it speaks. It processes speech directly, retaining information such as emotion and non-verbal cues that a text transcription can miss. It's aimed at researchers and developers building spoken AI applications, with local inference and self-hosted server options.
3KUpdated 3 weeks agoMIT
macOS#Batch processing#Hugging Face integration#Multilingual
OpenSuperWhisper is a local speech-to-text app for people who want to dictate or transcribe recordings on an Apple Silicon Mac. It supports Whisper and Parakeet, with model downloads available inside the app. The project is open source under the MIT license.
1.7KUpdated 2 days ago
Linux#Multimodal input#Quantization
RKLLM is a software stack for developers building local AI applications on Rockchip hardware. It uses the chip's neural processing unit (NPU) to run language and multimodal models on development boards, with support for the RK3588, RK3576, RK3562 and RV1126B series.