PocketPal AI is an open source assistant for people who want to run language models on a phone or tablet. It works on iOS, iPadOS and Android. Once you've downloaded a model, you can chat offline without an account, and your prompts, replies and documents stay on your device. The app is licensed under MIT.
It runs GGUF models such as Gemma, Qwen, Phi and Llama using llama.cpp. You can add a model from local storage or find one through Hugging Face, including gated models if you have an access token. Downloading models requires a connection; using a downloaded model does not. Inference can use the phone's CPU, GPU or a supported Qualcomm NPU, with CPU fallback. An in-app benchmark shows model speed and memory use on your device.
PocketPal also has on-device text-to-speech, including Kokoro voices. You can create Pals with their own model and instructions, including roleplay settings. Capable Pals can use built-in tools for calculations, date and time, and HTML rendering during a conversation. You can edit a message to generate a new response or retry with a different model.
PalsHub lets you find community-made Pals. It's separate from core offline chat. Benchmark results reach the AI Phone Leaderboard only if you choose to share them, and feedback leaves the device only when you submit it.
Claim this page and we'll verify you by hand. PocketPal AI gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find PocketPal AI?Promote it
Something wrong or outdated on this page?
layla-network.aiAI Characters and Roleplay
iOS · Android#Code execution#GGUF#llama.cpp backend
Layla is an offline AI assistant for Android and iOS that runs language models on your phone. It's for people who want private everyday chat or an AI companion with custom characters, memory and roleplay. Local conversations stay encrypted on your device; optional cloud mode sends requests to your chosen hosted provider.
2.8KUpdated 1 week agoAGPL-3.0
Android#GGUF#llama.cpp backend#Ollama integration
2.1KUpdated 8 months agoMIT
macOS · iOS#llama.cpp backend#Multimodal input#RAG
locallyai.appDesktop Chat Apps
macOS · iOS#MLX#Multilingual#Multimodal input
6.1KUpdated 5 days ago
macOS · iOS · Android#Hugging Face integration#Multimodal input#Quantization
Cactus is an on-device AI engine for developers building automation into mobile apps, wearables and embedded devices. Its Needle model handles tool calling locally, so a device can turn a request into an action without an internet connection. The focus is small devices, including smart home hardware, robots and microcontrollers.
24.8KUpdated 21 hours agoApache-2.0
iOS · Android#Agent Skills#Hugging Face integration#Multilingual
ChatterUI is an Android chat app for people who want to run a local LLM on their phone or use the same interface with a remote model. It supports assistant conversations and character chats, with controls for how chats are structured and how models generate replies. It's open source under AGPL-3.0.
LLM Farm runs large language models offline on iOS and macOS. It's for people who want on-device AI chat or need to compare how different models perform on Apple hardware before choosing one for a project. The app is open source under the MIT license, with ggml and llama.cpp handling local inference.
Locally AI is a native app for running language and vision models on recent iPhones, iPads, and Macs. It's for people who want a private AI assistant on their own device, with text, image processing, and voice conversations available without cloud processing. Once a model is downloaded, it works offline and doesn't require an account.
Google AI Edge Gallery is an open-source app for people who want to try generative AI on their own phone. It runs model inference locally, so offline chat, image analysis and audio tasks don't send your inputs to a server. It supports Android and iOS and uses the Apache 2.0 license.