OfflineLLM is an on-device AI chat app for Android users who want conversations to stay on their phone. It runs entirely offline and has no network permissions, so the app can't send prompts or responses to a cloud service. It requires Android 13 or later on a 64-bit ARM device.
You supply your own GGUF models, which run locally through llama.cpp. Supported choices include Gemma 3, Qwen3.5 and Gemma 4, with native chat support for Gemma 4 E2B and E4B. Model choice depends on available memory: the recommendations include small Gemma models for phones with 2 to 4 GB of RAM and larger models for devices with more memory.
Inference can use the CPU or optional Vulkan GPU acceleration, with automatic CPU fallback if GPU loading fails. GPU speed varies by phone, and some Mali devices can run faster on the CPU. Longer conversations reuse prior processing rather than processing the whole chat again for each reply.
The chat interface streams responses and supports Markdown, translation and read-aloud output through Android's text-to-speech system. You can search and rename conversations, export or restore chats as JSON, and use prompts for coding, writing or tutoring. A context indicator shows how much of the model's conversation capacity you're using.
Privacy controls include encrypted settings, an optional biometric lock and file overwriting before deletion. The app doesn't log prompts or responses. Release builds check their signing certificate and refuse to run if someone has repackaged them. The project uses the Apache 2.0 license.
Claim this page and we'll verify you by hand. OfflineLLM gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find OfflineLLM?Promote it
Something wrong or outdated on this page?
24.8KUpdated 21 hours agoApache-2.0
iOS · Android#Agent Skills#Hugging Face integration#Multilingual
Google AI Edge Gallery is an open-source app for people who want to try generative AI on their own phone. It runs model inference locally, so offline chat, image analysis and audio tasks don't send your inputs to a server. It supports Android and iOS and uses the Apache 2.0 license.
0Updated 3 weeks agoMIT
Android#GGUF#Hugging Face integration#llama.cpp backend
1.7KUpdated 1 day ago
macOS · Windows · Linux · iOS · Android#GGUF#Hugging Face integration#llama.cpp backend
2.8KUpdated 14 hours agoAGPL-3.0
Android#GGUF#llama.cpp backend#Ollama integration
20Updated 2 days agoAGPL-3.0
Android#GGUF#Hugging Face integration#llama.cpp backend
293Updated 1 day agoApache-2.0
Android#Multilingual#Ollama integration#OpenAI-compatible API
Airux Pocket AI runs text and vision models directly on arm64 Android phones running Android 8 or later. It's for people who want private AI chat on their phone, including offline conversations and photo analysis without cloud inference. You don't need an account. Chat history and downloaded models stay in the app's private storage.
Atomic Chat is a free local LLM app that also supplies models to coding assistants and AI agents. It builds on Jan by Menlo Research and runs on Apple Silicon Macs, Windows and Linux, with apps for iOS and Android. Local chats work offline without an account, and their data stays on your device.
ChatterUI is an Android chat app for people who want to run a local LLM on their phone or use the same interface with a remote model. It supports assistant conversations and character chats, with controls for how chats are structured and how models generate replies. It's open source under AGPL-3.0.
CloverPal is an Android app for people who want private AI chat and character roleplay on their phone. Its Offline Silly Tavern mode lets you create characters with their own avatars and system prompts, then have them converse in group chats. Model inference runs on-device, and chats and uploaded files stay on your phone.
Dictate Keyboard is an Android keyboard for people who prefer speaking to typing or find typing uncomfortable. It runs on Android 8.0 and later, with offline speech recognition on the phone and optional cloud or self-hosted providers. The app is open source under Apache 2.0 and builds on FlorisBoard.