Favicon of OfflineLLM

OfflineLLM

An offline AI chat app that runs GGUF models on Android through llama.cpp, with optional Vulkan GPU acceleration and no network permissions.

OfflineLLM is an on-device AI chat app for Android users who want conversations to stay on their phone. It runs entirely offline and has no network permissions, so the app can't send prompts or responses to a cloud service. It requires Android 13 or later on a 64-bit ARM device.

You supply your own GGUF models, which run locally through llama.cpp. Supported choices include Gemma 3, Qwen3.5 and Gemma 4, with native chat support for Gemma 4 E2B and E4B. Model choice depends on available memory: the recommendations include small Gemma models for phones with 2 to 4 GB of RAM and larger models for devices with more memory.

Inference can use the CPU or optional Vulkan GPU acceleration, with automatic CPU fallback if GPU loading fails. GPU speed varies by phone, and some Mali devices can run faster on the CPU. Longer conversations reuse prior processing rather than processing the whole chat again for each reply.

The chat interface streams responses and supports Markdown, translation and read-aloud output through Android's text-to-speech system. You can search and rename conversations, export or restore chats as JSON, and use prompts for coding, writing or tutoring. A context indicator shows how much of the model's conversation capacity you're using.

Privacy controls include encrypted settings, an optional biometric lock and file overwriting before deletion. The app doesn't log prompts or responses. Release builds check their signing certificate and refuse to run if someone has repackaged them. The project uses the Apache 2.0 license.

Similar to OfflineLLM