
Private LLM runs AI chat entirely on your iPhone, iPad, or Mac. It's an app for people who want to use language models without sending their conversations to a cloud service. After the first model download, it works offline and requires no account. Conversations stay on-device, with no tracking or logs.
Its model selection includes DeepSeek R1 Distill, Llama 3.3, Qwen3, Phi 4, and Google Gemma 3. A model picker recommends choices for your device, and the catalog lets you filter by available RAM and intended use. The app offers uncensored chat, so model choice also matters when comparing how it responds to different prompts.
Siri and Apple Shortcuts support extends the app beyond a chat window. You can use local AI to summarize text, generate writing, and pass responses to apps that support x-callback-url, without writing code. On macOS, built-in writing tools can rewrite, summarize, or correct selected text in other apps. They support English and major Western European languages.
Private LLM uses OmniQuant and GPTQ to reduce model memory needs while limiting the loss of output quality. It pairs those methods with Metal kernels tuned for individual models on Apple hardware. That approach is a technical distinction for readers comparing apps that run the same models locally.
Claim this page with an email at privatellm.app. Private LLM gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Private LLM?Promote it
Something wrong or outdated on this page?
2.3KUpdated 1 year agoMIT
macOS · iOS#MLX#Quantization#Works offline
Fullmoon is an open source chat app for people who want to run language models on their Apple devices and keep conversations local. It supports iOS, iPadOS, macOS and visionOS, with on-device inference optimized for Apple silicon. You can chat fully offline, and the app saves your chat history locally.
locallyai.appDesktop Chat Apps
macOS · iOS#MLX#Multilingual#Multimodal input
2.1KUpdated 8 months agoMIT
macOS · iOS#llama.cpp backend#Multimodal input#RAG
noemaai.comChat With Your Documents
macOS · iOS#GGUF#MCP#MLX
3.2KUpdated 4 days agoMIT
macOS · Windows · Linux · iOS · Android#GGUF#Human approval#LM Studio integration
23.8KUpdated 19 hours agoMPL-2.0
macOS · Windows · Linux · iOS · Android#Multilingual#Persistent memory#RAG
Locally AI is a native app for running language and vision models on recent iPhones, iPads, and Macs. It's for people who want a private AI assistant on their own device, with text, image processing, and voice conversations available without cloud processing. Once a model is downloaded, it works offline and doesn't require an account.
LLM Farm runs large language models offline on iOS and macOS. It's for people who want on-device AI chat or need to compare how different models perform on Apple hardware before choosing one for a project. The app is open source under the MIT license, with ggml and llama.cpp handling local inference.
Noema is a private local AI assistant for iPhone, iPad, Mac, and Vision Pro. It's for people who want to chat with models and work with their own files on Apple hardware without depending on a cloud service. Local chats can stay on your device, and file retrieval runs there too.
Off Grid AI runs language models on iOS, Android, macOS, Windows and Linux. You can chat, analyze documents and generate images on your own hardware. The mobile app uses the MIT license, while the desktop app uses AGPL.
Brave Leo is an AI assistant built into the Brave browser, with Bring Your Own Model support for people who want to use their own local or remote models while browsing. It can work with third-party APIs as well as Brave's hosted model choices. The browser runs on macOS, Windows, Linux, Android, and iOS.