
Atomic Chat is a free local LLM app that also supplies models to coding assistants and AI agents. It builds on Jan by Menlo Research and runs on Apple Silicon Macs, Windows and Linux, with apps for iOS and Android. Local chats work offline without an account, and their data stays on your device.
You can browse Hugging Face models, including Llama, Qwen, DeepSeek, Mistral and Gemma. The app supports GGUF, MLX and ONNX formats; Linux runs GGUF models only. Chats and projects keep conversations organized, while persistent memory carries context across sessions. Custom assistants let you give different conversations their own system prompts, and a preview panel displays generated HTML, CSS and JavaScript.
Its OpenAI-compatible API lets other software use the models running on your computer. Integrations include Cline, Goose, OpenCode and OpenHands, and MCP connections give agents access to tools such as files and web search. The local server accepts connections only from your own machine by default.
Desktop inference uses llama.cpp and an Apple Silicon MLX backend. CPU execution and CUDA or Vulkan GPU acceleration are available through llama.cpp. TurboQuant reduces memory used by the model's context cache, and supported models can use speculative decoding to increase generation speed. RAM recommendations range from 8 GB for 3B models to 32 GB for 13B models.
A cloud option is also available. The app can connect to providers such as OpenAI and Anthropic with your own API keys; those requests go to the selected provider.
Claim this page with an email at atomic.chat. Atomic Chat gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Atomic Chat?Promote it
Something wrong or outdated on this page?
2.1KUpdated 8 months agoMIT
macOS · iOS#llama.cpp backend#Multimodal input#RAG
LLM Farm runs large language models offline on iOS and macOS. It's for people who want on-device AI chat or need to compare how different models perform on Apple hardware before choosing one for a project. The app is open source under the MIT license, with ggml and llama.cpp handling local inference.
locallyai.appDesktop Chat Apps
macOS · iOS#MLX#Multilingual#Multimodal input
77.4KUpdated 1 year agoMIT
macOS · Windows · Linux · Docker#GGUF#llama.cpp backend#OpenAI-compatible API
44.7KUpdated 17 hours ago
macOS · Windows · Linux#Hugging Face integration#MCP#OpenAI-compatible API
1.3KUpdated 2 days agoApache-2.0
macOS · Windows · Linux · iOS · Android · Web#Code execution#LM Studio integration#MCP
qualcomm/GenieXInference Libraries and Bindings
macOS · Windows · Linux#GGUF#Hugging Face integration#llama.cpp backend
Locally AI is a native app for running language and vision models on recent iPhones, iPads, and Macs. It's for people who want a private AI assistant on their own device, with text, image processing, and voice conversations available without cloud processing. Once a model is downloaded, it works offline and doesn't require an account.
GPT4All is a local AI chatbot for people who want to run language models on their own desktop or laptop and keep conversations on their machine. Its LocalDocs feature lets you ask questions about your own documents without sending them to a cloud service. It suits developers, teams and individuals who want control over their models and data.
Jan gives people a ChatGPT-style chat interface for AI models running on their own computer. It's free and available for Windows, macOS and Linux. Local chats work offline, with the model and conversation kept on your machine. You can also use cloud models when you want access to a hosted provider.
Kai 9000 is an open-source personal AI assistant that can turn responses into interactive screens, including quizzes, recipe cards and dashboards. It's for people who want an assistant that remembers their preferences and can act through tools, with apps for Android, iOS, Windows, macOS, Linux and the web.
Nexa SDK is an on-device AI inference framework for developers building applications that process text, images or audio on users' hardware. It runs models locally across CPUs, GPUs and NPUs, with a shared interface for different backends. Its scope includes language and vision models, speech recognition, speech synthesis and image generation.