Favicon of chatty

chatty

A browser-based local LLM chat app that uses WebGPU, works offline after model download, and processes chats and documents on your hardware.

Chatty runs AI models directly in your browser using WebGPU. It's an open-source ChatGPT alternative for people who want familiar chat features while keeping conversations and document processing on their own hardware. Once a model has downloaded, you can chat offline.

You can use the hosted website or self-host the Next.js app, including through Docker. In either case, AI processing happens in the browser rather than on a server. Chatty uses WebLLM and supports models including Gemma, Llama 2, Llama 3, Mistral and DeepSeek R1 reasoning models.

File Q&A is fully local too: you can ask questions about PDFs, text documents and code files without sending their contents elsewhere. Document search uses local embeddings. Smaller models may handle this work less efficiently than larger ones.

The chat interface keeps conversation history and accepts custom instructions or memory to shape responses. It also supports voice input, response regeneration, Markdown rendering and code highlighting. You can export messages as JSON or Markdown.

Chrome and Edge support the WebGPU browser path by default; Firefox requires enabling it. A GPU with enough memory matters: the stated guidance is about 3 GB for 3B models and 6 GB for 7B models. Chatty is free under the MIT license, and its Docker packaging isn't optimized for production.

Similar to chatty