Chatty runs AI models directly in your browser using WebGPU. It's an open-source ChatGPT alternative for people who want familiar chat features while keeping conversations and document processing on their own hardware. Once a model has downloaded, you can chat offline.
You can use the hosted website or self-host the Next.js app, including through Docker. In either case, AI processing happens in the browser rather than on a server. Chatty uses WebLLM and supports models including Gemma, Llama 2, Llama 3, Mistral and DeepSeek R1 reasoning models.
File Q&A is fully local too: you can ask questions about PDFs, text documents and code files without sending their contents elsewhere. Document search uses local embeddings. Smaller models may handle this work less efficiently than larger ones.
The chat interface keeps conversation history and accepts custom instructions or memory to shape responses. It also supports voice input, response regeneration, Markdown rendering and code highlighting. You can export messages as JSON or Markdown.
Chrome and Edge support the WebGPU browser path by default; Firefox requires enabling it. A GPU with enough memory matters: the stated guidance is about 3 GB for 3B models and 6 GB for 7B models. Chatty is free under the MIT license, and its Docker packaging isn't optimized for production.
Claim this page and we'll verify you by hand. chatty gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find chatty?Promote it
Something wrong or outdated on this page?
636Updated 2 months agoMIT
Web#Hugging Face integration#Human approval#Multilingual
ChatLLM Web runs AI chat and a local agent workspace entirely in a WebGPU-capable browser. It's for people who want to work with their own text and code while keeping conversations, selected files, and generated artifacts on their device. The project is open source under the MIT license.
12KUpdated 12 months agoApache-2.0
macOS · Windows · Linux · Docker · Web#Code execution#llama.cpp backend#Multi-user access
38.7KUpdated 11 months agoApache-2.0
macOS · Windows · Linux · Docker · Web#Multimodal input#Ollama integration#OpenAI-compatible API
2.4KUpdated 13 hours agoApache-2.0
Docker · Web#Agent Skills#Code execution#Multi-user access
25.8KUpdated 4 months agoApache-2.0
macOS · Windows · Linux · Docker · Web#Hybrid search#llama.cpp backend#Multi-user access
22.2KUpdated 1 month agoMIT
macOS · Windows · Linux · Docker · Web#Hugging Face integration#Hybrid search#Ollama integration
h2oGPT is a self-hosted ChatGPT alternative for people who want to chat with local models and ask questions about their own documents. The project is archived and no longer maintained. It's open source under Apache 2.0, with support for Linux, macOS, Windows and Docker.
Langchain-Chatchat is a self-hosted application for asking questions about your own documents and using AI agents. It focuses on Chinese-language use and open models, with a fully offline setup that can keep documents and model processing on your hardware. Its code is open source under Apache 2.0.
Bionic GPT is a self-hosted AI agent platform for internal AI teams that need control over their models, company data and business systems. It runs on-premise, in a private cloud or in an air-gapped environment. Teams can build company-specific workflows on its existing workspace and agent runtime rather than assemble the underlying platform themselves.
kotaemon is a self-hosted document chat app for people who want to ask questions across their files and check where the answers came from. It runs in a browser on Windows, macOS or Linux, with Docker also supported. The project uses the Apache 2.0 license.
localGPT is a self-hosted AI document chat app for people who want to question and summarise files on their own hardware. Its local Ollama setup keeps documents and conversations on your machine. Answers include source passages, so you can check what the model used.