Favicon of ChatLLM Web

ChatLLM Web

A local AI chat and agent workspace that runs models in your browser with WebGPU. MIT licensed, with on-device storage and no account or API key.

Screenshot of ChatLLM Web website

ChatLLM Web runs AI chat and a local agent workspace entirely in a WebGPU-capable browser. It's for people who want to work with their own text and code while keeping conversations, selected files, and generated artifacts on their device. The project is open source under the MIT license.

The hosted website serves the application through Cloudflare Pages, while your device handles model inference. Model downloads come from Hugging Face and WebLLM library URLs. There's no application backend, account, API key, analytics, or telemetry. Its installable PWA caches the app and can reuse models after their first successful load.

The model studio checks browser capabilities and memory to recommend a compatible model. Choices include Qwen 3.5 2B, Llama 3.2 1B, Gemma 3 1B, and Qwen 2.5 Coder. You can compare compatibility and cache status, manage downloads, or import custom MLC models from approved sources. WebLLM runs generation through WebGPU, so model suitability depends on the device's GPU capabilities and available memory.

Agent mode uses Hermes models for function calling. Its sandboxed tools can read and search files you've selected, calculate, and create or edit saved artifacts. You review artifact changes as a line diff before saving them. The agent has no arbitrary disk access, shell execution, or network search.

For chat, you can attach text, Markdown, JSON, and common code files as direct context. Conversations and approved artifacts stay in browser storage; artifacts can be copied or downloaded.

Similar to ChatLLM Web