Favicon of WebLLM

WebLLM

A local LLM engine that runs models in the browser with WebGPU acceleration, OpenAI API compatibility, and an Apache 2.0 license.

Screenshot of WebLLM website

WebLLM runs language models directly in a user's browser, using WebGPU for GPU acceleration. It's an open-source engine for developers building web-based AI assistants and Chrome extensions that process prompts on the user's device rather than an inference server. The project uses the Apache 2.0 license.

Its OpenAI-compatible API lets developers use familiar interfaces with models running locally. It supports streaming responses and structured JSON output with custom schemas, with preliminary function-calling support. Applications can display text as it's generated or request structured results for other software to use. Seed controls support reproducible output, and logit controls let developers adjust generation behavior.

Supported model families include Llama, Phi, Gemma, Mistral, Qwen, and RedPajama. WebLLM also accepts custom models in MLC format and is a companion project to MLC LLM. That gives developers a route to use their own model variants in browser applications.

The TypeScript package works with an app's own interface. Web Worker and Service Worker support let model computation run separately from the interface, helping it stay responsive during generation. Browser storage can cache model files, and optional integrity checks verify downloaded configuration, WebAssembly, and tokenizer files against supplied hashes.

Similar to WebLLM