
WebLLM runs language models directly in a user's browser, using WebGPU for GPU acceleration. It's an open-source engine for developers building web-based AI assistants and Chrome extensions that process prompts on the user's device rather than an inference server. The project uses the Apache 2.0 license.
Its OpenAI-compatible API lets developers use familiar interfaces with models running locally. It supports streaming responses and structured JSON output with custom schemas, with preliminary function-calling support. Applications can display text as it's generated or request structured results for other software to use. Seed controls support reproducible output, and logit controls let developers adjust generation behavior.
Supported model families include Llama, Phi, Gemma, Mistral, Qwen, and RedPajama. WebLLM also accepts custom models in MLC format and is a companion project to MLC LLM. That gives developers a route to use their own model variants in browser applications.
The TypeScript package works with an app's own interface. Web Worker and Service Worker support let model computation run separately from the interface, helping it stay responsive during generation. Browser storage can cache model files, and optional integrity checks verify downloaded configuration, WebAssembly, and tokenizer files against supplied hashes.
Claim this page with an email at webllm.mlc.ai. WebLLM gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find WebLLM?Promote it
Something wrong or outdated on this page?
16.3KUpdated 1 week agoApache-2.0
Web#Hugging Face integration#Image-to-image#Multilingual
Transformers.js is a JavaScript library for developers building web apps that run AI models on the user's device. Inference happens in the browser, so an app doesn't need a separate model server to process its inputs. The library is open source under Apache 2.0.
3.5KUpdated 19 hours agoApache-2.0
macOS · Windows · Linux · iOS · Android · Web#Agent Skills#Hugging Face integration#Multimodal input
318Updated 3 weeks agoApache-2.0
macOS · Windows · Linux · iOS · Android · Web#Quantization
picoLLM is an on-device inference SDK for developers building apps that run compressed language models on users' hardware. It generates text locally, so prompts don't need to go to a cloud inference service. Its main distinction is Picovoice's compression method, which learns how to allocate precision across model weights rather than applying a fixed allocation.
1KUpdated 3 days agoMIT
iOS · Android#GGUF#llama.cpp backend#Multilingual
qualcomm/GenieXInference Libraries and Bindings
macOS · Windows · Linux#GGUF#Hugging Face integration#llama.cpp backend
5.1KUpdated 20 hours ago
macOS · Windows · Linux · iOS · Android · Web#MLX#Multimodal input#OpenAI-compatible API
ExecuTorch is PyTorch's runtime for developers building AI into mobile apps, desktop software and embedded devices. It runs models on the user's hardware, with support for Android, iOS, Linux, macOS and Windows, as well as microcontrollers. Developers can reuse a PyTorch model across targets, though hardware-specific deployments need their own exported model files.
LiteRT is Google's open-source framework for developers building AI into apps that run on users' own devices. It succeeds TensorFlow Lite and covers model conversion, optimization and local inference. It's licensed under Apache 2.0.
llama.rn brings llama.cpp into React Native apps so developers can run local LLM inference on iOS and Android. It's an MIT-licensed library for building AI features into a mobile app, with model processing on the device. It uses GGUF models and requires React Native's New Architecture.
Nexa SDK is an on-device AI inference framework for developers building applications that process text, images or audio on users' hardware. It runs models locally across CPUs, GPUs and NPUs, with a shared interface for different backends. Its scope includes language and vision models, speech recognition, speech synthesis and image generation.