WebLLM and Transformers.js: browser AI tutorial

Learn how a TypeScript browser app runs five AI demos locally, including Llama 3.2 chat, with cached models and WebGPU compatibility checks.

Player not loading? Watch on YouTube

The presenter walks through BrowserAI, a browser app with five demonstrations: image classification, Llama 3.2 chat, hand tracking, speech transcription and semantic search. The project uses front-end TypeScript without an inference backend. Models download into browser storage before use, and the presenter reuses cached downloads during the demos.

The metadata identifies the local LLM as Llama 3.2 1B with 4-bit quantization. The chat example summarizes a pasted Wikipedia article while the presenter points to increased GPU utilization as evidence of local processing. Other examples include an 80 MB image classifier and a 5 MB hand-tracking model. The Moonshine transcription demo processes a recording after the user stops it; the presenter reports 567 milliseconds for 12 seconds of audio and notes a transcription error. These timings describe the demonstrated session, not performance guarantees.

The code walkthrough explains WebGPU availability checks, an LLM worker using MLC Web-LLM, and image classification through Transformers.js. Some models need GPU acceleration, while others can use the CPU. The tutorial shows how developers can run models locally through a web interface, but browser support and download size constrain the approach. The presenter cautions that asking users to cache large models may be unsuitable for some applications, while describing browser inference as useful for shareable proofs of concept.