Transformers.js and WebGPU: browser AI explained

Learn how Transformers.js runs browser models with ONNX Runtime, WebGPU and persistent caching, plus why CPU fallbacks need application code.

Player not loading? Watch on YouTube

Nico Martin, an ML engineer at Hugging Face, explains how Transformers.js lets developers run models locally in a browser. The discussion covers text embeddings, audio transcription and background removal as well as language models. Its pipeline API handles preprocessing, inference and postprocessing so applications can pass task-specific inputs rather than construct tensors themselves.

Martin explains that an ONNX file alone does not guarantee compatibility: Transformers.js must also support the model's architecture. He describes WebGPU acceleration, including server runtime support introduced with version 4. Applications can check GPU availability and select CPU instead, but the library does not provide an automatic fallback list or an available-memory helper in the account given here.

Model downloads and hardware differences shape the on-device experience. Required files enter a persistent browser cache, but separate origins normally download their own copies. Martin describes cross-origin storage support as experimental. Large models can require gigabytes, and older phones or laptops may make inference too slow for the intended feature.

His offline AI agent demos use Gemma 4 E2B for tool calling, but he would not ship those demos as dependable customer applications. He suggests hybrid workflows with browser tools and cloud inference. Structured output and a custom WebGPU engine are work in progress; the reported 5x to 10x speedups come from early experiments.