Transformers.js tutorial: ONNX and browser AI pipelines

Learn how Transformers.js loads models and runs browser inference, with WebGPU or WASM options and examples of text generation and depth estimation.

Player not loading? Watch on YouTube

Nico explains how Transformers.js lets JavaScript applications run models locally through a shared pipeline API. The introduction covers browser text generation, speech recognition and background removal before explaining tensors, neural networks and model weights.

The tutorial separates ONNX model files from the runtime that executes their computation graphs. Transformers.js handles the surrounding work: downloading and caching assets, preparing inputs and converting model outputs into results an application can use. Nico describes browser caching through the Cache API and file-based caching in server environments.

The pipeline() walkthrough explains task IDs, model selection and download progress callbacks. A model ID usually points to a Hugging Face Hub repository, though local model paths depend on the environment setup. Browser inference can use WebGPU or WASM depending on availability. Nico says WebGPU is also recommended for server-side JavaScript runtimes starting with version four. Precision options include FP32, FP16, Q4 and Q4F16; quantization can reduce memory use but may cost accuracy.

Two examples explain the internal flow. Text generation repeatedly predicts and decodes tokens until a stop token or configured limit. Depth estimation prepares an image, runs one forward pass and returns relative depth information.