Player not loading? Watch on YouTube
This video compares LiteRT.js with TensorFlow.js and tests browser inference through a webcam motion capture demo. The speaker describes LiteRT.js as a JavaScript binding that brings LiteRT's on-device runtime to the browser through WebAssembly, with optimized CPU and WebGPU kernels. The demonstrated speed comparison switches LiteRT.js between CPU and WebGPU; it does not measure TensorFlow.js.
The architecture discussion covers XNNPACK for CPU inference, WebGPU for GPU workloads, and WebNN for neural processing hardware. The speaker notes that WebNN remains experimental in Chrome and Edge. Existing .tflite files can run directly, while PyTorch models have a conversion path through LiteRT Torch. He also discusses AI Edge Quantizer for reducing model sizes. Google's reported benchmark gains come with hardware, thermal throttling, and driver caveats.
The demo uses BlazePose's 33 body landmarks to drive a 3D character with Three.js. The presenter says it runs entirely client-side and offline. His measurements show about 38 FPS with 23.3 ms inference on CPU, compared with 120 FPS and 8.4 ms on WebGPU. These are results from his demo, not guarantees for other devices.
Captured movement can export as JSON or BVH for retargeting in Blender, though the presenter calls the animation unfinished and lacking detail. He also introduces LiteRT-LM.js as an announced future runtime for local LLM inference in the browser.