Pooled (SwarmLLM): distributed browser inference

Explore Pooled under its former SwarmLLM name: split a Qwen model across browsers using WebGPU and WebRTC, with latency and peer-trust trade-offs.

Player not loading? Watch on YouTube

The video covers Pooled under its former name, SwarmLLM. This open source project splits a 27-billion-parameter Qwen model across devices for browser-based inference. The speaker describes a MacBook and iPhone demo that produced 10.7 tokens per second. The MacBook held 62 transformer layers plus the embedding and output components; the phone held two layers.

Each browser downloads only its assigned weight ranges and caches them locally. WebGPU handles computation, while WebRTC carries a roughly 10KB hidden-state vector between peers. A signaling broker introduces the browsers but does not carry model traffic after connection. The video explains speculative decoding and reports a single-device generation rate of 16 tokens per second with it enabled. Its comparison with llama.cpp also shows substantially slower prompt processing in the browser.

Extra devices supply memory capacity, but each hop adds latency. The cited tests slowed when the room grew from three devices to 16. Trust is another limit: the speaker cites prompt-inversion research reporting 88.4% token recovery from intermediate activations, rather than presenting that figure as a measured attack on this demo. Remote computation is not verified.

The project remains experimental. Chrome on macOS is the only tested host; iPhone Safari can join as a worker. The speaker reports that Safari hosting on a Mac fails under memory pressure, while Firefox and Linux Chromium remain untested.