
Pooled runs a single open model across browser tabs on laptops, desktops and phones, combining their memory when the model won't fit on one device. It's for people who want local AI chat or a coding assistant using hardware they already have. It's open source under the MIT license and requires no account or per-device installation.
Each device runs part of the model on its GPU through WebGPU, with direct WebRTC connections between peers. Supported models include Qwen3 1.7B, Qwen 3.8 27B and Qwen 3.6 35B MoE. Devices download only their assigned model layers, either from Hugging Face or another room member, and cache them for later use. The website provides access to the app, but inference runs on the participating devices; no server processes the conversation.
Code mode can build and edit small apps, show a sandboxed preview, inspect console errors and attempt fixes. Everyone in the room can see the files and open the preview. Files stay in the host's browser or a chosen disk folder, and edits to that folder require the host's approval.
A local API bridge connects the room to tools such as Claude Code, Codex CLI, Continue and Open WebUI. It supports OpenAI Chat Completions, OpenAI Responses and Anthropic Messages APIs, including tool calling.
Rooms share questions and answers with their members. Intermediate model data also isn't private against a determined peer, so participating devices should belong to people you trust.
Claim this page with an email at pooled.run. Pooled gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Pooled?Promote it
Something wrong or outdated on this page?
23.2KUpdated 23 hours agoApache-2.0
macOS · Windows · Linux · iOS · Android · Web#OpenAI-compatible API
MLC LLM is an open-source compiler and deployment engine for developers who want to run language models on their own hardware or inside apps. Its main distinction is the range of devices it targets: the same underlying engine, MLCEngine, serves desktop, browser and mobile deployments. The project uses the Apache 2.0 license.
11.9KUpdated 4 days agoAGPL-3.0
macOS · Windows · Linux · Android · Docker · Web#GGUF#Hugging Face integration#llama.cpp backend
130KUpdated 37 minutes agoMIT
Web#Code execution#GGUF#Hugging Face integration
qualcomm/GenieXInference Libraries and Bindings
macOS · Windows · Linux#GGUF#Hugging Face integration#llama.cpp backend
407Updated 2 days agoMIT
macOS#MLX#Multimodal input#OpenAI-compatible API
12.5KUpdated 4 months agoApache-2.0
Docker · Web#Hugging Face integration#OpenAI-compatible API
KoboldCpp pairs local model inference with a browser interface built for chat, creative writing and roleplay. A fork of llama.cpp, it bundles KoboldAI Lite with tools for keeping character details and story context alongside your conversations. It's open source under AGPL-3.0.
llama.cpp runs language models on your own hardware and can serve them from a machine you control. It’s an MIT-licensed, open source inference engine for people building local AI apps, running a private model server, or using a model directly from the command line. It supports vision-language models too.
Nexa SDK is an on-device AI inference framework for developers building applications that process text, images or audio on users' hardware. It runs models locally across CPUs, GPUs and NPUs, with a shared interface for different backends. Its scope includes language and vision models, speech recognition, speech synthesis and image generation.
Slotstream runs Qwen3.8-Flash-Next on Apple Silicon Macs that don't have enough RAM to hold the whole model. It's aimed at people with 16 to 64 GB of memory who want local chat, image questions or a model backend for coding agents. Most model weights stay on the SSD, while frequently used expert networks stay in memory. The full model remains available.
OpenLLM is a self-hosted LLM server for developers who want to connect their applications to models running on their own hardware or servers. Its OpenAI-compatible API works with clients built for that interface, including the OpenAI Python client and LlamaIndex. The project is open source under the Apache License 2.0.