Pooled distributes local language models across browsers
October 7, 2026
Pooled combines the computing power of laptops and phones over WebRTC, letting open language models run in browsers even when they are too large for one device.
What this is about
Pooled is an open-source tool that distributes a language model across several devices. Participants open a room in a browser, and each laptop, desktop, or phone takes responsibility for part of the model layers. The project was previously called SwarmLLM and has since been renamed Pooled.
It is mainly relevant to people who want to test open models locally when their memory requirements exceed one device. The approach requires neither an account nor an installation on every participating device. For sensitive data, it still matters who joins the room and which content is visible.
What Pooled actually does
Pooled downloads only the assigned model layers to each device. Computation uses custom WGSL kernels through WebGPU. For every generated token, a small hidden-state vector travels between devices over direct WebRTC connections. The host assembles the output.
According to its documentation, the project supports several Qwen models. It lists about 17 GB of distributed memory for Qwen 3.8 27B and about 22.5 GB for Qwen 3.6 35B MoE. These are project figures, not independent measurements. Model pieces are cached locally. If a device leaves, its layers can be reassigned.
There is also a Code mode. It can create small web apps inside a sandboxed preview, inspect console errors, and propose changes. Access to a real folder requires host approval. A local bridge also lets clients that speak OpenAI- or Anthropic-compatible APIs use the room.
Why it matters
Local inference often fails not because of raw compute but because there is not enough memory. Pooled treats several existing devices as a shared shelf for model layers. That can enable tests without immediately buying a high-memory machine or renting a cloud GPU.
The browser also lowers the entry barrier: additional devices join through a link. The project publicly documents its architecture, threat model, and benchmarks and uses the MIT license. The API bridge matters for teams because existing coding or chat clients can use a Pooled room as a model endpoint.
Practical value depends heavily on the network. Every extra participant lengthens the route for each token. Corporate networks can block direct WebRTC links, requiring a TURN relay. Pooled is therefore better viewed as an inspectable experimentation tool than a replacement for a stable inference server.
In plain language
Imagine a very heavy picture book that no single backpack can carry. Pooled gives different chapters to several people. For each question, they pass a small note in a fixed order until the answer returns to the host. The farther apart those people are, the longer every round takes.
A practical example
A developer wants to test an open 27-billion-parameter model. Her laptop can spare only 10 GB, while two other devices can provide 4 GB and 5 GB. Together they exceed the roughly 17 GB requirement stated by the project. All three open the same room and download only their assigned layers.
She then connects a local coding client through the supplied API bridge. A source-code prompt runs across the three devices. Before real project files can be changed, she must approve folder access. She measures latency, stability, and power use against a smaller model running on the laptop alone. Pooled remains in the test setup only if the benefit outweighs the extra network delay.
Scope and limits
First, performance figures come from the project. Hardware, browser, model, and network latency can materially change results, and a native inference server may be faster.
Second, depending on room settings, the host and other participants may see prompts or answers. Confidential code belongs only in a room with tightly controlled membership. WebRTC and an optional relay do not replace organizational access controls.
Third, each request depends on all participating devices. Sleep mode, closed tabs, memory pressure, or weak Wi-Fi can interrupt a run. Anyone needing guaranteed availability, long context, or reproducible production performance should prefer an established local or hosted server.
The sensible next step is a trial with two personally controlled devices, a non-sensitive model, and public sample data. Record time to first token, throughput, failure behavior, and prompt visibility.
SEO & GEO keywords
Pooled, SwarmLLM, local AI, browser inference, WebGPU, WebRTC, distributed language models, open models, Qwen, coding agents, peer-to-peer inference
💡 In plain English
Pooled splits a large open language model across several browser devices. This can enable local testing, but speed and availability depend on the network and every participating device.
Key Takeaways
- →Pooled distributes model layers across devices with WebGPU and WebRTC.
- →Participants join through a browser link without an account.
- →A local bridge provides OpenAI- and Anthropic-compatible interfaces.
- →Prompts may be visible to other participants depending on room settings.
- →Project benchmarks should be verified on your own hardware.
FAQ
Does Pooled require installation on every device?
No. Joining the compute room only requires a compatible WebGPU browser; the optional API bridge runs separately on a local machine.
Does all data remain on one device?
No. Hidden model states travel between participating devices, and prompts or answers may be visible to room participants depending on settings.
Is Pooled suitable for production?
The project is primarily compelling for experiments. Production use requires separate testing of latency, failures, access controls, and browser compatibility.