WebAssembly lets you run a compiled C++ inference engine such as llama.cpp inside a browser tab, a serverless edge function or a plugin host, with no install and no server round-trip. That is attractive for privacy (prompts never leave the device), cost (the user's CPU does the work) and offline use. It is also easy to oversell. Wasm in the browser today runs on the CPU unless you add WebGPU, its memory model has hard limits, and its threads require server headers that break some sites.
This article explains how Wasm inference works from the instruction set up: why decoding is limited by memory bandwidth, how 128-bit SIMD kernels compute quantised dot products, what threads and cross-origin isolation require, how the 4 GB address space and 2 GB file limits shape model choice, and when to move to WebGPU. It ends with a sizing worked example and a checklist.
What WebAssembly gives an inference engine
A Wasm module is portable bytecode for a stack machine with typed integers and floats, a single contiguous linear memory, and no access to anything outside that memory except functions the host explicitly imports. The browser compiles it to native code ahead of execution, so tight loops run at a substantial fraction of native speed. The sandbox is the point: a model runtime can be shipped as one .wasm file and cannot read the user's files or open sockets on its own.
For inference, three features matter most. Fixed-width SIMD adds 128-bit vector instructions, and every serious engine depends on it. Threads pair a shared linear memory with atomics so several Web Workers can split a matrix multiply. Memory64 allows indexes beyond 32 bits; it shipped in Chrome 133 and Firefox 134 but not, at the time of writing, in Safari. Everything else is ordinary C or C++ compiled with Emscripten, which is why llama.cpp ports such as wllama exist at all. If GGUF and llama.cpp are new to you, read the GGUF format and llama.cpp first.
Why decode speed is a memory-bandwidth problem
Token generation is a sequence of matrix-vector products. For each new token, every weight in the model is read once and used for one multiply-add per element. Arithmetic intensity is roughly one or two operations per byte for 4-bit weights, so the processor spends its time waiting for memory, not computing. That gives a useful upper bound:
tokens_per_second <= effective_memory_bandwidth / bytes_of_weights_read_per_token
example (back-of-envelope, not a benchmark):
weights at ~4.5 bits/param, 1.5e9 params -> ~0.85 GB read per token
effective bandwidth seen by Wasm threads -> assume 15 GB/s
upper bound -> ~17 tokens/s, before overheadsTwo consequences follow. First, quantisation is the biggest lever you have: halving bytes per weight roughly doubles the ceiling, which is why 4-bit formats dominate in the browser; see quantisation for small models. Second, Wasm can only approach native speed on this workload; it cannot exceed it, and the gap comes from narrower vectors (128-bit, while native x86 builds can use 256- or 512-bit), bounds-checked memory and fewer threads. Prompt processing (prefill) is the opposite case: it multiplies matrices by matrices, is compute-bound, and suffers most from narrow SIMD. Long prompts on CPU Wasm are slow; plan for it. The GPU versus CPU comparison explains the same split on native hardware.
SIMD kernels for quantised dot products
Quantised formats store weights in small blocks with a shared scale. The inner loop multiplies int8 activations by int8 (or unpacked 4-bit) weights and accumulates in int32. With Wasm SIMD you widen bytes to 16 bits and use the pairwise multiply-add instruction i32x4.dot_i16x8_s:
#include <wasm_simd128.h>
#include <stdint.h>
// Dot product of two 32-element int8 blocks, the core of a Q8-style kernel.
static inline int32_t dot_i8_32(const int8_t *a, const int8_t *b) {
v128_t acc = wasm_i32x4_splat(0);
for (int i = 0; i < 32; i += 16) {
v128_t va = wasm_v128_load(a + i), vb = wasm_v128_load(b + i);
acc = wasm_i32x4_add(acc, wasm_i32x4_dot_i16x8(
wasm_i16x8_extend_low_i8x16(va), wasm_i16x8_extend_low_i8x16(vb)));
acc = wasm_i32x4_add(acc, wasm_i32x4_dot_i16x8(
wasm_i16x8_extend_high_i8x16(va), wasm_i16x8_extend_high_i8x16(vb)));
}
return wasm_i32x4_extract_lane(acc, 0) + wasm_i32x4_extract_lane(acc, 1)
+ wasm_i32x4_extract_lane(acc, 2) + wasm_i32x4_extract_lane(acc, 3);
}
// the block result is then scaled: sum += dot_i8_32(x, w) * scale_x * scale_w;Compile with emcc -O3 -msimd128 -pthread. The relaxed-SIMD proposal adds instructions such as a direct int8 dot product whose results are not guaranteed identical across CPUs for every input; engines that use it need a fallback build because support differs between browsers. Ship two binaries and pick at runtime by feature-detecting with a tiny validation module, the same way engines choose between single-thread and multi-thread builds.
Threads and cross-origin isolation
Wasm threads are Web Workers that share one WebAssembly.Memory backed by a SharedArrayBuffer. Since the Spectre mitigations, browsers only expose SharedArrayBuffer to pages that are cross-origin isolated, which requires two response headers on the top-level document:
Cross-Origin-Opener-Policy: same-origin
Cross-Origin-Embedder-Policy: require-corp
// in the page or worker, check before choosing a build
const canThread = self.crossOriginIsolated && typeof SharedArrayBuffer !== "undefined";
const threads = canThread ? Math.min(navigator.hardwareConcurrency || 4, 8) : 1;The price is that every cross-origin resource the page embeds, including images, scripts, iframes and fonts, must opt in with CORS or a Cross-Origin-Resource-Policy header, and popups to other origins lose their opener reference. Third-party ad, payment or login widgets often break. Common fixes are to serve the inference UI from a dedicated isolated route or subdomain, or to use COEP credentialless where supported. Without isolation, engines fall back to one thread, which typically costs several times the throughput.
Do not run inference on the main thread. Put the engine in a dedicated worker and stream tokens back with postMessage; the UI then stays responsive during a long prefill. More threads than physical cores rarely helps because decode is bandwidth-bound; leave headroom for the page itself.
Memory limits, sharding and the KV cache
Wasm32 linear memory is addressed with 32-bit indexes, so a module sees at most 4 GB, and the weights, KV cache, scratch buffers and the runtime heap all share it. Separately, a single ArrayBuffer used to load a file is limited in practice, and wllama's documentation caps a single model file at 2 GB for that reason. The standard workaround is to split the GGUF into chunks, which also lets the browser download them in parallel:
# split into 512 MB shards: model-00001-of-0000N.gguf ...
./llama-gguf-split --split-max-size 512M ./model-q4_k_m.gguf ./model-q4_k_mThe KV cache is the other large tenant of linear memory. Per token it costs 2 x layers x kv_heads x head_dim x bytes_per_value. A 24-layer model with 2 KV heads of dimension 64 in fp16 needs 2 x 24 x 2 x 64 x 2 = 12,288 bytes per token, so a 4,096-token context costs about 48 MB. Models without grouped-query attention can need ten times that; KV cache sizing walks through the formula. In practice a wasm32 build holds a model of a few gigabytes on desktop with a modest context, and less on mobile, where browsers cap memory lower; beyond that you need Memory64, which costs some speed for bounds checks and excludes Safari, or a smaller model.
Downloading a model on every visit is unacceptable, so persist the chunks in the Origin Private File System or the Cache API, call navigator.storage.persist() to reduce eviction risk, and check navigator.storage.estimate() before downloading. Browsers can still evict, so always handle a missing cache.
Runtimes in practice: wllama, ONNX Runtime Web and beyond
Three toolchains cover most browser use. wllama compiles llama.cpp to Wasm, loads GGUF directly and, from version 3, exposes an OpenAI-style chat API and, from 3.1, can offload layers to WebGPU when available. ONNX Runtime Web runs ONNX graphs with a Wasm execution provider configured through ort.env.wasm flags (numThreads, simd, proxy to run in a worker) and also offers WebGPU. WebLLM is WebGPU-first, with Wasm handling the runtime around compiled GPU kernels. A minimal wllama v3 session:
import { Wllama } from "@wllama/wllama";
const wllama = new Wllama({ default: "/wllama/wllama.wasm" }); // path to the shipped .wasm
await wllama.loadModelFromHF({
repo: "your-org/your-model-GGUF",
file: "model-q4_k_m-00001-of-00004.gguf", // first shard; the rest load automatically
});
const reply = await wllama.createChatCompletion({
messages: [{ role: "user", content: "Summarise this note in two lines: ..." }],
max_tokens: 128,
temperature: 0.3,
});Check the current README before copying option names; the wllama API changed shape between major versions. Outside the browser, Wasm runtimes such as Wasmtime and WasmEdge run the same kind of module server-side, and the WASI-NN proposal lets a module call a host-provided inference backend instead of doing the math itself. That trades portability of the math for native speed, and is useful for plugin systems where you want sandboxed glue around a trusted native engine.
Worked example: an offline note summariser
Suppose you want an in-browser note summariser that must work offline and keep text on the device. Requirements: 2,000-token inputs, 150-token summaries, usable on a mid-range laptop.
Start from memory. A 0.5 to 1.5 billion parameter instruct model at 4-bit quantisation weighs roughly 0.4 to 1 GB, fits wasm32 with room for a 4,096-token KV cache, and splits into two or three shards. Next, time. By the bandwidth bound above, decode on a multi-threaded build lands in the tens of tokens per second at best, so 150 tokens takes several seconds; acceptable with streaming. Prefill of 2,000 tokens is compute-bound and on narrow SIMD may take longer than the decode itself, so show a progress state and cap the input. Then deployment: serve the app with COOP and COEP, host shards on the same origin or a CDN that sends CORS headers, cache them in OPFS, and fall back to single-thread with a warning when isolation is unavailable. Finally, measure on real target devices; laptop CPUs vary by several times in memory bandwidth, and thermal throttling shows up after a minute of sustained decode.
If measurements miss the target, you have three moves in order of cost: a smaller or more aggressively quantised model, WebGPU offload where available, or a server fallback for long inputs. See edge deployment for small models for the broader decision.
Failure modes and trade-offs
- Silent single-thread fallback. A missing COEP header on one route drops throughput several-fold with no error; log
crossOriginIsolatedand thread count in telemetry. - Out-of-memory at load. Weights plus KV plus scratch exceed 4 GB, or the tab hits a browser memory cap on mobile; size context to the device.
- Shard and CORS mistakes. One shard served without CORS under COEP fails the whole load.
- Storage eviction. A cached model disappears under storage pressure and the user silently re-downloads gigabytes.
- Main-thread inference. The page freezes during prefill and browsers may flag it as unresponsive.
- Unverified feature builds. A relaxed-SIMD or Memory64 binary that fails to compile on one browser must fall back, not crash.
| Choice | Wasm CPU | WebGPU | Server API |
|---|---|---|---|
| Reach | Every modern browser | Most desktop browsers; varies on mobile | Everywhere with a network |
| Decode speed | Bandwidth-bound, modest | Usually several times faster | Fastest, scales with spend |
| Privacy | Data stays on device | Data stays on device | Data leaves device |
| Model size | Small; 4 GB address space | Bounded by GPU and browser limits | Any |
| Ops cost | CDN bandwidth only | CDN bandwidth only | GPU fleet |
What to do next
- Write down the target device class and measure its memory bandwidth; compute the decode ceiling for your candidate model sizes.
- Pick a 4-bit GGUF or quantised ONNX model that fits wasm32 with your context length, and split it into shards of 512 MB or less.
- Serve the page with COOP and COEP, verify
crossOriginIsolatedis true, and audit every cross-origin embed. - Run inference in a dedicated worker, stream tokens, and cap prompt length so prefill stays bounded.
- Persist shards in OPFS or the Cache API, request persistent storage, and handle eviction gracefully.
- Feature-detect SIMD, threads and WebGPU at runtime; ship fallback builds and record which path each session used.
- Benchmark on real low-end and high-end devices, including a sustained run to catch throttling, before promising latency.