A point of presence (PoP) is a small data centre a network or CDN provider runs close to users, often in hundreds of cities. Edge PoP LLM serving means doing part of the work of an LLM request in those sites instead of in one or two large GPU regions. The appeal is obvious: less distance, less latency. The catch is less obvious. GPUs at the edge are spread thin, large models do not fit there economically, and the latency an edge removes is often a small part of what users wait for.

This article breaks LLM latency into the parts an edge can and cannot shorten. It sets out which work belongs in a PoP, shows an edge router that serves from caches or a small model and escalates to origin, works through the capacity cost of many small GPU pools, and covers caching, model rollout and failure modes. All latency and capacity figures are assumptions chosen to illustrate the arithmetic; measure your own.

Choosing regions and failing over between them is covered in multi-region LLM serving. Running general code in PoPs, without GPUs, is covered in cloud edge compute architecture.

Where the latency goes

For a streamed response, two numbers matter. Time to first token (TTFT) is network round trips, plus queueing, plus prefill of the prompt. Total time is TTFT plus the number of output tokens times the time per output token (TPOT). An edge PoP shortens the network part. It terminates TLS near the user, so a new connection's handshakes cost a few milliseconds instead of a full intercontinental round trip each, and it carries traffic to origin on warm, long-lived backbone connections. Once streaming starts, tokens flow in a pipeline and distance adds its delay roughly once, not per token.

Assumption setOrigin onlyEdge model
Network to model160 ms20 ms
Prefill120 ms (large pool, fast GPUs)180 ms (small pool, smaller GPUs)
TPOT25 ms35 ms
TTFT280 ms200 ms
Total, 8-token answer0.48 s0.48 s
Total, 300-token answer7.8 s10.7 s

Under these assumptions the edge wins first-token time by 80 ms, ties on a short classification-style answer, and loses badly on a long answer because decode speed dominates. Running the same model at the edge on smaller, busier GPUs can make users wait longer. The edge helps most when the answer is short, when the response is cacheable, or when the work does not need a GPU at all.

The architecture

Edge PoP LLM serving: answer what you can near the user, stream the rest from originUsersbrowsers, apps, SDKsEdge PoP (one of many)TLS + authRate limitExact cacheSemantic cacheSmall GPU poolsmall models, short outputstenant-scoped caches; escalate on low confidenceOrigin regionlarge models, big poolsprefix / KV cachestream on backboneNeighbour PoPspill when local pool fullspillModel registryversioned weights pushed to PoPsMost of the latency saved is handshake and first-token time; long answers are dominated by decode speed.
Figure 1. An edge PoP terminates TLS, authenticates and rate-limits, answers from tenant-scoped caches or a small local model, and otherwise streams the request to an origin region over the provider's backbone.

What belongs at the edge

Sort each route by what it needs, not by where users are.

WorkAt the edge?Why
TLS, auth, quotas, request validationYesNo GPU; rejects bad traffic before it crosses the backbone
Exact-match response cacheYesDeterministic requests repeat; a hit is milliseconds
Semantic cacheSometimesNeeds an embedding model and careful thresholds and isolation
Short outputs: classification, routing, moderation, autocompleteOftenSmall models, few tokens, latency-sensitive
Embeddings for retrievalOftenSmall models, batchable, no decode loop
Long chat answers, agents, large modelsNoDecode-bound; needs big pools, prefix caches and the newest weights

Gateway features that need no GPU, such as quotas, key management, logging and fallbacks, are described in the AI gateway overview. Edge PoPs are a natural place to run them even when no model runs there.

An edge router

The edge router decides per request: reject, answer from cache, answer with a local small model, or send to origin. The sketch below is vendor-neutral; ctx stands for whatever your edge platform provides for key-value storage, a vector index, a local model pool and an origin client. The confidence gate is illustrative and must be calibrated on your own labelled traffic.

import hashlib, json

SEM_THRESHOLD = 0.95          # cosine similarity; tune per route on labelled pairs
EDGE_MAX_TOKENS = 64          # edge model only handles short answers
MIN_MEAN_LOGPROB = -0.5       # illustrative escalation gate; calibrate per route


def cache_key(tenant, req):
    """Everything that changes the answer must be in the key, including the model version."""
    canon = json.dumps({
        "model": req["model"], "version": MODEL_VERSION[req["model"]],
        "system": req.get("system", ""), "messages": req["messages"],
        "temperature": req.get("temperature", 1.0), "max_tokens": req.get("max_tokens"),
    }, sort_keys=True, separators=(",", ":"))
    return tenant + ":" + hashlib.sha256(canon.encode()).hexdigest()


async def handle(req, ctx):
    tenant = await ctx.authenticate(req)            # reject at the edge, not at origin
    await ctx.rate_limit(tenant)
    cacheable = req.get("temperature", 1.0) == 0 and not req.get("personalised", False)

    if cacheable:
        key = cache_key(tenant, req)
        if (hit := await ctx.kv.get(key)) is not None:
            return hit
        if req["route"] in SEMANTIC_ROUTES:
            emb = await ctx.embed(req["messages"][-1]["content"])
            match = await ctx.vectors.nearest(namespace=tenant, vector=emb)   # never cross tenants
            if match and match.score >= SEM_THRESHOLD and match.version == MODEL_VERSION[req["model"]]:
                return match.response

    if req["route"] in EDGE_ROUTES and ctx.local_pool.has_capacity():
        out = await ctx.local_pool.generate(req, max_tokens=EDGE_MAX_TOKENS)
        if out.finish_reason == "stop" and out.mean_logprob >= MIN_MEAN_LOGPROB:
            return out
        ctx.metrics.incr("edge_escalation", route=req["route"])      # fall through to origin

    # Sticky by prompt prefix so origin replicas reuse their prefix cache.
    return await ctx.origin.stream(req, affinity=hashlib.sha256(req.get("system", "").encode()).hexdigest())

Three rules are built in. Caches are keyed by tenant, and the semantic index is partitioned by tenant, so one customer's answer is never served to another. The model version is part of every key, so a rollout does not serve stale answers. And escalation is cheap: a weak edge answer costs a few milliseconds before the request goes to origin, not a wrong answer shown to the user. Routing between model tiers more generally is covered in LLM routing strategies.

Capacity: many small pools

One big pool needs less spare capacity than many small ones, because random peaks average out across a large population and not within a small one. Model concurrent requests as Poisson and provision every pool for its 99.9th percentile:

LayoutMean concurrent per poolp99.9 per poolSlots provisioned
1 origin pool1,0001,0981,098
50 PoPs20351,750
200 PoPs5132,600

The same traffic needs about 1.6 times the slots across 50 PoPs and 2.4 times across 200, before you count that every PoP must also hold each model's weights in GPU memory. Real traffic is burstier than Poisson and follows the sun, which makes the gap larger. Two responses help. Provision PoPs for typical load and spill peaks to a neighbour or to origin, accepting the slower path for a few requests. And keep the number of models at the edge small, because every additional model multiplies memory across every site.

Caching at the edge

Caching is where edges earn most of their value, and where they cause the most incidents.

  • Exact caching is safe when the request is deterministic (temperature 0 or a fixed seed the platform honours) and not personalised. The key must include the model version, system prompt and sampling parameters.
  • Semantic caching returns an earlier answer for a similar question. A threshold that is too low returns confident wrong answers: questions that differ in one negation or one number embed close together. Start high, measure false hits on labelled pairs, and enable it only on routes where a near-miss is harmless, such as FAQ-style help. Vendors offer this as a product; Fastly's AI Accelerator, for example, is described in its documentation as a semantic cache in front of LLM APIs.
  • Isolation is not optional. A shared cache across tenants leaks data the moment two customers ask similar questions about their own documents.
  • Expiry must follow the facts, not only time. Answers grounded in retrieved documents go stale when the documents change; tie them to a content version.
  • Prefix and KV caches live with the GPUs. Small edge pools see too little repeated traffic to benefit much, which is one more reason to send long conversations to origin with sticky routing.

Rolling out models to hundreds of sites

An origin region gets a new model version from one registry push. A fleet of PoPs needs it in hundreds of places, each with limited disk and GPU memory, over links shared with customer traffic. Treat the rollout like a CDN deployment. Pre-stage weights to every PoP before the switch, verify checksums, and flip a version pointer only when the PoP reports the weights loaded and a smoke test passed. Canary by PoP, starting with low-traffic sites, and compare escalation rate, cache-hit quality and latency against the previous version. Pin a conversation to one version for its lifetime, because a model switching mid-conversation produces visible inconsistencies. Keep the old version's weights until the new one has run cleanly for a full traffic cycle, so rollback is a pointer flip rather than a re-download.

Evaluation must cover the edge model on its own and the whole cascade together. A smaller, cheaper edge model that escalates more often can increase origin load and cost even while it looks fine in isolation. Techniques that speed up decode, such as speculative decoding, belong at origin, where the big pools and the draft-and-verify model pairs live.

Worked example: a writing assistant

Consider a writing assistant with three routes: inline autocomplete (5 to 15 tokens, latency-critical), a moderation check on every message, and a chat panel with long answers.

  1. Autocomplete and moderation run on a small model in each of the larger PoPs. Both have short outputs, so the 80 ms first-token gain is most of the user-visible latency.
  2. Autocomplete escalation is disabled: a slow suggestion is worse than none, so on low confidence or a full pool the PoP returns nothing.
  3. Moderation escalates to origin on low confidence, because a wrong pass is a safety issue and a few hundred extra milliseconds is not.
  4. Chat goes to origin, streamed over the backbone, sticky by system-prompt hash so the origin prefix cache stays warm. Its help-centre sub-route has a tenant-scoped semantic cache with a high threshold.
  5. Small PoPs run no GPUs and only terminate TLS, authenticate and route; their autocomplete traffic spills to the nearest large PoP.

Measure per route: edge hit rate, escalation rate, TTFT and total latency at p50 and p99, and origin load. If the edge model's escalation rate climbs after a rollout, the edge is adding latency to most requests while saving it on few.

Failure modes

  • Cross-tenant cache leak. A cache keyed only by prompt text returns one customer's answer to another.
  • Stale weights in a few PoPs. A failed pre-stage leaves some sites on the old model; users in those cities get different answers. Alert on version skew.
  • Escalation storms. A bad edge model escalates everything, and origin receives full traffic plus edge latency. Cap escalation rate and fall back to direct-to-origin routing.
  • Semantic false hits. Similar but different questions get the cached answer. Track user corrections and regenerations on cached responses.
  • Peak in one city. A local event saturates one PoP's pool. Without spill, users there queue while other pools sit idle.
  • Long streams through small PoPs. Requests for long outputs pinned to an edge model take longer than origin would. Enforce the output cap.

Trade-offs

Edge serving buys lower first-token latency, early rejection of bad traffic and cheap cache hits. It pays with GPU capacity fragmented across many sites, more model copies, a harder rollout, and a second model to evaluate and keep aligned with the first. Provider platforms such as Cloudflare Workers AI, which runs inference on GPUs inside its network, remove the hardware work but limit you to the models and runtimes they offer. Building your own edge tier gives control and costs operations. For most products the high-return steps are, in order: terminate and authenticate at the edge, cache exact deterministic answers, move short-output routes to small edge models, and leave long generation at origin.

What to do next

  1. Measure TTFT, TPOT and output length per route, and compute how much of each route's latency an edge could remove.
  2. Move TLS, authentication, quotas and validation to the edge first; no GPU is needed.
  3. Add an exact-match cache with tenant, model version and sampling parameters in the key.
  4. Pilot a small edge model on one short-output route, with a calibrated escalation gate and per-route metrics.
  5. Provision PoPs for typical load and design spill to neighbours and origin for peaks.
  6. Build a staged, canaried rollout with version pinning and skew alerts before running models in many PoPs.
  7. Re-evaluate the whole cascade after every edge or origin model change.
Key takeaway: Edge PoPs shorten handshakes and first-token time, but decode speed dominates long answers, so put authentication, quotas, tenant-scoped exact caches and short-output small models at the edge and keep long generation at origin. Many small GPU pools need far more spare capacity than one large pool, so provision for typical load and spill peaks. Roll models out like a CDN deployment, with staging, canaries and version pinning.