Once an application is split into many services, every client has to know where each one lives, speak each one's protocol, authenticate to each, and make many calls to draw one screen. The API gateway pattern puts a single entry point in front of the services. Clients talk to the gateway; the gateway handles the concerns that belong at the edge and routes or composes calls to the services behind it.

This article is about the pattern: when it pays off, which topology to use, how to compose responses under a deadline, how identity should flow through it, and how it goes wrong. The internals of a gateway, such as filter chains, resilience features and the control plane, are covered in API gateway architecture.

Advertisement

The problem the pattern solves

Picture a shopping app's home screen. It needs the user's profile, a product feed, an unread notification count and some recommendations, from four services. Without a gateway the app makes four calls over a mobile network. Each call pays its own round trip, which on a phone can easily be 100 to 200 ms. Each service must be reachable from the internet, terminate TLS, validate tokens and enforce rate limits. And the app's code is now coupled to the internal service layout, so splitting or renaming a service means shipping a new app version and waiting for users to update.

A gateway fixes the three problems together. Clients make one call to a stable public API. Internal services stay private and can be reorganised behind it. Cross-cutting edge work is done once. The cost is a new component on every request path, which must be fast, highly available and kept free of business logic.

What belongs in a gateway, and what does not

The useful test is whether a concern is about the edge or about the domain.

Belongs in the gatewayKeep out of the gateway
TLS termination, request size limitsBusiness rules such as pricing or eligibility
Authenticating the caller's tokenFine-grained authorization on domain objects
Per-client rate limits and quotasData transformations that encode domain knowledge
Routing by path, host or versionLong-running workflows and orchestration with state
Simple composition of a few readsWrites that span services (use a saga or a service)
Request ids, access logs, per-route metricsCaching that requires invalidation knowledge

The pull is always towards adding logic to the gateway, because it is the one place every request passes. A gateway that knows about discounts has become a shared monolith that every team must change and nobody owns. Keep composition to read-only fan-out with simple merging; anything that needs a transaction belongs in a service. Rate limiting design is covered in Rate limiter design.

Advertisement

Topologies

One gateway for everything. Simplest to run and right for a small number of client types. It degrades as clients diverge: the mobile app wants small payloads, partners want stable coarse resources, the web app wants something else, and the one gateway accumulates all of it.

Backend for frontend (BFF). One gateway per client experience, owned by the team that builds that client. Each BFF composes exactly what its client needs, and changes to one client do not risk another. The cost is duplication of edge concerns, which is usually handled by putting a thin shared edge layer (TLS, authentication, global rate limits) in front of the BFFs. The pattern is covered in Backend for frontend.

Per-domain gateways. Large organisations often put a gateway in front of each business domain (payments, catalogue) and let a top-level edge route between them. This scales ownership, at the cost of an extra hop.

API gateway as the single entry point: edge concerns, then compositionMobile appWeb appPartner APIAPI gatewayTLS, authenticate tokenrate limit per clientroute /v1/... to servicecompose under a deadlineinject internal identityper-route metricsprofile servicep99 80 msfeed servicep99 250 msnotificationsp99 60 msrecommendationsoptional, p99 400 msidentity providerissues tokens; JWKS keyskeys cachedClients make one round trip; the gateway fans out inside the data centre and returns acomplete or explicitly partial response when the deadline expires.
A gateway handles edge concerns once, then routes or fans out to private services and returns one response.

It also helps to separate the gateway from neighbours it is confused with. A layer 4 or 7 load balancer spreads traffic across replicas of one service but does not understand your API. A service mesh handles traffic between internal services (east-west) with sidecars or node proxies. A gateway handles traffic from outside clients (north-south) and owns the public API contract. Many systems have all three.

Request composition under a deadline

Composition is where the pattern earns its latency savings and where it most often causes outages. Three rules make it safe. Call independent services in parallel. Give the whole request a deadline and derive every downstream timeout from what remains of it. Decide in advance which parts are required and which are optional, and return an explicitly partial response when an optional part misses the deadline.

import asyncio, time
import httpx

DEADLINE_S = 0.300                      # total budget for the home screen

async def call(client, url, headers, deadline, required):
    remaining = deadline - time.monotonic()
    if remaining <= 0.005:              # not worth starting
        return None if not required else _fail(url)
    try:
        # wait_for bounds the whole call; httpx's own timeout is per phase
        r = await asyncio.wait_for(client.get(url, headers=headers), timeout=remaining)
        r.raise_for_status()
        return r.json()
    except (httpx.HTTPError, asyncio.TimeoutError):
        if required:
            raise
        return None                     # optional part degrades to absent

def _fail(url):
    raise httpx.TimeoutException(f"deadline exhausted before {url}")

async def home(client, user_hdrs):
    deadline = time.monotonic() + DEADLINE_S
    profile, feed, unread, recs = await asyncio.gather(
        call(client, "http://profile/v1/me",        user_hdrs, deadline, True),
        call(client, "http://feed/v1/home",         user_hdrs, deadline, True),
        call(client, "http://notify/v1/unread",     user_hdrs, deadline, False),
        call(client, "http://recs/v1/for-you",      user_hdrs, deadline, False),
        return_exceptions=False)
    return {"profile": profile, "feed": feed,
            "unread": unread, "recommendations": recs,
            "partial": unread is None or recs is None}

The response says it is partial, so the client can show a placeholder and retry the missing part later rather than treating a gap as an empty list. Pass the remaining deadline downstream as a header too, so services stop work the gateway has already abandoned. And do not retry inside the composition by default: if the gateway retries, the service retries its database and the client retries the gateway, three attempts at each of three layers is up to 27 times the load on the bottom layer exactly when it is struggling. Retry at one layer, with a budget, and only for idempotent operations (see Idempotency).

Identity at the edge and inside

The gateway is the natural place to authenticate the caller: validate the token's signature against cached public keys from the identity provider, and check expiry, issuer and audience. Rejecting bad requests here keeps them off every service.

What goes to the services next matters more. Strip any incoming headers that claim identity, such as a user id header, before routing, otherwise an external caller can simply set them. Then forward identity in a form services can trust: either the original token, which services validate themselves, or a short-lived internal token minted by the gateway or obtained through OAuth token exchange (RFC 8693) with a narrower audience. Plain headers are acceptable only if the network path from gateway to service is authenticated, for example with mutual TLS.

Authorization still belongs to the services. The gateway can check coarse scopes (does this token allow orders:read at all?) but only the orders service knows whether order 991 belongs to this user. A gateway-only authorization model fails the first time a service is reachable by any other path.

Versioning the public contract

The gateway owns the public API, so it is where versions live. Version in the path (/v1/orders) or in a header; path versions are easier to route, cache and see in logs. Within a major version, only make additive changes: new optional fields, new endpoints. Breaking changes get a new major version that runs alongside the old one, with the gateway routing each to the right backend or to an adapter that translates. Announce retirement dates, signal them in responses (the Sunset header, RFC 8594, exists for this) and track per-client traffic on old versions so you know who still has to move.

The gateway also lets internal services change freely. When the orders service splits into two, the gateway's routing changes and the public path does not.

Worked example: the home screen budget

Assume a phone round trip of 150 ms to the edge and the service p99 latencies in the diagram: profile 80 ms, feed 250 ms, notifications 60 ms, recommendations 400 ms.

Without a gateway, the app issues four requests. If it does them sequentially, the network alone is 4 x 150 = 600 ms before any service time. Even in parallel, the screen waits for the slowest call, about 150 + 400 = 550 ms at p99, and the phone holds four connections and four TLS sessions.

With a gateway and a 300 ms composition deadline, the phone pays one 150 ms round trip. The gateway's parallel calls inside the data centre finish when feed returns, about 250 ms at p99; recommendations is optional and is cut off at 300 ms when it is slow. The screen arrives in roughly 150 + 300 = 450 ms in the worst case, usually much less, and recommendations fills in later.

One more effect: the p99 of a fan-out is worse than the p99 of each call. If four independent calls each have a 1 percent chance of exceeding their p99, the chance that at least one does is 1 - 0.99^4, about 3.9 percent. Fan-out makes tail latency common, which is why the deadline and optional parts matter more than any single service's speed. Shedding work when overloaded is covered in Load shedding.

Failure modes

  • Gateway as monolith. Business logic accumulates until every release needs the gateway team. Keep a written list of what the gateway may do and review additions against it.
  • Single point of failure. Run it stateless across zones behind a load balancer, keep rate-limit state in a store that can fail open or closed by design, and size it for peak plus a zone loss.
  • Global config blast radius. One bad route or filter change affects every API. Treat config as code, validate it, and roll it out gradually with automatic rollback.
  • Retry storms and missing deadlines. Layered retries multiply load; timeouts longer than the client's own make the gateway hold dead requests.
  • Header spoofing. Identity headers accepted from outside allow impersonation; strip them at ingress.
  • Hidden fan-out cost. One public call becoming ten internal ones multiplies backend load; measure fan-out per route.

Trade-offs

ChoiceGainCost
Single gatewayOne place to operateDiverging clients collide in one codebase
BFF per clientClient teams move independentlyDuplicated edge work unless layered
Composition in gatewayFewer client round tripsGateway latency is now the slowest required call
Thin gateway, fat clientSimple gatewayChatty clients, coupling to internal APIs

What to do next

  1. List every public endpoint with its owning service, client types and calls per screen.
  2. Write down what your gateway is allowed to do, and move any business logic you find back into services.
  3. For each composed route, mark parts as required or optional and set one end-to-end deadline.
  4. Make sure downstream timeouts derive from the remaining deadline and that retries happen at exactly one layer.
  5. Strip identity headers at ingress and decide how services verify identity: token, exchanged token or mTLS.
  6. Add per-route metrics for latency, errors, fan-out and partial responses, plus per-client traffic on each API version.
  7. Roll out gateway config gradually, with validation and automatic rollback.
Key takeaway: An API gateway gives clients one stable entry point, hides the internal service layout, and does edge work such as TLS, authentication, rate limiting and routing once. Use a single gateway while clients are similar, and move to backends-for-frontends or per-domain gateways when they diverge. Compose reads in parallel under one deadline with explicitly optional parts, retry at one layer only, strip external identity headers and forward identity services can verify, and version the public contract at the gateway. Keep business logic out of it, or it becomes the monolith you split up to escape.