Most systems start with one API that every client calls. That works while there is one client. Once a web app, two mobile apps, a TV app and a partner integration all use it, the API ends up serving none of them well: the mobile app makes six round trips over a slow network to draw one screen, the web app receives fields it never displays, and every screen change needs a negotiation with the team that owns the shared API.
The Backend for Frontend (BFF) pattern answers this by giving each user experience its own thin server-side component, owned by the team that builds that experience. The BFF calls the domain services, combines and trims their responses, and returns exactly what its client needs. The pattern grew out of SoundCloud's move away from a single public API and was named and popularised by Sam Newman. This article covers where a BFF sits, how it differs from a gateway, how to write one that survives partial failure, how it handles authentication for browsers, and when it is the wrong choice.
The problem a BFF solves
A general-purpose API is designed around domain resources: products, prices, stock, reviews, users. A screen is designed around a task, and one screen usually needs parts of several resources. When the client assembles the screen itself, three costs appear. Latency: each dependent round trip over a mobile network costs a full network round-trip time, often 50 to 150 ms or more, and requests that depend on earlier results cannot be parallelised. Payload: resource-shaped responses carry every field for every client, which costs bandwidth and battery on phones. Coupling: every client release depends on the shared API exposing exactly the right shape, and every change to that API has to consider every client.
A BFF moves the assembly server-side, next to the services, where a round trip is typically about a millisecond instead of about a hundred. The client makes one call and receives one response designed for its screen. Just as important, the BFF is owned by the frontend team, so changing what a screen needs is a change inside one team's deployable rather than a cross-team request.
Where it sits in the architecture
A typical layout has four tiers. Clients connect to an edge layer: a CDN and an API gateway that terminate TLS, apply the web application firewall, rate limits and coarse authentication, and route by path or host. Behind the edge sit the BFFs, one per experience. Behind them are the domain services, which know nothing about screens.
The granularity rule most teams use is one BFF per user experience, not per device. If the iOS and Android apps show the same screens and are built by the same team, one mobile BFF serves both. If the web checkout and the kiosk checkout are built by different teams with different release cycles, they get separate BFFs even though both are web technology. Conway's law is the design input here: the BFF boundary should follow team ownership.
BFF, API gateway and GraphQL compared
| Aspect | API gateway | BFF | GraphQL layer |
|---|---|---|---|
| Owner | Platform team | The frontend team for that experience | Usually a platform or API team |
| Main job | Cross-cutting: TLS, auth checks, rate limits, routing | Aggregate, shape and apply client-specific policy | Let clients declare the shape they need |
| Business logic | None | Presentation logic only | Resolvers; should stay thin |
| Number of instances | One shared | One per experience | Usually one shared schema |
| Changes when a screen changes | No | Yes, by the same team | Often no; the query changes |
They combine rather than compete. The gateway stays in front of everything (see API gateway design). A GraphQL server can be the implementation of a BFF, and a single shared GraphQL layer can replace several BFFs when clients differ mostly in the fields they select rather than in their policies, error handling and authentication. The case for separate BFFs is strongest when clients need different behaviour, not just different fields.
A worked example: the mobile product page
A product screen in a mobile app needs a title and one small image from the catalog, a user-specific price, a stock flag and a review summary. Called directly, that is four requests, and the price call needs the user's segment from a fifth identity call first. With illustrative numbers of 120 ms round-trip time on a mobile network and two dependent rounds, the screen waits at least 240 ms for the network alone. The raw responses together are about 90 KB of JSON, because the catalog response carries every image size, every locale and the full description.
Through a mobile BFF the phone makes one request. The BFF runs the four calls in parallel inside the data centre, where each round trip is about a millisecond and service time dominates, and returns about 2 KB. With the same illustrative numbers, the phone waits one network round trip plus the slowest dependency, and the payload shrinks by more than 40 times. Here is the handler, written in TypeScript for Node 20.3 or later, which provides AbortSignal.any:
// Mobile BFF: GET /m/v2/product/:id (Node 20.3+ for AbortSignal.any, TypeScript)
type Dep<T> = { name: string; critical: boolean; timeoutMs: number; call: (s: AbortSignal) => Promise<T> };
async function run<T>(d: Dep<T>, parent: AbortSignal): Promise<{ ok: true; v: T } | { ok: false; err: string }> {
const signal = AbortSignal.any([parent, AbortSignal.timeout(d.timeoutMs)]);
try { return { ok: true, v: await d.call(signal) }; }
catch (e) { metrics.inc("bff_dep_failure", { dep: d.name }); return { ok: false, err: String(e) }; }
}
app.get("/m/v2/product/:id", async (req, res) => {
const deadline = AbortSignal.timeout(800); // whole-request budget
const id = req.params.id, user = req.session.userId;
const hdrs = { "x-request-id": req.id, authorization: `Bearer ${await tokens.forUser(user)}` };
const get = (url: string) => (s: AbortSignal) => fetch(url, { headers: hdrs, signal: s }).then(r => {
if (!r.ok) throw new Error(`${url} ${r.status}`); return r.json(); });
const [product, price, stock, reviews] = await Promise.all([
run({ name: "catalog", critical: true, timeoutMs: 300, call: get(`${CATALOG}/products/${id}`) }, deadline),
run({ name: "pricing", critical: true, timeoutMs: 250, call: get(`${PRICING}/prices/${id}?user=${user}`) }, deadline),
run({ name: "inventory", critical: false, timeoutMs: 200, call: get(`${INVENTORY}/stock/${id}`) }, deadline),
run({ name: "reviews", critical: false, timeoutMs: 200, call: get(`${REVIEWS}/summary/${id}`) }, deadline),
]);
if (!product.ok || !price.ok) return res.status(503).json({ error: "product_unavailable" });
res.set("cache-control", "private, max-age=30").json({ // shaped for the phone screen
id, title: product.v.title, image: product.v.images?.[0]?.small,
price: price.v.display, inStock: stock.ok ? stock.v.available > 0 : null,
rating: reviews.ok ? { avg: reviews.v.avg, count: reviews.v.count } : null,
degraded: [stock, reviews].some(r => !r.ok),
});
});The handler encodes several decisions that a BFF exists to make. Catalog and pricing are critical: without them the screen is meaningless, so their failure returns an error. Inventory and reviews are optional: their failure yields null fields and a degraded flag that tells the app to render a placeholder. Every dependency has its own timeout, and all of them are bounded by one 800 ms deadline for the whole request, so a slow dependency cannot hold the phone's request open indefinitely. The request ID is forwarded so a trace joins the phone's request to every downstream call.
Latency budgets and the tail
Fan-out makes tail latency worse. If each of four dependencies independently exceeds its 99th-percentile latency 1 percent of the time, the chance that at least one does is 1 - 0.99^4, about 3.9 percent. With ten dependencies it is about 9.6 percent. A BFF that waits for everything therefore has a much worse p99 than any service it calls.
The tools are the same as for any aggregator, applied deliberately. Give every downstream call a timeout shorter than the overall budget, and propagate the remaining deadline downstream in a header so services stop working on requests the BFF has already abandoned. Mark dependencies critical or optional and degrade on optional ones. Put a circuit breaker per dependency so a failing service is skipped quickly instead of consuming the budget on every request. Cache what is safe to cache: catalog data for seconds or minutes, never user-specific prices across users (see caching strategies and, for public responses, CDN design). Keep connection pools and concurrency limits per dependency, so one slow service cannot exhaust the connections the others need.
Authentication: the BFF as the token handler
For browser apps the BFF has become the recommended place to hold OAuth tokens. RFC 10017, OAuth 2.0 for Browser-Based Applications, published in August 2026 as BCP 212, describes the Backend for Frontend pattern as one of the architectures for browser-based OAuth clients and analyses its security properties against alternatives that keep tokens in JavaScript.
In that arrangement the BFF is a confidential OAuth client. It runs the authorization code flow with the identity provider, keeps the access and refresh tokens server-side, and gives the browser only a session cookie marked HttpOnly, Secure and SameSite. When the browser calls the BFF, the BFF looks up the session, attaches the access token to downstream calls and refreshes it when needed. Script injected into the page cannot read tokens it never sees, although it can still make requests through the user's session, so XSS defences remain necessary. Because the browser now authenticates with a cookie, the BFF must defend against cross-site request forgery with SameSite cookies plus a custom header or anti-CSRF token check on state-changing requests.
Authorization still belongs to the domain services. A BFF that decides whether a user may see an order is a second, drifting copy of the rules. The BFF forwards identity, and the service decides.
Failure modes and anti-patterns
- The BFF grows business logic. Discount rules, eligibility checks and workflow start appearing in the BFF because it is convenient for the frontend team. Within a year two BFFs disagree about a price. Keep BFFs to composition, formatting and client policy; push rules into services.
- Duplication across BFFs. Some overlap in glue code is the price of independence. Extract shared client libraries for calling services, not a shared BFF framework that recreates the coupling.
- One BFF for everyone. A single BFF serving web, mobile and partners is a general-purpose API with an extra hop.
- Old mobile versions. Web clients update on the next page load, but mobile apps stay in the field for months or years. The mobile BFF must keep supporting older response shapes, so version its routes (
/m/v2/) and track traffic per app version before removing one. - Fan-out storms. A popular screen that fans out to ten services multiplies traffic tenfold. Rate-limit per client at the edge (see rate limiter design) and cache in the BFF.
- The extra hop. A BFF adds one network hop and one more thing to deploy and page for. For a single client with simple screens, it is overhead.
Running BFFs in production
BFFs are I/O-bound: they wait on downstream calls far more than they compute. Use an asynchronous runtime, scale horizontally on concurrency rather than CPU, and keep them stateless apart from a session store. Define SLOs per BFF in terms the client team cares about, such as screen-level latency and the rate of degraded responses, and alert on dependency-level error rates too, so a degraded review service is visible before users complain.
Contract tests keep BFFs and services honest. Each BFF records the fields it actually reads from each service, and service pipelines run those contracts before deploying, so a service can remove a field no BFF uses without fear. Distributed tracing that joins the client request, the BFF span and every downstream span is the single most useful debugging tool, because most BFF incidents are really one slow dependency.
When not to build one
Skip the BFF if you have one client, if all clients need essentially the same data, or if the frontend team cannot own and operate a service. A well-designed resource API with sparse fieldsets, or a shared GraphQL layer, covers many of the same needs without new deployables. Build BFFs when clients differ in behaviour, when mobile latency or payload is a measurable problem, or when browser token handling pushes you toward a server-side session.
What to do next
- List your clients and the teams that own them; draw BFF boundaries along ownership.
- Pick one high-traffic screen, count its client round trips and payload, and measure the baseline.
- Build the BFF endpoint for that screen with per-dependency timeouts, an overall deadline and critical versus optional classification.
- Add tracing with a propagated request ID and dashboards for dependency error rates and degraded responses.
- For browser clients, move tokens into the BFF with HttpOnly session cookies and CSRF protection, following RFC 10017.
- Keep authorization in domain services and review BFF code for business rules that have crept in.
- Version mobile routes and track traffic by app version before deprecating any response shape.